❌

Reading view

Standardized pre-consultation by a large language model agent vs ophthalmology residents: a randomized clinical trial

npj Digital Medicine, Published online: 05 October 2026; doi:10.1038/s41746-026-03232-x

Standardized pre-consultation by a large language model agent vs ophthalmology residents: a randomized clinical trial
  •  

Transketolase-like 1 potentiates PD-1 blockade in hepatocellular carcinoma by glycolysis to prime dendritic cell lactylation

Signal Transduct Target Ther. 2026 Sep 28;11(1):418. doi: 10.1038/s41392-026-02875-2.

ABSTRACT

Hepatocellular carcinoma (HCC) exhibits a suboptimal response to immune checkpoint blockade (ICB) therapy; to overcome this resistance, we aimed to delineate key immune resistance factors via multi-omics analysis, develop strategies to block their immunosuppressive axes, and engineer a targeted nanosystem to enhance immunotherapy efficacy against PD-1 resistance in HCC. Using transcriptomic and proteomic data from anti-PD-1-treated HCC patients, along with functional validation in murine models and mechanistic molecular and cell biology studies, we identified transketolase-like 1 (TKTL1) as a dual-nature biomarker where overexpression predicted poor baseline prognosis yet enhanced response to ICB. Mechanistically, TKTL1 diverts glucose flux into glycolysis rather than pentose phosphate pathway (PPP), recruiting USP9X to deubiquitinate and stabilize HIF-1α, which upregulates HK2 to amplify glycolytic output and lactate accumulation. This metabolic rewiring orchestrates dual immunosuppressive circuits through HIF-1α-driven CCL4 secretion recruiting PD-L1high dendritic cells (DCs), coupled with lactate-induced TRIM28K408 lactylation that stabilizes PD-L1 by blocking ubiquitin-mediated degradation. We engineered a hepatoma-membrane-coated MnO₂ nanosystem (CQLH) co-delivering a TKTL1 inhibitor and lactate oxidase, which disrupted the TKTL1-HIF-1α-HK2 axis, depleted lactate, and reprogrammed the tumor microenvironment, thereby enhanced anti-PD-1 therapy to suppress tumor growth, especially in TKTL1high tumors. These findings define a critical "TKTL1-glycolysis-lactate-DC" axis driving anti-PD-1 sensitivity in HCC, position TKTL1 as both a potential biomarker for ICB response and a tractable therapeutic target, and demonstrate that the targeted CQLH nanosystem overcomes resistance and enhances anti-PD-1 efficacy, offering a precision immunotherapeutic strategy for TKTL1high HCC.

PMID:42802226 | PMC:PMC13616917 | DOI:10.1038/s41392-026-02875-2

  •  

Childhood asthma and the microbiome: from gut-lung axis mechanisms to precision prevention strategies

Front Immunol. 2026 Sep 2;17:1902053. doi: 10.3389/fimmu.2026.1902053. eCollection 2026.

ABSTRACT

Childhood asthma is a highly heterogeneous chronic respiratory disease, and its onset and progression are intricately linked to genetic susceptibility, environmental exposure, immune development, and the establishment of the early-life microbiome. In recent years, studies on the gut and respiratory microbiomes have suggested that the composition, metabolic functions, and interactions of microbial communities with the host immune system may be involved in the formation of asthma susceptibility, shaping of inflammatory phenotypes, and disease progression in children. The gut-lung axis, as an important pathway connecting gut microbiome, respiratory immunity, and systemic inflammatory responses, provides a new perspective for understanding the early mechanisms of childhood asthma. This article reviews the characteristics of the respiratory and gut microbiomes associated with childhood asthma, with a focus on the roles of the gut-lung axis, microbial metabolites, mucosal immune regulation, and environmental exposure. It also evaluates the research progress of probiotics, prebiotics, nutritional interventions, and novel microecological therapies. Additionally, the potential of microbial maturity, microbial metabolites, and immunophenotypes as biomarkers for risk prediction, phenotype stratification, and treatment response is analyzed. Furthermore, the role of multi-omics integration in supporting the identification of responsive populations, matching of intervention strategies, and dynamic monitoring of efficacy is discussed. Current evidence suggests that the microbiome offers promising targets for risk assessment and precision prevention of childhood asthma. However, relevant research still faces challenges such as ambiguous causality, high cohort heterogeneity, limited reproducibility of candidate biomarkers, inconsistent intervention outcomes, and insufficient evidence of long-term safety. At present, most biomarkers and multi-omics models remain in the stage of association discovery, lacking unified thresholds, cross-cohort validation, and biomarker-guided randomized controlled trials in children. Therefore, they cannot be routinely used for patient stratification or intervention selection. Future efforts should rely on standardized longitudinal birth cohorts, multi-omics integration, external validation, and high-quality clinical trials to clarify the incremental value of microbiome biomarkers over traditional clinical indicators and their clinical utility in the individualized management of childhood asthma.

PMID:42751182 | PMC:PMC13580037 | DOI:10.3389/fimmu.2026.1902053

  •  

Artificial Intelligence-Driven Multiomics and Clinical Investigation Identify Macrophage Migration Inhibitory Factor as a Pan-Cancer Biomarker

Phenomics. 2026 May 20;6(3):213-229. doi: 10.1007/s43657-026-00322-4. eCollection 2026 Jun.

ABSTRACT

Early cancer detection remains challenging due to the lack of reliable pan-cancer screening methods, particularly blood-based biomarkers. Using a novel three-tiered validation framework combining artificial intelligence (AI)-powered literature mining of 180,000 PubMed articles (1950-2024), multiomics integration across major databases, and extensive clinical validation, we identified macrophage migration inhibitory factor (MIF) as a promising blood-based biomarker for pan-cancer detection. Multiomics analysis revealed consistent MIF upregulation across 21 cancer types at the transcriptional level and across 12 cancer types at the protein level. Clinical validation in independent cohorts (n = 4,269) showed that serum MIF protein levels discriminated effectively between cancer patients and healthy controls (median AUC = 0.994) and between cancer and benign conditions (median AUC = 0.881). Notably, comparative analyses showed that MIF demonstrated superior or comparable performance to established cancer-specific markers, including AFP for hepatocellular carcinoma (MIF AUC = 0.885 vs. AFP AUC: 0.744-0.887) and CA125 for ovarian cancer (MIF AUC = 0.831 vs. CA125 AUC: 0.58-0.71). Meta-analysis of 28 cohorts (n = 5,347) confirmed the diagnostic efficacy of MIF (pooled AUC: 0.782). This cost-effective, blood-based ELISA approach establishes MIF as a valuable tool for broad applications in cancer screening.

SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s43657-026-00322-4.

PMID:42750739 | PMC:PMC13578188 | DOI:10.1007/s43657-026-00322-4

  •  

Integrated multi-omic profiling enables recurrence risk stratification beyond pathological stage in resected EGFR-mutant lung adenocarcinoma

J Thorac Oncol. 2026 Sep 16:104204. doi: 10.1016/j.jtho.2026.104204. Online ahead of print.

ABSTRACT

BACKGROUND: Early-stage EGFR-mutant lung adenocarcinoma (LUAD) demonstrates heterogeneous outcomes after curative surgery, yet adjuvant treatment decisions are guided by pathological stage alone. Following the ADAURA trial, adjuvant osimertinib is the standard of care for resected stage IB-IIIA EGFR-mutant LUAD; however, real-world data demonstrate that up to 40% of patients remain disease-free at five years without adjuvant osimertinib, underscoring the need for improved risk stratification.

PATIENTS AND METHODS: We performed integrated clinical, genomic and transcriptomic profiling of 400 patients with resected stage IA-IIIA EGFR-mutant LUAD. EGFR-mutant recurrence risk models integrating clinical, genomic and transcriptomic data were developed and validated across one internal and three external cohorts.

RESULTS: Genomic instability, including TP53 co-mutations, copy number alterations and APOBEC-associated mutational signatures, increased with pathological stage. RBM10 co-mutations were enriched in tumours with L858R mutations and correlated with upregulation of WNT signalling and epithelial-mesenchymal transition. Transcriptomic features outperformed clinical or genomic variables alone in predicting recurrence risk, and a multi-omic model demonstrated superior and reproducible performance, achieving a median concordance index of 75.4% across four independent validation cohorts. The multi-omic model stratified recurrence risk within individual pathological stages, including stage I disease, and identified patients most likely to benefit from adjuvant EGFR TKI.

CONCLUSIONS: These findings define the molecular heterogeneity of early-stage EGFR-mutant LUAD and support multi-omic risk stratification to inform adjuvant EGFR TKI decisions beyond pathological stage. Prospective validation in larger cohorts will be required to confirm these findings.

PMID:42749051 | DOI:10.1016/j.jtho.2026.104204

  •  

Factor IX Padua AAV gene therapy in adolescents with hemophilia B: a phase 1 trial

Nature Medicine, Published online: 16 September 2026; doi:10.1038/s41591-026-04636-8

In this single-arm phase 1 trial, an AAV gene therapy carrying the Padua variant of factor IX was well tolerated in 11 adolescents with hemophilia B and led to reductions in annualized bleeding rate.
  •  

Advancing cancer detection and treatment using longitudinal routine clinical data

Liu et al. develop Oncoformer, a multimodal transformer that reads routine laboratory tests and chest X-rays already collected in everyday care. Across more than 3.6 million individuals, it detects cancer, infers tumor stage, and stratifies treatment response and recurrence risk, pointing toward risk-adapted cancer care built on data already in hand.
  •  

Spatial proximity sequencing maps developmental dynamics in the germinal center

Sprox-seq enables spatial profiling of protein complexes, surface proteins, and mRNAs in intact tissues by combining proximity ligation with spatial transcriptomics. In human tonsils, Sprox-seq maps germinal center interaction networks, links CD21-CD35 complexes to proliferative programs, reveals interaction-based B cell state transitions, and directly captures B cell-follicular dendritic cell communication.
  •  

Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work

arXiv:2609.11977v1 Announce Type: new Abstract: Co-work agents execute complex workflows that combine information gathering, tool use, coding, and file manipulation across many model invocations. Because cost and latency accumulate over the full episode, their practical value depends not only on peak capability but also on how efficiently that capability is delivered. Yet many steps in everyday work emphasize state tracking, coordination, recovery, and follow-through rather than frontier-scale reasoning. We present Occamy-1.0, a cost-efficient co-work model obtained by further training the post-trained Qwen3.6-35B-A3B checkpoint. We construct execution-grounded data and environments, capture replayable long-horizon trajectories across multiple harnesses, and use staged post-training to develop and consolidate complementary execution capabilities. Across a broad suite of co-work benchmarks, Occamy-1.0 is consistently among the strongest comparably sized models and remains competitive with substantially larger frontier systems on several tasks. Under our stated evaluation and pricing protocol, its aggregate performance across four representative benchmarks places it at the low-cost knee of the observed cost--performance Pareto frontier. Supporting evaluations in tool calling, coding, and instruction following further show that this specialization preserves broad agentic capability. We release the model weights and a subset of the training data to support research on practical co-work agents and agentic post-training.
  •  

Toward Robust Personalized Alignment for LLMs: Mitigating Persona Drift in Multi-Turn Dialogue

arXiv:2609.12373v1 Announce Type: new Abstract: Persona drift remains a central challenge for personalized language models, as user profiles evolve over long interactions rather than remain permanently fixed. Models must therefore revise persistent persona states when preferences genuinely change, while avoiding updates driven by transient, ambiguous, or unresolved observations. We propose CORE, which separates turn-local evidence from persistent persona-state revision and selectively updates grounded user preferences through uncertainty-aware belief revision. We also introduce PERSIST, a held-out post-anchor benchmark for persona-state robustness under sequential interaction stress, covering ambiguity, conflict, and controlled social influence. Across ALOE, PersonaChat, and PERSIST, CORE improves personalized alignment and robustness, with complementary gains in normalized closed-slot state fidelity. Human evaluation and mechanistic controls further support explicit update control beyond stronger generation or persistent memory alone.
  •  

VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgets

arXiv:2609.12404v1 Announce Type: new Abstract: Learning from trial and error is a promising way to improve language agents on complex tasks such as computer control. Reflexion introduced verbal reinforcement learning, which turns failed trials into text that guides later attempts without updating model parameters. We introduce VRL-Bench, a harness for fair evaluation of trial-and-error learning under finite trial budgets. Across three models on MiniWoB and WebShop, we evaluate updates from several prominent verbal-memory methods spanning Reflexion and later work: each improves observed success over memory-free retry in some settings but reduces it in others. Replay experiments show that using reflection can reduce success rates, revealing a trade-off between exploiting experience and continued exploration. We propose VEX$^2$, a verbal exploration--exploitation scheduler that uses a language model to jointly select policies and allocate the remaining trial budget. VEX$^2$ is the only evaluated update to achieve positive observed success-rate gains over retry in all six settings.
  •  

EvoRS: On-Policy Self-Evolution of Reward Systems for Open-Ended Reinforcement Learning

arXiv:2609.12459v1 Announce Type: new Abstract: Open-ended reinforcement learning often relies on rubric-based rewards for tasks without directly verifiable answers. Yet the policy and reward system form a dynamic feedback loop: as the policy optimizes the current reward, an initially useful reward system may become unreliable due to reward hacking or reduced response discriminability. The reward system should therefore evolve rather than remain fixed during training. Existing dynamic-rubric methods adapt evaluation criteria, but reward failures can also arise from scoring mechanisms or signal composition. We introduce EvoRS, a self-evolving RL framework that evolves the reward system from on-policy experience, representing it as an executable Reward-DAG. Specifically, an agentic designer updates this system from on-policy rollouts and reward traces to maintain train-time reliability. Across writing and roleplay, EvoRS achieves the best quality under all three judges, outperforming the policy by \(2.107\) and \(4.767\) points, respectively, while reducing reward hacking and coverage failures and preserving reward informativeness. Ablations confirm that a comprehensive fixed reward system cannot remain reliable in open-ended tasks and must evolve throughout training.
  •  

When Does AI Augment Work? A Workflow-Level Framework for Human-Agent Collaboration

arXiv:2609.12482v1 Announce Type: new Abstract: We aim to characterise the value of artificial intelligence in the workplace. Current studies largely measure this value in terms of the current automation capabilities and public adoption of AI. However, such metrics ignore the greater impacts of human--agent collaboration in transforming the nature of work. To account for this, we must expand the scope of our analysis beyond atomised tasks of today, and instead focus on how AI can augment entire workflows of the future. To ground this analysis, we establish a precise definition of AI augmentation comprising six conditions, spanning durable net value, meaningful human control, accountability and recovery, and long-term human development through learning, career pathways, and job purpose. We elaborate on these conditions and apply the framework in a case study of AI-mediated social surveys. We conclude by outlining how organisations, researchers, and government leaders can use this framework to make sense of the future of work.
  •  

Tracing and Coordinating Cross-Layer Influence for Multimodal Model Merging

arXiv:2609.12897v1 Announce Type: new Abstract: Multimodal model merging aims to consolidate task experts into a single model that retains their complementary capabilities. Most unimodal model merging methods combine expert updates within individual layers, and multimodal approaches largely follow this design. However, an expert update changes the representations passed to subsequent layers, allowing its influence to propagate across depth and affect how visual and textual information interact. When visual and language updates are combined, later updates act on inputs already modified by earlier ones, coupling their effects. This poses two challenges: (1) how to characterize the multimodal influence of individual expert updates across depth, and (2) how to jointly combine expert updates based on their multimodal influence. To address these challenges, we propose TAC-Merge for tracing and coordinating cross-layer influence in multimodal model merging. It contains two modules, i.e., multimodal influence mapping (MIM) and coupled merge control (CMC). MIM constructs graphs of update effects and uses Ricci curvature together with expert predictions to define a shared fusion objective. CMC models interactions among coefficient adjustments and jointly optimizes regional weights to synthesize one shared model. Experiments across diverse multimodal tasks demonstrate the effectiveness of TAC-Merge in consolidating complementary expert capabilities and supporting generalization to unseen tasks.
  •  

How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks

arXiv:2609.13009v1 Announce Type: new Abstract: Low reported scores on leading physics benchmarks, including those featured in the Artificial Analysis Intelligence Index (2026), suggest that frontier language models still struggle with advanced physics, a demanding test of their scientific reasoning and quantitative problem-solving abilities. Yet this impression does not always align with domain experts' experiences using these models in their work. We revisit these reported findings by evaluating frontier models on six widely used physics benchmarks and auditing them with experts, focusing on text-only problems with verifiable final answers. For each subfield of physics, faculty and graduate researchers with relevant expertise carefully review problem statements, reference solutions, and model responses to distinguish genuine model errors from grader errors, incorrect reference solutions, and ambiguous or underspecified questions. Most audited cases initially evaluated as incorrect reflect these benchmarking issues rather than errors in the models' physics reasoning. We then ask experts to address these benchmarking issues by correcting erroneous reference solutions and repairing or excluding flawed questions. We find that GPT-5.6-Sol's measured mean@4 rises from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark, while its corrected pass@4 reaches 94.4% on the 54 retained CritPt challenges. Corrected scores are computed on the retained evaluation subsets following expert review. Scores on the audited subsets of UGPhysics, PRISM-Physics, and PHYBench also rise substantially after correction. These findings suggest that current benchmarks substantially understate frontier models' ability to solve well-posed physics problems. Near-saturation on these closed-ended tasks highlights the need for more demanding, expert-validated evaluations.
  •  

Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction

arXiv:2609.13082v1 Announce Type: new Abstract: Agentic systems offer a promising way to automate embodied benchmark construction, but existing approaches typically cover isolated stages or remain specialized to predefined environments and task families. More importantly, multi-step construction produces dependent intermediate artifacts that are often passed downstream without artifact-specific verification, allowing local defects to propagate into the final benchmark. We present Embodied-BenchForge, an agentic framework that transforms user-specified evaluation intents into complete embodied benchmark artifacts. It formulates construction as Closed-Loop Benchmark Synthesis, integrating forward artifact synthesis with backward verification and repair. Skill-Orchestrated Artifact Synthesis composes typed and reusable skills into executable workflows, while an artifact dependency graph records intermediate outputs and their dependencies. Requirement-Guided Verification and Repair applies artifact-specific contracts throughout construction and uses provenance to trigger local re-execution or upstream rollback when verification fails. Embodied-BenchForge constructs six benchmarks covering diverse embodied scenarios in the Offline EQA Track, together with one interactive benchmark containing 220 executable tasks in the Interactive Embodied Track. Evaluations of representative MLLMs and embodied agents show that the benchmarks distinguish model capabilities in both observation-based understanding and closed-loop execution. Quality assessment and ablations validate benchmark quality and the effectiveness of verification and repair, while repair and skill-reuse analyses demonstrate efficient localized recovery and cross-benchmark reusability.
  •  

PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews

arXiv:2609.11559v1 Announce Type: cross Abstract: Large language models (LLMs) and AI-enabled software increasingly participate in systematic-review decisions, yet the information needed to audit these workflows is reported inconsistently. We analyze SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations, to characterize changes in methods, review-stage use, evaluation and reported limitations. Automation has shifted toward LLM- and software-facing workflows, including stages that can alter the evidence base. Since 2023, 38.0% of software/product papers reported no evaluation, compared with 9.3% of LLM papers. Reporting coverage increased with LLM workflow complexity, yet 52% of positive-only LLM evaluations still reported an unmet reliability or performance requirement. From these patterns, we introduce PRISMA-LLM, an empirically grounded framework separating implementation disclosure from consequence-sensitive evaluation and limitation reporting.
  •  

The Anatomy and Boundary of Adaptation under Temporal Tabular Shift

arXiv:2609.12136v1 Announce Type: cross Abstract: Prequential adaptation of frozen tabular foundation models under temporal drift, with each label revealed only after prediction, helps some deployments and harms others, yet current practice does not predict which. We study the sources and limits of these gains. A diagnostic anatomy attributes gains to four recurring mechanisms under a streaming protocol that removes three optimistic biases and quantifies a fourth. Within an agnostic total-variation drift class, the target conditional is only partially identified: its identified-set diameter, the \emph{wall}, is irreducible from unlabeled data uniformly in sample size. A second, orthogonal $L^2$ projection wall quantifies what the frozen representation cannot express. Two canonical mechanism priors collapse the first wall. Under stated nuisance-rate conditions, the wall can be estimated from labeled historical windows at a $\sqrt N$ rate above the margin threshold $\gamma^\star=d_0/(2\alpha_s)$. At $\gamma=0$, the conditional lower-bound program depends on an open affinity estimate; the positive-margin lower branch also remains open. Semi-synthetic data illustrate the finite-sample mechanism with calibrated exponents. Stream-level proxies on eight industrial streams fall on the difficult side under a stated roughness bound, while the equality case $\gamma=\gamma^\star$ remains unresolved.
  •  

Neural Multichannel Distant Speaker Diarization with Heavy-tailed Source Separation Model

arXiv:2609.12154v1 Announce Type: cross Abstract: Distant speaker diarization remains challenging due to difficult acoustic environments, varying numbers of speakers and overlapping speech. Model-driven methods are proposed to exploit the speech source features in multi-channel recordings that help diarization. This paper generalizes a neural model that jointly learns to perform blind source separation and diarization over speech mixtures (neural FCASA) with heavy-tailed models. The popular Gaussian distribution has been applied for variance modeling in the original source separation model, which we replace with two families of heavy-tailed models (Leptokurtic Generalized Gaussian distribution and Student's t distribution) to better capture the heavy-tailedness in speech signals. Thanks to the Gaussian scale mixture model, we are able to unify the proposed method and the original one under the same form of learning objective. Our experiments show consistent large improvements in Diarization Error Rate (DER) and Jaccard Error Rate (JER) compared to the baseline on various corpora.
  •  
❌