❌

Normal view

Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC

Nature Medicine, Published online: 13 September 2026; doi:10.1038/s41591-026-04488-2

In a large international real-world study of non-small cell lung cancer, a multimodal explainable AI model outperformed established biomarkers for immunotherapy outcome prediction and improved physician decision-making.

Teclistamab versus lenalidomide-dexamethasone in high-risk smoldering multiple myeloma: a randomized phase 2 trial

Nature Medicine, Published online: 11 September 2026; doi:10.1038/s41591-026-04642-w

In the randomized phase 2 ImmunoPRISM trial, patients with high-risk smoldering multiple myeloma (MM) showed higher rates of complete clinical responses in response to treatment with teclistamab compared with lenalidomide–dexamethasone, although longer follow-up is required to determine durable prevention of progression to MM.

Computable longitudinal patient journeys from structured and unstructured EHR data

Nature Medicine, Published online: 10 September 2026; doi:10.1038/s41591-026-04695-x

A suite of large pre-trained language models accurately extracts computable clinical data from unstructured electronic health records and integrates the findings into knowledge graphs that can facilitate understanding patient trajectories and treatment responses in the real world.

“Small” Large Language Models in the Hospital: Evaluation Study on Real-World Data in a Resource-Constrained Setting

Background: Large language models (LLMs) are increasingly being deployed in health care, but their use and deployment in many real-world hospital environments pose significant challenges and concerns. In particular, state-of-the-art commercial models store or process data externally, which is often in conflict with ensuring patient data protection. At the same time, using LLMs locally is limited by the lack of available computing infrastructure. Small open-source LLMs that do not require substantial computing resources could offer a practical way to resolve these tensions, but their medical utility in real-world local contexts, especially in non-English languages, has not been sufficiently evaluated. Objective: This study aimed to evaluate the feasibility of small, locally deployable open-source LLMs for clinically relevant tasks in a resource-constrained hospital setting and to propose a reproducible framework for institution-specific evaluation before deployment. Methods: We evaluated 6 open-source LLMs ranging from 8B to 24B parameters (from the Mistral, Phi4, Falcon3, Llama3.1, and Meditron3 families) in a zero-shot setting across 7 tasks covering 4 clinical use cases: information extraction, medical text translation, text generation, and clinical decision support. We used deidentified French clinical data from a Swiss tertiary hospital, including discharge letters, clinical notes, and structured electronic health records. Performance was assessed using task-specific metrics, such as precision, recall, F1-score, embedding-based semantic similarity, recall-oriented understudy for gisting evaluation (ROUGE) score, readability indices, and human review by clinicians. Results: Model performance varied substantially between tasks. In the simplest retrieval task (needle-in-the-haystack), several models performed strongly, with Llama3.1 achieving an F1-score of 99.81% and Mistral-small achieving 99.71%. In contrast, performance was poor in more complex tasks. For detecting protected health information, the best-performing LLMs achieved only modest overall macro–F1-scores (0.33-0.34), substantially below a fine-tuned Robustly Optimized BERT Pretraining Approach (RoBERTa) baseline (0.94). In the task of extracting immune-related adverse events from discharge notes, the highest overall macro–F1-score was 0.35 with Phi4. For medical text translation, Phi4 ranked highest in embedding-based evaluation, whereas Meditron3-Phi4 performed the worst, with clinician reviews identifying hallucinations in 55% of its outputs. In the task of summarizing discharge letters, quality was low across all models, with the best penalized ROUGE score reaching only 0.169 with Llama3.1. In the tasks of generating patient-friendly discharge note summaries and clinical decision support, clinician ratings generally ranged from dissatisfied to neutral, and no model achieved consistently satisfactory performance. Conclusions: Small open-source LLMs appear feasible for simple retrieval-oriented tasks in local hospital deployments but are currently inadequate for more complex applications, such as clinical decision support, deidentification, extraction of adverse events, and medical summarization. These findings highlight the importance of locally grounded evaluation tailored to specific use cases and the need for robust institutional evaluation frameworks to ensure safe and reliable deployment. Trial Registration:

Neocortical long-range inhibition promotes cortical synchrony and sleep

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-10876-y

In mice, a sparse population of sleep-active long-range inhibitory neurons in the neocortex promote widespread cortical synchronization and sleep, revealing a cortical mechanism that contributes to the regulation of sleep.

Denisovans from southwestern China and their subsistence strategies

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-10997-4

Evidence from Bianfu Cave shows specialized hunting, expedient stone-tool production and extensive bone use of Denisovans, providing new insights into their ecology, behaviour and cultural legacy in eastern Asia.

A serpin–myeloid axis in pancreatic cancer heterogeneity and immune evasion

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-11002-8

SERPINE1 and SERPINB2-driven fibrin-rich niches locally programme immunosuppressive macrophages and exclude T cells, enabling spatially organized immune evasion in pancreatic ductal carcinoma.

TRI-611, a selective, brain-penetrant molecular glue degrader of ALK

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-10998-3

TRI-611 induces the degradation of ALK fusion proteins via a previously undescribed CRBN recruitment motif, and its preclinical anti-tumour activity highlights TRI-611 as a potential new way of treating ALK-positive non-small-cell lung cancer.

Proximity-guided graph learning reveals tumour-associated proximity antigens

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-11003-7

A proximity-mapping atlas defines tumour-associated proximity antigens, revealing disease-associated membrane spatial communities, and identifies EGFR–CDCP1 as a co-target pair that enhances tumour killing by multispecific therapeutics.

Consensus framework for the validation of generative AI: call for collaborators on the Validation Accords

Nature Medicine, Published online: 09 September 2026; doi:10.1038/s41591-026-04647-5

Consensus framework for the validation of generative AI: call for collaborators on the Validation Accords

Inhaled siRNA therapy targeting RAGE for pulmonary inflammation: a first-in-human randomized trial

Nature Medicine, Published online: 09 September 2026; doi:10.1038/s41591-026-04607-z

Following preclinical development, a randomized first-in-human trial showed that treatment with an inhaled, lung-epithelium-targeted siRNA directed against the receptor for advanced glycation end-products (RAGE) was safe and well tolerated and reduced RAGE levels in serum and bronchoalveolar lavage fluid.

Myeloperoxidase inhibition with mitiperstat in heart failure with preserved or mildly reduced ejection fraction: a randomized phase 2b trial

Nature Medicine, Published online: 09 September 2026; doi:10.1038/s41591-026-04615-z

In a phase 2 randomized clinical trial, treatment with the myeloperoxidase (MPO) inhibitor mitiperstat, intended to target the neutrophil–MPO inflammatory pathway, did not improve symptoms or exercise function in individuals with heart failure with preserved or mildly reduced ejection fraction.

Extended Reality Interventions for Osteoarthritis of the Knee and Recovery After Total Knee Arthroplasty: Systematic Review and Meta-Analyses

Background: Nonpharmacologic interventions are important for treating knee pain due to osteoarthritis or after total knee arthroplasty (TKA), and extended reality (XR) technology may enhance treatments for these indications. Objective: This systematic review aimed to evaluate XR interventions for pain due to knee osteoarthritis (KOA) or for recovery after TKA. Methods: Databases were searched through May 2023 and updated in December 2025. Eligible trials evaluated XR interventions to treat KOA pain or after TKA. We classified interventions by depth of immersion and clinical mechanism. We used the Grading of Recommendations Assessment, Development, and Evaluation (GRADE) criteria to determine the certainty of evidence for prioritized outcomes. Meta-analyses were performed when ≥3 studies evaluated similar comparisons, outcomes, and time points. Results: Eligible trials addressed KOA (k=12) or recovery after TKA (k=9). Sample sizes ranged from 36 to 306 participants, and most studies had a follow-up of ≤3 months. Nineteen studies assessed pain-related functioning and pain intensity, and 5 assessed adverse events (AEs). For KOA, 10 studies examined interactive digital rehabilitation (IDR), and 2 examined virtual reality (VR)–digitally augmented exercise (DAE). IDR for KOA may result in better pain-related functioning (low certainty of evidence [COE]; pooled standardized mean difference [SMD] −0.59, 95% CI −1.11 to −0.06; prediction interval [PI] −1.72 to 0.55; k=5) and lower pain intensity at 6‐8 weeks (low COE; pooled SMD −0.46, 95% CI −0.92 to 0.00; PI −1.39 to 0.47; k=4). VR-DAE for KOA (k=2) produced inconsistent results (very low COE). For post-TKA studies, 5 examined IDR, 2 examined VR-DAE, 1 examined VR-distraction, and 1 examined VR-psychoeducation. Post-TKA IDR may result in better pain-related functioning (low [k=4] and moderate COE [k=1]) but little to no difference in pain intensity (low-moderate COE; pooled SMD at 3‐4 months −0.12, 95% CI −0.75 to 0.52; PI –1.63 to 1.27; k=3). VR-psychoeducation probably results in lower pain at 4 weeks (moderate COE; k=1), and VR-distraction may result in 6 months (low COE; k=1), whereas VR-DAE produced mixed findings (k=2; very low COE). IDR was not associated with AEs, and VR may not be associated with AEs for KOA (high and low COE), though AE reporting was uncommon (k=5) and evidence was very uncertain for post-TKA. Conclusions: IDR may augment treatment for KOA and post-TKA recovery, and VR may benefit post-TKA rehabilitation. This review is the first to stratify by level of immersion, clinical mechanism, and follow-up duration and to systematically evaluate AEs. IDR may be ready for integration into KOA care, while use after TKA needs more evidence. Randomized controlled trials with implementation outcomes could determine how XR interventions can be used for KOA, whereas trials evaluating efficacy and AEs are needed before their use for post-TKA. Trial Registration: PROSPERO CRD42023439903; https://www.crd.york.ac.uk/PROSPERO/view/CRD42023439903

Identifying potential nonpulmonary vein triggers in persistent atrial fibrillation using digital twins and deep learning

npj Digital Medicine, Published online: 08 September 2026; doi:10.1038/s41746-026-03223-y

Identifying potential nonpulmonary vein triggers in persistent atrial fibrillation using digital twins and deep learning
❌