❌

Reading view

World models for biomedicine

Biological processes continuously adapt in response to intervention. Unlike static predictive models, biomedical world models support action-conditioned simulations for counterfactual reasoning, intervention design, and sequential planning. This perspective defines the key properties of biomedical world models and discusses the data, modeling, and evaluation challenges required to build them across biological and clinical scales.
  •  

Tahoe-100M: Mapping drug-induced molecular phenotypes at single-cell resolution

Tahoe-100M is an atlas of 100 million single-cell transcriptomes, capturing how 50 cancer cell lines respond to ∼1,100 drug-dose treatments. By pairing single-cell and molecular phenotypes at scale, the resource links drug mechanisms to cellular responses and provides an openly available substrate for training predictive models of cell behavior.
  •  

Broadly neutralizing antibodies in adult males living with HIV undergoing analytical treatment interruption: secondary and exploratory outcomes of the phase II randomized controlled RIO trial

Nature Medicine, Published online: 15 September 2026; doi:10.1038/s41591-026-04644-8

In the phase 2 RIO trial, there was delayed viral rebound and resistance to broadly neutralizing antibodies 3BNC117-LS and 10-1074-LS in adult males living with HIV undergoing analytical treatment interruption, and initial reservoir sensitivity to autologous antibodies was associated with a longer time to rebound.
  •  

The intrinsic cardiac nervous system is essential for cardiac function and survival

Two genetically distinct neuronal subtypes in the intrinsic cardiac nervous system are differentially required for baseline cardiac function and stress resilience to maintain cardiac stability and survival.
  •  

Anonymization of Portuguese Clinical Notes Using Large Language Models and Quantum-Enhanced Hybrid Architectures: Comparative Evaluation Study

Background: The widespread adoption of electronic health records (EHRs) has generated large-scale repositories of highly sensitive clinical information, emphasizing the need for robust anonymization strategies to enable secondary use for research while safeguarding patient privacy. Conventional rule-based and machine learning approaches for deidentifying medical text face limitations with the linguistic complexity, variability, and context dependence inherent to clinical documentation. Recent advances in large language models (LLMs), combined with emerging quantum computing paradigms, present novel opportunities to enhance the accuracy, scalability, and resilience of health care data anonymization. Objective: This study aims to evaluate the efficacy of LLM-based and quantum-enhanced hybrid architectures for medical text anonymization, assessing the effectiveness and computational efficiency across multiple entity types in Portuguese clinical notes. Methods: We constructed a gold-standard corpus of 1000 Portuguese outpatient clinical notes, manually annotated by 5 trained researchers for 5 protected-entity categories: patient names, dates, identifiers, organizations, and geographic locations. Four anonymization strategies were evaluated: 2 stand-alone LLMs (Llama-3.1-8B-instruct and Llama-3.3-70B-instruct) and 2 quantum-enhanced hybrid models (Dynex-QML with 8B and 70B base models) incorporating quantum optimization via Quadratic Unconstrained Binary Optimization (QUBO) formulations. The quantum-enhanced approach transforms the final attention layer of the LLM into a global constraint satisfaction problem solved via neuromorphic quantum annealing. Model performance was measured on a held-out test set of 500 notes using precision, recall, and -score metrics. Computational efficiency was quantified through end-to-end processing time. Results: The quantum-enhanced Dynex-QML-70B model achieved the highest overall performance with a macro-score of 0.855 (95% CI 0.823‐0.880), outperforming the stand-alone Llama-3.3-70B (0.726, 95% CI 0.704‐0.747), Dynex-QML-8B (0.733, 95% CI 0.709‐0.756), and Llama-3.1-8B (0.602, 95% CI 0.588‐0.615). Compared with Llama 3.3 70B, Dynex-QML (Llama 70B) improved macro-score by 0.128 (95% CI 0.091‐0.163; empirical 2-sided bootstrap
  •  
  •  
  •  

Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC

Nature Medicine, Published online: 13 September 2026; doi:10.1038/s41591-026-04488-2

In a large international real-world study of non-small cell lung cancer, a multimodal explainable AI model outperformed established biomarkers for immunotherapy outcome prediction and improved physician decision-making.
  •  

Smartphone-Based Monitoring of Quality of Life and Adverse Events After Neurosurgery: Prospective Cohort Study

Background: Postoperative outcome assessment is often based on discrete follow-up visits, limiting characterization of individual recovery trajectories, and the timely identification of adverse events (AEs). Longitudinal smartphone-based monitoring may overcome these limitations by enabling frequent, resource-efficient collection of patient-reported outcomes and complications throughout recovery. Such data may provide a more patient-centered understanding of the postoperative course and complement conventional clinical surveillance. Objective: This study aimed to evaluate the feasibility of smartphone-based longitudinal monitoring of quality of life, subjective well-being, and AEs after elective neurosurgery and compare postoperative recovery trajectories and agreement between patient- and clinician-reported AEs. Methods: This interim analysis of a prospective cohort study included adult patients undergoing elective lumbar decompression, lumbar fusion, supratentorial craniotomy, or infratentorial craniotomy at a Swiss tertiary referral center between June 2023 and January 2025. Participants used a smartphone app to longitudinally report subjective well-being (Subjective Well-Being Index; 0‐10), quality of life (EQ-5D-5L), and AEs for up to 1 year postoperatively. Complications were self-reported using the Therapy-Disability-Neurology (TDN) classification and retrospectively adjudicated by physicians. Descriptive analyses assessed data density, engagement, and concordance between patient- and clinician-reported events. Mixed-effects models were used to evaluate factors associated with postoperative well-being. Results: Of the 100 enrolled patients (median age 64.0, IQR 52.95‐71.6 years; n=45, 45% women), 86 (86%) provided postoperative data. During a median follow-up of 3.2 (IQR 0.2‐11.3) months, participants submitted 4354 longitudinal well-being entries. Patients reported 22 unique AEs, whereas physicians identified 44 AEs, with overlap for 9 (20.5%) events. Most physician-reported AEs were mild (30/44, 68.2%; TDN grade 1‐2), and no grade 4 or 5 events occurred. Patient-reported AEs primarily reflected symptomatic and functional impairments, whereas physician-reported events more often included clinically detected or subclinical findings. In mixed-effects models, time since surgery was associated with improved well-being, and no other factors were statistically significant. Conclusions: Smartphone-based postoperative monitoring was feasible in this elective neurosurgical cohort and generated dense longitudinal patient-reported data beyond routine follow-up. Patient and clinician AE reporting captured partly distinct aspects of postoperative recovery, suggesting that smartphone-based self-reporting may complement rather than replace clinical surveillance. Trial Registration: ClinicalTrials.gov NCT06352710; https://clinicaltrials.gov/study/NCT06352710
  •  

“Small” Large Language Models in the Hospital: Evaluation Study on Real-World Data in a Resource-Constrained Setting

Background: Large language models (LLMs) are increasingly being deployed in health care, but their use and deployment in many real-world hospital environments pose significant challenges and concerns. In particular, state-of-the-art commercial models store or process data externally, which is often in conflict with ensuring patient data protection. At the same time, using LLMs locally is limited by the lack of available computing infrastructure. Small open-source LLMs that do not require substantial computing resources could offer a practical way to resolve these tensions, but their medical utility in real-world local contexts, especially in non-English languages, has not been sufficiently evaluated. Objective: This study aimed to evaluate the feasibility of small, locally deployable open-source LLMs for clinically relevant tasks in a resource-constrained hospital setting and to propose a reproducible framework for institution-specific evaluation before deployment. Methods: We evaluated 6 open-source LLMs ranging from 8B to 24B parameters (from the Mistral, Phi4, Falcon3, Llama3.1, and Meditron3 families) in a zero-shot setting across 7 tasks covering 4 clinical use cases: information extraction, medical text translation, text generation, and clinical decision support. We used deidentified French clinical data from a Swiss tertiary hospital, including discharge letters, clinical notes, and structured electronic health records. Performance was assessed using task-specific metrics, such as precision, recall, F1-score, embedding-based semantic similarity, recall-oriented understudy for gisting evaluation (ROUGE) score, readability indices, and human review by clinicians. Results: Model performance varied substantially between tasks. In the simplest retrieval task (needle-in-the-haystack), several models performed strongly, with Llama3.1 achieving an F1-score of 99.81% and Mistral-small achieving 99.71%. In contrast, performance was poor in more complex tasks. For detecting protected health information, the best-performing LLMs achieved only modest overall macro–F1-scores (0.33-0.34), substantially below a fine-tuned Robustly Optimized BERT Pretraining Approach (RoBERTa) baseline (0.94). In the task of extracting immune-related adverse events from discharge notes, the highest overall macro–F1-score was 0.35 with Phi4. For medical text translation, Phi4 ranked highest in embedding-based evaluation, whereas Meditron3-Phi4 performed the worst, with clinician reviews identifying hallucinations in 55% of its outputs. In the task of summarizing discharge letters, quality was low across all models, with the best penalized ROUGE score reaching only 0.169 with Llama3.1. In the tasks of generating patient-friendly discharge note summaries and clinical decision support, clinician ratings generally ranged from dissatisfied to neutral, and no model achieved consistently satisfactory performance. Conclusions: Small open-source LLMs appear feasible for simple retrieval-oriented tasks in local hospital deployments but are currently inadequate for more complex applications, such as clinical decision support, deidentification, extraction of adverse events, and medical summarization. These findings highlight the importance of locally grounded evaluation tailored to specific use cases and the need for robust institutional evaluation frameworks to ensure safe and reliable deployment. Trial Registration:
  •  
  •  
❌