❌

Normal view

Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC

Nature Medicine, Published online: 13 September 2026; doi:10.1038/s41591-026-04488-2

In a large international real-world study of non-small cell lung cancer, a multimodal explainable AI model outperformed established biomarkers for immunotherapy outcome prediction and improved physician decision-making.

Smartphone-Based Monitoring of Quality of Life and Adverse Events After Neurosurgery: Prospective Cohort Study

Background: Postoperative outcome assessment is often based on discrete follow-up visits, limiting characterization of individual recovery trajectories, and the timely identification of adverse events (AEs). Longitudinal smartphone-based monitoring may overcome these limitations by enabling frequent, resource-efficient collection of patient-reported outcomes and complications throughout recovery. Such data may provide a more patient-centered understanding of the postoperative course and complement conventional clinical surveillance. Objective: This study aimed to evaluate the feasibility of smartphone-based longitudinal monitoring of quality of life, subjective well-being, and AEs after elective neurosurgery and compare postoperative recovery trajectories and agreement between patient- and clinician-reported AEs. Methods: This interim analysis of a prospective cohort study included adult patients undergoing elective lumbar decompression, lumbar fusion, supratentorial craniotomy, or infratentorial craniotomy at a Swiss tertiary referral center between June 2023 and January 2025. Participants used a smartphone app to longitudinally report subjective well-being (Subjective Well-Being Index; 0‐10), quality of life (EQ-5D-5L), and AEs for up to 1 year postoperatively. Complications were self-reported using the Therapy-Disability-Neurology (TDN) classification and retrospectively adjudicated by physicians. Descriptive analyses assessed data density, engagement, and concordance between patient- and clinician-reported events. Mixed-effects models were used to evaluate factors associated with postoperative well-being. Results: Of the 100 enrolled patients (median age 64.0, IQR 52.95‐71.6 years; n=45, 45% women), 86 (86%) provided postoperative data. During a median follow-up of 3.2 (IQR 0.2‐11.3) months, participants submitted 4354 longitudinal well-being entries. Patients reported 22 unique AEs, whereas physicians identified 44 AEs, with overlap for 9 (20.5%) events. Most physician-reported AEs were mild (30/44, 68.2%; TDN grade 1‐2), and no grade 4 or 5 events occurred. Patient-reported AEs primarily reflected symptomatic and functional impairments, whereas physician-reported events more often included clinically detected or subclinical findings. In mixed-effects models, time since surgery was associated with improved well-being, and no other factors were statistically significant. Conclusions: Smartphone-based postoperative monitoring was feasible in this elective neurosurgical cohort and generated dense longitudinal patient-reported data beyond routine follow-up. Patient and clinician AE reporting captured partly distinct aspects of postoperative recovery, suggesting that smartphone-based self-reporting may complement rather than replace clinical surveillance. Trial Registration: ClinicalTrials.gov NCT06352710; https://clinicaltrials.gov/study/NCT06352710

“Small” Large Language Models in the Hospital: Evaluation Study on Real-World Data in a Resource-Constrained Setting

Background: Large language models (LLMs) are increasingly being deployed in health care, but their use and deployment in many real-world hospital environments pose significant challenges and concerns. In particular, state-of-the-art commercial models store or process data externally, which is often in conflict with ensuring patient data protection. At the same time, using LLMs locally is limited by the lack of available computing infrastructure. Small open-source LLMs that do not require substantial computing resources could offer a practical way to resolve these tensions, but their medical utility in real-world local contexts, especially in non-English languages, has not been sufficiently evaluated. Objective: This study aimed to evaluate the feasibility of small, locally deployable open-source LLMs for clinically relevant tasks in a resource-constrained hospital setting and to propose a reproducible framework for institution-specific evaluation before deployment. Methods: We evaluated 6 open-source LLMs ranging from 8B to 24B parameters (from the Mistral, Phi4, Falcon3, Llama3.1, and Meditron3 families) in a zero-shot setting across 7 tasks covering 4 clinical use cases: information extraction, medical text translation, text generation, and clinical decision support. We used deidentified French clinical data from a Swiss tertiary hospital, including discharge letters, clinical notes, and structured electronic health records. Performance was assessed using task-specific metrics, such as precision, recall, F1-score, embedding-based semantic similarity, recall-oriented understudy for gisting evaluation (ROUGE) score, readability indices, and human review by clinicians. Results: Model performance varied substantially between tasks. In the simplest retrieval task (needle-in-the-haystack), several models performed strongly, with Llama3.1 achieving an F1-score of 99.81% and Mistral-small achieving 99.71%. In contrast, performance was poor in more complex tasks. For detecting protected health information, the best-performing LLMs achieved only modest overall macro–F1-scores (0.33-0.34), substantially below a fine-tuned Robustly Optimized BERT Pretraining Approach (RoBERTa) baseline (0.94). In the task of extracting immune-related adverse events from discharge notes, the highest overall macro–F1-score was 0.35 with Phi4. For medical text translation, Phi4 ranked highest in embedding-based evaluation, whereas Meditron3-Phi4 performed the worst, with clinician reviews identifying hallucinations in 55% of its outputs. In the task of summarizing discharge letters, quality was low across all models, with the best penalized ROUGE score reaching only 0.169 with Llama3.1. In the tasks of generating patient-friendly discharge note summaries and clinical decision support, clinician ratings generally ranged from dissatisfied to neutral, and no model achieved consistently satisfactory performance. Conclusions: Small open-source LLMs appear feasible for simple retrieval-oriented tasks in local hospital deployments but are currently inadequate for more complex applications, such as clinical decision support, deidentification, extraction of adverse events, and medical summarization. These findings highlight the importance of locally grounded evaluation tailored to specific use cases and the need for robust institutional evaluation frameworks to ensure safe and reliable deployment. Trial Registration:

Identifying potential nonpulmonary vein triggers in persistent atrial fibrillation using digital twins and deep learning

npj Digital Medicine, Published online: 08 September 2026; doi:10.1038/s41746-026-03223-y

Identifying potential nonpulmonary vein triggers in persistent atrial fibrillation using digital twins and deep learning

Sexual dimorphism in the complete Drosophila male central nervous system connectome

The Drosophila whole male central nervous system connectome enables end-to-end analysis of sensorimotor circuits. Comparison with existing female datasets shows that brain-wide wiring differences between the sexes are concentrated in higher centers.

The complete gustatory connectome of adult Drosophila reveals how taste guides feeding, foraging, and social behavior

The complete gustatory system and the full set of feeding motor neurons of adult Drosophila are mapped at synaptic resolution, enabling the identification of taste neuron types and their likely tastants and showing how valence-specific circuit motifs route taste to generate distinct feeding, neuroendocrine, and social outputs.

Dual downregulation of HRK and BAD promotes BCL-xL-mediated docetaxel resistance in prostate cancer

Oncogenesis, Published online: 20 August 2026; doi:10.1038/s41389-026-00652-y

Dual downregulation of HRK and BAD promotes BCL-xL-mediated docetaxel resistance in prostate cancer

Polyamines buffer labile iron to suppress ferroptosis

Polyamines buffer labile iron, restraining its reactivity to suppress ferroptosis. Their depletion raises labile iron and creates synthetic lethality with GPX4 loss, revealing polyamine metabolism as an endogenous regulator of iron homeostasis.

Integration of Digital Therapeutics Into Occupational Rehabilitation in Germany: Multilevel Simulation Study

Background: Expenditures for physiotherapy and extended outpatient physiotherapy (EAP) are increasing within Germany’s statutory accident insurance system (Berufsgenossenschaften), placing growing pressure on rehabilitation capacity and timely access to care. Digital health applications (DiGAs) are reimbursable nationwide and represent a novel component of routine rehabilitation pathways. However, their real-world system-level and economic effects in occupational rehabilitation remain insufficiently understood. Objective: This study aimed to evaluate how the integration of DiGAs into occupational rehabilitation pathways may influence costs, service capacity, and waiting times within routine care delivered by 5 German statutory accident insurance funds that cover 25.9 million insured individuals. Methods: Aggregated administrative data from 5 Berufsgenossenschaften (fiscal years 2023‐2024) were analyzed using a multilevel simulation framework combining (1) probabilistic cost-consequence modeling with Monte Carlo simulation (10,000 iterations), (2) an adherence-based adoption funnel distinguishing long-term engaged users (15%) and short-term users (85%) based on German claims data, and (3) a calibrated M/M/1 queuing model validated through discrete event simulation to estimate the effects on waiting times and system capacity. Primary outcomes included net financial impact, break-even thresholds, and changes in access-related performance metrics. Results: Combined physiotherapy and EAP expenditures reached €404 million (€1=US $1.18) in 2024, increasing by 10.1% year-over-year. The primary simulation (N=10,000 iterations) indicated mean annual net savings of €18.4 million (median €17.9 million) with a 90.7% probability of cost savings (95% uncertainty range: net cost of €8 million to net savings of €47.7 million). After incorporating adherence dynamics, the projected mean net savings were €16.2 million (95% CI €5-€29.8 million), corresponding to a 100% probability of positive financial impact within the modeled parameter space. Cost neutrality was maintained for DiGA prices up to €617.8 per prescription, nearly 40% above the base-case assumption of €450, indicating substantial economic robustness. Queuing analyses demonstrated that modest reductions in therapeutic demand decreased mean waiting times from 17.3 to 12.8 days (−26%), equivalent to approximately 120,000 cumulative patient waiting days saved annually across 26,705 EAP patients. The validation of discrete event simulation confirmed the magnitude and direction of analytic estimates. Conclusions: Under conservative assumptions, integrating digital therapeutics into occupational rehabilitation pathways is likely to generate both economic benefits and substantial system-level capacity gains. The break-even threshold of €617.80 per prescription provides a wide margin for pricing policy. Beyond cost effects, DiGAs may function as scalable capacity tools that alleviate systemic bottlenecks and improve timely access to rehabilitation services in capacity-constrained systems.
❌