❌

Normal view

“Small” Large Language Models in the Hospital: Evaluation Study on Real-World Data in a Resource-Constrained Setting

Background: Large language models (LLMs) are increasingly being deployed in health care, but their use and deployment in many real-world hospital environments pose significant challenges and concerns. In particular, state-of-the-art commercial models store or process data externally, which is often in conflict with ensuring patient data protection. At the same time, using LLMs locally is limited by the lack of available computing infrastructure. Small open-source LLMs that do not require substantial computing resources could offer a practical way to resolve these tensions, but their medical utility in real-world local contexts, especially in non-English languages, has not been sufficiently evaluated. Objective: This study aimed to evaluate the feasibility of small, locally deployable open-source LLMs for clinically relevant tasks in a resource-constrained hospital setting and to propose a reproducible framework for institution-specific evaluation before deployment. Methods: We evaluated 6 open-source LLMs ranging from 8B to 24B parameters (from the Mistral, Phi4, Falcon3, Llama3.1, and Meditron3 families) in a zero-shot setting across 7 tasks covering 4 clinical use cases: information extraction, medical text translation, text generation, and clinical decision support. We used deidentified French clinical data from a Swiss tertiary hospital, including discharge letters, clinical notes, and structured electronic health records. Performance was assessed using task-specific metrics, such as precision, recall, F1-score, embedding-based semantic similarity, recall-oriented understudy for gisting evaluation (ROUGE) score, readability indices, and human review by clinicians. Results: Model performance varied substantially between tasks. In the simplest retrieval task (needle-in-the-haystack), several models performed strongly, with Llama3.1 achieving an F1-score of 99.81% and Mistral-small achieving 99.71%. In contrast, performance was poor in more complex tasks. For detecting protected health information, the best-performing LLMs achieved only modest overall macro–F1-scores (0.33-0.34), substantially below a fine-tuned Robustly Optimized BERT Pretraining Approach (RoBERTa) baseline (0.94). In the task of extracting immune-related adverse events from discharge notes, the highest overall macro–F1-score was 0.35 with Phi4. For medical text translation, Phi4 ranked highest in embedding-based evaluation, whereas Meditron3-Phi4 performed the worst, with clinician reviews identifying hallucinations in 55% of its outputs. In the task of summarizing discharge letters, quality was low across all models, with the best penalized ROUGE score reaching only 0.169 with Llama3.1. In the tasks of generating patient-friendly discharge note summaries and clinical decision support, clinician ratings generally ranged from dissatisfied to neutral, and no model achieved consistently satisfactory performance. Conclusions: Small open-source LLMs appear feasible for simple retrieval-oriented tasks in local hospital deployments but are currently inadequate for more complex applications, such as clinical decision support, deidentification, extraction of adverse events, and medical summarization. These findings highlight the importance of locally grounded evaluation tailored to specific use cases and the need for robust institutional evaluation frameworks to ensure safe and reliable deployment. Trial Registration:

Breaking timescales with generative sampling of conformational transitions

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-11025-1

A generative committor-guided path-sampling framework reconstructs rare biomolecular transition pathways and reveals the underlying thermodynamics and kinetics without using predefined collective variables or brute-force sampling, at an acceptable computational cost.

Neocortical long-range inhibition promotes cortical synchrony and sleep

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-10876-y

In mice, a sparse population of sleep-active long-range inhibitory neurons in the neocortex promote widespread cortical synchronization and sleep, revealing a cortical mechanism that contributes to the regulation of sleep.

Denisovans from southwestern China and their subsistence strategies

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-10997-4

Evidence from Bianfu Cave shows specialized hunting, expedient stone-tool production and extensive bone use of Denisovans, providing new insights into their ecology, behaviour and cultural legacy in eastern Asia.

Proximity-guided graph learning reveals tumour-associated proximity antigens

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-11003-7

A proximity-mapping atlas defines tumour-associated proximity antigens, revealing disease-associated membrane spatial communities, and identifies EGFR–CDCP1 as a co-target pair that enhances tumour killing by multispecific therapeutics.

Myeloperoxidase inhibition with mitiperstat in heart failure with preserved or mildly reduced ejection fraction: a randomized phase 2b trial

Nature Medicine, Published online: 09 September 2026; doi:10.1038/s41591-026-04615-z

In a phase 2 randomized clinical trial, treatment with the myeloperoxidase (MPO) inhibitor mitiperstat, intended to target the neutrophil–MPO inflammatory pathway, did not improve symptoms or exercise function in individuals with heart failure with preserved or mildly reduced ejection fraction.

Designing a Gamified mHealth App for HIV Prevention and Comorbidities Among Malaysian Young Men Who Have Sex With Men: An Interdisciplinary Expert Panel Study

Background: New HIV cases among Malaysian men who have sex with men continue to rise, with young men who have sex with men (YMSM) accounting for 44% of new infections and experiencing high rates of comorbidities. Mobile health (mHealth) apps offer a promising approach to addressing these challenges by providing discreet access to health information, screening tools, and linkage to services. Given the near-universal smartphone ownership among Malaysian YMSM and high levels of mobile gaming engagement, gamified mHealth apps may be particularly effective in sustaining engagement and promoting HIV prevention behaviors and comorbidity management. However, realizing their full potential requires identifying the features and design principles that are most important to Malaysian YMSM and that can support sustained engagement and improve health outcomes in this vulnerable population. Objective: This study used an interdisciplinary expert panel approach to identify the key features, design principles, and gamification elements to be incorporated into MY-Hero, a gamified mHealth app designed to support HIV prevention, comorbidity management, and overall well-being among Malaysian YMSM. Methods: Guided by the integrated behavioral model (IBM), we conducted an interdisciplinary expert panel comprising experts in health, technology, and design sciences, as well as health care professionals with expertise in the Malaysian YMSM population, along with local lesbian, gay, bisexual, transgender, queer, and others (LGBTQ+) community leaders. Through a structured series of discussions and thematic analysis, we identified culturally appropriate design principles and key features for MY-Hero that align with the health needs and sociocultural context of Malaysian YMSM. Results: Three expert panel sessions were conducted with 9 experts. Several key themes emerged: (1) the importance of establishing clinical affiliation and facilitating linkage to HIV testing, pre-exposure prophylaxis (PrEP), and related services, including harm reduction services; (2) the need to enhance user interface (UI) and user experience (UX) by optimizing usability, interactivity, and engagement to maintain user interest; (3) the incorporation of customizable health content to tailor interventions based on individual characteristics and preferences; (4) the use of gamification mechanisms, such as reward systems and progression tracking to promote adoption and sustained engagement with health services; and (5) the importance of privacy and data security as critical design considerations for ensuring user safety, confidentiality, and trust. Conclusions: The findings highlight the importance of user-centered design, contextualized health content, and gamification mechanisms in mHealth tools for Malaysian YMSM. These insights provide practical guidance for developing culturally appropriate gamified mHealth interventions for vulnerable populations by informing strategies to enhance user engagement, reduce HIV prevention fatigue, and support the well-being of YMSM in stigmatized contexts.

Sexual dimorphism in the complete Drosophila male central nervous system connectome

The Drosophila whole male central nervous system connectome enables end-to-end analysis of sensorimotor circuits. Comparison with existing female datasets shows that brain-wide wiring differences between the sexes are concentrated in higher centers.

Localized PD-1 CAR T therapy reprograms neuroinflammation

PD-1+ T follicular helper-like cells drive B cell-associated pathology in multiple sclerosis. Programmable PD-1-targeting CAR T cells selectively eliminate these pathogenic CD4 T cells while delivering IL-10 at sites of inflammation, thereby suppressing neuroinflammation across preclinical models.

Large Language Model–Generated Patient Instructions for Prescriptions in Primary Health Care: Preclinical Algorithm Validation

Background: The application of generative artificial intelligence to simplify medication use instructions has the potential to enhance people’s health by improving treatment adherence. Objective: We evaluated the performance of large language models (LLMs) in generating medication usage instructions to complement prescriptions in primary health care. Methods: This randomized, blinded experimental preclinical study used prescription-inducing scenarios, assigned to 62 health care professionals, to validate instructions generated by LLMs during electronic prescriptions. The instructions were generated by ChatGPT-4.0 (OpenAI), Llama3.1-8B (Meta), and Llama3.1-8B-RAG (Meta) using retrieval-augmented generation based on patient information leaflets. Performance metrics assessed adequacy, completeness, clarity, language simplification, usefulness, and errors in the generated instructions, with scores to analyze overall and individual metrics. Results: The 3 models yielded high overall scores for producing qualified instructions (ChatGPT-4.0: median 88.4, IQR 22.8; Llama3.1-8B: median 66.5, IQR 50.9; Llama3.1-8B-RAG: median 79.9, IQR 34.4; Kruskal-Wallis test P=.003). Llama3.1-8B-RAG received evaluations with similar overall scores to ChatGPT-4.0 (post hoc test, P=.05) and similar to Llama3.1-8B (post hoc test, P=.44). ChatGPT-4.0 outperformed Llama3.1-8B (Bonferroni test, P<.001). Regarding specific domains, Llama3.1-8B-RAG received scores equivalent to those of ChatGPT-4.0 for adequacy (mean 6.24, SD 2.3 vs mean 6.82, SD 2.1; post hoc test, P=.54); completeness (mean 5.94, SD 2.2 vs 6.55, SD 1.9; post hoc test P=.38), clarity (mean 5.77, SD 2.4 vs mean 6.68, SD 1.9; post hoc test P=.09), and usefulness (mean 5.42, SD 2.4 vs mean 5.96, SD 2.2; post hoc test P=.63). ChatGPT-4.0 received higher scores in the language simplification criterion than Llama3.1-8B-RAG (mean 7.05, SD 1.5 vs mean 5.44, SD 2.6; post hoc test P<.001). Interrater variability in assigning scores ranged from 4.2% (n=3) to 85.8% (n=6) among primary health care professionals. Instructions leading to incorrect use of the medication had similar frequency among the models(ChatGPT-4.0: n=15, 22.7%; Llama3.1-8B: n=19, 22.8%; Llama3.1-8B-RAG: n=19, 22.8%; chi-square test P=.71). The frequencies of hallucination were similar (ChatGPT-4.0: n=7, 10.6%; Llama3.1-8B: n=9, 13.6%; Llama3.1-8B-RAG: n=6, 9.1%; chi-square test P=.67). Conclusions: The open-source LLM enhanced with external information presented similar performance to the closed-source model, except for ChatGPT4.0, which was superior in language simplification of messages. LLM generation demonstrated potential for instructing patients on medication use. Nonetheless, the introduction of this innovation into the electronic prescribing workflow demands prescriber validation for human oversight of the technology and requires a strategy for LLM performance governance.

Safety of Telemedicine Versus In-Person Care for Patients With Tracheal Devices: Propensity Score–Matched Cohort Study

Background: Patients with tracheal diseases often require long-term follow-up after tracheal device placement, with a risk of adverse events that may lead to emergency care and unplanned interventions. Telemedicine has been proposed as an alternative to in-person follow-up to improve access and continuity of care. Objective: The primary objective of this study was to compare the need for emergency department (ED) visits between telemedicine and in-person groups. Secondary objectives included comparing hospital readmissions, 30-day hospital readmissions, and unplanned interventions between groups. Methods: This retrospective, single-institution study included adult patients with tracheal devices who underwent telemedicine and in-person outpatient clinic visits between 2020 and 2024. To balance the groups, we used 1:1 propensity score matching. We collected demographic and clinical data and evaluated the need for ED visits, hospital readmissions, 30-day hospital readmissions, and unplanned interventions. Kaplan-Meier estimation of time to first ED visit was performed to assess outcomes after outpatient visits. Results: A total of 483 patients (n=277, 57% telemedicine and n=206, 43% in-person) underwent 2487 visits (1258 telemedicine and 1229 in-person). After propensity score matching, 336 patients remained (168 in each group). There were no significant differences in the need for ED visits, hospital readmissions, or unplanned interventions. The telemedicine group had significantly fewer 30-day hospital readmissions (odds ratio 0.38, 95% CI 0.16-0.87; =.02). Kaplan-Meier analysis indicated no statistically significant difference in ED-free visits. Conclusions: Telemedicine follow-up was associated with outcomes comparable to those of in-person follow-up in this cohort of adult patients with tracheal devices, with no evidence of an increased need for ED visits. In the matched analysis, telemedicine was associated with lower odds of 30-day hospital readmission.
❌