❌

Normal view

Large Language Model–Generated Patient Instructions for Prescriptions in Primary Health Care: Preclinical Algorithm Validation

Background: The application of generative artificial intelligence to simplify medication use instructions has the potential to enhance people’s health by improving treatment adherence. Objective: We evaluated the performance of large language models (LLMs) in generating medication usage instructions to complement prescriptions in primary health care. Methods: This randomized, blinded experimental preclinical study used prescription-inducing scenarios, assigned to 62 health care professionals, to validate instructions generated by LLMs during electronic prescriptions. The instructions were generated by ChatGPT-4.0 (OpenAI), Llama3.1-8B (Meta), and Llama3.1-8B-RAG (Meta) using retrieval-augmented generation based on patient information leaflets. Performance metrics assessed adequacy, completeness, clarity, language simplification, usefulness, and errors in the generated instructions, with scores to analyze overall and individual metrics. Results: The 3 models yielded high overall scores for producing qualified instructions (ChatGPT-4.0: median 88.4, IQR 22.8; Llama3.1-8B: median 66.5, IQR 50.9; Llama3.1-8B-RAG: median 79.9, IQR 34.4; Kruskal-Wallis test P=.003). Llama3.1-8B-RAG received evaluations with similar overall scores to ChatGPT-4.0 (post hoc test, P=.05) and similar to Llama3.1-8B (post hoc test, P=.44). ChatGPT-4.0 outperformed Llama3.1-8B (Bonferroni test, P<.001). Regarding specific domains, Llama3.1-8B-RAG received scores equivalent to those of ChatGPT-4.0 for adequacy (mean 6.24, SD 2.3 vs mean 6.82, SD 2.1; post hoc test, P=.54); completeness (mean 5.94, SD 2.2 vs 6.55, SD 1.9; post hoc test P=.38), clarity (mean 5.77, SD 2.4 vs mean 6.68, SD 1.9; post hoc test P=.09), and usefulness (mean 5.42, SD 2.4 vs mean 5.96, SD 2.2; post hoc test P=.63). ChatGPT-4.0 received higher scores in the language simplification criterion than Llama3.1-8B-RAG (mean 7.05, SD 1.5 vs mean 5.44, SD 2.6; post hoc test P<.001). Interrater variability in assigning scores ranged from 4.2% (n=3) to 85.8% (n=6) among primary health care professionals. Instructions leading to incorrect use of the medication had similar frequency among the models(ChatGPT-4.0: n=15, 22.7%; Llama3.1-8B: n=19, 22.8%; Llama3.1-8B-RAG: n=19, 22.8%; chi-square test P=.71). The frequencies of hallucination were similar (ChatGPT-4.0: n=7, 10.6%; Llama3.1-8B: n=9, 13.6%; Llama3.1-8B-RAG: n=6, 9.1%; chi-square test P=.67). Conclusions: The open-source LLM enhanced with external information presented similar performance to the closed-source model, except for ChatGPT4.0, which was superior in language simplification of messages. LLM generation demonstrated potential for instructing patients on medication use. Nonetheless, the introduction of this innovation into the electronic prescribing workflow demands prescriber validation for human oversight of the technology and requires a strategy for LLM performance governance.

Safety of Telemedicine Versus In-Person Care for Patients With Tracheal Devices: Propensity Score–Matched Cohort Study

Background: Patients with tracheal diseases often require long-term follow-up after tracheal device placement, with a risk of adverse events that may lead to emergency care and unplanned interventions. Telemedicine has been proposed as an alternative to in-person follow-up to improve access and continuity of care. Objective: The primary objective of this study was to compare the need for emergency department (ED) visits between telemedicine and in-person groups. Secondary objectives included comparing hospital readmissions, 30-day hospital readmissions, and unplanned interventions between groups. Methods: This retrospective, single-institution study included adult patients with tracheal devices who underwent telemedicine and in-person outpatient clinic visits between 2020 and 2024. To balance the groups, we used 1:1 propensity score matching. We collected demographic and clinical data and evaluated the need for ED visits, hospital readmissions, 30-day hospital readmissions, and unplanned interventions. Kaplan-Meier estimation of time to first ED visit was performed to assess outcomes after outpatient visits. Results: A total of 483 patients (n=277, 57% telemedicine and n=206, 43% in-person) underwent 2487 visits (1258 telemedicine and 1229 in-person). After propensity score matching, 336 patients remained (168 in each group). There were no significant differences in the need for ED visits, hospital readmissions, or unplanned interventions. The telemedicine group had significantly fewer 30-day hospital readmissions (odds ratio 0.38, 95% CI 0.16-0.87; =.02). Kaplan-Meier analysis indicated no statistically significant difference in ED-free visits. Conclusions: Telemedicine follow-up was associated with outcomes comparable to those of in-person follow-up in this cohort of adult patients with tracheal devices, with no evidence of an increased need for ED visits. In the matched analysis, telemedicine was associated with lower odds of 30-day hospital readmission.

Fibroblast growth factor receptor inhibition for succinate dehydrogenase-deficient gastrointestinal stromal tumors: a phase 2 trial

Nature Medicine, Published online: 26 May 2026; doi:10.1038/s41591-026-04376-9

In a multicenter phase 2 trial, the fibroblast growth factor receptor inhibitor rogaratinib showed encouraging clinical efficacy in patients with succinate dehydrogenase-deficient gastrointestinal stromal tumors, suggesting a potential new treatment option for this patient population and demonstrating that an epigenetic mechanism of oncogene activation can be successfully targeted with a tyrosine kinase inhibitor.

A framework for building a synthetic cell from the SynCell Asia Initiative

Nature Biotechnology, Published online: 26 May 2026; doi:10.1038/s41587-026-03153-w

Building a living cell from scratch requires overcoming a bottleneck that has remained unresolved despite decades of progress: orchestrating the spatiotemporal integration of core functional modules. To tackle this barrier, the SynCell Asia Initiative outlines a strategy for developing core functional modules followed by their systems-level integration through the establishment of a centralized, artificial intelligence (AI)-driven biofoundry.

Semaglutide versus placebo in individuals with poor weight loss after bariatric surgery: a double-blinded, randomized, placebo-controlled trial

Nature Medicine, Published online: 22 May 2026; doi:10.1038/s41591-026-04416-4

At week 68, in patients who experienced poor weight loss following bariatric surgery, semaglutide was associated with 18.0% weight loss compared to 0.4% weight gain in patients receiving placebo.

Cusp-singularity-enhanced Coriolis effect for sensitive chip-scale gyroscopes

Nature, Published online: 20 May 2026; doi:10.1038/s41586-026-10565-w

By using singularity physics to enable cubic-root scaling of frequency and phase modulations induced by the Coriolis effect to enhance the performance of chip-scale Coriolis vibratory gyroscopes, substantial improvements in signal-to-noise ratio and precision are demonstrated.

De novo design of quasisymmetric two-component protein cages

Nature, Published online: 20 May 2026; doi:10.1038/s41586-026-10464-0

Researchers designed two-component proteins forming quasisymmetric cages via geometric frustration, enabling tunable virus-like assemblies for cargo delivery, cellular uptake and studying intracellular diffusion and protein localization.

Neural representation of action symbols in primate frontal cortex

Nature, Published online: 20 May 2026; doi:10.1038/s41586-026-10297-x

A drawing-like task designed to study compositional generalization identifies a specific neural population in the ventral premotor cortex in primates that encodes action symbols.

Coupling dead cell recognition to Fcγ receptors augments anticancer immunity

Nature Cancer, Published online: 20 May 2026; doi:10.1038/s43018-026-01168-5

Castro-Dopico et al. report the design of reagents to bridge F-actin and Fcγ receptors, endowing a range of antigen-presenting cells with the ability to cross-present antigens from dead tumor cells and boosting antitumor immunity in preclinical models.

Combined ctDNA and serum PSA for dynamic monitoring of metastatic prostate cancer starting first-line treatment: a prospective national cohort study

Nature Cancer, Published online: 15 May 2026; doi:10.1038/s43018-026-01172-9

Jayaram et al. assessed the performance of the combination of circulating tumor DNA and serum prostate-specific antigen for dynamic monitoring of metastatic prostate cancer at first-line treatment in this prospective observational clinical trial.

A comparison of deep multiomics profiles across ethnicity, geography, and age

Multiomics profiling of healthy individuals reveals differences across molecular layers and key pathways related to immune, metabolic, and microbiome-linked processes across ethnicities, while geographic relocation reshapes these networks and influences aging trajectories.
❌