❌

Normal view

Natural Language Processing Identification of Nonprescribed Fentanyl Use in Electronic Health Records: Algorithm Development and Validation Study

Background: Overdose and suicide due to nonprescribed fentanyl use have increased significantly, yet health care systems lack reliable methods to identify patients who use nonprescribed fentanyl. codes are inconsistent and do not specify nonprescribed fentanyl use. Objective: This study aimed to develop natural language processing approaches to identifying nonprescribed fentanyl use in electronic health record (EHR) documentation. Methods: This retrospective study included Veterans Health Administration patients seen between April 5, 2023, and December 23, 2024. A term list was developed to identify fentanyl-related mentions in clinical text, and 250-character snippets surrounding identified mentions were extracted. Veterans (n=3878) were randomly sampled from 5 predefined groups based on the presence of 1 of 4 terms (“fent,” “blues,” “M30s,” and “tranq”) in their EHR documentation. Physician annotators classified snippets into “nonprescribed fentanyl use,” “prescribed fentanyl use,” or “other,” with interannotator agreement evaluated using the mean pairwise Cohen κ. Cross-validation folds were constructed at the patient level between training and test sets. Penalized logistic regression, Bio-ClinicalBERT, Llama 3-8B, and Mistral-7B were trained on labeled data and compared. Model performance was evaluated using precision, recall, and -scores for each class, with a focus on the nonprescribed fentanyl use class as the primary label of clinical interest using bootstrapped 95% CIs. A fairness analysis and Shapley additive explanations analysis were performed using Bio-ClinicalBERT. External validation was performed using Bio-ClinicalBERT on an independent sample of 200 snippets, each representing a unique patient from January 2025 to June 2026, with precision reported as the primary validation metric. Results: Of 7389 snippets, 9.6% (n=709) were classified as “nonprescribed fentanyl use,” 40.3% (n=2981) were classified as “prescribed fentanyl use,” and 50% (n=3699) were classified as “other.” Interannotator agreement was high (κ=0.822). Llama 3-8B achieved the highest -score for nonprescribed fentanyl use (0.87, 95% CI 0.83-0.92), followed by Mistral-7B (0.80, 95% CI 0.75-0.84), Bio-ClinicalBERT (0.80, 95% CI 0.74-0.85), and penalized logistic regression (0.74, 95% CI 0.73-0.75). Performance was consistent across demographic subgroups, with lower performance for the nonprescribed fentanyl use class observed in female and Hispanic subgroups. Shapley additive explanations analysis revealed clinically meaningful discriminating terms for each class, although subword tokens required contextual interpretation. External validation of Bio-ClinicalBERT demonstrated a precision of 0.79 for nonprescribed fentanyl use. Conclusions: Natural language processing can identify nonprescribed fentanyl use in EHR documentation, although model performance for this class was lower than overall model performance, reflecting the clinical complexity of identifying nonprescribed use and the variable ways in which clinicians document this problem. This approach may support risk prediction and targeting of interventions to patients exposed to nonprescribed fentanyl.

Large Language Models for Distress Rating in Korean Psycho-Oncology Interviews: Exploratory Clinician-Benchmarked Evaluation Study

Background: Psychiatric distress is common among patients with cancer; yet, systematic interview-based screening remains difficult to scale in routine clinical care. Large language models (LLMs) have shown promise as scalable tools for mental health assessment, but most existing evidence is derived from clinician-authored records, translated text, or proxy data. The performance characteristics, error patterns, and explanatory behaviors of contemporary LLMs when applied to authentic, non-English psychiatric interviews remain insufficiently characterized. Objective: This exploratory study evaluated how contemporary LLMs reproduced psycho-oncologists’ item-level symptom ratings from real-world Korean psycho-oncology interviews, focusing on concordance, directional bias, and clinician-adjudicated error characteristics. Methods: Between April 2024 and May 2025, 101 adults receiving oncologic care in South Korea underwent semistructured interviews. Board-certified psycho-oncologists provided real-time, time-stamped ratings of 39 items. Generative pretrained transformer 4o (GPT-4o), Claude 3.5 Sonnet, and Gemini 2.5 Flash generated item scores and brief rationales using identical Korean zero-shot rubrics. Concordance between clinician ratings was evaluated using ordinal and binary screening metrics. A paired Wilcoxon test compared patient-level total symptom burden. Binary mismatches were clinically adjudicated as ambiguous or definite overestimation or underestimation with an 8-etiology taxonomy. Model-generated rationales were further meta-evaluated using GPT-5.4, Claude Sonnet 4.6, and Gemini Pro 3.1 across 4 dimensions: citation (use of quoted supporting statements), structure (logical organization of the rationale), mapping (consistency between the rationale and the assigned item rating), and expansion (degree of interpretive elaboration beyond the explicit transcript content). The associations between meta-evaluation results and absolute error were examined using cross-classified mixed-effects models. Results: Out of 101 participants, 88 (87.1%) were predominantly female and had breast cancer as the primary cancer type (n=70, 69.3%). Across 3931 item-level ratings, all models showed good agreement with clinicians (intraclass correlation coefficient: 0.816‐0.872), with GPT-4o showing the highest agreement. Claude 3.5 and Gemini 2.5 yielded significantly higher patient-level symptom burden (adjusted

Comparative Efficacy of Different AI Systems for Polyp Detection by Size During Colonoscopy: Systematic Review and Network Meta-Analysis

Background: Colorectal cancer remains a leading cause of death despite being largely preventable through polypectomy. AI systems designed to enhance polyp detection during colonoscopy have shown promise, but the extent to which they improve detection of different-sized polyps remains unclear. Objective: This study compared the size-stratified efficacy of AI-assisted colonoscopy vs standard colonoscopy using the Hartung-Knapp-Sidik-Jonkman (HKSJ) method, and generated exploratory rankings while acknowledging all cross-platform comparisons are indirect. Methods: This systematic review and network meta-analysis (NMA) searched PubMed, Embase, Cochrane CENTRAL, and Web of Science from inception to July 25, 2026, supplemented by citation searching. We included randomized controlled trials (RCTs) comparing AI-assisted vs standard colonoscopy in adults (≥18 years of age), reporting mean polyp detection counts stratified by size (≤5 mm, 6-9 mm, and ≥10 mm). Two reviewers screened studies, extracted data, and assessed risk of bias using the Cochrane Risk of Bias 2.0. We conducted frequentist NMA using the HKSJ method with restricted maximum likelihood estimation, calculated 95% prediction intervals (PIs), and assessed heterogeneity using I2 and τ2. Certainty of evidence was rated using the GRADE (Grading of Recommendations Assessment, Development, and Evaluation) framework. Results: A total of 13 RCTs (4156 participants) compared 8 AI systems to standard colonoscopy, forming a network without direct AI comparisons. For diminutive polyps (≤5 mm), AI showed a modest advantage (standardized mean difference [SMD] 0.21, 95% CI 0.07 to 0.35, 95% PI –1.12 to 1.54), but substantial heterogeneity (I2=86.6%) and wide PI crossing the null indicated high uncertainty. EndoScreener showed the most consistent evidence (SMD 0.36, 95% CI 0.18-0.54). For small and large polyps, effects were minimal (SMD 0.02, 95% CI –0.02 to 0.06, 95% PI –0.03 to 0.07; SMD 0.01, 95% CI 0.00-0.02, 95% PI –0.01 to 0.03). GRADE certainty was very low for diminutive polyps and low for small and large polyps. Sensitivity analysis excluding Tianjin YuJin did not materially change findings. Conclusions: AI may modestly enhance diminutive polyp detection, but effects on small and large polyps are minimal, with no platform superiority. Given very low to low certainty, findings are hypothesis-generating. This exploratory NMA provides size-stratified comparisons that can inform future head-to-head trial design. Unlike prior reviews aggregating all polyp sizes, we show the overall AI benefit is driven by diminutive polyp detection, providing a framework for targeted deployment—prioritizing AI for diminutive polyp screening, with limited value for larger lesions. Head-to-head trials are urgently needed. Trial Registration: PROSPERO International Prospective Register of Systematic Reviews CRD420251266932; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251266932

“Small” Large Language Models in the Hospital: Evaluation Study on Real-World Data in a Resource-Constrained Setting

Background: Large language models (LLMs) are increasingly being deployed in health care, but their use and deployment in many real-world hospital environments pose significant challenges and concerns. In particular, state-of-the-art commercial models store or process data externally, which is often in conflict with ensuring patient data protection. At the same time, using LLMs locally is limited by the lack of available computing infrastructure. Small open-source LLMs that do not require substantial computing resources could offer a practical way to resolve these tensions, but their medical utility in real-world local contexts, especially in non-English languages, has not been sufficiently evaluated. Objective: This study aimed to evaluate the feasibility of small, locally deployable open-source LLMs for clinically relevant tasks in a resource-constrained hospital setting and to propose a reproducible framework for institution-specific evaluation before deployment. Methods: We evaluated 6 open-source LLMs ranging from 8B to 24B parameters (from the Mistral, Phi4, Falcon3, Llama3.1, and Meditron3 families) in a zero-shot setting across 7 tasks covering 4 clinical use cases: information extraction, medical text translation, text generation, and clinical decision support. We used deidentified French clinical data from a Swiss tertiary hospital, including discharge letters, clinical notes, and structured electronic health records. Performance was assessed using task-specific metrics, such as precision, recall, F1-score, embedding-based semantic similarity, recall-oriented understudy for gisting evaluation (ROUGE) score, readability indices, and human review by clinicians. Results: Model performance varied substantially between tasks. In the simplest retrieval task (needle-in-the-haystack), several models performed strongly, with Llama3.1 achieving an F1-score of 99.81% and Mistral-small achieving 99.71%. In contrast, performance was poor in more complex tasks. For detecting protected health information, the best-performing LLMs achieved only modest overall macro–F1-scores (0.33-0.34), substantially below a fine-tuned Robustly Optimized BERT Pretraining Approach (RoBERTa) baseline (0.94). In the task of extracting immune-related adverse events from discharge notes, the highest overall macro–F1-score was 0.35 with Phi4. For medical text translation, Phi4 ranked highest in embedding-based evaluation, whereas Meditron3-Phi4 performed the worst, with clinician reviews identifying hallucinations in 55% of its outputs. In the task of summarizing discharge letters, quality was low across all models, with the best penalized ROUGE score reaching only 0.169 with Llama3.1. In the tasks of generating patient-friendly discharge note summaries and clinical decision support, clinician ratings generally ranged from dissatisfied to neutral, and no model achieved consistently satisfactory performance. Conclusions: Small open-source LLMs appear feasible for simple retrieval-oriented tasks in local hospital deployments but are currently inadequate for more complex applications, such as clinical decision support, deidentification, extraction of adverse events, and medical summarization. These findings highlight the importance of locally grounded evaluation tailored to specific use cases and the need for robust institutional evaluation frameworks to ensure safe and reliable deployment. Trial Registration:

Commensal <i>Nakaseomyces glabratus</i> migrates into prostate tumors to accelerate cancer progression

Nature Cancer, Published online: 09 September 2026; doi:10.1038/s43018-026-01229-9

Lai et al. show that Nakaseomyces glabratus is enriched in fecal and tumor samples of patients with castration-resistant prostate cancer and that administration of the fungus accelerates cancer progression in prostate cancer-bearing castrated mice.

Within-family effect of ancestry on complex traits in a Mexican population

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-11039-9

This study uses a within-family design to identify significant ancestry differences in complex traits such as height and type 2 diabetes in a genetically diverse population from Mexico City.

Neocortical long-range inhibition promotes cortical synchrony and sleep

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-10876-y

In mice, a sparse population of sleep-active long-range inhibitory neurons in the neocortex promote widespread cortical synchronization and sleep, revealing a cortical mechanism that contributes to the regulation of sleep.

An operational perturbation proteomics-based virtual cell model

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-11001-9

Temporal protein-abundance measurements from systematically perturbed breast cancer cell lines were generated to develop ProteinTalks, a virtual cell model that functions as an operational tool for diverse drug discovery tasks.

Foaming photopolymers as a high-resolution biomimetic printing platform

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-10968-9

Deep-foam photolithography uses light-controlled polymer foaming to create high-resolution, multifunctional microstructures with tunable optical, wetting and fluid-handling properties for advanced manufacturing applications.

Denisovans from southwestern China and their subsistence strategies

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-10997-4

Evidence from Bianfu Cave shows specialized hunting, expedient stone-tool production and extensive bone use of Denisovans, providing new insights into their ecology, behaviour and cultural legacy in eastern Asia.

A serpin–myeloid axis in pancreatic cancer heterogeneity and immune evasion

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-11002-8

SERPINE1 and SERPINB2-driven fibrin-rich niches locally programme immunosuppressive macrophages and exclude T cells, enabling spatially organized immune evasion in pancreatic ductal carcinoma.

Mobile education builds resilience during shocks in five countries

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-10990-x

Remote phone-based targeted tutoring in five countries produced large, cost-effective learning gains during school closures, outperforming text messaging and working effectively when delivered by governments or non-governmental organizations to strengthen education resilience.

Ancient proteins identify various Denisovan remains from Southwest China

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-10976-9

Identification and proteomic analysis of bone fragments and teeth from an excavation in Southwest China provide insight into the evolution and phenotype of Denisovans and fill a geographical gap in their documented distribution.

Imaging cellular activity across all organs reveals body-wide circuits

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-10979-6

An imaging system developed to record cellular activity throughout the whole body of zebrafish captures cellular organ dynamics and identifies multiple distributed circuits.

TRI-611, a selective, brain-penetrant molecular glue degrader of ALK

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-10998-3

TRI-611 induces the degradation of ALK fusion proteins via a previously undescribed CRBN recruitment motif, and its preclinical anti-tumour activity highlights TRI-611 as a potential new way of treating ALK-positive non-small-cell lung cancer.

TM184C is a GPCR-like regulator of intercellular exchange and autophagy

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-10993-8

TM184C—an ancient G-protein-coupled receptor-like superdark protein involved in regulation of autophagy, intercellular connectivity and material exchange—underscores the promise of exploring the understudied human proteome and beyond.

Proximity-guided graph learning reveals tumour-associated proximity antigens

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-11003-7

A proximity-mapping atlas defines tumour-associated proximity antigens, revealing disease-associated membrane spatial communities, and identifies EGFR–CDCP1 as a co-target pair that enhances tumour killing by multispecific therapeutics.
❌