❌

Normal view

Received — 13 September 2026 ⏭ Journal of Medical Internet Research

Effectiveness of Wearable Digital Therapeutics in Improving Sleep Outcomes Among Individuals With Insomnia: Systematic Review and Meta-Analysis of Randomized Controlled Trials

Background: Wearable devices are increasingly used for sleep monitoring and as adjunctive treatment. Existing meta-analyses mostly pool composite digital therapies and rarely isolate stand-alone wearables or distinguish between objective and subjective end points. Whether stand-alone wearable interventions improve sleep outcomes in adults with insomnia, and which factors moderate treatment heterogeneity, remains unclear. Objective: This study aims to evaluate the effectiveness of wearable digital interventions on sleep outcomes in adults with insomnia versus control strategies and explore moderators of effectiveness, including device-wearing position, intervention duration, and control type, using meta-regression. Methods: This systematic review and meta-analysis was conducted in accordance with the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta‑Analyses) 2020 statement and the PRISMA-S (Preferred Reporting Items for Systematic Reviews and Meta‑Analyses Literature Search Extension) guideline. Five electronic databases and clinical trial registries were searched from inception to May 18, 2026. Eligible studies were randomized controlled trials (RCTs) evaluating wearable digital interventions in adults with insomnia compared with sham, waitlist, usual care, or active control conditions and had an intervention duration of at least 1 week. Study screening, data extraction, and risk-of-bias assessment were carried out independently by 2 reviewers. Pooled estimates were calculated using a restricted maximum likelihood random-effects model with the Hartung-Knapp-Sidik-Jonkman correction. Heterogeneity was assessed using the ² statistic, and 95% prediction intervals (PIs) were calculated for the primary analyses. The certainty of evidence was rated using the GRADE (Grading of Recommendations, Assessment, Development, and Evaluation) approach. Results: Sixteen RCTs (N=910) were included. Wearable digital interventions were associated with a significant reduction in objective sleep-onset latency (SOL; mean difference [MD] −4.52, 95% CI −8.38 to −0.67, PI −9.52 to 0.47 min) and a significant improvement in subjective sleep efficiency (SE; MD 2.00%, 95% CI 1.90%‐2.11%, PI 1.85%‐2.15%). Subjective total sleep time (TST) also showed a significant increase (MD 19.11, 95% CI 2.98‐35.24, PI −16.20 to 54.43 minutes). Meta-regression showed that control type, intervention duration, and device location did not explain the heterogeneity of the insomnia severity index (ISI) (=0). Sensitivity analysis confirmed the robustness of pooled ISI estimates, and an Egger test indicated no small-study effects (=.07). Certainty of evidence ranged from moderate to high. Conclusions: Wearable digital interventions provide selective benefits for objective SOL, subjective SE, and subjective TST in adults with insomnia, with no improvement in overall ISI. Despite statistically significant effects on several sleep parameters, wide PIs, substantial heterogeneity, and limited study numbers indicate preliminary, nonconclusive findings. Wearables should be viewed as affordable adjunctive tools requiring further validation, not substitutes for first-line cognitive behavioral therapy for insomnia. Large-scale, long-term RCTs with standardized protocols and patient-level external validation are required to consolidate the evidence base. Trial Registration: PROSPERO CRD420251038603; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251038603

Correction: From Metrics to Meaning in Neurological Rehabilitation: Clinicians’ Perspectives on Digital Metrics of Upper Limb Functioning—A Focus Group Study

Digital assessment technologies, such as optical motion capture and inertial measurement units, enable detailed kinematic analysis and continuous monitoring of upper limb activity in persons with neurological conditions. While such digital metrics of functioning are increasingly recognized in research, their uptake in clinical neurorehabilitation is limited. It remains unclear which digital metrics of functioning clinicians perceive as most meaningful and how these are integrated into patient-centered care. Understanding clinicians’ information needs and reasoning processes is a prerequisite for implementing digital assessment technology. To characterize how rehabilitation professionals perceive, prioritize, and integrate digital metrics of functioning into clinical reasoning and to identify features that would support their routine use. Three 90-minute focus groups were conducted in 3 Swiss neurorehabilitation centers, involving 11 clinicians with diverse professional backgrounds (5 physiotherapists, 4 occupational therapists, 1 movement scientist, and 1 medical practitioner). Participants discussed essential parameter domains and individually rated the relevance and meaningfulness of 17 kinematic metrics for the well-studied drinking task and 10 established arm use performance metrics. Verbatim transcripts were analyzed using reflexive thematic analysis, and rating data were summarized descriptively. Five main themes were identified. (1) Functional requirements to interpret movement quality and performance (active/passive range of motion (ROM), strength, selective muscle control, grasp) form the basis for interpreting movement. (2) Essential aspects of movement quality (smoothness, efficiency, compensatory movement) are valued when aligned with observable task execution. (3) Added value of real-world performance (hourly activity profiles, arm-use symmetry, functional workspace) represents the reference for patient-centered reasoning. (4) Individualizing what matters, including diagnosis-specific preferences, shapes assessment selection. (5) Blending clinical eye and reference data reflects clinicians’ reliance on visual judgment complemented by normative values. Intuitive metrics such as task duration, number of movement units, and ROM were favored, whereas confidence was lower in more complex metrics (e.g., jerk, inter-joint coordination). Clinicians value intuitive digital metrics of functioning when they are clearly linked to patient-centered outcomes and supported by normative references. The findings highlight the need for targeted educational strategies and digital competency training that help clinicians interpret digital metrics and integrate them with contextual information and clinical reasoning.

Evaluating Large Language Models in Clinical Audiology (AUDIOLOGYBENCH): Benchmark Development and Validation Study

Background: Large language models (LLMs) are increasingly being explored for clinical decision support, but their performance in audiology has not been systematically benchmarked using clinically grounded case materials and rubric-based safety evaluations. Objective: This study aimed to develop and evaluate AUDIOLOGYBENCH, a 3-tier benchmark for characterizing frontier LLM capability in clinical audiology along (1) curated domain knowledge, (2) literature-derived evidence, and (3) clinical reasoning under multimodal case input, with an explicit human audit of the automated adjudicator on the primary end point. Methods: The benchmark comprises 3139 objective items from educational resources, 3175 research article–derived items from peer-reviewed articles published between 2015 and 2025, and 67 multimodal clinical case studies graded against a standardized A-F rubric with 6 prespecified critical-error types that cap scores at D or F. Eight models were evaluated on the educational objective items: 4 frontier multimodal models (Gemini 2.5 Pro, Grok 4, OpenAI O3, and Claude Sonnet 4 Thinking) were evaluated on the research article–derived items, and on 804 case study evaluations. Adjudication used Gemini 2.5 Pro (objective and research-derived items) and Claude Opus 4.5 (case studies). The case study adjudicator was independently audited against PhD-level audiologist consensus on blinded subsamples, supplemented by a post-stratified human-calibrated sensitivity analysis. Results: A striking task-type dissociation emerged on case studies: clinical recommendations (Q3) achieved a mean score of 89.74 (SD 13.92, 95% CI 88.07‐91.41), a 98.1% (263/268) pass rate, and no dangerous recommendations; audiometric numerical interpretation (Q1) achieved a mean score of 67.89 (SD 18.47, 95% CI 65.68‐70.10), with a 35.4% (95/268) critical-error rate; and differential diagnosis (Q2) achieved a mean score of 67.79 (SD 15.33, 95% CI 65.95‐69.63). Question type, not model selection, dominated performance (eta-squared_H=0.333 vs 0.001; rank biserial ≥0.679). Interreviewer reliability between audiologists was high (quadratic-weighted κ of 0.78 and 0.85 across the 80-item and 50-item audits, respectively). When 2 audiologists regraded all 80 model Q1 responses with the diagnostic images available, the adjudicator’s per-item Q1 labels diverged from human judgment (κ=0.05; overflagging; sensitivity: 19/26, 73%; positive predictive value: 19/53, 36%), yet its reweighted Q1 critical-error rate (36.2%) was broadly consistent with the image-grounded human estimates (28%‐34%), suggesting no systematic inflation of the headline rate. The principal Q3>{Q1, Q2} ranking was preserved under post-stratified human calibration. Web-style multiple-choice items showed ceiling effects (>95% accuracy); short-answer prompts remained challenging (best 30%). Conclusions: Current frontier LLMs show strong recommendation generation but substantial limitations in audiometric numerical interpretation that are shared across models and that an automated adjudicator partially miscalibrated at the per-item level. AUDIOLOGYBENCH characterizes capability boundaries rather than certifying clinical readiness. Deployment of LLM-assisted audiology workflows requires structured human verification of all numerical findings and awareness of fabrication and severity misclassification failure modes documented here.

The Effectiveness of Digital Intervention on Psychological Resilience in Postoperative Breast Cancer Patients During Chemotherapy Intervals: Quasi-Experimental Study

Background: Patients with breast cancer during postoperative chemotherapy intervals commonly experience psychological distress and reduced resilience while recovering at home. Digital mindfulness interventions may provide accessible psychological support during this vulnerable period; however, evidence regarding tailored interventions for postoperative patients with breast cancer during chemotherapy intervals remains limited. Objective: This study aimed to examine the effectiveness of a digital intervention on psychological resilience in postoperative patients with breast cancer during chemotherapy intervals. Methods: A quasi-experimental study with repeated measures was conducted from October 2021 to June 2022. A total of 80 eligible participants were recruited from the Department of Breast Surgery at a tertiary hospital in Zhejiang Province, China, and 71 completed the study. The control group received routine discharge instructions and nursing follow-ups, whereas the intervention group additionally received an 8-week digital psychological resilience intervention. Outcomes were assessed at baseline (T0), 3 months post intervention (T1), and 6 months post intervention (T2). The measures included the Connor-Davidson Resilience Scale (CD-RISC), Hospital Anxiety and Depression Scale (HADS), Social Support Rating Scale (SSRS), Breast Cancer Survivor Self-Efficacy Scale (BCSSS), and Functional Assessment of Cancer Therapy-Breast (FACT-B). Independent-samples tests, chi-square tests, and repeated-measures ANOVA were performed using SPSS (version 26.0; IBM Corp). Results: No statistically significant baseline differences were observed between the two groups in the outcome measures. At T1, the intervention group had higher CD-RISC scores than the control group (mean 67.58, SD 11.41 vs mean 62.09, SD 10.18; =.036) and higher BCSSS scores (mean 42.36, SD 3.59 vs mean 39.23, SD 4.90; =.003). However, these between-group differences were no longer statistically significant at T2 (>.05). Significant time effects and group×time interaction effects were observed for both psychological resilience and self-efficacy (.05), although both scales showed significant time effects (

Performance of AI-Based Screening Tools for Obstructive Sleep Apnea Across Apnea-Hypopnea Index Thresholds: Systematic Review and Meta-Analysis

Background: Obstructive sleep apnea (OSA) is highly prevalent but remains substantially underdiagnosed. Polysomnography (PSG) is the reference standard, but its cost and limited availability constrain large-scale case identification. AI-based screening tools may support risk stratification and referral prioritization, but their diagnostic accuracy across apnea-hypopnea index (AHI) thresholds remains uncertain. Objective: This review aimed to systematically evaluate the diagnostic accuracy of AI-based OSA screening tools at AHI thresholds of ≥5, ≥15, and ≥30 events/hour, with emphasis on models using non-PSG–derived inputs. Methods: PubMed, Embase, Scopus, and Web of Science were searched for studies published from January 1, 2016, to May 3, 2026. Eligible studies included adults evaluated for suspected OSA or recruited from population-based cohorts, assessed AI-based models intended or interpretable for OSA screening, risk prediction, or screening-oriented severity classification, used PSG as the reference standard, and reported sufficient data to construct or reconstruct 2×2 contingency tables. Diagnostic accuracy was synthesized separately by AHI threshold and input source using bivariate random-effects models, with 95% CIs and prediction intervals (PIs). Risk of bias and certainty of evidence were assessed using QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies 2) and GRADE (Grading of Recommendations Assessment, Development, and Evaluation), respectively. Results: A total of 60 studies were included, of which 47 contributed data to the meta-analysis. At AHI thresholds of ≥5, ≥15, and ≥30 events/hour, pooled sensitivities were 0.94 (95% CI 0.92‐0.96; 95% PI 0.71‐0.99), 0.87 (95% CI 0.84‐0.89; 95% PI 0.66‐0.96), and 0.83 (95% CI 0.79‐0.87; 95% PI 0.61‐0.94), respectively; the corresponding specificities were 0.77 (95% CI 0.69‐0.84; 95% PI 0.30‐0.96), 0.81 (95% CI 0.75‐0.85; 95% PI 0.39‐0.96), and 0.91 (95% CI 0.87‐0.94; 95% PI 0.55‐0.99), respectively. The corresponding areas under the summary receiver operating characteristic curves were 0.943, 0.907, and 0.920. For non-PSG–derived tools, sensitivities were 0.92, 0.85, and 0.81, and specificities were 0.70, 0.74, and 0.85 at the 3 thresholds, respectively. For PSG-derived models, sensitivities were 0.96, 0.90, and 0.85, and specificities were 0.82, 0.88, and 0.96, respectively. Exploratory subgroup analyses suggested performance variation across selected study and model characteristics, including region, algorithmic framework, data source, and validation method. Conclusions: AI-based tools showed generally favorable screening performance for OSA across clinically relevant AHI thresholds, although wide PIs suggest variable performance across future comparable populations and settings. By synthesizing diagnostic accuracy across 3 AHI thresholds and distinguishing non-PSG–derived from PSG-derived models, this review extends previous broad or modality-specific reviews and offers a clinically interpretable, pathway-specific basis for linking model performance to intended use. The findings may clarify potential roles for non-PSG–derived tools in front-end screening and referral prioritization and for PSG-derived models in reduced-channel assessment and sleep-laboratory workflow support. Given substantial heterogeneity, limited external validation, and low or very low certainty of evidence, prospective validation is needed before routine implementation.

Generative AI Use, Perceived Usefulness, Perceived Risk, and Physician Burnout and Fulfillment Among Chinese Physicians: Mixed Methods Multiregional Study

Background: As generative AI (GenAI) becomes increasingly prevalent, its impact on physician mental health has garnered significant attention; yet, empirical evidence remains limited. Objective: This study aims to investigate the correlations between the usage frequency of GenAI, perceived usefulness (PU), and perceived risk (PR) of GenAI with physicians’ burnout and professional fulfillment. Methods: A mixed methods design was used, integrating a quantitative survey of physicians across 4 regions in China with in-depth qualitative interviews to elucidate the underlying psychological mechanisms. The quantitative component involved a cross-sectional survey of 961 physicians, with the questionnaire collecting data on demographic and professional characteristics, socioeconomic status, GenAI usage frequency, PU, and PR. Semistructured interviews with 10 physicians were used for in-depth mining. Multivariable logistic and linear regression models with province-level fixed effects were fitted to examine the association between usage of GenAI, PU, PR, and physicians’ burnout and fulfillment. Stratified analyses were further performed to explore the moderating effect of demographic and clinical characteristics. Results: Quantitative analysis revealed no direct correlation between GenAI usage frequency and burnout. However, PU was positively associated with professional fulfillment (odds ratio [OR] 1.56, 95% CI 1.17-2.08; P=.003), whereas PR was associated with a higher likelihood of burnout (OR 1.80, 95% CI 1.46-2.21; P<.001). Stratified analyses showed that for physicians working ≥3 night shifts per week, GenAI usage was associated with higher odds of burnout, although the estimate was imprecise (OR 13.96, 95% CI 2.40-81.04; P=.003). The qualitative findings further suggested that the benefits of using GenAI may be offset by the additional burden. The PU of GenAI was perceived to enhance professional fulfillment by bolstering self-efficacy, whereas the PR of GenAI was linked to heightened burnout rooted in unclear boundaries of responsibilities and rights, as well as challenges to professional identity. Conclusions: The GenAI revolution in medicine is as much a psychological transition as it is a technological one. GenAI use is not directly associated with improved psychological states among clinicians. The PU of GenAI relates to professional fulfillment, and the PR concerns correspond to elevated burnout. Sustaining clinician well-being during this digital shift thus parallels a dual requirement, balancing the potential for professional fulfillment tied to GenAI utility against the concurrent verification fatigue and legal uncertainty cluster around clinician burnout.

Clinicians’ Attitudes and Perceptions on the Adoption of AI in Mental Health Care: Scoping Review

Background: AI is increasingly being integrated into health care workflows, with growing interest in AI systems for documentation, screening, triage, monitoring, and decision support. Mental health care is particularly complex for AI implementation as clinical care largely depends on therapeutic relationships, contextual factors, empathic communication, and interpretation of subtle nonverbal cues. Although AI may offer opportunities to improve efficiency and access to care, its adoption is likely to depend on clinicians’ trust, ethical acceptability, safety, confidentiality, and clarity around professional responsibility. Objective: This scoping review aims to synthesize current evidence on mental health clinicians’ attitudes, perceptions, and beliefs regarding the use of AI tools within mental health care, including the perceived benefits, risks, acceptable use cases, and conditions considered necessary for implementation. Methods: A scoping review was conducted in accordance with Joanna Briggs Institute (JBI) guidance and PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) guidelines. Six databases (CINAHL, Embase, PsycINFO, PubMed, Scopus, Web of Science) were searched on May 5, 2026, for studies published from 2020 onwards. Studies were eligible if they examined clinicians’ attitudes, perceptions, or beliefs regarding AI in mental health care. Studies were screened by 2 reviewers, first by title and abstract, and then by full text. Data extraction included publication year, country, methodology, participant occupation, AI type, outcome domains, and relevant findings. Results: Database searches retrieved 12,356 records. Following screening, 35 records were included. Overall, clinicians demonstrated cautious optimism toward AI, particularly when positioned as supplementing instead of replacing clinical expertise. Perceived benefits centered on reducing administrative burden, supporting documentation, synthesizing large volumes of information, and improving access to care, particularly between sessions. However, clinicians also reported limited AI literacy and prior use, and were concerned about privacy, confidentiality, governance, data ownership, clinical safety, unsafe or inaccurate outputs, overreliance, unclear accountability, professional role boundaries, and impacts on therapeutic relationships. Clinicians emphasized the need for education, clear guidelines and governance, role clarity, co-designed systems, human oversight, and real-world, ongoing evaluation. Conclusions: This review is innovative in shifting focus from the technical performance of AI tools to the perspectives of the clinicians who will be expected to use, interpret, explain, and remain accountable for them in mental health care. Unlike previous reviews, this review provides a clinician-centered understanding of AI adoption, highlighting that acceptability depends not only on what AI can do but also on whether it can be integrated safely, ethically, and in ways that preserve professional judgment and therapeutic relationships. These findings suggest that implementation should begin with lower-risk (eg, administrative), clinician-facing applications, be supported by education and governance, and be evaluated in mental health settings before wider adoption. These insights provide practical direction for responsible, clinician-centered implementation across mental health services.

What Provider Frequently Asked Questions Miss: Evaluating Unmet Attention-Deficit/Hyperactivity Disorder Information Needs Through Comparison of Online Community Posts Using Large Language Model–Assisted Semantic Analysis in a Mixed Methods Study

Background: Attention-deficit/hyperactivity disorder (ADHD) is a prevalent neurodevelopmental disorder that affects the functioning and quality of life of individuals throughout their lifespan. Despite the extensive information available online, patients and caregivers continue to report unmet needs, particularly regarding diagnosis, treatment, medication effects, comorbidities, and long-term management strategies. Existing provider-generated frequently asked questions (FAQs) are widely used, but often fail to fully capture the concerns expressed in online communities. Objective: This study aimed to (1) evaluate the extent to which provider-generated ADHD FAQs cover questions from online communities, (2) identify unmet information needs by analyzing questions with low semantic similarity to FAQs, and (3) compare the response styles of provider-generated answers with community-generated answers through large language model (LLM)–assisted analysis. Methods: ADHD-related questions from a Korean online community were semantically compared with provider-generated FAQs using sentence embedding–based similarity analysis to assess coverage and identify matched versus unmatched questions. Unmatched questions underwent topic modeling using the LimTopic framework, integrating BERTopic with LLM-assisted summarization to uncover unmet needs. An LLM-assisted content analysis was conducted on the answers to the high-similarity FAQ–community question pairs, enabling an examination of the response styles used by each group when addressing the public. Results: Through similarity comparison using embedding models and manual verification, the paraphrase-multilingual-MiniLM-L12-v2 (MBERT) model, which achieved the highest -score of 0.45, was selected as the final embedding model. The optimal similarity threshold determined for this model was 0.766, and the coverage of questions with similarity above this threshold between FAQs and the online community was 52.09% (2598/4988). Most of the coverage was concentrated on 18 FAQs. Online community questions below the similarity threshold were reviewed by experts after LimTopic analysis, resulting in the identification of 12 categories of unmet consumer needs, including school and social support, treatment accessibility, psychological support, and comorbidity management. Response style analysis revealed significant differences between evidence and authority signaling and the actionability dimension. Conclusions: Provider-generated ADHD FAQs covered approximately half of consumers’ questions, revealing substantial gaps in information provision for patients with ADHD. Health information on ADHD should expand beyond basic medical knowledge to address consumers’ real-world experiences, including access to care, school and social support, evidence-based treatments, daily functioning strategies, psychological support, health care navigation, and comorbidity management. Integrating these elements into provider-generated FAQs can create more comprehensive, consumer-centered resources that better support informed decision-making and long-term self-management.

Cost-Effectiveness of Electronic Patient-Reported Outcome Measure Interventions in Cancer: Systematic Review and Parameter Extraction for Economic Modeling

Background: Complex digital interventions that integrate electronic patient-reported outcome measures (ePROM) into clinical practice in cancer have the potential to improve quality of life, increase survival, and reduce health resource use and costs. Such systems can help patients with cancer self-manage chemotherapy symptoms, reduce clinicians’ workloads through automated decision support, and resolve problems earlier. However, more research on the cost-effectiveness of ePROM monitoring is needed. Objective: This paper comprises two complementary components: (1) a systematic literature review summarizing and evaluating the quantitative and qualitative evidence related to the cost-effectiveness of ePROM monitoring and (2) a health economic model parameter extraction. We also conducted supplementary targeted searches and scoping to provide context to our findings. Methods: We searched Ovid (including MEDLINE and Embase), Scopus, and the International Health Technology Assessment Database for original English-language papers published on or before March 2025 using search strings that combined terms related to ePROMs, health economics, and cancer/oncology. We included papers reporting health economic–related outcomes for ePROM interventions designed for adult cancer populations and excluded screening tools and conference abstracts. Results: We included 34 publications from 27 unique studies and identified and analyzed 26 ePROM-integrated interventions within these. Most (23/26) of the included interventions explicitly described some form of alert handling and automated decision support based on remote ePROM monitoring. Of the 34 publications, 5 presented full cost-effectiveness analysis results, of which 3 were highly uncertain and lacked clear differences in costs and health outcomes between ePROMs and standard care; conversely, 2 presented strong evidence of cost-effectiveness due to quality-of-life improvements, reduced hospitalizations, and potentially more autonomy in health-related travel (eg, ePROM-monitored patients can drive or walk to the hospital instead of using taxis or ambulances). A further 5 publications reported partial health economic results (eg, cost-consequence and budget impact), of which 1 detected no difference in strategies; in contrast, 4 reported lower health resource use and costs of ePROMs, mainly due to hospitalization reductions. Overall, 12 of the 27 studies included a qualitative component but mostly focused on user experience and design-related themes; only 2 of these addressed economic-specific themes (eg, changes in workflow and resource use due to ePROM implementation and integration), indicating some potential for time saving due to ePROM monitoring. Conclusions: Some ePROM-integrated interventions demonstrated cost-effectiveness in cancer care, but the evidence base remains limited. Where evidence does exist, cost-effectiveness appears driven by reduced hospitalization and improved quality of life. Qualitative research within the included studies rarely addressed economic questions. We provide a detailed parameter extraction for use in future economic modeling and recommend research priorities, including quantitative mapping of ePROM symptom data onto health resource use patterns, and qualitative work exploring how ePROM implementation affects clinical workloads and patient-perspective costs.
Received — 10 September 2026 ⏭ Journal of Medical Internet Research

Effectiveness, Safety, and Workflow Burden of Large Language Model–Based Medical Report Generation: Systematic Review

Background: Systems based on large language models (LLMs), multimodal LLMs, and vision-language foundation models are increasingly being evaluated for medical report generation in imaging and related clinical workflows. Existing reviews have summarized technical architectures, radiology applications, readability, and benchmark performance, but clinical readiness remains uncertain because safety, human oversight, and workflow outcomes are sparsely and inconsistently reported. Objective: The aim of this study is to assess the effectiveness (expert acceptance and blinded preference), safety (clinically significant, omission, and commission errors), and workflow burden (reporting time, corrections, edit distance, and editing burden) of LLM-based medical report generation. Methods: We searched PubMed/MEDLINE, Embase, Web of Science Core Collection, Scopus, and the Cochrane Library for studies published from January 1, 2016, through May 15, 2026. Eligible studies evaluated LLMs, multimodal LLMs, or vision-language foundation models for image-to-report generation, impression generation from findings, report drafting, or structured reporting in imaging workflows. Two reviewers performed screening, extraction, risk-of-bias assessment, and Grading of Recommendations Assessment, Development, and Evaluation–informed narrative certainty assessment. Outcomes were clinically significant error rate, omission error rate, commission error rate, reporting time, edit burden, expert acceptance, and blinded expert preference. Meta-analysis was not performed because no comparable outcome had at least 2 studies with compatible task structure and analyzable data. Results: A total of 101 studies were included. Chest x-ray was the largest modality group (36 studies), followed by computed tomography, magnetic resonance imaging (MRI), ultrasound, endoscopy, pathology, ophthalmic, electrocardiographic, dental, and mixed-modality contexts. No study was judged at low risk of bias; 15 were moderate, 72 high, and 14 serious. Safety and workflow evidence remained heterogeneous and largely nonpoolable. In a chest x-ray study, AI report acceptance was similar to that of radiologist reports (6047/8580, 70.5% vs 6288/8580, 73.3%), but false-negative findings were slightly higher (1584/8580, 18.5% vs 1527/8580, 17.8%). In a clinician-collaboration chest x-ray study, AI reports were equivalent or preferred in 233 of 300 (77.7%) and 170 of 303 (56.1%) cases across 2 datasets; yet, clinically significant errors persisted. In a brain MRI study, AI assistance reduced reading time from 61 to 53 seconds, whereas impression drafting increased editing time and edit distance. Conclusions: This review shifts the synthesis from plausible report generation to clinically interpretable effectiveness, safety, and workflow effects. Expert acceptance and preference suggested assistive value in selected supervised settings, but these signals were limited by inconsistent reporting of clinically significant errors, omissions, commissions, and failed generations. Workflow effects were mixed, with some studies reporting shorter reading time or drafting support, and others reporting greater editing time or edit distance. The evidence remains too heterogeneous, biased, and sparse on case-level end points to support a pooled meta-analysis or autonomous clinical-readiness claims. Adoption should remain locally validated, clinician-supervised, and accompanied by standardized reporting of acceptance, preference, omissions, commissions, failed generations, reporting time, corrections, and editing burden. Trial Registration: PROSPERO CRD420261302844; https://www.crd.york.ac.uk/PROSPERO/view/CRD420261302844

Smartphone-Based Monitoring of Quality of Life and Adverse Events After Neurosurgery: Prospective Cohort Study

Background: Postoperative outcome assessment is often based on discrete follow-up visits, limiting characterization of individual recovery trajectories, and the timely identification of adverse events (AEs). Longitudinal smartphone-based monitoring may overcome these limitations by enabling frequent, resource-efficient collection of patient-reported outcomes and complications throughout recovery. Such data may provide a more patient-centered understanding of the postoperative course and complement conventional clinical surveillance. Objective: This study aimed to evaluate the feasibility of smartphone-based longitudinal monitoring of quality of life, subjective well-being, and AEs after elective neurosurgery and compare postoperative recovery trajectories and agreement between patient- and clinician-reported AEs. Methods: This interim analysis of a prospective cohort study included adult patients undergoing elective lumbar decompression, lumbar fusion, supratentorial craniotomy, or infratentorial craniotomy at a Swiss tertiary referral center between June 2023 and January 2025. Participants used a smartphone app to longitudinally report subjective well-being (Subjective Well-Being Index; 0‐10), quality of life (EQ-5D-5L), and AEs for up to 1 year postoperatively. Complications were self-reported using the Therapy-Disability-Neurology (TDN) classification and retrospectively adjudicated by physicians. Descriptive analyses assessed data density, engagement, and concordance between patient- and clinician-reported events. Mixed-effects models were used to evaluate factors associated with postoperative well-being. Results: Of the 100 enrolled patients (median age 64.0, IQR 52.95‐71.6 years; n=45, 45% women), 86 (86%) provided postoperative data. During a median follow-up of 3.2 (IQR 0.2‐11.3) months, participants submitted 4354 longitudinal well-being entries. Patients reported 22 unique AEs, whereas physicians identified 44 AEs, with overlap for 9 (20.5%) events. Most physician-reported AEs were mild (30/44, 68.2%; TDN grade 1‐2), and no grade 4 or 5 events occurred. Patient-reported AEs primarily reflected symptomatic and functional impairments, whereas physician-reported events more often included clinically detected or subclinical findings. In mixed-effects models, time since surgery was associated with improved well-being, and no other factors were statistically significant. Conclusions: Smartphone-based postoperative monitoring was feasible in this elective neurosurgical cohort and generated dense longitudinal patient-reported data beyond routine follow-up. Patient and clinician AE reporting captured partly distinct aspects of postoperative recovery, suggesting that smartphone-based self-reporting may complement rather than replace clinical surveillance. Trial Registration: ClinicalTrials.gov NCT06352710; https://clinicaltrials.gov/study/NCT06352710

Natural Language Processing Identification of Nonprescribed Fentanyl Use in Electronic Health Records: Algorithm Development and Validation Study

Background: Overdose and suicide due to nonprescribed fentanyl use have increased significantly, yet health care systems lack reliable methods to identify patients who use nonprescribed fentanyl. codes are inconsistent and do not specify nonprescribed fentanyl use. Objective: This study aimed to develop natural language processing approaches to identifying nonprescribed fentanyl use in electronic health record (EHR) documentation. Methods: This retrospective study included Veterans Health Administration patients seen between April 5, 2023, and December 23, 2024. A term list was developed to identify fentanyl-related mentions in clinical text, and 250-character snippets surrounding identified mentions were extracted. Veterans (n=3878) were randomly sampled from 5 predefined groups based on the presence of 1 of 4 terms (“fent,” “blues,” “M30s,” and “tranq”) in their EHR documentation. Physician annotators classified snippets into “nonprescribed fentanyl use,” “prescribed fentanyl use,” or “other,” with interannotator agreement evaluated using the mean pairwise Cohen κ. Cross-validation folds were constructed at the patient level between training and test sets. Penalized logistic regression, Bio-ClinicalBERT, Llama 3-8B, and Mistral-7B were trained on labeled data and compared. Model performance was evaluated using precision, recall, and -scores for each class, with a focus on the nonprescribed fentanyl use class as the primary label of clinical interest using bootstrapped 95% CIs. A fairness analysis and Shapley additive explanations analysis were performed using Bio-ClinicalBERT. External validation was performed using Bio-ClinicalBERT on an independent sample of 200 snippets, each representing a unique patient from January 2025 to June 2026, with precision reported as the primary validation metric. Results: Of 7389 snippets, 9.6% (n=709) were classified as “nonprescribed fentanyl use,” 40.3% (n=2981) were classified as “prescribed fentanyl use,” and 50% (n=3699) were classified as “other.” Interannotator agreement was high (κ=0.822). Llama 3-8B achieved the highest -score for nonprescribed fentanyl use (0.87, 95% CI 0.83-0.92), followed by Mistral-7B (0.80, 95% CI 0.75-0.84), Bio-ClinicalBERT (0.80, 95% CI 0.74-0.85), and penalized logistic regression (0.74, 95% CI 0.73-0.75). Performance was consistent across demographic subgroups, with lower performance for the nonprescribed fentanyl use class observed in female and Hispanic subgroups. Shapley additive explanations analysis revealed clinically meaningful discriminating terms for each class, although subword tokens required contextual interpretation. External validation of Bio-ClinicalBERT demonstrated a precision of 0.79 for nonprescribed fentanyl use. Conclusions: Natural language processing can identify nonprescribed fentanyl use in EHR documentation, although model performance for this class was lower than overall model performance, reflecting the clinical complexity of identifying nonprescribed use and the variable ways in which clinicians document this problem. This approach may support risk prediction and targeting of interventions to patients exposed to nonprescribed fentanyl.

Large Language Models for Distress Rating in Korean Psycho-Oncology Interviews: Exploratory Clinician-Benchmarked Evaluation Study

Background: Psychiatric distress is common among patients with cancer; yet, systematic interview-based screening remains difficult to scale in routine clinical care. Large language models (LLMs) have shown promise as scalable tools for mental health assessment, but most existing evidence is derived from clinician-authored records, translated text, or proxy data. The performance characteristics, error patterns, and explanatory behaviors of contemporary LLMs when applied to authentic, non-English psychiatric interviews remain insufficiently characterized. Objective: This exploratory study evaluated how contemporary LLMs reproduced psycho-oncologists’ item-level symptom ratings from real-world Korean psycho-oncology interviews, focusing on concordance, directional bias, and clinician-adjudicated error characteristics. Methods: Between April 2024 and May 2025, 101 adults receiving oncologic care in South Korea underwent semistructured interviews. Board-certified psycho-oncologists provided real-time, time-stamped ratings of 39 items. Generative pretrained transformer 4o (GPT-4o), Claude 3.5 Sonnet, and Gemini 2.5 Flash generated item scores and brief rationales using identical Korean zero-shot rubrics. Concordance between clinician ratings was evaluated using ordinal and binary screening metrics. A paired Wilcoxon test compared patient-level total symptom burden. Binary mismatches were clinically adjudicated as ambiguous or definite overestimation or underestimation with an 8-etiology taxonomy. Model-generated rationales were further meta-evaluated using GPT-5.4, Claude Sonnet 4.6, and Gemini Pro 3.1 across 4 dimensions: citation (use of quoted supporting statements), structure (logical organization of the rationale), mapping (consistency between the rationale and the assigned item rating), and expansion (degree of interpretive elaboration beyond the explicit transcript content). The associations between meta-evaluation results and absolute error were examined using cross-classified mixed-effects models. Results: Out of 101 participants, 88 (87.1%) were predominantly female and had breast cancer as the primary cancer type (n=70, 69.3%). Across 3931 item-level ratings, all models showed good agreement with clinicians (intraclass correlation coefficient: 0.816‐0.872), with GPT-4o showing the highest agreement. Claude 3.5 and Gemini 2.5 yielded significantly higher patient-level symptom burden (adjusted

Comparative Efficacy of Different AI Systems for Polyp Detection by Size During Colonoscopy: Systematic Review and Network Meta-Analysis

Background: Colorectal cancer remains a leading cause of death despite being largely preventable through polypectomy. AI systems designed to enhance polyp detection during colonoscopy have shown promise, but the extent to which they improve detection of different-sized polyps remains unclear. Objective: This study compared the size-stratified efficacy of AI-assisted colonoscopy vs standard colonoscopy using the Hartung-Knapp-Sidik-Jonkman (HKSJ) method, and generated exploratory rankings while acknowledging all cross-platform comparisons are indirect. Methods: This systematic review and network meta-analysis (NMA) searched PubMed, Embase, Cochrane CENTRAL, and Web of Science from inception to July 25, 2026, supplemented by citation searching. We included randomized controlled trials (RCTs) comparing AI-assisted vs standard colonoscopy in adults (≥18 years of age), reporting mean polyp detection counts stratified by size (≤5 mm, 6-9 mm, and ≥10 mm). Two reviewers screened studies, extracted data, and assessed risk of bias using the Cochrane Risk of Bias 2.0. We conducted frequentist NMA using the HKSJ method with restricted maximum likelihood estimation, calculated 95% prediction intervals (PIs), and assessed heterogeneity using I2 and τ2. Certainty of evidence was rated using the GRADE (Grading of Recommendations Assessment, Development, and Evaluation) framework. Results: A total of 13 RCTs (4156 participants) compared 8 AI systems to standard colonoscopy, forming a network without direct AI comparisons. For diminutive polyps (≤5 mm), AI showed a modest advantage (standardized mean difference [SMD] 0.21, 95% CI 0.07 to 0.35, 95% PI –1.12 to 1.54), but substantial heterogeneity (I2=86.6%) and wide PI crossing the null indicated high uncertainty. EndoScreener showed the most consistent evidence (SMD 0.36, 95% CI 0.18-0.54). For small and large polyps, effects were minimal (SMD 0.02, 95% CI –0.02 to 0.06, 95% PI –0.03 to 0.07; SMD 0.01, 95% CI 0.00-0.02, 95% PI –0.01 to 0.03). GRADE certainty was very low for diminutive polyps and low for small and large polyps. Sensitivity analysis excluding Tianjin YuJin did not materially change findings. Conclusions: AI may modestly enhance diminutive polyp detection, but effects on small and large polyps are minimal, with no platform superiority. Given very low to low certainty, findings are hypothesis-generating. This exploratory NMA provides size-stratified comparisons that can inform future head-to-head trial design. Unlike prior reviews aggregating all polyp sizes, we show the overall AI benefit is driven by diminutive polyp detection, providing a framework for targeted deployment—prioritizing AI for diminutive polyp screening, with limited value for larger lesions. Head-to-head trials are urgently needed. Trial Registration: PROSPERO International Prospective Register of Systematic Reviews CRD420251266932; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251266932

“Small” Large Language Models in the Hospital: Evaluation Study on Real-World Data in a Resource-Constrained Setting

Background: Large language models (LLMs) are increasingly being deployed in health care, but their use and deployment in many real-world hospital environments pose significant challenges and concerns. In particular, state-of-the-art commercial models store or process data externally, which is often in conflict with ensuring patient data protection. At the same time, using LLMs locally is limited by the lack of available computing infrastructure. Small open-source LLMs that do not require substantial computing resources could offer a practical way to resolve these tensions, but their medical utility in real-world local contexts, especially in non-English languages, has not been sufficiently evaluated. Objective: This study aimed to evaluate the feasibility of small, locally deployable open-source LLMs for clinically relevant tasks in a resource-constrained hospital setting and to propose a reproducible framework for institution-specific evaluation before deployment. Methods: We evaluated 6 open-source LLMs ranging from 8B to 24B parameters (from the Mistral, Phi4, Falcon3, Llama3.1, and Meditron3 families) in a zero-shot setting across 7 tasks covering 4 clinical use cases: information extraction, medical text translation, text generation, and clinical decision support. We used deidentified French clinical data from a Swiss tertiary hospital, including discharge letters, clinical notes, and structured electronic health records. Performance was assessed using task-specific metrics, such as precision, recall, F1-score, embedding-based semantic similarity, recall-oriented understudy for gisting evaluation (ROUGE) score, readability indices, and human review by clinicians. Results: Model performance varied substantially between tasks. In the simplest retrieval task (needle-in-the-haystack), several models performed strongly, with Llama3.1 achieving an F1-score of 99.81% and Mistral-small achieving 99.71%. In contrast, performance was poor in more complex tasks. For detecting protected health information, the best-performing LLMs achieved only modest overall macro–F1-scores (0.33-0.34), substantially below a fine-tuned Robustly Optimized BERT Pretraining Approach (RoBERTa) baseline (0.94). In the task of extracting immune-related adverse events from discharge notes, the highest overall macro–F1-score was 0.35 with Phi4. For medical text translation, Phi4 ranked highest in embedding-based evaluation, whereas Meditron3-Phi4 performed the worst, with clinician reviews identifying hallucinations in 55% of its outputs. In the task of summarizing discharge letters, quality was low across all models, with the best penalized ROUGE score reaching only 0.169 with Llama3.1. In the tasks of generating patient-friendly discharge note summaries and clinical decision support, clinician ratings generally ranged from dissatisfied to neutral, and no model achieved consistently satisfactory performance. Conclusions: Small open-source LLMs appear feasible for simple retrieval-oriented tasks in local hospital deployments but are currently inadequate for more complex applications, such as clinical decision support, deidentification, extraction of adverse events, and medical summarization. These findings highlight the importance of locally grounded evaluation tailored to specific use cases and the need for robust institutional evaluation frameworks to ensure safe and reliable deployment. Trial Registration:

Designing a Gamified mHealth App for HIV Prevention and Comorbidities Among Malaysian Young Men Who Have Sex With Men: An Interdisciplinary Expert Panel Study

Background: New HIV cases among Malaysian men who have sex with men continue to rise, with young men who have sex with men (YMSM) accounting for 44% of new infections and experiencing high rates of comorbidities. Mobile health (mHealth) apps offer a promising approach to addressing these challenges by providing discreet access to health information, screening tools, and linkage to services. Given the near-universal smartphone ownership among Malaysian YMSM and high levels of mobile gaming engagement, gamified mHealth apps may be particularly effective in sustaining engagement and promoting HIV prevention behaviors and comorbidity management. However, realizing their full potential requires identifying the features and design principles that are most important to Malaysian YMSM and that can support sustained engagement and improve health outcomes in this vulnerable population. Objective: This study used an interdisciplinary expert panel approach to identify the key features, design principles, and gamification elements to be incorporated into MY-Hero, a gamified mHealth app designed to support HIV prevention, comorbidity management, and overall well-being among Malaysian YMSM. Methods: Guided by the integrated behavioral model (IBM), we conducted an interdisciplinary expert panel comprising experts in health, technology, and design sciences, as well as health care professionals with expertise in the Malaysian YMSM population, along with local lesbian, gay, bisexual, transgender, queer, and others (LGBTQ+) community leaders. Through a structured series of discussions and thematic analysis, we identified culturally appropriate design principles and key features for MY-Hero that align with the health needs and sociocultural context of Malaysian YMSM. Results: Three expert panel sessions were conducted with 9 experts. Several key themes emerged: (1) the importance of establishing clinical affiliation and facilitating linkage to HIV testing, pre-exposure prophylaxis (PrEP), and related services, including harm reduction services; (2) the need to enhance user interface (UI) and user experience (UX) by optimizing usability, interactivity, and engagement to maintain user interest; (3) the incorporation of customizable health content to tailor interventions based on individual characteristics and preferences; (4) the use of gamification mechanisms, such as reward systems and progression tracking to promote adoption and sustained engagement with health services; and (5) the importance of privacy and data security as critical design considerations for ensuring user safety, confidentiality, and trust. Conclusions: The findings highlight the importance of user-centered design, contextualized health content, and gamification mechanisms in mHealth tools for Malaysian YMSM. These insights provide practical guidance for developing culturally appropriate gamified mHealth interventions for vulnerable populations by informing strategies to enhance user engagement, reduce HIV prevention fatigue, and support the well-being of YMSM in stigmatized contexts.

Extended Reality Interventions for Osteoarthritis of the Knee and Recovery After Total Knee Arthroplasty: Systematic Review and Meta-Analyses

Background: Nonpharmacologic interventions are important for treating knee pain due to osteoarthritis or after total knee arthroplasty (TKA), and extended reality (XR) technology may enhance treatments for these indications. Objective: This systematic review aimed to evaluate XR interventions for pain due to knee osteoarthritis (KOA) or for recovery after TKA. Methods: Databases were searched through May 2023 and updated in December 2025. Eligible trials evaluated XR interventions to treat KOA pain or after TKA. We classified interventions by depth of immersion and clinical mechanism. We used the Grading of Recommendations Assessment, Development, and Evaluation (GRADE) criteria to determine the certainty of evidence for prioritized outcomes. Meta-analyses were performed when ≥3 studies evaluated similar comparisons, outcomes, and time points. Results: Eligible trials addressed KOA (k=12) or recovery after TKA (k=9). Sample sizes ranged from 36 to 306 participants, and most studies had a follow-up of ≤3 months. Nineteen studies assessed pain-related functioning and pain intensity, and 5 assessed adverse events (AEs). For KOA, 10 studies examined interactive digital rehabilitation (IDR), and 2 examined virtual reality (VR)–digitally augmented exercise (DAE). IDR for KOA may result in better pain-related functioning (low certainty of evidence [COE]; pooled standardized mean difference [SMD] −0.59, 95% CI −1.11 to −0.06; prediction interval [PI] −1.72 to 0.55; k=5) and lower pain intensity at 6‐8 weeks (low COE; pooled SMD −0.46, 95% CI −0.92 to 0.00; PI −1.39 to 0.47; k=4). VR-DAE for KOA (k=2) produced inconsistent results (very low COE). For post-TKA studies, 5 examined IDR, 2 examined VR-DAE, 1 examined VR-distraction, and 1 examined VR-psychoeducation. Post-TKA IDR may result in better pain-related functioning (low [k=4] and moderate COE [k=1]) but little to no difference in pain intensity (low-moderate COE; pooled SMD at 3‐4 months −0.12, 95% CI −0.75 to 0.52; PI –1.63 to 1.27; k=3). VR-psychoeducation probably results in lower pain at 4 weeks (moderate COE; k=1), and VR-distraction may result in 6 months (low COE; k=1), whereas VR-DAE produced mixed findings (k=2; very low COE). IDR was not associated with AEs, and VR may not be associated with AEs for KOA (high and low COE), though AE reporting was uncommon (k=5) and evidence was very uncertain for post-TKA. Conclusions: IDR may augment treatment for KOA and post-TKA recovery, and VR may benefit post-TKA rehabilitation. This review is the first to stratify by level of immersion, clinical mechanism, and follow-up duration and to systematically evaluate AEs. IDR may be ready for integration into KOA care, while use after TKA needs more evidence. Randomized controlled trials with implementation outcomes could determine how XR interventions can be used for KOA, whereas trials evaluating efficacy and AEs are needed before their use for post-TKA. Trial Registration: PROSPERO CRD42023439903; https://www.crd.york.ac.uk/PROSPERO/view/CRD42023439903

Closing the Solidarity Gap Requires Closing the Accountability Gap for Patient-Facing AI

This commentary extends recent discussion of the solidarity gap associated with patient-facing AI by examining gaps in governance and risk allocation, health care professionals’ responsibilities in practice, and opportunities for professional stewardship and advocacy. We argue that equitable implementation requires shared accountability and meaningful health care professional participation in the design, evaluation, reimbursement, governance, and oversight of patient-facing AI before ambiguity results in patient harm.
❌