❌

Normal view

Received — 10 September 2026 ⏭ Journal of Medical Internet Research

Effectiveness, Safety, and Workflow Burden of Large Language Model–Based Medical Report Generation: Systematic Review

Background: Systems based on large language models (LLMs), multimodal LLMs, and vision-language foundation models are increasingly being evaluated for medical report generation in imaging and related clinical workflows. Existing reviews have summarized technical architectures, radiology applications, readability, and benchmark performance, but clinical readiness remains uncertain because safety, human oversight, and workflow outcomes are sparsely and inconsistently reported. Objective: The aim of this study is to assess the effectiveness (expert acceptance and blinded preference), safety (clinically significant, omission, and commission errors), and workflow burden (reporting time, corrections, edit distance, and editing burden) of LLM-based medical report generation. Methods: We searched PubMed/MEDLINE, Embase, Web of Science Core Collection, Scopus, and the Cochrane Library for studies published from January 1, 2016, through May 15, 2026. Eligible studies evaluated LLMs, multimodal LLMs, or vision-language foundation models for image-to-report generation, impression generation from findings, report drafting, or structured reporting in imaging workflows. Two reviewers performed screening, extraction, risk-of-bias assessment, and Grading of Recommendations Assessment, Development, and Evaluation–informed narrative certainty assessment. Outcomes were clinically significant error rate, omission error rate, commission error rate, reporting time, edit burden, expert acceptance, and blinded expert preference. Meta-analysis was not performed because no comparable outcome had at least 2 studies with compatible task structure and analyzable data. Results: A total of 101 studies were included. Chest x-ray was the largest modality group (36 studies), followed by computed tomography, magnetic resonance imaging (MRI), ultrasound, endoscopy, pathology, ophthalmic, electrocardiographic, dental, and mixed-modality contexts. No study was judged at low risk of bias; 15 were moderate, 72 high, and 14 serious. Safety and workflow evidence remained heterogeneous and largely nonpoolable. In a chest x-ray study, AI report acceptance was similar to that of radiologist reports (6047/8580, 70.5% vs 6288/8580, 73.3%), but false-negative findings were slightly higher (1584/8580, 18.5% vs 1527/8580, 17.8%). In a clinician-collaboration chest x-ray study, AI reports were equivalent or preferred in 233 of 300 (77.7%) and 170 of 303 (56.1%) cases across 2 datasets; yet, clinically significant errors persisted. In a brain MRI study, AI assistance reduced reading time from 61 to 53 seconds, whereas impression drafting increased editing time and edit distance. Conclusions: This review shifts the synthesis from plausible report generation to clinically interpretable effectiveness, safety, and workflow effects. Expert acceptance and preference suggested assistive value in selected supervised settings, but these signals were limited by inconsistent reporting of clinically significant errors, omissions, commissions, and failed generations. Workflow effects were mixed, with some studies reporting shorter reading time or drafting support, and others reporting greater editing time or edit distance. The evidence remains too heterogeneous, biased, and sparse on case-level end points to support a pooled meta-analysis or autonomous clinical-readiness claims. Adoption should remain locally validated, clinician-supervised, and accompanied by standardized reporting of acceptance, preference, omissions, commissions, failed generations, reporting time, corrections, and editing burden. Trial Registration: PROSPERO CRD420261302844; https://www.crd.york.ac.uk/PROSPERO/view/CRD420261302844

Smartphone-Based Monitoring of Quality of Life and Adverse Events After Neurosurgery: Prospective Cohort Study

Background: Postoperative outcome assessment is often based on discrete follow-up visits, limiting characterization of individual recovery trajectories, and the timely identification of adverse events (AEs). Longitudinal smartphone-based monitoring may overcome these limitations by enabling frequent, resource-efficient collection of patient-reported outcomes and complications throughout recovery. Such data may provide a more patient-centered understanding of the postoperative course and complement conventional clinical surveillance. Objective: This study aimed to evaluate the feasibility of smartphone-based longitudinal monitoring of quality of life, subjective well-being, and AEs after elective neurosurgery and compare postoperative recovery trajectories and agreement between patient- and clinician-reported AEs. Methods: This interim analysis of a prospective cohort study included adult patients undergoing elective lumbar decompression, lumbar fusion, supratentorial craniotomy, or infratentorial craniotomy at a Swiss tertiary referral center between June 2023 and January 2025. Participants used a smartphone app to longitudinally report subjective well-being (Subjective Well-Being Index; 0‐10), quality of life (EQ-5D-5L), and AEs for up to 1 year postoperatively. Complications were self-reported using the Therapy-Disability-Neurology (TDN) classification and retrospectively adjudicated by physicians. Descriptive analyses assessed data density, engagement, and concordance between patient- and clinician-reported events. Mixed-effects models were used to evaluate factors associated with postoperative well-being. Results: Of the 100 enrolled patients (median age 64.0, IQR 52.95‐71.6 years; n=45, 45% women), 86 (86%) provided postoperative data. During a median follow-up of 3.2 (IQR 0.2‐11.3) months, participants submitted 4354 longitudinal well-being entries. Patients reported 22 unique AEs, whereas physicians identified 44 AEs, with overlap for 9 (20.5%) events. Most physician-reported AEs were mild (30/44, 68.2%; TDN grade 1‐2), and no grade 4 or 5 events occurred. Patient-reported AEs primarily reflected symptomatic and functional impairments, whereas physician-reported events more often included clinically detected or subclinical findings. In mixed-effects models, time since surgery was associated with improved well-being, and no other factors were statistically significant. Conclusions: Smartphone-based postoperative monitoring was feasible in this elective neurosurgical cohort and generated dense longitudinal patient-reported data beyond routine follow-up. Patient and clinician AE reporting captured partly distinct aspects of postoperative recovery, suggesting that smartphone-based self-reporting may complement rather than replace clinical surveillance. Trial Registration: ClinicalTrials.gov NCT06352710; https://clinicaltrials.gov/study/NCT06352710

Natural Language Processing Identification of Nonprescribed Fentanyl Use in Electronic Health Records: Algorithm Development and Validation Study

Background: Overdose and suicide due to nonprescribed fentanyl use have increased significantly, yet health care systems lack reliable methods to identify patients who use nonprescribed fentanyl. codes are inconsistent and do not specify nonprescribed fentanyl use. Objective: This study aimed to develop natural language processing approaches to identifying nonprescribed fentanyl use in electronic health record (EHR) documentation. Methods: This retrospective study included Veterans Health Administration patients seen between April 5, 2023, and December 23, 2024. A term list was developed to identify fentanyl-related mentions in clinical text, and 250-character snippets surrounding identified mentions were extracted. Veterans (n=3878) were randomly sampled from 5 predefined groups based on the presence of 1 of 4 terms (“fent,” “blues,” “M30s,” and “tranq”) in their EHR documentation. Physician annotators classified snippets into “nonprescribed fentanyl use,” “prescribed fentanyl use,” or “other,” with interannotator agreement evaluated using the mean pairwise Cohen κ. Cross-validation folds were constructed at the patient level between training and test sets. Penalized logistic regression, Bio-ClinicalBERT, Llama 3-8B, and Mistral-7B were trained on labeled data and compared. Model performance was evaluated using precision, recall, and -scores for each class, with a focus on the nonprescribed fentanyl use class as the primary label of clinical interest using bootstrapped 95% CIs. A fairness analysis and Shapley additive explanations analysis were performed using Bio-ClinicalBERT. External validation was performed using Bio-ClinicalBERT on an independent sample of 200 snippets, each representing a unique patient from January 2025 to June 2026, with precision reported as the primary validation metric. Results: Of 7389 snippets, 9.6% (n=709) were classified as “nonprescribed fentanyl use,” 40.3% (n=2981) were classified as “prescribed fentanyl use,” and 50% (n=3699) were classified as “other.” Interannotator agreement was high (κ=0.822). Llama 3-8B achieved the highest -score for nonprescribed fentanyl use (0.87, 95% CI 0.83-0.92), followed by Mistral-7B (0.80, 95% CI 0.75-0.84), Bio-ClinicalBERT (0.80, 95% CI 0.74-0.85), and penalized logistic regression (0.74, 95% CI 0.73-0.75). Performance was consistent across demographic subgroups, with lower performance for the nonprescribed fentanyl use class observed in female and Hispanic subgroups. Shapley additive explanations analysis revealed clinically meaningful discriminating terms for each class, although subword tokens required contextual interpretation. External validation of Bio-ClinicalBERT demonstrated a precision of 0.79 for nonprescribed fentanyl use. Conclusions: Natural language processing can identify nonprescribed fentanyl use in EHR documentation, although model performance for this class was lower than overall model performance, reflecting the clinical complexity of identifying nonprescribed use and the variable ways in which clinicians document this problem. This approach may support risk prediction and targeting of interventions to patients exposed to nonprescribed fentanyl.

Large Language Models for Distress Rating in Korean Psycho-Oncology Interviews: Exploratory Clinician-Benchmarked Evaluation Study

Background: Psychiatric distress is common among patients with cancer; yet, systematic interview-based screening remains difficult to scale in routine clinical care. Large language models (LLMs) have shown promise as scalable tools for mental health assessment, but most existing evidence is derived from clinician-authored records, translated text, or proxy data. The performance characteristics, error patterns, and explanatory behaviors of contemporary LLMs when applied to authentic, non-English psychiatric interviews remain insufficiently characterized. Objective: This exploratory study evaluated how contemporary LLMs reproduced psycho-oncologists’ item-level symptom ratings from real-world Korean psycho-oncology interviews, focusing on concordance, directional bias, and clinician-adjudicated error characteristics. Methods: Between April 2024 and May 2025, 101 adults receiving oncologic care in South Korea underwent semistructured interviews. Board-certified psycho-oncologists provided real-time, time-stamped ratings of 39 items. Generative pretrained transformer 4o (GPT-4o), Claude 3.5 Sonnet, and Gemini 2.5 Flash generated item scores and brief rationales using identical Korean zero-shot rubrics. Concordance between clinician ratings was evaluated using ordinal and binary screening metrics. A paired Wilcoxon test compared patient-level total symptom burden. Binary mismatches were clinically adjudicated as ambiguous or definite overestimation or underestimation with an 8-etiology taxonomy. Model-generated rationales were further meta-evaluated using GPT-5.4, Claude Sonnet 4.6, and Gemini Pro 3.1 across 4 dimensions: citation (use of quoted supporting statements), structure (logical organization of the rationale), mapping (consistency between the rationale and the assigned item rating), and expansion (degree of interpretive elaboration beyond the explicit transcript content). The associations between meta-evaluation results and absolute error were examined using cross-classified mixed-effects models. Results: Out of 101 participants, 88 (87.1%) were predominantly female and had breast cancer as the primary cancer type (n=70, 69.3%). Across 3931 item-level ratings, all models showed good agreement with clinicians (intraclass correlation coefficient: 0.816‐0.872), with GPT-4o showing the highest agreement. Claude 3.5 and Gemini 2.5 yielded significantly higher patient-level symptom burden (adjusted

Comparative Efficacy of Different AI Systems for Polyp Detection by Size During Colonoscopy: Systematic Review and Network Meta-Analysis

Background: Colorectal cancer remains a leading cause of death despite being largely preventable through polypectomy. AI systems designed to enhance polyp detection during colonoscopy have shown promise, but the extent to which they improve detection of different-sized polyps remains unclear. Objective: This study compared the size-stratified efficacy of AI-assisted colonoscopy vs standard colonoscopy using the Hartung-Knapp-Sidik-Jonkman (HKSJ) method, and generated exploratory rankings while acknowledging all cross-platform comparisons are indirect. Methods: This systematic review and network meta-analysis (NMA) searched PubMed, Embase, Cochrane CENTRAL, and Web of Science from inception to July 25, 2026, supplemented by citation searching. We included randomized controlled trials (RCTs) comparing AI-assisted vs standard colonoscopy in adults (≥18 years of age), reporting mean polyp detection counts stratified by size (≤5 mm, 6-9 mm, and ≥10 mm). Two reviewers screened studies, extracted data, and assessed risk of bias using the Cochrane Risk of Bias 2.0. We conducted frequentist NMA using the HKSJ method with restricted maximum likelihood estimation, calculated 95% prediction intervals (PIs), and assessed heterogeneity using I2 and τ2. Certainty of evidence was rated using the GRADE (Grading of Recommendations Assessment, Development, and Evaluation) framework. Results: A total of 13 RCTs (4156 participants) compared 8 AI systems to standard colonoscopy, forming a network without direct AI comparisons. For diminutive polyps (≤5 mm), AI showed a modest advantage (standardized mean difference [SMD] 0.21, 95% CI 0.07 to 0.35, 95% PI –1.12 to 1.54), but substantial heterogeneity (I2=86.6%) and wide PI crossing the null indicated high uncertainty. EndoScreener showed the most consistent evidence (SMD 0.36, 95% CI 0.18-0.54). For small and large polyps, effects were minimal (SMD 0.02, 95% CI –0.02 to 0.06, 95% PI –0.03 to 0.07; SMD 0.01, 95% CI 0.00-0.02, 95% PI –0.01 to 0.03). GRADE certainty was very low for diminutive polyps and low for small and large polyps. Sensitivity analysis excluding Tianjin YuJin did not materially change findings. Conclusions: AI may modestly enhance diminutive polyp detection, but effects on small and large polyps are minimal, with no platform superiority. Given very low to low certainty, findings are hypothesis-generating. This exploratory NMA provides size-stratified comparisons that can inform future head-to-head trial design. Unlike prior reviews aggregating all polyp sizes, we show the overall AI benefit is driven by diminutive polyp detection, providing a framework for targeted deployment—prioritizing AI for diminutive polyp screening, with limited value for larger lesions. Head-to-head trials are urgently needed. Trial Registration: PROSPERO International Prospective Register of Systematic Reviews CRD420251266932; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251266932

“Small” Large Language Models in the Hospital: Evaluation Study on Real-World Data in a Resource-Constrained Setting

Background: Large language models (LLMs) are increasingly being deployed in health care, but their use and deployment in many real-world hospital environments pose significant challenges and concerns. In particular, state-of-the-art commercial models store or process data externally, which is often in conflict with ensuring patient data protection. At the same time, using LLMs locally is limited by the lack of available computing infrastructure. Small open-source LLMs that do not require substantial computing resources could offer a practical way to resolve these tensions, but their medical utility in real-world local contexts, especially in non-English languages, has not been sufficiently evaluated. Objective: This study aimed to evaluate the feasibility of small, locally deployable open-source LLMs for clinically relevant tasks in a resource-constrained hospital setting and to propose a reproducible framework for institution-specific evaluation before deployment. Methods: We evaluated 6 open-source LLMs ranging from 8B to 24B parameters (from the Mistral, Phi4, Falcon3, Llama3.1, and Meditron3 families) in a zero-shot setting across 7 tasks covering 4 clinical use cases: information extraction, medical text translation, text generation, and clinical decision support. We used deidentified French clinical data from a Swiss tertiary hospital, including discharge letters, clinical notes, and structured electronic health records. Performance was assessed using task-specific metrics, such as precision, recall, F1-score, embedding-based semantic similarity, recall-oriented understudy for gisting evaluation (ROUGE) score, readability indices, and human review by clinicians. Results: Model performance varied substantially between tasks. In the simplest retrieval task (needle-in-the-haystack), several models performed strongly, with Llama3.1 achieving an F1-score of 99.81% and Mistral-small achieving 99.71%. In contrast, performance was poor in more complex tasks. For detecting protected health information, the best-performing LLMs achieved only modest overall macro–F1-scores (0.33-0.34), substantially below a fine-tuned Robustly Optimized BERT Pretraining Approach (RoBERTa) baseline (0.94). In the task of extracting immune-related adverse events from discharge notes, the highest overall macro–F1-score was 0.35 with Phi4. For medical text translation, Phi4 ranked highest in embedding-based evaluation, whereas Meditron3-Phi4 performed the worst, with clinician reviews identifying hallucinations in 55% of its outputs. In the task of summarizing discharge letters, quality was low across all models, with the best penalized ROUGE score reaching only 0.169 with Llama3.1. In the tasks of generating patient-friendly discharge note summaries and clinical decision support, clinician ratings generally ranged from dissatisfied to neutral, and no model achieved consistently satisfactory performance. Conclusions: Small open-source LLMs appear feasible for simple retrieval-oriented tasks in local hospital deployments but are currently inadequate for more complex applications, such as clinical decision support, deidentification, extraction of adverse events, and medical summarization. These findings highlight the importance of locally grounded evaluation tailored to specific use cases and the need for robust institutional evaluation frameworks to ensure safe and reliable deployment. Trial Registration:

Designing a Gamified mHealth App for HIV Prevention and Comorbidities Among Malaysian Young Men Who Have Sex With Men: An Interdisciplinary Expert Panel Study

Background: New HIV cases among Malaysian men who have sex with men continue to rise, with young men who have sex with men (YMSM) accounting for 44% of new infections and experiencing high rates of comorbidities. Mobile health (mHealth) apps offer a promising approach to addressing these challenges by providing discreet access to health information, screening tools, and linkage to services. Given the near-universal smartphone ownership among Malaysian YMSM and high levels of mobile gaming engagement, gamified mHealth apps may be particularly effective in sustaining engagement and promoting HIV prevention behaviors and comorbidity management. However, realizing their full potential requires identifying the features and design principles that are most important to Malaysian YMSM and that can support sustained engagement and improve health outcomes in this vulnerable population. Objective: This study used an interdisciplinary expert panel approach to identify the key features, design principles, and gamification elements to be incorporated into MY-Hero, a gamified mHealth app designed to support HIV prevention, comorbidity management, and overall well-being among Malaysian YMSM. Methods: Guided by the integrated behavioral model (IBM), we conducted an interdisciplinary expert panel comprising experts in health, technology, and design sciences, as well as health care professionals with expertise in the Malaysian YMSM population, along with local lesbian, gay, bisexual, transgender, queer, and others (LGBTQ+) community leaders. Through a structured series of discussions and thematic analysis, we identified culturally appropriate design principles and key features for MY-Hero that align with the health needs and sociocultural context of Malaysian YMSM. Results: Three expert panel sessions were conducted with 9 experts. Several key themes emerged: (1) the importance of establishing clinical affiliation and facilitating linkage to HIV testing, pre-exposure prophylaxis (PrEP), and related services, including harm reduction services; (2) the need to enhance user interface (UI) and user experience (UX) by optimizing usability, interactivity, and engagement to maintain user interest; (3) the incorporation of customizable health content to tailor interventions based on individual characteristics and preferences; (4) the use of gamification mechanisms, such as reward systems and progression tracking to promote adoption and sustained engagement with health services; and (5) the importance of privacy and data security as critical design considerations for ensuring user safety, confidentiality, and trust. Conclusions: The findings highlight the importance of user-centered design, contextualized health content, and gamification mechanisms in mHealth tools for Malaysian YMSM. These insights provide practical guidance for developing culturally appropriate gamified mHealth interventions for vulnerable populations by informing strategies to enhance user engagement, reduce HIV prevention fatigue, and support the well-being of YMSM in stigmatized contexts.

Extended Reality Interventions for Osteoarthritis of the Knee and Recovery After Total Knee Arthroplasty: Systematic Review and Meta-Analyses

Background: Nonpharmacologic interventions are important for treating knee pain due to osteoarthritis or after total knee arthroplasty (TKA), and extended reality (XR) technology may enhance treatments for these indications. Objective: This systematic review aimed to evaluate XR interventions for pain due to knee osteoarthritis (KOA) or for recovery after TKA. Methods: Databases were searched through May 2023 and updated in December 2025. Eligible trials evaluated XR interventions to treat KOA pain or after TKA. We classified interventions by depth of immersion and clinical mechanism. We used the Grading of Recommendations Assessment, Development, and Evaluation (GRADE) criteria to determine the certainty of evidence for prioritized outcomes. Meta-analyses were performed when ≥3 studies evaluated similar comparisons, outcomes, and time points. Results: Eligible trials addressed KOA (k=12) or recovery after TKA (k=9). Sample sizes ranged from 36 to 306 participants, and most studies had a follow-up of ≤3 months. Nineteen studies assessed pain-related functioning and pain intensity, and 5 assessed adverse events (AEs). For KOA, 10 studies examined interactive digital rehabilitation (IDR), and 2 examined virtual reality (VR)–digitally augmented exercise (DAE). IDR for KOA may result in better pain-related functioning (low certainty of evidence [COE]; pooled standardized mean difference [SMD] −0.59, 95% CI −1.11 to −0.06; prediction interval [PI] −1.72 to 0.55; k=5) and lower pain intensity at 6‐8 weeks (low COE; pooled SMD −0.46, 95% CI −0.92 to 0.00; PI −1.39 to 0.47; k=4). VR-DAE for KOA (k=2) produced inconsistent results (very low COE). For post-TKA studies, 5 examined IDR, 2 examined VR-DAE, 1 examined VR-distraction, and 1 examined VR-psychoeducation. Post-TKA IDR may result in better pain-related functioning (low [k=4] and moderate COE [k=1]) but little to no difference in pain intensity (low-moderate COE; pooled SMD at 3‐4 months −0.12, 95% CI −0.75 to 0.52; PI –1.63 to 1.27; k=3). VR-psychoeducation probably results in lower pain at 4 weeks (moderate COE; k=1), and VR-distraction may result in 6 months (low COE; k=1), whereas VR-DAE produced mixed findings (k=2; very low COE). IDR was not associated with AEs, and VR may not be associated with AEs for KOA (high and low COE), though AE reporting was uncommon (k=5) and evidence was very uncertain for post-TKA. Conclusions: IDR may augment treatment for KOA and post-TKA recovery, and VR may benefit post-TKA rehabilitation. This review is the first to stratify by level of immersion, clinical mechanism, and follow-up duration and to systematically evaluate AEs. IDR may be ready for integration into KOA care, while use after TKA needs more evidence. Randomized controlled trials with implementation outcomes could determine how XR interventions can be used for KOA, whereas trials evaluating efficacy and AEs are needed before their use for post-TKA. Trial Registration: PROSPERO CRD42023439903; https://www.crd.york.ac.uk/PROSPERO/view/CRD42023439903

Closing the Solidarity Gap Requires Closing the Accountability Gap for Patient-Facing AI

This commentary extends recent discussion of the solidarity gap associated with patient-facing AI by examining gaps in governance and risk allocation, health care professionals’ responsibilities in practice, and opportunities for professional stewardship and advocacy. We argue that equitable implementation requires shared accountability and meaningful health care professional participation in the design, evaluation, reimbursement, governance, and oversight of patient-facing AI before ambiguity results in patient harm.
Received — 27 May 2026 ⏭ Journal of Medical Internet Research

Integration of Digital Therapeutics Into Occupational Rehabilitation in Germany: Multilevel Simulation Study

Background: Expenditures for physiotherapy and extended outpatient physiotherapy (EAP) are increasing within Germany’s statutory accident insurance system (Berufsgenossenschaften), placing growing pressure on rehabilitation capacity and timely access to care. Digital health applications (DiGAs) are reimbursable nationwide and represent a novel component of routine rehabilitation pathways. However, their real-world system-level and economic effects in occupational rehabilitation remain insufficiently understood. Objective: This study aimed to evaluate how the integration of DiGAs into occupational rehabilitation pathways may influence costs, service capacity, and waiting times within routine care delivered by 5 German statutory accident insurance funds that cover 25.9 million insured individuals. Methods: Aggregated administrative data from 5 Berufsgenossenschaften (fiscal years 2023‐2024) were analyzed using a multilevel simulation framework combining (1) probabilistic cost-consequence modeling with Monte Carlo simulation (10,000 iterations), (2) an adherence-based adoption funnel distinguishing long-term engaged users (15%) and short-term users (85%) based on German claims data, and (3) a calibrated M/M/1 queuing model validated through discrete event simulation to estimate the effects on waiting times and system capacity. Primary outcomes included net financial impact, break-even thresholds, and changes in access-related performance metrics. Results: Combined physiotherapy and EAP expenditures reached €404 million (€1=US $1.18) in 2024, increasing by 10.1% year-over-year. The primary simulation (N=10,000 iterations) indicated mean annual net savings of €18.4 million (median €17.9 million) with a 90.7% probability of cost savings (95% uncertainty range: net cost of €8 million to net savings of €47.7 million). After incorporating adherence dynamics, the projected mean net savings were €16.2 million (95% CI €5-€29.8 million), corresponding to a 100% probability of positive financial impact within the modeled parameter space. Cost neutrality was maintained for DiGA prices up to €617.8 per prescription, nearly 40% above the base-case assumption of €450, indicating substantial economic robustness. Queuing analyses demonstrated that modest reductions in therapeutic demand decreased mean waiting times from 17.3 to 12.8 days (−26%), equivalent to approximately 120,000 cumulative patient waiting days saved annually across 26,705 EAP patients. The validation of discrete event simulation confirmed the magnitude and direction of analytic estimates. Conclusions: Under conservative assumptions, integrating digital therapeutics into occupational rehabilitation pathways is likely to generate both economic benefits and substantial system-level capacity gains. The break-even threshold of €617.80 per prescription provides a wide margin for pricing policy. Beyond cost effects, DiGAs may function as scalable capacity tools that alleviate systemic bottlenecks and improve timely access to rehabilitation services in capacity-constrained systems.

Self-Reported Health Outcomes in Metabolic Health YouTube Comments: Cross-Sectional Study and Rule-Based Natural Language Processing Framework Development and Validation

Background: YouTube is increasingly used for healthcasting, the sharing of evidence-based dietary and lifestyle interventions by domain experts. In the metabolic health domain, channels focused on therapeutic carbohydrate restriction have accumulated audiences of millions. A distinctive feature is the comment section, where viewers share first-person accounts of health changes, constituting a unique source of real-world outcome data at scale. However, extracting structured health information from unstructured comments presents computational challenges. Objective: This observational, cross-sectional study aims to develop and validate a precision-optimized computational framework for extracting self-reported health outcomes from healthcasting YouTube comments and to characterize the prevalence, distribution across health aspects, and channel-level variation of reported outcomes across a large-scale metabolic health corpus. Methods: This study analyzed 43,111 unique YouTube comments from 110 videos across 11 therapeutic carbohydrate restriction-focused healthcasting channels (37,458 unique authors; data span November 2013 to January 2026; collected via YouTube data application programming interface version 3). The methodology comprised 3 construction phases and 5 validation studies. The construction phases were (1) exploratory corpus characterization, (2) iterative development of a 35-aspect hierarchical health outcome ontology, and (3) precision-optimized rule-based classification, validated through precision validation (stratified sample of n=500), recall estimation (n=510), external validation on 5 held-out channels (n=12,653 comments), large language model–assisted interrater reliability assessment, and transformer baseline comparison against Bidirectional Encoder Representations from Transformers (BERT) and Robustly Optimized BERT Pretraining Approach (ROBERTa) classifiers. A supplementary aspect–based sentiment analysis contextualized the positive-only design. Results: The framework identified 1790 positive health outcome reports (1790/43,111, 4.15% prevalence), achieving 97.6% (488/500) precision (95% CI 95.7%-98.6%) and estimated 56.2% recall (95% CI 43.4%-67.9%). The reports described 6674 positive outcomes, distributed across 35 health aspects and 18 named disease conditions extending beyond weight loss: pain and inflammation reduction (1137/6674, 17%), type 2 diabetes improvement (977/6674, 14.6%), skin health (784/6674, 11.8%), and psychological well-being (731/6674, 11%). Over half (3355/6674, 50.3%) spanned multiple research objectives. Significant channel-level variation was observed (χ²10=927.5; P<.001), with positive outcome rates ranging from 1.32% to 10.40% (odds ratio 8.68, 95% CI 7.10-10.61). Transformer baselines achieved higher recall but lower precision, confirming their advantage for high-confidence corpus generation. A supplementary aspect-based sentiment analysis indicated a positive-to-negative ratio of approximately 4.6:1 (n=1003), with negative experiences (59/495, 11.9%) predominantly involving gastrointestinal and cardiovascular concerns. Conclusions: This study presents, to our knowledge, the first validated, rule-based framework for extracting self-reported metabolic health outcomes from healthcasting YouTube comments at corpus scale. Unlike existing recall-oriented social media health classifiers, the precision-optimized design achieves the confidence threshold required for outcomes research without manual review. These findings demonstrate that expert-led health content comment sections constitute a scalable, complementary data source for monitoring real-world engagement with dietary interventions, with implications for public health surveillance, platform design, and health communication research.

Blockchain-Enabled Self-Sovereign Identity Applications in Health Care: Scoping Review

Background: Self-sovereign identity (SSI) provides a decentralized approach to digital identity management, enabling individuals to control their personal data without reliance on centralized authorities. Blockchain technology offers a tamper-resistant and distributed infrastructure that can support secure and verifiable identity systems. In health care, where identity fragmentation, privacy risks, and interoperability challenges persist, blockchain-enabled SSI (BC-SSI) has been proposed as a potential solution. However, existing research remains heterogeneous, with varying levels of technical maturity and limited evidence of real-world deployment. Objective: This study conducts a scoping review to systematically map BC-SSI applications in health care and to analyze their application domains, development stages, study aims, targeted challenges, and technological infrastructures. In addition, this study aims to identify structural gaps in current research and assess the readiness of BC-SSI systems for clinical deployment. Methods: This review followed the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) methodology. A comprehensive literature search conducted between September 2024 and August 2025 identified 37 peer-reviewed studies that met predefined inclusion criteria. Data were extracted and synthesized using descriptive and thematic analyses across application areas, system maturity, technological components, and reported challenges. Results: The findings indicate that BC-SSI research in health care remains at an early stage of maturity, with most studies proposing conceptual models or prototype implementations and limited real-world validation. Applications predominantly focus on identity verification, credential management, and privacy-preserving data exchange across domains such as electronic health records, mobile health, and access control systems. Commonly used technologies include decentralized identifiers, verifiable credentials, smart contracts, and privacy-enhancing mechanisms such as zero-knowledge proofs and selective disclosure. Despite rapid technical development, persistent challenges include interoperability limitations, governance gaps, usability concerns, and insufficient integration with health care infrastructures. Notably, a structural gap was identified between technological capability and system-level readiness for clinical deployment. Conclusions: BC-SSI technologies demonstrate potential for enabling secure, interoperable, and patient-centric identity management in health care. However, current research is predominantly technology-driven and lacks sufficient system-level validation. This study highlights the need for integrated architectural approaches, governance frameworks, and real-world evaluation to bridge the gap between conceptual innovation and clinical implementation. Advancing BC-SSI toward health care adoption will require coordinated progress across technical, organizational, and regulatory dimensions.

Large Language Model–Generated Patient Instructions for Prescriptions in Primary Health Care: Preclinical Algorithm Validation

Background: The application of generative artificial intelligence to simplify medication use instructions has the potential to enhance people’s health by improving treatment adherence. Objective: We evaluated the performance of large language models (LLMs) in generating medication usage instructions to complement prescriptions in primary health care. Methods: This randomized, blinded experimental preclinical study used prescription-inducing scenarios, assigned to 62 health care professionals, to validate instructions generated by LLMs during electronic prescriptions. The instructions were generated by ChatGPT-4.0 (OpenAI), Llama3.1-8B (Meta), and Llama3.1-8B-RAG (Meta) using retrieval-augmented generation based on patient information leaflets. Performance metrics assessed adequacy, completeness, clarity, language simplification, usefulness, and errors in the generated instructions, with scores to analyze overall and individual metrics. Results: The 3 models yielded high overall scores for producing qualified instructions (ChatGPT-4.0: median 88.4, IQR 22.8; Llama3.1-8B: median 66.5, IQR 50.9; Llama3.1-8B-RAG: median 79.9, IQR 34.4; Kruskal-Wallis test P=.003). Llama3.1-8B-RAG received evaluations with similar overall scores to ChatGPT-4.0 (post hoc test, P=.05) and similar to Llama3.1-8B (post hoc test, P=.44). ChatGPT-4.0 outperformed Llama3.1-8B (Bonferroni test, P<.001). Regarding specific domains, Llama3.1-8B-RAG received scores equivalent to those of ChatGPT-4.0 for adequacy (mean 6.24, SD 2.3 vs mean 6.82, SD 2.1; post hoc test, P=.54); completeness (mean 5.94, SD 2.2 vs 6.55, SD 1.9; post hoc test P=.38), clarity (mean 5.77, SD 2.4 vs mean 6.68, SD 1.9; post hoc test P=.09), and usefulness (mean 5.42, SD 2.4 vs mean 5.96, SD 2.2; post hoc test P=.63). ChatGPT-4.0 received higher scores in the language simplification criterion than Llama3.1-8B-RAG (mean 7.05, SD 1.5 vs mean 5.44, SD 2.6; post hoc test P<.001). Interrater variability in assigning scores ranged from 4.2% (n=3) to 85.8% (n=6) among primary health care professionals. Instructions leading to incorrect use of the medication had similar frequency among the models(ChatGPT-4.0: n=15, 22.7%; Llama3.1-8B: n=19, 22.8%; Llama3.1-8B-RAG: n=19, 22.8%; chi-square test P=.71). The frequencies of hallucination were similar (ChatGPT-4.0: n=7, 10.6%; Llama3.1-8B: n=9, 13.6%; Llama3.1-8B-RAG: n=6, 9.1%; chi-square test P=.67). Conclusions: The open-source LLM enhanced with external information presented similar performance to the closed-source model, except for ChatGPT4.0, which was superior in language simplification of messages. LLM generation demonstrated potential for instructing patients on medication use. Nonetheless, the introduction of this innovation into the electronic prescribing workflow demands prescriber validation for human oversight of the technology and requires a strategy for LLM performance governance.

Safety of Telemedicine Versus In-Person Care for Patients With Tracheal Devices: Propensity Score–Matched Cohort Study

Background: Patients with tracheal diseases often require long-term follow-up after tracheal device placement, with a risk of adverse events that may lead to emergency care and unplanned interventions. Telemedicine has been proposed as an alternative to in-person follow-up to improve access and continuity of care. Objective: The primary objective of this study was to compare the need for emergency department (ED) visits between telemedicine and in-person groups. Secondary objectives included comparing hospital readmissions, 30-day hospital readmissions, and unplanned interventions between groups. Methods: This retrospective, single-institution study included adult patients with tracheal devices who underwent telemedicine and in-person outpatient clinic visits between 2020 and 2024. To balance the groups, we used 1:1 propensity score matching. We collected demographic and clinical data and evaluated the need for ED visits, hospital readmissions, 30-day hospital readmissions, and unplanned interventions. Kaplan-Meier estimation of time to first ED visit was performed to assess outcomes after outpatient visits. Results: A total of 483 patients (n=277, 57% telemedicine and n=206, 43% in-person) underwent 2487 visits (1258 telemedicine and 1229 in-person). After propensity score matching, 336 patients remained (168 in each group). There were no significant differences in the need for ED visits, hospital readmissions, or unplanned interventions. The telemedicine group had significantly fewer 30-day hospital readmissions (odds ratio 0.38, 95% CI 0.16-0.87; =.02). Kaplan-Meier analysis indicated no statistically significant difference in ED-free visits. Conclusions: Telemedicine follow-up was associated with outcomes comparable to those of in-person follow-up in this cohort of adult patients with tracheal devices, with no evidence of an increased need for ED visits. In the matched analysis, telemedicine was associated with lower odds of 30-day hospital readmission.

Virtual Reality Interventions for Stress Reduction in the General Population: Systematic Review and Meta-Analysis of Randomized Controlled Trials

Background: Increasing mental demands across multiple life domains underscore the importance of effective individual stress management to mitigate the adverse health consequences of chronic stress. Growing evidence suggests that virtual reality (VR) interventions constitute an effective approach to stress reduction. Objective: This systematic review and meta-analysis aimed to examine and compare application areas of VR interventions for stress reduction in the general population and to identify potential predictors of effectiveness based on sample characteristics and intervention design. Methods: Five databases (MEDLINE, CINAHL, CENTRAL, PsycInfo, and Web of Science) were systematically searched for randomized controlled trials investigating the effectiveness of VR interventions for stress reduction in the general population. Studies were included if they primarily focused on stress reduction, included a neutral control condition, and reported a validated measurement of perceived stress. Trials targeting mental disorders or those conducted in the context of medical procedures were excluded. Two reviewers independently screened the literature, extracted data, and assessed the risk of bias using the Cochrane Collaboration’s tool. Effects were synthesized using pooled standardized mean differences, and relevant predictors were evaluated through subgroup analyses and meta-regressions. Results: A total of 55 relevant studies met the inclusion criteria, with 37 investigating single-session and 18 multisession interventions (ranging from 1 to 42 sessions over 2 days to 6 months). The meta-analysis included 39 studies with 4024 participants (sample sizes 24‐409; mean ages 19.2‐70.6 years). Intervention types included VR-based nature exposure (21), biophilic architectural elements (6), guided meditation (9), interactive tasks (4), and other approaches (1). On average, VR interventions significantly reduced perceived stress level (−0.55, 95% CI −0.70 to −0.40;

Patients’ Perspectives on the Implementation of AI in Radiological Diagnostics: Focus Group Study

Background: Rapid developments in artificial intelligence (AI) will enable its widespread use in radiological diagnostics in the near future. Patients will then be confronted with findings generated with the help of AI. Understanding patients’ perspectives on the use of this technology is one of the key factors for its successful implementation. Objective: This qualitative study aimed to gain insight into patients’ reasoning about the opportunities and risks of using AI in radiological diagnosis, to identify prerequisites for its acceptance, and to identify aspects that can promote trust in AI diagnoses, especially in scenarios of high personal concern. Methods: A total of 7 focus groups were conducted with 34 patients (n=15, 44% female participants) aged between 23 and 85 years (mean 49.06, SD 17.08 y), recruited using purposive sampling strategies. Each focus group was audiotaped, transcribed, and analyzed using the method of structured qualitative content analysis. Results: Study findings show that patients are open to the use of AI in radiological diagnostics. The basic prerequisites for this are (1) scientific evidence of safe outcomes that are more accurate and faster than those without AI; (2) recognizable added value in patient care; (3) transparency in the use of AI and disclosure to the patient; (4) comprehensive, binding measures for quality assurance; and (5) the use of AI solely to support the physician. However, the results indicate that further criteria are important for patients to be willing to choose a radiologist who uses AI and to trust AI diagnoses. In situations where they are personally affected, patients fear that physicians will place too much trust in the AI result and that the physician-patient relationship will become dehumanized. Therefore, the physicians’ abilities and functions that inspire trust from the patients’ perspective must come into play. These include (1) an independent diagnosis by the physician that takes into account not only the clinical context but also the individuality of the patient, (2) a comprehensible explanation of the pros and cons of using AI for patients and clear communication of AI output, and (3) a humane and empathetic physician-patient relationship, which shows that the physician continues to feel responsible for the patient. Conclusions: The results of the study underscore that a high quality of the entire “sociotechnical” system is an essential prerequisite for patient acceptance of the use of AI in radiological diagnostics and for trust in AI diagnoses. The further development of AI performance must go hand in hand with the creation of framework conditions for its use that meet patients’ expectations of the role of the physician and ensure a trust-building physician-patient relationship. The study provides valuable insights into how such integration of AI into radiological practice can be achieved.
  • ✇Journal of Medical Internet Research
  • From Design to Accountable Impact for Data Dashboards in Health Care Barbara-Jo Achuff
    Vornhagen and colleagues synthesize design practices for the development of health care dashboards and provide a timely reference for a rapidly expanding class of tools that increasingly mediate clinical decisions and quality improvement. The authors have defined 4 pillars of design with associated practices, establishing a practical approach to the thoughtful development of useful dashboards. Building on these pillars, this commentary proposes accountable dashboarding, defined as making explici
     

From Design to Accountable Impact for Data Dashboards in Health Care

Vornhagen and colleagues synthesize design practices for the development of health care dashboards and provide a timely reference for a rapidly expanding class of tools that increasingly mediate clinical decisions and quality improvement. The authors have defined 4 pillars of design with associated practices, establishing a practical approach to the thoughtful development of useful dashboards. Building on these pillars, this commentary proposes accountable dashboarding, defined as making explicit the causal chain linking data sources and governance to visualization, interpretation, action, and measurable outcomes. Dashboards could be treated as sociotechnical interventions in which upstream choices are inseparable from the user interface and downstream decision-making. Evaluation should extend beyond usability and satisfaction to include decision quality and behavioral proxies, unintended consequences, and patient-centered outcomes of dashboard-informed interventions.

Association Between Wearable Device Adoption and Health-Related Lifestyle Behaviors: Retrospective Cohort Study

Background: Wearable devices are increasingly adopted for personal health monitoring, but evidence on their long-term associations with health-related lifestyle behaviors in real-world population settings remains limited. Objective: This study examined the longitudinal association between wearable device adoption and engagement in health-related lifestyle behaviors using a nationally representative panel dataset from South Korea. Methods: We analyzed data from the 2016 and 2022 waves of the Korea Media Panel survey. Health-related lifestyle behaviors in the physical, social, and cultural domains were operationalized as estimated annual activity counts based on self-reported frequency measures. We used a difference-in-differences framework with generalized estimating equations to compare changes in these behaviors between new wearable adopters and nonadopters adjusting for demographic and socioeconomic characteristics. Relative changes were estimated using Poisson models with a log link, and subgroup analyses were conducted to explore variation across sociodemographic groups. As a sensitivity analysis, inverse probability of treatment weighting was additionally applied to assess the robustness of the findings to observed baseline imbalance. Results: Wearable device adoption was associated with greater increases in total, physical, and cultural health-related lifestyle activities over time. In the difference-in-differences model, adopters showed greater relative increases in total activity (rate ratio [RR] 1.24, 95% CI 1.08-1.35), physical activity (RR 1.36, 95% CI 1.12-1.64), and cultural activity (RR 1.78, 95% CI 1.31-2.42) than nonadopters. Subgroup analyses showed limited evidence of consistent heterogeneity and should be interpreted cautiously. Sensitivity analyses using inverse probability of treatment weighting showed overall patterns broadly similar to those of the primary analyses. Conclusions: In this nationally representative panel study, wearable device adoption was associated with greater increases in total, physical, and cultural health-related lifestyle activities over time, whereas no clear association was observed for social activity. These findings should be interpreted as associative rather than causal given the observational design and the inability to directly assess parallel trends.

Sequencing AI Automation and Data Interoperability in Oncology Using a Scenario-Planning Framework Coupled With Discrete-Event Simulation: Proof-of-Concept Study

Background: As oncology workflows integrate increasingly autonomous artificial intelligence (AI) agents, health systems face uncertainty regarding operational impacts. Traditional linear forecasting methods fail to capture second-order effects such as governance saturation, induced demand, and bottleneck migration. To navigate this complexity, the emerging field of medical futures studies requires methodologies that bridge qualitative strategic foresight with quantitative operational modeling. These system-level dynamics directly influence timely diagnosis, treatment delays, and overall health system resilience. Objective: This study aimed to develop a proof-of-concept framework coupling qualitative scenario planning with computational discrete-event simulation to stress-test oncology AI adoption strategies. Methods: We defined a strategic state space using 2 orthogonal axes, AI automation intensity and data interoperability, resulting in 4 distinct futures scenarios. We translated these qualitative narratives into a quantitative discrete-event simulation model of a 3-year operational horizon. The model quantified system performance (referral-to-treatment interval [RTTI] and throughput), volatility, and resource constraints across different adoption trajectories. Results: The scenario-planning phase yielded 4 operational archetypes (analog oncology, automation islands, interconnected clinicians, and AI-orchestrated care) with distinct constraints, risks, and failure modes. In the simulation, the fully integrated scenario maximized capacity (1244, SD 21.4 patients per year) and halved the mean RTTI to 14.9 (SD 0.3) days, a magnitude comparable to major pathway redesign interventions. Isolated automation without data infrastructure led to reduced system performance, increasing RTTI by 26% (37.1, SD 1.3 days) and reducing throughput to 647 (SD 10.1) patients per year due to administrative governance saturation. The model illustrated a structural bottleneck migration: successful upstream AI adoption shifted binding constraints from diagnostic scanners to downstream chemotherapy infusion units, whereas missing data interoperability resulted in governance constraints. Pathway optimization analysis indicated that a coordinated strategy prioritizing early improvements in data interoperability reduced transition volatility compared to an automation-first approach. Conclusions: Integrating qualitative scenario planning with quantitative simulations enabled a systematic evaluation of oncology AI adoption strategies. As a proof of concept, it offers a replicable framework for health leaders to model future scenarios of digital transformation in times of high uncertainty. Subsequent work should expand this methodology to incorporate financial and health equity dimensions, establishing simulation-based scenario planning as an important tool in medical futures studies.
❌