❌

Normal view

Received — 15 September 2026 ⏭ Journal of Medical Internet Research

Depictions of Depression in Generative AI Video Models: Mixed Methods Study of OpenAI’s Sora 2

Background: Generative AI video models are increasingly capable of producing complex depictions of mental health experiences, yet little is known about how these systems represent conditions such as depression. Because AI-generated content may reach people during vulnerable periods, understanding what visual narratives these models produce for sensitive concepts carries clinical relevance. Objective: This study aimed to characterize how OpenAI’s Sora 2 generative AI video model depicts depression and examine whether depictions differ between the consumer app and developer API access points, which differ in their product layer mediation. Methods: We generated 100 videos using the single-word prompt “Depression” across 2 access points: the consumer app (n=50, 50%) and developer API (n=50, 50%). Two trained coders independently coded narrative structure, visual environments, objects, figure demographics, and figure states. Interrater reliability was assessed using the Cohen κ, with dimensions showing insufficient agreement excluded from analysis. Computational features (visual aesthetics, audio, semantic content, and temporal dynamics) were extracted and compared between modalities using 2-tailed Welch tests with Benjamini-Hochberg false discovery rate correction. Results: App-generated videos exhibited a pronounced recovery bias: 78% (39/50) featured narrative arcs progressing from depressive states toward resolution compared with 14% (7/50) of API outputs. This divergence was reinforced across channels. App videos brightened over time (mean slope 2.90, SD 2.43 per second vs −0.18, SD 1.24 per second for the API; Cohen =1.59;

Design Guidelines for Online Health Forums: User-Centered Design Approach

Background: Online health forums are used widely, yet evidence of their effectiveness is inconsistent. Evidence-based forum design guidance grounded in theory and lived experience could improve the efficacy and outcomes of these forums for the many people using them worldwide. Objective: This study aimed to draw on the experience of online forum users and staff, and insights from existing research on technology design and self-determination theory, to generate a set of theoretically grounded guidelines for safe and well-being–supportive online forums. Methods: We conducted 54 semistructured interviews (36 forum users and 18 forum staff) and 4 design workshops with forum staff, combined with input from a multidisciplinary research team. Principles of qualitative framework analysis were used to adapt a preexisting framework for well-being–supportive technology to the context of online health forums. Results: The resulting design guidelines are framed around 4 overarching principles relating to the psychological needs for autonomy, competence, and relatedness as defined by self-determination theory, and the additional need for safety in online forums. Each principle is presented alongside pragmatic design heuristics and specific implementation strategies. Conclusions: User experiences of online health forums are mixed. We have drawn on self-determination theory to propose evidence-informed guidelines for service development and refinement, adaptable to specific user groups across diverse settings. International Registered Report Identifier (IRRID): RR2-https://doi.org/10.1136/bmjopen-2023-075142

Anonymization of Portuguese Clinical Notes Using Large Language Models and Quantum-Enhanced Hybrid Architectures: Comparative Evaluation Study

Background: The widespread adoption of electronic health records (EHRs) has generated large-scale repositories of highly sensitive clinical information, emphasizing the need for robust anonymization strategies to enable secondary use for research while safeguarding patient privacy. Conventional rule-based and machine learning approaches for deidentifying medical text face limitations with the linguistic complexity, variability, and context dependence inherent to clinical documentation. Recent advances in large language models (LLMs), combined with emerging quantum computing paradigms, present novel opportunities to enhance the accuracy, scalability, and resilience of health care data anonymization. Objective: This study aims to evaluate the efficacy of LLM-based and quantum-enhanced hybrid architectures for medical text anonymization, assessing the effectiveness and computational efficiency across multiple entity types in Portuguese clinical notes. Methods: We constructed a gold-standard corpus of 1000 Portuguese outpatient clinical notes, manually annotated by 5 trained researchers for 5 protected-entity categories: patient names, dates, identifiers, organizations, and geographic locations. Four anonymization strategies were evaluated: 2 stand-alone LLMs (Llama-3.1-8B-instruct and Llama-3.3-70B-instruct) and 2 quantum-enhanced hybrid models (Dynex-QML with 8B and 70B base models) incorporating quantum optimization via Quadratic Unconstrained Binary Optimization (QUBO) formulations. The quantum-enhanced approach transforms the final attention layer of the LLM into a global constraint satisfaction problem solved via neuromorphic quantum annealing. Model performance was measured on a held-out test set of 500 notes using precision, recall, and -score metrics. Computational efficiency was quantified through end-to-end processing time. Results: The quantum-enhanced Dynex-QML-70B model achieved the highest overall performance with a macro-score of 0.855 (95% CI 0.823‐0.880), outperforming the stand-alone Llama-3.3-70B (0.726, 95% CI 0.704‐0.747), Dynex-QML-8B (0.733, 95% CI 0.709‐0.756), and Llama-3.1-8B (0.602, 95% CI 0.588‐0.615). Compared with Llama 3.3 70B, Dynex-QML (Llama 70B) improved macro-score by 0.128 (95% CI 0.091‐0.163; empirical 2-sided bootstrap

Digital Phenotyping of Lifestyle Profiles and Mental Well-Being in German Adults: Prospective Longitudinal Cohort Study

Background: Digital phenotyping uses passively collected smartphone-sensing data to characterize everyday behavior in naturalistic settings, and has become an important approach for studying mental well-being. Most previous studies have examined associations between individual sensing variables and mental health. However, mental well-being is likely reflected not by isolated behaviors but by combinations of co-occurring daily behaviors that together form lifestyles. Person-centered approaches capable of identifying these behavioral configurations may, therefore, provide more interpretable digital phenotypes; yet, such approaches have rarely been applied to passive smartphone-sensing data. Objective: This study aimed to examine whether smartphone-captured behavioral and environmental data could be used to derive interpretable day-level and person-level lifestyle profiles, and whether person-level profiles were associated with mental well-being. We also tested whether Big Five personality traits—extraversion, agreeableness, conscientiousness, openness, and negative emotionality—moderated these associations. Methods: The study used a 2-week prospective longitudinal cohort design with a sample of 553 German adults (mean age 42.12, SD 12.89 years; 44.65% female) drawn from an initial sample recruited according to quotas designed to reflect the German population. Ten smartphone-sensing indicators captured 5 domains, including communication and social media app use, mobility, physical activity, environmental context, and phone-use intensity. Mental well-being was assessed using the Warwick–Edinburgh Mental Well-Being Scale, and personality was assessed using the 15-item Big Five Inventory–2 Extra-Short Form. We used multilevel latent profile analysis to identify day-level profiles nested within person-level profiles. Associations between profiles and mental well-being were tested using classification-error–adjusted mean comparisons and omnibus Wald tests. Moderation was examined using hierarchical regressions comparing models with and without profile-by-personality interactions. Results: Eight day-level profiles and 7 person-level profiles were identified. Day-level profiles reflected distinct combinations of smartphone-sensing indicators. Person-level profiles represented different distributions of these daily patterns. Profiles differed significantly only in positive functioning (Wald ²=13.39; =.04), not in overall mental well-being, positive affect, or satisfying interpersonal relationships. The physically active and unplugged profile had higher positive functioning than the mobile and always-on social profile (mean 3.94, SD 0.63 vs mean 3.61, SD 0.74; Cohen =0.47; 95% CI 0.21‐0.73). No other pairwise differences were significant. Sensitivity analyses excluding the smallest profile produced comparable results, supporting the robustness of the findings. Personality-by-profile interactions did not significantly improve prediction for any well-being outcome. Conclusions: The findings extend the field by showing that transparent, person-centered digital phenotypes can distinguish variation in positive functioning, although causal conclusions cannot be drawn. In real-world settings, such interpretable profiles could support understandable monitoring tools and, following prospective replication and validation, inform personalized multibehavior interventions that target combinations of behaviors rather than single behaviors in isolation.

Evaluation of the Square Eyes Model as a Screening Tool for Identifying Digital Technologies in Wearable Camera Images Among Children: Laboratory Study

Background: Accurate measurements of children’s digital technology use are essential for understanding its potential implications on health and well-being. Wearable cameras can provide such measurements, but image coding is a high burden for researchers. Machine learning–based object-recognition models have the potential to reduce this burden by identifying images containing technology. Objective: This study aims to evaluate the performance of an object recognition model, the Square Eyes model, as a screening tool for identifying technologies in wearable camera images among children for further human review, as well as to examine the potential influence of face-blurring methods on the model’s performance. Methods: This study used data collected on 48 children (aged 3‐14 y) during an approximately 1-hour laboratory session. The children performed various technology-related tasks while wearing a camera. A total of 221,226 images were coded by humans and processed through the Square Eyes model. The performance of the Square Eyes model as a screening tool was evaluated by (1) assessing agreement between the model and human coding; (2) evaluating the N-back algorithm, an algorithm embedded in the model aimed to flag images requiring human review; and (3) examining the potential influence of facial-blurring on model performance. Results: Humans detected technology in 92,745 (41.9%) images, and the Square Eyes model detected technologies with an overall accuracy of 78.0%. When considering specific technologies, agreement between the model and human coders was the highest for (n=19,148, 54.3%) and (n=8492, 44.5%) and lowest for smaller devices such as (n=2600, 31.3%) and (n=3685, 25.1%). The model’s N-back algorithm effectively flagged images that required further human review, with only 7144 (3.2%) images that were not flagged for screening containing a human-coded technology. An explorative analysis indicated that using a square face-blurring with border could have reduced the model’s ability to accurately detect technologies. Conclusions: The Square Eyes model demonstrated overall satisfying accuracy in detecting technologies and successfully flagged images that required further review by humans. These findings suggest that the model could be used as an effective screening tool for reducing the burden of human coding. However, the model could be improved to more accurately detect smaller devices, and the form of facial blurring in images should be considered.

Digitally Adapting LGBTQ-Affirmative Cognitive Behavioral Therapy for Chinese Men Who Have Sex With Men Living With HIV: User-Centered Design Approach

Background: Chinese men who have sex with men living with HIV (MSMLWH) experience substantial psychological distress driven by minority stress and HIV-related challenges. However, culturally tailored digital mental health interventions that address HIV-specific maladaptive cognitive schemas and culturally specific psychosocial stressors remain scarce in China. Objective: This study aimed to systematically adapt an evidence-based cognitive behavioral therapy (CBT) intervention Effective Skills to Empower Effective Men (ESTEEM) into a WeChat (Tencent) Mini-Program–based intervention (iESTEEM) specifically for Chinese MSMLWH and to evaluate its preliminary feasibility and usability. Methods: We used a three-phase user-centered design approach guided by the Assessment, Decision, Adaptation, Production, Topical Experts, Integration, Training, and Testing (ADAPT-ITT) framework. The study proceeded in three phases: (1) a qualitative needs assessment using semistructured interviews with 20 MSMLWH (mean age 23.25, SD 3.08 years); (2) systematic intervention adaptation and platform development, including theater testing (n=5); and (3) a 2-week pilot study involving 10 MSMLWH and five counselors to evaluate feasibility, usability, and acceptability through focus groups and objective platform analytics. Results: Phase 1 identified 3 major themes of psychological distress: persistent health anxiety fueled by catastrophizing, intersectional stigma internalization, the disclosure dilemma, and intimacy barriers rooted in defectiveness and shame schemas. Participants also prioritized anonymity and bite-sized learning. Guided by these findings, iESTEEM was developed as a counselor-assisted, privacy-preserving WeChat Mini-Program incorporating HIV-specific scenarios, multimodal learning modules, and a back-end risk-alert system. During the 2-week pilot, participants logged into the platform 14.1 (SD 6.7) times per person and completed 134.3 (SD 103.1) minutes of learning activities; all participants accessed module 1, and 90% (9/10) accessed modules 2‐5. Anxiety scores decreased from 8.9 (SD 2.3) to 7.2 (SD 3.0), whereas depression scores remained stable. All participants expressed a willingness to continue using the program and to recommend it to peers. Participants and counselors endorsed its contextual relevance, privacy protections, and clinical utility. Conclusions: This study provides a theory- and evidence-informed model for culturally adapting digital mental health interventions for highly stigmatized populations. By integrating lesbian, gay, bisexual, transgender, and queer (LGBTQ)-affirmative CBT principles, HIV-specific adaptations, and a privacy-preserving, counselor-assisted WeChat Mini-Program, iESTEEM demonstrated promising preliminary feasibility, acceptability, and engagement among Chinese MSMLWH. These findings support the potential of culturally tailored digital interventions to expand access to psychological support for this stigmatized population in resource-constrained settings. Ongoing randomized controlled trials will further evaluate its efficacy, implementation outcomes, and mechanism of action. Trial Registration: Chinese Clinical Trial Registry ChiCTR2400080263; https://www.chictr.org.cn/showproj.html?proj=216926
Received — 13 September 2026 ⏭ Journal of Medical Internet Research

Effectiveness of Wearable Digital Therapeutics in Improving Sleep Outcomes Among Individuals With Insomnia: Systematic Review and Meta-Analysis of Randomized Controlled Trials

Background: Wearable devices are increasingly used for sleep monitoring and as adjunctive treatment. Existing meta-analyses mostly pool composite digital therapies and rarely isolate stand-alone wearables or distinguish between objective and subjective end points. Whether stand-alone wearable interventions improve sleep outcomes in adults with insomnia, and which factors moderate treatment heterogeneity, remains unclear. Objective: This study aims to evaluate the effectiveness of wearable digital interventions on sleep outcomes in adults with insomnia versus control strategies and explore moderators of effectiveness, including device-wearing position, intervention duration, and control type, using meta-regression. Methods: This systematic review and meta-analysis was conducted in accordance with the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta‑Analyses) 2020 statement and the PRISMA-S (Preferred Reporting Items for Systematic Reviews and Meta‑Analyses Literature Search Extension) guideline. Five electronic databases and clinical trial registries were searched from inception to May 18, 2026. Eligible studies were randomized controlled trials (RCTs) evaluating wearable digital interventions in adults with insomnia compared with sham, waitlist, usual care, or active control conditions and had an intervention duration of at least 1 week. Study screening, data extraction, and risk-of-bias assessment were carried out independently by 2 reviewers. Pooled estimates were calculated using a restricted maximum likelihood random-effects model with the Hartung-Knapp-Sidik-Jonkman correction. Heterogeneity was assessed using the ² statistic, and 95% prediction intervals (PIs) were calculated for the primary analyses. The certainty of evidence was rated using the GRADE (Grading of Recommendations, Assessment, Development, and Evaluation) approach. Results: Sixteen RCTs (N=910) were included. Wearable digital interventions were associated with a significant reduction in objective sleep-onset latency (SOL; mean difference [MD] −4.52, 95% CI −8.38 to −0.67, PI −9.52 to 0.47 min) and a significant improvement in subjective sleep efficiency (SE; MD 2.00%, 95% CI 1.90%‐2.11%, PI 1.85%‐2.15%). Subjective total sleep time (TST) also showed a significant increase (MD 19.11, 95% CI 2.98‐35.24, PI −16.20 to 54.43 minutes). Meta-regression showed that control type, intervention duration, and device location did not explain the heterogeneity of the insomnia severity index (ISI) (=0). Sensitivity analysis confirmed the robustness of pooled ISI estimates, and an Egger test indicated no small-study effects (=.07). Certainty of evidence ranged from moderate to high. Conclusions: Wearable digital interventions provide selective benefits for objective SOL, subjective SE, and subjective TST in adults with insomnia, with no improvement in overall ISI. Despite statistically significant effects on several sleep parameters, wide PIs, substantial heterogeneity, and limited study numbers indicate preliminary, nonconclusive findings. Wearables should be viewed as affordable adjunctive tools requiring further validation, not substitutes for first-line cognitive behavioral therapy for insomnia. Large-scale, long-term RCTs with standardized protocols and patient-level external validation are required to consolidate the evidence base. Trial Registration: PROSPERO CRD420251038603; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251038603

Correction: From Metrics to Meaning in Neurological Rehabilitation: Clinicians’ Perspectives on Digital Metrics of Upper Limb Functioning—A Focus Group Study

Digital assessment technologies, such as optical motion capture and inertial measurement units, enable detailed kinematic analysis and continuous monitoring of upper limb activity in persons with neurological conditions. While such digital metrics of functioning are increasingly recognized in research, their uptake in clinical neurorehabilitation is limited. It remains unclear which digital metrics of functioning clinicians perceive as most meaningful and how these are integrated into patient-centered care. Understanding clinicians’ information needs and reasoning processes is a prerequisite for implementing digital assessment technology. To characterize how rehabilitation professionals perceive, prioritize, and integrate digital metrics of functioning into clinical reasoning and to identify features that would support their routine use. Three 90-minute focus groups were conducted in 3 Swiss neurorehabilitation centers, involving 11 clinicians with diverse professional backgrounds (5 physiotherapists, 4 occupational therapists, 1 movement scientist, and 1 medical practitioner). Participants discussed essential parameter domains and individually rated the relevance and meaningfulness of 17 kinematic metrics for the well-studied drinking task and 10 established arm use performance metrics. Verbatim transcripts were analyzed using reflexive thematic analysis, and rating data were summarized descriptively. Five main themes were identified. (1) Functional requirements to interpret movement quality and performance (active/passive range of motion (ROM), strength, selective muscle control, grasp) form the basis for interpreting movement. (2) Essential aspects of movement quality (smoothness, efficiency, compensatory movement) are valued when aligned with observable task execution. (3) Added value of real-world performance (hourly activity profiles, arm-use symmetry, functional workspace) represents the reference for patient-centered reasoning. (4) Individualizing what matters, including diagnosis-specific preferences, shapes assessment selection. (5) Blending clinical eye and reference data reflects clinicians’ reliance on visual judgment complemented by normative values. Intuitive metrics such as task duration, number of movement units, and ROM were favored, whereas confidence was lower in more complex metrics (e.g., jerk, inter-joint coordination). Clinicians value intuitive digital metrics of functioning when they are clearly linked to patient-centered outcomes and supported by normative references. The findings highlight the need for targeted educational strategies and digital competency training that help clinicians interpret digital metrics and integrate them with contextual information and clinical reasoning.

Evaluating Large Language Models in Clinical Audiology (AUDIOLOGYBENCH): Benchmark Development and Validation Study

Background: Large language models (LLMs) are increasingly being explored for clinical decision support, but their performance in audiology has not been systematically benchmarked using clinically grounded case materials and rubric-based safety evaluations. Objective: This study aimed to develop and evaluate AUDIOLOGYBENCH, a 3-tier benchmark for characterizing frontier LLM capability in clinical audiology along (1) curated domain knowledge, (2) literature-derived evidence, and (3) clinical reasoning under multimodal case input, with an explicit human audit of the automated adjudicator on the primary end point. Methods: The benchmark comprises 3139 objective items from educational resources, 3175 research article–derived items from peer-reviewed articles published between 2015 and 2025, and 67 multimodal clinical case studies graded against a standardized A-F rubric with 6 prespecified critical-error types that cap scores at D or F. Eight models were evaluated on the educational objective items: 4 frontier multimodal models (Gemini 2.5 Pro, Grok 4, OpenAI O3, and Claude Sonnet 4 Thinking) were evaluated on the research article–derived items, and on 804 case study evaluations. Adjudication used Gemini 2.5 Pro (objective and research-derived items) and Claude Opus 4.5 (case studies). The case study adjudicator was independently audited against PhD-level audiologist consensus on blinded subsamples, supplemented by a post-stratified human-calibrated sensitivity analysis. Results: A striking task-type dissociation emerged on case studies: clinical recommendations (Q3) achieved a mean score of 89.74 (SD 13.92, 95% CI 88.07‐91.41), a 98.1% (263/268) pass rate, and no dangerous recommendations; audiometric numerical interpretation (Q1) achieved a mean score of 67.89 (SD 18.47, 95% CI 65.68‐70.10), with a 35.4% (95/268) critical-error rate; and differential diagnosis (Q2) achieved a mean score of 67.79 (SD 15.33, 95% CI 65.95‐69.63). Question type, not model selection, dominated performance (eta-squared_H=0.333 vs 0.001; rank biserial ≥0.679). Interreviewer reliability between audiologists was high (quadratic-weighted κ of 0.78 and 0.85 across the 80-item and 50-item audits, respectively). When 2 audiologists regraded all 80 model Q1 responses with the diagnostic images available, the adjudicator’s per-item Q1 labels diverged from human judgment (κ=0.05; overflagging; sensitivity: 19/26, 73%; positive predictive value: 19/53, 36%), yet its reweighted Q1 critical-error rate (36.2%) was broadly consistent with the image-grounded human estimates (28%‐34%), suggesting no systematic inflation of the headline rate. The principal Q3>{Q1, Q2} ranking was preserved under post-stratified human calibration. Web-style multiple-choice items showed ceiling effects (>95% accuracy); short-answer prompts remained challenging (best 30%). Conclusions: Current frontier LLMs show strong recommendation generation but substantial limitations in audiometric numerical interpretation that are shared across models and that an automated adjudicator partially miscalibrated at the per-item level. AUDIOLOGYBENCH characterizes capability boundaries rather than certifying clinical readiness. Deployment of LLM-assisted audiology workflows requires structured human verification of all numerical findings and awareness of fabrication and severity misclassification failure modes documented here.

The Effectiveness of Digital Intervention on Psychological Resilience in Postoperative Breast Cancer Patients During Chemotherapy Intervals: Quasi-Experimental Study

Background: Patients with breast cancer during postoperative chemotherapy intervals commonly experience psychological distress and reduced resilience while recovering at home. Digital mindfulness interventions may provide accessible psychological support during this vulnerable period; however, evidence regarding tailored interventions for postoperative patients with breast cancer during chemotherapy intervals remains limited. Objective: This study aimed to examine the effectiveness of a digital intervention on psychological resilience in postoperative patients with breast cancer during chemotherapy intervals. Methods: A quasi-experimental study with repeated measures was conducted from October 2021 to June 2022. A total of 80 eligible participants were recruited from the Department of Breast Surgery at a tertiary hospital in Zhejiang Province, China, and 71 completed the study. The control group received routine discharge instructions and nursing follow-ups, whereas the intervention group additionally received an 8-week digital psychological resilience intervention. Outcomes were assessed at baseline (T0), 3 months post intervention (T1), and 6 months post intervention (T2). The measures included the Connor-Davidson Resilience Scale (CD-RISC), Hospital Anxiety and Depression Scale (HADS), Social Support Rating Scale (SSRS), Breast Cancer Survivor Self-Efficacy Scale (BCSSS), and Functional Assessment of Cancer Therapy-Breast (FACT-B). Independent-samples tests, chi-square tests, and repeated-measures ANOVA were performed using SPSS (version 26.0; IBM Corp). Results: No statistically significant baseline differences were observed between the two groups in the outcome measures. At T1, the intervention group had higher CD-RISC scores than the control group (mean 67.58, SD 11.41 vs mean 62.09, SD 10.18; =.036) and higher BCSSS scores (mean 42.36, SD 3.59 vs mean 39.23, SD 4.90; =.003). However, these between-group differences were no longer statistically significant at T2 (>.05). Significant time effects and group×time interaction effects were observed for both psychological resilience and self-efficacy (.05), although both scales showed significant time effects (

Performance of AI-Based Screening Tools for Obstructive Sleep Apnea Across Apnea-Hypopnea Index Thresholds: Systematic Review and Meta-Analysis

Background: Obstructive sleep apnea (OSA) is highly prevalent but remains substantially underdiagnosed. Polysomnography (PSG) is the reference standard, but its cost and limited availability constrain large-scale case identification. AI-based screening tools may support risk stratification and referral prioritization, but their diagnostic accuracy across apnea-hypopnea index (AHI) thresholds remains uncertain. Objective: This review aimed to systematically evaluate the diagnostic accuracy of AI-based OSA screening tools at AHI thresholds of ≥5, ≥15, and ≥30 events/hour, with emphasis on models using non-PSG–derived inputs. Methods: PubMed, Embase, Scopus, and Web of Science were searched for studies published from January 1, 2016, to May 3, 2026. Eligible studies included adults evaluated for suspected OSA or recruited from population-based cohorts, assessed AI-based models intended or interpretable for OSA screening, risk prediction, or screening-oriented severity classification, used PSG as the reference standard, and reported sufficient data to construct or reconstruct 2×2 contingency tables. Diagnostic accuracy was synthesized separately by AHI threshold and input source using bivariate random-effects models, with 95% CIs and prediction intervals (PIs). Risk of bias and certainty of evidence were assessed using QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies 2) and GRADE (Grading of Recommendations Assessment, Development, and Evaluation), respectively. Results: A total of 60 studies were included, of which 47 contributed data to the meta-analysis. At AHI thresholds of ≥5, ≥15, and ≥30 events/hour, pooled sensitivities were 0.94 (95% CI 0.92‐0.96; 95% PI 0.71‐0.99), 0.87 (95% CI 0.84‐0.89; 95% PI 0.66‐0.96), and 0.83 (95% CI 0.79‐0.87; 95% PI 0.61‐0.94), respectively; the corresponding specificities were 0.77 (95% CI 0.69‐0.84; 95% PI 0.30‐0.96), 0.81 (95% CI 0.75‐0.85; 95% PI 0.39‐0.96), and 0.91 (95% CI 0.87‐0.94; 95% PI 0.55‐0.99), respectively. The corresponding areas under the summary receiver operating characteristic curves were 0.943, 0.907, and 0.920. For non-PSG–derived tools, sensitivities were 0.92, 0.85, and 0.81, and specificities were 0.70, 0.74, and 0.85 at the 3 thresholds, respectively. For PSG-derived models, sensitivities were 0.96, 0.90, and 0.85, and specificities were 0.82, 0.88, and 0.96, respectively. Exploratory subgroup analyses suggested performance variation across selected study and model characteristics, including region, algorithmic framework, data source, and validation method. Conclusions: AI-based tools showed generally favorable screening performance for OSA across clinically relevant AHI thresholds, although wide PIs suggest variable performance across future comparable populations and settings. By synthesizing diagnostic accuracy across 3 AHI thresholds and distinguishing non-PSG–derived from PSG-derived models, this review extends previous broad or modality-specific reviews and offers a clinically interpretable, pathway-specific basis for linking model performance to intended use. The findings may clarify potential roles for non-PSG–derived tools in front-end screening and referral prioritization and for PSG-derived models in reduced-channel assessment and sleep-laboratory workflow support. Given substantial heterogeneity, limited external validation, and low or very low certainty of evidence, prospective validation is needed before routine implementation.

Generative AI Use, Perceived Usefulness, Perceived Risk, and Physician Burnout and Fulfillment Among Chinese Physicians: Mixed Methods Multiregional Study

Background: As generative AI (GenAI) becomes increasingly prevalent, its impact on physician mental health has garnered significant attention; yet, empirical evidence remains limited. Objective: This study aims to investigate the correlations between the usage frequency of GenAI, perceived usefulness (PU), and perceived risk (PR) of GenAI with physicians’ burnout and professional fulfillment. Methods: A mixed methods design was used, integrating a quantitative survey of physicians across 4 regions in China with in-depth qualitative interviews to elucidate the underlying psychological mechanisms. The quantitative component involved a cross-sectional survey of 961 physicians, with the questionnaire collecting data on demographic and professional characteristics, socioeconomic status, GenAI usage frequency, PU, and PR. Semistructured interviews with 10 physicians were used for in-depth mining. Multivariable logistic and linear regression models with province-level fixed effects were fitted to examine the association between usage of GenAI, PU, PR, and physicians’ burnout and fulfillment. Stratified analyses were further performed to explore the moderating effect of demographic and clinical characteristics. Results: Quantitative analysis revealed no direct correlation between GenAI usage frequency and burnout. However, PU was positively associated with professional fulfillment (odds ratio [OR] 1.56, 95% CI 1.17-2.08; P=.003), whereas PR was associated with a higher likelihood of burnout (OR 1.80, 95% CI 1.46-2.21; P<.001). Stratified analyses showed that for physicians working ≥3 night shifts per week, GenAI usage was associated with higher odds of burnout, although the estimate was imprecise (OR 13.96, 95% CI 2.40-81.04; P=.003). The qualitative findings further suggested that the benefits of using GenAI may be offset by the additional burden. The PU of GenAI was perceived to enhance professional fulfillment by bolstering self-efficacy, whereas the PR of GenAI was linked to heightened burnout rooted in unclear boundaries of responsibilities and rights, as well as challenges to professional identity. Conclusions: The GenAI revolution in medicine is as much a psychological transition as it is a technological one. GenAI use is not directly associated with improved psychological states among clinicians. The PU of GenAI relates to professional fulfillment, and the PR concerns correspond to elevated burnout. Sustaining clinician well-being during this digital shift thus parallels a dual requirement, balancing the potential for professional fulfillment tied to GenAI utility against the concurrent verification fatigue and legal uncertainty cluster around clinician burnout.

Clinicians’ Attitudes and Perceptions on the Adoption of AI in Mental Health Care: Scoping Review

Background: AI is increasingly being integrated into health care workflows, with growing interest in AI systems for documentation, screening, triage, monitoring, and decision support. Mental health care is particularly complex for AI implementation as clinical care largely depends on therapeutic relationships, contextual factors, empathic communication, and interpretation of subtle nonverbal cues. Although AI may offer opportunities to improve efficiency and access to care, its adoption is likely to depend on clinicians’ trust, ethical acceptability, safety, confidentiality, and clarity around professional responsibility. Objective: This scoping review aims to synthesize current evidence on mental health clinicians’ attitudes, perceptions, and beliefs regarding the use of AI tools within mental health care, including the perceived benefits, risks, acceptable use cases, and conditions considered necessary for implementation. Methods: A scoping review was conducted in accordance with Joanna Briggs Institute (JBI) guidance and PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) guidelines. Six databases (CINAHL, Embase, PsycINFO, PubMed, Scopus, Web of Science) were searched on May 5, 2026, for studies published from 2020 onwards. Studies were eligible if they examined clinicians’ attitudes, perceptions, or beliefs regarding AI in mental health care. Studies were screened by 2 reviewers, first by title and abstract, and then by full text. Data extraction included publication year, country, methodology, participant occupation, AI type, outcome domains, and relevant findings. Results: Database searches retrieved 12,356 records. Following screening, 35 records were included. Overall, clinicians demonstrated cautious optimism toward AI, particularly when positioned as supplementing instead of replacing clinical expertise. Perceived benefits centered on reducing administrative burden, supporting documentation, synthesizing large volumes of information, and improving access to care, particularly between sessions. However, clinicians also reported limited AI literacy and prior use, and were concerned about privacy, confidentiality, governance, data ownership, clinical safety, unsafe or inaccurate outputs, overreliance, unclear accountability, professional role boundaries, and impacts on therapeutic relationships. Clinicians emphasized the need for education, clear guidelines and governance, role clarity, co-designed systems, human oversight, and real-world, ongoing evaluation. Conclusions: This review is innovative in shifting focus from the technical performance of AI tools to the perspectives of the clinicians who will be expected to use, interpret, explain, and remain accountable for them in mental health care. Unlike previous reviews, this review provides a clinician-centered understanding of AI adoption, highlighting that acceptability depends not only on what AI can do but also on whether it can be integrated safely, ethically, and in ways that preserve professional judgment and therapeutic relationships. These findings suggest that implementation should begin with lower-risk (eg, administrative), clinician-facing applications, be supported by education and governance, and be evaluated in mental health settings before wider adoption. These insights provide practical direction for responsible, clinician-centered implementation across mental health services.

What Provider Frequently Asked Questions Miss: Evaluating Unmet Attention-Deficit/Hyperactivity Disorder Information Needs Through Comparison of Online Community Posts Using Large Language Model–Assisted Semantic Analysis in a Mixed Methods Study

Background: Attention-deficit/hyperactivity disorder (ADHD) is a prevalent neurodevelopmental disorder that affects the functioning and quality of life of individuals throughout their lifespan. Despite the extensive information available online, patients and caregivers continue to report unmet needs, particularly regarding diagnosis, treatment, medication effects, comorbidities, and long-term management strategies. Existing provider-generated frequently asked questions (FAQs) are widely used, but often fail to fully capture the concerns expressed in online communities. Objective: This study aimed to (1) evaluate the extent to which provider-generated ADHD FAQs cover questions from online communities, (2) identify unmet information needs by analyzing questions with low semantic similarity to FAQs, and (3) compare the response styles of provider-generated answers with community-generated answers through large language model (LLM)–assisted analysis. Methods: ADHD-related questions from a Korean online community were semantically compared with provider-generated FAQs using sentence embedding–based similarity analysis to assess coverage and identify matched versus unmatched questions. Unmatched questions underwent topic modeling using the LimTopic framework, integrating BERTopic with LLM-assisted summarization to uncover unmet needs. An LLM-assisted content analysis was conducted on the answers to the high-similarity FAQ–community question pairs, enabling an examination of the response styles used by each group when addressing the public. Results: Through similarity comparison using embedding models and manual verification, the paraphrase-multilingual-MiniLM-L12-v2 (MBERT) model, which achieved the highest -score of 0.45, was selected as the final embedding model. The optimal similarity threshold determined for this model was 0.766, and the coverage of questions with similarity above this threshold between FAQs and the online community was 52.09% (2598/4988). Most of the coverage was concentrated on 18 FAQs. Online community questions below the similarity threshold were reviewed by experts after LimTopic analysis, resulting in the identification of 12 categories of unmet consumer needs, including school and social support, treatment accessibility, psychological support, and comorbidity management. Response style analysis revealed significant differences between evidence and authority signaling and the actionability dimension. Conclusions: Provider-generated ADHD FAQs covered approximately half of consumers’ questions, revealing substantial gaps in information provision for patients with ADHD. Health information on ADHD should expand beyond basic medical knowledge to address consumers’ real-world experiences, including access to care, school and social support, evidence-based treatments, daily functioning strategies, psychological support, health care navigation, and comorbidity management. Integrating these elements into provider-generated FAQs can create more comprehensive, consumer-centered resources that better support informed decision-making and long-term self-management.

Cost-Effectiveness of Electronic Patient-Reported Outcome Measure Interventions in Cancer: Systematic Review and Parameter Extraction for Economic Modeling

Background: Complex digital interventions that integrate electronic patient-reported outcome measures (ePROM) into clinical practice in cancer have the potential to improve quality of life, increase survival, and reduce health resource use and costs. Such systems can help patients with cancer self-manage chemotherapy symptoms, reduce clinicians’ workloads through automated decision support, and resolve problems earlier. However, more research on the cost-effectiveness of ePROM monitoring is needed. Objective: This paper comprises two complementary components: (1) a systematic literature review summarizing and evaluating the quantitative and qualitative evidence related to the cost-effectiveness of ePROM monitoring and (2) a health economic model parameter extraction. We also conducted supplementary targeted searches and scoping to provide context to our findings. Methods: We searched Ovid (including MEDLINE and Embase), Scopus, and the International Health Technology Assessment Database for original English-language papers published on or before March 2025 using search strings that combined terms related to ePROMs, health economics, and cancer/oncology. We included papers reporting health economic–related outcomes for ePROM interventions designed for adult cancer populations and excluded screening tools and conference abstracts. Results: We included 34 publications from 27 unique studies and identified and analyzed 26 ePROM-integrated interventions within these. Most (23/26) of the included interventions explicitly described some form of alert handling and automated decision support based on remote ePROM monitoring. Of the 34 publications, 5 presented full cost-effectiveness analysis results, of which 3 were highly uncertain and lacked clear differences in costs and health outcomes between ePROMs and standard care; conversely, 2 presented strong evidence of cost-effectiveness due to quality-of-life improvements, reduced hospitalizations, and potentially more autonomy in health-related travel (eg, ePROM-monitored patients can drive or walk to the hospital instead of using taxis or ambulances). A further 5 publications reported partial health economic results (eg, cost-consequence and budget impact), of which 1 detected no difference in strategies; in contrast, 4 reported lower health resource use and costs of ePROMs, mainly due to hospitalization reductions. Overall, 12 of the 27 studies included a qualitative component but mostly focused on user experience and design-related themes; only 2 of these addressed economic-specific themes (eg, changes in workflow and resource use due to ePROM implementation and integration), indicating some potential for time saving due to ePROM monitoring. Conclusions: Some ePROM-integrated interventions demonstrated cost-effectiveness in cancer care, but the evidence base remains limited. Where evidence does exist, cost-effectiveness appears driven by reduced hospitalization and improved quality of life. Qualitative research within the included studies rarely addressed economic questions. We provide a detailed parameter extraction for use in future economic modeling and recommend research priorities, including quantitative mapping of ePROM symptom data onto health resource use patterns, and qualitative work exploring how ePROM implementation affects clinical workloads and patient-perspective costs.
Received — 10 September 2026 ⏭ Journal of Medical Internet Research

Effectiveness, Safety, and Workflow Burden of Large Language Model–Based Medical Report Generation: Systematic Review

Background: Systems based on large language models (LLMs), multimodal LLMs, and vision-language foundation models are increasingly being evaluated for medical report generation in imaging and related clinical workflows. Existing reviews have summarized technical architectures, radiology applications, readability, and benchmark performance, but clinical readiness remains uncertain because safety, human oversight, and workflow outcomes are sparsely and inconsistently reported. Objective: The aim of this study is to assess the effectiveness (expert acceptance and blinded preference), safety (clinically significant, omission, and commission errors), and workflow burden (reporting time, corrections, edit distance, and editing burden) of LLM-based medical report generation. Methods: We searched PubMed/MEDLINE, Embase, Web of Science Core Collection, Scopus, and the Cochrane Library for studies published from January 1, 2016, through May 15, 2026. Eligible studies evaluated LLMs, multimodal LLMs, or vision-language foundation models for image-to-report generation, impression generation from findings, report drafting, or structured reporting in imaging workflows. Two reviewers performed screening, extraction, risk-of-bias assessment, and Grading of Recommendations Assessment, Development, and Evaluation–informed narrative certainty assessment. Outcomes were clinically significant error rate, omission error rate, commission error rate, reporting time, edit burden, expert acceptance, and blinded expert preference. Meta-analysis was not performed because no comparable outcome had at least 2 studies with compatible task structure and analyzable data. Results: A total of 101 studies were included. Chest x-ray was the largest modality group (36 studies), followed by computed tomography, magnetic resonance imaging (MRI), ultrasound, endoscopy, pathology, ophthalmic, electrocardiographic, dental, and mixed-modality contexts. No study was judged at low risk of bias; 15 were moderate, 72 high, and 14 serious. Safety and workflow evidence remained heterogeneous and largely nonpoolable. In a chest x-ray study, AI report acceptance was similar to that of radiologist reports (6047/8580, 70.5% vs 6288/8580, 73.3%), but false-negative findings were slightly higher (1584/8580, 18.5% vs 1527/8580, 17.8%). In a clinician-collaboration chest x-ray study, AI reports were equivalent or preferred in 233 of 300 (77.7%) and 170 of 303 (56.1%) cases across 2 datasets; yet, clinically significant errors persisted. In a brain MRI study, AI assistance reduced reading time from 61 to 53 seconds, whereas impression drafting increased editing time and edit distance. Conclusions: This review shifts the synthesis from plausible report generation to clinically interpretable effectiveness, safety, and workflow effects. Expert acceptance and preference suggested assistive value in selected supervised settings, but these signals were limited by inconsistent reporting of clinically significant errors, omissions, commissions, and failed generations. Workflow effects were mixed, with some studies reporting shorter reading time or drafting support, and others reporting greater editing time or edit distance. The evidence remains too heterogeneous, biased, and sparse on case-level end points to support a pooled meta-analysis or autonomous clinical-readiness claims. Adoption should remain locally validated, clinician-supervised, and accompanied by standardized reporting of acceptance, preference, omissions, commissions, failed generations, reporting time, corrections, and editing burden. Trial Registration: PROSPERO CRD420261302844; https://www.crd.york.ac.uk/PROSPERO/view/CRD420261302844

Smartphone-Based Monitoring of Quality of Life and Adverse Events After Neurosurgery: Prospective Cohort Study

Background: Postoperative outcome assessment is often based on discrete follow-up visits, limiting characterization of individual recovery trajectories, and the timely identification of adverse events (AEs). Longitudinal smartphone-based monitoring may overcome these limitations by enabling frequent, resource-efficient collection of patient-reported outcomes and complications throughout recovery. Such data may provide a more patient-centered understanding of the postoperative course and complement conventional clinical surveillance. Objective: This study aimed to evaluate the feasibility of smartphone-based longitudinal monitoring of quality of life, subjective well-being, and AEs after elective neurosurgery and compare postoperative recovery trajectories and agreement between patient- and clinician-reported AEs. Methods: This interim analysis of a prospective cohort study included adult patients undergoing elective lumbar decompression, lumbar fusion, supratentorial craniotomy, or infratentorial craniotomy at a Swiss tertiary referral center between June 2023 and January 2025. Participants used a smartphone app to longitudinally report subjective well-being (Subjective Well-Being Index; 0‐10), quality of life (EQ-5D-5L), and AEs for up to 1 year postoperatively. Complications were self-reported using the Therapy-Disability-Neurology (TDN) classification and retrospectively adjudicated by physicians. Descriptive analyses assessed data density, engagement, and concordance between patient- and clinician-reported events. Mixed-effects models were used to evaluate factors associated with postoperative well-being. Results: Of the 100 enrolled patients (median age 64.0, IQR 52.95‐71.6 years; n=45, 45% women), 86 (86%) provided postoperative data. During a median follow-up of 3.2 (IQR 0.2‐11.3) months, participants submitted 4354 longitudinal well-being entries. Patients reported 22 unique AEs, whereas physicians identified 44 AEs, with overlap for 9 (20.5%) events. Most physician-reported AEs were mild (30/44, 68.2%; TDN grade 1‐2), and no grade 4 or 5 events occurred. Patient-reported AEs primarily reflected symptomatic and functional impairments, whereas physician-reported events more often included clinically detected or subclinical findings. In mixed-effects models, time since surgery was associated with improved well-being, and no other factors were statistically significant. Conclusions: Smartphone-based postoperative monitoring was feasible in this elective neurosurgical cohort and generated dense longitudinal patient-reported data beyond routine follow-up. Patient and clinician AE reporting captured partly distinct aspects of postoperative recovery, suggesting that smartphone-based self-reporting may complement rather than replace clinical surveillance. Trial Registration: ClinicalTrials.gov NCT06352710; https://clinicaltrials.gov/study/NCT06352710

Natural Language Processing Identification of Nonprescribed Fentanyl Use in Electronic Health Records: Algorithm Development and Validation Study

Background: Overdose and suicide due to nonprescribed fentanyl use have increased significantly, yet health care systems lack reliable methods to identify patients who use nonprescribed fentanyl. codes are inconsistent and do not specify nonprescribed fentanyl use. Objective: This study aimed to develop natural language processing approaches to identifying nonprescribed fentanyl use in electronic health record (EHR) documentation. Methods: This retrospective study included Veterans Health Administration patients seen between April 5, 2023, and December 23, 2024. A term list was developed to identify fentanyl-related mentions in clinical text, and 250-character snippets surrounding identified mentions were extracted. Veterans (n=3878) were randomly sampled from 5 predefined groups based on the presence of 1 of 4 terms (“fent,” “blues,” “M30s,” and “tranq”) in their EHR documentation. Physician annotators classified snippets into “nonprescribed fentanyl use,” “prescribed fentanyl use,” or “other,” with interannotator agreement evaluated using the mean pairwise Cohen κ. Cross-validation folds were constructed at the patient level between training and test sets. Penalized logistic regression, Bio-ClinicalBERT, Llama 3-8B, and Mistral-7B were trained on labeled data and compared. Model performance was evaluated using precision, recall, and -scores for each class, with a focus on the nonprescribed fentanyl use class as the primary label of clinical interest using bootstrapped 95% CIs. A fairness analysis and Shapley additive explanations analysis were performed using Bio-ClinicalBERT. External validation was performed using Bio-ClinicalBERT on an independent sample of 200 snippets, each representing a unique patient from January 2025 to June 2026, with precision reported as the primary validation metric. Results: Of 7389 snippets, 9.6% (n=709) were classified as “nonprescribed fentanyl use,” 40.3% (n=2981) were classified as “prescribed fentanyl use,” and 50% (n=3699) were classified as “other.” Interannotator agreement was high (κ=0.822). Llama 3-8B achieved the highest -score for nonprescribed fentanyl use (0.87, 95% CI 0.83-0.92), followed by Mistral-7B (0.80, 95% CI 0.75-0.84), Bio-ClinicalBERT (0.80, 95% CI 0.74-0.85), and penalized logistic regression (0.74, 95% CI 0.73-0.75). Performance was consistent across demographic subgroups, with lower performance for the nonprescribed fentanyl use class observed in female and Hispanic subgroups. Shapley additive explanations analysis revealed clinically meaningful discriminating terms for each class, although subword tokens required contextual interpretation. External validation of Bio-ClinicalBERT demonstrated a precision of 0.79 for nonprescribed fentanyl use. Conclusions: Natural language processing can identify nonprescribed fentanyl use in EHR documentation, although model performance for this class was lower than overall model performance, reflecting the clinical complexity of identifying nonprescribed use and the variable ways in which clinicians document this problem. This approach may support risk prediction and targeting of interventions to patients exposed to nonprescribed fentanyl.

Large Language Models for Distress Rating in Korean Psycho-Oncology Interviews: Exploratory Clinician-Benchmarked Evaluation Study

Background: Psychiatric distress is common among patients with cancer; yet, systematic interview-based screening remains difficult to scale in routine clinical care. Large language models (LLMs) have shown promise as scalable tools for mental health assessment, but most existing evidence is derived from clinician-authored records, translated text, or proxy data. The performance characteristics, error patterns, and explanatory behaviors of contemporary LLMs when applied to authentic, non-English psychiatric interviews remain insufficiently characterized. Objective: This exploratory study evaluated how contemporary LLMs reproduced psycho-oncologists’ item-level symptom ratings from real-world Korean psycho-oncology interviews, focusing on concordance, directional bias, and clinician-adjudicated error characteristics. Methods: Between April 2024 and May 2025, 101 adults receiving oncologic care in South Korea underwent semistructured interviews. Board-certified psycho-oncologists provided real-time, time-stamped ratings of 39 items. Generative pretrained transformer 4o (GPT-4o), Claude 3.5 Sonnet, and Gemini 2.5 Flash generated item scores and brief rationales using identical Korean zero-shot rubrics. Concordance between clinician ratings was evaluated using ordinal and binary screening metrics. A paired Wilcoxon test compared patient-level total symptom burden. Binary mismatches were clinically adjudicated as ambiguous or definite overestimation or underestimation with an 8-etiology taxonomy. Model-generated rationales were further meta-evaluated using GPT-5.4, Claude Sonnet 4.6, and Gemini Pro 3.1 across 4 dimensions: citation (use of quoted supporting statements), structure (logical organization of the rationale), mapping (consistency between the rationale and the assigned item rating), and expansion (degree of interpretive elaboration beyond the explicit transcript content). The associations between meta-evaluation results and absolute error were examined using cross-classified mixed-effects models. Results: Out of 101 participants, 88 (87.1%) were predominantly female and had breast cancer as the primary cancer type (n=70, 69.3%). Across 3931 item-level ratings, all models showed good agreement with clinicians (intraclass correlation coefficient: 0.816‐0.872), with GPT-4o showing the highest agreement. Claude 3.5 and Gemini 2.5 yielded significantly higher patient-level symptom burden (adjusted
❌