❌

Reading view

Effectiveness of Wearable Digital Therapeutics in Improving Sleep Outcomes Among Individuals With Insomnia: Systematic Review and Meta-Analysis of Randomized Controlled Trials

Background: Wearable devices are increasingly used for sleep monitoring and as adjunctive treatment. Existing meta-analyses mostly pool composite digital therapies and rarely isolate stand-alone wearables or distinguish between objective and subjective end points. Whether stand-alone wearable interventions improve sleep outcomes in adults with insomnia, and which factors moderate treatment heterogeneity, remains unclear. Objective: This study aims to evaluate the effectiveness of wearable digital interventions on sleep outcomes in adults with insomnia versus control strategies and explore moderators of effectiveness, including device-wearing position, intervention duration, and control type, using meta-regression. Methods: This systematic review and meta-analysis was conducted in accordance with the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta‑Analyses) 2020 statement and the PRISMA-S (Preferred Reporting Items for Systematic Reviews and Meta‑Analyses Literature Search Extension) guideline. Five electronic databases and clinical trial registries were searched from inception to May 18, 2026. Eligible studies were randomized controlled trials (RCTs) evaluating wearable digital interventions in adults with insomnia compared with sham, waitlist, usual care, or active control conditions and had an intervention duration of at least 1 week. Study screening, data extraction, and risk-of-bias assessment were carried out independently by 2 reviewers. Pooled estimates were calculated using a restricted maximum likelihood random-effects model with the Hartung-Knapp-Sidik-Jonkman correction. Heterogeneity was assessed using the ² statistic, and 95% prediction intervals (PIs) were calculated for the primary analyses. The certainty of evidence was rated using the GRADE (Grading of Recommendations, Assessment, Development, and Evaluation) approach. Results: Sixteen RCTs (N=910) were included. Wearable digital interventions were associated with a significant reduction in objective sleep-onset latency (SOL; mean difference [MD] −4.52, 95% CI −8.38 to −0.67, PI −9.52 to 0.47 min) and a significant improvement in subjective sleep efficiency (SE; MD 2.00%, 95% CI 1.90%‐2.11%, PI 1.85%‐2.15%). Subjective total sleep time (TST) also showed a significant increase (MD 19.11, 95% CI 2.98‐35.24, PI −16.20 to 54.43 minutes). Meta-regression showed that control type, intervention duration, and device location did not explain the heterogeneity of the insomnia severity index (ISI) (=0). Sensitivity analysis confirmed the robustness of pooled ISI estimates, and an Egger test indicated no small-study effects (=.07). Certainty of evidence ranged from moderate to high. Conclusions: Wearable digital interventions provide selective benefits for objective SOL, subjective SE, and subjective TST in adults with insomnia, with no improvement in overall ISI. Despite statistically significant effects on several sleep parameters, wide PIs, substantial heterogeneity, and limited study numbers indicate preliminary, nonconclusive findings. Wearables should be viewed as affordable adjunctive tools requiring further validation, not substitutes for first-line cognitive behavioral therapy for insomnia. Large-scale, long-term RCTs with standardized protocols and patient-level external validation are required to consolidate the evidence base. Trial Registration: PROSPERO CRD420251038603; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251038603
  •  

Correction: From Metrics to Meaning in Neurological Rehabilitation: Clinicians’ Perspectives on Digital Metrics of Upper Limb Functioning—A Focus Group Study

Digital assessment technologies, such as optical motion capture and inertial measurement units, enable detailed kinematic analysis and continuous monitoring of upper limb activity in persons with neurological conditions. While such digital metrics of functioning are increasingly recognized in research, their uptake in clinical neurorehabilitation is limited. It remains unclear which digital metrics of functioning clinicians perceive as most meaningful and how these are integrated into patient-centered care. Understanding clinicians’ information needs and reasoning processes is a prerequisite for implementing digital assessment technology. To characterize how rehabilitation professionals perceive, prioritize, and integrate digital metrics of functioning into clinical reasoning and to identify features that would support their routine use. Three 90-minute focus groups were conducted in 3 Swiss neurorehabilitation centers, involving 11 clinicians with diverse professional backgrounds (5 physiotherapists, 4 occupational therapists, 1 movement scientist, and 1 medical practitioner). Participants discussed essential parameter domains and individually rated the relevance and meaningfulness of 17 kinematic metrics for the well-studied drinking task and 10 established arm use performance metrics. Verbatim transcripts were analyzed using reflexive thematic analysis, and rating data were summarized descriptively. Five main themes were identified. (1) Functional requirements to interpret movement quality and performance (active/passive range of motion (ROM), strength, selective muscle control, grasp) form the basis for interpreting movement. (2) Essential aspects of movement quality (smoothness, efficiency, compensatory movement) are valued when aligned with observable task execution. (3) Added value of real-world performance (hourly activity profiles, arm-use symmetry, functional workspace) represents the reference for patient-centered reasoning. (4) Individualizing what matters, including diagnosis-specific preferences, shapes assessment selection. (5) Blending clinical eye and reference data reflects clinicians’ reliance on visual judgment complemented by normative values. Intuitive metrics such as task duration, number of movement units, and ROM were favored, whereas confidence was lower in more complex metrics (e.g., jerk, inter-joint coordination). Clinicians value intuitive digital metrics of functioning when they are clearly linked to patient-centered outcomes and supported by normative references. The findings highlight the need for targeted educational strategies and digital competency training that help clinicians interpret digital metrics and integrate them with contextual information and clinical reasoning.
  •  

Evaluating Large Language Models in Clinical Audiology (AUDIOLOGYBENCH): Benchmark Development and Validation Study

Background: Large language models (LLMs) are increasingly being explored for clinical decision support, but their performance in audiology has not been systematically benchmarked using clinically grounded case materials and rubric-based safety evaluations. Objective: This study aimed to develop and evaluate AUDIOLOGYBENCH, a 3-tier benchmark for characterizing frontier LLM capability in clinical audiology along (1) curated domain knowledge, (2) literature-derived evidence, and (3) clinical reasoning under multimodal case input, with an explicit human audit of the automated adjudicator on the primary end point. Methods: The benchmark comprises 3139 objective items from educational resources, 3175 research article–derived items from peer-reviewed articles published between 2015 and 2025, and 67 multimodal clinical case studies graded against a standardized A-F rubric with 6 prespecified critical-error types that cap scores at D or F. Eight models were evaluated on the educational objective items: 4 frontier multimodal models (Gemini 2.5 Pro, Grok 4, OpenAI O3, and Claude Sonnet 4 Thinking) were evaluated on the research article–derived items, and on 804 case study evaluations. Adjudication used Gemini 2.5 Pro (objective and research-derived items) and Claude Opus 4.5 (case studies). The case study adjudicator was independently audited against PhD-level audiologist consensus on blinded subsamples, supplemented by a post-stratified human-calibrated sensitivity analysis. Results: A striking task-type dissociation emerged on case studies: clinical recommendations (Q3) achieved a mean score of 89.74 (SD 13.92, 95% CI 88.07‐91.41), a 98.1% (263/268) pass rate, and no dangerous recommendations; audiometric numerical interpretation (Q1) achieved a mean score of 67.89 (SD 18.47, 95% CI 65.68‐70.10), with a 35.4% (95/268) critical-error rate; and differential diagnosis (Q2) achieved a mean score of 67.79 (SD 15.33, 95% CI 65.95‐69.63). Question type, not model selection, dominated performance (eta-squared_H=0.333 vs 0.001; rank biserial ≥0.679). Interreviewer reliability between audiologists was high (quadratic-weighted κ of 0.78 and 0.85 across the 80-item and 50-item audits, respectively). When 2 audiologists regraded all 80 model Q1 responses with the diagnostic images available, the adjudicator’s per-item Q1 labels diverged from human judgment (κ=0.05; overflagging; sensitivity: 19/26, 73%; positive predictive value: 19/53, 36%), yet its reweighted Q1 critical-error rate (36.2%) was broadly consistent with the image-grounded human estimates (28%‐34%), suggesting no systematic inflation of the headline rate. The principal Q3>{Q1, Q2} ranking was preserved under post-stratified human calibration. Web-style multiple-choice items showed ceiling effects (>95% accuracy); short-answer prompts remained challenging (best 30%). Conclusions: Current frontier LLMs show strong recommendation generation but substantial limitations in audiometric numerical interpretation that are shared across models and that an automated adjudicator partially miscalibrated at the per-item level. AUDIOLOGYBENCH characterizes capability boundaries rather than certifying clinical readiness. Deployment of LLM-assisted audiology workflows requires structured human verification of all numerical findings and awareness of fabrication and severity misclassification failure modes documented here.
  •  

The Effectiveness of Digital Intervention on Psychological Resilience in Postoperative Breast Cancer Patients During Chemotherapy Intervals: Quasi-Experimental Study

Background: Patients with breast cancer during postoperative chemotherapy intervals commonly experience psychological distress and reduced resilience while recovering at home. Digital mindfulness interventions may provide accessible psychological support during this vulnerable period; however, evidence regarding tailored interventions for postoperative patients with breast cancer during chemotherapy intervals remains limited. Objective: This study aimed to examine the effectiveness of a digital intervention on psychological resilience in postoperative patients with breast cancer during chemotherapy intervals. Methods: A quasi-experimental study with repeated measures was conducted from October 2021 to June 2022. A total of 80 eligible participants were recruited from the Department of Breast Surgery at a tertiary hospital in Zhejiang Province, China, and 71 completed the study. The control group received routine discharge instructions and nursing follow-ups, whereas the intervention group additionally received an 8-week digital psychological resilience intervention. Outcomes were assessed at baseline (T0), 3 months post intervention (T1), and 6 months post intervention (T2). The measures included the Connor-Davidson Resilience Scale (CD-RISC), Hospital Anxiety and Depression Scale (HADS), Social Support Rating Scale (SSRS), Breast Cancer Survivor Self-Efficacy Scale (BCSSS), and Functional Assessment of Cancer Therapy-Breast (FACT-B). Independent-samples tests, chi-square tests, and repeated-measures ANOVA were performed using SPSS (version 26.0; IBM Corp). Results: No statistically significant baseline differences were observed between the two groups in the outcome measures. At T1, the intervention group had higher CD-RISC scores than the control group (mean 67.58, SD 11.41 vs mean 62.09, SD 10.18; =.036) and higher BCSSS scores (mean 42.36, SD 3.59 vs mean 39.23, SD 4.90; =.003). However, these between-group differences were no longer statistically significant at T2 (>.05). Significant time effects and group×time interaction effects were observed for both psychological resilience and self-efficacy (.05), although both scales showed significant time effects (
  •  

Performance of AI-Based Screening Tools for Obstructive Sleep Apnea Across Apnea-Hypopnea Index Thresholds: Systematic Review and Meta-Analysis

Background: Obstructive sleep apnea (OSA) is highly prevalent but remains substantially underdiagnosed. Polysomnography (PSG) is the reference standard, but its cost and limited availability constrain large-scale case identification. AI-based screening tools may support risk stratification and referral prioritization, but their diagnostic accuracy across apnea-hypopnea index (AHI) thresholds remains uncertain. Objective: This review aimed to systematically evaluate the diagnostic accuracy of AI-based OSA screening tools at AHI thresholds of ≥5, ≥15, and ≥30 events/hour, with emphasis on models using non-PSG–derived inputs. Methods: PubMed, Embase, Scopus, and Web of Science were searched for studies published from January 1, 2016, to May 3, 2026. Eligible studies included adults evaluated for suspected OSA or recruited from population-based cohorts, assessed AI-based models intended or interpretable for OSA screening, risk prediction, or screening-oriented severity classification, used PSG as the reference standard, and reported sufficient data to construct or reconstruct 2×2 contingency tables. Diagnostic accuracy was synthesized separately by AHI threshold and input source using bivariate random-effects models, with 95% CIs and prediction intervals (PIs). Risk of bias and certainty of evidence were assessed using QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies 2) and GRADE (Grading of Recommendations Assessment, Development, and Evaluation), respectively. Results: A total of 60 studies were included, of which 47 contributed data to the meta-analysis. At AHI thresholds of ≥5, ≥15, and ≥30 events/hour, pooled sensitivities were 0.94 (95% CI 0.92‐0.96; 95% PI 0.71‐0.99), 0.87 (95% CI 0.84‐0.89; 95% PI 0.66‐0.96), and 0.83 (95% CI 0.79‐0.87; 95% PI 0.61‐0.94), respectively; the corresponding specificities were 0.77 (95% CI 0.69‐0.84; 95% PI 0.30‐0.96), 0.81 (95% CI 0.75‐0.85; 95% PI 0.39‐0.96), and 0.91 (95% CI 0.87‐0.94; 95% PI 0.55‐0.99), respectively. The corresponding areas under the summary receiver operating characteristic curves were 0.943, 0.907, and 0.920. For non-PSG–derived tools, sensitivities were 0.92, 0.85, and 0.81, and specificities were 0.70, 0.74, and 0.85 at the 3 thresholds, respectively. For PSG-derived models, sensitivities were 0.96, 0.90, and 0.85, and specificities were 0.82, 0.88, and 0.96, respectively. Exploratory subgroup analyses suggested performance variation across selected study and model characteristics, including region, algorithmic framework, data source, and validation method. Conclusions: AI-based tools showed generally favorable screening performance for OSA across clinically relevant AHI thresholds, although wide PIs suggest variable performance across future comparable populations and settings. By synthesizing diagnostic accuracy across 3 AHI thresholds and distinguishing non-PSG–derived from PSG-derived models, this review extends previous broad or modality-specific reviews and offers a clinically interpretable, pathway-specific basis for linking model performance to intended use. The findings may clarify potential roles for non-PSG–derived tools in front-end screening and referral prioritization and for PSG-derived models in reduced-channel assessment and sleep-laboratory workflow support. Given substantial heterogeneity, limited external validation, and low or very low certainty of evidence, prospective validation is needed before routine implementation.
  •  

Generative AI Use, Perceived Usefulness, Perceived Risk, and Physician Burnout and Fulfillment Among Chinese Physicians: Mixed Methods Multiregional Study

Background: As generative AI (GenAI) becomes increasingly prevalent, its impact on physician mental health has garnered significant attention; yet, empirical evidence remains limited. Objective: This study aims to investigate the correlations between the usage frequency of GenAI, perceived usefulness (PU), and perceived risk (PR) of GenAI with physicians’ burnout and professional fulfillment. Methods: A mixed methods design was used, integrating a quantitative survey of physicians across 4 regions in China with in-depth qualitative interviews to elucidate the underlying psychological mechanisms. The quantitative component involved a cross-sectional survey of 961 physicians, with the questionnaire collecting data on demographic and professional characteristics, socioeconomic status, GenAI usage frequency, PU, and PR. Semistructured interviews with 10 physicians were used for in-depth mining. Multivariable logistic and linear regression models with province-level fixed effects were fitted to examine the association between usage of GenAI, PU, PR, and physicians’ burnout and fulfillment. Stratified analyses were further performed to explore the moderating effect of demographic and clinical characteristics. Results: Quantitative analysis revealed no direct correlation between GenAI usage frequency and burnout. However, PU was positively associated with professional fulfillment (odds ratio [OR] 1.56, 95% CI 1.17-2.08; P=.003), whereas PR was associated with a higher likelihood of burnout (OR 1.80, 95% CI 1.46-2.21; P<.001). Stratified analyses showed that for physicians working ≥3 night shifts per week, GenAI usage was associated with higher odds of burnout, although the estimate was imprecise (OR 13.96, 95% CI 2.40-81.04; P=.003). The qualitative findings further suggested that the benefits of using GenAI may be offset by the additional burden. The PU of GenAI was perceived to enhance professional fulfillment by bolstering self-efficacy, whereas the PR of GenAI was linked to heightened burnout rooted in unclear boundaries of responsibilities and rights, as well as challenges to professional identity. Conclusions: The GenAI revolution in medicine is as much a psychological transition as it is a technological one. GenAI use is not directly associated with improved psychological states among clinicians. The PU of GenAI relates to professional fulfillment, and the PR concerns correspond to elevated burnout. Sustaining clinician well-being during this digital shift thus parallels a dual requirement, balancing the potential for professional fulfillment tied to GenAI utility against the concurrent verification fatigue and legal uncertainty cluster around clinician burnout.
  •  

Clinicians’ Attitudes and Perceptions on the Adoption of AI in Mental Health Care: Scoping Review

Background: AI is increasingly being integrated into health care workflows, with growing interest in AI systems for documentation, screening, triage, monitoring, and decision support. Mental health care is particularly complex for AI implementation as clinical care largely depends on therapeutic relationships, contextual factors, empathic communication, and interpretation of subtle nonverbal cues. Although AI may offer opportunities to improve efficiency and access to care, its adoption is likely to depend on clinicians’ trust, ethical acceptability, safety, confidentiality, and clarity around professional responsibility. Objective: This scoping review aims to synthesize current evidence on mental health clinicians’ attitudes, perceptions, and beliefs regarding the use of AI tools within mental health care, including the perceived benefits, risks, acceptable use cases, and conditions considered necessary for implementation. Methods: A scoping review was conducted in accordance with Joanna Briggs Institute (JBI) guidance and PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) guidelines. Six databases (CINAHL, Embase, PsycINFO, PubMed, Scopus, Web of Science) were searched on May 5, 2026, for studies published from 2020 onwards. Studies were eligible if they examined clinicians’ attitudes, perceptions, or beliefs regarding AI in mental health care. Studies were screened by 2 reviewers, first by title and abstract, and then by full text. Data extraction included publication year, country, methodology, participant occupation, AI type, outcome domains, and relevant findings. Results: Database searches retrieved 12,356 records. Following screening, 35 records were included. Overall, clinicians demonstrated cautious optimism toward AI, particularly when positioned as supplementing instead of replacing clinical expertise. Perceived benefits centered on reducing administrative burden, supporting documentation, synthesizing large volumes of information, and improving access to care, particularly between sessions. However, clinicians also reported limited AI literacy and prior use, and were concerned about privacy, confidentiality, governance, data ownership, clinical safety, unsafe or inaccurate outputs, overreliance, unclear accountability, professional role boundaries, and impacts on therapeutic relationships. Clinicians emphasized the need for education, clear guidelines and governance, role clarity, co-designed systems, human oversight, and real-world, ongoing evaluation. Conclusions: This review is innovative in shifting focus from the technical performance of AI tools to the perspectives of the clinicians who will be expected to use, interpret, explain, and remain accountable for them in mental health care. Unlike previous reviews, this review provides a clinician-centered understanding of AI adoption, highlighting that acceptability depends not only on what AI can do but also on whether it can be integrated safely, ethically, and in ways that preserve professional judgment and therapeutic relationships. These findings suggest that implementation should begin with lower-risk (eg, administrative), clinician-facing applications, be supported by education and governance, and be evaluated in mental health settings before wider adoption. These insights provide practical direction for responsible, clinician-centered implementation across mental health services.
  •  

What Provider Frequently Asked Questions Miss: Evaluating Unmet Attention-Deficit/Hyperactivity Disorder Information Needs Through Comparison of Online Community Posts Using Large Language Model–Assisted Semantic Analysis in a Mixed Methods Study

Background: Attention-deficit/hyperactivity disorder (ADHD) is a prevalent neurodevelopmental disorder that affects the functioning and quality of life of individuals throughout their lifespan. Despite the extensive information available online, patients and caregivers continue to report unmet needs, particularly regarding diagnosis, treatment, medication effects, comorbidities, and long-term management strategies. Existing provider-generated frequently asked questions (FAQs) are widely used, but often fail to fully capture the concerns expressed in online communities. Objective: This study aimed to (1) evaluate the extent to which provider-generated ADHD FAQs cover questions from online communities, (2) identify unmet information needs by analyzing questions with low semantic similarity to FAQs, and (3) compare the response styles of provider-generated answers with community-generated answers through large language model (LLM)–assisted analysis. Methods: ADHD-related questions from a Korean online community were semantically compared with provider-generated FAQs using sentence embedding–based similarity analysis to assess coverage and identify matched versus unmatched questions. Unmatched questions underwent topic modeling using the LimTopic framework, integrating BERTopic with LLM-assisted summarization to uncover unmet needs. An LLM-assisted content analysis was conducted on the answers to the high-similarity FAQ–community question pairs, enabling an examination of the response styles used by each group when addressing the public. Results: Through similarity comparison using embedding models and manual verification, the paraphrase-multilingual-MiniLM-L12-v2 (MBERT) model, which achieved the highest -score of 0.45, was selected as the final embedding model. The optimal similarity threshold determined for this model was 0.766, and the coverage of questions with similarity above this threshold between FAQs and the online community was 52.09% (2598/4988). Most of the coverage was concentrated on 18 FAQs. Online community questions below the similarity threshold were reviewed by experts after LimTopic analysis, resulting in the identification of 12 categories of unmet consumer needs, including school and social support, treatment accessibility, psychological support, and comorbidity management. Response style analysis revealed significant differences between evidence and authority signaling and the actionability dimension. Conclusions: Provider-generated ADHD FAQs covered approximately half of consumers’ questions, revealing substantial gaps in information provision for patients with ADHD. Health information on ADHD should expand beyond basic medical knowledge to address consumers’ real-world experiences, including access to care, school and social support, evidence-based treatments, daily functioning strategies, psychological support, health care navigation, and comorbidity management. Integrating these elements into provider-generated FAQs can create more comprehensive, consumer-centered resources that better support informed decision-making and long-term self-management.
  •  

Cost-Effectiveness of Electronic Patient-Reported Outcome Measure Interventions in Cancer: Systematic Review and Parameter Extraction for Economic Modeling

Background: Complex digital interventions that integrate electronic patient-reported outcome measures (ePROM) into clinical practice in cancer have the potential to improve quality of life, increase survival, and reduce health resource use and costs. Such systems can help patients with cancer self-manage chemotherapy symptoms, reduce clinicians’ workloads through automated decision support, and resolve problems earlier. However, more research on the cost-effectiveness of ePROM monitoring is needed. Objective: This paper comprises two complementary components: (1) a systematic literature review summarizing and evaluating the quantitative and qualitative evidence related to the cost-effectiveness of ePROM monitoring and (2) a health economic model parameter extraction. We also conducted supplementary targeted searches and scoping to provide context to our findings. Methods: We searched Ovid (including MEDLINE and Embase), Scopus, and the International Health Technology Assessment Database for original English-language papers published on or before March 2025 using search strings that combined terms related to ePROMs, health economics, and cancer/oncology. We included papers reporting health economic–related outcomes for ePROM interventions designed for adult cancer populations and excluded screening tools and conference abstracts. Results: We included 34 publications from 27 unique studies and identified and analyzed 26 ePROM-integrated interventions within these. Most (23/26) of the included interventions explicitly described some form of alert handling and automated decision support based on remote ePROM monitoring. Of the 34 publications, 5 presented full cost-effectiveness analysis results, of which 3 were highly uncertain and lacked clear differences in costs and health outcomes between ePROMs and standard care; conversely, 2 presented strong evidence of cost-effectiveness due to quality-of-life improvements, reduced hospitalizations, and potentially more autonomy in health-related travel (eg, ePROM-monitored patients can drive or walk to the hospital instead of using taxis or ambulances). A further 5 publications reported partial health economic results (eg, cost-consequence and budget impact), of which 1 detected no difference in strategies; in contrast, 4 reported lower health resource use and costs of ePROMs, mainly due to hospitalization reductions. Overall, 12 of the 27 studies included a qualitative component but mostly focused on user experience and design-related themes; only 2 of these addressed economic-specific themes (eg, changes in workflow and resource use due to ePROM implementation and integration), indicating some potential for time saving due to ePROM monitoring. Conclusions: Some ePROM-integrated interventions demonstrated cost-effectiveness in cancer care, but the evidence base remains limited. Where evidence does exist, cost-effectiveness appears driven by reduced hospitalization and improved quality of life. Qualitative research within the included studies rarely addressed economic questions. We provide a detailed parameter extraction for use in future economic modeling and recommend research priorities, including quantitative mapping of ePROM symptom data onto health resource use patterns, and qualitative work exploring how ePROM implementation affects clinical workloads and patient-perspective costs.
  •  
❌