❌

Reading view

Extent of Digital Health Fragmentation and Potential Implications for Antimicrobial Prescribing: Rapid Evidence Review

Background: Prior microbiology results, resistance patterns, and antimicrobial exposure are central to safe and effective antimicrobial prescribing. Digital health fragmentation refers to the dispersal of patient data across multiple electronic systems and the associated challenge of accessing complete information at the point of care. Antimicrobial prescribing for infections represents a critical use case to investigate the impact of digital health fragmentation on patient care. While interoperability has been studied in the context of patient safety, no review has described digital health fragmentation within the United Kingdom and examined its impact on antimicrobial prescribing and antimicrobial stewardship (AMS). Objective: This study aimed to (1) characterize the extent of digital health fragmentation in the United Kingdom, (2) summarize the available evidence on its impact on AMS and prescribing practices in high-income countries, and (3) identify potential solutions. Methods: A rapid review of the peer-reviewed literature was conducted following published guidance for rapid reviews and the PRISMA (Preferred Reporting Items of Systematic Reviews and Meta-Analyses) statement. MEDLINE ALL and PsycInfo were searched on August 19, 2025, using search terms relating to digital health fragmentation or interoperability, patient safety, and antimicrobial use. Searches were limited to English-language publications from 2015 (for characterizing the recent trends or current state of digital health fragmentation in the United Kingdom) or 2010 onward (for AMS-related impacts and solutions). Screening was conducted by 4 researchers following predefined inclusion and exclusion criteria. Extracted data were synthesized narratively through framework analysis. Study quality was appraised using the Mixed Methods Appraisal Tool. Results: Fourteen studies met the inclusion criteria. Ten studies described the extent and nature of digital health fragmentation in the United Kingdom. Digital health fragmentation affects a large number of patients and is linked to clinical care efficiency, quality, and safety risks, including limited access to external clinical records, missing or incomplete information, duplicate investigations, delays in decision‑making, and substantial time spent searching for data. Evidence specific to antimicrobial prescribing was limited (4 studies) but indicated that AMS relies on information spread across multiple systems, with poor interoperability disrupting workflows, hindering communication, and undermining stewardship activities. Only 1 study reported the development of a digital tool designed to address digital health fragmentation and support AMS. Conclusions: Digital health fragmentation negatively affects patient care across the United Kingdom, yet evidence on how it impacts AMS remains scarce. Given the urgency of the global antimicrobial resistance crisis, future research should therefore quantify the scale and impact of digital health fragmentation for AMS to inform investment and innovation in digital infrastructure and clinical-supportive solutions. Trial Registration: PROSPERO CRD420251126067; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251126067
  •  

Linking Dispense Data to Electronic Health Orders: Tutorial for Querying Commercial Pharmacy Databases to Support Systemwide Quality Improvement

Background: Assessing medication adherence is central to quality care, yet linking electronic health record (EHR) medication orders to outpatient pharmacy dispense data remains technically complex. Objective: This study aimed to present a generalized, reproducible tutorial for linking EHR medication orders to pharmacy dispense data that can be used to assess medication dispense proportions. Methods: We developed and validated a structured query approach to link EHR medication orders to external pharmacy dispense data using patient identifiers, medication-level identifiers, pharmacy identifiers, and temporal constraints. The tutorial emphasizes key design decisions, including handling multiple triggering events, deduplication across vendors, and managing formulation changes. A retrospective cohort of pediatric acute otitis media encounters (January 1, 2021, to January 1, 2024) was used as an illustrative example. Results: Overall, 98.3% (302/307) of pharmacies in the cohort returned at least 1 dispense record during the study period and were therefore classified as reporting pharmacies. Among 3404 orders, 2616 (76.9%) had a recorded dispense. Conclusions: EHR-integrated pharmacy data provide a feasible, timely proxy for assessing medication adherence. This tutorial provides a scalable framework for linking EHR and pharmacy data for medication adherence studies, while highlighting key methodological considerations for SQL coding.
  •  

Development and Preliminary Evaluation of a Conversational Agent Delivering Problem-Solving Therapy for Family Caregivers of Children With a Chronic Health Condition: Multiphase Mixed Methods Study

Background: Family caregivers of children with chronic health conditions experience substantial physical and mental health burdens, including burnout, anxiety, depression, fatigue, and sleep disturbances. Despite this need, validated digital mental health tools tailored to family caregivers remain limited. AI-powered conversational agents offer a promising approach for delivering on-demand, personalized mental health support, yet development and evaluation frameworks for this population are lacking. Objective: This paper describes the iterative development and formative evaluation of COCO (Caring of Caregivers Online), a conversational agent designed for family caregivers of children with chronic health conditions. COCO integrates problem-solving therapy (PST) and motivational interviewing (MI) within a human-in-the-loop development framework that progressed from rule-based interactions to a large language model (LLM)–powered conversational agent. Methods: COCO was developed across four phases: (1) caregiver persona and dialogue development based on PST and MI; (2) usability testing of a low-fidelity prototype with standardized patients in a single session of PST; (3) usability testing of a high-fidelity prototype with caregivers in a single session of PST (n=38); (4) integration of an LLM into COCO. The Wizard-of-Oz method was used across phases 2 and 3 to collect naturalistic dialogues and refine COCO’s conversational design. In phase 3, usability of COCO was assessed using the System Usability Scale (SUS). Caregiver emotions were measured before and after the session using 6 subscales of the PANAS-X. In phase 4, GPT-4 was integrated into COCO with few-shot learning and evaluated by research team members using the caregiver personas. Descriptive statistics were used to summarize quantitative measures. The MI principles and techniques used by COCO across the 4 phases were coded using the . Results: In phase 1, 4 gold-standard dialogues were developed using caregiver personas. In phase 2, standardized patients described COCO as validating and identified its problem-solving and on-demand support as helpful for caregivers. In phase 3, COCO-Wizard-of-Oz achieved a mean SUS score of 75.6% (SD 12.9%), reflecting acceptable usability. Participants demonstrated significant improvement in negative affect, sadness, guilt, and fatigue following PST sessions (
  •  

Authors’ Reply: Clarifying the Comparative Interpretation and Clinical Implications of Radiomics-Based AI for Pathological Response Prediction

This author reply responds to a Letter to the Editor commenting on our systematic review and meta‑analysis evaluating radiomics‑based artificial intelligence for predicting pathological response following neoadjuvant immunochemotherapy in non‑small‑cell lung cancer. We clarify several methodological points raised in the comment, including patient versus assessment counts in a cited study, cross‑study versus within‑patient comparisons of diagnostic metrics, and the sensitivity‑specificity trade‑off between artificial‑intelligence models and conventional response criteria (RECIST 1.1, PERCIST). We acknowledge two textual errors in the original discussion and confirm they do not affect primary pooled analyses. We further elaborate on eligibility constraints, heterogeneity across prediction time points, definitions of pathological complete response, and reporting standards such as DECIDE‑AI. Our core conclusion remains unchanged: radiomics‑based artificial intelligence shows promising predictive performance with a potential sensitivity advantage over RECIST 1.1, though definitive evidence requires prospective same‑patient, same‑time‑point validation studies.
  •  

Same-Patient, Same–Time Point Evidence for Radiomics Benchmarking

This letter examines the interpretation and clinical translation of a meta-analysis of radiomics-based artificial intelligence (AI) for predicting pathological response after neoadjuvant immunochemotherapy in resectable non–small cell lung cancer. We highlight that the conventional-response comparator combines PERCIST and RECIST 1.1 assessments from the same 36-patient cohort, whereas the AI estimates arise from unmatched cohorts; consequently, the reported denominator and Z tests do not establish comparative superiority. We further consider how restriction to patients who reached resection and variation in imaging timepoints narrow the clinical estimand and limit inference about earlier treatment redirection. Clarification of the RECIST and error-direction examples is also warranted. We propose same-patient comparisons in treatment-initiation cohorts using fixed imaging times, locked thresholds, explicit primary-tumor and nodal labels, and calibration and net-benefit analyses at prespecified clinical thresholds.
  •  

User Experiences With a Social Robot for Cardiometabolic Risk Assessment in a Community Setting in Uppsala, Sweden: Qualitative Semistructured Interview Study

Background: Cardiometabolic diseases (CMDs) are a major health concern worldwide, with people living in socioeconomically disadvantaged areas being disproportionately affected. Knowing one’s risk status for developing a CMD can help individuals in delaying or preventing the disease. Innovative tools and strategies such as the use of AI-based technology are needed to improve inclusivity, cost-effectiveness, and sustainability of screening processes. Objective: This study aimed to assess users’ perceptions of the usability and acceptance of the social robot “Furhat” for cardiometabolic risk assessment within a socioeconomically disadvantaged community in Uppsala, Sweden. Methods: In this qualitative study, 13 semistructured interviews were conducted with participants living in socially disadvantaged areas within Uppsala, after they completed a CMD risk assessment delivered by the social robot. The study was conducted in one of the socioeconomically disadvantaged areas in Uppsala from October 19 to 26, 2023. Participants were purposefully sampled at different community events via phone or on the street to achieve the desired variation regarding personal characteristics such as age, gender, or area of living. The used interview guide for this study consisted of questions regarding the demographics, participants’ experiences with the robot screening, and suggestions for improvements. A framework analysis approach was applied to the data. Results: Two themes were developed. The first theme includes participants’ perceptions of who would use the robot and why, if it was installed in the community. Participants perceived older individuals, immigrants, and, in some instances, women to be less interested in the robot screening or as having a harder time participating in it. However, there were between-group variations, especially among men and women, with participants often discounting opinions ascribed to their group by others. The second theme captures participants’ actual experiences of the interaction with the social robot once they have started the interaction. Although participants enjoyed the interaction with the robot, perceived it as nonjudgmental, and would recommend it to others, most would prefer conducting the screening with a human. Two major contributors for this preference were language barriers experienced during the robot-participant interaction, as well as their wish for more interactive, emotionally responsive, and communicative features of the robot. Conclusions: This study provides insights into users’ perceptions of the usability and acceptance of a social robot for risk assessment in a community context, which can help inform the further development of this technology for health screening in a community context. In the future, the robot could be further developed to conduct risk assessments in multiple languages and offer a broader service to users, such as providing appropriate information related to healthy and active living.
  •  

AI and the Reproduction of Health Inequity: Redistribution-Translation-Accumulation Framework

AI is increasingly embedded in the institutions and environments that shape health. Yet current frameworks for understanding its implications for health equity remain underdeveloped. The social determinants of health tradition provides a strong foundation, and recent work on digital determinants of health has begun to address the health implications of digital transformation. However, AI warrants distinct conceptual attention because of its triple role: it operates simultaneously as a determinant of health in its own right, as a mediator and moderator of existing determinants, and as an amplifier of advantage and disadvantage over time. This viewpoint proposes a redistribution-translation-accumulation framework for analyzing how AI may contribute to the reproduction of health inequity. The framework comprises 2 analytically distinct mechanisms and 1 cross-cutting temporal dynamic. Redistribution captures how AI reshapes the distribution of health-relevant resources and opportunities, including education, employment, and income, while AI itself becomes an unequally distributed determinant. Translation describes how AI changes the pathways through which social positions are converted into health outcomes. Proxy-based decision rules can formalize historical inequities, diagnostic algorithms may perform unevenly across populations due to unrepresentative training data, and AI-mediated information environments can alter institutional responsiveness. Accumulation is conceptualized not as a third parallel mechanism but as a temporal amplifier operating on both mechanisms: AI-driven feedback loops and institutional embedding can concentrate advantage and disadvantage over time, often without users’ awareness. The framework is offered as a hypothesis-generating heuristic and is directionally neutral: under specifiable design, deployment, and governance conditions, the same mechanisms can narrow rather than widen health gaps in high-income and low- and middle-income settings alike. The framework has direct implications for governance. Current approaches such as the EU AI Act’s Fundamental Rights Impact Assessment (FRIA) and Canada’s Algorithmic Impact Assessment (AIA) advance AI accountability but assess systems largely before or at deployment and do not systematically track distributional health consequences. Building on this framework, I propose a distributional impact assessment as a complementary tool for equity-oriented AI governance. Structured around the 3 RTA dimensions, it asks whether an AI system alters the distribution of health-relevant resources across groups (redistribution), changes how social positions are converted into health (translation), and risks concentrating disadvantage over time through feedback and institutional embedding (accumulation). It is operationalized with candidate indicators, data sources, responsible actors, and reassessment triggers and is illustrated through a retrospective worked example of a biased care-management algorithm.
  •  

Effects of Virtual Reality on Pain, Anxiety, and Fear During Thyroid Fine-Needle Aspiration Biopsy: Open-Label Randomized Controlled Trial

Background: Thyroid fine-needle aspiration biopsy (FNAB) is a commonly used diagnostic procedure in patients with suspected thyroid cancer; however, it may induce pain, anxiety, and fear during the procedure. Objective: This open-label randomized controlled trial aimed to evaluate the effect of virtual reality (VR) on pain as the primary outcome and anxiety and fear of pain as secondary outcomes in patients undergoing thyroid FNAB. Methods: The study was conducted between January 19, 2025, and April 30, 2025, at Gaziantep City Hospital, Türkiye. A total of 100 patients with suspected thyroid nodules were randomly assigned to either a VR intervention group (n=50) or a control group (n=50). Data were collected using a patient information form, the visual analog scale (VAS), the Beck Anxiety Inventory (BAI), and the Fear of Pain Questionnaire-III (FPQ-III). Results: After adjustment for baseline pain, previous thyroid mass diagnosis, and voice tone changes, the VR group had statistically significantly lower postintervention pain scores than the control group (adjusted mean 3.627 vs 4.493; F1,95=4.021; P=.048; partial η2=0.041). However, the unadjusted between-group comparison for pain was not statistically significant (P=.12), and the unadjusted effect size was small, with a 95% CI that crossed 0 (Cohen d=−0.33, 95% CI −0.72 to 0.07). No statistically significant adjusted between-group differences were observed for anxiety (P=.48) or fear of pain (P=.07). Unadjusted standardized between-group effect sizes were also small for anxiety (d=−0.17) and fear of pain (d=−0.11). Conclusions: The adjusted analysis suggested a small reduction in procedural pain with VR; however, the between-group difference was not statistically significant in the unadjusted analysis and reached statistical significance only after adjustment for baseline pain and 2 nonprespecified covariates selected on the basis of observed baseline imbalance. Moreover, the observed adjusted effect (f=0.207) was smaller than the minimum effect size the trial was powered to detect (f=0.283). No statistically significant adjusted between-group effects were found for anxiety or fear of pain. Therefore, the potential analgesic effect of VR should be interpreted cautiously and confirmed in larger, adequately powered trials. Trial Registration: ClinicalTrials.gov NCT06792929; https://clinicaltrials.gov/study/NCT06792929
  •  

Stage-Based Model of User Engagement Patterns in an Online Health Community for Cardiovascular Disease Management: Qualitative Interview Study

Background: Online health communities (OHCs) provide vital peer support and health information to individuals managing chronic conditions. However, sustained user engagement remains challenging, with many users reducing activity over time despite the ongoing benefits these platforms offer. While some individuals may improve or become more knowledgeable, sustained engagement remains critical because ongoing participation fosters trust, peer support, and continuous access to evolving health information that may persist beyond initial recovery or learning. Hence, it is important to consider how community needs evolve as users’ health needs and information-seeking behaviors change to inform how we support users through different stages of their health and participation journeys. Objective: The aim of the study is to examine user engagement in a large OHC, identifying perceived stage-based behaviors, motivation, barriers, and design opportunities that could facilitate progression between stages and enhance long-term participation. Methods: We conducted semistructured interviews with 19 members of the American Heart Association Support Network Community. Participants were patients or survivors managing various cardiovascular diseases. Using narrative thematic analysis, we examined users’ perceived engagement motivation, behavior, challenges, and design opportunities across different stages of their community involvement. Results: This study highlighted 4 distinct engagement stages: discovery (crisis-driven initial engagement), exploration (navigation and orientation), commitment (active engagement and information management), and integration (sustained engagement and mentorship). Key barriers included information architecture complexity, concerns about misinformation, limited support for role transitions, and decreased participation as health management improved. Participants identified opportunities through which OHCs could increase long-term engagement, including adaptive recommendation systems, health information literacy programs, structured role transition support, and alternative engagement modalities, such as synchronous interactions and health tracking tools. Conclusions: User engagement in OHCs is dynamic and evolves with changes in health status, knowledge, and personal circumstances. Supporting sustained engagement requires stage-appropriate interventions, including personalized content delivery, health information literacy education, structured pathways for role transitions, and diversified engagement options. These findings provide actionable insights for designing OHCs that better support users throughout their health journey.
  •  

Depictions of Depression in Generative AI Video Models: Mixed Methods Study of OpenAI’s Sora 2

Background: Generative AI video models are increasingly capable of producing complex depictions of mental health experiences, yet little is known about how these systems represent conditions such as depression. Because AI-generated content may reach people during vulnerable periods, understanding what visual narratives these models produce for sensitive concepts carries clinical relevance. Objective: This study aimed to characterize how OpenAI’s Sora 2 generative AI video model depicts depression and examine whether depictions differ between the consumer app and developer API access points, which differ in their product layer mediation. Methods: We generated 100 videos using the single-word prompt “Depression” across 2 access points: the consumer app (n=50, 50%) and developer API (n=50, 50%). Two trained coders independently coded narrative structure, visual environments, objects, figure demographics, and figure states. Interrater reliability was assessed using the Cohen κ, with dimensions showing insufficient agreement excluded from analysis. Computational features (visual aesthetics, audio, semantic content, and temporal dynamics) were extracted and compared between modalities using 2-tailed Welch tests with Benjamini-Hochberg false discovery rate correction. Results: App-generated videos exhibited a pronounced recovery bias: 78% (39/50) featured narrative arcs progressing from depressive states toward resolution compared with 14% (7/50) of API outputs. This divergence was reinforced across channels. App videos brightened over time (mean slope 2.90, SD 2.43 per second vs −0.18, SD 1.24 per second for the API; Cohen =1.59;
  •  

Design Guidelines for Online Health Forums: User-Centered Design Approach

Background: Online health forums are used widely, yet evidence of their effectiveness is inconsistent. Evidence-based forum design guidance grounded in theory and lived experience could improve the efficacy and outcomes of these forums for the many people using them worldwide. Objective: This study aimed to draw on the experience of online forum users and staff, and insights from existing research on technology design and self-determination theory, to generate a set of theoretically grounded guidelines for safe and well-being–supportive online forums. Methods: We conducted 54 semistructured interviews (36 forum users and 18 forum staff) and 4 design workshops with forum staff, combined with input from a multidisciplinary research team. Principles of qualitative framework analysis were used to adapt a preexisting framework for well-being–supportive technology to the context of online health forums. Results: The resulting design guidelines are framed around 4 overarching principles relating to the psychological needs for autonomy, competence, and relatedness as defined by self-determination theory, and the additional need for safety in online forums. Each principle is presented alongside pragmatic design heuristics and specific implementation strategies. Conclusions: User experiences of online health forums are mixed. We have drawn on self-determination theory to propose evidence-informed guidelines for service development and refinement, adaptable to specific user groups across diverse settings. International Registered Report Identifier (IRRID): RR2-https://doi.org/10.1136/bmjopen-2023-075142
  •  

Anonymization of Portuguese Clinical Notes Using Large Language Models and Quantum-Enhanced Hybrid Architectures: Comparative Evaluation Study

Background: The widespread adoption of electronic health records (EHRs) has generated large-scale repositories of highly sensitive clinical information, emphasizing the need for robust anonymization strategies to enable secondary use for research while safeguarding patient privacy. Conventional rule-based and machine learning approaches for deidentifying medical text face limitations with the linguistic complexity, variability, and context dependence inherent to clinical documentation. Recent advances in large language models (LLMs), combined with emerging quantum computing paradigms, present novel opportunities to enhance the accuracy, scalability, and resilience of health care data anonymization. Objective: This study aims to evaluate the efficacy of LLM-based and quantum-enhanced hybrid architectures for medical text anonymization, assessing the effectiveness and computational efficiency across multiple entity types in Portuguese clinical notes. Methods: We constructed a gold-standard corpus of 1000 Portuguese outpatient clinical notes, manually annotated by 5 trained researchers for 5 protected-entity categories: patient names, dates, identifiers, organizations, and geographic locations. Four anonymization strategies were evaluated: 2 stand-alone LLMs (Llama-3.1-8B-instruct and Llama-3.3-70B-instruct) and 2 quantum-enhanced hybrid models (Dynex-QML with 8B and 70B base models) incorporating quantum optimization via Quadratic Unconstrained Binary Optimization (QUBO) formulations. The quantum-enhanced approach transforms the final attention layer of the LLM into a global constraint satisfaction problem solved via neuromorphic quantum annealing. Model performance was measured on a held-out test set of 500 notes using precision, recall, and -score metrics. Computational efficiency was quantified through end-to-end processing time. Results: The quantum-enhanced Dynex-QML-70B model achieved the highest overall performance with a macro-score of 0.855 (95% CI 0.823‐0.880), outperforming the stand-alone Llama-3.3-70B (0.726, 95% CI 0.704‐0.747), Dynex-QML-8B (0.733, 95% CI 0.709‐0.756), and Llama-3.1-8B (0.602, 95% CI 0.588‐0.615). Compared with Llama 3.3 70B, Dynex-QML (Llama 70B) improved macro-score by 0.128 (95% CI 0.091‐0.163; empirical 2-sided bootstrap
  •  

Digital Phenotyping of Lifestyle Profiles and Mental Well-Being in German Adults: Prospective Longitudinal Cohort Study

Background: Digital phenotyping uses passively collected smartphone-sensing data to characterize everyday behavior in naturalistic settings, and has become an important approach for studying mental well-being. Most previous studies have examined associations between individual sensing variables and mental health. However, mental well-being is likely reflected not by isolated behaviors but by combinations of co-occurring daily behaviors that together form lifestyles. Person-centered approaches capable of identifying these behavioral configurations may, therefore, provide more interpretable digital phenotypes; yet, such approaches have rarely been applied to passive smartphone-sensing data. Objective: This study aimed to examine whether smartphone-captured behavioral and environmental data could be used to derive interpretable day-level and person-level lifestyle profiles, and whether person-level profiles were associated with mental well-being. We also tested whether Big Five personality traits—extraversion, agreeableness, conscientiousness, openness, and negative emotionality—moderated these associations. Methods: The study used a 2-week prospective longitudinal cohort design with a sample of 553 German adults (mean age 42.12, SD 12.89 years; 44.65% female) drawn from an initial sample recruited according to quotas designed to reflect the German population. Ten smartphone-sensing indicators captured 5 domains, including communication and social media app use, mobility, physical activity, environmental context, and phone-use intensity. Mental well-being was assessed using the Warwick–Edinburgh Mental Well-Being Scale, and personality was assessed using the 15-item Big Five Inventory–2 Extra-Short Form. We used multilevel latent profile analysis to identify day-level profiles nested within person-level profiles. Associations between profiles and mental well-being were tested using classification-error–adjusted mean comparisons and omnibus Wald tests. Moderation was examined using hierarchical regressions comparing models with and without profile-by-personality interactions. Results: Eight day-level profiles and 7 person-level profiles were identified. Day-level profiles reflected distinct combinations of smartphone-sensing indicators. Person-level profiles represented different distributions of these daily patterns. Profiles differed significantly only in positive functioning (Wald ²=13.39; =.04), not in overall mental well-being, positive affect, or satisfying interpersonal relationships. The physically active and unplugged profile had higher positive functioning than the mobile and always-on social profile (mean 3.94, SD 0.63 vs mean 3.61, SD 0.74; Cohen =0.47; 95% CI 0.21‐0.73). No other pairwise differences were significant. Sensitivity analyses excluding the smallest profile produced comparable results, supporting the robustness of the findings. Personality-by-profile interactions did not significantly improve prediction for any well-being outcome. Conclusions: The findings extend the field by showing that transparent, person-centered digital phenotypes can distinguish variation in positive functioning, although causal conclusions cannot be drawn. In real-world settings, such interpretable profiles could support understandable monitoring tools and, following prospective replication and validation, inform personalized multibehavior interventions that target combinations of behaviors rather than single behaviors in isolation.
  •  

Evaluation of the Square Eyes Model as a Screening Tool for Identifying Digital Technologies in Wearable Camera Images Among Children: Laboratory Study

Background: Accurate measurements of children’s digital technology use are essential for understanding its potential implications on health and well-being. Wearable cameras can provide such measurements, but image coding is a high burden for researchers. Machine learning–based object-recognition models have the potential to reduce this burden by identifying images containing technology. Objective: This study aims to evaluate the performance of an object recognition model, the Square Eyes model, as a screening tool for identifying technologies in wearable camera images among children for further human review, as well as to examine the potential influence of face-blurring methods on the model’s performance. Methods: This study used data collected on 48 children (aged 3‐14 y) during an approximately 1-hour laboratory session. The children performed various technology-related tasks while wearing a camera. A total of 221,226 images were coded by humans and processed through the Square Eyes model. The performance of the Square Eyes model as a screening tool was evaluated by (1) assessing agreement between the model and human coding; (2) evaluating the N-back algorithm, an algorithm embedded in the model aimed to flag images requiring human review; and (3) examining the potential influence of facial-blurring on model performance. Results: Humans detected technology in 92,745 (41.9%) images, and the Square Eyes model detected technologies with an overall accuracy of 78.0%. When considering specific technologies, agreement between the model and human coders was the highest for (n=19,148, 54.3%) and (n=8492, 44.5%) and lowest for smaller devices such as (n=2600, 31.3%) and (n=3685, 25.1%). The model’s N-back algorithm effectively flagged images that required further human review, with only 7144 (3.2%) images that were not flagged for screening containing a human-coded technology. An explorative analysis indicated that using a square face-blurring with border could have reduced the model’s ability to accurately detect technologies. Conclusions: The Square Eyes model demonstrated overall satisfying accuracy in detecting technologies and successfully flagged images that required further review by humans. These findings suggest that the model could be used as an effective screening tool for reducing the burden of human coding. However, the model could be improved to more accurately detect smaller devices, and the form of facial blurring in images should be considered.
  •  

Digitally Adapting LGBTQ-Affirmative Cognitive Behavioral Therapy for Chinese Men Who Have Sex With Men Living With HIV: User-Centered Design Approach

Background: Chinese men who have sex with men living with HIV (MSMLWH) experience substantial psychological distress driven by minority stress and HIV-related challenges. However, culturally tailored digital mental health interventions that address HIV-specific maladaptive cognitive schemas and culturally specific psychosocial stressors remain scarce in China. Objective: This study aimed to systematically adapt an evidence-based cognitive behavioral therapy (CBT) intervention Effective Skills to Empower Effective Men (ESTEEM) into a WeChat (Tencent) Mini-Program–based intervention (iESTEEM) specifically for Chinese MSMLWH and to evaluate its preliminary feasibility and usability. Methods: We used a three-phase user-centered design approach guided by the Assessment, Decision, Adaptation, Production, Topical Experts, Integration, Training, and Testing (ADAPT-ITT) framework. The study proceeded in three phases: (1) a qualitative needs assessment using semistructured interviews with 20 MSMLWH (mean age 23.25, SD 3.08 years); (2) systematic intervention adaptation and platform development, including theater testing (n=5); and (3) a 2-week pilot study involving 10 MSMLWH and five counselors to evaluate feasibility, usability, and acceptability through focus groups and objective platform analytics. Results: Phase 1 identified 3 major themes of psychological distress: persistent health anxiety fueled by catastrophizing, intersectional stigma internalization, the disclosure dilemma, and intimacy barriers rooted in defectiveness and shame schemas. Participants also prioritized anonymity and bite-sized learning. Guided by these findings, iESTEEM was developed as a counselor-assisted, privacy-preserving WeChat Mini-Program incorporating HIV-specific scenarios, multimodal learning modules, and a back-end risk-alert system. During the 2-week pilot, participants logged into the platform 14.1 (SD 6.7) times per person and completed 134.3 (SD 103.1) minutes of learning activities; all participants accessed module 1, and 90% (9/10) accessed modules 2‐5. Anxiety scores decreased from 8.9 (SD 2.3) to 7.2 (SD 3.0), whereas depression scores remained stable. All participants expressed a willingness to continue using the program and to recommend it to peers. Participants and counselors endorsed its contextual relevance, privacy protections, and clinical utility. Conclusions: This study provides a theory- and evidence-informed model for culturally adapting digital mental health interventions for highly stigmatized populations. By integrating lesbian, gay, bisexual, transgender, and queer (LGBTQ)-affirmative CBT principles, HIV-specific adaptations, and a privacy-preserving, counselor-assisted WeChat Mini-Program, iESTEEM demonstrated promising preliminary feasibility, acceptability, and engagement among Chinese MSMLWH. These findings support the potential of culturally tailored digital interventions to expand access to psychological support for this stigmatized population in resource-constrained settings. Ongoing randomized controlled trials will further evaluate its efficacy, implementation outcomes, and mechanism of action. Trial Registration: Chinese Clinical Trial Registry ChiCTR2400080263; https://www.chictr.org.cn/showproj.html?proj=216926
  •  

Effectiveness of Wearable Digital Therapeutics in Improving Sleep Outcomes Among Individuals With Insomnia: Systematic Review and Meta-Analysis of Randomized Controlled Trials

Background: Wearable devices are increasingly used for sleep monitoring and as adjunctive treatment. Existing meta-analyses mostly pool composite digital therapies and rarely isolate stand-alone wearables or distinguish between objective and subjective end points. Whether stand-alone wearable interventions improve sleep outcomes in adults with insomnia, and which factors moderate treatment heterogeneity, remains unclear. Objective: This study aims to evaluate the effectiveness of wearable digital interventions on sleep outcomes in adults with insomnia versus control strategies and explore moderators of effectiveness, including device-wearing position, intervention duration, and control type, using meta-regression. Methods: This systematic review and meta-analysis was conducted in accordance with the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta‑Analyses) 2020 statement and the PRISMA-S (Preferred Reporting Items for Systematic Reviews and Meta‑Analyses Literature Search Extension) guideline. Five electronic databases and clinical trial registries were searched from inception to May 18, 2026. Eligible studies were randomized controlled trials (RCTs) evaluating wearable digital interventions in adults with insomnia compared with sham, waitlist, usual care, or active control conditions and had an intervention duration of at least 1 week. Study screening, data extraction, and risk-of-bias assessment were carried out independently by 2 reviewers. Pooled estimates were calculated using a restricted maximum likelihood random-effects model with the Hartung-Knapp-Sidik-Jonkman correction. Heterogeneity was assessed using the ² statistic, and 95% prediction intervals (PIs) were calculated for the primary analyses. The certainty of evidence was rated using the GRADE (Grading of Recommendations, Assessment, Development, and Evaluation) approach. Results: Sixteen RCTs (N=910) were included. Wearable digital interventions were associated with a significant reduction in objective sleep-onset latency (SOL; mean difference [MD] −4.52, 95% CI −8.38 to −0.67, PI −9.52 to 0.47 min) and a significant improvement in subjective sleep efficiency (SE; MD 2.00%, 95% CI 1.90%‐2.11%, PI 1.85%‐2.15%). Subjective total sleep time (TST) also showed a significant increase (MD 19.11, 95% CI 2.98‐35.24, PI −16.20 to 54.43 minutes). Meta-regression showed that control type, intervention duration, and device location did not explain the heterogeneity of the insomnia severity index (ISI) (=0). Sensitivity analysis confirmed the robustness of pooled ISI estimates, and an Egger test indicated no small-study effects (=.07). Certainty of evidence ranged from moderate to high. Conclusions: Wearable digital interventions provide selective benefits for objective SOL, subjective SE, and subjective TST in adults with insomnia, with no improvement in overall ISI. Despite statistically significant effects on several sleep parameters, wide PIs, substantial heterogeneity, and limited study numbers indicate preliminary, nonconclusive findings. Wearables should be viewed as affordable adjunctive tools requiring further validation, not substitutes for first-line cognitive behavioral therapy for insomnia. Large-scale, long-term RCTs with standardized protocols and patient-level external validation are required to consolidate the evidence base. Trial Registration: PROSPERO CRD420251038603; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251038603
  •  

Correction: From Metrics to Meaning in Neurological Rehabilitation: Clinicians’ Perspectives on Digital Metrics of Upper Limb Functioning—A Focus Group Study

Digital assessment technologies, such as optical motion capture and inertial measurement units, enable detailed kinematic analysis and continuous monitoring of upper limb activity in persons with neurological conditions. While such digital metrics of functioning are increasingly recognized in research, their uptake in clinical neurorehabilitation is limited. It remains unclear which digital metrics of functioning clinicians perceive as most meaningful and how these are integrated into patient-centered care. Understanding clinicians’ information needs and reasoning processes is a prerequisite for implementing digital assessment technology. To characterize how rehabilitation professionals perceive, prioritize, and integrate digital metrics of functioning into clinical reasoning and to identify features that would support their routine use. Three 90-minute focus groups were conducted in 3 Swiss neurorehabilitation centers, involving 11 clinicians with diverse professional backgrounds (5 physiotherapists, 4 occupational therapists, 1 movement scientist, and 1 medical practitioner). Participants discussed essential parameter domains and individually rated the relevance and meaningfulness of 17 kinematic metrics for the well-studied drinking task and 10 established arm use performance metrics. Verbatim transcripts were analyzed using reflexive thematic analysis, and rating data were summarized descriptively. Five main themes were identified. (1) Functional requirements to interpret movement quality and performance (active/passive range of motion (ROM), strength, selective muscle control, grasp) form the basis for interpreting movement. (2) Essential aspects of movement quality (smoothness, efficiency, compensatory movement) are valued when aligned with observable task execution. (3) Added value of real-world performance (hourly activity profiles, arm-use symmetry, functional workspace) represents the reference for patient-centered reasoning. (4) Individualizing what matters, including diagnosis-specific preferences, shapes assessment selection. (5) Blending clinical eye and reference data reflects clinicians’ reliance on visual judgment complemented by normative values. Intuitive metrics such as task duration, number of movement units, and ROM were favored, whereas confidence was lower in more complex metrics (e.g., jerk, inter-joint coordination). Clinicians value intuitive digital metrics of functioning when they are clearly linked to patient-centered outcomes and supported by normative references. The findings highlight the need for targeted educational strategies and digital competency training that help clinicians interpret digital metrics and integrate them with contextual information and clinical reasoning.
  •  

Evaluating Large Language Models in Clinical Audiology (AUDIOLOGYBENCH): Benchmark Development and Validation Study

Background: Large language models (LLMs) are increasingly being explored for clinical decision support, but their performance in audiology has not been systematically benchmarked using clinically grounded case materials and rubric-based safety evaluations. Objective: This study aimed to develop and evaluate AUDIOLOGYBENCH, a 3-tier benchmark for characterizing frontier LLM capability in clinical audiology along (1) curated domain knowledge, (2) literature-derived evidence, and (3) clinical reasoning under multimodal case input, with an explicit human audit of the automated adjudicator on the primary end point. Methods: The benchmark comprises 3139 objective items from educational resources, 3175 research article–derived items from peer-reviewed articles published between 2015 and 2025, and 67 multimodal clinical case studies graded against a standardized A-F rubric with 6 prespecified critical-error types that cap scores at D or F. Eight models were evaluated on the educational objective items: 4 frontier multimodal models (Gemini 2.5 Pro, Grok 4, OpenAI O3, and Claude Sonnet 4 Thinking) were evaluated on the research article–derived items, and on 804 case study evaluations. Adjudication used Gemini 2.5 Pro (objective and research-derived items) and Claude Opus 4.5 (case studies). The case study adjudicator was independently audited against PhD-level audiologist consensus on blinded subsamples, supplemented by a post-stratified human-calibrated sensitivity analysis. Results: A striking task-type dissociation emerged on case studies: clinical recommendations (Q3) achieved a mean score of 89.74 (SD 13.92, 95% CI 88.07‐91.41), a 98.1% (263/268) pass rate, and no dangerous recommendations; audiometric numerical interpretation (Q1) achieved a mean score of 67.89 (SD 18.47, 95% CI 65.68‐70.10), with a 35.4% (95/268) critical-error rate; and differential diagnosis (Q2) achieved a mean score of 67.79 (SD 15.33, 95% CI 65.95‐69.63). Question type, not model selection, dominated performance (eta-squared_H=0.333 vs 0.001; rank biserial ≥0.679). Interreviewer reliability between audiologists was high (quadratic-weighted κ of 0.78 and 0.85 across the 80-item and 50-item audits, respectively). When 2 audiologists regraded all 80 model Q1 responses with the diagnostic images available, the adjudicator’s per-item Q1 labels diverged from human judgment (κ=0.05; overflagging; sensitivity: 19/26, 73%; positive predictive value: 19/53, 36%), yet its reweighted Q1 critical-error rate (36.2%) was broadly consistent with the image-grounded human estimates (28%‐34%), suggesting no systematic inflation of the headline rate. The principal Q3>{Q1, Q2} ranking was preserved under post-stratified human calibration. Web-style multiple-choice items showed ceiling effects (>95% accuracy); short-answer prompts remained challenging (best 30%). Conclusions: Current frontier LLMs show strong recommendation generation but substantial limitations in audiometric numerical interpretation that are shared across models and that an automated adjudicator partially miscalibrated at the per-item level. AUDIOLOGYBENCH characterizes capability boundaries rather than certifying clinical readiness. Deployment of LLM-assisted audiology workflows requires structured human verification of all numerical findings and awareness of fabrication and severity misclassification failure modes documented here.
  •  

The Effectiveness of Digital Intervention on Psychological Resilience in Postoperative Breast Cancer Patients During Chemotherapy Intervals: Quasi-Experimental Study

Background: Patients with breast cancer during postoperative chemotherapy intervals commonly experience psychological distress and reduced resilience while recovering at home. Digital mindfulness interventions may provide accessible psychological support during this vulnerable period; however, evidence regarding tailored interventions for postoperative patients with breast cancer during chemotherapy intervals remains limited. Objective: This study aimed to examine the effectiveness of a digital intervention on psychological resilience in postoperative patients with breast cancer during chemotherapy intervals. Methods: A quasi-experimental study with repeated measures was conducted from October 2021 to June 2022. A total of 80 eligible participants were recruited from the Department of Breast Surgery at a tertiary hospital in Zhejiang Province, China, and 71 completed the study. The control group received routine discharge instructions and nursing follow-ups, whereas the intervention group additionally received an 8-week digital psychological resilience intervention. Outcomes were assessed at baseline (T0), 3 months post intervention (T1), and 6 months post intervention (T2). The measures included the Connor-Davidson Resilience Scale (CD-RISC), Hospital Anxiety and Depression Scale (HADS), Social Support Rating Scale (SSRS), Breast Cancer Survivor Self-Efficacy Scale (BCSSS), and Functional Assessment of Cancer Therapy-Breast (FACT-B). Independent-samples tests, chi-square tests, and repeated-measures ANOVA were performed using SPSS (version 26.0; IBM Corp). Results: No statistically significant baseline differences were observed between the two groups in the outcome measures. At T1, the intervention group had higher CD-RISC scores than the control group (mean 67.58, SD 11.41 vs mean 62.09, SD 10.18; =.036) and higher BCSSS scores (mean 42.36, SD 3.59 vs mean 39.23, SD 4.90; =.003). However, these between-group differences were no longer statistically significant at T2 (>.05). Significant time effects and group×time interaction effects were observed for both psychological resilience and self-efficacy (.05), although both scales showed significant time effects (
  •  
❌