❌

Reading view

Beyond GPT-4: The Rapidly Evolving Potential of Large Language Models for Clinical Guideline Improvement

This commentary reviews the study by Jones et al, which evaluated whether GPT-4 could improve the readability of injectable medication guidelines while preserving important safety information. The study found that GPT-4 produced modest readability gains comparable to manual revision, but also introduced omissions and meaning changes in a minority of sections. These findings highlight both the potential and limitations of early large language models (LLMs) in clinical contexts. However, this study reflects the capabilities of a specific model in a rapidly evolving domain. Since the release of GPT-4, advances in multistep reasoning, model-critique workflows, and structured validation have substantially improved the ability of newer systems to detect omissions, maintain factual fidelity, and support controlled editing. As a result, some documented limitations may stem from the constraints of a single-model, single-pass workflow rather than intrinsic flaws in LLM-assisted guideline revision. This commentary highlights the need for evaluation frameworks that can keep pace with LLM progress and emphasizes that clinical oversight and user-centered testing remain essential. Updated research using contemporary models is needed to determine how emerging architectures can more safely support clarity, consistency, and maintenance of clinical guidelines.
  •  

Large Language Model–Based Analysis of Statin Therapy Discussions and Sentiment on Social Media: Cross-Sectional Observational Study

Background: Statin therapy, despite proven cardiovascular benefits, remains underused. Social media platforms may capture patient perspectives that are less visible in clinical encounters. Objective: This study aimed to characterize themes, sentiment, and decision-making factors related to statin therapy through large language model (LLM)–based analysis of Reddit discussions. Methods: This cross-sectional observational study analyzed English-language Reddit posts and comments mentioning statins from January 2022 to May 2025, identified via keyword-based Reddit application programming interface searches (≤1000 posts per keyword). A total of 5328 retrieved discussions (n=1661, 31.2% posts and n=3667, 68.8% keyword-containing comments) from public subreddits were included. Themes, sentiments (positive, neutral, or negative), guideline-informed clinical relevance, information-seeking behavior, adverse effect mentions, decision factors, and adherence-related content were extracted using an LLM-based pipeline. Results: Among 5328 discussions, prominent topics included adverse effects (n=1697, 31.9%), decision-making references related to laboratory results and physician advice (n=2767, 51.9% and n=2034, 38.2%, respectively), and alternative approaches (n=2485, 46.6%). Overall sentiment was neutral in 34% (n=1812) of discussions, negative in 30.9% (n=1646), and positive in 16.9% (n=900); the remainder were mixed or unclear. Statin-directed sentiment was neutral in 44.1% (n=2350) of discussions, negative in 25.2% (n=1343), and positive in 12.5% (n=666); the remainder did not express statin-directed sentiment. High clinical relevance was identified in 12.6% (n=672) of discussions. Adherence-related issues were mentioned in 29.8% (n=1587) of discussions. Among adverse effect mentions, muscle pain (n=129, 7.6%) and fatigue (n=110, 6.5%) were common. Conclusions: LLM-enabled analysis of Reddit discourse highlights substantial negative sentiment, adherence-related concerns, and adverse effect narratives surrounding statin therapy. These findings suggest opportunities for patient-centered communication and shared decision-making strategies that address symptom attribution, uncertainty, and information needs in digital information environments.
  •  

Health Equity Analysis of Awareness and Use of GetCheckedOnline, British Columbia’s Digital Intervention for Sexually Transmitted and Blood-Borne Infection Testing in 5 Urban, Suburban, and Rural Communities: Cross-Sectional Survey Study

Background: Digital sexually transmitted and blood-borne infection (STBBIs) testing services are used to improve testing access, but might replicate existing social inequities. Previous research has shown that the digital STBBI testing service has improved access to testing in British Columbia (BC), Canada. As part of the program’s continuous evaluation, we examined awareness and use of the service in 5 urban, suburban, and rural communities where the program has expanded. Objective: This study aimed to determine if social location is associated with differences in awareness and use of the service in 5 communities outside Vancouver, BC. Methods: From July to September 2022, we conducted a cross-sectional survey recruiting (in-person and online) sexually active people aged 16 years or older in 5 urban, suburban, and rural communities where had sample collection sites available at the time. We examined differences in awareness and use by age, gender identity, sexual identity, race/ethnicity, education, and income using logistic regression models informed by the Health Equity Measurement Framework. Results: Of the 1658 participants (n=1058, 63.8% in-person and n=600, 36.2% online), 35.3% (586/1658) were aware of and 19.5% (324/1658) had used it. Awareness and use were lower in the first and last age quartiles compared to the second quartile (>38 years: awareness odds ratio [OR] 0.23, 95% CI 0.17‐0.32; use OR 0.19, 95% CI 0.12‐0.28;
  •  

Context-Aware Sentence Classification of Radiology Reports Using Synthetic Data: Development and Validation Study

Background: Automated structuring of radiology reports is essential for data utilization and the development of medical artificial intelligence models. However, manual annotation by experts is labor-intensive, and processing real clinical data through commercial large language models (LLMs) presents significant privacy risks. These challenges are particularly pronounced for non-English languages like Japanese, where specialized medical corpora are scarce. While synthetic data generation offers a potential privacy-preserving alternative, its effectiveness in capturing complex clinical nuances—such as negation and contextual dependencies—to train robust classification models without any real-world training data has not been fully established. Objective: This study aimed to develop a context-aware sentence classification model for Japanese radiology reports using an entirely synthetic training pipeline, thereby eliminating reliance on real-world clinical data during the development phase. Furthermore, we sought to evaluate the generalizability of this approach by validating the model’s performance on diverse, multi-institutional, real-world reports. Methods: Japanese radiology reports (n=3104) were generated using GPT-4.1 and automatically annotated at the sentence level into 4 categories (background, positive finding, negative finding, and continuation) using GPT-4.1-mini. The synthetic data were partitioned into training (n=2670), validation (n=334), and test (n=100) sets. We fine-tuned several models, including lightweight local LLMs (Qwen3 and Llama 3.2 series) using low-rank adaptation and Japanese text classification models (Bidirectional Encoder Representations from Transformers [BERT]-base Japanese v3, Japanese Medical Robustly Optimized BERT Pretraining Approach [JMedRoBERTa]-base, and ModernBERT-Ja-130M). External validation was performed using 280 real-world reports (3477 sentences) from 7 institutions in the Japan Medical Image Database, with ground-truth labels established by board-certified radiologists. Evaluation metrics included accuracy, macro-averaged (macro ) score, and positive predictive value for positive findings (PPV_1). Results: All models achieved high performance on the synthetic test set (accuracy: 0.938‐0.951; macro -score: 0.924‐0.940). Overall performance declined on the external validation dataset (accuracy: 0.783‐0.813; macro -score: 0.761‐0.790), reflecting distributional differences between synthetic and real-world reports; however, PPV_1 remained stable and high across datasets (eg, 0.957 on the synthetic test set vs 0.952 on the external validation dataset for Qwen3 [4B]). Parsing errors occurred in LLM-based approaches (19‐260 sentences, 0.55%‐7.48% in the external dataset). Conclusions: This study demonstrates the feasibility of developing context-aware sentence classification models for Japanese radiology reports using a training pipeline based entirely on synthetic data. The stability of PPV_1 indicates that the models successfully captured the essential clinical terminology and linguistic patterns required to identify positive findings in real-world reports, despite the observed performance degradation during external validation. This approach substantially reduces manual annotation requirements and privacy risks, providing a scalable foundation for constructing structured radiology datasets to support the development of clinically relevant medical artificial intelligence models.
  •  

Views of People With Psychosis About Algorithm-Based Relapse Prediction and Data Sharing: Qualitative Study

Background: Preventing relapses of psychosis is difficult and important. Digital remote monitoring (DRM) systems are being developed and tested to support this. Increasingly, these systems use algorithm-based relapse prediction. Hence, understanding stakeholder views about algorithmic prediction is crucial. Existing qualitative work has explored health professionals’ views, but very few studies have examined the perspectives of people with psychosis on this topic. Objective: This paper aimed to provide an in-depth examination of the views of people with psychosis regarding algorithmic relapse prediction within a DRM system that incorporates active symptom monitoring and passive sensing data. Methods: People with psychosis (n=58) were recruited from 6 geographically distinct areas of the United Kingdom. They participated in semistructured qualitative interviews exploring their views about using a DRM system that predicts psychosis relapse based on a machine learning algorithm. Transcripts were analyzed using reflexive thematic analysis. People with lived experience of psychosis were involved extensively in study design, analysis, and reporting. Results: Findings were described across 4 themes. First, was a prominent theme. Participants emphasized that transparency about algorithm sensitivity and specificity is crucial and discussed the risks of the relapse prediction algorithm producing false positives (flagging that someone was relapsing when they were not) and false negatives (missing actual relapses). In both cases, participants said that errors may be partially mitigated through a approach (theme 2), with DRM blended with human oversight, from clinicians or a dedicated digital monitoring team, and calibrated based on service user, carer, and clinician feedback. The third theme, noted the interplay between users’ trust in the DRM system and their relationship with the clinical team. This theme described participants’ fears about potential overreactions (hospitalization or excessive medication) or underreactions (no additional support) from the clinical team in response to algorithm-generated relapse predictions. It emphasized the importance of retaining choice around the use of relapse detection algorithms and the sharing of personal data. The final theme described participants’ views about the , including facilitating early intervention, triaging care according to need, minimizing human bias in assessment, and efficiency in saving staff time. Conclusions: People with psychosis acknowledged potential benefits of algorithm-assisted relapse prediction for receiving timely or efficient care, but with several caveats. Algorithm-generated relapse alerts need to be sufficiently accurate and must be interpreted, with understanding of their limitations, by a trustworthy human who is aware of the relevant context. Algorithm-based relapse predictions should only be used with valid consent, in a way that promotes and respects the autonomy and voice of service users and avoids increasing the use of excessive restriction.
  •  

Child Vaccination Status and Behavioral and Social Drivers of Vaccination Among Their Caregivers in the Philippines: Cross-Sectional Survey Study Comparison of Household, Mobile, and Online Modes

Background: The World Health Organization recommends that countries routinely collect data on the behavioral and social drivers (BeSD) of vaccination to inform public health interventions that increase vaccine uptake. There is a need to identify data collection methods that can rapidly and inexpensively collect representative data, particularly in low- and middle-income countries. Objective: This study aimed to understand BeSD drivers of vaccination in the Philippines and assess the trade-offs between survey methods. We compared responses to household, mobile, and online surveys in terms of demographics, vaccination status, responses to BeSD questions, and cost. Methods: We conducted concurrent household, mobile (SMS text messaging and interactive voice response), and online surveys among caregivers of children 2 years of age and below in Regions V and XII of the Philippines, with sampling differing by survey method. We assessed, for each survey method, (1) respondent demographics (sex, age, region, and socioeconomic status) and (2) the weighted proportion of responses from caregivers of children who received at least one dose of diphtheria-pertussis-tetanus (DPT)–containing vaccine. We estimated the weighted proportion of each BeSD survey response option and calculated the financial cost (monetary outlays) per survey response from an implementer’s perspective by summing the costs incurred in each survey method and dividing by the number of responses received. Results: We surveyed a total of 1201 household respondents, 2153 mobile respondents, and 398 online respondents from January to March 2025. We found that online and mobile survey respondents were more likely to be male and have completed high school than household survey respondents. The weighted proportion of respondents indicating that their child had received at least one dose of DPT vaccine was 91.8% (n=1090; 95% CI 90%‐93.3%) for the household survey, 90.3% (n=1853) for the mobile survey, and 85% (n=346) for the online survey. With regard to vaccine demand, more than 85% of respondents in each survey method indicated that vaccines are very important, very safe, supported by family, and that they knew where to bring a child for vaccination. More than 30% of mobile and online survey respondents indicated that it was not easy to pay for vaccination. The financial cost to conduct the survey per survey response was US $2.61 for the online survey, US $6.93 for the mobile survey, and US $29.38 for the household survey. Conclusions: In the Philippines, household, mobile, and online survey methods reached caregivers of children who were unvaccinated against DPT, and these proportions were similar across survey methods. BeSD responses indicated high vaccine demand and challenges in caregivers’ cost to access vaccination. Determining the most appropriate survey method depends on trade-offs between representativeness and costs. However, areas with strong connectivity and high mobile device ownership can consider mobile and online methods as a lower-cost alternative to rapidly collect BeSD data.
  •  

Digital Health Technology Use Among Rehabilitation Professionals in China: Multi-Province Cross-Sectional Survey

Background: The rapid expansion of rehabilitation needs in China has intensified pressure on a workforce that remains unevenly distributed. Digital health technologies (DHTs) offer potential to increase service reach and efficiency. However, little is known about how rehabilitation professionals currently gather and document clinical information, nor about their readiness to integrate digital tools into routine practice within China’s rapidly digitalizing health system. Objective: This study aimed to describe how rehabilitation professionals in China collect subjective and objective clinical information, document patient data in routine practice, and assess their willingness to use DHTs in clinical settings. Methods: We conducted a multi-province observational cross-sectional survey using a culturally adapted questionnaire based on the World Health Organization Digital Health Interventions framework. The instrument assessed participant characteristics, information collection methods, documentation practices, and willingness to adopt digital functions across rehabilitation activities. Descriptive analyses and subgroup comparisons were performed on 324 complete responses from certified rehabilitation professionals. The multi-province cross-sectional online survey was conducted among licensed rehabilitation professionals in China with internet access. Participants were recruited through professional networks and social media platforms. Results: Respondents were drawn from 20 provincial-level administrative regions across China, including Fujian (n=72), Guangdong (n=77), and Shanxi (n=45), among others, with 82.7% (268/324) employed in public sector rehabilitation services. Traditional methods dominated clinical work. Face-to-face communication was used frequently for subjective assessment by 96.3% (312/324) of respondents, whereas digital channels such as email (22/324, 6.8%) and telephone (47/324, 14.5%) saw limited use. For objective information, visual observation (271/324, 83.7%) and manual measurement tools (195/324, 60.2%) remained the primary approaches, while motion capture technology (45/324, 13.8%) and wearable sensors (13/324, 4%) were rarely used. Documentation practices also relied heavily on analogue formats, with 82.1% (266/324) using handwritten notes and 60.2% (195/324) using paper templates. In contrast, willingness to adopt DHTs was consistently high, with 80.6% (261/324) of respondents indicating readiness to use digital systems for identity verification, 79.0% (256/324) for progress tracking, and 78.1% (253/324) for outcome measurement. Subgroup analyses revealed that educational level significantly influenced the adoption of advanced technologies, with master’s or doctoral degree holders reporting higher use of sensor-based assessment, motion capture, and wearable devices. In contrast, professional title and clinical specialty showed limited influence, with no significant differences observed for most digital health functions. Conclusions: Rehabilitation professionals in China demonstrate strong readiness to use DHTs, yet their routine practice remains largely paper-based and analogue. These findings provide evidence to inform implementation strategies, workforce training, and system-level planning aimed at accelerating digital transformation in rehabilitation services.
  •  

Goal Setting and Anchoring Effects on Meditation Using a Digital Platform: Large-Scale Digital Field Study

Background: Meditation has grown in popularity in recent years, but many people who try meditation often fail to establish a habit. Goal setting has been demonstrated to be an effective technique in behavior change in other health-related contexts but is understudied in the meditation context. Objective: This study had 2 objectives: (1) to assess the association between goal setting and the number of days people meditated and (2) to evaluate whether anchoring bias in the goal-setting question (via response option order) influences goal selection and subsequent meditation behavior. Methods: This large-scale quasi-experimental field study included 18,559 Spotify mobile users aged 18 years or older residing in Australia, Canada, New Zealand, the United Kingdom, or the United States who had listened to at least 5 minutes of meditation content from a specified teacher. The in-app experiment consisted of 2 goal-setting test conditions and an active control. In the test conditions, participants selected the number of days they intended to listen to content from the meditation teacher in the next 7 days. The conditions differed only in the order of goal response options (higher goals listed first vs last). The active control rated how much they liked the teacher but did not set a goal. Because responding was optional, selection bias is possible, and the design is quasi-experimental. Results: The act of setting any goal had a modest positive association with the number of days people meditated in both treatment condition 1 (=.08, 95% CI 0.01-0.16) and treatment condition 2 (=.08, 95% CI 0.002-0.15). People who committed to higher goals were also more likely to meditate more than those who committed to lower goals. Additionally, the distribution of goals between the treatment conditions varied (=84.24;
  •  

Innovations in Deaf Health Care Communication: Systematic Review of Sign Language Recognition Systems

Background: Deaf individuals often face communication challenges when interacting with those who can hear. Within health care settings, these challenges may pose risks to their safety, potentially resulting in misdiagnoses, treatment errors, and decreased quality of care. Objective: This study aims to systematically review the evidence on communication systems reported in the literature that use human-computer interaction techniques to support communication between deaf individuals who use sign language and hearing health professionals in health care settings. The review focuses on systems that are either currently in use or proposed for use in health care and that have been tested using human participants or videos of human users. Methods: A comprehensive search was performed via MEDLINE, Web of Science, ACM, IEEE Xplore, Scopus, and Google Scholar in March 2025. The inclusion criteria comprised studies developing a sign language recognition system within a health care context and testing with human users. Eligible studies underwent screening by 2 independent investigators (LRV and LMMSR or LFRdO and GTdSS), with any disagreements resolved by a senior researcher (MSM). Results: The search retrieved 21,778 publications, and screening of reference lists identified 2 additional studies, resulting in a total of 23 studies meeting the eligibility criteria. Most systems (15/23, 65.2%) were image-based, while 34.8% (8/23) relied on sensors (glove-based or depth-sensing). Applications varied across health care settings, including general hospital care (10/23, 43.5%), emergencies (8/23, 34.8%), and primary care (4/23, 17.4%). All systems were in the development and testing stage, with no data on security and psychological impacts. Accuracy ranged from 25% to 100% for image-based and 72% to 99.7% for sensor-based systems. Bidirectionality and facial expression recognition, crucial for effective communication, were largely overlooked. Conclusions: Image-based systems were more common than sensor-based ones, though both showed wide variability in accuracy in recognizing and interpreting signs. Most systems failed to address critical aspects such as bidirectional communication and the recognition of facial expressions, essential for effective communication. None fully addresses the requirements for integration into health care settings. These findings highlight the need for further research on implementation, usability, and impact on the quality of care for deaf patients. International Registered Report Identifier (IRRID): RR2-10.2196/55427
  •  

Misinformation in Social Media Narratives on Highly Pathogenic Avian Influenza: Systematic Content Analysis of Facebook and Instagram Posts

Background: Recurrent outbreaks of the highly pathogenic avian influenza (HPAI) A (H5N1) virus in farmed poultry, and reports of infections in dairy cattle herds in the United States since March 2024, have triggered concerns about the spillover threat to human populations and a subsequent influenza pandemic. The increasing threat that H5N1 poses to human health has led to more vigilant public health monitoring of these developments. In addition to intensifying surveillance, preventative strategies—like vaccinating those at higher risk—are being evaluated to help minimize infection and spread. Objective: Efforts to mitigate and respond to such an event will entail broad public health interventions including vaccination. However, analysis of the COVID-19 pandemic suggests that information quality can significantly impact the effectiveness of such measures by influencing public understanding and trust. Misinformation about H5N1 and other viruses circulating online often includes inaccurate information about transmission, prevention, and the severity of the viruses. By systematically analyzing these false narratives, public health authorities can better tailor their pandemic prevention, preparedness, and response strategies. Methods: In light of the emerging threat of H5N1, we analyzed the content of social media posts from Facebook (approximately 350,000) and Instagram (n=69,551) related to HPAI. Using 40 keywords associated with misinformation, we identified over 500 posts explicitly mentioning H5N1 and related terms for further systematic analysis. Posts were coded to identify targets and topics in the social media narratives. The “target” refers to the organization or person mentioned in the post, while the “topic” refers to the primary issue or subject being addressed. Results: Our content analysis identifies 7 main targets of misinformation, including government (149/544, 27%), health authorities (108/544, 20%), and international organizations (74/544, 14%). Also, from the 6 topics that have been identified, we found that the most widespread one was that authority figures purposefully engineer pandemics to achieve multiple political, economic, and other objectives (362/544, 67%) followed by societal destruction (121/544, 22%), and anti-vaccination (84/544, 15%). Other themes include societal destruction and religious allusions and prophecies. Conclusions: Our analysis of online content showed that H5N1 misinformation was primarily aimed at individuals or groups with differing degrees of political or institutional authority, such as government leaders and public health officials. These figures were often the focus due to their involvement in making health policy decisions and implementing public health measures. Decision-making entities and individuals were the target of various misinformation narratives. Results demonstrate the ongoing need for monitoring health misinformation to inform evolving public health responses to HPAI.
  •  

Young Adults’ Interactions With Food and Nutrition Content on Social Media and Implications for Intervention Design: Semistructured Interview Study

Background: Young adults increasingly rely on social media for nutrition information. However, little is known about (1) which types of eating-related content they actively engage with and why, and (2) how they interpret, evaluate, and incorporate this content into their everyday food choices and health behaviors. Objective: This qualitative study explored how UK young adults (aged 18-25 years) interact with food and nutrition content across social media platforms to inform the design of future social media interventions. Methods: Semistructured online interviews, guided by the Capability, Opportunity, Motivation–Behavior (COM-B) model, were conducted with 25 active social media users (18/25, 72% women, mean age 22.2, SD 1.9 years, ethnically diverse) in the United Kingdom between August and October 2024. The study design was informed by patient and public involvement to ensure relevance and acceptability. Data were analyzed using reflexive thematic analysis. To guide intervention development, key findings (coded as barriers and facilitators) were systematically mapped to the Theoretical Domains Framework, and the COM-B. Ethics approval was obtained from the University of Cambridge (24.368). Results: Five key themes were identified: (1) evolving engagement patterns (passive scrolling to active interaction and mixed feelings on algorithmic control), (2) conflicted information seeking (frustration with contradictory advice, varied strategies to assess credibility), (3) multifaceted behavioral impact (simultaneous positive impacts such as cooking inspiration and negative impacts such as restrictive eating triggers), (4) shifting goals (a movement from appearance-focused to health-centered goals; yet, vulnerability to body-image issues), and (5) intervention preferences (demand for credible professionals, customizable content, and privacy protection). Participants demonstrated a reactive learning process, developing “digital nutrition literacy” often after negative experiences. Social influences were identified as the most frequently cited domain (mapped to TDF [theoretical domains framework]/COM-B) shaping interactions with social media content. Conclusions: This study challenges assumptions of passive social media consumption, showing that young adults actively develop protective strategies yet remain vulnerable to misinformation. Digital interventions should leverage user agency and address diverse perceptions through customizable, credible content delivered with privacy and emotionally safe messaging. The COM-B and TDF mapping provide specific, evidence-based behavioral targets, particularly within the domain of Social Opportunity and Reflective Motivation, to guide the development of effective eHealth interventions.
  •  

Telehealth Delivery of the Homeostasis–Enrichment–Plasticity Approach for Premature Infants With Developmental Risks: Exploratory Feasibility Study

Background: Preterm delivery is an increasing worldwide health concern linked to increased neurodevelopmental risks. Early intervention is crucial for harnessing neuroplasticity to enhance developmental and functional performance outcomes; however, access to early intervention is frequently hindered by logistical, financial, and labor constraints. The Homeostasis–Enrichment–Plasticity (HEP) Approach is a family-centered early intervention model based on enriched environments, designed to improve infants’ sensory-motor, cognitive, and socio-emotional development. Objective: This study aimed to assess the feasibility, safety, acceptability, and outcomes sensitivity to change of implementing the HEP Approach through telehealth for premature infants at developmental risk. Methods: A pre-post exploratory feasibility study was performed, including 16 preterm infants (aged 4-12 months corrected age), of whom 14 completed the study. The 12-week intervention included weekly remote sessions focused on environmental enrichment, active exploration, and parental guidance. The feasibility and acceptability were evaluated using a 24-item questionnaire. Developmental outcomes were assessed with the Young Children’s Participation and Environment Measure, Ages and Stages Questionnaire (ASQ), Alberta Infant Motor Scale, Infant Motor Profile, and Depression Anxiety Stress Scales. Results: High adherence (14/14, 100%) and retention (14/16, 87.5%) rates demonstrated robust feasibility. Parents indicated 86%-100% agreement across all feasible criteria, affirming safety, satisfaction, and acceptability. No adverse incidents were reported. Changes were identified in participation (Young Children’s Participation and Environment Measure), motor development (Alberta Infant Motor Scale, Infant Motor Profile, and ASQ), communication and social-emotional domains (ASQ), and caregiver well-being (Depression Anxiety Stress Scales) (P<.05). Conclusions: The telehealth implementation of the HEP Approach demonstrated feasibility, safety, and strong acceptance among families, along with quantifiable developmental and psychosocial changes. These initial findings endorse the model’s viability as an accessible, family-oriented telehealth framework for infants born preterm. Future randomized controlled and longitudinal studies are necessary to validate intervention efficacy and scalability.
  •  

Feasibility and Acceptability of AI-Powered Tools for Early Autism Screening in Egypt: Semistructured Focus Group Study

Background: Autism spectrum disorder (ASD) is often underdiagnosed in low- and middle-income countries due to limited specialist access, sociocultural stigma, and fragmented screening systems. Artificial intelligence (AI)–powered screening tools may improve early detection by enabling low-cost, accessible assessments. However, adoption depends on stakeholder trust, ethical safeguards, and alignment with local health system capacities. Objective: This study explored the feasibility, acceptability, and perceived ethical and practical enablers and barriers to implementing AI-powered tools for early ASD screening in Egypt, with attention to urban–rural disparities and integration into existing care pathways. Methods: We used a qualitative design with semistructured focus group discussions with 49 participants (21 parents of children with ASD and 28 health care professionals) recruited from urban and rural governorates. Discussions were audio-recorded, transcribed verbatim, and analyzed using Braun and Clarke’s reflexive thematic analysis, supported by NVivo software (Lumivero). Methodological integrity was ensured through reflexivity, triangulation, and peer debriefing. Thematic saturation was monitored across groups, and participant diversity was prioritized across contexts. Results: Five themes emerged: (1) AI as a supportive tool rather than a replacement for clinicians, emphasizing scalability and assistance for nonspecialists; (2) the need for cultural and contextual adaptation to ensure local relevance; (3) privacy, trust, and transparency concerns, including data security, consent, and algorithmic opacity; (4) reducing diagnostic inequities by addressing urban–rural disparities and strengthening community-based deployment; and (5) the preference for hybrid AI–human models, with conditions for adoption including cultural sensitivity, human oversight, and digital literacy support. Counts (n/N) of parents and health care professionals contributing to each theme were used descriptively as indicators of pattern salience rather than as statistical estimates of prevalence. Participants expressed cautious optimism, with parents emphasizing accessibility and speed, while health care professionals highlighted concerns about reliability, cultural adaptation, and data governance. Conclusions: AI-powered ASD screening has potential to advance equitable early detection in underserved areas. Adoption requires transparent data governance, integration into hybrid human–AI models, culturally adaptive design, and targeted digital literacy initiatives. These findings provide an evidence-based roadmap for policymakers, technologists, and health system leaders to implement AI screening tools that are ethically sound, contextually relevant, and equity-focused.
  •  

Preferences for Personalized Text Message Appointment Reminders Among Outpatients in a Universal Health System: Cross-Sectional Study

Background: SMS text messaging reminders are widely used to reduce missed outpatient appointments; however, evidence remains limited regarding which types of reminder content patients prefer, particularly within East Asian universal health systems. In Taiwan, minimal financial barriers to care and unrestricted access to secondary and tertiary hospitals contribute to high outpatient visit volumes and persistent no-show rates. These contextual features underscore the need for behaviorally informed and demographically tailored reminder strategies rather than uniform messaging approaches. Objective: This study aimed to examine patient preferences for 6 theory-guided SMS appointment reminder types and to identify the predictors of reminder preference related to demographic characteristics and health care utilization, with the goal of informing personalized reminder design for a forthcoming randomized controlled trial. Methods: We conducted a cross-sectional online survey among adults in Taiwan with prior outpatient experience. Six SMS reminder prototypes were developed based on behavioral communication principles and validated by a multidisciplinary expert panel using item-level content validity indices. Participants selected their preferred SMS reminder type and reported sociodemographic characteristics and recent health care utilization. Bivariate associations were examined using chi-square tests and one-way ANOVA, with Benjamini-Hochberg false discovery rate correction applied to control for multiple testing. To identify independent predictors of SMS reminder preference while adjusting for potential confounding, we fitted a multinomial logistic regression model with all covariates entered simultaneously. Results: A total of 1095 respondents completed the survey. General reminders and messages referencing prior missed appointments were most frequently preferred, whereas empathy-based or relationally framed messages were selected less often. In false discovery rate–adjusted univariate analyses, both age and sex were associated with SMS reminder preference. However, in the fully adjusted multinomial logistic regression model, age emerged as the only statistically significant independent predictor. Participants younger than 50 years were significantly more likely to prefer alternative reminder message types compared with the general reminder (adjusted odds ratio 1.64, 95% CI 1.18‐2.28; =.003). Sex did not retain statistical significance after multivariable adjustment. Other sociodemographic characteristics and health care utilization variables, including education level, employment status, residential region, outpatient visit frequency, and recent missed appointments history, were not independently associated with reminder preference. Conclusions: Preferences for outpatient SMS reminder content vary systematically, with age representing the most robust independent predictor. Across the sample, concise and behavior-focused reminders were preferred over empathy-oriented or relational formats. These findings support age-informed tailoring of SMS reminder content and provide content-validated SMS prototypes for use in subsequent interventional research. The results offer formative evidence to guide the design of randomized trials aimed at reducing outpatient no-shows and improving the efficiency of ambulatory care delivery in Taiwan’s universal health care system.
  •  

Electrocardiogram-Based Mental Stress Detection Amid Everyday Activities Using Machine Learning: Model Development and Validation Study

Background: Frequent, sustained stress is linked to poor health and requires monitoring for early intervention. Electrocardiograms (ECG) are promising biomarkers because they can be recorded noninvasively and continuously using wearable devices. However, tracking stress with ECG is challenging because daily activities elicit responses similar to mental stress (MS), and various mental stimuli that individuals encounter complicate the use of machine learning (ML) models trained on a limited set of stressors. Objective: We (1) evaluated the ability of ML models to distinguish MS episodes from a composite “no-stress” background, including rest and low- to moderate-intensity activities; (2) assessed their generalizability to new stressors and participants; and (3) tested robustness to lower sampling rates and fewer features, to explore their suitability for lightweight wearables. Methods: We used a comprehensive ECG dataset sampled at 1000 hertz from 127 participants who underwent various mental stressors and engaged in diverse physical activities. A 30-second window was used to extract 55 features from time, frequency, nonlinear, and morphological domains. We trained a logistic regression (LR) model and an extreme gradient boosting (XGBoost) model, splitting the data into 60/20/20 for training, validation, and testing. Shapley additive explanation values were computed to explain model predictions. Additional analyses included leave-one-stressor-out; downsampling to 500, 250, and 125 hertz; a time-window sensitivity analysis; and reducing the number of features to as few as 5. Results: XGBoost achieved an area under the receiver operating characteristic curve (AUROC) of 0.741 (95% CI 0.701‐0.783) and an area under the precision-recall curve (AUPRC) of 0.706 (95% CI 0.658‐0.753), compared with 0.724 (95% CI 0.678‐0.772) and 0.691 (95% CI 0.639‐0.742) for LR. The mean performance difference between XGBoost and LR was 0.017 for AUROC (95% CI 0.001‐0.032) and 0.015 for AUPRC (95% CI −0.001 to 0.037; clustered bootstrap analysis using 2000 participant-level resamples), suggesting that LR performs comparably to the nonlinear XGBoost model. Both models were robust to downsampling and feature reduction (10 features retained >93% of performance). Extending the analysis window to 60 seconds improved model performance across all sampling rates, highlighting a trade-off between rapid detection and overall performance. When evaluating discrimination from physical activity, models achieved acceptable specificity for light physical activity (XGBoost: 0.787; LR: 0.794) but poor specificity for moderate physical activity (XGBoost: 0.418; LR: 0.444). Both models generalized to most unseen stressors, although performance varied across stressors, with limited transfer to the social-evaluative stressor. Feature importance analysis revealed fuzzy entropy and frequency-based features as key predictors. Conclusions: ML models can detect MS with high sensitivity and remain robust to lower sampling rates and fewer features. Generalization to novel stressors was stressor-dependent. Importantly, our results highlight challenges in distinguishing stress-related cardiac responses from those caused by physical exertion, revealing critical limitations of single-sensor ECG approaches for MS detection.
  •  

A Gamified Mobile Health Intervention to Promote Physical Activity, Executive Function, and Mental Health in College Students: Randomized Controlled Trial

Background: College students commonly experience suboptimal health conditions, including insufficient physical activity (PA), excessive body weight, and declining physical fitness. Traditional interventions face low adherence, while gamified mobile health (mHealth) programs may improve engagement and outcomes. Objective: This study aimed to evaluate the feasibility and effectiveness of a novel gamified, incentive-based mHealth intervention on primary outcomes (PA and adherence) and secondary outcomes (physical fitness, body composition, executive function [EF], and mental health). Methods: A 2-arm parallel-group randomized controlled trial (RCT) was conducted in 2025 at Yantai University with 160 college students (18‐25 years; BMI 18.5‐30.0) who were randomized 1:1 (computer-generated, sex-stratified blocks of 4; concealed allocation) to the intervention group (IG) or control group (CG; n=80 each); major exclusions were contraindications to exercise, severe physical/mental illness, recent PA interventions, or psychotropic medication use. Both used the same fitness watch–app system and identical PA targets (≥150 min moderate-to-vigorous physical activity [MVPA] per week or ≥900 metabolic equivalent-minutes [MET-min] per week); IG additionally received team-based gamification (competition, points/leaderboards, feedback, and rewards), while CG received monitoring only. PA and adherence were monitored throughout the 8-week intervention; other outcomes were assessed at baseline and 8 weeks (fitness, body composition, EF, and mental health). Open-label with blinded outcome assessors/analysts; intention-to-treat (ITT) with multiple imputation. Results: At 8 weeks, data were available for 154 participants (IG 78; CG 76); all 160 were analyzed per ITT. Compared to the CG, the IG demonstrated significantly higher mean levels in all primary PA outcomes over 8 weeks (daily steps: mean 10,356, SD 1245 versus 8242, SD 1087; Δ=2114; =1.81, 95% CI 1.44‐2.18;
  •  

Initial Insights Into an Institutional Secure Large Language Model for Magnetic Resonance Imaging Examination Requests: Retrospective Study

Background: Incomplete clinical details on magnetic resonance imaging (MRI) examination requests (MERs) can lead to suboptimal protocol selection. An institutional secure large language model (sLLM) with access to manually retrieved salient data from the electronic medical record (EMR) may improve request completeness and protocol accuracy across multiple MRI subspecialties. Objective: The objective of this study was to compare clinician MERs with sLLM-augmented MERs for information quality and to evaluate the protocoling accuracy of the sLLM versus board-certified radiologists across body, musculoskeletal, and neuroradiology MRI. Methods: This retrospective study included 608 random outpatient MRI examinations performed between September 2023 and July 2024 (body 206, musculoskeletal 203, neuroradiology 199). The cohort comprised 528 patients (mean 51.2 years, SD 19.2; range 4‐93; n=279, 52.8% women, n=249, 47.2% men). MERs without EMR access were excluded. A privately hosted Anthropic Claude 3.5 model (temperature 0) augmented each MER with manually retrieved salient EMR data and, via rule-based parsing, mapped the extracted elements onto predefined institutional criteria to recommend region or coverage and contrast use. Two experienced radiologists established a consensus reference standard. Two board-certified general radiologists (Rad 3 and Rad 4) and the sLLM were compared with this standard. Clinical information quality was graded using the Reason-for-Exam Imaging Reporting and Data System (RI-RADS). Interrater reliability was quantified with Gwet AC1. Paired accuracies were compared with the McNemar test to determine whether there was a statistically significant difference. Results: Interreader agreement for RI-RADS was almost perfect for sLLM-augmented MERs (AC1 0.97, 95% CI 0.94‐0.99) and moderate for clinician MERs (AC1 0.43, 95% CI 0.34‐0.52). Limited or deficient clinical information (RI-RADS C/D) fell to 0% to 0.7% (0/608 to 4/608) with sLLM augmentation vs 4.1% to 20.4% (25/608 to 124/608) for clinician MERs. Overall protocol accuracy was 93.1% (566/608; 95% CI 89.6‐96.6) for the sLLM, 91.4% (556/608; 95% CI 87.6‐95.3) for Rad 3, and 92.1% (560/608; 95% CI 88.4‐95.8) for Rad 4 (sLLM vs Rad 3 =.23 vs Rad 4 =.40). Region or coverage accuracy was similar (sLLM: 579/608, 95.2%; Rad 3: 585/608, 96.2%; Rad 4: 573/608, 94.2%; =.46 and =.36). Contrast decisions were more accurate using the sLLM at 94.4% (574/608; 95% CI 91.3‐97.5) vs Rad 3 at 92.1% (560/608; 95% CI 88.4‐95.8; =.027) and were not significantly different to Rad 4 at 92.9% (565/608; 95% CI 89.4‐96.4; =.16). Subspecialty analyses showed similar patterns, with the sLLM outperforming Rad 4 for musculoskeletal MRI contrast decisions (96.6% vs 91.1%; =.006) and matching readers elsewhere. Manual review indicated that sLLM improvements arose from EMR details not listed on the MER (infection/inflammation, tumor history, prior surgery). No clinically significant hallucinations were identified in a manual review of discordant cases. Conclusions: Across body, musculoskeletal, and neuroradiology MRI, sLLM-augmented examination requests improved clinical context and enhanced contrast selection while demonstrating accuracy comparable to general radiologists for region or coverage. Integrating sLLMs into routine vetting workflows may reduce manual workload in protocol selection for more efficient, standardized protocoling.
  •  

Social Media Intervention Based on the Information-Motivation-Behavioral Skills Model Promotes HIV Testing and Reduces High-Risk Behaviors Among Men Who Have Sex With Men in Resource-Limited Settings in China: Randomized Controlled Trial

Background: Social media intervention may enhance HIV prevention among men who have sex with men, but the effect of this intervention in resource-limited settings remains unclear. Objective: This randomized controlled trial evaluated whether a social media intervention grounded in the information-motivation-behavioral skills (IMB) model could be beneficial for HIV prevention among men who have sex with men in resource-limited settings. Methods: Participants were recruited in Nanning, China, between April 2023 and April 2024. Eligible participants were randomly assigned to either the social media intervention group or the routine HIV prevention services control group. Participants in the intervention group received a 3-month social media intervention, which included completing video-based tasks. Baseline surveys were conducted, followed by follow-up surveys every 3 months, for a total of 2 follow-ups. Outcomes included HIV testing uptake, high-risk behavior, AIDS-related knowledge, safe sex self-efficacy, and attitude. Results: A total of 180 eligible men who have sex with men were enrolled (90 per group). Follow-up rates were 97.8% (88/90) and 95.5% (86/90) for the intervention and control groups, respectively. At the follow-ups, the intervention group demonstrated significantly higher uptake of HIV testing, a lower proportion of participants reporting high-risk sexual behaviors, and higher condom use self-efficacy compared to the control group (all
  •  

Changes in Workplace Productivity and Estimated Cost Savings During Internet-Based Cognitive Behavioral Therapy in the Irish National Health Service: Naturalistic, Repeated-Measures, Retrospective Survey Study

Background: Depression and anxiety can significantly impact workplace productivity, for instance, by increasing absenteeism and presenteeism. This loss of productivity leads to diminished workplace economic outcomes. Internet-based cognitive behavioral therapy (iCBT) has emerged as a cost-effective intervention within workplace settings that improves workplace productivity loss due to depression and anxiety, but more generalizable evidence beyond the workplace, such as in a national health service setting, is lacking. Objective: This naturalistic, repeated-measures, retrospective study investigated the impact of iCBT on work productivity metrics using nationally representative data from patients enrolled in the Irish national health service (ie, the Health Service Executive). Methods: We analyzed repeated measures retrospective data from 7125 employed patients enrolled in iCBT at the Health Service Executive between March 2023 and May 2024. The Work Productivity and Activity Impairment questionnaire was used to measure absenteeism, presenteeism, overall productivity loss, and activity impairment. Secondary outcomes included depression (Patient Health Questionnaire-9) and anxiety (Generalized Anxiety Disorder-7). Patients were primarily 25 to 64 years old (n=5578, 78%), female (n=4956, 70%), and met clinical scoring criteria on the Patient Health Questionnaire-9 or Generalized Anxiety Disorder-7 (n=4774, 67%). Missing data were handled using multiple imputation. We used mixed-effects models to assess pre-post treatment changes in outcomes and then utilized Irish national salary estimates from 2022 to derive cost savings (in 2022 € values; €1=approximately US $1.05) based on productivity improvement during use of the iCBT program. Results: From baseline to follow-up, absenteeism reduced by 6.85% (
  •  

Supporting Access to Care Through Peripheral Devices and Patient-Generated Health Data: Qualitative Study

Background: In 2016, the US Department of Veterans Affairs (VA) implemented a national initiative to distribute video-enabled tablets and peripheral devices, such as blood pressure monitors and weighing scales, to patients facing geographic, clinical, or socioeconomic challenges. Such patients could potentially benefit from health monitoring in conjunction with video-based care, as peripheral devices offer opportunities to enrich care received during a video visit and support tracking of health-related data collected outside of clinical care, or patient-generated health data. However, little is known about experiences with the devices and how they could support improved access to care. Objective: We explored patients’ experiences with VA-issued peripheral devices and their impact on video-based care and health monitoring outside of clinical visits. Methods: We conducted in-depth semistructured interviews among patients who received VA-issued tablets and peripheral devices between 2023 and 2024. Purposive sampling was used to gather views based on gender, age, race or ethnicity, and rurality. Interviews were transcribed and analyzed using rapid qualitative analysis, guided by the Unified Theory of Acceptance and Use of Technology. Results: Among 25 patients, most received a blood pressure monitor (21/25, 84%), a weight scale (14/25, 56%), and/or a pulse oximetry device (12/25, 48%). The majority reported using their peripheral devices (23/25, 92%) and tablets (19/25, 76%) to monitor their vital signs and attend video visits. Qualitative analysis yielded ten themes reflecting experiences and impacts of the devices, organized by the Unified Theory of Acceptance and Use of Technology constructs: “effort expectancy” consisted of (1) familiar and easy to use devices and (2) challenges of Bluetooth pairing and measurement; “performance expectancy” consisted of (3) integration with video visits, (4) health monitoring for peace of mind, (5) perceptions of improved vital signs and lifestyle behaviors, (6) removing obstacles to in-person care, and (7) desiring an overall picture of health; “social influence” consisted of (8) fostering care team connections and (9) promoting awareness of tablets and peripheral devices; and “facilitating conditions” consisted of (10) supportive help desk infrastructure. Overall, patients described using peripheral devices during virtual visits by syncing data to the tablet for real-time access by their care team. They also reported manually tracking and sharing patient-generated health data with their care team. Despite some challenges with Bluetooth pairing, patients found the devices easy to use and contributed to improved health and motivation. Devices also reduced logistical burdens of in-person visits, especially for those with limited mobility, visual impairments, mental health needs, or transportation barriers. Conclusions: Patients perceive that peripheral devices can enhance video-based care and support health care access and chronic disease management. Patients reported benefits to health, behavior, and communication with care teams. To maximize the impact, program enhancements should prioritize device interoperability, accessible training, and expanded outreach. Trial Registration:
  •  
❌