❌

Reading view

Beyond GPT-4: The Rapidly Evolving Potential of Large Language Models for Clinical Guideline Improvement

This commentary reviews the study by Jones et al, which evaluated whether GPT-4 could improve the readability of injectable medication guidelines while preserving important safety information. The study found that GPT-4 produced modest readability gains comparable to manual revision, but also introduced omissions and meaning changes in a minority of sections. These findings highlight both the potential and limitations of early large language models (LLMs) in clinical contexts. However, this study reflects the capabilities of a specific model in a rapidly evolving domain. Since the release of GPT-4, advances in multistep reasoning, model-critique workflows, and structured validation have substantially improved the ability of newer systems to detect omissions, maintain factual fidelity, and support controlled editing. As a result, some documented limitations may stem from the constraints of a single-model, single-pass workflow rather than intrinsic flaws in LLM-assisted guideline revision. This commentary highlights the need for evaluation frameworks that can keep pace with LLM progress and emphasizes that clinical oversight and user-centered testing remain essential. Updated research using contemporary models is needed to determine how emerging architectures can more safely support clarity, consistency, and maintenance of clinical guidelines.
  •  

Large Language Model–Based Analysis of Statin Therapy Discussions and Sentiment on Social Media: Cross-Sectional Observational Study

Background: Statin therapy, despite proven cardiovascular benefits, remains underused. Social media platforms may capture patient perspectives that are less visible in clinical encounters. Objective: This study aimed to characterize themes, sentiment, and decision-making factors related to statin therapy through large language model (LLM)–based analysis of Reddit discussions. Methods: This cross-sectional observational study analyzed English-language Reddit posts and comments mentioning statins from January 2022 to May 2025, identified via keyword-based Reddit application programming interface searches (≤1000 posts per keyword). A total of 5328 retrieved discussions (n=1661, 31.2% posts and n=3667, 68.8% keyword-containing comments) from public subreddits were included. Themes, sentiments (positive, neutral, or negative), guideline-informed clinical relevance, information-seeking behavior, adverse effect mentions, decision factors, and adherence-related content were extracted using an LLM-based pipeline. Results: Among 5328 discussions, prominent topics included adverse effects (n=1697, 31.9%), decision-making references related to laboratory results and physician advice (n=2767, 51.9% and n=2034, 38.2%, respectively), and alternative approaches (n=2485, 46.6%). Overall sentiment was neutral in 34% (n=1812) of discussions, negative in 30.9% (n=1646), and positive in 16.9% (n=900); the remainder were mixed or unclear. Statin-directed sentiment was neutral in 44.1% (n=2350) of discussions, negative in 25.2% (n=1343), and positive in 12.5% (n=666); the remainder did not express statin-directed sentiment. High clinical relevance was identified in 12.6% (n=672) of discussions. Adherence-related issues were mentioned in 29.8% (n=1587) of discussions. Among adverse effect mentions, muscle pain (n=129, 7.6%) and fatigue (n=110, 6.5%) were common. Conclusions: LLM-enabled analysis of Reddit discourse highlights substantial negative sentiment, adherence-related concerns, and adverse effect narratives surrounding statin therapy. These findings suggest opportunities for patient-centered communication and shared decision-making strategies that address symptom attribution, uncertainty, and information needs in digital information environments.
  •  

Health Equity Analysis of Awareness and Use of GetCheckedOnline, British Columbia’s Digital Intervention for Sexually Transmitted and Blood-Borne Infection Testing in 5 Urban, Suburban, and Rural Communities: Cross-Sectional Survey Study

Background: Digital sexually transmitted and blood-borne infection (STBBIs) testing services are used to improve testing access, but might replicate existing social inequities. Previous research has shown that the digital STBBI testing service has improved access to testing in British Columbia (BC), Canada. As part of the program’s continuous evaluation, we examined awareness and use of the service in 5 urban, suburban, and rural communities where the program has expanded. Objective: This study aimed to determine if social location is associated with differences in awareness and use of the service in 5 communities outside Vancouver, BC. Methods: From July to September 2022, we conducted a cross-sectional survey recruiting (in-person and online) sexually active people aged 16 years or older in 5 urban, suburban, and rural communities where had sample collection sites available at the time. We examined differences in awareness and use by age, gender identity, sexual identity, race/ethnicity, education, and income using logistic regression models informed by the Health Equity Measurement Framework. Results: Of the 1658 participants (n=1058, 63.8% in-person and n=600, 36.2% online), 35.3% (586/1658) were aware of and 19.5% (324/1658) had used it. Awareness and use were lower in the first and last age quartiles compared to the second quartile (>38 years: awareness odds ratio [OR] 0.23, 95% CI 0.17‐0.32; use OR 0.19, 95% CI 0.12‐0.28;
  •  

Context-Aware Sentence Classification of Radiology Reports Using Synthetic Data: Development and Validation Study

Background: Automated structuring of radiology reports is essential for data utilization and the development of medical artificial intelligence models. However, manual annotation by experts is labor-intensive, and processing real clinical data through commercial large language models (LLMs) presents significant privacy risks. These challenges are particularly pronounced for non-English languages like Japanese, where specialized medical corpora are scarce. While synthetic data generation offers a potential privacy-preserving alternative, its effectiveness in capturing complex clinical nuances—such as negation and contextual dependencies—to train robust classification models without any real-world training data has not been fully established. Objective: This study aimed to develop a context-aware sentence classification model for Japanese radiology reports using an entirely synthetic training pipeline, thereby eliminating reliance on real-world clinical data during the development phase. Furthermore, we sought to evaluate the generalizability of this approach by validating the model’s performance on diverse, multi-institutional, real-world reports. Methods: Japanese radiology reports (n=3104) were generated using GPT-4.1 and automatically annotated at the sentence level into 4 categories (background, positive finding, negative finding, and continuation) using GPT-4.1-mini. The synthetic data were partitioned into training (n=2670), validation (n=334), and test (n=100) sets. We fine-tuned several models, including lightweight local LLMs (Qwen3 and Llama 3.2 series) using low-rank adaptation and Japanese text classification models (Bidirectional Encoder Representations from Transformers [BERT]-base Japanese v3, Japanese Medical Robustly Optimized BERT Pretraining Approach [JMedRoBERTa]-base, and ModernBERT-Ja-130M). External validation was performed using 280 real-world reports (3477 sentences) from 7 institutions in the Japan Medical Image Database, with ground-truth labels established by board-certified radiologists. Evaluation metrics included accuracy, macro-averaged (macro ) score, and positive predictive value for positive findings (PPV_1). Results: All models achieved high performance on the synthetic test set (accuracy: 0.938‐0.951; macro -score: 0.924‐0.940). Overall performance declined on the external validation dataset (accuracy: 0.783‐0.813; macro -score: 0.761‐0.790), reflecting distributional differences between synthetic and real-world reports; however, PPV_1 remained stable and high across datasets (eg, 0.957 on the synthetic test set vs 0.952 on the external validation dataset for Qwen3 [4B]). Parsing errors occurred in LLM-based approaches (19‐260 sentences, 0.55%‐7.48% in the external dataset). Conclusions: This study demonstrates the feasibility of developing context-aware sentence classification models for Japanese radiology reports using a training pipeline based entirely on synthetic data. The stability of PPV_1 indicates that the models successfully captured the essential clinical terminology and linguistic patterns required to identify positive findings in real-world reports, despite the observed performance degradation during external validation. This approach substantially reduces manual annotation requirements and privacy risks, providing a scalable foundation for constructing structured radiology datasets to support the development of clinically relevant medical artificial intelligence models.
  •  

Views of People With Psychosis About Algorithm-Based Relapse Prediction and Data Sharing: Qualitative Study

Background: Preventing relapses of psychosis is difficult and important. Digital remote monitoring (DRM) systems are being developed and tested to support this. Increasingly, these systems use algorithm-based relapse prediction. Hence, understanding stakeholder views about algorithmic prediction is crucial. Existing qualitative work has explored health professionals’ views, but very few studies have examined the perspectives of people with psychosis on this topic. Objective: This paper aimed to provide an in-depth examination of the views of people with psychosis regarding algorithmic relapse prediction within a DRM system that incorporates active symptom monitoring and passive sensing data. Methods: People with psychosis (n=58) were recruited from 6 geographically distinct areas of the United Kingdom. They participated in semistructured qualitative interviews exploring their views about using a DRM system that predicts psychosis relapse based on a machine learning algorithm. Transcripts were analyzed using reflexive thematic analysis. People with lived experience of psychosis were involved extensively in study design, analysis, and reporting. Results: Findings were described across 4 themes. First, was a prominent theme. Participants emphasized that transparency about algorithm sensitivity and specificity is crucial and discussed the risks of the relapse prediction algorithm producing false positives (flagging that someone was relapsing when they were not) and false negatives (missing actual relapses). In both cases, participants said that errors may be partially mitigated through a approach (theme 2), with DRM blended with human oversight, from clinicians or a dedicated digital monitoring team, and calibrated based on service user, carer, and clinician feedback. The third theme, noted the interplay between users’ trust in the DRM system and their relationship with the clinical team. This theme described participants’ fears about potential overreactions (hospitalization or excessive medication) or underreactions (no additional support) from the clinical team in response to algorithm-generated relapse predictions. It emphasized the importance of retaining choice around the use of relapse detection algorithms and the sharing of personal data. The final theme described participants’ views about the , including facilitating early intervention, triaging care according to need, minimizing human bias in assessment, and efficiency in saving staff time. Conclusions: People with psychosis acknowledged potential benefits of algorithm-assisted relapse prediction for receiving timely or efficient care, but with several caveats. Algorithm-generated relapse alerts need to be sufficiently accurate and must be interpreted, with understanding of their limitations, by a trustworthy human who is aware of the relevant context. Algorithm-based relapse predictions should only be used with valid consent, in a way that promotes and respects the autonomy and voice of service users and avoids increasing the use of excessive restriction.
  •  

Child Vaccination Status and Behavioral and Social Drivers of Vaccination Among Their Caregivers in the Philippines: Cross-Sectional Survey Study Comparison of Household, Mobile, and Online Modes

Background: The World Health Organization recommends that countries routinely collect data on the behavioral and social drivers (BeSD) of vaccination to inform public health interventions that increase vaccine uptake. There is a need to identify data collection methods that can rapidly and inexpensively collect representative data, particularly in low- and middle-income countries. Objective: This study aimed to understand BeSD drivers of vaccination in the Philippines and assess the trade-offs between survey methods. We compared responses to household, mobile, and online surveys in terms of demographics, vaccination status, responses to BeSD questions, and cost. Methods: We conducted concurrent household, mobile (SMS text messaging and interactive voice response), and online surveys among caregivers of children 2 years of age and below in Regions V and XII of the Philippines, with sampling differing by survey method. We assessed, for each survey method, (1) respondent demographics (sex, age, region, and socioeconomic status) and (2) the weighted proportion of responses from caregivers of children who received at least one dose of diphtheria-pertussis-tetanus (DPT)–containing vaccine. We estimated the weighted proportion of each BeSD survey response option and calculated the financial cost (monetary outlays) per survey response from an implementer’s perspective by summing the costs incurred in each survey method and dividing by the number of responses received. Results: We surveyed a total of 1201 household respondents, 2153 mobile respondents, and 398 online respondents from January to March 2025. We found that online and mobile survey respondents were more likely to be male and have completed high school than household survey respondents. The weighted proportion of respondents indicating that their child had received at least one dose of DPT vaccine was 91.8% (n=1090; 95% CI 90%‐93.3%) for the household survey, 90.3% (n=1853) for the mobile survey, and 85% (n=346) for the online survey. With regard to vaccine demand, more than 85% of respondents in each survey method indicated that vaccines are very important, very safe, supported by family, and that they knew where to bring a child for vaccination. More than 30% of mobile and online survey respondents indicated that it was not easy to pay for vaccination. The financial cost to conduct the survey per survey response was US $2.61 for the online survey, US $6.93 for the mobile survey, and US $29.38 for the household survey. Conclusions: In the Philippines, household, mobile, and online survey methods reached caregivers of children who were unvaccinated against DPT, and these proportions were similar across survey methods. BeSD responses indicated high vaccine demand and challenges in caregivers’ cost to access vaccination. Determining the most appropriate survey method depends on trade-offs between representativeness and costs. However, areas with strong connectivity and high mobile device ownership can consider mobile and online methods as a lower-cost alternative to rapidly collect BeSD data.
  •  

Digital Health Technology Use Among Rehabilitation Professionals in China: Multi-Province Cross-Sectional Survey

Background: The rapid expansion of rehabilitation needs in China has intensified pressure on a workforce that remains unevenly distributed. Digital health technologies (DHTs) offer potential to increase service reach and efficiency. However, little is known about how rehabilitation professionals currently gather and document clinical information, nor about their readiness to integrate digital tools into routine practice within China’s rapidly digitalizing health system. Objective: This study aimed to describe how rehabilitation professionals in China collect subjective and objective clinical information, document patient data in routine practice, and assess their willingness to use DHTs in clinical settings. Methods: We conducted a multi-province observational cross-sectional survey using a culturally adapted questionnaire based on the World Health Organization Digital Health Interventions framework. The instrument assessed participant characteristics, information collection methods, documentation practices, and willingness to adopt digital functions across rehabilitation activities. Descriptive analyses and subgroup comparisons were performed on 324 complete responses from certified rehabilitation professionals. The multi-province cross-sectional online survey was conducted among licensed rehabilitation professionals in China with internet access. Participants were recruited through professional networks and social media platforms. Results: Respondents were drawn from 20 provincial-level administrative regions across China, including Fujian (n=72), Guangdong (n=77), and Shanxi (n=45), among others, with 82.7% (268/324) employed in public sector rehabilitation services. Traditional methods dominated clinical work. Face-to-face communication was used frequently for subjective assessment by 96.3% (312/324) of respondents, whereas digital channels such as email (22/324, 6.8%) and telephone (47/324, 14.5%) saw limited use. For objective information, visual observation (271/324, 83.7%) and manual measurement tools (195/324, 60.2%) remained the primary approaches, while motion capture technology (45/324, 13.8%) and wearable sensors (13/324, 4%) were rarely used. Documentation practices also relied heavily on analogue formats, with 82.1% (266/324) using handwritten notes and 60.2% (195/324) using paper templates. In contrast, willingness to adopt DHTs was consistently high, with 80.6% (261/324) of respondents indicating readiness to use digital systems for identity verification, 79.0% (256/324) for progress tracking, and 78.1% (253/324) for outcome measurement. Subgroup analyses revealed that educational level significantly influenced the adoption of advanced technologies, with master’s or doctoral degree holders reporting higher use of sensor-based assessment, motion capture, and wearable devices. In contrast, professional title and clinical specialty showed limited influence, with no significant differences observed for most digital health functions. Conclusions: Rehabilitation professionals in China demonstrate strong readiness to use DHTs, yet their routine practice remains largely paper-based and analogue. These findings provide evidence to inform implementation strategies, workforce training, and system-level planning aimed at accelerating digital transformation in rehabilitation services.
  •  

Goal Setting and Anchoring Effects on Meditation Using a Digital Platform: Large-Scale Digital Field Study

Background: Meditation has grown in popularity in recent years, but many people who try meditation often fail to establish a habit. Goal setting has been demonstrated to be an effective technique in behavior change in other health-related contexts but is understudied in the meditation context. Objective: This study had 2 objectives: (1) to assess the association between goal setting and the number of days people meditated and (2) to evaluate whether anchoring bias in the goal-setting question (via response option order) influences goal selection and subsequent meditation behavior. Methods: This large-scale quasi-experimental field study included 18,559 Spotify mobile users aged 18 years or older residing in Australia, Canada, New Zealand, the United Kingdom, or the United States who had listened to at least 5 minutes of meditation content from a specified teacher. The in-app experiment consisted of 2 goal-setting test conditions and an active control. In the test conditions, participants selected the number of days they intended to listen to content from the meditation teacher in the next 7 days. The conditions differed only in the order of goal response options (higher goals listed first vs last). The active control rated how much they liked the teacher but did not set a goal. Because responding was optional, selection bias is possible, and the design is quasi-experimental. Results: The act of setting any goal had a modest positive association with the number of days people meditated in both treatment condition 1 (=.08, 95% CI 0.01-0.16) and treatment condition 2 (=.08, 95% CI 0.002-0.15). People who committed to higher goals were also more likely to meditate more than those who committed to lower goals. Additionally, the distribution of goals between the treatment conditions varied (=84.24;
  •  

Innovations in Deaf Health Care Communication: Systematic Review of Sign Language Recognition Systems

Background: Deaf individuals often face communication challenges when interacting with those who can hear. Within health care settings, these challenges may pose risks to their safety, potentially resulting in misdiagnoses, treatment errors, and decreased quality of care. Objective: This study aims to systematically review the evidence on communication systems reported in the literature that use human-computer interaction techniques to support communication between deaf individuals who use sign language and hearing health professionals in health care settings. The review focuses on systems that are either currently in use or proposed for use in health care and that have been tested using human participants or videos of human users. Methods: A comprehensive search was performed via MEDLINE, Web of Science, ACM, IEEE Xplore, Scopus, and Google Scholar in March 2025. The inclusion criteria comprised studies developing a sign language recognition system within a health care context and testing with human users. Eligible studies underwent screening by 2 independent investigators (LRV and LMMSR or LFRdO and GTdSS), with any disagreements resolved by a senior researcher (MSM). Results: The search retrieved 21,778 publications, and screening of reference lists identified 2 additional studies, resulting in a total of 23 studies meeting the eligibility criteria. Most systems (15/23, 65.2%) were image-based, while 34.8% (8/23) relied on sensors (glove-based or depth-sensing). Applications varied across health care settings, including general hospital care (10/23, 43.5%), emergencies (8/23, 34.8%), and primary care (4/23, 17.4%). All systems were in the development and testing stage, with no data on security and psychological impacts. Accuracy ranged from 25% to 100% for image-based and 72% to 99.7% for sensor-based systems. Bidirectionality and facial expression recognition, crucial for effective communication, were largely overlooked. Conclusions: Image-based systems were more common than sensor-based ones, though both showed wide variability in accuracy in recognizing and interpreting signs. Most systems failed to address critical aspects such as bidirectional communication and the recognition of facial expressions, essential for effective communication. None fully addresses the requirements for integration into health care settings. These findings highlight the need for further research on implementation, usability, and impact on the quality of care for deaf patients. International Registered Report Identifier (IRRID): RR2-10.2196/55427
  •  

Misinformation in Social Media Narratives on Highly Pathogenic Avian Influenza: Systematic Content Analysis of Facebook and Instagram Posts

Background: Recurrent outbreaks of the highly pathogenic avian influenza (HPAI) A (H5N1) virus in farmed poultry, and reports of infections in dairy cattle herds in the United States since March 2024, have triggered concerns about the spillover threat to human populations and a subsequent influenza pandemic. The increasing threat that H5N1 poses to human health has led to more vigilant public health monitoring of these developments. In addition to intensifying surveillance, preventative strategies—like vaccinating those at higher risk—are being evaluated to help minimize infection and spread. Objective: Efforts to mitigate and respond to such an event will entail broad public health interventions including vaccination. However, analysis of the COVID-19 pandemic suggests that information quality can significantly impact the effectiveness of such measures by influencing public understanding and trust. Misinformation about H5N1 and other viruses circulating online often includes inaccurate information about transmission, prevention, and the severity of the viruses. By systematically analyzing these false narratives, public health authorities can better tailor their pandemic prevention, preparedness, and response strategies. Methods: In light of the emerging threat of H5N1, we analyzed the content of social media posts from Facebook (approximately 350,000) and Instagram (n=69,551) related to HPAI. Using 40 keywords associated with misinformation, we identified over 500 posts explicitly mentioning H5N1 and related terms for further systematic analysis. Posts were coded to identify targets and topics in the social media narratives. The “target” refers to the organization or person mentioned in the post, while the “topic” refers to the primary issue or subject being addressed. Results: Our content analysis identifies 7 main targets of misinformation, including government (149/544, 27%), health authorities (108/544, 20%), and international organizations (74/544, 14%). Also, from the 6 topics that have been identified, we found that the most widespread one was that authority figures purposefully engineer pandemics to achieve multiple political, economic, and other objectives (362/544, 67%) followed by societal destruction (121/544, 22%), and anti-vaccination (84/544, 15%). Other themes include societal destruction and religious allusions and prophecies. Conclusions: Our analysis of online content showed that H5N1 misinformation was primarily aimed at individuals or groups with differing degrees of political or institutional authority, such as government leaders and public health officials. These figures were often the focus due to their involvement in making health policy decisions and implementing public health measures. Decision-making entities and individuals were the target of various misinformation narratives. Results demonstrate the ongoing need for monitoring health misinformation to inform evolving public health responses to HPAI.
  •  
❌