❌

Normal view

Received — 4 April 2026 ⏭ Journal of Medical Internet Research

Artificial Intelligence, Connected Care, and Enabling Digital Health Technologies in Rare Diseases With a Focus on Lysosomal Storage Disorders: Scoping Review

Background: Rare diseases affect more than 300 million people globally, and only about 5% have approved therapies. Lysosomal storage disorders (LSDs) exemplify the diagnostic and long-term care complexity typical of rare diseases, and digital health technologies (DHTs), especially artificial intelligence (AI) and connected care (CC), are emerging tools to support LSD management. Objective: We aimed to map and synthesize peer-reviewed and gray literature from the past decade on DHTs relevant for LSD care, with a primary analytic focus on AI-enabled and CC solutions and a contextual mapping of other enabling DHTs. Evidence distribution was charted by population, care-journey phase, and outcome domains to identify gaps, methodological limitations, and timely priorities relevant for research, clinical practice implementation, and policies. Methods: We conducted a scoping review guided by a population, concept, context framework and operationalized through a Population, Intervention, Comparison, and Outcome (PICO)-informed data-charting structure to map study characteristics and reported outcomes, without causal or effectiveness assumptions and without risk-of-bias assessment. We searched PubMed, Google Scholar, and ClinicalTrials.gov for studies published between October 2015 and September 2024, complemented by AI-assisted discovery tools for citation extension. Reproducibility logs (search strings, run dates, filters, and stepwise counts) were maintained. Of 1751 records retrieved, 245 were included. Evidence was charted by LSD population, intervention class (AI, CC, and other enabling DHTs), outcome domains (patient, health care, and societal), and phase of the care journey. Results: Among 245 included records, 92.2% (226/245) were peer-reviewed, and 7.8% (19/245) were gray literature; no completed and published randomized controlled trials or LSD-specific systematic reviews were identified, with evidence dominated by small, single-center observational studies. Overall, 40 peer-reviewed records reported AI-driven DHTs, 89 reported CC DHTs, and 144 reported other enabling DHTs (some multilabeled). Evidence was concentrated mostly in Gaucher and Fabry diseases. Nearly half of the mapped literature focused on screening and diagnosis, with fewer records addressing treatment intensification, rehabilitation, and end-of-life care. Outcomes were predominantly health care delivery performance measures, with fewer patient and societal outcomes. AI applications mainly supported diagnostic decision support, phenotyping, monitoring, tracking, and risk stratification; CC commonly involved telemedicine, remote monitoring, and patient-engagement platforms; enabling DHTs included interoperable data systems, registries, and digital infrastructures. Conclusions: The evidence base is appreciable for a niche field and reflects growing interest in AI and CC for LSD care, but heterogeneity and methodological limitations preclude inferences on effectiveness or routine implementation. This evidence map highlights relatively stronger areas and gaps, providing a structured foundation to inform timely expert consensus-building and research prioritization. Key priorities include interoperable data infrastructures and data availability, prospective multicenter evaluations, transparent reporting of algorithms and workflows, and implementation-relevant outcomes to support safe, equitable, and scalable adoption aligned with evolving European Union and global rare-disease priorities.

Predictive Value of Machine Learning for Poststroke Mortality Risk: Systematic Review and Meta-Analysis

Background: People with stroke face a high mortality risk, and an accurate prediction model is essential to the guidance of clinical decision-making in this population. Recently, with growing attention paid to machine learning (ML) in stroke care, some researchers have investigated the effectiveness of ML in predicting the mortality risk in stroke. However, systematic evidence is still lacking for its effectiveness. Objective: This systematic review aims to evaluate the value of ML in predicting the stroke mortality risk. The findings are expected to offer an evidence-based basis for developing and assessing clinical risk prediction tools. Methods: A search was made in Cochrane Library, PubMed, Embase, and Web of Science up to June 23, 2025, and studies that reported a complete performance of ML in predicting stroke mortality were included. Studies with only risk factors analyzed were excluded. The risk of bias of the included studies was assessed using PROBAST (Prediction model Risk of Bias Assessment Tool). Pooled risk ratios with 95% CIs and prediction intervals (PIs) were derived using the Hartung-Knapp-Sidik-Jonkman method under a random-effects model. Subgroup analyses were also conducted by model type, stroke type, patient source, and treatment background. Moreover, a metaregression was conducted on the C-index for out-of-hospital mortality at different time points to explore the influence of time factors on the model’s predictive performance. Results: Sixty-eight studies were included (23 predicting in-hospital mortality and 45 predicting out-of-hospital mortality), describing the development of 75 prediction models and 43 external validations. The follow-up period was 1 month to 15 years. For predicting in-hospital mortality, the external validation set had a pooled C-index of 0.727 (95% CI 0.677-0.781, 95% PI 0.521-1.000), with sensitivity and specificity of 0.64 (95% CI 0.57-0.70) and 0.74 (95% CI 0.70-0.77), respectively. For predicting out-of-hospital mortality, the pooled C-index was 0.847 (95% CI 0.808-0.887, 95% PI 0.750-0.956) in the external validation set, with sensitivity and specificity of 0.71 (95% CI 0.55-0.82) and 0.76 (95% CI 0.74-0.78), respectively. Comparatively, the overall pooled C-indexes were 0.788 (95% CI 0.766-0.810, 95% PI 0.621-0.999) and 0.812 (95% CI 0.798-0.826, 95% PI 0.693-0.952), respectively. The metaregression revealed a gradual decline in the predictive performance of the overall model and logistic regression model alone, whereas a random forest model maintained sustained performance. Age, National Institutes of Health Stroke Scale score, and stroke-related complications were the most frequently used variables for modeling. Conclusions: This is the first meta-analysis to demonstrate that ML-based prediction of stroke mortality is feasible. The performance of ML supports its role as an auxiliary tool for identifying high-risk populations, thereby optimizing clinical monitoring and resource allocation. However, due to substantial heterogeneity and a relatively high risk of bias in available studies, caution is warranted in real-world application. The effectiveness of ML may vary across settings, and external validation is recommended before broader implementation. Trial Registration: PROSPERO CRD420251086321; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251086321

Psychotherapists’ Trust, Distrust, and Generative AI Practices in Psychotherapy: Qualitative Study

Background: Generative artificial intelligence (GenAI) is increasingly used in mental health care, from client-facing chatbots to clinician-facing documentation aids. Psychotherapists’ willingness to rely on—or withhold reliance from—these tools has significant implications for care quality, yet little is known about how practicing clinicians calibrate trust and distrust in GenAI across tasks and contexts. Given that the therapeutic relationship is central to psychotherapy outcomes, understanding how GenAI intersects with this relational foundation is essential for responsible integration. Objective: This study aims to examine (1) psychotherapists’ experiences with, perceptions of, and trust or distrust in GenAI in therapeutic contexts and (2) how they perceive the role of GenAI within the therapeutic relationship and how their perceptions shape their trust and distrust in GenAI. Methods: We conducted a qualitative interview study using semistructured interviews with 18 actively practicing psychotherapists in the United States between January and May 2025. Participants were recruited through professional mailing lists, social media, and snowball sampling. Interviews (≈60 min each) were conducted via Zoom and explored psychotherapists’ experiences with, perceptions of, and trust or distrust in GenAI in therapeutic contexts. Data were analyzed using the general inductive approach, with iterative coding and team-based interpretation to identify themes. Results: Our findings show that psychotherapists’ GenAI adoption was highly individualized and contingent on maintaining professional role integrity—not merely technical oversight. Trust was sustained when GenAI operated in clinician-supervised, supportive roles for low-stakes tasks (eg, documentation and brainstorming), but diminished when control shifted, tasks involved high-stakes clinical judgment, or GenAI threatened to encroach on the authentic human connection central to therapy. Participants articulated conditions for trust that went beyond “human-in-the-loop” monitoring to include preservation of interpretive authority, ethical responsibility, and relational primacy. Distrust also extended to the broader sociotechnical ecosystem, including concerns about commercial incentives, insurance pressures, and the absence of clear organizational guidelines. Conclusions: Psychotherapists’ perspectives offer critical insights into GenAI’s current usages in their professional practices and the conditions under which they are willing to trust and distrust GenAI tools. Their experiences highlight the importance of maintaining clinician control, ensuring contextual appropriateness, and preserving the human connection central to psychotherapy. Future work should further examine how therapeutic orientation, professional experience, and client characteristics shape trust and distrust in GenAI. As GenAI becomes more embedded in mental health care, research is also needed to explore how specific GenAI system features can be responsibly designed to support clinical workflows and enhance therapeutic relationships. Organizational and policy frameworks will be essential to ensure responsible, ethically aligned, and human-centered GenAI deployment in psychotherapy.

Accuracy of Radiomics-Based Machine Learning for Predicting Risk of Recurrence in Non–Small Cell Lung Cancer: Systematic Review and Meta-Analysis

Background: During the diagnosis and treatment of non–small cell lung cancer (NSCLC), detecting the risk of its recurrence in an early phase is still challenging. Recent studies have investigated the radiomics-based machine learning (ML) models for detecting the risk of recurrence in NSCLC. However, there is still insufficient systematic evidence to prove its efficiency. Objective: This study is designed to systematically evaluate the effectiveness of radiomics-based ML in predicting the risk of recurrence in NSCLC, aiming to provide evidence-based support for the subsequent development of scoring tools to forecast recurrence risk. Methods: For acquiring research on radiomics-based models for forecasting the risk of recurrence in NSCLC, Cochrane Library, Web of Science, PubMed, and Embase were systematically retrieved, up to October 24, 2025. Studies on analyzing the recurrence of NSCLC using radiomics-based ML were included, while those in which only texture analysis was conducted or radiomics-based ML was not constructed were excluded. The Radiomics Quality Score (RQS) was used to appraise the eligible studies. Subgroup analyses were conducted according to the variables of the model, the background of treatment, the stage of lung cancer, and the pathological type. Results: Ultimately, 30 eligible studies in total were included, covering 7964 patients with NSCLC. According to the meta-analysis, the c-index of radiomics-based ML models for forecasting the risk of recurrence in NSCLC was 0.850 (95% CI 0.834‐0.866, 95% prediction interval [PI] 0.623‐1.004) in the training set. Specifically, the pooled c-index was 0.876 (95% CI 0.853‐0.900) among the patients receiving the stereotactic body radiation therapy and 0.825 (95% CI 0.804‐0.848) among those who received surgeries combined with other adjuvant treatment regimens. The c-index of the radiomics-based ML models combined with clinical features for forecasting the risk of recurrence in NSCLC was 0.833 (95% CI 0.822‐0.854, 95% PI 0.717‐0.945) in the training set. In contrast, the c-index of radiomics-based ML models for forecasting the risk of recurrence in NSCLC was 0.878 (95% CI 0.854‐0.902, 95% PI 0.681‐1.000) in the validation set. The c-index of radiomics-based ML models combined with clinical features for forecasting the risk of recurrence in NSCLC was 0.854 (95% CI 0.830‐0.878, 95% PI 0.655‐0.992) in the validation set. The average RQS across the included studies was 27.4%, revealing methodological limitations and an absence of standardization. Conclusions: This study is the first to confirm that radiomics-based ML models effectively predict the risk of recurrence in NSCLC. This study provides evidence-based support for the subsequent development or updating of radiomics-based ML models. However, the current methodological application of radiomics remains concerning. Therefore, in the future, research should standardize the workflow for implementing radiomics-based ML and incorporate multicenter imaging data to enhance its generalizability. Trial Registration: PROSPERO CRD42025631191; https://www.crd.york.ac.uk/PROSPERO/view/CRD42025631191

Strategy for Hepatitis B and C Virus Testing Campaigns Through Web Services and Digital Advertising in Japan: Nationwide Cross-Sectional Study With Correspondence Analysis

Background: Public awareness campaigns and testing promotion must be strengthened to eliminate infections with hepatitis B and C viruses (HBV and HCV, respectively) by 2030. Although public health campaigns using various forms of advertising are widely implemented, the most appropriate channels for viral hepatitis testing remain unclear. Objective: This study aims to identify web services and digital advertising channels appropriate for promoting HBV and HCV testing, segmented by prior testing history and the desire for hepatitis virus testing. Methods: A nationwide cross-sectional online survey of Japanese adults aged 20 to 69 years was conducted. The respondents answered questions regarding viral hepatitis testing status, routinely used web services (180 options), and exposure to digital advertising (25 options). Correspondence analysis was used to visualize relationships among testing segments, web services, and digital advertising. For individuals classified as “never having been tested and wishing to be tested,” channel-specific alignment was quantified using cosine θ. Sensitivity analyses were conducted by repeating the correspondence analysis after excluding respondents uncertain about their testing history and by fitting modified Poisson regression models with robust variance to estimate prevalence ratios and 95% CIs. Results: Of the 2000 respondents (1011 male and 989 female), 18% (n=359) reported prior HBV and HCV testing, and 22.1% (n=441) were unsure whether they had ever been tested. Web services characteristically associated with “never having been tested and wishing to be tested” included Lawson (convenience store: cosine =0.989) and Cosme (shopping: cosine =0.987). The corresponding digital advertising channels included in-store and storefront screens at Welcia (pharmacy chain: =0.994) and Lawson (cosine =0.937). Segment-specific patterns varied according to age group and sex. Sensitivity analyses excluding the unsure group showed similar patterns. Modified Poisson regression results were also consistent; for example, Lawson web service use was associated with a desire for hepatitis virus testing (prevalence ratio 1.75, 95% CI 1.22‐2.52). Conclusions: In Japan, the convenience store chain Lawson was a frequently used touchpoint, both online and offline, among individuals seeking viral hepatitis testing. Future studies are needed to determine whether implementing awareness-raising activities through Lawson can increase the uptake of testing and subsequent treatment.

The Role of Digital Biomarkers in Physiological Signal-Based Depression Assessment: Systematic Review and Meta-Analysis

Background: Digital biomarkers are increasingly being used to support depression assessment by providing objective, continuous, and real-time physiological and behavioral data. However, most existing studies have focused on individual biomarkers, such as sleep or cardiac parameters, while integrative evaluations that capture the multidimensional nature of depression remain limited. Objective: This systematic review evaluated digital biomarkers for depression and synthesized evidence on differences between individuals with depression and controls. Methods: Eligible studies included observational or interventional studies examining digital biomarkers for depression with validated outcome measures. We searched major international and Korean databases, including MEDLINE, PsycINFO, CINAHL, IEEE Xplore, Web of Science, Cochrane Library, KISS, RISS, KMbase, and KoreaMed, from inception to December 28, 2025. Risk of bias was assessed using the Quality Assessment of Diagnostic Accuracy Studies-2 tool and the Scottish Intercollegiate Guidelines Network checklist. Meta-analyses were conducted using random-effects models with the Hartung-Knapp-Sidik-Jonkman method, and other outcomes were narratively summarized. Results: The search yielded 39,617 records, of which 132 studies involving 57,852 participants met the inclusion criteria. These studies encompassed various digital biomarkers, including sleep, physical activity, cardiac measures, smartphone-derived data, speech, GPS data, and circadian rhythms. A meta-analysis of 22 studies (6947 participants) revealed that individuals with depression had significantly longer sleep onset latency (5 studies; n=292; +4.75 min, 95% CI 2.46-7.04; =.005; 95% prediction interval [PI] 0.01-10.27) and time in bed (3 studies; n=236; +31.81 min, 95% CI 18.22-45.39; =.01; 95% PI 2.28-55.16). Physical activity counts were also significantly lower (5 studies; n=462; standardized mean difference −0.71, 95% CI −1.33 to −0.09; =.03; 95% PI −2.18 to 0.71). Although individuals with depression showed a lower sleep efficiency, higher mean heart rate, and lower SD of normal-to-normal intervals, these differences were not statistically significant. Other digital markers yielded inconsistent results. Overall, these findings indicate that no single digital biomarker sufficiently captures depression-related changes. Instead, the results support the superiority of personalized, multimodal approaches. However, the generalizability of these findings is limited by the lack of standardized data collection protocols and high clinical heterogeneity across studies, as reflected in wide PIs. Conclusions: Certain digital biomarkers, particularly sleep onset latency and physical activity counts, showed consistent average differences between the depression and control groups. However, wide PIs indicate substantial variability across settings, suggesting that no single marker is sufficient for reliable detection. This study advances the field by providing a comprehensive meta-analysis of multidimensional digital biomarkers, establishing a quantitative foundation for objective depression screening and monitoring. These findings support the use of personalized, multimodal digital phenotyping approaches and highlight the need for standardized, clinically interpretable frameworks for real-world depression monitoring. Trial Registration: PROSPERO CRD42024518136; https://www.crd.york.ac.uk/PROSPERO/view/CRD42024518136

Effectiveness of the Components of a Digital Multiple Health Behavior Change Intervention Among Individuals Seeking Help Online (Coach): Factorial Randomized Trial

Background: Extant digital multiple health behavior change interventions have shown promise in various populations; however, evidence for a broader approach among the general population is lacking. Moreover, existing interventions often contain several components but are typically assessed as a whole, meaning it remains unclear to what extent individual components contribute to intervention effects and how they may interact to influence health outcomes. Objective: This study estimates the effects of 6 components of a digital health behavior change intervention on alcohol, diet, physical activity, and smoking outcomes among individuals searching for help online. Methods: A double-blind randomized factorial trial design with 6 two-level factors was used. Adults from the general public in Sweden who were seeking help to change their behaviors were recruited through web searches and social media. Participants were eligible if they were 18 years or older and had at least one health behavior classified as unhealthy. Effects of 6 components were estimated: screening/feedback, goal-setting/planning, motivation, skills/know-how, mindfulness, and self-authored SMS text messages. Primary outcomes were weekly alcohol consumption and frequency of heavy episodic drinking, average daily fruit and vegetable consumption, weekly moderate-to-vigorous physical activity, and 4-week point-prevalence smoking. Results: A total of 5419 individuals were randomized. Overall, the screening/feedback component was the most effective for changing health behaviors, along with goal-setting/planning and motivation to change. In particular, there was evidence that screening/feedback increased average daily portions of fruit and vegetables at 2 months (mean difference 0.17, compatibility interval [CoI] 0.09-0.25, probability of effect [POE] >99.9%) and at 4 months (mean difference 0.13, CoI 0.04-0.21, POE 99.9%) and reduced the frequency of heavy episodic drinking at 4 months (incidence rate ratio 0.91, CoI 0.81-1.03, POE 94.2%). Components also interacted to further improve health outcomes, most notably the combination of screening/feedback with motivation to change, which further increased fruit and vegetable consumption (2 months: mean difference 0.20, CoI 0.09-0.30, POE >99.9%; 4 months: mean difference 0.17, CoI 0.05-0.29, POE 99.8%). Conclusions: The results from this study contribute to the development of more effective interventions by providing novel insights into the effects of individual and pairwise components of complex digital health behavior change interventions. Trial Registration: ISRCTN Registry ISRCTN16420548; http://www.isrctn.com/ISRCTN16420548

Artificial Intelligence in Health Professions Education: Qualitative Study of Student Experiences

Background: Artificial intelligence (AI) is increasingly integrated into education and health care, raising questions about how students use these technologies and how AI influences their learning. In health education, understanding these trends is particularly important because student learning directly impacts future clinical skills. Objective: This study aimed to explore the use of AI tools by health sciences students at the University of Ottawa. More specifically, it sought to identify the most frequently used AI tools, describe students’ usage habits, determine which tools support knowledge acquisition and skill development, and gather students’ recommendations for effective strategies to raise awareness and train their peers on the responsible use of AI. Methods: A qualitative approach was used with students from 10 health professions who reported using AI in their studies. Data were collected through semistructured interviews and an open-ended qualitative online survey. Inductive thematic analysis within an interpretive paradigm was applied to capture patterns, perceptions, and emergent themes. Results: A total of 51 health professions students participated in the study. Most were women between the ages of 20 and 29 years. ChatGPT (OpenAI) emerged as the most frequently used AI tool. Students perceived AI as a complementary tool that facilitated knowledge acquisition, skill development, writing, and problem-solving. AI adoption was driven by curiosity, peer influence, and the desire to improve work efficiency. Students critically evaluated AI results, integrated the tools into their learning processes, and emphasized the importance of technical skills, critical thinking, and digital literacy. Peer learning, hands-on demonstrations, and access to online resources were recommended for effective AI training. Conclusions: This research demonstrates that health professions students actively use AI tools, particularly ChatGPT, to support learning, skill development, and academic tasks. Although AI is valuable as an educational aid and its use varies by student and context, this highlights the need for structured guidance, critical evaluation skills, and peer-supported training. These findings highlight the importance of thoughtfully integrating AI into educational programs to enhance learning outcomes, foster skill acquisition, and ensure responsible and effective adoption.

Well-Being and Cognitive Factors Influencing Health Care Workers’ Adherence to Internet-Based Stress Management: Mixed Methods Analysis of a Nonrandomized Controlled Study

Background: High stress levels are common among health care workers (HCWs), threatening their health and workforce stability. Internet-based mobile stress management (MSM) is a promising intervention for reducing work-related stress; however, poor adherence limits effectiveness. Exploring factors influencing HCWs’ adherence may thus aid in developing optimal interventions. Objective: The research aimed to investigate (1) how HCWs’ well-being and cognitive factors influenced MSM treatment adherence and (2) what HCWs’ specific needs for MSM were. Methods: This study was a convergent mixed methods secondary analysis of a nonrandomized controlled trial. HCWs who were currently employed, had internet access, had no serious medical problems, and were willing to participate were recruited by convenience sampling through an MSM project in a large Chinese general hospital from August 11, 2021, to January 31, 2022. Those intending to leave the hospital or with insufficient medical condition for follow-up were excluded. Quantitative data were collected from 157 HCWs (n=135, 86% female participants; mean age of 33.7, SD 4.9 y) electronically via Research Electronic Data Capture (REDCap). Measures included sociodemographic characteristics, the Fatigue Assessment Scale, the 14-item Perceived Stress Scale, a user experience questionnaire, an attitudes scale (perceived usefulness, feasibility, and enjoyment), and self-reported practice frequency. Qualitative data were collected via an open-ended question answered by 96 participants. Quantitative data were analyzed using hierarchical regression and structural equation modeling. Qualitative data were analyzed using reflexive-thematic analysis in NVivo (QSR International). Results: In the quantitative study (n=157), hierarchical regression analyses showed that fatigue was a significant negative predictor of adherence (b=−0.050, 95% CI −0.086 to −0.015; =−2.859; 2-tailed =.005), while user experience (b=0.074, 95% CI 0.042-0.106; =4.569; 2-tailed

Ethical Handling of Occupational Health and Safety Data in the Fire Service: Empirical Interview and Focus Group Study of Firefighter and Fire Service Leadership Privacy Preferences

Background: There are ongoing efforts to collect larger and higher-quality amounts of occupational health and safety data to better understand and prevent injuries and fatalities among high-risk workers, such as firefighters. Digital health systems including wearable technologies, mobile apps, or internet-based data collection platforms could collect large amounts of sensitive data, but there is little evidence on worker and employer perspectives on data privacy in the fire service. Objective: Our study examined firefighters’ and fire service leadership’s preferences regarding occupational health and safety data privacy. Methods: We conducted interviews and focus groups with career firefighters in Maryland and Virginia; interviews with union representatives and department-level leaders in each state; and interviews with national-level fire service leaders in advocacy, government, and research organizations (March to November 2023). Interviews and focus groups were audio recorded and transcribed. We analyzed transcripts using thematic analysis. Results: The sample included 31 career firefighters, 2 union leaders, 11 national leaders, and 21 department-level leaders (65 total participants from 35 interviews and 4 focus groups). We identified 4 themes: acceptability of data access, sharing, and reporting practices; data sharing and access preferences; appropriate use of firefighter data; and the need for improved communication. Leaders described firefighters’ concerns about job loss and loss of privacy. Firefighters expressed general preferences that their data be deidentified and not shared widely, and they identified mental health data as important but particularly sensitive information. Firefighters also expressed frustration about sharing data with researchers or their departments without knowing the purpose or outcomes. Both firefighters and leaders emphasized the need for enhanced communication and translation of data for firefighters. Conclusions: Fire service leaders held more concerns about the use and sharing of occupational health and safety data than firefighters, but both groups identified ways to further safeguard firefighter data and improve communication about health and safety data. Future fire service data collection should incorporate privacy protections, such as limiting the collection of identifiable information and restricting data access. Data collection should be accompanied by clear communication about the purpose of the data collection, how firefighter data will be used and accessed, and the interpretation of the results. Future digital health interventions should integrate these data privacy protections to respect firefighter preferences and contribute to acceptability and uptake.
❌