❌

Normal view

Received — 26 March 2026 ⏭ Journal of Medical Internet Research

Multimodal AI for Alzheimer Disease Diagnosis: Systematic Review of Datasets, Models, and Modalities

Background: Early detection of Alzheimer disease (AD) is essential for timely intervention; yet, diagnostic performance varies widely across modalities and datasets. Recent multimodal artificial intelligence (AI) models have made significant progress, but the evidence base remains fragmented due to heterogeneous datasets, modeling frameworks, and reporting quality. Objective: This systematic review aimed to analyze studies on multimodal AI models for AD diagnosis, prognosis, and risk prediction over 5 years. We evaluated dataset characteristics, modality combinations, modeling strategies, performance metrics, and methodological limitations. We further discuss real-world implications and translational pathways. Methods: Following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 guidelines, we systematically searched PubMed, IEEE Xplore, Scopus, ACM Digital Library, Cochrane, and arXiv, with the final datasets last searched on November 15, 2025. Studies applying multimodal machine learning or deep learning to AD, mild cognitive impairment, and dementia outcomes were included, whereas studies using a single modality or lacking sufficient methodological detail were excluded. QUADAS-2 (Revised Quality Assessment of Diagnostic Accuracy Studies tool) assessed risk of bias. Extracted performance results were synthesized across 4 major multimodal dataset families. Results: A total of 66 studies met the inclusion criteria. Across datasets, multimodal models consistently outperformed single-modal baselines. Alzheimer’s Disease Neuroimaging Initiative–based diagnosis achieved an average accuracy of 92.5% (SD 3.8%), while mild cognitive impairment–conversion models achieved an average area under the curve (AUC) of 0.922 (SD 0.045), and several fusion architectures reported AUCs above 0.95. In contrast, UK Biobank risk-prediction studies reported an average AUC of 0.84 (SD 0.056), and this reflects performance in large, population-based datasets. DementiaBank speech-language studies achieved an average AUC of 0.813 (SD 0.042), and cross-lingual AD detection achieved an accuracy of 77% (SD 6.5%). Self-collected multimodal datasets demonstrated average accuracies around 96% (SD 2.4%), but their generalizability is limited due to small sample sizes and single-center designs. Conclusions: This systematic review demonstrates that multimodal AI models consistently outperform single-modal models for AD diagnosis, prognosis, and risk prediction by integrating complementary biological, clinical, and behavioral information. Unlike prior reviews, this review provides a unified synthesis across heterogeneous clinical, imaging, genetic, and linguistic datasets, enabling cross-domain comparison of modeling strategies and performance. However, the generalizability of reported performance was limited due to substantial heterogeneity in dataset composition, outcome definitions, and validation, and prevalent risks of bias. By evaluating these factors, this review clarifies where current evidence is robust and where caution is warranted. The findings highlight the need for standardized multimodal benchmarks, transparent evaluation protocols, and clinically grounded model design to enable reliable real-world deployment. Overall, this work advances the field by framing multimodal AI not only as a performance-driven tool but also as a translational framework for equitable, interpretable, and scalable AD diagnosis. Trial Registration: PROSPERO CRD420251241895;

Suicidal Thoughts and Behaviors Among Chinese Adolescents in Relation to Negative Life Events, Internet Addiction, and Sexual Abuse: Cross-Sectional Study

Background: Increasing suicidal thoughts and behaviors (STB) among adolescents raise social concerns and have a well-recognized association with sexual abuse (SA). However, research regarding the mechanisms explaining the association between SA and STB remains limited. Objective: This study aims to examine the chained mediating effects of negative life events (NLE) and internet addiction (IA) between SA and STB among adolescents in China. Methods: This cross-sectional study used data from the Science Database of the People Mental Health survey conducted between March 2013 and December 2022 by the National Population Health Data Center of the National Research Institute for Family Planning. Through stratified sampling, 20,893 adolescents were recruited from 16 Chinese provinces. After excluding samples with missing relevant variables, 10,664 (55.89%; aged 16-17.9 y; n=5826, 54.63% women) adolescents were included in the final analysis. STB was the outcome variable, with NLE and IA as mediators, all assessed via a questionnaire that was uniformly administered by trained investigators in school settings. The Pearson χ test was used to analyze the association between SA and STB. Using a combination of multiple linear regression and bootstrap testing, the study constructed a chain mediation model to explore how SA influences STB in adolescents through NLE and IA. Results: The scores for SA, NLE, IA, and STB were 1.330 (SD 1.714), 51.960 (SD 23.822), 34.88 (SD 13.852), and 0.690 (SD 1.396), respectively. Multiple linear regression analysis indicated SA was associated with NLE (β=2.382, 95% CI 2.112‐2.653;

Willingness to Share Internet Use Data for Research on Early Disease Detection: Cross-Sectional Survey

Background: Preliminary research has suggested that internet use data could offer digital signals of early disease and has the potential to facilitate early detection and improve patient outcomes. However, there are significant challenges in linking individual-level internet use data with health outcomes. One key aspect is that the public might not be willing to share data for research or that selective data sharing might create bias in datasets and increase inequalities. Objective: Our study aimed to investigate the willingness of the public to share their internet use data for medical research and to identify key criteria that affect willingness to share. Methods: We conducted a web-based, cross-sectional online survey with 2390 UK adults with and without a history of cancer, heart disease, and depression using quota sampling. Participants were randomly assigned to explore willingness to share different types of internet use data for 1 of 3 health conditions (cancer, heart disease, and depression) and for provision of a pictorial example of internet use data. Logistic regression analysis (α=.05) for each condition was used to determine key factors of willingness to share, including sociodemographics and attitudes toward sharing. Open-ended comments regarding facilitators of sharing and concerns were analyzed thematically. Results: Willingness to share internet use data was high across conditions (74%‐77%, 95% CI 70.5%‐80.3%), especially for health app data (73%‐76%, 95% CI 69.8%‐79.1%). The pictorial example of browsing history did not affect willingness to share. For all conditions, factors consistently associated with willingness to share were perceived benefits (odds ratios [ORs] 5.692‐8.850; all

Determinants of the Uptake and Frequency of Use of a Web Portal Digital Health Intervention in Patients With Type 2 Diabetes and/or Coronary Heart Disease: Secondary Analysis of a Randomized Controlled Trial

Background: The targeted application and design of digital health interventions (DHIs) require an understanding of usage determinants. Usage includes uptake (initial use) and frequency (extent of use), but it is unclear whether both components are driven by the same determinants. Objective: This study aimed to examine the determinants of uptake and frequency of use and assess whether they differ. Methods: The investigated DHI was a web portal provided in an intervention for improving disease-related self-management. This study is a secondary analysis of intervention group data from a parallel-group randomized controlled trial. Eligibility criteria were being an adult and being diagnosed with type 2 diabetes and/or coronary heart disease. Sociodemographic, psychological, and health-related variables were examined as determinants. Determinants were analyzed using simple and multiple regression models. Uptake was analyzed using logistic regression, and frequency was analyzed using negative binomial regression with robust SEs. Frequency was analyzed for those who used the DHI at least once. Except for sociodemographic variables, all other variables were standardized to a range from 0 to 1. For simple regression, inflation of the α error due to multiple testing was controlled via the approach of Benjamini and Hochberg, and for multiple regression, it was controlled via the significance of the complete multiple regression model. Results: Of 462 intervention group members, 199 (43.1%) used the web portal at least once. After controlling for inflation of the α error, simple regression for uptake yielded significant effects for higher education (B=0.56, 95% CI 0.18-0.95; =.004), openness (B=1.08, 95% CI 0.33-1.83; =.005), intention regarding physical activity (B=2.28, 95% CI 1.30-3.26;

Improving Retrieval Augmented Generation for Health Care by Fine-Tuning Clinical Embedding Models: Development and Evaluation Study

Background: Embedding models are critical components of Retrieval Augmented Generation (RAG) systems for retrieving and searching unstructured medical data. However, existing models are predominantly trained on publicly available English datasets, limiting their effectiveness in non-English health care settings. More importantly, these models lack training on real-world clinical documents, leading to inaccurate context retrieval when integrated into RAG systems for health care applications. This gap is particularly pronounced in specialized medical documentation containing domain-specific terminology, abbreviations, and nuanced clinical language. Objective: This retrospective study aimed to develop and validate embedding models specifically trained on real-world clinical documents from multiple medical specialties to improve medical information retrieval (IR) and RAG system performance in both German and English language contexts. Methods: We fine-tuned embedding models, so-called sentence transformers, using the multilingual-e5-large architecture as a foundation. Training data consisted of approximately 11 million question-answer pairs synthetically generated from 400,000 diverse clinical documents from a large German tertiary hospital, spanning 163,840 patients and 282,728 clinical cases between 2018 and 2023. The large language model generated medically relevant questions and corresponding answers for each document. The dataset was additionally pseudonymized and translated into English to aim for broader applicability. Models were evaluated in 2 distinct scenarios: IR using questions with multiple relevant passages, and RAG system performance in both cross-patient and patient-centered contexts. Results: In the IR evaluation, the fine-tuned miracle model achieved a mAP@100 of 0.27, outperforming the multilingual-e5-large baseline (0.14) and state-of-the-art models such as bge-m3 (0.11). In the RAG evaluation, the model demonstrated robust performance comparable with the baseline in the constrained patient-centered scenario (BERTScore F1 0.781 vs 0.778) and showed moderate improvements in the unconstrained cross-patient setting (BLEURT 0.56 vs 0.53). Notably, the model trained on pseudonymized data achieved comparable retrieval performance (mAP@100 0.25) and the highest scores for patient-centered contextual precision (0.93). Performance gains were robust in the German dataset, while the translated English model demonstrated promising results as a proof of concept for cross-lingual transfer. Conclusions: By leveraging a comprehensive real-world dataset spanning multiple medical specialties and using large language models for synthetic question generation, we successfully created and validated domain-specific embedding models. These models can improve medical IR in large-scale search spaces and perform competitively in constrained RAG applications. By publishing the models trained on pseudonymized data, other health care institutions can integrate or adapt these embedding models to their needs. This work establishes a reproducible framework for developing domain-specific clinical embedding models, with the potential to improve data retrieval in medical settings.

Robot-Assisted Therapy for Upper Limb Rehabilitation After Stroke: Umbrella Review

Background: Stroke is a leading cause of long-term upper limb disability, severely impacting patients’ independence and quality of life. Robot-assisted therapy (RAT) has emerged as a promising, high-intensity rehabilitation alternative. However, conclusions from existing systematic reviews on its efficacy are inconsistent and often lack a holistic framework, limiting their use for guiding personalized clinical decisions. Objective: This study aims to systematically synthesize recent evidence on RAT for upper limb rehabilitation after stroke. Guided by the International Classification of Functioning, Disability and Health framework, it moves beyond singular outcomes to provide a multidimensional evaluation across body function, activity, and participation levels. The review aims to provide stratified guidance for clinical decision-making based on patient- and intervention-specific characteristics, thereby supporting evidence-based practice and informing future research. Methods: This study included systematic reviews and meta-analyses published from January 1, 2019, to December 26, 2025, comparing RAT with conventional therapy for upper limb rehabilitation after stroke. Overall, 6 databases, including PubMed, Web of Science, and Embase, were searched. Two reviewers (XZ and LZ) independently performed study selection, data extraction, and quality assessment using the AMSTAR 2 tool. The synthesis integrated outcome measures and subgroup analyses derived from the included studies. Results: This umbrella review included 21 meta-analyses encompassing 535 randomized controlled trials and 27,598 patients across acute, subacute, and chronic stroke stages. According to AMSTAR 2, 17 reviews were high quality, 3 moderate, and 1 critically low. The synthesis demonstrated that RAT was superior in improving upper limb motor function, but no statistically significant advantages were observed in activities of daily living compared to conventional therapy. Subgroup analyses revealed that treatment effects were influenced by stroke stage, upper limb motor impairment level, and robot type. Conclusions: RAT is an effective intervention for improving upper limb motor function after stroke. However, its benefits are primarily observed at the level of body function, with limited evidence for long-term maintenance. The current evidence is constrained by significant outcome heterogeneity and methodological limitations inherent to umbrella reviews. Future research should validate these findings in broader clinical practice, focus on translating functional gains into sustained improvements in daily activities and participation, and include cost-effectiveness evaluations. Trial Registration: PROSPERO CRD42024497183; https://www.crd.york.ac.uk/PROSPERO/view/CRD42024497183

Challenges of Standard Pediatric Epilepsy Monitoring and the Potential Benefits of Contactless Sensor Technologies: Exploratory Qualitative Study

Background: Epilepsy is a common neurological condition in children, and accurate detection of seizures and their frequency is essential for diagnosis and treatment. Standard monitoring using electroencephalography alongside clinical observation is often burdensome in pediatric settings, as electrodes can cause discomfort and restrict mobility. Contactless sensor technologies may offer a promising supplement by enabling monitoring without physical contact. Objective: This study aims to explore challenges in standard pediatric epilepsy monitoring from the perspective of health care professionals and examines the potential benefits and requirements of supplementary contactless sensor technologies in this setting. Methods: Participant observation of routine processes in standard pediatric epilepsy monitoring was conducted at a German university hospital. Field notes from 40 observed procedures were analyzed using structuring content analysis. Building on these findings, a focus group with pediatric neurologists, nurses, and medical technical assistants (n=6) explored the potential benefits and implementation requirements of contactless sensor technologies. Focus group data were analyzed using focus group illustration maps. Results: A reference workflow of standard pediatric epilepsy monitoring was derived, revealing psychosocial, medical, and organizational challenges faced by health care professionals. Electroencephalography recordings and clinical observation required considerable reassurance of patients and parents or carers, were vulnerable to movement artifacts and incomplete seizure documentation, and were labor- and resource-intensive. Focus group participants viewed contactless sensor technologies as a potentially valuable supplement by enabling continuous long-term monitoring with minimal additional burden. Conclusions: By identifying challenges associated with standard pediatric epilepsy monitoring, this study provides a foundation for the needs-based development and implementation of supplementary contactless sensor technologies. Such technologies should be tailored to the clinical setting and designed to address existing burdens, with the potential to complement standard monitoring. Trial Registration: Deutsches Register Klinischer Studien (DRKS) DRKS00027017; https://drks.de/search/de/trial/DRKS00027017

Coproducing an Online Platform for People With Long-Term Physical Health Conditions: Development and Usability Study

Background: There is relatively limited psychological support dedicated to people living with long-term physical health conditions and subthreshold depressive disorder. Online peer support may be an appropriate intervention to help bolster patients’ mental well-being to prevent progression of their symptoms to major depressive disorder. For interventions to be successfully integrated into the self-management routines of people with long-term physical health conditions, they should be co-designed to ensure that they align with the wants and needs of the target audience. Objective: This study aims to coproduce an online peer support intervention with people with lived experience, software experts, clinicians, and academics through an iterative process of co-design and subsequent co-validation through usability testing. Methods: We followed a 4-stage coproduction process: co-assess, co-design, co-validate, and co-deliver. Our research advisory group was actively involved in all stages, consisting of 1 coinvestigator and 6 people with lived experience of long-term physical and/or mental health comorbidities. The co-assess and co-design stages involved our participatory design panel, which included 10 members living with various long-term conditions. The participatory design panel participated in online focus groups to assess their unmet psychosocial needs and then co-designed the intervention prototype through online workshops with software developers. The co-validation stage involved an additional group of participants (n=12) with long-term physical health conditions. During co-validation, the prototype underwent usability testing, including think-aloud exercises and semistructured interviews. Content analysis identified the priorities for the iterative development that formed the basis of further research advisory group co-design workshops. The next stage, co-delivery, involved coproducing the protocol of a feasibility and acceptability randomized controlled trial. Results: Participants highlighted that a platform must feel safe and trustworthy for the space to support the mental well-being of those living with long-term health conditions. The participatory design panel co-designed a platform prototype to meet this need. During the co-validation stage, the think-aloud exercises identified common issues related to navigation challenges and feature glitches. Content analysis of the semistructured interviews confirmed that the community forum, resources, and other platform pages were appropriate and acceptable, but revealed usability concerns. Participants stressed the need for intuitive navigation and suggested new features that would enhance user experience. Facilitators and barriers to engagement were also noted, including the importance of fostering trust in the platform’s ethos and branding to create a safe space. Through iterative development and subsequent usability testing, the final prototype was approved. Conclusions: We have provided a worked example of a comprehensive, coproduction process where we worked alongside people with lived experience to successfully design an online peer support platform with embedded psychoeducation. The platform, called CommonGround, is ready to be evaluated in a feasibility randomized controlled trial.

The Relationship Between Electronic Health Literacy and Health-Related Quality of Life Among Chinese Older Adults: Cross-Sectional Study

Background: The rapid digitalization of health care has reshaped access to medical services. However, older adults often remain disadvantaged due to the digital divide. Electronic health literacy (EHL) is increasingly recognized as a determinant of health-related quality of life (HRQoL); however, its mechanisms and subgroup differences in China remain underexplored. Objective: This study aimed to examine the association between EHL and multidimensional HRQoL among Chinese older adults, with a focus on the mediating roles of attitudes toward own aging (ATOA) and self-efficacy (SE), and heterogeneity by age, residence, and lifestyle. Methods: A cross-sectional survey (July-November 2024) included 8364 adults aged ≥55 years from 4 provinces using stratified multistage sampling. HRQoL was measured by physical health (PH), mental health (MH), and life satisfaction (LS). EHL was assessed with the eHealth Literacy Scale (eHEALS), ATOA with the Philadelphia Geriatric Center Morale Scale subscale, and SE with the General Self-Efficacy Scale. Analyses used seemingly unrelated regressions, PROCESS (Andrew F. Hayes) macro mediation with 5000 bootstraps, and subgroup regressions. Results: EHL was positively associated with PH (=0.273;

The Current Landscape of Remote Digital Symptom Monitoring for Patients With Lung Cancer: Scoping Review

Background: Remote digital symptom monitoring systems (rSMS) have been increasingly used in recent years to monitor symptoms, health-related quality of life, and other patient-reported outcomes in lung cancer. Previous studies have demonstrated variability in study design, types of rSMS, and outcomes used to assess benefits for patients and health care systems. However, there remains a lack of synthesized evidence pertaining to the similarities and differences among rSMS, including their theoretical underpinnings, key functional components, and reported benefits and limitations. Objective: This review aims to identify and synthesize existing research to map the current landscape of rSMS in lung cancer, including the theoretical foundations for its development and implementation, as well as its types, applications, and outcomes. Methods: This scoping review followed the Joanna Briggs Institute scoping review framework and adhered to the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) guidelines. A comprehensive literature search was conducted from database inception to October 16, 2025, across 7 English-language databases and 3 Chinese-language databases (CNKI, WanFang, and SinoMed). Eligible studies were peer-reviewed original research articles examining rSMS among adults with lung cancer. Data were independently screened and extracted by 2 reviewers, with discrepancies resolved by a third reviewer. Quantitative data were extracted using a standardized form and synthesized descriptively. Content analysis was performed to analyze the qualitative data. Results: A total of 41 studies involving 11,765 patients and 85 health care providers were included. Twelve studies focused exclusively on advanced-stage lung cancer. Participants were generally middle-aged to older adults (mean ages 51‐74 y), with male participants typically comprising 30% to 50% across studies. Most studies were conducted in the United States (n=19). We identified 32 patient-reported outcome measures that were used either as core rSMS components or as study outcomes. Four common functional modules were observed across rSMS: data collection, data analysis, response systems, and patient education. Qualitative evidence was limited; the most frequently reported benefit was the promotion of patient-centered care. Health care providers raised concerns about uncertain effectiveness and increased workload. Conclusions: This scoping review highlights the promising role of rSMS in lung cancer care and provides a structured map of current evidence. It adds to prior literature in 3 ways. First, it summarizes how and how often theoretical frameworks are reported and applied in rSMS development and implementation. Second, it synthesizes and categorizes four common functional modules across systems. Third, it differentiates measures embedded as rSMS components from those used as evaluation outcomes. These contributions clarify current practices and methodological gaps and underscore the importance of theory-informed design, functional clarity, and stakeholder engagement in the development of patient-centered, clinically meaningful, and sustainable rSMS platforms. Trial Registration: OSF Registries 9637t; https://osf.io/9637t/overview
❌