❌

Reading view

AI-Enhanced Social Robotic Versus Computer-Based Virtual Patients for Clinical Reasoning Training in Medical Education: Observational Crossover Cohort Study

Background: Virtual patient (VP) simulations can be used to practice clinical reasoning (CR) in controlled learning environments. Traditional computer-based VP platforms often lack the authenticity and interactivity required for effective CR training. Artificial intelligence (AI)–enhanced social robotic VPs can enhance realism and engagement; however, quantitative evidence comparing them with conventional VP platforms remains limited. Objective: We compared medical students’ experience of an AI-enhanced social robotic versus a conventional computer-based VP platform regarding the extent to which the design characteristics of the respective platform facilitate CR skill training. Methods: This observational crossover cohort study involved 178 sixth-semester medical students at Karolinska Institutet, Stockholm, Sweden (response rate: 42.3%; 178 of 421 invited students; Spring 2024-Spring 2025), who experienced both a large language model–enhanced social robotic VP platform supporting dialogue (social artificial intelligence–enhanced robotic interface [SARI]) and a conventional computer-based VP platform (virtual interactive case [VIC]) during their clinical rotation within rheumatology. Platform order was determined by clinical rotation scheduling. VP design was evaluated using a validated questionnaire across 5 domains: authenticity, professional approach, coaching quality, learning effects, and overall judgment. Students’ CR training preferences were assessed using categorical responses and a Visual Analogue Scale, where a lower score favored SARI and a score of 5 indicated equal preference between platforms. Results: SARI outperformed VIC across all 5 VP design domains. Students rated SARI higher for authenticity (median 4.0, IQR 3.5-4.5 vs 3.0, IQR 2.5-3.5; P<.001 professional approach iqr vs>P<.001 coaching quality iqr vs>P<.001 learning effect iqr vs>P<.001 and overall judgment vs iqr>P<.001 students strongly preferred sari for cr training vs odds ratio ci>P<.001 with visual analogue scale scores confirming this preference iqr>P<.001 preferences were consistent across most subgroups prior vp experience and platform order in the difference was not significant that is students with vs or ci>P=.11) and students first introduced to VIC (55% vs 45%; OR 1.5; 95% CI 0.7-2.9; P=.33). Conclusions: Our findings provide the first quantitative evidence that AI-enhanced social robotic VPs offer superior design characteristics than conventional computer-based platforms for CR training in medical education. These results support the use of AI-driven social robots for VP simulations to better prepare medical students for real clinical encounters, and warrant future research on objective CR skill outcomes and long-term transfer to clinical practice. Unlike previous qualitative studies examining each platform separately, this study provides the first quantitative comparison of design characteristics between AI-enhanced social robotic and conventional computer-based VPs.
  •  

Monitoring of circulating tumor DNA allows early detection of disease relapse in patients with operable breast cancer

Mol Oncol. 2025 Nov 27. doi: 10.1002/1878-0261.70170. Online ahead of print.

ABSTRACT

Breast cancer is known for late recurrences, yet current follow-up lacks radiological or blood-based monitoring for systemic relapse. This study evaluated circulating tumor DNA (ctDNA) monitoring for early detection of systemic relapse after curative treatment. In this case-control study of 70 patients with operable breast cancer (35 with relapse and 35 without relapse), blood samples were collected every 6-12 months during a median 8.3-year follow-up. ctDNA was analyzed by targeted DNA sequencing using Oncomine™ Breast cfDNA Research Assay v2, and results were compared to genetic analysis of tumor and metastasis biopsies. ctDNA was detected at relapse in 19 of 35 (54%) patients with disease relapse and preceded clinical or radiological relapse detection in 17, with a median lead time of 10.3 months. In 13 (68%) patients, there was concordance with tumor mutations, and in seven patients, there was also concordance with metastasis. Among the relapse-free patients, seven were ctDNA-positive postsurgery, and only one of them had a match among the tumor variants. These findings suggest serial ctDNA analysis may enable earlier detection of systemic relapse in patients with operable breast cancer.

PMID:41307327 | DOI:10.1002/1878-0261.70170

  •  

Cognitive bias in LLM reasoning compromises interpretation of clinical oncology notes

arXiv:2511.20680v1 Announce Type: cross Abstract: Despite high performance on clinical benchmarks, large language models may reach correct conclusions through faulty reasoning, a failure mode with safety implications for oncology decision support that is not captured by accuracy-based evaluation. In this two-cohort retrospective study, we developed a hierarchical taxonomy of reasoning errors from GPT-4 chain-of-thought responses to real oncology notes and tested its clinical relevance. Using breast and pancreatic cancer notes from the CORAL dataset, we annotated 600 reasoning traces to define a three-tier taxonomy mapping computational failures to cognitive bias frameworks. We validated the taxonomy on 822 responses from prostate cancer consult notes spanning localized through metastatic disease, simulating extraction, analysis, and clinical recommendation tasks. Reasoning errors occurred in 23 percent of interpretations and dominated overall errors, with confirmation bias and anchoring bias most common. Reasoning failures were associated with guideline-discordant and potentially harmful recommendations, particularly in advanced disease management. Automated evaluators using state-of-the-art language models detected error presence but could not reliably classify subtypes. These findings show that large language models may provide fluent but clinically unsafe recommendations when reasoning is flawed. The taxonomy provides a generalizable framework for evaluating and improving reasoning fidelity before clinical deployment.
  •  

Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor

arXiv:2506.14652v2 Announce Type: replace-cross Abstract: In AI research and practice, rigor remains largely understood in terms of methodological rigor -- such as whether mathematical, statistical, or computational methods are correctly applied. We argue that this narrow conception of rigor has contributed to the concerns raised by the responsible AI community, including overblown claims about the capabilities of AI systems. Our position is that a broader conception of what rigorous AI research and practice should entail is needed. We believe such a conception -- in addition to a more expansive understanding of (1) methodological rigor -- should include aspects related to (2) what background knowledge informs what to work on (epistemic rigor); (3) how disciplinary, community, or personal norms, standards, or beliefs influence the work (normative rigor); (4) how clearly articulated the theoretical constructs under use are (conceptual rigor); (5) what is reported and how (reporting rigor); and (6) how well-supported the inferences from existing evidence are (interpretative rigor). In doing so, we also provide useful language and a framework for much-needed dialogue about the AI community's work by researchers, policymakers, journalists, and other stakeholders.
  •  

Smart spatial omics (S2-omics) optimizes region of interest selection to capture molecular heterogeneity in diverse tissues

Nat Cell Biol. 2025 Nov 26. doi: 10.1038/s41556-025-01811-w. Online ahead of print.

ABSTRACT

Spatial omics technologies have transformed biomedical research by enabling high-resolution molecular profiling while preserving the native tissue architecture. These advances provide unprecedented insights into tissue structure and function. However, the high cost and time-intensive nature of spatial omics experiments necessitate careful experimental design, particularly in selecting regions of interest (ROIs) from large tissue sections. Currently, ROI selection is performed manually, which introduces subjectivity, inconsistency and a lack of reproducibility. Previous studies have shown strong correlations between spatial molecular patterns and histological features, suggesting that readily available and cost-effective histology images can be leveraged to guide spatial omics experiments. Here we present Smart Spatial omics (S2-omics), an end-to-end workflow that automatically selects ROIs from histology images with the goal of maximizing molecular information content in the ROIs. Through comprehensive evaluations across multiple spatial omics platforms and tissue types, we demonstrate that S2-omics enables systematic and reproducible ROI selection and enhances the robustness and impact of downstream biological discovery.

PMID:41298871 | DOI:10.1038/s41556-025-01811-w

  •  

Information content as a health system screening tool for rare diseases

npj Digital Medicine, Published online: 25 November 2025; doi:10.1038/s41746-025-02096-x

Information content as a health system screening tool for rare diseases
  •  

Toward explainable AI approaches for breast imaging: adapting foundation models to diverse populations

arXiv:2511.17828v1 Announce Type: cross Abstract: Foundation models hold promise for specialized medical imaging tasks, though their effectiveness in breast imaging remains underexplored. This study leverages BiomedCLIP as a foundation model to address challenges in model generalization. BiomedCLIP was adapted for automated BI-RADS breast density classification using multi-modality mammographic data (synthesized 2D images, digital mammography, and digital breast tomosynthesis). Using 96,995 images, we compared single-modality (s2D only) and multi-modality training approaches, addressing class imbalance through weighted contrastive learning. Both approaches achieved similar accuracy (multi-modality: 0.74, single-modality: 0.73), with the multi-modality model offering broader applicability across different imaging modalities and higher AUC values consistently above 0.84 across BI-RADS categories. External validation on the RSNA and EMBED datasets showed strong generalization capabilities (AUC range: 0.80-0.93). GradCAM visualizations confirmed consistent and clinically relevant attention patterns, highlighting the models interpretability and robustness. This research underscores the potential of foundation models for breast imaging applications, paving the way for future extensions for diagnostic tasks.
  •  

Towards Automating Data Access Permissions in AI Agents

arXiv:2511.17959v1 Announce Type: cross Abstract: As AI agents attempt to autonomously act on users' behalf, they raise transparency and control issues. We argue that permission-based access control is indispensable in providing meaningful control to the users, but conventional permission models are inadequate for the automated agentic execution paradigm. We therefore propose automated permission management for AI agents. Our key idea is to conduct a user study to identify the factors influencing users' permission decisions and to encode these factors into an ML-based permission management assistant capable of predicting users' future decisions. We find that participants' permission decisions are influenced by communication context but importantly individual preferences tend to remain consistent within contexts, and align with those of other participants. Leveraging these insights, we develop a permission prediction model achieving 85.1% accuracy overall and 94.4% for high-confidence predictions. We find that even without using permission history, our model achieves an accuracy of 66.9%, and a slight increase of training samples (i.e., 1-4) can substantially increase the accuracy by 10.8%.
  •  

Reinforcement Learning for Portfolio Optimization with a Financial Goal and Defined Time Horizons

arXiv:2511.18076v1 Announce Type: cross Abstract: This research proposes an enhancement to the innovative portfolio optimization approach using the G-Learning algorithm, combined with parametric optimization via the GIRL algorithm (G-learning approach to the setting of Inverse Reinforcement Learning) as presented by. The goal is to maximize portfolio value by a target date while minimizing the investor's periodic contributions. Our model operates in a highly volatile market with a well-diversified portfolio, ensuring a low-risk level for the investor, and leverages reinforcement learning to dynamically adjust portfolio positions over time. Results show that we improved the Sharpe Ratio from 0.42, as suggested by recent studies using the same approach, to a value of 0.483 a notable achievement in highly volatile markets with diversified portfolios. The comparison between G-Learning and GIRL reveals that while GIRL optimizes the reward function parameters (e.g., lambda = 0.0012 compared to 0.002), its impact on portfolio performance remains marginal. This suggests that reinforcement learning methods, like G-Learning, already enable robust optimization. This research contributes to the growing development of reinforcement learning applications in financial decision-making, demonstrating that probabilistic learning algorithms can effectively align portfolio management strategies with investor needs.
  •  

The Alignment Paradox of Medical Large Language Models in Infertility Care: Decoupling Algorithmic Improvement from Clinical Decision-making Quality

arXiv:2511.18084v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly adopted in clinical decision support, yet aligning them with the multifaceted reasoning pathways of real-world medicine remains a major challenge. Using more than 8,000 infertility treatment records, we systematically evaluate four alignment strategies: Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), Group Relative Policy Optimization (GRPO), and In-Context Learning (ICL) through a dual-layer framework combining automatic benchmarks with blinded doctor-in-the-loop assessments. GRPO achieves the highest algorithmic accuracy across multiple decision layers, confirming the value of reinforcement-based optimization for structured prediction tasks. However, clinicians consistently prefer the SFT model, citing clearer reasoning processes (p = 0.035) and higher therapeutic feasibility (p = 0.019). In blinded pairwise comparisons, SFT attains the highest winning rate (51.2%), outperforming both GRPO (26.2%) and even physicians' original decisions (22.7%). These results reveal an alignment paradox: algorithmic improvements do not necessarily translate into higher clinical trust, and may diverge from human-centered preferences. Our findings highlight the need for alignment strategies that prioritize clinically interpretable and practically feasible reasoning, rather than solely optimizing decision-level accuracy.
  •  

Clinician-in-the-Loop Smart Home System to Detect Urinary Tract Infection Flare-Ups via Uncertainty-Aware Decision Support

arXiv:2511.18334v1 Announce Type: cross Abstract: Urinary tract infection (UTI) flare-ups pose a significant health risk for older adults with chronic conditions. These infections often go unnoticed until they become severe, making early detection through innovative smart home technologies crucial. Traditional machine learning (ML) approaches relying on simple binary classification for UTI detection offer limited utility to nurses and practitioners as they lack insight into prediction uncertainty, hindering informed clinical decision-making. This paper presents a clinician-in-the-loop (CIL) smart home system that leverages ambient sensor data to extract meaningful behavioral markers, train robust predictive ML models, and calibrate them to enable uncertainty-aware decision support. The system incorporates a statistically valid uncertainty quantification method called Conformal-Calibrated Interval (CCI), which quantifies uncertainty and abstains from making predictions ("I don't know") when the ML model's confidence is low. Evaluated on real-world data from eight smart homes, our method outperforms baseline methods in recall and other classification metrics while maintaining the lowest abstention proportion and interval width. A survey of 42 nurses confirms that our system's outputs are valuable for guiding clinical decision-making, underscoring their practical utility in improving informed decisions and effectively managing UTIs and other condition flare-ups in older adults.
  •  

Are Large Vision Language Models Truly Grounded in Medical Images? Evidence from Italian Clinical Visual Question Answering

arXiv:2511.19220v1 Announce Type: cross Abstract: Large vision language models (VLMs) have achieved impressive performance on medical visual question answering benchmarks, yet their reliance on visual information remains unclear. We investigate whether frontier VLMs demonstrate genuine visual grounding when answering Italian medical questions by testing four state-of-the-art models: Claude Sonnet 4.5, GPT-4o, GPT-5-mini, and Gemini 2.0 flash exp. Using 60 questions from the EuropeMedQA Italian dataset that explicitly require image interpretation, we substitute correct medical images with blank placeholders to test whether models truly integrate visual and textual information. Our results reveal striking variability in visual dependency: GPT-4o shows the strongest visual grounding with a 27.9pp accuracy drop (83.2% [74.6%, 91.7%] to 55.3% [44.1%, 66.6%]), while GPT-5-mini, Gemini, and Claude maintain high accuracy with modest drops of 8.5pp, 2.4pp, and 5.6pp respectively. Analysis of model-generated reasoning reveals confident explanations for fabricated visual interpretations across all models, suggesting varying degrees of reliance on textual shortcuts versus genuine visual analysis. These findings highlight critical differences in model robustness and the need for rigorous evaluation before clinical deployment.
  •  

AI-Assisted Cardiovascular Risk Assessment by General Practitioners in Resource-Constrained Indonesian Settings Using a Conceptual Prototype: Randomized Controlled Study

Background: Preventive strategies integrated with digital health and artificial intelligence (AI), have significant potential to mitigate the global burden of atherosclerotic cardiovascular disease (ASCVD). AI-enabled clinical decision support (CDS) systems increasingly provide patient-specific insights beyond traditional risk factors. Despite these advances, their capacity to enhance clinical decision-making in resource-constrained settings remains largely unexplored. Objective: We conducted a randomised controlled study to assess the effect of AI-based CDS on 10-year ASCVD risk assessment and management in primary prevention. Methods: In a three-way within-subject randomised design, doctors completed nine clinical vignettes representative of primary care presentations in a resource-constrained outpatient setting. For each vignette, participants assessed 10-year ASCVD risk and made management decisions using either a conceptual prototype of AI-based CDS, automated CDS, or no decision support. The conceptual prototype represented contemporary risk calculators based on traditional machine learning models (e.g., random forest, neural networks, logistic regression) that incorporate additional predictors alongside traditional risk factors. Primary outcomes were correct risk assessment and patient management (prescription of aspirin, statins, and anti-hypertensives; referral for advanced examinations). Decision-making time and perceptions about AI utility were also measured. Results: 102 doctors from all seven geographical regions of Indonesia participated. Most participants were 26–35 years (83%), 56% male, with a median of six years of clinical experience (IQR=4.75). AI-based CDS improved risk assessment by 27% (2(2, n=102) = 48.875, P<.001 compared to unassisted or one additional correct risk classification for every patients where doctors use ai needed treat nnt="3.7;" ci the prescription of statins also improved by n="102)=" p in pairwise comparisons assisted with ai-based cds correctly assessed significantly more cases adjusted and prescribed appropriate statin often medium effect size r=".35)" control. ai-assisted required less time marginal means sec vs f however improvements aspirin anti-hypertensives did not reach statistical significance. no improvement was observed referral decisions. participants generally viewed positively agreeing strongly that they would follow its recommendations indicating it if given access. believed could enhance efficiency assessment particularly high-volume primary care settings while noting need verify against clinical guidelines each patient. conclusions: coupled reduced decision-making highlight potential utility ascvd resource-constrained efficient healthcare resources is crucial. further research ascertain whether this online study translate real-world low-resource settings.>
  •  

ConCISE: A Reference-Free Conciseness Evaluation Metric for LLM-Generated Answers

arXiv:2511.16846v1 Announce Type: cross Abstract: Large language models (LLMs) frequently generate responses that are lengthy and verbose, filled with redundant or unnecessary details. This diminishes clarity and user satisfaction, and it increases costs for model developers, especially with well-known proprietary models that charge based on the number of output tokens. In this paper, we introduce a novel reference-free metric for evaluating the conciseness of responses generated by LLMs. Our method quantifies non-essential content without relying on gold standard references and calculates the average of three calculations: i) a compression ratio between the original response and an LLM abstractive summary; ii) a compression ratio between the original response and an LLM extractive summary; and iii) wordremoval compression, where an LLM removes as many non-essential words as possible from the response while preserving its meaning, with the number of tokens removed indicating the conciseness score. Experimental results demonstrate that our proposed metric identifies redundancy in LLM outputs, offering a practical tool for automated evaluation of response brevity in conversational AI systems without the need for ground truth human annotations.
  •  

Artificial Intelligence Index Report 2025

arXiv:2504.07139v3 Announce Type: replace Abstract: Welcome to the eighth edition of the AI Index report. The 2025 Index is our most comprehensive to date and arrives at an important moment, as AI's influence across society, the economy, and global governance continues to intensify. New in this year's report are in-depth analyses of the evolving landscape of AI hardware, novel estimates of inference costs, and new analyses of AI publication and patenting trends. We also introduce fresh data on corporate adoption of responsible AI practices, along with expanded coverage of AI's growing role in science and medicine. Since its founding in 2017 as an offshoot of the One Hundred Year Study of Artificial Intelligence, the AI Index has been committed to equipping policymakers, journalists, executives, researchers, and the public with accurate, rigorously validated, and globally sourced data. Our mission has always been to help these stakeholders make better-informed decisions about the development and deployment of AI. In a world where AI is discussed everywhere - from boardrooms to kitchen tables - this mission has never been more essential. The AI Index continues to lead in tracking and interpreting the most critical trends shaping the field - from the shifting geopolitical landscape and the rapid evolution of underlying technologies, to AI's expanding role in business, policymaking, and public life. Longitudinal tracking remains at the heart of our mission. In a domain advancing at breakneck speed, the Index provides essential context - helping us understand where AI stands today, how it got here, and where it may be headed next. Recognized globally as one of the most authoritative resources on artificial intelligence, the AI Index has been cited in major media outlets such as The New York Times, Bloomberg, and The Guardian; referenced in hundreds of academic papers; and used by policymakers and government agencies around the world.
  •  

Multimodal analysis of whole slide images in colorectal cancer

npj Digital Medicine, Published online: 24 November 2025; doi:10.1038/s41746-025-02095-y

Multimodal analysis of whole slide images in colorectal cancer
  •  

Integrative analysis of genomic and transcriptomic data informs precancer progression in the pancreas

bioRxiv [Preprint]. 2025 Nov 4:2025.11.03.686234. doi: 10.1101/2025.11.03.686234.

ABSTRACT

Pancreatic ductal adenocarcinoma (PDAC) arises from heterogeneous precursor lesions, including intraductal papillary mucinous neoplasms (IPMNs), but the features distinguishing indolent from progressive lesions remain unclear. We performed an integrative analysis of transcriptomic, genomic, and microenvironmental profiles of IPMNs to define multi-omic phenotypes. Using transfer learning, we projected IPMN-derived transcriptional programs onto spatial transcriptomic datasets from IPMNs and pancreatic intraepithelial neoplasias (PanINs). We identified two major phenotypes: one associated with cancer-associated fibroblasts and epithelial-to-mesenchymal transition, shared across IPMN, PanIN, and PDAC; and a second, glycolysis-enriched phenotype with a unique somatic mutation profile specific to IPMN. Spatial mapping further revealed grade-specific enrichment of transcriptional programs and distinct interactions with stromal and immune subtypes, underscoring the role of the precancer microenvironment in progression. These findings establish multi-omic phenotypes that unify genetic, transcriptional, and microenvironmental heterogeneity, providing a framework for distinguishing progressive from indolent precancers and a web-based public atlas for future exploration of these data and transcriptional phenotypes.

PMID:41279473 | PMC:PMC12637499 | DOI:10.1101/2025.11.03.686234

  •  

Knowledge-informed multimodal cfDNA analysis improves sensitivity and generalization in cancer detection

bioRxiv [Preprint]. 2025 Oct 21:2025.10.20.683167. doi: 10.1101/2025.10.20.683167.

ABSTRACT

Liquid biopsy offers a minimally invasive opportunity to detect and monitor cancers through analysis of cell-free DNA (cfDNA). However, current approaches face challenges of limited sensitivity at low tumor fractions, technical variability, and poor generalization across cohorts. Tumor-informed targeted methods offer high specificity but suffer from low sensitivity due to random sampling, tumor evolution and adaptation (including resistance mechanisms), and other sources of heterogeneity. Conversely, tumor-naive genome-wide methods can increase sensitivity but often sacrifice specificity, particularly at low tumor fractions. We developed Fragmentomics Analysis for Tumor Evaluation with AI (Fate-AI), a multimodal framework that integrates fragmentomic and methylation-derived features from low-pass whole-genome sequencing (LPWGS) and cell-free methylated DNA immunoprecipitation and high-throughput sequencing (cfMeDIP-seq). It employs a knowledge-informed strategy to select recurrently altered genomic regions and tissue-specific methylation loci to combine the advantages of tumor-naive approaches with the specificity of tumor-informed approaches. This approach derives robust per-sample normalized features that mitigate batch effects and enhance cross-cohort reproducibility. We evaluated Fate-AI on a total of 1,219 plasma samples spanning ten cancer types and healthy controls from multiple laboratories and sequencing centers, including 432 newly profiled cases (280 with both cfMeDIP-seq and LPWGS) together with 787 samples from four independent public datasets. Fate-AI achieved superior sensitivity and specificity compared to state-of-the-art methods, detecting tumor-derived signals at fractions as low as 10-5 in experimental dilutions. Fate-AI scores correlated with disease stage and tracked longitudinal progression, anticipating relapse months before clinical progression. Furthermore, Fate-AI enabled tissue-of-origin classification, with AUCs ranging from 0.84 to 0.97 across six cancer types. Collectively, our results demonstrate that Fate-AI provides a sensitive, generalizable, and clinically actionable platform for early detection, minimal residual disease monitoring, and tissue-of-origin classification, supporting its potential as a liquid biopsy framework in precision oncology.

PMID:41278930 | PMC:PMC12633305 | DOI:10.1101/2025.10.20.683167

  •  

Impact of Digital Interventions on the Treatment Burden of Patients With Chronic Conditions: Systematic Review

Background: Digital interventions can provide cost-effective, quality health care for patients with chronic conditions. Patients with chronic conditions often are burdened by a substantial load of adhering to a treatment regimen and suffer from impacts on their function and well-being. This treatment burden has consequences for treatment adherence and disease outcomes. Digital interventions have the potential to alleviate the burden, but they also may cause new challenges and an increased workload for the patient. Previous reviews have examined digital interventions or treatment burden separately, but there is a lack of systematic reviews on the intersection of digital interventions, treatment burden, and chronic conditions. Objective: This systematic review aimed to evaluate the evidence of how digital interventions impact the treatment burden experienced by people with chronic conditions, and to assess the quality of this evidence. Methods: We searched databases PubMed, Scopus, Web of Science, ACM, PubMed Central, and CINAHL for articles published between January 1, 2013, and June 17, 2025. We included studies that had key topics related to chronic conditions, treatment burden, and digital interventions. A total of 2 reviewers independently screened the articles in 2 stages, extracted data on study design, participant characteristics, intervention type, and treatment burden outcomes from included articles, and assessed their quality using the Critical Appraisal tools from the Joanna Briggs Institute. A convergent integrated approach was used for data synthesis and integration, where quantitative data were converted into qualitative data, and the qualitative and quantitative evidence were analyzed and categorized together. Results: We included 46 relevant studies in total. We categorized the interventions into 4 types: Telehealth, informational resources, self-management tools, and facilitated tools. The results of this study indicate that digital interventions mostly support patients with chronic conditions with their treatment burden, with minor concerns of increasing treatment burden. The main benefits are support with self-management, informational support, and easier ways to contact health care professionals. The main concerns were accessibility issues, time-consuming tools, and causing fear and anxiety. Conclusions: Our findings demonstrate how treatment burden is a relevant concept for future digital health care research and practice. Digital interventions can help patients with their treatment burden by supporting self-management, improving access to health care, improving patients’ experience, and addressing relevant concerns. More research is needed about conditions with low or medium initial treatment burden.
  •  
❌