❌

Normal view

  • ✇STAT
  • STAT+: Nature Medicine to investigate study that found cancer treatment is better in morning Angus Chen
    The notion that oncologists could boost immunotherapy responses simply by giving infusions in the morning, rather than late afternoon, is an attractive one. So when a clinical trial published in Nature Medicine this month showed that lung cancer patients treated in the morning had a massive reduction in the risk of progression compared to those treated in the afternoon, many scientists were intrigued, if skeptical. Now that study is coming under fire, as multiple scientists and sleuths raise
     

STAT+: Nature Medicine to investigate study that found cancer treatment is better in morning

21 February 2026 at 09:13

The notion that oncologists could boost immunotherapy responses simply by giving infusions in the morning, rather than late afternoon, is an attractive one. So when a clinical trial published in Nature Medicine this month showed that lung cancer patients treated in the morning had a massive reduction in the risk of progression compared to those treated in the afternoon, many scientists were intrigued, if skeptical.

Now that study is coming under fire, as multiple scientists and sleuths raise serious concerns about the data and point out inconsistencies in the trial.

These have called the study’s conclusions even further into question, which experts told STAT already lacked strong biological plausibility, and Nature Medicine appended a note on the study on Thursday that it is starting an investigation into the concerns.

Continue to STAT+ to read the full story…

© Jenny Kane/AP

Gastric cancer occurrence and heterogeneity: integration of clinical data, multi-omics and tumor microenvironment

Future Oncol. 2026 Feb 20:1-14. doi: 10.1080/14796694.2026.2631302. Online ahead of print.

ABSTRACT

Gastric cancer (GC) is an epithelial malignant tumor with high morbidity and mortality. In recent years, more and more studies have strengthened our understanding of how GC develops, including the origin of GC cells, precancerous lesions, gene mutations, transcriptional changes, protein translation and the tumor microenvironment. With the concept of accurate tumor therapy gradually applied to clinical practice, these data provide more reference and basis for early prevention, early screening, early detection and accurate treatment of GC.

PMID:41717787 | DOI:10.1080/14796694.2026.2631302

  • ✇cs.AI, q-bio.NC updates on arXiv.org
  • Intent Laundering: AI Safety Datasets Are Not What They Seem Shahriar Golchin · Marc Wetter
    arXiv:2602.16729v1 Announce Type: cross Abstract: We systematically evaluate the quality of widely used AI safety datasets from two perspectives: in isolation and in practice. In isolation, we examine how well these datasets reflect real-world attacks based on three key properties: driven by ulterior intent, well-crafted, and out-of-distribution. We find that these datasets overrely on "triggering cues": words or phrases with overt negative/sensitive connotations that are intended to trigger sa
     

Intent Laundering: AI Safety Datasets Are Not What They Seem

arXiv:2602.16729v1 Announce Type: cross Abstract: We systematically evaluate the quality of widely used AI safety datasets from two perspectives: in isolation and in practice. In isolation, we examine how well these datasets reflect real-world attacks based on three key properties: driven by ulterior intent, well-crafted, and out-of-distribution. We find that these datasets overrely on "triggering cues": words or phrases with overt negative/sensitive connotations that are intended to trigger safety mechanisms explicitly, which is unrealistic compared to real-world attacks. In practice, we evaluate whether these datasets genuinely measure safety risks or merely provoke refusals through triggering cues. To explore this, we introduce "intent laundering": a procedure that abstracts away triggering cues from attacks (data points) while strictly preserving their malicious intent and all relevant details. Our results indicate that current AI safety datasets fail to faithfully represent real-world attacks due to their overreliance on triggering cues. In fact, once these cues are removed, all previously evaluated "reasonably safe" models become unsafe, including Gemini 3 Pro and Claude Sonnet 3.7. Moreover, when intent laundering is adapted as a jailbreaking technique, it consistently achieves high attack success rates, ranging from 90% to over 98%, under fully black-box access. Overall, our findings expose a significant disconnect between how model safety is evaluated and how real-world adversaries behave.
  • ✇STAT
  • STAT+: Key study of Grail’s cancer detection test fails in setback for company Matthew Herper and Angus Chen
    A blood test for detecting cancer early being developed by the diagnostics firm Grail failed to meet its main goal in a giant study being conducted with England’s National Health Service, the company said Thursday. Grail’s test has been the standard bearer for new technologies that promise a blood test can be used to detect many different types of cancer early and eventually even to indicate to scientists where in the body to look for tumors. The company already sells its test, called Galleri
     

STAT+: Key study of Grail’s cancer detection test fails in setback for company

20 February 2026 at 06:38

A blood test for detecting cancer early being developed by the diagnostics firm Grail failed to meet its main goal in a giant study being conducted with England’s National Health Service, the company said Thursday.

Grail’s test has been the standard bearer for new technologies that promise a blood test can be used to detect many different types of cancer early and eventually even to indicate to scientists where in the body to look for tumors. The company already sells its test, called Galleri, for a list price of $1,000, although it is not yet approved by the Food and Drug Administration. Grail said Thursday it sold 185,000 tests in 2025, generating $136.8 million. 

The company’s shares were down 47% in after-hours trading.

Continue to STAT+ to read the full story…

© Adobe

Hunt Globally: Deep Research AI Agents for Drug Asset Scouting in Investing, Business Development, and Search & Evaluation

arXiv:2602.15019v1 Announce Type: new Abstract: Bio-pharmaceutical innovation has shifted: many new drug assets now originate outside the United States and are disclosed primarily via regional, non-English channels. Recent data suggests >85% of patent filings originate outside the U.S., with China accounting for nearly half of the global total; a growing share of scholarly output is also non-U.S. Industry estimates put China at ~30% of global drug development, spanning 1,200+ novel candidates. In this high-stakes environment, failing to surface "under-the-radar" assets creates multi-billion-dollar risk for investors and business development teams, making asset scouting a coverage-critical competition where speed and completeness drive value. Yet today's Deep Research AI agents still lag human experts in achieving high-recall discovery across heterogeneous, multilingual sources without hallucinations. We propose a benchmarking methodology for drug asset scouting and a tuned, tree-based self-learning Bioptic Agent aimed at complete, non-hallucinated scouting. We construct a challenging completeness benchmark using a multilingual multi-agent pipeline: complex user queries paired with ground-truth assets that are largely outside U.S.-centric radar. To reflect real deal complexity, we collected screening queries from expert investors, BD, and VC professionals and used them as priors to conditionally generate benchmark queries. For grading, we use LLM-as-judge evaluation calibrated to expert opinions. We compare Bioptic Agent against Claude Opus 4.6, OpenAI GPT-5.2 Pro, Perplexity Deep Research, Gemini 3 Pro + Deep Research, and Exa Websets. Bioptic Agent achieves 79.7% F1 versus 56.2% (Claude Opus 4.6), 50.6% (Gemini 3 Pro + Deep Research), 46.6% (GPT-5.2 Pro), 44.2% (Perplexity Deep Research), and 26.9% (Exa Websets). Performance improves steeply with additional compute, supporting the view that more compute yields better results.

MedScope: Incentivizing "Think with Videos" for Clinical Reasoning via Coarse-to-Fine Tool Calling

arXiv:2602.13332v1 Announce Type: cross Abstract: Long-form clinical videos are central to visual evidence-based decision-making, with growing importance for applications such as surgical robotics and related settings. However, current multimodal large language models typically process videos with passive sampling or weakly grounded inspection, which limits their ability to iteratively locate, verify, and justify predictions with temporally targeted evidence. To close this gap, we propose MedScope, a tool-using clinical video reasoning model that performs coarse-to-fine evidence seeking over long-form procedures. By interleaving intermediate reasoning with targeted tool calls and verification on retrieved observations, MedScope produces more accurate and trustworthy predictions that are explicitly grounded in temporally localized visual evidence. To address the lack of high-fidelity supervision, we build ClinVideoSuite, an evidence-centric, fine-grained clinical video suite. We then optimize MedScope with Grounding-Aware Group Relative Policy Optimization (GA-GRPO), which directly reinforces tool use with grounding-aligned rewards and evidence-weighted advantages. On full and fine-grained video understanding benchmarks, MedScope achieves state-of-the-art performance in both in-domain and out-of-domain evaluations. Our approach illuminates a path toward medical AI agents that can genuinely "think with videos" through tool-integrated reasoning. We will release our code, models, and data.

Rare, Yet Targetable: New Perspectives on Ampullary Carcinomas

Int J Mol Sci. 2026 Feb 6;27(3):1597. doi: 10.3390/ijms27031597.

ABSTRACT

Ampullary carcinoma (AC) is a rare gastrointestinal malignancy with dual intestinal and pancreatobiliary differentiation, complicating diagnosis, staging, and treatment. This review synthesizes current epidemiology, pathology, and multi-omic data to outline a pragmatic care pathway: lineage-first at presentation, mutation-fast at progression. Histology remains the primary classifier: the intestinal subtype generally aligns with colorectal regimens, whereas pancreatobiliary and mixed subtypes favor pancreaticobiliary therapy. In selected fit patients, modified FOLFIRINOX may address mixed phenotypes. Next-generation sequencing adds precision by identifying therapeutically relevant alterations, including ERBB2/HER2 amplifications, MSI-high/dMMR, BRAF V600E, and rare NTRK or RET fusions, while KRAS mutations are enriched in pancreatobiliary tumors. We recommend early application of a rapid-core panel (KRAS/BRAF, MSI/dMMR, ERBB2/HER2, RNA-based fusions) to capture high-impact targets, followed by comprehensive profiling at first progression. Liquid biopsy, plasma circulating tumor DNA (ctDNA), or bile-derived DNA may complement tissue and help identify the dominant lineage. Research priorities include ampulla-enriched umbrella trials, explicit AC subcohorts in tissue-agnostic studies, and ctDNA-informed endpoints. This lineage-first, mutation-fast paradigm supports precision care and evidence generation in AC.

PMID:41684016 | PMC:PMC12897727 | DOI:10.3390/ijms27031597

Clinical utility of OGN in pan-cancer: diagnostic biomarker and immune microenvironment regulator

12 February 2026 at 19:00

Transl Cancer Res. 2026 Jan 31;15(1):43. doi: 10.21037/tcr-2025-1499. Epub 2026 Jan 27.

ABSTRACT

BACKGROUND: Osteoglycin (OGN), an extracellular matrix protein, has emerging but poorly characterized roles in cancer. This study presents the first pan-cancer investigation of OGN's expression patterns, clinical significance, immune interactions, and functional mechanisms.

METHODS: Multi-omics data from Genotype Tissue Expression (GTEx), Cancer Cell Line Encyclopedia (CCLE), The Cancer Genome Atlas (TCGA), and Human Protein Atlas (HPA) databases were integrated. Differential expression was analyzed in normal tissues and tumor samples. Diagnostic utility was evaluated using area under the curve (AUC) of receiver operating characteristic (ROC) curve. Prognostic value was assessed via Kaplan-Meier [overall survival (OS); disease-specific survival (DSS); disease free interval (DFI); progression-free interval (PFI)] and Cox regression analyses. Immune microenvironment correlations were quantified using ESTIMATE, CIBERSORT, and gene set enrichment. Functional pathways were explored through gene set enrichment analysis (GSEA) and correlation with hallmark cancer signatures.

RESULTS: OGN was broadly expressed in normal tissues (brain, liver, kidney) but significantly downregulated in most tumor types (P<0.05, TCGA; validated at protein level, HPA). OGN demonstrated high diagnostic accuracy in pan-cancer (AUC: 0.703-0.990), achieving near-perfect performance in colon adenocarcinoma (COAD) (AUC: 0.966) and thyroid cancer (THCA) (AUC: 0.920). High OGN expression correlated with improved survival outcomes in thymoma (THYM) (OS/DSS) and cholangiocarcinoma (CHOL) (PFI/DFI), but worse prognosis in lung adenocarcinoma​/liver hepatocellular carcinoma​ (LUAD/LIHC), indicating cancer-type specificity. OGN expression strongly associated with immune cell infiltration (macrophages, natural killer cells, T cells), chemokine signaling, programmed death-ligand 1 (PD-L1) levels, microsatellite instability (MSI), and tumor mutation burden (TMB). GSEA revealed enrichment of OGN-linked genes in epithelial-mesenchymal transition (EMT), angiogenesis, JAK-STAT, and PI3K pathways across cancers.

CONCLUSIONS: Our pan-cancer analysis highlights OGN as a context-dependent regulator linking extracellular matrix (ECM) remodeling with immune and angiogenic signaling. Its pan-cancer dysregulation, diagnostic/prognostic value, and crosstalk with immune evasion mechanisms nominate OGN as a promising multi-functional biomarker and therapeutic target.

PMID:41674945 | PMC:PMC12885879 | DOI:10.21037/tcr-2025-1499

Advancing healthcare AI governance through a comprehensive maturity model based on systematic review

npj Digital Medicine, Published online: 11 February 2026; doi:10.1038/s41746-026-02418-7

Advancing healthcare AI governance through a comprehensive maturity model based on systematic review

Spatial and multi-omics transcriptomic dissects platinum resistance in lung adenocarcinoma: a five-gene predictive model with tumor microenvironment dynamics

9 February 2026 at 19:00

Chem Biol Interact. 2026 Feb 7:111952. doi: 10.1016/j.cbi.2026.111952. Online ahead of print.

ABSTRACT

The scarcity of reliable biomarkers and predictive models for platinum resistance in lung adenocarcinoma (LUAD) poses a significant clinical challenge. This study endeavors to identify molecular subtypes related to platinum resistance and construct a robust predictive model through multi-omics techniques. We performed integrative analysis of public datasets using advanced bioinformatics strategies, including spatial transcriptome deconvolution and consensus clustering. Bulk RNA deconvolution analysis was conducted to characterize tumor microenvironment heterogeneity. Feature selection was performed using the Supervised Principal Component (SuperPC) algorithm, followed by diagnostic model construction validated through receiver operating characteristic (ROC) analysis. Functional validation was performed through cytological experiments measuring cisplatin IC50 alterations following gene manipulation in LUAD cell lines. Consensus clustering revealed distinct LUAD subtypes, with Cluster1 demonstrating significant platinum resistance. We first subtyped the patients in the bulk transcriptome data based on consistency clustering, and then analyzed the differences between different platinum-resistant subtypes (Cluster 1 and Cluster 2), so as to screen 333 isotype-specific differentially expressed genes and 15 platinum resistance-related (PRR) genes were selected through machine learning. A refined 5-gene signature (ANKRD29/CACNA2D2/DSP/HSD17B6/SPP1) achieved exceptional predictive performance (AUC=0.9639). Spatial transcriptomics demonstrated compartmentalized expression patterns: SPP1/DSP localized to tumor niches, HSD17B6/CACNA2D2 to epithelial regions, and ANKRD29 depletion in stromal areas. Cellular colocalization analysis revealed malignant epithelial PH proximity to myeloid and mast cells. Functional validation confirmed that ANKRD29/CACNA2D2 overexpression sensitized A549/DDP cells to cisplatin, while DSP/SPP1/HSD17B6 overexpression induced resistance. Experiments in nude mice have shown that these genes are closely related to cisplatin resistance in LUAD. This study identifies the Cluster1 subtype and malignant epithelial PH as crucial determinants of platinum resistance in LUAD. Our innovative 5-gene predictive model exhibits clinical-grade diagnostic accuracy, and spatial transcriptomic characterization offers mechanistic insights into the dynamics of the tumor microenvironment.

PMID:41662930 | DOI:10.1016/j.cbi.2026.111952

Problems and Barriers Regarding the Admission, Financing, and Service Provision of Digital Health Apps: Qualitative Stakeholder Survey

Background: Since their introduction with the Digital Care Act in 2019, DiGA are a part of the German statutory healthcare system. In order to become a DiGA, mHealth apps have to complete a certification process covering both technical and evidence related aspects. After completion, DiGA are added to the DiGA-directory, containing a list of all reimbursable DiGA within German statutory health insurance (SHI). The first apps were added at the end of 2020 with the number steadily increasing. The novelty of the introduction leads to problems and barriers to optimal use along the way, which is studied from different stakeholder perspectives in this research article. Objective: The aim of the survey was to identify problems and barriers in the context of certification, financing and use of DiGA in Germany. Methods: We used semi-structured expert interviews to evaluate the perspective of stakeholders of the German healthcare system on DiGA. The interview guide was developed according to Helfferich, the interviews were transcribed and analyzed using the qualitative content approach by Mayring and Kuckartz. Results: We identified problems from stakeholder perspectives regarding the certification/admission, financing and service distribution regarding DiGA. The interviewed stakeholders reported problems with authorization of DiGA and the corresponding process. DiGA prices and the different negotiation positions were criticized, as well as financial challenges for smaller DiGA-manufacturers. Within service provision, technical problems, e. g., with activation codes or software surrounding DiGA-prescription were mentioned. Problems were also seen in insufficient knowledge and skills on the side of the patients as well as the medical providers. Conclusions: mHealth applications provide potentially disruptive innovations within the healthcare sector. Nevertheless, since the evidence-based and regulated use of this technology is relatively new there are still problems and barriers limiting the optimized, patient-centered use. This study provides an overview of problems in the context of DiGA in Germany from the stakeholder perspective. Since other countries showed interest in potentially adopting the German system, valuable implications can be drawn from this survey.

Quantifying Individual Health Status from Multi-omics Data by Health State Manifold

Phenomics. 2025 Dec 15;5(5):469-486. doi: 10.1007/s43657-024-00188-4. eCollection 2025 Oct.

ABSTRACT

Quantifying individual health status from increasingly accumulated omics data is essential for both early prevention and intervention of diseases, which attracts great attention from communities of biology and medicine. Most of the existing approaches mainly classify individuals into different catalogues or classes based on phenotypes and biomarkers. However, an individual's health status from a dynamical systems viewpoint can be viewed as a non-equilibrium steady state, which can generally be characterized by two key features, i.e. (1) homeostatic potential that represents the ability of homeostatic resilience to withstand perturbations or maintain functions at the current state/phenotype of this individual and (2) phenotypic potential that represents the state/phenotype of the individual on the whole process from health to disease. Here, we proposed a health state manifold (HSM) method derived from dynamic network biomarker method and diffusion map theory to quantify individual health status with the characterization of such two features in a robust and accurate manner based on multi-omics data. To verify our method, HSM method was applied to the quantification of diabetes mellitus (rat subjects) and the Roux-en-Y Gastric Bypass (human subjects) for both disease progression process and recovery process, which demonstrated its effectiveness and potential for personalized medicine and preventive medicine.

SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at 10.1007/s43657-024-00188-4.

PMID:41659741 | PMC:PMC12881232 | DOI:10.1007/s43657-024-00188-4

Spatial Multi-omics Analyses Reveal Diabetes Promotes Pancreatic Cancer Progression by Stimulating Cholesterol-Induced Neutrophil Extracellular Trap Formation

Cancer Res. 2026 Feb 9. doi: 10.1158/0008-5472.CAN-25-2854. Online ahead of print.

ABSTRACT

Pancreatic ductal adenocarcinoma (PDAC) patients with diabetes mellitus (DM) exhibit poor clinical outcomes. Metabolic reprogramming of both cancer cells and immune compartments plays a crucial role in shaping the anti-tumor immune response in PDAC. DM-induced metabolic alteration may disrupt the intricate crosstalk between immune cells and tumor-associated immune factors, profoundly influencing PDAC progression. Here, we performed an integrated, spatially resolved multi-omics study to investigate DM-associated, cell-specific metabolic remodeling within the PDAC tumor microenvironment. DM influenced interactions between tumor cells and immune cells, which accelerated PDAC growth in both humans and mice. PDAC patients with DM exhibited higher tumor-stage, poorer differentiation, and worse outcomes. Spatial metabolic and transcriptional profiling revealed that SREBP2-dependent cholesterol biosynthesis exacerbated PDAC progression. Increased cholesterol biosynthesis promoted neutrophil recruitment and accelerated formation of neutrophil extracellular traps (NETs) by stimulating the CXCL1-CXCR1/CXCR2 signaling axis, ultimately promoting PDAC growth. Inhibition of SREBP2, pharmacological blockade of CXCL1, or perturbation of NETs markedly reduced PDAC growth in diabetic mouse models. Together, these multi-omics analyses and follow-up mechanistic studies constitute an integrated approach that elucidates a metabolic mechanism by which diabetes promotes PDAC development by remodeling the tumor immune microenvironment and highlights a potential therapeutic strategy for PDAC with DM.

PMID:41661642 | DOI:10.1158/0008-5472.CAN-25-2854

Generating High-quality Privacy-preserving Synthetic Data

arXiv:2602.06390v1 Announce Type: cross Abstract: Synthetic tabular data enables sharing and analysis of sensitive records, but its practical deployment requires balancing distributional fidelity, downstream utility, and privacy protection. We study a simple, model agnostic post processing framework that can be applied on top of any synthetic data generator to improve this trade off. First, a mode patching step repairs categories that are missing or severely underrepresented in the synthetic data, while largely preserving learned dependencies. Second, a k nearest neighbor filter replaces synthetic records that lie too close to real data points, enforcing a minimum distance between real and synthetic samples. We instantiate this framework for two neural generative models for tabular data, a feed forward generator and a variational autoencoder, and evaluate it on three public datasets covering credit card transactions, cardiovascular health, and census based income. We assess marginal and joint distributional similarity, the performance of models trained on synthetic data and evaluated on real data, and several empirical privacy indicators, including nearest neighbor distances and attribute inference attacks. With moderate thresholds between 0.2 and 0.35, the post processing reduces divergence between real and synthetic categorical distributions by up to 36 percent and improves a combined measure of pairwise dependence preservation by 10 to 14 percent, while keeping downstream predictive performance within about 1 percent of the unprocessed baseline. At the same time, distance based privacy indicators improve and the success rate of attribute inference attacks remains largely unchanged. These results provide practical guidance for selecting thresholds and applying post hoc repairs to improve the quality and empirical privacy of synthetic tabular data, while complementing approaches that provide formal differential privacy guarantees.

Yunjue Agent Tech Report: A Fully Reproducible, Zero-Start In-Situ Self-Evolving Agent System for Open-Ended Tasks

arXiv:2601.18226v2 Announce Type: replace Abstract: Conventional agent systems often struggle in open-ended environments where task distributions continuously drift and external supervision is scarce. Their reliance on static toolsets or offline training lags behind these dynamics, leaving the system's capability boundaries rigid and unknown. To address this, we propose the In-Situ Self-Evolving paradigm. This approach treats sequential task interactions as a continuous stream of experience, enabling the system to distill short-term execution feedback into long-term, reusable capabilities without access to ground-truth labels. Within this framework, we identify tool evolution as the critical pathway for capability expansion, which provides verifiable, binary feedback signals. Within this framework, we develop Yunjue Agent, a system that iteratively synthesizes, optimizes, and reuses tools to navigate emerging challenges. To optimize evolutionary efficiency, we further introduce a Parallel Batch Evolution strategy. Empirical evaluations across five diverse benchmarks under a zero-start setting demonstrate significant performance gains over proprietary baselines. Additionally, complementary warm-start evaluations confirm that the accumulated general knowledge can be seamlessly transferred to novel domains. Finally, we propose a novel metric to monitor evolution convergence, serving as a function analogous to training loss in conventional optimization. We open-source our codebase, system traces, and evolved tools to facilitate future research in resilient, self-evolving intelligence.

Data-Centric Interpretability for LLM-based Multi-Agent Reinforcement Learning

arXiv:2602.05183v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly trained in complex Reinforcement Learning, multi-agent environments, making it difficult to understand how behavior changes over training. Sparse Autoencoders (SAEs) have recently shown to be useful for data-centric interpretability. In this work, we analyze large-scale reinforcement learning training runs from the sophisticated environment of Full-Press Diplomacy by applying pretrained SAEs, alongside LLM-summarizer methods. We introduce Meta-Autointerp, a method for grouping SAE features into interpretable hypotheses about training dynamics. We discover fine-grained behaviors including role-playing patterns, degenerate outputs, language switching, alongside high-level strategic behaviors and environment-specific bugs. Through automated evaluation, we validate that 90% of discovered SAE Meta-Features are significant, and find a surprising reward hacking behavior. However, through two user studies, we find that even subjectively interesting and seemingly helpful SAE features may be worse than useless to humans, along with most LLM generated hypotheses. However, a subset of SAE-derived hypotheses are predictively useful for downstream tasks. We further provide validation by augmenting an untrained agent's system prompt, improving the score by +14.2%. Overall, we show that SAEs and LLM-summarizer provide complementary views into agent behavior, and together our framework forms a practical starting point for future data-centric interpretability work on ensuring trustworthy LLM behavior throughout training.

Exploring AI-Augmented Sensemaking of Patient-Generated Health Data: A Mixed-Method Study with Healthcare Professionals in Cardiac Risk Reduction

arXiv:2602.05687v2 Announce Type: replace-cross Abstract: Individuals are increasingly generating substantial personal health and lifestyle data, e.g. through wearables and smartphones. While such data could transform preventative care, its integration into clinical practice is hindered by its scale, heterogeneity and the time pressure and data literacy of healthcare professionals (HCPs). We explore how large language models (LLMs) can support sensemaking of patient-generated health data (PGHD) with automated summaries and natural language data exploration. Using cardiovascular disease (CVD) risk reduction as a use case, 16 HCPs reviewed multimodal PGHD in a mixed-methods study with a prototype that integrated common charts, LLM-generated summaries, and a conversational interface. Findings show that AI summaries provided quick overviews that anchored exploration, while conversational interaction supported flexible analysis and bridged data-literacy gaps. However, HCPs raised concerns about transparency, privacy, and overreliance. We contribute empirical insights and sociotechnical design implications for integrating AI-driven summarization and conversation into clinical workflows to support PGHD sensemaking.
❌