❌

Reading view

AnyECG: Evolved ECG Foundation Model for Holistic Health Profiling

arXiv:2601.10748v1 Announce Type: cross Abstract: Background: Artificial intelligence enabled electrocardiography (AI-ECG) has demonstrated the ability to detect diverse pathologies, but most existing models focus on single disease identification, neglecting comorbidities and future risk prediction. Although ECGFounder expanded cardiac disease coverage, a holistic health profiling model remains needed. Methods: We constructed a large multicenter dataset comprising 13.3 million ECGs from 2.98 million patients. Using transfer learning, ECGFounder was fine-tuned to develop AnyECG, a foundation model for holistic health profiling. Performance was evaluated using external validation cohorts and a 10-year longitudinal cohort for current diagnosis, future risk prediction, and comorbidity identification. Results: AnyECG demonstrated systemic predictive capability across 1172 conditions, achieving an AUROC greater than 0.7 for 306 diseases. The model revealed novel disease associations, robust comorbidity patterns, and future disease risks. Representative examples included high diagnostic performance for hyperparathyroidism (AUROC 0.941), type 2 diabetes (0.803), Crohn disease (0.817), lymphoid leukemia (0.856), and chronic obstructive pulmonary disease (0.773). Conclusion: The AnyECG foundation model provides substantial evidence that AI-ECG can serve as a systemic tool for concurrent disease detection and long-term risk prediction.
  •  

MetaboNet: The Largest Publicly Available Consolidated Dataset for Type 1 Diabetes Management

arXiv:2601.11505v1 Announce Type: cross Abstract: Progress in Type 1 Diabetes (T1D) algorithm development is limited by the fragmentation and lack of standardization across existing T1D management datasets. Current datasets differ substantially in structure and are time-consuming to access and process, which impedes data integration and reduces the comparability and generalizability of algorithmic developments. This work aims to establish a unified and accessible data resource for T1D algorithm development. Multiple publicly available T1D datasets were consolidated into a unified resource, termed the MetaboNet dataset. Inclusion required the availability of both continuous glucose monitoring (CGM) data and corresponding insulin pump dosing records. Additionally, auxiliary information such as reported carbohydrate intake and physical activity was retained when present. The MetaboNet dataset comprises 3135 subjects and 1228 patient-years of overlapping CGM and insulin data, making it substantially larger than existing standalone benchmark datasets. The resource is distributed as a fully public subset available for immediate download at https://metabo-net.org/ , and with a Data Use Agreement (DUA)-restricted subset accessible through their respective application processes. For the datasets in the latter subset, processing pipelines are provided to automatically convert the data into the standardized MetaboNet format. A consolidated public dataset for T1D research is presented, and the access pathways for both its unrestricted and DUA-governed components are described. The resulting dataset covers a broad range of glycemic profiles and demographics and thus can yield more generalizable algorithmic performance than individual datasets.
  •  

The AI healthcare gold rush is here

AI companies are clustering around healthcare and fast.  In just the past week, OpenAI bought health startup Torch, Anthropic launched Claude for healthcare, and Sam Altman-backed MergeLabs closed a $250 million seed round at an $850 million valuation. The money and products are pouring into health and voice AI, but so are concerns about hallucination risks, inaccurate medical information, and […]
  •  

Cellular neighborhoods in cancer

Nat Cancer. 2026 Jan 16. doi: 10.1038/s43018-025-01107-w. Online ahead of print.

ABSTRACT

The concept of cellular neighborhoods, defined as recurring structures within the tissue with characteristic cell compositions and interactions, has transformed our understanding of the complexity and dynamics of tumor ecosystems. Recent advances in spatial omics and computational modeling have enabled high-resolution mapping of these neighborhoods, providing unprecedented insights into their roles in shaping tumor heterogeneity, evolution and therapeutic responses. Despite these advances, a unified framework for interpreting cellular neighborhoods remains lacking. This Perspective synthesizes emerging concepts and insights, focusing on the definition and classification of cellular neighborhoods in cancer, computational methods for identifying and comparing them, and their clinical relevance.

PMID:41545713 | DOI:10.1038/s43018-025-01107-w

  •  

Contaminating plasmid sequences and disrupted vector genomes in the liver following adeno-associated virus gene therapy

Nature Medicine, Published online: 16 January 2026; doi:10.1038/s41591-025-04073-z

Analyses of liver biopsies from a child with spinal muscular atrophy treated with adeno-associated virus gene therapy who developed hepatitis reveal contaminating manufacturing plasmids and disrupted vector genomes, possibly resulting from recombination events.
  •  

Exclusive eBook: How AGI Became a Consequential Conspiracy Theory

In this exclusive subscriber-only eBook, you’ll learn about how the idea that machines will be as smart as—or smarter than—humans has hijacked an entire industry.

by Will Douglas Heaven October 30, 2025

Table of Contents:

  • How Silicon Valley got AGI-pilled
  • The great AGI conspiracy
  • How AGI hijacked an industry
  • The great AGI conspiracy, concluded

Related Stories:

Access all subscriber-only eBooks:

  •  

Human-AI Co-design for Clinical Prediction Models

arXiv:2601.09072v1 Announce Type: new Abstract: Developing safe, effective, and practically useful clinical prediction models (CPMs) traditionally requires iterative collaboration between clinical experts, data scientists, and informaticists. This process refines the often small but critical details of the model building process, such as which features/patients to include and how clinical categories should be defined. However, this traditional collaboration process is extremely time- and resource-intensive, resulting in only a small fraction of CPMs reaching clinical practice. This challenge intensifies when teams attempt to incorporate unstructured clinical notes, which can contain an enormous number of concepts. To address this challenge, we introduce HACHI, an iterative human-in-the-loop framework that uses AI agents to accelerate the development of fully interpretable CPMs by enabling the exploration of concepts in clinical notes. HACHI alternates between (i) an AI agent rapidly exploring and evaluating candidate concepts in clinical notes and (ii) clinical and domain experts providing feedback to improve the CPM learning process. HACHI defines concepts as simple yes-no questions that are used in linear models, allowing the clinical AI team to transparently review, refine, and validate the CPM learned in each round. In two real-world prediction tasks (acute kidney injury and traumatic brain injury), HACHI outperforms existing approaches, surfaces new clinically relevant concepts not included in commonly-used CPMs, and improves model generalizability across clinical sites and time periods. Furthermore, HACHI reveals the critical role of the clinical AI team, such as directing the AI agent to explore concepts that it had not previously considered, adjusting the granularity of concepts it considers, changing the objective function to better align with the clinical objectives, and identifying issues of data bias and leakage.
  •  

Companion Agents: A Table-Information Mining Paradigm for Text-to-SQL

arXiv:2601.08838v1 Announce Type: cross Abstract: Large-scale Text-to-SQL benchmarks such as BIRD typically assume complete and accurate database annotations as well as readily available external knowledge, which fails to reflect common industrial settings where annotations are missing, incomplete, or erroneous. This mismatch substantially limits the real-world applicability of state-of-the-art (SOTA) Text-to-SQL systems. To bridge this gap, we explore a database-centric approach that leverages intrinsic, fine-grained information residing in relational databases to construct missing evidence and improve Text-to-SQL accuracy under annotation-scarce conditions. Our key hypothesis is that when a query requires multi-step reasoning over extensive table information, existing methods often struggle to reliably identify and utilize the truly relevant knowledge. We therefore propose to "cache" query-relevant knowledge on the database side in advance, so that it can be selectively activated at inference time. Based on this idea, we introduce Companion Agents (CA), a new Text-to-SQL paradigm that incorporates a group of agents accompanying database schemas to proactively mine and consolidate hidden inter-table relations, value-domain distributions, statistical regularities, and latent semantic cues before query generation. Experiments on BIRD under the fully missing evidence setting show that CA recovers +4.49 / +4.37 / +14.13 execution accuracy points on RSL-SQL / CHESS / DAIL-SQL, respectively, with larger gains on the Challenging subset +9.65 / +7.58 / +16.71. These improvements stem from CA's automatic database-side mining and evidence construction, suggesting a practical path toward industrial-grade Text-to-SQL deployment without reliance on human-curated evidence.
  •  

Triples and Knowledge-Infused Embeddings for Clustering and Classification of Scientific Documents

arXiv:2601.08841v1 Announce Type: cross Abstract: The increasing volume and complexity of scientific literature demand robust methods for organizing and understanding research documents. In this study, we explore how structured knowledge, specifically, subject-predicate-object triples, can enhance the clustering and classification of scientific papers. We propose a modular pipeline that combines unsupervised clustering and supervised classification over multiple document representations: raw abstracts, extracted triples, and hybrid formats that integrate both. Using a filtered arXiv corpus, we extract relational triples from abstracts and construct four text representations, which we embed using four state-of-the-art transformer models: MiniLM, MPNet, SciBERT, and SPECTER. We evaluate the resulting embeddings with KMeans, GMM, and HDBSCAN for unsupervised clustering, and fine-tune classification models for arXiv subject prediction. Our results show that full abstract text yields the most coherent clusters, but that hybrid representations incorporating triples consistently improve classification performance, reaching up to 92.6% accuracy and 0.925 macro-F1. We also find that lightweight sentence encoders (MiniLM, MPNet) outperform domain-specific models (SciBERT, SPECTER) in clustering, while SciBERT excels in structured-input classification. These findings highlight the complementary benefits of combining unstructured text with structured knowledge, offering new insights into knowledge-infused representations for semantic organization of scientific documents.
  •  

A Marketplace for AI-Generated Adult Content and Deepfakes

arXiv:2601.09117v1 Announce Type: cross Abstract: Generative AI systems increasingly enable the production of highly realistic synthetic media. Civitai, a popular community-driven platform for AI-generated content, operates a monetized feature called Bounties, which allows users to commission the generation of content in exchange for payment. To examine how this mechanism is used and what content it incentivizes, we conduct a longitudinal analysis of all publicly available bounty requests collected over a 14-month period following the platform's launch. We find that the bounty marketplace is dominated by tools that let users steer AI models toward content they were not trained to generate. At the same time, requests for content that is "Not Safe For Work" are widespread and have increased steadily over time, now comprising a majority of all bounties. Participation in bounty creation is uneven, with 20% of requesters accounting for roughly half of requests. Requests for "deepfake" - media depicting identifiable real individuals - exhibit a higher concentration than other types of bounties. A nontrivial subset of these requests involves explicit deepfakes despite platform policies prohibiting such content. These bounties disproportionately target female celebrities, revealing a pronounced gender asymmetry in social harm. Together, these findings show how monetized, community-driven generative AI platforms can produce gendered harms, raising questions about consent, governance, and enforcement.
  •  

Global Benchmark Database

arXiv:2405.10045v3 Announce Type: replace-cross Abstract: This paper presents Global Benchmark Database (GBD), a comprehensive suite of tools for provisioning and sustainably maintaining benchmark instances and their metadata. The availability of benchmark metadata is essential for many tasks in empirical research, e.g., for the data-driven compilation of benchmarks, the domain-specific analysis of runtime experiments, or the instance-specific selection of solvers. In this paper, we introduce the data model of GBD as well as its interfaces and provide examples of how to interact with them. We also demonstrate the integration of custom data sources and explain how to extend GBD with additional problem domains, instance formats and feature extractors.
  •  

Exploring the Secondary Risks of Large Language Models

arXiv:2506.12382v4 Announce Type: replace-cross Abstract: Ensuring the safety and alignment of Large Language Models is a significant challenge with their growing integration into critical applications and societal functions. While prior research has primarily focused on jailbreak attacks, less attention has been given to non-adversarial failures that subtly emerge during benign interactions. We introduce secondary risks a novel class of failure modes marked by harmful or misleading behaviors during benign prompts. Unlike adversarial attacks, these risks stem from imperfect generalization and often evade standard safety mechanisms. To enable systematic evaluation, we introduce two risk primitives verbose response and speculative advice that capture the core failure patterns. Building on these definitions, we propose SecLens, a black-box, multi-objective search framework that efficiently elicits secondary risk behaviors by optimizing task relevance, risk activation, and linguistic plausibility. To support reproducible evaluation, we release SecRiskBench, a benchmark dataset of 650 prompts covering eight diverse real-world risk categories. Experimental results from extensive evaluations on 16 popular models demonstrate that secondary risks are widespread, transferable across models, and modality independent, emphasizing the urgent need for enhanced safety mechanisms to address benign yet harmful LLM behaviors in real-world deployments.
  •  

GI-Bench: A Panoramic Benchmark Revealing the Knowledge-Experience Dissociation of Multimodal Large Language Models in Gastrointestinal Endoscopy Against Clinical Standards

arXiv:2601.08183v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) show promise in gastroenterology, yet their performance against comprehensive clinical workflows and human benchmarks remains unverified. To systematically evaluate state-of-the-art MLLMs across a panoramic gastrointestinal endoscopy workflow and determine their clinical utility compared with human endoscopists. We constructed GI-Bench, a benchmark encompassing 20 fine-grained lesion categories. Twelve MLLMs were evaluated across a five-stage clinical workflow: anatomical localization, lesion identification, diagnosis, findings description, and management. Model performance was benchmarked against three junior endoscopists and three residency trainees using Macro-F1, mean Intersection-over-Union (mIoU), and multi-dimensional Likert scale. Gemini-3-Pro achieved state-of-the-art performance. In diagnostic reasoning, top-tier models (Macro-F1 0.641) outperformed trainees (0.492) and rivaled junior endoscopists (0.727; p>0.05). However, a critical "spatial grounding bottleneck" persisted; human lesion localization (mIoU >0.506) significantly outperformed the best model (0.345; p
  •  

Circulating metabolites, genetics and lifestyle factors in relation to future risk of type 2 diabetes

Nat Med. 2026 Jan 14. doi: 10.1038/s41591-025-04105-8. Online ahead of print.

ABSTRACT

The human metabolome reflects complex metabolic states affected by genetic and environmental factors. However, metabolites associated with type 2 diabetes (T2D) risk and their determinants remain insufficiently characterized. Here we integrated blood metabolomic, genomic and lifestyle data from up to 23,634 initially T2D-free participants from ten cohorts. Of 469 metabolites examined, 235 were associated with incident T2D during up to 26 years of follow-up, including 67 associations not previously reported across bile acid, lipid, carnitine, urea cycle and arginine/proline, glycine and histidine pathways. Further genetic analyses linked these metabolites to signaling pathways and clinical traits central to T2D pathophysiology, including insulin resistance, glucose/insulin response, ectopic fat deposition, energy/lipid regulation and liver function. Lifestyle factors-particularly physical activity, obesity and diet-explained greater variations in T2D-associated versus non-associated metabolites, with specific metabolites revealed as potential mediators. Finally, a 44-metabolite signature improved T2D risk prediction beyond conventional factors. These findings provide a foundation for understanding T2D mechanisms and may inform precision prevention targeting specific metabolic pathways.

PMID:41535386 | DOI:10.1038/s41591-025-04105-8

  •  

Multi-omics to study chronic respiratory diseases and viral infections

Eur Respir Rev. 2026 Jan 14;35(179):240286. doi: 10.1183/16000617.0286-2024. Print 2026 Jan.

ABSTRACT

Despite recent advances, the underlying mechanisms of the development and progression of many chronic respiratory diseases remain to be elucidated. Factors such as heterogeneity and complexity of human diseases and difficulty interpreting large datasets hinder research into chronic respiratory diseases. Omics assesses the changes in specific biological entities, such as mRNA expression, epigenetics/epigenomics, genomics, proteomics, metagenomics and metabolomics, and provides valuable insights into the roles of these processes in chronic respiratory diseases. High-throughput omics at bulk, single-cell and spatial levels empower the exploration of disease-related changes through untargeted data-driven statistical methods. Multi-omics is the exploration and integration of multiple biological processes, which compared to a single-omics, can provide a substantially greater and more holistic overview of the pathogenic mechanisms that underpin complex diseases. Multi-omics analysis can comprehensively characterise the mechanisms that drive chronic respiratory diseases, capturing unique biological signatures and cellular interactions at different omics levels. Use of these methods has begun to identify key factors and biomarkers in chronic respiratory diseases. Here, we review current omics approaches and highlight recent advances in respiratory research achieved using multi-omics and integrative methods. Our review provides a valuable resource for researchers and clinicians in this area.

PMID:41534886 | DOI:10.1183/16000617.0286-2024

  •  

Complement-secreting CAFs are associated with better prognosis in pancreatic cancer: single-cell multiomics

Gut. 2026 Jan 13:gutjnl-2025-335683. doi: 10.1136/gutjnl-2025-335683. Online ahead of print.

ABSTRACT

BACKGROUND: Accumulating evidence has demonstrated that distinct tumour-promoting and tumour-restraining cancer-associated fibroblast (CAF) subtypes coexist in pancreatic ductal adenocarcinoma.

OBJECTIVE: To develop targeted CAF therapeutic strategies by reprogramming tumour-promoting CAF subtypes.

DESIGN: We leveraged multiomics technologies to systematically identify and characterise CAF subtypes transcriptionally, epigenetically and spatially and correlate them with clinicopathological features.

RESULTS: We found that complement-secreting CAFs (csCAFs), initially identified by our group and inflammatory CAFs (iCAFs) share significant overlap in their transcriptional profiles and chromatin accessibility. iCAFs specifically express transcription factors from the heme and oxidative homeostasis pathway and the activator protein 1 family, which are both involved in cellular response to oxidative stress. Notably, the composition of csCAFs among all CAFs declined during pancreatic carcinogenesis, while trajectory analysis showed that csCAFs could potentially differentiate into iCAFs. Spatially resolved analysis indicated that tumour regions with a higher csCAF composition were associated with lower levels of TGF-β ligands, fewer M2 tumour-associated macrophages and increased levels of lipid mediators. Additionally, we identified a spatially defined CXCL12-CXCR4 ligand-receptor interaction between csCAFs and T cells, but in distinct patterns between different metastatic organs. Patients with a higher composition of csCAFs have significantly longer overall survival and recurrence-free survival through multiplex immunohistochemistry and bulk RNA-seq deconvolution.

CONCLUSION: Our study demonstrates that csCAFs may represent an early-stage iCAF subtype and suggests a promising strategy for reprogramming iCAFs into csCAFs.

PMID:41534892 | DOI:10.1136/gutjnl-2025-335683

  •  

Evidence for Digital Health Tools Designed to Support the Triage of Musculoskeletal Conditions in Primary, Urgent, and Emergency Care Settings: Scoping Review

Background: The digital health research field is growing rapidly, and a summary of the available digital tools for triaging musculoskeletal conditions is needed. Effective and safe digital triage tools for musculoskeletal conditions could support patients in making informed care decisions, aid clinicians and patients in navigating care, and may contribute to reducing ED overcrowding and healthcare costs. Objective: To identify and describe digital health tools for use by adults to triage musculoskeletal conditions across primary, urgent, or emergency care settings. Methods: Our scoping review was conducted following the Johanna Briggs Institute recommendations for scoping reviews and Arksey & O’Malley’s framework. Systematic searches in MEDLINE (OVID), CINAHL (EBSCO), PsycINFO (EBSCO), Embase (OVID), Cochrane Library, Web of Science, OpenGrey, GoogleScholar, arXiv.org, medRxiv.org, and an extensive grey literature search were conducted with a librarian scientist from inception to Sept 18, 2025. Studies had to recruit adults (18+ years) with musculoskeletal conditions that identified a digital health tool designed to triage or diagnose in primary, urgent, or emergency care settings and report primary data to be included. Two reviewer pairs independently screened abstracts and full-text articles Relevant data were extracted in duplicate, and results were summarized descriptively. Results: The search yielded 5695 records, and we screened 189 full-text articles. Thirty-four studies (n=37,509 patients) met the inclusion criteria. The most common musculoskeletal conditions reported were rheumatoid/inflammatory arthritis (n=13, 38%). Nineteen (59%) studies reported on symptom checkers, 13 (44%) studies on triage/diagnosis tools, and 2 (6%) were studies of diagnostic predictor tools. There were 16 unique digital health tools. Two tools were built for triaging musculoskeletal conditions, and were not publicly available outside the UK National Health Service. Most tools were generic tools designed to screen for general health problems, including musculoskeletal conditions. The most common approach to evaluating performance (eg, accuracy) of the tools was to compare the concordance of the tool to a clinician diagnosis or triage recommendation. Sensitivity and specificity ranged from 39%-91% and 23%-80%, respectively. Reported accuracy of included tools ranged from 33% to 98%. Conclusions: Musculoskeletal conditions remain a blind spot for people designing, implementing, and evaluating digital health for triage: few tools were specifically designed for musculoskeletal conditions, and most existing tools performed poorly when applied to musculoskeletal populations. The evidence base supporting accuracy of digital health for triaging and diagnosing musculoskeletal conditions is weak, and tool performance was inconsistent and lacking transparency. We recommend health systems and clinicians use a multi-modal approach, integrating both digital health tools and clinical decision-making to safely triage and diagnose until a more robust tool for musculoskeletal conditions is available. Future tool developers need to use transparent, standardized processes that prioritize tool safety, clinical value, and trustworthiness when designing for clinicians and patients.
  •  
❌