❌

Reading view

Extrachromosomal DNA drives molecular and clinical heterogeneity in hepatocellular carcinoma: a multi-omics analysis and prognostic model development

Hum Genomics. 2026 Feb 3. doi: 10.1186/s40246-026-00927-w. Online ahead of print.

ABSTRACT

BACKGROUND: Extrachromosomal DNA (ecDNA) is an emerging hallmark of cancer that promotes tumor evolution and heterogeneity. However, the molecular characteristics and clinical significance of ecDNA in hepatocellular carcinoma (HCC) remain incompletely understood.

METHODS: The clinical outcomes, genomics, transcriptomics, proteomics, tumor microenvironment, and drug target landscapes of ecDNA-negative and ecDNA-positive HCC in the Cancer Genome Atlas (TCGA) were compared. Next, the least absolute shrinkage and selection operator (LASSO) and random survival forest (RSF) algorithms were used to screen the ecDNA gene signature. A nomogram was constructed and evaluated based on the risk score and clinicopathological features. Finally, the role of DNASE1L3 was validated through in vitro experiments.

RESULTS: EcDNA-positive tumors showed increased vascular invasion, higher AFP levels, and more TP53 mutations. These tumors displayed unique activation of proliferation pathways, decreased stromal infiltration, and heightened immune activation. Our validated six-gene signature (RNF186, BMP6, AOC1, FBLL1, MYBL2, and DNASE1L3) demonstrated strong prognostic value when combined with tumor stage in the nomogram. Notably, DNASE1L3 was downregulated in HCC, showed endothelial cell-specific expression, and suppressed the proliferation and migration of Hep3B2.1-7 cells.

CONCLUSION: Our study characterizes the molecular and clinical distinctions between ecDNA-negative and ecDNA-positive HCC and establishes a clinically applicable gene signature for patient prognosis. These findings advance our understanding of ecDNA-driven tumor heterogeneity and provide potential strategies for personalized HCC management.

PMID:41634868 | DOI:10.1186/s40246-026-00927-w

  •  

Integrative proteogenomics maps multifactorial aetiology, progression and therapeutic vulnerabilities in gastric cancer

Gut. 2026 Jan 30:gutjnl-2025-337247. doi: 10.1136/gutjnl-2025-337247. Online ahead of print.

ABSTRACT

BACKGROUND: Gastric cancer, with disproportionately higher incidence in East Asia, arises from complex host-microbiome-environment interactions beyond Helicobacter pylori (HP) infection. However, the molecular architecture linking environmental carcinogens, microbial succession and host response remains unclear.

OBJECTIVE: To delineate multifactorial aetiologies and clinically actionable subtypes/biomarkers of gastric cancer through integrative proteogenomic, microbial and environmental exposure profiling.

DESIGN: We established a multiomics atlas of paired tumour, adjacent mucosa tissues and blood from 154 treatment-naïve Taiwanese patients, integrating whole-exome sequencing, RNA-seq, proteome and phosphoproteome profiling with carcinogen signatures, HP status, microbiome composition and refined anatomical mapping. Cell-based functional assays tested carcinogen effects. Microbial subtype was assessed in an independent cohort.

RESULTS: A polycyclic-aromatic-hydrocarbon signature, dibenz[a,h]acridine, emerged as a high-risk exposure promoting invasion, immune suppression and poor survival, significantly exceeding nitrosamine-linked risk in this cohort. Multilayer integration defined three initiation ecologies: HP-driven inflammatory, non-HP microbiome-enriched immune-silent and HP-free microbially depleted states. Among HP-negative tumours, a Streptococcus-enriched subtype associated with tight-junction (CLDN18.2/ZO-1/OCLN) disruption and epithelial-mesenchymal transition, whereas a subset of clinically aggressive cases retained CLDN18.2-high epithelial-stable subtype for therapeutic accessibility. An independent cohort revealed gastric juice-derived Streptococcus anginosus abundance inversely correlated with tight-junction proteins. Anatomical mapping reveals location-specific, sex-specific, subtype-specific oncogenic networks and kinase activity, including CDK4 activation in clinical biomarker-negative tumours. Decision-tree models combining exposure and proteome-immune states refined recurrence and survival prediction beyond stage.

CONCLUSION: This proteogenomic framework defines exposure-informed and microbiome-informed gastric cancer subtypes, providing a molecular schema for patient stratification, prevention and actionable therapeutic vulnerabilities.

PMID:41617485 | DOI:10.1136/gutjnl-2025-337247

  •  

Federated Proximal Optimization for Privacy-Preserving Heart Disease Prediction: A Controlled Simulation Study on Non-IID Clinical Data

arXiv:2601.17183v1 Announce Type: cross Abstract: Healthcare institutions have access to valuable patient data that could be of great help in the development of improved diagnostic models, but privacy regulations like HIPAA and GDPR prevent hospitals from directly sharing data with one another. Federated Learning offers a way out to this problem by facilitating collaborative model training without having the raw patient data centralized. However, clinical datasets intrinsically have non-IID (non-independent and identically distributed) features brought about by demographic disparity and diversity in disease prevalence and institutional practices. This paper presents a comprehensive simulation research of Federated Proximal Optimization (FedProx) for Heart Disease prediction based on UCI Heart Disease dataset. We generate realistic non-IID data partitions by simulating four heterogeneous hospital clients from the Cleveland Clinic dataset (303 patients), by inducing statistical heterogeneity by demographic-based stratification. Our experimental results show that FedProx with proximal parameter mu=0.05 achieves 85.00% accuracy, which is better than both centralized learning (83.33%) and isolated local models (78.45% average) without revealing patient privacy. Through generous sheer ablation studies with statistical validation on 50 independent runs we demonstrate that proximal regularization is effective in curbing client drift in heterogeneous environments. This proof-of-concept research offers algorithmic insights and practical deployment guidelines for real-world federated healthcare systems, and thus, our results are directly transferable to hospital IT-administrators, implementing privacy-preserving collaborative learning.
  •  

The Limits of AI Data Transparency Policy: Three Disclosure Fallacies

arXiv:2601.18127v1 Announce Type: cross Abstract: Data transparency has emerged as a rallying cry for addressing concerns about AI: data quality, privacy, and copyright chief among them. Yet while these calls are crucial for accountability, current transparency policies often fall short of their intended aims. Similar to nutrition facts for food, policies aimed at nutrition facts for AI currently suffer from a limited consideration of research on effective disclosures. We offer an institutional perspective and identify three common fallacies in policy implementations of data disclosures for AI. First, many data transparency proposals exhibit a specification gap between the stated goals of data transparency and the actual disclosures necessary to achieve such goals. Second, reform attempts exhibit an enforcement gap between required disclosures on paper and enforcement to ensure compliance in fact. Third, policy proposals manifest an impact gap between disclosed information and meaningful changes in developer practices and public understanding. Informed by the social science on transparency, our analysis identifies affirmative paths for transparency that are effective rather than merely symbolic.
  •  

Multimodal digital biopsy for preoperative prediction of occult peritoneal metastasis in gastric cancer

npj Digital Medicine, Published online: 26 January 2026; doi:10.1038/s41746-025-02268-9

Multimodal digital biopsy for preoperative prediction of occult peritoneal metastasis in gastric cancer
  •  

PyHealth 2.0: A Comprehensive Open-Source Toolkit for Accessible and Reproducible Clinical Deep Learning

arXiv:2601.16414v1 Announce Type: cross Abstract: Difficulty replicating baselines, high computational costs, and required domain expertise create persistent barriers to clinical AI research. To address these challenges, we introduce PyHealth 2.0, an enhanced clinical deep learning toolkit that enables predictive modeling in as few as 7 lines of code. PyHealth 2.0 offers three key contributions: (1) a comprehensive toolkit addressing reproducibility and compatibility challenges by unifying 15+ datasets, 20+ clinical tasks, 25+ models, 5+ interpretability methods, and uncertainty quantification including conformal prediction within a single framework that supports diverse clinical data modalities - signals, imaging, and electronic health records - with translation of 5+ medical coding standards; (2) accessibility-focused design accommodating multimodal data and diverse computational resources with up to 39x faster processing and 20x lower memory usage, enabling work from 16GB laptops to production systems; and (3) an active open-source community of 400+ members lowering domain expertise barriers through extensive documentation, reproducible research contributions, and collaborations with academic health systems and industry partners, including multi-language support via RHealth. PyHealth 2.0 establishes an open-source foundation and community advancing accessible, reproducible healthcare AI. Available at pip install pyhealth.
  •  

DeepEra: A Deep Evidence Reranking Agent for Scientific Retrieval-Augmented Generated Question Answering

arXiv:2601.16478v1 Announce Type: cross Abstract: With the rapid growth of scientific literature, scientific question answering (SciQA) has become increasingly critical for exploring and utilizing scientific knowledge. Retrieval-Augmented Generation (RAG) enhances LLMs by incorporating knowledge from external sources, thereby providing credible evidence for scientific question answering. But existing retrieval and reranking methods remain vulnerable to passages that are semantically similar but logically irrelevant, often reducing factual reliability and amplifying hallucinations.To address this challenge, we propose a Deep Evidence Reranking Agent (DeepEra) that integrates step-by-step reasoning, enabling more precise evaluation of candidate passages beyond surface-level semantics. To support systematic evaluation, we construct SciRAG-SSLI (Scientific RAG - Semantically Similar but Logically Irrelevant), a large-scale dataset comprising about 300K SciQA instances across 10 subjects, constructed from 10M scientific corpus. The dataset combines naturally retrieved contexts with systematically generated distractors to test logical robustness and factual grounding. Comprehensive evaluations confirm that our approach achieves superior retrieval performance compared to leading rerankers. To our knowledge, this work is the first to comprehensively study and empirically validate innegligible SSLI issues in two-stage RAG frameworks.
  •  

An Optimized Decision Tree-Based Framework for Explainable IoT Anomaly Detection

arXiv:2601.14305v1 Announce Type: cross Abstract: The increase in the number of Internet of Things (IoT) devices has tremendously increased the attack surface of cyber threats thus making a strong intrusion detection system (IDS) with a clear explanation of the process essential towards resource-constrained environments. Nevertheless, current IoT IDS systems are usually traded off with detection quality, model elucidability, and computational effectiveness, thus the deployment on IoT devices. The present paper counteracts these difficulties by suggesting an explainable AI (XAI) framework based on an optimized Decision Tree classifier with both local and global importance methods: SHAP values that estimate feature attribution using local explanations, and Morris sensitivity analysis that identifies the feature importance in a global view. The proposed system attains the state of art on the test performance with 99.91% accuracy, F1-score of 99.51% and Cohen Kappa of 0.9960 and high stability is confirmed by a cross validation mean accuracy of 98.93%. Efficiency is also enhanced in terms of computations to provide faster inferences compared to those that are generalized in ensemble models. SrcMac has shown as the most significant predictor in feature analyses according to SHAP and Morris methods. Compared to the previous work, our solution eliminates its major drawback lack because it allows us to apply it to edge devices and, therefore, achieve real-time processing, adhere to the new regulation of transparency in AI, and achieve high detection rates on attacks of dissimilar classes. This combination performance of high accuracy, explainability, and low computation make the framework useful and reliable as a resource-constrained IoT security problem in real environments.
  •  

OpenNovelty: An LLM-powered Agentic System for Verifiable Scholarly Novelty Assessment

arXiv:2601.01576v2 Announce Type: replace-cross Abstract: Evaluating novelty is critical yet challenging in peer review, as reviewers must assess submissions against a vast, rapidly evolving literature. This report presents OpenNovelty, an LLM-powered agentic system for transparent, evidence-based novelty analysis. The system operates through four phases: (1) extracting the core task and contribution claims to generate retrieval queries; (2) retrieving relevant prior work based on extracted queries via semantic search engine; (3) constructing a hierarchical taxonomy of core-task-related work and performing contribution-level full-text comparisons against each contribution; and (4) synthesizing all analyses into a structured novelty report with explicit citations and evidence snippets. Unlike naive LLM-based approaches, \textsc{OpenNovelty} grounds all assessments in retrieved real papers, ensuring verifiable judgments. We deploy our system on 500+ ICLR 2026 submissions with all reports publicly available on our website, and preliminary analysis suggests it can identify relevant prior work, including closely related papers that authors may overlook. OpenNovelty aims to empower the research community with a scalable tool that promotes fair, consistent, and evidence-backed peer review.
  •  

Japanese AI Agent System on Human Papillomavirus Vaccination: System Design

arXiv:2601.10718v1 Announce Type: new Abstract: Human papillomavirus (HPV) vaccine hesitancy poses significant public health challenges, particularly in Japan where proactive vaccination recommendations were suspended from 2013 to 2021. The resulting information gap is exacerbated by misinformation on social media, and traditional ways cannot simultaneously address individual queries while monitoring population-level discourse. This study aimed to develop a dual-purpose AI agent system that provides verified HPV vaccine information through a conversational interface while generating analytical reports for medical institutions based on user interactions and social media. We implemented a system comprising: a vector database integrating academic papers, government sources, news media, and social media; a Retrieval-Augmented Generation chatbot using ReAct agent architecture with multi-tool orchestration across five knowledge sources; and an automated report generation system with modules for news analysis, research synthesis, social media sentiment analysis, and user interaction pattern identification. Performance was assessed using a 0-5 scoring scale. For single-turn evaluation, the chatbot achieved mean scores of 4.83 for relevance, 4.89 for routing, 4.50 for reference quality, 4.90 for correctness, and 4.88 for professional identity (overall 4.80). Multi-turn evaluation yielded higher scores: context retention 4.94, topic coherence 5.00, and overall 4.98. The report generation system achieved completeness 4.00-5.00, correctness 4.00-5.00, and helpfulness 3.67-5.00, with reference validity 5.00 across all periods. This study demonstrates the feasibility of an integrated AI agent system for bidirectional HPV vaccine communication. The architecture enables verified information delivery with source attribution while providing systematic public discourse analysis, with a transferable framework for adaptation to other medical contexts.
  •  

AnyECG: Evolved ECG Foundation Model for Holistic Health Profiling

arXiv:2601.10748v1 Announce Type: cross Abstract: Background: Artificial intelligence enabled electrocardiography (AI-ECG) has demonstrated the ability to detect diverse pathologies, but most existing models focus on single disease identification, neglecting comorbidities and future risk prediction. Although ECGFounder expanded cardiac disease coverage, a holistic health profiling model remains needed. Methods: We constructed a large multicenter dataset comprising 13.3 million ECGs from 2.98 million patients. Using transfer learning, ECGFounder was fine-tuned to develop AnyECG, a foundation model for holistic health profiling. Performance was evaluated using external validation cohorts and a 10-year longitudinal cohort for current diagnosis, future risk prediction, and comorbidity identification. Results: AnyECG demonstrated systemic predictive capability across 1172 conditions, achieving an AUROC greater than 0.7 for 306 diseases. The model revealed novel disease associations, robust comorbidity patterns, and future disease risks. Representative examples included high diagnostic performance for hyperparathyroidism (AUROC 0.941), type 2 diabetes (0.803), Crohn disease (0.817), lymphoid leukemia (0.856), and chronic obstructive pulmonary disease (0.773). Conclusion: The AnyECG foundation model provides substantial evidence that AI-ECG can serve as a systemic tool for concurrent disease detection and long-term risk prediction.
  •  

Contaminating plasmid sequences and disrupted vector genomes in the liver following adeno-associated virus gene therapy

Nature Medicine, Published online: 16 January 2026; doi:10.1038/s41591-025-04073-z

Analyses of liver biopsies from a child with spinal muscular atrophy treated with adeno-associated virus gene therapy who developed hepatitis reveal contaminating manufacturing plasmids and disrupted vector genomes, possibly resulting from recombination events.
  •  

Circulating metabolites, genetics and lifestyle factors in relation to future risk of type 2 diabetes

Nat Med. 2026 Jan 14. doi: 10.1038/s41591-025-04105-8. Online ahead of print.

ABSTRACT

The human metabolome reflects complex metabolic states affected by genetic and environmental factors. However, metabolites associated with type 2 diabetes (T2D) risk and their determinants remain insufficiently characterized. Here we integrated blood metabolomic, genomic and lifestyle data from up to 23,634 initially T2D-free participants from ten cohorts. Of 469 metabolites examined, 235 were associated with incident T2D during up to 26 years of follow-up, including 67 associations not previously reported across bile acid, lipid, carnitine, urea cycle and arginine/proline, glycine and histidine pathways. Further genetic analyses linked these metabolites to signaling pathways and clinical traits central to T2D pathophysiology, including insulin resistance, glucose/insulin response, ectopic fat deposition, energy/lipid regulation and liver function. Lifestyle factors-particularly physical activity, obesity and diet-explained greater variations in T2D-associated versus non-associated metabolites, with specific metabolites revealed as potential mediators. Finally, a 44-metabolite signature improved T2D risk prediction beyond conventional factors. These findings provide a foundation for understanding T2D mechanisms and may inform precision prevention targeting specific metabolic pathways.

PMID:41535386 | DOI:10.1038/s41591-025-04105-8

  •  

A nowhere-to-hide mechanism ensures complete piRNA-directed DNA methylation

Nature, Published online: 14 January 2026; doi:10.1038/s41586-025-09940-w

In mice, a SPOCD1–TPR-dependent ‘nowhere-to-hide’ mechanism is required for complete non-stochastic piRNA-directed LINE1 DNA methylation by preventing transposons from escaping surveillance within heterochromatin.
  •  
❌