❌

Normal view

Artificial intelligence in hepatopathy diagnosis and treatment: Big data analytics, deep learning, and clinical prediction models

World J Gastroenterol. 2025 Dec 14;31(46):111176. doi: 10.3748/wjg.v31.i46.111176.

ABSTRACT

Artificial intelligence (AI) is rapidly transforming the landscape of hepatology by enabling automated data interpretation, early disease detection, and individualized treatment strategies. Chronic liver diseases, including non-alcoholic fatty liver disease, cirrhosis, and hepatocellular carcinoma, often progress silently and pose diagnostic challenges due to reliance on invasive biopsies and operator-dependent imaging. This review explores the integration of AI across key domains such as big data analytics, deep learning-based image analysis, histopathological interpretation, biomarker discovery, and clinical prediction modeling. AI algorithms have demonstrated high accuracy in liver fibrosis staging, hepatocellular carcinoma detection, and non-alcoholic fatty liver disease risk stratification, while also enhancing survival prediction and treatment response assessment. For instance, convolutional neural networks trained on portal venous-phase computed tomography have achieved area under the curves up to 0.92 for significant fibrosis (F2-F4) and 0.89 for advanced fibrosis, with magnetic resonance imaging-based models reporting comparable performance. Advanced methodologies such as federated learning preserve patient privacy during cross-center model training, and explainable AI techniques promote transparency and clinician trust. Despite these advancements, clinical adoption remains limited by challenges including data heterogeneity, algorithmic bias, regulatory uncertainty, and lack of real-time integration into electronic health records. Looking forward, the convergence of multi-omics, imaging, and clinical data through interpretable and validated AI frameworks holds great promise for precision liver care. Continued efforts in model standardization, ethical oversight, and clinician-centered deployment will be essential to realize the full potential of AI in hepatopathy diagnosis and treatment.

PMID:41479639 | PMC:PMC12754151 | DOI:10.3748/wjg.v31.i46.111176

A clinically validated 3D deep learning approach for quantifying vascular invasion in pancreatic cancer

npj Digital Medicine, Published online: 31 December 2025; doi:10.1038/s41746-025-02260-3

A clinically validated 3D deep learning approach for quantifying vascular invasion in pancreatic cancer

Quantifying the global eco-footprint of wearable healthcare electronics

Nature, Published online: 31 December 2025; doi:10.1038/s41586-025-09819-w

An integrated systems engineering framework based on life-cycle inventories is used to quantify the global eco-footprint of wearable healthcare electronics and identify effective mitigation strategies.

Artificial Intelligence Applications in the Diagnosis, Treatment, and Prognosis of Hepatocellular Carcinoma

Gut Liver. 2025 Dec 31. doi: 10.5009/gnl250268. Online ahead of print.

ABSTRACT

The global burden of hepatocellular carcinoma (HCC) has shifted from viral to nonviral etiologies. However, successful antiviral therapy does not fully eliminate the risk of HCC, underscoring the demand for more effective surveillance strategies. Current screening methods, such as semiannual ultrasonography and the measurement of α-fetoprotein levels, offer suboptimal sensitivity for early detection. A cost-effective, reliable surveillance approach remains an unmet need. The Barcelona Clinic Liver Cancer staging system provides a framework to guide HCC therapy; yet, some gray zone exists, particularly for patients with intermediate-stage disease. Although tyrosine kinase inhibitors and immunotherapies have transformed the therapeutic landscape, their efficacies vary among patients, highlighting the necessity for personalized treatment strategies. In response to these challenges, artificial intelligence (AI) approaches have emerged as transformative tools in healthcare. By processing complex, nonlinear relationships and uncovering hidden patterns in clinical data, AI methods offer capabilities beyond those of traditional statistical methods. Furthermore, AI-driven multi-omics analysis holds promise for identifying novel biomarkers, thereby advancing precision medicine for HCC patients. This review introduces the potential of AI applications in enhancing the diagnosis, treatment, and prognosis of HCC.

PMID:41472345 | DOI:10.5009/gnl250268

Artificial Intelligence Applications in the Diagnosis, Treatment, and Prognosis of Hepatocellular Carcinoma

Gut Liver. 2025 Dec 31. doi: 10.5009/gnl250268. Online ahead of print.

ABSTRACT

The global burden of hepatocellular carcinoma (HCC) has shifted from viral to nonviral etiologies. However, successful antiviral therapy does not fully eliminate the risk of HCC, underscoring the demand for more effective surveillance strategies. Current screening methods, such as semiannual ultrasonography and the measurement of α-fetoprotein levels, offer suboptimal sensitivity for early detection. A cost-effective, reliable surveillance approach remains an unmet need. The Barcelona Clinic Liver Cancer staging system provides a framework to guide HCC therapy; yet, some gray zone exists, particularly for patients with intermediate-stage disease. Although tyrosine kinase inhibitors and immunotherapies have transformed the therapeutic landscape, their efficacies vary among patients, highlighting the necessity for personalized treatment strategies. In response to these challenges, artificial intelligence (AI) approaches have emerged as transformative tools in healthcare. By processing complex, nonlinear relationships and uncovering hidden patterns in clinical data, AI methods offer capabilities beyond those of traditional statistical methods. Furthermore, AI-driven multi-omics analysis holds promise for identifying novel biomarkers, thereby advancing precision medicine for HCC patients. This review introduces the potential of AI applications in enhancing the diagnosis, treatment, and prognosis of HCC.

PMID:41472345 | DOI:10.5009/gnl250268

PIC-SURE: an open-source platform for integrating clinical and genomic data

npj Digital Medicine, Published online: 30 December 2025; doi:10.1038/s41746-025-02284-9

PIC-SURE: an open-source platform for integrating clinical and genomic data

Bidirectional RAG: Safe Self-Improving Retrieval-Augmented Generation Through Multi-Stage Validation

arXiv:2512.22199v1 Announce Type: new Abstract: Retrieval-Augmented Generation RAG systems enhance large language models by grounding responses in external knowledge bases, but conventional RAG architectures operate with static corpora that cannot evolve from user interactions. We introduce Bidirectional RAG, a novel RAG architecture that enables safe corpus expansion through validated write back of high quality generated responses. Our system employs a multi stage acceptance layer combining grounding verification (NLI based entailment, attribution checking, and novelty detection to prevent hallucination pollution while enabling knowledge accumulation. Across four datasets Natural Questions, TriviaQA, HotpotQA, Stack Overflow with three random seeds 12 experiments per system, Bidirectional RAG achieves 40.58% average coverage nearly doubling Standard RAG 20.33% while adding 72% fewer documents than naive write back 140 vs 500. Our work demonstrates that self improving RAG is feasible and safe when governed by rigorous validation, offering a practical path toward RAG systems that learn from deployment.

SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence

arXiv:2512.22334v1 Announce Type: new Abstract: We introduce SciEvalKit, a unified benchmarking toolkit designed to evaluate AI models for science across a broad range of scientific disciplines and task capabilities. Unlike general-purpose evaluation platforms, SciEvalKit focuses on the core competencies of scientific intelligence, including Scientific Multimodal Perception, Scientific Multimodal Reasoning, Scientific Multimodal Understanding, Scientific Symbolic Reasoning, Scientific Code Generation, Science Hypothesis Generation and Scientific Knowledge Understanding. It supports six major scientific domains, spanning from physics and chemistry to astronomy and materials science. SciEvalKit builds a foundation of expert-grade scientific benchmarks, curated from real-world, domain-specific datasets, ensuring that tasks reflect authentic scientific challenges. The toolkit features a flexible, extensible evaluation pipeline that enables batch evaluation across models and datasets, supports custom model and dataset integration, and provides transparent, reproducible, and comparable results. By bridging capability-based evaluation and disciplinary diversity, SciEvalKit offers a standardized yet customizable infrastructure to benchmark the next generation of scientific foundation models and intelligent agents. The toolkit is open-sourced and actively maintained to foster community-driven development and progress in AI4Science.

Why AI Safety Requires Uncertainty, Incomplete Preferences, and Non-Archimedean Utilities

arXiv:2512.23508v1 Announce Type: new Abstract: How can we ensure that AI systems are aligned with human values and remain safe? We can study this problem through the frameworks of the AI assistance and the AI shutdown games. The AI assistance problem concerns designing an AI agent that helps a human to maximise their utility function(s). However, only the human knows these function(s); the AI assistant must learn them. The shutdown problem instead concerns designing AI agents that: shut down when a shutdown button is pressed; neither try to prevent nor cause the pressing of the shutdown button; and otherwise accomplish their task competently. In this paper, we show that addressing these challenges requires AI agents that can reason under uncertainty and handle both incomplete and non-Archimedean preferences.

Interpretable Link Prediction in AI-Driven Cancer Research: Uncovering Co-Authorship Patterns

arXiv:2512.22181v1 Announce Type: cross Abstract: Artificial intelligence (AI) is transforming cancer diagnosis and treatment. The intricate nature of this disease necessitates the collaboration of diverse stakeholders with varied expertise to ensure the effectiveness of cancer research. Despite its importance, forming effective interdisciplinary research teams remains challenging. Understanding and predicting collaboration patterns can help researchers, organizations, and policymakers optimize resources and foster impactful research. We examined co-authorship networks as a proxy for collaboration within AI-driven cancer research. Using 7,738 publications (2000-2017) from Scopus, we constructed 36 overlapping co-authorship networks representing new, persistent, and discontinued collaborations. We engineered both attribute-based and structure-based features and built four machine learning classifiers. Model interpretability was performed using Shapley Additive Explanations (SHAP). Random forest achieved the highest recall for all three types of examined collaborations. The discipline similarity score emerged as a crucial factor, positively affecting new and persistent patterns while negatively impacting discontinued collaborations. Additionally, high productivity and seniority were positively associated with discontinued links. Our findings can guide the formation of effective research teams, enhance interdisciplinary cooperation, and inform strategic policy decisions.

Fairness Evaluation of Risk Estimation Models for Lung Cancer Screening

arXiv:2512.22242v1 Announce Type: cross Abstract: Lung cancer is the leading cause of cancer-related mortality in adults worldwide. Screening high-risk individuals with annual low-dose CT (LDCT) can support earlier detection and reduce deaths, but widespread implementation may strain the already limited radiology workforce. AI models have shown potential in estimating lung cancer risk from LDCT scans. However, high-risk populations for lung cancer are diverse, and these models' performance across demographic groups remains an open question. In this study, we drew on the considerations on confounding factors and ethically significant biases outlined in the JustEFAB framework to evaluate potential performance disparities and fairness in two deep learning risk estimation models for lung cancer screening: the Sybil lung cancer risk model and the Venkadesh21 nodule risk estimator. We also examined disparities in the PanCan2b logistic regression model recommended in the British Thoracic Society nodule management guideline. Both deep learning models were trained on data from the US-based National Lung Screening Trial (NLST), and assessed on a held-out NLST validation set. We evaluated AUROC, sensitivity, and specificity across demographic subgroups, and explored potential confounding from clinical risk factors. We observed a statistically significant AUROC difference in Sybil's performance between women (0.88, 95% CI: 0.86, 0.90) and men (0.81, 95% CI: 0.78, 0.84, p

Harnessing Large Language Models for Biomedical Named Entity Recognition

arXiv:2512.22738v1 Announce Type: cross Abstract: Background and Objective: Biomedical Named Entity Recognition (BioNER) is a foundational task in medical informatics, crucial for downstream applications like drug discovery and clinical trial matching. However, adapting general-domain Large Language Models (LLMs) to this task is often hampered by their lack of domain-specific knowledge and the performance degradation caused by low-quality training data. To address these challenges, we introduce BioSelectTune, a highly efficient, data-centric framework for fine-tuning LLMs that prioritizes data quality over quantity. Methods and Results: BioSelectTune reformulates BioNER as a structured JSON generation task and leverages our novel Hybrid Superfiltering strategy, a weak-to-strong data curation method that uses a homologous weak model to distill a compact, high-impact training dataset. Conclusions: Through extensive experiments, we demonstrate that BioSelectTune achieves state-of-the-art (SOTA) performance across multiple BioNER benchmarks. Notably, our model, trained on only 50% of the curated positive data, not only surpasses the fully-trained baseline but also outperforms powerful domain-specialized models like BioMedBERT.

Heterogeneity in Multi-Agent Reinforcement Learning

arXiv:2512.22941v1 Announce Type: cross Abstract: Heterogeneity is a fundamental property in multi-agent reinforcement learning (MARL), which is closely related not only to the functional differences of agents, but also to policy diversity and environmental interactions. However, the MARL field currently lacks a rigorous definition and deeper understanding of heterogeneity. This paper systematically discusses heterogeneity in MARL from the perspectives of definition, quantification, and utilization. First, based on an agent-level modeling of MARL, we categorize heterogeneity into five types and provide mathematical definitions. Second, we define the concept of heterogeneity distance and propose a practical quantification method. Third, we design a heterogeneity-based multi-agent dynamic parameter sharing algorithm as an example of the application of our methodology. Case studies demonstrate that our method can effectively identify and quantify various types of agent heterogeneity. Experimental results show that the proposed algorithm, compared to other parameter sharing baselines, has better interpretability and stronger adaptability. The proposed methodology will help the MARL community gain a more comprehensive and profound understanding of heterogeneity, and further promote the development of practical algorithms.

Multi-agent Self-triage System with Medical Flowcharts

arXiv:2511.12439v2 Announce Type: replace Abstract: Online health resources and large language models (LLMs) are increasingly used as a first point of contact for medical decision-making, yet their reliability in healthcare remains limited by low accuracy, lack of transparency, and susceptibility to unverified information. We introduce a proof-of-concept conversational self-triage system that guides LLMs with 100 clinically validated flowcharts from the American Medical Association, providing a structured and auditable framework for patient decision support. The system leverages a multi-agent framework consisting of a retrieval agent, a decision agent, and a chat agent to identify the most relevant flowchart, interpret patient responses, and deliver personalized, patient-friendly recommendations, respectively. Performance was evaluated at scale using synthetic datasets of simulated conversations. The system achieved 95.29% top-3 accuracy in flowchart retrieval (N=2,000) and 99.10% accuracy in flowchart navigation across varied conversational styles and conditions (N=37,200). By combining the flexibility of free-text interaction with the rigor of standardized clinical protocols, this approach demonstrates the feasibility of transparent, accurate, and generalizable AI-assisted self-triage, with potential to support informed patient decision-making while improving healthcare resource utilization.

Taming Data Challenges in ML-based Security Tasks: Lessons from Integrating Generative AI

arXiv:2507.06092v3 Announce Type: replace-cross Abstract: Machine learning-based supervised classifiers are widely used for security tasks, and their improvement has been largely focused on algorithmic advancements. We argue that data challenges that negatively impact the performance of these classifiers have received limited attention. We address the following research question: Can developments in Generative AI (GenAI) address these data challenges and improve classifier performance? We propose augmenting training datasets with synthetic data generated using GenAI techniques to improve classifier generalization. We evaluate this approach across 7 diverse security tasks using 6 state-of-the-art GenAI methods and introduce a novel GenAI scheme called Nimai that enables highly controlled data synthesis. We find that GenAI techniques can significantly improve the performance of security classifiers, achieving improvements of up to 32.6% even in severely data-constrained settings (only ~180 training samples). Furthermore, we demonstrate that GenAI can facilitate rapid adaptation to concept drift post-deployment, requiring minimal labeling in the adjustment process. Despite successes, our study finds that some GenAI schemes struggle to initialize (train and produce data) on certain security tasks. We also identify characteristics of specific tasks, such as noisy labels, overlapping class distributions, and sparse feature vectors, which hinder performance boost using GenAI. We believe that our study will drive the development of future GenAI tools designed for security tasks.

Digital Health Technologies Applied in Patients With Early Cognitive Change: Scoping Review

Background: Background: Digital health technologies have the potential to revolutionize the screening, diagnostic support, monitoring and intervention of early cognitive change. However, the full spectrum of their application and the existing evidence base in this specific patient population have not been systematically delineated. Objective: Objective: To review and synthesize digital health technologies' applications, roles, and challenges in patients with early cognitive changes. Methods: Methods: This scoping review followed the enhanced Arksey & O'Malley Framework and PRISMA-ScR guidelines. A comprehensive search of four databases (PubMed, Embase, Web of Science, and Cochrane Library) was conducted from their inception until May 31, 2024. Studies were selected and data were extracted using the Population-Concept-Context framework, focusing on digital health interventions for patients with early cognitive changes. Results: Results: A total of 163 articles were included in this review, revealing a notable increase in the use of digital health technologies for patients with early cognitive changes since 2020. Of the studies, 162 focused on Mild Cognitive Impairment(95.1%), 10 on Subjective Cognitive Decline (6.1%), and 7 examined caregiver support (4.3%). The technologies were categorized into six groups: Smartphone/Computer Application, Virtual Reality, Artificial Intelligence/Big Data, Robotics, the Internet of Things, and Telemedicine. The clinical outcomes demonstrated statistically significant improvements in cognitive performance, and patient-reported outcomes including overall well-being, quality of life and social engagement. But digital health technologies also present implementation challenges, such as Virtual Reality induced vestibular symptoms, connectivity issues in telemedicine, and unintended negative consequences. Conclusions: Conclusion: This review affirms the efficacy of digital health technologies (DHTs) in screening, diagnosing, intervening in, and monitoring early cognitive changes. However, challenges such as cost, technical complexity, and user engagement currently impede broader adoption. In conclusion, while DHTs demonstrably enhance healthcare professional efficiency in managing early cognitive impairment—by facilitating clinical decision-making, optimizing patient management, and enabling personalized care—overcoming existing implementation barriers remains critical. Furthermore, rigorous assessment of their long-term effects through future research is essential. Collectively, these findings underscore the substantial potential of DHTs to transform cognitive health management and patient care, offering valuable insights for healthcare professionals, researchers, and policymakers to optimize these solutions. Clinical Trial: Trial Registration: No Trial Registration.

Computational network models for forecasting and control of mental health trajectories in digital applications

npj Digital Medicine, Published online: 30 December 2025; doi:10.1038/s41746-025-02252-3

Computational network models for forecasting and control of mental health trajectories in digital applications

Metabolic signatures in gastroenteropancreatic neuroendocrine neoplasms: unraveling diagnostic and prognostic insights

Front Endocrinol (Lausanne). 2025 Dec 11;16:1676021. doi: 10.3389/fendo.2025.1676021. eCollection 2025.

ABSTRACT

Gastroenteropancreatic neuroendocrine neoplasms (GEP-NENs) are a heterogeneous group of tumors characterized by diverse biological behaviors and variable clinical outcomes. Recent advances have highlighted the important role of metabolic reprogramming in tumorigenesis, progression, and therapeutic resistance in GEP-NENs. In this review, we synthesize the current evidence on metabolic biomarkers and altered metabolic pathways-particularly those involving glucose, lipid, and amino acid metabolism. Key biomarkers such as GLUT-1, FASN, and enzymes involved in ferroptosis, cholesterol biosynthesis, and amino acid catabolism demonstrate strong associations with tumor aggressiveness, hypoxia, and mTOR signaling. Moreover, metabolomic profiling and functional studies suggest that metabolic markers may inform prognosis and predict response to targeted therapies such as Everolimus. Although promising, the clinical translation of these markers is still limited and requires further validation in large, subtype-specific cohorts. Our findings highlight the importance of integrating metabolic profiling into the diagnostic and therapeutic landscape of GEP-NENs. Future research should prioritize biomarker standardization, multi-omics integration, and the development of metabolism-based therapeutic strategies tailored to tumor subtype and differentiation grade.

PMID:41458541 | PMC:PMC12738315 | DOI:10.3389/fendo.2025.1676021

❌