❌

Reading view

Application of artificial intelligence in hepatology

Front Digit Health. 2026 Sep 16;8:1851723. doi: 10.3389/fdgth.2026.1851723. eCollection 2026.

ABSTRACT

Artificial intelligence (AI) is being applied across diagnostic and therapeutic workflows in hepatology. This narrative review summarizes recent advances in AI for liver disease. In medical imaging and digital pathology, computer vision enables automated quantitative analysis of ultrasound, CT, MRI, and histologic images, with the aim of improving the consistency of lesion detection, disease staging, and prognostic assessment. In biomarker research, machine learning can analyze high-dimensional liquid-biopsy and multi-omics data to develop diagnostic and prognostic models; some have outperformed conventional markers in their study cohorts. Electronic health records (EHRs) and large language models (LLMs) are also being investigated for clinical decision support and personalized management. However, most reported evidence remains retrospective, and clinical adoption is limited by data heterogeneity, poor interpretability, uncertain generalizability, and regulatory requirements. Progress will require standardized datasets, external and prospective validation, clinically relevant endpoints, and human-centered implementation before gains in model performance can be translated into better patient outcomes.

PMID:42819088 | PMC:PMC13624904 | DOI:10.3389/fdgth.2026.1851723

  •  

SPTLC2-driven sphingolipid reprogramming of neutrophils impairs anti-tumour immunity and drives liver cancer progression

Oncogene, Published online: 28 September 2026; doi:10.1038/s41388-026-03996-2

SPTLC2-driven sphingolipid reprogramming of neutrophils impairs anti-tumour immunity and drives liver cancer progression
  •  

DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations

arXiv:2605.24539v1 Announce Type: new Abstract: Agent harness evolution improves frozen language-model agents by modifying the executable structures around them. We study this paradigm as a form of sample-efficient fast adaptation: instead of updating model weights, an agent can acquire task-specific competence by changing its external harness, while leaving the base model's general capabilities intact. Prior work shows that self-generated rollouts can support harness search, suggesting that agents may acquire new task competence through practice. Yet in long-horizon stochastic environments, self-practice becomes fragile: rewards are sparse, outcomes are high-variance, and failures are hard to attribute to concrete harness mechanisms. We introduce DemoEvolve, a demonstration-bootstrapped approach to harness evolution. When reward-only search is too broad and noisy, competent human trajectories serve as expert reference experience for the coding proposer, guiding harness-level diagnosis and editing. Experiments on Liar's Dice show that self-rollout evolution can work when episodes are short and failures are attributable. In contrast, Balatro exposes a harder long-horizon stochastic regime, where self-rollout evolution is misled by sparse feedback and candidate-selection noise, while tutorial-like textual knowledge alone does not yield stable improvement. Under the same limited budget, DemoEvolve produces more effective and auditable harness edits and achieves better performance. Overall, demonstrations make sparse-feedback harness evolution more diagnosable, localizable, and stable.
  •  

FileGram: Grounding Agent Personalization in File-System Behavioral Traces

arXiv:2604.04901v1 Announce Type: cross Abstract: Coworking AI agents operating within local file systems are rapidly emerging as a paradigm in human-AI interaction; however, effective personalization remains limited by severe data constraints, as strict privacy barriers and the difficulty of jointly collecting multimodal real-world traces prevent scalable training and evaluation, and existing methods remain interaction-centric while overlooking dense behavioral traces in file-system operations; to address this gap, we propose FileGram, a comprehensive framework that grounds agent memory and personalization in file-system behavioral traces, comprising three core components: (1) FileGramEngine, a scalable persona-driven data engine that simulates realistic workflows and generates fine-grained multimodal action sequences at scale; (2) FileGramBench, a diagnostic benchmark grounded in file-system behavioral traces for evaluating memory systems on profile reconstruction, trace disentanglement, persona drift detection, and multimodal grounding; and (3) FileGramOS, a bottom-up memory architecture that builds user profiles directly from atomic actions and content deltas rather than dialogue summaries, encoding these traces into procedural, semantic, and episodic channels with query-time abstraction; extensive experiments show that FileGramBench remains challenging for state-of-the-art memory systems and that FileGramEngine and FileGramOS are effective, and by open-sourcing the framework, we hope to support future research on personalized memory-centric file-system agents.
  •  

Artificial Intelligence for Predicting Immunotherapy Efficacy in Non-Small Cell Lung Cancer

J Inflamm Res. 2026 Mar 17;19:581764. doi: 10.2147/JIR.S581764. eCollection 2026.

ABSTRACT

Immune checkpoint inhibitors (ICIs) have significantly improved the clinical outcomes for patients with non-small cell lung cancer (NSCLC). However, patient heterogeneity and the limitations of current biomarkers contribute to variations in therapeutic responses. Identifying potential beneficiaries of immunotherapy and predicting efficacy remain critical challenges. In recent years, artificial intelligence (AI) has become increasingly applied in cancer treatment, particularly for modeling clinical data and predicting patient prognosis. By integrating multi-omics data such as radiomics, pathomics, genomics, transcriptomics, proteomics, and microbiomics, AI enables comprehensive biomarker discovery and facilitates prediction of immunotherapy responses and potential toxicities in NSCLC patients. Despite these advancements, challenges such as data standardization, limited interpretability, and technical barriers persist. This review summarizes the application of AI in predicting immunotherapy efficacy for NSCLC patients and discusses the challenges and future directions in the context of precision medicine.

PMID:41867453 | PMC:PMC13005593 | DOI:10.2147/JIR.S581764

  •  

Insulin resistance prediction from wearables and routine blood biomarkers

Nature, Published online: 16 March 2026; doi:10.1038/s41586-026-10179-2

A machine-learning model that integrates data from wearable devices (such as smartwatches) with blood biomarkers and demographic data can predict whether someone has insulin resistance, enabling timely lifestyle interventions to prevent progression to type 2 diabetes.
  •  

HEARTS: Benchmarking LLM Reasoning on Health Time Series

arXiv:2603.06638v1 Announce Type: cross Abstract: The rise of large language models (LLMs) has shifted time series analysis from narrow analytics to general-purpose reasoning. Yet, existing benchmarks cover only a small set of health time series modalities and tasks, failing to reflect the diverse domains and extensive temporal dependencies inherent in real-world physiological modeling. To bridge these gaps, we introduce HEARTS (Health Reasoning over Time Series), a unified benchmark for evaluating hierarchical reasoning capabilities of LLMs over general health time series. HEARTS integrates 16 real-world datasets across 12 health domains and 20 signal modalities, and defines a comprehensive taxonomy of 110 tasks grouped into four core capabilities: Perception, Inference, Generation, and Deduction. Evaluating 14 state-of-the-art LLMs on more than 20K test samples reveals intriguing findings. First, LLMs substantially underperform specialized models, and their performance is only weakly related to general reasoning scores. Moreover, LLMs often rely on simple heuristics and struggle with multi-step temporal reasoning. Finally, performance declines with increasing temporal complexity, with similar failure modes within model families, indicating that scaling alone is insufficient. By making these gaps measurable, HEARTS provides a standardized testbed and living benchmark for developing next-generation LLM agents capable of reasoning over diverse health signals.
  •  
❌