❌

Normal view

Application of artificial intelligence in hepatology

Front Digit Health. 2026 Sep 16;8:1851723. doi: 10.3389/fdgth.2026.1851723. eCollection 2026.

ABSTRACT

Artificial intelligence (AI) is being applied across diagnostic and therapeutic workflows in hepatology. This narrative review summarizes recent advances in AI for liver disease. In medical imaging and digital pathology, computer vision enables automated quantitative analysis of ultrasound, CT, MRI, and histologic images, with the aim of improving the consistency of lesion detection, disease staging, and prognostic assessment. In biomarker research, machine learning can analyze high-dimensional liquid-biopsy and multi-omics data to develop diagnostic and prognostic models; some have outperformed conventional markers in their study cohorts. Electronic health records (EHRs) and large language models (LLMs) are also being investigated for clinical decision support and personalized management. However, most reported evidence remains retrospective, and clinical adoption is limited by data heterogeneity, poor interpretability, uncertain generalizability, and regulatory requirements. Progress will require standardized datasets, external and prospective validation, clinically relevant endpoints, and human-centered implementation before gains in model performance can be translated into better patient outcomes.

PMID:42819088 | PMC:PMC13624904 | DOI:10.3389/fdgth.2026.1851723

A nonlinear multi-omics data integration and classification model based on pathway self-attention and graph convolutional networks

Yi Chuan. 2026 Sep;48(9):931-945. doi: 10.16288/j.yczz.25-275.

ABSTRACT

The abundance of omics data has significantly advanced the development of multi-omics data integration techniques. Non-linear embedding approaches for data integration have gradually become the mainstream in multi-omics research, as these approaches can substantially improve cancer analysis by enhancing the quality of the embeddings. However, current multi-omics data integration methods are typically confined to omics measurements, neglecting domain-specific prior knowledge encompassing biological pathways. In this study, we proposed a multi-omics integrated classification model, PathTransGCN, based on pathway self-attention and graph convolutional networks (GCN). The model integrated biological pathway information into multi-omics data analysis with the aim of enhancing the accuracy of cancer classification. Multi-omics data for breast cancer (BRCA), non-small cell lung cancer (NSCLC), and low-grade glioma (LGG) were obtained from The Cancer Genome Atlas (TCGA) and UCSC Xena databases. These data included gene mutations, DNA methylation, copy number variations, and gene expression, and were used to assess the model's generalizability across different cancers. First, PathTransGCN employed a pathway self-attention module to learn latent representations of samples across different pathways, thereby obtaining multi-omics integration vectors. Concurrently, a patient similarity network (PSN) was constructed using the similarity network fusion (SNF) approach. Second, the integrated vectors and the PSN were jointly fed into a GCN for end-to-end training, enabling precise classification of cancer subtypes. Through multi-omics data analysis of the BRCA dataset, PathTransGCN outperformed several popular algorithms (such as MoGCN and DeePathNet) in the five-class classification of cancer subtypes, achieving an accuracy rate of 87.6% and an F1 score of 86.4%. Moreover, the model demonstrated robust generalization capabilities across both NSCLC and LGG datasets, while effectively identifying key disease-associated biomarkers at the pathway level. Experimental results demonstrate that PathTransGCN exhibits outstanding performance in integrating omics data and delivering interpretable classification outcomes, presenting significant potential for clinical applications.

PMID:42751828 | DOI:10.16288/j.yczz.25-275

Online Video Agent Harness for Long Video Understanding

arXiv:2609.12818v1 Announce Type: cross Abstract: Long video understanding often behaves like a visual needle-in-a-haystack problem: query-relevant evidence is sparsely distributed across long temporal spans, while packing dense frames into a single VLM context incurs \textit{context rot} and high cost. Existing video agents often rely on query-agnostic offline preprocessing or ad hoc tool sets, which can miss query-specific details and waste computation. In this work, we present VideoXAgent, a purely online video-agent harness for long video understanding that starts from the given video file and user query, plans and decomposes the task, invokes specialized expert tools on demand, and aggregates multimodal evidence to produce a final answer while resolving conflicts among observations. To support this on-demand invocation, we design a suite of heterogeneous expert tools guided by a data-driven taxonomy of atomic capabilities, spanning scripts, VLMs, and domain models (e.g., detection, OCR, ASR, face recognition). The harness further enforces objective evidence prompting and budget-aware control to curb hallucination and non-termination. Across Video-MME-Long, LongVideoBench-Long, LVBench, and MINERVA, VideoXAgent is competitive with frontier LMMs and video agents under a smaller context footprint---about 50k tokens of agent context per sample, even on hour-long videos. In particular, on complex video-reasoning benchmarks such as MINERVA, it matches this level while using only about 15\% of the context of a 1,024-frame dense-packing baseline. Notably, the harness remains effective with a visually weak or even text-only orchestrator, suggesting that strong long-video understanding can emerge from progressive agentic evidence seeking rather than from packing the full video into a single context. Project page: https://go-agent-x.github.io/video_agent_harness/

AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training

arXiv:2507.01663v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a pivotal technology in the post-training phase of large language models (LLMs). Traditional task-collocated RL frameworks suffer from significant scalability bottlenecks, while task-separated RL frameworks face challenges in managing complex dataflows and resolving resource idling. Furthermore, most existing frameworks are tightly coupled with LLM training or inference engines, making them difficult to support custom-designed engines. To address these challenges, we propose AsyncFlow, an asynchronous streaming RL framework tailored for efficient post-training. Specifically, we introduce a distributed data storage and transfer module that provides panoramic data management and fine-grained scheduling capabilities in a fully streamed manner. This architecture inherently enables automated pipeline overlapping among RL tasks and dynamic load-balancing. Moreover, we propose an asynchronous producer-consumer workflow, which is engineered to minimize computational idleness by strategically deferring the parameter update process within staleness thresholds. Finally, the core capabilities of AsyncFlow are architecturally decoupled from underlying training and inference engines and encapsulated by service-oriented user interfaces, offering a modular and customizable user experience. Extensive experiments demonstrate an average throughput of 1.59x compared to the state-of-the-art baseline. The architecture presented in this work provides actionable insights for designing next-generation RL training systems.

RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems

arXiv:2609.09657v1 Announce Type: new Abstract: Existing emotional support conversation systems mainly focus on one-on-one seeker-supporter interactions and individual emotional states, leaving interpersonal relations in multi-party scenarios underexplored. In this work, we introduce relation-aware emotional support conversation, a new task that evaluates whether LLMs can capture and utilize the evolving dynamics of relationships to offer more effective emotional support. We construct RESCUE (Relation-aware Emotional Support Conversation Understanding and Evaluation Benchmark) from real couple and family interview conversations, containing 191 samples, 7,079 annotated turns, and 1,064.8 minutes of video. Based on rich annotations of socio-emotional and support-related dynamics, RESCUE defines six tasks that evaluate two core capabilities required for relation-aware emotional support: Relational Understanding and Relation-Sensitive Support. Experiments with ten LLMs show that current models perform relatively well on tasks relying on local emotional or intervention cues, but struggle with relation-intensive tasks such as relation pattern prediction, viewpoint prediction, and support strategy prediction. These findings reveal the limitations of current LLMs in modeling interpersonal relations and making relation-sensitive support decisions.

AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration

arXiv:2605.20025v2 Announce Type: replace Abstract: Automating scientific discovery requires more than generating papers from ideas. Real research is iterative: hypotheses are challenged from multiple perspectives, experiments fail and inform the next attempt, and lessons accumulate across cycles. Existing autonomous research systems often model this process as a linear pipeline: they rely on single-agent reasoning, stop when execution fails, and do not carry experience across runs. We present AutoResearchClaw, a multi-agent autonomous research pipeline built on five mechanisms: structured multi-agent debate for hypothesis generation and result analysis, a self-healing executor with a \textsc{Pivot}/\textsc{Refine} decision loop that transforms failures into information, verifiable result reporting that prevents fabricated numbers and hallucinated citations, human-in-the-loop collaboration with seven intervention modes spanning full autonomy to step-by-step oversight, and cross-run evolution that converts past mistakes into future safeguards. On ARC-Bench, a 25-topic experiment-stage benchmark, AutoResearchClaw outperforms AI Scientist v2 by 54.7%. A human-in-the-loop ablation across seven intervention modes reveals that precise, targeted collaboration at high-leverage decision points consistently outperforms both full autonomy and exhaustive step-by-step oversight. We position AutoResearchClaw as a research amplifier that augments rather than replaces human scientific judgment. Code is available at https://github.com/aiming-lab/AutoResearchClaw.

M$^\star$: Every Task Deserves Its Own Memory Harness

arXiv:2604.11811v2 Announce Type: replace-cross Abstract: Large language model agents rely on specialized memory systems to accumulate and reuse knowledge during extended interactions. Recent architectures typically adopt a fixed memory design tailored to specific domains, such as semantic retrieval for conversations or skills reused for coding. However, a memory system optimized for one purpose frequently fails to transfer to others. To address this limitation, we introduce M$^\star$, a method that automatically discovers task-optimized memory harnesses through executable program evolution. Specifically, M$^\star$ models an agent memory system as a memory program written in Python. This program encapsulates the data Schema, the storage Logic, and the agent workflow Instructions. We optimize these components jointly using a reflective code evolution method; this approach employs a population-based search strategy and analyzes evaluation failures to iteratively refine the candidate programs. We evaluate M$^\star$ on four distinct benchmarks spanning conversation, embodied planning, and expert reasoning. Our results demonstrate that M$^\star$ improves performance over existing fixed-memory baselines robustly across all evaluated tasks. Furthermore, the evolved memory programs exhibit structurally distinct processing mechanisms for each domain. This finding indicates that specializing the memory mechanism for a given task explores a broad design space and provides a superior solution compared to general-purpose memory paradigms.

GPX8<sup>+</sup> cancer-associated fibroblast-derived lactate contributes to lenvatinib resistance by facilitating BRPF1 expression through histone H3 lysine 18 lactylation in hepatocellular carcinoma

Oncogene, Published online: 22 May 2026; doi:10.1038/s41388-026-03711-1

GPX8+ cancer-associated fibroblast-derived lactate contributes to lenvatinib resistance by facilitating BRPF1 expression through histone H3 lysine 18 lactylation in hepatocellular carcinoma

Exploring the prognostic role of senescence-related genes in gastric cancer through multi-omics integration and machine learning

Hum Genomics. 2026 May 9. doi: 10.1186/s40246-026-00979-y. Online ahead of print.

ABSTRACT

Cellular senescence plays a context-dependent role in gastric cancer (GC), functioning both through tumor-suppressive arrest and the tumor-promoting senescence-associated secretory phenotype. However, its systematic integration into prognostic models remains limited. Here, we develop a novel interpretable framework to identify and validate a robust senescence-related gene signature for GC prognosis. We first introduce a dual-model interpretable feature selection strategy that integrates a biologically informed Kolmogorov-Arnold Network with a tabular foundation model to identify cancer-associated senescence genes. From the initial candidates, an ensemble of ten machine learning algorithms distills a core 4-gene signature to construct a Senescence Risk Score (SRS). The SRS proves to be a powerful and independent prognostic indicator, effectively stratifies patients into high- and low-risk groups with distinct overall survival across multiple cohorts. High-risk patients exhibit an "immune-hot" but potentially dysfunctional tumor microenvironment, characterized by enriched immune cell infiltration, elevated checkpoint expression, and distinct metabolic reprogramming favoring pathways such as angiogenesis and epithelial-mesenchymal transition (EMT). Furthermore, the SRS correlates with differential somatic mutation profiles and suggests potential sensitivity to specific chemotherapeutic agents. In vitro functional assays confirmed the oncogenic role of SERPINE1, a top-ranked core gene, in promoting GC cell proliferation. Regulatory network analysis revealed potential upstream transcription factors and miRNAs governing the signature. Collectively, we present a validated senescence-related prognostic signature that enables effective risk stratification of patients with gastric cancer.

PMID:42106891 | DOI:10.1186/s40246-026-00979-y

A Generative Foundation Model for Multimodal Histopathology

arXiv:2604.03635v1 Announce Type: cross Abstract: Accurate diagnosis and treatment of complex diseases require integrating histological, molecular, and clinical data, yet in practice these modalities are often incomplete owing to tissue scarcity, assay cost, and workflow constraints. Existing computational approaches attempt to impute missing modalities from available data but rely on task-specific models trained on narrow, single source-target pairs, limiting their generalizability. Here we introduce MuPD (Multimodal Pathology Diffusion), a generative foundation model that embeds hematoxylin and eosin (H&E)-stained histology, molecular RNA profiles, and clinical text into a shared latent space through a diffusion transformer with decoupled cross-modal attention. Pretrained on 100 million histology image patches, 1.6 million text-histology pairs, and 10.8 million RNA-histology pairs spanning 34 human organs, MuPD supports diverse cross-modal synthesis tasks with minimal or no task-specific fine-tuning. For text-conditioned and image-to-image generation, MuPD synthesizes histologically faithful tissue architectures, reducing Fr\'echet inception distance (FID) scores by 50% relative to domain-specific models and improving few-shot classification accuracy by up to 47% through synthetic data augmentation. For RNA-conditioned histology generation, MuPD reduces FID by 23% compared with the next-best method while preserving cell-type distributions across five cancer types. As a virtual stainer, MuPD translates H&E images to immunohistochemistry and multiplex immunofluorescence, improving average marker correlation by 37% over existing approaches. These results demonstrate that a single, unified generative model pretrained across heterogeneous pathology modalities can substantially outperform specialized alternatives, providing a scalable computational framework for multimodal histopathology.

Gray Anchoring: a New Computational Theory for Biological Color Constancy

arXiv:2410.08823v3 Announce Type: replace Abstract: It is still challenging for computer vision to imitate human color perception, e.g., color constancy, which is a fundamental perceptual ability in humans to perceive, interpret and interact with their surroundings. Among others, the anchoring theory provides impressive insights for human lightness perception, yet the specific anchoring rules underlying color constancy have remained contentious for decades. In this work, we introduced a novel computational theory - gray-anchoring (GA) theory - to explain how the early stage of visual system contributes to color constancy and demonstrate how our GA rule applies to the chromatic domain by identifying gray surfaces within complex scenes. Furthermore, we also demonstrate the potential neural implementation of gray-anchoring by quantitatively analyzing the computational flows of concentric double-opponent (DO) cells in V1. The simulational results show that the concentric DO cells have the ability to identify gray surfaces within color-biased scenes and these gray surfaces can then be used by the higher-level cortices to easily estimate the illuminant. This finding offers not only a clear functional explanation of the concentric DO receptive fields of this cell type in the visual system but also an effective and efficient solution to computational color constancy for computer vision.

LiteInception: A Lightweight and Interpretable Deep Learning Framework for General Aviation Fault Diagnosis

arXiv:2604.01725v1 Announce Type: new Abstract: General aviation fault diagnosis and efficient maintenance are critical to flight safety; however, deploying deep learning models on resource-constrained edge devices poses dual challenges in computational capacity and interpretability. This paper proposes LiteInception--a lightweight interpretable fault diagnosis framework designed for edge deployment. The framework adopts a two-stage cascaded architecture aligned with standard maintenance workflows: Stage 1 performs high-recall fault detection, and Stage 2 conducts fine-grained fault classification on anomalous samples, thereby decoupling optimization objectives and enabling on-demand allocation of computational resources. For model compression, a multi-method fusion strategy based on mutual information, gradient analysis, and SE attention weights is proposed to reduce the input sensor channels from 23 to 15, and a 1+1 branch LiteInception architecture is introduced that compresses InceptionTime parameters by 70%, accelerates CPU inference by over 8x, with less than 3% F1 loss. Furthermore, knowledge distillation is introduced as a precision-recall regulation mechanism, enabling the same lightweight model to adapt to different scenarios--such as safety-critical and auxiliary diagnosis--by switching training strategies. Finally, a dual-layer interpretability framework integrating four attribution methods is constructed, providing traceable evidence chains of "which sensor x which time period." Experiments on the NGAFID dataset demonstrate a fault detection accuracy of 81.92% with 83.24% recall, and a fault identification accuracy of 77.00%, validating the framework's favorable balance among efficiency, accuracy, and interpretability.

UniMixer: A Unified Architecture for Scaling Laws in Recommendation Systems

arXiv:2604.00590v2 Announce Type: replace-cross Abstract: In recent years, the scaling laws of recommendation models have attracted increasing attention, which govern the relationship between performance and parameters/FLOPs of recommenders. Currently, there are three mainstream architectures for achieving scaling in recommendation models, namely attention-based, TokenMixer-based, and factorization-machine-based methods, which exhibit fundamental differences in both design philosophy and architectural structure. In this paper, we propose a unified scaling architecture for recommendation systems, namely \textbf{UniMixer}, to improve scaling efficiency and establish a unified theoretical framework that unifies the mainstream scaling blocks. By transforming the rule-based TokenMixer to an equivalent parameterized structure, we construct a generalized parameterized feature mixing module that allows the token mixing patterns to be optimized and learned during model training. Meanwhile, the generalized parameterized token mixing removes the constraint in TokenMixer that requires the number of heads to be equal to the number of tokens. Furthermore, we establish a unified scaling module design framework for recommender systems, which bridges the connections among attention-based, TokenMixer-based, and factorization-machine-based methods. To further boost scaling ROI, a lightweight UniMixing module is designed, \textbf{UniMixing-Lite}, which further compresses the model parameters and computational cost while significantly improve the model performance. The scaling curves are shown in the following figure. Extensive offline and online experiments are conducted to verify the superior scaling abilities of \textbf{UniMixer}.

Single-cell multiomics uncovers an endothelial mechanosensitive PIEZO1-IL-33 axis driving pulmonary fibrosis

Nat Commun. 2026 Mar 20;17(1):2655. doi: 10.1038/s41467-026-70193-w.

ABSTRACT

Pulmonary fibrosis represents a progressive interstitial lung disease marked by excessive extracellular matrix deposition and architectural distortion. Vascular endothelial cells critically contribute to fibrogenesis through paracrine secretion of pro-fibrotic mediators, yet their mechanobiological regulation remains elusive. Using integrated single-cell multi-omics profiling of human pulmonary fibrosis specimens and experimental fibrosis models induced by bleomycin or silica, we identify mechanosensitive Piezo1 upregulation in Endothelial cells as a hallmark of fibrotic progression. Endothelial-specific Piezo1 knockout significantly attenuates Bleomycin-induced fibrotic remodeling in male mice, establishing its pathogenic necessity. Mechanistically, PIEZO1 activation promotes pulmonary fibrosis development via CAPN2-mediated STAT3 phosphorylation, which may regulate the secretion of the pro-fibrotic molecule interleukin-33. These findings suggest that the endothelial PIEZO1-CAPN2-STAT3-IL33 axis is a potential therapeutic target for PF intervention.

PMID:41862476 | PMC:PMC13004862 | DOI:10.1038/s41467-026-70193-w

Contextual Counterfactual Credit Assignment for Multi-Agent Reinforcement Learning in LLM Collaboration

arXiv:2603.06859v1 Announce Type: cross Abstract: Cooperative multi-agent reinforcement learning (MARL) systems powered by large language models (LLMs) are frequently optimized via sparse terminal-only feedback. This shared signal entangles upstream decisions, obstructing accurate decision-level credit assignment. To address this trajectory-level diffusion, we introduce Contextual Counterfactual Credit Assignment (\textbf{\texttt{C3}}). Instead of distributing rewards across an entire episode, \textbf{\texttt{C3}} isolates the causal impact of individual messages by freezing the exact transcript-derived context, evaluating context-matched alternatives via fixed-continuation replay, and applying a leave-one-out (LOO) baseline. This localized intervention extracts unbiased, low-variance marginal advantages for standard policy-gradient optimization. Evaluated across five mathematical and coding benchmarks under matched budgets, \textbf{\texttt{C3}} improves terminal performance over established baselines. Mechanistic diagnostics further show that these gains are accompanied by higher credit fidelity, lower contextual variance, and stronger inter-agent causal dependence. Our code is available at https://github.com/EIT-EAST-Lab/C3.

Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization

arXiv:2506.17252v4 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) has emerged as an effective approach for aligning large language models (LLMs) with human preferences. However, its performance is highly dependent on the quality of the underlying human preference data. To address this bottleneck, prior work has explored various data selection strategies, but these methods often overlook the impact of the evolving states of the language model during the optimization process. In this paper, we introduce a novel problem: Sample Scheduling for DPO, which aims to dynamically and adaptively schedule training samples based on the model's evolving batch-wise states throughout preference optimization. To solve this problem, we propose SamS, an efficient and effective algorithm that adaptively selects samples in each training batch based on the LLM's learning feedback to maximize the potential generalization performance. Notably, without modifying the core DPO algorithm, simply integrating SamS significantly improves performance across tasks, with minimal additional computational overhead. This work points to a promising new direction for improving LLM alignment through batch-wise sample selection, with potential generalization to RLHF and broader supervised learning paradigms.
❌