Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
EvoRS: On-Policy Self-Evolution of Reward Systems for Open-Ended Reinforcement Learning
arXiv:2609.12459v1 Announce Type: new Abstract: Open-ended reinforcement learning often relies on rubric-based rewards for tasks without directly verifiable answers. Yet the policy and reward system form a dynamic feedback loop: as the policy optimizes the current reward, an initially useful reward system may become unreliable due to reward hacking or reduced response discriminability. The reward system should therefore evolve rather than remain fixed during training. Existing dynamic-rubric me
-
Omics In Lung
-
From the invasive front to organotropic pre-metastatic niches: spatial immune regulatory networks governing cholangiocarcinoma dissemination and metastasis-intercepting immunotherapy
Front Immunol. 2026 Aug 20;17:1919864. doi: 10.3389/fimmu.2026.1919864. eCollection 2026.ABSTRACTCholangiocarcinoma is an aggressive biliary tract malignancy in which metastatic relapse and primary or acquired resistance to immunotherapy remain major causes of mortality. Although immune checkpoint inhibitors have improved first-line treatment for advanced biliary tract cancer, most patients do not achieve durable benefit, indicating that immune failure is not explained by a single checkpoint pat
From the invasive front to organotropic pre-metastatic niches: spatial immune regulatory networks governing cholangiocarcinoma dissemination and metastasis-intercepting immunotherapy
Front Immunol. 2026 Aug 20;17:1919864. doi: 10.3389/fimmu.2026.1919864. eCollection 2026.
ABSTRACT
Cholangiocarcinoma is an aggressive biliary tract malignancy in which metastatic relapse and primary or acquired resistance to immunotherapy remain major causes of mortality. Although immune checkpoint inhibitors have improved first-line treatment for advanced biliary tract cancer, most patients do not achieve durable benefit, indicating that immune failure is not explained by a single checkpoint pathway. In this Review, we propose a spatial immune-regulatory continuum for cholangiocarcinoma dissemination. Most direct single-cell and spatial evidence currently derives from intrahepatic cholangiocarcinoma, and its applicability to perihilar and distal disease remains to be established. This continuum begins in the tumor core and invasive front, where malignant cells, cancer-associated fibroblasts, tumor-associated macrophages, endothelial and lymphatic cells, regulatory T cells, immature neutrophils and excluded or dysfunctional cytotoxic T cells form a pro-invasive ecosystem. It then extends through extracellular vesicles, soluble mediators and lymphovascular routes that may educate organotropic pre-metastatic niches. Finally, lymph node, lung, liver, peritoneal and bone microenvironments provide organ-specific extracellular matrix, myeloid and stromal programs that enable immune evasion and metastatic colonization. By integrating clinical evidence, multi-omics studies, single-cell and spatial transcriptomics, extracellular vesicle biology, pre-metastatic niche concepts and emerging therapeutic strategies, we argue that cholangiocarcinoma metastasis should be targeted before overt dissemination whenever possible. In this Review, "metastasis-intercepting immunotherapy" is used as an author-defined conceptual framework for strategies intended to prevent or disrupt the immune-stromal conditions that enable dissemination and colonization, rather than merely shrink established metastatic lesions. Metastasis-intercepting immunotherapy will likely require rational combinations that reprogram the invasive front, restore dendritic-cell-mediated antigen presentation, block tumor-stroma-myeloid circuits, disrupt EV-mediated communication that may contribute to niche formation and select patients using spatial biomarkers rather than bulk immune markers alone.
PMID:42694469 | PMC:PMC13539491 | DOI:10.3389/fimmu.2026.1919864
-
cs.AI, q-bio.NC updates on arXiv.org
-
FrontierOR: Benchmarking LLMs' Capacity for Efficient Algorithm Design in Large-Scale Optimization
arXiv:2605.25246v2 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for optimization modeling and solver-code generation, yet practical operations research and optimization problems often require a harder capability: designing scalable algorithms that exploit problem structure and outperform direct formulation-and-solve baselines. Existing benchmarks are limited to small or simplified examples far below real-world scale and complexity. We introduce FrontierOR, amo
FrontierOR: Benchmarking LLMs' Capacity for Efficient Algorithm Design in Large-Scale Optimization
-
cs.AI, q-bio.NC updates on arXiv.org
-
IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning
arXiv:2605.23997v1 Announce Type: cross Abstract: Multimodal large language models via reinforcement learning (RL) have demonstrated remarkable capabilities in complex visual reasoning tasks, yet they remain limited in long-horizon multimodal scenarios, often suffering from visual hallucination and logical error. Current methods typically pre-encode high-dimensional visual scenes into discrete textual proxies to facilitate downstream reasoning. As the reasoning chain unfolds, however, the inher
IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning
-
cs.AI, q-bio.NC updates on arXiv.org
-
Agent Learning via Early Experience
arXiv:2510.08558v3 Announce Type: replace Abstract: A long-term goal of language agents is to learn and improve through their own experience, ultimately outperforming humans in complex, real-world tasks. However, training agents from experience data with reinforcement learning remains difficult in many environments, which either lack verifiable rewards (e.g., websites) or require inefficient long-horizon rollouts (e.g., multi-turn tool use). As a result, most current agents rely on supervised f
Agent Learning via Early Experience
-
cs.AI, q-bio.NC updates on arXiv.org
-
Beyond Retrieval: Modeling Confidence Decay and Deterministic Agentic Platforms in Generative Engine Optimization
arXiv:2604.03656v1 Announce Type: new Abstract: Generative Engine Optimization (GEO) is rapidly reshaping digital marketing paradigms in the era of Large Language Models (LLMs). However, current GEO strategies predominantly rely on Retrieval-Augmented Generation (RAG), which inherently suffers from probabilistic hallucinations and the "zero-click" paradox, failing to establish sustainable commercial trust. In this paper, we systematically deconstruct the probabilistic flaws of existing RAG-base
Beyond Retrieval: Modeling Confidence Decay and Deterministic Agentic Platforms in Generative Engine Optimization
-
cs.AI, q-bio.NC updates on arXiv.org
-
BAAI Cardiac Agent: An intelligent multimodal agent for automated reasoning and diagnosis of cardiovascular diseases from cardiac magnetic resonance imaging
arXiv:2604.04078v1 Announce Type: cross Abstract: Cardiac magnetic resonance (CMR) is a cornerstone for diagnosing cardiovascular disease. However, it remains underutilized due to complex, time-consuming interpretation across multi-sequences, phases, quantitative measures that heavily reliant on specialized expertise. Here, we present BAAI Cardiac Agent, a multimodal intelligent system designed for end-to-end CMR interpretation. The agent integrates specialized cardiac expert models to perform
BAAI Cardiac Agent: An intelligent multimodal agent for automated reasoning and diagnosis of cardiovascular diseases from cardiac magnetic resonance imaging
-
cs.AI, q-bio.NC updates on arXiv.org
-
Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics
arXiv:2510.09901v2 Announce Type: replace Abstract: Computing has long served as a cornerstone of scientific discovery. Recently, a paradigm shift has emerged with the rise of large language models (LLMs), introducing autonomous systems, referred to as agents, that accelerate discovery across varying levels of autonomy. These language agents provide a flexible and versatile framework that orchestrates interactions with human scientists, natural language, computer language and code, and physics.
Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics
-
cs.AI, q-bio.NC updates on arXiv.org
-
SelfGrader: Stable Jailbreak Detection for Large Language Models using Token-Level Logits
arXiv:2604.01473v1 Announce Type: cross Abstract: Large Language Models (LLMs) are powerful tools for answering user queries, yet they remain highly vulnerable to jailbreak attacks. Existing guardrail methods typically rely on internal features or textual responses to detect malicious queries, which either introduce substantial latency or suffer from the randomness in text generation. To overcome these limitations, we propose SelfGrader, a lightweight guardrail method that formulates jailbreak
SelfGrader: Stable Jailbreak Detection for Large Language Models using Token-Level Logits
-
cs.AI, q-bio.NC updates on arXiv.org
-
$V_0$: A Generalist Value Model for Any Policy at State Zero
arXiv:2602.03584v2 Announce Type: replace-cross Abstract: Policy gradient methods rely on a baseline to measure the relative advantage of an action, ensuring the model reinforces behaviors that outperform its current average capability. In the training of Large Language Models (LLMs) using Actor-Critic methods (e.g., PPO), this baseline is typically estimated by a Value Model (Critic) often as large as the policy model itself. However, as the policy continuously evolves, the value model require
$V_0$: A Generalist Value Model for Any Policy at State Zero
-
(Multiomics OR Omics) AND (Lung OR gastric OR Hepatocellular)
-
Research on the compatibility mechanism of the Tingli Dazao Xiefei Decoction by multi-organ metabolomics strategy
J Ethnopharmacol. 2026 Mar 21:121548. doi: 10.1016/j.jep.2026.121548. Online ahead of print.ABSTRACTETHNOPHARMACOLOGICAL RELEVANCE: The Tingli Dazao Xiefei Decoction (TD) is a traditional phlegm-eliminating prescription composed of Descurainia sophia (L.) Webb. ex Prantl (TLZ) and Ziziphus jujuba Mill. (DZ), which can relieve lung, heart and kidney injury in asthma. TLZ acts as the monarch drug in the TD. Based on the research mode of "material basis of traditional Chinese medicinal properties c
Research on the compatibility mechanism of the Tingli Dazao Xiefei Decoction by multi-organ metabolomics strategy
J Ethnopharmacol. 2026 Mar 21:121548. doi: 10.1016/j.jep.2026.121548. Online ahead of print.
ABSTRACT
ETHNOPHARMACOLOGICAL RELEVANCE: The Tingli Dazao Xiefei Decoction (TD) is a traditional phlegm-eliminating prescription composed of Descurainia sophia (L.) Webb. ex Prantl (TLZ) and Ziziphus jujuba Mill. (DZ), which can relieve lung, heart and kidney injury in asthma. TLZ acts as the monarch drug in the TD. Based on the research mode of "material basis of traditional Chinese medicinal properties can be divided and combined", we have confirmed that the flavonoid glycosides components /the oligosaccharide components/the fatty oil component (FG/Oli/FO) are effective components of TLZ. However, the compatibility mechanism of the TD, and the contribution of the effective components of TLZ to the efficacy were still unclear.
AIM OF THE STUDY: To clarify the compatibility mechanism of TD, and the contribution of the effective components of TLZ to the efficacy from a comprehensive perspective of lung, heart, and kidney.
METHODS: First, we chose the asthma model corresponding to the efficacy of TD in purging the lungs and relieving asthma, and the rats were divided into the normal (NC) group, model (M) group, dexamethasone (DEX) group, and treatment groups of TD/TLZ/DZ/FO+DZ/Oli+DZ/FG+DZ. Second, metabolomics and network pharmacology were applied to elucidate the comprehensive protective effect of TD/FG+DZ/Oli+DZ/FO+DZ. Third, the multi-omics results were validated using Western blotting, RT-qPCR, flow cytometry, and immunofluorescence.
RESULTS: FO+DZ/Oli+DZ/FG+DZ had different degrees of protective effects against lung/heart/kidney injury in asthma. In metabolomics research, the principal component analysis (PCA) and cluster analysis results showed that the TLZ group was closer to TD group than DZ group, the FO+DZ and Oli+DZ group clustered with TD/NC groups in the lung and kidney, and the FO+DZ and FG+DZ group clustered with TD/NC groups in the heart. Pathway enrichment analysis suggested that the comprehensive protective effect of TLZ and its effective components combined with DZ on lung/heart/kidney may be achieved by regulating the arginine and proline metabolism, alanine, aspartate and glutamate metabolism, and unsaturated fatty acid biosynthesis. Multi-organ metabolomics and network pharmacology revealed consistent biological functions in KEGG pathways. Validation experiment showed that TLZ and its effective components combined with DZ could reverse the abnormal expression of proteins and RNA related to inflammation, airway remodeling, excitotoxicity, and energy-supply, apoptosis at different levels. Furthermore, FO+DZ may reduce asthma damage by inhibiting the FABP4/PPAR-γ/NF-κB signaling pathway.
CONCLUSION: TLZ played the key role in TD, and FO had the best therapeutic effect on each organ; the efficacy of Oli was mainly reflected in reducing lung and kidney damage, and FG was mainly involved in enhancing energy metabolism in the heart. These findings proved that traditional Chinese medicine could exert comprehensive efficacy in a 'multi-components trigger multi-channel' way.
PMID:41871629 | DOI:10.1016/j.jep.2026.121548
-
Omics In Lung
-
Research on the compatibility mechanism of the Tingli Dazao Xiefei Decoction by multi-organ metabolomics strategy
J Ethnopharmacol. 2026 Mar 21:121548. doi: 10.1016/j.jep.2026.121548. Online ahead of print.ABSTRACTETHNOPHARMACOLOGICAL RELEVANCE: The Tingli Dazao Xiefei Decoction (TD) is a traditional phlegm-eliminating prescription composed of Descurainia sophia (L.) Webb. ex Prantl (TLZ) and Ziziphus jujuba Mill. (DZ), which can relieve lung, heart and kidney injury in asthma. TLZ acts as the monarch drug in the TD. Based on the research mode of "material basis of traditional Chinese medicinal properties c
Research on the compatibility mechanism of the Tingli Dazao Xiefei Decoction by multi-organ metabolomics strategy
J Ethnopharmacol. 2026 Mar 21:121548. doi: 10.1016/j.jep.2026.121548. Online ahead of print.
ABSTRACT
ETHNOPHARMACOLOGICAL RELEVANCE: The Tingli Dazao Xiefei Decoction (TD) is a traditional phlegm-eliminating prescription composed of Descurainia sophia (L.) Webb. ex Prantl (TLZ) and Ziziphus jujuba Mill. (DZ), which can relieve lung, heart and kidney injury in asthma. TLZ acts as the monarch drug in the TD. Based on the research mode of "material basis of traditional Chinese medicinal properties can be divided and combined", we have confirmed that the flavonoid glycosides components /the oligosaccharide components/the fatty oil component (FG/Oli/FO) are effective components of TLZ. However, the compatibility mechanism of the TD, and the contribution of the effective components of TLZ to the efficacy were still unclear.
AIM OF THE STUDY: To clarify the compatibility mechanism of TD, and the contribution of the effective components of TLZ to the efficacy from a comprehensive perspective of lung, heart, and kidney.
METHODS: First, we chose the asthma model corresponding to the efficacy of TD in purging the lungs and relieving asthma, and the rats were divided into the normal (NC) group, model (M) group, dexamethasone (DEX) group, and treatment groups of TD/TLZ/DZ/FO+DZ/Oli+DZ/FG+DZ. Second, metabolomics and network pharmacology were applied to elucidate the comprehensive protective effect of TD/FG+DZ/Oli+DZ/FO+DZ. Third, the multi-omics results were validated using Western blotting, RT-qPCR, flow cytometry, and immunofluorescence.
RESULTS: FO+DZ/Oli+DZ/FG+DZ had different degrees of protective effects against lung/heart/kidney injury in asthma. In metabolomics research, the principal component analysis (PCA) and cluster analysis results showed that the TLZ group was closer to TD group than DZ group, the FO+DZ and Oli+DZ group clustered with TD/NC groups in the lung and kidney, and the FO+DZ and FG+DZ group clustered with TD/NC groups in the heart. Pathway enrichment analysis suggested that the comprehensive protective effect of TLZ and its effective components combined with DZ on lung/heart/kidney may be achieved by regulating the arginine and proline metabolism, alanine, aspartate and glutamate metabolism, and unsaturated fatty acid biosynthesis. Multi-organ metabolomics and network pharmacology revealed consistent biological functions in KEGG pathways. Validation experiment showed that TLZ and its effective components combined with DZ could reverse the abnormal expression of proteins and RNA related to inflammation, airway remodeling, excitotoxicity, and energy-supply, apoptosis at different levels. Furthermore, FO+DZ may reduce asthma damage by inhibiting the FABP4/PPAR-γ/NF-κB signaling pathway.
CONCLUSION: TLZ played the key role in TD, and FO had the best therapeutic effect on each organ; the efficacy of Oli was mainly reflected in reducing lung and kidney damage, and FG was mainly involved in enhancing energy metabolism in the heart. These findings proved that traditional Chinese medicine could exert comprehensive efficacy in a 'multi-components trigger multi-channel' way.
PMID:41871629 | DOI:10.1016/j.jep.2026.121548
-
cs.AI, q-bio.NC updates on arXiv.org
-
Reinforcement Learning with Promising Tokens for Large Language Models
arXiv:2602.03195v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has emerged as a key paradigm for aligning and optimizing large language models (LLMs). Standard approaches treat the LLM as the policy and apply RL directly over the full vocabulary space. However, this formulation includes the massive tail of contextually irrelevant tokens in the action space, which could distract the policy from focusing on decision-making among the truly reasonable tokens. In this work, we