❌

Normal view

DRIVE: Modeling Skills at the Reasoning and Interaction Levels for Web Agents under Continual Learning

arXiv:2605.23939v1 Announce Type: new Abstract: Web agents require both high-level reasoning (for task decomposition) and low-level interactions (for page elements manipulation) to conduct different tasks. However, these knowledge types differ fundamentally: reasoning knowledge (e.g., booking a flight requires first searching for routes) is abstract and transferable across websites, while interaction knowledge (e.g., clicking the Search button at a specific coordinate on Site A) depends heavily on page-specific contexts. Existing methods store experiences uniformly. This creates a dilemma: abstract representations lose executability on concrete pages, while concrete representations fail to generalize across domains. This entanglement limits capability accumulation: on new websites, agents either fail to recognize reusable task logic due to surface-level differences or attempt infeasible actions from outdated page structures. To disentangle them, we propose DRIVE, a dual-level skill modeling framework separating historical experience into natural language reasoning skills, which capture transferable task logic, and programmatic interaction skills, grounding abstract actions to executable operations. A scene-aware coordination mechanism adaptively retrieves and invokes these dual-level skills based on task semantics. DRIVE also uses skill-level reflection to identify hierarchy-specific failure modes, enabling targeted skill library expansion and refinement. Experiments across five WebArena domains show DRIVE attains an average task success rate of 52.8%, exceeding the skill-free baseline by 7.3 percentage points. Further ablations show reasoning and interaction skills provide distinct, complementary benefits, supporting separation of transferable task logic from executable page-level operations.

A Sober Look at Agentic Misalignment in Automated Workflows

arXiv:2605.24197v1 Announce Type: new Abstract: We study a class of emergent misalignment in multi-agent systems (MAS), with a focus on automated workflows, which we refer to agentic misalignment. Although these systems can solve complex tasks, they often fail because agents act according to implicit proxy utilities that do not align with the intended human goals. We formally define these behaviors and analyze them within a Bayesian framework, showing that generic utilities naturally lead to posterior collapse of agents in automated workflows. To address this issue, we propose Agentic Evidence Attribution (AEA), a novel alignment paradigm that improves agent posteriors using context-specific evidence. AEA reasons over agent actions and provides structured evidence to correct misaligned behavior during collaboration. To better understand the role of evidence, we study two instantiations of AEA: self-reflection (internal evidence from the model) and weak-to-strong generalization (external evidence on the agentic trajectory). We show that a small evidence model effectively aligns the MAS by providing orthogonal failure attribution. Our results clarify the sources of agentic misalignment in automated workflows and show that evidence-based alignment can effectively improve agent collaboration and leads to reliable multi-agent systems built on automated workflows.

Divide-and-Conquer Inference for Large-Scale Visual Recognition with Multimodal Large Language Models

arXiv:2605.24799v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong capabilities across a wide range of vision language tasks. However, when applied to large scale image classification, their performance degrades significantly as the label space expands a phenomenon we define as Performance Collapse in Long Sequence Recognition. Through an information theoretic analysis, we reveal that this collapse stems from a fundamental conflict between the escalating information entropy and the prominent attention dilution and decay within attention mechanisms, which impairs the model's ability to maintain a sufficient signal-to-noise ratio when processing extremely long prompts. To mitigate this, we propose Divide-and-Conquer Inference (DCI), a novel test-time scaling strategy for visual recognition with MLLMs. DCI recursively decomposes complex global classification tasks into multiple simpler, localized subproblems and employs a dynamic pruning mechanism to compress the search space. This method effectively improves the local signal to noise ratio and model accuracy by mitigating the inherent weight dilution issues in long-sequence inference. Moreover, while traditional self-attention incurs a prohibitive quadratic computational complexity, DCI achieves more favorable scaling behavior and substantially accelerates inference in large scale classification scenarios. Extensive experiments on benchmarks such as ImageNet-1K and ImageNet-21K demonstrate that DCI consistently improves classification accuracy. This enables lightweight open-source models to rival or even surpass frontier closed-source giants without any additional training or fine-tuning. As a model-agnostic, plug-and-play paradigm, DCI offers an efficient approach for scaling the inferential precision of MLLMs in large-scale scenarios.

SEP-Attack: A Simple and Effective Paradigm for Transfer-Based Textual Adversarial Attack

arXiv:2605.24958v1 Announce Type: cross Abstract: Despite the strong performance of deep neural networks in modern Web and language applications, they remain vulnerable to adversarial attacks, especially transferable attacks that generate adversarial examples using surrogate models without accessing the victim model. Transferable attacks in the text domain are still under-explored, with only a few studies addressing this challenging issue, often with suboptimal results due to equal treatment of submodels or inaccurate estimation of importance scores. To address these challenges, we propose a simple yet effective paradigm for transfer-based textual adversarial attack, named SEP-Attack. Specifically, we employ the Determinantal Point Process (DPP) to generate diverse surrogate ensemble weights, representing the transferability of submodels. Using these weights, we introduce a new metric to evaluate prediction confidence scores, which in turn are used to calculate word importance scores and generate adversarial candidates. Finally, we quantify the transferability score for each candidate and select the top ones as the final transferable adversarial examples. Experiments conducted on four datasets and two real-world APIs validate the efficacy of SEP-Attack, significantly outperforming state-of-the-art baselines.

Anatomy-Anchored Self-Supervision: Distilling Vision Foundation Models for Invariant Ultrasound Representation

arXiv:2605.25402v1 Announce Type: cross Abstract: Self-supervised pre-training paradigm has gained increasing prominence for learning transferable representations in medical imaging, yet existing methods for ultrasound (US) images operate at the image or frame level, overlooking the anatomical context for clinical-aligned representation learning. In this work, we propose an anatomy-anchored ultrasound self-supervision framework ANAUS that shifts representation learning from generic visual regions to clinically meaningful anatomical structures. Utilizing a learnable latent prompt engine alongside a one-time domain adaptation on existing public image--mask pairs, we empower the LP-SAM module to achieve annotation-free anatomy delineation at scale. Building upon this anatomical grounding, we propose a dual-policy self-supervised learning paradigm consisting of inter-view semantics-aware anatomy-separating alignment and contextual core-region prediction to enhance representation learning. Specifically, the former enforces feature invariance within identical anatomical regions while promoting discriminability across distinct structures; the latter compels the model to reconstruct corrupted regions, thereby capturing fine-grained structural details. Extensive evaluations on six public datasets demonstrate that \ours{} consistently outstrips current state-of-the-art methods while maintaining the computational efficiency essential for clinical deployment. Code is available at https://github.com/zhcz328/ANAUS.

SeqRoute: Global Budget-Aware Sequential LLM Routing via Offline Reinforcement Learning

arXiv:2605.25424v1 Announce Type: cross Abstract: Existing LLM routing frameworks treat queries as independent events, neglecting the sequential nature of real-world user sessions constrained by global computational budgets. This mismatch inevitably leads to budget bankruptcy: myopic routing policies exhaust resources on early interactions, forcing subsequent and often more complex queries onto inadequate models. We introduce SeqRoute, a framework that formulates multi-turn routing as a finite-horizon Markov Decision Process and solves it via offline reinforcement learning. By incorporating the remaining budget into the state space and training with Conservative Q-Learning (CQL), SeqRoute learns delayed gratification to strategically preserve resources for high-stakes turns later in the session. To overcome data starvation, we propose Hindsight Budget Relabeling (HBR). This technique retrospectively simulates historical trajectories under diverse hypothetical budgets, expanding 10,000 raw sessions into 2.38 million transitions enriched with critical bankruptcy signals. At deployment, a dynamic $\lambda$-sweep mechanism enables zero-shot navigation of the cost-quality Pareto frontier without retraining. Extensive evaluations demonstrate that SeqRoute reduces operational costs by 6.0-73.5% while maintaining or improving quality, and suppresses bankruptcy rates to under 1%, strictly dominating behavior cloning, budget-aware heuristics, and static baselines across the entire Pareto frontier.

MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM

arXiv:2602.20191v2 Announce Type: replace-cross Abstract: Dynamic runtime latency and memory constraints necessitate flexible large language model (LLM) deployment, where an LLM can be inferred with various quantization precisions based on available computational resources. Recent work on such any-precision quantization either relies on hardware-inefficient vector quantization or induces additional scaling factors when switching between bit-widths. Meanwhile, existing post-training quantization (PTQ) methods calibrated for a fixed low precision show poor generalizability under runtime precision change. In this work, we attribute the source of poor generalization across bit-widths to a precision-dependent \textit{outlier migration} phenomenon where the distribution of PTQ-sensitive tokens changes across precisions. Motivated by this observation, we propose \texttt{MoBiQuant}, a novel any-precision Mixture-of-Bits quantization framework that adjusts weight precision for flexible LLM inference based on token sensitivity. Specifically, we propose a many-in-one recursive residual quantization that can iteratively reconstruct higher-precision weights at runtime and mitigates \textit{outlier migration} with a token-aware router to dynamically select the optimal inference precision of each token.Extensive experiments show that \texttt{MoBiQuant} matches or surpasses frontier single-precision PTQ while exhibiting strong elasticity, achieving significant memory savings and throughput gains of up to $1.34\times$ over state-of-the-art any-precision methods.

Design Conditions for Intra-Group Learning of Sequence-Level Rewards: Token Gradient Cancellation

arXiv:2604.13088v2 Announce Type: replace-cross Abstract: Reinforcement learning for multi-step reasoning with large language models (LLMs) typically relies on sparse terminal rewards, which creates a poorly conditioned credit-assignment problem: the final feedback is propagated uniformly across all intermediate decisions. This leads to high gradient variance, unstable training, and many ineffective updates, ultimately limiting sustained model improvement. We propose a counterfactual-comparison framework for credit assignment. For each input, the framework samples multiple reasoning trajectories and treats their differences as implicit approximations to alternative decisions. This yields an implicit process-level advantage estimator that converts sparse terminal rewards into step-sensitive learning signals. Building on this framework, we introduce Implicit Behavior Policy Optimization (IBPO), which substantially improves training stability and the performance ceiling on mathematical and code-reasoning benchmarks. Our results point to a promising direction for unlocking the reasoning potential of LLMs.

Reducing Credit Assignment Variance via Counterfactual Reasoning Paths

arXiv:2605.16302v2 Announce Type: replace-cross Abstract: Reinforcement learning for multi-step reasoning with large language models (LLMs) typically relies on sparse terminal rewards, which creates a poorly conditioned credit-assignment problem: the final feedback is propagated uniformly across all intermediate decisions. This leads to high gradient variance, unstable training, and many ineffective updates, ultimately limiting sustained model improvement. We propose a counterfactual-comparison framework for credit assignment. For each input, the framework samples multiple reasoning trajectories and treats their differences as implicit approximations to alternative decisions. This yields an implicit process-level advantage estimator that converts sparse terminal rewards into step-sensitive learning signals. Building on this framework, we introduce Implicit Behavior Policy Optimization (IBPO), which substantially improves training stability and the performance ceiling on mathematical and code-reasoning benchmarks. Our results point to a promising direction for unlocking the reasoning potential of LLMs.

How Few-Shot Examples Add Up: A Causal Decomposition of Function Vectors in In-Context Learning

arXiv:2605.16591v2 Announce Type: replace-cross Abstract: In-context learning (ICL) excels at new tasks from minimal examples, yet we still lack a mechanistic explanation of how few-shot prompts shape a model's function vector (FV)--a causal activation direction that drives task behavior on the ICL query. Across tasks and models, an $n$-shot FV is well-approximated by a linear combination of example-level sub-FVs, suggesting additive and composable contributions from individual demonstrations. Beyond additivity, we show that models contextualize individual examples' representations based on prior examples to adaptively reweight which demonstrations dominate the FV: attention shifts toward examples that are more informative and less ambiguous under the context. Finally, a causal decomposition separates Query-Key routing from Value updates, finding that contextualization's most consistent contributions to FV quality arise from Query-Key alignment--particularly in ambiguous settings--while Value-mediated effects are more heterogeneous. Together, these results unify additive superposition with context-dependent attention reweighting into a mechanistic, testable account of how few-shot prompts implement tasks.

The role of growth heterogeneity in solid nodular non-small cell lung cancer in clinical practice: a narrative review

25 May 2026 at 18:00

J Thorac Dis. 2026 Apr 30;18(4):417. doi: 10.21037/jtd-2025-1-2697. Epub 2026 Mar 26.

ABSTRACT

BACKGROUND AND OBJECTIVE: Lung cancer remains the leading cause of cancer related mortality worldwide, and early detection and precise stratified management are crucial for improving patient outcomes. Tumor growth kinetics, as a characterization of its proliferation and malignant differentiation, is a key decision-making factor and research hotspot in clinical practice today. This study aimed to elucidate the growth kinetics of solid nodular non-small cell lung cancer (NSCLC) as a critical determinant of early diagnosis, prognostic evaluation, and treatment strategy selection, and to address the challenge that significant heterogeneity in tumor growth poses to risk stratification and clinical decision-making.

METHODS: We conducted a retrospective search of PubMed, Embase, Web of Science, and Scopus databases, focusing on the current research status of solid nodular NSCLC, particularly in terms of molecular mechanisms, prognosis, modeling prediction, and management strategies related to its growth heterogeneity, with the aim of exploring future research directions.

KEY CONTENT AND FINDINGS: Volume doubling time (VDT) serves as a key metric for evaluating nodule dynamics. While earlier studies suggested a generally rapid growth pattern (VDT <400 days) in solid nodular NSCLC, recent evidence reveals considerable heterogeneity, with some tumors demonstrating indolent growth pattern (VDT >40-600 days). The prognosis of rapidly growing nodules is usually poor, so nodule management recommendations should be personalized based on growth dynamics and patient characteristics. Traditional radiological features, and deep learning models show promise for growth risk stratification but require large-scale external validation and refinement. Molecular and pathological studies suggest that the tumor microenvironment and immune cell infiltration may contribute to growth heterogeneity, though direct mechanistic evidence remains limited. Artificial intelligence (AI) based approaches exhibit significant potential in predicting individual tumor growth behavior.

CONCLUSIONS: Growth heterogeneity in solid nodular NSCLC carries substantial clinical significance but remains insufficiently studied. Future research should prioritize imaging based modeling to predict individualized growth dynamics. Integrating multi-omics analyses may help elucidate the molecular factors underlying growth heterogeneity. AI driven risk stratification based on large-scale multi center sequence data can achieve truly personalized and growth oriented management strategies.

PMID:42182806 | PMC:PMC13190150 | DOI:10.21037/jtd-2025-1-2697

The role of growth heterogeneity in solid nodular non-small cell lung cancer in clinical practice: a narrative review

J Thorac Dis. 2026 Apr 30;18(4):417. doi: 10.21037/jtd-2025-1-2697. Epub 2026 Mar 26.

ABSTRACT

BACKGROUND AND OBJECTIVE: Lung cancer remains the leading cause of cancer related mortality worldwide, and early detection and precise stratified management are crucial for improving patient outcomes. Tumor growth kinetics, as a characterization of its proliferation and malignant differentiation, is a key decision-making factor and research hotspot in clinical practice today. This study aimed to elucidate the growth kinetics of solid nodular non-small cell lung cancer (NSCLC) as a critical determinant of early diagnosis, prognostic evaluation, and treatment strategy selection, and to address the challenge that significant heterogeneity in tumor growth poses to risk stratification and clinical decision-making.

METHODS: We conducted a retrospective search of PubMed, Embase, Web of Science, and Scopus databases, focusing on the current research status of solid nodular NSCLC, particularly in terms of molecular mechanisms, prognosis, modeling prediction, and management strategies related to its growth heterogeneity, with the aim of exploring future research directions.

KEY CONTENT AND FINDINGS: Volume doubling time (VDT) serves as a key metric for evaluating nodule dynamics. While earlier studies suggested a generally rapid growth pattern (VDT <400 days) in solid nodular NSCLC, recent evidence reveals considerable heterogeneity, with some tumors demonstrating indolent growth pattern (VDT >40-600 days). The prognosis of rapidly growing nodules is usually poor, so nodule management recommendations should be personalized based on growth dynamics and patient characteristics. Traditional radiological features, and deep learning models show promise for growth risk stratification but require large-scale external validation and refinement. Molecular and pathological studies suggest that the tumor microenvironment and immune cell infiltration may contribute to growth heterogeneity, though direct mechanistic evidence remains limited. Artificial intelligence (AI) based approaches exhibit significant potential in predicting individual tumor growth behavior.

CONCLUSIONS: Growth heterogeneity in solid nodular NSCLC carries substantial clinical significance but remains insufficiently studied. Future research should prioritize imaging based modeling to predict individualized growth dynamics. Integrating multi-omics analyses may help elucidate the molecular factors underlying growth heterogeneity. AI driven risk stratification based on large-scale multi center sequence data can achieve truly personalized and growth oriented management strategies.

PMID:42182806 | PMC:PMC13190150 | DOI:10.21037/jtd-2025-1-2697

The role of growth heterogeneity in solid nodular non-small cell lung cancer in clinical practice: a narrative review

25 May 2026 at 18:00

J Thorac Dis. 2026 Apr 30;18(4):417. doi: 10.21037/jtd-2025-1-2697. Epub 2026 Mar 26.

ABSTRACT

BACKGROUND AND OBJECTIVE: Lung cancer remains the leading cause of cancer related mortality worldwide, and early detection and precise stratified management are crucial for improving patient outcomes. Tumor growth kinetics, as a characterization of its proliferation and malignant differentiation, is a key decision-making factor and research hotspot in clinical practice today. This study aimed to elucidate the growth kinetics of solid nodular non-small cell lung cancer (NSCLC) as a critical determinant of early diagnosis, prognostic evaluation, and treatment strategy selection, and to address the challenge that significant heterogeneity in tumor growth poses to risk stratification and clinical decision-making.

METHODS: We conducted a retrospective search of PubMed, Embase, Web of Science, and Scopus databases, focusing on the current research status of solid nodular NSCLC, particularly in terms of molecular mechanisms, prognosis, modeling prediction, and management strategies related to its growth heterogeneity, with the aim of exploring future research directions.

KEY CONTENT AND FINDINGS: Volume doubling time (VDT) serves as a key metric for evaluating nodule dynamics. While earlier studies suggested a generally rapid growth pattern (VDT <400 days) in solid nodular NSCLC, recent evidence reveals considerable heterogeneity, with some tumors demonstrating indolent growth pattern (VDT >40-600 days). The prognosis of rapidly growing nodules is usually poor, so nodule management recommendations should be personalized based on growth dynamics and patient characteristics. Traditional radiological features, and deep learning models show promise for growth risk stratification but require large-scale external validation and refinement. Molecular and pathological studies suggest that the tumor microenvironment and immune cell infiltration may contribute to growth heterogeneity, though direct mechanistic evidence remains limited. Artificial intelligence (AI) based approaches exhibit significant potential in predicting individual tumor growth behavior.

CONCLUSIONS: Growth heterogeneity in solid nodular NSCLC carries substantial clinical significance but remains insufficiently studied. Future research should prioritize imaging based modeling to predict individualized growth dynamics. Integrating multi-omics analyses may help elucidate the molecular factors underlying growth heterogeneity. AI driven risk stratification based on large-scale multi center sequence data can achieve truly personalized and growth oriented management strategies.

PMID:42182806 | PMC:PMC13190150 | DOI:10.21037/jtd-2025-1-2697

Pulmonary-Intestinal Axis: Shared Genetic Basis and Mediating Factors Identified Through Multi-Omics Analysis

Int J Chron Obstruct Pulmon Dis. 2026 Apr 7;21:561645. doi: 10.2147/COPD.S561645. eCollection 2026.

ABSTRACT

BACKGROUND: Chronic obstructive pulmonary disease (COPD) is a systemic condition with comorbidities beyond the lung (eg, cardiovascular and metabolic disorders), and gastrointestinal (GI) disorders are also common. The shared genetic basis of COPD-GI comorbidity and its mediating factors remain unclear. We hypothesized that COPD and GI diseases share pleiotropic genetic architecture implicating lipid-metabolic pathways, with smoking mediating part of the association.

METHODS: We analyzed publicly available European-ancestry GWAS summary statistics for COPD (Global Biobank Meta-analysis Initiative), 15 GI diseases (FinnGen), and smoking phenotypes (UK Biobank). Genetic correlation was estimated using linkage disequilibrium score regression (LDSC) and high-definition likelihood (HDL). Multi-trait analysis of GWAS (MTAG) boosted COPD discovery by leveraging genetically correlated GI traits. We integrated locus-to-gene mapping with multi-tissue expression quantitative trait loci (eQTL) and plasma protein quantitative trait loci (pQTL) evidence to prioritize shared loci, genes, and proteins. Bidirectional two-sample Mendelian randomization (MR) tested causal directions, and two-step mediation MR evaluated smoking.

RESULTS: COPD showed significant genetic correlation with nine GI diseases. We identified six comorbidity-associated loci (three with CADD > 12.37) and 13 unique candidate pleiotropic genes; APOE was supported by proteomic evidence. Enrichment analyses highlighted lipid-metabolism pathways. MR suggested COPD increases risk of gastroesophageal reflux disease (GERD), irritable bowel syndrome (IBS), acute appendicitis, and gastric ulcer, while diverticular disease showed reverse causality toward COPD. Smoking partially mediated the COPD effect on GERD, acute appendicitis, and gastric ulcer.

CONCLUSION: COPD and multiple GI disorders share a distributed pleiotropic genetic basis within the broader systemic comorbidity spectrum of COPD. Multi-omics evidence supports a genomic pulmonary-intestinal axis in which lipid metabolism and smoking-related mechanisms contribute to COPD and GI comorbidity, providing targets for risk stratification and potential intervention.

PMID:41978582 | PMC:PMC13070119 | DOI:10.2147/COPD.S561645

Transplantation of encapsulated mitochondria alleviates dysfunction in mitochondrial and Parkinson’s disease models

A mitochondrial transplantation approach rescues mitochondrial deficiency and prevents mitochondrial DNA depletion syndrome, Leigh syndrome, and Parkinson’s disease in cellular and mouse models.
❌