❌

Normal view

Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism

arXiv:2605.23945v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) has become a key post-training paradigm for improving model quality. However, the synchronous three-stage RLHF pipeline is often bottlenecked by the generation stage, where response-length skew causes the effective batch size to shrink rapidly during decoding, leaving GPUs underutilized while a few long responses remain unfinished. Mainstream frameworks employ a static tensor parallelism (TP) configuration that cannot adapt to changing batch characteristics, leaving substantial performance headroom unexplored. We propose PAT, an adaptive TP method that dynamically reconfigures TP during the generation stage of each RLHF iteration. PAT introduces two key techniques. First, a predictor-guided online reconfiguration method decides both the reconfiguration point and the target TP configuration based on offline profiling, triggering reconfiguration only when the predicted latency benefit outweighs the reconfiguration overhead. Second, a lightweight online reconfiguration mechanism updates only the states and layouts affected by TP changes: it adapts unfinished decoding states through a cost-model-based choice between KV-cache migration and recomputation, performs in-place weight resharding, and reuses cached communication groups. We implement PAT on top of SGLang and integrate it with the VeRL framework. Evaluations on LLaMA3.1-8B and Qwen3-14B using DeepScaleR show that PAT reduces generation latency by up to 34.6% and end-to-end RLHF training iteration latency by up to 27.2% compared to the original VeRL setup.

CR1(+) tumor-associated macrophages orchestrate an immunosuppressive niche in hepatocellular carcinoma: a genetic and multi-omics dissection

J Transl Med. 2026 May 25. doi: 10.1186/s12967-026-08301-z. Online ahead of print.

ABSTRACT

BACKGROUND: Hepatocellular carcinoma (HCC) remains a major global health burden and a leading cause of cancer-related mortality. Advanced disease is characterized by a profoundly immunosuppressive tumor microenvironment (TME) and limited durable responses to therapy. However, the upstream genetic determinants that drive tumor-associated macrophage (TAM) dysfunction in HCC remain poorly defined. Using an integrative genetic and multi-omics framework, we investigated complement receptor 1 (CR1) as a candidate regulator of this immunosuppressive niche.

METHODS: We combined Mendelian randomization (MR) and metabolite mediation analyses with bulk, single-cell, and spatial transcriptomics to define the role of CR1 in HCC. Public datasets included the TCGA-HCC cohort, a single-cell RNA-sequencing dataset comprising 53,474 high-quality cells from 21 samples, and two spatially profiled HCC sections. Clinical validation was performed in 30 paired HCC and adjacent liver tissues. Functional assays were conducted in THP-1-derived macrophages using CR1 gain- and loss-of-function approaches, phagocytosis assays, and macrophage-CD8+ T-cell co-culture experiments.

RESULTS: MR analyses implicated CR1 in HCC susceptibility at both the protein and transcript levels. pQTL analysis linked genetically predicted circulating CR1 levels to HCC risk (IVW OR = 1.403, p = 0.017), and mediation analysis identified specific metabolites as candidate intermediates. Integrative multi-omics analyses showed that CR1 was preferentially enriched in TAMs, spatially co-localized with the M2 marker CD206, and associated with reduced CD8+ T-cell infiltration, enhanced T-cell exhaustion signatures, advanced clinicopathological features, and poorer survival. In 30 paired clinical samples, CR1-high tumors exhibited increased M2-like macrophage accumulation and reduced CD8+ T-cell infiltration. Functionally, CR1 overexpression drove macrophages toward an M2-like phenotype, enhanced phagocytic activity, increased PD-L1 expression, and suppressed CD8+ T-cell proliferation as well as IFN-gamma and granzyme B production, whereas CR1 knockdown produced the opposite phenotype.

CONCLUSIONS: Our study provides the first integrated genetic, spatial, and functional evidence that CR1+ TAMs constitute a clinically relevant immunoregulatory axis in HCC. These findings extend current understanding of complement-associated immunosuppression beyond canonical complement cascade activity and support CR1 as a candidate biomarker and therapeutic target for macrophage reprogramming, with potential translational relevance for combination strategies involving immune checkpoint blockade.

PMID:42185899 | DOI:10.1186/s12967-026-08301-z

CR1(+) tumor-associated macrophages orchestrate an immunosuppressive niche in hepatocellular carcinoma: a genetic and multi-omics dissection

J Transl Med. 2026 May 25. doi: 10.1186/s12967-026-08301-z. Online ahead of print.

ABSTRACT

BACKGROUND: Hepatocellular carcinoma (HCC) remains a major global health burden and a leading cause of cancer-related mortality. Advanced disease is characterized by a profoundly immunosuppressive tumor microenvironment (TME) and limited durable responses to therapy. However, the upstream genetic determinants that drive tumor-associated macrophage (TAM) dysfunction in HCC remain poorly defined. Using an integrative genetic and multi-omics framework, we investigated complement receptor 1 (CR1) as a candidate regulator of this immunosuppressive niche.

METHODS: We combined Mendelian randomization (MR) and metabolite mediation analyses with bulk, single-cell, and spatial transcriptomics to define the role of CR1 in HCC. Public datasets included the TCGA-HCC cohort, a single-cell RNA-sequencing dataset comprising 53,474 high-quality cells from 21 samples, and two spatially profiled HCC sections. Clinical validation was performed in 30 paired HCC and adjacent liver tissues. Functional assays were conducted in THP-1-derived macrophages using CR1 gain- and loss-of-function approaches, phagocytosis assays, and macrophage-CD8+ T-cell co-culture experiments.

RESULTS: MR analyses implicated CR1 in HCC susceptibility at both the protein and transcript levels. pQTL analysis linked genetically predicted circulating CR1 levels to HCC risk (IVW OR = 1.403, p = 0.017), and mediation analysis identified specific metabolites as candidate intermediates. Integrative multi-omics analyses showed that CR1 was preferentially enriched in TAMs, spatially co-localized with the M2 marker CD206, and associated with reduced CD8+ T-cell infiltration, enhanced T-cell exhaustion signatures, advanced clinicopathological features, and poorer survival. In 30 paired clinical samples, CR1-high tumors exhibited increased M2-like macrophage accumulation and reduced CD8+ T-cell infiltration. Functionally, CR1 overexpression drove macrophages toward an M2-like phenotype, enhanced phagocytic activity, increased PD-L1 expression, and suppressed CD8+ T-cell proliferation as well as IFN-gamma and granzyme B production, whereas CR1 knockdown produced the opposite phenotype.

CONCLUSIONS: Our study provides the first integrated genetic, spatial, and functional evidence that CR1+ TAMs constitute a clinically relevant immunoregulatory axis in HCC. These findings extend current understanding of complement-associated immunosuppression beyond canonical complement cascade activity and support CR1 as a candidate biomarker and therapeutic target for macrophage reprogramming, with potential translational relevance for combination strategies involving immune checkpoint blockade.

PMID:42185899 | DOI:10.1186/s12967-026-08301-z

CR1(+) tumor-associated macrophages orchestrate an immunosuppressive niche in hepatocellular carcinoma: a genetic and multi-omics dissection

J Transl Med. 2026 May 25. doi: 10.1186/s12967-026-08301-z. Online ahead of print.

ABSTRACT

BACKGROUND: Hepatocellular carcinoma (HCC) remains a major global health burden and a leading cause of cancer-related mortality. Advanced disease is characterized by a profoundly immunosuppressive tumor microenvironment (TME) and limited durable responses to therapy. However, the upstream genetic determinants that drive tumor-associated macrophage (TAM) dysfunction in HCC remain poorly defined. Using an integrative genetic and multi-omics framework, we investigated complement receptor 1 (CR1) as a candidate regulator of this immunosuppressive niche.

METHODS: We combined Mendelian randomization (MR) and metabolite mediation analyses with bulk, single-cell, and spatial transcriptomics to define the role of CR1 in HCC. Public datasets included the TCGA-HCC cohort, a single-cell RNA-sequencing dataset comprising 53,474 high-quality cells from 21 samples, and two spatially profiled HCC sections. Clinical validation was performed in 30 paired HCC and adjacent liver tissues. Functional assays were conducted in THP-1-derived macrophages using CR1 gain- and loss-of-function approaches, phagocytosis assays, and macrophage-CD8+ T-cell co-culture experiments.

RESULTS: MR analyses implicated CR1 in HCC susceptibility at both the protein and transcript levels. pQTL analysis linked genetically predicted circulating CR1 levels to HCC risk (IVW OR = 1.403, p = 0.017), and mediation analysis identified specific metabolites as candidate intermediates. Integrative multi-omics analyses showed that CR1 was preferentially enriched in TAMs, spatially co-localized with the M2 marker CD206, and associated with reduced CD8+ T-cell infiltration, enhanced T-cell exhaustion signatures, advanced clinicopathological features, and poorer survival. In 30 paired clinical samples, CR1-high tumors exhibited increased M2-like macrophage accumulation and reduced CD8+ T-cell infiltration. Functionally, CR1 overexpression drove macrophages toward an M2-like phenotype, enhanced phagocytic activity, increased PD-L1 expression, and suppressed CD8+ T-cell proliferation as well as IFN-gamma and granzyme B production, whereas CR1 knockdown produced the opposite phenotype.

CONCLUSIONS: Our study provides the first integrated genetic, spatial, and functional evidence that CR1+ TAMs constitute a clinically relevant immunoregulatory axis in HCC. These findings extend current understanding of complement-associated immunosuppression beyond canonical complement cascade activity and support CR1 as a candidate biomarker and therapeutic target for macrophage reprogramming, with potential translational relevance for combination strategies involving immune checkpoint blockade.

PMID:42185899 | DOI:10.1186/s12967-026-08301-z

TABQAWORLD: Optimizing Multimodal Reasoning for Multi-Turn Table Question Answering

arXiv:2604.03393v1 Announce Type: new Abstract: Multimodal reasoning has emerged as a powerful framework for enhancing reasoning capabilities of reasoning models. While multi-turn table reasoning methods have improved reasoning accuracy through tool use and reward modeling, they rely on fixed text serialization for table state readouts. This introduces representation errors in table encoding that significantly accumulate over multiple turns. Such accumulation is alleviated by tabular grounding methods in the expense of inference compute and cost, rendering real world deployment impractical. To address this, we introduce TABQAWORLD, a table reasoning framework that jointly optimizes tabular action through representation and estimation. For representation, TABQAWORLD employs an action-conditioned multimodal selection policy, which dynamically switches between visual and textual representations to maximize table state readout reliability. For estimation, TABQAWORLD optimizes stepwise reasoning trajectory through table metadata including dimension, data types and key values, safely planning trajectory and compressing low-complexity actions to reduce conversation turns and latency. Designed as a training-free framework, empirical evaluations show that TABQAWORLD achieves state-of-the-art performance with 4.87% accuracy improvements over baselines, with 5.42% accuracy gain and 33.35% inference latency reduction over static settings, establishing a new standard for reliable and efficient table reasoning.

Four Generations of Quantum Biomedical Sensors

arXiv:2603.29944v2 Announce Type: replace-cross Abstract: Quantum sensing technologies offer transformative potential for ultra-sensitive biomedical sensing, yet their clinical translation remains constrained by classical noise limits and a reliance on macroscopic ensembles. We propose a unifying generational framework to organize the evolving landscape of quantum biosensors based on their utilization of quantum resources. First-generation devices utilize discrete energy levels for signal transduction but follow classical scaling laws. Second-generation sensors exploit quantum coherence to reach the standard quantum limit, while third-generation architectures leverage entanglement and spin squeezing to approach Heisenberg-limited precision. We further define an emerging fourth generation characterized by the end-to-end integration of quantum sensing with quantum learning and variational circuits, enabling adaptive inference directly within the quantum domain. By analyzing critical parameters such as bandwidth matching and sensor-tissue proximity, we identify key technological bottlenecks and propose a roadmap for transitioning from measuring physical observables to extracting structured biological information with quantum-enhanced intelligence.

Pyruvate is a natural suppressor of interferon signaling by inducing STAT1 protein pyruvylation

Yibo et al. identify protein pyruvylation as a post-translational modification that can modulate immune signaling and host antiviral response.

Four Generations of Quantum Biomedical Sensors

arXiv:2603.29944v1 Announce Type: cross Abstract: Quantum sensing technologies offer transformative potential for ultra-sensitive biomedical sensing, yet their clinical translation remains constrained by classical noise limits and a reliance on macroscopic ensembles. We propose a unifying generational framework to organize the evolving landscape of quantum biosensors based on their utilization of quantum resources. First-generation devices utilize discrete energy levels for signal transduction but follow classical scaling laws. Second-generation sensors exploit quantum coherence to reach the standard quantum limit, while third-generation architectures leverage entanglement and spin squeezing to approach Heisenberg-limited precision. We further define an emerging fourth generation characterized by the end-to-end integration of quantum sensing with quantum learning and variational circuits, enabling adaptive inference directly within the quantum domain. By analyzing critical parameters such as bandwidth matching and sensor-tissue proximity, we identify key technological bottlenecks and propose a roadmap for transitioning from measuring physical observables to extracting structured biological information with quantum-enhanced intelligence.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models

arXiv:2506.09082v4 Announce Type: replace-cross Abstract: The rise of vision foundation models (VFMs) calls for systematic evaluation. A common approach pairs VFMs with large language models (LLMs) as general-purpose heads, followed by evaluation on broad Visual Question Answering (VQA) benchmarks. However, this protocol has two key blind spots: (i) the instruction tuning data may not align with VQA test distributions, meaning a wrong prediction can stem from such data mismatch rather than a VFM' visual shortcomings; (ii) VQA benchmarks often require multiple visual abilities, making it hard to tell whether errors stem from lacking all required abilities or just a single critical one. To address these gaps, we introduce AVA-Bench, the first benchmark that explicitly disentangles 14 Atomic Visual Abilities (AVAs) -- foundational skills like localization, depth estimation, and spatial understanding that collectively support complex visual reasoning tasks. By decoupling AVAs and matching training and test distributions within each, AVA-Bench pinpoints exactly where a VFM excels or falters. Applying AVA-Bench to leading VFMs thus reveals distinctive "ability fingerprints," turning VFM selection from educated guesswork into principled engineering. Notably, we find that a 0.5B LLM yields similar VFM rankings as a 7B LLM while cutting GPU hours by 8x, enabling more efficient evaluation. By offering a comprehensive and transparent benchmark, we hope AVA-Bench lays the foundation for the next generation of VFMs.

ToolTree: Efficient LLM Agent Tool Planning via Dual-Feedback Monte Carlo Tree Search and Bidirectional Pruning

arXiv:2603.12740v1 Announce Type: new Abstract: Large Language Model (LLM) agents are increasingly applied to complex, multi-step tasks that require interaction with diverse external tools across various domains. However, current LLM agent tool planning methods typically rely on greedy, reactive tool selection strategies that lack foresight and fail to account for inter-tool dependencies. In this paper, we present ToolTree, a novel Monte Carlo tree search-inspired planning paradigm for tool planning. ToolTree explores possible tool usage trajectories using a dual-stage LLM evaluation and bidirectional pruning mechanism that enables the agent to make informed, adaptive decisions over extended tool-use sequences while pruning less promising branches before and after the tool execution. Empirical evaluations across both open-set and closed-set tool planning tasks on 4 benchmarks demonstrate that ToolTree consistently improves performance while keeping the highest efficiency, achieving an average gain of around 10\% compared to the state-of-the-art planning paradigm.

Asynchronous Verified Semantic Caching for Tiered LLM Architectures

arXiv:2602.13165v2 Announce Type: replace-cross Abstract: Large language models (LLMs) now sit in the critical path of search, assistance, and agentic workflows, making semantic caching essential for reducing inference cost and latency. Production deployments typically use a tiered static-dynamic design: a static cache of curated, offline vetted responses mined from logs, backed by a dynamic cache populated online. In practice, both tiers are commonly governed by a single embedding similarity threshold, which induces a hard tradeoff: conservative thresholds miss safe reuse opportunities, while aggressive thresholds risk serving semantically incorrect responses. We introduce Krites, an asynchronous, LLM-judged caching policy that expands static coverage without changing serving decisions. On the critical path, Krites behaves exactly like a standard static threshold policy. When the nearest static neighbor of the prompt falls just below the static threshold, Krites asynchronously invokes an LLM judge to verify whether the static response is acceptable for the new prompt. Approved matches are promoted into the dynamic cache, allowing future repeats and paraphrases to reuse curated static answers and expanding static reach over time. In trace-driven simulations on conversational and search workloads, Krites increases the fraction of requests served with curated static answers (direct static hits plus verified promotions) by up to3.9 times for conversational traffic and search-style queries relative to tuned baselines, with unchanged critical path latency.

The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable Reward

arXiv:2509.07430v4 Announce Type: replace-cross Abstract: A central paradox in fine-tuning Large Language Models (LLMs) with Reinforcement Learning with Verifiable Reward (RLVR) is the frequent degradation of multi-attempt performance (Pass@k) despite improvements in single-attempt accuracy (Pass@1). This is often accompanied by catastrophic forgetting, where models lose previously acquired skills. While various methods have been proposed, the choice and function of the divergence term have been surprisingly unexamined as a proactive solution. We argue that standard RLVR objectives -- both those using the mode-seeking reverse KL-divergence and those forgoing a divergence term entirely -- lack a crucial mechanism for knowledge retention. The reverse-KL actively accelerates this decay by narrowing the policy, while its absence provides no safeguard against the model drifting from its diverse knowledge base. We propose a fundamental shift in perspective: using the divergence term itself as the solution. Our framework, Diversity-Preserving Hybrid RL (DPH-RL), leverages mass-covering f-divergences (like forward-KL and JS-divergence) to function as a rehearsal mechanism. By continuously referencing the initial policy, this approach forces the model to maintain broad solution coverage. Extensive experiments on math and SQL generation demonstrate that DPH-RL not only resolves the Pass@k degradation but improves both Pass@1 and Pass@k in- and out-of-domain. Additionally, DPH-RL is more training-efficient because it computes f-divergence using generator functions, requiring only sampling from the initial policy and no online reference model. Our work highlights a crucial, overlooked axis for improving RLVR, demonstrating that the proper selection of a divergence measure is a powerful tool for building more general and diverse reasoning models.

CausalFlip: A Benchmark for LLM Causal Judgment Beyond Semantic Matching

arXiv:2602.20094v1 Announce Type: new Abstract: As large language models (LLMs) witness increasing deployment in complex, high-stakes decision-making scenarios, it becomes imperative to ground their reasoning in causality rather than spurious correlations. However, strong performance on traditional reasoning benchmarks does not guarantee true causal reasoning ability of LLMs, as high accuracy may still arise from memorizing semantic patterns instead of analyzing the underlying true causal structures. To bridge this critical gap, we propose a new causal reasoning benchmark, CausalFlip, designed to encourage the development of new LLM paradigm or training algorithms that ground LLM reasoning in causality rather than semantic correlation. CausalFlip consists of causal judgment questions built over event triples that could form different confounder, chain, and collider relations. Based on this, for each event triple, we construct pairs of semantically similar questions that reuse the same events but yield opposite causal answers, where models that rely heavily on semantic matching are systematically driven toward incorrect predictions. To further probe models' reliance on semantic patterns, we introduce a noisy-prefix evaluation that prepends causally irrelevant text before intermediate causal reasoning steps without altering the underlying causal relations or the logic of the reasoning process. We evaluate LLMs under multiple training paradigms, including answer-only training, explicit Chain-of-Thought (CoT) supervision, and a proposed internalized causal reasoning approach that aims to mitigate explicit reliance on correlation in the reasoning process. Our results show that explicit CoT can still be misled by spurious semantic correlations, where internalizing reasoning steps yields substantially improved causal grounding, suggesting that it is promising to better elicit the latent causal reasoning capabilities of base LLMs.
❌