❌

Normal view

KDM4A drives TGCT metastasis by inducing focal adhesion disassembly via STAT1-mediated <i>CCL3</i> transcriptional activation

Oncogene, Published online: 04 October 2026; doi:10.1038/s41388-026-04002-5

KDM4A drives TGCT metastasis by inducing focal adhesion disassembly via STAT1-mediated CCL3 transcriptional activation

Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

arXiv:2609.11318v2 Announce Type: replace Abstract: Deep research agents are increasingly capable of web search, tool use, multimodal evidence analysis, and information synthesis. However, existing benchmarks mainly evaluate medium-horizon exploration and rarely test whether agents can sustain long, dependency-heavy research processes. We introduce Mr. LHDR (Multimodal real-world Long-Horizon Deep Research), a benchmark for evaluating real-world deep research over long, irreducible chains of interdependent evidence across eight categories. Each question is constructed from a hidden Node-Relation graph and requires an average of 12.1 necessary intermediate conclusions with a mean dependency depth of 10.4 before reaching a short, unique, and verifiable answer. Questions incorporate multimodal evidence, including images, maps, PDFs, logos, charts, tables, and video frames, with at least one non-text element that changes the reasoning state. Mr. LHDR evaluates both final answers and the correctness of intermediate conclusions under annotated dependencies. We evaluate general models, deep research systems, and agent frameworks using Overall Accuracy (OA), Strict Accuracy (SA), Checklist Score (CS), and Dependency-Aware Checklist Score (DACS). Results show that even the strongest system achieves only 43.1% OA and 34.3% SA, indicating that final-answer accuracy substantially overestimates complete research success. Removing images reduces DACS by 12.6 points, demonstrating the importance of multimodal evidence, while SA consistently declines as reasoning chains become longer. These findings reveal sustained, dependency-consistent evidence integration, rather than isolated fact retrieval, as a key bottleneck for current deep research agents.

MCAT-mediated mitochondrial fatty acid metabolism regulates Lauren subtype divergence and suppresses gastric cancer progression through ROS/P53-dependent mitophagy and ferroptosis

Cell Death Differ. 2026 Sep 8. doi: 10.1038/s41418-026-01867-7. Online ahead of print.

ABSTRACT

Gastric cancer (GC) displays marked heterogeneity under the Lauren classification, yet the metabolic determinants of subtype divergence remain unclear. Here, we identify Malonyl-CoA:ACP transacylase (MCAT), a Lauren subtype-associated gene encoding a key mitochondrial fatty acid synthesis (mtFAS) enzyme, as a subtype-specific tumor suppressor in GC. Integrative multi-omics profiling revealed that MCAT expression is enriched in intestinal-type GC and correlates with favorable prognosis. Mechanistically, MCAT overexpression drives metabolic reprogramming through mitochondrial free fatty acid overload, suppressing β-oxidation while elevating mitochondrial reactive oxygen species (ROS), which triggers P53 phosphorylation at Ser15. This event concurrently activates PINK1/Parkin-mediated mitophagy and suppresses the SLC7A11/GPX4 axis to induce ferroptosis. Genetic rescue experiments confirmed that P53-Ser15 phosphorylation is essential for both mitophagy and ferroptosis induction. Endogenous MCAT levels are sufficient to determine basal ROS/P53/mitophagy/ferroptosis axis activity, and knockdown in high-expressing cells reverses these phenotypes, supporting a physiological, threshold-dependent role. In vivo, MCAT overexpression suppresses tumor growth and enhances mitophagy and ferroptosis markers. Collectively, these findings establish MCAT as a metabolic switch that links mtFAS to ROS/P53-dependent cell death, providing a potential biomarker and therapeutic target for GC.

PMID:42711380 | DOI:10.1038/s41418-026-01867-7

A programmed cell death learning signature predicts immunotherapy response and identifies AP1S1 as a regulator of immune exclusion in breast cancer

Chin J Cancer Res. 2026 Aug 30;38(4):480-500. doi: 10.21147/j.issn.1000-9604.2026.04.08.

ABSTRACT

OBJECTIVE: Breast cancer remains a leading cause of global cancer mortality, characterized by profound heterogeneity. While immune checkpoint blockade (ICB) has transformed oncology, its efficacy in breast cancer is often hindered by "immune-cold" microenvironments and immune exclusion. Programmed cell death (PCD) is a critical regulator of tumor immune microenvironment (TIME). However, its role in the breast cancer immune microenvironment remains poorly understood.

METHODS: We integrated multi-omics data from six breast cancer cohorts (N=3,764) to develop a programmed cell death learning signature (PCDsig) using over 100 machine learning combinations. The model was benchmarked against 29 published signatures. Single-cell transcriptomic analysis decoded the immune landscape and cellular crosstalk. The role of adaptor-related protein complex 1 subunit sigma 1 (AP1S1) was validated through a clinical cohort, in vitro functional assays, and in vivo syngeneic mouse models.

RESULTS: PCDsig significantly stratified patient prognosis across all cohorts, consistently outperforming 29 existing models. High PCDsig scores correlated with immune-excluded phenotypes, reduced CD8+ T cell infiltration, and lower immunophenoscores. Single-cell analysis revealed that high-PCDsig tumors utilize vascular endothelial growth factor A (VEGFA) signaling to foster an immunosuppressive microenvironment. AP1S1 was identified as the core driver of immune exclusion. And our clinical cohort supported the immune exclusion effect of AP1S1. AP1S1 knockdown impaired tumor progression in vitro and fundamentally remodeled the tumor immune ecosystem in vivo. Combining AP1S1 inhibition with anti-programmed cell death ligand 1 (anti-PD-L1) therapy exerted profound synergistic effects, driven by massive infiltration and functional activation of cytotoxic Granzyme B (GZMB)+CD8+ T cells.

CONCLUSIONS: Our study establishes the PCDsig we developed is a potential prognostic and predictive biomarker for breast cancer. We provide the first evidence of AP1S1 as a core immunomodulatory oncogene that mediates immune exclusion. Targeting AP1S1 represents a highly promising strategy to sensitize cold breast tumors to ICB, offering a new perspective for precision immunotherapy.

PMID:42712842 | PMC:PMC13551362 | DOI:10.21147/j.issn.1000-9604.2026.04.08

MCAT-mediated mitochondrial fatty acid metabolism regulates Lauren subtype divergence and suppresses gastric cancer progression through ROS/P53-dependent mitophagy and ferroptosis

Cell Death Differ. 2026 Sep 8. doi: 10.1038/s41418-026-01867-7. Online ahead of print.

ABSTRACT

Gastric cancer (GC) displays marked heterogeneity under the Lauren classification, yet the metabolic determinants of subtype divergence remain unclear. Here, we identify Malonyl-CoA:ACP transacylase (MCAT), a Lauren subtype-associated gene encoding a key mitochondrial fatty acid synthesis (mtFAS) enzyme, as a subtype-specific tumor suppressor in GC. Integrative multi-omics profiling revealed that MCAT expression is enriched in intestinal-type GC and correlates with favorable prognosis. Mechanistically, MCAT overexpression drives metabolic reprogramming through mitochondrial free fatty acid overload, suppressing β-oxidation while elevating mitochondrial reactive oxygen species (ROS), which triggers P53 phosphorylation at Ser15. This event concurrently activates PINK1/Parkin-mediated mitophagy and suppresses the SLC7A11/GPX4 axis to induce ferroptosis. Genetic rescue experiments confirmed that P53-Ser15 phosphorylation is essential for both mitophagy and ferroptosis induction. Endogenous MCAT levels are sufficient to determine basal ROS/P53/mitophagy/ferroptosis axis activity, and knockdown in high-expressing cells reverses these phenotypes, supporting a physiological, threshold-dependent role. In vivo, MCAT overexpression suppresses tumor growth and enhances mitophagy and ferroptosis markers. Collectively, these findings establish MCAT as a metabolic switch that links mtFAS to ROS/P53-dependent cell death, providing a potential biomarker and therapeutic target for GC.

PMID:42711380 | DOI:10.1038/s41418-026-01867-7

A programmed cell death learning signature predicts immunotherapy response and identifies AP1S1 as a regulator of immune exclusion in breast cancer

Chin J Cancer Res. 2026 Aug 30;38(4):480-500. doi: 10.21147/j.issn.1000-9604.2026.04.08.

ABSTRACT

OBJECTIVE: Breast cancer remains a leading cause of global cancer mortality, characterized by profound heterogeneity. While immune checkpoint blockade (ICB) has transformed oncology, its efficacy in breast cancer is often hindered by "immune-cold" microenvironments and immune exclusion. Programmed cell death (PCD) is a critical regulator of tumor immune microenvironment (TIME). However, its role in the breast cancer immune microenvironment remains poorly understood.

METHODS: We integrated multi-omics data from six breast cancer cohorts (N=3,764) to develop a programmed cell death learning signature (PCDsig) using over 100 machine learning combinations. The model was benchmarked against 29 published signatures. Single-cell transcriptomic analysis decoded the immune landscape and cellular crosstalk. The role of adaptor-related protein complex 1 subunit sigma 1 (AP1S1) was validated through a clinical cohort, in vitro functional assays, and in vivo syngeneic mouse models.

RESULTS: PCDsig significantly stratified patient prognosis across all cohorts, consistently outperforming 29 existing models. High PCDsig scores correlated with immune-excluded phenotypes, reduced CD8+ T cell infiltration, and lower immunophenoscores. Single-cell analysis revealed that high-PCDsig tumors utilize vascular endothelial growth factor A (VEGFA) signaling to foster an immunosuppressive microenvironment. AP1S1 was identified as the core driver of immune exclusion. And our clinical cohort supported the immune exclusion effect of AP1S1. AP1S1 knockdown impaired tumor progression in vitro and fundamentally remodeled the tumor immune ecosystem in vivo. Combining AP1S1 inhibition with anti-programmed cell death ligand 1 (anti-PD-L1) therapy exerted profound synergistic effects, driven by massive infiltration and functional activation of cytotoxic Granzyme B (GZMB)+CD8+ T cells.

CONCLUSIONS: Our study establishes the PCDsig we developed is a potential prognostic and predictive biomarker for breast cancer. We provide the first evidence of AP1S1 as a core immunomodulatory oncogene that mediates immune exclusion. Targeting AP1S1 represents a highly promising strategy to sensitize cold breast tumors to ICB, offering a new perspective for precision immunotherapy.

PMID:42712842 | PMC:PMC13551362 | DOI:10.21147/j.issn.1000-9604.2026.04.08

A Comprehensive Review of Radiomics in Pulmonary Nodule Management: Clinical Applications and Standardization Dilemmas

23 June 2026 at 18:00

Curr Med Imaging. 2026 Jun 22. doi: 10.2174/0115734056460566260609044755. Online ahead of print.

ABSTRACT

Lung cancer is the most common and fatal malignant tumour. Early detection and treatment are likely to reduce mortality, but most pulmonary nodules identified during routine health checks are harmless. Consequently, a clear distinction between benign and malignant nodules is vital to improve early detection and reduce unnecessary interventions. Radiomics, a new omics technology, can be used to extract high-dimensional quantitative features from medical images, providing a profound understanding of tumour pathophysiology. Radiomics has attracted the attention of medical researchers since its formal definition by the Dutch researcher Lambin et al. in 2012. The number of research papers on radiomics has grown tremendously over the past few years. At present, it is used to predict pulmonary nodule malignancy, for noninvasive risk stratification, for integration with genomics to identify genetic mutations associated with lung cancer, and for evaluation of therapeutic responses. With this review, we summarise the literature on radiomics of pulmonary nodules, discuss how it could be used in nodule management, and address the current challenges and future directions for improving precision oncology.

PMID:42333843 | DOI:10.2174/0115734056460566260609044755

How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning

arXiv:2605.23926v1 Announce Type: new Abstract: Reasoning-capable large language models solve hard problems by emitting long chains of thought, paying heavily in latency, GPU time, and energy. Casual inspection of their traces reveals extensive reformulation, verification, and circular self-reflection, yet how much of this deliberation is actually necessary has never been measured at scale or explained from first principles. This paper closes both gaps. We formalise reasoning redundancy directly in terms of the reasoning model itself: the redundancy of a correct trace is the largest fraction of its trailing segmented steps that can be truncated while $\pi$, forced to terminate thinking and emit a final answer, still produces the correct answer. A large-scale quantification across four frontier reasoning models and two mathematical benchmarks shows that step-level redundancy is consistently high -- between 61% and 93% across the 8 (model, benchmark) conditions we study, with the median critical prefix equal to a single segmented step in six of the eight conditions -- that the finding is robust to the choice of judge family, and that although $\rho$ decreases with problem difficulty on MATH-500, all four models remain substantially redundant ($\rho \in [46\%, 85\%]$) even on the hardest Level-5 problems. We then prove that this redundancy is a structural consequence of length-agnostic outcome rewards, not a model-specific artefact: under any such reward, no finite expected stopping time is optimal. The result holds regardless of RL algorithm, base model, data distribution, or whether the policy is obtained via RL or distillation; over-thinking is therefore not a bug to be patched in individual models but a structural property of how current reasoning models are trained. Code: https://github.com/zhiyuanZhai20/how-much-thinking-is-enough

SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills

arXiv:2605.24117v1 Announce Type: new Abstract: Large language model (LLM) agents accumulate rich episodic trajectories while solving real-world tasks, but it remains unclear whether such experience can be distilled into reusable procedural skills. We introduce SkillEvolBench, a diagnostic benchmark for evaluating this step from experience reuse to skill formation. It contains 180 tasks across six real-world agent environments, organized into role-conditioned task families with shared latent procedures. Agents learn from acquisition tasks, update an external skill library using compacted trajectories and verifier feedback, and then face frozen deployment tasks testing context shift, adversarial shortcuts, and composition. By comparing self-generated and curated-start skill evolution against no-skill and raw-trajectory controls, SkillEvolBench separates procedural abstraction from base capability, curated prior knowledge, and direct reuse of episodic traces. Across ten model configurations and three agent harnesses, we find that current agents often adapt locally but rarely form robust reusable skills. Skill-based conditions can improve acquisition or replay, and individual models sometimes gain on specific deployment axes, but these gains are unstable under frozen deployment. Raw-trajectory reuse frequently outperforms distilled skills, suggesting that current abstraction procedures discard contextual and procedural cues that remain useful for future tasks. Capacity and cost analyses further show that writing more skills or larger Tier-3 resource libraries is not sufficient: additional updates can improve coverage while introducing episode-specific drift and procedural clutter. These findings position SkillEvolBench as a testbed for measuring when one-off experience becomes durable procedural knowledge rather than task-local memory.

Detecting Unfaithful Chain-of-Thought via Circuit-Guided Internal-External Discrepancy

arXiv:2605.25603v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning improves the problem-solving ability of large language models (LLMs), but generated reasoning traces may not faithfully reflect the model's actual decision process. Existing CoT unfaithfulness detectors mainly rely on external signals from generated rationales, such as textual plausibility or answer consistency, while overlooking evidence from the model's internal computation. Although recent circuit tracing methods provide a way to obtain model-internal evidence by tracing how information flows through model components during reasoning, constructing full reasoning circuits for long CoTs is costly and difficult to scale. To address these challenges, we propose Circuit-guided Internal-External Discrepancy Scorer (CIE-Scorer), a framework for instance-level CoT unfaithfulness detection. The key idea is that faithful reasoning traces should align with the model's computational process, whereas unfaithful traces may diverge from it. CIE-Scorer efficiently traces compact sentence-level circuits from informative reasoning tokens, constructs internal and external reasoning graphs, and measures their discrepancy using Fused Gromov--Wasserstein distance. Experiments on four datasets from FaithCoT-Bench show that CIE-Scorer achieves state-of-the-art performance while reducing the cost of circuit construction, demonstrating the effectiveness of combining mechanistic interpretability signals with external reasoning traces for CoT unfaithfulness detection.

Inference-Time Alignment of Diffusion Models via Trust-Region Iterative Twisted Sequential Monte Carlo

arXiv:2605.25123v1 Announce Type: cross Abstract: We study inference-time alignment for diffusion-based generative models, aiming to steer a base model toward high-reward outputs without updating its weights. Recent Sequential Monte Carlo (SMC)-based steering methods approximate reward-tilted target distributions in a principled way, but their proposals remain largely tied to the base sampler. Since reward information is mainly used after propagation through particle reweighting and resampling, these methods can require large particle budgets and suffer from weight degeneracy and high-variance estimates. One way to reduce variance and improve particle efficiency is to iteratively learn twisting functions that provide look-ahead guidance, as in twisted SMC. However, existing learnable twisting methods are developed mainly for classical sequential inference and can be unstable when applied to diffusion-based alignment with high-dimensional state spaces and terminal, noisy, or black-box rewards. We propose Trust-Region Iterative Twisted Sequential Monte Carlo (TRI-TSMC), a trust-region framework for learning twisting functions in SMC-based inference-time alignment. Each iteration computes an exact KL-constrained update in path space, which admits a closed-form solution by tempered importance reweighting, and projects this target back to the parameterized twisted family by weighted maximum likelihood. Theoretically, we formalize the value-function interpretation of the optimal twisting function and show that it yields a zero-variance sampler. We prove that the trust-region update follows an escort path toward the target distribution, that the weighted maximum-likelihood update is a forward-KL projection, and that the path reduces residual importance-weight variance. Empirically, TRI-TSMC improves primary alignment objectives on discrete diffusion text generation and text-to-image generation under matched inference-time budgets.

Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent Games

arXiv:2605.04906v2 Announce Type: replace Abstract: While Large Language Models (LLMs) excel in certain reasoning tasks, they struggle in multi-agent games where the final outcome depends on the joint strategies of all agents. In multi-agent games, the non-stationarity of other agents brings significant challenges on the evaluation of the reasoning process and the credit assignment over multiple reasoning steps. Existing single-agent reinforcement learning (RL) approaches and their multi-agent extensions fail to address these challenges as they do not incorporate other agents in the reasoning process. In this work, we propose Strat-Reasoner, a novel RL-based framework that improves LLMs' strategic reasoning ability in multi-agent games. We introduce a novel recursive reasoning paradigm where an agent's reasoning also integrates other agents' reasoning processes. To provide effective reward signals for the intermediate reasoning sequences, we employ a centralized Chain-of-Thought (CoT) comparison module to evaluate the reasoning quality. Finally, we compute an accurate hybrid advantage and develop a group-relative RL approach to optimize the LLM policy. Experimental results show that Strat-Reasoner substantially improves strategic abilities of underlying LLMs, achieving 22.1\% average performance improvements across various multi-agent games. Code is publicly available at https://github.com/ydhe1012/Strat-Reasoner.

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

arXiv:2602.18527v2 Announce Type: replace-cross Abstract: Current audio-visual large language models (AV-LLMs) are predominantly restricted to 2D perception, relying on RGB video and monaural audio. This design choice introduces a fundamental dimensionality mismatch that precludes reliable source localization and spatial reasoning in complex 3D environments. We address this limitation by presenting JAEGER, a framework that extends AV-LLMs to 3D space, to enable joint spatial grounding and reasoning through the integration of RGB-D observations and multi-channel first-order ambisonics. A core contribution of our work is the neural intensity vector (Neural IV), a learned spatial audio representation that encodes robust directional cues to enhance direction-of-arrival estimation, even in adverse acoustic scenarios with overlapping sources. To facilitate large-scale training and systematic evaluation, we propose SpatialSceneQA, a benchmark of 61k instruction-tuning samples curated from simulated physical environments. Extensive experiments demonstrate that our approach consistently surpasses 2D-centric baselines across diverse spatial perception and reasoning tasks, underscoring the necessity of explicit 3D modelling for advancing AI in physical environments. Our source code, pre-trained model checkpoints, and datasets are available at https://github.com/liuzhan22/JAEGER.

Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses

arXiv:2605.02900v2 Announce Type: replace-cross Abstract: Embodied Artificial Intelligence (Embodied AI) integrates perception, cognition, planning, and interaction into agents that operate in open-world, safety-critical environments. As these systems gain autonomy and enter domains such as transportation, healthcare, and industrial or assistive robotics, ensuring their safety becomes both technically challenging and socially indispensable. Unlike digital AI systems, embodied agents must act under uncertain sensing, incomplete knowledge, and dynamic human-robot interactions, where failures can directly lead to physical harm. This survey provides a comprehensive and structured review of safety research in embodied AI, examining attacks and defenses across the full embodied pipeline, from perception and cognition to planning, action and interaction, and agentic system. We introduce a multi-level taxonomy that unifies fragmented lines of work and connects embodied-specific safety findings with broader advances in vision, language, and multimodal foundation models. Our review synthesizes insights from over 500 papers spanning adversarial, backdoor, jailbreak, and hardware-level attacks; attack detection, safe training and robust inference; and risk-aware human-agent interaction. This analysis reveals several overlooked challenges, including the fragility of multimodal perception fusion, the instability of planning under jailbreak attacks, and the trustworthiness of human-agent interaction in open-ended scenarios. By organizing the field into a coherent framework and identifying critical research gaps, this survey provides a roadmap for building embodied agents that are not only capable and autonomous but also safe, robust, and reliable in real-world deployment.

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models

arXiv:2605.12374v4 Announce Type: replace-cross Abstract: Visual latent reasoning lets a multimodal large language model (MLLM) create intermediate visual evidence as continuous tokens, avoiding external tools or image generators. However, existing methods usually follow an output-as-input latent paradigm and yield unstable gains. We identify evidence for a feature-space mismatch that can contribute to this instability: dominant visual-latent models build on pre-norm MLLMs and reuse decoder hidden states as predicted latent inputs, even though these states occupy a substantially different norm regime from the input embeddings the model was trained to consume (Xie et al., 2025; Li et al., 2026; Team et al., 2026). This mismatch can make direct latent feedback unreliable. Motivated by this diagnosis, we propose GAP, a Granular Alignment Paradigm for visual latent modeling. GAP aligns visual latent reasoning at three levels: feature-level alignment maps decoder outputs into input-compatible visual latents through a lightweight PCA-aligned latent head; context-level alignment grounds latent targets with inspectable auxiliary visual supervision; and capacity-guided alignment assigns latent supervision selectively to examples where the base MLLM struggles. On Qwen2.5-VL 7B, the resulting model achieves the best mean aggregate perception and reasoning performance among our supervised variants. Inference-time intervention probing further suggests that generated latents provide task-relevant visual signal beyond merely adding token slots.

Personalized neoadjuvant treatment regimen selection in locally advanced rectal cancer based on regimen-specific response modeling

npj Digital Medicine, Published online: 26 May 2026; doi:10.1038/s41746-026-02798-w

Personalized neoadjuvant treatment regimen selection in locally advanced rectal cancer based on regimen-specific response modeling

A Learning-Based Cooperative Coevolution Framework for Heterogeneous Large-Scale Global Optimization

arXiv:2604.01241v1 Announce Type: cross Abstract: Cooperative Coevolution (CC) effectively addresses Large-Scale Global Optimization (LSGO) via decomposition but struggles with the emerging class of Heterogeneous LSGO (H-LSGO) problems arising from real-world applications, where subproblems exhibit diverse dimensions and distinct landscapes. The prevailing CC paradigm, relying on a fixed low-dimensional optimizer, often fails to navigate this heterogeneity. To address this limitation, we propose the Learning-Based Heterogeneous Cooperative Coevolution Framework (LH-CC). By formulating the optimization process as a Markov Decision Process, LH-CC employs a meta-agent to adaptively select the most suitable optimizer for each subproblem. We also introduce a flexible benchmark suite to generate diverse H-LSGO problem instances. Extensive experiments on 3000-dimensional problems with complex coupling relationships demonstrate that LH-CC achieves superior solution quality and computational efficiency compared to state-of-the-art baselines. Furthermore, the framework exhibits robust generalization across varying problem instances, optimization horizons, and optimizers. Our findings reveal that dynamic optimizer selection is a pivotal strategy for solving complex H-LSGO problems.

ExpertFlow: Efficient Mixture-of-Experts Inference via Predictive Expert Caching and Token Scheduling

arXiv:2410.17954v2 Announce Type: replace Abstract: Sparse Mixture-of-Experts (MoE) models can outperform dense large language models at similar computation by activating only a small set of experts per token. However, stacking many expert modules introduces substantial parameter memory, which makes MoE models difficult to deploy in memory-constrained environments such as single-GPU devices. Offloading alleviates this issue by storing inactive experts in CPU memory and loading them on demand, but existing methods remain limited: static caches disregard input-dependent routing, and methods that train separate models to predict expert usage ahead of time are often inaccurate or require significant training cost. We propose ExpertFlow, a lightweight MoE inference system that addresses this routing dependency through three coordinated components: 1) a transformer-based routing path predictor that estimates expert usage across all MoE layers in a single forward pass, 2) a token scheduler that groups tokens with similar predicted routes to improve expert utilization, and 3) a predictive expert cache that loads only the required experts while correcting mispredictions at runtime. Together, these components enable efficient expert loading and execution, reducing GPU memory usage by up to 93.72% and improving inference throughput by up to 10x over strong offloading baselines on a single GPU.

Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills

arXiv:2603.25158v3 Announce Type: replace Abstract: Equipping Large Language Model (LLM) agents with domain-specific skills is critical for tackling complex tasks. Yet, manual authoring creates a severe scalability bottleneck. Conversely, automated skill generation often yields fragile or fragmented results because it either relies on shallow parametric knowledge or sequentially overfits to non-generalizable trajectory-local lessons. To overcome this, we introduce Trace2Skill, a framework that mirrors how human experts author skills: by holistically analyzing broad execution experience before distilling it into a single, comprehensive guide. Instead of reacting sequentially to individual trajectories, Trace2Skill dispatches a parallel fleet of sub-agents to analyze a diverse pool of executions. It extracts trajectory-specific lessons and hierarchically consolidates them into a unified, conflict-free skill directory via inductive reasoning. Trace2Skill supports both deepening existing human-written skills and creating new ones from scratch. Experiments in challenging domains, such as spreadsheet, VisionQA and math reasoning, show that Trace2Skill significantly improves upon strong baselines, including Anthropic's official xlsx skills. Crucially, this trajectory-grounded evolution does not merely memorize task instances or model-specific quirks: evolved skills transfer across LLM scales and generalize to OOD settings. For example, skills evolved by Qwen3.5-35B on its own trajectories improved a Qwen3.5-122B agent by up to 57.65 absolute percentage points on WikiTableQuestions. Ultimately, our results demonstrate that complex agent experience can be packaged into highly transferable, declarative skills -- requiring no parameter updates, no external retrieval modules, and utilizing open-source models as small as 35B parameters.
❌