❌

Reading view

KT4EQG: Personalized Exercise Question Generation via Knowledge Tracing

arXiv:2605.23933v1 Announce Type: cross Abstract: Educational Question Generation (EQG) aims to synthesize customized exercise questions that enhance student learning. An effective EQG system should ideally personalize questions for each student by modeling the student's knowledge state and generating questions that provide the greatest learning benefit. However, few existing EQG approaches are able to achieve such fine-grained personalization. In this paper, we explore how EQG can benefit from knowledge tracing (KT), which models students' knowledge states based on historical performance and predicts future performance. We propose KT4EQG, a personalized EQG framework that generates effective questions for individual students under the guidance of a KT model. Specifically, KT4EQG seeks to maximize a student's potential improvement in overall knowledge mastery by leveraging the KT model to select the most suitable knowledge concept for the student to practice. An LLM-based question generator is then trained to produce a question faithfully grounded in the selected concept. Experimental results on XES3G5M and MOOCRadar show that KT4EQG consistently generates more effective questions than methods with limited or no personalization.
  •  

DBPnet: Damper Characteristics-Based Bayesian Physics-Informed Neural Network for Wheel Load Estimation

arXiv:2605.24860v1 Announce Type: cross Abstract: Advanced driver assistance systems (ADAS) play an important role in modern automotive intelligence, significantly enhancing vehicle safety and stability. The performance of ADAS critically relies on accurate and reliable vehicle state estimation, particularly from vehicle dynamic sensors. Among these signals, wheel load is a key variable for chassis control and safety-critical functions, yet it remains difficult to estimate robustly due to complex suspension geometry, nonlinear dynamics, and measurement noise. To address this issue, we propose DBPnet, a Bayesian physics-informed neural network (PINN) with a physics-aware embedding module inspired by damper characteristics. First, this paper presents a suspension linkage-level modeling (SLLM) approach that constructs a nonlinear instantaneous dynamic model by explicitly considering the complex geometric structure of the suspension. Building upon SLLM, Bayesian inference is integrated into the PINN to effectively cope with noise and uncertainty in the vehicle chassis system, thereby improving the model's robustness. Then, a physics-informed loss function is employed to ensure consistency with fundamental physical principles, while the damper characteristics-inspired embedding module extracts temporal variation features of input signals and incorporates them into each layer of the PINN, ensuring that physical observations guide the neural network without being constrained by fixed physical models. Extensive evaluations on high-fidelity simulations and real-world experiments demonstrate that our DBPnet consistently achieves lower RMSE and MaxError than baseline methods. These results highlight the potential of our DBPnet to advance wheel load estimation and contribute to the development of more reliable ADAS actuator functions.
  •  

Hide to Guide: Learning via Semantic Masking

arXiv:2605.25198v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a powerful paradigm for improving language models on reasoning-intensive tasks, but its effectiveness is often limited by exploration. For example, models often fail on hard problems, leaving little useful reward signal. External expert traces offer a natural source of guidance, yet they may also expose reward-relevant content along the critical path to the verifier target, such as final answers, intermediate values, executable implementations, or answer-related entities. This content can create an unintended reward hacking channel, allowing the policy to obtain reward by copying the trace rather than learning the underlying reasoning or agentic behavior. Existing guided-RL methods reduce this risk by using partial trajectories, but they mainly control how much expert information is shown heuristically rather than which parts should be hidden. To this end, we propose Semantic Masked Expert Policy Optimization (SMEPO), a fine-grained semantic masking strategy for expert-guided RLVR. Instead of truncating traces coarsely or revealing them unchanged, SMEPO masks reward-relevant semantic spans along the critical path while preserving the expert's decomposition, plan, and procedural structure. This turns hard problems from reasoning from scratch into a fill-in-the-blank process: the policy can follow the expert's problem-solving route, but must still reconstruct the missing values, code, or entities by itself. SMEPO is simple to apply and requires no changes to the reward function or RL objective. Across diverse domains, including math, code, and agentic search, SMEPO improves accuracy by up to 3.2 points over GRPO and reduces training time by up to 4.2x. The code is available at https://github.com/mit-han-lab/SMEPO.
  •  

Benchmarking Pathology Foundation Models for Spatial Domain Understanding

arXiv:2605.25764v1 Announce Type: cross Abstract: Pathology foundation models (PFMs) have emerged as a core approach for learning transferable representations from whole slide images (WSIs), and they are typically benchmarked through downstream clinical endpoints. While such task level evaluations are indispensable, they offer limited insight into what the representations themselves encode, particularly whether PFM embeddings can distinguish meaningful tissue regions and capture their spatial relationships. We present SpaPath-Bench, a representation level benchmark designed to diagnose spatial representation capability in PFMs. SpaPath-Bench formulates spatial domain identification (SDI) on paired whole slide image and spatial transcriptomics (ST) data as a diagnostic task. It curates 42 public paired WSI and ST slides, enables large scale evaluation across 19 encoders and seven SDI methods, and measures partition quality using three complementary criteria: unsupervised spatial coherence, transcriptomics referenced agreement, and expert referenced agreement. Across 83K runs, SpaPath-Bench reveals that different pretraining paradigms capture distinct aspects of tissue spatial architecture, and it provides practical guidance for building the next generation of spatially aware computational pathology models. Code and data pipelines are publicly available at https://bokai-zhao.github.io/SpaPath-benchboard/.
  •  

Topology-Driven Transferability Estimation of Medical Foundation Models for Segmentation

arXiv:2602.23916v2 Announce Type: replace-cross Abstract: The advent of large-scale self-supervised learning (SSL) has produced a vast zoo of medical foundation models. However, selecting optimal medical foundation models for specific segmentation tasks remains a computational bottleneck. Existing Transferability Estimation (TE) metrics, primarily designed for classification, rely on global statistical assumptions and fail to capture the topological complexity essential for dense prediction. We propose a novel Topology-Driven Transferability Estimation framework that evaluates manifold tractability rather than statistical overlap. Our approach introduces three components: (1) Global Representation Topology Divergence (GRTD), utilizing Minimum Spanning Trees to quantify feature-label structural isomorphism; (2) Local Boundary-Aware Topological Consistency (LBTC), which assesses manifold separability specifically at critical anatomical boundaries; and (3) Task-Adaptive Fusion, which dynamically integrates global and local metrics based on the semantic cardinality of the target task. Validated on the large-scale OpenMind benchmark across diverse anatomical targets and SSL foundation models, our approach significantly outperforms state-of-the-art baselines by around 31% relative improvement in the weighted Kendall metric, providing a robust, training-free proxy for efficient model selection without the cost of fine-tuning. The code will be made publicly available upon acceptance.
  •  

Integrative multi-omics and experimental validation reveal UBE2C as a central hub gene and prognostic biomarker in hepatocellular carcinoma

Int Immunopharmacol. 2026 May 19;183:116866. doi: 10.1016/j.intimp.2026.116866. Online ahead of print.

ABSTRACT

Hepatocellular carcinoma (HCC) is a lethal malignancy with a high recurrence rate and limited treatment options. Ubiquitin-conjugating enzyme E2 C (UBE2C) is implicated in various cancers, yet its impact on the HCC immune landscape remains incompletely understood. Herein, hub genes in HCC were identified, by integrating co-expression networks and protein-protein interaction analyses, from the TCGA, GEO, and CPTAC databases. Their expression was analysed using a single-cell transcriptomic database and verified in HCC tissues and cell lines via quantitative reverse transcription-PCR and immunoblotting. Functional roles of UBE2C were assessed using in vitro knockdown experiments and an in vivo subcutaneous tumour model. The tumour immune microenvironment was profiled using spatial transcriptomics, RNA-seq data, and ssGSEA. A prognostic nomogram was constructed based on multivariate Cox regression. UBE2C was identified as a significantly upregulated hub gene in HCC. Single-cell RNA-seq revealed predominant expression of UBE2C in hepatocytes, with dynamic upregulation along differentiation trajectories. UBE2C knockdown suppressed proliferation, induced apoptosis, and inhibited tumour growth. Spatial transcriptomics highlighted UBE2C-high regions within proliferative niches exhibiting immunosuppressive traits-including TGFB1 enrichment, impaired CXCL9-CXCR3 signalling, and exclusion of cytotoxic T cells-which were reduced in immunotherapy responders. UBE2C expression correlated with immune checkpoint genes and specific immune cell subsets. A UBE2C-based nomogram integrating T stage and tumour stage robustly predicted patient survival, and miR-300 and miR-381-3p were identified as potential upstream regulators. These findings establish UBE2C as a key driver of HCC progression and a biomarker for prognosis and immunotherapy stratification.

PMID:42155390 | DOI:10.1016/j.intimp.2026.116866

  •  

UBTF-HSP90A-MIF stress circuit drives lenvatinib resistance and immune exclusion in hepatocellular carcinoma

J Adv Res. 2026 Apr 5:S2090-1232(26)00280-8. doi: 10.1016/j.jare.2026.04.002. Online ahead of print.

ABSTRACT

INTRODUCTION: The clinical benefit of combining lenvatinib with PD-1 blockade in HCC is frequently constrained by adaptive resistance and the development of an immune-cold tumor microenvironment.

OBJECTIVES: This study aimed to elucidate the molecular mechanisms underlying adaptive resistance and immune exclusion during lenvatinib-PD-1 therapy in HCC, with a particular focus on a UBTF/HSP90A/MIF regulatory circuit. We examined whether genetic or pharmacologic targeting of macrophage migration inhibitory factor (MIF) could restore lenvatinib sensitivity, remodel the tumor immune microenvironment, and serve as a predictive biomarker in clinical cohorts.

METHODS: Paired lenvatinib-sensitive and -resistant HCC models were interrogated using integrated multi-omic and functional approaches, including RNA sequencing, promoter pull-down assays, ChIP, luciferase reporter assays, PLA, and flow cytometry. Key findings were validated in patient-derived organoids and xenografts, as well as in an immunocompetent hydrodynamic HCC mouse model. Clinical relevance was evaluated in independent cohorts treated with lenvatinib plus anti-PD-1 therapy.

RESULTS: UBTF directly bound to and transcriptionally activated the HSP90A promoter, resulting in increased HSP90A expression and stabilization of MIF. MIF signaling through CD74 co-activated the PI3K-AKT and MAPK pathways, sustaining tumor cell proliferation under lenvatinib pressure. Single-cell RNA sequencing and multiplex immunohistochemistry revealed macrophage enrichment and CD8+ T-cell exclusion in resistant tumors. Genetic ablation of Mif (Alb-Cre; Mifflox/flox) or pharmacologic inhibition with 4-IPP (4-Iodo-6-phenylpyrimidine) restored lenvatinib sensitivity, reprogrammed the tumor immune microenvironment, and, when combined with PD-1 blockade, achieved superior tumor control and prolonged survival. In clinical datasets, low pretreatment MIF expression was associated with improved responses to lenvatinib plus PD-1 therapy.

CONCLUSIONS: These findings define a UBTF/HSP90A/MIF axis linking proteostasis and cytokine signaling to immune-metabolic dysfunction and lenvatinib resistance in HCC. MIF emerges as both a mechanistic driver and a predictive biomarker, supporting prospective evaluation of therapeutic strategies combining lenvatinib-PD-1 with MIF- or HSP90A-targeted interventions to personalize TKI-ICI therapy.

PMID:41946392 | DOI:10.1016/j.jare.2026.04.002

  •  

Protective Effects of the Ethyl Acetate Fraction from Madeng'ai on Lipopolysaccharide-Induced Acute Lung Injury in Mice: Insights from Integrated Multi-Omics Analysis

J Ethnopharmacol. 2026 Apr 4:121650. doi: 10.1016/j.jep.2026.121650. Online ahead of print.

ABSTRACT

ETHNOPHARMACOLOGICAL RELEVANCE: Madeng'ai (MDA) is a traditional medicinal plant of the Dong ethnic group. Its roots have been widely used in folk medicine for clearing heat and removing toxins, alleviating swelling and relieving pain, dispersing blood stasis and arresting bleeding, as well as promoting wound healing. It is taxonomically classified as a variety of Potentilla freyniana Bornm.

AIM OF THE STUDY: Acute lung injury (ALI) is a life-threatening pulmonary disorder associated with high mortality, underscoring the urgent need to explore novel therapeutic strategies. This study aimed to evaluate the protective effects of the ethyl acetate fraction of MDA (MEA) against LPS-induced ALI in mice and to investigate its underlying mechanisms.

MATERIALS AND METHODS: LC-MS/MS was employed to tentatively identify the bioactive components of MEA. A mouse model of ALI was established by LPS induction. The protective effects of MEA were evaluated through assessments of lung histopathology, inflammatory cytokine levels, and oxidative stress markers. The underlying mechanisms were systematically investigated by integrating transcriptomics, metabolomics, network pharmacology, molecular docking, and Western blotting.

RESULTS: MEA significantly attenuated LPS-induced pulmonary pathological lesions, pulmonary edema, and excessive inflammatory responses in ALI mice. Comprehensive bioinformatics analyses predicted potential mechanisms involving oxidative stress and the regulation of metabolic pathways. Experimental validation via Western blotting confirmed that MEA inhibited TLR4-mediated inflammatory signaling and modulated the PI3K/AKT pathway, thereby exerting multi-pathway protective effects against ALI.

CONCLUSIONS: Collectively, this study confirms that MEA, as a traditional herbal extract, holds potential as an adjuvant therapeutic agent for ALI, providing experimental evidence for the modernization and development of ethnic medicines.

PMID:41941987 | DOI:10.1016/j.jep.2026.121650

  •  

Protective Effects of the Ethyl Acetate Fraction from Madeng'ai on Lipopolysaccharide-Induced Acute Lung Injury in Mice: Insights from Integrated Multi-Omics Analysis

J Ethnopharmacol. 2026 Apr 4:121650. doi: 10.1016/j.jep.2026.121650. Online ahead of print.

ABSTRACT

ETHNOPHARMACOLOGICAL RELEVANCE: Madeng'ai (MDA) is a traditional medicinal plant of the Dong ethnic group. Its roots have been widely used in folk medicine for clearing heat and removing toxins, alleviating swelling and relieving pain, dispersing blood stasis and arresting bleeding, as well as promoting wound healing. It is taxonomically classified as a variety of Potentilla freyniana Bornm.

AIM OF THE STUDY: Acute lung injury (ALI) is a life-threatening pulmonary disorder associated with high mortality, underscoring the urgent need to explore novel therapeutic strategies. This study aimed to evaluate the protective effects of the ethyl acetate fraction of MDA (MEA) against LPS-induced ALI in mice and to investigate its underlying mechanisms.

MATERIALS AND METHODS: LC-MS/MS was employed to tentatively identify the bioactive components of MEA. A mouse model of ALI was established by LPS induction. The protective effects of MEA were evaluated through assessments of lung histopathology, inflammatory cytokine levels, and oxidative stress markers. The underlying mechanisms were systematically investigated by integrating transcriptomics, metabolomics, network pharmacology, molecular docking, and Western blotting.

RESULTS: MEA significantly attenuated LPS-induced pulmonary pathological lesions, pulmonary edema, and excessive inflammatory responses in ALI mice. Comprehensive bioinformatics analyses predicted potential mechanisms involving oxidative stress and the regulation of metabolic pathways. Experimental validation via Western blotting confirmed that MEA inhibited TLR4-mediated inflammatory signaling and modulated the PI3K/AKT pathway, thereby exerting multi-pathway protective effects against ALI.

CONCLUSIONS: Collectively, this study confirms that MEA, as a traditional herbal extract, holds potential as an adjuvant therapeutic agent for ALI, providing experimental evidence for the modernization and development of ethnic medicines.

PMID:41941987 | DOI:10.1016/j.jep.2026.121650

  •  

Scale over Preference: The Impact of AI-Generated Content on Online Content Ecology

arXiv:2604.01690v1 Announce Type: new Abstract: The rapid proliferation of Artificial Intelligence-Generated Content (AIGC) is fundamentally restructuring online content ecologies, necessitating a rigorous examination of its behavioral and distributional implications. Leveraging a comprehensive longitudinal dataset comprising tens of millions of users from a leading Chinese video-sharing platform, this study elucidated the distinct creation and consumption behaviors characterizing AIGC versus Human-Generated Content (HGC). We identified a prevalent scale-over-preference dynamic, wherein AIGC creators achieve aggregate engagement comparable to HGC creators through high-volume production, despite a marked consumer preference for HGC. Deeper analysis uncovered the ability of the algorithmic content distribution mechanism in moderating these competing interests regarding AIGC. These findings advocated for the implementation of AIGC-sensitive distribution algorithms and precise governance frameworks to ensure the long-term health of the online content platforms.
  •  

VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing

arXiv:2603.29852v1 Announce Type: cross Abstract: We introduce VectorGym, a comprehensive benchmark suite for Scalable Vector Graphics (SVG) that spans generation from text and sketches, complex editing, and visual understanding. VectorGym addresses the lack of realistic, challenging benchmarks aligned with professional design workflows. Our benchmark comprises four tasks with expert human-authored annotations: the novel Sketch2SVG task (VG-Sketch); a new SVG editing dataset (VG-Edit) featuring complex, multi-step edits with higher-order primitives; Text2SVG generation (VG-Text); and SVG captioning (VG-Cap). Unlike prior benchmarks that rely on synthetic edits, VectorGym provides gold-standard human annotations that require semantic understanding and design intent. We also propose a multi-task reinforcement learning approach that jointly optimizes across all four tasks using rendering-based rewards. Our method, built on GRPO with curriculum learning, trains a Qwen3-VL 8B model that achieves state-of-the-art performance among open-source models, surpassing much larger models including Qwen3-VL 235B and matching GPT-4o. We also introduce a VLM-as-a-Judge metric for SVG generation, validated through human correlation studies. Our evaluation of frontier VLMs reveals significant performance gaps, positioning VectorGym as a rigorous framework for advancing visual code generation. VectorGym is publicly available on huggingface.co/datasets/ServiceNow/VectorGym.
  •  

TRIM21-mediated degradation of HILPDA overcomes anti-PD-1 immunotherapy resistance in breast cancer by limiting PD-L1 palmitoylation

Oncogene, Published online: 24 March 2026; doi:10.1038/s41388-026-03728-6

TRIM21-mediated degradation of HILPDA overcomes anti-PD-1 immunotherapy resistance in breast cancer by limiting PD-L1 palmitoylation
  •  

OpenSage: Self-programming Agent Generation Engine

arXiv:2602.16891v2 Announce Type: replace Abstract: Agent development kits (ADKs) provide effective platforms and tooling for constructing agents, and their designs are critical to the constructed agents' performance, especially the functionality for agent topology, tools, and memory. However, current ADKs either lack sufficient functional support or rely on humans to manually design these components, limiting agents' generalizability and overall performance. We propose OpenSage, the first ADK that enables LLMs to automatically create agents with self-generated topology and toolsets while providing comprehensive and structured memory support. OpenSage offers effective functionality for agents to create and manage their own sub-agents and toolkits. It also features a hierarchical, graph-based memory system for efficient management and a specialized toolkit tailored to software engineering tasks. Extensive experiments across three state-of-the-art benchmarks with various backbone models demonstrate the advantages of OpenSage over existing ADKs. We also conduct rigorous ablation studies to demonstrate the effectiveness of our design for each component. We believe OpenSage can pave the way for the next generation of agent development, shifting the focus from human-centered to AI-centered paradigms.
  •  

HKDC1-Mediated Polyamine Rewiring Drives Lenvatinib Resistance and Immune Escape in Hepatocellular Carcinoma

Clin Mol Hepatol. 2026 Mar 11. doi: 10.3350/cmh.2025.1269. Online ahead of print.

ABSTRACT

BACKGROUND/AIMS: Lenvatinib resistance and immune exclusion limit outcomes in HCC. We hypothesized that metabolic rewiring orchestrates resistance to lenvatinib and PD-1 blockade.

METHODS: We established LS/LR HCC models and employed multi-omics (proteomics/RNA-seq), ChIP, luciferase, and RIP assays to map HKDC1 regulation. Tumor immunity was profiled by scRNA-seq, mIHC, and flow cytometry. SPD + lenvatinib efficacy was tested in cell lines, patient-derived organoids/xenografts. Tested therapy effect in an immunocompetent hydrodynamic HCC model with hepatocyte-specific Hkdc1 deletion; and analyzed a postoperative cohort (n = 40) treated with lenvatinib + PD-1.

RESULTS: HKDC1, upregulated in LR HCC, was transcriptionally activated by USF1 and promoted SMS-mediated polyamine rewiring. This impaired CD8⁺ T-cell metabolism, reversible by HKDC1 knockdown or spermidine (SPD). SPD synergized with lenvatinib, triggering autophagy and suppressing tumor growth in vitro and in vivo. High HKDC1 predicted poor response and survival in patients receiving lenvatinib + aPD-1.

CONCLUSIONS: A USF1/HKDC1/SMS axis couples polyamine metabolism to immune dysfunction and lenvatinib resistance. HKDC1 is a predictive biomarker and therapeutic node and support polyamine-axis modulation to sensitize HCC to lenvatinib plus PD-1 therapy.

PMID:41812646 | DOI:10.3350/cmh.2025.1269

  •  

Towards Efficient Federated Learning of Networked Mixture-of-Experts for Mobile Edge Computing

arXiv:2511.01743v2 Announce Type: replace-cross Abstract: Recent advancements in large artificial intelligence models (LAMs) are driving significant innovations in mobile edge computing within next-generation wireless networks. However, the substantial demands for computational resources and larges-cale training data required to train LAMs conflict with the limited storage and computational capacity of edge devices, posing significant challenges to training and deploying LAMs at the edge. In this work, we introduce the Networked Mixture-of-Experts (NMoE) system, in which clients perform inference collaboratively by distributing tasks to suitable neighbors based on their expertise and aggregate the returned results. For training the NMoE, we propose a federated learning framework that integrates both supervised and self-supervised learning to balance personalization and generalization, while preserving communication efficiency and data privacy. We conduct extensive experiments to demonstrate the efficacy of the proposed NMoE system, providing insights for the NMoE training algorithms.
  •  

CrystaL: Spontaneous Emergence of Visual Latents in MLLMs

arXiv:2602.20980v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable performance by integrating powerful language backbones with large-scale visual encoders. Among these, latent Chain-of-Thought (CoT) methods enable implicit reasoning in continuous hidden states, facilitating seamless vision-language integration and faster inference. However, existing heuristically predefined supervision signals in latent CoT provide limited guidance for preserving critical visual information in intermediate latent states. To address this limitation, we propose CrystaL (Crystallized Latent Reasoning), a single-stage framework with two paths to process intact and corrupted images, respectively. By explicitly aligning the attention patterns and prediction distributions across the two paths, CrystaL crystallizes latent representations into task-relevant visual semantics, without relying on auxiliary annotations or external modules. Extensive experiments on perception-intensive benchmarks demonstrate that CrystaL consistently outperforms state-of-the-art baselines, achieving substantial gains in fine-grained visual understanding while maintaining robust reasoning capabilities.
  •  

CubeComposer: Spatio-Temporal Autoregressive 4K 360{\deg} Video Generation from Perspective Video

arXiv:2603.04291v1 Announce Type: cross Abstract: Generating high-quality 360{\deg} panoramic videos from perspective input is one of the crucial applications for virtual reality (VR), whereby high-resolution videos are especially important for immersive experience. Existing methods are constrained by computational limitations of vanilla diffusion models, only supporting $\leq$ 1K resolution native generation and relying on suboptimal post super-resolution to increase resolution. We introduce CubeComposer, a novel spatio-temporal autoregressive diffusion model that natively generates 4K-resolution 360{\deg} videos. By decomposing videos into cubemap representations with six faces, CubeComposer autoregressively synthesizes content in a well-planned spatio-temporal order, reducing memory demands while enabling high-resolution output. Specifically, to address challenges in multi-dimensional autoregression, we propose: (1) a spatio-temporal autoregressive strategy that orchestrates 360{\deg} video generation across cube faces and time windows for coherent synthesis; (2) a cube face context management mechanism, equipped with a sparse context attention design to improve efficiency; and (3) continuity-aware techniques, including cube-aware positional encoding, padding, and blending to eliminate boundary seams. Extensive experiments on benchmark datasets demonstrate that CubeComposer outperforms state-of-the-art methods in native resolution and visual quality, supporting practical VR application scenarios. Project page: https://lg-li.github.io/project/cubecomposer
  •  

R1-Code-Interpreter: LLMs Reason with Code via Supervised and Multi-stage Reinforcement Learning

arXiv:2505.21668v3 Announce Type: replace Abstract: Practical guidance on training Large Language Models (LLMs) to leverage Code Interpreter across diverse tasks remains lacking. We present R1-Code-Interpreter, an extension of a text-only LLM trained via multi-turn supervised fine-tuning (SFT) and reinforcement learning (RL) to autonomously generate multiple code queries during step-by-step reasoning. Unlike prior RL + tool-use efforts focused on narrow domains such as math or retrieval, we curate 144 diverse reasoning and planning tasks and show that training a general-purpose Code Interpreter across them presents significant challenges due to task heterogeneity and scarcity of effective samples. To address this, we introduce a multi-stage curriculum learning approach that partitions training samples by measured improvement potential. The RL training prioritizes samples with higher potential and gradually shifts to lower-potential ones, increasing the average RL gains from merely +3.4% to +9.3% across Qwen-2.5 models (3/7/14B). Our final model, R1-CI-14B, improves average accuracy on the 37 test tasks from 44.1% to 72.4%, outperforming text-only GPT-4o (58.6%) and GPT-4o with Code Interpreter (70.9%). Notably, R1-CI-14B also exhibits emergent self-checking behavior through code generation. Datasets, Codes, and Models are available at https://github.com/yongchao98/R1-Code-Interpreter and https://huggingface.co/yongchao98.
  •  

Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents

arXiv:2510.24702v2 Announce Type: replace-cross Abstract: Public research results on large-scale supervised finetuning of AI agents remain relatively rare, since the collection of agent training data presents unique challenges. In this work, we argue that the bottleneck is not a lack of underlying data sources, but that a large variety of data is fragmented across heterogeneous formats, tools, and interfaces. To this end, we introduce the agent data protocol (ADP), a light-weight representation language that serves as an "interlingua" between agent datasets in diverse formats and unified agent training pipelines downstream. The design of ADP is expressive enough to capture a large variety of tasks, including API/tool use, browsing, coding, software engineering, and general agentic workflows, while remaining simple to parse and train on without engineering at a per-dataset level. In experiments, we unified a broad collection of 13 existing agent training datasets into ADP format, and converted the standardized ADP data into training-ready formats for multiple agent frameworks. We performed SFT on these data, and demonstrated an average performance gain of ~20% over corresponding base models, and delivers state-of-the-art or near-SOTA performance on standard coding, browsing, tool use, and research benchmarks, without domain-specific tuning. All code and data are released publicly, in the hope that ADP could help lower the barrier to standardized, scalable, and reproducible agent training.
  •  
❌