❌

Reading view

Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World

arXiv:2605.26086v1 Announce Type: new Abstract: Large language model agents are increasingly envisioned as always-on personal assistants with access to anything relevant in the user's digital world. Yet current systems operate over only narrow slices of that world, limiting context-sensitive reasoning and effective assistance. Existing benchmarks similarly provide only partial user state and therefore fail to capture performance in such a broad, always-on setting. To address this gap, we introduce Claw-Anything, a benchmark that expands agent context along three dimensions: long-horizon activity histories, interdependent backend services, and integrated GUI and CLI interaction across multiple devices. To instantiate this setting, we simulate months of user activity through multi-round event injection, producing complex world states and realistic noise, including irrelevant events and conflicting signals. Agents must reason over rich contextual environments while remaining robust to such noise. This expanded scope also enables the evaluation of proactive assistance, requiring agents to anticipate user needs and deliver timely recommendations. Experiments show that GPT-5.5 achieves only 34.5% pass@1, substantially below prior benchmarks, underscoring a gap between current agent capabilities and the demands of always-on personal assistance. Alongside the benchmark, we release an automated data-generation pipeline that yields 2,000 training environments and improves the base model by 23.7%, demonstrating its utility of scalable data infrastructure.
  •  

MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research

arXiv:2605.26114v1 Announce Type: new Abstract: We present MobileGym, a browser-hosted, lightweight, fully controllable environment for everyday mobile use, targeting interaction fidelity without replicating proprietary backends. It enables two capabilities previously out of reach for everyday apps: verifiable outcome signals through deterministic state-based judging over structured JSON state, and scalable online RL through low-cost parallel rollouts. The full environment state is captured, configured, forked, and compared as structured JSON, and a single server can host hundreds of parallel instances, with about 400 MB memory per instance and about 3 s cold start. A layered state model and a declarative task-definition framework keep state programmability and task creation practical at scale, and a single programmatic judging mechanism delivers both deterministic evaluation verdicts and dense RL rewards. The accompanying MobileGym-Bench provides 416 parameterized task templates, including 256 test and 160 train templates, over 28 apps, with deterministic judges and a structured AnswerSheet protocol that avoids free-text matching failures. In a Sim-to-Real case study, GRPO on Qwen3-VL-4B-Instruct gains +12.8 percentage points on the 256-task test set, and on a 59-task real-device signal subset, real-device execution retains 95.1% of the simulation-side training gain. Project page: https://mobilegym.github.io.
  •  

Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning

arXiv:2602.10090v3 Announce Type: replace Abstract: Recent advances in large language model (LLM) have empowered autonomous agents to perform multi-turn interactions with tools and environments. However, scaling such agent training is limited by the lack of diverse and reliable environments. In this paper, we propose Agent World Model (AWM), a fully synthetic environment generation pipeline. Using this pipeline, we scale to 1,000 environments covering everyday scenarios, in which agents can interact with rich toolsets and obtain high-quality observations. Notably, these environments are code-driven and backed by databases, providing more reliable and consistent state transitions than environments simulated by LLMs. Moreover, they enable more efficient agent interaction compared with collecting trajectories from realistic environments. To demonstrate the effectiveness of this resource, we perform large-scale reinforcement learning for multi-turn tool-use agents. Thanks to the fully executable environments and accessible database states, we can also design reliable reward functions. Experiments on three benchmarks show that training exclusively in synthetic environments, rather than benchmark-specific ones, yields strong out-of-distribution generalization. The code is available at https://github.com/Snowflake-Labs/agent-world-model.
  •  

HiTeC: Hierarchical Contrastive Learning on Text-Attributed Hypergraph with Semantic-Aware Augmentation

arXiv:2508.03104v3 Announce Type: replace-cross Abstract: Contrastive learning (CL) has become a dominant paradigm for self-supervised hypergraph learning, enabling effective training without costly labels. However, node entities in real-world hypergraphs are often associated with rich textual information, which has been largely ignored in prior works. Directly applying existing CL-based methods to such text-attributed hypergraphs (TAHGs) leads to three key limitations: (1) The common use of graph-agnostic text encoders fails to capture the correlations between textual semantics and hypergraph topology, resulting in less expressive representations. (2) Their reliance on random data augmentations introduces noise and weakens the contrastive signals. (3) The primary focus on node- and hyperedge-level contrastive signals limits the ability to capture long-range dependencies, which is essential for effective representation learning. To address these challenges, we introduce HiTeC, a two-stage hierarchical contrastive learning framework for effective self-supervised learning on TAHGs. In the first stage, we pre-train the text encoder with a structure-aware contrastive objective to overcome the graph-agnostic nature of conventional methods. In the second stage, we begin by introducing semantic-aware augmentations, including structure-contextualized text augmentation and semantic-aware hyperedge dropping, to facilitate informative view generation. Subsequently, we propose a multi-scale contrastive loss with an $s$-walk-based subgraph-level objective to capture long-range dependencies. Extensive experiments on six real-world datasets validate the effectiveness of our proposed method.
  •  

Multi-omics analysis identified serum B4GALT1 as a prognostic factor for small cell lung cancer

J Thorac Dis. 2026 Apr 30;18(4):353. doi: 10.21037/jtd-2025-1-2610. Epub 2026 Mar 20.

ABSTRACT

BACKGROUND: Small cell lung cancer (SCLC) is an aggressive neuroendocrine tumor characterized by rapid progression, early metastasis, and high mortality, with limited effective long-term treatment options. B4GALT1, a β-1,4-galactosyltransferase, has been implicated in the malignant progression of various cancers, but its specific role and underlying mechanisms in SCLC remain largely unexplored. We conducted a multi-omics analysis and clinical sample study to explore the function of B4GALT1 in SCLC.

METHODS: This study comprehensively investigated the expression pattern, functional significance, and clinical relevance of B4GALT1 in SCLC. We conducted multi-omics analyses, including single-cell data processing, InferCNV analysis, and immune infiltration analysis, to explore the association between B4GALT1 and the immune microenvironment of SCLC and patient survival. To determine B4GALT1 as a potential circulating biomarker, quantitative data-independent acquisition (DIA) proteomics analysis was performed on serum samples from SCLC patients and healthy controls. Enzyme-linked immunosorbent assay (ELISA) was used to further verify the differential expression of serum B4GALT1 in a larger cohort of SCLC patients, to evaluate its diagnostic, prognostic, and treatment response predictive value.

RESULTS: Multi-omics analysis revealed that B4GALT1 expression was significantly associated with patient survival. The expression of B4GALT1 positively correlated with macrophage infiltration in the tumor and negatively correlated with CD4+ T cells in the tumor. There was a negative correlation in inactivated naïve B cells, eosinophils, and CD4 naïve T cells, while it showed a positive correlation in dendritic cells, M0/M1/M2 macrophages, natural killer (NK) cells, CD8 T cells, follicular helper T cells, and regulatory T cells. ELISA results showed that serum protein B4GALT1 expression was higher in patients with SCLC than in healthy controls. Elevated serum B4GALT1 protein levels correlated with poor treatment outcomes in patients with SCLC undergoing chemoradiotherapy.

CONCLUSIONS: Our findings establish B4GALT1 as a critical prognostic, diagnostic, and predictive biomarker in SCLC, with its expression closely linked to the tumor immune microenvironment and treatment response. Targeting B4GALT1 or its related pathways may represent a novel therapeutic strategy, and serum B4GALT1 holds promise as a liquid biopsy marker for SCLC patient stratification, monitoring, and guiding treatment decisions.

PMID:42182735 | PMC:PMC13190155 | DOI:10.21037/jtd-2025-1-2610

  •  

Multi-omics analysis identified serum B4GALT1 as a prognostic factor for small cell lung cancer

J Thorac Dis. 2026 Apr 30;18(4):353. doi: 10.21037/jtd-2025-1-2610. Epub 2026 Mar 20.

ABSTRACT

BACKGROUND: Small cell lung cancer (SCLC) is an aggressive neuroendocrine tumor characterized by rapid progression, early metastasis, and high mortality, with limited effective long-term treatment options. B4GALT1, a β-1,4-galactosyltransferase, has been implicated in the malignant progression of various cancers, but its specific role and underlying mechanisms in SCLC remain largely unexplored. We conducted a multi-omics analysis and clinical sample study to explore the function of B4GALT1 in SCLC.

METHODS: This study comprehensively investigated the expression pattern, functional significance, and clinical relevance of B4GALT1 in SCLC. We conducted multi-omics analyses, including single-cell data processing, InferCNV analysis, and immune infiltration analysis, to explore the association between B4GALT1 and the immune microenvironment of SCLC and patient survival. To determine B4GALT1 as a potential circulating biomarker, quantitative data-independent acquisition (DIA) proteomics analysis was performed on serum samples from SCLC patients and healthy controls. Enzyme-linked immunosorbent assay (ELISA) was used to further verify the differential expression of serum B4GALT1 in a larger cohort of SCLC patients, to evaluate its diagnostic, prognostic, and treatment response predictive value.

RESULTS: Multi-omics analysis revealed that B4GALT1 expression was significantly associated with patient survival. The expression of B4GALT1 positively correlated with macrophage infiltration in the tumor and negatively correlated with CD4+ T cells in the tumor. There was a negative correlation in inactivated naïve B cells, eosinophils, and CD4 naïve T cells, while it showed a positive correlation in dendritic cells, M0/M1/M2 macrophages, natural killer (NK) cells, CD8 T cells, follicular helper T cells, and regulatory T cells. ELISA results showed that serum protein B4GALT1 expression was higher in patients with SCLC than in healthy controls. Elevated serum B4GALT1 protein levels correlated with poor treatment outcomes in patients with SCLC undergoing chemoradiotherapy.

CONCLUSIONS: Our findings establish B4GALT1 as a critical prognostic, diagnostic, and predictive biomarker in SCLC, with its expression closely linked to the tumor immune microenvironment and treatment response. Targeting B4GALT1 or its related pathways may represent a novel therapeutic strategy, and serum B4GALT1 holds promise as a liquid biopsy marker for SCLC patient stratification, monitoring, and guiding treatment decisions.

PMID:42182735 | PMC:PMC13190155 | DOI:10.21037/jtd-2025-1-2610

  •  

RDFace: A Benchmark Dataset for Rare Disease Facial Image Analysis under Extreme Data Scarcity and Phenotype-Aware Synthetic Generation

arXiv:2604.03454v1 Announce Type: cross Abstract: Rare diseases often manifest with distinctive facial phenotypes in children, offering valuable diagnostic cues for clinicians and AI-assisted screening systems. However, progress in this field is severely limited by the scarcity of curated, ethically sourced facial data and the high similarity among phenotypes across different conditions. To address these challenges, we introduce RDFace, a curated benchmark dataset comprising 456 pediatric facial images spanning 103 rare genetic conditions (average 4.4 samples per condition). Each ethically verified image is paired with standardized metadata. RDFace enables the development and evaluation of data-efficient AI models for rare disease diagnosis under real-world low-data constraints. We benchmark multiple pretrained vision backbones using cross-validation and explore synthetic augmentation with DreamBooth and FastGAN. Generated images are filtered via facial landmark similarity to maintain phenotype fidelity and merged with real data, improving diagnostic accuracy by up to 13.7% in ultra-low-data regimes. To assess semantic validity, phenotype descriptions generated by a vision-language model from real and synthetic images achieve a report similarity score of 0.84. RDFace establishes a transparent, benchmark-ready dataset for equitable rare disease AI research and presents a scalable framework for evaluating both diagnostic performance and the integrity of synthetic medical imagery.
  •  

TSPO: Breaking the Double Homogenization Dilemma in Multi-turn Search Policy Optimization

arXiv:2601.22776v2 Announce Type: replace Abstract: Multi-turn tool-integrated reasoning enables Large Language Models (LLMs) to solve complex tasks through iterative information retrieval. However, current reinforcement learning (RL) frameworks for search-augmented reasoning predominantly rely on sparse outcome-level rewards, leading to a "Double Homogenization Dilemma." This manifests as (1) Process homogenization, where the thinking, reasoning, and tooling involved in generation are ignored. (2) Intra-group homogenization, coarse-grained outcome rewards often lead to inefficiencies in intra-group advantage estimation with methods like Group Relative Policy Optimization (GRPO) during sampling. To address this, we propose Turn-level Stage-aware Policy Optimization (TSPO). TSPO introduces the First-Occurrence Latent Reward (FOLR) mechanism, allocating partial rewards to the step where the ground-truth answer first appears, thereby preserving process-level signals and increasing reward variance within groups without requiring external reward models or any annotations. Extensive experiments demonstrate that TSPO significantly outperforms state-of-the-art baselines, achieving average performance gains of 24% and 13.6% on Qwen2.5-3B and 7B models, respectively. Code is available at https://github.com/Flipped-May/TSPO.
  •  

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning

arXiv:2601.11109v3 Announce Type: replace-cross Abstract: Vision-as-inverse-graphics, the concept of reconstructing images into editable programs, remains challenging for Vision-Language Models (VLMs), which inherently lack fine-grained spatial grounding in one-shot settings. To address this, we introduce VIGA (Vision-as-Inverse-Graphics Agent), an interleaved multimodal reasoning framework where symbolic logic and visual perception actively cross-verify each other. VIGA operates through a tightly coupled code-render-inspect loop: synthesizing symbolic programs, projecting them into visual states, and inspecting discrepancies to guide iterative edits. Equipped with high-level semantic skills and an evolving multimodal memory, VIGA sustains evidence-based modifications over long horizons. This training-free, task-agnostic framework seamlessly supports 2D document generation, 3D reconstruction, multi-step 3D editing, and 4D physical interaction. Finally, we introduce BlenderBench, a challenging visual-to-code benchmark. Empirically, VIGA substantially improves accuracy compared with one-shot baselines in BlenderGym (35.32%), SlideBench (117.17%) and our proposed BlenderBench (124.70%).
  •  

Integrative Multi-Omics and Single-Cell Analysis Reveal THOC3 and THOC7 as Oncogenic RNA Processing Regulators in Lung Adenocarcinoma

Int J Med Sci. 2026 Mar 9;23(4):1408-1430. doi: 10.7150/ijms.128975. eCollection 2026.

ABSTRACT

Lung adenocarcinoma (LUAD) remains a leading cause of cancer-related mortality worldwide. Although the transcription-export (TREX) complex plays a central role in RNA maturation and nuclear export, the clinical and biological relevance of individual THO Complex Subunit (including THOC1, THOC2, THOC3, THOC5, THOC6, and THOC7) in LUAD is not well defined. We performed integrative analyses combining bulk transcriptomics from TCGA/GTEx and independent GEO cohorts, survival modeling, DNA methylation profiling, protein-level annotation from public resources, protein-protein interaction network analysis, immune infiltration estimation (TIMER), and single-cell RNA sequencing (scRNA-seq) to evaluate the relevance of THOC3 and THOC7 in LUAD. Across TCGA and external GEO validation datasets, THOC3 and THOC7 were consistently upregulated in LUAD and associated with poorer overall and disease-free survival, whereas other THO complex members showed weaker or inconsistent associations. Given these comparatively consistent and reproducible signals, we therefore prioritized THOC3 and THOC7 for downstream multi-layer analyses. Epigenetic profiling and interaction network analyses placed both genes within conserved RNA processing and export programs linked to genome maintenance pathways. Single-cell transcriptomic analysis provided additional resolution, demonstrating predominant enrichment of THOC3 and THOC7 in malignant epithelial clusters, with THOC3 aligning with transcriptional programs associated with DNA replication and repair, and THOC7 with proliferative and checkpoint-related states. Notably, expression of both genes was also detectable in myeloid and neutrophil subsets, and THOC7 expression remained elevated in recurrent LUAD samples, indicating association with aggressive and treatment-resistant disease states. Collectively, by integrating bulk, single-cell, epigenetic, and immune profiling across multiple independent cohorts, this study identifies THOC3 and THOC7 as reproducible molecular correlates of aggressive LUAD phenotypes. These highlight dysregulated RNA export programs as potential biomarkers of poor prognosis and motivate future functional studies to assess RNA export dependencies in LUAD.

PMID:41938520 | PMC:PMC13048885 | DOI:10.7150/ijms.128975

  •  

Integrative Multi-Omics and Single-Cell Analysis Reveal THOC3 and THOC7 as Oncogenic RNA Processing Regulators in Lung Adenocarcinoma

Int J Med Sci. 2026 Mar 9;23(4):1408-1430. doi: 10.7150/ijms.128975. eCollection 2026.

ABSTRACT

Lung adenocarcinoma (LUAD) remains a leading cause of cancer-related mortality worldwide. Although the transcription-export (TREX) complex plays a central role in RNA maturation and nuclear export, the clinical and biological relevance of individual THO Complex Subunit (including THOC1, THOC2, THOC3, THOC5, THOC6, and THOC7) in LUAD is not well defined. We performed integrative analyses combining bulk transcriptomics from TCGA/GTEx and independent GEO cohorts, survival modeling, DNA methylation profiling, protein-level annotation from public resources, protein-protein interaction network analysis, immune infiltration estimation (TIMER), and single-cell RNA sequencing (scRNA-seq) to evaluate the relevance of THOC3 and THOC7 in LUAD. Across TCGA and external GEO validation datasets, THOC3 and THOC7 were consistently upregulated in LUAD and associated with poorer overall and disease-free survival, whereas other THO complex members showed weaker or inconsistent associations. Given these comparatively consistent and reproducible signals, we therefore prioritized THOC3 and THOC7 for downstream multi-layer analyses. Epigenetic profiling and interaction network analyses placed both genes within conserved RNA processing and export programs linked to genome maintenance pathways. Single-cell transcriptomic analysis provided additional resolution, demonstrating predominant enrichment of THOC3 and THOC7 in malignant epithelial clusters, with THOC3 aligning with transcriptional programs associated with DNA replication and repair, and THOC7 with proliferative and checkpoint-related states. Notably, expression of both genes was also detectable in myeloid and neutrophil subsets, and THOC7 expression remained elevated in recurrent LUAD samples, indicating association with aggressive and treatment-resistant disease states. Collectively, by integrating bulk, single-cell, epigenetic, and immune profiling across multiple independent cohorts, this study identifies THOC3 and THOC7 as reproducible molecular correlates of aggressive LUAD phenotypes. These highlight dysregulated RNA export programs as potential biomarkers of poor prognosis and motivate future functional studies to assess RNA export dependencies in LUAD.

PMID:41938520 | PMC:PMC13048885 | DOI:10.7150/ijms.128975

  •  

User-preference alignment with uncertainty-aware interactive rectification for liver organ and tumor segmentation and analysis from CT images

npj Digital Medicine, Published online: 03 April 2026; doi:10.1038/s41746-026-02544-2

User-preference alignment with uncertainty-aware interactive rectification for liver organ and tumor segmentation and analysis from CT images
  •  

X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving

arXiv:2603.19979v2 Announce Type: replace-cross Abstract: Scalable and reliable evaluation is increasingly critical in the end-to-end era of autonomous driving, where vision--language--action (VLA) policies directly map raw sensor streams to driving actions. Yet, current evaluation pipelines still rely heavily on real-world road testing, which is costly, biased toward limited scenario coverage, and difficult to reproduce. These challenges motivate a real-world simulator that can generate realistic future observations under proposed actions, while remaining controllable and stable over long horizons. We present X-World, an action-conditioned multi-camera generative world model that simulates future observations directly in video space. Given synchronized multi-view camera history and a future action sequence, X-World generates future multi-camera video streams that follow the commanded actions. To ensure reproducible and editable scene rollouts, X-World further supports optional controls over dynamic traffic agents and static road elements, and retains a text-prompt interface for appearance-level control (e.g., weather and time of day). Beyond world simulation, X-World also enables video style transfer by conditioning on appearance prompts while preserving the underlying action and scene dynamics. At the core of X-World is a multi-view latent video generator designed to explicitly encourage cross-view geometric consistency and temporal coherence under diverse control signals. Experiments show that X-World achieves high-quality multi-view video generation with (i) strong view consistency across cameras, (ii) stable temporal dynamics over long rollouts, and (iii) high controllability with strict action following and faithful adherence to optional scene controls. These properties make X-World a practical foundation for scalable and reproducible evaluation.
  •  
  •  

Improving Safety Alignment via Balanced Direct Preference Optimization

arXiv:2603.22829v1 Announce Type: new Abstract: With the rapid development and widespread application of Large Language Models (LLMs), their potential safety risks have attracted widespread attention. Reinforcement Learning from Human Feedback (RLHF) has been adopted to enhance the safety performance of LLMs. As a simple and effective alternative to RLHF, Direct Preference Optimization (DPO) is widely used for safety alignment. However, safety alignment still suffers from severe overfitting, which limits its actual performance. This paper revisits the overfitting phenomenon from the perspective of the model's comprehension of the training data. We find that the Imbalanced Preference Comprehension phenomenon exists between responses in preference pairs, which compromises the model's safety performance. To address this, we propose Balanced Direct Preference Optimization (B-DPO), which adaptively modulates optimization strength between preferred and dispreferred responses based on mutual information. A series of experimental results show that B-DPO can enhance the safety capability while maintaining the competitive general capabilities of LLMs on various mainstream benchmarks compared to state-of-the-art methods. \color{red}{Warning: This paper contains examples of harmful texts, and reader discretion is recommended.
  •  

SketchGraphNet: A Memory-Efficient Hybrid Graph Transformer for Large-Scale Sketch Corpora Recognition

arXiv:2603.07521v1 Announce Type: cross Abstract: This work investigates large-scale sketch recognition from a graph-native perspective, where free-hand sketches are directly modeled as structured graphs rather than raster images or stroke sequences. We propose SketchGraphNet, a hybrid graph neural architecture that integrates local message passing with a memory-efficient global attention mechanism, without relying on auxiliary positional or structural encodings. To support systematic evaluation, we construct SketchGraph, a large-scale benchmark comprising 3.44 million graph-structured sketches across 344 categories, with two variants (A and R) to reflect different noise conditions. Each sketch is represented as a spatiotemporal graph with normalized stroke-order attributes. On SketchGraph-A and SketchGraph-R, SketchGraphNet achieves Top-1 accuracies of 83.62% and 87.61%, respectively, under a unified training configuration. MemEffAttn further reduces peak GPU memory by over 40% and training time by more than 30% compared with Performer-based global attention, while maintaining comparable accuracy.
  •  

Evaluating LLM-Based Grant Proposal Review via Structured Perturbations

arXiv:2603.08281v1 Announce Type: cross Abstract: As AI-assisted grant proposals outpace manual review capacity in a kind of ``Malthusian trap'' for the research ecosystem, this paper investigates the capabilities and limitations of LLM-based grant reviewing for high-stakes evaluation. Using six EPSRC proposals, we develop a perturbation-based framework probing LLM sensitivity across six quality axes: funding, timeline, competency, alignment, clarity, and impact. We compare three review architectures: single-pass review, section-by-section analysis, and a 'Council of Personas' ensemble emulating expert panels. The section-level approach significantly outperforms alternatives in both detection rate and scoring reliability, while the computationally expensive council method performs no better than baseline. Detection varies substantially by perturbation type, with alignment issues readily identified but clarity flaws largely missed by all systems. Human evaluation shows LLM feedback is largely valid but skewed toward compliance checking over holistic assessment. We conclude that current LLMs may provide supplementary value within EPSRC review but exhibit high variability and misaligned review priorities. We release our code and any non-protected data.
  •  

RLJP: Legal Judgment Prediction via First-Order Logic Rule-enhanced with Large Language Models

arXiv:2505.21281v2 Announce Type: replace Abstract: Legal Judgment Prediction (LJP) is a pivotal task in legal AI. Existing semantic-enhanced LJP models integrate judicial precedents and legal knowledge for high performance. But they neglect legal reasoning logic, a critical component of legal judgments requiring rigorous logical analysis. Although some approaches utilize legal reasoning logic for high-quality predictions, their logic rigidity hinders adaptation to case-specific logical frameworks, particularly in complex cases that are lengthy and detailed. This paper proposes a rule-enhanced legal judgment prediction framework based on first-order logic (FOL) formalism and comparative learning (CL) to develop an adaptive adjustment mechanism for legal judgment logic and further enhance performance in LJP. Inspired by the process of human exam preparation, our method follows a three-stage approach: first, we initialize judgment rules using the FOL formalism to capture complex reasoning logic accurately; next, we propose a Confusion-aware Contrastive Learning (CACL) to dynamically optimize the judgment rules through a quiz consisting of confusable cases; finally, we utilize the optimized judgment rules to predict legal judgments. Experimental results on two public datasets show superior performance across all metrics. The code is publicly available{https://anonymous.4open.science/r/RLJP-FDF1}.
  •  

LEDOM: Reverse Language Model

arXiv:2507.01335v3 Announce Type: replace-cross Abstract: Autoregressive language models are trained exclusively left-to-right. We explore the complementary factorization, training right-to-left at scale, and ask what reasoning patterns emerge when a model conditions on future context to predict the past. We train LEDOM, an open-source purely reverse autoregressive language model (2B/7B parameters, 435B tokens), and find it develops capabilities distinct from forward models, including abductive inference, question synthesis, and natural resolution of the reversal curse. We then explore one application of the reverse model: combining forward likelihood $P(y \mid x)$ with reverse posterior $P(x \mid y)$ through noisy channel duality. We propose Reverse Reward, which reranks forward outputs using reverse posterior estimates, and prove that bidirectional scoring penalizes hallucinated reasoning chains whose backward reconstruction degrades. Reverse Reward yields gains of up to 6.6\% on AIME 2024 and 15\% on AMC 2023 across multiple strong baselines. We release all models, code, and data here.
  •  

CM2: Reinforcement Learning with Checklist Rewards for Multi-Turn and Multi-Step Agentic Tool Use

arXiv:2602.12268v2 Announce Type: replace Abstract: AI agents are increasingly used to solve real-world tasks by reasoning over multi-turn user interactions and invoking external tools. However, applying reinforcement learning to such settings remains difficult: realistic objectives often lack verifiable rewards and instead emphasize open-ended behaviors; moreover, RL for multi-turn, multi-step agentic tool use is still underexplored; and building and maintaining executable tool environments is costly, limiting scale and coverage. We propose CM2, an RL framework that replaces verifiable outcome rewards with checklist rewards. CM2 decomposes each turn's intended behavior into fine-grained binary criteria with explicit evidence grounding and structured metadata, turning open-ended judging into more stable classification-style decisions. To balance stability and informativeness, our method adopts a strategy of sparse reward assignment but dense evaluation criteria. Training is performed in a scalable LLM-simulated tool environment, avoiding heavy engineering for large tool sets. Experiments show that CM2 consistently improves over supervised fine-tuning. Starting from an 8B Base model and training on an 8k-example RL dataset, CM2 improves over the SFT counterpart by 8 points on tau^-Bench, by 10 points on BFCL-V4, and by 12 points on ToolSandbox. The results match or even outperform similarly sized open-source baselines, including the judging model. CM2 thus provides a scalable recipe for optimizing multi-turn, multi-step tool-using agents without relying on verifiable rewards. Code provided by the open-source community: https://github.com/namezhenzhang/CM2-RLCR-Tool-Agent.
  •  
❌