❌

Normal view

SAM: State-Adaptive Memory for Long-Horizon Reasoning Agent

arXiv:2605.24468v1 Announce Type: new Abstract: Long-horizon agentic reasoning requires large language models to act over long interaction histories containing thoughts, tool calls, observations, and partial conclusions. The challenge is not merely that these histories grow long, but that information needed for the current decision may be scattered across distant steps and only become relevant later. Existing approaches address this difficulty by truncating the interaction history, compressing it into shorter surrogates, or retrieving selected parts of it for reuse, but they do not explicitly model how access to past interaction should adapt to the agent's evolving state. We instead cast long-horizon reasoning as a problem of state-adaptive memory. To this end, we propose State-Adaptive Memory~(SAM), a standalone framework that consolidates ongoing interaction into compact memory cues while preserving raw trajectory pages for intent-driven recall. These cues are not treated as replacements for history; rather, they serve as lightweight handles that allow the agent to reconstruct temporally distant information according to its current needs, without retraining the underlying backbone. We further optimize the memory module through expert-guided supervision and reinforcement learning, aligning it with trajectory-level utility. Across BrowseComp, BrowseComp-ZH, WideSearch, and HLE, SAM consistently outperforms strong baselines over diverse agent backbones. Our results suggest that explicit memory modeling provides a simple and effective foundation for long-horizon agentic reasoning.

AgentFugue: Agent Scaling for Long-Horizon Tasks through Collective Reasoning

arXiv:2605.24486v1 Announce Type: new Abstract: Recent progress on long-horizon agentic tasks has been driven largely by scaling up individual agents through stronger models, better tools, and more effective scaffolding. In contrast, much less is understood about scaling out: whether multiple peer agents, all targeting the same task, can become an additional source of capability without relying on explicit role specialization or workflow orchestration. We study this question and propose AgentFugue, a collective reasoning framework built around a shared reasoning hub. As peer agents explore the same task in parallel, the hub records concise notes on what each agent has established, attempted, or ruled out, and enables each agent to selectively access what other agents have discovered in a form useful for its current search. This design turns otherwise isolated trajectories into a connected ecology of reusable intermediate reasoning without requiring centralized planning. We instantiate the hub as a plug-in communication layer, trained with supervised fine-tuning and end-to-end reinforcement learning. Across the challenging long-horizon settings we study, AgentFugue improves over strong baselines. Our results suggest that collective reasoning can turn scaling out peer agent systems into a distinct source of capability gains, rather than merely a way of spending more compute.
  • ✇cs.AI, q-bio.NC updates on arXiv.org
  • Extreme Region Policy Distillation Changyu Chen · Xiting Wang · Rui Yan
    arXiv:2605.25582v1 Announce Type: cross Abstract: Reinforcement learning for large language models faces a fundamental trade-off between sample efficiency and asymptotic performance: strictly on-policy methods discard trajectories after a single update, while off-policy reuse introduces distribution mismatch that existing trust-region techniques mitigate primarily by enforcing conservative optimization, often leaving rich training signals underutilized. To investigate this, we perform extensive
     

Extreme Region Policy Distillation

arXiv:2605.25582v1 Announce Type: cross Abstract: Reinforcement learning for large language models faces a fundamental trade-off between sample efficiency and asymptotic performance: strictly on-policy methods discard trajectories after a single update, while off-policy reuse introduces distribution mismatch that existing trust-region techniques mitigate primarily by enforcing conservative optimization, often leaving rich training signals underutilized. To investigate this, we perform extensive off-policy updates on fixed data. Our experiments reveal that aggressive multi-step optimization brings rapid initial gains, but excessive updates cause trajectory probabilities to deviate and entropy to collapse, with performance plateauing early. Tightening KL constraints merely lowers the ceiling without resolving the degradation. This motivates Extreme Region Policy Distillation (ERPD), a two-stage framework that decouples sample efficiency from KL efficiency. The first stage performs weakly constrained off-policy optimization on fixed data to maximally extract training signals. The resulting policy provides token-level supervision. In the second stage, we distill these signals into the base policy under trust-region constraints, filtering harmful drift while preserving useful signals. The distilled policy achieves comparable or better performance with substantially smaller KL divergence, indicating that much of the first-stage divergence was spent on unnecessary drift rather than genuine improvement. Crucially, ERPD accommodates both strong and weak teachers: when aggressive optimization yields no stronger policy, even degenerate teachers provide effective supervision via alternative signal construction strategies. We validate ERPD on mathematical reasoning, showing gains for strong base models where on-policy training plateaus, and reliable improvements with weak teachers.

The 2025 lung cancer landscape: advances in screening, molecular taxonomy and therapeutic strategy: a narrative review

Transl Lung Cancer Res. 2026 Mar 23;15(3):62. doi: 10.21037/tlcr-2025-1-1477. Epub 2026 Mar 18.

ABSTRACT

BACKGROUND AND OBJECTIVE: In 2025, lung cancer research advanced rapidly across the disease continuum, from population-level risk assessment and screening to mechanistic studies of early carcinogenesis and therapeutic innovation in perioperative and metastatic settings. A key shift moved beyond a smoking-centred paradigm toward a multidimensional risk framework reflecting the growing burden among never-smokers and the roles of air pollution, occupational exposures, and systemic metabolic-inflammatory states. This narrative review aims to synthesize influential 2025 evidence across prevention, diagnosis, treatment, and survivorship, and to identify convergent themes and translational gaps relevant to clinical practice and policy.

METHODS: We performed a narrative synthesis of influential lung cancer studies published in major international journals in 2025. Evidence was organized along a clinically oriented pathway spanning carcinogenesis and screening, precision diagnosis, treatment optimization in resectable and advanced disease, and survivorship, emphasizing practice-informing trials, high-impact translational research, and implementation-relevant technologies.

KEY CONTENT AND FINDINGS: Lineage tracing, single-cell and spatial omics, and evolutionary inference refined concepts of field cancerization, clonal selection, and copy-number-driven fitness. In small-cell lung cancer, evidence further supported neuronal coupling and synapse-like programs as potentially tractable vulnerabilities. Clinically, low-dose computed tomography (CT) strategies and data-informed nodule thresholds aimed to balance under-detection against over-surveillance harms. In diagnostics, artificial intelligence (AI) models increasingly inferred molecular features from routine histopathology ("virtual molecular testing") and should be regarded as decision support requiring prospective validation, population calibration, and explicit failure-mode reporting. Multimodal approaches integrating imaging with circulating tumor DNA (ctDNA) improved feasibility in tissue-limited settings, but clinical utility remains contingent on assay standardization and pathway-level implementation. In resectable disease, longer follow-up consolidated neoadjuvant chemo-immunotherapy for selected patients, while ctDNA kinetics emerged as a candidate biomarker for response-adaptive escalation and de-escalation. In advanced non-small cell lung cancer (NSCLC), phase III evidence for antibody-drug conjugates and bispecific antibodies began reshaping sequencing, while highlighting challenges in toxicity, access, affordability, and immature overall survival in several programs.

CONCLUSIONS: The 2025 landscape reflects coordinated progress in risk conceptualization, biology, diagnostics, and therapeutics, yet gaps in validation, standardization, and real-world deliverability persist. Priorities include prospective evaluation of AI- and ctDNA-enabled pathways, toxicity-informed sequencing, and equitable implementation aligned with health-system capacity.

PMID:41982682 | PMC:PMC13071762 | DOI:10.21037/tlcr-2025-1-1477

Integrative fragmentomic and mutational signature profile of plasma cfDNA for early lung cancer detection

NPJ Precis Oncol. 2026 Apr 15. doi: 10.1038/s41698-026-01416-y. Online ahead of print.

ABSTRACT

Detecting lung cancer effectively in the general population is essential for optimizing treatment outcomes and improving the 5-year survival rate. While low-dose computed tomography (LDCT) is the current standard, it has limitations in broader populations. We developed a blood-based multi-omics model using whole-genome cell-free DNA (cfDNA) features to distinguish lung cancer from non-cancer individuals. This study included 1600 patients and an equal number of non-cancer controls, divided into training and validation cohorts. The model achieved an area under the curve (AUC) of 95.59% for the training cohort and 95.74% for the validation cohort. The model consistently performed well across various cancer stages and histological subtypes. To further validate the performance of the model, an external validation cohort was utilized. Notably, it also effectively differentiated non-cancer samples from cancer samples in the external validation cohort, with 85.9% sensitivity and 94.78% specificity. Importantly, in simulated population screenings, our ctDNA assay outperformed both LDCT and a previously established method. This suggests its potential utility in wider lung cancer screening programs, possibly complementing the LDCT approach. In conclusion, our ctDNA assay emerges as a promising and highly sensitive tool for the early detection and categorization of lung cancer.

PMID:41986614 | DOI:10.1038/s41698-026-01416-y

❌