❌

Reading view

M$^\star$: Every Task Deserves Its Own Memory Harness

arXiv:2604.11811v2 Announce Type: replace-cross Abstract: Large language model agents rely on specialized memory systems to accumulate and reuse knowledge during extended interactions. Recent architectures typically adopt a fixed memory design tailored to specific domains, such as semantic retrieval for conversations or skills reused for coding. However, a memory system optimized for one purpose frequently fails to transfer to others. To address this limitation, we introduce M$^\star$, a method that automatically discovers task-optimized memory harnesses through executable program evolution. Specifically, M$^\star$ models an agent memory system as a memory program written in Python. This program encapsulates the data Schema, the storage Logic, and the agent workflow Instructions. We optimize these components jointly using a reflective code evolution method; this approach employs a population-based search strategy and analyzes evaluation failures to iteratively refine the candidate programs. We evaluate M$^\star$ on four distinct benchmarks spanning conversation, embodied planning, and expert reasoning. Our results demonstrate that M$^\star$ improves performance over existing fixed-memory baselines robustly across all evaluated tasks. Furthermore, the evolved memory programs exhibit structurally distinct processing mechanisms for each domain. This finding indicates that specializing the memory mechanism for a given task explores a broad design space and provides a superior solution compared to general-purpose memory paradigms.
  •  
  •  

ClawArena: Benchmarking AI Agents in Evolving Information Environments

arXiv:2604.04202v1 Announce Type: cross Abstract: AI agents deployed as persistent assistants must maintain correct beliefs as their information environment evolves. In practice, evidence is scattered across heterogeneous sources that often contradict one another, new information can invalidate earlier conclusions, and user preferences surface through corrections rather than explicit instructions. Existing benchmarks largely assume static, single-authority settings and do not evaluate whether agents can keep up with this complexity. We introduce ClawArena, a benchmark for evaluating AI agents in evolving information environments. Each scenario maintains a complete hidden ground truth while exposing the agent only to noisy, partial, and sometimes contradictory traces across multi-channel sessions, workspace files, and staged updates. Evaluation is organized around three coupled challenges: multi-source conflict reasoning, dynamic belief revision, and implicit personalization, whose interactions yield a 14-category question taxonomy. Two question formats, multi-choice (set-selection) and shell-based executable checks, test both reasoning and workspace grounding. The current release contains 64 scenarios across 8 professional domains, totaling 1{,}879 evaluation rounds and 365 dynamic updates. Experiments on five agent frameworks and five language models show that both model capability (15.4% range) and framework design (9.2%) substantially affect performance, that self-evolving skill frameworks can partially close model-capability gaps, and that belief revision difficulty is determined by update design strategy rather than the mere presence of updates. Code is available at https://github.com/aiming-lab/ClawArena.
  •  

Preserving Forgery Artifacts: AI-Generated Video Detection at Native Scale

arXiv:2604.04634v1 Announce Type: cross Abstract: The rapid advancement of video generation models has enabled the creation of highly realistic synthetic media, raising significant societal concerns regarding the spread of misinformation. However, current detection methods suffer from critical limitations. They rely on preprocessing operations like fixed-resolution resizing and cropping. These operations not only discard subtle, high-frequency forgery traces but also cause spatial distortion and significant information loss. Furthermore, existing methods are often trained and evaluated on outdated datasets that fail to capture the sophistication of modern generative models. To address these challenges, we introduce a comprehensive dataset and a novel detection framework. First, we curate a large-scale dataset of over 140K videos from 15 state-of-the-art open-source and commercial generators, along with Magic Videos benchmark designed specifically for evaluating ultra-realistic synthetic content. In addition, we propose a novel detection framework built on the Qwen2.5-VL Vision Transformer, which operates natively at variable spatial resolutions and temporal durations. This native-scale approach effectively preserves the high-frequency artifacts and spatiotemporal inconsistencies typically lost during conventional preprocessing. Extensive experiments demonstrate that our method achieves superior performance across multiple benchmarks, underscoring the critical importance of native-scale processing and establishing a robust new baseline for AI-generated video detection.
  •  

TSPO: Breaking the Double Homogenization Dilemma in Multi-turn Search Policy Optimization

arXiv:2601.22776v2 Announce Type: replace Abstract: Multi-turn tool-integrated reasoning enables Large Language Models (LLMs) to solve complex tasks through iterative information retrieval. However, current reinforcement learning (RL) frameworks for search-augmented reasoning predominantly rely on sparse outcome-level rewards, leading to a "Double Homogenization Dilemma." This manifests as (1) Process homogenization, where the thinking, reasoning, and tooling involved in generation are ignored. (2) Intra-group homogenization, coarse-grained outcome rewards often lead to inefficiencies in intra-group advantage estimation with methods like Group Relative Policy Optimization (GRPO) during sampling. To address this, we propose Turn-level Stage-aware Policy Optimization (TSPO). TSPO introduces the First-Occurrence Latent Reward (FOLR) mechanism, allocating partial rewards to the step where the ground-truth answer first appears, thereby preserving process-level signals and increasing reward variance within groups without requiring external reward models or any annotations. Extensive experiments demonstrate that TSPO significantly outperforms state-of-the-art baselines, achieving average performance gains of 24% and 13.6% on Qwen2.5-3B and 7B models, respectively. Code is available at https://github.com/Flipped-May/TSPO.
  •  

A Data-driven Approach for Biomarker Discovery based on U-centered Distance Correlation Network: Multi-omics Warning Signals for Non-small Cell Lung Cancer

Comb Chem High Throughput Screen. 2026 Mar 27. doi: 10.2174/0113862073445368260131002109. Online ahead of print.

ABSTRACT

INTRODUCTION/OBJECTIVE: Lung cancer is the leading cause of cancer-related mortality worldwide, and non-small cell lung cancer (NSCLC) accounts for the majority of cases. Alterations in metabolic activities play important roles in NSCLC development, wherein related genes and metabolites interact with each other, involving multiple forms.

METHODS: To comprehensively understand the pathogenic mechanisms and improve the performance of clinical early, precise diagnosis, this study proposed a data-driven approach for biomarker discovery based on U-centered distance correlation network (DCN) to investigate NSCLC metabolism-related reactions. In DCN, changes in molecular relationships during NSCLC initiation and progression are measured using the t-statistics of U-centered distance correlation for network construction, in which prospective warning signals representing NSCLC onset can be identified without human intervention. Additionally, the network construction criterion in DCN can precisely and effectively capture both linear and nonlinear molecular relationships in simple and biologically relevant manners.

RESULTS: DCN was successfully employed to analyze NSCLC metabolism-related metabolomics and genomics datasets. Statistical analyses confirmed that compared with other algorithms, the gene and metabolite biomarker panels identified by DCN provided more reliable diagnostic capabilities for clinical NSCLC detection. Biological analyses revealed that disturbed energy metabolism and lipid metabolism occurred during tumor cell proliferation and growth in NSCLC patients.

DISCUSSION: The gene ASPA and metabolite aspartic acid were significantly decreased in NSCLC samples, suggesting that the corresponding amino acid metabolic activities were intricately linked to NSCLC progression.

CONCLUSION: These findings demonstrated that DCN can further facilitate NSCLC studies to improve clinical outcomes in patients.

PMID:41937706 | DOI:10.2174/0113862073445368260131002109

  •  

A Data-driven Approach for Biomarker Discovery based on U-centered Distance Correlation Network: Multi-omics Warning Signals for Non-small Cell Lung Cancer

Comb Chem High Throughput Screen. 2026 Mar 27. doi: 10.2174/0113862073445368260131002109. Online ahead of print.

ABSTRACT

INTRODUCTION/OBJECTIVE: Lung cancer is the leading cause of cancer-related mortality worldwide, and non-small cell lung cancer (NSCLC) accounts for the majority of cases. Alterations in metabolic activities play important roles in NSCLC development, wherein related genes and metabolites interact with each other, involving multiple forms.

METHODS: To comprehensively understand the pathogenic mechanisms and improve the performance of clinical early, precise diagnosis, this study proposed a data-driven approach for biomarker discovery based on U-centered distance correlation network (DCN) to investigate NSCLC metabolism-related reactions. In DCN, changes in molecular relationships during NSCLC initiation and progression are measured using the t-statistics of U-centered distance correlation for network construction, in which prospective warning signals representing NSCLC onset can be identified without human intervention. Additionally, the network construction criterion in DCN can precisely and effectively capture both linear and nonlinear molecular relationships in simple and biologically relevant manners.

RESULTS: DCN was successfully employed to analyze NSCLC metabolism-related metabolomics and genomics datasets. Statistical analyses confirmed that compared with other algorithms, the gene and metabolite biomarker panels identified by DCN provided more reliable diagnostic capabilities for clinical NSCLC detection. Biological analyses revealed that disturbed energy metabolism and lipid metabolism occurred during tumor cell proliferation and growth in NSCLC patients.

DISCUSSION: The gene ASPA and metabolite aspartic acid were significantly decreased in NSCLC samples, suggesting that the corresponding amino acid metabolic activities were intricately linked to NSCLC progression.

CONCLUSION: These findings demonstrated that DCN can further facilitate NSCLC studies to improve clinical outcomes in patients.

PMID:41937706 | DOI:10.2174/0113862073445368260131002109

  •  

DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning

arXiv:2604.01765v1 Announce Type: cross Abstract: Recently, world-action models (WAM) have emerged to bridge vision-language-action (VLA) models and world models, unifying their reasoning and instruction-following capabilities and spatio-temporal world modeling. However, existing WAM approaches often focus on modeling 2D appearance or latent representations, with limited geometric grounding-an essential element for embodied systems operating in the physical world. We present DriveDreamer-Policy, a unified driving world-action model that integrates depth generation, future video generation, and motion planning within a single modular architecture. The model employs a large language model to process language instructions, multi-view images, and actions, followed by three lightweight generators that produce depth, future video, and actions. By learning a geometry-aware world representation and using it to guide both future prediction and planning within a unified framework, the proposed model produces more coherent imagined futures and more informed driving actions, while maintaining modularity and controllable latency. Experiments on the Navsim v1 and v2 benchmarks demonstrate that DriveDreamer-Policy achieves strong performance on both closed-loop planning and world generation tasks. In particular, our model reaches 89.2 PDMS on Navsim v1 and 88.7 EPDMS on Navsim v2, outperforming existing world-model-based approaches while producing higher-quality future video and depth predictions. Ablation studies further show that explicit depth learning provides complementary benefits to video imagination and improves planning robustness.
  •  

Beyond Matching to Tiles: Bridging Unaligned Aerial and Satellite Views for Vision-Only UAV Navigation

arXiv:2603.22153v2 Announce Type: replace-cross Abstract: Recent advances in cross-view geo-localization (CVGL) methods have shown strong potential for supporting unmanned aerial vehicle (UAV) navigation in GNSS-denied environments. However, existing work predominantly focuses on matching UAV views to onboard map tiles, which introduces an inherent trade-off between accuracy and storage overhead, and overlooks the importance of the UAV's heading during navigation. Moreover, the substantial discrepancies and varying overlaps in cross-view scenarios have been insufficiently considered, limiting their generalization to real-world scenarios. In this paper, we present Bearing-UAV, a purely vision-driven cross-view navigation method that jointly predicts UAV absolute location and heading from neighboring features, enabling accurate, lightweight, and robust navigation in the wild. Our method leverages global and local structural features and explicitly encodes relative spatial relationships, making it robust to cross-view variations, misalignment, and feature-sparse conditions. We also present Bearing-UAV-90k, a multi-city benchmark for evaluating cross-view localization and navigation. Extensive experiments show encouraging results that Bearing-UAV yields lower localization error than previous matching/retrieval paradigm across diverse terrains. Our code and dataset will be made publicly available.
  •  

DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving

arXiv:2601.01528v2 Announce Type: replace-cross Abstract: Video generation models, as one form of world models, have emerged as one of the most exciting frontiers in AI, promising agents the ability to imagine the future by modeling the temporal evolution of complex scenes. In autonomous driving, this vision gives rise to driving world models: generative simulators that imagine ego and agent futures, enabling scalable simulation, safe testing of corner cases, and rich synthetic data generation. Yet, despite fast-growing research activity, the field lacks a rigorous benchmark to measure progress and guide priorities. Existing evaluations remain limited: generic video metrics overlook safety-critical imaging factors; trajectory plausibility is rarely quantified; temporal and agent-level consistency is neglected; and controllability with respect to ego conditioning is ignored. Moreover, current datasets fail to cover the diversity of conditions required for real-world deployment. To address these gaps, we present DrivingGen, the first comprehensive benchmark for generative driving world models. DrivingGen combines a diverse evaluation dataset curated from both driving datasets and internet-scale video sources, spanning varied weather, time of day, geographic regions, and complex maneuvers, with a suite of new metrics that jointly assess visual realism, trajectory plausibility, temporal coherence, and controllability. Benchmarking 14 state-of-the-art models reveals clear trade-offs: general models look better but break physics, while driving-specific ones capture motion realistically but lag in visual quality. DrivingGen offers a unified evaluation framework to foster reliable, controllable, and deployable driving world models, enabling scalable simulation, planning, and data-driven decision-making.
  •  
❌