❌

Reading view

OASES: Outcome-Aligned Search-Evaluation Co-Training for Agentic Search

arXiv:2604.03675v3 Announce Type: replace Abstract: Agentic search enables language models to solve knowledge-intensive tasks by adaptively acquiring external evidence over multiple steps. Reinforcement learning with verifiable rewards (RLVR) has emerged as a widely adopted training paradigm for search agents, yet outcome-only rewards are sparse and provide limited credit assignment for intermediate search actions. Existing process-reward methods therefore seek to densify supervision through proxy signals, external evaluators, or likelihood-based information gain. However, proxy rewards can deviate from the final outcome objective, while fixed evaluators can become stale as the search policy evolves, leading to unreliable process supervision. To address these challenges, we propose OASES, an Outcome-Aligned Search-Evaluation Supervision framework for agentic search. OASES derives outcome-aligned process rewards by evaluating how well each intermediate search state supports answering the original question. It further co-trains the search policy and the state evaluator on policy, allowing the evaluator to adapt to evolving search behavior and provide more reliable process rewards. Experiments on five multi-hop QA benchmarks show that OASES consistently outperforms strong RL baselines, with further analyses confirming the benefits of outcome-aligned process rewards and search-evaluation co-training.
  •  

Raspberry aqueous extract ameliorates MAFLD in mice by regulating gut microbiota and purine metabolism

Front Nutr. 2026 Apr 30;13:1818086. doi: 10.3389/fnut.2026.1818086. eCollection 2026.

ABSTRACT

INTRODUCTION: Metabolism-associated fatty liver disease (MAFLD) has emerged as a severe worldwide public health burden with insufficient available clinical therapeutic strategies, which underscores the urgent demand for safe, natural dietary interventions. Raspberry (Rubus idaeus L.), a typical food-medicine homologous fruit abundant in diverse bioactive components including anthocyanins, flavonoids and polysaccharides, possesses prominent nutritional and medicinal potential.

METHODS: In this study, raspberry aqueous extract (RE) was prepared to comprehensively investigate its ameliorative effects and underlying molecular mechanisms against MAFLD. MAFLD animal model was established in C57BL/6 mice via 12-week high-fat diet (HFD) feeding. From the 9th week, model mice were intragastrically administered with RE at doses of 1 g/kg/d and 2 g/kg/d for continuous intervention. Integrated multi-omics analyses including 16S rRNA microbial sequencing, serum/hepatic biochemical detection, histopathological examination, in vivo microbial colonization assay, and in vitro cellular and metabolomic experiments were performed to systematically clarify the regulatory mechanism.

RESULTS: RE treatment markedly improved the core pathological phenotypes of MAFLD mice, and significantly mitigated hepatic steatosis and hepatocellular injury. 16S rRNA sequencing demonstrated that RE remodeled the gut microbial dysbiosis, specifically elevating the abundance of beneficial genus Ileibacterium and suppressing pathogenic microbial taxa. Meanwhile, RE strengthened intestinal mucosal barrier integrity by upregulating tight junction protein expression, and activated hepatic purine metabolic reprogramming to boost the levels of critical metabolites including inosine and ADP. Spearman correlation analysis verified the significantly positive correlation between Ileibacterium abundance and hepatic inosine content, and both factors were closely correlated with the remission of MAFLD pathological indicators. In vivo colonization experiments further validated that Ileibacterium intervention alone remarkably alleviated hepatic lipid deposition and liver damage in MAFLD mice. In vitro strain metabolomics confirmed that Ileibacterium could directly biosynthesize and secrete inosine extracellularly. Furthermore, in vitro AML12 hepatocyte experiments revealed that 100 ΞΌM inosine remarkably relieved palmitic acid-induced lipotoxicity via reducing intracellular lipid overload, reactive oxygen species (ROS) accumulation and mitochondrial dysfunction, alongside modulating the expression of lipid metabolism, inflammatory and autophagy-related genes.

DISCUSSION: Collectively, our results elucidate that raspberry aqueous extract alleviates experimental MAFLD through the gut microbiota-purine metabolism-inosine regulatory axis, in which Ileibacterium and inosine act as the core synergistic mediators. This study provides solid preclinical experimental evidence for the development and application of raspberry as a promising functional food for the prevention and nutritional intervention of MAFLD.

PMID:42146077 | PMC:PMC13171365 | DOI:10.3389/fnut.2026.1818086

  •  

PRAISE: Prefix-Based Rollout Reuse in Agentic Search Training

arXiv:2604.03675v1 Announce Type: new Abstract: In agentic search, large language models (LLMs) are trained to perform multi-turn retrieval and reasoning for complex tasks such as multi-hop question answering (QA). However, current search-based Reinforcement Learning (RL) methods suffer from two core limitations: expensive long-horizon rollouts are under-utilized during training, and supervision is typically available only at the final answer, resulting in severe reward sparsity. We present Prefix-based Rollout reuse for Agentic search with Intermediate Step rEwards (PRAISE), a framework for improving both data efficiency and credit assignment in agentic search training. Given a complete search trajectory, PRAISE extracts prefix states at different search turns, elicits intermediate answers from them, and uses these prefixes both to construct additional training trajectories and to derive step-level rewards from performance differences across prefixes. Our method uses a single shared model for both search policy learning and prefix answer evaluation, enabling joint optimization without extra human annotations or a separate reward model. Experiments on multi-hop QA benchmarks show that PRAISE consistently improves performance over strong baselines.
  •  

Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows

arXiv:2512.13168v4 Announce Type: replace Abstract: We introduce FinWorkBench (a.k.a. Finch), a benchmark for evaluating agents on real-world, enterprise-grade finance and accounting workflows that interleave data entry, structuring, formatting, web search, cross-file retrieval, calculation, modeling, validation, translation, visualization, and reporting. Finch is built from authentic enterprise workspaces from Enron (15,000 files and 500,000 emails) and other financial institutions spanning 2000 to 2025, preserving the in-the-wild messiness of multimodal artifacts such as tables and charts across diverse domains including budgeting, trading, and asset management. We propose a workflow construction process that combines LLM-assisted mining of workflows from authentic enterprise environments with expert annotation. Specifically, we use LLM-assisted, expert-verified derivation of workflows from real-world email threads and spreadsheet version histories, followed by meticulous workflow annotation requiring more than 700 hours of expert effort. This process yields 172 composite workflows with 384 tasks, involving 1,710 spreadsheets with 27 million cells, along with PDFs and other artifacts, capturing the intrinsically messy, long-horizon, knowledge-intensive, and collaborative nature of enterprise work. We conduct both human and automated evaluations of frontier AI systems, including GPT 5.1, Claude Sonnet/Opus 4.5, Gemini 3 Pro, Grok 4, and Qwen 3 Max. GPT 5.1 Pro spends an average of 16.8 minutes per workflow yet passes only 38.4% of workflows. Comprehensive case studies further highlight the challenges that real-world enterprise workflows pose for AI agents.
  •  

STRUCTUREDAGENT: Planning with AND/OR Trees for Long-Horizon Web Tasks

arXiv:2603.05294v2 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have enabled agentic systems for sequential decision-making. Such agents must perceive their environment, reason across multiple time steps, and take actions that optimize long-term objectives. However, existing web agents struggle on complex, long-horizon tasks due to limited in-context memory for tracking history, weak planning abilities, and greedy behaviors that lead to premature termination. To address these challenges, we propose STRUCTUREDAGENT, a hierarchical planning framework with two core components: (1) an online hierarchical planner that uses dynamic AND/OR trees for efficient search and (2) a structured memory module that tracks and maintains candidate solutions to improve constraint satisfaction in information-seeking tasks. The framework also produces interpretable hierarchical plans, enabling easier debugging and facilitating human intervention when needed. Our results on WebVoyager, WebArena, and custom shopping benchmarks show that STRUCTUREDAGENT improves performance on long-horizon web-browsing tasks compared to standard LLM-based agents.
  •  

Visual Prompt Guided Unified Pushing Policy

arXiv:2602.19193v1 Announce Type: cross Abstract: As one of the simplest non-prehensile manipulation skills, pushing has been widely studied as an effective means to rearrange objects. Existing approaches, however, typically rely on multi-step push plans composed of pre-defined pushing primitives with limited application scopes, which restrict their efficiency and versatility across different scenarios. In this work, we propose a unified pushing policy that incorporates a lightweight prompting mechanism into a flow matching policy to guide the generation of reactive, multimodal pushing actions. The visual prompt can be specified by a high-level planner, enabling the reuse of the pushing policy across a wide range of planning problems. Experimental results demonstrate that the proposed unified pushing policy not only outperforms existing baselines but also effectively serves as a low-level primitive within a VLM-guided planning framework to solve table-cleaning tasks efficiently.
  •  
❌