❌

Normal view

ToolMind Technical Report: A Large-Scale, Reasoning-Enhanced Tool-Use Dataset

arXiv:2511.15718v2 Announce Type: replace Abstract: Large Language Model (LLM) agents have developed rapidly in recent years to solve complex real-world problems using external tools. However, the scarcity of high-quality trajectories still hinders the development of stronger LLM agents. Most existing works on multi-turn dialogue synthesis validate correctness only at the trajectory level, which may overlook turn-level errors that can propagate during training and degrade model performance. To address these limitations, we introduce ToolMind, a large-scale, high-quality tool-agentic dataset with 160k synthetic data instances generated using over 20k tools and 200k augmented open-source data instances. Our data synthesis pipeline first constructs a function graph based on parameter correlations and then uses a multi-agent framework to simulate realistic user-assistant-tool interactions. Beyond trajectory-level validation, we employ fine-grained turn-level filtering to remove erroneous or suboptimal steps, ensuring that only high-quality reasoning traces are retained. This approach mitigates error amplification during training while preserving self-corrective reasoning signals essential for robust tool-use learning. Models fine-tuned on ToolMind show significant improvements over baselines on several benchmarks.

SciAgent: A Unified Multi-Agent System for Generalistic Scientific Reasoning

arXiv:2511.08151v2 Announce Type: replace Abstract: Recent advances in large language models have enabled AI systems to achieve expert-level performance on domain-specific scientific tasks, yet these systems remain narrow and handcrafted. We introduce SciAgent, a unified multi-agent system designed for generalistic scientific reasoning-the ability to adapt reasoning strategies across disciplines and difficulty levels. SciAgent organizes problem solving as a hierarchical process: a Coordinator Agent interprets each problem's domain and complexity, dynamically orchestrating specialized Worker Systems, each composed of interacting reasoning Sub-agents for symbolic deduction, conceptual modeling, numerical computation, and verification. These agents collaboratively assemble and refine reasoning pipelines tailored to each task. Across mathematics and physics Olympiads (IMO, IMC, IPhO, CPhO), SciAgent consistently attains or surpasses human gold-medalist performance, demonstrating both domain generality and reasoning adaptability. Additionally, SciAgent has been tested on the International Chemistry Olympiad (IChO) and selected problems from the Humanity's Last Exam (HLE) benchmark, further confirming the system's ability to generalize across diverse scientific domains. This work establishes SciAgent as a concrete step toward generalistic scientific intelligence-AI systems capable of coherent, cross-disciplinary reasoning at expert levels.

SciAgent: A Unified Multi-Agent System for Generalistic Scientific Reasoning

arXiv:2511.08151v1 Announce Type: new Abstract: Recent advances in large language models have enabled AI systems to achieve expert-level performance on domain-specific scientific tasks, yet these systems remain narrow and handcrafted. We introduce SciAgent, a unified multi-agent system designed for generalistic scientific reasoning-the ability to adapt reasoning strategies across disciplines and difficulty levels. SciAgent organizes problem solving as a hierarchical process: a Coordinator Agent interprets each problem's domain and complexity, dynamically orchestrating specialized Worker Systems, each composed of interacting reasoning Sub-agents for symbolic deduction, conceptual modeling, numerical computation, and verification. These agents collaboratively assemble and refine reasoning pipelines tailored to each task. Across mathematics and physics Olympiads (IMO, IMC, IPhO, CPhO), SciAgent consistently attains or surpasses human gold-medalist performance, demonstrating both domain generality and reasoning adaptability. Additionally, SciAgent has been tested on the International Chemistry Olympiad (IChO) and selected problems from the Humanity's Last Exam (HLE) benchmark, further confirming the system's ability to generalize across diverse scientific domains. This work establishes SciAgent as a concrete step toward generalistic scientific intelligence-AI systems capable of coherent, cross-disciplinary reasoning at expert levels.

No-Human in the Loop: Agentic Evaluation at Scale for Recommendation

arXiv:2511.03051v1 Announce Type: new Abstract: Evaluating large language models (LLMs) as judges is increasingly critical for building scalable and trustworthy evaluation pipelines. We present ScalingEval, a large-scale benchmarking study that systematically compares 36 LLMs, including GPT, Gemini, Claude, and Llama, across multiple product categories using a consensus-driven evaluation protocol. Our multi-agent framework aggregates pattern audits and issue codes into ground-truth labels via scalable majority voting, enabling reproducible comparison of LLM evaluators without human annotation. Applied to large-scale complementary-item recommendation, the benchmark reports four key findings: (i) Anthropic Claude 3.5 Sonnet achieves the highest decision confidence; (ii) Gemini 1.5 Pro offers the best overall performance across categories; (iii) GPT-4o provides the most favorable latency-accuracy-cost tradeoff; and (iv) GPT-OSS 20B leads among open-source models. Category-level analysis shows strong consensus in structured domains (Electronics, Sports) but persistent disagreement in lifestyle categories (Clothing, Food). These results establish ScalingEval as a reproducible benchmark and evaluation protocol for LLMs as judges, with actionable guidance on scaling, reliability, and model family tradeoffs.

Clinical performance evaluation of a plasma dual-target methylation test for the detection of primary liver cancer: a multicenter study

Primary liver cancer (PLC) is a global health concern. The plasma dual-target methylation (PDTM) test, which interrogates the methylation status of GNB4 and Riplet, exhibits a commendable ability to discriminate ...

Key Lipid Reprogramming Revealed in Gastric Signet Ring Cell Carcinoma by Spatial Mass Spectrometry Metabolomics

J Am Soc Mass Spectrom. 2025 Aug 6;36(8):1598-1608. doi: 10.1021/jasms.4c00505. Epub 2025 Jul 2.

ABSTRACT

Gastric signet ring cell carcinoma (GSRC) is an aggressive subtype of gastric cancer (GC) with a poor prognosis. The lack of a systematic molecular and metabolic heterogeneity overview has led to slow progress in clinical practice. This study used mass spectrometry imaging (MSI) to investigate the metabolic landscape of GSRC in GC tissue with various differentiation grades. Our comprehensive spatial profiling of metabolites and lipids unveiled distinct metabolic signatures across different tissue subregions. A substantial number of lipidomic biomarkers associated with GSRC were identified, including phosphatidylethanolamine N-methyl (PE-NMe), phosphatidylethanolamine (PE), sphingomyelin (SM), diacylglycerol (DG), phosphatidic acid (PA), and phosphatidylcholine (PC), which may provide insights into its pathogenesis and potential therapeutic targets. Furthermore, multi-omics network analysis revealed intricate metabolic pathways involved in GSRC progression. Our findings highlight the importance of understanding the metabolic heterogeneity of GSRC and pave the way for future studies exploring its clinical implications and therapeutic strategies.

PMID:40600435 | DOI:10.1021/jasms.4c00505

Advancements in liquid biopsy for breast Cancer: Molecular biomarkers and clinical applications

Cancer Treat Rev. 2025 Jun 14;139:102979. doi: 10.1016/j.ctrv.2025.102979. Online ahead of print.

ABSTRACT

Breast cancer is characterized by significant molecular heterogeneity; therefore, there are distinct clinical features, treatment modalities, and prognostic outcomes across its various molecular subtypes. In the era of precision medicine, liquid biopsy has emerged as a convenient and minimally invasive technique capable of dynamically representing the comprehensive tumor gene spectrum. This review systematically elaborates the clinical value of liquid biopsy as a breakthrough tool for precision diagnosis and treatment in breast cancer through dynamic detection of key biomarkers, including circulating tumor DNA (ctDNA), circulating tumor cells (CTCs), exosomes, and non-coding RNA (ncRNA). Specific genetic mutations and methylation signatures in ctDNA can be applied to early breast cancer screening, minimal residual disease monitoring, and tracking drug resistance mechanisms. CTCs enumeration (≥1/7.5 mL in early-stage cancer or ≥ 5/7.5 mL in metastatic cancer) and PD-L1 expression levels demonstrate direct correlations with prognostic stratification and the efficacy of immunotherapy. As the specificity and sensitivity of liquid biopsy continue to improve, personalized treatment strategies, informed by biomarker analysis and targeted precision therapies, have unveiled new avenues of hope for patients with breast cancer. However, several challenges persist in the practical application of liquid biopsy. Despite persistent challenges, such as insufficient standardization and difficulties in resolving low-abundance variants, future advancements should focus on multi-omics integration and AI-driven technological breakthroughs to overcome bottlenecks in clinical translation. This review summarizes cutting-edge liquid biopsy technologies for identifying clinically significant molecular biomarkers, focusing on discussing critical challenges in the strategies to advance precision oncology applications for optimized treatment guidance and disease surveillance in breast cancer.

PMID:40540857 | DOI:10.1016/j.ctrv.2025.102979

Purine salvage–associated metabolites as biomarkers for early diagnosis of esophageal squamous cell carcinoma: a diagnostic model–based study

Cell Death Discovery, Published online: 14 March 2024; doi:10.1038/s41420-024-01896-6

Purine salvage–associated metabolites as biomarkers for early diagnosis of esophageal squamous cell carcinoma: a diagnostic model–based study
❌