❌

Normal view

SciAgent: A Unified Multi-Agent System for Generalistic Scientific Reasoning

arXiv:2511.08151v1 Announce Type: new Abstract: Recent advances in large language models have enabled AI systems to achieve expert-level performance on domain-specific scientific tasks, yet these systems remain narrow and handcrafted. We introduce SciAgent, a unified multi-agent system designed for generalistic scientific reasoning-the ability to adapt reasoning strategies across disciplines and difficulty levels. SciAgent organizes problem solving as a hierarchical process: a Coordinator Agent interprets each problem's domain and complexity, dynamically orchestrating specialized Worker Systems, each composed of interacting reasoning Sub-agents for symbolic deduction, conceptual modeling, numerical computation, and verification. These agents collaboratively assemble and refine reasoning pipelines tailored to each task. Across mathematics and physics Olympiads (IMO, IMC, IPhO, CPhO), SciAgent consistently attains or surpasses human gold-medalist performance, demonstrating both domain generality and reasoning adaptability. Additionally, SciAgent has been tested on the International Chemistry Olympiad (IChO) and selected problems from the Humanity's Last Exam (HLE) benchmark, further confirming the system's ability to generalize across diverse scientific domains. This work establishes SciAgent as a concrete step toward generalistic scientific intelligence-AI systems capable of coherent, cross-disciplinary reasoning at expert levels.

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework

arXiv:2511.05385v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models' (LLMs) reliability. For flexibility, agentic RAG employs autonomous, multi-round retrieval and reasoning to resolve queries. Although recent agentic RAG has improved via reinforcement learning, they often incur substantial token overhead from search and reasoning processes. This trade-off prioritizes accuracy over efficiency. To address this issue, this work proposes TeaRAG, a token-efficient agentic RAG framework capable of compressing both retrieval content and reasoning steps. 1) First, the retrieved content is compressed by augmenting chunk-based semantic retrieval with a graph retrieval using concise triplets. A knowledge association graph is then built from semantic similarity and co-occurrence. Finally, Personalized PageRank is leveraged to highlight key knowledge within this graph, reducing the number of tokens per retrieval. 2) Besides, to reduce reasoning steps, Iterative Process-aware Direct Preference Optimization (IP-DPO) is proposed. Specifically, our reward function evaluates the knowledge sufficiency by a knowledge matching mechanism, while penalizing excessive reasoning steps. This design can produce high-quality preference-pair datasets, supporting iterative DPO to improve reasoning conciseness. Across six datasets, TeaRAG improves the average Exact Match by 4% and 2% while reducing output tokens by 61% and 59% on Llama3-8B-Instruct and Qwen2.5-14B-Instruct, respectively. Code is available at https://github.com/Applied-Machine-Learning-Lab/TeaRAG.

Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-In-The-Loop LLM

arXiv:2410.14879v4 Announce Type: replace-cross Abstract: Passive tracking methods, such as phone and wearable sensing, have become dominant in monitoring human behaviors in modern ubiquitous computing studies. While there have been significant advances in machine-learning approaches to translate periods of raw sensor data to model momentary behaviors, (e.g., physical activity recognition), there still remains a significant gap in the translation of these sensing streams into meaningful, high-level, context-aware insights that are required for various applications (e.g., summarizing an individual's daily routine). To bridge this gap, experts often need to employ a context-driven sensemaking process in real-world studies to derive insights. This process often requires manual effort and can be challenging even for experienced researchers due to the complexity of human behaviors. We conducted three rounds of user studies with 21 experts to explore solutions to address challenges with sensemaking. We follow a human-centered design process to identify needs and design, iterate, build, and evaluate Vital Insight (VI), a novel, LLM-assisted, prototype system to enable human-in-the-loop inference (sensemaking) and visualizations of multi-modal passive sensing data from smartphones and wearables. Using the prototype as a technology probe, we observe experts' interactions with it and develop an expert sensemaking model that explains how experts move between direct data representations and AI-supported inferences to explore, question, and validate insights. Through this iterative process, we also synthesize and discuss a list of design implications for the design of future AI-augmented visualization systems to better assist experts' sensemaking processes in multi-modal health sensing data.

Language Ranker: A Lightweight Ranking framework for LLM Decoding

arXiv:2510.21883v1 Announce Type: cross Abstract: Conventional research on large language models (LLMs) has primarily focused on refining output distributions, while paying less attention to the decoding process that transforms these distributions into final responses. Recent advances, such as scaling the computation of inference time with reward models, have underscored the importance of decoding, but these methods often suffer from high computational costs and limited applicability. In this paper, we revisit LLM generation through the lens of recommender systems, conceptualizing the decoding process as analogous to the ranking stage in recommendation pipelines. From this perspective, we observe that both traditional decoding methods and reward models exhibit clear limitations such as redundancy. Motivated by this insight, we propose Language Ranker, a novel framework that introduces a lightweight module to rerank candidate responses using features extracted by the base model. Experiments across a wide range of tasks show that Language Ranker achieves performance comparable to large-scale reward models, while requiring only

Cell-free epigenomes enhanced fragmentomics-based model for early detection of lung cancer

Clin Transl Med. 2025 Feb;15(2):e70225. doi: 10.1002/ctm2.70225.

ABSTRACT

BACKGROUND: Lung cancer is a leading cause of cancer mortality, highlighting the need for innovative non-invasive early detection methods. Although cell-free DNA (cfDNA) analysis shows promise, its sensitivity in early-stage lung cancer patients remains a challenge. This study aimed to integrate insights from epigenetic modifications and fragmentomic features of cfDNA using machine learning to develop a more accurate lung cancer detection model.

METHODS: To address this issue, a multi-centre prospective cohort study was conducted, with participants harbouring suspicious malignant lung nodules and healthy volunteers recruited from two clinical centres. Plasma cfDNA was analysed for its epigenetic and fragmentomic profiles using chromatin immunoprecipitation sequencing, reduced representation bisulphite sequencing and low-pass whole-genome sequencing. Machine learning algorithms were then employed to integrate the multi-omics data, aiding in the development of a precise lung cancer detection model.

RESULTS: Cancer-related changes in cfDNA fragmentomics were significantly enriched in specific genes marked by cell-free epigenomes. A total of 609 genes were identified, and the corresponding cfDNA fragmentomic features were utilised to construct the ensemble model. This model achieved a sensitivity of 90.4% and a specificity of 83.1%, with an AUC of 0.94 in the independent validation set. Notably, the model demonstrated exceptional sensitivity for stage I lung cancer cases, achieving 95.1%. It also showed remarkable performance in detecting minimally invasive adenocarcinoma, with a sensitivity of 96.2%, highlighting its potential for early detection in clinical settings.

CONCLUSIONS: With feature selection guided by multiple epigenetic sequencing approaches, the cfDNA fragmentomics-based machine learning model demonstrated outstanding performance in the independent validation cohort. These findings highlight its potential as an effective non-invasive strategy for the early detection of lung cancer.

KEYPOINTS: Our study elucidated the regulatory relationships between epigenetic modifications and their effects on fragmentomic features. Identifying epigenetically regulated genes provided a critical foundation for developing the cfDNA fragmentomics-based machine learning model. The model demonstrated exceptional clinical performance, highlighting its substantial potential for translational application in clinical practice.

PMID:39909829 | PMC:PMC11798665 | DOI:10.1002/ctm2.70225

❌