❌

Reading view

SciAgent: A Unified Multi-Agent System for Generalistic Scientific Reasoning

arXiv:2511.08151v1 Announce Type: new Abstract: Recent advances in large language models have enabled AI systems to achieve expert-level performance on domain-specific scientific tasks, yet these systems remain narrow and handcrafted. We introduce SciAgent, a unified multi-agent system designed for generalistic scientific reasoning-the ability to adapt reasoning strategies across disciplines and difficulty levels. SciAgent organizes problem solving as a hierarchical process: a Coordinator Agent interprets each problem's domain and complexity, dynamically orchestrating specialized Worker Systems, each composed of interacting reasoning Sub-agents for symbolic deduction, conceptual modeling, numerical computation, and verification. These agents collaboratively assemble and refine reasoning pipelines tailored to each task. Across mathematics and physics Olympiads (IMO, IMC, IPhO, CPhO), SciAgent consistently attains or surpasses human gold-medalist performance, demonstrating both domain generality and reasoning adaptability. Additionally, SciAgent has been tested on the International Chemistry Olympiad (IChO) and selected problems from the Humanity's Last Exam (HLE) benchmark, further confirming the system's ability to generalize across diverse scientific domains. This work establishes SciAgent as a concrete step toward generalistic scientific intelligence-AI systems capable of coherent, cross-disciplinary reasoning at expert levels.
  •  

Targeted inhibition of gastric adenocarcinoma by nano-curcumin liposomes: Insights from combined machine learning and experimental analyses into the mechanisms of cuproptosis and metabolic reprogramming

Int J Pharm. 2025 Nov 9:126368. doi: 10.1016/j.ijpharm.2025.126368. Online ahead of print.

ABSTRACT

PURPOSE: Gastric adenocarcinoma is a highly aggressive malignancy characterized by a complex tumor microenvironment. Nano-curcumin liposomes hold great potential in inhibiting tumor growth and survival, as well as inducing cuproptosis and oxidative stress. Although the anticancer properties of curcumin have been demonstrated, the specific mechanisms by which curcumin inhibites gastric adenocarcinoma through cuproptosis remains unclear. This study investigated how nano-curcumin liposomes mediated the inhibition of gastric adenocarcinoma cell proliferation and survival via cuproptosis.

METHODS: This study utilized the gastric adenocarcinoma cell line AGS to establish 2D and 3D in vitro gastric adenocarcinoma models. Furthermore, we prepared nano-curcumin liposomes to investigate their effects and regulatory mechanisms on AGS gastric adenocarcinoma models. A series of in vitro assays, including flow cytometry, CCK-8, scratch assays and morphological assessments, were performed to evaluate the effects of nano-curcumin liposomes on cell apoptosis, proliferation and migration. Additionally, bioinformatics and machine learning methods were employed to identify key targets that inhibited gastric adenocarcinoma growth and survival associated with nano-curcumin liposomes, which were further validated through RT-qPCR and omics analysis. Computer simulations were also conducted to assess the stability of binding interactions between curcumin and key target proteins.

RESULTS: Cellular experiments demonstrated that nano-curcumin liposomes significantly inhibited proliferation and invasive capacity of gastric adenocarcinoma cells while promoting cellular oxidative stress. Bioinformatics and machine learning analyses identified FDX1, GPX4, SERPINE1 and SLC27A5 as key targets. RT-qPCR results confirmed that nano-curcumin liposomes significantly downregulated the expression of these targets. Molecular dynamics simulations indicated that curcumin could form stable binding interactions with key protein targets.

CONCLUSION: This study revealed that nano-curcumin liposomes inhibited growth and survival of gastric adenocarcinoma cells by interfering with the expression of FDX1, GPX4, SERPINE1 and SLC27A5, which were closely linked to copper-induced oxidative stress. Nano-curcumin liposomes downregulated the expression of FDX1 and GPX4, disrupted mitochondrial energy metabolism, and induced oxidative stress, thereby promoting tumor-associated programmed cell death linked to cuproptosis. Furthermore, by downregulating SERPINE1, nano-curcumin liposomes modulated cell adhesion and migration, inhibiting the invasive and metastatic potential of tumor cells. Finally, downregulation of SLC27A5 altered tumor metabolism and cellular homeostasis, induced oxidative stress, and disrupted intracellular environmental stability, thereby suppressing the growth of gastric adenocarcinoma.

PMID:41218732 | DOI:10.1016/j.ijpharm.2025.126368

  •  

Self-Correction Distillation for Structured Data Question Answering

arXiv:2511.07998v1 Announce Type: cross Abstract: Structured data question answering (QA), including table QA, Knowledge Graph (KG) QA, and temporal KG QA, is a pivotal research area. Advances in large language models (LLMs) have driven significant progress in unified structural QA frameworks like TrustUQA. However, these frameworks face challenges when applied to small-scale LLMs since small-scale LLMs are prone to errors in generating structured queries. To improve the structured data QA ability of small-scale LLMs, we propose a self-correction distillation (SCD) method. In SCD, an error prompt mechanism (EPM) is designed to detect errors and provide customized error messages during inference, and a two-stage distillation strategy is designed to transfer large-scale LLMs' query-generation and error-correction capabilities to small-scale LLM. Experiments across 5 benchmarks with 3 structured data types demonstrate that our SCD achieves the best performance and superior generalization on small-scale LLM (8B) compared to other distillation methods, and closely approaches the performance of GPT4 on some datasets. Furthermore, large-scale LLMs equipped with EPM surpass the state-of-the-art results on most datasets.
  •  

SCoTT: Strategic Chain-of-Thought Tasking for Wireless-Aware Robot Navigation in Digital Twins

arXiv:2411.18212v3 Announce Type: replace-cross Abstract: Path planning under wireless performance constraints is a complex challenge in robot navigation. However, naively incorporating such constraints into classical planning algorithms often incurs prohibitive search costs. In this paper, we propose SCoTT, a wireless-aware path planning framework that leverages vision-language models (VLMs) to co-optimize average path gains and trajectory length using wireless heatmap images and ray-tracing data from a digital twin (DT). At the core of our framework is Strategic Chain-of-Thought Tasking (SCoTT), a novel prompting paradigm that decomposes the exhaustive search problem into structured subtasks, each solved via chain-of-thought prompting. To establish strong baselines, we compare classical A* and wireless-aware extensions of it, and derive DP-WA*, an optimal, iterative dynamic programming algorithm that incorporates all path gains and distance metrics from the DT, but at significant computational cost. In extensive experiments, we show that SCoTT achieves path gains within 2% of DP-WA* while consistently generating shorter trajectories. Moreover, SCoTT's intermediate outputs can be used to accelerate DP-WA* by reducing its search space, saving up to 62% in execution time. We validate our framework using four VLMs, demonstrating effectiveness across both large and small models, thus making it applicable to a wide range of compact models at low inference cost. We also show the practical viability of our approach by deploying SCoTT as a ROS node within Gazebo simulations. Finally, we discuss data acquisition pipelines, compute requirements, and deployment considerations for VLMs in 6G-enabled DTs, underscoring the potential of natural language interfaces for wireless-aware navigation in real-world applications.
  •  

Digital Health Technologies for Screening and Identifying Unmet Social Needs: Scoping Review

Background: Social determinants of health (SDOH) strongly influence clinical outcomes. Social needs are the individual-level, actionable facets of the broader SDOH framework, including food security, stable housing, and access to essential services. When these needs go unmet, they adversely affect wellbeing and quality of care. Systematically detecting social needs is therefore critical, and emerging digital tools now offer efficient, scalable approaches for screening and identification. Objective: This scoping review aims to examine digital health technology (DHT) use or interventions documented for screening and identifying unmet social needs within high-need populations. We explore trends, effects, challenges, and limitations of identified technologies. Methods: Following PRISMA-ScR guidelines, we searched databases including MEDLINE, Embase, Scopus, ACM Digital Library, and Web of Science for studies published from 2010 to 2025. Eligible studies used technology to screen for and identify unmet social needs in populations with health and socioeconomic challenges. Data extraction focused on the types of technology, screening processes, and social needs identified. Results: Our findings highlight a limited yet evolving landscape of technological applications. We identified 14 studies using tools like self-assessment surveys, tablet-based systems, and electronic portals. These tools were applied across diverse groups, such as refugees and patients in emergency departments. Innovative approaches, such as chatbots and multi-dimensional risk appraisal systems for older adults, showed potential. However, challenges included single-site studies, small samples, and integration issues with medical records. The effectiveness of these tools in screening for unmet social needs shows mixed outcomes. Conclusions: DHTs play a pivotal role in improving the identification of unmet social needs. The findings underscore the need for broader, more integrated research to fully understand the impact of technology-based assessments and screening processes for social needs. Future efforts should focus on facilitated screening using technology both within and outside of the visit, ensuring the linkage to appropriate resources and care.
  •  

STAT+: Chinese government’s support for biotech fuels huge rally

Want to stay on top of the science and politics driving biotech today? Sign up to get our biotech newsletter in your inbox.

Good morning, we just had our first snow of the season in Chicago, I just ordered a pie for Thanksgiving, and I’m still in denial that the year is almost ending.

Onto the news today.

Continue to STAT+ to read the full story…

© PHILIPPE LOPEZ/AFP/Getty Images

  •  

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework

arXiv:2511.05385v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models' (LLMs) reliability. For flexibility, agentic RAG employs autonomous, multi-round retrieval and reasoning to resolve queries. Although recent agentic RAG has improved via reinforcement learning, they often incur substantial token overhead from search and reasoning processes. This trade-off prioritizes accuracy over efficiency. To address this issue, this work proposes TeaRAG, a token-efficient agentic RAG framework capable of compressing both retrieval content and reasoning steps. 1) First, the retrieved content is compressed by augmenting chunk-based semantic retrieval with a graph retrieval using concise triplets. A knowledge association graph is then built from semantic similarity and co-occurrence. Finally, Personalized PageRank is leveraged to highlight key knowledge within this graph, reducing the number of tokens per retrieval. 2) Besides, to reduce reasoning steps, Iterative Process-aware Direct Preference Optimization (IP-DPO) is proposed. Specifically, our reward function evaluates the knowledge sufficiency by a knowledge matching mechanism, while penalizing excessive reasoning steps. This design can produce high-quality preference-pair datasets, supporting iterative DPO to improve reasoning conciseness. Across six datasets, TeaRAG improves the average Exact Match by 4% and 2% while reducing output tokens by 61% and 59% on Llama3-8B-Instruct and Qwen2.5-14B-Instruct, respectively. Code is available at https://github.com/Applied-Machine-Learning-Lab/TeaRAG.
  •  
  •  

Evaluating Control Protocols for Untrusted AI Agents

arXiv:2511.02997v1 Announce Type: new Abstract: As AI systems become more capable and widely deployed as agents, ensuring their safe operation becomes critical. AI control offers one approach to mitigating the risk from untrusted AI agents by monitoring their actions and intervening or auditing when necessary. Evaluating the safety of these protocols requires understanding both their effectiveness against current attacks and their robustness to adaptive adversaries. In this work, we systematically evaluate a range of control protocols in SHADE-Arena, a dataset of diverse agentic environments. First, we evaluate blue team protocols, including deferral to trusted models, resampling, and deferring on critical actions, against a default attack policy. We find that resampling for incrimination and deferring on critical actions perform best, increasing safety from 50% to 96%. We then iterate on red team strategies against these protocols and find that attack policies with additional affordances, such as knowledge of when resampling occurs or the ability to simulate monitors, can substantially improve attack success rates against our resampling strategy, decreasing safety to 17%. However, deferring on critical actions is highly robust to even our strongest red team strategies, demonstrating the importance of denying attack policies access to protocol internals.
  •  

No-Human in the Loop: Agentic Evaluation at Scale for Recommendation

arXiv:2511.03051v1 Announce Type: new Abstract: Evaluating large language models (LLMs) as judges is increasingly critical for building scalable and trustworthy evaluation pipelines. We present ScalingEval, a large-scale benchmarking study that systematically compares 36 LLMs, including GPT, Gemini, Claude, and Llama, across multiple product categories using a consensus-driven evaluation protocol. Our multi-agent framework aggregates pattern audits and issue codes into ground-truth labels via scalable majority voting, enabling reproducible comparison of LLM evaluators without human annotation. Applied to large-scale complementary-item recommendation, the benchmark reports four key findings: (i) Anthropic Claude 3.5 Sonnet achieves the highest decision confidence; (ii) Gemini 1.5 Pro offers the best overall performance across categories; (iii) GPT-4o provides the most favorable latency-accuracy-cost tradeoff; and (iv) GPT-OSS 20B leads among open-source models. Category-level analysis shows strong consensus in structured domains (Electronics, Sports) but persistent disagreement in lifestyle categories (Clothing, Food). These results establish ScalingEval as a reproducible benchmark and evaluation protocol for LLMs as judges, with actionable guidance on scaling, reliability, and model family tradeoffs.
  •  

Explaining Decisions in ML Models: a Parameterized Complexity Analysis (Part I)

arXiv:2511.03545v1 Announce Type: new Abstract: This paper presents a comprehensive theoretical investigation into the parameterized complexity of explanation problems in various machine learning (ML) models. Contrary to the prevalent black-box perception, our study focuses on models with transparent internal mechanisms. We address two principal types of explanation problems: abductive and contrastive, both in their local and global variants. Our analysis encompasses diverse ML models, including Decision Trees, Decision Sets, Decision Lists, Boolean Circuits, and ensembles thereof, each offering unique explanatory challenges. This research fills a significant gap in explainable AI (XAI) by providing a foundational understanding of the complexities of generating explanations for these models. This work provides insights vital for further research in the domain of XAI, contributing to the broader discourse on the necessity of transparency and accountability in AI systems.
  •  

FP-AbDiff: Improving Score-based Antibody Design by Capturing Nonequilibrium Dynamics through the Underlying Fokker-Planck Equation

arXiv:2511.03113v1 Announce Type: cross Abstract: Computational antibody design holds immense promise for therapeutic discovery, yet existing generative models are fundamentally limited by two core challenges: (i) a lack of dynamical consistency, which yields physically implausible structures, and (ii) poor generalization due to data scarcity and structural bias. We introduce FP-AbDiff, the first antibody generator to enforce Fokker-Planck Equation (FPE) physics along the entire generative trajectory. Our method minimizes a novel FPE residual loss over the mixed manifold of CDR geometries (R^3 x SO(3)), compelling locally-learned denoising scores to assemble into a globally coherent probability flow. This physics-informed regularizer is synergistically integrated with deep biological priors within a state-of-the-art SE(3)-equivariant diffusion framework. Rigorous evaluation on the RAbD benchmark confirms that FP-AbDiff establishes a new state-of-the-art. In de novo CDR-H3 design, it achieves a mean Root Mean Square Deviation of 0.99 {\AA} when superposing on the variable region, a 25% improvement over the previous state-of-the-art model, AbX, and the highest reported Contact Amino Acid Recovery of 39.91%. This superiority is underscored in the more challenging six-CDR co-design task, where our model delivers consistently superior geometric precision, cutting the average full-chain Root Mean Square Deviation by ~15%, and crucially, achieves the highest full-chain Amino Acid Recovery on the functionally dominant CDR-H3 loop (45.67%). By aligning generative dynamics with physical laws, FP-AbDiff enhances robustness and generalizability, establishing a principled approach for physically faithful and functionally viable antibody design.
  •  

LGM: Enhancing Large Language Models with Conceptual Meta-Relations and Iterative Retrieval

arXiv:2511.03214v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit strong semantic understanding, yet struggle when user instructions involve ambiguous or conceptually misaligned terms. We propose the Language Graph Model (LGM) to enhance conceptual clarity by extracting meta-relations-inheritance, alias, and composition-from natural language. The model further employs a reflection mechanism to validate these meta-relations. Leveraging a Concept Iterative Retrieval Algorithm, these relations and related descriptions are dynamically supplied to the LLM, improving its ability to interpret concepts and generate accurate responses. Unlike conventional Retrieval-Augmented Generation (RAG) approaches that rely on extended context windows, our method enables large language models to process texts of any length without the need for truncation. Experiments on standard benchmarks demonstrate that the LGM consistently outperforms existing RAG baselines.
  •  

Hybrid Fact-Checking that Integrates Knowledge Graphs, Large Language Models, and Search-Based Retrieval Agents Improves Interpretable Claim Verification

arXiv:2511.03217v1 Announce Type: cross Abstract: Large language models (LLMs) excel in generating fluent utterances but can lack reliable grounding in verified information. At the same time, knowledge-graph-based fact-checkers deliver precise and interpretable evidence, yet suffer from limited coverage or latency. By integrating LLMs with knowledge graphs and real-time search agents, we introduce a hybrid fact-checking approach that leverages the individual strengths of each component. Our system comprises three autonomous steps: 1) a Knowledge Graph (KG) Retrieval for rapid one - hop lookups in DBpedia, 2) an LM-based classification guided by a task-specific labeling prompt, producing outputs with internal rule-based logic, and 3) a Web Search Agent invoked only when KG coverage is insufficient. Our pipeline achieves an F1 score of 0.93 on the FEVER benchmark on the Supported/Refuted split without task- specific fine - tuning. To address Not enough information cases, we conduct a targeted reannotation study showing that our approach frequently uncovers valid evidence for claims originally labeled as Not Enough Information (NEI), as confirmed by both expert annotators and LLM reviewers. With this paper, we present a modular, opensource fact-checking pipeline with fallback strategies and generalization across datasets.
  •  

REFA: Reference Free Alignment for multi-preference optimization

arXiv:2412.16378v4 Announce Type: replace-cross Abstract: To mitigate reward hacking from response verbosity, modern preference optimization methods are increasingly adopting length normalization (e.g., SimPO, ORPO, LN-DPO). While effective against this bias, we demonstrate that length normalization itself introduces a failure mode: the URSLA shortcut. Here models learn to satisfy the alignment objective by prematurely truncating low-quality responses rather than learning from their semantic content. To address this, we introduce REFA, a new alignment framework that proposes probabilistic control on a structural token that controls termination. Our core innovation is a new class of regularizers that operate directly on the probability of the End-of-Sequence (EOS) token, a previously unexploited control lever. This token-level intervention provides a principled solution to the URSLA shortcut, ensuring genuine quality improvements. Furthermore, it unlocks a versatile mechanism for managing the alignment-efficiency tradeoff, enabling practitioners to fine-tune models that adhere to specific token budgets. Empirically, REFA achieves a 60.29% win rate and a 52.17% length-controlled win rate on AlpacaEval2 with Llama-3-8B-Instruct, demonstrating the power of our token-level control paradigm.
  •  

Harnessing multi-omics approaches to decipher tumor evolution and improve diagnosis and therapy in lung cancer

Biomark Res. 2025 Nov 5;13(1):140. doi: 10.1186/s40364-025-00859-y.

ABSTRACT

With the advancement of novel technologies such as whole-genome sequencing, single-cell sequencing, and spatial transcriptomics, single-omics analyses have already promoted the research of tumorigenesis as well as development and have partly elucidated the evolutionary processes of lung cancer. However, it is still difficult to distinguish these confounding features via single dimensional approaches due to the complexity, heterogeneity and cell-cell interactions with the immune microenvironment in lung cancer. Multi-omics approaches provide a holistic framework for constructing detailed tumor ecosystem landscapes, thereby facilitating the development of a more robust classification system for precision diagnosis and treatment, and aiding in the discovery of novel cancer biomarkers. In this review, we summarize the potential and applications of multi-omics approaches in characterizing intratumor heterogeneity and the tumor microenvironment throughout the course of lung cancer development. By further discussing the discovery and application of diagnostic and therapeutic biomarkers across precancerous lesions, early-stage lung cancer, tumor progression, metastasis, and therapy resistance, we outline the current challenges and future prospects of using multi-omics to identify reliable biomarkers. Moreover, we emphasize that integrative multi-omics models hold great promise for elucidating the complex interactions within the lung cancer ecosystem, thereby contributing to improved diagnostic accuracy, optimized therapeutic strategies, and better patient outcomes.

PMID:41194170 | PMC:PMC12590604 | DOI:10.1186/s40364-025-00859-y

  •  

Improving dataset transparency in dermatologic Artificial Intelligence using a dataset nutrition label

npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02125-9

Biased and poorly documented dermatology datasets pose risks to the development of safe and generalizable artificial intelligence (AI) tools. We created a Dataset Nutrition Label (DNL) for multiple dermatology datasets to support transparent and responsible data use. The DNL offers a structured, digestible summary of key attributes, including metadata, limitations, and risks, enabling data users to better assess suitability and proactively address potential sources of bias in datasets.
  •  
❌