❌

Reading view

SciAgent: A Unified Multi-Agent System for Generalistic Scientific Reasoning

arXiv:2511.08151v1 Announce Type: new Abstract: Recent advances in large language models have enabled AI systems to achieve expert-level performance on domain-specific scientific tasks, yet these systems remain narrow and handcrafted. We introduce SciAgent, a unified multi-agent system designed for generalistic scientific reasoning-the ability to adapt reasoning strategies across disciplines and difficulty levels. SciAgent organizes problem solving as a hierarchical process: a Coordinator Agent interprets each problem's domain and complexity, dynamically orchestrating specialized Worker Systems, each composed of interacting reasoning Sub-agents for symbolic deduction, conceptual modeling, numerical computation, and verification. These agents collaboratively assemble and refine reasoning pipelines tailored to each task. Across mathematics and physics Olympiads (IMO, IMC, IPhO, CPhO), SciAgent consistently attains or surpasses human gold-medalist performance, demonstrating both domain generality and reasoning adaptability. Additionally, SciAgent has been tested on the International Chemistry Olympiad (IChO) and selected problems from the Humanity's Last Exam (HLE) benchmark, further confirming the system's ability to generalize across diverse scientific domains. This work establishes SciAgent as a concrete step toward generalistic scientific intelligence-AI systems capable of coherent, cross-disciplinary reasoning at expert levels.
  •  

Targeted inhibition of gastric adenocarcinoma by nano-curcumin liposomes: Insights from combined machine learning and experimental analyses into the mechanisms of cuproptosis and metabolic reprogramming

Int J Pharm. 2025 Nov 9:126368. doi: 10.1016/j.ijpharm.2025.126368. Online ahead of print.

ABSTRACT

PURPOSE: Gastric adenocarcinoma is a highly aggressive malignancy characterized by a complex tumor microenvironment. Nano-curcumin liposomes hold great potential in inhibiting tumor growth and survival, as well as inducing cuproptosis and oxidative stress. Although the anticancer properties of curcumin have been demonstrated, the specific mechanisms by which curcumin inhibites gastric adenocarcinoma through cuproptosis remains unclear. This study investigated how nano-curcumin liposomes mediated the inhibition of gastric adenocarcinoma cell proliferation and survival via cuproptosis.

METHODS: This study utilized the gastric adenocarcinoma cell line AGS to establish 2D and 3D in vitro gastric adenocarcinoma models. Furthermore, we prepared nano-curcumin liposomes to investigate their effects and regulatory mechanisms on AGS gastric adenocarcinoma models. A series of in vitro assays, including flow cytometry, CCK-8, scratch assays and morphological assessments, were performed to evaluate the effects of nano-curcumin liposomes on cell apoptosis, proliferation and migration. Additionally, bioinformatics and machine learning methods were employed to identify key targets that inhibited gastric adenocarcinoma growth and survival associated with nano-curcumin liposomes, which were further validated through RT-qPCR and omics analysis. Computer simulations were also conducted to assess the stability of binding interactions between curcumin and key target proteins.

RESULTS: Cellular experiments demonstrated that nano-curcumin liposomes significantly inhibited proliferation and invasive capacity of gastric adenocarcinoma cells while promoting cellular oxidative stress. Bioinformatics and machine learning analyses identified FDX1, GPX4, SERPINE1 and SLC27A5 as key targets. RT-qPCR results confirmed that nano-curcumin liposomes significantly downregulated the expression of these targets. Molecular dynamics simulations indicated that curcumin could form stable binding interactions with key protein targets.

CONCLUSION: This study revealed that nano-curcumin liposomes inhibited growth and survival of gastric adenocarcinoma cells by interfering with the expression of FDX1, GPX4, SERPINE1 and SLC27A5, which were closely linked to copper-induced oxidative stress. Nano-curcumin liposomes downregulated the expression of FDX1 and GPX4, disrupted mitochondrial energy metabolism, and induced oxidative stress, thereby promoting tumor-associated programmed cell death linked to cuproptosis. Furthermore, by downregulating SERPINE1, nano-curcumin liposomes modulated cell adhesion and migration, inhibiting the invasive and metastatic potential of tumor cells. Finally, downregulation of SLC27A5 altered tumor metabolism and cellular homeostasis, induced oxidative stress, and disrupted intracellular environmental stability, thereby suppressing the growth of gastric adenocarcinoma.

PMID:41218732 | DOI:10.1016/j.ijpharm.2025.126368

  •  

HybridGuard: Enhancing Minority-Class Intrusion Detection in Dew-Enabled Edge-of-Things Networks

arXiv:2511.07793v1 Announce Type: cross Abstract: Securing Dew-Enabled Edge-of-Things (EoT) networks against sophisticated intrusions is a critical challenge. This paper presents HybridGuard, a framework that integrates machine learning and deep learning to improve intrusion detection. HybridGuard addresses data imbalance through mutual information based feature selection, ensuring that the most relevant features are used to improve detection performance, especially for minority attack classes. The framework leverages Wasserstein Conditional Generative Adversarial Networks with Gradient Penalty (WCGAN-GP) to further reduce class imbalance and enhance detection precision. It adopts a two-phase architecture called DualNetShield to support advanced traffic analysis and anomaly detection, improving the granular identification of threats in complex EoT environments. HybridGuard is evaluated on the UNSW-NB15, CIC-IDS-2017, and IOTID20 datasets, where it demonstrates strong performance across diverse attack scenarios and outperforms existing solutions in adapting to evolving cybersecurity threats. This approach establishes HybridGuard as an effective tool for protecting EoT networks against modern intrusions.
  •  

Self-Correction Distillation for Structured Data Question Answering

arXiv:2511.07998v1 Announce Type: cross Abstract: Structured data question answering (QA), including table QA, Knowledge Graph (KG) QA, and temporal KG QA, is a pivotal research area. Advances in large language models (LLMs) have driven significant progress in unified structural QA frameworks like TrustUQA. However, these frameworks face challenges when applied to small-scale LLMs since small-scale LLMs are prone to errors in generating structured queries. To improve the structured data QA ability of small-scale LLMs, we propose a self-correction distillation (SCD) method. In SCD, an error prompt mechanism (EPM) is designed to detect errors and provide customized error messages during inference, and a two-stage distillation strategy is designed to transfer large-scale LLMs' query-generation and error-correction capabilities to small-scale LLM. Experiments across 5 benchmarks with 3 structured data types demonstrate that our SCD achieves the best performance and superior generalization on small-scale LLM (8B) compared to other distillation methods, and closely approaches the performance of GPT4 on some datasets. Furthermore, large-scale LLMs equipped with EPM surpass the state-of-the-art results on most datasets.
  •  

Clinical Uncertainty Impacts Machine Learning Evaluations

arXiv:2509.22242v2 Announce Type: replace Abstract: Clinical dataset labels are rarely certain as annotators disagree and confidence is not uniform across cases. Typical aggregation procedures, such as majority voting, obscure this variability. In simple experiments on medical imaging benchmarks, accounting for the confidence in binary labels significantly impacts model rankings. We therefore argue that machine-learning evaluations should explicitly account for annotation uncertainty using probabilistic metrics that directly operate on distributions. These metrics can be applied independently of the annotations' generating process, whether modeled by simple counting, subjective confidence ratings, or probabilistic response models. They are also computationally lightweight, as closed-form expressions have linear-time implementations once examples are sorted by model score. We thus urge the community to release raw annotations for datasets and to adopt uncertainty-aware evaluation so that performance estimates may better reflect clinical data.
  •  

SCoTT: Strategic Chain-of-Thought Tasking for Wireless-Aware Robot Navigation in Digital Twins

arXiv:2411.18212v3 Announce Type: replace-cross Abstract: Path planning under wireless performance constraints is a complex challenge in robot navigation. However, naively incorporating such constraints into classical planning algorithms often incurs prohibitive search costs. In this paper, we propose SCoTT, a wireless-aware path planning framework that leverages vision-language models (VLMs) to co-optimize average path gains and trajectory length using wireless heatmap images and ray-tracing data from a digital twin (DT). At the core of our framework is Strategic Chain-of-Thought Tasking (SCoTT), a novel prompting paradigm that decomposes the exhaustive search problem into structured subtasks, each solved via chain-of-thought prompting. To establish strong baselines, we compare classical A* and wireless-aware extensions of it, and derive DP-WA*, an optimal, iterative dynamic programming algorithm that incorporates all path gains and distance metrics from the DT, but at significant computational cost. In extensive experiments, we show that SCoTT achieves path gains within 2% of DP-WA* while consistently generating shorter trajectories. Moreover, SCoTT's intermediate outputs can be used to accelerate DP-WA* by reducing its search space, saving up to 62% in execution time. We validate our framework using four VLMs, demonstrating effectiveness across both large and small models, thus making it applicable to a wide range of compact models at low inference cost. We also show the practical viability of our approach by deploying SCoTT as a ROS node within Gazebo simulations. Finally, we discuss data acquisition pipelines, compute requirements, and deployment considerations for VLMs in 6G-enabled DTs, underscoring the potential of natural language interfaces for wireless-aware navigation in real-world applications.
  •  

Digital Health Technologies for Screening and Identifying Unmet Social Needs: Scoping Review

Background: Social determinants of health (SDOH) strongly influence clinical outcomes. Social needs are the individual-level, actionable facets of the broader SDOH framework, including food security, stable housing, and access to essential services. When these needs go unmet, they adversely affect wellbeing and quality of care. Systematically detecting social needs is therefore critical, and emerging digital tools now offer efficient, scalable approaches for screening and identification. Objective: This scoping review aims to examine digital health technology (DHT) use or interventions documented for screening and identifying unmet social needs within high-need populations. We explore trends, effects, challenges, and limitations of identified technologies. Methods: Following PRISMA-ScR guidelines, we searched databases including MEDLINE, Embase, Scopus, ACM Digital Library, and Web of Science for studies published from 2010 to 2025. Eligible studies used technology to screen for and identify unmet social needs in populations with health and socioeconomic challenges. Data extraction focused on the types of technology, screening processes, and social needs identified. Results: Our findings highlight a limited yet evolving landscape of technological applications. We identified 14 studies using tools like self-assessment surveys, tablet-based systems, and electronic portals. These tools were applied across diverse groups, such as refugees and patients in emergency departments. Innovative approaches, such as chatbots and multi-dimensional risk appraisal systems for older adults, showed potential. However, challenges included single-site studies, small samples, and integration issues with medical records. The effectiveness of these tools in screening for unmet social needs shows mixed outcomes. Conclusions: DHTs play a pivotal role in improving the identification of unmet social needs. The findings underscore the need for broader, more integrated research to fully understand the impact of technology-based assessments and screening processes for social needs. Future efforts should focus on facilitated screening using technology both within and outside of the visit, ensuring the linkage to appropriate resources and care.
  •  

STAT+: Chinese government’s support for biotech fuels huge rally

Want to stay on top of the science and politics driving biotech today? Sign up to get our biotech newsletter in your inbox.

Good morning, we just had our first snow of the season in Chicago, I just ordered a pie for Thanksgiving, and I’m still in denial that the year is almost ending.

Onto the news today.

Continue to STAT+ to read the full story…

© PHILIPPE LOPEZ/AFP/Getty Images

  •  

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework

arXiv:2511.05385v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models' (LLMs) reliability. For flexibility, agentic RAG employs autonomous, multi-round retrieval and reasoning to resolve queries. Although recent agentic RAG has improved via reinforcement learning, they often incur substantial token overhead from search and reasoning processes. This trade-off prioritizes accuracy over efficiency. To address this issue, this work proposes TeaRAG, a token-efficient agentic RAG framework capable of compressing both retrieval content and reasoning steps. 1) First, the retrieved content is compressed by augmenting chunk-based semantic retrieval with a graph retrieval using concise triplets. A knowledge association graph is then built from semantic similarity and co-occurrence. Finally, Personalized PageRank is leveraged to highlight key knowledge within this graph, reducing the number of tokens per retrieval. 2) Besides, to reduce reasoning steps, Iterative Process-aware Direct Preference Optimization (IP-DPO) is proposed. Specifically, our reward function evaluates the knowledge sufficiency by a knowledge matching mechanism, while penalizing excessive reasoning steps. This design can produce high-quality preference-pair datasets, supporting iterative DPO to improve reasoning conciseness. Across six datasets, TeaRAG improves the average Exact Match by 4% and 2% while reducing output tokens by 61% and 59% on Llama3-8B-Instruct and Qwen2.5-14B-Instruct, respectively. Code is available at https://github.com/Applied-Machine-Learning-Lab/TeaRAG.
  •  

AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science

arXiv:2502.16395v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to automate data analysis through executable code generation. Yet, data science tasks often admit multiple statistically valid solutions, e.g. different modeling strategies, making it critical to understand the reasoning behind analyses, not just their outcomes. While manual review of LLM-generated code can help ensure statistical soundness, it is labor-intensive and requires expertise. A more scalable approach is to evaluate the underlying workflows-the logical plans guiding code generation. However, it remains unclear how to assess whether an LLM-generated workflow supports reproducible implementations. To address this, we present AIRepr, an Analyst-Inspector framework for automatically evaluating and improving the reproducibility of LLM-generated data analysis workflows. Our framework is grounded in statistical principles and supports scalable, automated assessment. We introduce two novel reproducibility-enhancing prompting strategies and benchmark them against standard prompting across 15 analyst-inspector LLM pairs and 1,032 tasks from three public benchmarks. Our findings show that workflows with higher reproducibility also yield more accurate analyses, and that reproducibility-enhancing prompts substantially improve both metrics. This work provides a foundation for transparent, reliable, and efficient human-AI collaboration in data science. Our code is publicly available.
  •  
  •  

A Mega-Study of Digital Twins Reveals Strengths, Weaknesses and Opportunities for Further Improvement

arXiv:2509.19088v3 Announce Type: replace-cross Abstract: Digital representations of individuals ("digital twins") promise to transform social science and decision-making. Yet it remains unclear whether such twins truly mirror the people they emulate. We conducted 19 preregistered studies with a representative U.S. panel and their digital twins, each constructed from rich individual-level data, enabling direct comparisons between human and twin behavior across a wide range of domains and stimuli (including never-seen-before ones). Twins reproduced individual responses with 75% accuracy and seemingly low correlation with human answers (approximately 0.2). However, this apparently high accuracy was no higher than that achieved by generic personas based on demographics only. In contrast, correlation improved when twins incorporated detailed personal information, even outperforming traditional machine learning benchmarks that require additional data. Twins exhibited systematic strengths and weaknesses - performing better in social and personality domains, but worse in political ones - and were more accurate for participants with higher education, higher income, and moderate political views and religious attendance. Together, these findings delineate both the promise and the current limits of digital twins: they capture some relative differences among individuals but not yet the unique judgments of specific people. All data and code are publicly available to support the further development and evaluation of digital twin pipelines.
  •  

Evaluating Control Protocols for Untrusted AI Agents

arXiv:2511.02997v1 Announce Type: new Abstract: As AI systems become more capable and widely deployed as agents, ensuring their safe operation becomes critical. AI control offers one approach to mitigating the risk from untrusted AI agents by monitoring their actions and intervening or auditing when necessary. Evaluating the safety of these protocols requires understanding both their effectiveness against current attacks and their robustness to adaptive adversaries. In this work, we systematically evaluate a range of control protocols in SHADE-Arena, a dataset of diverse agentic environments. First, we evaluate blue team protocols, including deferral to trusted models, resampling, and deferring on critical actions, against a default attack policy. We find that resampling for incrimination and deferring on critical actions perform best, increasing safety from 50% to 96%. We then iterate on red team strategies against these protocols and find that attack policies with additional affordances, such as knowledge of when resampling occurs or the ability to simulate monitors, can substantially improve attack success rates against our resampling strategy, decreasing safety to 17%. However, deferring on critical actions is highly robust to even our strongest red team strategies, demonstrating the importance of denying attack policies access to protocol internals.
  •  

No-Human in the Loop: Agentic Evaluation at Scale for Recommendation

arXiv:2511.03051v1 Announce Type: new Abstract: Evaluating large language models (LLMs) as judges is increasingly critical for building scalable and trustworthy evaluation pipelines. We present ScalingEval, a large-scale benchmarking study that systematically compares 36 LLMs, including GPT, Gemini, Claude, and Llama, across multiple product categories using a consensus-driven evaluation protocol. Our multi-agent framework aggregates pattern audits and issue codes into ground-truth labels via scalable majority voting, enabling reproducible comparison of LLM evaluators without human annotation. Applied to large-scale complementary-item recommendation, the benchmark reports four key findings: (i) Anthropic Claude 3.5 Sonnet achieves the highest decision confidence; (ii) Gemini 1.5 Pro offers the best overall performance across categories; (iii) GPT-4o provides the most favorable latency-accuracy-cost tradeoff; and (iv) GPT-OSS 20B leads among open-source models. Category-level analysis shows strong consensus in structured domains (Electronics, Sports) but persistent disagreement in lifestyle categories (Clothing, Food). These results establish ScalingEval as a reproducible benchmark and evaluation protocol for LLMs as judges, with actionable guidance on scaling, reliability, and model family tradeoffs.
  •  

Explaining Decisions in ML Models: a Parameterized Complexity Analysis (Part I)

arXiv:2511.03545v1 Announce Type: new Abstract: This paper presents a comprehensive theoretical investigation into the parameterized complexity of explanation problems in various machine learning (ML) models. Contrary to the prevalent black-box perception, our study focuses on models with transparent internal mechanisms. We address two principal types of explanation problems: abductive and contrastive, both in their local and global variants. Our analysis encompasses diverse ML models, including Decision Trees, Decision Sets, Decision Lists, Boolean Circuits, and ensembles thereof, each offering unique explanatory challenges. This research fills a significant gap in explainable AI (XAI) by providing a foundational understanding of the complexities of generating explanations for these models. This work provides insights vital for further research in the domain of XAI, contributing to the broader discourse on the necessity of transparency and accountability in AI systems.
  •  

Digital Transformation Chatbot (DTchatbot): Integrating Large Language Model-based Chatbot in Acquiring Digital Transformation Needs

arXiv:2511.02842v1 Announce Type: cross Abstract: Many organisations pursue digital transformation to enhance operational efficiency, reduce manual efforts, and optimise processes by automation and digital tools. To achieve this, a comprehensive understanding of their unique needs is required. However, traditional methods, such as expert interviews, while effective, face several challenges, including scheduling conflicts, resource constraints, inconsistency, etc. To tackle these issues, we investigate the use of a Large Language Model (LLM)-powered chatbot to acquire organisations' digital transformation needs. Specifically, the chatbot integrates workflow-based instruction with LLM's planning and reasoning capabilities, enabling it to function as a virtual expert and conduct interviews. We detail the chatbot's features and its implementation. Our preliminary evaluation indicates that the chatbot performs as designed, effectively following predefined workflows and supporting user interactions with areas for improvement. We conclude by discussing the implications of employing chatbots to elicit user information, emphasizing their potential and limitations.
  •  
❌