Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
Evaluating Control Protocols for Untrusted AI Agents
arXiv:2511.02997v1 Announce Type: new Abstract: As AI systems become more capable and widely deployed as agents, ensuring their safe operation becomes critical. AI control offers one approach to mitigating the risk from untrusted AI agents by monitoring their actions and intervening or auditing when necessary. Evaluating the safety of these protocols requires understanding both their effectiveness against current attacks and their robustness to adaptive adversaries. In this work, we systematica
-
cs.AI, q-bio.NC updates on arXiv.org
-
No-Human in the Loop: Agentic Evaluation at Scale for Recommendation
arXiv:2511.03051v1 Announce Type: new Abstract: Evaluating large language models (LLMs) as judges is increasingly critical for building scalable and trustworthy evaluation pipelines. We present ScalingEval, a large-scale benchmarking study that systematically compares 36 LLMs, including GPT, Gemini, Claude, and Llama, across multiple product categories using a consensus-driven evaluation protocol. Our multi-agent framework aggregates pattern audits and issue codes into ground-truth labels via s
No-Human in the Loop: Agentic Evaluation at Scale for Recommendation
-
cs.AI, q-bio.NC updates on arXiv.org
-
Explaining Decisions in ML Models: a Parameterized Complexity Analysis (Part I)
arXiv:2511.03545v1 Announce Type: new Abstract: This paper presents a comprehensive theoretical investigation into the parameterized complexity of explanation problems in various machine learning (ML) models. Contrary to the prevalent black-box perception, our study focuses on models with transparent internal mechanisms. We address two principal types of explanation problems: abductive and contrastive, both in their local and global variants. Our analysis encompasses diverse ML models, includin
Explaining Decisions in ML Models: a Parameterized Complexity Analysis (Part I)
-
cs.AI, q-bio.NC updates on arXiv.org
-
FP-AbDiff: Improving Score-based Antibody Design by Capturing Nonequilibrium Dynamics through the Underlying Fokker-Planck Equation
arXiv:2511.03113v1 Announce Type: cross Abstract: Computational antibody design holds immense promise for therapeutic discovery, yet existing generative models are fundamentally limited by two core challenges: (i) a lack of dynamical consistency, which yields physically implausible structures, and (ii) poor generalization due to data scarcity and structural bias. We introduce FP-AbDiff, the first antibody generator to enforce Fokker-Planck Equation (FPE) physics along the entire generative traj
FP-AbDiff: Improving Score-based Antibody Design by Capturing Nonequilibrium Dynamics through the Underlying Fokker-Planck Equation
-
cs.AI, q-bio.NC updates on arXiv.org
-
LGM: Enhancing Large Language Models with Conceptual Meta-Relations and Iterative Retrieval
arXiv:2511.03214v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit strong semantic understanding, yet struggle when user instructions involve ambiguous or conceptually misaligned terms. We propose the Language Graph Model (LGM) to enhance conceptual clarity by extracting meta-relations-inheritance, alias, and composition-from natural language. The model further employs a reflection mechanism to validate these meta-relations. Leveraging a Concept Iterative Retrieval Algorithm
LGM: Enhancing Large Language Models with Conceptual Meta-Relations and Iterative Retrieval
-
cs.AI, q-bio.NC updates on arXiv.org
-
Hybrid Fact-Checking that Integrates Knowledge Graphs, Large Language Models, and Search-Based Retrieval Agents Improves Interpretable Claim Verification
arXiv:2511.03217v1 Announce Type: cross Abstract: Large language models (LLMs) excel in generating fluent utterances but can lack reliable grounding in verified information. At the same time, knowledge-graph-based fact-checkers deliver precise and interpretable evidence, yet suffer from limited coverage or latency. By integrating LLMs with knowledge graphs and real-time search agents, we introduce a hybrid fact-checking approach that leverages the individual strengths of each component. Our sys
Hybrid Fact-Checking that Integrates Knowledge Graphs, Large Language Models, and Search-Based Retrieval Agents Improves Interpretable Claim Verification
-
cs.AI, q-bio.NC updates on arXiv.org
-
REFA: Reference Free Alignment for multi-preference optimization
arXiv:2412.16378v4 Announce Type: replace-cross Abstract: To mitigate reward hacking from response verbosity, modern preference optimization methods are increasingly adopting length normalization (e.g., SimPO, ORPO, LN-DPO). While effective against this bias, we demonstrate that length normalization itself introduces a failure mode: the URSLA shortcut. Here models learn to satisfy the alignment objective by prematurely truncating low-quality responses rather than learning from their semantic co
REFA: Reference Free Alignment for multi-preference optimization
-
Omics In Lung
-
Harnessing multi-omics approaches to decipher tumor evolution and improve diagnosis and therapy in lung cancer
Biomark Res. 2025 Nov 5;13(1):140. doi: 10.1186/s40364-025-00859-y.ABSTRACTWith the advancement of novel technologies such as whole-genome sequencing, single-cell sequencing, and spatial transcriptomics, single-omics analyses have already promoted the research of tumorigenesis as well as development and have partly elucidated the evolutionary processes of lung cancer. However, it is still difficult to distinguish these confounding features via single dimensional approaches due to the complexity,
Harnessing multi-omics approaches to decipher tumor evolution and improve diagnosis and therapy in lung cancer
Biomark Res. 2025 Nov 5;13(1):140. doi: 10.1186/s40364-025-00859-y.
ABSTRACT
With the advancement of novel technologies such as whole-genome sequencing, single-cell sequencing, and spatial transcriptomics, single-omics analyses have already promoted the research of tumorigenesis as well as development and have partly elucidated the evolutionary processes of lung cancer. However, it is still difficult to distinguish these confounding features via single dimensional approaches due to the complexity, heterogeneity and cell-cell interactions with the immune microenvironment in lung cancer. Multi-omics approaches provide a holistic framework for constructing detailed tumor ecosystem landscapes, thereby facilitating the development of a more robust classification system for precision diagnosis and treatment, and aiding in the discovery of novel cancer biomarkers. In this review, we summarize the potential and applications of multi-omics approaches in characterizing intratumor heterogeneity and the tumor microenvironment throughout the course of lung cancer development. By further discussing the discovery and application of diagnostic and therapeutic biomarkers across precancerous lesions, early-stage lung cancer, tumor progression, metastasis, and therapy resistance, we outline the current challenges and future prospects of using multi-omics to identify reliable biomarkers. Moreover, we emphasize that integrative multi-omics models hold great promise for elucidating the complex interactions within the lung cancer ecosystem, thereby contributing to improved diagnostic accuracy, optimized therapeutic strategies, and better patient outcomes.
PMID:41194170 | PMC:PMC12590604 | DOI:10.1186/s40364-025-00859-y
-
npj Digital Medicine
-
Improving dataset transparency in dermatologic Artificial Intelligence using a dataset nutrition label
npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02125-9Biased and poorly documented dermatology datasets pose risks to the development of safe and generalizable artificial intelligence (AI) tools. We created a Dataset Nutrition Label (DNL) for multiple dermatology datasets to support transparent and responsible data use. The DNL offers a structured, digestible summary of key attributes, including metadata, limitations, and risks, enabling data users to better ass
Improving dataset transparency in dermatologic Artificial Intelligence using a dataset nutrition label
npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02125-9
Biased and poorly documented dermatology datasets pose risks to the development of safe and generalizable artificial intelligence (AI) tools. We created a Dataset Nutrition Label (DNL) for multiple dermatology datasets to support transparent and responsible data use. The DNL offers a structured, digestible summary of key attributes, including metadata, limitations, and risks, enabling data users to better assess suitability and proactively address potential sources of bias in datasets.-
npj Digital Medicine
-
Evaluating clinical AI summaries with large language models as judges
npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02005-2Evaluating clinical AI summaries with large language models as judges
Evaluating clinical AI summaries with large language models as judges
npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02005-2
Evaluating clinical AI summaries with large language models as judges-
Journal of Medical Internet Research
-
Combining International Standards to Develop Clinical Decision Support for Parent Smoking Cessation in Pediatrics
Smoking has severe health consequences, and secondhand smoke (SHS) exposure among children increases the risk of sudden infant death syndrome, chronic respiratory diseases, such as asthma, and lung cancer in adulthood. For many parents, pediatricians are the primary source of interaction with the healthcare system. Nevertheless, in pediatric settings, appropriate tobacco treatments are rarely, if ever, provided to parents who smoke. To best address tobacco use among parents, it is ideal to devel
Combining International Standards to Develop Clinical Decision Support for Parent Smoking Cessation in Pediatrics
-
Journal of Medical Internet Research
-
Key Features of Digital Phenotyping for Monitoring Mental Disorders: Systematic Review
Background: The COVID-19 pandemic has intensified mental health issues globally, highlighting the urgent need for remote mental health monitoring. Digital phenotyping using smart devices has emerged as a promising approach, but it remains unclear which features are essential for predicting depression and anxiety. Objective: This systematic review aimed to identify the types of features collected through smart packages—integrated systems combining smartphones with wearable devices such as Actiwat
Key Features of Digital Phenotyping for Monitoring Mental Disorders: Systematic Review
-
Nature Medicine
-
A multimodal whole-slide foundation model for pathology
Nature Medicine, Published online: 05 November 2025; doi:10.1038/s41591-025-03982-3Pretrained using 335,645 whole-slide images, a foundation model is developed to provide representations for slide- and patient-level tasks. It is capable of performing clinical tasks and generating reports even in data-scarce scenarios, such as rare cancer diagnosis and survival prediction, without requiring further fine-tuning.
A multimodal whole-slide foundation model for pathology
Nature Medicine, Published online: 05 November 2025; doi:10.1038/s41591-025-03982-3
Pretrained using 335,645 whole-slide images, a foundation model is developed to provide representations for slide- and patient-level tasks. It is capable of performing clinical tasks and generating reports even in data-scarce scenarios, such as rare cancer diagnosis and survival prediction, without requiring further fine-tuning.-
Nature - Issue - nature.com science feeds
-
Fair human-centric image dataset for ethical AI benchmarking
Nature, Published online: 05 November 2025; doi:10.1038/s41586-025-09716-2The Fair Human-Centric Image Benchmark (FHIBE, pronounced ‘Feebee’)—an image dataset that implements best practices for consent, privacy, compensation, safety, diversity and utility—can be used responsibly as a fairness evaluation dataset for many human-centric computer vision applications.
Fair human-centric image dataset for ethical AI benchmarking
Nature, Published online: 05 November 2025; doi:10.1038/s41586-025-09716-2
The Fair Human-Centric Image Benchmark (FHIBE, pronounced ‘Feebee’)—an image dataset that implements best practices for consent, privacy, compensation, safety, diversity and utility—can be used responsibly as a fairness evaluation dataset for many human-centric computer vision applications.-
cs.AI, q-bio.NC updates on arXiv.org
-
AI Diffusion in Low Resource Language Countries
arXiv:2511.02752v1 Announce Type: cross Abstract: Artificial intelligence (AI) is diffusing globally at unprecedented speed, but adoption remains uneven. Frontier Large Language Models (LLMs) are known to perform poorly on low-resource languages due to data scarcity. We hypothesize that this performance deficit reduces the utility of AI, thereby slowing adoption in Low-Resource Language Countries (LRLCs). To test this, we use a weighted regression model to isolate the language effect from socio
AI Diffusion in Low Resource Language Countries
-
cs.AI, q-bio.NC updates on arXiv.org
-
MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning
arXiv:2511.02805v1 Announce Type: cross Abstract: Typical search agents concatenate the entire interaction history into the LLM context, preserving information integrity but producing long, noisy contexts, resulting in high computation and memory costs. In contrast, using only the current turn avoids this overhead but discards essential information. This trade-off limits the scalability of search agents. To address this challenge, we propose MemSearcher, an agent workflow that iteratively maint
MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning
-
cs.AI, q-bio.NC updates on arXiv.org
-
How can we assess human-agent interactions? Case studies in software agent design
arXiv:2510.09801v2 Announce Type: replace Abstract: LLM-powered agents are both a promising new technology and a source of complexity, where choices about models, tools, and prompting can affect their usefulness. While numerous benchmarks measure agent accuracy across domains, they mostly assume full automation, failing to represent the collaborative nature of real-world use cases. In this paper, we make two major steps towards the rigorous assessment of human-agent interactions. First, we prop
How can we assess human-agent interactions? Case studies in software agent design
-
Nature - Issue - nature.com science feeds
-
Antibody drugs show promise for treating bird flu and HIV
Nature, Published online: 05 November 2025; doi:10.1038/d41586-025-03540-4Scientists are developing antibodies to track the evolution of these viruses and better treat infections.
Antibody drugs show promise for treating bird flu and HIV
Nature, Published online: 05 November 2025; doi:10.1038/d41586-025-03540-4
Scientists are developing antibodies to track the evolution of these viruses and better treat infections.-
cs.AI, q-bio.NC updates on arXiv.org
-
From Passive to Proactive: A Multi-Agent System with Dynamic Task Orchestration for Intelligent Medical Pre-Consultation
arXiv:2511.01445v1 Announce Type: new Abstract: Global healthcare systems face critical challenges from increasing patient volumes and limited consultation times, with primary care visits averaging under 5 minutes in many countries. While pre-consultation processes encompassing triage and structured history-taking offer potential solutions, they remain limited by passive interaction paradigms and context management challenges in existing AI systems. This study introduces a hierarchical multi-ag
From Passive to Proactive: A Multi-Agent System with Dynamic Task Orchestration for Intelligent Medical Pre-Consultation
-
cs.AI, q-bio.NC updates on arXiv.org
-
Digital Twin based Automatic Reconfiguration of Robotic Systems in Smart Environments
arXiv:2511.00094v1 Announce Type: cross Abstract: Robotic systems have become integral to smart environments, enabling applications ranging from urban surveillance and automated agriculture to industrial automation. However, their effective operation in dynamic settings - such as smart cities and precision farming - is challenged by continuously evolving topographies and environmental conditions. Traditional control systems often struggle to adapt quickly, leading to inefficiencies or operation