Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
Evaluating Control Protocols for Untrusted AI Agents
arXiv:2511.02997v1 Announce Type: new Abstract: As AI systems become more capable and widely deployed as agents, ensuring their safe operation becomes critical. AI control offers one approach to mitigating the risk from untrusted AI agents by monitoring their actions and intervening or auditing when necessary. Evaluating the safety of these protocols requires understanding both their effectiveness against current attacks and their robustness to adaptive adversaries. In this work, we systematica
-
cs.AI, q-bio.NC updates on arXiv.org
-
No-Human in the Loop: Agentic Evaluation at Scale for Recommendation
arXiv:2511.03051v1 Announce Type: new Abstract: Evaluating large language models (LLMs) as judges is increasingly critical for building scalable and trustworthy evaluation pipelines. We present ScalingEval, a large-scale benchmarking study that systematically compares 36 LLMs, including GPT, Gemini, Claude, and Llama, across multiple product categories using a consensus-driven evaluation protocol. Our multi-agent framework aggregates pattern audits and issue codes into ground-truth labels via s
No-Human in the Loop: Agentic Evaluation at Scale for Recommendation
-
cs.AI, q-bio.NC updates on arXiv.org
-
Explaining Decisions in ML Models: a Parameterized Complexity Analysis (Part I)
arXiv:2511.03545v1 Announce Type: new Abstract: This paper presents a comprehensive theoretical investigation into the parameterized complexity of explanation problems in various machine learning (ML) models. Contrary to the prevalent black-box perception, our study focuses on models with transparent internal mechanisms. We address two principal types of explanation problems: abductive and contrastive, both in their local and global variants. Our analysis encompasses diverse ML models, includin
Explaining Decisions in ML Models: a Parameterized Complexity Analysis (Part I)
-
cs.AI, q-bio.NC updates on arXiv.org
-
FP-AbDiff: Improving Score-based Antibody Design by Capturing Nonequilibrium Dynamics through the Underlying Fokker-Planck Equation
arXiv:2511.03113v1 Announce Type: cross Abstract: Computational antibody design holds immense promise for therapeutic discovery, yet existing generative models are fundamentally limited by two core challenges: (i) a lack of dynamical consistency, which yields physically implausible structures, and (ii) poor generalization due to data scarcity and structural bias. We introduce FP-AbDiff, the first antibody generator to enforce Fokker-Planck Equation (FPE) physics along the entire generative traj
FP-AbDiff: Improving Score-based Antibody Design by Capturing Nonequilibrium Dynamics through the Underlying Fokker-Planck Equation
-
cs.AI, q-bio.NC updates on arXiv.org
-
LGM: Enhancing Large Language Models with Conceptual Meta-Relations and Iterative Retrieval
arXiv:2511.03214v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit strong semantic understanding, yet struggle when user instructions involve ambiguous or conceptually misaligned terms. We propose the Language Graph Model (LGM) to enhance conceptual clarity by extracting meta-relations-inheritance, alias, and composition-from natural language. The model further employs a reflection mechanism to validate these meta-relations. Leveraging a Concept Iterative Retrieval Algorithm
LGM: Enhancing Large Language Models with Conceptual Meta-Relations and Iterative Retrieval
-
cs.AI, q-bio.NC updates on arXiv.org
-
Hybrid Fact-Checking that Integrates Knowledge Graphs, Large Language Models, and Search-Based Retrieval Agents Improves Interpretable Claim Verification
arXiv:2511.03217v1 Announce Type: cross Abstract: Large language models (LLMs) excel in generating fluent utterances but can lack reliable grounding in verified information. At the same time, knowledge-graph-based fact-checkers deliver precise and interpretable evidence, yet suffer from limited coverage or latency. By integrating LLMs with knowledge graphs and real-time search agents, we introduce a hybrid fact-checking approach that leverages the individual strengths of each component. Our sys
Hybrid Fact-Checking that Integrates Knowledge Graphs, Large Language Models, and Search-Based Retrieval Agents Improves Interpretable Claim Verification
-
cs.AI, q-bio.NC updates on arXiv.org
-
REFA: Reference Free Alignment for multi-preference optimization
arXiv:2412.16378v4 Announce Type: replace-cross Abstract: To mitigate reward hacking from response verbosity, modern preference optimization methods are increasingly adopting length normalization (e.g., SimPO, ORPO, LN-DPO). While effective against this bias, we demonstrate that length normalization itself introduces a failure mode: the URSLA shortcut. Here models learn to satisfy the alignment objective by prematurely truncating low-quality responses rather than learning from their semantic co
REFA: Reference Free Alignment for multi-preference optimization
-
Omics In Lung
-
Harnessing multi-omics approaches to decipher tumor evolution and improve diagnosis and therapy in lung cancer
Biomark Res. 2025 Nov 5;13(1):140. doi: 10.1186/s40364-025-00859-y.ABSTRACTWith the advancement of novel technologies such as whole-genome sequencing, single-cell sequencing, and spatial transcriptomics, single-omics analyses have already promoted the research of tumorigenesis as well as development and have partly elucidated the evolutionary processes of lung cancer. However, it is still difficult to distinguish these confounding features via single dimensional approaches due to the complexity,
Harnessing multi-omics approaches to decipher tumor evolution and improve diagnosis and therapy in lung cancer
Biomark Res. 2025 Nov 5;13(1):140. doi: 10.1186/s40364-025-00859-y.
ABSTRACT
With the advancement of novel technologies such as whole-genome sequencing, single-cell sequencing, and spatial transcriptomics, single-omics analyses have already promoted the research of tumorigenesis as well as development and have partly elucidated the evolutionary processes of lung cancer. However, it is still difficult to distinguish these confounding features via single dimensional approaches due to the complexity, heterogeneity and cell-cell interactions with the immune microenvironment in lung cancer. Multi-omics approaches provide a holistic framework for constructing detailed tumor ecosystem landscapes, thereby facilitating the development of a more robust classification system for precision diagnosis and treatment, and aiding in the discovery of novel cancer biomarkers. In this review, we summarize the potential and applications of multi-omics approaches in characterizing intratumor heterogeneity and the tumor microenvironment throughout the course of lung cancer development. By further discussing the discovery and application of diagnostic and therapeutic biomarkers across precancerous lesions, early-stage lung cancer, tumor progression, metastasis, and therapy resistance, we outline the current challenges and future prospects of using multi-omics to identify reliable biomarkers. Moreover, we emphasize that integrative multi-omics models hold great promise for elucidating the complex interactions within the lung cancer ecosystem, thereby contributing to improved diagnostic accuracy, optimized therapeutic strategies, and better patient outcomes.
PMID:41194170 | PMC:PMC12590604 | DOI:10.1186/s40364-025-00859-y
-
npj Digital Medicine
-
Improving dataset transparency in dermatologic Artificial Intelligence using a dataset nutrition label
npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02125-9Biased and poorly documented dermatology datasets pose risks to the development of safe and generalizable artificial intelligence (AI) tools. We created a Dataset Nutrition Label (DNL) for multiple dermatology datasets to support transparent and responsible data use. The DNL offers a structured, digestible summary of key attributes, including metadata, limitations, and risks, enabling data users to better ass
Improving dataset transparency in dermatologic Artificial Intelligence using a dataset nutrition label
npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02125-9
Biased and poorly documented dermatology datasets pose risks to the development of safe and generalizable artificial intelligence (AI) tools. We created a Dataset Nutrition Label (DNL) for multiple dermatology datasets to support transparent and responsible data use. The DNL offers a structured, digestible summary of key attributes, including metadata, limitations, and risks, enabling data users to better assess suitability and proactively address potential sources of bias in datasets.-
npj Digital Medicine
-
Evaluating clinical AI summaries with large language models as judges
npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02005-2Evaluating clinical AI summaries with large language models as judges
Evaluating clinical AI summaries with large language models as judges
npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02005-2
Evaluating clinical AI summaries with large language models as judges