❌

Reading view

Evaluating Control Protocols for Untrusted AI Agents

arXiv:2511.02997v1 Announce Type: new Abstract: As AI systems become more capable and widely deployed as agents, ensuring their safe operation becomes critical. AI control offers one approach to mitigating the risk from untrusted AI agents by monitoring their actions and intervening or auditing when necessary. Evaluating the safety of these protocols requires understanding both their effectiveness against current attacks and their robustness to adaptive adversaries. In this work, we systematically evaluate a range of control protocols in SHADE-Arena, a dataset of diverse agentic environments. First, we evaluate blue team protocols, including deferral to trusted models, resampling, and deferring on critical actions, against a default attack policy. We find that resampling for incrimination and deferring on critical actions perform best, increasing safety from 50% to 96%. We then iterate on red team strategies against these protocols and find that attack policies with additional affordances, such as knowledge of when resampling occurs or the ability to simulate monitors, can substantially improve attack success rates against our resampling strategy, decreasing safety to 17%. However, deferring on critical actions is highly robust to even our strongest red team strategies, demonstrating the importance of denying attack policies access to protocol internals.
  •  

No-Human in the Loop: Agentic Evaluation at Scale for Recommendation

arXiv:2511.03051v1 Announce Type: new Abstract: Evaluating large language models (LLMs) as judges is increasingly critical for building scalable and trustworthy evaluation pipelines. We present ScalingEval, a large-scale benchmarking study that systematically compares 36 LLMs, including GPT, Gemini, Claude, and Llama, across multiple product categories using a consensus-driven evaluation protocol. Our multi-agent framework aggregates pattern audits and issue codes into ground-truth labels via scalable majority voting, enabling reproducible comparison of LLM evaluators without human annotation. Applied to large-scale complementary-item recommendation, the benchmark reports four key findings: (i) Anthropic Claude 3.5 Sonnet achieves the highest decision confidence; (ii) Gemini 1.5 Pro offers the best overall performance across categories; (iii) GPT-4o provides the most favorable latency-accuracy-cost tradeoff; and (iv) GPT-OSS 20B leads among open-source models. Category-level analysis shows strong consensus in structured domains (Electronics, Sports) but persistent disagreement in lifestyle categories (Clothing, Food). These results establish ScalingEval as a reproducible benchmark and evaluation protocol for LLMs as judges, with actionable guidance on scaling, reliability, and model family tradeoffs.
  •  

Explaining Decisions in ML Models: a Parameterized Complexity Analysis (Part I)

arXiv:2511.03545v1 Announce Type: new Abstract: This paper presents a comprehensive theoretical investigation into the parameterized complexity of explanation problems in various machine learning (ML) models. Contrary to the prevalent black-box perception, our study focuses on models with transparent internal mechanisms. We address two principal types of explanation problems: abductive and contrastive, both in their local and global variants. Our analysis encompasses diverse ML models, including Decision Trees, Decision Sets, Decision Lists, Boolean Circuits, and ensembles thereof, each offering unique explanatory challenges. This research fills a significant gap in explainable AI (XAI) by providing a foundational understanding of the complexities of generating explanations for these models. This work provides insights vital for further research in the domain of XAI, contributing to the broader discourse on the necessity of transparency and accountability in AI systems.
  •  

Digital Transformation Chatbot (DTchatbot): Integrating Large Language Model-based Chatbot in Acquiring Digital Transformation Needs

arXiv:2511.02842v1 Announce Type: cross Abstract: Many organisations pursue digital transformation to enhance operational efficiency, reduce manual efforts, and optimise processes by automation and digital tools. To achieve this, a comprehensive understanding of their unique needs is required. However, traditional methods, such as expert interviews, while effective, face several challenges, including scheduling conflicts, resource constraints, inconsistency, etc. To tackle these issues, we investigate the use of a Large Language Model (LLM)-powered chatbot to acquire organisations' digital transformation needs. Specifically, the chatbot integrates workflow-based instruction with LLM's planning and reasoning capabilities, enabling it to function as a virtual expert and conduct interviews. We detail the chatbot's features and its implementation. Our preliminary evaluation indicates that the chatbot performs as designed, effectively following predefined workflows and supporting user interactions with areas for improvement. We conclude by discussing the implications of employing chatbots to elicit user information, emphasizing their potential and limitations.
  •  

Mathematical exploration and discovery at scale

arXiv:2511.02864v1 Announce Type: cross Abstract: AlphaEvolve is a generic evolutionary coding agent that combines the generative capabilities of LLMs with automated evaluation in an iterative evolutionary framework that proposes, tests, and refines algorithmic solutions to challenging scientific and practical problems. In this paper we showcase AlphaEvolve as a tool for autonomously discovering novel mathematical constructions and advancing our understanding of long-standing open problems. To demonstrate its breadth, we considered a list of 67 problems spanning mathematical analysis, combinatorics, geometry, and number theory. The system rediscovered the best known solutions in most of the cases and discovered improved solutions in several. In some instances, AlphaEvolve is also able to generalize results for a finite number of input values into a formula valid for all input values. Furthermore, we are able to combine this methodology with Deep Think and AlphaProof in a broader framework where the additional proof-assistants and reasoning systems provide automated proof generation and further mathematical insights. These results demonstrate that large language model-guided evolutionary search can autonomously discover mathematical constructions that complement human intuition, at times matching or even improving the best known results, highlighting the potential for significant new ways of interaction between mathematicians and AI systems. We present AlphaEvolve as a powerful new tool for mathematical discovery, capable of exploring vast search spaces to solve complex optimization problems at scale, often with significantly reduced requirements on preparation and computation time.
  •  

FP-AbDiff: Improving Score-based Antibody Design by Capturing Nonequilibrium Dynamics through the Underlying Fokker-Planck Equation

arXiv:2511.03113v1 Announce Type: cross Abstract: Computational antibody design holds immense promise for therapeutic discovery, yet existing generative models are fundamentally limited by two core challenges: (i) a lack of dynamical consistency, which yields physically implausible structures, and (ii) poor generalization due to data scarcity and structural bias. We introduce FP-AbDiff, the first antibody generator to enforce Fokker-Planck Equation (FPE) physics along the entire generative trajectory. Our method minimizes a novel FPE residual loss over the mixed manifold of CDR geometries (R^3 x SO(3)), compelling locally-learned denoising scores to assemble into a globally coherent probability flow. This physics-informed regularizer is synergistically integrated with deep biological priors within a state-of-the-art SE(3)-equivariant diffusion framework. Rigorous evaluation on the RAbD benchmark confirms that FP-AbDiff establishes a new state-of-the-art. In de novo CDR-H3 design, it achieves a mean Root Mean Square Deviation of 0.99 {\AA} when superposing on the variable region, a 25% improvement over the previous state-of-the-art model, AbX, and the highest reported Contact Amino Acid Recovery of 39.91%. This superiority is underscored in the more challenging six-CDR co-design task, where our model delivers consistently superior geometric precision, cutting the average full-chain Root Mean Square Deviation by ~15%, and crucially, achieves the highest full-chain Amino Acid Recovery on the functionally dominant CDR-H3 loop (45.67%). By aligning generative dynamics with physical laws, FP-AbDiff enhances robustness and generalizability, establishing a principled approach for physically faithful and functionally viable antibody design.
  •  

LGM: Enhancing Large Language Models with Conceptual Meta-Relations and Iterative Retrieval

arXiv:2511.03214v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit strong semantic understanding, yet struggle when user instructions involve ambiguous or conceptually misaligned terms. We propose the Language Graph Model (LGM) to enhance conceptual clarity by extracting meta-relations-inheritance, alias, and composition-from natural language. The model further employs a reflection mechanism to validate these meta-relations. Leveraging a Concept Iterative Retrieval Algorithm, these relations and related descriptions are dynamically supplied to the LLM, improving its ability to interpret concepts and generate accurate responses. Unlike conventional Retrieval-Augmented Generation (RAG) approaches that rely on extended context windows, our method enables large language models to process texts of any length without the need for truncation. Experiments on standard benchmarks demonstrate that the LGM consistently outperforms existing RAG baselines.
  •  

Hybrid Fact-Checking that Integrates Knowledge Graphs, Large Language Models, and Search-Based Retrieval Agents Improves Interpretable Claim Verification

arXiv:2511.03217v1 Announce Type: cross Abstract: Large language models (LLMs) excel in generating fluent utterances but can lack reliable grounding in verified information. At the same time, knowledge-graph-based fact-checkers deliver precise and interpretable evidence, yet suffer from limited coverage or latency. By integrating LLMs with knowledge graphs and real-time search agents, we introduce a hybrid fact-checking approach that leverages the individual strengths of each component. Our system comprises three autonomous steps: 1) a Knowledge Graph (KG) Retrieval for rapid one - hop lookups in DBpedia, 2) an LM-based classification guided by a task-specific labeling prompt, producing outputs with internal rule-based logic, and 3) a Web Search Agent invoked only when KG coverage is insufficient. Our pipeline achieves an F1 score of 0.93 on the FEVER benchmark on the Supported/Refuted split without task- specific fine - tuning. To address Not enough information cases, we conduct a targeted reannotation study showing that our approach frequently uncovers valid evidence for claims originally labeled as Not Enough Information (NEI), as confirmed by both expert annotators and LLM reviewers. With this paper, we present a modular, opensource fact-checking pipeline with fallback strategies and generalization across datasets.
  •  

RAG-IT: Retrieval-Augmented Instruction Tuning for Automated Financial Analysis

arXiv:2412.08179v2 Announce Type: replace-cross Abstract: Financial analysis relies heavily on the interpretation of earnings reports to assess company performance and guide decision-making. Traditional methods for generating such analyses demand significant financial expertise and are often time-consuming. With the rapid advancement of Large Language Models (LLMs), domain-specific adaptations have emerged for financial tasks such as sentiment analysis and entity recognition. This paper introduces RAG-IT (Retrieval-Augmented Instruction Tuning), a novel framework designed to automate the generation of earnings report analyses through an LLM fine-tuned specifically for the financial domain. Our approach integrates retrieval augmentation with instruction-based fine-tuning to enhance factual accuracy, contextual relevance, and domain adaptability. We construct a comprehensive financial instruction dataset derived from extensive financial documents and earnings reports to guide the LLM's adaptation to specialized financial reasoning. Experimental results demonstrate that RAG-IT outperforms general-purpose open-source models and achieves performance comparable to commercial systems like GPT-3.5 on financial report generation tasks. This research highlights the potential of retrieval-augmented instruction tuning to streamline and elevate financial analysis automation, advancing the broader field of intelligent financial reporting.
  •  

REFA: Reference Free Alignment for multi-preference optimization

arXiv:2412.16378v4 Announce Type: replace-cross Abstract: To mitigate reward hacking from response verbosity, modern preference optimization methods are increasingly adopting length normalization (e.g., SimPO, ORPO, LN-DPO). While effective against this bias, we demonstrate that length normalization itself introduces a failure mode: the URSLA shortcut. Here models learn to satisfy the alignment objective by prematurely truncating low-quality responses rather than learning from their semantic content. To address this, we introduce REFA, a new alignment framework that proposes probabilistic control on a structural token that controls termination. Our core innovation is a new class of regularizers that operate directly on the probability of the End-of-Sequence (EOS) token, a previously unexploited control lever. This token-level intervention provides a principled solution to the URSLA shortcut, ensuring genuine quality improvements. Furthermore, it unlocks a versatile mechanism for managing the alignment-efficiency tradeoff, enabling practitioners to fine-tune models that adhere to specific token budgets. Empirically, REFA achieves a 60.29% win rate and a 52.17% length-controlled win rate on AlpacaEval2 with Llama-3-8B-Instruct, demonstrating the power of our token-level control paradigm.
  •  

Evaluating Large Language Models for Detecting Antisemitism

arXiv:2509.18293v2 Announce Type: replace-cross Abstract: Detecting hateful content is a challenging and important problem. Automated tools, like machine-learning models, can help, but they require continuous training to adapt to the ever-changing landscape of social media. In this work, we evaluate eight open-source LLMs' capability to detect antisemitic content, specifically leveraging in-context definition. We also study how LLMs understand and explain their decisions given a moderation policy as a guideline. First, we explore various prompting techniques and design a new CoT-like prompt, Guided-CoT, and find that injecting domain-specific thoughts increases performance and utility. Guided-CoT handles the in-context policy well, improving performance and utility by reducing refusals across all evaluated models, regardless of decoding configuration, model size, or reasoning capability. Notably, Llama 3.1 70B outperforms fine-tuned GPT-3.5. Additionally, we examine LLM errors and introduce metrics to quantify semantic divergence in model-generated rationales, revealing notable differences and paradoxical behaviors among LLMs. Our experiments highlight the differences observed across LLMs' utility, explainability, and reliability. Code and resources available at: https://github.com/idramalab/quantify-llm-explanations
  •  

Site-specific DNA insertion into the human genome with engineered recombinases

Nature Biotechnology, Published online: 06 November 2025; doi:10.1038/s41587-025-02895-3

Engineered DNA recombinases efficiently and specifically insert genetic cargos without the use of landing pads.
  •  

Harnessing multi-omics approaches to decipher tumor evolution and improve diagnosis and therapy in lung cancer

Biomark Res. 2025 Nov 5;13(1):140. doi: 10.1186/s40364-025-00859-y.

ABSTRACT

With the advancement of novel technologies such as whole-genome sequencing, single-cell sequencing, and spatial transcriptomics, single-omics analyses have already promoted the research of tumorigenesis as well as development and have partly elucidated the evolutionary processes of lung cancer. However, it is still difficult to distinguish these confounding features via single dimensional approaches due to the complexity, heterogeneity and cell-cell interactions with the immune microenvironment in lung cancer. Multi-omics approaches provide a holistic framework for constructing detailed tumor ecosystem landscapes, thereby facilitating the development of a more robust classification system for precision diagnosis and treatment, and aiding in the discovery of novel cancer biomarkers. In this review, we summarize the potential and applications of multi-omics approaches in characterizing intratumor heterogeneity and the tumor microenvironment throughout the course of lung cancer development. By further discussing the discovery and application of diagnostic and therapeutic biomarkers across precancerous lesions, early-stage lung cancer, tumor progression, metastasis, and therapy resistance, we outline the current challenges and future prospects of using multi-omics to identify reliable biomarkers. Moreover, we emphasize that integrative multi-omics models hold great promise for elucidating the complex interactions within the lung cancer ecosystem, thereby contributing to improved diagnostic accuracy, optimized therapeutic strategies, and better patient outcomes.

PMID:41194170 | PMC:PMC12590604 | DOI:10.1186/s40364-025-00859-y

  •  

Improving dataset transparency in dermatologic Artificial Intelligence using a dataset nutrition label

npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02125-9

Biased and poorly documented dermatology datasets pose risks to the development of safe and generalizable artificial intelligence (AI) tools. We created a Dataset Nutrition Label (DNL) for multiple dermatology datasets to support transparent and responsible data use. The DNL offers a structured, digestible summary of key attributes, including metadata, limitations, and risks, enabling data users to better assess suitability and proactively address potential sources of bias in datasets.
  •  
  •  

Liquid biopsy in gastrointestinal oncology: clinical applications and translational integration of ctDNA, CTCs, and sEVs

Oncol Rev. 2025 Oct 20;19:1702932. doi: 10.3389/or.2025.1702932. eCollection 2025.

ABSTRACT

BACKGROUND AND AIMS: Liquid biopsy offers a minimally invasive tool to detect actionable mutations, monitor minimal residual disease (MRD), and guide therapy in gastrointestinal (GI) cancers. We critically review the clinical utility of circulating tumor DNA (ctDNA), circulating tumor cells (CTCs), and small extracellular vesicles (sEVs) across GI malignancies and propose a framework for their integration into clinical practice.

METHODS: We synthesized evidence from over 200 studies, including prospective trials and translational research, to assess diagnostic accuracy, prognostic value, and clinical actionability of each biomarker type in esophageal, gastric, colorectal, pancreatic, hepatocellular, and biliary cancers.

RESULTS: ctDNA has shown strong potential for MRD detection and treatment monitoring, particularly in colorectal and pancreatic cancer. CTCs offer insights into metastatic risk and therapeutic resistance, while sEVs provide molecular cargo relevant to immunomodulation and disease progression. Emerging microfluidics and AI-driven multi-omics approaches may overcome current limitations.

CONCLUSION: The integration of liquid biopsy technologies into GI oncology holds promise for early detection and precision therapy. We propose a five-phase clinical roadmap and outine the key research gaps that need to be addressed before widespread implementation in routine care.

PMID:41190015 | PMC:PMC12580207 | DOI:10.3389/or.2025.1702932

  •  

Combining International Standards to Develop Clinical Decision Support for Parent Smoking Cessation in Pediatrics

Smoking has severe health consequences, and secondhand smoke (SHS) exposure among children increases the risk of sudden infant death syndrome, chronic respiratory diseases, such as asthma, and lung cancer in adulthood. For many parents, pediatricians are the primary source of interaction with the healthcare system. Nevertheless, in pediatric settings, appropriate tobacco treatments are rarely, if ever, provided to parents who smoke. To best address tobacco use among parents, it is ideal to develop scalable solutions that are coordinated across health systems, community partners, and national services within pediatric settings. We describe our experience developing and implementing a parent tobacco treatment platform (PTTP) within a pediatric institution that leverages multiple international standards to support interoperability, with the overarching goal of providing a model for how such work can be approached. The clinical decision support (CDS) system includes clinician- and patient-facing components, connects parents to three different treatment options (nicotine replacement therapy, text-based counseling, and telephonic counseling), and incorporates three international standards (Fast Healthcare Interoperability Resources [FHIR], SMART on FHIR, and CDS Hooks). FHIR is used across all components. SMART on FHIR is limited to the clinician-facing tool, and CDS Hooks is used in the patient-facing portion. While healthcare interoperability standards supported a significant portion of the overall system, non-standard technologies and enhancements of existing standards were also required. Further, no connections with community partners could use existing interoperability standards. Over one year, the CDS was used in 194,946 visits, identified 7,847 parents who smoke, and connected 2,954 parents to 6,320 distinct treatment services, a significant improvement compared to prior efforts. Our project demonstrates that building CDS systems using international standards, such as SMART on FHIR, FHIR, and CDS Hooks, is possible, but challenges remain. Limits in the CDS Hooks standard to support common workflows and a lack of communication standards used by 3rd parties outside the healthcare system represent areas for future work. To support these requirements, additional EHR-specific records and communication mechanisms were required.
  •  

Key Features of Digital Phenotyping for Monitoring Mental Disorders: Systematic Review

Background: The COVID-19 pandemic has intensified mental health issues globally, highlighting the urgent need for remote mental health monitoring. Digital phenotyping using smart devices has emerged as a promising approach, but it remains unclear which features are essential for predicting depression and anxiety. Objective: This systematic review aimed to identify the types of features collected through smart packages—integrated systems combining smartphones with wearable devices such as Actiwatches, smartbands, and smartwatches—and to determine which features should be considered essential for mental health monitoring based on the type of device used. Methods: A systematic review was conducted. Searches were performed across Web of Science, PubMed, and Scopus on February 5, 2025. Inclusion criteria comprised quantitative studies involving adults (≥19 years) using smart devices to predict depression or anxiety based on passive data collection. Studies focusing solely on smartphones or qualitative designs were excluded. Risk of bias was assessed using the Mixed Methods Appraisal Tool and Quality Criteria Checklist. Data were synthesized descriptively, and the relative contribution of each feature was further assessed by calculating coverage (proportion of studies using a feature) and importance among used (proportion identifying it as important when used). These metrics were visualized in quadrant-based scatter plots to identify consistently important features across devices. Results: From 1,382 records, 22 studies across 11 countries were included. The overall synthesis identified a core feature package—accelerometer (ACC), steps, Heart Rate (HR), and sleep. Device-specific analyses revealed further nuances: In Actiwatch studies, ACC and activity were consistently important, but sleep features were rarely examined. In smartbands, HR, steps, sleep, and phone usage were essential, while Global Positioning System (GPS), Electrodermal Activity (EDA), and skin temperature (TEMP) showed high importance when used, suggesting opportunities for broader adoption. In smartwatch studies, sleep and HR emerged as core features, whereas steps and ACC were widely used but often not identified as important. Conclusions: This systematic review identified a core feature package comprising ACC, steps, HR, and sleep that consistently contributes to mood disorder prediction across devices. At the same time, device-specific differences were observed: Actiwatch studies mainly emphasized ACC and activity but underutilized sleep features; smartbands highlighted HR, steps, sleep, and phone usage, with EDA, TEMP, and GPS showing additional promise; and smartwatches most reliably leveraged sleep and HR, while steps and ACC were widely used yet less effective. These findings suggest that while a shared core set of features exists, optimizing digital phenotyping requires tailoring feature selection to the characteristics of each device type. To advance this field, improving data accessibility, particularly in smartwatch ecosystems, and adopting standardized reporting frameworks will be essential to enhance comparability, reproducibility, and future meta-analytic integration. Clinical Trial: Open Science Framework https://osf.io/nz7k8
  •  

Curated and harmonised transcriptomics datasets of interstitial lung diseases

Data Brief. 2025 Oct 14;63:112139. doi: 10.1016/j.dib.2025.112139. eCollection 2025 Dec.

ABSTRACT

This study provides manually curated and homogenised transcriptomics data of interstitial lung disease (ILD) patients retrieved from the NCBI Gene Expression Omnibus and European Nucleotide Archive repositories. The compendium includes 30 transcriptomics datasets generated with DNA microarrays and RNA sequencing (RNA-seq) technologies for a total of 1371 samples. All the datasets underwent metadata curation and harmonisation, data quality check, and preprocessing with standardised procedures. Furthermore, a robust data model was developed to standardise phenotypic data, thereby enhancing comparability across heterogeneous datasets. Gene expression data and lists of differentially expressed genes computed between ILD and healthy samples are provided. Among the ILDs included in this study, idiopathic pulmonary fibrosis (IPF) is the most represented worldwide. Co-expression networks of IPF and healthy samples were inferred, which are also included in this study. This study enhances the Findability, Accessibility, Interoperability, and Reusability (FAIR) of publicly available transcriptomic datasets related to ILDs. The resulting resource provides a integrated platform for the implementation and validation of systems biology and pharmacology approaches, facilitating the development of novel diagnostic and therapeutic strategies for ILDs.

PMID:41189603 | PMC:PMC12581653 | DOI:10.1016/j.dib.2025.112139

  •  
❌