❌

Normal view

The Evaluation Gap in Medicine, AI and LLMs: Navigating Elusive Ground Truth & Uncertainty via a Probabilistic Paradigm

arXiv:2601.05500v1 Announce Type: new Abstract: Benchmarking the relative capabilities of AI systems, including Large Language Models (LLMs) and Vision Models, typically ignores the impact of uncertainty in the underlying ground truth answers from experts. This ambiguity is particularly consequential in medicine where uncertainty is pervasive. In this paper, we introduce a probabilistic paradigm to theoretically explain how high certainty in ground truth answers is almost always necessary for even an expert to achieve high scores, whereas in datasets with high variation in ground truth answers there may be little difference between a random labeller and an expert. Therefore, ignoring uncertainty in ground truth evaluation data can result in the misleading conclusion that a non-expert has similar performance to that of an expert. Using the probabilistic paradigm, we thus bring forth the concepts of expected accuracy and expected F1 to estimate the score an expert human or system can achieve given ground truth answer variability. Our work leads to the recommendation that when establishing the capability of a system, results should be stratified by probability of the ground truth answer, typically measured by the agreement rate of ground truth experts. Stratification becomes critical when the overall performance drops below a threshold of 80%. Under stratified evaluation, performance comparison becomes more reliable in high certainty bins, mitigating the effect of the key confounding factor -- uncertainty.

Safety Not Found (404): Hidden Risks of LLM-Based Robotics Decision Making

arXiv:2601.05529v1 Announce Type: new Abstract: One mistake by an AI system in a safety-critical setting can cost lives. As Large Language Models (LLMs) become integral to robotics decision-making, the physical dimension of risk grows; a single wrong instruction can directly endanger human safety. This paper addresses the urgent need to systematically evaluate LLM performance in scenarios where even minor errors are catastrophic. Through a qualitative evaluation of a fire evacuation scenario, we identified critical failure cases in LLM-based decision-making. Based on these, we designed seven tasks for quantitative assessment, categorized into: Complete Information, Incomplete Information, and Safety-Oriented Spatial Reasoning (SOSR). Complete information tasks utilize ASCII maps to minimize interpretation ambiguity and isolate spatial reasoning from visual processing. Incomplete information tasks require models to infer missing context, testing for spatial continuity versus hallucinations. SOSR tasks use natural language to evaluate safe decision-making in life-threatening contexts. We benchmark various LLMs and Vision-Language Models (VLMs) across these tasks. Beyond aggregate performance, we analyze the implications of a 1% failure rate, highlighting how "rare" errors escalate into catastrophic outcomes. Results reveal serious vulnerabilities: several models achieved a 0% success rate in ASCII navigation, while in a simulated fire drill, models instructed robots to move toward hazardous areas instead of emergency exits. Our findings lead to a sobering conclusion: current LLMs are not ready for direct deployment in safety-critical systems. A 99% accuracy rate is dangerously misleading in robotics, as it implies one out of every hundred executions could result in catastrophic harm. We demonstrate that even state-of-the-art models cannot guarantee safety, and absolute reliance on them creates unacceptable risks.

A Survey of Agentic AI and Cybersecurity: Challenges, Opportunities and Use-case Prototypes

arXiv:2601.05293v1 Announce Type: cross Abstract: Agentic AI marks an important transition from single-step generative models to systems capable of reasoning, planning, acting, and adapting over long-lasting tasks. By integrating memory, tool use, and iterative decision cycles, these systems enable continuous, autonomous workflows in real-world environments. This survey examines the implications of agentic AI for cybersecurity. On the defensive side, agentic capabilities enable continuous monitoring, autonomous incident response, adaptive threat hunting, and fraud detection at scale. Conversely, the same properties amplify adversarial power by accelerating reconnaissance, exploitation, coordination, and social-engineering attacks. These dual-use dynamics expose fundamental gaps in existing governance, assurance, and accountability mechanisms, which were largely designed for non-autonomous and short-lived AI systems. To address these challenges, we survey emerging threat models, security frameworks, and evaluation pipelines tailored to agentic systems, and analyze systemic risks including agent collusion, cascading failures, oversight evasion, and memory poisoning. Finally, we present three representative use-case implementations that illustrate how agentic AI behaves in practical cybersecurity workflows, and how design choices shape reliability, safety, and operational effectiveness.

Streamlining evidence based clinical recommendations with large language models

arXiv:2505.10282v2 Announce Type: replace-cross Abstract: Clinical evidence underpins informed healthcare decisions, yet integrating it into real-time practice remains challenging due to intensive workloads, complex procedures, and time constraints. This study presents Quicker, an LLM-powered system that automates evidence synthesis and generates clinical recommendations following standard guideline development workflows. Quicker delivers an end-to-end pipeline from clinical questions to recommendations and supports customized decision-making through integrated tools and interactive interfaces. To evaluate how closely Quicker can reproduce guideline development processes, we constructed Q2CRBench-3, a benchmark derived from guideline development records for three diseases. Experiments show that Quicker produces precise question decomposition, expert-aligned retrieval, and near-comprehensive screening. Quicker assistance improved the accuracy of extracted study data, and its recommendations were more comprehensive and coherent than clinician-written ones. In system-level testing, Quicker working with one participant reduced recommendation development to 20-40 min. Overall, the findings demonstrate Quicker's potential to enhance the speed and reliability of evidence-based clinical decision-making.

CliCARE: Grounding Large Language Models in Clinical Guidelines for Decision Support over Longitudinal Cancer Electronic Health Records

arXiv:2507.22533v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) hold significant promise for improving clinical decision support and reducing physician burnout by synthesizing complex, longitudinal cancer Electronic Health Records (EHRs). However, their implementation in this critical field faces three primary challenges: the inability to effectively process the extensive length and fragmented nature of patient records for accurate temporal analysis; a heightened risk of clinical hallucination, as conventional grounding techniques such as Retrieval-Augmented Generation (RAG) do not adequately incorporate process-oriented clinical guidelines; and unreliable evaluation metrics that hinder the validation of AI systems in oncology. To address these issues, we propose CliCARE, a framework for Grounding Large Language Models in Clinical Guidelines for Decision Support over Longitudinal Cancer Electronic Health Records. The framework operates by transforming unstructured, longitudinal EHRs into patient-specific Temporal Knowledge Graphs (TKGs) to capture long-range dependencies, and then grounding the decision support process by aligning these real-world patient trajectories with a normative guideline knowledge graph. This approach provides oncologists with evidence-grounded decision support by generating a high-fidelity clinical summary and an actionable recommendation. We validated our framework using large-scale, longitudinal data from a private Chinese cancer dataset and the public English MIMIC-IV dataset. In these settings, CliCARE significantly outperforms baselines, including leading long-context LLMs and Knowledge Graph-enhanced RAG methods. The clinical validity of our results is supported by a robust evaluation protocol, which demonstrates a high correlation with assessments made by oncologists.

Benchmarking LLM-based Agents for Single-cell Omics Analysis

arXiv:2508.13201v2 Announce Type: replace-cross Abstract: The surge in multimodal single-cell omics data exposes limitations in traditional, manually defined analysis workflows. AI agents offer a paradigm shift, enabling adaptive planning, executable code generation, traceable decisions, and real-time knowledge fusion. However, the lack of a comprehensive benchmark critically hinders progress. We introduce a novel benchmarking evaluation system to rigorously assess agent capabilities in single-cell omics analysis. This system comprises: a unified platform compatible with diverse agent frameworks and LLMs; multidimensional metrics assessing cognitive program synthesis, collaboration, execution efficiency, bioinformatics knowledge integration, and task completion quality; and 50 diverse real-world single-cell omics analysis tasks spanning multi-omics, species, and sequencing technologies. Our evaluation reveals that Grok-3-beta achieves state-of-the-art performance among tested agent frameworks. Multi-agent frameworks significantly enhance collaboration and execution efficiency over single-agent approaches through specialized role division. Attribution analyses of agent capabilities identify that high-quality code generation is crucial for task success, and self-reflection has the most significant overall impact, followed by retrieval-augmented generation (RAG) and planning. This work highlights persistent challenges in code generation, long-context handling, and context-aware knowledge retrieval, providing a critical empirical foundation and best practices for developing robust AI agents in computational biology.

Personalizing Treatment for Pancreatic Ductal Adenocarcinoma: The Emerging Role of Minimal Residual Disease in Perioperative Decision-Making

Cancers (Basel). 2025 Dec 27;18(1):94. doi: 10.3390/cancers18010094.

ABSTRACT

Pancreatic ductal adenocarcinoma (PDAC) is a highly aggressive malignancy with poor long-term survival despite advances in surgical techniques, systemic therapies, and perioperative management. High rates of systemic recurrence following curative-intent resection suggest that many patients harbor minimal residual disease (MRD), microscopic tumor burden that persists postoperatively and remains undetectable by conventional diagnostic tools. Recent advances in liquid biopsy technologies, particularly circulating tumor DNA (ctDNA) analysis, alongside detailed characterization of the PDAC mutational landscape, offer a promising non-invasive approach for MRD detection. Emerging evidence indicates that MRD status can serve as a sensitive prognostic biomarker, identify patients at high risk of relapse, and guide personalized perioperative therapy, including optimization of adjuvant treatment. This review summarizes current knowledge on the biology and detection of MRD in PDAC, its implications for perioperative risk stratification and treatment decision-making, and discusses future directions for integrating MRD assessment into clinical practice to enable more precise, individualized patient management.

PMID:41514607 | PMC:PMC12784771 | DOI:10.3390/cancers18010094

Integrative Genomic and AI Approaches to Lung Cancer and Implications for Disease Prevention in Former Smokers

Int J Mol Sci. 2026 Jan 4;27(1):521. doi: 10.3390/ijms27010521.

ABSTRACT

Tobacco smoking accounts for nearly 90% of lung cancer deaths worldwide, yet the mechanisms underlying persistent cancer risk in former smokers are not fully understood. Epidemiological evidence shows that more than 40% of lung cancers develop over 15 years after cessation, demonstrating that while some smoking-induced molecular alterations resolve rapidly, others remain as long-lasting scars that promote carcinogenesis. This review synthesizes longitudinal and cross-sectional genomic, epigenomic, and transcriptomic studies of airway and lung tissues to distinguish persistent from nonpersistent smoking-induced molecular alterations. Persistent alterations include somatic mutations in TP53 and KRAS, DNA methylation at tumor suppressor loci, dysregulated noncoding RNAs, chromosomal instability, and epigenetic age acceleration. Nonpersistent changes, such as acute inflammatory responses and detoxification pathways, generally normalize within months to several years following cessation. Multi-omics profiling reveals coordinated patterns of dysregulation consistent with field cancerization in former smokers. In addition, the integration of multi-omics data with artificial intelligence may enable composite molecular signatures for stratifying high-risk former smokers, link molecular persistence to clinical outcomes, and inform chemoprevention strategies. Collectively, these observations clarify which molecular alterations sustain long-term cancer risk despite smoking cessation and highlight opportunities for precision prevention and earlier detection in high-risk populations.

PMID:41516393 | PMC:PMC12786486 | DOI:10.3390/ijms27010521

❌