Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
Making LLMs Reliable When It Matters Most: A Five-Layer Architecture for High-Stakes Decisions
arXiv:2511.07669v1 Announce Type: new Abstract: Current large language models (LLMs) excel in verifiable domains where outputs can be checked before action but prove less reliable for high-stakes strategic decisions with uncertain outcomes. This gap, driven by mutually reinforcing cognitive biases in both humans and artificial intelligence (AI) systems, threatens the defensibility of valuations and sustainability of investments in the sector. This report describes a framework emerging from sy
-
cs.AI, q-bio.NC updates on arXiv.org
-
Clinical Uncertainty Impacts Machine Learning Evaluations
arXiv:2509.22242v2 Announce Type: replace Abstract: Clinical dataset labels are rarely certain as annotators disagree and confidence is not uniform across cases. Typical aggregation procedures, such as majority voting, obscure this variability. In simple experiments on medical imaging benchmarks, accounting for the confidence in binary labels significantly impacts model rankings. We therefore argue that machine-learning evaluations should explicitly account for annotation uncertainty using prob
Clinical Uncertainty Impacts Machine Learning Evaluations
-
cs.AI, q-bio.NC updates on arXiv.org
-
TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework
arXiv:2511.05385v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models' (LLMs) reliability. For flexibility, agentic RAG employs autonomous, multi-round retrieval and reasoning to resolve queries. Although recent agentic RAG has improved via reinforcement learning, they often incur substantial token overhead from search and reasoning processes. This trade-off prioritizes accuracy over efficiency. To address this issue,
TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework
-
cs.AI, q-bio.NC updates on arXiv.org
-
TOBUGraph: Knowledge Graph-Based Retrieval for Enhanced LLM Performance Beyond RAG
arXiv:2412.05447v3 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) is one of the leading and most widely used techniques for enhancing LLM retrieval capabilities, but it still faces significant limitations in commercial use cases. RAG primarily relies on the query-chunk text-to-text similarity in the embedding space for retrieval and can fail to capture deeper semantic relationships across chunks, is highly sensitive to chunking strategies, and is prone to hallucinat
TOBUGraph: Knowledge Graph-Based Retrieval for Enhanced LLM Performance Beyond RAG
-
Nature Medicine
-
Author Correction: Global burden of chikungunya virus infections and the potential benefit of vaccination campaigns
Nature Medicine, Published online: 10 November 2025; doi:10.1038/s41591-025-04065-zAuthor Correction: Global burden of chikungunya virus infections and the potential benefit of vaccination campaigns
Author Correction: Global burden of chikungunya virus infections and the potential benefit of vaccination campaigns
Nature Medicine, Published online: 10 November 2025; doi:10.1038/s41591-025-04065-z
Author Correction: Global burden of chikungunya virus infections and the potential benefit of vaccination campaigns-
Nature Biotechnology - Issue - nature.com science feeds
-
Publisher Correction: Deep-learning-based virtual screening of antibacterial compounds
Nature Biotechnology, Published online: 07 November 2025; doi:10.1038/s41587-025-02941-0Publisher Correction: Deep-learning-based virtual screening of antibacterial compounds
Publisher Correction: Deep-learning-based virtual screening of antibacterial compounds
Nature Biotechnology, Published online: 07 November 2025; doi:10.1038/s41587-025-02941-0
Publisher Correction: Deep-learning-based virtual screening of antibacterial compounds-
cs.AI, q-bio.NC updates on arXiv.org
-
Hybrid Fact-Checking that Integrates Knowledge Graphs, Large Language Models, and Search-Based Retrieval Agents Improves Interpretable Claim Verification
arXiv:2511.03217v1 Announce Type: cross Abstract: Large language models (LLMs) excel in generating fluent utterances but can lack reliable grounding in verified information. At the same time, knowledge-graph-based fact-checkers deliver precise and interpretable evidence, yet suffer from limited coverage or latency. By integrating LLMs with knowledge graphs and real-time search agents, we introduce a hybrid fact-checking approach that leverages the individual strengths of each component. Our sys
Hybrid Fact-Checking that Integrates Knowledge Graphs, Large Language Models, and Search-Based Retrieval Agents Improves Interpretable Claim Verification
-
Nature Biotechnology - Issue - nature.com science feeds
-
Site-specific DNA insertion into the human genome with engineered recombinases
Nature Biotechnology, Published online: 06 November 2025; doi:10.1038/s41587-025-02895-3Engineered DNA recombinases efficiently and specifically insert genetic cargos without the use of landing pads.
Site-specific DNA insertion into the human genome with engineered recombinases
Nature Biotechnology, Published online: 06 November 2025; doi:10.1038/s41587-025-02895-3
Engineered DNA recombinases efficiently and specifically insert genetic cargos without the use of landing pads.-
npj Digital Medicine
-
Improving dataset transparency in dermatologic Artificial Intelligence using a dataset nutrition label
npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02125-9Biased and poorly documented dermatology datasets pose risks to the development of safe and generalizable artificial intelligence (AI) tools. We created a Dataset Nutrition Label (DNL) for multiple dermatology datasets to support transparent and responsible data use. The DNL offers a structured, digestible summary of key attributes, including metadata, limitations, and risks, enabling data users to better ass
Improving dataset transparency in dermatologic Artificial Intelligence using a dataset nutrition label
npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02125-9
Biased and poorly documented dermatology datasets pose risks to the development of safe and generalizable artificial intelligence (AI) tools. We created a Dataset Nutrition Label (DNL) for multiple dermatology datasets to support transparent and responsible data use. The DNL offers a structured, digestible summary of key attributes, including metadata, limitations, and risks, enabling data users to better assess suitability and proactively address potential sources of bias in datasets.-
npj Digital Medicine
-
Evaluating clinical AI summaries with large language models as judges
npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02005-2Evaluating clinical AI summaries with large language models as judges
Evaluating clinical AI summaries with large language models as judges
npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02005-2
Evaluating clinical AI summaries with large language models as judges-
MRD
-
Liquid biopsy in gastrointestinal oncology: clinical applications and translational integration of ctDNA, CTCs, and sEVs
Oncol Rev. 2025 Oct 20;19:1702932. doi: 10.3389/or.2025.1702932. eCollection 2025.ABSTRACTBACKGROUND AND AIMS: Liquid biopsy offers a minimally invasive tool to detect actionable mutations, monitor minimal residual disease (MRD), and guide therapy in gastrointestinal (GI) cancers. We critically review the clinical utility of circulating tumor DNA (ctDNA), circulating tumor cells (CTCs), and small extracellular vesicles (sEVs) across GI malignancies and propose a framework for their integration i
Liquid biopsy in gastrointestinal oncology: clinical applications and translational integration of ctDNA, CTCs, and sEVs
Oncol Rev. 2025 Oct 20;19:1702932. doi: 10.3389/or.2025.1702932. eCollection 2025.
ABSTRACT
BACKGROUND AND AIMS: Liquid biopsy offers a minimally invasive tool to detect actionable mutations, monitor minimal residual disease (MRD), and guide therapy in gastrointestinal (GI) cancers. We critically review the clinical utility of circulating tumor DNA (ctDNA), circulating tumor cells (CTCs), and small extracellular vesicles (sEVs) across GI malignancies and propose a framework for their integration into clinical practice.
METHODS: We synthesized evidence from over 200 studies, including prospective trials and translational research, to assess diagnostic accuracy, prognostic value, and clinical actionability of each biomarker type in esophageal, gastric, colorectal, pancreatic, hepatocellular, and biliary cancers.
RESULTS: ctDNA has shown strong potential for MRD detection and treatment monitoring, particularly in colorectal and pancreatic cancer. CTCs offer insights into metastatic risk and therapeutic resistance, while sEVs provide molecular cargo relevant to immunomodulation and disease progression. Emerging microfluidics and AI-driven multi-omics approaches may overcome current limitations.
CONCLUSION: The integration of liquid biopsy technologies into GI oncology holds promise for early detection and precision therapy. We propose a five-phase clinical roadmap and outine the key research gaps that need to be addressed before widespread implementation in routine care.
PMID:41190015 | PMC:PMC12580207 | DOI:10.3389/or.2025.1702932
-
Journal of Medical Internet Research
-
Combining International Standards to Develop Clinical Decision Support for Parent Smoking Cessation in Pediatrics
Smoking has severe health consequences, and secondhand smoke (SHS) exposure among children increases the risk of sudden infant death syndrome, chronic respiratory diseases, such as asthma, and lung cancer in adulthood. For many parents, pediatricians are the primary source of interaction with the healthcare system. Nevertheless, in pediatric settings, appropriate tobacco treatments are rarely, if ever, provided to parents who smoke. To best address tobacco use among parents, it is ideal to devel
Combining International Standards to Develop Clinical Decision Support for Parent Smoking Cessation in Pediatrics
-
Nature Medicine
-
A multimodal whole-slide foundation model for pathology
Nature Medicine, Published online: 05 November 2025; doi:10.1038/s41591-025-03982-3Pretrained using 335,645 whole-slide images, a foundation model is developed to provide representations for slide- and patient-level tasks. It is capable of performing clinical tasks and generating reports even in data-scarce scenarios, such as rare cancer diagnosis and survival prediction, without requiring further fine-tuning.
A multimodal whole-slide foundation model for pathology
Nature Medicine, Published online: 05 November 2025; doi:10.1038/s41591-025-03982-3
Pretrained using 335,645 whole-slide images, a foundation model is developed to provide representations for slide- and patient-level tasks. It is capable of performing clinical tasks and generating reports even in data-scarce scenarios, such as rare cancer diagnosis and survival prediction, without requiring further fine-tuning.-
Nature - Issue - nature.com science feeds
-
Fair human-centric image dataset for ethical AI benchmarking
Nature, Published online: 05 November 2025; doi:10.1038/s41586-025-09716-2The Fair Human-Centric Image Benchmark (FHIBE, pronounced ‘Feebee’)—an image dataset that implements best practices for consent, privacy, compensation, safety, diversity and utility—can be used responsibly as a fairness evaluation dataset for many human-centric computer vision applications.
Fair human-centric image dataset for ethical AI benchmarking
Nature, Published online: 05 November 2025; doi:10.1038/s41586-025-09716-2
The Fair Human-Centric Image Benchmark (FHIBE, pronounced ‘Feebee’)—an image dataset that implements best practices for consent, privacy, compensation, safety, diversity and utility—can be used responsibly as a fairness evaluation dataset for many human-centric computer vision applications.-
cs.AI, q-bio.NC updates on arXiv.org
-
How can we assess human-agent interactions? Case studies in software agent design
arXiv:2510.09801v2 Announce Type: replace Abstract: LLM-powered agents are both a promising new technology and a source of complexity, where choices about models, tools, and prompting can affect their usefulness. While numerous benchmarks measure agent accuracy across domains, they mostly assume full automation, failing to represent the collaborative nature of real-world use cases. In this paper, we make two major steps towards the rigorous assessment of human-agent interactions. First, we prop
How can we assess human-agent interactions? Case studies in software agent design
-
cs.AI, q-bio.NC updates on arXiv.org
-
Digital Twin based Automatic Reconfiguration of Robotic Systems in Smart Environments
arXiv:2511.00094v1 Announce Type: cross Abstract: Robotic systems have become integral to smart environments, enabling applications ranging from urban surveillance and automated agriculture to industrial automation. However, their effective operation in dynamic settings - such as smart cities and precision farming - is challenged by continuously evolving topographies and environmental conditions. Traditional control systems often struggle to adapt quickly, leading to inefficiencies or operation
Digital Twin based Automatic Reconfiguration of Robotic Systems in Smart Environments
-
cs.AI, q-bio.NC updates on arXiv.org
-
Will Humanity Be Rendered Obsolete by AI?
arXiv:2510.22814v2 Announce Type: replace Abstract: This article analyzes the existential risks artificial intelligence (AI) poses to humanity, tracing the trajectory from current AI to ultraintelligence. Drawing on Irving J. Good and Nick Bostrom's theoretical work, plus recent publications (AI 2027; If Anyone Builds It, Everyone Dies), it explores AGI and superintelligence. Considering machines' exponentially growing cognitive power and hypothetical IQs, it addresses the ethical and existenti
Will Humanity Be Rendered Obsolete by AI?
-
cs.AI, q-bio.NC updates on arXiv.org
-
The Denario project: Deep knowledge AI agents for scientific discovery
arXiv:2510.26887v1 Announce Type: new Abstract: We present Denario, an AI multi-agent system designed to serve as a scientific research assistant. Denario can perform many different tasks, such as generating ideas, checking the literature, developing research plans, writing and executing code, making plots, and drafting and reviewing a scientific paper. The system has a modular architecture, allowing it to handle specific tasks, such as generating an idea, or carrying out end-to-end scientific
The Denario project: Deep knowledge AI agents for scientific discovery
-
cs.AI, q-bio.NC updates on arXiv.org
-
Frame Semantic Patterns for Identifying Underreporting of Notifiable Events in Healthcare: The Case of Gender-Based Violence
arXiv:2510.26969v1 Announce Type: cross Abstract: We introduce a methodology for the identification of notifiable events in the domain of healthcare. The methodology harnesses semantic frames to define fine-grained patterns and search them in unstructured data, namely, open-text fields in e-medical records. We apply the methodology to the problem of underreporting of gender-based violence (GBV) in e-medical records produced during patients' visits to primary care units. A total of eight pattern
Frame Semantic Patterns for Identifying Underreporting of Notifiable Events in Healthcare: The Case of Gender-Based Violence
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Systematic Literature Review of Spatio-Temporal Graph Neural Network Models for Time Series Forecasting and Classification
arXiv:2410.22377v3 Announce Type: replace-cross Abstract: In recent years, spatio-temporal graph neural networks (GNNs) have attracted considerable interest in the field of time series analysis, due to their ability to capture, at once, dependencies among variables and across time points. The objective of this systematic literature review is hence to provide a comprehensive overview of the various modeling approaches and application domains of GNNs for time series classification and forecasting