Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
No-Human in the Loop: Agentic Evaluation at Scale for Recommendation
arXiv:2511.03051v1 Announce Type: new Abstract: Evaluating large language models (LLMs) as judges is increasingly critical for building scalable and trustworthy evaluation pipelines. We present ScalingEval, a large-scale benchmarking study that systematically compares 36 LLMs, including GPT, Gemini, Claude, and Llama, across multiple product categories using a consensus-driven evaluation protocol. Our multi-agent framework aggregates pattern audits and issue codes into ground-truth labels via s
-
cs.AI, q-bio.NC updates on arXiv.org
-
Explaining Decisions in ML Models: a Parameterized Complexity Analysis (Part I)
arXiv:2511.03545v1 Announce Type: new Abstract: This paper presents a comprehensive theoretical investigation into the parameterized complexity of explanation problems in various machine learning (ML) models. Contrary to the prevalent black-box perception, our study focuses on models with transparent internal mechanisms. We address two principal types of explanation problems: abductive and contrastive, both in their local and global variants. Our analysis encompasses diverse ML models, includin
Explaining Decisions in ML Models: a Parameterized Complexity Analysis (Part I)
-
cs.AI, q-bio.NC updates on arXiv.org
-
Digital Transformation Chatbot (DTchatbot): Integrating Large Language Model-based Chatbot in Acquiring Digital Transformation Needs
arXiv:2511.02842v1 Announce Type: cross Abstract: Many organisations pursue digital transformation to enhance operational efficiency, reduce manual efforts, and optimise processes by automation and digital tools. To achieve this, a comprehensive understanding of their unique needs is required. However, traditional methods, such as expert interviews, while effective, face several challenges, including scheduling conflicts, resource constraints, inconsistency, etc. To tackle these issues, we inve
Digital Transformation Chatbot (DTchatbot): Integrating Large Language Model-based Chatbot in Acquiring Digital Transformation Needs
-
cs.AI, q-bio.NC updates on arXiv.org
-
Generative Artificial Intelligence in Bioinformatics: A Systematic Review of Models, Applications, and Methodological Advances
arXiv:2511.03354v1 Announce Type: cross Abstract: Generative artificial intelligence (GenAI) has become a transformative approach in bioinformatics that often enables advancements in genomics, proteomics, transcriptomics, structural biology, and drug discovery. To systematically identify and evaluate these growing developments, this review proposed six research questions (RQs), according to the preferred reporting items for systematic reviews and meta-analysis methods. The objective is to evalu
Generative Artificial Intelligence in Bioinformatics: A Systematic Review of Models, Applications, and Methodological Advances
-
cs.AI, q-bio.NC updates on arXiv.org
-
REFA: Reference Free Alignment for multi-preference optimization
arXiv:2412.16378v4 Announce Type: replace-cross Abstract: To mitigate reward hacking from response verbosity, modern preference optimization methods are increasingly adopting length normalization (e.g., SimPO, ORPO, LN-DPO). While effective against this bias, we demonstrate that length normalization itself introduces a failure mode: the URSLA shortcut. Here models learn to satisfy the alignment objective by prematurely truncating low-quality responses rather than learning from their semantic co
REFA: Reference Free Alignment for multi-preference optimization
-
Nature Biotechnology - Issue - nature.com science feeds
-
Drugmakers share data to feed voracious foundation models
Nature Biotechnology, Published online: 06 November 2025; doi:10.1038/s41587-025-02901-8Big pharma shares its machine learning models with biotechs, but awaits definitive data on success of artificial intelligence-generated drugs.
Drugmakers share data to feed voracious foundation models
Nature Biotechnology, Published online: 06 November 2025; doi:10.1038/s41587-025-02901-8
Big pharma shares its machine learning models with biotechs, but awaits definitive data on success of artificial intelligence-generated drugs.-
Nature Biotechnology - Issue - nature.com science feeds
-
Site-specific DNA insertion into the human genome with engineered recombinases
Nature Biotechnology, Published online: 06 November 2025; doi:10.1038/s41587-025-02895-3Engineered DNA recombinases efficiently and specifically insert genetic cargos without the use of landing pads.
Site-specific DNA insertion into the human genome with engineered recombinases
Nature Biotechnology, Published online: 06 November 2025; doi:10.1038/s41587-025-02895-3
Engineered DNA recombinases efficiently and specifically insert genetic cargos without the use of landing pads.-
STAT

-
STAT+: What’s FDA plotting for therapy chatbot regulation?
You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday. What to know about the FDA’s therapy bots meeting The Food and Drug Administration is considering whether and how to regulate therapy chatbots that are based on large language models. Today, the agency’s Digital Health Advisory Committee is meeting to consider the topic. In a new story, I explai
STAT+: What’s FDA plotting for therapy chatbot regulation?
You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday.
What to know about the FDA’s therapy bots meeting
The Food and Drug Administration is considering whether and how to regulate therapy chatbots that are based on large language models. Today, the agency’s Digital Health Advisory Committee is meeting to consider the topic. In a new story, I explain what’s going on, including some fresh insider intel.
The FDA wants to provide more clarity to developers of generative AI medical devices about what needs regulatory green light and how to get it. The agency is also also worried about LLM-based therapy bots that can provide unpredictable outputs. Regulators are aware about the growing concerns around general purpose bots like ChatGPT, which have been linked to delusions and allegedly to suicides.
Continue to STAT+ to read the full story…


© Sarah Silbiger/Getty Images
-
npj Digital Medicine
-
Improving dataset transparency in dermatologic Artificial Intelligence using a dataset nutrition label
npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02125-9Biased and poorly documented dermatology datasets pose risks to the development of safe and generalizable artificial intelligence (AI) tools. We created a Dataset Nutrition Label (DNL) for multiple dermatology datasets to support transparent and responsible data use. The DNL offers a structured, digestible summary of key attributes, including metadata, limitations, and risks, enabling data users to better ass
Improving dataset transparency in dermatologic Artificial Intelligence using a dataset nutrition label
npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02125-9
Biased and poorly documented dermatology datasets pose risks to the development of safe and generalizable artificial intelligence (AI) tools. We created a Dataset Nutrition Label (DNL) for multiple dermatology datasets to support transparent and responsible data use. The DNL offers a structured, digestible summary of key attributes, including metadata, limitations, and risks, enabling data users to better assess suitability and proactively address potential sources of bias in datasets.-
npj Digital Medicine
-
Evaluating clinical AI summaries with large language models as judges
npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02005-2Evaluating clinical AI summaries with large language models as judges
Evaluating clinical AI summaries with large language models as judges
npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02005-2
Evaluating clinical AI summaries with large language models as judges-
MRD
-
Liquid biopsy in gastrointestinal oncology: clinical applications and translational integration of ctDNA, CTCs, and sEVs
Oncol Rev. 2025 Oct 20;19:1702932. doi: 10.3389/or.2025.1702932. eCollection 2025.ABSTRACTBACKGROUND AND AIMS: Liquid biopsy offers a minimally invasive tool to detect actionable mutations, monitor minimal residual disease (MRD), and guide therapy in gastrointestinal (GI) cancers. We critically review the clinical utility of circulating tumor DNA (ctDNA), circulating tumor cells (CTCs), and small extracellular vesicles (sEVs) across GI malignancies and propose a framework for their integration i
Liquid biopsy in gastrointestinal oncology: clinical applications and translational integration of ctDNA, CTCs, and sEVs
Oncol Rev. 2025 Oct 20;19:1702932. doi: 10.3389/or.2025.1702932. eCollection 2025.
ABSTRACT
BACKGROUND AND AIMS: Liquid biopsy offers a minimally invasive tool to detect actionable mutations, monitor minimal residual disease (MRD), and guide therapy in gastrointestinal (GI) cancers. We critically review the clinical utility of circulating tumor DNA (ctDNA), circulating tumor cells (CTCs), and small extracellular vesicles (sEVs) across GI malignancies and propose a framework for their integration into clinical practice.
METHODS: We synthesized evidence from over 200 studies, including prospective trials and translational research, to assess diagnostic accuracy, prognostic value, and clinical actionability of each biomarker type in esophageal, gastric, colorectal, pancreatic, hepatocellular, and biliary cancers.
RESULTS: ctDNA has shown strong potential for MRD detection and treatment monitoring, particularly in colorectal and pancreatic cancer. CTCs offer insights into metastatic risk and therapeutic resistance, while sEVs provide molecular cargo relevant to immunomodulation and disease progression. Emerging microfluidics and AI-driven multi-omics approaches may overcome current limitations.
CONCLUSION: The integration of liquid biopsy technologies into GI oncology holds promise for early detection and precision therapy. We propose a five-phase clinical roadmap and outine the key research gaps that need to be addressed before widespread implementation in routine care.
PMID:41190015 | PMC:PMC12580207 | DOI:10.3389/or.2025.1702932
-
Journal of Medical Internet Research
-
Combining International Standards to Develop Clinical Decision Support for Parent Smoking Cessation in Pediatrics
Smoking has severe health consequences, and secondhand smoke (SHS) exposure among children increases the risk of sudden infant death syndrome, chronic respiratory diseases, such as asthma, and lung cancer in adulthood. For many parents, pediatricians are the primary source of interaction with the healthcare system. Nevertheless, in pediatric settings, appropriate tobacco treatments are rarely, if ever, provided to parents who smoke. To best address tobacco use among parents, it is ideal to devel
Combining International Standards to Develop Clinical Decision Support for Parent Smoking Cessation in Pediatrics
-
Nature Medicine
-
A multimodal whole-slide foundation model for pathology
Nature Medicine, Published online: 05 November 2025; doi:10.1038/s41591-025-03982-3Pretrained using 335,645 whole-slide images, a foundation model is developed to provide representations for slide- and patient-level tasks. It is capable of performing clinical tasks and generating reports even in data-scarce scenarios, such as rare cancer diagnosis and survival prediction, without requiring further fine-tuning.
A multimodal whole-slide foundation model for pathology
Nature Medicine, Published online: 05 November 2025; doi:10.1038/s41591-025-03982-3
Pretrained using 335,645 whole-slide images, a foundation model is developed to provide representations for slide- and patient-level tasks. It is capable of performing clinical tasks and generating reports even in data-scarce scenarios, such as rare cancer diagnosis and survival prediction, without requiring further fine-tuning.-
Nature - Issue - nature.com science feeds
-
Fair human-centric image dataset for ethical AI benchmarking
Nature, Published online: 05 November 2025; doi:10.1038/s41586-025-09716-2The Fair Human-Centric Image Benchmark (FHIBE, pronounced ‘Feebee’)—an image dataset that implements best practices for consent, privacy, compensation, safety, diversity and utility—can be used responsibly as a fairness evaluation dataset for many human-centric computer vision applications.
Fair human-centric image dataset for ethical AI benchmarking
Nature, Published online: 05 November 2025; doi:10.1038/s41586-025-09716-2
The Fair Human-Centric Image Benchmark (FHIBE, pronounced ‘Feebee’)—an image dataset that implements best practices for consent, privacy, compensation, safety, diversity and utility—can be used responsibly as a fairness evaluation dataset for many human-centric computer vision applications.-
cs.AI, q-bio.NC updates on arXiv.org
-
Causal Graph Neural Networks for Healthcare
arXiv:2511.02531v1 Announce Type: cross Abstract: Healthcare artificial intelligence systems routinely fail when deployed across institutions, with documented performance drops and perpetuation of discriminatory patterns embedded in historical data. This brittleness stems, in part, from learning statistical associations rather than causal mechanisms. Causal graph neural networks address this triple crisis of distribution shift, discrimination, and inscrutability by combining graph-based represe
Causal Graph Neural Networks for Healthcare
-
cs.AI, q-bio.NC updates on arXiv.org
-
TabTune: A Unified Library for Inference and Fine-Tuning Tabular Foundation Models
arXiv:2511.02802v1 Announce Type: cross Abstract: Tabular foundation models represent a growing paradigm in structured data learning, extending the benefits of large-scale pretraining to tabular domains. However, their adoption remains limited due to heterogeneous preprocessing pipelines, fragmented APIs, inconsistent fine-tuning procedures, and the absence of standardized evaluation for deployment-oriented metrics such as calibration and fairness. We present TabTune, a unified library that sta
TabTune: A Unified Library for Inference and Fine-Tuning Tabular Foundation Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
How can we assess human-agent interactions? Case studies in software agent design
arXiv:2510.09801v2 Announce Type: replace Abstract: LLM-powered agents are both a promising new technology and a source of complexity, where choices about models, tools, and prompting can affect their usefulness. While numerous benchmarks measure agent accuracy across domains, they mostly assume full automation, failing to represent the collaborative nature of real-world use cases. In this paper, we make two major steps towards the rigorous assessment of human-agent interactions. First, we prop
How can we assess human-agent interactions? Case studies in software agent design
-
cs.AI, q-bio.NC updates on arXiv.org
-
AutoPDL: Automatic Prompt Optimization for LLM Agents
arXiv:2504.04365v5 Announce Type: replace-cross Abstract: The performance of large language models (LLMs) depends on how they are prompted, with choices spanning both the high-level prompting pattern (e.g., Zero-Shot, CoT, ReAct, ReWOO) and the specific prompt content (instructions and few-shot demonstrations). Manually tuning this combination is tedious, error-prone, and specific to a given LLM and task. Therefore, this paper proposes AutoPDL, an automated approach to discovering good LLM agen
AutoPDL: Automatic Prompt Optimization for LLM Agents
-
cs.AI, q-bio.NC updates on arXiv.org
-
Diffusion Models at the Drug Discovery Frontier: A Review on Generating Small Molecules versus Therapeutic Peptides
arXiv:2511.00209v1 Announce Type: cross Abstract: Diffusion models have emerged as a leading framework in generative modeling, showing significant potential to accelerate and transform the traditionally slow and costly process of drug discovery. This review provides a systematic comparison of their application in designing two principal therapeutic modalities: small molecules and therapeutic peptides. We analyze how a unified framework of iterative denoising is adapted to the distinct molecular
Diffusion Models at the Drug Discovery Frontier: A Review on Generating Small Molecules versus Therapeutic Peptides
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation
arXiv:2510.19755v3 Announce Type: replace-cross Abstract: Diffusion Models have become a cornerstone of modern generative AI for their exceptional generation quality and controllability. However, their inherent \textit{multi-step iterations} and \textit{complex backbone networks} lead to prohibitive computational overhead and generation latency, forming a major bottleneck for real-time applications. Although existing acceleration techniques have made progress, they still face challenges such as