Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
No-Human in the Loop: Agentic Evaluation at Scale for Recommendation
arXiv:2511.03051v1 Announce Type: new Abstract: Evaluating large language models (LLMs) as judges is increasingly critical for building scalable and trustworthy evaluation pipelines. We present ScalingEval, a large-scale benchmarking study that systematically compares 36 LLMs, including GPT, Gemini, Claude, and Llama, across multiple product categories using a consensus-driven evaluation protocol. Our multi-agent framework aggregates pattern audits and issue codes into ground-truth labels via s
-
cs.AI, q-bio.NC updates on arXiv.org
-
Explaining Decisions in ML Models: a Parameterized Complexity Analysis (Part I)
arXiv:2511.03545v1 Announce Type: new Abstract: This paper presents a comprehensive theoretical investigation into the parameterized complexity of explanation problems in various machine learning (ML) models. Contrary to the prevalent black-box perception, our study focuses on models with transparent internal mechanisms. We address two principal types of explanation problems: abductive and contrastive, both in their local and global variants. Our analysis encompasses diverse ML models, includin
Explaining Decisions in ML Models: a Parameterized Complexity Analysis (Part I)
-
cs.AI, q-bio.NC updates on arXiv.org
-
Digital Transformation Chatbot (DTchatbot): Integrating Large Language Model-based Chatbot in Acquiring Digital Transformation Needs
arXiv:2511.02842v1 Announce Type: cross Abstract: Many organisations pursue digital transformation to enhance operational efficiency, reduce manual efforts, and optimise processes by automation and digital tools. To achieve this, a comprehensive understanding of their unique needs is required. However, traditional methods, such as expert interviews, while effective, face several challenges, including scheduling conflicts, resource constraints, inconsistency, etc. To tackle these issues, we inve
Digital Transformation Chatbot (DTchatbot): Integrating Large Language Model-based Chatbot in Acquiring Digital Transformation Needs
-
cs.AI, q-bio.NC updates on arXiv.org
-
Generative Artificial Intelligence in Bioinformatics: A Systematic Review of Models, Applications, and Methodological Advances
arXiv:2511.03354v1 Announce Type: cross Abstract: Generative artificial intelligence (GenAI) has become a transformative approach in bioinformatics that often enables advancements in genomics, proteomics, transcriptomics, structural biology, and drug discovery. To systematically identify and evaluate these growing developments, this review proposed six research questions (RQs), according to the preferred reporting items for systematic reviews and meta-analysis methods. The objective is to evalu
Generative Artificial Intelligence in Bioinformatics: A Systematic Review of Models, Applications, and Methodological Advances
-
cs.AI, q-bio.NC updates on arXiv.org
-
REFA: Reference Free Alignment for multi-preference optimization
arXiv:2412.16378v4 Announce Type: replace-cross Abstract: To mitigate reward hacking from response verbosity, modern preference optimization methods are increasingly adopting length normalization (e.g., SimPO, ORPO, LN-DPO). While effective against this bias, we demonstrate that length normalization itself introduces a failure mode: the URSLA shortcut. Here models learn to satisfy the alignment objective by prematurely truncating low-quality responses rather than learning from their semantic co
REFA: Reference Free Alignment for multi-preference optimization
-
Nature Biotechnology - Issue - nature.com science feeds
-
Drugmakers share data to feed voracious foundation models
Nature Biotechnology, Published online: 06 November 2025; doi:10.1038/s41587-025-02901-8Big pharma shares its machine learning models with biotechs, but awaits definitive data on success of artificial intelligence-generated drugs.
Drugmakers share data to feed voracious foundation models
Nature Biotechnology, Published online: 06 November 2025; doi:10.1038/s41587-025-02901-8
Big pharma shares its machine learning models with biotechs, but awaits definitive data on success of artificial intelligence-generated drugs.-
Nature Biotechnology - Issue - nature.com science feeds
-
Site-specific DNA insertion into the human genome with engineered recombinases
Nature Biotechnology, Published online: 06 November 2025; doi:10.1038/s41587-025-02895-3Engineered DNA recombinases efficiently and specifically insert genetic cargos without the use of landing pads.
Site-specific DNA insertion into the human genome with engineered recombinases
Nature Biotechnology, Published online: 06 November 2025; doi:10.1038/s41587-025-02895-3
Engineered DNA recombinases efficiently and specifically insert genetic cargos without the use of landing pads.-
STAT

-
STAT+: What’s FDA plotting for therapy chatbot regulation?
You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday. What to know about the FDA’s therapy bots meeting The Food and Drug Administration is considering whether and how to regulate therapy chatbots that are based on large language models. Today, the agency’s Digital Health Advisory Committee is meeting to consider the topic. In a new story, I explai
STAT+: What’s FDA plotting for therapy chatbot regulation?
You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday.
What to know about the FDA’s therapy bots meeting
The Food and Drug Administration is considering whether and how to regulate therapy chatbots that are based on large language models. Today, the agency’s Digital Health Advisory Committee is meeting to consider the topic. In a new story, I explain what’s going on, including some fresh insider intel.
The FDA wants to provide more clarity to developers of generative AI medical devices about what needs regulatory green light and how to get it. The agency is also also worried about LLM-based therapy bots that can provide unpredictable outputs. Regulators are aware about the growing concerns around general purpose bots like ChatGPT, which have been linked to delusions and allegedly to suicides.
Continue to STAT+ to read the full story…


© Sarah Silbiger/Getty Images
-
npj Digital Medicine
-
Improving dataset transparency in dermatologic Artificial Intelligence using a dataset nutrition label
npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02125-9Biased and poorly documented dermatology datasets pose risks to the development of safe and generalizable artificial intelligence (AI) tools. We created a Dataset Nutrition Label (DNL) for multiple dermatology datasets to support transparent and responsible data use. The DNL offers a structured, digestible summary of key attributes, including metadata, limitations, and risks, enabling data users to better ass
Improving dataset transparency in dermatologic Artificial Intelligence using a dataset nutrition label
npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02125-9
Biased and poorly documented dermatology datasets pose risks to the development of safe and generalizable artificial intelligence (AI) tools. We created a Dataset Nutrition Label (DNL) for multiple dermatology datasets to support transparent and responsible data use. The DNL offers a structured, digestible summary of key attributes, including metadata, limitations, and risks, enabling data users to better assess suitability and proactively address potential sources of bias in datasets.-
npj Digital Medicine
-
Evaluating clinical AI summaries with large language models as judges
npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02005-2Evaluating clinical AI summaries with large language models as judges
Evaluating clinical AI summaries with large language models as judges
npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02005-2
Evaluating clinical AI summaries with large language models as judges