❌

Reading view

No-Human in the Loop: Agentic Evaluation at Scale for Recommendation

arXiv:2511.03051v1 Announce Type: new Abstract: Evaluating large language models (LLMs) as judges is increasingly critical for building scalable and trustworthy evaluation pipelines. We present ScalingEval, a large-scale benchmarking study that systematically compares 36 LLMs, including GPT, Gemini, Claude, and Llama, across multiple product categories using a consensus-driven evaluation protocol. Our multi-agent framework aggregates pattern audits and issue codes into ground-truth labels via scalable majority voting, enabling reproducible comparison of LLM evaluators without human annotation. Applied to large-scale complementary-item recommendation, the benchmark reports four key findings: (i) Anthropic Claude 3.5 Sonnet achieves the highest decision confidence; (ii) Gemini 1.5 Pro offers the best overall performance across categories; (iii) GPT-4o provides the most favorable latency-accuracy-cost tradeoff; and (iv) GPT-OSS 20B leads among open-source models. Category-level analysis shows strong consensus in structured domains (Electronics, Sports) but persistent disagreement in lifestyle categories (Clothing, Food). These results establish ScalingEval as a reproducible benchmark and evaluation protocol for LLMs as judges, with actionable guidance on scaling, reliability, and model family tradeoffs.
  •  

Explaining Decisions in ML Models: a Parameterized Complexity Analysis (Part I)

arXiv:2511.03545v1 Announce Type: new Abstract: This paper presents a comprehensive theoretical investigation into the parameterized complexity of explanation problems in various machine learning (ML) models. Contrary to the prevalent black-box perception, our study focuses on models with transparent internal mechanisms. We address two principal types of explanation problems: abductive and contrastive, both in their local and global variants. Our analysis encompasses diverse ML models, including Decision Trees, Decision Sets, Decision Lists, Boolean Circuits, and ensembles thereof, each offering unique explanatory challenges. This research fills a significant gap in explainable AI (XAI) by providing a foundational understanding of the complexities of generating explanations for these models. This work provides insights vital for further research in the domain of XAI, contributing to the broader discourse on the necessity of transparency and accountability in AI systems.
  •  

Digital Transformation Chatbot (DTchatbot): Integrating Large Language Model-based Chatbot in Acquiring Digital Transformation Needs

arXiv:2511.02842v1 Announce Type: cross Abstract: Many organisations pursue digital transformation to enhance operational efficiency, reduce manual efforts, and optimise processes by automation and digital tools. To achieve this, a comprehensive understanding of their unique needs is required. However, traditional methods, such as expert interviews, while effective, face several challenges, including scheduling conflicts, resource constraints, inconsistency, etc. To tackle these issues, we investigate the use of a Large Language Model (LLM)-powered chatbot to acquire organisations' digital transformation needs. Specifically, the chatbot integrates workflow-based instruction with LLM's planning and reasoning capabilities, enabling it to function as a virtual expert and conduct interviews. We detail the chatbot's features and its implementation. Our preliminary evaluation indicates that the chatbot performs as designed, effectively following predefined workflows and supporting user interactions with areas for improvement. We conclude by discussing the implications of employing chatbots to elicit user information, emphasizing their potential and limitations.
  •  

Generative Artificial Intelligence in Bioinformatics: A Systematic Review of Models, Applications, and Methodological Advances

arXiv:2511.03354v1 Announce Type: cross Abstract: Generative artificial intelligence (GenAI) has become a transformative approach in bioinformatics that often enables advancements in genomics, proteomics, transcriptomics, structural biology, and drug discovery. To systematically identify and evaluate these growing developments, this review proposed six research questions (RQs), according to the preferred reporting items for systematic reviews and meta-analysis methods. The objective is to evaluate impactful GenAI strategies in methodological advancement, predictive performance, and specialization, and to identify promising approaches for advanced modeling, data-intensive discovery, and integrative biological analysis. RQ1 highlights diverse applications across multiple bioinformatics subfields (sequence analysis, molecular design, and integrative data modeling), which demonstrate superior performance over traditional methods through pattern recognition and output generation. RQ2 reveals that adapted specialized model architectures outperformed general-purpose models, an advantage attributed to targeted pretraining and context-aware strategies. RQ3 identifies significant benefits in the bioinformatics domains, focusing on molecular analysis and data integration, which improves accuracy and reduces errors in complex analysis. RQ4 indicates improvements in structural modeling, functional prediction, and synthetic data generation, validated by established benchmarks. RQ5 suggests the main constraints, such as the lack of scalability and biases in data that impact generalizability, and proposes future directions focused on robust evaluation and biologically grounded modeling. RQ6 examines that molecular datasets (such as UniProtKB and ProteinNet12), cellular datasets (such as CELLxGENE and GTEx) and textual resources (such as PubMedQA and OMIM) broadly support the training and generalization of GenAI models.
  •  

REFA: Reference Free Alignment for multi-preference optimization

arXiv:2412.16378v4 Announce Type: replace-cross Abstract: To mitigate reward hacking from response verbosity, modern preference optimization methods are increasingly adopting length normalization (e.g., SimPO, ORPO, LN-DPO). While effective against this bias, we demonstrate that length normalization itself introduces a failure mode: the URSLA shortcut. Here models learn to satisfy the alignment objective by prematurely truncating low-quality responses rather than learning from their semantic content. To address this, we introduce REFA, a new alignment framework that proposes probabilistic control on a structural token that controls termination. Our core innovation is a new class of regularizers that operate directly on the probability of the End-of-Sequence (EOS) token, a previously unexploited control lever. This token-level intervention provides a principled solution to the URSLA shortcut, ensuring genuine quality improvements. Furthermore, it unlocks a versatile mechanism for managing the alignment-efficiency tradeoff, enabling practitioners to fine-tune models that adhere to specific token budgets. Empirically, REFA achieves a 60.29% win rate and a 52.17% length-controlled win rate on AlpacaEval2 with Llama-3-8B-Instruct, demonstrating the power of our token-level control paradigm.
  •  

Site-specific DNA insertion into the human genome with engineered recombinases

Nature Biotechnology, Published online: 06 November 2025; doi:10.1038/s41587-025-02895-3

Engineered DNA recombinases efficiently and specifically insert genetic cargos without the use of landing pads.
  •  

STAT+: What’s FDA plotting for therapy chatbot regulation?

You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday.

What to know about the FDA’s therapy bots meeting

The Food and Drug Administration is considering whether and how to regulate therapy chatbots that are based on large language models. Today, the agency’s Digital Health Advisory Committee is meeting to consider the topic. In a new story, I explain what’s going on, including some fresh insider intel.

The FDA wants to provide more clarity to developers of generative AI medical devices about what needs regulatory green light and how to get it. The agency is also also worried about LLM-based therapy bots that can provide unpredictable outputs. Regulators are aware about the growing concerns around general purpose bots like ChatGPT, which have been linked to delusions and allegedly to suicides.

Continue to STAT+ to read the full story…

© Sarah Silbiger/Getty Images

  •  

Improving dataset transparency in dermatologic Artificial Intelligence using a dataset nutrition label

npj Digital Medicine, Published online: 05 November 2025; doi:10.1038/s41746-025-02125-9

Biased and poorly documented dermatology datasets pose risks to the development of safe and generalizable artificial intelligence (AI) tools. We created a Dataset Nutrition Label (DNL) for multiple dermatology datasets to support transparent and responsible data use. The DNL offers a structured, digestible summary of key attributes, including metadata, limitations, and risks, enabling data users to better assess suitability and proactively address potential sources of bias in datasets.
  •  
  •  

Liquid biopsy in gastrointestinal oncology: clinical applications and translational integration of ctDNA, CTCs, and sEVs

Oncol Rev. 2025 Oct 20;19:1702932. doi: 10.3389/or.2025.1702932. eCollection 2025.

ABSTRACT

BACKGROUND AND AIMS: Liquid biopsy offers a minimally invasive tool to detect actionable mutations, monitor minimal residual disease (MRD), and guide therapy in gastrointestinal (GI) cancers. We critically review the clinical utility of circulating tumor DNA (ctDNA), circulating tumor cells (CTCs), and small extracellular vesicles (sEVs) across GI malignancies and propose a framework for their integration into clinical practice.

METHODS: We synthesized evidence from over 200 studies, including prospective trials and translational research, to assess diagnostic accuracy, prognostic value, and clinical actionability of each biomarker type in esophageal, gastric, colorectal, pancreatic, hepatocellular, and biliary cancers.

RESULTS: ctDNA has shown strong potential for MRD detection and treatment monitoring, particularly in colorectal and pancreatic cancer. CTCs offer insights into metastatic risk and therapeutic resistance, while sEVs provide molecular cargo relevant to immunomodulation and disease progression. Emerging microfluidics and AI-driven multi-omics approaches may overcome current limitations.

CONCLUSION: The integration of liquid biopsy technologies into GI oncology holds promise for early detection and precision therapy. We propose a five-phase clinical roadmap and outine the key research gaps that need to be addressed before widespread implementation in routine care.

PMID:41190015 | PMC:PMC12580207 | DOI:10.3389/or.2025.1702932

  •  

Combining International Standards to Develop Clinical Decision Support for Parent Smoking Cessation in Pediatrics

Smoking has severe health consequences, and secondhand smoke (SHS) exposure among children increases the risk of sudden infant death syndrome, chronic respiratory diseases, such as asthma, and lung cancer in adulthood. For many parents, pediatricians are the primary source of interaction with the healthcare system. Nevertheless, in pediatric settings, appropriate tobacco treatments are rarely, if ever, provided to parents who smoke. To best address tobacco use among parents, it is ideal to develop scalable solutions that are coordinated across health systems, community partners, and national services within pediatric settings. We describe our experience developing and implementing a parent tobacco treatment platform (PTTP) within a pediatric institution that leverages multiple international standards to support interoperability, with the overarching goal of providing a model for how such work can be approached. The clinical decision support (CDS) system includes clinician- and patient-facing components, connects parents to three different treatment options (nicotine replacement therapy, text-based counseling, and telephonic counseling), and incorporates three international standards (Fast Healthcare Interoperability Resources [FHIR], SMART on FHIR, and CDS Hooks). FHIR is used across all components. SMART on FHIR is limited to the clinician-facing tool, and CDS Hooks is used in the patient-facing portion. While healthcare interoperability standards supported a significant portion of the overall system, non-standard technologies and enhancements of existing standards were also required. Further, no connections with community partners could use existing interoperability standards. Over one year, the CDS was used in 194,946 visits, identified 7,847 parents who smoke, and connected 2,954 parents to 6,320 distinct treatment services, a significant improvement compared to prior efforts. Our project demonstrates that building CDS systems using international standards, such as SMART on FHIR, FHIR, and CDS Hooks, is possible, but challenges remain. Limits in the CDS Hooks standard to support common workflows and a lack of communication standards used by 3rd parties outside the healthcare system represent areas for future work. To support these requirements, additional EHR-specific records and communication mechanisms were required.
  •  

A multimodal whole-slide foundation model for pathology

Nature Medicine, Published online: 05 November 2025; doi:10.1038/s41591-025-03982-3

Pretrained using 335,645 whole-slide images, a foundation model is developed to provide representations for slide- and patient-level tasks. It is capable of performing clinical tasks and generating reports even in data-scarce scenarios, such as rare cancer diagnosis and survival prediction, without requiring further fine-tuning.
  •  

Fair human-centric image dataset for ethical AI benchmarking

Nature, Published online: 05 November 2025; doi:10.1038/s41586-025-09716-2

The Fair Human-Centric Image Benchmark (FHIBE, pronounced ‘Feebee’)—an image dataset that implements best practices for consent, privacy, compensation, safety, diversity and utility—can be used responsibly as a fairness evaluation dataset for many human-centric computer vision applications.
  •  

Causal Graph Neural Networks for Healthcare

arXiv:2511.02531v1 Announce Type: cross Abstract: Healthcare artificial intelligence systems routinely fail when deployed across institutions, with documented performance drops and perpetuation of discriminatory patterns embedded in historical data. This brittleness stems, in part, from learning statistical associations rather than causal mechanisms. Causal graph neural networks address this triple crisis of distribution shift, discrimination, and inscrutability by combining graph-based representations of biomedical data with causal inference principles to learn invariant mechanisms rather than spurious correlations. This Review examines methodological foundations spanning structural causal models, disentangled causal representation learning, and techniques for interventional prediction and counterfactual reasoning on graphs. We analyse applications demonstrating clinical value across psychiatric diagnosis through brain network analysis, cancer subtyping via multi-omics causal integration, continuous physiological monitoring with mechanistic interpretation, and drug recommendation correcting prescription bias. These advances establish foundations for patient-specific Causal Digital Twins, enabling in silico clinical experimentation, with integration of large language models for hypothesis generation and causal graph neural networks for mechanistic validation. Substantial barriers remain, including computational requirements precluding real-time deployment, validation challenges demanding multi-modal evidence triangulation beyond cross-validation, and risks of causal-washing where methods employ causal terminology without rigorous evidentiary support. We propose tiered frameworks distinguishing causally-inspired architectures from causally-validated discoveries and identify critical research priorities making causal rather than purely associational claims.
  •  

TabTune: A Unified Library for Inference and Fine-Tuning Tabular Foundation Models

arXiv:2511.02802v1 Announce Type: cross Abstract: Tabular foundation models represent a growing paradigm in structured data learning, extending the benefits of large-scale pretraining to tabular domains. However, their adoption remains limited due to heterogeneous preprocessing pipelines, fragmented APIs, inconsistent fine-tuning procedures, and the absence of standardized evaluation for deployment-oriented metrics such as calibration and fairness. We present TabTune, a unified library that standardizes the complete workflow for tabular foundation models through a single interface. TabTune provides consistent access to seven state-of-the-art models supporting multiple adaptation strategies, including zero-shot inference, meta-learning, supervised fine-tuning (SFT), and parameter-efficient fine-tuning (PEFT). The framework automates model-aware preprocessing, manages architectural heterogeneity internally, and integrates evaluation modules for performance, calibration, and fairness. Designed for extensibility and reproducibility, TabTune enables consistent benchmarking of adaptation strategies of tabular foundation models. The library is open source and available at https://github.com/Lexsi-Labs/TabTune .
  •  

How can we assess human-agent interactions? Case studies in software agent design

arXiv:2510.09801v2 Announce Type: replace Abstract: LLM-powered agents are both a promising new technology and a source of complexity, where choices about models, tools, and prompting can affect their usefulness. While numerous benchmarks measure agent accuracy across domains, they mostly assume full automation, failing to represent the collaborative nature of real-world use cases. In this paper, we make two major steps towards the rigorous assessment of human-agent interactions. First, we propose PULSE, a framework for more efficient human-centric evaluation of agent designs, which comprises collecting user feedback, training an ML model to predict user satisfaction, and computing results by combining human satisfaction ratings with model-generated pseudo-labels. Second, we deploy the framework on a large-scale web platform built around the open-source software agent OpenHands, collecting in-the-wild usage data across over 15k users. We conduct case studies around how three agent design decisions -- choice of LLM backbone, planning strategy, and memory mechanisms -- impact developer satisfaction rates, yielding practical insights for software agent design. We also show how our framework can lead to more robust conclusions about agent design, reducing confidence intervals by 40% compared to a standard A/B test. Finally, we find substantial discrepancies between in-the-wild results and benchmark performance (e.g., the anti-correlation between results comparing claude-sonnet-4 and gpt-5), underscoring the limitations of benchmark-driven evaluation. Our findings provide guidance for evaluations of LLM agents with humans and identify opportunities for better agent designs.
  •  

AutoPDL: Automatic Prompt Optimization for LLM Agents

arXiv:2504.04365v5 Announce Type: replace-cross Abstract: The performance of large language models (LLMs) depends on how they are prompted, with choices spanning both the high-level prompting pattern (e.g., Zero-Shot, CoT, ReAct, ReWOO) and the specific prompt content (instructions and few-shot demonstrations). Manually tuning this combination is tedious, error-prone, and specific to a given LLM and task. Therefore, this paper proposes AutoPDL, an automated approach to discovering good LLM agent configurations. Our approach frames this as a structured AutoML problem over a combinatorial space of agentic and non-agentic prompting patterns and demonstrations, using successive halving to efficiently navigate this space. We introduce a library implementing common prompting patterns using the PDL prompt programming language. AutoPDL solutions are human-readable, editable, and executable PDL programs that use this library. This approach also enables source-to-source optimization, allowing human-in-the-loop refinement and reuse. Evaluations across three tasks and seven LLMs (ranging from 3B to 70B parameters) show consistent accuracy gains ($9.21\pm15.46$ percentage points), up to 67.5pp, and reveal that selected prompting strategies vary across models and tasks.
  •  

Diffusion Models at the Drug Discovery Frontier: A Review on Generating Small Molecules versus Therapeutic Peptides

arXiv:2511.00209v1 Announce Type: cross Abstract: Diffusion models have emerged as a leading framework in generative modeling, showing significant potential to accelerate and transform the traditionally slow and costly process of drug discovery. This review provides a systematic comparison of their application in designing two principal therapeutic modalities: small molecules and therapeutic peptides. We analyze how a unified framework of iterative denoising is adapted to the distinct molecular representations, chemical spaces, and design objectives of each modality. For small molecules, these models excel at structure-based design, generating novel, pocket-fitting ligands with desired physicochemical properties, yet face the critical hurdle of ensuring chemical synthesizability. Conversely, for therapeutic peptides, the focus shifts to generating functional sequences and designing de novo structures, where the primary challenges are achieving biological stability against proteolysis, ensuring proper folding, and minimizing immunogenicity. Despite these distinct challenges, both domains face shared hurdles: the need for more accurate scoring functions, the scarcity of high-quality experimental data, and the crucial requirement for experimental validation. We conclude that the full potential of diffusion models will be unlocked by bridging these modality-specific gaps and integrating them into automated, closed-loop Design-Build-Test-Learn (DBTL) platforms, thereby shifting the paradigm from chemical exploration to the targeted creation of novel therapeutics.
  •  

A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation

arXiv:2510.19755v3 Announce Type: replace-cross Abstract: Diffusion Models have become a cornerstone of modern generative AI for their exceptional generation quality and controllability. However, their inherent \textit{multi-step iterations} and \textit{complex backbone networks} lead to prohibitive computational overhead and generation latency, forming a major bottleneck for real-time applications. Although existing acceleration techniques have made progress, they still face challenges such as limited applicability, high training costs, or quality degradation. Against this backdrop, \textbf{Diffusion Caching} offers a promising training-free, architecture-agnostic, and efficient inference paradigm. Its core mechanism identifies and reuses intrinsic computational redundancies in the diffusion process. By enabling feature-level cross-step reuse and inter-layer scheduling, it reduces computation without modifying model parameters. This paper systematically reviews the theoretical foundations and evolution of Diffusion Caching and proposes a unified framework for its classification and analysis. Through comparative analysis of representative methods, we show that Diffusion Caching evolves from \textit{static reuse} to \textit{dynamic prediction}. This trend enhances caching flexibility across diverse tasks and enables integration with other acceleration techniques such as sampling optimization and model distillation, paving the way for a unified, efficient inference framework for future multimodal and interactive applications. We argue that this paradigm will become a key enabler of real-time and efficient generative AI, injecting new vitality into both theory and practice of \textit{Efficient Generative Intelligence}.
  •  
❌