❌

Normal view

Anthropic buys biotech startup Coefficient Bio in $400M deal: Reports

4 April 2026 at 04:28
Anthropic has purchased the stealth biotech AI startup Coefficient Bio in a $400 million stock deal, according to The Information and Eric Newcomer.

Integrating liquid biopsies and artificial intelligence for early cancer detection: A systematic review and meta-analysis

Eur J Cancer. 2026 Mar 24;239:116699. doi: 10.1016/j.ejca.2026.116699. Online ahead of print.

ABSTRACT

INTRODUCTION: The latest generation of liquid biopsies incorporates multi-omic features, including genomics, methylomics, and fragmentomics. Machine learning (ML) approaches have been proposed to synthesize these complex biological data for the development of diagnostic classifiers. This study aims to evaluate the integration of ML with circulating cell-free DNA (cfDNA) analysis for early cancer detection.

METHODS: Medline, Embase, Cochrane, and Web of Science were searched in July 2025. Eligible studies combined ML and cfDNA features to distinguish cancer patients (stages I-III) from non-cancer controls. Summary diagnostic performance metrics and their 95% confidence intervals (CI) were calculated.

RESULTS: The study included 109 articles permitting analyses for lung (n = 34), liver (n = 29), colorectal (n = 28), pancreatic (n = 16), breast (n = 17), esophageal (n = 12), ovarian (n = 13), gastric (n = 9), head and neck (n = 4), and mixed (n = 27) cancer types. Specificity was consistently high across all tumor types and stages (94%-99%). Sensitivity ranged from 72% to 92% for stage I-III, 44-91% for stage I, 71-98% for stage II and 83-99% for stage III. In the pooled study population, neural networks (90%, 95% CI: 81%-95%), random forest (86%, 95% CI: 77%-92%) and heterogeneous ensemble learning (85%, 95% CI: 79%-89%) demonstrated the highest sensitivity. The stratified analysis by classifier feature revealed 86% (95% CI: 80%-90%) sensitivity for fragmentation and 81% (95% CI: 76%-85%) for methylation, with 92%-96% specificity.

CONCLUSION: ML and cfDNA profiling show potential for early cancer detection, with ensemble methods, neural networks and random forests achieving the best overall performance. Fragmentomic features provide the highest sensitivity.

PMID:41930854 | DOI:10.1016/j.ejca.2026.116699

Integrating liquid biopsies and artificial intelligence for early cancer detection: A systematic review and meta-analysis

Eur J Cancer. 2026 Mar 24;239:116699. doi: 10.1016/j.ejca.2026.116699. Online ahead of print.

ABSTRACT

INTRODUCTION: The latest generation of liquid biopsies incorporates multi-omic features, including genomics, methylomics, and fragmentomics. Machine learning (ML) approaches have been proposed to synthesize these complex biological data for the development of diagnostic classifiers. This study aims to evaluate the integration of ML with circulating cell-free DNA (cfDNA) analysis for early cancer detection.

METHODS: Medline, Embase, Cochrane, and Web of Science were searched in July 2025. Eligible studies combined ML and cfDNA features to distinguish cancer patients (stages I-III) from non-cancer controls. Summary diagnostic performance metrics and their 95% confidence intervals (CI) were calculated.

RESULTS: The study included 109 articles permitting analyses for lung (n = 34), liver (n = 29), colorectal (n = 28), pancreatic (n = 16), breast (n = 17), esophageal (n = 12), ovarian (n = 13), gastric (n = 9), head and neck (n = 4), and mixed (n = 27) cancer types. Specificity was consistently high across all tumor types and stages (94%-99%). Sensitivity ranged from 72% to 92% for stage I-III, 44-91% for stage I, 71-98% for stage II and 83-99% for stage III. In the pooled study population, neural networks (90%, 95% CI: 81%-95%), random forest (86%, 95% CI: 77%-92%) and heterogeneous ensemble learning (85%, 95% CI: 79%-89%) demonstrated the highest sensitivity. The stratified analysis by classifier feature revealed 86% (95% CI: 80%-90%) sensitivity for fragmentation and 81% (95% CI: 76%-85%) for methylation, with 92%-96% specificity.

CONCLUSION: ML and cfDNA profiling show potential for early cancer detection, with ensemble methods, neural networks and random forests achieving the best overall performance. Fragmentomic features provide the highest sensitivity.

PMID:41930854 | DOI:10.1016/j.ejca.2026.116699

Immune endotypes in tuberculosis: Keys to decoding disease complexity

J Intern Med. 2026 Apr 3. doi: 10.1111/joim.70092. Online ahead of print.

ABSTRACT

Tuberculosis (TB) remains a major global health challenge, with multi-drug antibiotic regimens as the current standard of care. While effective at killing Mycobacterium tuberculosis, these treatments do not resolve persistent inflammation, prevent lung damage, or reverse immune dysregulation that contribute to poor outcomes and disease recurrence. Precision medicine offers a promising alternative but requires deeper insight into disease mechanisms to enable tailored interventions. This comprehensive review introduces the concept of immune endotyping to define the underlying disease mechanisms as tools to decode clinical and immunological heterogeneity in TB. TB displays a wide spectrum of clinical phenotypes, from latent or asymptomatic infection to mild or severe disease with characteristic non-cavitary or cavitary lung pathology. Instead, distinct immune endotypes capture the diverse biological pathways that shape disease progression and treatment response. Similar clinical presentations may arise from different immune dysfunctions, underscoring the need to move beyond broad phenotypic classifications. Advances in multi-omics and computational analyses uncover immune signatures that enable stratification for host-directed therapies (HDTs) targeting hyperinflammation, immunosuppression, coagulopathy or metabolic exhaustion. Integrating clinical, radiological, and immunological data through multimodal profiling is essential for developing personalized interventions. We also explore how endotyping has transformed treatment in other diseases, offering valuable insights for TB. Additionally, we present examples of how putative immune endotypes may be targeted with appropriate HDTs. In summary, this review underscores the potential of immune endotypes to advance precision medicine in TB, moving beyond one-size-fits-all treatment to improve outcomes, especially in severe and drug-resistant cases.

PMID:41930636 | DOI:10.1111/joim.70092

Integrating liquid biopsies and artificial intelligence for early cancer detection: A systematic review and meta-analysis

Eur J Cancer. 2026 Mar 24;239:116699. doi: 10.1016/j.ejca.2026.116699. Online ahead of print.

ABSTRACT

INTRODUCTION: The latest generation of liquid biopsies incorporates multi-omic features, including genomics, methylomics, and fragmentomics. Machine learning (ML) approaches have been proposed to synthesize these complex biological data for the development of diagnostic classifiers. This study aims to evaluate the integration of ML with circulating cell-free DNA (cfDNA) analysis for early cancer detection.

METHODS: Medline, Embase, Cochrane, and Web of Science were searched in July 2025. Eligible studies combined ML and cfDNA features to distinguish cancer patients (stages I-III) from non-cancer controls. Summary diagnostic performance metrics and their 95% confidence intervals (CI) were calculated.

RESULTS: The study included 109 articles permitting analyses for lung (n = 34), liver (n = 29), colorectal (n = 28), pancreatic (n = 16), breast (n = 17), esophageal (n = 12), ovarian (n = 13), gastric (n = 9), head and neck (n = 4), and mixed (n = 27) cancer types. Specificity was consistently high across all tumor types and stages (94%-99%). Sensitivity ranged from 72% to 92% for stage I-III, 44-91% for stage I, 71-98% for stage II and 83-99% for stage III. In the pooled study population, neural networks (90%, 95% CI: 81%-95%), random forest (86%, 95% CI: 77%-92%) and heterogeneous ensemble learning (85%, 95% CI: 79%-89%) demonstrated the highest sensitivity. The stratified analysis by classifier feature revealed 86% (95% CI: 80%-90%) sensitivity for fragmentation and 81% (95% CI: 76%-85%) for methylation, with 92%-96% specificity.

CONCLUSION: ML and cfDNA profiling show potential for early cancer detection, with ensemble methods, neural networks and random forests achieving the best overall performance. Fragmentomic features provide the highest sensitivity.

PMID:41930854 | DOI:10.1016/j.ejca.2026.116699

Integrating liquid biopsies and artificial intelligence for early cancer detection: A systematic review and meta-analysis

Eur J Cancer. 2026 Mar 24;239:116699. doi: 10.1016/j.ejca.2026.116699. Online ahead of print.

ABSTRACT

INTRODUCTION: The latest generation of liquid biopsies incorporates multi-omic features, including genomics, methylomics, and fragmentomics. Machine learning (ML) approaches have been proposed to synthesize these complex biological data for the development of diagnostic classifiers. This study aims to evaluate the integration of ML with circulating cell-free DNA (cfDNA) analysis for early cancer detection.

METHODS: Medline, Embase, Cochrane, and Web of Science were searched in July 2025. Eligible studies combined ML and cfDNA features to distinguish cancer patients (stages I-III) from non-cancer controls. Summary diagnostic performance metrics and their 95% confidence intervals (CI) were calculated.

RESULTS: The study included 109 articles permitting analyses for lung (n = 34), liver (n = 29), colorectal (n = 28), pancreatic (n = 16), breast (n = 17), esophageal (n = 12), ovarian (n = 13), gastric (n = 9), head and neck (n = 4), and mixed (n = 27) cancer types. Specificity was consistently high across all tumor types and stages (94%-99%). Sensitivity ranged from 72% to 92% for stage I-III, 44-91% for stage I, 71-98% for stage II and 83-99% for stage III. In the pooled study population, neural networks (90%, 95% CI: 81%-95%), random forest (86%, 95% CI: 77%-92%) and heterogeneous ensemble learning (85%, 95% CI: 79%-89%) demonstrated the highest sensitivity. The stratified analysis by classifier feature revealed 86% (95% CI: 80%-90%) sensitivity for fragmentation and 81% (95% CI: 76%-85%) for methylation, with 92%-96% specificity.

CONCLUSION: ML and cfDNA profiling show potential for early cancer detection, with ensemble methods, neural networks and random forests achieving the best overall performance. Fragmentomic features provide the highest sensitivity.

PMID:41930854 | DOI:10.1016/j.ejca.2026.116699

Immune endotypes in tuberculosis: Keys to decoding disease complexity

J Intern Med. 2026 Apr 3. doi: 10.1111/joim.70092. Online ahead of print.

ABSTRACT

Tuberculosis (TB) remains a major global health challenge, with multi-drug antibiotic regimens as the current standard of care. While effective at killing Mycobacterium tuberculosis, these treatments do not resolve persistent inflammation, prevent lung damage, or reverse immune dysregulation that contribute to poor outcomes and disease recurrence. Precision medicine offers a promising alternative but requires deeper insight into disease mechanisms to enable tailored interventions. This comprehensive review introduces the concept of immune endotyping to define the underlying disease mechanisms as tools to decode clinical and immunological heterogeneity in TB. TB displays a wide spectrum of clinical phenotypes, from latent or asymptomatic infection to mild or severe disease with characteristic non-cavitary or cavitary lung pathology. Instead, distinct immune endotypes capture the diverse biological pathways that shape disease progression and treatment response. Similar clinical presentations may arise from different immune dysfunctions, underscoring the need to move beyond broad phenotypic classifications. Advances in multi-omics and computational analyses uncover immune signatures that enable stratification for host-directed therapies (HDTs) targeting hyperinflammation, immunosuppression, coagulopathy or metabolic exhaustion. Integrating clinical, radiological, and immunological data through multimodal profiling is essential for developing personalized interventions. We also explore how endotyping has transformed treatment in other diseases, offering valuable insights for TB. Additionally, we present examples of how putative immune endotypes may be targeted with appropriate HDTs. In summary, this review underscores the potential of immune endotypes to advance precision medicine in TB, moving beyond one-size-fits-all treatment to improve outcomes, especially in severe and drug-resistant cases.

PMID:41930636 | DOI:10.1111/joim.70092

The Digital Twin Counterfactual Framework: A Validation Architecture for Simulated Potential Outcomes

arXiv:2604.01325v1 Announce Type: new Abstract: The fundamental problem of causal inference - that the counterfactual outcome for any individual is never observed - has shaped the entire methodology of the field. Every existing approach substitutes assumptions for missing data: ignorability, parallel trends, exclusion restrictions. None produces the counterfactual itself. This paper proposes the Digital Twin Counterfactual Framework (DTCF): rather than estimating the counterfactual statistically, we simulate it using a digital twin and subject the simulation to a hierarchical validation regime. We formalize the digital twin simulator as a stochastic mapping within the potential outcomes framework and introduce a hierarchy of twin fidelity assumptions - from marginal fidelity through joint fidelity to structural fidelity - each unlocking a progressively richer class of estimands. The central contribution is threefold. First, a five-level validation architecture converts the unfalsifiable claim that the simulator produces correct counterfactuals into falsifiable tests against observable data. Second, a formal decomposition separates causal quantities into those that are marginally validated (ATE, CATE, QTE - testable through observable-arm comparison) and those that are copula-dependent (the ITE distribution, probability of benefit/harm, variance of treatment effects - permanently reliant on the unobservable within-individual dependence structure). Third, bounding, sensitivity, and uncertainty quantification tools make the copula dependence explicit. The DTCF does not resolve the fundamental problem of causal inference. What it provides is a framework in which marginal causal claims become increasingly testable, joint causal claims become explicitly assumption-indexed, and the gap between the two is formally characterized.
  • ✇cs.AI, q-bio.NC updates on arXiv.org
  • Semantic Modeling for World-Centered Architectures Andrei Mantsivoda · Darya Gavrilina
    arXiv:2604.01359v1 Announce Type: new Abstract: We introduce world-centered multi-agent systems (WMAS) as an alternative to traditional agent-centered architectures, arguing that structured domains such as enterprises and institutional systems require a shared, explicit world representation to ensure semantic consistency, explainability, and long-term stability. We classify worlds along dimensions including ontological explicitness, normativity, etc. In WMAS, learning and coordination operate o
     

Semantic Modeling for World-Centered Architectures

arXiv:2604.01359v1 Announce Type: new Abstract: We introduce world-centered multi-agent systems (WMAS) as an alternative to traditional agent-centered architectures, arguing that structured domains such as enterprises and institutional systems require a shared, explicit world representation to ensure semantic consistency, explainability, and long-term stability. We classify worlds along dimensions including ontological explicitness, normativity, etc. In WMAS, learning and coordination operate over a shared world model rather than isolated agent-local representations, enabling global consistency and verifiable system behavior. We propose semantic models as a mathematical formalism for representing such worlds. Finally, we present the Ontobox platform as a realization of WMAS.

A Role-Based LLM Framework for Structured Information Extraction from Healthy Food Policies

arXiv:2604.01529v1 Announce Type: new Abstract: Current Large Language Model (LLM) approaches for information extraction (IE) in the healthy food policy domain are often hindered by various factors, including misinformation, specifically hallucinations, misclassifications, and omissions that result from the structural diversity and inconsistency of policy documents. To address these limitations, this study proposes a role-based LLM framework that automates the IE from unstructured policy data by assigning specialized roles: an LLM policy analyst for metadata and mechanism classification, an LLM legal strategy specialist for identifying complex legal approaches, and an LLM food system expert for categorizing food system stages. This framework mimics expert analysis workflows by incorporating structured domain knowledge, including explicit definitions of legal mechanisms and classification criteria, into role-specific prompts. We evaluate the framework using 608 healthy food policies from the Healthy Food Policy Project (HFPP) database, comparing its performance against zero-shot, few-shot, and chain-of-thought (CoT) baselines using Llama-3.3-70B. Our proposed framework demonstrates superior performance in complex reasoning tasks, offering a reliable and transparent methodology for automating IE from health policies.

PHMForge: A Scenario-Driven Agentic Benchmark for Industrial Asset Lifecycle Maintenance

arXiv:2604.01532v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed for complex tool-orchestration tasks, yet existing benchmarks fail to capture the rigorous demands of industrial domains where incorrect decisions carry significant safety and financial consequences. To address this critical gap, we introduce PHMForge, the first comprehensive benchmark specifically designed to evaluate LLM agents on Prognostics and Health Management (PHM) tasks through realistic interactions with domain-specific MCP servers. Our benchmark encompasses 75 expert-curated scenarios spanning 7 industrial asset classes (turbofan engines, bearings, electric motors, gearboxes, aero-engines) across 5 core task categories: Remaining Useful Life (RUL) Prediction, Fault Classification, Engine Health Analysis, Cost-Benefit Analysis, and Safety/Policy Evaluation. To enable rigorous evaluation, we construct 65 specialized tools across two MCP servers and implement execution-based evaluators with task-commensurate metrics: MAE/RMSE for regression, F1-score for classification, and categorical matching for health assessments. Through extensive evaluation of leading frameworks (ReAct, Cursor Agent, Claude Code) paired with frontier LLMs (Claude Sonnet 4.0, GPT-4o, Granite-3.0-8B), we find that even top-performing configurations achieve only 68\% task completion, with systematic failures in tool orchestration (23\% incorrect sequencing), multi-asset reasoning (14.9 percentage point degradation), and cross-equipment generalization (42.7\% on held-out datasets). We open-source our complete benchmark, including scenario specifications, ground truth templates, tool implementations, and evaluation scripts, to catalyze research in agentic industrial AI.

MM-ReCoder: Advancing Chart-to-Code Generation with Reinforcement Learning and Self-Correction

arXiv:2604.01600v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have recently demonstrated promising capabilities in multimodal coding tasks such as chart-to-code generation. However, existing methods primarily rely on supervised fine-tuning (SFT), which requires the model to learn code patterns through chart-code pairs but does not expose the model to a code execution environment. Moreover, while self-correction through execution feedback offers a potential route to improve coding quality, even state-of-the-art MLLMs have been shown to struggle with effective self-correction. In this work, we introduce MM-ReCoder, a chart-to-code generation model trained with reinforcement learning (RL) and equipped with self-correction ability. We propose a two-stage multi-turn self-correction RL strategy based on Group Relative Policy Optimization (GRPO). The first stage enhances the model's self-correction ability via rolling out a shared first turn, while the second stage improves the coding capability with full-trajectory optimization. MM-ReCoder learns to produce more accurate and executable code through the interaction with the environment and by iteratively correcting its own outputs. Our results on three chart-to-code benchmarks demonstrate the state-of-the-art performance of MM-ReCoder.

Exploring Robust Multi-Agent Workflows for Environmental Data Management

arXiv:2604.01647v1 Announce Type: new Abstract: Embedding LLM-driven agents into environmental FAIR data management is compelling - they can externalize operational knowledge and scale curation across heterogeneous data and evolving conventions. However, replacing deterministic components with probabilistic workflows changes the failure mode: LLM pipelines may generate plausible but incorrect outputs that pass superficial checks and propagate into irreversible actions such as DOI minting and public release. We introduce EnviSmart, a production data management system deployed on campus-wide storage infrastructure for environmental research. EnviSmart treats reliability as an architectural property through two mechanisms: a three-track knowledge architecture that externalizes behaviors (governance constraints), domain knowledge (retrievable context), and skills (tool-using procedures) as persistent, interlocking artifacts; and a role-separated multi-agent design where deterministic validators and audited handoffs restore fail-stop semantics at trust boundaries before irreversible steps. We compare two production deployments. The University's GIS Center Ecological Archive (849 curated datasets) serves as a single-agent baseline. SF2Bench, a compound flooding benchmark comprising 2,452 monitoring stations and 8,557 published files spanning 39 years, validates the multi-agent workflow. The multi-agent approach improved both efficiency - completed by a single operator in two days with repeated artifact reuse across deployments - and reliability: audited handoffs detected and blocked a coordinate transformation error affecting all 2,452 stations before publication. A representative incident (ISS-004) demonstrated boundary-based containment with 10-minute detection latency, zero user exposure, and 80-minute resolution. This paper has been accepted at PEARC 2026.

GenGait: A Transformer-Based Model for Human Gait Anomaly Detection and Normative Twin Generation

arXiv:2604.01997v1 Announce Type: new Abstract: Gait analysis provides an objective characterization of locomotor function and is widely used to support diagnosis and rehabilitation monitoring across neurological and orthopedic disorders. Deep learning has been increasingly applied to this domain, yet most approaches rely on supervised classifiers trained on disease-labeled data, limiting generalization to heterogeneous pathological presentations. This work proposes a label-free framework for joint-level anomaly detection and kinematic correction based on a Transformer masked autoencoder trained exclusively on normative gait sequences from 150 adults, acquired with a markerless multi-camera motion-capture system. At inference, a two-pass procedure is applied to potentially pathological input sequences, first it estimates joint inconsistency scores by occluding individual joints and measuring deviations from the learned normative prior. Then, it withholds the flagged joints from the encoder input and reconstructs the full skeleton from the remaining spatiotemporal context, yielding corrected kinematic trajectories at the flagged positions. Validation on 10 held-out normative participants, who mimicked seven simulated gait abnormalities, showed accurate localization of biomechanically inconsistent joints, a significant reduction in angular deviation across all analyzed joints with large effect sizes, and preservation of normative kinematics. The proposed approach enables interpretable, subject-specific localization of gait impairments without requiring disease labels. Video is available at https://youtu.be/Rcm3jqR5pN4.

VISTA: Visualization of Token Attribution via Efficient Analysis

arXiv:2604.02217v1 Announce Type: new Abstract: Understanding how Large Language Models (LLMs) process information from prompts remains a significant challenge. To shed light on this "black box," attention visualization techniques have been developed to capture neuron-level perceptions and interpret how models focus on different parts of input data. However, many existing techniques are tailored to specific model architectures, particularly within the Transformer family, and often require backpropagation, resulting in nearly double the GPU memory usage and increased computational cost. A lightweight, model-agnostic approach for attention visualization remains lacking. In this paper, we introduce a model-agnostic token importance visualization technique to better understand how generative AI systems perceive and prioritize information from input text, without incurring additional computational cost. Our method leverages perturbation-based strategies combined with a three-matrix analytical framework to generate relevance maps that illustrate token-level contributions to model predictions. The framework comprises: (1) the Angular Deviation Matrix, which captures shifts in semantic direction; (2) the Magnitude Deviation Matrix, which measures changes in semantic intensity; and (3) the Dimensional Importance Matrix, which evaluates contributions across individual vector dimensions. By systematically removing each token and measuring the resulting impact across these three complementary dimensions, we derive a composite importance score that provides a nuanced and mathematically grounded measure of token significance. To support reproducibility and foster wider adoption, we provide open-source implementations of all proposed and utilized explainability techniques, with code and resources publicly available at https://github.com/Infosys/Infosys-Responsible-AI-Toolkit

When to ASK: Uncertainty-Gated Language Assistance for Reinforcement Learning

arXiv:2604.02226v1 Announce Type: new Abstract: Reinforcement learning (RL) agents often struggle with out-of-distribution (OOD) scenarios, leading to high uncertainty and random behavior. While language models (LMs) contain valuable world knowledge, larger ones incur high computational costs, hindering real-time use, and exhibit limitations in autonomous planning. We introduce Adaptive Safety through Knowledge (ASK), which combines smaller LMs with trained RL policies to enhance OOD generalization without retraining. ASK employs Monte Carlo Dropout to assess uncertainty and queries the LM for action suggestions only when uncertainty exceeds a set threshold. This selective use preserves the efficiency of existing policies while leveraging the language model's reasoning in uncertain situations. In experiments on the FrozenLake environment, ASK shows no improvement in-domain, but demonstrates robust navigation in transfer tasks, achieving a reward of 0.95. Our findings indicate that effective neuro-symbolic integration requires careful orchestration rather than simple combination, highlighting the need for sufficient model scale and effective hybridization mechanisms for successful OOD generalization.

De Jure: Iterative LLM Self-Refinement for Structured Extraction of Regulatory Rules

arXiv:2604.02276v1 Announce Type: new Abstract: Regulatory documents encode legally binding obligations that LLM-based systems must respect. Yet converting dense, hierarchically structured legal text into machine-readable rules remains a costly, expert-intensive process. We present De Jure, a fully automated, domain-agnostic pipeline for extracting structured regulatory rules from raw documents, requiring no human annotation, domain-specific prompting, or annotated gold data. De Jure operates through four sequential stages: normalization of source documents into structured Markdown; LLM-driven semantic decomposition into structured rule units; multi-criteria LLM-as-a-judge evaluation across 19 dimensions spanning metadata, definitions, and rule semantics; and iterative repair of low-scoring extractions within a bounded regeneration budget, where upstream components are repaired before rule units are evaluated. We evaluate De Jure across four models on three regulatory corpora spanning finance, healthcare, and AI governance. On the finance domain, De Jure yields consistent and monotonic improvement in extraction quality, reaching peak performance within three judge-guided iterations. De Jure generalizes effectively to healthcare and AI governance, maintaining high performance across both open- and closed-source models. In a downstream compliance question-answering evaluation via RAG, responses grounded in De Jure extracted rules are preferred over prior work in 73.8% of cases at single-rule retrieval depth, rising to 84.0% under broader retrieval, confirming that extraction fidelity translates directly into downstream utility. These results demonstrate that explicit, interpretable evaluation criteria can substitute for human annotation in complex regulatory domains, offering a scalable and auditable path toward regulation-grounded LLM alignment.

Low-Burden LLM-Based Preference Learning: Personalizing Assistive Robots from Natural Language Feedback for Users with Paralysis

arXiv:2604.01463v1 Announce Type: cross Abstract: Physically Assistive Robots (PARs) require personalized behaviors to ensure user safety and comfort. However, traditional preference learning methods, like exhaustive pairwise comparisons, cause severe physical and cognitive fatigue for users with profound motor impairments. To solve this, we propose a low-burden, offline framework that translates unstructured natural language feedback directly into deterministic robotic control policies. To safely bridge the gap between ambiguous human speech and robotic code, our pipeline uses Large Language Models (LLMs) grounded in the Occupational Therapy Practice Framework (OTPF). This clinical reasoning decodes subjective user reactions into explicit physical and psychological needs, which are then mapped into transparent decision trees. Before deployment, an automated "LLM-as-a-Judge" verifies the code's structural safety. We validated this system in a simulated meal preparation study with 10 adults with paralysis. Results show our natural language approach significantly reduces user workload compared to traditional baselines. Additionally, independent clinical experts confirmed the generated policies are safe and accurately reflect user preferences.

ProdCodeBench: A Production-Derived Benchmark for Evaluating AI Coding Agents

arXiv:2604.01527v1 Announce Type: cross Abstract: Benchmarks that reflect production workloads are better for evaluating AI coding agents in industrial settings, yet existing benchmarks differ from real usage in programming language distribution, prompt style and codebase structure. This paper presents a methodology for curating production-derived benchmarks, illustrated through ProdCodeBench - a benchmark built from real sessions with a production AI coding assistant. We detail our data collection and curation practices including LLM-based task classification, test relevance validation, and multi-run stability checks which address challenges in constructing reliable evaluation signals from monorepo environments. Each curated sample consists of a verbatim prompt, a committed code change and fail-to-pass tests spanning seven programming languages. Our systematic analysis of four foundation models yields solve rates from 53.2% to 72.2% revealing that models making greater use of work validation tools, such as executing tests and invoking static analysis, achieve higher solve rates. This suggests that iterative verification helps achieve effective agent behavior and that exposing codebase-specific verification mechanisms may significantly improve the performance of externally trained agents operating in unfamiliar environments. We share our methodology and lessons learned to enable other organizations to construct similar production-derived benchmarks.

RefinementEngine: Automating Intent-to-Device Filtering Policy Deployment under Network Constraints

arXiv:2604.01627v1 Announce Type: cross Abstract: Translating security intent into deployable network enforcement rules and maintaining their effectiveness despite evolving cyber threats remains a largely manual process in most Security Operations Centers (SOCs). In large and heterogeneous networks, this challenge is complicated by topology-dependent reachability constraints and device-specific security control capabilities, making the process slow, error-prone, and a recurring source of misconfigurations. This paper presents RefinementEngine, an engine that automates the refinement of high-level security intents into low-level, deployment-ready configurations. Given a network topology, devices, and available security controls, along with high-level intents and Cyber Threat Intelligence (CTI) reports, RefinementEngine automatically generates settings that implement the desired intent, counter reported threats, and can be directly deployed on target security controls. The proposed approach is validated through real-world use cases on packet and web filtering policies derived from actual CTI reports, demonstrating both correctness, practical applicability, and adaptability to new data.
❌