❌

Normal view

A Decision-Theoretic Approach for Managing Misalignment

arXiv:2512.15584v1 Announce Type: new Abstract: When should we delegate decisions to AI systems? While the value alignment literature has developed techniques for shaping AI values, less attention has been paid to how to determine, under uncertainty, when imperfect alignment is good enough to justify delegation. We argue that rational delegation requires balancing an agent's value (mis)alignment with its epistemic accuracy and its reach (the acts it has available). This paper introduces a formal, decision-theoretic framework to analyze this tradeoff precisely accounting for a principal's uncertainty about these factors. Our analysis reveals a sharp distinction between two delegation scenarios. First, universal delegation (trusting an agent with any problem) demands near-perfect value alignment and total epistemic trust, conditions rarely met in practice. Second, we show that context-specific delegation can be optimal even with significant misalignment. An agent's superior accuracy or expanded reach may grant access to better overall decision problems, making delegation rational in expectation. We develop a novel scoring framework to quantify this ex ante decision. Ultimately, our work provides a principled method for determining when an AI is aligned enough for a given context, shifting the focus from achieving perfect alignment to managing the risks and rewards of delegation under uncertainty.

aiXiv: A Next-Generation Open Access Ecosystem for Scientific Discovery Generated by AI Scientists

arXiv:2508.15126v2 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have enabled AI agents to autonomously generate scientific proposals, conduct experiments, author papers, and perform peer reviews. Yet this flood of AI-generated research content collides with a fragmented and largely closed publication ecosystem. Traditional journals and conferences rely on human peer review, making them difficult to scale and often reluctant to accept AI-generated research content; existing preprint servers (e.g. arXiv) lack rigorous quality-control mechanisms. Consequently, a significant amount of high-quality AI-generated research lacks appropriate venues for dissemination, hindering its potential to advance scientific progress. To address these challenges, we introduce aiXiv, a next-generation open-access platform for human and AI scientists. Its multi-agent architecture allows research proposals and papers to be submitted, reviewed, and iteratively refined by both human and AI scientists. It also provides API and MCP interfaces that enable seamless integration of heterogeneous human and AI scientists, creating a scalable and extensible ecosystem for autonomous scientific discovery. Through extensive experiments, we demonstrate that aiXiv is a reliable and robust platform that significantly enhances the quality of AI-generated research proposals and papers after iterative revising and reviewing on aiXiv. Our work lays the groundwork for a next-generation open-access ecosystem for AI scientists, accelerating the publication and dissemination of high-quality AI-generated research content. Code: https://github.com/aixiv-org aiXiv: https://aixiv.science

Systems pharmacology approaches decipher the anti-cancer efficacy of ethnopharmacological agents in hepatocellular carcinoma

Sci Rep. 2025 Dec 17;15(1):43996. doi: 10.1038/s41598-025-27744-w.

ABSTRACT

Hepatocellular carcinoma (HCC) poses a significant global health burden with limited therapeutic efficacy. Chinese herbal medicines (CHMs) offer multi-target potential, yet their systematic screening and mechanistic elucidation remain challenging. We established a high-throughput multi-omics platform integrating transcriptomics, proteomics, and deep learning (autoencoder and multiple kernel learning) to screen 187 medicinal plants. Five CHMs candidates were identified and shown to modulate hub genes (e.g., AKR1B10, HMGCR, THBS1) and key pathways (TNF/IL-17/MAPK, apoptosis, ferroptosis). Proteomic validation and functional assays confirmed their roles in suppressing proliferation, migration, and inducing apoptosis in HCC cells. This study provides a robust, data-driven pipeline for natural anti-HCC drug discovery, linking specific hub genes to CHM efficacy and offering novel insights into precision ethnopharmacology.

PMID:41408124 | DOI:10.1038/s41598-025-27744-w

Leveraging LLMs for Structured Data Extraction from Unstructured Patient Records

arXiv:2512.13700v1 Announce Type: new Abstract: Manual chart review remains an extremely time-consuming and resource-intensive component of clinical research, requiring experts to extract often complex information from unstructured electronic health record (EHR) narratives. We present a secure, modular framework for automated structured feature extraction from clinical notes leveraging locally deployed large language models (LLMs) on institutionally approved, Health Insurance Portability and Accountability Act (HIPPA)-compliant compute infrastructure. This system integrates retrieval augmented generation (RAG) and structured response methods of LLMs into a widely deployable and scalable container to provide feature extraction for diverse clinical domains. In evaluation, the framework achieved high accuracy across multiple medical characteristics present in large bodies of patient notes when compared against an expert-annotated dataset and identified several annotation errors missed in manual review. This framework demonstrates the potential of LLM systems to reduce the burden of manual chart review through automated extraction and increase consistency in data capture, accelerating clinical research.

ValuePilot: A Two-Phase Framework for Value-Driven Decision-Making

arXiv:2512.13716v1 Announce Type: new Abstract: Personalized decision-making is essential for human-AI interaction, enabling AI agents to act in alignment with individual users' value preferences. As AI systems expand into real-world applications, adapting to personalized values beyond task completion or collective alignment has become a critical challenge. We address this by proposing a value-driven approach to personalized decision-making. Human values serve as stable, transferable signals that support consistent and generalizable behavior across contexts. Compared to task-oriented paradigms driven by external rewards and incentives, value-driven decision-making enhances interpretability and enables agents to act appropriately even in novel scenarios. We introduce ValuePilot, a two-phase framework consisting of a dataset generation toolkit (DGT) and a decision-making module (DMM). DGT constructs diverse, value-annotated scenarios from a human-LLM collaborative pipeline. DMM learns to evaluate actions based on personal value preferences, enabling context-sensitive, individualized decisions. When evaluated on previously unseen scenarios, DMM outperforms strong LLM baselines, including GPT-5, Claude-Sonnet-4, Gemini-2-flash, and Llama-3.1-70b, in aligning with human action choices. Our results demonstrate that value-driven decision-making is an effective and extensible engineering pathway toward building interpretable, personalized AI agents.

Enhancing Transparency and Traceability in Healthcare AI: The AI Product Passport

arXiv:2512.13702v1 Announce Type: cross Abstract: Objective: To develop the AI Product Passport, a standards-based framework improving transparency, traceability, and compliance in healthcare AI via lifecycle-based documentation. Materials and Methods: The AI Product Passport was developed within the AI4HF project, focusing on heart failure AI tools. We analyzed regulatory frameworks (EU AI Act, FDA guidelines) and existing standards to design a relational data model capturing metadata across AI lifecycle phases: study definition, dataset preparation, model generation/evaluation, deployment/monitoring, and passport generation. MLOps/ModelOps concepts were integrated for operational relevance. Co-creation involved feedback from AI4HF consortium and a Lisbon workshop with 21 diverse stakeholders, evaluated via Mentimeter polls. The open-source platform was implemented with Python libraries for automated provenance tracking. Results: The AI Product Passport was designed based on existing standards and methods with well-defined lifecycle management and role-based access. Its implementation is a web-based platform with a relational data model supporting auditable documentation. It generates machine- and human-readable reports, customizable for stakeholders. It aligns with FUTURE-AI principles (Fairness, Universality, Traceability, Usability, Robustness, Explainability), ensuring fairness, traceability, and usability. Exported passports detail model purpose, data provenance, performance, and deployment context. GitHub-hosted backend/frontend codebases enhance accessibility. Discussion and Conclusion: The AI Product Passport addresses transparency gaps in healthcare AI, meeting regulatory and ethical demands. Its open-source nature and alignment with standards foster trust and adaptability. Future enhancements include FAIR data principles and FHIR integration for improved interoperability, promoting responsible AI deployment.

Graph AI generates neurological hypotheses validated in molecular, organoid, and clinical systems

arXiv:2512.13724v1 Announce Type: cross Abstract: Neurological diseases are the leading global cause of disability, yet most lack disease-modifying treatments. We present PROTON, a heterogeneous graph transformer that generates testable hypotheses across molecular, organoid, and clinical systems. To evaluate PROTON, we apply it to Parkinson's disease (PD), bipolar disorder (BD), and Alzheimer's disease (AD). In PD, PROTON linked genetic risk loci to genes essential for dopaminergic neuron survival and predicted pesticides toxic to patient-derived neurons, including the insecticide endosulfan, which ranked within the top 1.29% of predictions. In silico screens performed by PROTON reproduced six genome-wide $\alpha$-synuclein experiments, including a split-ubiquitin yeast two-hybrid system (normalized enrichment score [NES] = 2.30, FDR-adjusted $p

Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong Baseline

arXiv:2512.13731v1 Announce Type: cross Abstract: Mathematical Expression Recognition (MER) has made significant progress in recognizing simple expressions, but the robust recognition of complex mathematical expressions with many tokens and multiple lines remains a formidable challenge. In this paper, we first introduce CMER-Bench, a carefully constructed benchmark that categorizes expressions into three difficulty levels: easy, moderate, and complex. Leveraging CMER-Bench, we conduct a comprehensive evaluation of existing MER models and general-purpose multimodal large language models (MLLMs). The results reveal that while current methods perform well on easy and moderate expressions, their performance degrades significantly when handling complex mathematical expressions, mainly because existing public training datasets are primarily composed of simple samples. In response, we propose MER-17M and CMER-3M that are large-scale datasets emphasizing the recognition of complex mathematical expressions. The datasets provide rich and diverse samples to support the development of accurate and robust complex MER models. Furthermore, to address the challenges posed by the complicated spatial layout of complex expressions, we introduce a novel expression tokenizer, and a new representation called Structured Mathematical Language, which explicitly models the hierarchical and spatial structure of expressions beyond LaTeX format. Based on these, we propose a specialized model named CMERNet, built upon an encoder-decoder architecture and trained on CMER-3M. Experimental results show that CMERNet, with only 125 million parameters, significantly outperforms existing MER models and MLLMs on CMER-Bench.

TF-MCL: Time-frequency Fusion and Multi-domain Cross-Loss for Self-supervised Depression Detection

arXiv:2512.13736v1 Announce Type: cross Abstract: In recent years, there has been a notable increase in the use of supervised detection methods of major depressive disorder (MDD) based on electroencephalogram (EEG) signals. However, the process of labeling MDD remains challenging. As a self-supervised learning method, contrastive learning could address the shortcomings of supervised learning methods, which are unduly reliant on labels in the context of MDD detection. However, existing contrastive learning methods are not specifically designed to characterize the time-frequency distribution of EEG signals, and their capacity to acquire low-semantic data representations is still inadequate for MDD detection tasks. To address the problem of contrastive learning method, we propose a time-frequency fusion and multi-domain cross-loss (TF-MCL) model for MDD detection. TF-MCL generates time-frequency hybrid representations through the use of a fusion mapping head (FMH), which efficiently remaps time-frequency domain information to the fusion domain, and thus can effectively enhance the model's capacity to synthesize time-frequency information. Moreover, by optimizing the multi-domain cross-loss function, the distribution of the representations in the time-frequency domain and the fusion domain is reconstructed, thereby improving the model's capacity to acquire fusion representations. We evaluated the performance of our model on the publicly available datasets MODMA and PRED+CT and show a significant improvement in accuracy, outperforming the existing state-of-the-art (SOTA) method by 5.87% and 9.96%, respectively.

Assessing High-Risk Systems: An EU AI Act Verification Framework

arXiv:2512.13907v1 Announce Type: cross Abstract: A central challenge in implementing the AI Act and other AI-relevant regulations in the EU is the lack of a systematic approach to verify their legal mandates. Recent surveys show that this regulatory ambiguity is perceived as a significant burden, leading to inconsistent readiness across Member States. This paper proposes a comprehensive framework designed to help close this gap by organising compliance verification along two fundamental dimensions: the type of method (controls vs. testing) and the target of assessment (data, model, processes, and final product). Additionally, our framework maps core legal requirements to concrete verification activities, serving as a vital bridge between policymakers and practitioners, and aligning legal text with technical standards and best practices. The proposed approach aims to reduce interpretive uncertainty, promote consistency in assessment practices, and support the alignment of regulatory, ethical, and technical perspectives across the AI lifecycle.

A data-physics hybrid generative model for patient-specific post-stroke motor rehabilitation using wearable sensor data

arXiv:2512.14329v1 Announce Type: cross Abstract: Dynamic prediction of locomotor capacity after stroke is crucial for tailoring rehabilitation, yet current assessments provide only static impairment scores and do not indicate whether patients can safely perform specific tasks such as slope walking or stair climbing. Here, we develop a data-physics hybrid generative framework that reconstructs an individual stroke survivor's neuromuscular control from a single 20 m level-ground walking trial and predicts task-conditioned locomotion across rehabilitation scenarios. The system combines wearable-sensor kinematics, a proportional-derivative physics controller, a population Healthy Motion Atlas, and goal-conditioned deep reinforcement learning with behaviour cloning and generative adversarial imitation learning to generate physically plausible, patient-specific gait simulations for slopes and stairs. In 11 stroke survivors, the personalized controllers preserved idiosyncratic gait patterns while improving joint-angle and endpoint fidelity by 4.73% and 12.10%, respectively, and reducing training time to 25.56% relative to a physics-only baseline. In a multicentre pilot involving 21 inpatients, clinicians who used our locomotion predictions to guide task selection and difficulty obtained larger gains in Fugl-Meyer lower-extremity scores over 28 days of standard rehabilitation than control clinicians (mean change 6.0 versus 3.7 points). These findings indicate that our generative, task-predictive framework can augment clinical decision-making in post-stroke gait rehabilitation and provide a template for dynamically personalized motor recovery strategies.

A Multicenter Benchmark of Multiple Instance Learning Models for Lymphoma Subtyping from HE-stained Whole Slide Images

arXiv:2512.14640v1 Announce Type: cross Abstract: Timely and accurate lymphoma diagnosis is essential for guiding cancer treatment. Standard diagnostic practice combines hematoxylin and eosin (HE)-stained whole slide images with immunohistochemistry, flow cytometry, and molecular genetic tests to determine lymphoma subtypes, a process requiring costly equipment, skilled personnel, and causing treatment delays. Deep learning methods could assist pathologists by extracting diagnostic information from routinely available HE-stained slides, yet comprehensive benchmarks for lymphoma subtyping on multicenter data are lacking. In this work, we present the first multicenter lymphoma benchmarking dataset covering four common lymphoma subtypes and healthy control tissue. We systematically evaluate five publicly available pathology foundation models (H-optimus-1, H0-mini, Virchow2, UNI2, Titan) combined with attention-based (AB-MIL) and transformer-based (TransMIL) multiple instance learning aggregators across three magnifications (10x, 20x, 40x). On in-distribution test sets, models achieve multiclass balanced accuracies exceeding 80% across all magnifications, with all foundation models performing similarly and both aggregation methods showing comparable results. The magnification study reveals that 40x resolution is sufficient, with no performance gains from higher resolutions or cross-magnification aggregation. However, on out-of-distribution test sets, performance drops substantially to around 60%, highlighting significant generalization challenges. To advance the field, larger multicenter studies covering additional rare lymphoma subtypes are needed. We provide an automated benchmarking pipeline to facilitate such future research.

COMMA: A Communicative Multimodal Multi-Agent Benchmark

arXiv:2410.07553v5 Announce Type: replace Abstract: The rapid advances of multimodal agents built on large foundation models have largely overlooked their potential for language-based communication between agents in collaborative tasks. This oversight presents a critical gap in understanding their effectiveness in real-world deployments, particularly when communicating with humans. Existing agentic benchmarks fail to address key aspects of inter-agent communication and collaboration, particularly in scenarios where agents have unequal access to information and must work together to achieve tasks beyond the scope of individual capabilities. To fill this gap, we introduce COMMA: a novel puzzle benchmark designed to evaluate the collaborative performance of multimodal multi-agent systems through language communication. Our benchmark features a variety of multimodal puzzles, providing a comprehensive evaluation across four key categories of agentic capability in a communicative collaboration setting. Our findings reveal surprising weaknesses in state-of-the-art models, including strong proprietary models like GPT-4o and reasoning models like o4-mini. Many chain of thought reasoning models such as R1-Onevision and LLaVA-CoT struggle to outperform even a random baseline in agent-agent collaboration, indicating a potential growth area in their communication abilities.

A Knowledge Graph-based Retrieval-Augmented Generation Framework for Algorithm Selection in the Facility Layout Problem

arXiv:2509.18054v2 Announce Type: replace-cross Abstract: Selecting a solution algorithm for the Facility Layout Problem (FLP), an NP-hard optimization problem with multiobjective trade-off, is a complex task that requires deep expert knowledge. The performance of a given algorithm depends on the specific characteristics of the problem, such as the number of facilities, objectives, and constraints. This creates a need for a data-driven recommendation method to guide algorithm selection in automated design systems. This paper introduces a new recommendation method to make this expertise accessible, based on a Knowledge Graph-Based Retrieval-Augmented Generation (KG-RAG) framework. In this framework, a domain-specific knowledge graph (KG) is constructed from the literature. The method then employs a multifaceted retrieval mechanism to gather relevant evidence from this KG using three distinct approaches: precise graph-based search, flexible vector-based search, and cluster-based high-level search. The retrieved evidence is utilized by a Large Language Model (LLM) to generate algorithm recommendations based on data-driven reasoning. This KG-RAG framework is tested on a use case consisting of six problems comprising of complex multi-objective and multi-constraint FLP case. The results are compared with the Gemini 1.5 Flash chatbot. The results show that KG-RAG achieves an average reasoning score of 4.7 out of 5 compared to 3.3 for the baseline chatbot.

Single-cell and spatial transcriptomic characterization of pulmonary pleomorphic carcinoma

Commun Biol. 2025 Dec 16;8(1):1773. doi: 10.1038/s42003-025-09162-w.

ABSTRACT

Pulmonary pleomorphic carcinoma (PPC) is a rare subtype of lung cancer that comprises both epithelial and sarcomatoid components. The molecular basis of PPC, including the cellular dynamics of its components, remains largely unknown. To elucidate potential therapeutic targets for PPC, we perform a multi-omics analysis incorporating digital spatial profiling and single-cell RNA sequencing (scRNA-seq). PPC exhibits diverse driver gene alterations, including MET exon 14 skipping mutation (METex14) and ALK fusion. In spatial transcriptomics, MET gene and protein are overexpressed exclusively within the epithelial component and not in the sarcomatoid component, even in patients harboring METex14. Epithelial-mesenchymal transition (EMT)-related transcriptional changes, along with extracellular matrix (ECM) remodeling between the epithelial and sarcomatoid components, are observed. scRNA-seq identifies cell populations within the epithelial component that contribute to the malignant transformation and differentiation of the sarcomatoid component. They are characterized by an intermediate EMT state with ECM remodeling signature, suggesting their potential as novel therapeutic targets for PPC.

PMID:41402584 | PMC:PMC12708732 | DOI:10.1038/s42003-025-09162-w

New Perspectives on Gastric Inflammaging: Integrating Multi-Omics Mechanisms and Gerotherapeutic Strategies in Chronic Gastritis

Aging Dis. 2025 Dec 15. doi: 10.14336/AD.2025.1444. Online ahead of print.

ABSTRACT

Chronic gastritis (CG) is a highly prevalent, age-associated inflammatory disorder of gastric mucosa and a key precursor of gastric cancer in older adults. Beyond Helicobacter pylori infection and environmental insults, accumulating evidence indicates that chronic, low-grade inflammation coupled with aging biology, "gastric inflammaging", plays a central role in driving mucosal degeneration, atrophy, and malignant transformation. Here, we synthesize current mechanistic and multi-omics evidence to conceptualize CG as a tractable model of organ-specific inflammaging. We first summarize how hallmarks of aging-including cellular senescence and the senescence-associated secretory phenotype (SASP), mitochondrial dysfunction, impaired autophagy, immune exhaustion, and microbiome dysbiosis-converge to create a self-perpetuating inflammatory microenvironment in the stomach. We then review emerging single-cell and spatial multi-omics studies that delineate senescence-inflammation niches and reveal how these molecular neighborhoods relate to disease stage and cancer risk. Finally, we discuss therapeutic implications, highlighting geroscience-guided interventions such as senolytics/senomorphics, inflammasome and cGAS-STING pathway modulators, microbiota- and metabolite-targeted strategies, lifestyle interventions, and natural products, and propose a precision framework linking inflammaging biomarkers to patient stratification and clinical endpoints. Reframing CG as a gastric inflammaging model may provide a prototype for organ-specific healthy aging strategies and near-term gerotherapeutic trials aimed at extending healthspan.

PMID:41400573 | DOI:10.14336/AD.2025.1444

Opinion: We crunched the numbers on drug discovery in the U.S. vs. China. The results were alarming

16 December 2025 at 17:30

This essay is part of a First Opinion series on the future of the National Institutes of Health and American science.

Around the world, nations with robust research and development infrastructure race to create therapeutics that meet the needs of their residents. Simply put, they dictate research priorities based on need. During the Covid pandemic, the United States was one of the first countries to gain access to vaccines to protect its citizens.

Read the rest…

© Adobe

Generative AI Mental Health Chatbots as Therapeutic Tools: Systematic Review and Meta-Analysis of Their Role in Reducing Mental Health Issues

Background: To date, there is no comprehensive paper that systematically synthesizes the effect of generative AI chatbot’s impact on mental health. Can generative AI chatbots help reduce our psychological distress? Objective: To comprehensively assess existing evidence, a systematic review and meta-analysis is essential to evaluate the overall effectiveness, identify gaps, and guide future research in this evolving field. This paper aims to: 1) synthesize current evidence on generative AI chatbot interventions targeting mental health issues, 2) quantify the effectiveness of these interventions via a meta-analysis of randomized controlled trials (RCTs), and examine key moderators of intervention effectiveness. Methods: This systematic review included 26 studies for narrative synthesis, out of which 12 randomized controlled trials were included in the meta-analysis. Results: The systematic synthesis revealed that 1) generative AI-chatbot interventions mostly took place in non-WEIRD countries (Western, Educated, Industrialized, Rich, and Democratic) and 2) there is a lack of studies focusing on young children and older adults. The meta-analysis showed a statistically significant effect (ES = 0.36, p = .039), which means that generative AI chatbots are, on average, effective in reducing negative mental health issues. Among moderators, we found statistically significant and higher effect sizes among interventions that have an active control group, conducted in WEIRD countries, recruited non-clinical populations, older age, majority female, non-personalized, with human assistance, and social-oriented. Conclusions: In conclusion, this comprehensive review has highlighted the potential of generative AI chatbots in addressing anxiety, depression, negative mood, and stress. The findings indicate that generative AI interventions are particularly beneficial in WEIRD countries, among non-clinical populations, older adults, and females. Human-assisted and social-oriented programs, as opposed to fully autonomous or task-oriented ones, demonstrate greater effectiveness. Meanwhile, non-personalized chatbots appear to yield more effective outcomes than personalized systems.

Robustness of Probabilistic Models to Low-Quality Data: A Multi-Perspective Analysis

arXiv:2512.11912v1 Announce Type: new Abstract: A systematic, comparative investigation into the effects of low-quality data reveals a stark spectrum of robustness across modern probabilistic models. We find that autoregressive language models, from token prediction to sequence-to-sequence tasks, are remarkably resilient (for GPT-2, test NLL increases modestly from 2.87 to 3.59 despite 50% token corruption). By contrast, under the same levels of data corruption, class-conditional diffusion models degrade catastrophically (image-label consistency plummets by 56.81% relative to baseline), while classifiers show a moderate impact that diminishes with dataset scale. To explain these discrepancies, we analyze the results through a multi-perspective lens, integrating information theory, PAC learning, and gradient dynamics. These analyses suggest that robustness is heavily influenced by two key principles: the richness of conditioning information, which constrains the learning problem, and the absolute information content of the training data, which allows the signal from correct information to dominate statistical noise.
❌