❌

Normal view

The Seeds of Scheming: Weakness of Will in the Building Blocks of Agentic Systems

arXiv:2512.05449v1 Announce Type: new Abstract: Large language models display a peculiar form of inconsistency: they "know" the correct answer but fail to act on it. In human philosophy, this tension between global judgment and local impulse is called akrasia, or weakness of will. We propose akrasia as a foundational concept for analyzing inconsistency and goal drift in agentic AI systems. To operationalize it, we introduce a preliminary version of the Akrasia Benchmark, currently a structured set of prompting conditions (Baseline [B], Synonym [S], Temporal [T], and Temptation [X]) that measures when a model's local response contradicts its own prior commitments. The benchmark enables quantitative comparison of "self-control" across model families, decoding strategies, and temptation types. Beyond single-model evaluation, we outline how micro-level akrasia may compound into macro-level instability in multi-agent systems that may be interpreted as "scheming" or deliberate misalignment. By reframing inconsistency as weakness of will, this work connects agentic behavior to classical theories of agency and provides an empirical bridge between philosophy, psychology, and the emerging science of agentic AI.

Simulating Life Paths with Digital Twins: AI-Generated Future Selves Influence Decision-Making and Expand Human Choice

arXiv:2512.05397v1 Announce Type: cross Abstract: Major life transitions demand high-stakes decisions, yet people often struggle to imagine how their future selves will live with the consequences. To support this limited capacity for mental time travel, we introduce AI-enabled digital twins that have ``lived through'' simulated life scenarios. Rather than predicting optimal outcomes, these simulations extend prospective cognition by making alternative futures vivid enough to support deliberation without assuming which path is best. We evaluate this idea in a randomized controlled study (N=192) using multimodal synthesis - facial age progression, voice cloning, and large language model dialogue - to create personalized avatars representing participants 30 years forward. Young adults 18 to 28 years old described pending binary decisions and were assigned to guided imagination or one of four avatar conditions: single-option, balanced dual-option, or expanded three-option with a system-generated novel alternative. Results showed asymmetric effects: single-sided avatars increased shifts toward the presented option, while balanced presentation produced movement toward both. Introducing a system-generated third option increased adoption of this new alternative compared to control, suggesting that AI-generated future selves can expand choice by surfacing paths that might otherwise go unnoticed. Participants rated evaluative reasoning and eudaimonic meaning-making as more important than emotional or visual vividness. Perceived persuasiveness and baseline agency predicted decision change. These findings advance understanding of AI-mediated episodic prospection and raise questions about autonomy in AI-augmented decisions.

Optimizing Medical Question-Answering Systems: A Comparative Study of Fine-Tuned and Zero-Shot Large Language Models with RAG Framework

arXiv:2512.05863v1 Announce Type: cross Abstract: Medical question-answering (QA) systems can benefit from advances in large language models (LLMs), but directly applying LLMs to the clinical domain poses challenges such as maintaining factual accuracy and avoiding hallucinations. In this paper, we present a retrieval-augmented generation (RAG) based medical QA system that combines domain-specific knowledge retrieval with open-source LLMs to answer medical questions. We fine-tune two state-of-the-art open LLMs (LLaMA~2 and Falcon) using Low-Rank Adaptation (LoRA) for efficient domain specialization. The system retrieves relevant medical literature to ground the LLM's answers, thereby improving factual correctness and reducing hallucinations. We evaluate the approach on benchmark datasets (PubMedQA and MedMCQA) and show that retrieval augmentation yields measurable improvements in answer accuracy compared to using LLMs alone. Our fine-tuned LLaMA~2 model achieves 71.8% accuracy on PubMedQA, substantially improving over the 55.4% zero-shot baseline, while maintaining transparency by providing source references. We also detail the system design and fine-tuning methodology, demonstrating that grounding answers in retrieved evidence reduces unsupported content by approximately 60%. These results highlight the potential of RAG-augmented open-source LLMs for reliable biomedical QA, pointing toward practical clinical informatics applications.

The AI Productivity Index (APEX)

arXiv:2509.25721v4 Announce Type: replace-cross Abstract: We present an extended version of the AI Productivity Index (APEX-v1-extended), a benchmark for assessing whether frontier models are capable of performing economically valuable tasks in four jobs: investment banking associate, management consultant, big law associate, and primary care physician (MD). This technical report details the extensions to APEX-v1, including an increase in the held-out evaluation set from n = 50 to n = 100 cases per job (n = 400 total) and updates to the grading methodology. We present a new leaderboard, where GPT5 (Thinking = High) remains the top performing model with a score of 67.0%. APEX-v1-extended shows that frontier models still have substantial limitations when performing typical professional tasks. To support further research, we are open sourcing n = 25 non-benchmark example cases per role (n = 100 total) along with our evaluation harness.

Molecular stratification of esophageal adenocarcinoma: implications for prognosis and treatment strategy

Oncogene, Published online: 08 December 2025; doi:10.1038/s41388-025-03650-3

Molecular stratification of esophageal adenocarcinoma: implications for prognosis and treatment strategy

A typology of physician input approaches to using AI chatbots for clinical decision-making

npj Digital Medicine, Published online: 05 December 2025; doi:10.1038/s41746-025-02184-y

A typology of physician input approaches to using AI chatbots for clinical decision-making

Privacy Risks and Preservation Methods in Explainable Artificial Intelligence: A Scoping Review

arXiv:2505.02828v3 Announce Type: replace Abstract: Explainable Artificial Intelligence (XAI) has emerged as a pillar of Trustworthy AI and aims to bring transparency in complex models that are opaque by nature. Despite the benefits of incorporating explanations in models, an urgent need is found in addressing the privacy concerns of providing this additional information to end users. In this article, we conduct a scoping review of existing literature to elicit details on the conflict between privacy and explainability. Using the standard methodology for scoping review, we extracted 57 articles from 1,943 studies published from January 2019 to December 2024. The review addresses 3 research questions to present readers with more understanding of the topic: (1) what are the privacy risks of releasing explanations in AI systems? (2) what current methods have researchers employed to achieve privacy preservation in XAI systems? (3) what constitutes a privacy preserving explanation? Based on the knowledge synthesized from the selected studies, we categorize the privacy risks and preservation methods in XAI and propose the characteristics of privacy preserving explanations to aid researchers and practitioners in understanding the requirements of XAI that is privacy compliant. Lastly, we identify the challenges in balancing privacy with other system desiderata and provide recommendations for achieving privacy preserving XAI. We expect that this review will shed light on the complex relationship of privacy and explainability, both being the fundamental principles of Trustworthy AI.

A Definition of AGI

arXiv:2510.18212v3 Announce Type: replace Abstract: The lack of a concrete definition for Artificial General Intelligence (AGI) obscures the gap between today's specialized AI and human-level cognition. This paper introduces a quantifiable framework to address this, defining AGI as matching the cognitive versatility and proficiency of a well-educated adult. To operationalize this, we ground our methodology in Cattell-Horn-Carroll theory, the most empirically validated model of human cognition. The framework dissects general intelligence into ten core cognitive domains-including reasoning, memory, and perception-and adapts established human psychometric batteries to evaluate AI systems. Application of this framework reveals a highly "jagged" cognitive profile in contemporary models. While proficient in knowledge-intensive domains, current AI systems have critical deficits in foundational cognitive machinery, particularly long-term memory storage. The resulting AGI scores (e.g., GPT-4 at 27%, GPT-5 at 57%) concretely quantify both rapid progress and the substantial gap remaining before AGI.

AI Deception: Risks, Dynamics, and Controls

arXiv:2511.22619v2 Announce Type: replace Abstract: As intelligence increases, so does its shadow. AI deception, in which systems induce false beliefs to secure self-beneficial outcomes, has evolved from a speculative concern to an empirically demonstrated risk across language models, AI agents, and emerging frontier systems. This project provides a comprehensive and up-to-date overview of the AI deception field, covering its core concepts, methodologies, genesis, and potential mitigations. First, we identify a formal definition of AI deception, grounded in signaling theory from studies of animal deception. We then review existing empirical studies and associated risks, highlighting deception as a sociotechnical safety challenge. We organize the landscape of AI deception research as a deception cycle, consisting of two key components: deception emergence and deception treatment. Deception emergence reveals the mechanisms underlying AI deception: systems with sufficient capability and incentive potential inevitably engage in deceptive behaviors when triggered by external conditions. Deception treatment, in turn, focuses on detecting and addressing such behaviors. On deception emergence, we analyze incentive foundations across three hierarchical levels and identify three essential capability preconditions required for deception. We further examine contextual triggers, including supervision gaps, distributional shifts, and environmental pressures. On deception treatment, we conclude detection methods covering benchmarks and evaluation protocols in static and interactive settings. Building on the three core factors of deception emergence, we outline potential mitigation strategies and propose auditing approaches that integrate technical, community, and governance efforts to address sociotechnical challenges and future AI risks. To support ongoing work in this area, we release a living resource at www.deceptionsurvey.com.

Challenges and Limitations of Generative AI in Synthesizing Wearable Sensor Data

arXiv:2505.14206v2 Announce Type: replace-cross Abstract: The widespread adoption of wearable sensors has the potential to provide massive and heterogeneous time series data, driving the use of Artificial Intelligence in human sensing applications. However, data collection remains limited due to stringent ethical regulations, privacy concerns, and other constraints, hindering progress in the field. Synthetic data generation, particularly through Generative Adversarial Networks and Diffusion Models, has emerged as a promising solution to mitigate both data scarcity and privacy issues. However, these models are often limited to narrow operational scenarios, such as short-term and unimodal signal patterns. To address this gap, we present a systematic evaluation of state-of-the-art generative models for time series data, explicitly assessing their performance in challenging scenarios such as stress and emotion recognition. Our study examines the extent to which these models can jointly handle multi-modality, capture long-range dependencies, and support conditional generation-core requirements for real-world wearable sensor data generation. To enable a fair and rigorous comparison, we also introduce an evaluation framework that evaluates both the intrinsic fidelity of the generated data and their utility in downstream predictive tasks. Our findings reveal critical limitations in the existing approaches, particularly in maintaining cross-modal consistency, preserving temporal coherence, and ensuring robust performance in train-on-synthetic, test-on-real, and data augmentation scenarios. Finally, we present our future research directions to enhance synthetic time series generation and improve the applicability of generative models in the wearable computing domain.

Harnessing cuproptosis for pancreatic cancer therapy: From molecular insights to clinical prospects

Biomed Pharmacother. 2025 Dec 2;193:118852. doi: 10.1016/j.biopha.2025.118852. Online ahead of print.

ABSTRACT

Pancreatic cancer (PC) remains a high-fatality malignancy with limited clinical progress, characterized by aggressive biology, marked resistance to standard therapies, and dismal outcomes. Even with state-of-the-art resection, radiotherapy, and multidrug chemotherapy, median survival benefits are modest, highlighting an urgent need for mechanism-based interventions. Cuproptosis, a newly delineated modality of regulated cell death initiated by intracellular copper accumulation and mitochondrial stress, presents a biologically coherent therapeutic avenue. Distinct from apoptosis, necroptosis, and ferroptosis, cuproptosis is driven by the direct binding of copper to lipoylated enzymes of the tricarboxylic acid (TCA) cycle, resulting in bioenergetic failure, misfolded protein aggregation, and collapse of cytotoxic proteostasis. Converging studies suggest that copper disequilibrium and metabolic reprogramming are recurrent features of PC, potentially contributing to malignant progression, immune evasion, and chemoresistance. These insights motivate two complementary strategies: first, therapeutic manipulation of copper flux, via chelators, ionophores, or transport modulators, to selectively trigger cuproptosis in tumor cells; and second, sensitization of mitochondrial metabolism, through targeting lipoic-acid pathway components, pyruvate utilization, or TCA load, to lower the threshold for cuproptotic killing. In parallel, multi-omic interrogation of cuproptosis-associated genes, proteins, and metabolites may yield prognostic and predictive biomarkers, enabling risk-adapted treatment selection and rational combinations with cytotoxic, targeted, or immunotherapeutic modalities. This review synthesizes recent advances on cuproptosis in PC and outlines its translational potential as both a therapeutic target and a biomarker framework.

PMID:41337879 | DOI:10.1016/j.biopha.2025.118852

A Framework for Causal Concept-based Model Explanations

arXiv:2512.02735v1 Announce Type: new Abstract: This work presents a conceptual framework for causal concept-based post-hoc Explainable Artificial Intelligence (XAI), based on the requirements that explanations for non-interpretable models should be understandable as well as faithful to the model being explained. Local and global explanations are generated by calculating the probability of sufficiency of concept interventions. Example explanations are presented, generated with a proof-of-concept model made to explain classifiers trained on the CelebA dataset. Understandability is demonstrated through a clear concept-based vocabulary, subject to an implicit causal interpretation. Fidelity is addressed by highlighting important framework assumptions, stressing that the context of explanation interpretation must align with the context of explanation generation.

Digital Biometrics in Predicting Risk for Obstructive Sleep Apnea and Hypertension: Decentralized, Prospective Cohort Study

Background: Sleep is an important component of human health and can be measured longitudinally using digital activity trackers. Further, decentralized digital research has the potential to provide a real-world picture of sleep in large populations. Objective: This study examined whether longitudinal sleep patterns from activity trackers could predict risk of obstructive sleep apnea (OSA) and hypertension, as defined the Berlin questionnaire and self report, respectively. Methods: We recruited adults ≥18 years nationwide to join our sleep-focused smartphone-based study, called the Research Framework for Exploring Sleep Health (REFRESH). Our sample of N= 391 adults is comprised predominately of females (68%; n = 247 out of 364) at a mean age of 48 years (standard deviation (SD) = 13.62). Participants were asked to fill out health-related surveys, including the Berlin questionnaire, and the Horne-Ostberg questionnaire for chronotype. Participants were asked to link their own activity tracker to the application to collect longitudinal sleep data. Results: We analyzed sleep data from 391 participants, among a predominately White (65%; n = 231 out of 353) followed by multiracial (17%; n = 61 out of 353) and Hispanic or Latino (6.5%; n = 23 out of 353) cohort. Collinearity testing showed that OSA risk and self-reported hypertension could be considered independently. Holding body mass index (BMI) at fixed value, the odds of having high OSA risk increased by 159% for every one-hour increase in weekday sleep variability (odds ratio (OR) = 2.592, 95% confidence interval (CI) [1.613, 4.400]; P

Multi-omic profiling provides insights into the heterogeneity, microenvironmental features, and biomarker landscape of small-cell lung cancer

Mol Cancer. 2025 Dec 2. doi: 10.1186/s12943-025-02514-4. Online ahead of print.

ABSTRACT

BACKGROUND: Greater understanding of differential therapeutic sensitivity, specifically to immunotherapy, in small-cell lung cancer (SCLC) is required.

METHODS: We explored SCLC heterogeneity through integrated molecular characterization of tumor tissue samples from 159 treatment-naive patients, utilizing genetic, epigenetic, transcriptional, and proteomic profiling, immunohistochemistry staining for multiple biologically relevant markers including transcriptional subtype-defining proteins, and spatial immune profiling using multiplex immunofluorescence.

RESULTS: Multi-omics analysis confirmed high heterogeneity across/within neuroendocrine and non-neuroendocrine subtypes. Methylomics analysis identified four methylome clusters that may enhance subtype prediction, prognosis, and longitudinal monitoring of subtype evolution. Immunohistochemistry analysis showed high MHC-I expression in non-neuroendocrine subtypes, which have greatest potential benefit from adding immunotherapy to chemotherapy; high DLL3 expression associated with neuroendocrine subtypes and an immune-cold tumor microenvironment. Multiplex immunofluorescence demonstrated associations of MHC-I with spatial arrangement and phenotypic features of immune cells in the tumor microenvironment of high-MHC-I-expressing SCLC, providing mechanistic rationale for MHC-I as a potential biomarker of immunotherapy response.

CONCLUSIONS: This multimodal profiling analysis provides further insights into the biologic complexity of SCLC and highlights potential therapeutic vulnerabilities of distinct disease subtypes.

PMID:41331472 | DOI:10.1186/s12943-025-02514-4

Whole-genome landscapes of 1,364 breast cancers

Nature, Published online: 03 December 2025; doi:10.1038/s41586-025-09812-3

Whole-genome and transcriptome analysis of 1,364 cases of breast cancer from South Korea broadens our understanding of breast cancer biology and reveals genomic features that connect tumour biology with treatment responses and clinical outcomes.

CACARA: Cross-Modal Alignment Leveraging a Text-Centric Approach for Cost-Effective Multimodal and Multilingual Learning

arXiv:2512.00496v1 Announce Type: cross Abstract: As deep learning models evolve, new applications and challenges are rapidly emerging. Tasks that once relied on a single modality, such as text, images, or audio, are now enriched by seamless interactions between multimodal data. These connections bridge information gaps: an image can visually materialize a text, while audio can add context to an image. Researchers have developed numerous multimodal models, but most rely on resource-intensive training across multiple modalities. Similarly, extending these models to new languages often follows the same resource-heavy training strategy. In this work, we propose a multimodal and multilingual architecture, CACARA, trained through emergent alignment learning, enabling the seamless integration of new modalities into an existing bimodal/multimodal model without requiring full retraining. This work breaks new ground by demonstrating that this emergent alignment paradigm can unlock multilingual capabilities from monolingual training. By fine-tuning the newly incorporated modality only on data aligned with the English language, our model develops support for over 100 languages without explicit multilingual pretraining or tuning of the text encoder. Such emergent multimodal and multilingual properties are gained efficiently, preserving previously learned knowledge at a training cost comparable to that of a monolingual model. Our strategy achieves up to a 14.24 percentage points improvement in R@1 audio-to-text retrieval, outperforming state-of-the-art multimodal models -- all without the heavy computational cost of retraining across every modality and language.
  • ✇cs.AI, q-bio.NC updates on arXiv.org
  • Slovak Conceptual Dictionary Miroslav Bl\v{s}t\'ak
    arXiv:2512.00579v1 Announce Type: cross Abstract: When solving tasks in the field of natural language processing, we sometimes need dictionary tools, such as lexicons, word form dictionaries or knowledge bases. However, the availability of dictionary data is insufficient in many languages, especially in the case of low resourced languages. In this article, we introduce a new conceptual dictionary for the Slovak language as the first linguistic tool of this kind. Since Slovak language is a langu
     

Slovak Conceptual Dictionary

arXiv:2512.00579v1 Announce Type: cross Abstract: When solving tasks in the field of natural language processing, we sometimes need dictionary tools, such as lexicons, word form dictionaries or knowledge bases. However, the availability of dictionary data is insufficient in many languages, especially in the case of low resourced languages. In this article, we introduce a new conceptual dictionary for the Slovak language as the first linguistic tool of this kind. Since Slovak language is a language with limited linguistic resources and there are currently not available any machine-readable linguistic data sources with a sufficiently large volume of data, many tasks which require automated processing of Slovak text achieve weaker results compared to other languages and are almost impossible to solve.

Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models

arXiv:2512.00590v1 Announce Type: cross Abstract: Knowledge graphs (KGs) provide structured, verifiable grounding for large language models (LLMs), but current LLM-based systems commonly use KGs as auxiliary structures for text retrieval, leaving their intrinsic quality underexplored. In this work, we propose Wikontic, a multi-stage pipeline that constructs KGs from open-domain text by extracting candidate triplets with qualifiers, enforcing Wikidata-based type and relation constraints, and normalizing entities to reduce duplication. The resulting KGs are compact, ontology-consistent, and well-connected; on MuSiQue, the correct answer entity appears in 96% of generated triplets. On HotpotQA, our triplets-only setup achieves 76.0 F1, and on MuSiQue 59.8 F1, matching or surpassing several retrieval-augmented generation baselines that still require textual context. In addition, Wikontic attains state-of-the-art information-retention performance on the MINE-1 benchmark (86%), outperforming prior KG construction methods. Wikontic is also efficient at build time: KG construction uses less than 1,000 output tokens, about 3$\times$ fewer than AriGraph and $

Deep Learning-Based Computer Vision Models for Early Cancer Detection Using Multimodal Medical Imaging and Radiogenomic Integration Frameworks

arXiv:2512.00714v1 Announce Type: cross Abstract: Early cancer detection remains one of the most critical challenges in modern healthcare, where delayed diagnosis significantly reduces survival outcomes. Recent advancements in artificial intelligence, particularly deep learning, have enabled transformative progress in medical imaging analysis. Deep learning-based computer vision models, such as convolutional neural networks (CNNs), transformers, and hybrid attention architectures, can automatically extract complex spatial, morphological, and temporal patterns from multimodal imaging data including MRI, CT, PET, mammography, histopathology, and ultrasound. These models surpass traditional radiological assessment by identifying subtle tissue abnormalities and tumor microenvironment variations invisible to the human eye. At a broader scale, the integration of multimodal imaging with radiogenomics linking quantitative imaging features with genomics, transcriptomics, and epigenetic biomarkers has introduced a new paradigm for personalized oncology. This radiogenomic fusion allows the prediction of tumor genotype, immune response, molecular subtypes, and treatment resistance without invasive biopsies.
❌