❌

Reading view

DeFi TrustBoost: Blockchain and AI for Trustworthy Decentralized Financial Decisions

arXiv:2512.00142v1 Announce Type: cross Abstract: This research introduces the Decentralized Finance (DeFi) TrustBoost Framework, which combines blockchain technology and Explainable AI to address challenges faced by lenders underwriting small business loan applications from low-wealth households. The framework is designed with a strong emphasis on fulfilling four crucial requirements of blockchain and AI systems: confidentiality, compliance with data protection laws, resistance to adversarial attacks, and compliance with regulatory audits. It presents a technique for tamper-proof auditing of automated AI decisions and a strategy for on-chain (inside-blockchain) and off-chain data storage to facilitate collaboration within and across financial organizations.
  •  

Rethinking Lung Cancer Screening: AI Nodule Detection and Diagnosis Outperforms Radiologists, Leading Models, and Standards Beyond Size and Growth

arXiv:2512.00281v1 Announce Type: cross Abstract: Early detection of malignant lung nodules is critical, but its dependence on size and growth in screening inherently delays diagnosis. We present an AI system that redefines lung cancer screening by performing both detection and malignancy diagnosis directly at the nodule level on low-dose CT scans. To address limitations in dataset scale and explainability, we designed an ensemble of shallow deep learning and feature-based specialized models. Trained and evaluated on 25,709 scans with 69,449 annotated nodules, the system outperforms radiologists, Lung-RADS, and leading AI models (Sybil, Brock, Google, Kaggle). It achieves an area under the receiver operating characteristic curve (AUC) of 0.98 internally and 0.945 on an independent cohort. With 0.5 false positives per scan at 99.3\% sensitivity, it addresses key barriers to AI adoption. Critically, it outperforms radiologists across all nodule sizes and stages, excelling in stage 1 cancers, and all growth-based metrics, including the least accurate: Volume-Doubling Time. It also surpasses radiologists by up to one year in diagnosing indeterminate and slow-growing nodules.
  •  

MedCondDiff: Lightweight, Robust, Semantically Guided Diffusion for Medical Image Segmentation

arXiv:2512.00350v1 Announce Type: cross Abstract: We introduce MedCondDiff, a diffusion-based framework for multi-organ medical image segmentation that is efficient and anatomically grounded. The model conditions the denoising process on semantic priors extracted by a Pyramid Vision Transformer (PVT) backbone, yielding a semantically guided and lightweight diffusion architecture. This design improves robustness while reducing both inference time and VRAM usage compared to conventional diffusion models. Experiments on multi-organ, multi-modality datasets demonstrate that MedCondDiff delivers competitive performance across anatomical regions and imaging modalities, underscoring the potential of semantically guided diffusion models as an effective class of architectures for medical imaging tasks.
  •  

SelfAI: Building a Self-Training AI System with LLM Agents

arXiv:2512.00403v1 Announce Type: cross Abstract: Recent work on autonomous scientific discovery has leveraged LLM-based agents to integrate problem specification, experiment planning, and execution into end-to-end systems. However, these frameworks are often confined to narrow application domains, offer limited real-time interaction with researchers, and lack principled mechanisms for determining when to halt exploration, resulting in inefficiencies, reproducibility challenges, and under-utilized human expertise. To address these gaps, we propose \textit{SelfAI}, a general multi-agent platform that combines a User Agent for translating high-level research objectives into standardized experimental configurations, a Cognitive Agent powered by LLMs with optimal stopping criteria to iteratively refine hyperparameter searches, and an Experiment Manager responsible for orchestrating parallel, fault-tolerant training workflows across heterogeneous hardware while maintaining a structured knowledge base for continuous feedback. We further introduce two novel evaluation metrics, Score and $\text{AUP}_D$, to quantify discovery efficiency and search diversity. Across regression, NLP, computer vision, scientific computing, medical imaging, and drug discovery benchmarks, SelfAI consistently achieves strong performance and reduces redundant trials compared to classical Bayesian optimization and LLM-based baselines, while enabling seamless interaction with human researchers.
  •  

Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models

arXiv:2512.00590v1 Announce Type: cross Abstract: Knowledge graphs (KGs) provide structured, verifiable grounding for large language models (LLMs), but current LLM-based systems commonly use KGs as auxiliary structures for text retrieval, leaving their intrinsic quality underexplored. In this work, we propose Wikontic, a multi-stage pipeline that constructs KGs from open-domain text by extracting candidate triplets with qualifiers, enforcing Wikidata-based type and relation constraints, and normalizing entities to reduce duplication. The resulting KGs are compact, ontology-consistent, and well-connected; on MuSiQue, the correct answer entity appears in 96% of generated triplets. On HotpotQA, our triplets-only setup achieves 76.0 F1, and on MuSiQue 59.8 F1, matching or surpassing several retrieval-augmented generation baselines that still require textual context. In addition, Wikontic attains state-of-the-art information-retention performance on the MINE-1 benchmark (86%), outperforming prior KG construction methods. Wikontic is also efficient at build time: KG construction uses less than 1,000 output tokens, about 3$\times$ fewer than AriGraph and $
  •  

Human Decision-making is Susceptible to AI-driven Manipulation

arXiv:2502.07663v3 Announce Type: replace Abstract: AI systems are increasingly intertwined with daily life, assisting users with various tasks and guiding decision-making. This integration introduces risks of AI-driven manipulation, where such systems may exploit users' cognitive biases and emotional vulnerabilities to steer them toward harmful outcomes. Through a randomized between-subjects experiment with 233 participants, we examined human susceptibility to such manipulation in financial (e.g., purchases) and emotional (e.g., conflict resolution) decision-making contexts. Participants interacted with one of three AI agents: a neutral agent (NA) optimizing for user benefit without explicit influence, a manipulative agent (MA) designed to covertly influence beliefs and behaviors, or a strategy-enhanced manipulative agent (SEMA) equipped with established psychological tactics, allowing it to select and apply them adaptively during interactions to reach its hidden objectives. By analyzing participants' preference ratings, we found significant susceptibility to AI-driven manipulation. Particularly across both decision-making domains, interacting with the manipulative agents significantly increased the odds of rating hidden incentives higher than optimal options (Financial, MA: OR=5.24, SEMA: OR=7.96; Emotional, MA: OR=5.52, SEMA: OR=5.71) compared to the NA group. Notably, we found no clear evidence that employing psychological strategies (SEMA) was overall more effective than simple manipulative objectives (MA) on our primary outcomes. Hence, AI-driven manipulation could become widespread even without requiring sophisticated tactics and expertise. While our findings are preliminary and derived from hypothetical, low-stakes scenarios, we highlight a critical vulnerability in human-AI interactions, emphasizing the need for ethical safeguards and regulatory frameworks to protect human autonomy.
  •  

The Unified Cognitive Consciousness Theory for Language Models: Anchoring Semantics, Thresholds of Activation, and Emergent Reasoning

arXiv:2506.02139v5 Announce Type: replace Abstract: We propose semantic anchoring, a unified account of how large language models turn pretrained capacity into goal-directed behavior: external structure (in-context examples, retrieval, or light tuning) binds the model's latent patterns to desired targets. Unified Contextual Control Theory (UCCT) formalizes this via anchoring strength $S = \rho_d - d_r - \log k$, where $\rho_d$ measures target cohesion in representation space, $d_r$ measures mismatch from prior knowledge, and $k$ is the anchor budget. UCCT predicts threshold-like performance flips and strictly generalizes in-context learning, reading retrieval and fine-tuning as anchoring variants. Three controlled studies provide evidence. Experiment 1 demonstrates cross-domain anchoring rebinding strong priors in text and vision. Experiment 2 varies representational familiarity via numeral bases (base-10/8/9) at fixed complexity, yielding ordered thresholds and transfer patterns tracking $\rho_d$, $d_r$, and $S$. Experiment 3 establishes a geometry-to-behavior correlate: layer-wise peak anchoring and trajectory area predict few-shot thresholds $\theta_{50}$. UCCT offers testable theory and practical metrics for optimizing prompts, retrieval, and tuning.
  •  

Life-Code: Central Dogma Modeling with Multi-Omics Sequence Unification

arXiv:2502.07299v3 Announce Type: replace-cross Abstract: The interactions between DNA, RNA, and proteins are fundamental to biological processes, as illustrated by the central dogma of molecular biology. Although modern biological pre-trained models have achieved great success in analyzing these macromolecules individually, their interconnected nature remains underexplored. This paper follows the guidance of the central dogma to redesign both the data and model pipeline and offers a comprehensive framework, Life-Code, that spans different biological functions. As for data flow, we propose a unified pipeline to integrate multi-omics data by reverse-transcribing RNA and reverse-translating amino acids into nucleotide-based sequences. As for the model, we design a codon tokenizer and a hybrid long-sequence architecture to encode the interactions between coding and non-coding regions through masked modeling pre-training. To model the translation and folding process with coding sequences, Life-Code learns protein structures of the corresponding amino acids by knowledge distillation from off-the-shelf protein language models. Such designs enable Life-Code to capture complex interactions within genetic sequences, providing a more comprehensive understanding of multi-omics with the central dogma. Extensive experiments show that Life-Code achieves state-of-the-art results on various tasks across three omics, highlighting its potential for advancing multi-omics analysis and interpretation.
  •  

The AI Productivity Index (APEX)

arXiv:2509.25721v3 Announce Type: replace-cross Abstract: We present an extended version of the AI Productivity Index (APEX-v1-extended), a benchmark for assessing whether frontier models are capable of performing economically valuable tasks in four jobs: investment banking associate, management consultant, big law associate, and primary care physician (MD). This technical report details the extensions to APEX-v1, including an increase in the held-out evaluation set from n = 50 to n = 100 cases per job (n = 400 total) and updates to the grading methodology. We present a new leaderboard, where GPT5 (Thinking = High) remains the top performing model with a score of 67.0%. APEX-v1-extended shows that frontier models still have substantial limitations when performing typical professional tasks. To support further research, we are open sourcing n = 25 non-benchmark example cases per role (n = 100 total) along with our evaluation harness.
  •  

Maximizing the efficiency of human feedback in AI alignment: a comparative analysis

arXiv:2511.12796v2 Announce Type: replace-cross Abstract: Reinforcement Learning from Human Feedback (RLHF) relies on preference modeling to align machine learning systems with human values, yet the popular approach of random pair sampling with Bradley-Terry modeling is statistically limited and inefficient under constrained annotation budgets. In this work, we explore alternative sampling and evaluation strategies for preference inference in RLHF, drawing inspiration from areas such as game theory, statistics, and social choice theory. Our best-performing method, Swiss InfoGain, employs a Swiss tournament system with a proxy mutual-information-gain pairing rule, which significantly outperforms all other methods in constrained annotation budgets while also being more sample-efficient. Even in high-resource settings, we can identify superior alternatives to the Bradley-Terry baseline. Our experiments demonstrate that adaptive, resource-aware strategies reduce redundancy, enhance robustness, and yield statistically significant improvements in preference learning, highlighting the importance of balancing alignment quality with human workload in RLHF pipelines.
  •  

Exosome-Mediated RUNX3 DNA Delivery for Lung Cancer Therapy

ACS Appl Mater Interfaces. 2025 Dec 1. doi: 10.1021/acsami.5c15987. Online ahead of print.

ABSTRACT

Gene therapy represents a promising strategy for treating lung cancer, with the potential to inhibit the proliferation of cancerous cells and induce apoptosis. However, current gene therapy for lung cancer encounters challenges with delivery, targeting, and safety, such as off-target effects, immune responses, and the necessity for better delivery methods. Here, we introduce gene therapy using the key regulator in lung adenocarcinoma, runt-related transcription factor 3 (RUNX3), within exosomes (Exos), which are known for their biocompatibility and ability to selectively target cancer cells. We packaged the RUNX3 plasmid DNA into human exosomes (hExo-Rs), designed to target and induce apoptosis in cancer cells, resulting in a viability decrease to 43.3%. Normal fibroblasts remained viable at 96.0%, confirming the safety of hExo-Rs for future therapies. We delivered hExo-Rs to cancer spheroids, examined their effects, and found that cytokines from treated cells promote M1 macrophage polarization, emphasizing their potential for immunotherapy. We developed a hydrogel platform for the targeted 14-day release of RUNX3 pDNA by attaching hExo-Rs to gelatin using microbial transglutaminase, which enables the selective decrease in cancer cell viability and confirms apoptosis. Our demonstration of RUNX3 gene therapy with Exos presents selective anticancer effectiveness and the promise of clinical use through localized, sustained release using the hydrogel.

PMID:41325015 | DOI:10.1021/acsami.5c15987

  •  

STAT+: A drug that was ‘engineered with AI’ enters Phase 3 testing

Want to stay on top of the science and politics driving biotech today? Sign up to get our biotech newsletter in your inbox.

Good morning. My colleagues will be in Orlando later this week for the American Society of Hematology meeting. Sign up for their newsletter to stay on top of the news from the conference. We’re also holding an event there on Friday.

Fight around hospital drug discount program escalates with new lawsuit

The American Hospital Association and several hospital systems have filed a lawsuit against the Trump administration, seeking to halt an upcoming pilot program for a controversial drug discount program.

Continue to STAT+ to read the full story…

© Adobe

  •  

Opinion: Racial bias in medicine can be as simple as dismissing Black patients as a ‘hard stick’

I was moments away from a routine screening colonoscopy when it happened again. The warm and professional pre-procedure nurse began preparing for intravenous insertion. She tied the tourniquet loosely around my arm, took a quick glance, and untied it within seconds. “I can’t find a vein. You must be dehydrated,” she said, moving immediately to the back of my hand.

I paused. I didn’t feel dehydrated. Yes, I had followed the bowel prep instructions, consuming only liquids the day before, but I had no signs of dehydration. I knew my body. I knew my veins.

Read the rest…

© AIZAR RALDES/AFP via Getty Images

  •  

Acceptability of Health Information Technology by Health Care Professionals: Where We Are Now and How We Can Fill the Gap

Digital health is expected to improve efficiency and quality of health. Health information technologies (HIT) imply allocated time, appropriate training and new types of responsibility whose physical and mental impact on healthcare professionals (HCPs) has emerged as an important issue. The present review provides updated data and opinions about such potential impact and discusses the relevance of ongoing programs established to better characterize barriers and facilitators of HIT implementation. The extent of Internet-based healthcare information and digital apps imposes new responsibilities on HCPs in helping patients select reliable sources and incorporate them in the understanding and self-management of the disease. Several reviews also identified exhaustion, depersonalization, workload, over-alerting, poor work-life integration and job unsatisfaction as potential drivers of electronic health report (EHR)-associated clinician burnout and/or HIT unacceptability. Paradoxically, the increasing use of generative artificial intelligence in the decision-making process may in turn introduce an additional layer of complexity due to required specific skills and associated cognitive overload and stress. Regarding EHRs, various approaches like more proportionate use, better adequation of available commercial tools, or multidisciplinary workflows within the clinic and building of new specialty-specific tools are expected to reduce clinician burden. The ongoing E-health Efficiency Evaluation (E3) project has been developed to define the factors and dimensions impacting overall digital environment and to identify relevant ways of optimizing its acceptability by HCPs. The way of preventing and alleviating the adverse effects of digital health is a major challenge that all HIT stakeholders should be aware of.
  •  

AI-Enhanced Social Robotic Versus Computer-Based Virtual Patients for Clinical Reasoning Training in Medical Education: Observational Crossover Cohort Study

Background: Virtual patient (VP) simulations can be used to practice clinical reasoning (CR) in controlled learning environments. Traditional computer-based VP platforms often lack the authenticity and interactivity required for effective CR training. Artificial intelligence (AI)–enhanced social robotic VPs can enhance realism and engagement; however, quantitative evidence comparing them with conventional VP platforms remains limited. Objective: We compared medical students’ experience of an AI-enhanced social robotic versus a conventional computer-based VP platform regarding the extent to which the design characteristics of the respective platform facilitate CR skill training. Methods: This observational crossover cohort study involved 178 sixth-semester medical students at Karolinska Institutet, Stockholm, Sweden (response rate: 42.3%; 178 of 421 invited students; Spring 2024-Spring 2025), who experienced both a large language model–enhanced social robotic VP platform supporting dialogue (social artificial intelligence–enhanced robotic interface [SARI]) and a conventional computer-based VP platform (virtual interactive case [VIC]) during their clinical rotation within rheumatology. Platform order was determined by clinical rotation scheduling. VP design was evaluated using a validated questionnaire across 5 domains: authenticity, professional approach, coaching quality, learning effects, and overall judgment. Students’ CR training preferences were assessed using categorical responses and a Visual Analogue Scale, where a lower score favored SARI and a score of 5 indicated equal preference between platforms. Results: SARI outperformed VIC across all 5 VP design domains. Students rated SARI higher for authenticity (median 4.0, IQR 3.5-4.5 vs 3.0, IQR 2.5-3.5; P<.001 professional approach iqr vs>P<.001 coaching quality iqr vs>P<.001 learning effect iqr vs>P<.001 and overall judgment vs iqr>P<.001 students strongly preferred sari for cr training vs odds ratio ci>P<.001 with visual analogue scale scores confirming this preference iqr>P<.001 preferences were consistent across most subgroups prior vp experience and platform order in the difference was not significant that is students with vs or ci>P=.11) and students first introduced to VIC (55% vs 45%; OR 1.5; 95% CI 0.7-2.9; P=.33). Conclusions: Our findings provide the first quantitative evidence that AI-enhanced social robotic VPs offer superior design characteristics than conventional computer-based platforms for CR training in medical education. These results support the use of AI-driven social robots for VP simulations to better prepare medical students for real clinical encounters, and warrant future research on objective CR skill outcomes and long-term transfer to clinical practice. Unlike previous qualitative studies examining each platform separately, this study provides the first quantitative comparison of design characteristics between AI-enhanced social robotic and conventional computer-based VPs.
  •  

Harnessing Single-Cell RNA-Seq for Computational Drug Repurposing in Cancer Immunotherapy

Pharmaceuticals (Basel). 2025 Nov 20;18(11):1769. doi: 10.3390/ph18111769.

ABSTRACT

Immune checkpoint inhibitors (ICIs) have revolutionized cancer treatment and show notable success in some cancer types such as non-small cell lung cancer, melanoma and colorectal cancers, while they demonstrate relatively low response rate in others, such as esophageal cancers. Due to the heterogeneous nature of the tumor microenvironment and patient-to-patient variability, there remains a need to improve ICI response rates. Combining ICIs with therapies that can overcome resistance is a promising strategy. Compared to de novo drug development, drug repurposing offers a faster and more cost-effective approach to identifying such combination candidates. A variety of computational drug repurposing tools leverage genomics and/or transcriptomic data. As single-cell RNA sequencing (scRNA-seq) technology becomes available, it enables precise targeting of cancer-driving cellular components. In this review, we highlight current computational drug repurposing tools utilizing scRNA-seq data and demonstrate the application of two such tools, scDrug and scDrugPrio, on an esophageal squamous cell carcinoma dataset to identify potential drug candidates for combination with ICI therapy to enhance treatment response. scDrug focuses on predicting tumor cell-specific cytotoxicity, while scDrugPrio prioritizes drugs by reversing gene signatures associated with ICI non-responsiveness across diverse tumor microenvironment cell types. Together, this review underscores the importance of a multi-faceted approach in computational drug repurposing and highlights its potential for identifying drugs that enhance ICI treatment. Future work can expand the application of these strategies to multi-omics and spatial transcriptomics datasets, as well as personalized patient samples, to further refine drug repurposing involving ICI therapy.

PMID:41305010 | PMC:PMC12655618 | DOI:10.3390/ph18111769

  •  

Harnessing Single-Cell RNA-Seq for Computational Drug Repurposing in Cancer Immunotherapy

Pharmaceuticals (Basel). 2025 Nov 20;18(11):1769. doi: 10.3390/ph18111769.

ABSTRACT

Immune checkpoint inhibitors (ICIs) have revolutionized cancer treatment and show notable success in some cancer types such as non-small cell lung cancer, melanoma and colorectal cancers, while they demonstrate relatively low response rate in others, such as esophageal cancers. Due to the heterogeneous nature of the tumor microenvironment and patient-to-patient variability, there remains a need to improve ICI response rates. Combining ICIs with therapies that can overcome resistance is a promising strategy. Compared to de novo drug development, drug repurposing offers a faster and more cost-effective approach to identifying such combination candidates. A variety of computational drug repurposing tools leverage genomics and/or transcriptomic data. As single-cell RNA sequencing (scRNA-seq) technology becomes available, it enables precise targeting of cancer-driving cellular components. In this review, we highlight current computational drug repurposing tools utilizing scRNA-seq data and demonstrate the application of two such tools, scDrug and scDrugPrio, on an esophageal squamous cell carcinoma dataset to identify potential drug candidates for combination with ICI therapy to enhance treatment response. scDrug focuses on predicting tumor cell-specific cytotoxicity, while scDrugPrio prioritizes drugs by reversing gene signatures associated with ICI non-responsiveness across diverse tumor microenvironment cell types. Together, this review underscores the importance of a multi-faceted approach in computational drug repurposing and highlights its potential for identifying drugs that enhance ICI treatment. Future work can expand the application of these strategies to multi-omics and spatial transcriptomics datasets, as well as personalized patient samples, to further refine drug repurposing involving ICI therapy.

PMID:41305010 | DOI:10.3390/ph18111769

  •  

Cognitive bias in LLM reasoning compromises interpretation of clinical oncology notes

arXiv:2511.20680v1 Announce Type: cross Abstract: Despite high performance on clinical benchmarks, large language models may reach correct conclusions through faulty reasoning, a failure mode with safety implications for oncology decision support that is not captured by accuracy-based evaluation. In this two-cohort retrospective study, we developed a hierarchical taxonomy of reasoning errors from GPT-4 chain-of-thought responses to real oncology notes and tested its clinical relevance. Using breast and pancreatic cancer notes from the CORAL dataset, we annotated 600 reasoning traces to define a three-tier taxonomy mapping computational failures to cognitive bias frameworks. We validated the taxonomy on 822 responses from prostate cancer consult notes spanning localized through metastatic disease, simulating extraction, analysis, and clinical recommendation tasks. Reasoning errors occurred in 23 percent of interpretations and dominated overall errors, with confirmation bias and anchoring bias most common. Reasoning failures were associated with guideline-discordant and potentially harmful recommendations, particularly in advanced disease management. Automated evaluators using state-of-the-art language models detected error presence but could not reliably classify subtypes. These findings show that large language models may provide fluent but clinically unsafe recommendations when reasoning is flawed. The taxonomy provides a generalizable framework for evaluating and improving reasoning fidelity before clinical deployment.
  •  
❌