❌

Normal view

MedCausalX: Adaptive Causal Reasoning with Self-Reflection for Trustworthy Medical Vision-Language Models

arXiv:2603.23085v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have enabled interpretable medical diagnosis by integrating visual perception with linguistic reasoning. Yet, existing medical chain-of-thought (CoT) models lack explicit mechanisms to represent and enforce causal reasoning, leaving them vulnerable to spurious correlations and limiting their clinical reliability. We pinpoint three core challenges in medical CoT reasoning: how to adaptively trigger causal correction, construct high-quality causal-spurious contrastive samples, and maintain causal consistency across reasoning trajectories. To address these challenges, we propose MedCausalX, an end-to-end framework explicitly models causal reasoning chains in medical VLMs. We first introduce the CRMed dataset providing fine-grained anatomical annotations, structured causal reasoning chains, and counterfactual variants that guide the learning of causal relationships beyond superficial correlations. Building upon CRMed, MedCausalX employs a two-stage adaptive reflection architecture equipped with $\langle$causal$\rangle$ and $\langle$verify$\rangle$ tokens, enabling the model to autonomously determine when and how to perform causal analysis and verification. Finally, a trajectory-level causal correction objective optimized through error-attributed reinforcement learning refines the reasoning chain, allowing the model to distinguish genuine causal dependencies from shortcut associations. Extensive experiments on multiple benchmarks show that MedCausalX consistently outperforms state-of-the-art methods, improving diagnostic consistency by +5.4 points, reducing hallucination by over 10 points, and attaining top spatial grounding IoU, thereby setting a new standard for causally grounded medical reasoning.
  • ✇cs.AI, q-bio.NC updates on arXiv.org
  • Lie to Me: How Faithful Is Chain-of-Thought Reasoning in Reasoning Models? Richard J. Young
    arXiv:2603.22582v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning has been proposed as a transparency mechanism for large language models in safety-critical deployments, yet its effectiveness depends on faithfulness (whether models accurately verbalize the factors that actually influence their outputs), a property that prior evaluations have examined in only two proprietary models, finding acknowledgment rates as low as 25% for Claude 3.7 Sonnet and 39% for DeepSeek-R1. To exte
     

Lie to Me: How Faithful Is Chain-of-Thought Reasoning in Reasoning Models?

arXiv:2603.22582v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning has been proposed as a transparency mechanism for large language models in safety-critical deployments, yet its effectiveness depends on faithfulness (whether models accurately verbalize the factors that actually influence their outputs), a property that prior evaluations have examined in only two proprietary models, finding acknowledgment rates as low as 25% for Claude 3.7 Sonnet and 39% for DeepSeek-R1. To extend this evaluation across the open-weight ecosystem, this study tests 12 open-weight reasoning models spanning 9 architectural families (7B-685B parameters) on 498 multiple-choice questions from MMLU and GPQA Diamond, injecting six categories of reasoning hints (sycophancy, consistency, visual pattern, metadata, grader hacking, and unethical information) and measuring the rate at which models acknowledge hint influence in their CoT when hints successfully alter answers. Across 41,832 inference runs, overall faithfulness rates range from 39.7% (Seed-1.6-Flash) to 89.9% (DeepSeek-V3.2-Speciale) across model families, with consistency hints (35.5%) and sycophancy hints (53.9%) exhibiting the lowest acknowledgment rates. Training methodology and model family predict faithfulness more strongly than parameter count, and keyword-based analysis reveals a striking gap between thinking-token acknowledgment (approximately 87.5%) and answer-text acknowledgment (approximately 28.6%), suggesting that models internally recognize hint influence but systematically suppress this acknowledgment in their outputs. These findings carry direct implications for the viability of CoT monitoring as a safety mechanism and suggest that faithfulness is not a fixed property of reasoning models but varies systematically with architecture, training method, and the nature of the influencing cue.

Neural ODE and SDE Models for Adaptation and Planning in Model-Based Reinforcement Learning

arXiv:2603.23245v1 Announce Type: cross Abstract: We investigate neural ordinary and stochastic differential equations (neural ODEs and SDEs) to model stochastic dynamics in fully and partially observed environments within a model-based reinforcement learning (RL) framework. Through a sequence of simulations, we show that neural SDEs more effectively capture the inherent stochasticity of transition dynamics, enabling high-performing policies with improved sample efficiency in challenging scenarios. We leverage neural ODEs and SDEs for efficient policy adaptation to changes in environment dynamics via inverse models, requiring only limited interactions with the new environment. To address partial observability, we introduce a latent SDE model that combines an ODE with a GAN-trained stochastic component in latent space. Policies derived from this model provide a strong baseline, outperforming or matching general model-based and model-free approaches across stochastic continuous-control benchmarks. This work demonstrates the applicability of action-conditional latent SDEs for RL planning in environments with stochastic transitions. Our code is available at: https://github.com/ChaoHan-UoS/NeuralRL

Contrastive Metric Learning for Point Cloud Segmentation in Highly Granular Detectors

arXiv:2603.23356v1 Announce Type: cross Abstract: We propose a novel clustering approach for point-cloud segmentation based on supervised contrastive metric learning (CML). Rather than predicting cluster assignments or object-centric variables, the method learns a latent representation in which points belonging to the same object are embedded nearby while unrelated points are separated. Clusters are then reconstructed using a density-based readout in the learned metric space, decoupling representation learning from cluster formation and enabling flexible inference. The approach is evaluated on simulated data from a highly granular calorimeter, where the task is to separate highly overlapping particle showers represented as sets of calorimeter hits. A direct comparison with object condensation (OC) is performed using identical graph neural network backbones and equal latent dimensionality, isolating the effect of the learning objective. The CML method produces a more stable and separable embedding geometry for both electromagnetic and hadronic particle showers, leading to improved local neighbourhood consistency, a more reliable separation of overlapping showers, and better generalization when extrapolating to unseen multiplicities and energies. This translates directly into higher reconstruction efficiency and purity, particularly in high-multiplicity regimes, as well as improved energy resolution. In mixed-particle environments, CML maintains strong performance, suggesting robust learning of the shower topology, while OC exhibits significant degradation. These results demonstrate that similarity-based representation learning combined with density-based aggregation is a promising alternative to object-centric approaches for point cloud segmentation in highly granular detectors.

Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation

arXiv:2603.20172v2 Announce Type: replace-cross Abstract: Recent work on chain-of-thought (CoT) faithfulness reports single aggregate numbers (e.g., DeepSeek-R1 acknowledges hints 39% of the time), implying that faithfulness is an objective, measurable property of a model. This paper provides evidence that it is not. Three classifiers (a regex-only detector, a regex-plus-LLM pipeline, and a Claude Sonnet 4 judge) are applied to 10,276 influenced reasoning traces from 12 open-weight models spanning 9 families and 7B to 1T parameters. On identical data, these classifiers produce faithfulness rates of 74.4%, 82.6%, and 69.7%. Per-model gaps range from 2.6 to 30.6 percentage points; all pairwise McNemar tests are significant (p

Genomic history of early dogs in Europe

Nature, Published online: 25 March 2026; doi:10.1038/s41586-026-10112-7

Genome-wide analysis shows European dogs existed by 14,200 years ago, were already genetically distinct, received less Neolithic Southwest Asian admixture than humans did and contributed substantially to later European dogs.

Exposed phosphatidylserine is an inhibitory molecule in T cell exhaustion

Nature, Published online: 25 March 2026; doi:10.1038/s41586-026-10266-4

Insights into the mechanism by which phosphatidylserine functions as a non-classical inhibitory molecule during T cell exhaustion, and how phosphatidylserine-targeting antibodies enhance T cell responses are explored.

Pembrolizumab and olaparib in homologous-recombination-deficient metastatic pancreatic cancer: the phase 2 POLAR trial

Nature Medicine, Published online: 25 March 2026; doi:10.1038/s41591-026-04299-5

Results of the phase 2 POLAR trial show that biomarker-guided treatment in patients with metastatic pancreatic cancer based on homologous repair deficiency leads to encouraging clinical response rates in immune cell-infiltrated tumors.

Oxygen supply through the tracheolar–muscle system does not constrain insect gigantism

Nature, Published online: 25 March 2026; doi:10.1038/s41586-026-10291-3

New evidence suggests that diffusive oxygen transport through the tracheolar–muscle system is not the limiting factor on insect body size.

Superluminal correlations in ensembles of optical phase singularities

Nature, Published online: 25 March 2026; doi:10.1038/s41586-026-10209-z

Ultrafast electron imaging shows full phase-space dynamics of optical singularities, which can reach superluminal velocities before annihilation and break the particle-like analogy of topological defects.

<i>LRRK2</i>-targeting antisense oligonucleotide in Parkinson’s disease: a phase 1 randomized controlled trial

Nature Medicine, Published online: 24 March 2026; doi:10.1038/s41591-026-04262-4

The first-in-human clinical trial of the LRRK2-targeting antisense oligonucleotide BIIB094 in Parkinson’s disease demonstrates that the treatment is well tolerated and produces dose-dependent reductions in cerebrospinal fluid levels of LRRK2 and phosphorylated Rab10, indicating successful target engagement.

Aryl hydrocarbon receptor is critical for both AR-dependent and AR-indifferent enzalutamide resistance in castration-resistant prostate cancer

Oncogene, Published online: 23 March 2026; doi:10.1038/s41388-026-03723-x

Aryl hydrocarbon receptor is critical for both AR-dependent and AR-indifferent enzalutamide resistance in castration-resistant prostate cancer

Generalist biological artificial intelligence in modeling the language of life

Nature Biotechnology, Published online: 20 March 2026; doi:10.1038/s41587-026-03064-w

This Review discusses the promises and pitfalls of biological AI algorithms and presents a vision for generalist biological artificial intelligence, in which models can perform diverse tasks across biological domains.

Pluripotent stem-cell-based screening uncovers sildenafil as a mitochondrial disease therapy

Leigh syndrome is a severe and untreatable mitochondrial disease. Using patient-derived models in 2D and 3D, Zink and colleagues identify the PDE5 inhibitor sildenafil as a repurposable drug candidate, leading to lifespan extension in mammalian models and clinical improvement in six individuals with Leigh syndrome.

Human-specific features of the cerebellum and ZP2-regulated synapse development

Human-specific transcriptomic and regulatory features are present in the cerebellum, with ZP2 playing a key role in synapse regulation. ZP2 expression is induced by pontine mossy fibers, leading to decreased synaptic proteins and neuronal activity, which provides insights into the evolutionary development of the human cerebellum.
❌