❌

Reading view

Uncertainty Gating for Cost-Aware Explainable Artificial Intelligence

arXiv:2603.29915v1 Announce Type: new Abstract: Post-hoc explanation methods are widely used to interpret black-box predictions, but their generation is often computationally expensive and their reliability is not guaranteed. We propose epistemic uncertainty as a low-cost proxy for explanation reliability: high epistemic uncertainty identifies regions where the decision boundary is poorly defined and where explanations become unstable and unfaithful. This insight enables two complementary use cases: `improving worst-case explanations' (routing samples to cheap or expensive XAI methods based on expected explanation reliability), and `recalling high-quality explanations' (deferring explanation generation for uncertain samples under constrained budget). Across four tabular datasets, five diverse architectures, and four XAI methods, we observe a strong negative correlation between epistemic uncertainty and explanation stability. Further analysis shows that epistemic uncertainty distinguishes not only stable from unstable explanations, but also faithful from unfaithful ones. Experiments on image classification confirm that our findings generalize beyond tabular data.
  •  

Byzantine-Robust and Communication-Efficient Distributed Training: Compressive and Cyclic Gradient Coding

arXiv:2603.28780v1 Announce Type: cross Abstract: In this paper, we study the problem of distributed training (DT) under Byzantine attacks with communication constraints. While prior work has developed various robust aggregation rules at the server to enhance robustness to Byzantine attacks, the existing methods suffer from a critical limitation in that the solution error does not diminish when the local gradients sent by different devices vary considerably, as a result of data heterogeneity among the subsets held by different devices. To overcome this limitation, we propose a novel DT method, cyclic gradient coding-based DT (LAD). In LAD, the server allocates the entire training dataset to the devices before training begins. In each iteration, it assigns computational tasks redundantly to the devices using cyclic gradient coding. Each honest device then computes local gradients on a fixed number of data subsets and encodes the local gradients before transmitting to the server. The server aggregates the coded vectors from the honest devices and the potentially incorrect messages from Byzantine devices using a robust aggregation rule. Leveraging the redundancy of computation across devices, the convergence performance of LAD is analytically characterized, demonstrating improved robustness against Byzantine attacks and significantly lower solution error. Furthermore, we extend LAD to a communication-efficient variant, compressive and cyclic gradient coding-based DT (Com-LAD), which further reduces communication overhead under constrained settings. Numerical results validate the effectiveness of the proposed methods in enhancing both Byzantine resilience and communication efficiency.
  •  

Zero-Shot Coordination in Ad Hoc Teams with Generalized Policy Improvement and Difference Rewards

arXiv:2510.16187v2 Announce Type: replace-cross Abstract: Real-world multi-agent systems may require ad hoc teaming, where an agent must coordinate with other previously unseen teammates to solve a task in a zero-shot manner. Prior work often either selects a pretrained policy based on an inferred model of the new teammates or pretrains a single policy that is robust to potential teammates. Instead, we propose to leverage all pretrained policies in a zero-shot transfer setting. We formalize this problem as an ad hoc multi-agent Markov decision process and present a solution that uses two key ideas, generalized policy improvement and difference rewards, for efficient and effective knowledge transfer between different teams. We empirically demonstrate that our algorithm, Generalized Policy improvement for Ad hoc Teaming (GPAT), successfully enables zero-shot transfer to new teams in three simulated environments: cooperative foraging, predator-prey, and Overcooked. We also demonstrate our algorithm in a real-world multi-robot setting.
  •  

Sample-Efficient Hypergradient Estimation for Decentralized Bi-Level Reinforcement Learning

arXiv:2603.14867v3 Announce Type: replace-cross Abstract: Many strategic decision-making problems, such as environment design for warehouse robots, can be naturally formulated as bi-level reinforcement learning (RL), where a leader agent optimizes its objective while a follower solves a Markov decision process (MDP) conditioned on the leader's decisions. In many situations, a fundamental challenge arises when the leader cannot intervene in the follower's optimization process; it can only observe the optimization outcome. We address this decentralized setting by deriving the hypergradient of the leader's objective, i.e., the gradient of the leader's strategy that accounts for changes in the follower's optimal policy. Unlike prior hypergradient-based methods that require extensive data for repeated state visits or rely on gradient estimators whose complexity can increase substantially with the high-dimensional leader's decision space, we leverage the Boltzmann covariance trick to derive an alternative hypergradient formulation. This enables efficient hypergradient estimation solely from interaction samples, even when the leader's decision space is high-dimensional. Additionally, to our knowledge, this is the first method that enables hypergradient-based optimization for 2-player Markov games in decentralized settings. Experiments highlight the impact of hypergradient updates and demonstrate our method's effectiveness in both discrete and continuous state tasks.
  •  

Investigating the replicability of the social and behavioural sciences

Nature, Published online: 01 April 2026; doi:10.1038/s41586-025-10078-y

A large-scale study on the replicability of claims from social and behavioural science journals reports that about half of the results replicate in the same patterns as the original study.
  •  

Reproducibility and robustness of economics and political science research

Nature, Published online: 01 April 2026; doi:10.1038/s41586-026-10251-x

Robustness checks and reproduction of analyses with existing and updated data based on 110 articles in economics and political science journals with data and code-sharing requirements found high levels of robustness and reproducibility and determined that robustness was not dependent on author characteristics or data availability.
  •  

Systems-Level Analysis of HPAI H5N1 Infection in Ducks: Integrating Transcriptomic, Proteomic, and Phosphoproteomic Data

Int J Mol Sci. 2026 Mar 23;27(6):2884. doi: 10.3390/ijms27062884.

ABSTRACT

Ducks, once considered mere reservoirs, now serve as both victims and amplifiers of persistent highly pathogenic avian influenza (HPAI) virus cycles in wild populations. The molecular pathogenesis of HPAI is shaped by complex, dysregulated molecular networks, necessitating a systems biology approach that integrates computational modeling of host-pathogen interactions. Despite recent advances, a comprehensive understanding of the signaling pathways, molecular mechanisms, and hub genes driving HPAI H5N1 pathogenesis in avian hosts remains incomplete. This study addresses this gap by employing an integrated multi-omics strategy-combining transcriptomic, proteomic, and phosphoproteomic analyses-to map the signaling networks and key host factors involved in HPAI H5N1 infection in duck lung tissue. Our network analysis revealed activation of RIG-I-like receptor, toll-like receptor, NOD-like receptor, NF-κB, and JAK/STAT signaling pathways. Phosphoproteomic profiling independently confirmed the activation of these pathways, supporting the integrated network findings. Key regulatory hub genes identified include STAT1, DDX58 (RIG-I), MYD88, NFKBIA, NFKB1, IRF7, SOCS3, ACTB, TLR4, TLR7, IL-6, CASP1, and CASP8, which form a central hub in duck antiviral immunity. Some of these genes may represent promising targets for therapeutic or vaccine development against avian influenza. Collectively, this work delineates the critical signaling pathways and hub genes underlying HPAI H5N1 pathogenesis in ducks through comprehensive multi-omics integration.

PMID:41898742 | PMC:PMC13026356 | DOI:10.3390/ijms27062884

  •  

Improving Retrieval Augmented Generation for Health Care by Fine-Tuning Clinical Embedding Models: Development and Evaluation Study

Background: Embedding models are critical components of Retrieval Augmented Generation (RAG) systems for retrieving and searching unstructured medical data. However, existing models are predominantly trained on publicly available English datasets, limiting their effectiveness in non-English health care settings. More importantly, these models lack training on real-world clinical documents, leading to inaccurate context retrieval when integrated into RAG systems for health care applications. This gap is particularly pronounced in specialized medical documentation containing domain-specific terminology, abbreviations, and nuanced clinical language. Objective: This retrospective study aimed to develop and validate embedding models specifically trained on real-world clinical documents from multiple medical specialties to improve medical information retrieval (IR) and RAG system performance in both German and English language contexts. Methods: We fine-tuned embedding models, so-called sentence transformers, using the multilingual-e5-large architecture as a foundation. Training data consisted of approximately 11 million question-answer pairs synthetically generated from 400,000 diverse clinical documents from a large German tertiary hospital, spanning 163,840 patients and 282,728 clinical cases between 2018 and 2023. The large language model generated medically relevant questions and corresponding answers for each document. The dataset was additionally pseudonymized and translated into English to aim for broader applicability. Models were evaluated in 2 distinct scenarios: IR using questions with multiple relevant passages, and RAG system performance in both cross-patient and patient-centered contexts. Results: In the IR evaluation, the fine-tuned miracle model achieved a mAP@100 of 0.27, outperforming the multilingual-e5-large baseline (0.14) and state-of-the-art models such as bge-m3 (0.11). In the RAG evaluation, the model demonstrated robust performance comparable with the baseline in the constrained patient-centered scenario (BERTScore F1 0.781 vs 0.778) and showed moderate improvements in the unconstrained cross-patient setting (BLEURT 0.56 vs 0.53). Notably, the model trained on pseudonymized data achieved comparable retrieval performance (mAP@100 0.25) and the highest scores for patient-centered contextual precision (0.93). Performance gains were robust in the German dataset, while the translated English model demonstrated promising results as a proof of concept for cross-lingual transfer. Conclusions: By leveraging a comprehensive real-world dataset spanning multiple medical specialties and using large language models for synthetic question generation, we successfully created and validated domain-specific embedding models. These models can improve medical IR in large-scale search spaces and perform competitively in constrained RAG applications. By publishing the models trained on pseudonymized data, other health care institutions can integrate or adapt these embedding models to their needs. This work establishes a reproducible framework for developing domain-specific clinical embedding models, with the potential to improve data retrieval in medical settings.
  •  

Detecting outliers of pursuit eye movements: a preliminary analysis of autism spectrum disorder

arXiv:2603.22705v2 Announce Type: new Abstract: Background: Autism spectrum disorder (ASD) is characterized by significant clinical and biological heterogeneity. Conventional group-mean analyses of eye movements often mask individual atypicalities, potentially overlooking critical pathological signatures. This study aimed to identify idiosyncratic oculomotor patterns in ASD using an "outlier analysis" of smooth pursuit eye movement (SPEM). Methods: We recorded SPEM during a slow Lissajous pursuit task in 18 adults with ASD and 39 typically developed (TD) individuals. To quantify individual deviations, we derived an "outlier score" based on the Mahalanobis distance. This score was calculated from a feature vector, optimized via Principal Component Analysis (PCA), comprising the temporal lag ($\Delta$t) and the spatial deviation ($\Delta$s). An outlier was statistically defined as a score exceeding $\sqrt{10}$ (approximately 3.16$\sigma$) relative to the TD normative distribution. Results: While the TD group exhibited a low outlier rate of 5.1%, the ASD group demonstrated a significantly higher prevalence of 38.9% (7/18) (binomial P = 0.0034). Furthermore, the mean outlier score was significantly elevated in the ASD group (3.00 $\pm$ 2.62) compared to the TD group (1.52 $\pm$ 0.80; P = 0.002). Notably, these extreme deviations were captured even when conventional mean-based comparisons showed limited sensitivity. Conclusions: Our outlier analysis successfully visualized the high degree of idiosyncratic atypicality in ASD oculomotor control. By shifting the focus from group averages to individual deviations, this approach provides a sensitive metric for capturing the inherent heterogeneity of ASD, offering a potential baseline for identifying clinical subtypes.
  •  

Dynamical Systems Theory Behind a Hierarchical Reasoning Model

arXiv:2603.22871v1 Announce Type: new Abstract: Current large language models (LLMs) primarily rely on linear sequence generation and massive parameter counts, yet they severely struggle with complex algorithmic reasoning. While recent reasoning architectures, such as the Hierarchical Reasoning Model (HRM) and Tiny Recursive Model (TRM), demonstrate that compact recursive networks can tackle these tasks, their training dynamics often lack rigorous mathematical guarantees, leading to instability and representational collapse. We propose the Contraction Mapping Model (CMM), a novel architecture that reformulates discrete recursive reasoning into continuous Neural Ordinary and Stochastic Differential Equations (NODEs/NSDEs). By explicitly enforcing the convergence of the latent phase point to a stable equilibrium state and mitigating feature collapse with a hyperspherical repulsion loss, the CMM provides a mathematically grounded and highly stable reasoning engine. On the Sudoku-Extreme benchmark, a 5M-parameter CMM achieves a state-of-the-art accuracy of 93.7 %, outperforming the 27M-parameter HRM (55.0 %) and 5M-parameter TRM (87.4 %). Remarkably, even when aggressively compressed to an ultra-tiny footprint of just 0.26M parameters, the CMM retains robust predictive power, achieving 85.4 % on Sudoku-Extreme and 82.2 % on the Maze benchmark. These results establish a new frontier for extreme parameter efficiency, proving that mathematically rigorous latent dynamics can effectively replace brute-force scaling in artificial reasoning.
  •  

Genomic history of early dogs in Europe

Nature, Published online: 25 March 2026; doi:10.1038/s41586-026-10112-7

Genome-wide analysis shows European dogs existed by 14,200 years ago, were already genetically distinct, received less Neolithic Southwest Asian admixture than humans did and contributed substantially to later European dogs.
  •  

Correction: LXRα limits TGFβ-dependent hepatocellular carcinoma associated fibroblast differentiation

Oncogenesis, Published online: 18 March 2026; doi:10.1038/s41389-026-00610-8

Correction: LXRα limits TGFβ-dependent hepatocellular carcinoma associated fibroblast differentiation
  •  

FC-Track: Overlap-Aware Post-Association Correction for Online Multi-Object Tracking

arXiv:2603.12758v1 Announce Type: cross Abstract: Reliable multi-object tracking (MOT) is essential for robotic systems operating in complex and dynamic environments. Despite recent advances in detection and association, online MOT methods remain vulnerable to identity switches caused by frequent occlusions and object overlap, where incorrect associations can propagate over time and degrade tracking reliability. We present a lightweight post-association correction framework (FC-Track) for online MOT that explicitly targets overlap-induced mismatches during inference. The proposed method suppresses unreliable appearance updates under high-overlap conditions using an Intersection over Area (IoA)-based filtering strategy, and locally corrects detection-to-tracklet mismatches through appearance similarity comparison within overlapped tracklet pairs. By preventing short-term mismatches from propagating, our framework effectively mitigates long-term identity switches without resorting to global optimization or re-identification. The framework operates online without global optimization or re-identification, making it suitable for real-time robotic applications. We achieve 81.73 MOTA, 82.81 IDF1, and 66.95 HOTA on the MOT17 test set with a running speed of 5.7 FPS, and 77.52 MOTA, 80.90 IDF1, and 65.67 HOTA on the MOT20 test set with a running speed of 0.6 FPS. Specifically, our framework FC-Track produces only 29.55% long-term identity switches, which is substantially lower than existing online trackers. Meanwhile, our framework maintains state-of-the-art performance on the MOT20 benchmark.
  •  

Team RAS in 10th ABAW Competition: Multimodal Valence and Arousal Estimation Approach

arXiv:2603.13056v1 Announce Type: cross Abstract: Continuous emotion recognition in terms of valence and arousal under in-the-wild (ITW) conditions remains a challenging problem due to large variations in appearance, head pose, illumination, occlusions, and subject-specific patterns of affective expression. We present a multimodal method for valence-arousal estimation ITW. Our method combines three complementary modalities: face, behavior, and audio. The face modality relies on GRADA-based frame-level embeddings and Transformer-based temporal regression. We use Qwen3-VL-4B-Instruct to extract behavior-relevant information from video segments, while Mamba is used to model temporal dynamics across segments. The audio modality relies on WavLM-Large with attention-statistics pooling and includes a cross-modal filtering stage to reduce the influence of unreliable or non-speech segments. To fuse modalities, we explore two fusion strategies: a Directed Cross-Modal Mixture-of-Experts Fusion Strategy that learns interactions between modalities with adaptive weighting, and a Reliability-Aware Audio-Visual Fusion Strategy that combines visual features at the frame-level while using audio as complementary context. The results are reported on the Aff-Wild2 dataset following the 10th Affective Behavior Analysis in-the-Wild (ABAW) challenge protocol. Experiments demonstrate that the proposed multimodal fusion strategy achieves a Concordance Correlation Coefficient (CCC) of 0.658 on the Aff-Wild2 development set.
  •  

Noninvasive biomarkers in thymic epithelial tumors: a systematic review of cfDNA/ctDNA detection, molecular profiling, and organoid-based monitoring

J Thorac Dis. 2026 Feb 28;18(2):171. doi: 10.21037/jtd-2025-1-2467. Epub 2026 Feb 26.

ABSTRACT

BACKGROUND: Thymic epithelial tumors (TETs), including thymomas and thymic carcinomas, are rare malignancies with limited treatment options and no established biomarkers for surveillance. Circulating cell-free DNA (cfDNA) and circulating tumor DNA (ctDNA) provide a non-invasive method for understanding tumor biology, detecting minimal residual disease (MRD), and possibly identifying recurrence. While this approach has added to the management of other solid tumors, its role in TETs remains poorly defined. The objective of this review was to evaluate the feasibility, molecular insights, and clinical utility of cfDNA and ctDNA for diagnosis, molecular profiling, and recurrence monitoring in TETs.

METHODS: This systematic review summarizes the current evidence on cfDNA and ctDNA in TETs. Studies were identifies through systematic searches of PubMed, Embase, Web of Science, MEDLINE, Cochrane Library, and American Society of Clinical Oncology (ASCO) meeting abstracts from inception through July 2025. Eligible studies reported cfDNA or ctDNA analysis in patients with histologically confirmed thymoma or thymic carcinoma, and excluded reviews, commentaries, abstracts without full text, and non-blood based liquid biopsy studies. Data extraction included patient characteristics, assay platforms, mutational findings, and clinical applications. Data were synthesized narratively due to methodological heterogeneity. No formal risk of bias assessment was performed because of the small number of included studies.

RESULTS: Six studies involving 289 patients met inclusion criteria. ctDNA detection was feasible across all studies, with detection rates ranging from 46% to 80%. Recurrent alterations included TP53, CDKN2A/B, KIT, and other variants. Liquid biopsy enabled genomic profiling at diagnosis and dynamic monitoring during treatment. Notably, several studies have suggested that disease recurrence may be detectable through liquid biopsy prior to the appearance of radiographic changes on conventional imaging. Despite these promising observations, evidence remains limited by small sample size, variability in assay methods, and short follow up duration.

CONCLUSIONS: Liquid biopsy approaches based on cfDNA and ctDNA have shown applicability in TETs and provide clinically relevant molecular information in settings where tissue-based analysis is limited. Tumor informed ctDNA strategies show particular promise for postoperative monitoring and longitudinal disease assessment, whereas broader clinical adoption remains investigational. Further prospective, multicenter studies are needed to establish standardized workflows and clarify the role of liquid biopsy across diagnostic, therapeutic, and surveillance contexts in TETs.

PMID:41816481 | PMC:PMC12972770 | DOI:10.21037/jtd-2025-1-2467

  •  

The dynamic basis of G-protein recognition and activation by a GPCR

Nature, Published online: 11 March 2026; doi:10.1038/s41586-026-10228-w

Conventional and time-resolved cryo-electron microscopy reveal how NTSR1 dynamically engages and releases different G proteins, capturing over 20 intermediates and uncovering key mechanistic steps in GDP- and GTP-driven activation, subtype selectivity and distinct dissociation pathways.
  •  
❌