❌

Normal view

Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC

Nature Medicine, Published online: 13 September 2026; doi:10.1038/s41591-026-04488-2

In a large international real-world study of non-small cell lung cancer, a multimodal explainable AI model outperformed established biomarkers for immunotherapy outcome prediction and improved physician decision-making.

Forest carbon protocols underestimate climate-driven carbon loss risks

Nature, Published online: 20 May 2026; doi:10.1038/s41586-026-10571-y

The buffer pool designed to compensate for unintended carbon losses from the largest forest climate mitigation programme in the United States is too small when considering the impact of future climate change scenarios.

An Agent-Based Framework for the Automatic Validation of Mathematical Optimization Models

arXiv:2511.16383v2 Announce Type: replace Abstract: Recently, using Large Language Models (LLMs) to generate optimization models from natural language descriptions has became increasingly popular. However, a major open question is how to validate that the generated models are correct and satisfy the requirements defined in the natural language description. In this work, we propose a novel agent-based method for automatic validation of optimization models that builds upon and extends methods from software testing to address optimization modeling . This method consists of several agents that initially generate a problem-level testing API, then generate tests utilizing this API, and, lastly, generate mutations specific to the optimization model (a well-known software testing technique assessing the fault detection power of the test suite). In this work, we detail this validation method and show, through both theory and experiments, the high quality of validation provided by this agent ensemble in terms of the well-known software testing measure called mutation coverage.

Determinants of the Uptake and Frequency of Use of a Web Portal Digital Health Intervention in Patients With Type 2 Diabetes and/or Coronary Heart Disease: Secondary Analysis of a Randomized Controlled Trial

Background: The targeted application and design of digital health interventions (DHIs) require an understanding of usage determinants. Usage includes uptake (initial use) and frequency (extent of use), but it is unclear whether both components are driven by the same determinants. Objective: This study aimed to examine the determinants of uptake and frequency of use and assess whether they differ. Methods: The investigated DHI was a web portal provided in an intervention for improving disease-related self-management. This study is a secondary analysis of intervention group data from a parallel-group randomized controlled trial. Eligibility criteria were being an adult and being diagnosed with type 2 diabetes and/or coronary heart disease. Sociodemographic, psychological, and health-related variables were examined as determinants. Determinants were analyzed using simple and multiple regression models. Uptake was analyzed using logistic regression, and frequency was analyzed using negative binomial regression with robust SEs. Frequency was analyzed for those who used the DHI at least once. Except for sociodemographic variables, all other variables were standardized to a range from 0 to 1. For simple regression, inflation of the α error due to multiple testing was controlled via the approach of Benjamini and Hochberg, and for multiple regression, it was controlled via the significance of the complete multiple regression model. Results: Of 462 intervention group members, 199 (43.1%) used the web portal at least once. After controlling for inflation of the α error, simple regression for uptake yielded significant effects for higher education (B=0.56, 95% CI 0.18-0.95; =.004), openness (B=1.08, 95% CI 0.33-1.83; =.005), intention regarding physical activity (B=2.28, 95% CI 1.30-3.26;

Human-in-the-Loop LLM Grading for Handwritten Mathematics Assessments

arXiv:2603.13083v1 Announce Type: cross Abstract: Providing timely and individualised feedback on handwritten student work is highly beneficial for learning but difficult to achieve at scale. This challenge has become more pressing as generative AI undermines the reliability of take-home assessments, shifting emphasis toward supervised, in-class evaluation. We present a scalable, end-to-end workflow for LLM-assisted grading of short, pen-and-paper assessments. The workflow spans (1) constructing solution keys, (2) developing detailed rubric-style grading keys used to guide the LLM, and (3) a grading procedure that combines automated scanning and anonymisation, multi-pass LLM scoring, automated consistency checks, and mandatory human verification. We deploy the system in two undergraduate mathematics courses using six low-stakes in-class tests. Empirically, LLM assistance reduces grading time by approximately 23% while achieving agreement comparable to, and in several cases tighter than, fully manual grading. Occasional model errors occur but are effectively contained by the hybrid design. Overall, our results show that carefully embedded human-in-the-loop LLM grading can substantially reduce workload while maintaining fairness and accuracy.

Alignment-Aware and Reliability-Gated Multimodal Fusion for Unmanned Aerial Vehicle Detection Across Heterogeneous Thermal-Visual Sensors

arXiv:2603.08208v1 Announce Type: cross Abstract: Reliable unmanned aerial vehicle (UAV) detection is critical for autonomous airspace monitoring but remains challenging when integrating sensor streams that differ substantially in resolution, perspective, and field of view. Conventional fusion methods-such as wavelet-, Laplacian-, and decision-level approaches-often fail to preserve spatial correspondence across modalities and suffer from annotation of inconsistencies, limiting their robustness in real-world settings. This study introduces two fusion strategies, Registration-aware Guided Image Fusion (RGIF) and Reliability-Gated Modality-Attention Fusion (RGMAF), designed to overcome these limitations. RGIF employs Enhanced Correlation Coefficient (ECC)-based affine registration combined with guided filtering to maintain thermal saliency while enhancing structural detail. RGMAF integrates affine and optical-flow registration with a reliability-weighted attention mechanism that adaptively balances thermal contrast and visual sharpness. Experiments were conducted on the Multi-Sensor and Multi-View Fixed-Wing (MMFW)-UAV dataset comprising 147,417 annotated air-to-air frames collected from infrared, wide-angle, and zoom sensors. Among single-modality detectors, YOLOv10x demonstrated the most stable cross-domain performance and was selected as the detection backbone for evaluating fused imagery. RGIF improved the visual baseline by 2.13% mAP@50 (achieving 97.65%), while RGMAF attained the highest recall of 98.64%. These findings show that registration-aware and reliability-adaptive fusion provides a robust framework for integrating heterogeneous modalities, substantially enhancing UAV detection performance in multimodal environments.
  • ✇cs.AI, q-bio.NC updates on arXiv.org
  • Reinforcement Learning with Symbolic Reward Machines Thomas Krug · Daniel Neider
    arXiv:2603.03068v1 Announce Type: cross Abstract: Reward Machines (RMs) are an established mechanism in Reinforcement Learning (RL) to represent and learn sparse, temporally extended tasks with non-Markovian rewards. RMs rely on high-level information in the form of labels that are emitted by the environment alongside the observation. However, this concept requires manual user input for each environment and task. The user has to create a suitable labeling function that computes the labels. Thes
     

Reinforcement Learning with Symbolic Reward Machines

arXiv:2603.03068v1 Announce Type: cross Abstract: Reward Machines (RMs) are an established mechanism in Reinforcement Learning (RL) to represent and learn sparse, temporally extended tasks with non-Markovian rewards. RMs rely on high-level information in the form of labels that are emitted by the environment alongside the observation. However, this concept requires manual user input for each environment and task. The user has to create a suitable labeling function that computes the labels. These limitations lead to poor applicability in widely adopted RL frameworks. We propose Symbolic Reward Machines (SRMs) together with the learning algorithms QSRM and LSRM to overcome the limitations of RMs. SRMs consume only the standard output of the environment and process the observation directly through guards that are represented by symbolic formulas. In our evaluation, our SRM methods outperform the baseline RL approaches and generate the same results as the existing RM methods. At the same time, our methods adhere to the widely used environment definition and provide interpretable representations of the task to the user.

CASR-Net: An Image Processing-focused Deep Learning-based Coronary Artery Segmentation and Refinement Network for X-ray Coronary Angiogram

arXiv:2510.27315v2 Announce Type: replace-cross Abstract: Early detection of coronary artery disease (CAD) is critical for reducing mortality and improving patient treatment planning. While angiographic image analysis from X-rays is a common and cost-effective method for identifying cardiac abnormalities, including stenotic coronary arteries, poor image quality can significantly impede clinical diagnosis. We present the Coronary Artery Segmentation and Refinement Network (CASR-Net), a three-stage pipeline comprising image preprocessing, segmentation, and refinement. A novel multichannel preprocessing strategy combining CLAHE and an improved Ben Graham method provides incremental gains, increasing Dice Score Coefficient (DSC) by 0.31-0.89% and Intersection over Union (IoU) by 0.40-1.16% compared with using the techniques individually. The core innovation is a segmentation network built on a UNet with a DenseNet121 encoder and a Self-organized Operational Neural Network (Self-ONN) based decoder, which preserves the continuity of narrow and stenotic vessel branches. A final contour refinement module further suppresses false positives. Evaluated with 5-fold cross-validation on a combination of two public datasets that contain both healthy and stenotic arteries, CASR-Net outperformed several state-of-the-art models, achieving an IoU of 61.43%, a DSC of 76.10%, and clDice of 79.36%. These results highlight a robust approach to automated coronary artery segmentation, offering a valuable tool to support clinicians in diagnosis and treatment planning.

GPT-5 vs Other LLMs in Long Short-Context Performance

arXiv:2602.14188v1 Announce Type: cross Abstract: With the significant expansion of the context window in Large Language Models (LLMs), these models are theoretically capable of processing millions of tokens in a single pass. However, research indicates a significant gap between this theoretical capacity and the practical ability of models to robustly utilize information within long contexts, especially in tasks that require a comprehensive understanding of numerous details. This paper evaluates the performance of four state-of-the-art models (Grok-4, GPT-4, Gemini 2.5, and GPT-5) on long short-context tasks. For this purpose, three datasets were used: two supplementary datasets for retrieving culinary recipes and math problems, and a primary dataset of 20K social media posts for depression detection. The results show that as the input volume on the social media dataset exceeds 5K posts (70K tokens), the performance of all models degrades significantly, with accuracy dropping to around 50-53% for 20K posts. Notably, in the GPT-5 model, despite the sharp decline in accuracy, its precision remained high at approximately 95%, a feature that could be highly effective for sensitive applications like depression detection. This research also indicates that the "lost in the middle" problem has been largely resolved in newer models. This study emphasizes the gap between the theoretical capacity and the actual performance of models on complex, high-volume data tasks and highlights the importance of metrics beyond simple accuracy for practical applications.
❌