❌

Normal view

GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration

arXiv:2605.24636v2 Announce Type: new Abstract: While large language models (LLMs) hold transformative potential for medicine, their reasoning robustness and safety in real-world clinical scenarios remain critically underexplored, particularly in dentistry. Here we introduce GlobalDentBench, the first multinational dental benchmark, featuring a taxonomy that encompasses 14 dental specialties across 88 countries and regions spanning six continents. The benchmark comprises 8,978 expert-validated questions across three formats (multiple-choice, short-answer, and case-based questions) and assesses three progressive reasoning levels: knowledge recall (L1), routine reasoning (L2), and individualized reasoning (L3). To ensure data quality, the automated construction framework was calibrated by six senior dentists, achieving expert agreement rates of 99.98% for multiple-choice and short-answer questions and 96.78% for the more complex case-based questions. Evaluation of 12 frontier LLMs on GlobalDentBench revealed a sharp, stepwise performance degradation with increasing reasoning complexity. Specifically, accuracy plummeted from 81.34% on multiple-choice to 64.53% on short-answer and 22.34% on case-based questions, while declining markedly from 74.01% at L1 to 55.64% at L2 and 35.71% at L3. More critically, risk analysis of real-world dental cases demonstrated an alarming overall unsafe rate of 31.01% in LLM-generated clinical recommendations, with 4.51% posing risks of irreversible patient harm and risks particularly pronounced in specialties such as orthodontics. These findings expose fundamental limitations in the medical reasoning and safety of current LLMs. Consequently, GlobalDentBench provides a scalable foundation for trustworthy clinical AI evaluation, underscoring the urgent need for rigorous validation before the safe deployment of these models in healthcare.

CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents

arXiv:2605.25624v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has driven breakthroughs in domains such as math, tool-use, and software engineering, yet its extension to computer-use agents (CUAs) has been bottlenecked by the scarcity of scalable training data with deterministic rewards. Constructing such data for CUAs requires consistent task instruction, executable environment, and verifiable reward. However, hand-curated benchmarks achieve high reward fidelity but cover few applications and LLM-as-judge-based datasets scale broadly but lack reliable verification. We present CUA-Gym, a scalable pipeline that co-generates task instructions, environment states, and reward functions. Concretely, a Generator agent constructs the initial and golden environment states, and a separate Discriminator agent writes the reward function from the task specification. An orchestrator agent drives the two through iterative rounds upon execution. Generated tuples then pass a final filter combining LLM majority voting and agent rollouts, ensuring quality beyond the per-task adversarial loop. To address the scarcity of training environments, we further synthesize CUA-Gym-Hub, a broad suite of high-fidelity mock web applications grounded in real-world software-use distributions, expanding the scale of CUA RLVR data by magnitude. Using this pipeline, we construct CUA-Gym, a dataset of 32,112 verified RLVR training tuples grounded in 110 environments. Trained with GSPO on CUA-Gym, our CUA-Gym-A3B and CUA-Gym-A17B achieve 62.1% and 72.6% on OSWorld-Verified, outperforming prior open-source CUAs at comparable scales, with performance scaling smoothly in both data volume and environment diversity. The same checkpoints also improve on the held-out WebArena benchmark, indicating transfer beyond the training environments. We will open-source the full synthesis pipeline, dataset, CUA-Gym-Hub environments, and models.

Small Models, Strong Priors: Architectural Inductive Bias for Parameter-Efficient Neural PDE Solvers

arXiv:2605.25949v1 Announce Type: cross Abstract: Neural PDE solvers have followed the scaling trajectory of vision and language, with recent foundation models reaching billions of parameters. We argue that scale is a poor substitute for architectural inductive bias in this domain: structured priors deliver outsized parameter efficiency, and the pattern of where they succeed and fail is itself informative about what they capture. We instantiate this argument in WaveLiT, an architecture combining a discrete wavelet transform for lossless multi-resolution tokenization, an augmented linear attention block, a shared-weight multiscale feature pyramid, and a wavelet-domain auxiliary loss. Bespoke 1-10M-parameter WaveLiT models compete with foundation models of 100-1000$\times$ their size across eight TheWell benchmarks, with the largest gains on wave and acoustic-dominated benchmarks where the wavelet-multiscale prior fits the dominant dynamical structure and small per-step errors do not compound geometrically under rollout. Trained jointly across all eight benchmarks, a 10M-parameter foundation variant exhibits a structured, physically interpretable transfer pattern -- strongest where the wavelet-multiscale prior matches the dynamics, weakest on chaotic advection-dominated flows. The entire pipeline trains on a single GPU. The results suggest that small-model PDE performance is shaped by architectural inductive bias rather than scale, and that the structure of a prior's failures is a useful empirical signal about its content.

SentGraph: Hierarchical Sentence Graph for Multi-hop Retrieval-Augmented Question Answering

arXiv:2601.03014v3 Announce Type: replace-cross Abstract: Traditional Retrieval-Augmented Generation (RAG) effectively supports single-hop question answering with large language models but faces significant limitations in multi-hop question answering tasks, which require combining evidence from multiple documents. Existing chunk-based retrieval often provides irrelevant and logically incoherent context, leading to incomplete evidence chains and incorrect reasoning during answer generation. To address these challenges, we propose SentGraph, a sentence-level graph-based RAG framework that explicitly models fine-grained logical relationships between sentences for multi-hop question answering. Specifically, we construct a hierarchical sentence graph offline by first adapting Rhetorical Structure Theory to distinguish nucleus and satellite sentences, and then organizing them into topic-level subgraphs with cross-document entity bridges. During online retrieval, SentGraph performs graph-guided evidence selection and path expansion to retrieve fine-grained sentence-level evidence. Extensive experiments on four multi-hop question answering benchmarks demonstrate the effectiveness of SentGraph, validating the importance of explicitly modeling sentence-level logical dependencies for multi-hop reasoning.

Nonlinear atomic tunnelling boosted by bright squeezed vacuum

Nature, Published online: 20 May 2026; doi:10.1038/s41586-026-10485-9

Bright squeezed vacuum light boosts nonlinear atomic tunnelling ionization more than 20-fold compared with coherent light, enabling quantum control of strong-field processes without increasing classical intensity.

Pulmonary-Intestinal Axis: Shared Genetic Basis and Mediating Factors Identified Through Multi-Omics Analysis

Int J Chron Obstruct Pulmon Dis. 2026 Apr 7;21:561645. doi: 10.2147/COPD.S561645. eCollection 2026.

ABSTRACT

BACKGROUND: Chronic obstructive pulmonary disease (COPD) is a systemic condition with comorbidities beyond the lung (eg, cardiovascular and metabolic disorders), and gastrointestinal (GI) disorders are also common. The shared genetic basis of COPD-GI comorbidity and its mediating factors remain unclear. We hypothesized that COPD and GI diseases share pleiotropic genetic architecture implicating lipid-metabolic pathways, with smoking mediating part of the association.

METHODS: We analyzed publicly available European-ancestry GWAS summary statistics for COPD (Global Biobank Meta-analysis Initiative), 15 GI diseases (FinnGen), and smoking phenotypes (UK Biobank). Genetic correlation was estimated using linkage disequilibrium score regression (LDSC) and high-definition likelihood (HDL). Multi-trait analysis of GWAS (MTAG) boosted COPD discovery by leveraging genetically correlated GI traits. We integrated locus-to-gene mapping with multi-tissue expression quantitative trait loci (eQTL) and plasma protein quantitative trait loci (pQTL) evidence to prioritize shared loci, genes, and proteins. Bidirectional two-sample Mendelian randomization (MR) tested causal directions, and two-step mediation MR evaluated smoking.

RESULTS: COPD showed significant genetic correlation with nine GI diseases. We identified six comorbidity-associated loci (three with CADD > 12.37) and 13 unique candidate pleiotropic genes; APOE was supported by proteomic evidence. Enrichment analyses highlighted lipid-metabolism pathways. MR suggested COPD increases risk of gastroesophageal reflux disease (GERD), irritable bowel syndrome (IBS), acute appendicitis, and gastric ulcer, while diverticular disease showed reverse causality toward COPD. Smoking partially mediated the COPD effect on GERD, acute appendicitis, and gastric ulcer.

CONCLUSION: COPD and multiple GI disorders share a distributed pleiotropic genetic basis within the broader systemic comorbidity spectrum of COPD. Multi-omics evidence supports a genomic pulmonary-intestinal axis in which lipid metabolism and smoking-related mechanisms contribute to COPD and GI comorbidity, providing targets for risk stratification and potential intervention.

PMID:41978582 | PMC:PMC13070119 | DOI:10.2147/COPD.S561645

❌