❌

Reading view

GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration

arXiv:2605.24636v2 Announce Type: new Abstract: While large language models (LLMs) hold transformative potential for medicine, their reasoning robustness and safety in real-world clinical scenarios remain critically underexplored, particularly in dentistry. Here we introduce GlobalDentBench, the first multinational dental benchmark, featuring a taxonomy that encompasses 14 dental specialties across 88 countries and regions spanning six continents. The benchmark comprises 8,978 expert-validated questions across three formats (multiple-choice, short-answer, and case-based questions) and assesses three progressive reasoning levels: knowledge recall (L1), routine reasoning (L2), and individualized reasoning (L3). To ensure data quality, the automated construction framework was calibrated by six senior dentists, achieving expert agreement rates of 99.98% for multiple-choice and short-answer questions and 96.78% for the more complex case-based questions. Evaluation of 12 frontier LLMs on GlobalDentBench revealed a sharp, stepwise performance degradation with increasing reasoning complexity. Specifically, accuracy plummeted from 81.34% on multiple-choice to 64.53% on short-answer and 22.34% on case-based questions, while declining markedly from 74.01% at L1 to 55.64% at L2 and 35.71% at L3. More critically, risk analysis of real-world dental cases demonstrated an alarming overall unsafe rate of 31.01% in LLM-generated clinical recommendations, with 4.51% posing risks of irreversible patient harm and risks particularly pronounced in specialties such as orthodontics. These findings expose fundamental limitations in the medical reasoning and safety of current LLMs. Consequently, GlobalDentBench provides a scalable foundation for trustworthy clinical AI evaluation, underscoring the urgent need for rigorous validation before the safe deployment of these models in healthcare.
  •  

SPA-Cache: Singular Proxies for Adaptive Caching in Diffusion Language Models

arXiv:2602.02544v2 Announce Type: replace-cross Abstract: While Diffusion Language Models (DLMs) offer a flexible, arbitrary-order alternative to the autoregressive paradigm, their non-causal nature precludes standard KV caching, forcing costly hidden state recomputation at every decoding step. Existing DLM caching approaches reduce this cost by selective hidden state updates; however, they are still limited by (i) costly token-wise update identification heuristics and (ii) rigid, uniform budget allocation that fails to account for heterogeneous hidden state dynamics. To address these challenges, we present SPA-Cache that jointly optimizes update identification and budget allocation in DLM cache. First, we derive a low-dimensional singular proxy that enables the identification of update-critical tokens in a low-dimensional subspace, substantially reducing the overhead of update identification. Second, we introduce an adaptive strategy that allocates fewer updates to stable layers without degrading generation quality. Together, these contributions significantly improve the efficiency of DLMs, yielding up to an $8\times$ throughput improvement over vanilla decoding and a $2$--$4\times$ speedup over existing caching baselines.
  •  

eHealth Literacy and Type 2 Diabetes Prevention Among At-Risk Populations: Mechanistic Systematic Review Using Theory-Driven Thematic Analysis

Background: Type 2 diabetes (T2D) is emerging as a growing global public health crisis. Early and effective interventions can reduce T2D incidence among at-risk populations. Compared with traditional approaches, digital health technologies offer promising opportunities for prevention, with eHealth literacy (eHL) emerging as a critical determinant of digital prevention outcomes. Objective: This systematic review aims to synthesize and explain the pathways and mechanisms through which eHL supports T2D prevention among at-risk populations. Methods: We searched Scopus, Web of Science, and PubMed databases for English-language original research published between January 1, 2000, and August 14, 2025. Studies included were prevention research involving eHL engagement among populations at risk for T2D. Nonoriginal literature, such as editorials and abstracts, as well as research protocols, was excluded. The findings were synthesized using a thematic analysis approach, integrating the Theoretical Domains Framework with the eHL model. Two reviewers independently screened literature and extracted data, and discrepancies were resolved by a third reviewer. The Mixed Methods Appraisal Tool was used to assess risk of bias. Results: This review included 28 studies (n=13,100), mostly quantitative and published within the past decade, targeting people with prediabetes, prior gestational diabetes, and overweight/metabolic risk. Study quality was moderate to high (Mixed Methods Appraisal Tool 60%‐100%) with no high risk of bias. eHL supported prevention mainly through knowledge (28/28), behavioral regulation (16/28), social influences (15/28), environmental resources (12/28), and goals (11/28), while emotions, memory, attention, decision process, and beliefs about competence were rarely addressed. Health literacy (27/28), information literacy (20/28), and communicative eHL (20/28) were most common; critical eHL and media literacy were not addressed. Studies reported positive outcomes: high engagement, weight loss (≥5%), improved glycemic markers, and enhanced lifestyle behaviors. Conclusions: This is the first systematic exploration of eHL mechanism pathways in T2D prevention via theoretical mapping. We found interventions yield positive effects despite highly uneven mechanism application: extant research relies excessively on knowledge and behavioral pathways while underemphasizing emotional support, autonomy, and critical evaluation—factors linked to long-term adherence. We provide a mechanism-based framework and identify critical gaps, including the absence of focus on critical eHL and media literacy. This review is limited by substantial variation across studies that did not allow for meta-analysis and by the limited evidence base on eHL. Future interventions should explore and test emotional and autonomy support, information discernment training, and accessibility optimization in T2D prevention. These comprehensive, equity-focused intervention approaches will help ensure that eHL becomes a truly effective public health tool that benefits everyone, especially at-risk and vulnerable populations. Trial Registration: PROSPERO CRD42025630395; https://www.crd.york.ac.uk/PROSPERO/view/CRD42025630395
  •  

Resource-Efficient Personal Large Language Models Fine-Tuning with Collaborative Edge Computing

arXiv:2408.10746v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have unlocked a plethora of powerful applications at the network edge, such as intelligent personal assistants. Data privacy and security concerns have prompted a shift towards edge-based fine-tuning of personal LLMs, away from cloud reliance. However, this raises issues of computational intensity and resource scarcity, hindering training efficiency and feasibility. While current studies investigate parameter-efficient fine-tuning (PEFT) techniques to mitigate resource constraints, our analysis indicates that these techniques are not sufficiently resource-efficient for edge devices. To tackle these challenges, we propose Pluto and Charon (PAC), a time and memory efficient collaborative edge AI framework for personal LLMs fine-tuning. PAC breaks the resource wall of personal LLMs fine-tuning with a sophisticated algorithm-system co-design. (1) Algorithmically, PAC implements a personal LLMs fine-tuning technique that is efficient in terms of parameters, time, and memory. It utilizes Parallel Adapters to circumvent the need for a full backward pass through the LLM backbone. Additionally, an activation cache mechanism further streamlining the process by negating the necessity for repeated forward passes across multiple epochs. (2) Systematically, PAC leverages edge devices in close proximity, pooling them as a collective resource for in-situ personal LLMs fine-tuning, utilizing a hybrid data and pipeline parallelism to orchestrate distributed training. The use of the activation cache eliminates the need for forward pass through the LLM backbone,enabling exclusive fine-tuning of the Parallel Adapters using data parallelism. Extensive evaluation based on prototype implementation demonstrates that PAC remarkably outperforms state-of-the-art approaches, achieving up to 8.64x end-to-end speedup and up to 88.16% reduction in memory footprint.
  •  
❌