❌

Normal view

Integration of Digital Therapeutics Into Occupational Rehabilitation in Germany: Multilevel Simulation Study

Background: Expenditures for physiotherapy and extended outpatient physiotherapy (EAP) are increasing within Germany’s statutory accident insurance system (Berufsgenossenschaften), placing growing pressure on rehabilitation capacity and timely access to care. Digital health applications (DiGAs) are reimbursable nationwide and represent a novel component of routine rehabilitation pathways. However, their real-world system-level and economic effects in occupational rehabilitation remain insufficiently understood. Objective: This study aimed to evaluate how the integration of DiGAs into occupational rehabilitation pathways may influence costs, service capacity, and waiting times within routine care delivered by 5 German statutory accident insurance funds that cover 25.9 million insured individuals. Methods: Aggregated administrative data from 5 Berufsgenossenschaften (fiscal years 2023‐2024) were analyzed using a multilevel simulation framework combining (1) probabilistic cost-consequence modeling with Monte Carlo simulation (10,000 iterations), (2) an adherence-based adoption funnel distinguishing long-term engaged users (15%) and short-term users (85%) based on German claims data, and (3) a calibrated M/M/1 queuing model validated through discrete event simulation to estimate the effects on waiting times and system capacity. Primary outcomes included net financial impact, break-even thresholds, and changes in access-related performance metrics. Results: Combined physiotherapy and EAP expenditures reached €404 million (€1=US $1.18) in 2024, increasing by 10.1% year-over-year. The primary simulation (N=10,000 iterations) indicated mean annual net savings of €18.4 million (median €17.9 million) with a 90.7% probability of cost savings (95% uncertainty range: net cost of €8 million to net savings of €47.7 million). After incorporating adherence dynamics, the projected mean net savings were €16.2 million (95% CI €5-€29.8 million), corresponding to a 100% probability of positive financial impact within the modeled parameter space. Cost neutrality was maintained for DiGA prices up to €617.8 per prescription, nearly 40% above the base-case assumption of €450, indicating substantial economic robustness. Queuing analyses demonstrated that modest reductions in therapeutic demand decreased mean waiting times from 17.3 to 12.8 days (−26%), equivalent to approximately 120,000 cumulative patient waiting days saved annually across 26,705 EAP patients. The validation of discrete event simulation confirmed the magnitude and direction of analytic estimates. Conclusions: Under conservative assumptions, integrating digital therapeutics into occupational rehabilitation pathways is likely to generate both economic benefits and substantial system-level capacity gains. The break-even threshold of €617.80 per prescription provides a wide margin for pricing policy. Beyond cost effects, DiGAs may function as scalable capacity tools that alleviate systemic bottlenecks and improve timely access to rehabilitation services in capacity-constrained systems.

What do early data from Utah’s Doctronic AI pilot show?

26 May 2026 at 21:58

You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday.

Good morning health tech readers!

Read the rest…

© Adobe

PALoRA: Projection-Adaptive LoRA for Preserving Reasoning in Large Language Models

arXiv:2605.24549v1 Announce Type: new Abstract: Efficiently updating Large Language Models (LLMs) with new or evolving factual knowledge remains a central challenge, as even parameter-efficient adaptation can erode previously acquired reasoning abilities. This tension reflects a plasticity-stability dilemma: models must incorporate new knowledge while preserving skill-critical representations. In this work, we study this trade-off through the spectral structure of multilayer perceptron weight matrices. We show, both theoretically and empirically, that information essential for reasoning is not localized only in dominant singular directions, but is instead distributed across the singular spectrum. Motivated by this observation, we introduce PALoRA, a two-stage framework for knowledge injection with reduced interference. PALoRA first trains a Singular Value Fine-Tuning (SVF) expert on a reasoning dataset and uses its learned singular scaling vector as a frozen geometric probe to identify components that are critical for the target skill. It then performs factual knowledge injection with Low-Rank Adaptation (LoRA) under a structural orthogonality constraint, ensuring that updates avoid the identified skill-relevant subspace. Across Llama 3.1 8B and Mistral 7B, and across mathematical, coding, and scientific reasoning benchmarks, PALoRA preserves on average 95% of the SVF expert's reasoning performance while maintaining competitive factual recall. It consistently improves skill retention over prior spectral Parameter-Efficient Fine-Tuning (PEFT) methods while adding less than 0.006% parameter overhead.

When Can We Trust Early Warnings? Leakage-Excluded Early Outcome Prediction from LMS Interaction Logs

arXiv:2605.25794v1 Announce Type: new Abstract: Early-warning models built from Learning Management System (LMS) logs aim to predict end-of-course outcomes early enough to enable timely learner support. However, reported "early" performance is often inflated by temporal leakage. This occurs when the pipeline uses information that would not yet be available at the time of prediction. We formalize cutoff-based early outcome prediction under a temporal availability constraint and introduce LEAP (Leakage-Excluded Early-Availability Protocol), which enforces cutoff-first truncation prior to joins and aggregation and audits feature provenance to prevent post-cutoff evidence from entering the benchmark. We instantiate LEAP on the public Open University Learning Analytics Dataset (OULAD) as a multi-step protocol for leakage-controlled evaluation across weekly cutoffs. Using several standard learning methods, we evaluate performance using ROC-AUC, PR-AUC, Brier score, and F1@0.5. Results show improving performance as the observation window expands, with a marked gain around week~3; Random Forest performs best at the earliest cutoffs, while Gradient Boosting dominates thereafter. Leakage ablations further show that temporal violations, especially through assessment information, can inflate apparent "early" performance.

Teaching Through Analogies: A Modular Pipeline for Educational Analogy Generation

arXiv:2605.24211v1 Announce Type: cross Abstract: Analogies help learners understand unfamiliar concepts by relating them to known concepts. Despite recent advances, large language models (LLMs) continue to struggle to generate analogies of comparable quality to those produced by humans. We present a modular pipeline for educational analogy generation, decomposing the task into four stages: source finding, sub-concept generation, explanation generation, and evaluation. Grounded in Structure Mapping Theory, the pipeline enables systematic, stage-by-stage analysis of how model choice and input configuration affect analogy quality. We evaluate 12 state-of-the-art LLMs across six model families on two datasets with structured sub-concept annotations (SCAR and ParallelPARC), alongside seven embedding models for closed-setting retrieval. Our results show that sub-concepts substantially improve explanation quality and closed setting retrieval precision but provide limited benefit in open-ended source generation. We further introduce an LLM-as-a-judge evaluation methodology and validate its scoring against human annotations from seven annotators, finding that Claude Sonnet 4.6 aligns more reliably with human rankings than with fine-grained absolute scores. Taken together, our findings reveal cross-stage interactions that isolated studies cannot capture, and highlight sub-concept grounding as a key driver of analogy quality generation.

TRAFA: Anticipating User Actions to Reduce Errors in Procedural Tasks with Predictive Feedback

arXiv:2605.24526v1 Announce Type: cross Abstract: Interactive assistance systems typically provide feedback after an action has been completed, supporting error recovery but not preventing the error itself. We present TRAFA, a real-time predictive feedback system for procedural tasks that intervenes before errors are committed. TRAFA operationalizes predictive feedback through a Track-Forecast-Act framework that tracks hand and object state, forecasts user motion conditioned on scene context, and triggers feedback when a predicted action is likely to violate task constraints. We instantiate this pipeline in a sequential assembly setting and evaluate it through both technical benchmarking and a controlled user study against conventional reactive feedback. Our results show that predictive feedback improves task accuracy and efficiency while maintaining a comparable number of feedback events. These findings position feedback timing as a key dimension in system design and show how real-time anticipation can be integrated into interactive systems to prevent errors before they occur.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering

arXiv:2605.24703v1 Announce Type: cross Abstract: Large language models (LLMs) and time-series language models (TSLMs) are increasingly applied to time-series question answering (TSQA). Unlike text-only QA, TSQA requires models to ground answers in temporal signals whose patterns may occur at different scales, specific time locations, or across separated intervals. However, existing benchmarks are typically organized by task types or high-level reasoning categories, making it difficult to diagnose the underlying signal-level capabilities driving model performance. We introduce TS-Skill, a controlled benchmark for evaluating three composable analytical skills in TSQA: temporal scale selection (SK1), temporal localization (SK2), and cross-interval integration (SK3). TS-Skill provides timestamp-aware questions, broad domain coverage, and human-validated QA quality. To construct the benchmark at scale, we develop SKEvol, a skill-guided agentic framework that combines domain-aware time-series seed generation, skill-controlled question generation, metadata- and code-assisted answer construction, multi-phase signal-grounded verification, and human-in-the-loop curation. Experiments on ten state-of-the-art LLMs and TSLMs reveal substantial and uneven capability gaps across SK1-SK3. In particular, SK3 remains consistently challenging for non-agent models, whereas tool-augmented agents show a selective advantage on standalone SK3. These findings demonstrate that skill-level evaluation can uncover temporal reasoning failures that are obscured by aggregate TSQA scores.

Everything at Every Scale: Scale-Invariant Diffusion with Continuous Super-Resolution

arXiv:2605.26032v1 Announce Type: cross Abstract: Creating images from noise is image generation; reconstructing fine details from coarse inputs is super-resolution. Despite their practical differences, both can be understood as reversing information loss across scales. We introduce $\textbf{SKILD}$, a $\textbf{S}$cale-invariant $\textbf{K}$-Space $\textbf{I}$mage $\textbf{L}$earning $\textbf{D}$iffusion model that unifies generation and continuous super-resolution within a single unconditional framework. Both natural images and critical physical systems exhibit scale invariance, and we leverage it to design a forward process that attenuates image content from fine to coarse scales while injecting spectrum-matched Gaussian noise, making scale an explicit coordinate of the diffusion dynamics. The same trained reverse process performs generation and continuous super-resolution by varying only the starting timestep: $\textit{no task-specific architecture, no conditioning branch, no classifier-free guidance, no retraining per scale factor}$. Empirically, SKILD reaches FID $2.65$ and Inception Score $9.63$ on unconditional CIFAR-10, performs $2\times$--$8\times$ super-resolution on ImageNet from a single unconditional checkpoint while outperforming conditional models across perceptual metrics, and reconstructs critical Ising models whose connected four-point correlations closely track the ground truth.

Teaching large language models to reason like expert diagnosticians

arXiv:2509.12194v2 Announce Type: replace Abstract: Differential diagnosis is an iterative process that integrates patient information with broader medical knowledge. Clinical case series such as the NEJM Clinicopathologic Conferences (CPCs), published continuously since 1923, feature expert physicians who demonstrate diagnostic reasoning to peers, and have been used for decades to evaluate AI. However, prior AI evaluations have largely focused on final diagnostic accuracy rather than nuanced clinical reasoning. Here, we introduce Dr. CaBot, an agentic AI system that emulates an expert diagnostician by generating written and narrated slide-based presentations from an initial case description alone. CaBot recently generated the first AI diagnosis published in the 100+ year history of the NEJM CPCs. In blinded evaluations, physicians misclassified the source of the differential (CaBot vs. physician-written) in 46/62 (74%) of trials and rated them favorably across quality dimensions. When tasked with solving cases for 72 patients with undiagnosed disease from the NIH Undiagnosed Diseases Network, CaBot identified the working diagnosis in 50/72 (69%) of cases from referral notes alone. To promote transparency and research, we also developed CPC-Bench, a physician-validated benchmark based on 7,102 CPCs and 47,648 questions across 10 tasks. We show that CaBot outperforms frontier models on CPC-Bench, and release both CaBot and CPC-Bench publicly to foster progress in clinical AI.

JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures

arXiv:2602.17162v2 Announce Type: replace Abstract: Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature". While effective at capturing local syntax, these generative paradigms prioritize token-level reconstruction over high-level functional context. We introduce JEPA-DNA, a model-agnostic continual training framework that integrates a Joint-Embedding Predictive Architecture (JEPA) with traditional generative objectives. By supervising global sequence embeddings in a latent space, JEPA-DNA forces models to predict the functional representations of masked genomic segments, shifting the learning signal from token recovery to semantic alignment. We evaluate JEPA-DNA on 17 diverse genomic benchmark tasks, demonstrating consistent gains in linear probing and zero-shot performance regardless of the underlying GFM architecture or generative objective. Our framework establishes a new state-of-the-art for GFMs, surpassing the best existing models by bridging generative precision with latent semantic grounding. Through extensive ablation studies, we further characterize the synergistic interplay between generative and latent objectives. Our code is publicly available at https://github.com/NVIDIA-Digital-Bio/JEPA-DNA.

Reprogramming temozolomide response in glioblastoma through regulated and immunogenic cell death modalities

Cell Death Discovery, Published online: 26 May 2026; doi:10.1038/s41420-026-03151-6

Reprogramming temozolomide response in glioblastoma through regulated and immunogenic cell death modalities

Dual ctDNA and CTCs analysis for minimal residual disease detection and relapse monitoring in early-stage breast cancer

NPJ Breast Cancer. 2026 May 20. doi: 10.1038/s41523-026-00970-9. Online ahead of print.

ABSTRACT

Early detection of minimal residual disease (MRD) by liquid biopsy could enable earlier relapse identification in early-stage breast cancer (BC), but most approaches are single-analyte. We prospectively studied 58 patients with early-stage BC, including HR+/HER2-, triple-negative, and HER2+ subtypes, treated with neoadjuvant chemotherapy or primary surgery plus adjuvant therapy. Patient-specific droplet digital PCR assays were designed for both ctDNA and CTCs analysis, with one truncal tumour mutation tracked per patient for detection in both analytes, in serial high-volume blood samples collected at diagnosis before treatment, approximately 1 month after surgery, and at 6-month intervals during follow-up in patients at high risk of relapse. We developed a composite Liquid biopsy Minimal Residual Disease (L-MRD) score integrating six clinical and molecular variables. At baseline (pre-treatment), ctDNA and/or CTCs were detected in 67.2% of patients and remained positive post-surgery in 53.6%. During follow-up, MRD detection anticipated all relapses (100% sensitivity; median lead time 27.95 months) and the L-MRD score stratified 5-year recurrence risk (2.2% vs 41.6%; HR = 6.6; P = 0.04) with strong performance (AUC = 0.891; NPV = 97.8%). This dual-analyte, longitudinal MRD strategy improves relapse prediction and supports further research regarding the clinical utility of personalised surveillance in early-stage BC.

PMID:42161983 | DOI:10.1038/s41523-026-00970-9

❌