❌

Reading view

Performance of AI-Based Screening Tools for Obstructive Sleep Apnea Across Apnea-Hypopnea Index Thresholds: Systematic Review and Meta-Analysis

Background: Obstructive sleep apnea (OSA) is highly prevalent but remains substantially underdiagnosed. Polysomnography (PSG) is the reference standard, but its cost and limited availability constrain large-scale case identification. AI-based screening tools may support risk stratification and referral prioritization, but their diagnostic accuracy across apnea-hypopnea index (AHI) thresholds remains uncertain. Objective: This review aimed to systematically evaluate the diagnostic accuracy of AI-based OSA screening tools at AHI thresholds of ≥5, ≥15, and ≥30 events/hour, with emphasis on models using non-PSG–derived inputs. Methods: PubMed, Embase, Scopus, and Web of Science were searched for studies published from January 1, 2016, to May 3, 2026. Eligible studies included adults evaluated for suspected OSA or recruited from population-based cohorts, assessed AI-based models intended or interpretable for OSA screening, risk prediction, or screening-oriented severity classification, used PSG as the reference standard, and reported sufficient data to construct or reconstruct 2×2 contingency tables. Diagnostic accuracy was synthesized separately by AHI threshold and input source using bivariate random-effects models, with 95% CIs and prediction intervals (PIs). Risk of bias and certainty of evidence were assessed using QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies 2) and GRADE (Grading of Recommendations Assessment, Development, and Evaluation), respectively. Results: A total of 60 studies were included, of which 47 contributed data to the meta-analysis. At AHI thresholds of ≥5, ≥15, and ≥30 events/hour, pooled sensitivities were 0.94 (95% CI 0.92‐0.96; 95% PI 0.71‐0.99), 0.87 (95% CI 0.84‐0.89; 95% PI 0.66‐0.96), and 0.83 (95% CI 0.79‐0.87; 95% PI 0.61‐0.94), respectively; the corresponding specificities were 0.77 (95% CI 0.69‐0.84; 95% PI 0.30‐0.96), 0.81 (95% CI 0.75‐0.85; 95% PI 0.39‐0.96), and 0.91 (95% CI 0.87‐0.94; 95% PI 0.55‐0.99), respectively. The corresponding areas under the summary receiver operating characteristic curves were 0.943, 0.907, and 0.920. For non-PSG–derived tools, sensitivities were 0.92, 0.85, and 0.81, and specificities were 0.70, 0.74, and 0.85 at the 3 thresholds, respectively. For PSG-derived models, sensitivities were 0.96, 0.90, and 0.85, and specificities were 0.82, 0.88, and 0.96, respectively. Exploratory subgroup analyses suggested performance variation across selected study and model characteristics, including region, algorithmic framework, data source, and validation method. Conclusions: AI-based tools showed generally favorable screening performance for OSA across clinically relevant AHI thresholds, although wide PIs suggest variable performance across future comparable populations and settings. By synthesizing diagnostic accuracy across 3 AHI thresholds and distinguishing non-PSG–derived from PSG-derived models, this review extends previous broad or modality-specific reviews and offers a clinically interpretable, pathway-specific basis for linking model performance to intended use. The findings may clarify potential roles for non-PSG–derived tools in front-end screening and referral prioritization and for PSG-derived models in reduced-channel assessment and sleep-laboratory workflow support. Given substantial heterogeneity, limited external validation, and low or very low certainty of evidence, prospective validation is needed before routine implementation.
  •  

Low-Cost Labels, Reliable Choices: Rollout-Calibrated Hyper-Heuristics for Job Shop Scheduling

arXiv:2605.23957v1 Announce Type: new Abstract: Learning-assisted hyper-heuristics can select among dispatching rules while preserving the feasibility and interpretability of constructive Job Shop Scheduling Problem (JSSP) heuristics. Their main computational cost lies in label generation rather than model fitting, since each supervised label usually requires rolling out candidate rules from a partial schedule. We study this label-cost problem together with a reliability problem: a learned selector should not switch away from a strong default rule unless the predicted gain is credible. The proposed selector uses regret-normalized rollout labels, a contextual KNN uncertainty estimate, and a gate that acts only when the predicted improvement exceeds an uncertainty-adjusted margin. We also vary rollout depth and breadth to measure the cost-quality trade-off. On synthetic JSSP instances, the gated selector achieves the lowest mean RPD among learned selectors, remains close to the best fixed dispatching rule, and reduces Random-HH mean RPD by more than an order of magnitude.
  •  

D3S2: Diffusion-Guided Dataset Distillation for Semantic Segmentation

arXiv:2605.25022v1 Announce Type: cross Abstract: Dataset distillation (DD) aims to compress large-scale datasets into compact synthetic sets while preserving training efficacy. However, existing studies mainly focus on image classification, leaving dense prediction tasks such as semantic segmentation largely underexplored. In this work, we identify three key challenges for segmentation DD: (i) long-tailed class imbalance, (ii) the need for strict pixel-wise alignment between images and dense labels, and (iii) the high computational cost of optimizing high-resolution data with complex models. To address these challenges, we propose D3S2, a Diffusion-guided Dataset Distillation framework for Semantic Segmentation. Our method adopts a two-stage design. In Class-Balanced Mask Selection, we construct a representative mask set via a greedy strategy that prioritizes underrepresented classes. In Diffusion-Guided Image Synthesis, we employ a pretrained layout-to-image diffusion model to generate images conditioned on the selected masks, naturally ensuring spatial alignment. To further enhance the training utility of synthesized data, we introduce guided diffusion sampling with two complementary objectives: a segmentation-consistency loss for pixel-level alignment, and a class-wise feature matching loss for aligning per-class feature statistics across layers. Extensive experiments demonstrate the superiority of D3S2. Notably, at an extremely compression rate of 1%, our method achieves 24.99% and 35.49% mIoU on ADE20K and COCO-Stuff with Mask2Former (Swin-S), outperforming random selection by 9.34% and 5.70%, respectively.
  •  
  •  

Spatial multi-omics unveils the monoclonal origin, neuroendocrine plasticity, and microenvironment niches in combined small-cell lung cancer

Cell Rep Med. 2026 Apr 10:102741. doi: 10.1016/j.xcrm.2026.102741. Online ahead of print.

ABSTRACT

Combined small-cell lung cancer (cSCLC) is an aggressive subtype of SCLC with mixed histologic components. Despite heterogeneity and poorer prognosis than de novo SCLC, cSCLC is managed as SCLC because molecular insight into biology, lineage plasticity, and tumor microenvironment (TME) is limited. We perform spatial whole-exome sequencing, spatial transcriptomics, and single-nucleus RNA sequencing across 19 treatment-naive cSCLC tumors. Different histologic components share a monoclonal origin, whereas divergence associates with distinct mutation and copy-number alteration patterns. Our results define spatially exclusive or interspersed tumor domains with distinct TME and immune landscapes; fibroblast-rich boundaries enriched for an aggressive fibroblast subtype may shape TME and treatment responses. We identify lineage plasticity, including adenocarcinoma-to-SCLC transdifferentiation and SCLC-subtype coexistence, and develop cSCLC Detector, a sensitive mutation-based assay improving cSCLC detection in tissue and liquid biopsies. These findings illuminate cSCLC evolution and heterogeneity, underscoring the need for tailored diagnostic and therapeutic strategies for this aggressive subtype.

PMID:41966692 | DOI:10.1016/j.xcrm.2026.102741

  •  

Spatial multi-omics unveils the monoclonal origin, neuroendocrine plasticity, and microenvironment niches in combined small-cell lung cancer

Cell Rep Med. 2026 Apr 10:102741. doi: 10.1016/j.xcrm.2026.102741. Online ahead of print.

ABSTRACT

Combined small-cell lung cancer (cSCLC) is an aggressive subtype of SCLC with mixed histologic components. Despite heterogeneity and poorer prognosis than de novo SCLC, cSCLC is managed as SCLC because molecular insight into biology, lineage plasticity, and tumor microenvironment (TME) is limited. We perform spatial whole-exome sequencing, spatial transcriptomics, and single-nucleus RNA sequencing across 19 treatment-naive cSCLC tumors. Different histologic components share a monoclonal origin, whereas divergence associates with distinct mutation and copy-number alteration patterns. Our results define spatially exclusive or interspersed tumor domains with distinct TME and immune landscapes; fibroblast-rich boundaries enriched for an aggressive fibroblast subtype may shape TME and treatment responses. We identify lineage plasticity, including adenocarcinoma-to-SCLC transdifferentiation and SCLC-subtype coexistence, and develop cSCLC Detector, a sensitive mutation-based assay improving cSCLC detection in tissue and liquid biopsies. These findings illuminate cSCLC evolution and heterogeneity, underscoring the need for tailored diagnostic and therapeutic strategies for this aggressive subtype.

PMID:41966692 | DOI:10.1016/j.xcrm.2026.102741

  •  

Countering Catastrophic Forgetting of Large Language Models for Better Instruction Following via Weight-Space Model Merging

arXiv:2604.01538v1 Announce Type: cross Abstract: Large language models have been adopted in the medical domain for clinical documentation to reduce clinician burden. However, studies have reported that LLMs often "forget" a significant amount of instruction-following ability when fine-tuned using a task-specific medical dataset, a critical challenge in adopting general-purpose LLMs for clinical applications. This study presents a model merging framework to efficiently adapt general-purpose LLMs to the medical domain by countering this forgetting issue. By merging a clinical foundation model (GatorTronLlama) with a general instruct model (Llama-3.1-8B-Instruct) via interpolation-based merge methods, we seek to derive a domain-adapted model with strong performance on clinical tasks while retaining instruction-following ability. Comprehensive evaluation across medical benchmarks and five clinical generation tasks (e.g., radiology and discharge summarization) shows that merged models can effectively mitigate catastrophic forgetting, preserve clinical domain expertise, and retain instruction-following ability. In addition, our model merging strategies demonstrate training efficiency, achieving performance on par with fully fine-tuned baselines under severely constrained supervision (e.g., 64-shot vs. 256-shot). Consequently, weight-space merging constitutes a highly scalable solution for adapting open-source LLMs to clinical applications, facilitating broader deployment in resource-constrained healthcare environments.
  •  

Inactivating <i>SnRK1β1A</i> promotes broad-spectrum disease resistance in rice

Nature, Published online: 25 March 2026; doi:10.1038/s41586-026-10273-5

SnRK1β1A in rice promotes susceptibility to multiple fungal diseases, and disrupting this infection-inducible gene confers broad-spectrum resistance without compromising growth or yield under normal field conditions.
  •  

Androgen activity in the male embryonic hindbrain drives lethal PFA ependymoma

Nature, Published online: 25 March 2026; doi:10.1038/s41586-026-10264-6

Androgen activity in the male embryonic hindbrain prolongs hindbrain differentiation in male individuals and drives sex differences in the incidence and prognosis of posterior fossa type A (PFA) ependymoma, an aggressive childhood brain tumour.
  •  

ARL-Tangram: Unleash the Resource Efficiency in Agentic Reinforcement Learning

arXiv:2603.13019v1 Announce Type: cross Abstract: Agentic reinforcement learning (RL) has emerged as a transformative workload in cloud clusters, enabling large language models (LLMs) to solve complex problems through interactions with real world. However, unlike traditional RL, agentic RL demands substantial external cloud resources, e.g., CPUs for code execution and GPUs for reward models, that exist outside the primary training cluster. Existing agentic RL framework typically rely on static over-provisioning, i.e., resources are often tied to long-lived trajectories or isolated by tasks, which leads to severe resource inefficiency. We propose the action-level orchestration, and incorporate it into ARL-Tangram, a unified resource management system that enables fine-grained external resource sharing and elasticity. ARL-Tangram utilizes a unified action-level formulation and an elastic scheduling algorithm to minimize action completion time (ACT) while satisfying heterogeneous resource constraints. Further, heterogeneous resource managers are tailored to efficiently support the action-level execution on resources with heterogeneous characteristics and topologies. Evaluation on real-world agentic RL tasks demonstrates that ARL-Tangram improves average ACT by up to 4.3$\times$, speeds up the step duration of RL training by up to 1.5$\times$, and saves the external resources by up to 71.2$\%$. This system has been deployed to support the training of the MiMo series models.
  •  
❌