❌

Reading view

AI-generated data contamination erodes pathological variability and diagnostic reliability

arXiv:2601.12946v1 Announce Type: cross Abstract: Generative artificial intelligence (AI) is rapidly populating medical records with synthetic content, creating a feedback loop where future models are increasingly at risk of training on uncurated AI-generated data. However, the clinical consequences of this AI-generated data contamination remain unexplored. Here, we show that in the absence of mandatory human verification, this self-referential cycle drives a rapid erosion of pathological variability and diagnostic reliability. By analysing more than 800,000 synthetic data points across clinical text generation, vision-language reporting, and medical image synthesis, we find that models progressively converge toward generic phenotypes regardless of the model architecture. Specifically, rare but critical findings, including pneumothorax and effusions, vanish from the synthetic content generated by AI models, while demographic representations skew heavily toward middle-aged male phenotypes. Crucially, this degradation is masked by false diagnostic confidence; models continue to issue reassuring reports while failing to detect life-threatening pathology, with false reassurance rates tripling to 40%. Blinded physician evaluation confirms that this decoupling of confidence and accuracy renders AI-generated documentation clinically useless after just two generations. We systematically evaluate three mitigation strategies, finding that while synthetic volume scaling fails to prevent collapse, mixing real data with quality-aware filtering effectively preserves diversity. Ultimately, our results suggest that without policy-mandated human oversight, the deployment of generative AI threatens to degrade the very healthcare data ecosystems it relies upon.
  •  

Stereo-seq V2: Spatial mapping of total RNA on FFPE sections with high resolution

Stereo-seq V2 facilitates single-cell-resolution spatial RNA mapping in FFPE samples through random primer capture, uncovering ncRNAs, host-pathogen transcriptome profiling, and spatial immune repertoires in situ.
  •  

Self-Correction Distillation for Structured Data Question Answering

arXiv:2511.07998v1 Announce Type: cross Abstract: Structured data question answering (QA), including table QA, Knowledge Graph (KG) QA, and temporal KG QA, is a pivotal research area. Advances in large language models (LLMs) have driven significant progress in unified structural QA frameworks like TrustUQA. However, these frameworks face challenges when applied to small-scale LLMs since small-scale LLMs are prone to errors in generating structured queries. To improve the structured data QA ability of small-scale LLMs, we propose a self-correction distillation (SCD) method. In SCD, an error prompt mechanism (EPM) is designed to detect errors and provide customized error messages during inference, and a two-stage distillation strategy is designed to transfer large-scale LLMs' query-generation and error-correction capabilities to small-scale LLM. Experiments across 5 benchmarks with 3 structured data types demonstrate that our SCD achieves the best performance and superior generalization on small-scale LLM (8B) compared to other distillation methods, and closely approaches the performance of GPT4 on some datasets. Furthermore, large-scale LLMs equipped with EPM surpass the state-of-the-art results on most datasets.
  •  

Stereo-seq V2: Spatial mapping of total RNA on FFPE sections with high resolution

Cell. 2025 Aug 22:S0092-8674(25)00922-5. doi: 10.1016/j.cell.2025.08.008. Online ahead of print.

ABSTRACT

Performing total RNA profiling on formalin-fixed, paraffin-embedded (FFPE) samples, the predominant sample conservation method in clinical practice, remains challenging for current spatial transcriptomics techniques. Here, we introduce Stereo-seq V2, which employs random primers to capture and sequence RNAs in situ on FFPE sections and provides single-cell resolution. The random-priming-based strategy offers unbiased transcript capturing and uniform gene body coverage, which increase the sensitivity to marker genes, the efficiency of non-polyadenylation (poly(A)) RNA profiling, and immune repertoire coverage. We demonstrated the robust performance of Stereo-seq V2 on clinical FFPE samples using triple-negative breast cancer (TNBC) sections and identified tumor-specific alternative splicing events. In a Mycobacterium tuberculosis (Mtb)-infected mouse model, we monitored gene expression dynamics of host and pathogen transcriptomes simultaneously by utilizing Stereo-seq V2. We also assembled immune repertoires and identified Mtb-specific BCR clones, which could also be observed in human tuberculous lung samples. These results highlight Stereo-seq V2's potential in biomedical research and personalized medicine.

PMID:40882628 | DOI:10.1016/j.cell.2025.08.008

  •  
❌