❌

Reading view

scBaseCount: An AI agent-curated, standardized, auto-updated single-cell data repository

scBaseCount is presently the largest public single-cell RNA-seq repository, containing over 502 million cells across 27 organisms and 75 tissues. An AI agent autonomously discovers, annotates, and uniformly reprocesses all 10× Genomics datasets in the SRA, creating a harmonized, continually updated resource for studying the diversity of cell biology and training AI models.
  •  

Tahoe-100M: Mapping drug-induced molecular phenotypes at single-cell resolution

Tahoe-100M is an atlas of 100 million single-cell transcriptomes, capturing how 50 cancer cell lines respond to ∼1,100 drug-dose treatments. By pairing single-cell and molecular phenotypes at scale, the resource links drug mechanisms to cellular responses and provides an openly available substrate for training predictive models of cell behavior.
  •  

Predicting cellular responses to perturbation across diverse contexts with State

Modeling perturbation effects across large single-cell populations requires flexibility to capture heterogeneity. By training over sets of cells in a shared embedding space, State outperforms baselines at generalizing effects to new contexts. Cell-Eval, the framework used for this comparison, provides a comprehensive benchmark for future models.
  •  

Virtual Cell Challenge 2026: Benchmarking zero-shot generalization across cellular contexts

The Virtual Cell Challenge returns in 2026 with a more demanding test of biological generalization: zero-shot prediction across multiple independent cellular contexts. Participants will build models to predict gene knockdown responses in a new Arc-generated dataset comprising unseen cell lines. The goal is to determine whether the best models can meaningfully close the gap between preclinical experimental predictions and human biology.
  •  
❌