❌

Normal view

scBaseCount: An AI agent-curated, standardized, auto-updated single-cell data repository

scBaseCount is presently the largest public single-cell RNA-seq repository, containing over 502 million cells across 27 organisms and 75 tissues. An AI agent autonomously discovers, annotates, and uniformly reprocesses all 10× Genomics datasets in the SRA, creating a harmonized, continually updated resource for studying the diversity of cell biology and training AI models.

Tahoe-100M: Mapping drug-induced molecular phenotypes at single-cell resolution

Tahoe-100M is an atlas of 100 million single-cell transcriptomes, capturing how 50 cancer cell lines respond to ∼1,100 drug-dose treatments. By pairing single-cell and molecular phenotypes at scale, the resource links drug mechanisms to cellular responses and provides an openly available substrate for training predictive models of cell behavior.

Predicting cellular responses to perturbation across diverse contexts with State

Modeling perturbation effects across large single-cell populations requires flexibility to capture heterogeneity. By training over sets of cells in a shared embedding space, State outperforms baselines at generalizing effects to new contexts. Cell-Eval, the framework used for this comparison, provides a comprehensive benchmark for future models.

Virtual Cell Challenge 2026: Benchmarking zero-shot generalization across cellular contexts

26 August 2026 at 08:00
The Virtual Cell Challenge returns in 2026 with a more demanding test of biological generalization: zero-shot prediction across multiple independent cellular contexts. Participants will build models to predict gene knockdown responses in a new Arc-generated dataset comprising unseen cell lines. The goal is to determine whether the best models can meaningfully close the gap between preclinical experimental predictions and human biology.
❌