CellForge
ML evaluation platform for reproducible model generalization benchmarks
A machine-learning evaluation harness with reproducible splits, model adapters, benchmark manifests, and analysis tools for comparing generalization across changing contexts.
Strong aggregate model scores can hide failures on unseen perturbations, donors, cell types, or biological contexts.
Dataset registry and split engine define reproducible IID and OOD conditions.
Model adapters compare foundation models with strong simple baselines through one interface.
Biological metrics include correlation, prediction error, DEG recovery, top-k overlap, and direction of effect.
AnnData result artifacts connect benchmark failures to CELLxGENE exploration.
The portfolio proof is the benchmark table itself: simple baselines and foundation models compared under IID and OOD conditions with reproducible run manifests.