While single-cell foundation models (SCFMs) have shown promise across various downstream tasks, their generalization performance in label scarce settings remains a critical bottleneck. The absence of systematic benchmarks for these low resource scenarios hinders their translation to real world biomedical research. To bridge this gap, we present CellBench-LS, a comprehensive frame work designed to rigorously evaluate SCFMs generalization under low-supervision conditions. This framework employ a stratified evaluation protocol to systematically compare traditional methods and foundation models. We evaluate their zero-shot representational abilities on cell clustering and batch correction tasks, and apply lightweight fine-tuning of task-specific heads for predictive tasks, such as celltype annotation, expression reconstruction, and perturbation prediction.
