benchmarks.bio

Omics · Capabilities

scBench results

Agentic analysis of real single-cell RNA-seq datasets across platforms and task categories.

Published score: Full benchmark (195 evaluations) · short horizon · data updated 2026-09-20

Published scBench scores
RankModelHarnessProviderScore
1GPT-6 AstraCodexOpenAI64.8%
2GPT-6 AstraPiOpenAI64.4%
3GPT-5.6 SolPiOpenAI62.1%
4GPT-5.6 TerraPiOpenAI62.0%
5Grok 4.6Grok BuildSpaceXAI61.5%
6GPT-5.6 SolCodexOpenAI60.3%
7Claude Opus 5Claude CodeAnthropic60.1%
8Grok 4.6PiSpaceXAI59.7%
9Claude Sonnet 5PiAnthropic59.5%
10Claude Opus 5PiAnthropic59.5%
11DeepSeek V4.1 FlashPiDeepSeek59.5%
12Claude Opus 4.8Claude CodeAnthropic58.0%
13GPT-5.5CodexOpenAI57.8%
14Claude Opus 4.8PiAnthropic57.3%
15Gemini 3.5 FlashPiGoogle56.9%
16Claude Sonnet 5Claude CodeAnthropic56.7%
17GPT-5.5PiOpenAI56.7%
18Grok 4.5PiSpaceXAI56.6%
19Kimi K3PiMoonshot55.7%
20GPT-5.6 TerraCodexOpenAI55.0%
21Claude Opus 4.7Claude CodeAnthropic54.0%
22Claude Opus 4.7PiAnthropic53.9%
23GPT-5.6 LunaPiOpenAI51.9%
24GPT-5.6 LunaCodexOpenAI49.2%

Methodology and provenance

Scores use a 0–100 percentage scale. See the benchmark repository, hosted paper PDF, and original publication for task construction, grading, aggregation, and uncertainty details.