benchmarks.bio

Omics · Capabilities

scBench-Long results

Long-horizon single-cell analysis tasks requiring sustained empirical work and biological reasoning.

Published score: Verified merged subset (22 evaluations); Original full release: 21 evaluations · long horizon · data updated 2026-09-20

Published scBench-Long scores
RankModelHarnessProviderScore
1GPT-5.6 SolCodexOpenAI43.9%
2GPT-6 AstraCodexOpenAI42.4%
3GPT-5.6 SolPiOpenAI40.9%
4GPT-6 AstraPiOpenAI39.4%
5Claude Opus 5PiAnthropic36.4%
6GPT-5.6 TerraPiOpenAI34.9%
7GPT-5.6 TerraCodexOpenAI31.8%
8Claude Opus 5Claude CodeAnthropic30.3%
9GPT-5.6 LunaPiOpenAI28.8%
10Grok 4.6Grok BuildSpaceXAI27.3%
11DeepSeek V4.1 FlashPiDeepSeek27.3%
12Claude Opus 4.8PiAnthropic25.8%
13Grok 4.5PiSpaceXAI25.8%
14Grok 4.6PiSpaceXAI24.2%
15Gemini 3.5 FlashPiGoogle24.2%
16GPT-5.5PiOpenAI22.7%
17Claude Opus 4.8Claude CodeAnthropic21.2%
18Claude Sonnet 5PiAnthropic19.7%
19Claude Sonnet 5Claude CodeAnthropic16.7%
20GPT-5.5CodexOpenAI16.7%
21Kimi K3PiMoonshot16.7%
22GPT-5.6 LunaCodexOpenAI12.1%
23Claude Opus 4.7Claude CodeAnthropic10.6%

Methodology and provenance

Scores use a 0–100 percentage scale. See the benchmark repository, hosted paper PDF, and original publication for task construction, grading, aggregation, and uncertainty details.