benchmarks.bio

Omics · Capabilities

SpatialBench-Long results

Long-horizon agentic analysis of spatial biology datasets requiring multi-step biological reasoning.

Published score: Verified subset (22 evaluations); Full benchmark: 24 evaluations · long horizon · data updated 2026-09-20

Published SpatialBench-Long scores
RankModelHarnessProviderScore
1GPT-5.6 TerraCodexOpenAI37.9%
2GPT-5.6 SolCodexOpenAI36.4%
3GPT-5.6 SolPiOpenAI36.4%
4GPT-6 AstraCodexOpenAI34.9%
5GPT-6 AstraPiOpenAI34.9%
6DeepSeek V4.1 FlashPiDeepSeek34.9%
7GPT-5.5CodexOpenAI33.3%
8Claude Opus 5PiAnthropic31.8%
9Grok 4.6Grok BuildSpaceXAI30.3%
10Gemini 3.5 FlashPiGoogle30.3%
11GPT-5.5PiOpenAI30.3%
12GPT-5.6 TerraPiOpenAI30.3%
13Claude Opus 4.8Claude CodeAnthropic28.8%
14Claude Opus 4.8PiAnthropic28.8%
15Claude Opus 5Claude CodeAnthropic28.8%
16Kimi K3PiMoonshot28.8%
17GPT-5.6 LunaCodexOpenAI27.3%
18Claude Sonnet 5PiAnthropic24.2%
19Grok 4.5PiSpaceXAI22.7%
20GPT-5.6 LunaPiOpenAI21.2%
21Grok 4.6PiSpaceXAI19.7%
22Claude Sonnet 5Claude CodeAnthropic15.2%

Methodology and provenance

Scores use a 0–100 percentage scale. See the benchmark repository, hosted paper PDF, and original publication for task construction, grading, aggregation, and uncertainty details.