benchmarks.bio

Omics · Capabilities

SpatialBench results

Agentic analysis of real spatial transcriptomics datasets across platforms and biological task types.

Published score: Verified subset (115 evaluations); Full benchmark: 159 evaluations · short horizon · data updated 2026-09-20

Published SpatialBench scores
RankModelHarnessProviderScore
1Grok 4.6Grok BuildSpaceXAI77.4%
2DeepSeek V4.1 FlashPiDeepSeek75.4%
3GPT-5.6 SolCodexOpenAI74.5%
4GPT-5.6 SolPiOpenAI73.0%
5Claude Opus 5Claude CodeAnthropic71.0%
6Grok 4.5PiSpaceXAI71.0%
7GPT-5.6 TerraPiOpenAI70.4%
8GPT-6 AstraCodexOpenAI70.1%
9GPT-6 AstraPiOpenAI70.1%
10Grok 4.6PiSpaceXAI69.9%
11GPT-5.6 TerraCodexOpenAI69.6%
12Claude Sonnet 5PiAnthropic69.3%
13Claude Opus 4.8Claude CodeAnthropic69.0%
14Gemini 3.5 FlashPiGoogle67.0%
15Claude Opus 5PiAnthropic66.7%
16Claude Opus 4.8PiAnthropic66.4%
17GPT-5.5PiOpenAI65.2%
18GPT-5.5CodexOpenAI64.9%
19Claude Sonnet 5Claude CodeAnthropic64.3%
20Claude Opus 4.7Claude CodeAnthropic64.1%
21Claude Opus 4.7PiAnthropic63.8%
22Kimi K3PiMoonshot61.5%
23GPT-5.6 LunaCodexOpenAI60.6%
24GPT-5.6 LunaPiOpenAI57.1%

Methodology and provenance

Scores use a 0–100 percentage scale. See the benchmark repository, hosted paper PDF, and original publication for task construction, grading, aggregation, and uncertainty details.