<!-- Generated by scripts/build-backrooms-results-table.mjs. Do not edit directly; see docs/updating-results.md. -->

# SpatialBench-Long results

> Long-horizon agentic analysis of spatial biology datasets requiring multi-step biological reasoning.

Data updated: 2026-09-20

Artifact generated: 2026-09-21T21:43:49.599Z

- Domain: Omics
- Horizon: long
- Assessment: Capabilities
- Published score scope: Verified subset (22 evaluations); Full benchmark: 24 evaluations

Scores use a 0–100 percentage scale. Consult the repository and paper for benchmark-specific task construction, grading, aggregation, and uncertainty methodology. Results compare model-and-harness combinations.

| Rank | Model | Harness | Provider | Score |
| ---: | --- | --- | --- | ---: |
| 1 | GPT-5.6 Terra | Codex | OpenAI | 37.9% |
| 2 | GPT-5.6 Sol | Codex | OpenAI | 36.4% |
| 3 | GPT-5.6 Sol | Pi | OpenAI | 36.4% |
| 4 | GPT-6 Astra | Codex | OpenAI | 34.9% |
| 5 | GPT-6 Astra | Pi | OpenAI | 34.9% |
| 6 | DeepSeek V4.1 Flash | Pi | DeepSeek | 34.9% |
| 7 | GPT-5.5 | Codex | OpenAI | 33.3% |
| 8 | Claude Opus 5 | Pi | Anthropic | 31.8% |
| 9 | Grok 4.6 | Grok Build | SpaceXAI | 30.3% |
| 10 | Gemini 3.5 Flash | Pi | Google | 30.3% |
| 11 | GPT-5.5 | Pi | OpenAI | 30.3% |
| 12 | GPT-5.6 Terra | Pi | OpenAI | 30.3% |
| 13 | Claude Opus 4.8 | Claude Code | Anthropic | 28.8% |
| 14 | Claude Opus 4.8 | Pi | Anthropic | 28.8% |
| 15 | Claude Opus 5 | Claude Code | Anthropic | 28.8% |
| 16 | Kimi K3 | Pi | Moonshot | 28.8% |
| 17 | GPT-5.6 Luna | Codex | OpenAI | 27.3% |
| 18 | Claude Sonnet 5 | Pi | Anthropic | 24.2% |
| 19 | Grok 4.5 | Pi | SpaceXAI | 22.7% |
| 20 | GPT-5.6 Luna | Pi | OpenAI | 21.2% |
| 21 | Grok 4.6 | Pi | SpaceXAI | 19.7% |
| 22 | Claude Sonnet 5 | Claude Code | Anthropic | 15.2% |

## Sources and downloads

- [Paper PDF](/papers/spatialbench-long.pdf)
- [Original publication](https://latch.bio/spatialbench-long)
- [Repository](https://github.com/latchbio/spatialbench-long)
- [Interactive benchmark](https://benchmarks.bio/spatial-long)
- [JSON](https://benchmarks.bio/benchmarks/spatialbench-long/latest.json)
- [CSV](https://benchmarks.bio/benchmarks/spatialbench-long/latest.csv)
- [V1 cross-benchmark leaderboard](https://benchmarks.bio/results/v1/)
