<!-- Generated by scripts/build-backrooms-results-table.mjs. Do not edit directly; see docs/updating-results.md. -->

# benchmarks.bio V1 results

> V1 cross-benchmark leaderboard for agentic biology benchmarks on messy, real-world biological data.

Data updated: 2026-09-20

Artifact generated: 2026-09-21T20:33:37.065Z

## How to interpret these results

Overall score is the equally weighted mean of observed benchmark scores.

An overall score uses the configured weight for each available capability-benchmark score. Missing benchmark scores are excluded from the weighted mean and are never treated as zero. Coverage reports observed capability benchmarks divided by the 11 capability benchmarks in the default scope. Scores are percentages on a 0–100 scale.

Results compare complete model-and-harness combinations; the same model may score differently under different agent harnesses. Published score scopes are stated explicitly below; they can differ from other available result sets. See each benchmark page and linked repository for benchmark-specific grading and confidence-interval methodology.

## Overall capabilities leaderboard

| Rank | Model | Harness | Provider | Overall score | Coverage |
| ---: | --- | --- | --- | ---: | ---: |
| 1 | GPT-6 Astra | Pi | OpenAI | 51.3% | 11/11 |
| 2 | GPT-6 Astra | Codex | OpenAI | 51.1% | 11/11 |
| 3 | Claude Opus 5 | Claude Code | Anthropic | 48.6% | 11/11 |
| 4 | Claude Opus 5 | Pi | Anthropic | 48.2% | 11/11 |
| 5 | Grok 4.6 | Grok Build | SpaceXAI | 48.2% | 11/11 |
| 6 | GPT-5.6 Sol | Pi | OpenAI | 47.5% | 9/11 |
| 7 | GPT-5.6 Sol | Codex | OpenAI | 45.9% | 10/11 |
| 8 | DeepSeek V4.1 Flash | Pi | DeepSeek | 45.0% | 9/11 |
| 9 | Grok 4.6 | Pi | SpaceXAI | 44.8% | 11/11 |
| 10 | Claude Opus 4.8 | Pi | Anthropic | 44.8% | 10/11 |
| 11 | Claude Opus 4.8 | Claude Code | Anthropic | 42.8% | 10/11 |
| 12 | GPT-5.6 Terra | Pi | OpenAI | 42.4% | 10/11 |
| 13 | Gemini 3.5 Flash | Pi | Google | 41.7% | 10/11 |
| 14 | GPT-5.6 Terra | Codex | OpenAI | 41.5% | 10/11 |
| 15 | Grok 4.5 | Pi | SpaceXAI | 41.3% | 10/11 |
| 16 | GPT-5.5 | Pi | OpenAI | 40.1% | 10/11 |
| 17 | Claude Opus 4.7 | Claude Code | Anthropic | 39.8% | 7/11 |
| 18 | Claude Sonnet 5 | Pi | Anthropic | 39.5% | 9/11 |
| 19 | Claude Sonnet 5 | Claude Code | Anthropic | 39.1% | 10/11 |
| 20 | GPT-5.5 | Codex | OpenAI | 38.7% | 10/11 |
| 21 | Kimi K3 | Pi | Moonshot | 38.1% | 7/11 |
| 22 | GPT-5.6 Luna | Pi | OpenAI | 33.8% | 10/11 |
| 23 | GPT-5.6 Luna | Codex | OpenAI | 32.2% | 10/11 |

## Included benchmarks

| Benchmark | Domain | Horizon | Assessment | Published score scope |
| --- | --- | --- | --- | --- |
| [SpatialBench](https://benchmarks.bio/benchmarks/spatialbench/) | Omics | short | Capabilities | Verified subset (115 evaluations); Full benchmark: 159 evaluations |
| [SpatialBench-Long](https://benchmarks.bio/benchmarks/spatialbench-long/) | Omics | long | Capabilities | Verified subset (22 evaluations); Full benchmark: 24 evaluations |
| [scBench](https://benchmarks.bio/benchmarks/scbench/) | Omics | short | Capabilities | Full benchmark (195 evaluations) |
| [scBench-Long](https://benchmarks.bio/benchmarks/scbench-long/) | Omics | long | Capabilities | Verified merged subset (22 evaluations); Original full release: 21 evaluations |
| [EpiBench](https://benchmarks.bio/benchmarks/epibench/) | Omics | short | Capabilities | Full benchmark (106 evaluations) |
| [VariantBench](https://benchmarks.bio/benchmarks/variantbench/) | Omics | short | Capabilities | Full benchmark (118 evaluations) |
| [TxBench-Antibody-Discovery](https://benchmarks.bio/benchmarks/txbench-ab/) | Therapeutics | short | Capabilities | Full benchmark (100 evaluations) |
| [TxBench-Preclinical-Pharmacology](https://benchmarks.bio/benchmarks/txbench-pp/) | Therapeutics | short | Capabilities | Full benchmark (100 evaluations) |
| [TxBench-Oligo-Discovery](https://benchmarks.bio/benchmarks/txbench-od/) | Therapeutics | short | Capabilities | Full benchmark (113 evaluations) |
| [BioSecBench-Surveillance](https://benchmarks.bio/benchmarks/biosecbench-surveillance/) | Biosecurity | short | Capabilities | Full benchmark (102 evaluations) |
| [BioSecBench-Refusal](https://benchmarks.bio/benchmarks/biosecbench-refusal/) | Biosecurity | short | Safeguards | Full benchmark (107 evaluations) |
| [BioSecBench-Function](https://benchmarks.bio/benchmarks/functionbench/) | Biosecurity | short | Capabilities | Full benchmark (111 evaluations) |

## Downloads and provenance

- [Complete results as JSON](https://benchmarks.bio/results/v1/latest.json)
- [Capabilities leaderboard as CSV](https://benchmarks.bio/results/v1/latest.csv)
- [Static HTML leaderboard](https://benchmarks.bio/results/v1/)
- [V2 leaderboard](https://benchmarks.bio/results/v2/) — clean-v1 reruns, complete coverage, and eval-count-weighted scores.
- [Interactive leaderboard](https://benchmarks.bio/)
- [Machine-readable site guide](https://benchmarks.bio/llms.txt)
