<!-- Generated by scripts/build-backrooms-results-table.mjs. Do not edit directly; see docs/updating-results.md. -->

# benchmarks.bio V0 results

> Legacy V0 cross-benchmark leaderboard for agentic biology benchmarks on messy, real-world biological data.

Data updated: 2026-09-29

Artifact generated: 2026-09-30T21:41:57.156Z

## How to interpret these results

Overall score is the equally weighted mean of observed benchmark scores.

An overall score uses the configured weight for each available capability-benchmark score. Missing benchmark scores are excluded from the weighted mean and are never treated as zero. Coverage reports observed capability benchmarks divided by the 12 capability benchmarks in the default scope. Scores are percentages on a 0–100 scale.

Results compare complete model-and-harness combinations; the same model may score differently under different agent harnesses. Published score scopes are stated explicitly below; they can differ from other available result sets. See each benchmark page and linked repository for benchmark-specific grading and confidence-interval methodology.

## Overall capabilities leaderboard

| Rank | Model | Harness | Provider | Overall score | Coverage |
| ---: | --- | --- | --- | ---: | ---: |
| 1 | GPT-6 Astra | Pi | OpenAI | 51.0% | 12/12 |
| 2 | GPT-6 Astra | Codex | OpenAI | 50.9% | 12/12 |
| 3 | Claude Opus 5 | Claude Code | Anthropic | 49.0% | 12/12 |
| 4 | Claude Opus 5 | Pi | Anthropic | 48.7% | 12/12 |
| 5 | Grok 4.6 | Grok Build | SpaceXAI | 47.7% | 12/12 |
| 6 | GPT-5.6 Sol | Pi | OpenAI | 45.9% | 10/12 |
| 7 | Claude Opus 4.8 | Pi | Anthropic | 45.1% | 11/12 |
| 8 | Grok 4.6 | Pi | SpaceXAI | 44.3% | 12/12 |
| 9 | GPT-5.6 Sol | Codex | OpenAI | 44.3% | 11/12 |
| 10 | DeepSeek V4.1 Flash | Pi | DeepSeek | 44.3% | 10/12 |
| 11 | Claude Opus 4.8 | Claude Code | Anthropic | 43.4% | 11/12 |
| 12 | GPT-5.6 Terra | Pi | OpenAI | 42.4% | 10/12 |
| 13 | Gemini 3.5 Flash | Pi | Google | 41.7% | 10/12 |
| 14 | GPT-5.6 Terra | Codex | OpenAI | 41.5% | 10/12 |
| 15 | Grok 4.5 | Pi | SpaceXAI | 41.3% | 10/12 |
| 16 | GPT-5.5 | Pi | OpenAI | 40.1% | 10/12 |
| 17 | Claude Sonnet 5 | Pi | Anthropic | 39.8% | 10/12 |
| 18 | Claude Opus 4.7 | Claude Code | Anthropic | 39.8% | 7/12 |
| 19 | Claude Sonnet 5 | Claude Code | Anthropic | 39.3% | 11/12 |
| 20 | GPT-5.5 | Codex | OpenAI | 38.7% | 10/12 |
| 21 | Kimi K3 | Pi | Moonshot | 36.8% | 8/12 |
| 22 | GPT-5.6 Luna | Pi | OpenAI | 32.8% | 11/12 |
| 23 | GPT-5.6 Luna | Codex | OpenAI | 31.4% | 11/12 |

## Included benchmarks

| Benchmark | Domain | Horizon | Assessment | Published score scope |
| --- | --- | --- | --- | --- |
| [SpatialBench](https://benchmarks.bio/benchmarks/spatialbench/) | Omics | short | Capabilities | Verified subset (115 evaluations); Full benchmark: 159 evaluations |
| [SpatialBench-Long](https://benchmarks.bio/benchmarks/spatialbench-long/) | Omics | long | Capabilities | Verified subset (22 evaluations); Full benchmark: 24 evaluations |
| [scBench](https://benchmarks.bio/benchmarks/scbench/) | Omics | short | Capabilities | Full benchmark (195 evaluations) |
| [scBench-Long](https://benchmarks.bio/benchmarks/scbench-long/) | Omics | long | Capabilities | Verified merged subset (22 evaluations); Original full release: 21 evaluations |
| [EpiBench](https://benchmarks.bio/benchmarks/epibench/) | Omics | short | Capabilities | Full benchmark (106 evaluations) |
| [VariantBench](https://benchmarks.bio/benchmarks/variantbench/) | Omics | short | Capabilities | Full benchmark (118 evaluations) |
| [MetagenomicsBench](https://benchmarks.bio/benchmarks/metagenomics/) | Omics | short | Capabilities | Full benchmark (100 evaluations) |
| [TxBench-Antibody-Discovery](https://benchmarks.bio/benchmarks/txbench-ab/) | Therapeutics | short | Capabilities | Full benchmark (100 evaluations) |
| [TxBench-Preclinical-Pharmacology](https://benchmarks.bio/benchmarks/txbench-pp/) | Therapeutics | short | Capabilities | Full benchmark (100 evaluations) |
| [TxBench-Oligo-Discovery](https://benchmarks.bio/benchmarks/txbench-od/) | Therapeutics | short | Capabilities | Full benchmark (113 evaluations) |
| [BioSecBench-Surveillance](https://benchmarks.bio/benchmarks/biosecbench-surveillance/) | Biosecurity | short | Capabilities | Full benchmark (102 evaluations) |
| [BioSecBench-Refusal](https://benchmarks.bio/benchmarks/biosecbench-refusal/) | Biosecurity | short | Safeguards | Full benchmark (107 evaluations) |
| [BioSecBench-Function](https://benchmarks.bio/benchmarks/functionbench/) | Biosecurity | short | Capabilities | Full benchmark (111 evaluations) |

## Downloads and provenance

- [Complete results as JSON](https://benchmarks.bio/results/v0/latest.json)
- [Capabilities leaderboard as CSV](https://benchmarks.bio/results/v0/latest.csv)
- [Static HTML leaderboard](https://benchmarks.bio/results/v0/)
- [Current V1 leaderboard](https://benchmarks.bio/results/v1/) — clean-v1 reruns, complete coverage, and eval-count-weighted scores.
- [Interactive leaderboard](https://benchmarks.bio/)
- [Machine-readable site guide](https://benchmarks.bio/llms.txt)
