benchmarks.bio publishes agentic AI benchmark results across three domains: omics, therapeutics, and biosecurity. Every benchmark is built from a snapshot of real experimental data captured immediately before a target analysis step, paired with a deterministic grader that checks whether an agent recovered the actual biological result rather than whether it produced plausible-looking output. Results are reported per model and per harness rather than per model alone, because the harness an agent runs under changes its score substantially. The overall score in the table below is the equally weighted mean of a configuration's observed scores across the 11 capability benchmarks; a missing benchmark score is excluded from that mean and is never counted as zero. V1 and V2 cover different evaluation sets and should not be read as a single time series.
| Rank | Model | Harness | Provider | Overall score |
|---|---|---|---|---|
| 1 | GPT-6 Astra | Pi | OpenAI | 51.3% |
| 2 | GPT-6 Astra | Codex | OpenAI | 51.1% |
| 3 | Claude Opus 5 | Claude Code | Anthropic | 48.6% |
| 4 | Claude Opus 5 | Pi | Anthropic | 48.2% |
| 5 | Grok 4.6 | Grok Build | SpaceXAI | 48.2% |
| 6 | GPT-5.6 Sol | Pi | OpenAI | 47.5% |
| 7 | GPT-5.6 Sol | Codex | OpenAI | 45.9% |
| 8 | DeepSeek V4.1 Flash | Pi | DeepSeek | 45.0% |
| 9 | Grok 4.6 | Pi | SpaceXAI | 44.8% |
| 10 | Claude Opus 4.8 | Pi | Anthropic | 44.8% |
| 11 | Claude Opus 4.8 | Claude Code | Anthropic | 42.8% |
| 12 | GPT-5.6 Terra | Pi | OpenAI | 42.4% |
| 13 | Gemini 3.5 Flash | Pi | 41.7% | |
| 14 | GPT-5.6 Terra | Codex | OpenAI | 41.5% |
| 15 | Grok 4.5 | Pi | SpaceXAI | 41.3% |
| 16 | GPT-5.5 | Pi | OpenAI | 40.1% |
| 17 | Claude Opus 4.7 | Claude Code | Anthropic | 39.8% |
| 18 | Claude Sonnet 5 | Pi | Anthropic | 39.5% |
| 19 | Claude Sonnet 5 | Claude Code | Anthropic | 39.1% |
| 20 | GPT-5.5 | Codex | OpenAI | 38.7% |
| 21 | Kimi K3 | Pi | Moonshot | 38.1% |
| 22 | GPT-5.6 Luna | Pi | OpenAI | 33.8% |
| 23 | GPT-5.6 Luna | Codex | OpenAI | 32.2% |
Included benchmarks
- SpatialBench — Omics; published score: Verified subset (115 evaluations); Full benchmark: 159 evaluations
- SpatialBench-Long — Omics; published score: Verified subset (22 evaluations); Full benchmark: 24 evaluations
- scBench — Omics; published score: Full benchmark (195 evaluations)
- scBench-Long — Omics; published score: Verified merged subset (22 evaluations); Original full release: 21 evaluations
- EpiBench — Omics; published score: Full benchmark (106 evaluations)
- VariantBench — Omics; published score: Full benchmark (118 evaluations)
- TxBench-Antibody-Discovery — Therapeutics; published score: Full benchmark (100 evaluations)
- TxBench-Preclinical-Pharmacology — Therapeutics; published score: Full benchmark (100 evaluations)
- TxBench-Oligo-Discovery — Therapeutics; published score: Full benchmark (113 evaluations)
- BioSecBench-Surveillance — Biosecurity; published score: Full benchmark (102 evaluations)
- BioSecBench-Refusal — Biosecurity; published score: Full benchmark (107 evaluations)
- BioSecBench-Function — Biosecurity; published score: Full benchmark (111 evaluations)