This page reports published MetagenomicsBench scores from benchmarks.bio. MetagenomicsBench contains 100 evaluations built from real experimental data, each one graded deterministically against the biological result the original analysis reached, so a run counts as a pass only when the agent recovers that result rather than when it produces plausible-looking output. Results are reported per model and per harness rather than per model alone, because the harness an agent runs under changes its score substantially. Scores use a 0-100 percentage scale and are not comparable across benchmarks; see the linked repository and paper for task construction, grading, and confidence-interval methodology.
| Rank | Model | Harness | Provider | Score |
|---|---|---|---|---|
| 1 | Claude Opus 5 | Pi | Anthropic | 54.0% |
| 2 | Claude Opus 5 | Claude Code | Anthropic | 53.3% |
| 3 | Claude Opus 4.8 | Claude Code | Anthropic | 50.0% |
| 4 | GPT-6 Astra | Codex | OpenAI | 49.7% |
| 5 | Claude Opus 4.8 | Pi | Anthropic | 48.3% |
| 6 | GPT-6 Astra | Pi | OpenAI | 47.0% |
| 7 | Claude Sonnet 5 | Pi | Anthropic | 42.3% |
| 8 | Grok 4.6 | Grok Build | SpaceXAI | 42.0% |
| 9 | Claude Sonnet 5 | Claude Code | Anthropic | 41.0% |
| 10 | Grok 4.6 | Pi | SpaceXAI | 38.7% |
| 11 | DeepSeek V4.1 Flash | Pi | DeepSeek | 38.0% |
| 12 | Claude Opus 4.7 | Pi | Anthropic | 35.3% |
| 13 | GPT-5.6 Sol | Pi | OpenAI | 31.3% |
| 14 | GPT-5.6 Sol | Codex | OpenAI | 28.7% |
| 15 | Kimi K3 | Pi | Moonshot | 27.3% |
| 16 | GPT-5.6 Luna | Codex | OpenAI | 24.3% |
| 17 | GPT-5.6 Luna | Pi | OpenAI | 22.7% |
Methodology and provenance
Scores use a 0–100 percentage scale. See the benchmark repository, hosted paper PDF, and original publication for task construction, grading, aggregation, and uncertainty details.