benchmarks.bio

Omics · Capabilities

MetagenomicsBench results

Agentic microbial community analysis tasks spanning shotgun, amplicon, long-read and spatially integrated metagenomics.

Published score: Full benchmark (100 evaluations) · short horizon · data updated 2026-09-29

This page reports published MetagenomicsBench scores from benchmarks.bio. MetagenomicsBench contains 100 evaluations built from real experimental data, each one graded deterministically against the biological result the original analysis reached, so a run counts as a pass only when the agent recovers that result rather than when it produces plausible-looking output. Results are reported per model and per harness rather than per model alone, because the harness an agent runs under changes its score substantially. Scores use a 0-100 percentage scale and are not comparable across benchmarks; see the linked repository and paper for task construction, grading, and confidence-interval methodology.

Published MetagenomicsBench scores
RankModelHarnessProviderScore
1Claude Opus 5PiAnthropic54.0%
2Claude Opus 5Claude CodeAnthropic53.3%
3Claude Opus 4.8Claude CodeAnthropic50.0%
4GPT-6 AstraCodexOpenAI49.7%
5Claude Opus 4.8PiAnthropic48.3%
6GPT-6 AstraPiOpenAI47.0%
7Claude Sonnet 5PiAnthropic42.3%
8Grok 4.6Grok BuildSpaceXAI42.0%
9Claude Sonnet 5Claude CodeAnthropic41.0%
10Grok 4.6PiSpaceXAI38.7%
11DeepSeek V4.1 FlashPiDeepSeek38.0%
12Claude Opus 4.7PiAnthropic35.3%
13GPT-5.6 SolPiOpenAI31.3%
14GPT-5.6 SolCodexOpenAI28.7%
15Kimi K3PiMoonshot27.3%
16GPT-5.6 LunaCodexOpenAI24.3%
17GPT-5.6 LunaPiOpenAI22.7%

Methodology and provenance

Scores use a 0–100 percentage scale. See the benchmark repository, hosted paper PDF, and original publication for task construction, grading, aggregation, and uncertainty details.