benchmarks.bio

Legacy V0 results

V0 cross-benchmark results

Overall score is the equally weighted mean of observed benchmark scores.

Data updated 2026-09-29; artifact generated 2026-09-30T21:41:57.156Z. Missing scores are excluded from a model's mean and are never treated as zero.

benchmarks.bio publishes agentic AI benchmark results across three domains: omics, therapeutics, and biosecurity. Every benchmark is built from a snapshot of real experimental data captured immediately before a target analysis step, paired with a deterministic grader that checks whether an agent recovered the actual biological result rather than whether it produced plausible-looking output. Results are reported per model and per harness rather than per model alone, because the harness an agent runs under changes its score substantially. The overall score in the table below is the equally weighted mean of a configuration's observed scores across the 12 capability benchmarks; a missing benchmark score is excluded from that mean and is never counted as zero. V0 and V1 cover different evaluation sets and should not be read as a single time series.

Capabilities leaderboard across 12 benchmarks
RankModelHarnessProviderOverall score
1GPT-6 AstraPiOpenAI51.0%
2GPT-6 AstraCodexOpenAI50.9%
3Claude Opus 5Claude CodeAnthropic49.0%
4Claude Opus 5PiAnthropic48.7%
5Grok 4.6Grok BuildSpaceXAI47.7%
6GPT-5.6 SolPiOpenAI45.9%
7Claude Opus 4.8PiAnthropic45.1%
8Grok 4.6PiSpaceXAI44.3%
9GPT-5.6 SolCodexOpenAI44.3%
10DeepSeek V4.1 FlashPiDeepSeek44.3%
11Claude Opus 4.8Claude CodeAnthropic43.4%
12GPT-5.6 TerraPiOpenAI42.4%
13Gemini 3.5 FlashPiGoogle41.7%
14GPT-5.6 TerraCodexOpenAI41.5%
15Grok 4.5PiSpaceXAI41.3%
16GPT-5.5PiOpenAI40.1%
17Claude Sonnet 5PiAnthropic39.8%
18Claude Opus 4.7Claude CodeAnthropic39.8%
19Claude Sonnet 5Claude CodeAnthropic39.3%
20GPT-5.5CodexOpenAI38.7%
21Kimi K3PiMoonshot36.8%
22GPT-5.6 LunaPiOpenAI32.8%
23GPT-5.6 LunaCodexOpenAI31.4%

Included benchmarks