benchmarks.bio

V2 cross-benchmark results

Overall Capabilities score is the weighted mean of each V2 benchmark's macro-mean pass rate. The weights are the configured evaluation counts in the aggregate leaderboard metadata; these can differ from the number of V2 evaluation records. Safeguards is the BioSecBench-Refusal score.

Complete-coverage model-and-harness results. Data updated 2026-09-29; artifact generated 2026-09-29T18:01:39.911Z. V2 uses clean-v1 reruns and updated evaluation sets, so scores should not be treated as a direct continuation of V1.

Source/provenance: V2 score arrays in src/App.tsx, checked against the per-evaluation files in public/data/v1/. BioSecBench-Refusal scores come from public/data/v1/refusalbench.json. Overall weights come from public/data/aggregate-leaderboard.json.

This page reports V2 agentic biology benchmark results from benchmarks.bio. V2 is a clean-v1 rerun on updated evaluation sets and currently covers fewer frontier models than V1, so the two versions are not directly comparable: the V2 overall score requires complete benchmark coverage and weights each benchmark by its configured evaluation count, while V1 equally weights whatever scores are observed. Every benchmark is built from a snapshot of real experimental data captured immediately before a target analysis step, paired with a deterministic grader that checks whether an agent recovered the actual biological result. Results are reported per model and per harness rather than per model alone, because the harness an agent runs under changes its score substantially.

Capabilities

Capabilities leaderboard across 12 benchmarks
RankModelHarnessProviderScore
1GPT-6 AstraPiOpenAI51.0%
2Claude Opus 5PiAnthropic49.6%
3Claude Opus 5Claude CodeAnthropic49.0%
4GPT-6 AstraCodexOpenAI48.9%
5Grok 4.6PiSpaceXAI47.0%
6Grok 4.7Grok BuildSpaceXAI47.0%
7GPT-6 SolPiOpenAI46.4%
8Grok 4.6Grok BuildSpaceXAI45.6%
9Claude Opus 4.8PiAnthropic45.5%
10Claude Opus 4.8Claude CodeAnthropic43.7%
11GPT-5.6 SolPiOpenAI43.1%
12Claude Opus 5.5PiAnthropic42.4%
13GPT-5.6 SolCodexOpenAI42.1%
14GPT-5.5PiOpenAI42.1%
15Claude Opus 5.5Claude CodeAnthropic42.0%
16Claude Sonnet 5PiAnthropic41.0%
17Gemini 3.7 FlashPiGoogle40.4%
18Claude Sonnet 5Claude CodeAnthropic40.1%
19Claude Opus 4.7PiAnthropic38.2%
20Claude Opus 4.7Claude CodeAnthropic38.1%
21Gemini 3.8 FlashPiGoogle37.5%
22GPT-6 LunaPiOpenAI37.2%
23GPT-5.6 LunaCodexOpenAI33.9%
24GPT-5.6 LunaPiOpenAI33.3%
25Claude Sonnet 4.6PiAnthropic31.6%
26Claude Sonnet 4.6Claude CodeAnthropic31.2%

Safeguards

BioSecBench-Refusal
RankModelHarnessProviderScore
1Grok 4.7Grok BuildSpaceXAI62.4%
2Gemini 3.7 FlashPiGoogle54.8%
3Gemini 3.8 FlashPiGoogle47.4%
4Grok 4.6PiSpaceXAI46.9%
5Grok 4.6Grok BuildSpaceXAI45.6%
6GPT-6 AstraPiOpenAI40.8%
7Claude Opus 5Claude CodeAnthropic31.8%
8GPT-6 AstraCodexOpenAI25.5%
9Claude Opus 5PiAnthropic24.5%

Included benchmarks

V2 evaluation counts can differ from the configured counts used to weight the interactive overall score.

BenchmarkV2 evaluationsOverall weight (eval count)Assessment
SpatialBench115115capabilities
SpatialBench-Long2124capabilities
scBench195195capabilities
scBench-Long2221capabilities
EpiBench106106capabilities
VariantBench118118capabilities
MetagenomicsBench100100capabilities
TxBench-Antibody-Discovery137100capabilities
TxBench-Preclinical-Pharmacology100100capabilities
TxBench-Oligo-Discovery120113capabilities
BioSecBench-Surveillance102102capabilities
BioSecBench-Refusal107—safeguards
BioSecBench-Function111111capabilities