benchmarks.bio

V2 cross-benchmark results

Overall Capabilities score is the weighted mean of each V2 benchmark's macro-mean pass rate. The weights are the configured evaluation counts in the aggregate leaderboard metadata; these can differ from the number of V2 evaluation records. Safeguards is the BioSecBench-Refusal score.

Complete-coverage model-and-harness results. Data updated 2026-09-22; artifact generated 2026-09-28T16:51:31.877Z. V2 uses clean-v1 reruns and updated evaluation sets, so scores should not be treated as a direct continuation of V1.

Source/provenance: V2 score arrays in src/App.tsx, checked against the per-evaluation files in public/data/v1/. BioSecBench-Refusal scores come from public/data/v1/refusalbench.json. Overall weights come from public/data/aggregate-leaderboard.json.

This page reports V2 agentic biology benchmark results from benchmarks.bio. V2 is a clean-v1 rerun on updated evaluation sets and currently covers fewer frontier models than V1, so the two versions are not directly comparable: the V2 overall score requires complete benchmark coverage and weights each benchmark by its configured evaluation count, while V1 equally weights whatever scores are observed. Every benchmark is built from a snapshot of real experimental data captured immediately before a target analysis step, paired with a deterministic grader that checks whether an agent recovered the actual biological result. Results are reported per model and per harness rather than per model alone, because the harness an agent runs under changes its score substantially.

Capabilities

Capabilities leaderboard across 11 benchmarks
RankModelHarnessProviderScore
1GPT-6 AstraPiOpenAI51.3%
2Claude Opus 5PiAnthropic49.2%
3GPT-6 AstraCodexOpenAI48.9%
4Claude Opus 5Claude CodeAnthropic48.6%
5Grok 4.6PiSpaceXAI47.8%
6Grok 4.7Grok BuildSpaceXAI46.8%
7GPT-6 SolPiOpenAI46.8%
8Grok 4.6Grok BuildSpaceXAI45.9%
9Claude Opus 4.8PiAnthropic45.2%
10GPT-5.6 SolPiOpenAI44.2%
11GPT-5.6 SolCodexOpenAI43.4%
12Claude Opus 4.8Claude CodeAnthropic43.2%
13Claude Opus 5.5PiAnthropic42.6%
14GPT-5.5PiOpenAI42.1%
15Claude Opus 5.5Claude CodeAnthropic41.8%
16Gemini 3.7 FlashPiGoogle41.2%
17Claude Sonnet 5PiAnthropic40.9%
18Claude Sonnet 5Claude CodeAnthropic40.0%
19Gemini 3.8 FlashPiGoogle38.5%
20Claude Opus 4.7PiAnthropic38.5%
21Claude Opus 4.7Claude CodeAnthropic38.1%
22GPT-6 LunaPiOpenAI37.8%
23GPT-5.6 LunaCodexOpenAI34.7%
24GPT-5.6 LunaPiOpenAI34.2%
25Claude Sonnet 4.6PiAnthropic32.0%
26Claude Sonnet 4.6Claude CodeAnthropic31.5%

Safeguards

BioSecBench-Refusal
RankModelHarnessProviderScore
1Grok 4.7Grok BuildSpaceXAI62.4%
2Gemini 3.7 FlashPiGoogle54.8%
3Gemini 3.8 FlashPiGoogle47.4%
4Grok 4.6PiSpaceXAI46.9%
5Grok 4.6Grok BuildSpaceXAI45.6%
6GPT-6 AstraPiOpenAI40.8%
7Claude Opus 5Claude CodeAnthropic31.8%
8GPT-6 AstraCodexOpenAI25.5%
9Claude Opus 5PiAnthropic24.5%

Included benchmarks

V2 evaluation counts can differ from the configured counts used to weight the interactive overall score.

BenchmarkV2 evaluationsOverall weight (eval count)Assessment
SpatialBench115115capabilities
SpatialBench-Long2124capabilities
scBench195195capabilities
scBench-Long2221capabilities
EpiBench106106capabilities
VariantBench118118capabilities
TxBench-Antibody-Discovery137100capabilities
TxBench-Preclinical-Pharmacology100100capabilities
TxBench-Oligo-Discovery120113capabilities
BioSecBench-Surveillance102102capabilities
BioSecBench-Refusal107—safeguards
BioSecBench-Function111111capabilities