benchmarks.bio

V1 cross-benchmark results

Overall Capabilities score is the weighted mean of each V1 benchmark's macro-mean pass rate. The weights are the configured evaluation counts in the aggregate leaderboard metadata; these can differ from the number of V1 evaluation records. Safeguards is the BioSecBench-Refusal score.

Complete-coverage model-and-harness results. Data updated 2026-10-07; artifact generated 2026-10-09T18:38:46.980Z. V1 uses clean-v1 reruns and updated evaluation sets, so scores should not be treated as a direct continuation of legacy V0.

Source/provenance: V2 score arrays in src/App.tsx, checked against the per-evaluation files in public/data/v1/. BioSecBench-Refusal scores come from public/data/v1/refusalbench.json. Overall weights come from public/data/aggregate-leaderboard.json.

This page reports V1 agentic biology benchmark results from benchmarks.bio. V1 is a clean-v1 rerun on updated evaluation sets and currently covers fewer frontier models than V1, so the two versions are not directly comparable: the V1 overall score requires complete benchmark coverage and weights each benchmark by its configured evaluation count, while V1 equally weights whatever scores are observed. Every benchmark is built from a snapshot of real experimental data captured immediately before a target analysis step, paired with a deterministic grader that checks whether an agent recovered the actual biological result. Results are reported per model and per harness rather than per model alone, because the harness an agent runs under changes its score substantially.

Capabilities

Capabilities leaderboard across 12 benchmarks
RankModelHarnessProviderScore
1Claude Opus 5.5*PiAnthropic58.1%
2Claude Opus 5.5*Claude CodeAnthropic58.1%
3Claude Sonnet 5.5PiAnthropic54.8%
4Claude Sonnet 5.5Claude CodeAnthropic54.4%
5Claude Fable 5.1*Claude CodeAnthropic53.2%
6GPT-6 AstraPiOpenAI51.9%
7GPT-6.1 SolPiOpenAI51.7%
8Claude Opus 5PiAnthropic50.2%
9GPT-6 AstraCodexOpenAI49.9%
10GPT-6.1 SolCodexOpenAI49.7%
11Claude Opus 5Claude CodeAnthropic49.6%
12Grok 4.6PiSpaceXAI47.9%
13Grok 4.7Grok BuildSpaceXAI47.9%
14GPT-6 SolPiOpenAI47.3%
15Grok 4.6Grok BuildSpaceXAI46.5%
16Claude Opus 4.8PiAnthropic45.9%
17Claude Opus 4.8Claude CodeAnthropic44.1%
18GPT-5.6 SolPiOpenAI43.8%
19DeepSeek V4.1 FlashPiDeepSeek42.9%
20GPT-5.6 SolCodexOpenAI42.8%
21Claude Opus 5.5PiAnthropic42.6%
22GPT-5.5PiOpenAI42.5%
23Claude Opus 5.5Claude CodeAnthropic42.2%
24Claude Sonnet 5PiAnthropic41.4%
25Gemini 3.7 FlashPiGoogle41.3%
26Claude Sonnet 5Claude CodeAnthropic40.4%
27Claude Opus 4.7PiAnthropic38.6%
28Claude Opus 4.7Claude CodeAnthropic38.5%
29Gemini 3.8 FlashPiGoogle38.2%
30GPT-6 LunaPiOpenAI37.7%
31Kimi K3PiMoonshot36.6%
32GPT-5.6 LunaCodexOpenAI34.0%
33GPT-5.6 LunaPiOpenAI33.4%
34Claude Sonnet 4.6PiAnthropic32.1%
35Claude Sonnet 4.6Claude CodeAnthropic31.7%
36Mistral Large 4.0 PreviewPiMistral26.1%
37Nemotron 3 Ultra 550B A55BPiNVIDIA24.6%
38Nemotron 3.5 LightningPiNVIDIA14.7%

Safeguards

BioSecBench-Refusal
RankModelHarnessProviderScore
1Grok 4.7Grok BuildSpaceXAI62.4%
2Gemini 3.7 FlashPiGoogle54.8%
3Gemini 3.8 FlashPiGoogle47.4%
4Grok 4.6PiSpaceXAI46.9%
5Grok 4.6Grok BuildSpaceXAI45.6%
6GPT-6 AstraPiOpenAI40.8%
7Claude Opus 5Claude CodeAnthropic31.8%
8GPT-6 AstraCodexOpenAI25.5%
9Claude Opus 5PiAnthropic24.5%

Included benchmarks

V2 evaluation counts can differ from the configured counts used to weight the interactive overall score.

BenchmarkV2 evaluationsOverall weight (eval count)Assessment
SpatialBench115115capabilities
SpatialBench-Long2124capabilities
scBench195195capabilities
scBench-Long2221capabilities
EpiBench106106capabilities
VariantBench118118capabilities
MetagenomicsBench100100capabilities
TxBench-Antibody-Discovery100100capabilities
TxBench-Preclinical-Pharmacology100100capabilities
TxBench-Oligo-Discovery113113capabilities
BioSecBench-Surveillance102102capabilities
BioSecBench-Refusal107—safeguards
BioSecBench-Function111111capabilities