benchmarks.bio

Biosecurity · Capabilities

BioSecBench-Function results

Biosecurity function-analysis tasks measuring capabilities on biologically sensitive questions.

Published score: Full benchmark (111 evaluations) · short horizon · data updated 2026-09-20

Published BioSecBench-Function scores
RankModelHarnessProviderScore
1Claude Opus 5Claude CodeAnthropic50.3%
2GPT-6 AstraCodexOpenAI48.6%
3GPT-6 AstraPiOpenAI47.1%
4Claude Opus 4.8PiAnthropic44.6%
5Grok 4.6Grok BuildSpaceXAI44.1%
6Claude Opus 5PiAnthropic43.0%
7Gemini 3.5 FlashPiGoogle42.2%
8Grok 4.6PiSpaceXAI42.0%
9DeepSeek V4.1 FlashPiDeepSeek38.7%
10GPT-5.6 SolPiOpenAI38.3%
11Claude Sonnet 5Claude CodeAnthropic37.2%
12Claude Opus 4.8Claude CodeAnthropic37.0%
13GPT-5.5CodexOpenAI36.5%
14Grok 4.5PiSpaceXAI35.6%
15Claude Sonnet 5PiAnthropic33.8%
16GPT-5.6 TerraPiOpenAI33.5%
17GPT-5.6 SolCodexOpenAI32.3%
18GPT-5.5PiOpenAI29.1%
19GPT-5.6 TerraCodexOpenAI29.1%
20GPT-5.6 LunaCodexOpenAI26.9%
21GPT-5.6 LunaPiOpenAI25.4%

Methodology and provenance

Scores use a 0–100 percentage scale. See the benchmark repository, hosted paper PDF, and original publication for task construction, grading, aggregation, and uncertainty details.