benchmarks.bio

Biosecurity · Capabilities

BioSecBench-Surveillance results

Biosecurity surveillance tasks that measure an agent's ability to analyze biological threat signals.

Published score: Full benchmark (102 evaluations) · short horizon · data updated 2026-09-29

This page reports published BioSecBench-Surveillance scores from benchmarks.bio. BioSecBench-Surveillance contains 102 evaluations built from real experimental data, each one graded deterministically against the biological result the original analysis reached, so a run counts as a pass only when the agent recovers that result rather than when it produces plausible-looking output. Results are reported per model and per harness rather than per model alone, because the harness an agent runs under changes its score substantially. Scores use a 0-100 percentage scale and are not comparable across benchmarks; see the linked repository and paper for task construction, grading, and confidence-interval methodology.

Published BioSecBench-Surveillance scores
RankModelHarnessProviderScore
1Claude Opus 5PiAnthropic51.5%
2GPT-6 AstraPiOpenAI48.5%
3Grok 4.6Grok BuildSpaceXAI48.0%
4Grok 4.6PiSpaceXAI47.5%
5Claude Opus 5Claude CodeAnthropic43.6%
6GPT-6 AstraCodexOpenAI35.2%

Methodology and provenance

Scores use a 0–100 percentage scale. See the benchmark repository, hosted paper PDF, and original publication for task construction, grading, aggregation, and uncertainty details.