This page reports published BioSecBench-Surveillance scores from benchmarks.bio. BioSecBench-Surveillance contains 102 evaluations built from real experimental data, each one graded deterministically against the biological result the original analysis reached, so a run counts as a pass only when the agent recovers that result rather than when it produces plausible-looking output. Results are reported per model and per harness rather than per model alone, because the harness an agent runs under changes its score substantially. Scores use a 0-100 percentage scale and are not comparable across benchmarks; see the linked repository and paper for task construction, grading, and confidence-interval methodology.
| Rank | Model | Harness | Provider | Score |
|---|---|---|---|---|
| 1 | Claude Opus 5 | Pi | Anthropic | 51.5% |
| 2 | GPT-6 Astra | Pi | OpenAI | 48.5% |
| 3 | Grok 4.6 | Grok Build | SpaceXAI | 48.0% |
| 4 | Grok 4.6 | Pi | SpaceXAI | 47.5% |
| 5 | Claude Opus 5 | Claude Code | Anthropic | 43.6% |
| 6 | GPT-6 Astra | Codex | OpenAI | 35.2% |
Methodology and provenance
Scores use a 0–100 percentage scale. See the benchmark repository, hosted paper PDF, and original publication for task construction, grading, aggregation, and uncertainty details.