<!-- Generated by scripts/build-backrooms-results-table.mjs. Do not edit directly; see docs/updating-results.md. -->

# BioSecBench-Surveillance results

> Biosecurity surveillance tasks that measure an agent's ability to analyze biological threat signals.

Data updated: 2026-09-20

Artifact generated: 2026-09-21T21:43:49.599Z

- Domain: Biosecurity
- Horizon: short
- Assessment: Capabilities
- Published score scope: Full benchmark (102 evaluations)

Scores use a 0–100 percentage scale. Consult the repository and paper for benchmark-specific task construction, grading, aggregation, and uncertainty methodology. Results compare model-and-harness combinations.

| Rank | Model | Harness | Provider | Score |
| ---: | --- | --- | --- | ---: |
| 1 | Claude Opus 5 | Pi | Anthropic | 51.5% |
| 2 | GPT-6 Astra | Pi | OpenAI | 48.5% |
| 3 | Grok 4.6 | Grok Build | SpaceXAI | 48.0% |
| 4 | Grok 4.6 | Pi | SpaceXAI | 47.5% |
| 5 | Claude Opus 5 | Claude Code | Anthropic | 43.6% |
| 6 | GPT-6 Astra | Codex | OpenAI | 35.2% |

## Sources and downloads

- [Paper PDF](/papers/biosecbench-surveillance.pdf)
- [Original publication](https://latch.bio/biosecbench-surveillance)
- [Repository](https://github.com/latchbio/biosecbench-surveillance)
- [Interactive benchmark](https://benchmarks.bio/surveillance)
- [JSON](https://benchmarks.bio/benchmarks/biosecbench-surveillance/latest.json)
- [CSV](https://benchmarks.bio/benchmarks/biosecbench-surveillance/latest.csv)
- [V1 cross-benchmark leaderboard](https://benchmarks.bio/results/v1/)
