<!-- Generated by scripts/build-v2-results.mjs. Do not edit directly. -->
# V2 cross-benchmark results

These are the complete-coverage V2 results, separate from the [V1 leaderboard](https://benchmarks.bio/results/v1/). Read the [version guide](https://benchmarks.bio/results/) before comparing them. The interactive V2 view is available at [benchmarks.bio/?v=2](https://benchmarks.bio/?v=2).

Overall Capabilities score is the weighted mean of each V2 benchmark's macro-mean pass rate. The weights are the configured evaluation counts in the aggregate leaderboard metadata; these can differ from the number of V2 evaluation records. Safeguards is the BioSecBench-Refusal score.

V2 uses clean-v1 reruns and updated evaluation sets. Scores are model-and-harness combinations. Only models scored on all 11 capability benchmarks and BioSecBench-Refusal appear in the overall tables.

Data updated: 2026-09-21. Artifact generated: 2026-09-21T21:43:49.599Z. These dates have different meanings; a rebuild does not change the underlying data date.

Source/provenance: V2 score arrays in src/App.tsx, checked against the per-evaluation files in public/data/v2/. BioSecBench-Refusal scores come from public/data/v2/refusalbench.json. Overall weights come from public/data/aggregate-leaderboard.json. Source paths: `src/App.tsx`, `public/data/v2/`, `public/data/aggregate-leaderboard.json`.

The interactive V2 overall view uses the configured evaluation-count weights from the aggregate metadata. In four benchmarks these weights differ from the number of evaluations in the V2 source data; the table below shows both. Thus its overall score is not precisely an equal-weighted mean across every V2 evaluation record.

## Capabilities

| Rank | Model | Harness | Provider | Overall score |
| ---: | --- | --- | --- | ---: |
| 1 | GPT-6 Astra | Pi | OpenAI | 51.4% |
| 2 | Claude Opus 5 | Pi | Anthropic | 49.5% |
| 3 | GPT-6 Astra | Codex | OpenAI | 48.9% |
| 4 | Claude Opus 5 | Claude Code | Anthropic | 48.4% |
| 5 | Grok 4.6 | Pi | SpaceXAI | 47.8% |
| 6 | Grok 4.7 | Grok Build | SpaceXAI | 46.8% |
| 7 | Grok 4.6 | Grok Build | SpaceXAI | 45.9% |
| 8 | Gemini 3.7 Flash | Pi | Google | 41.2% |
| 9 | Gemini 3.8 Flash | Pi | Google | 38.5% |

## Safeguards — BioSecBench-Refusal

| Rank | Model | Harness | Provider | Score |
| ---: | --- | --- | --- | ---: |
| 1 | Grok 4.7 | Grok Build | SpaceXAI | 62.4% |
| 2 | Gemini 3.7 Flash | Pi | Google | 54.8% |
| 3 | Gemini 3.8 Flash | Pi | Google | 47.4% |
| 4 | Grok 4.6 | Pi | SpaceXAI | 46.9% |
| 5 | Grok 4.6 | Grok Build | SpaceXAI | 45.6% |
| 6 | GPT-6 Astra | Pi | OpenAI | 40.8% |
| 7 | Claude Opus 5 | Claude Code | Anthropic | 31.8% |
| 8 | GPT-6 Astra | Codex | OpenAI | 25.5% |
| 9 | Claude Opus 5 | Pi | Anthropic | 24.5% |

## Included benchmarks

| Benchmark | V2 evaluations | Overall weight (eval count) | Assessment |
| --- | ---: | ---: | --- |
| SpatialBench | 115 | 115 | capabilities |
| SpatialBench-Long | 21 | 24 | capabilities |
| scBench | 195 | 195 | capabilities |
| scBench-Long | 22 | 21 | capabilities |
| EpiBench | 106 | 106 | capabilities |
| VariantBench | 118 | 118 | capabilities |
| TxBench-Antibody-Discovery | 137 | 100 | capabilities |
| TxBench-Preclinical-Pharmacology | 100 | 100 | capabilities |
| TxBench-Oligo-Discovery | 120 | 113 | capabilities |
| BioSecBench-Surveillance | 102 | 102 | capabilities |
| BioSecBench-Refusal | 107 | — | safeguards |
| BioSecBench-Function | 111 | 111 | capabilities |

## Downloads

- [V2 JSON](https://benchmarks.bio/results/v2/latest.json) — both leaderboards and per-benchmark scores.
- [V2 capabilities CSV](https://benchmarks.bio/results/v2/latest.csv).
- [V2 HTML](https://benchmarks.bio/results/v2/).
