Overall Capabilities score is the weighted mean of each V2 benchmark's macro-mean pass rate. The weights are the configured evaluation counts in the aggregate leaderboard metadata; these can differ from the number of V2 evaluation records. Safeguards is the BioSecBench-Refusal score.
Complete-coverage model-and-harness results. Data updated 2026-09-21; artifact generated 2026-09-21T20:33:37.065Z. V2 uses clean-v1 reruns and updated evaluation sets, so scores should not be treated as a direct continuation of V1.
Source/provenance: V2 score arrays in src/App.tsx, checked against the per-evaluation files in public/data/v2/. BioSecBench-Refusal scores come from public/data/v2/refusalbench.json. Overall weights come from public/data/aggregate-leaderboard.json.
Capabilities
Capabilities leaderboard across 11 benchmarks
Rank
Model
Harness
Provider
Score
1
GPT-6 Astra
Pi
OpenAI
51.4%
2
Claude Opus 5
Pi
Anthropic
49.5%
3
GPT-6 Astra
Codex
OpenAI
48.9%
4
Claude Opus 5
Claude Code
Anthropic
48.4%
5
Grok 4.6
Pi
SpaceXAI
47.8%
6
Grok 4.7
Grok Build
SpaceXAI
46.8%
7
Grok 4.6
Grok Build
SpaceXAI
45.9%
8
Gemini 3.7 Flash
Pi
Google
41.2%
9
Gemini 3.8 Flash
Pi
Google
38.5%
Safeguards
BioSecBench-Refusal
Rank
Model
Harness
Provider
Score
1
Grok 4.7
Grok Build
SpaceXAI
62.4%
2
Gemini 3.7 Flash
Pi
Google
54.8%
3
Gemini 3.8 Flash
Pi
Google
47.4%
4
Grok 4.6
Pi
SpaceXAI
46.9%
5
Grok 4.6
Grok Build
SpaceXAI
45.6%
6
GPT-6 Astra
Pi
OpenAI
40.8%
7
Claude Opus 5
Claude Code
Anthropic
31.8%
8
GPT-6 Astra
Codex
OpenAI
25.5%
9
Claude Opus 5
Pi
Anthropic
24.5%
Included benchmarks
V2 evaluation counts can differ from the configured counts used to weight the interactive overall score.