V1 — broader coverage
Broader historical model coverage. The default interactive Overall/Capabilities view uses V1. Available benchmark scores are equally weighted; missing scores are excluded rather than treated as zero.
23 model-and-harness rows · data updated 2026-09-20
V2 — newer, narrower coverage
Newer clean-v1 reruns and updated evaluation sets, currently tested on fewer frontier models. The V2 overall view requires complete coverage and uses configured evaluation-count weights. BioSecBench-Refusal defaults to V2 in the interactive site.
9 complete-coverage model-and-harness rows · data updated 2026-09-21
Source/provenance: V2 score arrays in src/App.tsx, checked against the per-evaluation files in public/data/v2/. BioSecBench-Refusal scores come from public/data/v2/refusalbench.json. Overall weights come from public/data/aggregate-leaderboard.json.