benchmarks.bio

Choose a results version

V1 and V2 use different evaluations and scoring rules, so their ranks are not directly comparable. V2 is newer, but V1 covers more historical models. We publish them separately rather than combine their scores.

The interactive site's Overall/Capabilities view defaults to V1; BioSecBench-Refusal defaults to V2.

V1 — broader coverage

Broader historical model coverage. The default interactive Overall/Capabilities view uses V1. Available benchmark scores are equally weighted; missing scores are excluded rather than treated as zero.

23 model-and-harness rows · data updated 2026-09-20

V2 — newer, narrower coverage

Newer clean-v1 reruns and updated evaluation sets, currently tested on fewer frontier models. The V2 overall view requires complete coverage and uses configured evaluation-count weights. BioSecBench-Refusal defaults to V2 in the interactive site.

9 complete-coverage model-and-harness rows · data updated 2026-09-21

Source/provenance: V2 score arrays in src/App.tsx, checked against the per-evaluation files in public/data/v2/. BioSecBench-Refusal scores come from public/data/v2/refusalbench.json. Overall weights come from public/data/aggregate-leaderboard.json.