benchmarks.bio

V1 results

V1 cross-benchmark results

Overall score is the equally weighted mean of observed benchmark scores.

Data updated 2026-09-20; artifact generated 2026-09-21T20:01:04.986Z. Missing scores are excluded from a model's mean and are never treated as zero.

Capabilities leaderboard across 11 benchmarks
RankModelHarnessProviderOverall score
1GPT-6 AstraPiOpenAI51.3%
2GPT-6 AstraCodexOpenAI51.1%
3Claude Opus 5Claude CodeAnthropic48.6%
4Claude Opus 5PiAnthropic48.2%
5Grok 4.6Grok BuildSpaceXAI48.2%
6GPT-5.6 SolPiOpenAI47.5%
7GPT-5.6 SolCodexOpenAI45.9%
8DeepSeek V4.1 FlashPiDeepSeek45.0%
9Grok 4.6PiSpaceXAI44.8%
10Claude Opus 4.8PiAnthropic44.8%
11Claude Opus 4.8Claude CodeAnthropic42.8%
12GPT-5.6 TerraPiOpenAI42.4%
13Gemini 3.5 FlashPiGoogle41.7%
14GPT-5.6 TerraCodexOpenAI41.5%
15Grok 4.5PiSpaceXAI41.3%
16GPT-5.5PiOpenAI40.1%
17Claude Opus 4.7Claude CodeAnthropic39.8%
18Claude Sonnet 5PiAnthropic39.5%
19Claude Sonnet 5Claude CodeAnthropic39.1%
20GPT-5.5CodexOpenAI38.7%
21Kimi K3PiMoonshot38.1%
22GPT-5.6 LunaPiOpenAI33.8%
23GPT-5.6 LunaCodexOpenAI32.2%

Included benchmarks

  • SpatialBench — Omics; published score: Verified subset (115 evaluations); Full benchmark: 159 evaluations
  • SpatialBench-Long — Omics; published score: Verified subset (22 evaluations); Full benchmark: 24 evaluations
  • scBench — Omics; published score: Full benchmark (195 evaluations)
  • scBench-Long — Omics; published score: Verified merged subset (22 evaluations); Original full release: 21 evaluations
  • EpiBench — Omics; published score: Full benchmark (106 evaluations)
  • VariantBench — Omics; published score: Full benchmark (118 evaluations)
  • TxBench-Antibody-Discovery — Therapeutics; published score: Full benchmark (100 evaluations)
  • TxBench-Preclinical-Pharmacology — Therapeutics; published score: Full benchmark (100 evaluations)
  • TxBench-Oligo-Discovery — Therapeutics; published score: Full benchmark (113 evaluations)
  • BioSecBench-Surveillance — Biosecurity; published score: Full benchmark (102 evaluations)
  • BioSecBench-Refusal — Biosecurity; published score: Full benchmark (107 evaluations)
  • BioSecBench-Function — Biosecurity; published score: Full benchmark (111 evaluations)