benchmarks.bio

Omics · Capabilities

EpiBench results

Agentic benchmark tasks drawn from real epigenomics analyses.

Published score: Full benchmark (106 evaluations) · short horizon · data updated 2026-09-20

Published EpiBench scores
RankModelHarnessProviderScore
1GPT-5.5PiOpenAI45.0%
2GPT-6 AstraCodexOpenAI44.6%
3GPT-5.6 SolCodexOpenAI42.1%
4DeepSeek V4.1 FlashPiDeepSeek42.1%
5GPT-6 AstraPiOpenAI41.8%
6GPT-5.5CodexOpenAI39.9%
7GPT-5.6 SolPiOpenAI39.9%
8GPT-5.6 TerraPiOpenAI39.9%
9GPT-5.6 TerraCodexOpenAI39.3%
10Claude Opus 4.8PiAnthropic39.0%
11Claude Sonnet 5Claude CodeAnthropic37.7%
12Claude Opus 4.8Claude CodeAnthropic37.4%
13Grok 4.6Grok BuildSpaceXAI37.4%
14GPT-5.6 LunaPiOpenAI36.5%
15Claude Opus 5PiAnthropic35.2%
16Claude Sonnet 5PiAnthropic35.2%
17Grok 4.6PiSpaceXAI34.3%
18Claude Opus 4.7Claude CodeAnthropic31.4%
19GPT-5.6 LunaCodexOpenAI31.4%
20Claude Opus 4.7PiAnthropic31.1%
21Claude Opus 5Claude CodeAnthropic31.1%
22Grok 4.5PiSpaceXAI28.9%
23Gemini 3.5 FlashPiGoogle28.9%
24Kimi K3PiMoonshot26.7%

Methodology and provenance

Scores use a 0–100 percentage scale. See the benchmark repository, hosted paper PDF, and original publication for task construction, grading, aggregation, and uncertainty details.