benchmarks.bio

Therapeutics · Capabilities

TxBench-Antibody-Discovery results

Agentic benchmark tasks for antibody discovery and development decisions.

Published score: Full benchmark (100 evaluations) · short horizon · data updated 2026-09-20

Published TxBench-Antibody-Discovery scores
RankModelHarnessProviderScore
1GPT-6 AstraPiOpenAI56.9%
2GPT-6 AstraCodexOpenAI55.2%
3Claude Opus 5Claude CodeAnthropic53.0%
4Grok 4.6PiSpaceXAI51.9%
5Claude Opus 5PiAnthropic51.4%
6Grok 4.6Grok BuildSpaceXAI51.0%
7Claude Opus 4.8PiAnthropic45.6%
8Gemini 3.5 FlashPiGoogle42.6%
9Claude Opus 4.8Claude CodeAnthropic42.0%
10Claude Opus 4.7Claude CodeAnthropic39.0%
11DeepSeek V4.1 FlashPiDeepSeek39.0%
12Grok 4.5PiSpaceXAI38.9%
13Claude Sonnet 5Claude CodeAnthropic37.3%
14Claude Sonnet 5PiAnthropic33.9%
15GPT-5.6 SolPiOpenAI33.8%
16GPT-5.5PiOpenAI33.6%
17GPT-5.6 SolCodexOpenAI33.3%
18GPT-5.5CodexOpenAI31.5%
19GPT-5.6 TerraPiOpenAI31.0%
20GPT-5.6 TerraCodexOpenAI29.7%
21GPT-5.6 LunaPiOpenAI19.4%
22GPT-5.6 LunaCodexOpenAI18.2%

Methodology and provenance

Scores use a 0–100 percentage scale. See the benchmark repository, hosted paper PDF, and original publication for task construction, grading, aggregation, and uncertainty details.