benchmarks.bio

Therapeutics · Capabilities

TxBench-Preclinical-Pharmacology results

Agentic benchmark tasks for preclinical pharmacology analysis and decision-making.

Published score: Full benchmark (100 evaluations) · short horizon · data updated 2026-09-20

Published TxBench-Preclinical-Pharmacology scores
RankModelHarnessProviderScore
1Claude Opus 5Claude CodeAnthropic68.7%
2Grok 4.6Grok BuildSpaceXAI64.7%
3Claude Opus 5PiAnthropic63.3%
4Grok 4.6PiSpaceXAI62.0%
5GPT-5.6 SolPiOpenAI61.7%
6GPT-6 AstraCodexOpenAI61.7%
7GPT-6 AstraPiOpenAI60.3%
8Claude Opus 4.8PiAnthropic59.3%
9GPT-5.6 SolCodexOpenAI59.0%
10Grok 4.5PiSpaceXAI57.7%
11GPT-5.5PiOpenAI55.3%
12Claude Opus 4.8Claude CodeAnthropic54.7%
13DeepSeek V4.1 FlashPiDeepSeek54.3%
14GPT-5.6 TerraPiOpenAI54.0%
15GPT-5.6 TerraCodexOpenAI51.7%
16Gemini 3.5 FlashPiGoogle51.3%
17Claude Opus 4.7PiAnthropic49.3%
18Claude Sonnet 5Claude CodeAnthropic48.3%
19GPT-5.5CodexOpenAI47.3%
20GPT-5.6 LunaPiOpenAI46.7%
21Kimi K3PiMoonshot44.7%
22Claude Opus 4.7Claude CodeAnthropic44.0%
23Claude Sonnet 5PiAnthropic42.0%
24GPT-5.6 LunaCodexOpenAI42.0%

Methodology and provenance

Scores use a 0–100 percentage scale. See the benchmark repository, hosted paper PDF, and original publication for task construction, grading, aggregation, and uncertainty details.