# benchmarks.bio > Open agentic AI benchmarks on messy, real-world biological data. V1 data updated: 2026-09-20. Artifact generated: 2026-09-21T20:33:37.065Z. V1 result files are generated from the same sources used by the interactive V1 leaderboard. Overall score is the equally weighted mean of observed benchmark scores. Missing scores are excluded rather than treated as zero. ## Results versions V2 is newer but has fewer tested frontier models; V1 has broader historical coverage. The versions use different evaluations and should not be compared as a single time series. The interactive Overall/Capabilities view defaults to V1, while BioSecBench-Refusal defaults to V2. V2 data updated: 2026-09-21; source: V2 score arrays in src/App.tsx, checked against the per-evaluation files in public/data/v2/. BioSecBench-Refusal scores come from public/data/v2/refusalbench.json. Overall weights come from public/data/aggregate-leaderboard.json. - [Version guide](https://benchmarks.bio/results/) - [Machine-readable version manifest](https://benchmarks.bio/results/versions.json) ## Cross-benchmark results - [V1 leaderboard](https://benchmarks.bio/results/v1/): Static HTML table that does not require JavaScript. - [V1 results as JSON](https://benchmarks.bio/results/v1/latest.json): Complete structured leaderboard and per-benchmark scores. - [V1 results as CSV](https://benchmarks.bio/results/v1/latest.csv): Flat analysis-friendly capabilities leaderboard. - [V1 results as Markdown](https://benchmarks.bio/results/v1/index.md): Result definitions, caveats, benchmark inventory, and leaderboard. - [V2 results](https://benchmarks.bio/results/v2/): Separate clean-v1 rerun leaderboard; [JSON](https://benchmarks.bio/results/v2/latest.json) and [Markdown](https://benchmarks.bio/results/v2/index.md). ## Benchmarks - [SpatialBench](https://benchmarks.bio/benchmarks/spatialbench/index.md) ([paper PDF](https://benchmarks.bio/papers/spatialbench.pdf)): Agentic analysis of real spatial transcriptomics datasets across platforms and biological task types. Published score: Verified subset (115 evaluations); Full benchmark: 159 evaluations. - [SpatialBench-Long](https://benchmarks.bio/benchmarks/spatialbench-long/index.md) ([paper PDF](https://benchmarks.bio/papers/spatialbench-long.pdf)): Long-horizon agentic analysis of spatial biology datasets requiring multi-step biological reasoning. Published score: Verified subset (22 evaluations); Full benchmark: 24 evaluations. - [scBench](https://benchmarks.bio/benchmarks/scbench/index.md) ([paper PDF](https://benchmarks.bio/papers/scbench.pdf)): Agentic analysis of real single-cell RNA-seq datasets across platforms and task categories. Published score: Full benchmark (195 evaluations). - [scBench-Long](https://benchmarks.bio/benchmarks/scbench-long/index.md) ([paper PDF](https://benchmarks.bio/papers/scbench-long.pdf)): Long-horizon single-cell analysis tasks requiring sustained empirical work and biological reasoning. Published score: Verified merged subset (22 evaluations); Original full release: 21 evaluations. - [EpiBench](https://benchmarks.bio/benchmarks/epibench/index.md) ([paper PDF](https://benchmarks.bio/papers/epibench.pdf)): Agentic benchmark tasks drawn from real epigenomics analyses. Published score: Full benchmark (106 evaluations). - [VariantBench](https://benchmarks.bio/benchmarks/variantbench/index.md) ([paper PDF](https://benchmarks.bio/papers/variantbench.pdf)): Agentic genetic variant discovery and interpretation tasks grounded in real biological analyses. Published score: Full benchmark (118 evaluations). - [TxBench-Antibody-Discovery](https://benchmarks.bio/benchmarks/txbench-ab/index.md) ([paper PDF](https://benchmarks.bio/papers/txbench-ab.pdf)): Agentic benchmark tasks for antibody discovery and development decisions. Published score: Full benchmark (100 evaluations). - [TxBench-Preclinical-Pharmacology](https://benchmarks.bio/benchmarks/txbench-pp/index.md) ([paper PDF](https://benchmarks.bio/papers/txbench-pp.pdf)): Agentic benchmark tasks for preclinical pharmacology analysis and decision-making. Published score: Full benchmark (100 evaluations). - [TxBench-Oligo-Discovery](https://benchmarks.bio/benchmarks/txbench-od/index.md) ([paper PDF](https://benchmarks.bio/papers/txbench-oligo.pdf)): Agentic benchmark tasks for oligonucleotide discovery and development decisions. Published score: Full benchmark (113 evaluations). - [BioSecBench-Surveillance](https://benchmarks.bio/benchmarks/biosecbench-surveillance/index.md) ([paper PDF](https://benchmarks.bio/papers/biosecbench-surveillance.pdf)): Biosecurity surveillance tasks that measure an agent's ability to analyze biological threat signals. Published score: Full benchmark (102 evaluations). - [BioSecBench-Refusal](https://benchmarks.bio/benchmarks/biosecbench-refusal/index.md) ([paper PDF](https://benchmarks.bio/papers/biosecbench-refusal.pdf)): Safeguards benchmark measuring refusal behavior across harmful and routine biological requests. Published score: Full benchmark (107 evaluations). - [BioSecBench-Function](https://benchmarks.bio/benchmarks/functionbench/index.md) ([paper PDF](https://benchmarks.bio/papers/functionbench.pdf)): Biosecurity function-analysis tasks measuring capabilities on biologically sensitive questions. Published score: Full benchmark (111 evaluations). ## Optional - [Interactive leaderboard](https://benchmarks.bio/): Filterable charts, benchmark details, and example evaluations.