BixBench3 Tests AI Agents on Full-Study Computational Biology
Edison Scientific's BixBench3 evaluates frontier agents on 20 study-scale computational biology tasks using large raw datasets and artifact-level grading.
BixBench3 moves biology-agent evaluation from short tasks to full studies
Edison Scientific introduced BixBench3 on August 26, 2026, a benchmark designed to test whether AI agents can execute computational-biology analyses at the scale of complete research studies rather than short isolated exercises.
The benchmark contains 20 tasks derived from published computational-biology papers. An agent receives raw data, a research objective and a high-level methodological plan, then must build the analysis pipeline and produce structured artifacts that support the study's findings.
Across the benchmark, agents work with an average of 67GB of raw data per task, with individual tasks ranging from 7GB to 241GB. The collection spans transcriptomics, epigenomics, proteomics, genomics and microbiome analyses, and six tasks combine multiple omics types.
Artifact-level grading exposes where long pipelines break
Each task asks for 4–14 graded artifacts. BixBench3 contains 138 artifacts in total, ranging from direct outputs such as read-count matrices to downstream analyses such as differential-expression results and enriched biological pathways.
The benchmark compares agent-produced artifacts with reference artifacts from the original published studies. Overall task score is the proportion of requested artifacts judged to have been faithfully reproduced.
This is deliberately strict. Edison notes that a scientifically valid alternative method can still score poorly if it produces outputs that differ from the specified published workflow. BixBench3 therefore measures execution of a defined research objective and analysis plan, not whether an agent can independently choose the best scientific question or method.
Frontier agents reach roughly half the available score
Edison evaluated 13 frontier language models across the 20 tasks. The benchmark authors report an average score of 0.48 for GPT-5.6 Sol, followed closely by Kimi K3 at 0.47, while performance varied substantially by task, data size and analysis depth.
The runs were unusually long and computationally expensive for an AI benchmark. Edison reports an average attempt lasting 6.8 hours, processing 102 million tokens and costing $43. The largest run consumed 1.07 billion tokens, ran for 24 hours and cost $525.
These are benchmark-author results under Edison's harness, datasets, model configurations and pricing assumptions. They should not be interpreted as universal measures of scientific ability or proof that the tested models can autonomously conduct a reliable biological study.
The benchmark highlights a long-horizon reliability gap
Performance generally fell as agents moved deeper through multistage analysis pipelines. Edison also found premature termination, repetitive retry loops, environment setup problems and incomplete data among the failure modes associated with weaker runs.
The important signal is not that agents have solved computational biology. It is that some frontier systems can now complete meaningful portions of very long scientific workflows while still failing often enough that expert review, reproducibility checks and validation remain essential.
BixBench3's public benchmark materials give AI-for-science teams a more demanding way to measure end-to-end execution, data handling and long-horizon reliability than short-form question answering or single-tool bioinformatics tests.
This article is built from the source material below. Open the originals for full context and the latest updates.