Analysis
Analysis

BixBench3 Tests AI Agents on Full-Study Computational Biology

Published Aug 26, 2026 Sources checked Aug 27, 2026

Edison Scientific's BixBench3 evaluates frontier agents on 20 study-scale computational biology tasks using large raw datasets and artifact-level grading.

BixBench3 moves biology-agent evaluation from short tasks to full studies

Edison Scientific introduced BixBench3 on August 26, 2026, a benchmark designed to test whether AI agents can execute computational-biology analyses at the scale of complete research studies rather than short isolated exercises.

The benchmark contains 20 tasks derived from published computational-biology papers. An agent receives raw data, a research objective and a high-level methodological plan, then must build the analysis pipeline and produce structured artifacts that support the study's findings.

Across the benchmark, agents work with an average of 67GB of raw data per task, with individual tasks ranging from 7GB to 241GB. The collection spans transcriptomics, epigenomics, proteomics, genomics and microbiome analyses, and six tasks combine multiple omics types.

Artifact-level grading exposes where long pipelines break

Each task asks for 4–14 graded artifacts. BixBench3 contains 138 artifacts in total, ranging from direct outputs such as read-count matrices to downstream analyses such as differential-expression results and enriched biological pathways.

The benchmark compares agent-produced artifacts with reference artifacts from the original published studies. Overall task score is the proportion of requested artifacts judged to have been faithfully reproduced.

This is deliberately strict. Edison notes that a scientifically valid alternative method can still score poorly if it produces outputs that differ from the specified published workflow. BixBench3 therefore measures execution of a defined research objective and analysis plan, not whether an agent can independently choose the best scientific question or method.

Frontier agents reach roughly half the available score

Edison evaluated 13 frontier language models across the 20 tasks. The benchmark authors report an average score of 0.48 for GPT-5.6 Sol, followed closely by Kimi K3 at 0.47, while performance varied substantially by task, data size and analysis depth.

The runs were unusually long and computationally expensive for an AI benchmark. Edison reports an average attempt lasting 6.8 hours, processing 102 million tokens and costing $43. The largest run consumed 1.07 billion tokens, ran for 24 hours and cost $525.

These are benchmark-author results under Edison's harness, datasets, model configurations and pricing assumptions. They should not be interpreted as universal measures of scientific ability or proof that the tested models can autonomously conduct a reliable biological study.

The benchmark highlights a long-horizon reliability gap

Performance generally fell as agents moved deeper through multistage analysis pipelines. Edison also found premature termination, repetitive retry loops, environment setup problems and incomplete data among the failure modes associated with weaker runs.

The important signal is not that agents have solved computational biology. It is that some frontier systems can now complete meaningful portions of very long scientific workflows while still failing often enough that expert review, reproducibility checks and validation remain essential.

BixBench3's public benchmark materials give AI-for-science teams a more demanding way to measure end-to-end execution, data handling and long-horizon reliability than short-form question answering or single-tool bioinformatics tests.

Sources

This article is built from the source material below. Open the originals for full context and the latest updates.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books