Reviewed current signal · 2026-08-26
BixBench3 measures research-study-scale computational biology agents
Reviewed through September 18, 2026
2026-08-26 · Reviewed current signal
BixBench3 measures research-study-scale computational biology agents
- Era
- Current reviewed signal
- Theme
- AI for science
- Evidence form
- Preprint
- Source of record
- arXiv / FutureHouse
- Source tier
- A
- Impact
- High
- School / paradigm
- Not recorded — current signals carry no formal school
- Application
- Computational biology, bioinformatics, scientific agents, and long-running data analysis
- Researchers
- Not recorded
Understand
Plain-language record, transferred from the reviewed source module.
What changed. Across 20 studies, 138 artifacts, and 13 frontier models, the best mean score was 0.48. Average performance fell to 0.10 above 100 GB and to 0.24 for artifacts at dependency depth three or greater.
Technique / discovery. Artifact dependency graphs, raw-data execution, expert-calibrated programmatic grading, and 24-hour containerized runs.
Apply
Professional implication, only where the reviewed record states one.
Why it matters. Complete scientific workflows expose error accumulation, data-scale limits, and recovery failures hidden by isolated question-answer benchmarks.
Application. Computational biology, bioinformatics, scientific agents, and long-running data analysis
Verify
Evidence status, stated limitations, and the external sources this record actually carries.
Evidence maturity. Preprint (source tier A)
Identified bottleneck. Large datasets, long dependency chains, environment setup, premature termination, retry loops, and scientific alternative-method grading remain hard.
Caveat / evidence note. New arXiv preprint. Tasks prescribe the question and high-level method, programmatic grading can penalize valid alternatives, and source-paper errors may propagate.
Review status. Reviewed. User requested: Yes.
Reproduce
A reproduction tutorial is linked only when one exists for this exact record.
A reproduction tutorial is not yet available for this entry. The closest reviewed material is Agent planning and cognitive architectures and AI for science.
Cite or share
Related
- 2026-08-26Planetary Prediction Engine automates geospatial data-to-model workflows
- 2026-05-27AI is transforming scientific discovery
- 2026-06-30Claude Science, an AI workbench for scientists
- 2026-06-30DrugDiscoveryBench
- 2026-03-10From games to biology and beyond: 10 years of AlphaGo
- 2026-02-10IsoDDE moves beyond structure prediction
