Reviewed current signal · 2026-08-26

    BixBench3 measures research-study-scale computational biology agents

    Reviewed through September 18, 2026

    2026-08-26 · Reviewed current signal

    BixBench3 measures research-study-scale computational biology agents

    Era
    Current reviewed signal
    Theme
    AI for science
    Evidence form
    Preprint
    Source of record
    arXiv / FutureHouse
    Source tier
    A
    Impact
    High
    School / paradigm
    Not recorded — current signals carry no formal school
    Application
    Computational biology, bioinformatics, scientific agents, and long-running data analysis
    Researchers
    Not recorded

    Understand

    Plain-language record, transferred from the reviewed source module.

    What changed. Across 20 studies, 138 artifacts, and 13 frontier models, the best mean score was 0.48. Average performance fell to 0.10 above 100 GB and to 0.24 for artifacts at dependency depth three or greater.

    Technique / discovery. Artifact dependency graphs, raw-data execution, expert-calibrated programmatic grading, and 24-hour containerized runs.

    Apply

    Professional implication, only where the reviewed record states one.

    Why it matters. Complete scientific workflows expose error accumulation, data-scale limits, and recovery failures hidden by isolated question-answer benchmarks.

    Application. Computational biology, bioinformatics, scientific agents, and long-running data analysis

    Verify

    Evidence status, stated limitations, and the external sources this record actually carries.

    Evidence maturity. Preprint (source tier A)

    Identified bottleneck. Large datasets, long dependency chains, environment setup, premature termination, retry loops, and scientific alternative-method grading remain hard.

    Caveat / evidence note. New arXiv preprint. Tasks prescribe the question and high-level method, programmatic grading can penalize valid alternatives, and source-paper errors may propagate.

    Review status. Reviewed. User requested: Yes.

    Reproduce

    A reproduction tutorial is linked only when one exists for this exact record.

    A reproduction tutorial is not yet available for this entry. The closest reviewed material is Agent planning and cognitive architectures and AI for science.

    Cite or share

    Related