Historical milestone · 2025

    Measuring AI ability to complete long tasks

    Reviewed through September 18, 2026

    2025 · Historical milestone

    Measuring AI ability to complete long tasks

    Era
    2020s
    Theme
    Reliability, uncertainty & evaluation
    Evidence form
    Empirical benchmark + trend model
    School / paradigm
    Agent evaluation / task horizons
    Institution / context
    METR
    Researchers
    Megan Kinniment; collaborators

    Understand

    Plain-language record, transferred from the reviewed source module.

    Theory or experimental setup. Estimated the human-expert duration of software tasks that frontier agents could complete with 50 percent success and fitted a historical trend.

    Result / historical claim. Made task duration a measurable axis of agent capability and reported rapid growth on the selected task distribution.

    Apply

    Professional implication, only where the reviewed record states one.

    The checked-in record does not state a separate professional application for this entry. The topic page places it in the wider research lineage: .

    Verify

    Evidence status, stated limitations, and the external sources this record actually carries.

    Evidence form. Empirical benchmark + trend model

    Limitation / debate. The task mix was software-heavy, duration estimates were uncertain, and trend extrapolation does not prove deployment reliability.

    Source status. This milestone row does not carry a primary-source URL in the approved export, and we do not have a verified link for it in our own research. We do not guess one.

    No primary-source URL is recorded for this entry in our reviewed data. Rather than manufacture a citation, we link the Implement Agentic research page that carries the record.

    Reproduce

    A reproduction tutorial is linked only when one exists for this exact record.

    A reproduction tutorial is not yet available for this entry. The closest reviewed material is .

    Cite or share

    APA-like: This historical record carries a year only, and no author or publisher of record in the checked-in data. An APA reference would have to invent that metadata.

    BibTeX: BibTeX requires an author and publication venue. Historical lineage entries store a narrative record and its source link, not structured authorship, so the field would be fabricated.

    Related