Reviewed current signal · 2026-04-13

    HORIZON: diagnosing long-horizon agent failures

    Reviewed through September 18, 2026

    2026-04-13 · Reviewed current signal

    HORIZON: diagnosing long-horizon agent failures

    Era
    Current reviewed signal
    Theme
    Agent reliability & evaluation
    Evidence form
    Preprint
    Source of record
    arXiv
    Source tier
    A
    Impact
    High
    School / paradigm
    Not recorded — current signals carry no formal school
    Application
    General long-horizon agents
    Researchers
    Not recorded

    Understand

    Plain-language record, transferred from the reviewed source module.

    What changed. HORIZON collected more than 3,100 trajectories across four domains and proposed a trajectory-grounded LLM-judge pipeline for attributing failure modes, with reported human-judge agreement of kappa 0.84.

    Technique / discovery. Cross-domain task construction, horizon-controlled evaluation, trace analysis, and validated LLM judging.

    Apply

    Professional implication, only where the reviewed record states one.

    Why it matters. The field is moving from pass/fail scores toward causal diagnosis of where an agent's trajectory breaks.

    Application. General long-horizon agents

    Verify

    Evidence status, stated limitations, and the external sources this record actually carries.

    Evidence maturity. Preprint (source tier A)

    Identified bottleneck. Longer tasks compound state, planning, and recovery errors; automated judges may still share model biases.

    Caveat / evidence note. Preprint and evolving benchmark; model names and harnesses age quickly.

    Review status. Reviewed. User requested: Yes.

    Reproduce

    A reproduction tutorial is linked only when one exists for this exact record.

    A reproduction tutorial is not yet available for this entry. The closest reviewed material is Reliability, uncertainty, and evaluation and Multi-agent coordination.

    Cite or share

    Related