Reviewed current signal · 2026-04-13
HORIZON: diagnosing long-horizon agent failures
Reviewed through September 18, 2026
2026-04-13 · Reviewed current signal
HORIZON: diagnosing long-horizon agent failures
- Era
- Current reviewed signal
- Theme
- Agent reliability & evaluation
- Evidence form
- Preprint
- Source of record
- arXiv
- Source tier
- A
- Impact
- High
- School / paradigm
- Not recorded — current signals carry no formal school
- Application
- General long-horizon agents
- Researchers
- Not recorded
Understand
Plain-language record, transferred from the reviewed source module.
What changed. HORIZON collected more than 3,100 trajectories across four domains and proposed a trajectory-grounded LLM-judge pipeline for attributing failure modes, with reported human-judge agreement of kappa 0.84.
Technique / discovery. Cross-domain task construction, horizon-controlled evaluation, trace analysis, and validated LLM judging.
Apply
Professional implication, only where the reviewed record states one.
Why it matters. The field is moving from pass/fail scores toward causal diagnosis of where an agent's trajectory breaks.
Application. General long-horizon agents
Verify
Evidence status, stated limitations, and the external sources this record actually carries.
Evidence maturity. Preprint (source tier A)
Identified bottleneck. Longer tasks compound state, planning, and recovery errors; automated judges may still share model biases.
Caveat / evidence note. Preprint and evolving benchmark; model names and harnesses age quickly.
Review status. Reviewed. User requested: Yes.
Reproduce
A reproduction tutorial is linked only when one exists for this exact record.
A reproduction tutorial is not yet available for this entry. The closest reviewed material is Reliability, uncertainty, and evaluation and Multi-agent coordination.
Cite or share
Related
- 2026-04-20HiL-Bench: does an agent know when to ask for help?
- 2025-10-06Gaia2 and Agents Research Environments
- 2026-08-27Double-blind model evaluation protects both proprietary weights and confidential prompts
- 2026-05-06Teaching AI agents to ask better questions with a world model
- 2026-07-23TERMINAL-BENCH 3.0: Harder Tasks for Better Agents
- 2026-08-26OpenAI-Hugging Face incident exposes multi-agent containment failures
