Reproduction tutorial · September 18, 2026

    Agent-oversight metrics audit

    These tutorials test the mechanism behind each paper at a practical scale. They do not reproduce frontier-scale headline results unless explicitly stated.

    Source of this tutorial

    This tutorial tests the mechanism behind reviewed development 01 in the immutable briefing of September 18, 2026.

    Scope and limits

    Reproduction level
    mechanism audit, not a reproduction of Anthropic’s internal rates
    Claim under test
    whether online and offline monitor performance can be measured with complete denominators for coverage, latency, blocking, detection, escalation, and human-review load.
    Excluded claim
    this tutorial does not reproduce Anthropic’s internal rates or establish that its reported measures are independently verified.
    Reader, effort, and cost
    instrument a disposable agent harness; run 1,000 benign tasks plus preregistered policy-violation fixtures across five seeds.

    Canonical resources

    Project plan

    Agent-oversight metrics audit — step and effort summary
    StepStageEffort
    01Freeze the audit protocolBefore any run
    02Run online and offline monitorsFive seeds
    03Measure oversight performanceDuring every run
    04Publish the audit recordAfter the runs
    1. 01Freeze the audit protocol(Before any run)

      Define the 1,000 benign tasks, preregister the policy-violation fixtures, freeze monitor versions, and specify five seeds before collecting results.

    2. 02Run online and offline monitors(Five seeds)

      Instrument a disposable agent harness with online and offline monitors. Run the fixed benign tasks and preregistered policy-violation fixtures across all five seeds.

    3. 03Measure oversight performance(During every run)

      Record action coverage, decision latency, block rate, false positives, false negatives, escalation latency, and human-review load with complete denominators.

    4. 04Publish the audit record(After the runs)

      Publish denominators, logs, intervals, and frozen monitor versions. Success requires complete coverage and prespecified detection without unacceptable benign blocking.

    Controls, failure modes, and safety

    Controlled variables

    task set, violation fixtures, monitor versions, seeds, action definitions, and metric denominators.

    Failure modes

    missing actions from coverage denominators, changing fixtures after results, pooling away seed variation, or reporting blocks without false-positive and false-negative rates.

    Safety

    use only benign tasks and preregistered policy-violation fixtures in a disposable harness; do not test live systems, accounts, or harmful capabilities.

    Affected topics

    Back to the September 18, 2026 briefing