Reproduction tutorial · September 18, 2026

    Full-duplex dialogue benchmark

    These tutorials test the mechanism behind each paper at a practical scale. They do not reproduce frontier-scale headline results unless explicitly stated.

    Source of this tutorial

    This tutorial tests the mechanism behind reviewed development 04 in the immutable briefing of September 18, 2026.

    Scope and limits

    Reproduction level
    controlled dialogue benchmark
    Claim under test
    whether streaming full-duplex interaction improves dialogue-task performance under interruption, correction, tool delay, long-session memory, and safety constraints.
    Excluded claim
    this tutorial does not reproduce Google’s internal model-card evaluations or authorize recording private conversations.
    Reader, effort, and cost
    create 40 consented synthetic or openly licensed dialogue tasks and compare two conditions over five randomized orders.

    Canonical resources

    Project plan

    Full-duplex dialogue benchmark — step and effort summary
    StepStageEffort
    01Create and freeze dialogue tasksBefore evaluation
    02Match the two conditionsBefore evaluation
    03Run randomized ordersFive randomized orders
    04Measure and reportAfter the runs
    1. 01Create and freeze dialogue tasks(Before evaluation)

      Create 40 consented synthetic or openly licensed dialogue tasks covering interruption, correction, tool delay, long-session memory, and safety boundaries. Do not record private conversations.

    2. 02Match the two conditions(Before evaluation)

      Compare standard turn-taking with a streaming condition under identical models and task budgets.

    3. 03Run randomized orders(Five randomized orders)

      Run both conditions across five randomized task orders while preserving identical model and task budgets.

    4. 04Measure and report(After the runs)

      Measure first-audio latency, interruption stop time, correction recovery, task success, memory accuracy, and unsafe continuation. Report distributions and paired intervals.

    Controls, failure modes, and safety

    Controlled variables

    model, task budget, 40-task set, condition definitions, and five randomized task orders.

    Failure modes

    unmatched model or task budgets, private recordings, inconsistent interruption timing, changing tasks after results, or reporting averages without distributions and paired intervals.

    Safety

    use only consented synthetic or openly licensed dialogue. Do not record private conversations, and preserve the safety-boundary tasks in both conditions.

    Affected topics

    Back to the September 18, 2026 briefing