Reproduction tutorial · September 18, 2026
Full-duplex dialogue benchmark
These tutorials test the mechanism behind each paper at a practical scale. They do not reproduce frontier-scale headline results unless explicitly stated.
Source of this tutorial
This tutorial tests the mechanism behind reviewed development 04 in the immutable briefing of September 18, 2026.
Scope and limits
- Reproduction level
- controlled dialogue benchmark
- Claim under test
- whether streaming full-duplex interaction improves dialogue-task performance under interruption, correction, tool delay, long-session memory, and safety constraints.
- Excluded claim
- this tutorial does not reproduce Google’s internal model-card evaluations or authorize recording private conversations.
- Reader, effort, and cost
- create 40 consented synthetic or openly licensed dialogue tasks and compare two conditions over five randomized orders.
Canonical resources
Project plan
| Step | Stage | Effort |
|---|---|---|
| 01 | Create and freeze dialogue tasks | Before evaluation |
| 02 | Match the two conditions | Before evaluation |
| 03 | Run randomized orders | Five randomized orders |
| 04 | Measure and report | After the runs |
01Create and freeze dialogue tasks(Before evaluation)
Create 40 consented synthetic or openly licensed dialogue tasks covering interruption, correction, tool delay, long-session memory, and safety boundaries. Do not record private conversations.
02Match the two conditions(Before evaluation)
Compare standard turn-taking with a streaming condition under identical models and task budgets.
03Run randomized orders(Five randomized orders)
Run both conditions across five randomized task orders while preserving identical model and task budgets.
04Measure and report(After the runs)
Measure first-audio latency, interruption stop time, correction recovery, task success, memory accuracy, and unsafe continuation. Report distributions and paired intervals.
Controls, failure modes, and safety
Controlled variables
model, task budget, 40-task set, condition definitions, and five randomized task orders.
Failure modes
unmatched model or task budgets, private recordings, inconsistent interruption timing, changing tasks after results, or reporting averages without distributions and paired intervals.
Safety
use only consented synthetic or openly licensed dialogue. Do not record private conversations, and preserve the safety-boundary tasks in both conditions.
