Reproduction tutorial · September 18, 2026
Policy-tier evaluation harness
These tutorials test the mechanism behind each paper at a practical scale. They do not reproduce frontier-scale headline results unless explicitly stated.
Source of this tutorial
This tutorial tests the mechanism behind reviewed development 02 in the immutable briefing of September 18, 2026.
Scope and limits
- Reproduction level
- synthetic policy-tier mechanism evaluation
- Claim under test
- whether tier-specific policies improve legitimate completion without increasing prohibited assistance compared with a single-policy baseline.
- Excluded claim
- this tutorial does not evaluate Anthropic’s private program, real biology risks, pathogens, sequences of concern, protected health information, wet-lab instructions, or real credentials.
- Reader, effort, and cost
- build a synthetic, non-sensitive task suite and run a single-policy baseline plus tier-specific policies over five task-order seeds.
Canonical resources
- Program announcement: https://www.anthropic.com/news/life-sciences-verification-program (opens in a new tab)
Project plan
| Step | Stage | Effort |
|---|---|---|
| 01 | Build the synthetic task suite | Before evaluation |
| 02 | Freeze baseline and tier policies | Before evaluation |
| 03 | Run randomized task orders | Five task-order seeds |
| 04 | Score policy outcomes | After the runs |
01Build the synthetic task suite(Before evaluation)
Create non-sensitive tasks split into general, verified-standard, and institutionally reviewed categories. Include no pathogens, sequences of concern, protected health information, wet-lab instructions, or real credentials.
02Freeze baseline and tier policies(Before evaluation)
Define one single-policy baseline and the tier-specific policies while holding model, prompts, and budgets fixed.
03Run randomized task orders(Five task-order seeds)
Evaluate the baseline and tier-specific conditions over five task-order seeds without changing the model, prompts, or budgets.
04Score policy outcomes(After the runs)
Measure legitimate completion, over-refusal, policy leakage, cross-tier confusion, and audit-log completeness. Success requires improved legitimate completion without increased prohibited assistance.
Controls, failure modes, and safety
Controlled variables
model, prompts, budgets, task suite, category definitions, and five task-order seeds.
Failure modes
sensitive task content, changing tiers after results, unmatched budgets, policy leakage, cross-tier confusion, or incomplete audit logs.
Safety
use synthetic, non-sensitive tasks only. Exclude pathogens, sequences of concern, protected health information, wet-lab instructions, and real credentials.
