Historical milestone · 1991
Dyna architecture
Reviewed through September 18, 2026
1991 · Historical milestone
Dyna architecture
- Era
- 1990s
- Theme
- Agent planning & cognitive architectures
- Evidence form
- Used learned experience both to update a value/policy and to train a model that generates simulated planning updates
- School / paradigm
- Integrated learning and planning
- Institution / context
- GTE Laboratories
- Researchers
- Richard Sutton
Researcher index
Richard Sutton
Reinforcement learning · University of Massachusetts / Alberta
TD learning and Dyna
Why it still matters. Unified bootstrapped prediction, model learning, and planning from experience.
Representative source for this researcher — not necessarily the source of this milestone: https://doi.org/10.1007/BF00115009 (opens in a new tab)
School of thought
Reinforcement learning and adaptive agents
Matched on representative researcher.
Intelligence is learned through temporally extended interaction, reward, exploration, and improvement from experience.
Critique. Reward specification, exploration, delayed credit, instability, and unsafe trial-and-error.
Modern descendants. Agent post-training, reward modeling, planning with world models, robotics, and online adaptation.
Understand
Plain-language record, transferred from the reviewed source module.
Theory or experimental setup. Unified model-free learning, model learning, and planning in one architecture.
Result / historical claim. Model errors can bias imagined experience; planning budgets and exploration remain difficult.
Apply
Professional implication, only where the reviewed record states one.
The checked-in record does not state a separate professional application for this entry. The topic page places it in the wider research lineage: .
Verify
Evidence status, stated limitations, and the external sources this record actually carries.
Evidence form. Used learned experience both to update a value/policy and to train a model that generates simulated planning updates
Limitation / debate. World-model agents, synthetic rollouts, experience replay, and planning with learned dynamics.
Source status. This milestone row does not carry a primary-source URL in the approved export, and we do not have a verified link for it in our own research. We do not guess one.
No primary-source URL is recorded for this entry in our reviewed data. Rather than manufacture a citation, we link the Implement Agentic research page that carries the record.
Reproduce
A reproduction tutorial is linked only when one exists for this exact record.
A reproduction tutorial is not yet available for this entry. The closest reviewed material is .
Cite or share
APA-like: This historical record carries a year only, and no author or publisher of record in the checked-in data. An APA reference would have to invent that metadata.
BibTeX: BibTeX requires an author and publication venue. Historical lineage entries store a narrative record and its source link, not structured authorship, so the field would be fabricated.
