Historical milestone · 1991

    Dyna architecture

    Reviewed through September 18, 2026

    1991 · Historical milestone

    Dyna architecture

    Era
    1990s
    Theme
    Agent planning & cognitive architectures
    Evidence form
    Used learned experience both to update a value/policy and to train a model that generates simulated planning updates
    School / paradigm
    Integrated learning and planning
    Institution / context
    GTE Laboratories
    Researchers
    Richard Sutton

    Researcher index

    Richard Sutton

    Reinforcement learning · University of Massachusetts / Alberta

    TD learning and Dyna

    Why it still matters. Unified bootstrapped prediction, model learning, and planning from experience.

    Representative source for this researcher — not necessarily the source of this milestone: https://doi.org/10.1007/BF00115009 (opens in a new tab)

    School of thought

    Reinforcement learning and adaptive agents

    Matched on representative researcher.

    Intelligence is learned through temporally extended interaction, reward, exploration, and improvement from experience.

    Critique. Reward specification, exploration, delayed credit, instability, and unsafe trial-and-error.

    Modern descendants. Agent post-training, reward modeling, planning with world models, robotics, and online adaptation.

    Understand

    Plain-language record, transferred from the reviewed source module.

    Theory or experimental setup. Unified model-free learning, model learning, and planning in one architecture.

    Result / historical claim. Model errors can bias imagined experience; planning budgets and exploration remain difficult.

    Apply

    Professional implication, only where the reviewed record states one.

    The checked-in record does not state a separate professional application for this entry. The topic page places it in the wider research lineage: .

    Verify

    Evidence status, stated limitations, and the external sources this record actually carries.

    Evidence form. Used learned experience both to update a value/policy and to train a model that generates simulated planning updates

    Limitation / debate. World-model agents, synthetic rollouts, experience replay, and planning with learned dynamics.

    Source status. This milestone row does not carry a primary-source URL in the approved export, and we do not have a verified link for it in our own research. We do not guess one.

    No primary-source URL is recorded for this entry in our reviewed data. Rather than manufacture a citation, we link the Implement Agentic research page that carries the record.

    Reproduce

    A reproduction tutorial is linked only when one exists for this exact record.

    A reproduction tutorial is not yet available for this entry. The closest reviewed material is .

    Cite or share

    APA-like: This historical record carries a year only, and no author or publisher of record in the checked-in data. An APA reference would have to invent that metadata.

    BibTeX: BibTeX requires an author and publication venue. Historical lineage entries store a narrative record and its source link, not structured authorship, so the field would be fabricated.

    Related