Historical milestone · 1988

    Temporal-difference learning

    Reviewed through September 18, 2026

    1988 · Historical milestone

    Temporal-difference learning

    Era
    1980s
    Theme
    Machine-learning foundations
    Evidence form
    Predicted future reward and updated estimates from successive predictions without waiting for a final outcome
    School / paradigm
    Reinforcement learning
    Institution / context
    University of Massachusetts Amherst
    Researchers
    Richard Sutton

    Researcher index

    Richard Sutton

    Reinforcement learning · University of Massachusetts / Alberta

    TD learning and Dyna

    Why it still matters. Unified bootstrapped prediction, model learning, and planning from experience.

    Representative source for this researcher — not necessarily the source of this milestone: https://doi.org/10.1007/BF00115009 (opens in a new tab)

    School of thought

    Reinforcement learning and adaptive agents

    Matched on representative researcher.

    Intelligence is learned through temporally extended interaction, reward, exploration, and improvement from experience.

    Critique. Reward specification, exploration, delayed credit, instability, and unsafe trial-and-error.

    Modern descendants. Agent post-training, reward modeling, planning with world models, robotics, and online adaptation.

    Understand

    Plain-language record, transferred from the reviewed source module.

    Theory or experimental setup. Unified Monte Carlo and dynamic-programming ideas and enabled online learning from incomplete trajectories.

    Result / historical claim. Function approximation, off-policy learning, delayed credit, and instability complicate large systems.

    Apply

    Professional implication, only where the reviewed record states one.

    The checked-in record does not state a separate professional application for this entry. The topic page places it in the wider research lineage: .

    Verify

    Evidence status, stated limitations, and the external sources this record actually carries.

    Evidence form. Predicted future reward and updated estimates from successive predictions without waiting for a final outcome

    Limitation / debate. Value learning, reward models, agent learning from trajectories, and bootstrapped critics.

    Source status. This milestone row does not carry a primary-source URL in the approved export, and we do not have a verified link for it in our own research. We do not guess one.

    No primary-source URL is recorded for this entry in our reviewed data. Rather than manufacture a citation, we link the Implement Agentic research page that carries the record.

    Reproduce

    A reproduction tutorial is linked only when one exists for this exact record.

    A reproduction tutorial is not yet available for this entry. The closest reviewed material is .

    Cite or share

    APA-like: This historical record carries a year only, and no author or publisher of record in the checked-in data. An APA reference would have to invent that metadata.

    BibTeX: BibTeX requires an author and publication venue. Historical lineage entries store a narrative record and its source link, not structured authorship, so the field would be fabricated.

    Related