Historical milestone · 2017

    Deep reinforcement learning from human preferences

    Reviewed through September 18, 2026

    2017 · Historical milestone

    Deep reinforcement learning from human preferences

    Era
    2010s
    Theme
    Safety, security & alignment
    Evidence form
    Human-in-the-loop experiments
    School / paradigm
    Preference learning / alignment
    Institution / context
    OpenAI; DeepMind
    Researchers
    Paul Christiano; collaborators

    Understand

    Plain-language record, transferred from the reviewed source module.

    Theory or experimental setup. Learned reward functions from pairwise human comparisons and optimized agents on simulated control and Atari tasks.

    Result / historical claim. Demonstrated that sparse preference feedback could train complex behavior without a hand-specified reward.

    Apply

    Professional implication, only where the reviewed record states one.

    The checked-in record does not state a separate professional application for this entry. The topic page places it in the wider research lineage: Safety, security, and alignment.

    Verify

    Evidence status, stated limitations, and the external sources this record actually carries.

    Evidence form. Human-in-the-loop experiments

    Limitation / debate. Learned rewards can be incomplete or exploitable and depend on labeler consistency and coverage.

    Source status. The source link below is the verified link our reviewed topic research already carries for this milestone.

    Reproduce

    A reproduction tutorial is linked only when one exists for this exact record.

    A reproduction tutorial is not yet available for this entry. The closest reviewed material is Safety, security, and alignment.

    Cite or share

    APA-like: This historical record carries a year only, and no author or publisher of record in the checked-in data. An APA reference would have to invent that metadata.

    BibTeX: BibTeX requires an author and publication venue. Historical lineage entries store a narrative record and its source link, not structured authorship, so the field would be fabricated.

    Related