Historical milestone · 1992
Q-learning convergence
Reviewed through September 18, 2026
1992 · Historical milestone
Q-learning convergence
- Era
- 1990s
- Theme
- Machine-learning foundations
- Evidence form
- Off-policy temporal-difference control learned action values from sampled transitions without a model of the environment
- School / paradigm
- Reinforcement learning
- Institution / context
- Cambridge University
- Researchers
- Christopher Watkins; Peter Dayan
School of thought
Reinforcement learning and adaptive agents
Matched on representative researcher.
Intelligence is learned through temporally extended interaction, reward, exploration, and improvement from experience.
Critique. Reward specification, exploration, delayed credit, instability, and unsafe trial-and-error.
Modern descendants. Agent post-training, reward modeling, planning with world models, robotics, and online adaptation.
Understand
Plain-language record, transferred from the reviewed source module.
Theory or experimental setup. Proved convergence under tabular assumptions and established a canonical model-free control algorithm.
Result / historical claim. Large state spaces, function approximation, exploration, and nonstationarity break simple guarantees.
Apply
Professional implication, only where the reviewed record states one.
The checked-in record does not state a separate professional application for this entry. The topic page places it in the wider research lineage: .
Verify
Evidence status, stated limitations, and the external sources this record actually carries.
Evidence form. Off-policy temporal-difference control learned action values from sampled transitions without a model of the environment
Limitation / debate. Agent post-training, tool-use policies, offline RL, and value-guided decision making.
Source status. This milestone row does not carry a primary-source URL in the approved export, and we do not have a verified link for it in our own research. We do not guess one.
No primary-source URL is recorded for this entry in our reviewed data. Rather than manufacture a citation, we link the Implement Agentic research page that carries the record.
Reproduce
A reproduction tutorial is linked only when one exists for this exact record.
A reproduction tutorial is not yet available for this entry. The closest reviewed material is .
Cite or share
APA-like: This historical record carries a year only, and no author or publisher of record in the checked-in data. An APA reference would have to invent that metadata.
BibTeX: BibTeX requires an author and publication venue. Historical lineage entries store a narrative record and its source link, not structured authorship, so the field would be fabricated.
