Historical milestone · 2022

    Training compute-optimal large language models

    Reviewed through September 18, 2026

    2022 · Historical milestone

    Training compute-optimal large language models

    Era
    2020s
    Theme
    Machine-learning foundations
    Evidence form
    Controlled scaling study
    School / paradigm
    Scaling laws / empirical optimization
    Institution / context
    Google DeepMind
    Researchers
    Jordan Hoffmann; collaborators

    Understand

    Plain-language record, transferred from the reviewed source module.

    Theory or experimental setup. Varied model size and training-token count under fixed compute budgets and trained a 70-billion-parameter test model on substantially more data.

    Result / historical claim. Reported that many large language models were undertrained and that jointly scaling parameters and data improved compute efficiency.

    Apply

    Professional implication, only where the reviewed record states one.

    The checked-in record does not state a separate professional application for this entry. The topic page places it in the wider research lineage: .

    Verify

    Evidence status, stated limitations, and the external sources this record actually carries.

    Evidence form. Controlled scaling study

    Limitation / debate. The fitted relationship was empirical, regime-dependent, and did not settle data quality, rights, or downstream reliability.

    Source status. This milestone row does not carry a primary-source URL in the approved export, and we do not have a verified link for it in our own research. We do not guess one.

    No primary-source URL is recorded for this entry in our reviewed data. Rather than manufacture a citation, we link the Implement Agentic research page that carries the record.

    Reproduce

    A reproduction tutorial is linked only when one exists for this exact record.

    A reproduction tutorial is not yet available for this entry. The closest reviewed material is .

    Cite or share

    APA-like: This historical record carries a year only, and no author or publisher of record in the checked-in data. An APA reference would have to invent that metadata.

    BibTeX: BibTeX requires an author and publication venue. Historical lineage entries store a narrative record and its source link, not structured authorship, so the field would be fabricated.

    Related