Historical milestone · 2022

    Holistic Evaluation of Language Models

    Reviewed through September 18, 2026

    2022 · Historical milestone

    Holistic Evaluation of Language Models

    Era
    2020s
    Theme
    Reliability, uncertainty & evaluation
    Evidence form
    Benchmark + taxonomy
    School / paradigm
    Multi-metric evaluation
    Institution / context
    Stanford CRFM
    Researchers
    Percy Liang; collaborators

    Understand

    Plain-language record, transferred from the reviewed source module.

    Theory or experimental setup. Evaluated language models across many scenarios and metrics including accuracy, calibration, robustness, fairness, bias, toxicity, and efficiency.

    Result / historical claim. Made model choice visibly multidimensional rather than reducible to one leaderboard number.

    Apply

    Professional implication, only where the reviewed record states one.

    The checked-in record does not state a separate professional application for this entry. The topic page places it in the wider research lineage: .

    Verify

    Evidence status, stated limitations, and the external sources this record actually carries.

    Evidence form. Benchmark + taxonomy

    Limitation / debate. Coverage was necessarily incomplete and vulnerable to benchmark reuse, version drift, and disputed constructs.

    Source status. This milestone row does not carry a primary-source URL in the approved export, and we do not have a verified link for it in our own research. We do not guess one.

    No primary-source URL is recorded for this entry in our reviewed data. Rather than manufacture a citation, we link the Implement Agentic research page that carries the record.

    Reproduce

    A reproduction tutorial is linked only when one exists for this exact record.

    A reproduction tutorial is not yet available for this entry. The closest reviewed material is .

    Cite or share

    APA-like: This historical record carries a year only, and no author or publisher of record in the checked-in data. An APA reference would have to invent that metadata.

    BibTeX: BibTeX requires an author and publication venue. Historical lineage entries store a narrative record and its source link, not structured authorship, so the field would be fabricated.

    Related