Historical milestone · 2022
Holistic Evaluation of Language Models
Reviewed through September 18, 2026
2022 · Historical milestone
Holistic Evaluation of Language Models
- Era
- 2020s
- Theme
- Reliability, uncertainty & evaluation
- Evidence form
- Benchmark + taxonomy
- School / paradigm
- Multi-metric evaluation
- Institution / context
- Stanford CRFM
- Researchers
- Percy Liang; collaborators
Understand
Plain-language record, transferred from the reviewed source module.
Theory or experimental setup. Evaluated language models across many scenarios and metrics including accuracy, calibration, robustness, fairness, bias, toxicity, and efficiency.
Result / historical claim. Made model choice visibly multidimensional rather than reducible to one leaderboard number.
Apply
Professional implication, only where the reviewed record states one.
The checked-in record does not state a separate professional application for this entry. The topic page places it in the wider research lineage: .
Verify
Evidence status, stated limitations, and the external sources this record actually carries.
Evidence form. Benchmark + taxonomy
Limitation / debate. Coverage was necessarily incomplete and vulnerable to benchmark reuse, version drift, and disputed constructs.
Source status. This milestone row does not carry a primary-source URL in the approved export, and we do not have a verified link for it in our own research. We do not guess one.
No primary-source URL is recorded for this entry in our reviewed data. Rather than manufacture a citation, we link the Implement Agentic research page that carries the record.
Reproduce
A reproduction tutorial is linked only when one exists for this exact record.
A reproduction tutorial is not yet available for this entry. The closest reviewed material is .
Cite or share
APA-like: This historical record carries a year only, and no author or publisher of record in the checked-in data. An APA reference would have to invent that metadata.
BibTeX: BibTeX requires an author and publication venue. Historical lineage entries store a narrative record and its source link, not structured authorship, so the field would be fabricated.
