Historical milestone · 2004

    MapReduce

    Reviewed through September 18, 2026

    2004 · Historical milestone

    MapReduce

    Era
    2000s
    Theme
    Infrastructure, efficiency & open ecosystems
    Evidence form
    Programming model and runtime automatically partitioned, scheduled, retried, and combined large key-value computations across commodity clusters
    School / paradigm
    Distributed systems
    Institution / context
    Google
    Researchers
    Jeffrey Dean; Sanjay Ghemawat

    Understand

    Plain-language record, transferred from the reviewed source module.

    Theory or experimental setup. Made data-parallel processing accessible and demonstrated terabyte-scale jobs on roughly 1,800 machines.

    Result / historical claim. Batch abstraction is inefficient for iterative learning and hides costs such as shuffles, stragglers, and data locality.

    Apply

    Professional implication, only where the reviewed record states one.

    The checked-in record does not state a separate professional application for this entry. The topic page places it in the wider research lineage: Infrastructure, efficiency, and open ecosystems.

    Verify

    Evidence status, stated limitations, and the external sources this record actually carries.

    Evidence form. Programming model and runtime automatically partitioned, scheduled, retried, and combined large key-value computations across commodity clusters

    Limitation / debate. Distributed data preparation, fault-tolerant training infrastructure, and workflow orchestration.

    Source status. The source link below is the verified link our reviewed topic research already carries for this milestone.

    Reproduce

    A reproduction tutorial is linked only when one exists for this exact record.

    A reproduction tutorial is not yet available for this entry. The closest reviewed material is Infrastructure, efficiency, and open ecosystems.

    Cite or share

    APA-like: This historical record carries a year only, and no author or publisher of record in the checked-in data. An APA reference would have to invent that metadata.

    BibTeX: BibTeX requires an author and publication venue. Historical lineage entries store a narrative record and its source link, not structured authorship, so the field would be fabricated.

    Related