Data Curation and Dedup: Why Less Data Beats More of It
Data curation and deduplication — not architecture tweaks — are what separate a model that generalizes from one that memorizes noise. FineWeb and Llama 3 both converged on filtering, near-duplicate removal, and deliberate domain mixing as the real quality lever behind their training corpora. The signal: labs are spending their engineering hours on editorial pipelines, not just bigger scrapes. For PMs, the same three levers — quality scoring, dedup, and mixing ratio — apply directly to whatever fine-tuning dataset your team is building right now.