Synthetic Data Won't Collapse Your Model — Sloppy Data Will
Synthetic data — text and examples generated by one model to train another — has become a core input at frontier AI labs as the supply of fresh, high-quality human text runs thin. Researchers are split between two camps: synthetic data as the fix for the looming data wall, and model collapse as the hidden cost of models increasingly training on machine-generated output. The split matters now because the labs pulling ahead aren't avoiding synthetic data, they're building verification layers around it, while smaller players risk quietly degrading their models by skipping that step. For PMs, the real exposure isn't in the lab's training pipeline — it's in whether you're watching your own vendor's outputs for the early signs of homogenization.