Constitutional AI and RLAIF Are Killing the Rater Army
Constitutional AI and RLAIF (reinforcement learning from AI feedback) let labs replace much of the human-rater pipeline with a model judging outputs against a written set of principles, rather than thousands of contractors ranking responses by hand. Anthropic pioneered the approach in 2022 and Google DeepMind's follow-up RLAIF research found AI-judged training could roughly match human-judged training on preference tasks. The bigger signal is a cost and speed curve: rater pipelines scale with headcount and queue time, AI-judged pipelines scale with compute, which is why labs can now run safety-tuning passes far more often. For PMs, this changes the math on whether your own fine-tuning or moderation layer still needs a human-labeling vendor for routine judgment calls, or just a well-written rulebook.