Safety Tuning and Red-Teaming: What's Tested, What's Sold
Safety tuning is the two-layer process labs use to make models refuse harmful requests: refusal behavior trained directly into the weights via reinforcement learning, plus separate classifier systems that screen inputs and outputs after generation. Red-teaming — human specialists and, increasingly, automated adversarial models — is how labs find the gaps in both layers before and after release. The bigger signal as of mid-2026: independent grading from the Future of Life Institute's Summer 2026 Index puts the best-performing lab at a C+, while binding audit requirements under the EU AI Act apply to only a handful of the largest systemic-risk models. For PMs, every refusal your users hit is a tuning decision made by someone else's safety team, and the over-refusal rate is now your product's problem to manage, not just the vendor's.