How AI Benchmarks Really Work, and Why to Doubt the Leaderboard
AI benchmarks like MMLU, GPQA, and HumanEval each test a narrow, mechanically-graded proxy task — multiple choice, exact match, or unit tests — not general capability. As of mid-2026, many older benchmarks have saturated (top models clustering 88-99%) while contamination and funding conflicts of interest have eroded trust in what leaderboard rank actually proves. The gap between benchmark performance and real-world task performance is now the central skepticism story in AI evaluation. For PMs, a vendor's leaderboard screenshot should shortlist candidates, never decide a contract.