Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
No Priors: AI, Machine Learning, Tech, & Startups
Jun 26, 2026
Current AI evaluation frameworks fail because they ignore 'test-time compute,' treating model capability as a static number rather than a function of budget. Noam Brown argues that as models scale, performance on complex tasks doesn't plateau for weeks, making traditional benchmark grids misleading. To accurately measure progress, the industry must shift to plotting performance against compute cost.
Key insight: Modern frontier models can continue to improve on complex tasks for up to 100 million tokens of inference, meaning traditional static benchmarks are failing to capture the true ceiling of their capabilities.