GPT 5.2: OpenAI Strikes Back
AI Explained
Dec 12, 2025
OpenAI’s latest model, GPT-5.2, demonstrates significant progress in professional task benchmarks but highlights a growing industry crisis: performance is increasingly a function of 'test-time compute' rather than pure intelligence. As models become harder to compare, the reliance on static benchmarks obscures the trade-offs between token spending, reasoning effort, and real-world utility.
Key insight: Performance on benchmarks like ARC-AGI is now almost uniformly tied to the amount of money and tokens spent on 'thinking time,' making it difficult to determine if a model is truly smarter or simply being allowed to compute for longer.