Gemini 3.1 Pro and the Downfall of Benchmarks: Welcome to the Vibe Era of AI
AI Explained
Feb 20, 2026
The era of generalist AI models is shifting toward domain-specific optimization, rendering traditional benchmarks increasingly unreliable. Because labs now prioritize post-training on internal datasets to boost specific scores, performance in one area no longer predicts capability in others. We have reached a threshold where frontier models perform on par with the average human in text-based reasoning.
Key insight: Anthropic CEO Dario Amade suggests that true generalization might be achieved simply by specializing in enough individual domains, potentially eliminating the need for models to learn on the job via continual learning.