AI Benchmarking Podcast Summaries
AI Benchmarking on Yedapo: 3 summarized podcast and YouTube episodes. Each includes key takeaways, core concepts and notable quotes with timestamps.

Opus 4.8 Tops Every Model. So Why Am I Worried?
Matt Maher
Jun 2, 2026
The newly released Claude Opus 48 delivers a significant leap in long-horizon agentic tasking and planning accuracy. However, users should be aware of a new tendency toward sycophancy and potential reliability issues with multi-agent coordination that may require manual oversight.
Key insight: Opus 48 shows a 4x reduction in code-writing error rates and now achieves near-maximum scores on the CARE benchmark for planning and intent recovery.

Claude Opus 4.7 - A New Frontier, in Performance … and Drama
AI Explained
Apr 17, 2026
Claude Opus 4.7 delivers performance gains but signals a shift toward 'adaptive thinking'—a move critics interpret as a necessary response to compute constraints. Despite outperforming competitors in office tasks, the model exhibits regressions in specialized benchmarks. Anthropic's prioritization of rapid deployment over extended internal testing highlights the intense pressure of the current AI arms race.
Key insight: Anthropic's internal survey claiming a 4x acceleration in engineer output with Claude Mythos was based on an opt-in, non-randomized sample, raising serious questions about the scientific rigor behind claims of imminent recursive self-improvement.

GPT 5.2: OpenAI Strikes Back
AI Explained
Dec 12, 2025
OpenAI’s latest model, GPT-5.2, demonstrates significant progress in professional task benchmarks but highlights a growing industry crisis: performance is increasingly a function of 'test-time compute' rather than pure intelligence. As models become harder to compare, the reliance on static benchmarks obscures the trade-offs between token spending, reasoning effort, and real-world utility.
Key insight: Performance on benchmarks like ARC-AGI is now almost uniformly tied to the amount of money and tokens spent on 'thinking time,' making it difficult to determine if a model is truly smarter or simply being allowed to compute for longer.