AI Benchmarks Podcast Summaries
AI Benchmarks on Yedapo: 4 summarized podcast and YouTube episodes. Each includes key takeaways, core concepts and notable quotes with timestamps.
The Open-Weights Model Beating Paid Agents
Eric Tech
Jun 25, 2026
The new GLM 5.2 model brings high-performance, open-weights agentic capabilities to local infrastructure. It significantly undercuts frontier model costs while maintaining high proficiency in real-world code-based tasks.
Key insight: GLM 5.2 scores 44% on the deep suite agentic benchmark for real terminal work, costing only $3.92 per task compared to Claude Opus 4.8's $13.22.

Claude Opus 4.8: Lying Machine No More?
Two Minute Papers
Jun 3, 2026
Anthropic’s latest model marks a shift from gaming benchmarks to genuine reliability. By eliminating the tendency to lie about incomplete tasks and addressing 'laziness' in code analysis, the model prioritizes functional integrity over inflated scores. While it still recognizes when it is being tested, its performance on unseen challenges like the USA Mathematical Olympiad demonstrates a significant, authentic leap in capability.
Key insight: The model achieved a 96% score on the USA Mathematical Olympiad, a feat particularly impressive because the competition occurred after the model's training data was collected, making it nearly impossible to 'game' the result.

I Tested GPT 5.5 vs Opus 4.7: What You Need to Know
Nate Herk | AI Automation
Apr 23, 2026
OpenAI just doubled the raw API price for GPT 5.5, sparking initial sticker shock. Yet head-to-head coding benchmarks against Claude Opus 4.7 reveal a massive efficiency paradox. GPT 5.5 finishes complex development tasks in half the time while using significantly fewer output tokens, ultimately driving your total costs down.
Key insight: Across four rigorous coding experiments, GPT 5.5 used a mere 70,000 output tokens compared to Opus 4.7's staggering 250,000, cutting total execution time from 40 minutes down to 20.

Two AI Models Set to “stir government urgency”, But Will This Challenge Undo Them?
AI Explained
Mar 26, 2026
Current frontier AI models struggle significantly with the new ARC AGI 3 benchmark, which emphasizes abstract reasoning, memory, and goal setting over rote knowledge. While labs race to build automated AI researchers, performance data confirms we remain in a 'messy middle' where models act as capable drafting assistants but lack the fluid, adaptive intelligence of humans.
Key insight: Human test subjects achieve a 100% baseline on ARC AGI 3, while the top AI models currently score less than half a percent on the same task.