Multimodal AI Podcast Summaries
Multimodal AI on Yedapo: 2 summarized podcast and YouTube episodes. Each includes key takeaways, core concepts and notable quotes with timestamps.

DeepSeek's Deleted Vision Paper Is Nuts...
bycloud
Jul 8, 2026
DeepSeek's new approach solves the 'reference gap' in multimodal models by allowing AI to 'point' at images using bounding boxes and coordinates. By interleaving visual primitives into its chain-of-thought, the model effectively grounds its reasoning in space rather than relying solely on ambiguous language descriptions, significantly outperforming frontier models in topological and counting tasks.
Key insight: For maze navigation, while most frontier models hover around 50% accuracy, this new architecture hits 66.9% by treating visual reasoning like a 'scratchpad' where the model draws points to track progress.

🔴 ¡Pruebo la VOZ AVANZADA de CHATGPT! (Espectacular) feat. @Lahiperactina
Dot CSV
Aug 1, 2024
The host explores OpenAI's new multimodal Advanced Voice Mode, demonstrating its ability to handle real-time interruptions, emotional modulation, and complex roleplay. While the technology represents a significant leap in conversational fluidity and latency, it remains an early alpha with strict safety guardrails that limit its ability to perform specific tasks like voice mimicry or complex reasoning.
Key insight: The model can process audio input and generate audio output directly without transcribing to text first, allowing for parallel processing of different data streams like code and speech simultaneously.