DeepSeek's Deleted Vision Paper Is Nuts...
bycloud
Jul 8, 2026
DeepSeek's new approach solves the 'reference gap' in multimodal models by allowing AI to 'point' at images using bounding boxes and coordinates. By interleaving visual primitives into its chain-of-thought, the model effectively grounds its reasoning in space rather than relying solely on ambiguous language descriptions, significantly outperforming frontier models in topological and counting tasks.
Key insight: For maze navigation, while most frontier models hover around 50% accuracy, this new architecture hits 66.9% by treating visual reasoning like a 'scratchpad' where the model draws points to track progress.