Gemma Podcast Summaries
Gemma on Yedapo: 2 summarized podcast and YouTube episodes. Each includes key takeaways, core concepts and notable quotes with timestamps.

Gemma-4 26B A4B + vLLM: Best MoE Model of 2026: Running Locally
Fahd Mirza
Apr 4, 2026
Google's Mixture of Experts architecture shatters the trade-off between model size and inference speed. Host Fahad Mirza demonstrates how activating just eight experts per token allows a massive 26B parameter model to deliver elite reasoning while maintaining the agility of a 4B model.
Key insight: The model utilizes 128 experts plus a shared expert across 30 layers, but only activates 4 billion parameters during any single inference pass for maximum efficiency.

Gemma 4 E2B + Hermes Agent + vLLM: Multimodal AI Stack Locally for Free
Fahd Mirza
Apr 3, 2026
Fahad Mirza demonstrates how Google's Gemma E2B integrates with Hermis agent to create a fully local, multimodal powerhouse. This 2-billion parameter model shatters the myth that high-tier vision and audio capabilities require massive server farms. By leveraging VLLM, Mirza proves that edge devices can now execute complex autonomous agency with minimal hardware.
Key insight: Despite its tiny footprint, Gemma E2B successfully handles audio transcription in over 100 languages and performs vision-based OCR while consuming less than 8GB of VRAM.