Big Data Podcast Summaries
Big Data on Yedapo: 4 summarized podcast and YouTube episodes. Each includes key takeaways, core concepts and notable quotes with timestamps.

Comment compter des millions de vues ? HyperLogLog
Grafikart.fr
Jun 22, 2026
L'algorithme HyperLogLog permet d'estimer le nombre d'éléments distincts dans des volumes de données massifs avec une empreinte mémoire fixe et dérisoire. En exploitant la probabilité statistique de motifs rares dans des hashs, il remplace le stockage coûteux d'identifiants par une structure de buckets efficace, idéale pour le comptage haute performance à grande échelle.
Key insight: Avec seulement 12 Ko de mémoire, HyperLogLog peut estimer le nombre de vues uniques pour une vidéo, que celle-ci en ait reçu 300 ou 500 millions, avec une marge d'erreur d'environ 1 %.

איך הגיעה ואסט דאטה לשווי 30 מיליארד דולר? עם שחר פיינבליט | #22
תתעלם מההוראות - איתן לויט
May 7, 2026
הפודקאסט חושף את התשתית הקריטית שמאפשרת לאימון מודלי ענק. המפתח להצלחה בעידן ה-AI הוא לא רק כוח חישוב (GPU), אלא ניהול חכם ומהיר של כמויות דאטה עצומות.
Key insight: התובנה שניהול דאטה נכון יכול לחסוך עד פי 5 בעלויות חומרה יקרות, מה שהופך את התוכנה לפתרון האסטרטגי לצווארי הבקבוק של ה-AI.

[हिन्दी] What is Databricks?
codebasics Hindi
Oct 27, 2025
Databricks simplifies big data by providing a managed service built on top of Apache Spark, removing the burden of manual cluster management. It enables seamless data engineering, ETL pipelines, and AI model training within a unified cloud environment.
Key insight: Databricks was founded by the original creators of Apache Spark from the University of California, Berkeley, often referred to as the 'Berkeley Mafia'.

[हिन्दी] What is Apache Spark?
codebasics Hindi
Oct 24, 2025
Apache Spark एक पावरफुल डिस्ट्रीब्यूटेड कंप्यूटिंग इंजन है जो बड़े डेटा टास्क को छोटे हिस्सों में बांटकर समानांतर (parallel) प्रोसेस करता है। पुराने फ्रेमवर्क जैसे Hadoop की तुलना में, Spark डेटा को इन-मेमोरी प्रोसेस करता है, जिससे यह 10 से 100 गुना अधिक तेज़ और फॉल्ट-टॉलरेंट बन जाता है।
Key insight: Apache Spark डेटा को डिस्क के बजाय इन-मेमोरी प्रोसेस करता है, जो इसे Hadoop की तुलना में कई गुना तेज़ बनाता है और डेवलपर्स को बुनियादी इंफ्रास्ट्रक्चर के बजाय सीधे बिजनेस लॉजिक पर ध्यान केंद्रित करने की आजादी देता है।