The AI Concepts Podcast
The AI Concepts Podcast is my attempt to turn the complex world of artificial intelligence into bite-sized, easy-to-digest episodes. Imagine a space where you can pick any AI topic and immediately grasp it, like flipping through an Audio Lexicon - but even better! Using vivid analogies and storytelling, I guide you through intricate ideas, helping you create mental images that stick. Whether you’re a tech enthusiast, business leader, technologist or just curious, my episodes bridge the gap between cutting-edge AI and everyday understanding. Dive in and let your imagination bring these concepts to life!
The AI Concepts Podcast is my attempt to turn the complex world of artificial intelligence into bite-sized, easy-to-digest episodes. Imagine a space where you can pick any AI topic and immediately grasp it, like flipping through an Audio Lexicon - but even better! Using vivid analogies and storytelling, I guide you through intricate ideas, helping you create mental images that stick. Whether you’re a tech enthusiast, business leader, technologist or just curious, my episodes bridge the gap between cutting-edge AI and everyday understanding. Dive in and let your imagination bring these concepts to life!
Episodes
2 hours ago
2 hours ago
13 min
This episode separates two concepts that are often confused: state, which keeps track of what is happening now, and memory, which allows information from the past to become useful later. We explore how applications create continuity around a model, and why good memory is not about remembering everything, but remembering the right things at the right time.
2 hours ago
13 min
2 hours ago
2 hours ago
12 min
This episode breaks down what orchestration actually means, from sequencing and routing to retries, parallel execution and human approvals, and explores the different ways those flows can be managed. Most importantly, we separate orchestration from the model itself and show why it is really about controlling how work moves through an application.
2 hours ago
12 min
Aug 21, 2026
Module 7: The LLM Application Loop
Aug 21, 2026
Aug 21, 2026
10 min
Who actually decides what happens next inside an LLM application? This episode explores the difference between decisions made by code and decisions made by the model, and why real applications often use both. We follow the loop that emerges when a model can request information or actions, receive the results and decide what to do next, revealing what developers mean when they talk about “owning the loop” and setting the stage for orchestration.
Aug 21, 2026
10 min
Aug 19, 2026
Module 7: Why Do We Need LLM Frameworks?
Aug 19, 2026
Aug 19, 2026
10 min
If LLM applications can be built with regular code and APIs, why do frameworks exist at all? This episode explores what happens as a simple application grows and starts needing retrieval, memory, routing, multiple models, retries and tracing. We look at what frameworks actually take off the developer’s plate, when those abstractions become useful, and why sometimes plain code is still the better choice.
Aug 19, 2026
10 min
Aug 18, 2026
Aug 18, 2026
10 min
What actually sits behind an LLM application? This episode takes one simple request and follows it beneath the surface, revealing how the model, application code, APIs, external data and context work together to produce something genuinely useful. As the request gets more complex, we begin to see why concepts like memory, tools and orchestration enter the picture. It is a practical look at what we are really building when we say we are building with LLMs.
Aug 18, 2026
10 min
Jun 11, 2026
Jun 11, 2026
8 min
This episode closes out Module 6 by tackling the question that has been getting louder since large context windows arrived. If a model can hold hundreds of thousands or even millions of tokens at once, do we still need all the architecture we just spent this module building? We explore why RAG was never just about fitting text into a small prompt, what retrieval is actually doing that a large context window cannot, and how the shift from compression to curation changes what good RAG looks like today. We cover when long context is genuinely the better tool, when retrieval still matters deeply, and why in most real enterprise systems the best answer is both working together. The episode closes with the argument that RAG is not disappearing. It is maturing. And everything we built in this module is part of that stronger foundation. By the end you will have a clear and honest picture of where these two approaches fit, and why understanding both puts you well ahead of most people working in this space.
Jun 11, 2026
8 min
Jun 9, 2026
Jun 9, 2026
8 min
This episode addresses the category of questions that vector search fundamentally cannot answer, questions about relationships between things. We explore what a knowledge graph is and why traversing connections between entities requires a completely different data structure than semantic similarity search. We break down Microsoft's GraphRAG approach, how it extracts entities and relationships from documents during indexing, uses community detection to identify clusters of related knowledge, and generates summaries that enable global queries across an entire corpus rather than just local document retrieval. We cover the cost improvements brought by LazyGraphRAG, the hybrid vector-plus-graph pattern most production teams are moving toward, Neo4j as the go-to graph database, and a lighter-weight entity extraction approach for teams not ready for a full knowledge graph. By the end you will understand when relationships matter more than text and how to build systems that can answer both kinds of questions.
Jun 9, 2026
8 min
Jun 9, 2026
Jun 9, 2026
7 min
This episode addresses a retrieval failure that has nothing to do with your index and everything to do with the query itself. We explore the vocabulary gap between how people ask questions and how documents are written, and why even strong embedding models cannot always bridge it. We break down three techniques that fix the query before the search runs: query rewriting to reformulate casual language into formal search terms, HyDE which generates a hypothetical answer and uses that as the search query instead of the question, and multi-query expansion which generates multiple phrasings to cast a wider retrieval net. We also cover step-back prompting for queries that need broader conceptual grounding before searching. By the end you will understand why the question itself is often the highest-leverage thing to improve in a retrieval pipeline.
Jun 9, 2026
7 min
Jun 9, 2026
Jun 9, 2026
7 min
This episode addresses the fundamental tension between retrieval precision and generation context. We explore why small chunks produce tight embeddings that retrieve well but leave the model without enough surrounding information, and why large chunks give the model context but dilute the embedding and hurt search quality. We break down parent-child indexing as the solution that decouples these two problems entirely, how child chunks handle the search and parent chunks handle the generation, and how to structure the hierarchy for documents of different complexity. We cover practical implementations in LlamaIndex and LangChain and close with guidance on when this pattern earns its place in a pipeline. By the end you will understand how to stop choosing between finding the right thing and giving the model enough to work with.
Jun 9, 2026
7 min
Jun 9, 2026
Jun 9, 2026
9 min
This episode addresses the gap between finding candidate chunks and finding the right ones. We explore the bi-encoder bottleneck, why compressing text into a single vector for comparison loses critical nuance, and how cross-encoders fix this by reading the query and document together in a single forward pass. We introduce ColBERT as a powerful middle ground between speed and accuracy through token-level late interaction, walk through the production tooling landscape including Cohere Rerank, BGE models, and RAGatouille, and close by stitching hybrid search and reranking into a complete three-stage retrieval funnel. By the end you will understand why two-stage retrieval is now the standard architecture for any serious RAG pipeline.
Jun 9, 2026
9 min




