We consume more knowledge through video than ever, in lectures, podcasts, meetings and technical tutorials. Finding one two-minute explanation inside a two-hour recording still means scrubbing the timeline and hoping.
Standard RAG works well on PDFs and text, but video is unstructured and multimodal. Transcribing only the audio loses slides, diagrams, code on screen and speaker names. Analyzing only frames loses the spoken explanation.
This session walks through a Video RAG system that treats audio, on-screen text and visuals as three independent tracks and merges them into one searchable index:
- Audio: Whisper transcription with timestamps.
- On-screen text: dense OCR that captures slides, code, UI text and speaker names.
- Visuals: LLM analysis of smartly selected keyframes, with deduplication to keep costs under control on long videos.
Ask a question and you get a grounded answer, the exact timestamps, and a playable clip of the relevant segment. The system can also answer "which video talks about X?" across a whole library.
The session includes a live demo and the real problems I hit along the way: keyframe budgets, noisy OCR, memory limits and cost control.
Key takeaways
- How to design a multimodal ingestion pipeline instead of transcript-only RAG
- How to chunk and index audio, OCR and visual context together
- Practical ways to control vision-model cost on long videos
- Patterns you can apply to any media-heavy platform