All talks
Talk

Beyond Text: A Multimodal RAG System Across Video, Audio, Images & Text

What’s the most fearless animal on Earth? That's the question I used to test my latest project and it led to something bigger: chatting with your data now actually means all your data. I've been experimenting with Multimodal (Gemini Embedding 2) in a RAG system that goes far beyond PDFs and text. It creates unified embeddings across: Video - find exact moments visually, Audio - search spoken content, Images - understand context in visuals, Text - fully integrated with everything. The real value isn't just the tech it's what you can build with it. Think video editors that work via prompts, or research tools that connect interviews, images, and documents instantly.