YouTube15 Sept 2026
13m

Realtime Voice Agents with Frontier Intelligence — Bohan Li, EliseAI

Podcast cover

AI Engineer

Real-time voice agents require a cascaded architecture that mirrors self-driving car systems, specifically separating perception, planning, and control layers to balance intelligence with low latency. Perception relies on a streaming speculative transcriber that layers fast, real-time detection with slower, context-aware batch processing to ensure accuracy. Planning efficiency improves by offloading tool calling to background agents, which prevents unnecessary inference round trips and allows the main agent to maintain context. Finally, text-to-speech latency is minimized through a prefix cache that streams pre-generated audio while the model completes its response, creating a seamless user experience. By integrating these components, voice agents can handle complex tasks like scheduling appointments in healthcare and housing with natural, human-like responsiveness. EliseAI leverages this framework to build reliable, intelligent systems for critical service sectors.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise