YouTube15 Sept 2026
15m

Voice Agents Can Just Do Things — Charlie Guo, OpenAI

Podcast cover

AI Engineer

Voice agents are evolving beyond simple conversational interfaces, shifting toward three distinct interaction modes: speech-to-speech, speech-to-action, and event-to-speech. Rather than relying solely on verbal responses, developers can leverage existing software structures to enable voice-driven form filling, creative tool manipulation, and proactive system notifications. Modern real-time models, such as GPT Realtime 2, facilitate this by processing native audio as tokens, eliminating the latency and information loss associated with traditional transcription-based chains. This approach preserves critical emotional context and tone while allowing models to "think" before responding. By integrating these modes, developers can create more accessible, efficient, and intuitive experiences that move beyond the turn-based chatbot paradigm, ultimately positioning voice as a primary interface for future AGI interactions.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise