Speech Recognition Is Not a Solved Problem — Pavan Muddireddy
Machine Learning Street Talk (MLST)
Mistral AI research scientist Pavan Muddireddy explores the evolution of audio-native AI models and their integration into the broader AI stack. The discussion focuses on the shift toward end-to-end speech models that bypass traditional cascaded systems, effectively reducing error propagation and latency. Key innovations include the use of continuous latent embeddings and flow-matching techniques for generation, which offer superior control over latency-quality trade-offs compared to discrete token approaches. Pavan highlights the importance of Direct Preference Optimization (DPO) in mitigating hallucinations and degenerate loops in autoregressive models. While voice interfaces are becoming increasingly ubiquitous, they function best as auxiliary tools alongside visual media, particularly for complex tasks like coding or email triage, where visual feedback remains essential for maintaining cognitive clarity and operational precision.
Sign in to continue reading, translating and more.
Open full episode in Podwise
