YouTube22 Aug 2026
15m

Preferences Over Benchmarks: Model Routing — Archana Kamath & Tyler Gillam, DigitalOcean

Podcast cover

AI Engineer

Model routing offers a superior alternative to relying on a single frontier model, which often leads to excessive costs, poor task fit, and operational risks. Instead of chasing benchmark leaders, organizations should implement routing systems that select the optimal model based on specific request requirements like latency, cost, and task complexity. DigitalOcean’s inference router demonstrates this approach by utilizing an open-source, purpose-built model that routes requests in under 200 milliseconds. Live demonstrations show that this strategy can reduce session costs by up to 3x while maintaining performance parity with premium models. By integrating evaluation loops and custom rules, teams can continuously refine their routing logic, ensuring that each task—from simple classification to complex code generation—is handled by the most efficient and effective model available without vendor lock-in.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise