Why this matters in practice
Production voice AI depends on latency, reliability, evaluation, and workflow integration—not conversational quality alone.
Chapters
- 00:00Intro
- 00:26Meet the panel
- 01:23Daily, WebRTC, and 20 years of building voice
- 03:49How Smallest AI made real-time TTS work
- 05:23Why Decagon moved into voice
- 08:16How Vapi and Retell think about the stack
- 11:14Forward deployed vs solutions engineering
- 16:48Voice agent architecture, explained simply
- 22:11Cascade vs speech-to-speech: the real tradeoff
- 28:11Hybrid pipelines and mixing models
- 32:16Accents, multilingual, and getting Singlish right
- 35:32Prompts vs workflows, and the latency fight
- 44:26How you actually evaluate a voice agent
- 49:04Simulation-based evals
- 49:49What production metrics really look like
- 51:53Building a QA framework that scales
- 54:53Evaluating speech-to-speech
- 58:10Open source benchmarks
- 59:57Why picking a voice is so subjective
- 01:02:03Personalization and custom voices
- 01:03:20Voice quality is solved, GPU efficiency is the new war
- 01:06:23Why outbound calls work better than you'd think
- 01:10:38Deploying in regulated industries (HIPAA, retention, audits)
- 01:12:43Turn-taking, the hardest unsolved problem in voice
- 01:18:53Where voice agents go in the next year
- 01:27:55Audience Q&A: inside Smallest's Hydra model
- 01:32:15The deployment problems nobody has solved yet
- 01:37:46Closing thoughts and thanks
Related topics
More episodes
- Spencer Whitman - Gray Swan AI's $200M Plan to Secure AI SystemsSpencer Whitman — Gray Swan AI
- Tony Gentilcore - Glean, the $7.2B Startup Sam Altman Warned Investors AboutTony Gentilcore — Glean
- Russ Salakhutdinov - Kimi K3 CEO’s PhD Advisor Predicts the Future of AI AgentsRuss Salakhutdinov — Sooth Labs
