Decagon, Retell, Vapi, Smallest AI and Daily

Voice AI - The Next Frontier | Decagon, Retell, Vapi, Smallest AI, Daily

Voice agents are one of the hottest use cases in enterprise right now, but also one of the hardest to actually take live. Getting latency low enough to feel human without dumbing down the responses, making reliable tool calls to a CRM without dropping the customer mid-call, building fallback models for when Anthropic or OpenAI are running hot. None of it is as simple as the demos make it look.

Last week I hosted a fireside chat with five eng leaders who deal with this stuff every day: Basia Sudol (Head of Enterprise Solutions, Decagon), Varun Singh (CPTO, Daily), Steven Diaz (FDE Manager, Vapi), Tyler D'Silva (Founding FDE, Retell AI), and Sudarshan Kamath (Founder, Smallest AI).

We get into why nobody serious is shipping real-time voice-to-voice yet, why LLMs forget the middle of your prompt (and what that does to your architecture), why a giant prompt quietly destroys your unit economics, and why voice agent costs are now being compared directly against human labor.

Plus the stuff nobody warns you about: turn-taking, HIPAA constraints, why outbound is easier than inbound, why getting an exec to actually like the voice can be harder than any model problem, and more!

Voice AI - The Next Frontier | Decagon, Retell, Vapi, Smallest AI, Daily
Episode still: Decagon, Retell, Vapi, Smallest AI and Daily

Why this matters in practice

Production voice AI depends on latency, reliability, evaluation, and workflow integration—not conversational quality alone.

Chapters

  1. 00:00Intro
  2. 00:26Meet the panel
  3. 01:23Daily, WebRTC, and 20 years of building voice
  4. 03:49How Smallest AI made real-time TTS work
  5. 05:23Why Decagon moved into voice
  6. 08:16How Vapi and Retell think about the stack
  7. 11:14Forward deployed vs solutions engineering
  8. 16:48Voice agent architecture, explained simply
  9. 22:11Cascade vs speech-to-speech: the real tradeoff
  10. 28:11Hybrid pipelines and mixing models
  11. 32:16Accents, multilingual, and getting Singlish right
  12. 35:32Prompts vs workflows, and the latency fight
  13. 44:26How you actually evaluate a voice agent
  14. 49:04Simulation-based evals
  15. 49:49What production metrics really look like
  16. 51:53Building a QA framework that scales
  17. 54:53Evaluating speech-to-speech
  18. 58:10Open source benchmarks
  19. 59:57Why picking a voice is so subjective
  20. 01:02:03Personalization and custom voices
  21. 01:03:20Voice quality is solved, GPU efficiency is the new war
  22. 01:06:23Why outbound calls work better than you'd think
  23. 01:10:38Deploying in regulated industries (HIPAA, retention, audits)
  24. 01:12:43Turn-taking, the hardest unsolved problem in voice
  25. 01:18:53Where voice agents go in the next year
  26. 01:27:55Audience Q&A: inside Smallest's Hydra model
  27. 01:32:15The deployment problems nobody has solved yet
  28. 01:37:46Closing thoughts and thanks

Related topics

More episodes