Forward Deployed

AI Agents in the Enterprise | Sierra, Mercor, Intercom, Turing | $2.8B+ Raised

(We know the audio quality isn't great on this one :( But the conversation is still well worth it!)

Last week I hosted a fireside chat on what it actually takes to build AI agents in the enterprise with Natalie Meurer (Head of Agent Eng, Sierra), Harsh Trivedi (founding engineer, Mercor), Juhi Parekh (GM, Turing), and Kevin Lynch (Senior FDE, Fin).

We get into why new models aren't always better (and why you can't just swap in the latest release and assume your agent improves), how the data-labeling/RL environment business might only have a couple years left, why real-time voice-to-voice models still aren't production-ready, how cheaper inference is still causing prices to go up, how baking in a constellation of models into enterprise agents is so important for reliability,

and much, much, more!

AI Agents in the Enterprise | Sierra, Mercor, Intercom, Turing | $2.8B+ Raised
Episode still

Why this matters in practice

Enterprise-agent adoption depends on choosing bounded workflows, connecting the right operating context, and giving teams a clear way to evaluate results.

Chapters

  1. 00:00Intro
  2. 00:21Meet the panel
  3. 01:57What everyone's actually using agents for day to day
  4. 06:10The reality of forward deployed work
  5. 09:28What agents couldn't do a year ago that they can now
  6. 12:29Why you have to tell agents what NOT to do
  7. 16:26What a harness actually is
  8. 22:03RL environments explained
  9. 28:46Does the data-labeling and RL environment business even last?
  10. 37:02Why benchmarks don't tell you what works in production
  11. 38:12Agent engineering vs forward deployed engineering
  12. 41:38Deploying into 100-year-old enterprise systems
  13. 44:25Why AI adoption is an org problem, not a tech problem
  14. 45:36Hiring for judgment when engineers aren't really coding anymore
  15. 48:16Why agents are a new kind of software
  16. 50:46The first 90 days of an enterprise deployment
  17. 53:20Why compliance environments break normal testing
  18. 56:58Layering AI on AI to get to 99% accuracy
  19. 01:00:56New models aren't always better — the swap problem
  20. 01:02:35Improving agents without waiting for a new model
  21. 01:06:36Does agent performance secretly degrade over time?
  22. 01:09:40Why one model is never enough: the constellation approach
  23. 01:11:14Building resilience when inference providers go down
  24. 01:13:48When fine-tuning actually makes sense
  25. 01:14:53Why voice-to-voice still isn't production-ready
  26. 01:16:25The cascaded pipeline that real voice agents use
  27. 01:21:45Audience Q&A: managing change inside the enterprise
  28. 01:23:24Why inference getting cheaper makes things more expensive
  29. 01:26:54Charging for outcomes instead of conversations
  30. 01:30:19What the real moat is when everyone uses the same models
  31. 01:37:04Synthetic data and where the data wall actually is
  32. 01:38:50Closing thoughts

Related topics

More episodes