A working voice AI demo takes a strong engineer about a weekend. The demo talks, books a test meeting, and everyone on the call gets excited. Six months later, a surprising number of those projects are quietly dead: the prompts have drifted, the CRM sync broke when someone renamed a field, and nobody owns the phone number that stopped connecting.
A preliminary MIT Media Lab report found that ninety-five percent of generative AI pilots fail to show measurable financial returns within six months. The failure is almost never the model. It is the stack around it, and more precisely, the absence of anyone running that stack.
We build and operate voice AI systems for a living, across multiple runtimes, so this post is the map we wish every buyer had: the six layers a production deployment runs on, what each one hides, and how to decide who should run them.
6
Layers in every production voice AI deployment, whoever runs them.
<1s
Voice to voice response time production stacks now target.
95%
Of generative AI pilots fail to show returns within six months, per a widely cited MIT study.
The six layers of the voice AI stack
The demo you watched is the top layer. Underneath it sit five more: the runtime that generates the voice, the telephony that carries it, the orchestration that decides what happens next, the integrations that write it all down, and the operations work that keeps yesterday’s agent good tomorrow.
The production voice AI stack
With Kaigen
We run all five layers for you
Built, monitored, and tuned weekly. Assess, Build, Deploy, Optimize.
None of these layers is optional. The only decision is who runs them. Here is what each one involves once real calls are flowing.
The conversation is the product
Everything below this layer exists so that a buyer can ask a layered question at 9pm and get a useful answer in under a second. The bar is unforgiving: sub second response, clean interruption handling, and language that matches the caller, including a mid sentence switch from English to Hindi. Our post on state of the art voice agent capabilities walks this layer in detail, and the fastest way to judge it is to hear a live agent rather than read about one.
The trap at this layer is that it is the only one a demo exercises. A buying process that stops here is how teams end up owning the five layers they never saw.
Voice runtime: rented, never outsourced
The runtime converts speech to text, reasons with a language model, and speaks the reply. Retell, ElevenLabs, Vapi, and Bland are runtimes, and the good ones are genuinely good: fast pipelines, clean SDKs, fair per minute pricing. We run several in production and compared the layer honestly in Kaigen Labs vs Retell AI and Kaigen Labs vs ElevenLabs Agents.
What renting a runtime does not remove: model versions change under you, latency varies by region, voices get deprecated, and every provider has outages. A production deployment needs someone watching those changes and, ideally, a second runtime to fail over to. That work does not appear in the demo or on the pricing page.
Telephony and numbers: the layer nobody budgets for
Calls ride on carriers, SIP trunks, and phone numbers with reputations. Pickup rates move with caller ID, local presence, and whether carriers have quietly flagged your number. India routes promotional and transactional traffic through separate number series with registration requirements; the United States prices consent mistakes per call. The regulatory half of this layer is mapped in our voice AI compliance guide.
Telephony fails quietly. Nothing errors; connect rates drift down until someone notices the pipeline thinned. Owning this layer means watching deliverability the way an email team watches sender reputation.
Orchestration: where deployments are won
One call is rarely the motion. The motion is a text before the call, a voicemail with an immediate follow up, an email with the document discussed, and a retry at a different hour, all sharing one memory of the buyer. Widely cited lead response research shows contact odds collapse within minutes of an inquiry, so orchestration is also the layer that decides whether speed exists at all. We wrote up the sequencing playbook in multi channel automation.
On a DIY build this layer is yours entirely: queues, state machines, retry logic, and the memory that keeps channel three from repeating what the buyer said on channel one. It is the largest single block of engineering in the stack, and the least visible in a demo.
Integrations: the connector is not the work
Every platform lists CRM connectors. The work is behind the connector: field mapping, trigger logic, idempotency, and the Tuesday your CRM admin renames a property and the sync silently stops. Calendars add timezone math and double booking edge cases. The difference between a demo integration and a production one is who notices when it breaks, and how fast.
Operations: the layer that decides month six
Prompts drift as your offer changes. New objections appear that the agent handles badly. A model upgrade changes tone. Regulations move. The teams whose agents are better in month six than at launch all do the same thing: they read transcripts, run evals on real conversations, and tune weekly. The ones whose agents quietly degrade skipped exactly that. This layer is not a task; it is a standing job, and it is the reason we treat managed operation as the product rather than an add on.
Build the stack, or buy the outcome
Both paths are legitimate. The honest fork is not about talent; it is about whether operating this stack is a job your company should own.
WHEN BUILDING FITS
Build it in house if…
- Voice is core product, not a sales channel
- You have engineers who can own it as a standing job
- You want control over every layer and will pay for it in headcount
- Your volume justifies a dedicated operating team
WHEN MANAGED FITS
Have us run it if…
- You want the outcome, not a second engineering roadmap
- Speed to lead and follow up decide your revenue
- You need multi channel, CRM write back, and failover on day one
- You want someone accountable for month six, not month one
The economics of that fork, including what a seat produces versus what software produces, are modeled in the economics of voice AI.
Pressure test any vendor, including us
Q1
Who tunes prompts in month six?
If the answer is a name on your payroll, budget for it. If it is nobody, that is the whole risk.
Q2
What happens when the runtime has an outage?
A single provider stack cannot fail over to itself. Ask what the fallback is and who flips it.
Q3
Who owns the CRM sync when a field changes?
Schema changes are routine. If they break the pipeline for a week, the integration was a demo.
Q4
Where do compliance updates live?
Disclosure rules and number regulations move. Someone has to track them per market you call.
KEY TAKEAWAYS
- Every production voice AI deployment runs six layers; the demo shows you one.
- The runtime is rentable. Telephony health, orchestration, integrations, and operations are jobs, not purchases.
- Orchestration is where deployments are won: speed and multi channel follow up live there.
- Month six is decided at the operations layer, by whoever reads transcripts and tunes weekly.
- Build if voice is your product. Buy the outcome if voice is your sales channel.
FAQ
What is a voice AI stack?
The full set of systems a production voice agent runs on: the conversation layer buyers experience, the voice runtime (speech recognition, language model, synthesis), telephony and phone numbers, orchestration across channels, integrations into CRM and calendars, and the ongoing operations work of monitoring and tuning.
Can I build a production voice agent myself?
Yes. Runtimes like Retell, Vapi, and ElevenLabs make the first working agent achievable in days for a strong team. The build is not the hard part: production means owning telephony health, multi channel orchestration, integration maintenance, and weekly tuning as a standing job.
What is the difference between a voice AI platform and a managed service?
A platform rents you the runtime layer and hands you the other five. A managed service like Kaigen Labs operates the whole stack: it builds the agent, runs the channels and integrations, monitors every call, and tunes the system weekly, with accountability for outcomes rather than uptime alone.
Why do so many voice AI pilots fail?
A preliminary MIT Media Lab report put the failure rate of generative AI pilots at ninety-five percent within six months, and the pattern is consistent: the model performs, but nobody owns the operational layer, so prompts drift, integrations break silently, and the pilot never compounds into results.
How fast can a production voice agent go live?
On the Kaigen Method (Assess, Build, Deploy, Optimize) most pilots are live in two to three weeks, connected to the CRM and calendar, starting on a focused slice of volume with a human in the loop and measured against your existing baseline.
MAP YOUR STACK
Want this map drawn for your own motion?
Twenty minute call with the Kaigen team. Bring how you sell today; we sketch the six layers for your case and tell you honestly which ones you could run yourself.
Book a 20 minute audit →Comparing runtimes for the build path? Start with our neutral operator guides, Retell AI vs Vapi and Retell AI vs Bland AI. Weighing a runtime against the managed path instead? See Kaigen Labs vs Vapi, Kaigen Labs vs Bland AI, or how a managed deployment works end to end.




