
Build vs Buy AI Voice Agent: A 2026 Decision Framework (With the Hidden 3rd Option)
Most build-vs-buy guides give you a binary choice. The honest 2026 answer is a three-way matrix, and the option nobody talks about (hiring an agency to build custom flows on top of a platform) wins for roughly a quarter of teams we audit.
Building an AI voice agent from scratch in 2026 costs $250K, $2M in Year 1 and takes 4–9 months. Buying SaaS (Vapi, Retell, Bland) costs $5K, $100K and goes live in 5–14 days. A third path, hiring an agency to build custom flows on top of a platform, costs $30K, $150K and ships in 4–10 weeks.
Quick answer (the 3-option honest verdict):
- Buy SaaS (Vapi/Retell/Bland) wins for ~65% of teams. Year 1: $5K, $100K, live in 5–14 days.
- Agency-built-on-platform wins for ~25%. Year 1: $30K, $150K, live in 4–10 weeks, fully custom flows.
- Pure custom build wins for ~5%. Year 1: $250K, $2M, live in 4–9 months, only sane above 500K min/mo.
- Hybrid (buy platform + in-house custom layer) covers the final ~5%, usually F500 plus regulated industries.
Build, Buy, or Blend? The Three Paths Side by Side
Every existing build vs buy AI voice agent guide presents two options. In real-world deployments, three paths dominate: buy a hosted platform, build from scratch with your own engineers, or blend the two by hiring an agency to ship custom flows on top of Vapi, Retell, or Bland. They have wildly different cost curves, timelines, and risk profiles, and the decision usually isn't close once you cost it honestly.
| Path | Year 1 Cost | Time to Production | Customization | Best For |
|---|---|---|---|---|
| Buy SaaS | $5K, $100K | 5–14 days | Template + prompt + tools | SMB to mid-market, standard flows |
| Agency-built-on-platform | $30K, $150K | 4–10 weeks | Full custom flows on hosted runtime | Mid-market with unique workflows |
| Pure custom build | $250K, $2M | 4–9 months | Total: every layer is yours | F500, >500K min/mo, voice = core product |

Two notes on the table. First, "Year 1 cost" for buy assumes 50K–200K minutes per month; below that you're at the low end, above it you're closing on the agency tier. Second, the agency tier shows the all-in spend including platform passthrough, not just the agency fee. You'll see the math in H2 #5.
The procurement-team framing of "make or buy AI" misses the third path entirely. McKinsey calls it "build, buy, or blend" for a reason. When we measured Retell vs Vapi vs Bland end-to-end on a real workload, every winning deployment we shipped in 2025 fell into the blend bucket, not the binary. (See the Retell vs Vapi vs Bland comparison for the platform-side numbers; the Master of Code "true cost of AI voice agents" teardown anchors the cost ranges in the table above.)
Most build-vs-buy guides give you a binary choice. The honest 2026 answer is a three-way matrix.
How Much Does It Really Cost to Buy an AI Voice Agent?
The true cost of buying SaaS voice AI is $0.15–$0.33 per minute, which is 2–4× the advertised rate once ASR (automatic speech recognition), TTS (text-to-speech), LLM, telephony, and platform fees stack up. Year-1 cost for 100K minutes/month runs roughly $20K, $50K depending on vendor and voice choice. The headline price is honest; it just isn't what you'll actually pay.
Here's the May 2026 verified math when you go to buy an AI voice agent off the shelf.
| Vendor | Advertised | True all-in /min | Source (retrieved 2026-05-17) |
|---|---|---|---|
| Vapi | $0.05 | $0.18–$0.33 | vapi.ai/pricing |
| Retell AI | $0.07 | $0.13–$0.31 | retellai.com/pricing |
| Bland | $0.09–$0.11 | $0.15–$0.30 | bland.ai/pricing |
| Synthflow | $0.08–$0.13 (bundled) | $0.10–$0.18 | Synthflow pricing page |
Why the gap exists: the advertised number is the platform's slice. On top of that you're paying for the STT provider (Deepgram is typical, ~$0.0043/min), the LLM (GPT-4o-mini at $0.15/$0.60 per 1M tokens, or Claude Haiku at similar), the TTS engine (ElevenLabs Turbo at $0.06/min for premium voices, or vendor-bundled at lower quality), the SIP trunk ($0.014/min via Twilio), and per-number rental ($1–$5/month each). Add post-call analytics, knowledge-base storage, and a dedicated-support tier and you land where the table says.
Worked Year-1 examples at the voice ai agent cost ladder most teams hit:
- 10K min/mo (early pilot): roughly $2,000–$4,000 total spend. SaaS wins by an absurd margin against any build.
- 100K min/mo (mid-market deployment): $20K, $50K all-in. Still SaaS territory unless you have a wildly unique use case.
- 500K min/mo (call-center scale): $90K, $180K. Now you're starting to see why teams pull out the build spreadsheet.
For the full itemized breakdown by vendor and feature, see our itemized voice-agent pricing breakdown.
The $0.05/min headline price is real. So is the $0.28/min bill at month-end.
What Does It Actually Cost to Build an AI Voice Agent From Scratch?
Building a production AI voice agent from scratch costs $250K, $2M in Year 1, takes 4–9 months, and needs a team of 3–5 senior engineers with ASR, LLM, and telephony experience. The itemized TCO (total cost of ownership) lands at roughly 60% engineering salaries, 15% infrastructure and APIs, 10% compliance, 15% opportunity cost, and almost every internal build doubles its original estimate.
Engineers love quoting "we can build this in 8 weeks for $80K." Here's the itemized TCO every vendor blog hides, line by line, with US-loaded salaries assumed (we'll show the India/EU adjustment after).
| Line item | Year 1 cost (US) | Notes |
|---|---|---|
| Engineering team (3 senior × $180K) | $540,000 | India team: ~$180K. EU mid: ~$360K. |
| Infrastructure (compute, storage, observability) | $24,000 | AWS/GCP, logs, traces, on-call alerting |
| ASR API passthrough (Deepgram or Whisper) | $15,000–$60,000 | Volume-dependent |
| TTS API passthrough (ElevenLabs) | $15,000–$60,000 | Premium voices double this |
| Telephony (Twilio Programmable Voice) | $0.014/min × volume | $17K at 100K min/mo |
| Compliance audit (HIPAA / SOC 2 / PCI scope) | $20,000–$80,000 | Plus annual renewal |
| Opportunity cost (6 mo forgone roadmap) | Variable, usually $200K+ | What else those engineers don't ship |
| Year 1 total (typical mid-case) | $650K, $850K | Often blows past $1M when delays hit |
The "2× rule" is real: every internal voice-AI build we've audited came in at roughly twice the original estimate, almost always because the team underbudgeted compliance plumbing and turn-taking edge cases. Master of Code documented the same multiplier in their teardown, and Caller.Digital's TCO model lands in the same band. Plan for it, or quit the build early.
Don't forget the LLM orchestration layer: the framework that ties STT, retrieval, tool calls, and TTS into a turn-taking loop. You'll evaluate LangGraph, the OpenAI Agents SDK, or a custom finite-state machine. (We benchmark the top picks in AI agent frameworks for builders, and the deployment side in agentic AI deployment platform.) None of this is rocket science, but every layer is another integration test, another flaky failure mode, another on-call page at 2 a.m.
State management gets its own line item. Agents that forget the caller's name two turns in feel broken to users. We covered the trade-offs in memory for AI agents; short version, you'll either lean on Redis with custom serialization or rent a managed memory layer for $200–$2,000/month.
Here's the cumulative cost curve over three years comparing all three paths:
"Cumulative Cost (Year 1–3): Build vs Buy vs Agency-Built"
Data table
| "Months from kickoff" | "Pure Build" | "Buy SaaS" | "Agency-Built on Platform" |
|---|---|---|---|
| "Month 6" | 350000 | 15000 | 45000 |
| "Month 12" | 650000 | 45000 | 90000 |
| "Month 18" | 800000 | 75000 | 135000 |
| "Month 24" | 950000 | 105000 | 180000 |
| "Month 30" | 1100000 | 135000 | 225000 |
| "Month 36" | 1250000 | 165000 | 270000 |
Assumptions: 100K minutes/month workload, US-loaded engineering salaries, vendor pricing per May 2026 retrieval. Build crossover with agency-built only occurs around Month 32 when call volume exceeds 500K min/mo (see H2 #6).
Every internal voice-AI build comes in at twice the original estimate. Plan for it or quit early.
The Hidden Third Path: Agency-Built on Platform
Here's the option every vendor blog skips: hire an agency to build your custom AI voice agent flows on top of an existing platform (Vapi, Retell, or Bland). You skip the engineering hiring trap, ship in 4–10 weeks, and own the same business logic you'd own in a custom build, without renting a five-person ASR-LLM-TTS bench to maintain it.
Definition: an outside engineering team writes your flows, prompts, integrations, evals, and CRM hooks on top of a hosted voice platform. You pay a project fee for the build, then per-minute platform passthrough for runtime. The agency owns the code; you own the agent.
The cost math:
- Project fee: $30K, $80K depending on flow complexity (number of integrations, regulatory scope, eval rigor)
- Platform passthrough: $0.15–$0.30/min at the all-in rates from H2 #3
- Year-1 outcome: $50K, $150K total, roughly 4–10 weeks to first production call
Who it fits (the ~25%):
- Mid-market with unique workflows that don't fit a Vapi template
- Regulated industries (healthcare, finance, legal) needing audit-ready evals + compliance documentation
- Teams with zero in-house voice-AI engineers and no appetite to hire a specialist bench for a single project
- Multi-vertical agencies and franchisors who need 5+ distinct flows ships in parallel
Who it does NOT fit (the honesty gate, read this if you're considering it):
- Lone founder building MVP: go DIY on Vapi or Retell directly. The agency fee won't earn itself back at MVP volume. (Pair it with n8n for low-code agent workflows if you want orchestration without code.)
- F500 with a 50-engineer voice-AI team: build it. You have the bench, the volume, and the IP rationale.
- Sub-300ms latency requirement: only a co-located custom build hits that. No platform does, no matter who writes the flows.
- Use cases requiring full model ownership (you're training a proprietary LLM on call transcripts): you need the build path for ML access. The platforms abstract that layer away.
If your team lands in the agency-built bucket, you can talk to a custom voice agent development partner about scoping. No urgency, no hard pitch. Most teams who reach out spend 2–4 weeks just mapping their call flows before any contract conversation, and that's the right pace.
Most mid-market teams want a custom voice agent. Few want to hire a five-person ASR-LLM-TTS team to build one.
When Does Building Actually Win? The Break-Even Math
Building a custom AI voice agent only pays back when monthly call volume exceeds roughly 500,000 minutes AND the use case requires fewer than 5 integrations AND voice AI is a core product differentiator, not a commodity feature. Below that threshold, buying SaaS or hiring an agency on a platform beats building by 2–4× on three-year TCO. The formula is sharper than most teams admit.
In plain English:
Build wins when monthly minutes > 500,000 AND use case requires <5 integrations AND voice AI is a core differentiator (not commodity).
Worked example #1, 50K min/mo SMB clinic chain. Buy wins by ~4× over three years. Buy SaaS: ~$15K Year 1, ~$45K cumulative Year 3. Pure build: ~$650K Year 1, ~$1.25M cumulative Year 3. The clinic chain doesn't have voice-AI engineers, doesn't have the call volume, doesn't have a regulatory edge case the platforms can't handle. SaaS isn't just cheaper. It's the only sane answer. (See voice agents for restaurants for the equivalent math on the food vertical, which lands in the same place.)
Worked example #2, 2M min/mo enterprise contact center. Build breaks even with agency-built at roughly Month 28 and wins by ~1.6× by Year 3, only if the use case is repeatable across the contact center's verticals. Pure build: ~$650K Year 1, ~$3.5M Year 3 cumulative (mostly engineering ongoing). Buy SaaS at $0.20/min average: ~$4.8M Year 3. The build wins on math here, but only because the volume amortizes the engineering team.
The hybrid edge case covers the final ~5%: F500 firms running >5M min/mo, regulated industries with proprietary ML layers, or teams whose differentiator is genuinely the model itself. They buy the platform for telephony + STT + TTS commodity layers and build the LLM and post-call ML in-house. This is what Caller.Digital identifies as the "above 500K min/mo and below 5 integrations" threshold, where repeatability is everything.
Below 500K minutes, you're paying engineers to rebuild what Vapi already shipped.
Build wins on math at 500,000 minutes a month. Below that, you're paying engineers to rebuild what Vapi already shipped.
What Hidden Integration Work Will Blow Up Your Build Estimate?
Hidden integration work in a custom voice agent build includes A2P 10DLC (application-to-person 10-digit long code) registration, STIR/SHAKEN call attestation, TCPA (Telephone Consumer Protection Act) consent capture, jurisdictional recording-disclosure logic, CRM webhook authentication, and PII redaction. Platforms like Vapi, Retell, and Bland solve every line of this on day one. Your 8-week estimate forgot 6 weeks of telephony compliance plumbing.
Here's the AI phone agent scaffolding nobody puts on the original spec:
- A2P 10DLC registration: 2–6 weeks, $50–$1,000 in fees, required for any US SMS-adjacent flow (appointment reminders, callback confirmations). See Twilio's A2P 10DLC docs for the registration matrix.
- STIR/SHAKEN attestation: carrier-dependent caller-ID signing. Without it, your outbound calls are flagged "Spam Likely" in 30–60% of mobile cases. Ongoing maintenance, not one-time.
- TCPA consent capture: recording disclosure, do-not-call list integration, written opt-in storage. The FCC's 2024 AI-robocall ruling extended TCPA to AI voice; non-compliance is $500–$1,500 per call.
- State-by-state recording disclosure: 12 US states require two-party consent (California, Florida, Illinois, etc.), plus GDPR for any EU caller. Your agent needs to know where the caller is before the second sentence.
- CRM webhook auth + retry logic: OAuth refresh, exponential backoff, dead-letter queues. Sounds trivial; isn't.
- PII redaction + post-call summary storage: credit-card numbers, SSNs, medical details scrubbed before storage. HIPAA wants this; SOC 2 auditors will too.
Add these together and you've added 4–8 weeks of work that doesn't show up in the original "we can build this in 8 weeks" estimate. Platforms ship every line of it pre-wired.
Your 8-week build estimate forgot 6 weeks of telephony compliance plumbing.
Why Do Custom Voice Builds Sound Slower Than Platforms?
The total latency budget for natural-feeling voice AI is under 700ms end-to-end: STT (150–250ms) + LLM first-token (200–400ms) + TTS first-byte (100–200ms) + network/jitter (100–250ms). Platforms hit this median; custom builds frequently land at 1.2–2.0s because teams underestimate network and cold-start latency. Callers start talking over the agent at 400ms of silence, so the difference is audible.
The arithmetic most build estimates ignore:
| Stage | Latency (typical) | What slips |
|---|---|---|
| STT first chunk | 150–250ms | Streaming vs batch; cold model = +200ms |
| LLM first-token | 200–400ms | Cold start adds 400ms+; long system prompts add 100ms |
| TTS first-byte | 100–200ms | Premium voices slower; non-streaming = +500ms |
| Network RTT | 50–150ms | Worse if your AWS region isn't adjacent to caller's carrier |
| Jitter buffer | 50–100ms | The slice teams forget |
| Total budget | 550–1,100ms | Above 1.2s and the call feels off |

Why platforms hit 700–900ms median: pipeline optimization (interleaved STT/LLM streaming), network adjacency to telephony carriers, and warm model pools that never cold-start. Why in-house builds routinely land at 1.2–2.0s: teams don't budget the network/jitter slice, cold-start LLM calls add 400ms on the first turn of every call, and most home-grown TTS implementations don't stream first-byte audio before the full response is ready. Close's data on the 0.4s caller-talk-over threshold is the most-cited number in voice AI for a reason. It's the line where natural turn-taking breaks. Deepgram's latency benchmarks confirm STT first-chunk floors near 150ms.
Vapi hits 700ms because they own the network adjacency. Your AWS-region build won't.
Can You Really Build a Voice Agent in 15 Lines of Code?
Modern voice AI platforms like Vapi and Retell expose 10–20 line SDK calls that create a production-ready voice agent. The same functionality from scratch (STT, LLM orchestration, TTS, telephony, turn-taking) typically takes 4–9 months of engineering work. Counterintuitively, showing the code makes the buy case stronger, not weaker.
Here's a working Vapi Web SDK voice agent in 15 lines (verified against docs.vapi.ai/quickstart/web, retrieved 2026-05-17):
import Vapi from "@vapi-ai/web";
const vapi = new Vapi(process.env.VAPI_PUBLIC_KEY);
await vapi.start({
model: {
provider: "openai",
model: "gpt-4o-mini",
systemPrompt: "You are a friendly clinic receptionist. Book appointments, answer FAQs, escalate emergencies.",
},
voice: { provider: "11labs", voiceId: "rachel" },
transcriber: { provider: "deepgram", model: "nova-2" },
firstMessage: "Hi, this is the clinic. How can I help today?",
});
vapi.on("call-end", (data) => console.log("Call ended:", data.summary));Every line in that snippet replaces months of in-house work. The transcriber field is your STT integration. The voice field is your entire TTS pipeline including streaming first-byte handling. systemPrompt is the entire turn-taking and tool-calling layer (Vapi handles barge-in, endpointing, and tool routing under the hood). call-end gives you the post-call summary you'd otherwise build a Whisper + GPT-4 pipeline for.
This is what your custom build is reproducing for 4–9 months. The closer you read the SDK, the harder it gets to justify the build.
Counterintuitively, the more you understand the code, the stronger the buy case looks.
When Is Building a Voice Agent the Wrong Answer?
Building a voice agent is the wrong answer when your team has no voice-AI shipping experience, your estimate is under $150K Year-1, your timeline is under 8 weeks, your use case maps cleanly to an existing platform template, or your call volume is under 500K minutes/month. Hit any two of these and the build path is malpractice, not engineering.
Your engineering team will pitch building. They're not lying. They can build it. The question is whether they should. Watch for these five anti-pattern signals:
- "We can just wire up Whisper and GPT-4." Nobody who's shipped voice in production says this. The hard parts are turn-taking, barge-in detection, and TTS streaming, none of which are in that sentence.
- Year-1 estimate under $150K. Either the team doesn't understand the compliance scope, hasn't priced the telephony layer, or assumed they wouldn't need observability. All three are red flags.
- No engineer on the team has shipped voice before. Voice AI failure modes are different from chat. Echo cancellation, jitter buffering, barge-in: these aren't web-dev problems.
- Timeline under 8 weeks. Even the platforms that ship pre-built primitives need 2–4 weeks to wire integrations + CRM + evals. From scratch in 8 weeks means cutting compliance, cutting evals, or both.
- "We'll build the TTS in-house." No, you won't. ElevenLabs and Cartesia each have hundreds of engineers and dedicated audio model researchers. Buy what commodified.
Why engineers pitch building anyway: career capital, NIH ("not invented here") bias, headcount justification, the genuine joy of greenfield engineering. None of those are bad reasons to want to build. They're just bad reasons to actually build.
Reframe the decision: build what differentiates you, buy what commodified. The LLM, ASR, and TTS layers have commodified in 2025–2026. Your differentiation lives in flows, prompts, evals, and CRM integration. Don't rebuild the layers Vapi, Retell, OpenAI, and Deepgram already shipped.
Build what differentiates. Buy what commodified in 2025.
The Honest Decision Matrix
After building voice AI for contact center clients both ways, our reasoned split is: ~65% of teams should buy SaaS (Vapi, Retell, Bland), ~25% should hire an agency to build on a platform, ~5% need a hybrid (platform + custom in-house layer), and ~5% should build from scratch. Pure custom build only pays back above 500K minutes/month with a dedicated voice-AI team. These are our recommendations after auditing the math on 4 client decisions in 2025–2026, not industry statistics.
| Split | Path | Best for | Year-1 cost |
|---|---|---|---|
| ~65% | Buy SaaS | SMB to mid-market, standard flows, <500K min/mo | $5K, $100K |
| ~25% | Agency-built-on-platform | Unique flows, regulated industries, no voice-AI bench | $30K, $150K |
| ~5% | Hybrid (platform + in-house ML) | F500, regulated, proprietary model layer | $200K, $600K |
| ~5% | Pure custom build | Voice AI is core product, >500K min/mo, has the bench | $250K, $2M |

The four-node decision tree we walk clients through:
- Do you process >500K minutes/month? No → buy or agency-built. Yes → continue.
- Is voice AI a core product differentiator (not a feature)? No → buy or agency-built. Yes → continue.
- Do you have an in-house voice-AI engineering team (3+ shipped voice products)? No → agency-built or hybrid. Yes → continue.
- Do you need <300ms total latency or proprietary model training? No → hybrid. Yes → pure build.
Most teams stop at node 1 or node 2, which is why 90% of the matrix lands in buy or agency-built.
If you fit the 25% bucket, here's how we build on Retell or Vapi for clients. Start with your call flows, not with a contract.
Build what differentiates you. Buy what commodified. For most teams in 2026, that means buying the platform and customizing the flows.
The Verdict
- Three paths exist, not two. Buy SaaS for ~65% of teams. Agency-built-on-platform for ~25%. Pure build only above 500K min/mo with a real bench.
- The break-even formula is simple: monthly minutes > 500K AND <5 integrations AND voice AI is core. Miss any of the three and the build math doesn't work.
- Decision rule: build what differentiates you, buy what commodified. In 2026, that means buying the platform layer almost every time.
If you're in the 25% bucket where agency-built makes sense, start by mapping your call flows on a whiteboard, then talk to a partner who has shipped on Retell or Vapi. No urgency. The right move is to scope before you sign.
FAQ
Is it better to build or buy an AI voice agent?
For roughly 65% of teams, buying a SaaS voice platform (Vapi, Retell, Bland) is better. It's 5–14 days to production at $5K, $100K Year 1, versus 4–9 months and $250K, $2M for a custom build. Building only wins above 500K minutes/month when voice AI is a core product differentiator and you have a 3+ engineer voice-AI bench in-house.
How much does it cost to build an AI voice agent from scratch?
Building from scratch costs $250K, $2M in Year 1 for US-loaded teams. The breakdown is roughly $540K in engineering (3 senior engineers), $24K infrastructure, $30K, $120K in ASR and TTS API passthrough, $0.014/min telephony, and $20K, $80K for HIPAA/SOC 2 compliance audits. Most builds come in at 2× the original estimate due to compliance and edge-case work that's underbudgeted.
How long does it take to build a production voice AI agent?
Building from scratch takes 4–9 months for a production-ready agent with proper compliance, evals, and observability. Buying a SaaS platform like Vapi or Retell takes 5–14 days. The hidden third path (hiring an agency to build custom flows on an existing platform) takes 4–10 weeks. The gap is mostly the telephony and compliance scaffolding (A2P 10DLC, STIR/SHAKEN, TCPA) that platforms solve on day one.
Can I build a voice agent on Vapi or Retell instead of from scratch?
Yes, this is the agency-built-on-platform path. You (or an outside agency) build custom flows, prompts, integrations, and evals on top of Vapi, Retell, or Bland's hosted runtime. Year-1 cost is $30K, $150K total ($30K, $80K project fee plus $0.15–$0.30/minute passthrough), and you ship in 4–10 weeks instead of 4–9 months. You own the business logic; the platform handles STT, TTS, telephony, and turn-taking.
What's the break-even call volume for building voice AI?
The break-even threshold is around 500,000 minutes per month, and only if the use case requires fewer than 5 integrations and voice AI is a core product differentiator. Below 500K min/mo, SaaS or agency-built-on-platform beats building by 2–4× on three-year TCO. Above 500K min/mo with the right team and use case, a pure build breaks even with agency-built around Month 28 and wins by Year 3.
Is it cheaper to use a SaaS voice platform or build my own?
For nearly every team under 500K minutes/month, SaaS is cheaper by a wide margin, usually 4–10× cheaper on three-year TCO. The true SaaS rate is $0.15–$0.33 per minute (2–4× the advertised number once ASR, TTS, LLM, and telephony stack up), but that still beats $650K, $1.25M of engineering, infrastructure, and compliance costs you'd carry building it yourself.
What engineering skills do I need to build a voice agent?
You need at least 3 senior engineers with combined experience in real-time audio streaming, ASR integration (Deepgram or Whisper), LLM orchestration (LangGraph, OpenAI Agents SDK, or custom state machines), TTS streaming (ElevenLabs or Cartesia), SIP telephony (Twilio Programmable Voice or Telnyx), turn-taking and barge-in detection, and compliance work (TCPA, HIPAA, SOC 2). If your team hasn't shipped voice in production, the learning curve adds 2–4 months to the timeline.
How does latency differ between custom builds and platforms?
Platforms hit 700–900ms end-to-end median (Vapi, Retell, Bland). Custom in-house builds routinely land at 1.2–2.0 seconds because teams don't budget the network and jitter slices and cold-start LLM calls add 400ms on every first turn. Above 1.2 seconds, callers start talking over the agent and turn-taking breaks. The 700ms ceiling matters because platforms own network adjacency to telephony carriers. Your AWS-region build won't.
When does building voice AI make financial sense?
Building makes financial sense when monthly call volume exceeds 500,000 minutes, the use case requires fewer than 5 integrations, voice AI is a core product differentiator (not a feature), and you have a 3+ engineer in-house voice-AI team. Miss any one of these and the build math doesn't work. This is roughly 5% of teams we audit, usually F500 contact centers or AI-native companies where voice is the product.
What is the hidden cost of building voice AI in-house?
The hidden costs are A2P 10DLC registration (2–6 weeks, $50–$1,000), STIR/SHAKEN attestation, TCPA consent capture, state-by-state recording disclosure logic, CRM webhook authentication, PII redaction, post-call summary storage, observability and on-call infrastructure, and 6 months of opportunity cost on your forgone product roadmap. Platforms solve every line of this on day one. The 2× build-estimate rule exists because teams forget these line items.