✦ SUB-500MS END-TO-END TELEPHONY STACK 2026
Real-Time Conversational
Voice Agent Cost Calculator
Model per-minute telephony costs: Deepgram Nova-3 STT ($0.0043/min) + Gemini Flash LLM ($0.015/min) + Cartesia Sonic TTS ($0.012/min) + SIP trunking.
7 STT · 6 LLM · 6 TTS providers
Latency cascade with conversational threshold
Full vs. human call center ROI comparison
Voice Agent Stack Configuration
Per-Minute Token Estimate
Input tokens/min (user + system)
800 tok
Output tokens/min (agent response)
200 tok
TTS characters/min
750 chars
Composite Cost per Minute
$0.0099 / min
Monthly at 50,000 min: $495.00
STT cost$0.0043/min
LLM cost$0.0002/min
TTS cost$0.0053/min
Turnaround Latency Cascade
Total Human-Perceived Latency
525 ms ✓
Conversational-grade (<600ms) users won't notice lag
ROI vs Human Call Center
At $0.0099/min AI vs $2.00/min human agent: 99.5% cost reduction. Saves ~$99,505/month at 50,000 min volume.
STT Deep Dive
Speech-to-Text Provider Comparison 2026
| Provider / Model | Price per Minute | Streaming Latency | Word Error Rate | Free Tier | Key Strength |
|---|---|---|---|---|---|
| Deepgram Nova-3 | $0.0043 / min | ~220 ms | 5.5% WER | 12,000 min/mo | Best accuracy + lowest latency combo |
| Google STT v2 (Chirp) | $0.004 / min | ~300 ms | 6.1% WER | 60 min/mo | Multi-language (135+ langs), Google ecosystem |
| Azure Speech-to-Text | $0.0035 / min | ~280 ms | 6.8% WER | 5 hrs/mo | Cheapest, Azure native, enterprise compliance |
| AssemblyAI Universal-2 | $0.0065 / min | ~400 ms | 5.8% WER | None | LeMUR AI post-processing, best for transcripts |
| OpenAI Whisper Large-v3 | $0.006 / min | ~550 ms | 5.0% WER | None | Highest accuracy, best for non-English languages |
| Rev AI | $0.020 / min | ~600 ms | 4.2% WER | 300 min trial | Human fallback review, legal transcription grade |
TTS Deep Dive
Text-to-Speech Voice Engine Comparison 2026
| Provider / Model | Price per 1k Chars | TTFB (Latency) | Voice Quality | Custom Voices | Best Use Case |
|---|---|---|---|---|---|
| Cartesia Sonic | $0.007 | ~95 ms | 4.8/5.0 Natural | Yes (voice clone) | Low-latency real-time agents, IVR |
| Cartesia Sonic 2 Pro | $0.015 | ~85 ms | 5.0/5.0 Premium | Yes (instant clone) | Premium real-time voice with brand voice |
| Deepgram Aura-2 | $0.015 | ~130 ms | 4.8/5.0 Natural | Limited | Deepgram-native stack, single-vendor |
| PlayHT Turbo | $0.012 | ~180 ms | 4.8/5.0 Natural | Yes | Cost-balanced with good quality |
| ElevenLabs Turbo v2.5 | $0.030 | ~220 ms | 5.0/5.0 Human-like | Yes (voice clone) | Highest brand quality, marketing calls |
| OpenAI TTS HD | $0.030 | ~350 ms | 4.8/5.0 Clear | 6 preset voices | OpenAI-native stack integration |
Recommended Stacks
Pre-Built Voice Agent Stack Blueprints
| Stack Name | STT | LLM | TTS | Cost / Min | Latency | Best For |
|---|---|---|---|---|---|---|
| Ultra-Budget | Azure STT $0.0035/min |
Gemini Flash-Lite $0.075/1M |
Cartesia Sonic $0.007/1k |
~$0.0087/min | ~485 ms | High-volume inbound IVR, simple FAQs |
| Cost-Optimized | Deepgram Nova-3 $0.0043/min |
Gemini 2.0 Flash $0.10/1M |
Cartesia Sonic $0.007/1k |
~$0.0099/min | ~525 ms | Sales AI, outbound calls, scheduling bots |
| Balanced Quality | Deepgram Nova-3 $0.0043/min |
Claude 3.5 Haiku $0.80/1M |
PlayHT Turbo $0.012/1k |
~$0.016/min | ~555 ms | Customer support, complex reasoning tasks |
| Premium Brand | Deepgram Nova-3 $0.0043/min |
Claude 3.5 Haiku $0.80/1M |
ElevenLabs Turbo $0.030/1k |
~$0.033/min | ~690 ms | Luxury brand AI, high-touch sales |
| Max Capability | OpenAI Whisper $0.006/min |
GPT-4o $2.50/1M |
OpenAI TTS HD $0.030/1k |
~$0.058/min | ~1,280 ms | Research, edge cases, maximum accuracy |
Architecture Guide
Voice Agent Latency Optimization Techniques
Conversation Turn Architecture
# End-to-end voice agent turn timing:
1. VAD (Voice Activity Detection)
Silero VAD: ~20ms → detect user finished
Azure / Deepgram: built-in endpointing
2. STT Streaming
Partial transcripts arrive mid-sentence
Use Nova-3 with endpointing_ms=350
3. LLM Streaming Tokens
Begin TTS synthesis on first token!
Don't wait for full LLM completion
4. TTS Streaming Playback
Chunk audio: begin playback at ~2kb
Target: TTFB < 200ms
Cost Reduction Patterns
# 1. Cache common agent responses
# "What are your hours?" → pre-baked audio
# Saves 100% of LLM + TTS cost on hits
# 2. Reduce system prompt tokens
system_prompt = "concise instructions"
# 500-token prompts → save $0.00005/turn
# At 3M turns/month → saves $150/month
# 3. Turn-based LLM batching
# Only call LLM when VAD confirms end-of-turn
# Avoids speculative LLM calls on partial input
# 4. Deepgram Aura TTS for Deepgram STT users
# Single WebSocket = lower connection overhead
The Latency Cascade That Kills Conversational Quality
The biggest mistake in voice AI: waiting for the full LLM response before starting TTS. With GPT-4o generating 200 tokens at ~80 tok/sec, that's 2.5 seconds of silence before your user hears the first word. Stream LLM tokens directly into TTS send the first sentence fragment to TTS as soon as the first punctuation mark arrives. This alone cuts perceived latency by 50–60%.
FAQ
Voice Agent Pricing Common Questions
How much does Deepgram Nova-3 cost per minute?
Deepgram Nova-3 costs $0.0043 per minute of audio transcribed ($0.258/hr). At 50,000 minutes/month that's $215/month for STT alone. The free tier includes 12,000 minutes/month. Nova-3 achieves approximately 5.5% Word Error Rate on mixed domain audio with streaming latency of ~220ms. It supports custom vocabulary boosting, diarization (speaker separation), and 30+ languages.
What is the cheapest TTS engine for real-time voice agents?
Cartesia Sonic is the cheapest production-quality TTS for real-time agents at $0.007 per 1,000 characters approximately 4.3x cheaper than ElevenLabs Turbo v2.5 or OpenAI TTS HD ($0.030/1k). More importantly, Cartesia achieves ~95ms Time-to-First-Byte latency, the fastest in the industry. At 750 chars/min × 50,000 min/month, Cartesia costs $262/month vs $1,125/month for ElevenLabs at equivalent volume.
What is voice agent latency cascade?
Latency cascade is the cumulative response delay from user speech-end to agent audio-start: (1) VAD end-of-turn detection: ~200ms. (2) STT transcription latency: 220–550ms. (3) LLM time-to-first-token: 190–380ms. (4) TTS first-byte latency: 85–350ms. A best-in-class stack (Nova-3 + Gemini Flash + Cartesia) achieves ~525ms total. The human conversational threshold is ~600ms above this, pauses feel awkward. Above 800ms, users start speaking over the agent.
Which LLM is best for real-time voice agents?
For real-time voice, Gemini 2.0 Flash is the leading choice: 210ms TTFT, strong instruction-following, and just $0.00016/min in LLM fees at typical conversation token rates. For compliance-critical telephony (HIPAA, finance), Claude 3.5 Haiku is preferred. GPT-4o ($0.0024/min in LLM fees) should be avoided in high-volume real-time voice pipelines unless maximum reasoning capability is required it's 15x more expensive than Gemini Flash for equivalent voice tasks.
How much does AI voice replace human call center agents?
A human call center agent in the US fully-loaded (salary + benefits + overhead + management) costs approximately $1.50–$3.00 per minute. A cost-optimized AI voice stack (Deepgram + Gemini Flash + Cartesia) costs $0.0099 per minute a 99.3–99.7% cost reduction. At 50,000 minutes/month, AI costs ~$495 vs $75,000–$150,000 for human agents. The caveat: AI agents handle approximately 60–80% of inbound calls autonomously; complex escalations still require human handoff.