🎙️ Audio Transcription Benchmarks

Cheapest Speech-to-Text APIs in 2026: Cost Per Minute Ranked

The complete unit economics guide for AI voice bots, call center transcription, and meeting note agents. Compare latency, word error rates (WER), and volume discounts.

#1

Deepgram Nova-3 BEST FOR REAL-TIME VOICE BOTS

The industry standard for voice agents (Vapi, Retell, LiveKit). Delivers sub-250ms streaming latency over WebSockets, industry-lowest Word Error Rate on accented speech, and native speaker diarization.

Latency: ~220ms • WER: 6.84% • Streaming: Native bi-directional WebSockets
Cost / Audio Minute
$0.0043
$0.258 / Audio Hour
#2

Groq Whisper Large v3 CHEAPEST BATCH INFERENCE

Powered by Groq's custom LPU hardware. Transcribes audio at an astonishing 250x realtime speed. The absolute lowest-cost endpoint available for processing pre-recorded podcasts, voicemails, and audio files.

Speed: 250x realtime • Model: Whisper Large v3 • Provider: Groq API
Cost / Audio Minute
$0.0018
$0.111 / Audio Hour
#3

OpenAI Whisper Large v3 HIGHEST MULTILINGUAL QUALITY

OpenAI's official managed transcription endpoint. Excellent translation capabilities across 98 languages, reliable punctuation, and built-in timestamp generation. Batch processing only.

Max File: 25MB • Languages: 98 languages • Provider: OpenAI Direct
Cost / Audio Minute
$0.0060
$0.360 / Audio Hour
#4

AssemblyAI Nano / Best BEST AUDIO INTELLIGENCE

Specialized for telephony and conversation intelligence. Provides automatic sentiment analysis, PII redaction, topic detection, and chapter generation alongside accurate transcripts.

Features: PII Redaction, Diarization • Streaming: Supported • Provider: AssemblyAI
Cost / Audio Minute
$0.0065
$0.390 / Audio Hour

Top Speech-to-Text Endpoints Ranked (2026 Master Table)

Provider & Model Cost / Minute Cost / Hour 1,000 Hours Cost Streaming Latency Word Error Rate (WER) Best For
Groq Whisper Large v3 $0.0018 $0.111 $111 Batch only (250x RT) 7.2% Massive batch audio archives
Deepgram Nova-3 $0.0043 $0.258 $258 ~220ms 6.84% Real-time voice bots, LiveKit
Deepgram Nova-2 $0.0043 $0.258 $258 ~260ms 7.80% Legacy telephony pipelines
OpenAI Whisper Large v3 $0.0060 $0.360 $360 Batch only 7.4% Multilingual translation
AssemblyAI Best $0.0065 $0.390 $390 ~350ms 7.1% PII redaction, call analytics
Gladia Audio Intelligence $0.0090 $0.540 $540 ~300ms 7.5% Multi-speaker overlap meetings
Google Cloud Speech-to-Text $0.0160 $0.960 $960 ~400ms 8.4% Google Workspace integrations

How to Choose Between Streaming vs Batch STT

For conversational AI voice bots (e.g. AI receptionists or support agents), Time to First Transcript Token (TTFT) is everything. If your user says "hello" and transcription takes 800ms, your entire conversational turn will feel sluggish and unnatural. In this scenario, Deepgram Nova-3 via streaming WebSockets is the only viable choice under $0.005/min.

Conversely, if your workload is transcribing recorded customer service calls or user voicemails after the fact, streaming low-latency is completely irrelevant. Routing recorded files to Groq Whisper Large v3 at $0.0018/min reduces your transcription expenses by nearly 60% with zero quality degradation.

Frequently Asked Questions: Speech-to-Text Pricing

Does Deepgram charge extra for speaker diarization?
Deepgram includes speaker diarization in standard Nova-3 pricing at no extra per-minute surcharge. In contrast, older legacy cloud providers like Google Cloud and AWS charge an additional $0.006/min for speaker identification.
What audio format yields the lowest latency in streaming STT?
Raw Linear PCM 16-bit 16kHz mono audio yields the fastest transcription speeds because the server does not need to waste CPU cycles decoding compressed MP3 or AAC codecs before feeding audio to the acoustic model.
Can I self-host Whisper for cheaper than $0.0018/minute?
Unless you are saturating multiple dedicated A100 or L4 GPUs 24 hours a day with constant audio streams, self-hosting faster-whisper or vLLM audio will almost certainly cost more than Groq's $0.0018/min rate once idle server time and DevOps engineering hours are accounted for.