Groq API Pricing Sheet & LPU Inference Rates 2026
Groq LPUs deliver 300-1,200 tokens/sec at ultra-competitive per-token economics. Verified daily against official cloud provider documentation.
Groq Cloud Production Pricing Matrix
Real-time unit costs, context sizes, operational throughput, and optimal use-case recommendations.
| Service / Model Tier | Primary Rate | Secondary / Output Rate | Context / Capacity | Throughput / Latency | Optimal Workload |
|---|---|---|---|---|---|
| Llama 3.3 70B Versatile | $0.59 | $0.79 | 128,000 | ~300 tok/s | Flagship open weights reasoning |
| Llama 3.1 8B Instant | $0.05 | $0.08 | 128,000 | ~1,250 tok/s | Ultra-fast classification and tool loops |
| DeepSeek R1 Distill Llama 70B | $0.75 | $0.99 | 128,000 | ~250 tok/s | Deep reasoning distilled with open weights |
| Whisper Large v3 (Audio) | $0.111 / hour | $0.111 / hour | N/A | 216x real-time | Fastest speech-to-text API in production |
Architectural & Financial Billing Nuances
Committed Use & Volume Discounts
Enterprise accounts spending over $2,000/mo typically qualify for 20% to 45% volume concessions or reserved throughput capacity agreements.
SLA & Latency Guarantees
Standard tiers offer 99.9% availability. Dedicated instances provide private VPC endpoints, zero noisy-neighbor degradation, and sub-100ms TTFT guarantees.
Frequently Asked Questions: Groq Cloud
Developer guidance on API compatibility, rate limit increases, and billing optimization.
Why is Groq so fast compared to GPU inference?
Groq uses Language Processing Units (LPUs) with massive on-chip SRAM bandwidth (up to 80TB/s), eliminating the memory-bus bottleneck of standard GPUs.
Does Groq have rate limits on free or pay-as-you-go tiers?
Yes, Groq offers a free tier with strict TPM/RPM caps, and an On-Demand Pay-As-You-Go tier with higher limits and enterprise SLAs.
Is Groq API compatible with the OpenAI SDK?
Yes, setting `base_url='https://api.groq.com/openai/v1'` makes it a 100% drop-in replacement.