59 Models →
Official 2026 Rate Card

Groq API Pricing Sheet & LPU Inference Rates 2026

Groq LPUs deliver 300-1,200 tokens/sec at ultra-competitive per-token economics. Verified daily against official cloud provider documentation.

Groq Cloud Production Pricing Matrix

Real-time unit costs, context sizes, operational throughput, and optimal use-case recommendations.

Service / Model Tier Primary Rate Secondary / Output Rate Context / Capacity Throughput / Latency Optimal Workload
Llama 3.3 70B Versatile$0.59$0.79128,000~300 tok/sFlagship open weights reasoning
Llama 3.1 8B Instant$0.05$0.08128,000~1,250 tok/sUltra-fast classification and tool loops
DeepSeek R1 Distill Llama 70B$0.75$0.99128,000~250 tok/sDeep reasoning distilled with open weights
Whisper Large v3 (Audio)$0.111 / hour$0.111 / hourN/A216x real-timeFastest speech-to-text API in production

Architectural & Financial Billing Nuances

Committed Use & Volume Discounts

Enterprise accounts spending over $2,000/mo typically qualify for 20% to 45% volume concessions or reserved throughput capacity agreements.

SLA & Latency Guarantees

Standard tiers offer 99.9% availability. Dedicated instances provide private VPC endpoints, zero noisy-neighbor degradation, and sub-100ms TTFT guarantees.

Frequently Asked Questions: Groq Cloud

Developer guidance on API compatibility, rate limit increases, and billing optimization.

Why is Groq so fast compared to GPU inference?

Groq uses Language Processing Units (LPUs) with massive on-chip SRAM bandwidth (up to 80TB/s), eliminating the memory-bus bottleneck of standard GPUs.

Does Groq have rate limits on free or pay-as-you-go tiers?

Yes, Groq offers a free tier with strict TPM/RPM caps, and an On-Demand Pay-As-You-Go tier with higher limits and enterprise SLAs.

Is Groq API compatible with the OpenAI SDK?

Yes, setting `base_url='https://api.groq.com/openai/v1'` makes it a 100% drop-in replacement.