⚡ Wafer-Scale AI Inference

Cerebras API Pricing & Rates

World record AI inference speeds powered by the CS-3 Wafer-Scale Engine. Up to 2,100 tokens/sec on Llama 3.1 8B at $0.10 per 1M tokens.

Llama 3.1 8B (In/Out)
$0.10
Per 1M tokens (2,100 tok/s)
Llama 3.3 70B (In/Out)
$0.60
Per 1M tokens (~450 tok/s)
Free Tier Allowance
1M / day
No credit card needed
Max Speed Record
2,140
Tokens per second

🚀 Live Speed & Generation Time Calculator

Compare the real-time generation latency of Cerebras Wafer-Scale Engine against standard GPU cloud instances:

Generation Length (Output Tokens) 1,000 tokens
Cerebras Time (2,100 tok/s)
0.48s
Groq Time (450 tok/s)
2.22s
Standard GPU (75 tok/s)
13.33s
Cerebras Cost (8B)
$0.00010

📋 Cerebras Official Rate Card & Model Specs

Model Input / 1M Output / 1M Context Measured Throughput Rate Limits
Llama 3.1 8B $0.10 $0.10 8k ~2,100 tok/s 30 RPM / 60k TPM (Free)
Llama 3.3 70B $0.60 $0.60 8k ~450 tok/s Pay-as-you-go available
Llama 3.1 70B $0.60 $0.60 8k ~450 tok/s Pay-as-you-go

💻 1-Minute OpenAI SDK Integration

Cerebras is 100% drop-in compatible with the official OpenAI Python library:

from openai import OpenAI client = OpenAI( base_url="https://api.cerebras.ai/v1", api_key="your_cerebras_api_key_here" ) response = client.chat.completions.create( model="llama3.1-8b", messages=[{"role": "user", "content": "Explain quantum computing in 3 sentences."}] ) print(response.choices[0].message.content)

Frequently Asked Questions

How much does the Cerebras inference API cost?
Cerebras charges $0.10 per 1M tokens (input and output) for Llama 3.1 8B, and $0.60 per 1M input / $0.60 per 1M output tokens for Llama 3.3 70B, accompanied by a generous free tier of up to 1M tokens per day.
Why is Cerebras capable of 2,100 tokens per second?
Cerebras runs on the CS-3 Wafer-Scale Engine (WSE-3), a single giant chip with 900,000 AI cores and 44GB of on-chip SRAM providing 21 petabytes per second of memory bandwidth, eliminating external memory bus bottlenecks entirely.
Is Cerebras API compatible with OpenAI client libraries?
Yes, Cerebras provides a 100% OpenAI-compatible endpoint. Simply set base_url='https://api.cerebras.ai/v1' and pass your Cerebras API key.