Cerebras charges $0.10 per 1M tokens (input and output) for Llama 3.1 8B, and $0.60 per 1M input / $0.60 per 1M output tokens for Llama 3.3 70B, accompanied by a generous free tier of up to 1M tokens per day.
Why is Cerebras capable of 2,100 tokens per second?
Cerebras runs on the CS-3 Wafer-Scale Engine (WSE-3), a single giant chip with 900,000 AI cores and 44GB of on-chip SRAM providing 21 petabytes per second of memory bandwidth, eliminating external memory bus bottlenecks entirely.
Is Cerebras API compatible with OpenAI client libraries?
Yes, Cerebras provides a 100% OpenAI-compatible endpoint. Simply set base_url='https://api.cerebras.ai/v1' and pass your Cerebras API key.