Average Surrounding Context (Tokens per FIM call)1,200 tokens
Monthly Input Tokens
18.0M
Codestral Monthly Bill
$4.05
GitHub Copilot Flat
$10.00
Cost vs Flat Copilot
59.5% Cheaper
âš¡ How Codestral Fill-in-the-Middle (FIM) Works
Traditional chat models can only append tokens to the end of a prompt. Codestral has native support for inserting code between a prefix and suffix, vital for cursor-based IDE code editing:
Codestral 22B is priced at $0.20 per 1 million input tokens and $0.60 per 1 million output tokens on Mistral's commercial La Plateforme API.
What is the Fill-in-the-Middle (FIM) endpoint?
Codestral supports native FIM via the /v1/fim/completions endpoint. Developers provide prompt (code before cursor) and suffix (code after cursor), allowing the model to generate insertions with sub-200ms latency without rewriting surrounding syntax.
How does Codestral compare with Qwen 2.5 Coder 32B?
Codestral offers an enormous 256k context window compared to Qwen 32B's 32k/128k native window and has a dedicated low-latency FIM endpoint, while Qwen 2.5 Coder 32B is slightly cheaper ($0.18/$0.35) and leads slightly on Python SWE-bench Verified benchmarks.