💻 Code Generation & FIM Specialist

Codestral 22B Pricing & FIM Guide

Mistral's 22B generative code specialist with a massive 256k context window and dedicated Fill-In-the-Middle endpoint for instant IDE autocompletions.

Input / 1M Tokens
$0.20
La Plateforme rate
Output / 1M Tokens
$0.60
Blended code generation
Context Window
256,000
Industry leading for 22B
FIM Latency
~180ms
Sub-second IDE completions

🧮 Developer IDE Autocomplete Cost Calculator

Estimate the monthly API expense of using Codestral for inline autocompletions (via Continue.dev, Cursor, or VS Code plugins):

Daily Inline Completions Requested 500 requests/day
Average Surrounding Context (Tokens per FIM call) 1,200 tokens
Monthly Input Tokens
18.0M
Codestral Monthly Bill
$4.05
GitHub Copilot Flat
$10.00
Cost vs Flat Copilot
59.5% Cheaper

âš¡ How Codestral Fill-in-the-Middle (FIM) Works

Traditional chat models can only append tokens to the end of a prompt. Codestral has native support for inserting code between a prefix and suffix, vital for cursor-based IDE code editing:

curl https://api.mistral.ai/v1/fim/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $MISTRAL_API_KEY" \ -d '{ "model": "codestral-2405", "prompt": "def calculate_discount(price, tier):\n if tier == \"VIP\":\n return price * 0.8\n", "suffix": " return price\n", "temperature": 0.1 }'

Frequently Asked Questions

What is the API pricing for Codestral 22B?
Codestral 22B is priced at $0.20 per 1 million input tokens and $0.60 per 1 million output tokens on Mistral's commercial La Plateforme API.
What is the Fill-in-the-Middle (FIM) endpoint?
Codestral supports native FIM via the /v1/fim/completions endpoint. Developers provide prompt (code before cursor) and suffix (code after cursor), allowing the model to generate insertions with sub-200ms latency without rewriting surrounding syntax.
How does Codestral compare with Qwen 2.5 Coder 32B?
Codestral offers an enormous 256k context window compared to Qwen 32B's 32k/128k native window and has a dedicated low-latency FIM endpoint, while Qwen 2.5 Coder 32B is slightly cheaper ($0.18/$0.35) and leads slightly on Python SWE-bench Verified benchmarks.