Calculate exact bottom-line dollars saved with Anthropic, OpenAI, and DeepSeek prompt caching. Model cache hit rates, prefix sizes, and annual gross margin expansion.
| Provider & Model | Standard Input | Cache Write Rate | Cache Read Rate | Effective Discount | Minimum Cache Size | Cache TTL Window |
|---|---|---|---|---|---|---|
| Anthropic Claude 3.7 Sonnet | $3.00 / 1M | $3.75 / 1M | $0.30 / 1M | 90.0% Off | 1,024 tokens | 5 minutes (refreshed on hit) |
| OpenAI GPT-4o | $2.50 / 1M | $2.50 / 1M | $1.25 / 1M | 50.0% Off | 1,024 tokens | Automatic (typically 5-10 min) |
| Anthropic Claude 3.5 Haiku | $0.80 / 1M | $1.00 / 1M | $0.08 / 1M | 90.0% Off | 2,048 tokens | 5 minutes (refreshed on hit) |
| DeepSeek V3 | $0.14 / 1M | $0.14 / 1M | $0.014 / 1M | 90.0% Off | 1,024 tokens | Automatic persistent caching |
| DeepSeek R1 Reasoning | $0.55 / 1M | $0.55 / 1M | $0.14 / 1M | 74.5% Off | 1,024 tokens | Automatic persistent caching |
To maximize prompt caching ROI, you must structure your prompt payloads with strict determinism. AI providers match prefixes sequentially from the very first token. If a single dynamic character (such as an unpredictable timestamp or user ID) is placed near the top of the prompt, the entire remaining context cache is invalidated.
By placing all static tokens at the beginning of the message array and placing dynamic user inputs strictly at the end, your application guarantees that the first 5,000 to 50,000 tokens achieve an unbroken cache hit on every single invocation.