Why does a 25-turn conversation cost 10x more than 25 individual queries? Model compounding message payloads, prompt caching savings, and context truncation ceilings.
| Turn Number | Input Tokens Sent | Cumulative Chat History | Naive Turn Cost | Prompt Cached Turn Cost | Cost Multiplier vs Turn 1 |
|---|
Every modern foundation model API (Anthropic Messages API, OpenAI Chat Completions, Google Gemini) is completely stateless. The model does not retain conversational memory across HTTP connections. To give the user the experience of a coherent conversation, client libraries must append each subsequent query and assistant response to an ever-growing messages: [] array.
If a user has sent $N$ messages, the total input tokens processed by your API key across the entire conversation is:
Because the summation term grows as $O(N^2)$, the marginal cost of late-stage conversation turns skyrockets. In an unmonitored customer support agent or AI coding assistant, a single user having an extended 40-turn troubleshooting session can consume over 300,000 tokens—surpassing the cost of 50 standard one-shot search queries.
cache_control: {"type": "ephemeral"} breakpoint on the message history or system prompt. On OpenAI and DeepSeek, caching is automatic if the prefix matches at least 1,024 tokens. DeepSeek V3 offers up to 90% cache discounts ($0.014/M input read).