đź§  Google Reasoning Model

Gemini 2.0 Flash Thinking Pricing & Economics

Google's experimental reasoning architecture with visible chain-of-thought, combined with an enormous 1,048,576-token context window at high speed.

Input Rate / 1M Tokens
$0.15
Prompts under 128k
Output & Thinking Rate / 1M
$0.60
Includes chain-of-thought
Context Window
1,048,576
1M tokens supported
Free Tier in AI Studio
10 RPM
Free prototyping quota

Architecture & Reasoning Economics

Gemini 2.0 Flash Thinking bridges the gap between lightweight flash models and heavy reasoning engines (like OpenAI o1 or DeepSeek R1). Rather than requiring a dedicated heavy frontier model, Google trains a specialized reasoning head atop the ultra-fast Gemini 2.0 Flash foundation.

This allows the model to produce internal chain-of-thought tokens at approximately 120 tokens/sec—more than triple the speed of traditional reasoning models—while pricing thinking tokens at a micro-rate of $0.60/1M.

Frequently Asked Questions

Can I hide thinking tokens in the API response?
Yes. In the Google Generative AI SDK, you can set thinking_config: {"include_thoughts": false} to receive only the final synthesized response, though thinking tokens are still billed as part of output compute.
Does Gemini 2.0 Flash Thinking support image and multimodal reasoning?
Yes! Unlike OpenAI o1 (text-only), Gemini 2.0 Flash Thinking can reason across visual diagrams, complex charts, screenshots, and video streams.