Google's experimental reasoning architecture with visible chain-of-thought, combined with an enormous 1,048,576-token context window at high speed.
Gemini 2.0 Flash Thinking bridges the gap between lightweight flash models and heavy reasoning engines (like OpenAI o1 or DeepSeek R1). Rather than requiring a dedicated heavy frontier model, Google trains a specialized reasoning head atop the ultra-fast Gemini 2.0 Flash foundation.
This allows the model to produce internal chain-of-thought tokens at approximately 120 tokens/sec—more than triple the speed of traditional reasoning models—while pricing thinking tokens at a micro-rate of $0.60/1M.
thinking_config: {"include_thoughts": false} to receive only the final synthesized response, though thinking tokens are still billed as part of output compute.