← Back to Infographics

Capability Requires Energy

Within a given model, per-query energy scales roughly linearly with the number of tokens generated. Reasoning models emit a large hidden chain-of-thought before the visible answer, so a single query can consume many times the energy of a direct reply. Drag the slider to see the scaling.

400
Energy per query
0.6 Wh
Efficiency
1,667
queries / kWh
vs a direct answer
One query, split by what you see
visible
hidden reasoning
The visible answer stays small; the hidden chain-of-thought is what grows.
Where real models land (energy per query, log scale)
0.1 Wh 1 Wh 10 Wh 50 Wh

The slider illustrates within-model scaling only, calibrated at ~0.0015 Wh per generated token (so a ~200-token answer ≈ 0.3 Wh); actual consumption depends on model size, hardware, and serving stack. Reference points are independently sourced estimates for the specific models measured: median Gemini at 0.24 Wh (Google, 2025); GPT-4o ~0.3 Wh (Epoch AI, 2025); Claude 3.7 at 0.84–17 Wh and DeepSeek-R1 at 2.25–33.6 Wh (Jegham et al., 2025; IEA). Measurement boundaries differ, so values are indicative and not strictly comparable across models.