Beyond attention · Part 1

AI's energy and water bill, and where hybrid models fit

September 25, 2026

This is the first of four posts on a quiet shift in how large language models are built. Most of today’s models are transformers, which rely on a mechanism called attention. A growing number are hybrids that mix a few attention layers with many cheaper “state” layers. Why care? One reason is money. Another is the physical footprint of AI in electricity, water, and carbon. This post covers both, separates what is measured from what is forecast, and ends with what hybrids can and cannot do about it.

Beyond attention · Part 1 · At a glance

Data center electricity could roughly double by 2030. Memory per conversation is one lever.

Global data center electricity, terawatt hours

2024, estimate415 TWh
2025, estimate485 TWh
2030, forecast950 TWh

2024 and 2025 are IEA estimates of past use. 2030 is the IEA Base Case forecast.

0.24 Wh

Google's measured median Gemini text prompt, May 2025

25×

Energy per response for reasoning versus plain chat (ML.ENERGY)

90%

Less memory per 128K conversation when 4 of 40 layers stay attention (our arithmetic)

Architecture is one lever among many. Hybrids shrink the memory each conversation needs, but no public fleet data yet shows the energy saved at scale.

Sources IEA 2025; IEA 2026; Google 2025; ML.ENERGY.andrewhendel.com

The macro picture

The scale of spending is hard to overstate. The International Energy Agency (IEA) reports in its 2026 Key Questions on Energy and AI that big technology companies spent more than USD 400 billion on capital expenditure in 2025, and that this is expected to rise another 75% in 2026. The IEA notes that the capital spending of just five technology companies is now larger than global investment in oil and gas production. Its own estimates imply a forecast of USD 3.9 trillion of cumulative data center investment between 2026 and 2030.

The chip side shows the same curve. NVIDIA reported data center revenue of USD 89.0 billion for the quarter ended July 26, 2026, up 117% from a year earlier.

The macro picture

Spending, electricity, and memory prices all moved by large multiples

Capex for 2025 is reported. Data center electricity for 2025 is an IEA estimate, since companies had not yet published full year data. The memory price change covers late 2024 to early 2026.

$400B+

Big tech capital spending in 2025, with a further 75% rise expected in 2026

IEA 2026

485 TWh

Estimated global data center electricity use in 2025, about 1.5% of world demand

IEA 2026

10x

Rise in memory prices between late 2024 and early 2026, an order of magnitude

IEA 2026

Source IEA, Key Questions on Energy and AI (2026).

The buildings are changing too. The IEA says, based on satellite tracking, that AI focused data centers have more than tripled in capacity in 18 months, and that the power density of AI servers rose 11 times between 2020 and 2025. Grid transformers take two to three years to deliver, gas turbines around five, and grid connections five to ten years in many places.

Memory is the other bottleneck, and it matters for this series. Modern AI chips sit next to stacks of high bandwidth memory (HBM), a fast and expensive kind of DRAM. The IEA reports a shortage of HBM that is “anticipated to persist through at least the end of 2027,” and notes that making HBM takes three times the wafer capacity of conventional memory for the same number of gigabytes. Market leading GPU systems had less than 150 GB of memory in 2023. The IEA, citing SemiAnalysis, expects more than 1,000 GB by 2027. A design that needs less memory per user stretches a scarce resource.

Electricity, measured versus forecast

The IEA’s April 2025 Energy and AI report estimated that data centers used about 415 TWh in 2024, around 1.5% of the world’s electricity. The 2026 update estimates that use grew more than 15% in 2025 to about 485 TWh. Both are estimates built from company reports and models, not a meter on every building. Everything after 2025 is a forecast. The IEA’s Base Case has data centers roughly doubling to 950 TWh by 2030, around 3% of global demand.

Electricity

Data center electricity could roughly double by 2030

Global data center electricity use in terawatt hours. 2024 and 2025 are IEA estimates of past use. 2030 is a forecast (IEA Base Case), not a measurement.

2024Estimate of past use
415 TWh
2025Estimate of past use
485 TWh
2030Forecast, IEA Base Case
950 TWh (forecast)
Show as table
ItemValue
2024 (Estimate of past use)415 TWh
2025 (Estimate of past use)485 TWh
2030 (Forecast, IEA Base Case)950 TWh (forecast)
The 2030 bar is a projection and carries a wide band of uncertainty. Sources IEA, Energy and AI (2025); IEA, Key Questions on Energy and AI (2026).

Locally, the share is far higher. Lawrence Berkeley National Laboratory (LBNL) estimates that US data centers used 176 TWh in 2023, 4.4% of US electricity, and forecasts 6.7% to 12% by 2028. The Electric Power Research Institute (EPRI) forecasts 9% to 17% by 2030 and says Virginia is already the one state where data centers use more than 20% of electricity.

Why do these ranges differ by a factor of two? Method and boundary. LBNL builds its estimate bottom up from shipments of chips and equipment. EPRI builds from project pipelines, and cautions that announced capacity “should be treated as a pipeline indicator rather than a near-term peak forecast.” Overhead matters too. Power usage effectiveness (PUE) is the ratio of total facility power to the power reaching the computers. Google reports a fleet PUE of 1.09 in its 2026 environmental report, much lower than the industry average it cites. By our arithmetic, a per query estimate that assumes a PUE of 1.5 instead of 1.1 comes out about 35% higher on that factor alone.

The per query numbers, and why they disagree

Several groups have put a number on a single chatbot reply. They agree for a simple text prompt, then diverge sharply.

  • Google measured its median Gemini Apps text prompt in production in May 2025 at 0.24 Wh and 0.26 mL of water. Its boundary includes idle capacity, host processors, and facility overhead. Google also reports a 33x drop in energy per median prompt over one year.
  • OpenAI’s chief executive stated that an average ChatGPT query uses about 0.34 Wh. No method was published.
  • Epoch AI estimated about 0.3 Wh for a typical GPT-4o reply from first principles, and about 2.5 Wh for a 10,000 token input and roughly 40 Wh for a 100,000 token input.
  • Microsoft Research estimated a median of 0.31 Wh for a frontier model, rising 13x to 3.91 Wh for queries 15x longer.

The disagreements have simple causes. Boundaries differ (just the GPU, the whole server, or the whole building). Medians differ from averages, because a long tail of long prompts pulls the average up. Batch size and precision can move energy per token 3 to 5x on the same model, according to the ML.ENERGY Leaderboard. Most of all, the task differs. A “query” can be a short answer or an hour of agent work.

Per query energy

What kind of query it is matters more than which model answers

GPU electricity per query in watt hours, measured on open models with the standardized AI Energy Score method. GPU only, so real data center totals are higher. Agentic values are indicative of order of magnitude.

Medium language model
0.05 Wh
Large mixture of experts
0.31 Wh
Agentic
1.14 Wh
Reasoning
7.6 Wh
Agentic with reasoning
50 Wh
Show as table
ItemValue
Medium language model0.05 Wh
Large mixture of experts0.31 Wh
Agentic1.14 Wh
Reasoning7.6 Wh
Agentic with reasoning50 Wh
Source IEA, Key Questions on Energy and AI (2026), Figure 2.1.

The IEA notes that if every conventional web search became a simple AI text query, it would use less than 4 TWh a year, small against 485 TWh. The growth comes from other workloads, and no company publishing per query figures has broken its use down by workload.

Water and carbon

Water numbers are even harder to compare, because people count different things. Withdrawal is water taken from a source. Consumption is water that does not come back, mostly evaporated in cooling towers. Some estimates also count the water used by the power plants that supply the electricity.

A widely cited study by Li, Ren, and colleagues estimated that training GPT-3 in Microsoft’s US data centers could directly evaporate 700,000 liters of fresh water, and forecast that global AI could account for 4.2 to 6.6 billion cubic meters of water withdrawal in 2027. LBNL estimates that US data centers directly consumed about 66 billion liters of water in 2023, while the indirect water footprint through electricity was nearly 800 billion liters. By our arithmetic, that makes the power plant share about 12 times the onsite share. This explains a gap between two per reply figures. Google’s 0.26 mL counts onsite cooling only. Mistral’s life cycle assessment of a 400 token reply, 45 mL, includes water in power generation and hardware manufacturing. The ratio is about 170 to 1, and both can be correct.

Company totals are rising. Google reports its water consumption rose 34% in 2025, to 41 billion liters.

Carbon follows the same pattern. Google reports total ambition based emissions of about 14.5 million tonnes of CO2 equivalent in 2025, up 18% on the year and 81% since 2019, even as its operational emissions fell. Microsoft reports total emissions up 25% year over year in fiscal 2025, driven mainly by data center expansion and by an accounting change on renewable energy certificates. In Meta’s 2024 figures, 99% of market based emissions were Scope 3, the supply chain, and capital goods like servers and buildings alone were about two thirds of the total.

That points to embodied carbon, the emissions from making chips, servers, and buildings. The IEA estimates that a typical high performance server takes more than 10 MWh to manufacture versus more than 80 MWh to run over five years. On an average grid, running the hardware dominates. As operators buy more clean power, making the hardware becomes a larger share, which is the argument of Gupta et al. This matters for the memory story. Every extra HBM stack has a manufacturing footprint as well as a power draw.

When a language model generates text, it produces one token at a time (a token is a word or a piece of a word). For each token, the chip has to read the model’s weights and its memory of the conversation so far out of HBM. It does very little math per byte it reads.

Moving data is far more expensive than computing on it. In his ISSCC 2014 keynote, Mark Horowitz put a DRAM access at 1 to 2 nanojoules, a couple of orders of magnitude above an on-chip operation at about 10 picojoules (45 nm figures). A simulation study of accelerator designs finds that this generation phase, called decode, is memory bound and dominates inference energy. No public source we found gives the share of production GPU energy spent on memory movement, so treat that link as mechanism, not a measured fleet number.

The memory of the conversation is the KV cache. For every past token, each attention layer stores a key vector and a value vector, so the cache grows with every token of context, for every user. The vLLM paper reports 800 KB per token for a 13 billion parameter model. By our arithmetic, Llama 3 70B needs about 40 GiB of KV cache for one 128,000 token conversation, more than half the memory of an 80 GB H100.

That growth raises energy per token, not just token count. Memory is finite, so a bigger cache per user means fewer users fit on the chip at once. Batch size is how many requests share a pass over the weights, so a smaller batch spreads each weight read over fewer tokens.

Architecture and energy

Longer context raises the energy of every token, not only the number of tokens

The mechanism as described by the ML.ENERGY measurements on H100 and B200 GPUs. Figures in step 4 are the authors' measurements for one model.

  1. 1Context gets longer

    Long documents, long chats, agent traces, and reasoning models that write thousands of hidden tokens before answering.

  2. 2The KV cache grows with every token

    Each attention layer keeps a key and a value for every past token, per user, in HBM.

  3. 3Fewer users fit at once

    Memory that holds one long conversation could have held many short ones, so the batch shrinks.

  4. 4Each token costs more energy

    Weight reads are shared by fewer tokens. For Qwen 3 32B the authors report 1.5x to 2.1x higher energy per token on reasoning work, depending on the batch setting.

Sources ML.ENERGY Leaderboard v3.0 (2026); Chung et al. (2026).

The same study found that reasoning tasks used 25 times more energy per response than ordinary chat, from about ten times more output tokens combined with higher energy per token. The AI Energy Score team found that turning reasoning on for the same model raised GPU energy by 150 to 700 times.

How hybrid models can ease part of it

A state space model, such as Mamba, replaces the growing KV cache with a fixed size state, a compressed summary of the past that is updated for each new token. It does not grow with context. The catch is recall. A fixed summary loses detail, and studies show transformers are much better at copying and retrieving exact details from context. Hybrids keep a few attention layers for exact lookup and make the rest state layers. Part 2 explains the designs in detail.

In production hybrids the attention share is small. NVIDIA’s Nemotron-H uses roughly 8% attention layers. IBM’s Granite 4.0 H mixes Mamba-2 and transformer blocks at 9 to 1, and IBM claims over 70% less RAM for long inputs and many concurrent users. AI21 reports a 4 GB KV cache for Jamba at 256,000 tokens, versus 32 GB for Mixtral.

We can check the direction ourselves from the published Granite 4.0 H-Small configuration, which has 4 attention layers and 36 Mamba-2 layers.

Memory per conversation

A hybrid's memory per conversation grows far more slowly with context

Our arithmetic from the Granite 4.0 H-Small configuration, in GiB per sequence at 16 bit precision. The hybrid line is the real layout (4 attention layers plus fixed Mamba-2 state). The attention line is a counterfactual with all 40 layers as attention, same heads. Weights are excluded.

All attention (counterfactual)Hybrid (Granite 4.0 H-Small)
051015208K32K64K96K128KContext length (tokens)Memory per sequence (GiB)All attention (counterfactual)Hybrid (Granite 4.0 H-Small)
Show as table
Context length (tokens)All attention (counterfactual)Hybrid (Granite 4.0 H-Small)
8,1921.3 GiB0.2 GiB
32,7685 GiB0.6 GiB
65,53610 GiB1.1 GiB
98,30415 GiB1.6 GiB
131,07220 GiB2.1 GiB
Derived from a published config, not measured on hardware. Source Granite 4.0 H-Small model card and config.

At 128,000 tokens the counterfactual needs about 20 GiB per conversation and the hybrid about 2.1 GiB, roughly 90% less. The Mamba-2 state is about 74 MiB whether the conversation is 1,000 tokens or a million. Following the chain in the figure above, less memory per user means more users per chip, bigger batches, and fewer bytes moved per token. NVIDIA says Nemotron-H 56B generates 2.4 times more output tokens per second per GPU than comparable transformers at long inputs. Those are throughput claims, not energy measurements. Part 3 works through the serving cost math.

The honest limits

Architecture is one lever among many. Grid mix sets how much carbon each kilowatt hour carries. Cooling design sets water use and PUE. Utilization, meaning how busy the chips are kept, can matter as much as the model. Hardware generations help too. ML.ENERGY found B200 GPUs beat H100s on energy in 88% of comparisons, with a median 35% reduction, and that a mixture of experts model used 3.56 times less energy per token than a dense model of similar size. Quantization, batching, and smaller task specific models all move the number.

No public fleet measurement shows hybrid energy savings at scale. The evidence is memory arithmetic, lab throughput benchmarks run by the model builders, and the general link between memory traffic and energy. That is a reasonable basis for expecting savings at long context, not proof of them.

Savings are uneven. The Falcon-H1 team notes that “Transformers are slightly faster at shorter context lengths.” For short chats the weights dominate and a hybrid saves little. And not everyone is convinced. MiniMax, which shipped a hybrid, returned to full attention for its next model, citing weaker multi step reasoning at scale and less mature infrastructure.

Finally, efficiency can be consumed by more use. Google cut energy per median prompt 33x in a year, yet by our arithmetic from its own report, its data center electricity rose about 38% in 2025. The IEA describes this as a Jevons dynamic, where cheaper tokens invite more tokens. Hybrids could make long context and reasoning cheap enough that people use far more of both. That may be a good trade, but it is not automatically a smaller footprint.

In short, the part of AI’s energy bill that grows with context length is the part hybrids target, and memory is both an energy cost and a supply bottleneck. Whether that shows up on the grid depends on everything else in this post. Part 4 looks at what fixed size state means for the chips themselves.

Sources

  1. Key Questions on Energy and AI, International Energy Agency, 2026.
  2. Energy and AI, International Energy Agency, April 10, 2025.
  3. NVIDIA Announces Financial Results for Second Quarter Fiscal 2027, NVIDIA, August 26, 2026.
  4. 2024 United States Data Center Energy Usage Report, Lawrence Berkeley National Laboratory, December 2024.
  5. Powering Intelligence, Updated U.S. Data Center Scenarios, EPRI, 2026.
  6. Measuring the environmental impact of delivering AI at Google Scale, Elsworth et al., Google, August 21, 2025.
  7. The Gentle Singularity, Sam Altman, June 2025.
  8. How much energy does ChatGPT use?, Epoch AI, February 7, 2025.
  9. Energy Use of AI Inference, Efficiency Pathways, and Test-Time Scaling, Oviedo et al., Microsoft Research, revised June 2026.
  10. Diagnosing Inference Energy Consumption with the ML.ENERGY Leaderboard v3.0, ML.ENERGY Initiative, January 29, 2026.
  11. Where Do the Joules Go? Diagnosing Inference Energy Consumption, Chung, Wu, Ma, Chowdhury, January 30, 2026.
  12. AI Energy Score v2, Refreshed Leaderboard, now with Reasoning, Hugging Face, December 4, 2025.
  13. Making AI Less “Thirsty”, Li, Yang, Islam, Ren, Communications of the ACM, March 2025.
  14. Our contribution to a global environmental standard for AI, Mistral AI, July 22, 2025.
  15. 2026 Environmental Report, Google, June 30, 2026.
  16. 2026 Environmental Sustainability Report, Microsoft, July 9, 2026.
  17. 2025 Sustainability Report, Meta, September 2025.
  18. Chasing Carbon, The Elusive Environmental Footprint of Computing, Gupta et al., IEEE HPCA, 2021.
  19. Computing’s Energy Problem (and what we can do about it), Mark Horowitz, IEEE ISSCC, February 2014.
  20. Prefill vs. Decode Bottlenecks, SRAM-Frequency Tradeoffs and the Memory-Bandwidth Ceiling, Uppsala University, December 26, 2025.
  21. Efficient Memory Management for Large Language Model Serving with PagedAttention, Kwon et al., SOSP, September 2023.
  22. Mamba, Linear-Time Sequence Modeling with Selective State Spaces, Gu and Dao, December 1, 2023.
  23. Repeat After Me, Transformers are Better than State Space Models at Copying, Jelassi et al., February 1, 2024.
  24. Nemotron-H, A Family of Accurate and Efficient Hybrid Mamba-Transformer Models, NVIDIA, April 4, 2025.
  25. IBM Granite 4.0 announcement, IBM, October 2, 2025.
  26. granite-4.0-h-small model card and config, IBM, October 2025.
  27. Jamba, A Hybrid Transformer-Mamba Language Model, AI21, March 28, 2024.
  28. Falcon-H1 blog, TII, May 20, 2025.
  29. Why Did M2 End Up as a Full Attention Model?, Haohai Sun, MiniMax, October 29, 2025.