Which silicon serves inference decides the energy-per-token curve, but total datacenter power still roughly doubles by 2030 inference workload serving mix which silicon serves? GPU serving B200 ~1,000W / part ASIC serving Sohu math blocks <1/2 voltage* no independent tokens-per-watt ? inference serving slice — energy per served token ? 2024 2030 high low Etched served-power path — verifiable datapoints 2 MW San Jose office 10 MW 80k sq ft site GW-scale ambition 2027 total datacenter electricity (IEA, incl. training) — 2024 vs 2030 415 TWh ~945 TWh ~1.5% of world use roughly 2.3x 2024 accelerated AI servers ~30%/yr — roughly half the net increase rack density 10–15 kW to 50–150 kW; binding constraint shifts to grid lead times node-level gains do not cancel aggregate growth slice energy / token falls aggregate power still ~doubles
mix 55% slice energy / token mid

Move toward ASIC-heavy serving: the slice's energy per token falls at the node level — the aggregate datacenter reservoir above the outcome does not shrink. The ASIC chip-level efficiency is an Etched claim (asterisk), not independently verified.