Google’s TPU Economics 👀
Google GCP CEO Thomas Kurian
@ThomasOrTK just said that their payback period on AI servers is <2 years, and half that on their own silicon. 😊 (at
@GoldmanSachs today)
But it’s not clear whether he means revenue payback or profit/cash-flow payback. 🤔
Let’s assume he means revenue payback, the more realistic of the two in my opinion. 😐 I’d also interpret the statement as payback for just the servers and not the entire datacenter.
This means that they monetize TPU clusters at about $30M per facility MW.
How do I get there?🤓
Currently, the list price for Ironwood ranges from $5.40/hr (with a 3 yr commitment) to $12/hr on demand. Google says 10MW of IT/pod power can support 9,216 TPU chips. However, this likely does not include the other energy needs (chillers, pumps, transformers, lighting, etc.) So assuming a very efficient modern Google AI data center with 1.10-1.15 PUE (power usage effectiveness), it means about 800-840 Ironwood TPUs per facility MW.
Assuming 820 TPUs per MW, at the 3 yr per hour rate of $5.40, that’s $39M/MW at 100% utilisation. I likely wouldn’t lower the utilization assumption because everyone is sold out everywhere. 😮
But I’d lower the rate for these two reasons: (1) TK paired his statement with a mention of 5 yr contracts, and (2) it’s possible volume enterprise contracts (eg for Anthropic) will be at a further discount with a longer commitment.
Here there is not list price to use for 5 yr contracts, so extrapolating a diminishing returns curve on the discount from the listed schedule (from on-demand to 1 yr to 3 yrs) suggests $4.50/hr for 5 year commitments (16% discount from the 3 year $5.40/hr rate). Add on top an additional 10% discount for scaled customers and you get $4.08/hr.
Thats how I estimate the ~$29M revenue per facility MW of TPUs.
TKs statement also has presents some suggestions on margin. Alphabet depreciates their servers and networking equipment over 6 years. With 2yr/1yr payback on AI servers/TPU respectively, it means every $2 server investment produces $1/$2 per year respectively. Depreciation for AI servers generally consume 33% of each revenue dollar ($2 investment divided by 6 years depreciation as percentage of $1 per year) whereas it consumes 17% of each revenue dollar for TPUs ($2 investment divided by 6 years as percentage of $2 per year). So there’s a 16 percentage point gross margin advantage for TPUs.
TPU revenue covers the capex in one year for assets that are depreciating over six years. Of course there are operating costs, but once you get long term commitments, it becomes an attractive asset. No wonder Google is spending on capex.
This whole post is speculative, but the margin part especially so. 😝 But it will be interesting to build out this model with more information to come. Tell me what I missed!