Compute futures only fail if every new GPU makes the last one worthless.
The market keeps proving the opposite.
A100s shipped ~6 years ago are still running at high utilization. CoreWeave has said its A100 capacity remains fully booked. SemiAnalysis’ 1-year H100 rental index rose nearly 40% even as Blackwell ramped.
The important point:
Compute futures do not require every GPU generation to hold the same price forever.
They require compute capacity to retain measurable economic value over time.
And GPUs increasingly look less like disposable technology and more like an economic fleet.
At ~$6/H100-hour for premium capacity, workloads need high value per token.
At ~$2–3/hour for contracted, spot, or older capacity, an entirely different set of workloads becomes economical.
That creates a hardware cascade:
→ Frontier training: newest systems — Rubin, GB300 NVL72
→ Real-time inference: B200, H200, H100
→ Batch/asynchronous inference: H100, A100
→ Long-tail workloads: L40S, L4, T4, V100, prosumer GPUs
Not every workload needs the best chip on Earth.
Think about airlines.
New widebodies fly the highest-value long-haul routes. Older aircraft move into lower-cost routes where their economics still work.
GPUs can behave similarly.
Frontier training optimizes for cluster performance and interconnect.
Real-time inference optimizes for latency.
Batch inference optimizes for throughput per dollar.
Embeddings, speech-to-text, recommendations, vision, smaller models, and on-prem workloads may simply need hardware that is materially faster than a CPU.
This is also Jevons paradox at work:
As compute becomes cheaper, we don't necessarily consume less of it.
We find more things worth computing.
Two constraints make this especially important.
1. Power is binding.
With fixed MW allocations and multiyear interconnection queues, the question increasingly becomes:
Which available GPU system produces the most economic value per MW?
2. Capital basis matters.
A fully depreciated GPU does not need to earn the return required by newly financed hardware.
It may only need to cover power + operations.
That means older hardware can undercut new hardware and still generate cash.
So obsolescence isn't triggered because NVIDIA announces a new chip.
It happens when:
operating cost > economic value produced
or
replacement economics become overwhelmingly superior.
That is why compute can support a real forward curve.
Oil markets distinguish grades.
Power markets distinguish nodes and delivery hours.
Compute markets can distinguish:
hardware class
performance
location
latency
energy efficiency
availability
A new GPU generation can reprice the curve.
It doesn't eliminate the curve.
Not every workload needs the best chip on Earth. And not every new GPU makes the last one worthless.