training and inference are pulling from the same fleet but the fungibility only runs one way, since a frontier post-training run needs coherent clusters with real interconnect and inference doesn’t, which means training can always pull capacity down from serving but serving can never backfill training. so if you’re sprinting a post-training run, and gpt-5.6 landing 8 days after fable came back off the commerce hold looks like a sprint to me, you’re pulling capacity that was already allocated to serving, which is the whole thing, because a run that was scheduled is already carved out of the revenue plan and doesn’t cost you anything you were counting on.
that capacity doesn’t book anything for the quarter even though it’s still depreciating and still drawing power, and you can’t really rent your way around it either because nobody is selling you 30MW of coherent training capacity on six weeks notice. so the only fleet you can actually move is your own.
then july comes and the run finishes, capacity goes back to serving, the model ships, and Friar tells investors the run rate is at $40b and jumped 20% in July with business customers up 32%. that’s arr not recognized revenue so it isn’t apples to apples with the quarter, but the direction is clear. which to me says the soft quarter and hot july are the same story. the compute went into building the next model instead of serving customers, so q2 ate the cost and july got the payoff
a ~400M gap (OAI’s missing revenue assuming 25% QoQ rev growth) annualizes to somewhere around 20-40MW, which against a fleet measured in gigawatts is a couple percent of the pool that can actually run the job.
Arkady basically priced this on the
$nbis call, $40-$50M/MW for time-bounded needs up to six months, specifically post-training sprints and model releases. that tier exists because the burst is real and recurring
i’m not saying this is what happened, i don’t have the segment data, but a soft quarter followed immediately by a sharp reacceleration is exactly the shape you’d get if the compute went somewhere else for a few weeks and the alternative reading requires demand to have softened and then un-softened inside of six weeks, which is a stranger thing to believe than a quarter where compute was reallocated from inference to an unscheduled training run