my biggest takeaways from ClusterMax 3.0 (there are many, always a banger from
@JordanNanos and team). tldr demand for small managed clusters is rising exponentially and integration with the powered shell is more important as racks scale. DM if you are building/investing/tinkering:
- managing the cluster is a huge expense and complicated (172 ways an Nvidia GPU can fail). the market is decidedly moving in two directions: 1. labs buying bare metal at massive scale (100s of MWs) with talent and resources to manage own clusters vs. 2. start ups raising $10s to $100s of millions to spend on compute that need a neocloud to manage the clusters. paradox here! even CoreWeave has become focused on long-term bare metal contracts, which are easy to fund with an IG counterparty, rather than managed clusters, which have higher margins but which the market is less keen to fund (!)
- even as heavyweights take higher % of compute deployed (self managed), managed clusters for "smaller" actors is *still growing exponentially*. we experienced this first hand when we saw the demand for our small cluster at
@CrucibleCap
- serving open source is more lucrative than many realize - open source could lead to managed clusters taking on more share against frontier lab bare metal
- access to debt is the name of the game, rise in insurance that protects neoclouds and lenders against contract termination
- the Blackwell -> Rubin transition is much smoother for operators than Hopper -> Blackwell - VR adoption is ahead of schedule as a result.
- "to identify failures, we recommend monitoring dashboards. A good dashboard shows the failed component, the affected jobs, the current scheduler state, and when each check last ran." -> we use
@Aravolta_X25 monitor our cluster at
@CrucibleCap
- **** there is lots more to do be done on cybersecurity, this is a structural risk for the industry *****
- agentic coding is putting margin pressure on managed GPU clusters because customers that use AI tools are more willing to take bare metal or lightly managed cluster, but AI amplifies skill gaps. trust the experts if you need to !
- all documentation is still human-first, not agent-first, many systems are worse because they are not built for agents.
@nirvanalabsai is building bare metal infrastructure for agents
- all neoclouds need to have an endpoints business going forward as more and more customers are looking for the simplest way to purchase vanilla open source tokens for their workflows. Coreweave's new managed inference platform announced in March is now at $100M ARR
- there is a reason that CoreWeave is in a league of its own and why its separated from hundreds, literally, of startup neoclouds. startup clouds take notes! CoreWeave has automated its lifecycle controller for provisioning, testing and monitoring - keeps stats on all failures and uses intelligence to predict faults
- integration of colo + bare metal is going to be more important: as the physical demands of compute continue increasing, in CoreWeave’s view, “how the building operates is now part of how the computer operates-they are not two individual things anymore”
-another reminder that AMD's software stack is well behind Nvidia's, with datapoints, and GB300 stukk iffers best perf/dollar despite higher pricing
ClusterMAX 3.0 is here!
ClusterMAX 3.0 debuts with a comprehensive review of the neocloud industry, covering 77 providers.
We increase our market view to cover 323 providers, up from 209 in ClusterMAX 2.0, 169 in ClusterMAX 1.0, and 124 in the original AI Neocloud Playbook and Anatomy article.
We have now interviewed well over 200 end users of neoclouds as part of this research.
We update our itemized list of criteria across 10 categories, and update our direct descriptions of our expectations for Slurm, Kubernetes, Standalone Machines, Monitoring Dashboards, and Health Checks. All of this content is live on our website. We encourage providers to use these lists when developing their offerings. We still consider these lists as an amalgamation of our experience interviewing end users, making them representative of the features that end users expect from their cloud providers.
Nebius joins CoreWeave in the Platinum tier. While CoreWeave still sets the technical bar for others to follow, Nebius is now established as a provider that consistently commands a premium pricing over others. Strong business decisions by Nebius have put them in a position to serve an entire class of neolabs at seller’s prices.
Google Cloud joins Oracle in the Gold tier. Azure moves to Silver, Fluidstack moves to Unavailable, and Crusoe drops to Bronze. Lambda, Firmus and TensorWave remain in Silver, while GMI moves up to Silver from Bronze.
Many companies drop from Silver (or Gold) to Bronze or lower. We raise the bar this round as only 19 neoclouds globally achieve a Medallion rating.
We establish a tier between Bronze and Underperforming: the Participation Ribbon tier. 15 providers join this rating, which more accurately describes our opinion that they do the bare minimum to get by.