CUDO delivers the full AI infrastructure stack for enterprise-scale AI, built on a power-first strategy and validated against NVIDIA reference architectures.

United Kingdom
What does it take to get AI infrastructure live? Land. Power. Data centres. Networks. Compute. And, behind all of it, people who can make those pieces work together. That’s the challenge we’re building CUDO around. It means the work here rarely sits neatly in one box. Our teams are solving problems across the infrastructure stack, making decisions quickly and taking real ownership of getting things delivered. So rather than tell you what it’s like to work at CUDO, we asked some of the people doing it. As CUDO continues to grow, we’re growing the team behind it too. We’ll be sharing more opportunities to join us over the coming weeks. For now, meet some of the people building it.
1
1
192
Two identical GPU clusters. Same hardware. Same generation. Very different results in production. Because GPU count and spec sheets only tell you what the hardware could do. They don’t tell you what it will sustain once a real workload is running. Fabric design. Collective communication. Storage throughput. Cooling. Power. Straggler detection. Runtime configuration. Day-to-day operations. They all determine how much of the performance you paid for actually reaches the workload. The difference can be significant. Published examples covered in our latest blog show comparable-era training environments achieving dramatically different levels of useful GPU utilisation, while inference performance can shift materially on unchanged hardware simply through changes to the serving stack and configuration. That changes the questions buyers should be asking. Not just: How many GPUs? Which generation? But: What throughput can this system sustain under load, and who is responsible for making sure it does? Hardware sets the floor. Throughput is built. Read why identical GPU clusters diverge in production, and what to look for beyond the spec sheet. cudocompute.com/blog/why-ide…
1
4
3,938
Are we planning the UK's AI infrastructure around where compute is going, or only where it is today? This week, our Chief Business Development Officer and Co-Founder, @CudoPete, joined industry leaders for the opening keynote panel of The Xcelerated Compute Show, exploring UK sovereign compute in the age of frontier AI. The conversation moved across sovereignty, AI Growth Zones, power, domestic capacity and what the UK needs to build if it wants to capture more of the economic value created by AI. But one of Pete's points looked beyond the infrastructure challenge immediately in front of us. Right now, the direction of travel is towards increasingly power-dense compute. Pete pointed to the prospect of one megawatt racks within the next three to five years. That creates an obvious question for an industry already working through constraints around land, power and grid capacity: how much further does that curve go? Pete challenged the assumption that it only moves in one direction. Emerging areas such as biological compute could eventually change the relationship between compute output and power density altogether. We could spend the next few years moving towards unprecedented densities, only for new architectures to begin pushing that requirement back down. And when the infrastructure we are planning today will be operating for decades, that matters. As Pete put it: “How do we get ahead of this? And how do we change the future, not just in the UK, but across the world?” The UK has the research capability, universities and talent to play a role in answering that question. So while we absolutely need to solve for today's constraints, from deliverable power to the right locations for AI infrastructure, we also need to think further ahead. What does the next generation of compute require? What could come after it? And are the infrastructure decisions we make today flexible enough for both? Land takes years. Power takes years. Compute can change much faster. Planning for AI infrastructure means planning for that difference.
2
323
The individual pieces of AI infrastructure can all look viable. The deployment can still be years away. Land can be secured without deliverable power. Power can be announced without an energisation date. Data centre capacity can exist without the network, hardware or operational capability needed to turn it into usable compute. That is the coordination problem sitting underneath the AI infrastructure build-out. And it is showing up in the decisions companies are making. In our Land. Power. Compute. research, 96% of AI infrastructure decision makers said power constraints or geopolitics had already caused them to change infrastructure plans. Because hyperscale AI does not buy land, power, data centres, networks and compute on separate timelines. It buys capacity against a deployment window. The question is not whether each piece exists. It is whether they can be brought together in the right place, at the right time, with economics that work. That is the conversation we are taking to DCD Connect London on 16–17 September. @CudoPete, Kir Devaser, Kaitlin Argeaux, Mark Richman, Dean Fletcher, Barry Kick and Michael Jones will be there across both days. If you are working through where capacity is going, what will actually be deliverable and what has to line up to get compute live, find one of us. And if you think the industry is looking at the wrong constraints, we especially want to hear why.
2
222
A power contract for a facility is a 10-15 year commitment. GPU roadmaps move on a much shorter cycle. How do you future-proof a facility when the hardware moves faster than you can build one? Barry Kick, our AI Business Development Director, puts it this way: “Efficient delivery of solutions that are close to value, powered by infrastructure that has the optimal technical and operational fit, that is the winning formula.” You commit to a site, a substation and a power contract for the rest of the decade. Then you design the building for hardware that will be two generations old by the time the substation is energised, and may not exist yet when the drawings are signed. You don't future-proof that by forecasting. Nobody is going to correctly call the rack density, the cooling regime and the power draw of a 2031 accelerator from a 2026 site plan. What you can do is replace forecasting with situational awareness. Know what is actually deliverable rather than what is announced, where power genuinely exists, when a grid connection will realistically land, and what the build constraints are on that specific site. Then design so that a hardware change costs you a retrofit and not a rebuild. That challenge came through in our Land. Power. Compute. report too. Across 701 infrastructure decision-makers in the UK, US and Europe, power availability was the most cited deployment blocker. Because ultimately, the constraint is the chain, not the chip.
1
4
1,097
Compute capacity gets the headlines. Power availability is what actually decides who builds, where, and how fast. Power availability and grid capacity now rank as the leading factors in AI infrastructure location decisions. But having power on paper isn't enough. A site needs the capacity to support the workload, a realistic route to getting that power online and the resilience to keep it running. Land without a clear power pathway is an option, not an asset. As @WSP’s Anushka Devaser puts it: “The biggest misunderstanding is assuming securing power is only about price or MW capacity.” AI infrastructure needs more than an available MW figure. It needs firm, resilient power that can support the workload in practice. That’s why land, power and compute need to be considered together. Decisions around one quickly affect what’s possible with the others. Read our Land. Power. Compute. research below.
1
1
4
260
Announced power and power you can actually draw on are rarely the same number. That is what we want to talk about in Austin this week. Power availability is the most cited deployment blocker across the US, UK and Europe. It stands at 32.1% according to the 701 AI infrastructure decision makers we surveyed for Land. Power. Compute. Which puts the risk in the chain, not the chip: whether the power behind the GPUs arrives inside your window. That gap is our subject because we carry it. CUDO delivers the whole chain under one contract, land, power, data centre, network and compute on an SLA, inside the window you buy in. One accountable party rather than four. Two views we are taking with us, on 2027 capacity: the number that matters is not what you have reserved, it is the date it energises. On placement: the workload and the deliverable power decide where compute sits, not land availability. If you see it differently, we would like to hear it in the comments or in person. Dean Fletcher, our Strategic Alliance Director, and Mohamed Alibi, our Senior HPC Engineer, are at Metro Connect Fall and Datacloud USA.
4
208
21.3% → 55.2% MFU. That’s more than a 2× difference in useful work on comparable-era hardware. GPT-3 175B achieved 21.3% MFU. MegaScale reached 55.2%. The difference is not just the GPU. It’s the system around it. DeepSeek-V3 is a useful example. It ran on 2,048 NVIDIA H800 GPUs, hardware built on the same Hopper architecture as the H100 but with NVLink bandwidth reduced from 900 GB/s to 400 GB/s for export compliance. That single difference in interconnect reshaped the entire training strategy. The team avoided tensor parallelism because it was inefficient under the constrained NVLink, built a bidirectional pipeline scheme to overlap computation and communication, and relied on eight 400G InfiniBand NICs per node to compensate for the loss of intra-node bandwidth. The chip was nearly the same. The configuration determined what the cluster could actually do. Even after that hardware-aware co-design, third-party analysis estimates the run landed at around 23% model FLOPs utilization, a figure the DeepSeek team never had to disclose and one that rarely appears in headline infrastructure specifications. GPU model and count only tell part of the story. What matters is how effectively the entire system turns that hardware into useful compute. Read the article below.
1
2
3
262
While the GPU shortage gets the headlines, having access to GPUs doesn’t mean much if the infrastructure isn’t ready to deploy them. Power is becoming one of the biggest constraints. Ben Stirk, Head of Enterprise and AI at @Arkdatacentres, put it plainly in our Land. Power. Compute. report: “If you don’t already have power secured for a site, or if you haven’t already placed the orders to get it there, you aren’t just behind. You aren’t even in the race.” The reason is simple. GPUs can be procured on relatively short cycles. Power can’t. Securing enough power for high-density AI infrastructure can take years, which means decisions being made now can determine where organisations are actually able to deploy and scale later. Our research backs this up. Across 701 AI infrastructure decision-makers surveyed with Censuswide, power availability is the leading factor influencing infrastructure location at 33.7%. And when we looked at what is driving the cost of compute, 71% of selections were infrastructure-related. GPU pricing ranked sixth. For organisations planning AI infrastructure, the challenge is no longer just securing the compute. It’s making sure the land, power and infrastructure are ready for it. That’s what Land. Power. Compute. looks at: the constraints shaping where AI gets built, what it costs and how quickly it can scale. Read the research below.
1
1
8
252
Two identical GPU clusters. Same datasheet. Different outcomes in production. Any competent buyer can read the spec sheet and sign the purchase order. What separates the two clusters is everything that happens after: the fabric tuned to the topology, the collectives configured for the actual job, storage feeding data fast enough to keep the GPUs busy, stragglers caught before they cost you a day, the thermal envelope held under load. In serving, it's the same story, the inference stack and precision matched to the workload, and a software stack mature enough to realise what the silicon promised. The two diverge because that work is never identical. The hardware is bought. The throughput is built. Read the full breakdown below.
1
5
228
CUDO Compute retweeted
Advanced AI capabilities are no longer available only through closed third-party APIs. Enterprises can choose where models run, how they are adapted and what happens to the data. That gives organisations greater control over a capability that is fundamental to how they operate.
1
2
5
262
Future-proofing AI data centres means reducing the cost of being wrong. When they are built with modularity in mind, expansion or upgrades can often be done without interrupting existing operations, helping maintain availability. Modular data centres are often designed with energy efficiency in mind, featuring optimized power distribution and cooling systems tailored to each module's specific needs. Modular growth offers a more adaptable approach by aligning capital deployment with supply availability. Rather than committing to 200 MW of infrastructure that may sit partially stranded while awaiting GPU allocation, operators can deploy 25–50 MW pods as hardware becomes available. Neither is universally right, but phased deployment matches CapEx to actual deployment velocity and reduces the risk of your power and cooling architecture aging out before the site is full. Future-proofing does not mean predicting every change. It means structuring the investment so that getting one assumption wrong does not compromise the entire deployment. Operators can and should design for change. Read more in our blog below.
1
5
259
As enterprises connect AI systems to documents, databases and operational tools, the risk changes. A malicious instruction no longer needs to come directly from a user. It can be hidden in a webpage, email or file that an AI system has been asked to process. If the system treats that content as trusted, the model could disclose sensitive information, manipulate an output or attempt an unauthorised action. That is prompt injection, and no single filter or carefully written system prompt can eliminate the risk. The practical response is a defence-in-depth approach: • Treat external content as untrusted • Give models and agents the minimum access they need • Isolate sensitive workloads and data • Validate outputs before they reach other systems • Require human approval for consequential actions • Continuously test how the system behaves under attack This makes security architecture an essential part of the infrastructure decision. The question is no longer: “Which model should we deploy?” “It’s a case of zero trust regardless of how good the model is.” - Mohamed Alibi, Senior HPC Engineer at CUDO Compute. At CUDO, we believe enterprise AI infrastructure should be evaluated on the control and resilience it provides, not compute performance alone. What safeguards are you building around your AI systems before they move into production?
4
245
Why can you buy two identical B200 clusters and get different throughput every time? The same GPU generation and nominal interconnect specifications do not guarantee the same real-world results. Hardware establishes the potential, but how closely a cluster gets to that potential depends on what happens after procurement: architecture, configuration, workload optimisation and day-to-day operations.
1
1
7
257
Configuration matters too. When training a 70B model on eight H100 GPUs, a poorly configured NCCL setup can consume 20–30% of each step in communication. Tuning it to the cluster topology can reduce that to 5–8% potentially saving 10–25% of the training cost at scale. Then there is resilience. @Meta's Llama 3 training run used 16,384 H100 GPUs over 54 days and recorded 419 unexpected interruptions, roughly one every three hours with more than half related to GPUs or HBM3 memory. Despite this, the team maintained more than 90% effective training time. That was an operational achievement, not simply a feature of the hardware.
2
3
80
At CUDO Compute, performance is measured in production, not on a specification sheet. Hardware is only one part of the equation. The architecture, configuration and operational expertise behind a cluster determine how consistently it delivers at scale. Read the full article below. cudocompute.com/blog/why-ide…
3
59