The Superintelligence Cloud

San Francisco, CA
Pinned Tweet
AI infrastructure is the largest industrial buildout of our lifetime. Lambda is assembling the leadership team to match the opportunity ahead. Today, Lambda welcomes global infrastructure operator Michel Combes as CEO and former AT&T Communications CEO John Donovan as Chairman of the Board. Co-founder Stephen Balaban takes on the CTO role full-time, shaping the technology that will define the next decade of AI compute. Read the exclusive from @business: bloomberg.com/news/articles/…
8
8
56
12,522
AI compute is not becoming a commodity. The industry is building vertically integrated infrastructure, reaching from software and networking to the physical sites and power that make large-scale compute possible. The stakes are physical. Lambda Co-Founder & CTO Stephen Balaban makes that case on the Main Stage at Yotta 2026. He will also discuss our target of 3 GW of AI compute capacity by 2030.
2
4
14
1,971
The argument in one line. AI infrastructure is moving from traditional data centers to gigawatt-scale campuses, and the companies building it are changing with it.
1
215
Day 2 raises a different question. If you are building an AI-native enterprise, when do you rent compute, when do you reserve dedicated capacity, and when does owning it become the competitive advantage? Lambda President of Cloud Services David Ward joins a panel on exactly that. It takes place Wednesday, September 30, 11:40 a.m.–12:25 p.m., in the Future of Compute track. Panel page at yotta-event.com/the-ai-nativ…
1
162
Stephen Balaban (@stephenbalaban) has been through 5 pivots in 14 years. He started with facial recognition and AI filters before building the workstation and GPU compute business. The interview is a look at the long, non-linear path behind Lambda, and the decisions that shaped it along the way. Listen to the full conversation with @LambdaAPI’s CTO below on the @FoundersInArms podcast with @immad and @rajatsuri. foundersinarms.substack.com/…
1
5
27
2,206
Ask a 3D vision-language model what's near the table and in front of the curtain, and it might guess "sewing machine." The right answer is a tray rack. CVP (UC San Diego + Lambda, WACV 2026) fixes this with a target-affinity token for task-relevant objects and an allocentric grid for global context. Against Video-3D-LLM: • SQA3D EM: 58.6 → 62.3 • Scan2Cap CIDEr: 83.8 → 90.5 • Better on all 5 benchmarks tested Full results across ScanQA, SQA3D, ScanRefer, Multi3DRefer, and Scan2Cap, plus how the central/peripheral split works: lambda.ai/blog/cvp-spatial-r…
7
19
2,174
.@stephenbalaban on why Lambda prices compute like a utility, dollars per GPU, not per token, in Georgia Butler's latest Tokenomics piece for @dcdnews: "…in the same way that a utility provider might look at selling dollars per kilowatt hour, we look at dollars per GPU."
DCD Magazine issue 62 out now: The coming wave dlvr.it/TVWlln
1
6
1,095
AI for molecular dynamics has a data problem. The trajectories you need to train on are expensive. EGInterpolator (ICLR 2026, with Stanford) learns molecular structure first from abundant conformer data, then uses scarce MD data to learn motion. On the DRUGS benchmark, it reduced the gap to reference simulations by 73% for bond angles, 78% for bond lengths, and 24% for torsional motion. The structure-first ablation also matters. Removing pretraining increased mean JSD from 0.173 to 0.332 for bond angles and from 0.142 to 0.386 for bond lengths.
3
3
3
884
Lambda retweeted
Wearable AR has had a decade of attempts and still hasn't found the product. The pattern is familiar: a technology looks inevitable long before it's actually usable. Worth remembering the next time something seems like it should obviously work. More from our conversation with @stephenbalaban of Lambda on Founders in Arms. Link in bio.
12
3
29
17,533
Lambda retweeted
More GPUs do not guarantee more progress. Better orchestration does. With @lambdaAPI, we increased GPU utilization from ~20% to 43% and cut queue starvation by 74%. Make every GPU count. #SPREEAI #Lambda #AIInfrastructure
GPU utilization increased from ~20% to 43% on a reservation of 96 NVIDIA H100 GPUs, while cutting queue starvation by 74%, with blocked jobs falling from roughly 10 per day to around 4. That’s the concrete result @SpreeAI saw after fixing their orchestration. When a unified diffusion model requires 80–100 GB of memory, you can’t simply throw workloads at a cluster and expect to use those GPUs efficiently. SPREEAI was dealing with workload fragmentation, ad-hoc submissions, and storage I/O blocking that left expensive GPUs idle. Working with Lambda’s ML engineering team, they implemented MLflow-based experiment orchestration with structured queuing and workload matching. They also connected Lambda’s Prometheus APIs to Grafana for real-time visibility into utilization gaps. The video testimonial covers how they diagnosed the bottlenecks and what the remediation looked like.
1
1
1
516
GPU utilization increased from ~20% to 43% on a reservation of 96 NVIDIA H100 GPUs, while cutting queue starvation by 74%, with blocked jobs falling from roughly 10 per day to around 4. That’s the concrete result @SpreeAI saw after fixing their orchestration. When a unified diffusion model requires 80–100 GB of memory, you can’t simply throw workloads at a cluster and expect to use those GPUs efficiently. SPREEAI was dealing with workload fragmentation, ad-hoc submissions, and storage I/O blocking that left expensive GPUs idle. Working with Lambda’s ML engineering team, they implemented MLflow-based experiment orchestration with structured queuing and workload matching. They also connected Lambda’s Prometheus APIs to Grafana for real-time visibility into utilization gaps. The video testimonial covers how they diagnosed the bottlenecks and what the remediation looked like.
1
15
1,827
Huh, this is rather fascinating. Lambda Labs, the AI neocloud company, originally sold an HD camera-enabled baseball cap. The idea, as I surmise, was to acquire images for convolutional neural networks, a different technology from the now more well known large language models.
People ask us, how did you know cloud compute would be so important back in 2015? And I say I didn’t. I don’t know shit. 1517 invested in a hat company.
1
4
623
Evaluating image editing models with a single score hides the nuance behind a lower score. EdiVal-Agent (ICLR 2026) splits evaluation across instruction following, content consistency, and visual quality. Its agentic judge reached 81.3% agreement with human judgments, compared with 75.2% for a VLM-only evaluator and 68.9% for CLIP_dir. lambda.ai/blog/edival-agent-…
1
7
7
1,018
The failure that stands out is how quickly multi-turn editing exposes weak preservation. Each new instruction has to land without undoing earlier edits or changing unrelated content. Errors compound even when single-turn results look strong.
1
4
458
EdiVal-IF measures instruction following. EdiVal-CC covers content consistency, while EdiVal-VQ checks visual quality. The paper came from a collaboration between UT Austin and UCLA, with Microsoft and Lambda. arxiv.org/abs/2509.13399
231
Lambda retweeted
!!!
12
2
48
2,237
Build or buy AI? Wrong question. Lambda's @boborado is on the main stage at AI Infra Summit 2026 with @JLL CTO Yao Morin, @usbank EVP & Chief AI Officer Prashant Mehrotra, and @carrier Chief Data & AI Officer Arun Nandi, and the room's landing on the same answer: it's build AND buy. New term coined: valuemaxxing. Using value per task as a lens to govern resource allocation, while architecting around agility (easy model swaps) and availability (SLAs and model lifecycle).
1
3
14
1,072
MLPerf Inference v6.1 is out, and we have two firsts: the first agentic inference workload on datacenter hardware, and the first MLPerf deployment of a model over a trillion parameters. On 4x NVIDIA Blackwell Ultra GPUs, we posted the leading Offline throughput on GPT-OSS 120B among all Blackwell Ultra GPU submissions, plus an 8.85% throughput gain over v6.0 on identical hardware. That's six months of pure software optimization. On NVIDIA HGX B200, we swapped in Kimi K2.6 for the open-division agentic benchmark, running a 1T+ parameter model where the reference workload expects 27B. Same harness, no memory ceiling. Full results and methodology in the blog. Link below. lambda.ai/blog/mlperf-infere…
6
6
19
1,356
VLM inference made its debut in our lineup too. Qwen3-VL on NVIDIA HGX B200 came in 28.5% faster (Offline) than the fastest v6.0 submission on comparable hardware. The agentic number that matters most: 1,007 of 1,007 replay turns completed, zero failures, 86.83% BFCL v4 accuracy, at a fraction of the latency of the 27B reference deployments.
1
1
355
MLPerf keeps expanding into what enterprise teams actually run: reasoning, multimodal, now agentic. Lambda's infrastructure is ready for all three. 1-Click Clusters, 16 to 1,536+ GPUs, no contracts required." lambda.ai/1-click-clusters
1
199