2 years ago, we achieved the first big milestone with
@exolabs.
We clustered 2 MacBooks to run Llama 405B.
It felt like magic.
The consensus was running this model was only possible in a data center.
We ran it on consumer hardware, on 2 M3 Max MacBook Pros.
Most people thought it was a gimmick. It only ran at 2 tok/sec!
But, we believed that improvements to the software, hardware, and models would all compound.
So that maybe in a few years, we thought, this would improve 10x in software, 10x in hardware, 10x models = 1000x.
That was the vision.
We imagined a world where you would have frontier intelligence running quietly on your desk.
Today is the day that vision became reality.
The M5 Ultra is a 10x step-change improvement vs the M3 Max we originally clustered.
That, compounded with software improvements like RDMA over Thunderbolt, MTP and better kernels, and high intelligence density models like Qwen 3.8 27B, means we now have 1,000x better Local AI than when we started.
I am so grateful to the small group of people at Apple (including
@doogie69 @awnihannun @angeloskath @DiganiJagrit @doogie69) who believed in this vision and had the foresight as well as the courage to take a swing at this early on. I'm confident they're just getting warmed up (looking forward to 4-bit / 8-bit compute units in M7 Ultra🤞).
With the M5 Ultra Mac Studio, we are going to have unmetered tokens running at API speeds at effectively zero marginal cost, running on your desk, so quietly and consuming so little power you won't even notice it.
Local AI is good now.
exo featured on Apple's new M5 Ultra Mac Studio and M6 / M5 Pro Mac Mini pages.
Over the past year, we have worked closely with Apple on low-latency RDMA networking over Thunderbolt 5, enabling clusters of Macs to run massive models like Kimi K3 and GLM-5.3 at API speeds.
With RDMA, aggregate memory bandwidth across Macs scales ~linearly.
A cluster of 4 x M5 Ultra Mac Studios scale to an aggregate memory bandwidth of ~4.8TB/s. Previously, these were speeds only achievable with data center GPUs.
Our vision is a data center on every desk. Apple Silicon's superior memory unit economics, power efficiency, and out of the box experience for Local AI make that possible.
Thank you to
@angeloskath,
@awnihannun,
@twid and countless others at Apple who tirelessly to bring this technology to the world.