Frontier AI on local hardware. EXO 1.0 is now open-source (Apache 2.0): github.com/exo-explore/exo

The future of AI is open source and decentralized
"Exo's use of Llama 405B and consumer-grade devices to run inference at scale on the edge shows that the future of AI is open source and decentralized." - @mo_baioumy
35
117
984
469,692
💛
Just @alexocheema and @exolabs rocking it, and with yet more coming down the pike reuters.com/business/retail-…
2
1
11
4,298
EXO Labs retweeted
I interviewed DHH, the creator of Omarchy - The Agent Operating system + Ruby on Rails, one of the most important frameworks of our time. He's a father, builder, race-car driver, and such an inspiring person. One of the most positive people I've had the honor of meeting.
83
122
1,546
184,869
EXO Labs retweeted
In the Lab rn! Exo is cooking!
58
10
515
32,993
EXO Labs retweeted
My GTM lead asked me to share this. For serious business people pls. tally.so/r/lbx4l6
2
1
16
3,283
EXO Labs retweeted
Replying to @Austen
@johnternus has the clear path to owning sovereign AI on the desktop and phone thanks to Jobs making uber own silicon AND ramping up support for @exolabs AND putting 128gigs+ in machines Predict they buy @perplexity_ai
12
11
175
126,579
EXO Labs retweeted
Apple's AI strategy is Apple Silicon. The iPhone, iPad, MacBook, Mac Mini and Mac Studio all use the same hardware architecture. Apple Silicon is energy efficient, quiet, and the memory unit economics are incredible. Apple has leaned hard into Local AI, now serving every segment of Local AI, from SLMs running on an iPhone/iPad to 2T models running on a cluster of 4 x 512GB M5 Ultra with RDMA using @exolabs (the cluster has 2TB memory @ 4.8TB/s). Most inference will run locally, and Apple wants it to run on their silicon. The main issue I see right now is the software. We need a stable, canonical inference engine rather than 10 unstable, incomplete ones. Who is solving this?
Replying to @Austen
@johnternus has the clear path to owning sovereign AI on the desktop and phone thanks to Jobs making uber own silicon AND ramping up support for @exolabs AND putting 128gigs+ in machines Predict they buy @perplexity_ai
43
58
888
105,952
EXO Labs retweeted
hiring a founding designer for @exolabs in-person in SF $20k referral fee jobs.ashbyhq.com/exo/2d9292b…
8
9
66
9,802
EXO Labs retweeted
Interesting article in @theinformation. Macs are Apple's fastest growing business right now, driven by massive AI demand. Apple's AI strategy IS the Mac. Trillion dollar opportunity imho. theinformation.com/articles/…
16
22
193
11,388
Our patch enabling direct USB-C networking between Apple Silicon Macs and DGX Spark has landed in the mainline Linux Kernel. This simplifies heterogeneous Local AI setups like P/D disaggregation over an ordinary USB-C cable. github.com/torvalds/linux/co…
93
208
1,696
292,396
Before the patch, Linux detected the Mac but used the generic CDC-NCM config. This expected an interrupt endpoint but the Mac doesn't provide that, so the driver failed to attach and no network device would appear. Previously there were hacky workarounds like manually patching the kernel. This small patch makes it so Linux exposes usable network interfaces out of the box. github.com/torvalds/linux/co…
6
6
87
13,823
macOS 27
13
3
164
24,762
EXO Labs retweeted
why do I have four mac studios? because my inference cost is zero dollars. > 291B parameters > tensor-sharded across four ultra chips > wired in a thunderbolt 5 ring > at 80Gbps with RDMA sorry anthropic
54
6
231
37,238
EXO Labs retweeted
Someone posted this on reddit and now it’s #1 on r/LocalLLM! Answering questions about exo and local AI in the thread.
exo featured on Apple's new M5 Ultra Mac Studio and M6 / M5 Pro Mac Mini pages. Over the past year, we have worked closely with Apple on low-latency RDMA networking over Thunderbolt 5, enabling clusters of Macs to run massive models like Kimi K3 and GLM-5.3 at API speeds. With RDMA, aggregate memory bandwidth across Macs scales ~linearly. A cluster of 4 x M5 Ultra Mac Studios scale to an aggregate memory bandwidth of ~4.8TB/s. Previously, these were speeds only achievable with data center GPUs. Our vision is a data center on every desk. Apple Silicon's superior memory unit economics, power efficiency, and out of the box experience for Local AI make that possible. Thank you to @angeloskath, @awnihannun, @twid and countless others at Apple who tirelessly to bring this technology to the world.
6
8
106
15,380
EXO Labs retweeted
exo is currently top of r/LocalLLM (in response to this tweet). answering questions there in detail on my alt: Longjumping_Crow_597. feel free to ask any questions, exo or local AI related. link: teddit.net/r/LocalLLM/commen…
exo featured on Apple's new M5 Ultra Mac Studio and M6 / M5 Pro Mac Mini pages. Over the past year, we have worked closely with Apple on low-latency RDMA networking over Thunderbolt 5, enabling clusters of Macs to run massive models like Kimi K3 and GLM-5.3 at API speeds. With RDMA, aggregate memory bandwidth across Macs scales ~linearly. A cluster of 4 x M5 Ultra Mac Studios scale to an aggregate memory bandwidth of ~4.8TB/s. Previously, these were speeds only achievable with data center GPUs. Our vision is a data center on every desk. Apple Silicon's superior memory unit economics, power efficiency, and out of the box experience for Local AI make that possible. Thank you to @angeloskath, @awnihannun, @twid and countless others at Apple who tirelessly to bring this technology to the world.
4
5
53
9,483
EXO Labs retweeted
2 years ago, we achieved the first big milestone with @exolabs. We clustered 2 MacBooks to run Llama 405B. It felt like magic. The consensus was running this model was only possible in a data center. We ran it on consumer hardware, on 2 M3 Max MacBook Pros. Most people thought it was a gimmick. It only ran at 2 tok/sec! But, we believed that improvements to the software, hardware, and models would all compound. So that maybe in a few years, we thought, this would improve 10x in software, 10x in hardware, 10x models = 1000x. That was the vision. We imagined a world where you would have frontier intelligence running quietly on your desk. Today is the day that vision became reality. The M5 Ultra is a 10x step-change improvement vs the M3 Max we originally clustered. That, compounded with software improvements like RDMA over Thunderbolt, MTP and better kernels, and high intelligence density models like Qwen 3.8 27B, means we now have 1,000x better Local AI than when we started. I am so grateful to the small group of people at Apple (including @doogie69 @awnihannun @angeloskath @DiganiJagrit @doogie69) who believed in this vision and had the foresight as well as the courage to take a swing at this early on. I'm confident they're just getting warmed up (looking forward to 4-bit / 8-bit compute units in M7 Ultra🤞). With the M5 Ultra Mac Studio, we are going to have unmetered tokens running at API speeds at effectively zero marginal cost, running on your desk, so quietly and consuming so little power you won't even notice it. Local AI is good now.
exo featured on Apple's new M5 Ultra Mac Studio and M6 / M5 Pro Mac Mini pages. Over the past year, we have worked closely with Apple on low-latency RDMA networking over Thunderbolt 5, enabling clusters of Macs to run massive models like Kimi K3 and GLM-5.3 at API speeds. With RDMA, aggregate memory bandwidth across Macs scales ~linearly. A cluster of 4 x M5 Ultra Mac Studios scale to an aggregate memory bandwidth of ~4.8TB/s. Previously, these were speeds only achievable with data center GPUs. Our vision is a data center on every desk. Apple Silicon's superior memory unit economics, power efficiency, and out of the box experience for Local AI make that possible. Thank you to @angeloskath, @awnihannun, @twid and countless others at Apple who tirelessly to bring this technology to the world.
48
61
782
69,278
exo featured on Apple's new M5 Ultra Mac Studio and M6 / M5 Pro Mac Mini pages. Over the past year, we have worked closely with Apple on low-latency RDMA networking over Thunderbolt 5, enabling clusters of Macs to run massive models like Kimi K3 and GLM-5.3 at API speeds. With RDMA, aggregate memory bandwidth across Macs scales ~linearly. A cluster of 4 x M5 Ultra Mac Studios scale to an aggregate memory bandwidth of ~4.8TB/s. Previously, these were speeds only achievable with data center GPUs. Our vision is a data center on every desk. Apple Silicon's superior memory unit economics, power efficiency, and out of the box experience for Local AI make that possible. Thank you to @angeloskath, @awnihannun, @twid and countless others at Apple who tirelessly to bring this technology to the world.
63
117
1,365
257,858
EXO Labs retweeted
Let’s go boys. Can’t wait to hook this up to my M3 Ultra and cluster with @exolabs! Arvis is going to enjoy some new memory!
2
18
2,059
EXO Labs retweeted
1
3
28
4,252
EXO Labs retweeted
Guess what? local ai for everyone (: we really did it local.ai
194
60
2,350
381,536