Pluralis is a research lab focused on collectively-owned AI.

Today we're releasing Agora: the first ever pretraining stack that allows non-collocated consumer GPUs to be competitive with centralized clusters Agora is 15x faster than Megatron-LM in this setting and is only 1.5x less efficient in terms of tokens per unit compute than TorchTitan on H100s, despite running on devices that have no NVLink or InfiniBand support.
29
45
314
93,223
We are fortunate to have 3 papers accepted to this years NeurIPS. The first is AsyncMesh - a technique for asynchronous updates across both data and pipeline parallelism simultaneously. Second is NuMuon, which shows the low-rank structure we observe in Subspace Networks is also present and can be exploited under Muon-style optimisers. The third introduces a technique to finetune open-weight models in the Protocol Learning setup. This is significant as prior to this work, this kind of distributed training required architectural changes, and hence starting from scratch.
7
9
65
6,014
Pluralis Research retweeted
My position is typically to focus on getting things to work rather than making noise about what needs to be done, but enough people have messaged me in last 24 hours that felt I should write some brief thoughts. The tl;dr is I completely agree with the sentiment but think no-one aside from Pluralis is earnestly moving towards a realistic solution. The reality is without some mechanism to pool resources for the creation of the models - with hardware that can't be taken away - the situation remains exactly the same, with a faint hope that Nvidia pours enough money into various challengers that the open <> closed model gap stays narrower for longer. I see very few scenarios where an individual group can pull ahead to the frontier (and even if they do, they will be subject to the same dystopian regulation the labs are now speedrunning towards). It must be a community driven effort to achieve this and hence we require a substrate to pool compute, data, and expertise effectively. This looks like p2p/distributed/multiparty training using hardware that has already been purchased to do local AI, thats sitting idle half the time, is physically controlled by individuals and can't be taken away. Repurposing this hardware for creation of the models and not just consumption is absolutely critical. I don't view most of the other things in this list as meaningfully moving the needle towards a good outcome. RL post-trains as a service.. ok great you made it easier for enterprise clients to use open models. Inference providers serving openweights... ok you host models and make money I don't see what you've changed. Open harnesses... my concern was never that I'd be locked out of the harness layer. Local inference... I don't really care about qwen3.8 27b running at 200 tok/s I want deepseek v4 pro running locally, and I want a better model than deepseek v4 pro. I wanna own and impact the process not the output. We don't have this yet because its hugely technically challenging - it requires solving fundamental research problems around low-bandwidth training and a re-write of most layers of the training stack. This is what we've been doing with extreme urgency for the last year and what we will continue to do until its working. Once such a system exists however, we will look back and find it strange that private companies ever felt they would be able to monopolize and control something this fundamental.
If you are reading this, the only possible resistance is to build decentralized, permissionless, and sovereign AI. Open source, self-hosting tools, RL post-training, local inference, open harnesses, we must build it all. We must fight overcentralization with an open stack.
10
23
92
28,671
This exceeded all expectations and continued to train with 40+ nodes in a single stage, 370 across the swarm and GPUs in multiple continents. TPS reached over 250k tok/s. Compute wise, this local hardware swarm is equivalent to 40 colocated B200's and would run $150k per month, potentially more depending on the provider. Thankyou to omk, mehmet and everyone else who helped us track down bottlenecks mid-run.
A 13B stress-test has just finished in Agora and instead of winding down we're going to ramp the contributor count until it blows up. If you have access to good local ML hardware, point your agent at agora-testing.pluralis.ai and join the pre-train for a few hours. Currently at 165 consumer GPU's.
1
9
51
4,725
A 13B stress-test has just finished in Agora and instead of winding down we're going to ramp the contributor count until it blows up. If you have access to good local ML hardware, point your agent at agora-testing.pluralis.ai and join the pre-train for a few hours. Currently at 165 consumer GPU's.
3
16
75
10,885
Our founder @AlexanderLong at @StanfordSBA with @jbrukh of @coinfund, discussing whether decentralized training can match datacenter training speed and how we can ensure the future of models remains open and accessible to everyone.
Replying to @AlexanderLong
@AlexanderLong of @Pluralis in conversation with @jbrukh of @coinfund on whether decentralized training can actually perform 👇 piped.video/watch?v=WRpf1hY9…
1
5
25
4,811
60 page tech report detailing Agora and the run it just powered now on Arxiv. Spans many different fields, from communication-efficient model parallelism, asynchronous optimization, to fault-tolerant systems design.
5
15
66
6,497
RL post-training on Macs 14 Macs across 4 countries generate every rollout for the run. Everything's running over the internet. No wire between any of them. As far as we can tell, this is the first RL post-training run with its whole rollout fleet on consumer Macs.
12
24
171
35,503
Two challenges remain. Each Mac holds the whole model, so we can only train models that fit on one Mac, and the trainer sits in a single cluster. Agora gets past both. It just pretrained Pluralis-8B across hundreds of consumer GPUs, the first pipeline-parallel training run over the open internet as far as we can tell.
1
4
3,386
Put Stoa (our system) and Agora together, and the whole loop could run in the open: large-model inference across the Macs, large-model training on Agora. The compute is already out there, idle, more than the clusters behind today's frontier models combined. As the best models are drifting behind closed APIs, training them on hardware people already own, owned by the people who train them, is how we get them back. Blog: pluralis.ai/blog/rl-post-tra… Code: github.com/PluralisResearch/…
3
2
14
1,136
Pluralis Research retweeted
On the back of our ICML workshop on Protocol Learning, our official workshop proposal has been accepted at NeurIPS, titled "Open, Collaborative, and Decentralized Training of Foundation Models". There is an exceptional set of speakers, and the call for papers is coming soon! In collaboration with @benjamintherien, @Ana_koloskova, @ebelilov, @niclane, #kaja, #Aakanksha, @Pluralis
2
6
32
5,040
The cheat sheet for our Protocol Learning workshop, brought to you by Pluralis Research Scientist @hmdolatabadi Videos coming soon!
We had an amazing time yesterday at the Protocol Learning Workshop at #ICML2026 in Seoul! It was a full day of learning from researchers working across distributed optimisation, decentralised training, and large-scale LLM training. Here’s a brief recap 🧵 1/n
2
15
3,746
Second Protocol Learning Workshop at ICML is a wrap. A packed day of talks and posters covering large scale distributed training and open source AI. A movement is growing.
1
5
50
3,695
Pluralis Research retweeted
Poster session was engaging and sparked a lot of discussions. Thanks for all the speakers and poster presenters! @Pluralis
2
14
1,096