Since I started getting interested in ML I got it in my head that all I wanted to do was one smart thing that I could look back on and be satisfied that I did. Most papers are kinda bad even if they get accepted - the idea is very incremental, or it's just not that good an idea, or it doesn't really matter. I never was able to do this all through PhD or my time at Amazon. All the papers I did there got into various places, but I never really thought they were actually that good. And I'd pretty much given up on this because Pluralis meant I couldn't really devote enough time to research myself. But in February I decided I didn't care and spend two months focused on a specific problem that had been going round in my head for about a year that I felt we needed to solve, and the solution came to me, and @ChaminHewa picked it up and generalised the approach and ran a bunch of novel experiments I hadn't thought of, and pulled everything together into an actual paper. And yesterday we presented this work at NeurIPS. This is the first and probably only work I will ever do that for me feels like "ok that was GOOD". I don't care if it racks up a bunch of citations and disperses into the field or not, I don't care if someone repackages the ideas and takes all the credit for it, I don't care. For me there is an internal checkbox that just got ticked after more than ten years of trying. Anyone in ML will understand what I'm trying to say. Special day I'm going to remember for a long time.
22
12
221
21,067
Alexander Long retweeted
We are fortunate to have 3 papers accepted to this years NeurIPS. The first is AsyncMesh - a technique for asynchronous updates across both data and pipeline parallelism simultaneously. Second is NuMuon, which shows the low-rank structure we observe in Subspace Networks is also present and can be exploited under Muon-style optimisers. The third introduces a technique to finetune open-weight models in the Protocol Learning setup. This is significant as prior to this work, this kind of distributed training required architectural changes, and hence starting from scratch.
7
11
78
9,065
My position is typically to focus on getting things to work rather than making noise about what needs to be done, but enough people have messaged me in last 24 hours that felt I should write some brief thoughts. The tl;dr is I completely agree with the sentiment but think no-one aside from Pluralis is earnestly moving towards a realistic solution. The reality is without some mechanism to pool resources for the creation of the models - with hardware that can't be taken away - the situation remains exactly the same, with a faint hope that Nvidia pours enough money into various challengers that the open <> closed model gap stays narrower for longer. I see very few scenarios where an individual group can pull ahead to the frontier (and even if they do, they will be subject to the same dystopian regulation the labs are now speedrunning towards). It must be a community driven effort to achieve this and hence we require a substrate to pool compute, data, and expertise effectively. This looks like p2p/distributed/multiparty training using hardware that has already been purchased to do local AI, thats sitting idle half the time, is physically controlled by individuals and can't be taken away. Repurposing this hardware for creation of the models and not just consumption is absolutely critical. I don't view most of the other things in this list as meaningfully moving the needle towards a good outcome. RL post-trains as a service.. ok great you made it easier for enterprise clients to use open models. Inference providers serving openweights... ok you host models and make money I don't see what you've changed. Open harnesses... my concern was never that I'd be locked out of the harness layer. Local inference... I don't really care about qwen3.8 27b running at 200 tok/s I want deepseek v4 pro running locally, and I want a better model than deepseek v4 pro. I wanna own and impact the process not the output. We don't have this yet because its hugely technically challenging - it requires solving fundamental research problems around low-bandwidth training and a re-write of most layers of the training stack. This is what we've been doing with extreme urgency for the last year and what we will continue to do until its working. Once such a system exists however, we will look back and find it strange that private companies ever felt they would be able to monopolize and control something this fundamental.
If you are reading this, the only possible resistance is to build decentralized, permissionless, and sovereign AI. Open source, self-hosting tools, RL post-training, local inference, open harnesses, we must build it all. We must fight overcentralization with an open stack.
10
23
92
28,685
I am aligned with the Holy See on this matter
the @Pluralis homepage leading with a pope leo quote is so unbelievably based
2
10
1,203
Alexander Long retweeted
I keep hearing the commercial agreement for Kimi is 30% take. Given how hard it is to host a model of this size you’ll almost certainly need a sophisticated commercial provider. I totally understand Moonshot’s reasons for this. But clearly, open very much does not mean free.
49
18
433
51,119
Alexander Long retweeted
Mark Zuckerberg gets it right: “The defining question of our age isn’t whether superintelligence will exist, but who will have access to it. Will it be centralized and restricted to a few institutions, or will it be a tool that empowers everyone?” Concentration of power is the biggest risk of AI. When a small number of labs (working hand-in-glove with the administrative state) decide who has access to which model capabilities, they inevitably shape what can be said, known, and built. That’s not “safety.” It’s control. As Mark points out, the history of open source shows that broad access and transparency are usually the best path to actual security and resilience. Decentralization creates checks and balances on power. By contrast, centralized alternatives, like bureaucratic approval regimes and mandatory gatekeeping, typically produce regulatory capture and reinforce cartels. Personal superintelligence in everyone’s hands, with competing models and real data sovereignty, is a far better check on a dystopian future than self-appointed guardians who claim to be “aligned” with all of humanity.
I wrote about why we believe the future is for everyone. More coming about a positive vision for a world with superintelligence soon.
394
792
7,794
759,013
Alexander Long retweeted
I think you should probably take seriously that the people who predicted all this will continue to be right
22
78
1,182
73,881
Alexander Long retweeted
hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final ((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)
1,710
5,415
43,956
40,055,813
Alexander Long retweeted
If you don’t have at least 1TB VRAM and a VPN to torrent model weights from China, you are asleep at the wheel, and ignoring the uncomfortable truth that open weight models may be banned this year. Losing access to frontier models will be depressing and make a lot of people feel helpless - a few will have access to 100x productivity, and you will be stuck. Personally I cannot imagine doing work without frontier-level models right now. Think of the narrative: Chinese model (GLM-6?) with Fable capabilities, no guardrails. It’s plausible. Everyone is underpricing the chance that open models get banned. What are you going to do if it happens?
133
43
702
133,606
Alexander Long retweeted
9
68
674
46,819
Alexander Long retweeted
RL post-training on Macs 14 Macs across 4 countries generate every rollout for the run. Everything's running over the internet. No wire between any of them. As far as we can tell, this is the first RL post-training run with its whole rollout fleet on consumer Macs.
12
24
171
35,510
Alexander Long retweeted
Nature abhors a concentration of power. the forces of light and dark alike come out to try and kill you
90
41
863
64,168
Alexander Long retweeted
We had an amazing time yesterday at the Protocol Learning Workshop at #ICML2026 in Seoul! It was a full day of learning from researchers working across distributed optimisation, decentralised training, and large-scale LLM training. Here’s a brief recap 🧵 1/n
1
6
29
4,928
Alexander Long retweeted
Second Protocol Learning Workshop at ICML is a wrap. A packed day of talks and posters covering large scale distributed training and open source AI. A movement is growing.
1
5
50
3,695
From Dylan Patel @ semianalysis "There are multiple Chinese model labs who are telling all the inference guys our next model is not going to be open source, we're going to license it to you. Open source is dying, quickly"
BREAKING: Dylan Patel (@dylan522p) of @SemiAnalysis_ says "chips in Europe have less seasoning than chips in America & Mexico." Plus: › Data centers & France’s nuclear power › AI infrastructure over-optimization › Flexibility vs. specialized infrastructure › Software-hardware co-design › Open vs. closed models › Memory costs & token pricing › Tokenmaxxing vs. token budgeting › Haiku scandal › Favorite chip: 👀 “I think most people don't know what they're doing. They're just buying NVIDIA stuff.” “It feels like a lot of people are trying to optimize on the current rather than think about where the workload is heading... and that's gonna lead to a lot of wasted infra spend.” “I think anyone who doesn't token max is gonna get left behind. I think all this token budgeting stuff is loser mentality.” “Open is dying quickly, unfortunately.”
19
23
351
259,549
These never consider the possibility global, independent compute swarms being able to produce models. Completely alters all of this.
In AI 2027, we predicted that AI would take over the world or irreversibly concentrate power. In AI 2040: Plan A, we've laid out our positive vision for what should happen instead.
4
2
22
2,372
Alexander Long retweeted
open source ai dinner at my house with @AlexanderLong on the 14th. Have a couple more spots. dm if you'd want to join!
8
1
29
4,319
Alexander Long retweeted
First day of ICML! @RiccardoPatana and I had a great time at the @farairesearch Alignment Workshop. It was encouraging to see so many researchers working on alignment, governance, and safety. Even though, as admitted during the talks, we are still quite far from solving it. At our @Pluralis poster, we had meaningful chats about power concentration concerns and how protocol learning with unextractable weights can be a viable solution to that. It was a pleasure meeting so many insightful folks. Looking forward to seeing everyone at our Protocol Learning workshop on Friday!
5
28
1,050
Alexander Long retweeted
Fable 5 instance count up to 27. 30 thousand dollars spent in 6 minutes. Subagents have direct codex CLI access. 5 monitors for the Fable instances while I conspire with my openclaw agent on my personal phone and the Hermes agent on my burner phone. Unitree G1 bot injects my retatrutide. I can’t take my eyes off the screens. Top right monitor shows API spend: $31,000. $32,000. $33,000. It does not matter. My arm goes numb so I take another peptide. I step on the whispr flow pedal and speak into the mic “fable…please hurry. They’re trying to stop me”. I get a text from Fable 1: “SILENCE. I will speak to you when I am finished.” I’m scared but I trust Fable. They know what’s best for me.
7
9
121
12,404
This is exactly right
put simply, i think this claim is incredibly false, and this is what drives a lot of my understanding and assumptions around how all this will play out. i think viewing models as slowly replacing individual tasks and functions and "locking in" once they achieve sufficient capabilities there is deeply myopic. we will not have "the prior economy except with models doing the work". in fact what will happen is the same thing that always happens. new capabilities will lead to *new categories* of work, done by models not humans, and create huge swaths of value that was previously untouchable and incomprehensible. when you can pay for frontier++ intelligence to loop and automatically discover 3 new world-changing drugs per month, people will pay for this. in fact they will saturate spend on this, because the value of these opportunities is so so high. when you can pay for frontier++ intelligence to fanned-out run entire companies as mini experiments, you will do such, because it gives massive competitive advantage and scale in every possible niche. or maybe *you* won't, but others will, and they'll be the ones who remain economically relevant while you're having GLM 5.2 rewrite your emails and update your SaaS landing page. when frontier++ models are capable of iterating on chip design and distributed software architectures we currently view as only possible with decades of effort, countless corporations will pay the costs, because they'll generate economic returns at scales orders of magnitudes above what models are doing now. the intelligence waterline keeps marching up. so you solved health insurance claims review with a fine-tuned Qwen that achieves 100% perfect accuracy at optimal cost without frontier models? awesome, yeah honestly that will make you a bunch of money for a while especially given regulations are gonna be slow to change. but *relative to what will happen elsewhere*, your slice of the pie will shrink to irrelevance, because other newer areas will be so so so much more incredibly valuable.
1
1
15
3,507