A small step for mankind, a massive leap for decentralised training... for agency. In the space of 9 months, @tplr_ai went from 1.2B -> 72B. It's never been easy, and has broken everyone on the team multiple times. But I speak for all of us when I say it is the most rewarding thing we have ever done. We have a fraction of the resources. We don't have the PhDs. But Bittensor shows you it doesn't matter. Innovation happens at the edge. We innovate through scarcity. The ones who rewrite the rules are never the ones with the most. They're the ones who refuse to accept the limits they were handed. Bittensor is prophecy. Subnets (@covenant_ai and others) are the tools through which that prophecy is manifested. Next stop: TRILLIONS.
We just completed the largest decentralised LLM pre-training run in history: Covenant-72B. Permissionless, on Bittensor subnet 3. 72B parameters. ~1.1T tokens. Commodity internet. No centralized cluster. No whitelist. Anyone with GPUs could join or leave freely. 1/n
42
33
266
74,964
Sam Dare retweeted
The heads of the biggest AI labs want to agree among themselves on when everyone should slow down. @TheEconomist's piece on that push ends with Covenant-72B, the model we finished training in March on GPUs contributed over the internet, as a reason such agreements may be hard to enforce. What the piece doesn't say is that once training no longer requires one giant, tightly connected cluster, the power to build new models no longer has to sit with a few companies. Globally distributed training and open models are essential tools for keeping that power from concentrating in the frontier labs. They let many more participants, from nations to companies, own a share of the future of AI and take part in building it. economist.com/international/…
2
4
7
658
Sam Dare retweeted
1
3
8
898
Sam Dare retweeted
Good morning! ☀️
9
6
291
13,107
Sam Dare retweeted
This is an elitist statement. That there’s no reason to be this fast, and therefore we have to slow down! It’s truly remarkable the level of intelligence I have available on my phone with a $200/month pro plan. However, it’s way too expensive for this to be widely available to billions of people around the world. Tao is worried about the nonlinear nature of progress, except that technological breakthroughs, almost by definition are nonlinear. Otherwise, we would have run out of resources in the world long time ago. People forget how recently we figured out how to produce food at today’s scale. In the 1800s, an important source of fertilizer was an island off the coast of South America. Birds had spent centuries dropping their shit there, building up deposits that people mined for fertilizer. Then came a remarkable invention by the German scientist Haber, which Bosch helped scale industrially. By combining nitrogen from the air with hydrogen, using energy, we could manufacture ammonia for fertilizer. Fertilizer from thin air. There’s another lesson in there about the nature of technology. The same technology that helped create an abundance of food also enabled the production of explosives at scale.
Mathematician Terence Tao: "we have to slow down AI. the pace is insane, and there's no reason to be this fast — no reason at all" It's amazing how willing we are to change everything without any idea what happens afterward These are extremely nonlinear dynamics
13
11
167
17,193
Sam Dare retweeted
i predict math ai progress will organically slow down millennium prize problems are non-renewable resources for pr, soon to be depleted then the big labs won't be incentivized to spend major effort on math ai similarly, deepmind moved on from go ai (alpha go/alpha zero) not long after beating lee sedol (world champ). they could have kept training stronger and stronger go ai, but who would appreciate it? and i think being even better at abstract math won't transfer to other ai priorities (rsi, knowledge work, robotics, etc.) regardless, the work of pro mathematicians will shift from solving problems to understanding ai solutions
Mathematician Terence Tao: "we have to slow down AI. the pace is insane, and there's no reason to be this fast — no reason at all" It's amazing how willing we are to change everything without any idea what happens afterward These are extremely nonlinear dynamics
46
6
154
48,785
Sam Dare retweeted
Can autoresearch agents find better training data, not just better training code? Here’s our new EMNLP paper: AutoData: Agentic Search for Pre-training Data Selection [1/6]
7
13
119
18,743
Sam Dare retweeted
Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied! We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
30
99
843
148,468
Pre-training over the internet with unreliable peers requires reimagining fault tolerance from first principles. This opens the door to greater economic gains by way of been able to admit spot instances in our runs. The internet is the data center
We’ve been researching fault tolerance in Crucible, Templar’s pre-training platform. The goal: keep training through node failures and make better use of unreliable workers and spot instances, without idling an entire model replica when one stage goes down. 1/n
4
17
2,107
Sam Dare retweeted
Lean Verified Transformers (srush.github.io/lean-transfo…) In which we prove a bunch of Transformer invariants from scratch in Lean, and speculate about how hard it would be to do that for the rest of the world's code.
27
126
1,038
65,617
- Pipeline Parallel Training over broadband - 30B “The gradients must flow” tplr.ai/dashboard @tplr_ai
1
15
850
Skip it , dont fail
1
12
1,758
We are cooking some pretty interesting products around post training and inference @tplr_ai . Would love to talk to engineers and researchers to understand where you bottlenecks are . The rewards are my eternal gratitude and some credits. DM if you're up for it.
7
2
18
1,832
Sam Dare retweeted
3
9
1,155
Sam Dare retweeted
2
2
7
1,249
Sam Dare retweeted
The cybersecurity guards in frontier models perfectly illustrate the weird incentives in software engineering. They equally block vulnerability construction _and_ attempts to improve the correctness of software because they are both the same task.
3
1
16
1,322
To our knowledge this is the first economic validation of decentralised pretraining . But what’s the end goal? The future is heterogenous compute . @tplr_ai will connect stranded computed into a virtual fabric . The internet is the datacenter .
Crucible, Templar's pre-training platform, has completed its first production end-to-end training runs. The latest trained an 8B model on 50.53B tokens across 48 distributed A100s, at an estimated $0.1202 per million tokens of GPU rental. The run reached 48.3% effective MFU. At AWS p4de Capacity Blocks pricing, a 48-A100 cluster operating at the literature-derived 65% compute ceiling comes to an estimated $0.1686 per million tokens. Crucible's measured $0.1202 was about 29% lower after its low-bandwidth overhead. The comparison excludes R2 storage and operations. The full writeup shows the method and accounting. Blog: tplr.ai/publications/blog/in…
4
16
2,838
Sam Dare retweeted
2
12
1,423
1/ Here’s a demo showing how I generated tokens for free on @OpenRouter using the @FireworksAI_HQ endpoint for GLM-5.3
2
6
20
6,300
Sam Dare retweeted
these guys built an infinite movie generation machine... @fal's post-trained Minimax H3 Max is 50x faster than the original: 5 seconds of video out of 3 seconds of compute it generates faster than you can watch it and i found the open-sourced version: huggingface.co/FastVideo/Fas…
Minimax H3 Max has generates video faster than you can watch it so I hooked it to a twitch livestream! Now you can watch infinite interdimensional cable - link to the stream below
12
4
39
9,376