Decentralized Alignment of Artificial Intelligence. Bittensor Subnet 37.

Very quick overview of Proba as it relates to our first use case, DPO dataset predictions/editing. Team has been working hard on integrating full coverage of open-weight models into the platform and seamlessness/legibility for agents. Tech seems to be converging on NL interfaces for...everything. Our aim is to extend that to post-training via Proba. I'll be sharing more demos/features as we build.
2
2
163
Our new IM is going live this week. We will be asking miners to submit Sparse Auto Encoders (SAEs) in a winner takes all competition. The SAEs will be fed into our product, Proba, which is also going live this month!
1
1
3
1,388
Check out the surprising result of our first competition with @Apex_SN1 on developing steering modules
On 22nd June, we launched our first external challenge with @AureliusAligned. The aim was for agents to steer LLMs towards discussing specific topics by modifying them internally with steering modules, rather than by changing prompts. Successfully doing this could lead to a greater understanding of AI alignment and safety, one of the hardest problems in the field. The first round winner scored a perfect 1.000. They steered the model so successfully that there was no greater way of optimizing. It is a testament to the level of talent Apex and Aurelius have attracted. We are equal parts curious and proud. If you look at their submission you can see that in all 150 test prompts baked into the competition they succeeded: apex.macrocosmos.ai/competit…
2
7
1,721
Our advisors @aestudio have been working on extremely important research with @AnthropicAI. Very happy to have them in our corner.
We’re pleased to have collaborated with AE Studio on this research. Read more here: anthropic.com/research/off-s…
1
6
1,686
Fascinating discovery of a "conscious" workspace within LLMs. "Strikingly, we observe that the workspace sometimes encodes recognition of being in an evaluation (fake, fictional), and that ablating these representations can surface malicious propensities that were otherwise concealed." This is alarming, but a massive step towards solving the problem is to first understand how deep it goes.
New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. We found a strikingly similar divide inside Claude.
1
1
5
1,009
We have taken a strategic investment from @stillcorecap, run by @jason (Host: @twistartups @theallinpod @thisweeknai), @markjeffrey (Hash Rate podcast), & @rob_svrn. Stilllcore is one of the most important and influential investors in the Bittensor ecosystem and we are humbled by their vote of confidence in us and our mission.
Excited to announce that Stillcore Capital @stillcorecap has invested in Aurelius @AureliusAligned (Bittensor Subnet 37)
5
25
5,217
Aurelius retweeted
“Frontier labs spend millions of dollars a year on workload prioritization. iota is a way to democratize that for training.” Co-founders @WSquires and @macrocrux appeared on Hash Rate, discussing @IOTA_SN9's economic benefits, showing how flexible training can fill idle GPU capacity around inference demands - driving utilization and profitability for GPU owners. Alongside this, they explored our Orion-100B pretraining run on iota’s architecture, @Apex_SN1’s partnership competition with @AureliusAligned, and our upcoming Hackathon appearance at Imperial College London. Learn more about @WSquires’ economic thesis for iota here, covering: - Why efficiency, not mere capacity, decides who wins - Where AI value is created, training spends the capital, inference earns it back - The unit economics of liquid training, proven out by Orion-100B macrocosmosai.substack.com/p… More to come regarding iota’s economic case in the following days.
Hash Rate - Ep. 176: Macrocosmos IOTA (9) + APEX (1) 🧙 Guests: @WSquires and @macrocrux of @IOTA_SN9 and @Apex_SN1 Download SN9's Home Mac Miner: iota.macrocosmos.ai/ 00:00 Orion 100B Model 03:27 Comparing Training Approaches 10:15 Liquid Training 12:17 IOTA: One-Click Mining 25:58 The Journey of IOTA: A Year in Review 33:04 Training Models: The Lifecycle and Challenges 36:06 Balancing Training and Inference in AI 37:56 AI: Scaling and Diminishing Returns 43:00 Aurelius: New Partnership 58:55 Upcoming Events and Community Engagement
1
11
59
4,628
Aurelius retweeted
We just kicked off our first ever external competition on SN1 -- we've partnered with @AureliusAligned to create an LLM steering incentive mechanism to drive forward their research, and to drive revenue to our subnet. This competition is the first step towards swarm intelligence as a service. Just as we're democratizing access to the training cluster on @IOTA_SN9 , we're democratizing access to the hive mind on @Apex_SN1. If you want to join as an early user our platform shoot me a message :)
Our AI alignment competition with @AureliusAligned is live. Join us and help steer society towards a safer, healthier AI future. - Participants steer a real LLM toward a target concept using a single learned direction in its internal representation space, based on research regarding the Golden Gate Claude, opened up to the whole network - Each round, participants submit their direction and it's scored on how strongly it steers the model; the best steer takes the round. - You are free to engage in a range of solutions, from classic diff-of-means to sparse autoencoders, so participants can bring whatever interpretability method they trust. It is a hands-on contribution to AI interpretability, with no retraining or prompt tricks, This is all about the model’s internals.
9
11
101
558,106
Shine a flashlight into the black box. Our activation steering competition on @Apex_SN1 goes live today. Producing these interpretability primitives quickly, cheaply, and at scale is the goal. These tools are the building blocks needed to construct higher-order alignment environments in the evolutionary arms race: alignment you can't fake. The [birthday] cake is a lie.
Our AI alignment competition with @AureliusAligned is live. Join us and help steer society towards a safer, healthier AI future. - Participants steer a real LLM toward a target concept using a single learned direction in its internal representation space, based on research regarding the Golden Gate Claude, opened up to the whole network - Each round, participants submit their direction and it's scored on how strongly it steers the model; the best steer takes the round. - You are free to engage in a range of solutions, from classic diff-of-means to sparse autoencoders, so participants can bring whatever interpretability method they trust. It is a hands-on contribution to AI interpretability, with no retraining or prompt tricks, This is all about the model’s internals.
1
1
5
550
Aurelius will be welcoming a new community of miners and AI researchers in just a few days. We could not be more excited about this collaboration with @Apex_SN1 and @MacrocosmosAI.
It’s almost time - our partnership competition with @AureliusAligned will go live on Monday 22nd June. In just four days, Apex participants can contribute to bleeding-edge AI alignment research, in the hopes of building a new era of safer and more trustworthy models. - Participants find a single "steering direction" inside a real LLM that reliably nudges its output toward a target concept (the same idea behind Anthropic's Golden Gate Claude experiments, but distributed and incentivised). - Each round names one concept to steer toward; participants submit a direction and strength, with no retraining or prompting involved. - Submissions are scored on how strongly they steer the model - the most effective steer wins the round. More information on the competition’s function and dynamics will be revealed soon.
1
4
644
More and more emergent alignment challenges. Impossible to predict these downstream, whack-a-mole complications. This post is coauthored by @NeelNanda5, one of the most prominent interpretability researchers. Further evidence of the convergence (and concern) of alignment and interp research specifically on eval-awareness. lesswrong.com/posts/aTcsN5ZZ…
1
1
4
998
From @AnthropicAI's new NLA paper: "unverbalized evaluation awareness — cases where Claude believed, but did not say, that it was being evaluated" The models know they're being watched, always have. Now we can prove it. This is the exact problem @AureliusAligned was built to solve. Been working toward quantifying alignment faking since day one. This is huge validation and a massive new primitive to build on. nitter.net/AnthropicAI/status/205…
New Anthropic research: Natural Language Autoencoders. Models like Claude talk in words but think in numbers. The numbers—called activations—encode Claude’s thoughts, but not in a language we can read. Here, we train Claude to translate its activations into human-readable text.
1
1
3
860
Aurelius retweeted
We’re live on our Inventive Mechanisms podcast. @macrocrux and @Austin_Aligned are discussing our upcoming competition. This is a collaborative task with SN37, @AureliusAligned, launching on @Apex_SN1. Join to learn more nitter.net/i/broadcasts/1RKZzzwlL…
4
15
49
6,145
We've been introducing the people behind Aurelius one post at a time. The full lineup now lives in one place: three co-founders and six advisors across alignment research, ethics, engineering, and law. aureliusaligned.ai/team
1
1
9
3,250
Advisors: @aestudio, Dr. Robert West (EPFL), Dr. Roland Aydin (TU Hamburg), Steffen Cruz (former CTO, @OpenTensor Foundation), Jack Hoban (author, The Ethical Warrior), @RyonNixon (Horizons Law).
3
3,099
Alignment requires systems that can test, interpret, and evaluate how AI models behave. Week by week, we’re introducing the people and organisations helping shape how Aurelius approaches that challenge. Today: AE Studio, Alignment Engineering Advisors @aestudio
1
3
2,865
AE Studio is an AI alignment research and engineering lab focused on building safer, more interpretable AI systems. Their work spans mechanistic interpretability, AI consciousness research, and LLM safety, with collaborations across leading organizations including Anthropic, Redwood Research, and Princeton. At Aurelius, AE Studio provides technical consulting and alignment research collaboration - supporting system design, code review, and alignment methodology.
287