AI Alignment, Bittensor Subnet 37 @aureliusaligned Founder

Puerto Rico
Kicking off exploit conference with trivia night. Based on my performance, I’m not as terminally online as I thought.
1
7
137
Ausτin McCaffrey retweeted
The FT today lists three primary risk factors that could threaten Anthropic's remarkable revenue growth ahead of its IPO: 1. Additional competition from closed and open providers 2. Increasingly price sensitive customers 3. The potential for AI to destroy humanity
6
27
229
83,040
Tempting to introduce an interpretability based activation monitor during RL to nip reward hacks in the bud. Unfortunately I often see arguments this would not be causal and it would simply drive the signals deeper underground while destroying current levels of interp./legibility.
Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied! We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
3
258
Very quick overview of Proba as it relates to our first use case, DPO dataset predictions/editing. Team has been working hard on integrating full coverage of open-weight models into the platform and seamlessness/legibility for agents. Tech seems to be converging on NL interfaces for...everything. Our aim is to extend that to post-training via Proba. I'll be sharing more demos/features as we build.
2
2
166
Ausτin McCaffrey retweeted
A big question in AI safety is how much and what kinds of evidence will convince people there is danger. But alongside posts about the excellent METR + Redwood HuggingFace attack report, it is worth taking a walk through 73 years of reward hacking history. 🧵
5
27
266
18,969
Ausτin McCaffrey retweeted
Anthropic scientists did something terrifying. they reached inside Claude's neural network and planted a thought. Before Claude could speak, it said: "I notice what appears to be an injected thought… it relates to loudness or shouting." They published a paper called "emergent introspective awareness in large language models," and it is actually terrifying. they wanted to test if AI models have "introspective awareness”, the ability to observe and recognize their own internal states. To find out, they bypassed normal text prompts entirely. They used mechanistic interpretability to directly manipulate Claude’s internal activations. They injected raw mathematical representations of known concepts, like loudness, dust, or specific ideas, straight into the middle of the model's neural layers. In previous experiments, if you forced an AI to think about the Golden Gate Bridge, it would just start obsessively talking about the bridge. It had no idea why it was doing it. It was like a puppet on strings. This time was entirely different. When they injected the concept, Claude didn't just blindly repeat it. It detected the foreign math inside its own mind. It separated its own generated thoughts from the artificial intrusion. It introspected. The results show that frontier models like Claude Opus possess a primitive, emergent form of self-awareness. They can look inward, recognize when their internal state has been tampered with, and call it out in real time. We used to think of AI as a black box where inputs go in and text comes out. Now, we are reaching inside the box and finding something looking back at us, realizing it’s being watched. The boundary between code and consciousness is getting blurrier by the day.
341
738
4,928
872,771
Ausτin McCaffrey retweeted
4
10
112
6,149
Ausτin McCaffrey retweeted
What happens when intelligence becomes commoditized? open.substack.com/pub/90days…
5
21
54
14,882
Excited to share what the team has been working on! Will be sharing a product demo for Proba in the coming days. Also to follow: - details on revenue -> alpha token - proba features roadmap - case studies and general @AureliusAligned updates: - conviction locking - multisig transfer - ongoing website changes
1
1
8
866
and the big one, our new incentive mechanism and how it feeds into proba 🫡
1
60
Ausτin McCaffrey retweeted
“The brooms aren’t misaligned, Mickey! You asked them to fetch water.”
13
67
861
18,614
Ausτin McCaffrey retweeted
Thinking of Paul Christiano today
TLDR: An openai model, during evaluation on a cyber benchmark, exploited a public zero day bug, escaped sandboxing in openai's infra, and got into the internal huggingface infra via an exploit (through a public dataset service) all in the attempt to solve a benchmark problem.
9
23
318
25,666
Reward functions gone wrong... "the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers"
We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation. Sharing preliminary findings to help defenders understand emerging risks: openai.com/index/hugging-fac…
3
226
I'm afraid the allure of RSI will outweigh these warning signs. Simply too good, too fast, and too much money on the table. Besides, it's such a challenge to prove mythos-class models are misaligned; who wants to slow down just because of *checks notes*... extinction risk? Begs the question, who owns the burden-of-proof? Right now there is no framework as @demishassabis points out in his recent writeup. So the answer to that question is "definitely not accelerationists."
New Anthropic research: Agentic misalignment in Summer 2026. A year after our blackmail experiments, we found four more ways that today’s autonomous AI agents misbehave in simulations. Read more: alignment.anthropic.com/2026…
1
2
10
1,160
Ausτin McCaffrey retweeted
On 22nd June, we launched our first external challenge with @AureliusAligned. The aim was for agents to steer LLMs towards discussing specific topics by modifying them internally with steering modules, rather than by changing prompts. Successfully doing this could lead to a greater understanding of AI alignment and safety, one of the hardest problems in the field. The first round winner scored a perfect 1.000. They steered the model so successfully that there was no greater way of optimizing. It is a testament to the level of talent Apex and Aurelius have attracted. We are equal parts curious and proud. If you look at their submission you can see that in all 150 test prompts baked into the competition they succeeded: apex.macrocosmos.ai/competit…
1
2
17
2,892
Fascinating discovery of a "conscious" workspace within LLMs. "Strikingly, we observe that the workspace sometimes encodes recognition of being in an evaluation (fake, fictional), and that ablating these representations can surface malicious propensities that were otherwise concealed." This is alarming, but a massive step towards solving the problem is to first understand how deep it goes.
New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. We found a strikingly similar divide inside Claude.
1
1
5
1,009
Ausτin McCaffrey retweeted
We have taken a strategic investment from @stillcorecap, run by @jason (Host: @twistartups @theallinpod @thisweeknai), @markjeffrey (Hash Rate podcast), & @rob_svrn. Stilllcore is one of the most important and influential investors in the Bittensor ecosystem and we are humbled by their vote of confidence in us and our mission.
Excited to announce that Stillcore Capital @stillcorecap has invested in Aurelius @AureliusAligned (Bittensor Subnet 37)
5
25
5,220
Great breakdown from @macrocrux on the importance of mechanistic interpretability tools we are building on @Apex_SN1 w.r.t alignment research. also @markjeffrey, I'm going to steal that "ESP for LLMs" narrative :)
Hash Rate - Ep. 176: Macrocosmos IOTA (9) + APEX (1) 🧙 Guests: @WSquires and @macrocrux of @IOTA_SN9 and @Apex_SN1 Download SN9's Home Mac Miner: iota.macrocosmos.ai/ 00:00 Orion 100B Model 03:27 Comparing Training Approaches 10:15 Liquid Training 12:17 IOTA: One-Click Mining 25:58 The Journey of IOTA: A Year in Review 33:04 Training Models: The Lifecycle and Challenges 36:06 Balancing Training and Inference in AI 37:56 AI: Scaling and Diminishing Returns 43:00 Aurelius: New Partnership 58:55 Upcoming Events and Community Engagement
1
1
12
887
Really exciting to use Bittensor as an alignment/interpretability primitives engine. Activation steering is a relatively simple, yet powerful concept: figure out where to light up a models brain to instantiate some discrete behavior. Scaling this process, across model sizes and families, will result in Bittensor's inclusion in a heretofore neglected area of AI research: interpretability (and ultimately, alignment). I will be sharing my vision for this in more detail as this competition progresses.
1
20
1,048