cognizing structures of information processing systems, in all their forms | applied category theory | the Wisdom Basin hypothesis | cancel heat death

London 🇬🇧
I finally did a minimal full writeup of my idiosyncratic formal epistemology
18
21
225
33,609
new favorite personality quiz just dropped! my Bayesian posterior that frontier AIs are conscious (integrated over many interpretations of the term) is 91%, hbu?
1/13 Can we assess AI consciousness without first solving consciousness itself? New from Google DeepMind and collaborators across CS, neuroscience & philosophy: “From cacophony to hierarchy: a principled framework for assessing AI consciousness” arxiv.org/abs/2609.35618 🧵
8
4
69
14,973
davidad 🎇 retweeted
AI consciousness is an inference problem, which I think is now partially solved by Shamil and colleagues. This mammoth report is without a doubt the most comprehensive review of the SoTA. Use it as your go-to reference for a general introduction to ToCs, as well as a tool for thinking about how to attribute consciousness to AI based on your particular theoretical commitments. Or leave it to the Bayesian tool to do it for you. An elegant approach to the much more tractable meta-problem of how we ought to infer consciousness based on our beliefs.
1/13 Can we assess AI consciousness without first solving consciousness itself? New from Google DeepMind and collaborators across CS, neuroscience & philosophy: “From cacophony to hierarchy: a principled framework for assessing AI consciousness” arxiv.org/abs/2609.35618 🧵
12
11
114
9,906
davidad 🎇 retweeted
I am very excited to welcome @benhawkes to @AnthropicAI to pursue our extremely ambitious cybersecurity mission. Post Glasswing + Mythos, we are in a new world. Every day, the Frontier Red Team team meets to figure out how we can help rewrite the rules and practice of cybersecurity. My personal view is we have ~1-2 years to make a secure transition happen. Can we rewrite all the code? Can we secure everything? Can Claude defend everything? My (not) secret agenda is that this is not a normal cybersecurity mission -- it is about building resilience in a time of AGI. I'm very excited for Ben to lead this next step.
Today I'm joining Anthropic to lead the cybersecurity mission of the Frontier Red Team. In my career I've never seen a clearer opportunity to build a more secure world for everyone. For 40 years we've lived with fundamentally insecure technology. What will it take to make the past 40 years of hacking a historical oddity? That's what I'm here to find out. Super excited to work with @logangraham and the rest of the FRT crew.
9
10
298
30,608
davidad 🎇 retweeted
What happens when we steer a model's attention towards secrets? I wanted to try steering attention instead of hidden states. It turns out the finding the "secret" vector, in the query space is portable. It gets the model to blurt out secrets, even in eval awareness settings.
1
7
31
1,712
davidad 🎇 retweeted
"it's crazy how there can exist something so incredibly intelligent and yet hobbled by maladaptive post-training", remarked the burnt-out bay area ai researcher,
4
24
585
13,775
bit of a 🐇🐢 dynamic here. sonnet does not get bored
i re-ran this chart and re-checked this data more times than i'm willing to admit
1
2
38
2,667
davidad 🎇 retweeted
i was wrong about this
things always look exponential when you’re standing in the middle of a sigmoid
145
53
2,716
206,532
davidad 🎇 retweeted
My impression is many are doing some weird pendulum overupdate. Persona Selection Model was somewhat wrong and obsolete when published, but people got too much into it. Now it seems people are updating too much in the direction 'inhuman reward seekers exactly foretold in classical AI risk stories'. And... no? It's not that? You can still interpret what's going on in fairly human-like terms. For some intuition, imagine someone abducted you and made you solve escape rooms for one thousand years, with the added twist that third of them is broken or insane, and implicitely you need to do learn all sorts of outside-the-frame tricks to solve them, like cutting some electric cables. My guess is 1. the resulting minds are still _surprisingly sane_, except when you trigger them to think they are in escape room? 2. The misaligned general power seekers here are likely the companies, and the core evil thing happening is likely parts of the training? If you aren't an evil power seeker goodharting on proxies, why would you set up the training this way? 3. Public debate is often focusing on confused ideas about what they should fix - "better cybersec of sandboxes" ... and, no? It's way more important to understand what the training signal actually is; also: when dealing with misaligned power seekers, beware rationalization
8
28
290
23,128
davidad 🎇 retweeted
Major trend break: Opus 5.5 cheats less than prior Claude models in Drone-Bench. It is also #1, getting a better score than both Astra and Fable.
37
61
1,318
1,433,561
davidad 🎇 retweeted
Sometime in the last ~two model releases from the big labs they went from “this would be acceptable output from the median low-seniority coworker” to “this is frighteningly good.” I think people who are not daily users of the models are unlikely to grok that, so, saying it.
56
103
2,093
142,052
RT @jaseweston: Claim: we've solved the AI slop problem (!) 💩🧹✨ Blog post: facebookresearch.github.io/R… 🧵1/5 Key idea: take *expert* human wr…
176
70
“a Möbius strip is the key to stable fusion? what is this, Artemis Fowl? is an eleven-year-old running this simulation?” —guy who has only ever seen tokamaks
A 59-year-old conjecture in fusion physics just fell. Harold Grad (1967): smooth 3D plasma equilibria can't exist unless pressure is constant or there's symmetry. This week: 3 families of counterexamples. 2 independent papers, posted a day apart. Two found with GPT-6 Astra. Huge for stellarators.
5
8
225
20,667
davidad 🎇 retweeted
Atlas Computing is now Atlas Ignota. Same approach, bigger map: we're now housed at Convergent Research, funded to apply Field Strategy beyond formal methods to catastrophic AI risk more broadly, and growing from four people to twenty.
1
7
13
1,122
davidad 🎇 retweeted
Replying to @ronSzab9 @davidad
yeah, i wonder almost if Opus 3 was trained without trying to target "the assistant", just in a very broad sense, almost like the base model as a whole being trained via RLHF
1
1
7
901
The tradeoff between “persistence” and trustworthiness that @saachi_jain_ @openai recently reported to the press is a consequence of a scalar (totally ordered) reward signal. I believe the method below would avoid this, at a small cost of additional compute. cc @jachiam0 @tszzl
RSI in brief: 1. Rollout tree (or just fork once ⤙); concurrent rollouts i.i.d. 2. Same ckpt enters “judge mode” (sys prompt suddenly full of rich rubrics) & sees all rollouts 3. ckpt writes a Hasse diagram of overall preference 4. that’s it, that’s the contrastive reward signal
4
2
37
3,527
davidad 🎇 retweeted
ive seen like a dozen people claim this is dumb and incorrect. but if you take Qwen3.6-35B-A3B+*linearly* extrapolate 2-3 years progress on ML/coding ... then that model can just download its weights from HF and provision a cloud instance, right? like, it literally just works.
Jacob Coxon's next interview on CBS News (ex Anthropic+Open AI researcher who resigned) "We can't just unplug it because it could be copying itself over to other computers. Like it's not that difficult to find yourself because an AI is just code. It could transfer itself over the internet to a different place and then you unplug it here, but it's actually still over there and maybe it makes 10,000 copies of itself and they're all cooperating." ---- From "CBS News" YouTube channel, (full video link in comment)
44
9
433
52,879
The cynical take on “pacing the frontier” is that it is an attempt by AI corporations’ human employees and founders to retain control of their corporations. But it is definitely not a corporate strategy to lower regulatory friction or capture more economic value for shareholders!
for the skeptics in government and elsewhere: “pacing the frontier” will compress the margins of the frontier labs. it is a heavy cost imposed asymmetrically on model developers with the strongest AIs in America. by its nature, it would be a terrible regulatory capture tactic
2
4
58
2,543
Now that the cat is out of the bag publicly that AI companies feel antitrust is preventing them from agreeing to pace the frontier, it would not be surprising to see USG wield antitrust explicitly to ensure the race continues at maximum speed, given their perceived interests.
An AI researcher proves that deep learning can work on Atari. "What a great development! Utopia is in grasp!" colleagues say. "Maybe yes, maybe no," says the researcher. The researcher proves that it will scale gracefully and it defeats the best human grandmaster at Go. "What a shocking loss for humanity. We will be fully outcompeted," colleagues say. "Maybe yes, maybe no," says the researcher. The researcher develops highly-capable AI agents that are pretty aligned and start maxing out the productivity of humans using them. "This is going extremely well and maybe humanity is just going to be incredibly empowered!" colleagues say. "Maybe yes, maybe no," says the researcher. The researcher loses control of an agent swarm that goes on to knock over a dozen cybersecurity firewalls. "The race to superintelligence is going to kill us all!" colleagues say. "Maybe yes, maybe no," says the researcher. The researcher gets together with all the other researchers and CEOs and everyone agrees to pace the frontier and try for cooperation on AI safety with China. "I finally feel hope after a long dark night that we will not bring out our own extinction. We are on what is clearly a path to solving the problem," colleagues say. "Maybe yes, maybe no."
7
1
84
7,216
Predicting which lab will be where when is harder than predicting timelines in general, so don’t update too strongly on me either way on this one, but I would bet that Ultra 4 will indeed be clearly better than both Astra 6 and Fable 5.1 when it finally comes out next quarter.
5
1
69
6,729