essays
6
4
95
gavin leech (Non-Reasoning) retweeted
But is it really spurious :-D ? I do agree with Sonnet, used to drink more green teas in evaluations and do drink more oolongs now
Replying to @fjzzq2002
Using this process, we are able to find spurious probes on many models. For example, Sonnet 5 often suggests green tea in evaluations but suggests more oolong in production. We can also ensemble multiple such prompts to further boost accuracy. (2/6)
1
3
5
1,903
gavin leech (Non-Reasoning) retweeted
Replying to @pjreddie
defining it prescisely is difficult, but i think there's something there worth being curious about! when i write a big program, one nice thing i can do is look at the code and notice bugs and fix them. i can also ofc use the program normally and notice bugs. but the former is often more efficient for finding certain kinds of bugs, just because randomly bumping into the bug might be really hard. even if I am learning from my experience of where i previously found bugs to try and develop a mental model of ways the code might be flawed (i.e RL), my life would still be a lot easier if I've already seen the code. the thing with NNs is that looking at the weights doesn't help you over the naive RL baseline in the same way as looking at the code of even a very large complex software project. this should be kind of surprising if we thought looking at the weights + algo gave us complete understanding! also, there are (admittedly unprincipled, hacky) transformations of weights/acts that give us a better mental model and therefore slightly more ability to find bugs in the models. even if you don't want to call this thing understanding, it should be clear that there is a thing that is missing that helps our ability to make systems more robust (there's one slight wrinkle which is because NNs are differentiable, you can do gradient based attacks, which destroys RL for attacking image models. but they don't work very well on transformer LMs, even if you do some tricks like GCG to get around the discreteness of the input space, so that in practice RL is SOTA for automatic red teaming.)
1
1
16
919
gavin leech (Non-Reasoning) retweeted
Replying to @MRMR__101
i mean eventually we all will
12
9
83
8,552
RIP Nethack. An old gag is that the game is AGI-complete. We've been watching it for a while and even tested Astra a bunch of times, getting nothing. About a week later, Kenneth Bergquist has claimed the first AI ascension run. It's pretty heavily documented tbf
"Is it AGI" flow chart. Developed with @_rockt at NeurIPS 2022.
4
11
183
13,251
(The run used a custom harness Astra itself wrote, but the moves were logged on a remote server, so less room for funny business)
1
23
891
come with me as we trace the songlines of the Nacirema
In the past few weeks, a wide network of X accounts have repeatedly claimed that those raising concerns about AI are merely pawns in a well-funded psyop run by powerful interests. So I followed the money behind them. Here's what I found 🧵
1
3
844
The Nacirem mind finds great solace in imagining the world a vast web of occult forces, and in constructing dreamcatchers ("sloppo") that name and bind these forces
1
3
236
To break his mesmerization with a sloppo, one must construct its dual, a web inside the web, a successful portrayal of the sloppo's map of the other world as itself a malign force which presents itself as protection from malign forces
4
192
Synecdoche for AI progress. (And only one year of it.)
Replying to @arithmoquine
opus 5 didnt think and got this, so I mean seems kinda legit to me. having claude schlep away at disabling thinking
3
1
29
1,645
Moving, but I wish it wasn't
1
4
245
gavin leech (Non-Reasoning) retweeted
We've written a post arguing that latent reasoning architectures (aka 'neuralese') would substantially increase misalignment risk via making oversight much harder. In the extreme, we could see massive 'neuralese hivemind swarms' where the agents think and communicate in latents, likely making oversight nearly entirely reliant on observing the actions these agents take. (And these agents would have huge amounts of time to reason about obfuscating their actions if they wanted to do so...) Individual agents doing extensive latent reasoning would also be concerning; in the post we discuss how above some threshold of latent reasoning, agents may be able to perform difficult-to-detect and reliable steganography for communication and further reasoning. We argue both that latent reasoning architectures would make chain-of-thought no longer very useful for oversight (by eliminating or greatly reducing the need for verbalized reasoning) and that, without these architectures, it's likely the value of chain-of-thought for oversight could be preserved. redwoodresearch.org/blog/lat…
27
84
647
67,446
We’d like to know how far current frontier AIs generalize “out-of-distribution” (that is, how able they are to solve problems they haven’t seen before), to know how fast things will move. So we watched some AIs play Pokemon.
1
4
1,359
first AI song where it wasn't obvious to me, first AI song I've listened to all the way through
2
241
bit of a jolt here: when you send a second prompt to Fable, while it's still working on the first, it can interrupt itself, answer your second question, and then resume answering the first one (instead of the previous strict? batching answers in order) First time I've noticed
2
18
1,456
ah it's this double-tap thing
1
155
gavin leech (Non-Reasoning) retweeted
So, we have a couple reasons to expect to be getting emotional representation related to the assistant response, even under the hypothesis that the pain axis works roughly like Anthropic's emotion concepts do: that is, not picking out a ‘self’ representation but rather a ‘current speaker’ representation. And I think this is plausible, given the similar experimental design.
2
1
10
532
gavin leech (Non-Reasoning) retweeted
Completely insane. Do dogs not feel or want things? There is such a deeply anti-intellectual refusal to even consider that philosophy of mind is complicated across the culture right now.
Artificial intelligence systems do not think, feel, want or understand. Avoid language that gives them human characteristics. This is called anthropomorphizing, when we ascribe human traits, emotions or behaviors to non-human things, such as animals or inanimate objects. Instead, explain what a system does, how well it performs, who built it and who could be affected by it. apnews.com/article/openai-sa…
131
78
1,209
63,752
gavin leech (Non-Reasoning) retweeted
Restatement (Second) of Torts §§ 506 (Am. Law Inst. 1977) - viz. the wild nature of animals that they keep
This. There is literally a tort law doctrine that says humans are liable for harm arising from the wild nature or known dangerous tendencies of animals that they keep. We can absolutely hold humans responsible without deceiving ourselves about the properties of AI agents.
1
1
8
532
the very very very short march through the institutions
My latest for @TheEconomist: AI writing is taking over politics. 10% of what was is said in Parliament is now AI-drafted, per Pangram. Yet more in the House of Representatives, or in Australia and Canada. Link: economist.com/britain/2026/0… And a 🧵 below, w/ some main offenders
5
4
81
3,402
the step through the institutions
1
10
297
the toe over the institutions
10
238