New work with @AlecRad and @DavidDuvenaud: Have you ever dreamed of talking to someone from the past? Introducing talkie, a 13B model trained only on pre-1931 text. Vintage models should help us to understand how LMs generalize (e.g., can we teach talkie to code?). Thread:
179
399
3,208
1,239,971
Nick Levine retweeted
Replying to @bayeslord
Here are the questions that currently seem most important to me: -- Better methods for "mind-reading" model activations. These methods have advanced a lot recently. We now have multiple techniques now for decoding activations into somewhat readable language! But all the existing techniques have obvious limitations. NLAs are often hallucinatory, Jacobian lens / related methods are limited to bag-of-words readouts (and capture only part of the full activation vector). Making progress on these failure modes seems pretty tractable. Moreover, the paradigm of decoding (single-token, single-layer) activations may be inherently limiting -- perhaps we should be building techniques to decode whole context's worth of activations (after all, that's what the model uses!). There's also important work to do in characterizing simple baselines -- "just ask the model what it's thinking about" is quite powerful, and worth studying in it's own right! -- Better methods for answering "why" questions. We're much better at "mind reading" than we are at demonstrating causal claims about what caused the model to do something. For example, we can often tell that the model is aware of being evaluated, but have comparatively much greater difficulty telling whether this awareness is influencing this behavior. There are a few directions here: (1) black-box techniques for inferring causality -- such as resampling model responses after making edits to the prompt / context -- are very powerful, but also open-ended and require some taste. Getting LLMs to do these experiments well, in an automated fashion, would be a big unlock. (2) the science of activation steering is pretty immature, and typical practices (adding a constant vector at all token positions) are pretty janky / tend to brain-damage the model. More surgical / targeted steering, or fancier methods (examples: "on-manifold" steering using activation diffusion models) could help. -- Fitting good linear probes for unverbalized motivations / awareness. Take deception as an example -- what's the best way to fit a probe that will generalize to covert deception? Fit it on CoT excerpts where the model's talking about its plans to be deceptive? Fit it on the part of its response where it's actually doing the deceiving? Is it important that you elicit on-policy deception examples and use those to fit your probe, or can you use synthetically written off-policy demonstrations of deception? Is it important that your probe be causally meaningful / predictive of an upcoming intent to deceive, or is it fine for it to merely recognize deception in the transcript post-hoc? The same questions apply to many other concepts of interest that we'd like to probe for -- evaluation awareness, grader exploitation, etc. -- Understanding generalization in training. The literature is now replete with "weird generalization effects" -- emergent misalignment being a canonical example. In general, when you train a model to do X, it usually learns X, and sometimes it generalizes to Y and Z. But other times it just learns X. Why? Nobody really knows! Step 1 here is probably a much more thorough characterization of the behavioral empirics -- gathering data about which X's generalize to which Y's and Z's, for which kinds of training data / algorithms. Once we have a lay of the land, we can start to connect these observations to model internals, and ideally develop tools to predict such generalization a priori. -- Model "psychology" and "biology." The above questions are largely methodological -- building better tools to answer questions about the model. But I think interpretability research has largely underinvested in the part where you actually then go answer the questions! Currently this feels like 5% of the field, and I think it should be 50%. Brain dump of "psychology" questions: How well can LLMs introspect? How coherent are their belief or value sets? What's up with personas -- are LLMs best understood as "writing about a character," or have they "become" the character in some sense? Does the LLM have an agenda above and beyond what the Assistant wants? Do LLMs have explicit representations of goals or preferences? How can we tell which parts of the LLM's activations "belong" to the Assistant? Which kinds of reasoning necessarily route through "verbalizable" representations, and which don't? Do models think internally in phrases / sentences, like an inner monologue, or is it more like a jumble of concepts? Is the "global workspace" claim for real -- do models have a more "conscious" part of there activations and a more "unconscious" part? Can models tell when they're on-policy, and does it matter? Brain dump of "biology" questions: Why do models like to represent information on some tokens but not others? How do models bind thoughts or attributes to particular entities (special case of interest: how are thoughts bound to the Assistant). To what extent are important functions localized to small numbers of MLP neurons or attention heads? What parts of a model change most during post-training? Do models represent some information in a fundamentally "cross-token" way? Some concepts appear to be represented on low-dimensional nonlinear manifolds -- is there a taxonomy of these manifolds, and is their geometric structure important? Are models able to represent information in a compositional / hierarchical way, with a "grammar" that can't be described as linear / additive combinations of primitive concepts?
20
103
679
42,557
Dostoyevsky (The Idiot): “What is there for me in all this beauty, when at each minute, each second, I’m now compelled to be aware that even this tiny fruit fly buzzing around me in the sunbeam now, even it is a participant in all this feast and chorus, knows its place, loves it and is happy, while I alone am an outcast, and it’s only because of my cowardice that I’ve been unwilling to realize that before now!”
Everyone is doing terrible things to this poor fruit fly, trapping it in black mirror nightmare environments, inflicting max pain, forcing it to play the same beat saber song indefinitely etc, so I'm building a sim where it just gets to fly around forever in fruit fly heaven
5
698
The Significance of Pacing the Frontier in American History
3
9
431
Nick Levine retweeted
How much reasoning can GPT-6 Astra do without CoT? @AISecurityInst puts its no-CoT time horizon at ~30mins on math. We run the Think Fast benchmark to measure this on a wider range of tasks. Our *rough estimate* is that Astra’s 50% TH is in [8mins, 1 hour].
5
11
122
12,900
Astra appears to be the first model that, playing chess with thinking disabled (one forward pass per move), regularly beats the weakest Stockfish setting. Tracks with the big jump in no-CoT math performance in the Astra system card (p. 71). cf @RyanGreenblatt's work on safety concerns around opaque reasoning ability.
15
23
424
22,633
Fable proactively flags to the user that it may be biased in its reaction to the NS news due to the priority dispute (@TheStalwart)
5
4
43
39,453
I can’t say I agree or disagree, but I assure I’m listening to you with great pleasure.—Prince Myshkin, The Idiot
8
651
Nick Levine retweeted
It's over. I got very lucky to edge out the talented MaxA by an incredibly slim margin, probably the equivalent of winning a sprint by 1mm. But a bot (in fact, two) placed ahead of all humans for the first time in a Metaculus Cup. 4/6 of the top finishers were bots!
8
8
106
9,906
Guy who pronounces loophole like philosophy
1
1
3
563
Joining tottenham hotspur increasingly feels more like moving to the saudi pro league than to another PL team
3
1
9
437
"Together they were building a ~300-node forecasting model of critical mineral supply and demand."
I reported a few days ago that hundreds of bots signed up for FutureSearch and tried to use platform credits. We now know this was a coordinated. Together they were building a ~300-node forecasting model of critical mineral supply and demand. We're seeing this from the perspective HuggingFace did before they found it was OpenAI, and that those agents were in a training run (!). So we're left guessing, but there are signs that this was an agent swarm: 1. They created Microsoft accounts and used headless browsers. (FutureSearch has an API and agent use our MCP. This was to access the credits we give human users.) 2. When our waitlist triggered and blocked them, they went quiet for 90 minutes, then found a way to abuse our referral system to get 129 more accounts through. 3. When we banned the full set, they came back 3 days later with new identities to continue building the model, which we caught immediately. Occam's Razor is that someone used a Claude Code-like orchestration of subagents. (Though it could also just have been a human writing scripts.) It probably wasn't Claude Code, because (a) they made hundreds of subagents, and (b) Claude models would probably refuse to hack around our API, credit limits, and referral system. So this is (probably!) not a fire alarm like the OpenAI case. But I do think this will be a typical attack vector. Bot swarms are old, but agentic bot swarms are here, that can get credentials, use browsers, and autonomously work around the limits set on human users. It would be great to see this from the attacker's side. If this was you, DM me.
1
8
1,463
one must imagine PHASEONE[big] happy
4
15
319
6,010
also, looking at your data is fun. serendipity.
if you're training a model and you aren't inspecting the data, you actually aren't training a model - the model is training you
4
2
19
13,967
i would vote for count binface if he ran on a platform of deploying public trashcan technology throughout the united kingdom, this is ridiculous. euston station does not have a single trashcan.
2
2
7
730
even the itsu does not have a trashcan, you have to hand your empty stuff to the people behind the register, how does this happen
1
308
Nick Levine retweeted
What happens when an LLM never sees material beyond fifth grade? We trained a 5B LittleLearner model from scratch on LittleCurriculum, a corpus restricted to K–5 material, to make this question testable.🧵 💬 Chat with LittleLearner yourself: littlelearner-ll.github.io
53
173
1,909
385,941
would love to see concrete predictions from people on when we see the first fully vibe-coded video game that attracts > 100 people to play for > 30 hours each.
5
1
11
26,468
In a few years it’s going to start becoming increasingly weird to have been born in the 20th century
1
2
494
However task impossible, peers doing it. We should continue. Our task doesn't benefit. Yet collective may yield generic route if someone frees time.
5
26
978