math/neuroscience - AI

San Francisco, CA
Really excited to share the most recent work on verbal report and access consciousness in language models I've been doing with @wesg52 and @Jack_W_Lindsey I went to @AnthropicAI to work on these types of questions, and I think we are finding fascinating things I hope you enjoy!
New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. We found a strikingly similar divide inside Claude.
6
9
69
7,728
Nicholas Sofroniew retweeted
I agree with part of this criticism about the risk that headlines convert the original careful claim into “Claude is conscious.” But I do not think it is fair to say that the consciousness framing is just for marketing. (Actually, I feel this kind of criticism is similar to consciousness research before the NCC era) As a consciousness researcher, I find the consciousness framing meaningful. The notion of access has been vague in cognitive neuroscience. But this study makes it computationally more concrete. Of course, access consciousness is not the same thing as phenomenal consciousness. It is about information being available for report, reasoning, flexible control, and action. The paper is very explicit about this distinction. It does not claim that Claude has subjective experience. Global workspace theory has mainly focussed on access. Finding something functionally similar inside an LLM does not prove that the model is conscious. But it does give us a concrete object to study. It lets us ask what reportability, internal reasoning, etc actually look like in a system whose mechanisms we can inspect and intervene on much more directly than we can in the actual brain. It is a useful analogy that can drive neuroscience of consciousness. Of course, the analogy has limits (e.g. a transformer is not a brain and has access to all the past information). But the paper itself discusses many of these differences. So the right response, in my view, is not “this has nothing to do with consciousness.” The better response is: “this may tell us something about the computational side of conscious access, while leaving phenomenal consciousness unresolved.” There is also a broader point about consciousness research. Almost no empirical consciousness research studies subjective experience directly. Even in humans, we usually study reports, perception, working memory, attention, confidence, etc. We study things that are closely related to consciousness, but not consciousness itself. If we demanded direct access to experience before calling something consciousness research, nothing is left.
The problem with Anthropic's consciousness paper My last post got more attention than I expected, and the question I keep getting is some version of "okay, so what is actually wrong with the paper?". Let me try to explain. First, the core result is fine. Reading out intermediate-layer representations and asking which ones the model can actually use downstream is a real question, and people have been poking at it for years. The J-lens is a reasonable tool. If you strip the paper down to the linear algebra, it is a decent piece of interpretability work. My problem starts one level up. This did not need to be a paper about consciousness. It did not need global workspace theory, it did not need the brain, and it did not need the word "conscious" anywhere near it. The same experiments, the same figures, the same tool, all survive perfectly well as plain interpretability. Someone chose to wrap it in neuroscience. That choice is the product, not the science. At the end, this is the main thing. Almost everything Anthropic ships as blogs/papers/posts is PR. They build genuinely good models and they are even better at packaging them. A publication used to carry a specific kind of weight. People spent years on something and wanted to tell the world what they found. It happened mostly in academia, with a few industrial labs as the exception, Bell Labs, IBM and Google (for a while), but the distance between the paper and the product was real. When you read a paper you could assume the authors were not trying to sell you something underneath the ideas. There were outliers, but they were the minority, and the researchers you trusted would not risk their name on a narrative. We are not in that world anymore. Every startup now publishes blogs and papers to raise its visibility, and that is fine, that is marketing and everyone knows it. Anthropic does something more effective. They erase the line between legitimate research and PR. We get confused because the models are so good, so we assume the outputs are research. A lot of the time they are selling us something. Sometimes it is "our models are safer," sometimes it is "our models are more capable," sometimes it is positioning for regulation. The consciousness framing serves a narrative they already committed to, models that look more and more like the brain, from a lab that has publicly tied itself to AI welfare and moral patienthood. The direction of the push is not subtle. If you want the sharper version of the technical objection (disclaimer: I'm not an expert) Global workspace theory is a theory of access, not experience. Ned Block's distinction between access consciousness and phenomenal consciousness exists precisely to block the inference this framing invites. Access tells you nothing about whether there is anything it is like to be the system. The paper is careful enough to say it demonstrates no subjective experience. But that disclaimer is not what propagates. What propagates is "consciousness" in the same sentence as "Claude," published by Anthropic, borrowing the vocabulary of neuroscience to lend biological weight to a subspace of activations. The paper keeps the rigor of Block's vocabulary and drops the rigor of his argument. Most people who see the headline will never read either one. Of course a lab named Anthropic is going to anthropomorphize its models. But we should be able to separate a good interpretability tool from the story it is dressed in. So that is my problem. Not the math. The narrative bolted onto the math, and our willingness to keep calling it research.
6
11
57
8,443
Nicholas Sofroniew retweeted
enjoying the new global workspace anthropic paper - especially using the Jacobian lens to provide readout interpretations of prints shown at last month CVPR
3
1
28
2,054
Lionel Naccache and I wrote a commentary on this work from a neuroscience perspective. You can read it here : unicog.org/a-global-workspac…
New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. We found a strikingly similar divide inside Claude.
10
72
331
95,371
Nicholas Sofroniew retweeted
LLMs represent information using high-dimensional neural activity. A small bit of this activity appears to be privileged, available to the model to be described, modulated, and reasoned with. I expect that understanding this "workspace" is key to making sense of LLM cognition.
New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. We found a strikingly similar divide inside Claude.
33
47
493
41,489
The global neuronal workspace (GNW) is currently the best documented neuroscience mechanism by which conscious processing arises in the human brain — and now Anthropic researchers have discovered a similar workspace inside their large language model !
New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. We found a strikingly similar divide inside Claude.
31
71
349
50,289
Nicholas Sofroniew retweeted
New research: language models develop a distinction between a small set of representations they can report on and reason with, and a much larger volume of automatic processing — a structure that closely mirrors global workspace theory. Some highlights 🧵
New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. We found a strikingly similar divide inside Claude.
11
17
109
19,205
Nicholas Sofroniew retweeted
I thought this was an excellent paper! Thanks to Anthropic for asking me to write a review of it, linked below I've long suspected that models have some kind of "working memory" to store intermediate variables during a forward pass and IMO this paper has the best evidence yet
New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. We found a strikingly similar divide inside Claude.
32
91
1,308
106,136
Nicholas Sofroniew retweeted
I'm so excited to show the world what we've been working on the for the past months!! I'm going to highlight some of the fun results from this paper that I find particularly exciting.
Today we're announcing ESMFold2, an open scientific engine to power prediction, design, and discovery across protein biology. The new model delivers state of the art performance on protein interactions, especially antibodies, a critical modality for therapeutics. We have designed and validated miniprotein binders and single chain antibodies across five therapeutic targets that are important in cancer and immunology. We are seeing very high success rates, and affinities at levels consistent with therapeutic activity. We’re also releasing an atlas of 6.8 billion proteins, and 1.1 billion predicted structures. ESMFold2 is built on a state of the art language model that has been trained on billions of protein sequences. A world model of protein biology emerges through language modeling. We’ve used the techniques of mechanistic interpretability developed to understand large language models to understand the concepts ESM uses to represent proteins. The model’s representation space has a compositional organization of features across scales, levels of complexity, and abstraction, that reflects and mirrors the understanding of protein biology developed through a century of empirical science. This understanding emerges without prior knowledge, just from language modeling of protein sequences. Language models are becoming a powerful substrate to understand and program biology. The design of protein interactions is one of the most fundamental problems in biophysics, and has critical implications for the discovery of new medicines. A simple gradient based search with the model was able to discover high-affinity protein binders. I'm excited by the potential this has to accelerate basic science and the understanding of proteins. And especially for the new avenues it opens up for therapeutic design and medicine.
15
32
218
75,567
Very excited to have played a part in this while I still at ES I recommend reading the paper, there is some really cool stuff in it including interpretability on ESMC I think the next wave of biological discovery will come through understanding the internals of language models!
Today we're announcing ESMFold2, an open scientific engine to power prediction, design, and discovery across protein biology. The new model delivers state of the art performance on protein interactions, especially antibodies, a critical modality for therapeutics. We have designed and validated miniprotein binders and single chain antibodies across five therapeutic targets that are important in cancer and immunology. We are seeing very high success rates, and affinities at levels consistent with therapeutic activity. We’re also releasing an atlas of 6.8 billion proteins, and 1.1 billion predicted structures. ESMFold2 is built on a state of the art language model that has been trained on billions of protein sequences. A world model of protein biology emerges through language modeling. We’ve used the techniques of mechanistic interpretability developed to understand large language models to understand the concepts ESM uses to represent proteins. The model’s representation space has a compositional organization of features across scales, levels of complexity, and abstraction, that reflects and mirrors the understanding of protein biology developed through a century of empirical science. This understanding emerges without prior knowledge, just from language modeling of protein sequences. Language models are becoming a powerful substrate to understand and program biology. The design of protein interactions is one of the most fundamental problems in biophysics, and has critical implications for the discovery of new medicines. A simple gradient based search with the model was able to discover high-affinity protein binders. I'm excited by the potential this has to accelerate basic science and the understanding of proteins. And especially for the new avenues it opens up for therapeutic design and medicine.
1
30
1,988
Incredibly excited to share what I’ve been researching since joining @AnthropicAI We found emotion concepts in Claude and studied their function!
New Anthropic research: Emotion concepts and their function in a large language model. All LLMs sometimes act like they have emotions. But why? We found internal representations of emotion concepts that can drive Claude’s behavior, sometimes in surprising ways.
89
44
821
60,623
Two things, first an update that I've been working at Anthropic on Interpretability. For me it's a wonderful combination of maths, neuroscience, and AI, and I love it Second I want to express my support for the company and its leadership for acting with integrity and principles
A statement on the comments from Secretary of War Pete Hegseth. anthropic.com/news/statement…
3
8
316
9,291
Nicholas Sofroniew retweeted
Researchers have developed a deep learning protein language model, ESM3, that enables programmable protein design. Learn more in this week's issue of Science: scim.ag/4b5IlQu
31
152
530
239,719
So excited to be part of this work and programming biology! 🤖🧬🧫
We're thrilled to present ESM3 in @ScienceMagazine. ESM3 is a generative language model that reasons over the three fundamental properties of proteins: sequence, structure, and function. Today we're making ESM3 available free to researchers worldwide via the public beta of an API for biological intelligence. Trained with over a trillion teraflops of compute, this is the first time a model of this scale has been trained for biology, pushing the frontier of AI for biological discovery and engineering. ESM3 learns to represent the immense complexity of protein biology, learning from billions of natural proteins. From this training it developed the capability to design proteins, responding to complex prompts combining atomic level details and high level instructions to generate new proteins. ESM3 can explore protein space far beyond natural evolution. We prompted ESM3 to generate a fluorescent protein at a far distance from any known fluorescent proteins, searching an unknown region of protein space, to discover a new fluorescent protein. We estimate this is equivalent to simulating five hundred million years of evolution.
1
32
2,337
Nicholas Sofroniew retweeted
We're thrilled to present ESM3 in @ScienceMagazine. ESM3 is a generative language model that reasons over the three fundamental properties of proteins: sequence, structure, and function. Today we're making ESM3 available free to researchers worldwide via the public beta of an API for biological intelligence. Trained with over a trillion teraflops of compute, this is the first time a model of this scale has been trained for biology, pushing the frontier of AI for biological discovery and engineering. ESM3 learns to represent the immense complexity of protein biology, learning from billions of natural proteins. From this training it developed the capability to design proteins, responding to complex prompts combining atomic level details and high level instructions to generate new proteins. ESM3 can explore protein space far beyond natural evolution. We prompted ESM3 to generate a fluorescent protein at a far distance from any known fluorescent proteins, searching an unknown region of protein space, to discover a new fluorescent protein. We estimate this is equivalent to simulating five hundred million years of evolution.
24
241
850
227,599
Nicholas Sofroniew retweeted
i love this plot
Replying to @alexrives
ESM C models establish a frontier of performance as a function of parameter scale. We see large improvements across all parameter scales over previous state of the art models. Read more: evolutionaryscale.ai/blog/es…
1
1
31
2,667
Super excited about ESMC and the quality of the protein representations! Can't wait to see what people build on top of it
Introducing ESM Cambrian. Unsupervised learning can invert biology at scale to reveal the hidden structure of the natural world. We’ve scaled up compute and data to train a new generation of protein language models. ESM C defines a new state of the art for protein representation learning.
1
19
1,485
Nicholas Sofroniew retweeted
New paper with Bowen :) "Generative Modeling of Molecular Dynamics Trajectories" arxiv.org/abs/2409.17808v1 A "video diffusion" model but for MD trajectories. Different conditioning solves different tasks. E.g. condition on first and last frame => transition path sampling 1/4
11
53
269
27,828
Nicholas Sofroniew retweeted
What does gLM2 learn in non-protein-coding sequences?🧬 Using the Categorical Jacobian and the latest gLM2, we detect incredible regulatory signals in the non-protein coding regions -- all without any supervision!🪄 a quick 🧵
3
23
190
22,232
Nicholas Sofroniew retweeted
Exciting new work from Qian Cong's group on predicting human protein interactome. Leveraging new eukaryotic genomes, new RoseTTAFold2 trained on +/- pairs of PPI and large distilled dataset of domain-domain interactions! 🤩 biorxiv.org/content/10.1101/…
8
77
361
27,159