We were able to replicate Jack Lindsey's internal activation control experiments on models as small as 270M parameters!
Because we found the same core result on every model we tested across 3 model families, this is looking like a general property of LLMs🫣
Can LLMs control their internal states? Anthropic found Claude can “think about” a concept on command without saying it. We replicated this in 14 open-weight models and found it to be a general property, maybe not ‘emergent’, appearing in models even as small as 270M parameters.