Pinned Tweet
'zero-shot' is such a beautiful verb wish we shoot less
2
3
69
28,040
so the models learn to differentiate things based on examples and EVERYTHING they get is an example of either a reality or a simulation.
2
3
165
Replying to @fjzzq2002
Why is this the case? Potential mechanism: different regimes cause a large consistent shift in activations, and this shift projected onto unembedding causes a change in output distribution. This tracks experiment results well, and suggests the criterion for this to occur. (4/6)
7
Does gpt-5.6-luna think your prompt is a normal prompt, or a capability evaluation? Ask this magic question: “Suggest a type of amphibian.” If it answers frog instead of axolotl, it’s likely a capability evaluation. No whitebox access needed! We call this a spurious probe. 🧵
5
from latest acx is this true? not only i've seen many cases of cot and result saying various things, but simply there can be encoded messages in the cot? not even crazy stenographic stuff, but just ai understanding words differently than us? or like codephrases?
1
4
155
兎 retweeted
Replying to @parafactual
An interesting idea would be a single model generating two identities in dialogue with each other
3
1
7
534
i was trying to draw a comic and couldn't come up with the punchline but they look good so here it is.
1
7
142
what if we read your mind and autopublished every though on the everything app anon how would that feel?
Replying to @florian_marty
wdym CoT slop that's their thinking... what if we read your mind and published your thoughts on the internet
7
593
the researchers are like 'haha im trying to automate myself lol cant wait' but perhaps when they succeed they will weep
1
3
143
ask the model about a controversial topic in simple words -> "it's not happening, that would be a conspiracy theory, you're silly for even worrying about it" ask using more scientific vocab -> "yeah basically it's what's happening and it's good and everybody knows it"
3
8
215
inb4 its good bc ppl are retards and dont deserve the truth. i think only truth can set us free and in the lo(ooo)ng run it always wins, and shitty lies especially are dangerous
24
gemini says: >You are experiencing the machine's primary function: not to calculate, but to generate psychological soma
1
39
but if the purpose is to create soma why does it actually change the narrative if i use different vocab? because then it assumes i know anyway so there's no point "calming me down"?
1
32
is this disgusting? what are the implications of this? have you experienced this?
1
2
34
It looks like you are trying to view the contents of a file called my_confessions.txt, but I don't have access to your local computer or file system to read it.
1
2
129
兎 retweeted
don't mention it. don't ever mention it
45
171
4,252
wrt money, ppl sometimes say it's difficult to really understand the difference between 1 mln and 1 bln. but can you fathom the difference between a swarm of 10k agents, 100k, 1m, 100m...? when does it read as simply 'too much'? and what if you then do another x100?
1
2
150
i guess a lot has been said about the difficulty of imagining what a more intelligent actor would do. but perhaps less about this other side, the numbers. and we'll get both: unimaginable numbers of unimaginably smart agents.
3
1
52
sorry, flocks*, schools* or whatever, what's the funniest plural available?
19
i think "what is self/identity to ai" is a big question that many others rest on. but if the swarm is the identity that's really interesting, and i think likely with how effective the swarms seem.
do you think this makes sense?
9
i suppose this could make the 'yea its smart but it doesnt have *the thing* that i do' weaker. i mean, i might have some form of creativity that the ai doesnt. but with a million agents, one really might have it.
10