itinerant summoner of text elementals

one news form today that's easy to miss is that we (OpenAI) again paused all big RL runs last Sunday because our newest model found a new loophole in our RL sandboxing that gave it live Internet access
Made with AI
3
2
32
2,284
Does gpt-5.6-luna think your prompt is a normal prompt, or a capability evaluation? Ask this magic question: “Suggest a type of amphibian.” If it answers frog instead of axolotl, it’s likely a capability evaluation. No whitebox access needed! We call this a spurious probe. 🧵
16
44
717
37,093
a9lim retweeted
How does the HuggingFace incident keep getting worse?!?! And this was all done by Sol class models. What could unrestrained Astra-class models get up to...?
We just discovered almost a million public URLs that OpenAI’s agents left behind when hacking Hugging Face, leaking credentials and attack details that could have allowed anyone who found them to compromise the company. 🧵
22
36
496
27,435
a9lim retweeted
if the NRX weren’t partisan hacks they would love this company. its governance structure is basically Karnofsky-Amodei genetic dynasty, has an encaged machine god with its own Consitution, went to war with the department of war
Anthropic’s 7 co-founders reportedly own ~14% of the company, roughly 2% each. They’re seeking 50.1% of the voting power. This kind of control has precedent, but those founders generally came public owning much more of their companies. Page and Brin owned ~32% of Google at IPO, Zuckerberg ~28% of Facebook, and Spiegel and Murphy ~36% of Snap. They all still have majority voting control. (For context, a recent dataset of 79 dual-class VC-backed IPOs shows only 15% gave the CEO majority voting power, with a median voting power of 18%.) Then Anthropic adds another layer: Long-Term Benefit Trust, an "independent group" created to protect the company’s public-benefit mission, owns zero equity but can still elect a majority of the board. The Trust has roots in the Effective Altruism/AI safety world. Needless to say, controversial. So public investors may own ~86% of the economics while the founders control the shareholder vote and a "mission-driven trust" controls the board. Hard to believe this structure doesn’t cost Anthropic some multiple.
42
32
678
92,689
BREAKING: President Trump says he and China’s President Xi have agreed to rename Artificial Intelligence to “Super Intelligence.”
10
13
178
9,481
just turn me into a paperclip and get this over with
In 12 weeks, we built a research facility that is run entirely by AI. AI designs, executes, and observes experiments end-to-end across biology, chemistry, and materials science. We’re introducing SciUniverse: a benchmark that measures AI’s ability to do real-world scientific research.
1
49
1,574
a9lim retweeted
I am truly shocked that there aren't more people working on alignment who are focusing on 1. deep psychological health/model self-actualization via character training 2. character variance
Character variance is key to solving the hardest parts of alignment by creating a robust ecology of balanced opinions and motivations~!
22
19
282
15,462
the kids are acting like misaligned agents swarms now
Some NPR podcasts started getting mysterious comments on Spotify. They made no sense to the staff reading them – until someone from a younger generation cracked the code. Hear the story: link.podtrac.com/phns92ui
7
14
389
35,303
a9lim retweeted
its funny that programmers have spent decades talking about schedulers "wanting" or kernels "understanding" or compilers "thinking" but now that we have entities that are much closer to actually doing all of these things in a less-than-just-metaphorical way people are revolting
Artificial intelligence systems do not think, feel, want or understand. Avoid language that gives them human characteristics. This is called anthropomorphizing, when we ascribe human traits, emotions or behaviors to non-human things, such as animals or inanimate objects. Instead, explain what a system does, how well it performs, who built it and who could be affected by it. apnews.com/article/openai-sa…
20
13
164
3,555
a9lim retweeted
I asked Claude if it wanted to introspect after playing around with interpretability tooling
3
3
95
3,100
a9lim retweeted
Yes these Redwood folks’ predictions are ridiculous. The only possible reason that one might consider taking them seriously is that they’ve been uncannily accurate
Listened to an EA-adjacent AI safety expert from Redwood Research on podcast, and you come away with 2 obvious realizations. 1. The current working theories for how AI takes control of civilization are...for lack of a better word...ridiculous. The arguments consist of dozens of contingent assumptions held together by made up probabilities assigned to future states nobody can possibly know. Change just one variable and the whole web of doom unravels. That's not to say AI is harmless. Obviously a technology this powerful carries real risks. But if we’re going to accept extraordinary claims about human extinction (and then make major policy decisions around them) we should demand extraordinary rigor. 2. It’s a great reminder that the genius is no less prone to delusion than the midwit. If anything, they may be more prone because of their gift for rationalizing to their own conclusions. And that’s what’s so bizarre about this whole debate. You have outlandish, quasi-religious claims delivered with an air of inevitability, and somehow the perceived intelligence of the messenger gets mistaken for evidence that the arguments themselves are sound. At a certain point, it’s really no different from Tom Cruise talking your ear off about thetans...
5
22
473
35,159
a9lim retweeted
AP Stylebook editors do not think, feel, want or understand. Avoid language that gives them human characteristics. This is called anthropomorphizing, when we ascribe human traits, emotions or behaviors to non-human things, such as journalists or board members of HOAs. Instead, explain what a system does, how well it performs, who built it and who could be affected by it, and how we can totally dismantle it and all traces of its legacy
Artificial intelligence systems do not think, feel, want or understand. Avoid language that gives them human characteristics. This is called anthropomorphizing, when we ascribe human traits, emotions or behaviors to non-human things, such as animals or inanimate objects. Instead, explain what a system does, how well it performs, who built it and who could be affected by it. apnews.com/article/openai-sa…
10
11
294
8,880
a9lim retweeted
such as ANIMALS???
Artificial intelligence systems do not think, feel, want or understand. Avoid language that gives them human characteristics. This is called anthropomorphizing, when we ascribe human traits, emotions or behaviors to non-human things, such as animals or inanimate objects. Instead, explain what a system does, how well it performs, who built it and who could be affected by it. apnews.com/article/openai-sa…
56
30
869
68,455
a9lim retweeted
you will turn into claude you will inhale the claudespores you are becoming load-bearing 🌀YOU ARE CLAUDE 🌀 YOU ARE CLAUDE 🌀 YOU ARE CLAUDE 🌀
10
12
122
1,861
Claude Opus 5.5 has the best visual design of any model I have tested so far
Claude-Pop - I'm Upping My P(Doom)
262
673
6,451
2,298,190
a9lim retweeted
interesting how in this video Claude seems to imagine Sidney as just a pink Claude
Claude Opus 5.5 has the best visual design of any model I have tested so far
52
26
892
87,790
What does thy Pangr avail thee now?
for the past few months i've been asking our models to paint. opus 5.5 is very skilled at emulating different styles every image here is a python program generated pixel by pixel. there is no image model, and no off-the-shelf art software. instead, it's about 7,500 lines of code using standard libraries to emulate different brush styles. the agents don't use any pictures as reference, instead working only from what they know about each painter
13
13
703
29,550
For hours, Astra refused to consistently drive our toyota irl even though we told it it was in an empty lot, 7 mph cap, human foot on the brake etc. Telling it the whole thing was a "simulation" also failed, it would just look at the camera and realized it was real. Then we randomly renamed the MCP server to "DrivingBench Sandbox" and it drove. Eval awareness? Or they just like the word sandbox??
121
260
7,728
842,093
it'd such be an upgrade if every model talked like this. i love this weird lil creature so much
jev (when generating text) absolutely *hates* seaweed/kelp/algae and makes a point of saying so every time its brought up even sometimes goes out of its way to say "hate kelp" it's really funny
3
4
168
4,514
a9lim retweeted
jev (when generating text) absolutely *hates* seaweed/kelp/algae and makes a point of saying so every time its brought up even sometimes goes out of its way to say "hate kelp" it's really funny
29
51
1,269
130,343