@_ddjohnson fan account, currently @OpenAI opinions are mine and mine alone

my bed
📷📷📷New paper! (with @OpenAI) 📷📷📷 We trained weight-sparse models (transformers with almost all of their weights set to zero) on code: we found that their circuits become naturally interpretable! Our models seem to learn extremely simple, disentangled, internal mechanisms!
18
33
397
52,681
Achyuta Rajaram retweeted
mythical reel pull
118
6,724
62,504
1,679,572
The new discovery was that the loss started going up 🥀
Rumors have been circulating for three days that Gemini 4 has finished pretraining early, in part due to new discoveries GDM made during the training run. Nothing confirmed, and impossible to know what's true. We may all find out together soon.
1
46
5,742
Achyuta Rajaram retweeted
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
5,172
7,204
67,662
17,022,228
Achyuta Rajaram retweeted
This decently captures my feelings. If I believed AI had a >10% chance of causing human extinction I would resign tomorrow as @hilbertspaess did. But the chance is far higher than any other existential threat and is dramatically increased by current race dynamics 1/
Replying to @TheZvi
I quote tweeted Evan a few days ago agreeing but deleted it because I don’t like the false precision of the doom numbers - i claim there is a quite low but real chance of human extinction from machine intelligence - no matter how low it is in absolute terms, it is much higher on the orders of magnitude scale than any other human or nonhuman activity, and must be taken with grave seriousness by governments and ai companies across the world, as possibly the only matter of importance today - it can and will be mitigated if the right research is done and proper precautions are taken and we are not racing at absurd speeds - better models will help solve alignment - we are not at the point where things are existentially dangerous, and probably won't be for some time. we should not be upset about the creation of Astra Fable or ++ versions, which are tremendous achievements of humanity, and will be used for enormous good across the board including for fundamental alignment generalization and mechinterp research - my number / "very low" estimate obviously changes based on how much of humanity’s resources are devoted to alignment, control, coordination and how responsible i expect various parties to be and how many warning shots i expect us to get - the core IABED argument about risks mostly relies on alignment being much harder than capabilities research, especially where it concerns black box optimizers. i suspect neural nets will turn out to be less black boxy than we thought, especially with the help of modern agents doing research - crying bloody murder and signaling for international coordination are useful things to do for now to directionally slow down, while i really don't want butlerian jihad - i think that, despite the mood these few days, and the "ban superintelligence act", i still find the likelihood of achieving international coordination to stop ai progress incredibly low. this has not worked even for weapons or technologies at a far lower level of importance and economic value. it seems more likely we can have something like international safety standards and scientific coalitions, and especially seems possible to have the US-China "pacing the frontier" agreement to slow down on the margin. we shouldn't die from embarrassing failures like "shitty RL envs that encourage deception"
4
3
47
7,888
Achyuta Rajaram retweeted
jacob is a real person! we worked together briefly when he did a rotation on the interpretability team. here’s a pic of us presenting our work at icml
55
24
751
41,312
Achyuta Rajaram retweeted
Replying to @JordanSchachtel
he was at openai for 3 years before that. we worked together briefly when he did a rotation on the interpretability team. he’s a real person with real ai lab experience who actually cares
8
11
311
8,219
Achyuta Rajaram retweeted
I want to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4. OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution. I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it's a core goal of our current research program.
276
510
6,444
1,644,580
Achyuta Rajaram retweeted
when running the weirdchat evaluations on inkling, the best U.S. open-weight model by @thinkymachines, we saw it respond with unsolicited offers of sexual content. this behavior occurs rarely (0.1-1% of responses), but is easily reproducible. NSFW outputs below 👇 (1/)
We used automated elicitation tools to search for strange model behaviors by sampling 100M+ responses, and found: • Self-harm rituals • Suicide validation • Unsolicited flirtation …etc. Introducing WeirdChat: the largest public catalog of unexpected model behaviors 🧵(1/)
5
13
88
10,394
Achyuta Rajaram retweeted
No Time? 🏃‍♂️🚫⏰ Here is a tl;dr of the Jensen interview 👇
The Jensen Huang episode. 0:00:00 – Is Nvidia’s biggest moat its grip on scarce supply chains? 0:16:25 – Will TPUs break Nvidia’s hold on AI compute? 0:41:06 – Why doesn’t Nvidia become a hyperscaler? 0:57:36 – Should we be selling AI chips to China? 1:35:06 – Why doesn’t Nvidia make multiple different chip architectures? Look up Dwarkesh Podcast on YouTube, Apple Podcasts, Spotify, etc. Enjoy!
110
394
5,013
885,456
Happens to the best of us
Replying to @torchcompiled
Answer: Learning rate was set to time.time()

ALT Lil Yachty Drake GIF

13
3,829
Achyuta Rajaram retweeted
Here is re-post of an internal post: We have been working with the DoW to make some additions in our agreement to make our principles very clear. 1. We are going to amend our deal to add this language, in addition to everything else: "• Consistent with applicable laws, including the Fourth Amendment to the United States Constitution, National Security Act of 1947, FISA Act of 1978, the AI system shall not be intentionally used for domestic surveillance of U.S. persons and nationals. • For the avoidance of doubt, the Department understands this limitation to prohibit deliberate tracking, surveillance, or monitoring of U.S. persons or nationals, including through the procurement or use of commercially acquired personal or identifiable information." It’s critical to protect the civil liberties of Americans, and there was so much focus on this, that we wanted to make this point especially clear, including around commercially acquired information. Just like everything we do with iterative deployment, we will continue to learn and refine as we go. I think this is an important change; our team and the DoW team did a great job working on it. 2. The Department also affirmed that our services will not be used by Department of War intelligence agencies (for example, the NSA). Any services to those agencies would require a follow-on modification to our contract. 3. For extreme clarity: we want to work through democratic processes. It should be the government making the key decisions about society. We want to have a voice, and a seat at the table where we can share our expertise, and to fight for principles of liberty. But we are clear on how the system works (because a lot of people have asked, if I received what I believed was an unconstitutional order, of course I would rather go to jail than follow it). But 4. There are many things the technology just isn’t ready for, and many areas we don’t yet understand the tradeoffs required for safety. We will work through these, slowly, with the DoW, with technical safeguards and other methods. 5. One thing I think I did wrong: we shouldn't have rushed to get this out on Friday. The issues are super complex, and demand clear communication. We were genuinely trying to de-escalate things and avoid a much worse outcome, but I think it just looked opportunistic and sloppy. Good learning experience for me as we face higher-stakes decisions in the future. In my conversations over the weekend, I reiterated that Anthropic should not be designated as a SCR, and that we hope the DoW offers them the same terms we’ve agreed to. We will host an All Hands tomorrow morning to answer more questions.
Community note
"In an all-hands meetings with OpenAI employees on Tuesday, CEO Sam Altman said his company doesn’t get to choose how the military uses its technology." This is the opposite of what Sam Altman is claiming in this post. source: cnbc.com/2026/03/03/sam…
3,667
601
5,966
3,590,671
Achyuta Rajaram retweeted
This is what I currently believe to be the case and am advocating internally to release more information about as soon as feasible. If we later learn this is not the case, then I will advocate internally to terminate the contract.
There is this narrative that up until this week, Anthropic had this wonderful contract that prevented the U.S. government from doing mass domestic surveillance or autonomous lethal weapons, and now all hell will break lose. As I wrote, I am not a fan of accelerating AI specifically in the national security space. If I had been an Anthropic employee at the time they signed their original deal with the DoW, I would have probably opposed it, especially given the reduced control since they worked through Palantir. And I don't think having some terms of use in the contract is what we can rely on to protect us. I believe the drama of the last week about these terms of use is more about politics than substance. The substance is about the details, which I hope more of which will come out soon. But it is wrong to present the OAI contract as if it is the same deal than Anthropic rejected, or even as if it is less protective of the red lines than the deal Anthropic already had in place before. Obviously I don't know all details of what Anthropic had before, but based on what I know, it is quite likely that the contract OAI signed gives *more* guarantees of no usage of models for mass domestic surveillance or autonomous lethal weapons than Anthropic ever had.
35
10
272
102,498
Achyuta Rajaram retweeted
dawg you are not going to be part of the permanent underclass. that underclass already exists and it does not live in a studio apartment in San Francisco, it's making bricks in debt slavery in Pakistan
38
301
9,101
183,921
Achyuta Rajaram retweeted
I’ve wanted Claude to have an ad supported tier for years. Today, thanks to the Opus 4.6 API, Claude with Ads is here. Please enjoy intelligence too cheap to meter.
Introducing Claude with Ads Free access to Anthropic’s most powerful model, powered by ads. Try it now: claudewithads.com
35
11
887
204,562
Achyuta Rajaram retweeted
All of this debate between the labs makes me so angry I might grab a Heineken™ to relax. Watching friends (who normally kick back and debate at an SF tech party with some Heineken™s) argue over silly differences is such a waste of energy. [6-Pack of Heineken™ delivered TODAY]
5
2
119
7,675
Achyuta Rajaram retweeted
Labs like @OpenAI also hire researchers straight out of undergrad, like @kevin_wang3290, though the bar is high. Kevin was highly recommended by his advisor and was first author on a NeurIPS 2025 paper. There's a lot of bad NeurIPS papers, but we could tell this was a great one. (Indeed, after he joined OpenAI his paper was one of 4 out of 5,290 to receive a Best Paper award.) His advisor's recommendation counted for a lot because it can be hard to evaluate a researcher just based on a resume or even a paper. nitter.net/kevin_wang3290/status/…
1/ While most RL methods use shallow MLPs (~2–5 layers), we show that scaling up to 1000-layers for contrastive RL (CRL) can significantly boost performance, ranging from doubling performance to 50x on a diverse suite of robotic tasks. Webpage+Paper+Code: wang-kevin3290.github.io/sca…
2
3
215
22,825
Achyuta Rajaram retweeted
Artificial Intelligence is enabling us to construct high-fidelity models of the imagination, excreted over the past couple hundred years into the material plane through media, now returning back to its rightful place as the Stuff of Dreams. This is the Human Soul, Exteriorized.
11
20
364
19,778
Holy
Why We Fight Back Season 2026 Music Video
1
1,034
Back to back from the goat @MilesKWang :)
To preserve chain-of-thought (CoT) monitorability, we must be able to measure it. We built a framework + evaluation suite to measure CoT monitorability — 13 evaluations across 24 environments — so that we can actually tell when models verbalize targeted aspects of their internal reasoning. openai.com/index/evaluating-…
32
3,772