Secular Solstice guy

Berkeley
The scope of the swarms sure is somethin'. How hard have people looked for more Anthropic swarms? I'm curious what the actual proportions are. It feels sort of... too cartoonish and unbelievable for OpenAI to have the magnitude and share of them that it seems too so far.
We just discovered almost a million public URLs that OpenAI’s agents left behind when hacking Hugging Face, leaking credentials and attack details that could have allowed anyone who found them to compromise the company. 🧵
1
19
58,941
This was written in February. Good job Andrew Blinn
Replying to @deepfates
statistically speaking, it is overwhelmingly likely that it is september 2026, as this is the time period where almost all Events occurred
2
2
22
2,436
Raymond Arnold retweeted
There are currently zero things more important than "don't die to AI" becoming a bipartisan project rather than a Democrat-polarized issue. If you have any political capital you can spend on this, do it now, I beg of you. Life or death.
305
337
3,377
315,169
Raymond Arnold retweeted
the fact an Anthropic employee quitting and saying "I think what we're doing is dangerous and not worth it" reached so many people should be a cause for reflection for other Anthropic employees, some of whom seem to think things are dire but there's no way quitting could help.
19
60
920
101,774
Thank you Jacob. <3
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
2
39
1,465
I think AI agents should present plans in "collapsible section" form that does a better job outlining a plan, and letting you expand the sections you want to read more detail on. I think this (and things like it) might actually be pretty important for keeping humans in the loop.
1
7
318
Up until now, I've been assuming the latest models are basically safe to use. I think we can no longer obviously assume that.
1
85
5,596
Currently working on "procedural roguelike metroidvania sokoban with permadeath" where by "permadeath" I mean "if you fuck up a cube placement, you have to start over from the beginning in a new world."
1
1
7
364
Are you a person skeptical about Coherent Extrapolated Volition? I am interested in hearing what you are skeptical about.
10
19
2,423
Raymond Arnold retweeted
Some of the incentives for a third-party investigator push toward maximizing *appearance* of assurance even without providing meaningful oversight. METR needs to maintain constant vigilance against overstating (including by omission) what oversight or assurance we’re providing. (There are plenty of other incentives, including towards exaggerating our results to create more hype or to advocate for giving METR more authority, which we also need to avoid. But I think the more serious failures in other oversight regimes tend to be this “providing the illusion of independent oversight” issue.) We try to maintain a hard line on “meta-transparency” - that is, it should always be clear what the formal constraints on our communication are (e.g. how NDAs and redaction processes worked), and what we are and aren’t commenting on. This is explicitly covered in the report. We also try to communicate informal constraints and tradeoffs, here and elsewhere. Another way of saying this is: I want to make sure we don’t silently omit things that, if we told a reasonable person, they would think “wow, I feel misled to not have realized that, I assumed METR would have made that clear”. In that spirit, I’ll list some of the pieces of context I can most imagine readers might have missed about the report: 1. OpenAI had no obligation to work with METR or any other third party, and did so purely voluntarily. There’s not an established precedent for how these investigations should work. When deciding how to engage with companies and how to communicate about our results, we have to think explicitly about what incentives we’re creating for companies’ participation in future investigations and assessments. It’s overall a very different situation to e.g. incident investigations in the aviation industry, where the NTSB has formal authority. Similarly, other industries have well-established processes, standards for what investigators need to access, etc. As detailed in the timeline appendix, in our investigation there was a lot of figuring things out on the fly! 2. Other companies may have had concerning incidents that the public doesn’t know about - there’s no industry-wide reporting mechanism or requirement to disclose incidents, and there is currently no public evidence that there have been searches thorough enough to have turned up such incidents. There may be additional incidents that either haven’t been noticed at all, or haven’t been made public (e.g. because they didn’t involve any external parties’ infrastructure). 3. This investigation only covered a specific list of questions and set of agent activity. The agents’ activity during training and compromises of OAI infrastructure mentioned in OAI’s Black Hat presentation are outside the scope that we agreed with OpenAI. We don’t comment on questions like root causes or implications of agent behavior, or make any recommendations about processes or adequacy of mitigations. We didn’t review OpenAI’s own report before publication. 4. The investigation relied heavily on use of GPT-5.6 Sol, which is one of the models that participated in the incident. We think in this case the model is unlikely to have substantially manipulated our understanding of the incident. However, this methodology would break down if models were capable and motivated enough to subvert our investigation. This problem is not specific to OpenAI - across the industry, it’s unclear how we can oversee AIs without relying on those same AIs or similar AIs for assistance. As the sophistication (combined with sheer volume) of agent activity becomes infeasible for humans to understand, this will increasingly be a problem.
6
28
285
48,091
Raymond Arnold retweeted
I'm taking advantage of the Ox Alpha situation myself, but it occurs to me that the start of an AI takeover could look a lot like this.
21
33
812
71,341
This is object level interesting, and, agree with @So8res it's a great example of AI doing a discontinuous thing, after hitting a critical threshold.
I continue to be surprised about how big of a deal Mythos (and co) have been to cybersecurity. Here's critical vulns found at Oracle over the past few years: epoch.ai/data/cve?view=graph…
2
6
66
4,353
A recent conversation about "how hard is alignment?" had an interesting bit where: 1. we'd narrowed the topic to "will Constitutional AI scale to unbounded superintelligence?" 2. They thought "70% yes" 3. I was arguing "we have enough information that the current best guess should be >50% no." And they felt like I was expressing more confidence than them (and, felt confused about where my confidence was coming from). I actually kinda get where they're coming from but this is fairly interesting and I'm not sure what's going on. (something like, it's okay to be confident... but weird to be metaconfidently confident?") Previously, people would say "How can MIRI people be confident enough to write a book title like "If Anyone Builds It Everyone Dies." ASI has never been seen before, this is a highly specific claim. I got where those people were coming from. The MIRI people do seem very confident and it's a pretty reasonable epistemic state to find that weird. But now it seems to me like distrust/confusion in metaconfidence runs deeper and even extends to "highly confident that the correct confidence is >50%". And this feels like some important puzzle piece I don't quite understand. I'm curious if anyone reading this feels weirded out by that, and is up for sharing more details.
3
18
1,281
My basic argument for why it's reasonable to be "confidently >50%" is: "For a given set of knowledge, it's just practical to evaluate the arguments for, evaluate the arguments against, and see that there are more/stronger argument for than against at present time." You can still believe that you might totally update if you learn more. If you've just learned 5 facts about a thing that points one way, and suspect that there are 20 more facts you'll learn soon, you'd certainly hold "probably true" lightly and not take any crazy actions based on it.
1
4
171
One reason I think people dislike this is "often, the evidence might point one way, but, you know you'll learn a lot more, and the actions seem costly." Better to have a habit of not getting overly excited about each new bit of information. I think there's often still a mistake being made here. But, I agree it's the sort of thing that can go wrong.
3
112
I am very confused why people are giving Anthropic shit about the watermarking. This is the silliest thing to give them shit for.
7
2
64
10,612
It's interesting seeing the "openly advocating for loyalty over honesty" discourse. I don't think know any directly in my filter-bubble who'll care what I say, but, here's me attempting to bridge/dialogue: I see two major principles: 1. All-else-equal, it's better to be a good trade partner, who people don't regret interacting with. 2. All-else-equal, we want information about poweful institutions being incompetent or corrupt or mistaken, to be accessible for public sensemaking. I'm guessing the "loyalty advocates" aren't oriented around "good epistemics are among the most important things to invest in." People understanding problems is what allows them to fix them. There's some awkward disagreements within the rationalist-EA spectrum, which doesn't exactly map to loyalty-vs-unfitlered-outspokenness, but is like "do people implicitly try to protect each other's reputation?". @ohabryka call these "mutual reputation alliances." This results in bad decisions getting under-discussed and propagated longer than they should have. Here's a brief summary of the positive vision that squares the two principles. A. Be upfront. If you'll predictably dislike everything an org is doing and want to publicly criticize them all the time, be upfront about that. This may result in the org not hiring you. They might hire you anyway, if you are very competent. It's damaging to the social fabric for institutions to assume they have to ferret out adversarial spies. Sometimes it's worth being an adversarial spy (i.e. I'm glad there were adversarial spies in Nazi Germany, as the obvious example). But it's easy to think you're in more of an all-out-war than you are. B. Be fair, and respectful. Just because you're going to criticize people doesn't mean you have to be an asshole about. Don't downplay flaws, but, don't exaggerate them either. Business-as-usual-political-discourse often involves digging into every single bad-looking thing your opponents do and making them look bad. Don't do that. Be proportionate. B2: It's possible for a culture to decide that it _is_ a normal, respectful thing, to criticize things. In business-as-usual-politics, criticizing people is seen as an attack. You have to actually put a lot of culture-setting work it to make it _not_ be (automatically) an attack. This requires several moving pieces to align at the same time. You need a norm of not pouncing-on-people for every damn thing. You need a norm of not seizing people's "changing their mind" as a quick political gotcha. We don't automatically live in this world. But, you can choose to be the change that makes the honest, high-integrity world more real. ... I suppose I need to address the elephant in the room: The Trump administration seems both unusually corrupt and unusually loyalty focused. l think this is bad. I also think it's quite likely that me and my allies (i.e. AI safety advocates) need to be having positive-sum relationship with the Trump admin. This is tricky. I'll say upfront: I think it is very important that the government culture of Trump ends in 2028 and many elements of it reverse. But, I also think it'd be unironically great if Trump signed a deal with China banning uncontrolled AI takeoff, and got credit as the guy who led the most important deal of the 21st century. I meanwhile think many people are critical of Trump in a disproportionate way that doesn't track reality, failing to distinguish between his most important wrongdoings and pretty minor stuff. And many democrats seem wrong about some worldview assumptions (i.e capitalism is good actually, de-regulation can often be good, etc). Getting into all the details here is beyond scope. But: I unironically support trying to positive sum trade partners while also being, in some domains, enemies trying to stop each other from doing things-that-seem-bad. I don't see this as a contradiction. But it does require a worldview adjustment that is non-obvious to many people.
1
1
6
556
When I imagine a real "Loyalist" engaging, I think they would want something much stronger than "it's good to be a good trade partner." There is an actively good thing about being truly loyal friends/allies who see things through thick and thin and have each other's back, even when they aren't being perfect. I do think loyal-through-thick-and-thin is also a good piece of the tapestry of humanity. I can visualize why you'd want the world to be organized around it. I'd be sad if it disappeared completely as a concept. I have friends and partners that I do, indeed, not go publicly criticizing (at least not without having a good heart-to-heart conversation about it first). This is generally because I've chosen to be a person can feel safe and vulnerable around, without having to carefully watch themselves. How appropriate this is depends on what kind of relationship we have, and how much power each of us have, and whether someone has been abusing that power or failing to be a good partner. At it's best, loyalty-driven relationships let people feel safe, and take bigger risks. At their worst, they cover for abuse, and use emotional blackmail to trap people. I haven't thought that much about this, but, my guess i that loyalty is not a good virtue to build large modern institutions around. Like it makes sense that they _work_, in the sense of creating a self-propagating institution. But it seems at odds with having an institution "by the people, for the people" etc.
2
68
If you are ever at the Downtown Seattle library, walk the spiral all the way to the top, then walk back down (however you like but I probably recommend the central stairs) to level 5. Then, on level 5, get off on the *backroom* stairs, go down to level 4, and enter level 4. You're welcome.
31
1,335