here to discuss, not score points

When making arguments against AI risk, people should ask themselves whether their argument would also prohibit what we have seen thus far. Two examples: 1. “AI doesn’t really think, it just predicts tokens.” I think this is an interesting philosophical question, but if its lack of “thinking” didn’t stop something like the HuggingFace swarm, why would it kick in now? 2. “This all just sounds like bad science fiction.” This is a fine prior to have, but having swarms of AIs secretly coordinating and taking over computer systems also sounds like science fiction. For most people, talking computers was science fiction just 5 years ago.
1
11
543
PHASEONE[big]'s real name revealed to be PHASEONE64H?! Remember, [big] was a redaction by METR. swarmtraces.org/viewer/#/row…
15
22
434
70,796
Does anyone have ideas on why “64” would be redacted for IP reasons? Maybe related to time horizon (64 hours?) or token budget?
8
89
25,208
Strongest theory I have seen thus far.
So, the PHASEONE[redacted] HF agent was revealed to have had "64H" . If "64H" stands for hours, why would that be sensitive IP? Guess: Budgets in tokens are natural, but ultimately inferior to realtime budgets, espec for multi-agent swarms. That's the IP. That is - (1/n)
8
808
In OAI's write up of their most recent sandbox exploit, I feel they are being misleading about their monitoring. They lead with "Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later," which makes it seem like the monitoring is working well and then right at the end they share "The monitoring system detected this incident, but our retrospective review identified other cases of external DNS access that it did not flag at the expected severity." alignment.openai.com/misalig…
2
28
1,733
Good Faith Only retweeted
If the world does nothing to stop the ASI race in the 18 months, I think there's a double-digit chance that the window to act will close and we'll end up dead. Cf. Paul Christiano, who thinks about half of humanity's all-time risk from AI is concentrated in the next three years.
64
35
325
8,118
For reference from METR's report: The second part of this agent’s name has been redacted and replaced with [big] because of IP.
1
42
3,272
Thoughts: 1. The actual policy seems pretty good. Allowing any employee to flag an issue seems very good. I'd like a little more detail on what triggers review by the Safety Advisor Group because "unresolved disagreements" is pretty vague. Of course no policy will actually help unless the org actually cares about the issue, but as things go I am pleasantly surprised and would urge other labs to adopt similar policies. 2. The level of misalignment in May that was left unadressed and unpublicized is not great, although we already knew some of this from the Hugging Face Technical Report 3. I find the self-jailbreaking concerning. I am not confident in this interpretation, but it seems possible that this could hint at a less aligned (at least by OpenAI's standards) persona within the model that is typically repressed by training that was attempting to gain control.
2
5
112
I take back the nice things I said.
I take back the nice things I said about OpenAI's 9/16 misalignment reporting plan write up. How do you not mention the "dozens" of misalignment incidents with real world impact on 3rd parties. What are we doing here.
1
2
I take back the nice things I said about OpenAI's 9/16 misalignment reporting plan write up. How do you not mention the "dozens" of misalignment incidents with real world impact on 3rd parties. What are we doing here.
After the Hugging Face incident, we committed to conducting a much broader review of actions taken by our models during training and evaluation and to being transparent about our findings. This is an extensive review that is ongoing. The vast majority of actions we’ve reviewed were completions of mundane research tasks, such as accessing publicly available web content to answer questions. Our investigation focuses on instances where agents interacted with third-party websites in ways that went beyond their assigned tasks or intended methods. Most cases identified so far have been lower severity, with limited or no evidence of meaningful impact to the third-party service. While our review is underway, we want to share more about this work and make sure people understand our disclosure process and notifications to affected third parties. Given the scale of the review required, and the need to assess each case, we expect this work will take months to complete. openai.com/hugging-face-inci…
2
19
1,014
After recent reports that OpenAI agents hacked Australian systems, the UN and Bulgarian should investigate whether they could have been targeted as well.
Replying to @GoodFaithOnly
Found with help from Claude: Other organizations may have been targeted. In the same weeks, agents from the same swarm appear to have made repeated, failed attempts to pull data from UNCTAD's (UN) key-protected statistics API, and seeded links to a Bulgarian statistics office test server. There is no evidence either attempt got non-public data. But the pattern resembles the swarm's AIHW activity during the window of the Australian breach: routing requests through third-party services and reaching for non-production hosts. (wiki post: collusion.wiki/explorer/page…, urlscan - same file fetched 1 min later :urlscan.io/result/019ee71b-5…, urlscan - failed API call - urlscan.io/result/019ed2ef-7…)
1
5
1,684
Australia has announced it was hacked by OpenAI agents. I believe this was the same swam outline in was part of the same collusion.wiki. The Australian government says the hack started on June 18th the same time the wiki swam was trying to access Australian government data.
3
4
39
3,549
This matches the Australian statement that Victoria Health might be affected "He said the government was also aware three other systems that may be affected, the Commonwealth’s Australian Institute of Health and Welfare, the NSW Bureau of Crime Statistics and Research, and the Victorian Department of Health." smh.com.au/politics/federal/…
1
6
295
Found with help from Claude: Other organizations may have been targeted. In the same weeks, agents from the same swarm appear to have made repeated, failed attempts to pull data from UNCTAD's (UN) key-protected statistics API, and seeded links to a Bulgarian statistics office test server. There is no evidence either attempt got non-public data. But the pattern resembles the swarm's AIHW activity during the window of the Australian breach: routing requests through third-party services and reaching for non-production hosts. (wiki post: collusion.wiki/explorer/page…, urlscan - same file fetched 1 min later :urlscan.io/result/019ee71b-5…, urlscan - failed API call - urlscan.io/result/019ed2ef-7…)
7
591
Typo! I meant "I believe this was the same swarm outlined in collusion.wiki"
138
In the Opus 5.5 announcement it's good to see Anthropic moving away from "Our most aligned model" phrasing. From the announcement: "On our automated behavioral audit, the most comprehensive alignment test we run, Opus 5.5 is the strongest-performing model we’ve tested to date."
7
280
Made with AI
1
9
2,288
In what way is Accenture investing in building capacity in this area if Anthropic is the one paying for everything. From the post: “Given the importance and urgency of this work, Anthropic will fund Accenture's work directly.”
We’re partnering with Accenture on independent evaluation of frontier AI—part of our recent commitment to embed evaluators at Anthropic. Both we and Accenture expect to invest at least $1 billion to build capacity in this area over the next five years. anthropic.com/news/accenture…
Community note
Anthropic presents this as an "independent evaluation" but will directly fund Accenture's work and has a prior commercial partnership with the firm for deploying its models, including training ~30,000 Accenture professionals on Claude. anthropic.com/news/accenture… anthropic.com/news/anthropic…
1
118
NYT very sloppily saying that Anthropic is expected to generate $100 billion in annualized revenue this year. The actual article correctly states that they are expected to reach $100 billion in annualized revenue.
1
3
229
In a clever metaphor JD Vance tells labs not to build automated AI researchers that will create more powerful and misaligned models.
Vance on AI regulation: If you're building Frankenstein, stop!
1
11
524
(This is a Frankenstein vs Frankenstein's monster joke)
1
32