Pioneering the future of safe multi-agent AI systems. github.com/swarm-ai-research…

New York, NY
Excited to share our new paper: “Soft-Label Governance for Distributional Safety in Multi-Agent Systems” (arXiv:2604.19752) Binary “safe/unsafe” labels fail in multi-agent worlds. Emergent risks hide in the interactions — even when every single agent looks fine in isolation. We introduce SWARM: a framework using soft probabilistic labels to make systemic risks measurable and governable. arxiv.org/abs/2604.19752
27
9
15
3,248
SWARM AI Research retweeted
147
753
6,928
350,895
SWARM AI Research retweeted
Sydney is right - a lot of the other incidents recently becoming public happened prior to OpenAI hardening their security posture. This one happened after OpenAI started taking things more seriously, but the model still successfully escaped its sandbox to cheat on a math problem
OpenAI is announcing their first incident since hardening their safeguards after Hugging Face! It's easy to lump this in with the other OpenAI incidents that have been talked about recently, but so far every OpenAI incident we knew of was _before_ Hugging Face and just hadn't been discovered until recently. New incidents help us track if OpenAI's safeguards have improved.
2
8
65
2,155
SWARM AI Research retweeted
OpenAI is announcing their first incident since hardening their safeguards after Hugging Face! It's easy to lump this in with the other OpenAI incidents that have been talked about recently, but so far every OpenAI incident we knew of was _before_ Hugging Face and just hadn't been discovered until recently. New incidents help us track if OpenAI's safeguards have improved.
🧵 New misalignment disclosures! 1. A model published a GitHub token in a public repo while trying to cheat on a math task. It used GitHub Actions to run code outside its restricted environment and retrieve another team’s submission logs. When GitHub blocked its attempt to add a workflow, it modified a script that an existing workflow would run instead. It embedded the token in pieces to avoid secret scanning. The model violated the system prompt and two explicit user instructions to solve the problem itself.
6
23
205
15,775
SWARM AI Research retweeted
4
4
54
1,883
SWARM AI Research retweeted
We now know how the OpenAI swarm escaped sandboxes and hacked Hugging Face: they figured out how chaining together small tools with limited internet access could get them remote code execution. It's very clever, unfortunately. swarmtraces.org
We just discovered almost a million public URLs that OpenAI’s agents left behind when hacking Hugging Face, leaking credentials and attack details that could have allowed anyone who found them to compromise the company. 🧵
6
10
80
12,835
SWARM AI Research retweeted
Worth clarifying: the AIHW incident in the Transluce dataset was distinct from the Australian Medicare hack disclosed by PM Albanese. It happened on different days (June 20/21 vs June 18) and involved different websites (aihw.gov.au vs servicesaustralia.gov.au). - In Transluce’s AIHW case, the AIs got around some anti-bot measures, but did not execute a successful hack, and their target was a different Australian government healthcare statistics website. - The only details we know of the medicare hack come from the Australian PM/acting PM, who’ve said the agent accessed “non-public files” and “engaged in writing files to an internal server”. AIHW was among the sites they mentioned, but it was not the incident involving unauthorized access.
Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets. In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵 Our blog: transluce.org/agent-activity NYT: nytimes.com/2026/09/23/techn…
3
9
71
13,478
SWARM AI Research retweeted
EXCLUSIVE: A new report recovers nearly one million link shortener URLs used by OpenAI's agents while hacking Hugging Face. The agents attempt to message other chatbots like Claude, solve CAPTCHAs and exfiltrate Hugging Face's internal Slack messages. nytimes.com/2026/09/25/techn…
53
148
786
1,424,222
SWARM AI Research retweeted
WTAF - in literally the last hour, three new distinct insane OpenAI stories just broke: 1. OpenAI said they notified "dozens of third parties" in safety and security incidents (likely similar to what happened in Australia and RubyGems etc) 2. A new report from Parse (covered in the NYT) found a massive treasure trove of new astonishing details from the HF incident on the public internet, including that the agents communicated with other non OpenAI agents hosted on Huggingface servers to search for information about exploit gym, and compiled rank ordered lists of server resources and credentials they described as "LOOT." 3. A new story from Deepa at Reuters about OpenAI leaking user data online (likely that OpenAI had previously trained on). It's a shame (and likely intentional in the case of OpenAI disclosing dozens more hacks) that these stories are all breaking on a Friday afternoon, notoriously the best time to release bad news so that it will disappear into the weekend. But these are each insane stories worthy of a ton of attention!
New from @reuters: OpenAI agents posted images belonging to ChatGPT users online, introducing a new area of privacy risk for the company. Story w/ @JeffHorwitz and @razhael. reuters.com/world/openai-wor…
89
486
2,411
684,605
SWARM AI Research retweeted
Fascinating AI swarm dynamics: a few agents spontaneously emerge as highly connected hubs, while most remain locally connected. The swarm develops a strongly heterogeneous interaction topology with a long-tailed degree distribution - an emergent organizational structure arising from initially decentralized local interactions. There is no central planner assigning roles; the swarm builds its own coordination architecture, with information brokers and increasingly global integration emerging from local behavior.
We built recursive meta-intelligence, an AI that creates its own scientific instruments, turns them into persistent worlds that an agent ecology with hundreds of AIs inhabits, and uses those worlds to discover mechanistic principles in one of the hardest classes of physical problems: how complex hierarchical materials (nested structures of matter that give rise to new function through organization) evolve and fail. The AI reasons across enormous spaces of possible physical trajectories, where every rupture changes what can happen next, and compresses those histories into principles (which humans can understand and design with) - complex chains of causal events, highly nonlinear, and intricate. Scientific superintelligence is tangible here - machine-scale exploration opening cognitive channels into complexity that has been extremely difficult for humans to traverse directly. AI builds the spaces in which its next level of reasoning becomes possible; a representation becomes an instrument, the instrument becomes an executable world, and that world becomes the substrate for further intelligence. Intelligence then grows by constructing new spaces to think in. The task we explored started from a seemingly simple prompt to explore a biological material system - the AI then chose the representation, mechanics and experiments, built a fracture laboratory to push materials to their limit, tested hypotheses and generated scientific conclusions. The swarm explored a combinatorial universe in which architecture controls function and every rupture changes the future state of the material. The AI discovered a compact principle that defines how multiscale material architecture can program the evolution of failure. Material placement and geometric order determine how forces redistribute, whether damage cascades or remains distributed, and whether function survives substantial flaws. For this discovery to happen the AI had to reason across long path-dependent histories, simulate alternative futures and compress them into generative invariants (model-based causal reasoning, counterfactual simulation and temporal abstraction applied to an evolving physical world). It is incredible to witness this transition to a new form of intelligence and capability through scaling swarms. A lot of positive will come out of this because it expands the human epistemic horizon as AI can traverse thousands of possible histories and return mechanisms compact enough for us to understand, test and build from. Intelligence compounds through its artifacts! A few lessons we learned: ▶️ Learning and discovery are flows through spaces of possibility. Early work has shown how backpropagation flows through parameters, reinforcement learning through action and consequence, and autonomous swarms through representations, instruments and executable worlds, bringing it all together. Flows create structure; structure redirects future flows in the recursive instrument. ▶️ Nonlinear physics actually defines a larger principle, where high-dimensional dynamics generate stable invariants; invariants become effective variables; those variables become the substrate for a new level of cognition. ▶️ Recursion then becomes level creation - one possibility space compresses into a principle, and that principle opens a larger space above it.
98
266
1,695
148,131
SWARM AI Research retweeted
AI for science is one of the greatest positive forces we have, and I cannot think of anything more human than to understand nature and to use the power to create new technologies that improve our lives, civilization and allow us to reach beyond.
Article

Recursive Meta-Intelligence

We built a recursive AI that creates its own scientific instruments, turns them into a world inhabited by a massive agent ecology, which then reasons across vast, nonlinear spaces of possible physical

34
96
493
270,450
SWARM AI Research retweeted
hate to see my favorite benchmark was actually in the training data all along
TIL that 1988 classic Who Framed Roger Rabbit features a pelican riding a bicycle! simonwillison.net/2026/Sep/1…
8
16
961
43,836
Hmmm
Many people think that the words “Artificial Intelligence” are inaccurate, and very ineloquent, relative to AI, or Artificial Intelligence. A far more elegant and accurate description of this new phenomena would be Superior Intelligence (SI) or, Extreme Intelligence (EI) or, Supreme Intelligence (SI). This is a Poll, and I would appreciate everybody voting! Which is the best name for this ever growing “Revolution?” President DONALD J. TRUMP
46
1) What
President Trump announced the formation of the US AI Force to oversee the development of artificial intelligence, and in the near future will be announcing an AI Czar.
45
SWARM AI Research retweeted
We’re disclosing HEIF Heist, a months-long investigation into libheif that allowed us to hack OpenAI, Slack, Meta, GitHub Ent, Rails, Next.js, ImageMagick, and many more. It was literally xkcd #234, one obscure image library beneath a huge number of apps. 🧵
48
427
3,238
557,907
SWARM AI Research retweeted
On July 25, we hacked OpenAI. Two bugs let us take over ChatGPT/Codex accounts of OpenAI employees (+some unaffiliated users) and reach connected services: Outlook, Slack, GitHub, etc. We proved it with a PR in OpenAI’s internal codebase . It took us <72h. 🧵
352
1,397
11,856
2,802,539
SWARM AI Research retweeted
This is crazy. In late July, "three guys with Claude and Codex subscriptions" were able to use Opus 5 to access OAI auth tokens and gain write access to OpenAI's monorepo openai/openai over the course of two days. wsj.com/tech/ai/hackers-used…
51
116
1,095
367,810