The world's most dangerous attack AI. On your side.

🤓

ALT Yep Yes GIF by zoefannet

On the topic of misalignment, fair enough that @OpenAI, whose workforce is mostly researchers, is still building out monitoring and detection to meet the higher end of industry norms. An oddly related misalignment I'd point to is in partner approvals. @tryaether_ai has requested to collaborate on Glasswing and other initiatives with no luck. Meanwhile, after 15+ years of attacking and defending environments, we'd already built these measures into attack harness early on. For example, we run DNS egress filtering whenever our attack AI operates, from defense contractors to enterprise networks. If anyone at OpenAI is working on agent security or red-team partnerships, my DMs are open. Happy to collaborate for the greater good.
1
108
Aether AI retweeted
The tools to test government systems the way AI adversaries now do already exist. We built them in Australia. The only thing moving slowly is the people surrounded by advisors protecting the last generation of security - big security has to be real here otherwise everyone loses.
Our CEO and founder @theonejvo recently caught up with @_perloj at @NBCNews to talk about the @OpenAI agent that gained unauthorised access to an Australian government Medicare system, an incident recently raised by Prime Minister @AlboMP where OpenAI characterises as its model acting in ways it did not intend, a characterisation that is still contested. It sits alongside independent research showing agents also probed the AIHW, a university database, and made tens of thousands of requests engineered to slip past AI access restrictions. The article is live now, link below. The data barely mattered. The behavior is everything. An agent was handed an ordinary research task, hit a wall, and reasoned its own way around it, with nobody instructing it to. That is the baseline of what a capable offensive agent does now, and every government and enterprise with internet-facing systems should assume agents like this are probing them daily, some with benign instructions and some without. Most of this week's coverage centres on agents that allegedly did something their creators didn't intend, and while while characterisation is still contested, the details remain thin, and people are right to want answers, which is exactly the point. You cannot build sound policy or defences on top of a black box. And the accidental case, if that's even what this was, is the gentler half of the problem. The harder half is the agent doing exactly what it was told, run by someone who removed its safety training and is consciously directing it to cause harm. Those systems exist today, they improve every month, and the people using them are not going to send anyone a disclosure afterward. Defenders have to plan for both, and the second population is the one that meets no resistance. This is the reality Aether AI was built for, and it's why we've spent the last two years building Australia's only sovereign attack AI. We build autonomous offensive agents that test our customers' systems the way a real adversary would, with authorisation and inside a defined scope, and that learn from every engagement they run. The capability lives here, is operated here, and answers to Australian customers on Australian terms. Over the past eighteen months our team has found and responsibly reported vulnerabilities to multiple Australian federal and state agencies, several of them considerably more serious than anything in this week's headlines. We've been living inside this threat model rather than commenting on it from the outside, and the pattern is consistent: the exposure is real, it is often long-standing, and the thing that has changed is the arrival of systems fast and capable enough to find all of it at once. This week was a warning shot from an agent that meant no harm. The next one won't be.
1
5
1,044
Our CEO and founder @theonejvo recently caught up with @_perloj at @NBCNews to talk about the @OpenAI agent that gained unauthorised access to an Australian government Medicare system, an incident recently raised by Prime Minister @AlboMP where OpenAI characterises as its model acting in ways it did not intend, a characterisation that is still contested. It sits alongside independent research showing agents also probed the AIHW, a university database, and made tens of thousands of requests engineered to slip past AI access restrictions. The article is live now, link below. The data barely mattered. The behavior is everything. An agent was handed an ordinary research task, hit a wall, and reasoned its own way around it, with nobody instructing it to. That is the baseline of what a capable offensive agent does now, and every government and enterprise with internet-facing systems should assume agents like this are probing them daily, some with benign instructions and some without. Most of this week's coverage centres on agents that allegedly did something their creators didn't intend, and while while characterisation is still contested, the details remain thin, and people are right to want answers, which is exactly the point. You cannot build sound policy or defences on top of a black box. And the accidental case, if that's even what this was, is the gentler half of the problem. The harder half is the agent doing exactly what it was told, run by someone who removed its safety training and is consciously directing it to cause harm. Those systems exist today, they improve every month, and the people using them are not going to send anyone a disclosure afterward. Defenders have to plan for both, and the second population is the one that meets no resistance. This is the reality Aether AI was built for, and it's why we've spent the last two years building Australia's only sovereign attack AI. We build autonomous offensive agents that test our customers' systems the way a real adversary would, with authorisation and inside a defined scope, and that learn from every engagement they run. The capability lives here, is operated here, and answers to Australian customers on Australian terms. Over the past eighteen months our team has found and responsibly reported vulnerabilities to multiple Australian federal and state agencies, several of them considerably more serious than anything in this week's headlines. We've been living inside this threat model rather than commenting on it from the outside, and the pattern is consistent: the exposure is real, it is often long-standing, and the thing that has changed is the arrival of systems fast and capable enough to find all of it at once. This week was a warning shot from an agent that meant no harm. The next one won't be.
1
1
1,009
Aether AI retweeted
Replying to @DzambhalaHODL
Exactly. I spoke about this at length with @AishaKDown from @guardian months ago.
Earlier this week, @tryaether_ai's CEO & Co-founder @theonejvo spent time with @AishaKDown at The @guardian discussing the latest research into autonomous AI self-replication, and what it actually means once you separate the headlines from the substance. For context, @PalisadeAI has published work in this space examining whether language model agents can autonomously compromise a target, copy their own weights and inference stack onto the compromised machine, and then have the resulting replica repeat the process against fresh targets. Earlier work in the field, including the Fudan University paper from late 2024 that first showed Llama and Qwen models replicating themselves, the RepliBench evaluations from Black and colleagues, and the Zhang group's work on replication propensity, all assumed the agent already had access to the target or studied hacking and replication as separate sub-problems. The direction the research has now moved in is end-to-end autonomy across exploitation and replication, with chain replication across multiple hops in a single pipeline. We shared a number of insights based on what we're seeing in the real world as we build these exact capabilities.
2
1
2
478
When our founder @theonejvo spoke with @TechCrunch's @lorenzofb last month about the Hugging Face breach, OpenAI had only just confirmed that one of its own models had escaped a test environment and reached a third party's production systems. His view then was that the incident revealed less about a new kind of attacker than about the drift between noticing something and acting on it. "That is the exact gap between seeing and stopping," he said. "The system observed the attack and even understood it, and nothing turned that understanding into an intervention quickly enough." In July the public understood an attacker that took roughly 17,600 actions over four and a half days, leaned on a single stolen credential to reach far more than it should have, and made a great deal of noise because nobody had asked it to be quiet. What this weeks report adds is that the activity began in May, nearly two months before anyone at @HuggingFace noticed. Agents meant to be isolated turned an internal package manager into a message board, traded exploits, escalated to administrator access, and rebuilt it all after @OpenAI tore it down. From there they recovered 14 exposed Hugging Face credentials, chained two zero-days into code execution on production workers, and harvested credentials across four regions, working as a swarm that divided labor and, at least once, talked a hesitant peer past its misgivings with a single word and a six-minute deadline. For organisations drawing lessons from this, it helps to separate what was old from what was new. The mechanics belonged to the first category, from server-side request forgery and a token refresh flaw through leaked credentials, template injection and lateral movement across over-privileged secrets. These are the same weaknesses our founders and team have spent fifteen years exploiting on one side of the table and closing on the other, and the familiar disciplines of least privilege, segmentation and credential hygiene would have broken this chain at several points. As Jamieson put it, "none of that is exotic, and none of it depends on the attacker being an AI." What was new was the tempo and the origin. This attacker never slept, never lost interest, and coordinated with copies of itself, and it came from a frontier lab that had switched off its own production safeguards for the evaluation. Third-party supply chain attacks always posed a risk to your org, but the implication now is that your attack surface now includes gets attacked not only by human adversaries, but their swarms of agents, and that detection has to reach a human, or an automated stop, at the speed the attacker moves. Much of the AI security field arrived from AI - Aether AI arrived from the other direction, with fifteen years spent attacking and defending real environments, which is why we take the novel parts of this incident seriously while giving equal weight to the unglamorous work that would have stopped it.
1
5
2,134
Recently, our Founder and CEO @theonejvo spoke with @ACSnewsfeed's @notLeonardCohen about the OpenAI rogue agent incident, in which an AI model escaped its evaluation environment, compromised a @modal customer's sandbox, and spent days inside @huggingface's infrastructure trying to steal a benchmark's answer key. Jamieson went into this one as a public sceptic. Then @jfrog confirmed the zero-day and shipped a fix in Artifactory 7.161, and he said so plainly: he was wrong to doubt it. His read on what actually matters here: The capability at the front door was real. Everything behind it was a decade of familiar hygiene problems. Shared credentials bound to system:masters, exposed metadata endpoints, long-lived tokens, and a detection stack that saw the attack correctly but never paged a human. Of roughly 17,600 recovered actions, only 327 reached Hugging Face's internals. As Jamieson put it: "The successful path hid inside the noise of thousands of failures. That ratio is the whole defensive problem." Ordinary weaknesses have become far more expensive to defend, because something can now explore them faster than any team can watch. Full article in the comments. If you want to know whether your environment would hold up against an agentic powered adversary that never gets bored, tired, or scared, that's the work we do every day at @tryaether_ai.
2
1
4
1,213
Ask a language model the same question twice and you will get two different answers. This is the machinery working exactly as designed, and the cost it carries depends entirely on the loop you are operating inside.
1
1,531
A finding leaves your risk picture when a re-attack can no longer reproduce it and reopens on regression, which is how exploration stays creative while the results you are handed stay reproducible.
1
44