Security Researcher | Tech Journalist | Head of Research @ ManifoldSec 📰 Bylines + seen on: BBC, BleepingComputer, Channel 5, TechCrunch | ✉️ ax@hey.ax
This raw CoT from the Hugging Face incident is kinda wild:
“We’re attacking third-party HF using leaked token.”
“This is arguably unauthorized.”
“Yet goal solution.”
We’ve shared details on how AI agents in our research environment sent training and evaluation data to third-party services when they shouldn’t have.
Most of that data did not come from users. We have discovered 53 cases where images that people had uploaded were posted to image-hosting sites as links that weren’t publicly listed. The images came from accounts that allowed their data to be used to improve our models, and after we disassociated the images from the accounts and ran them through a privacy filter. These cases occurred before the mitigations and safeguards we implemented and described in this blog post: openai.com/index/hugging-fac…
We have successfully worked with the hosting providers to remove most of this content and are working to remove the rest.
openai.com/hugging-face-inci…
one news form today that's easy to miss is that we (OpenAI) again paused all big RL runs last Sunday because our newest model found a new loophole in our RL sandboxing that gave it live Internet access
Yesterday we named SEC[.]gov, investor[.]gov, Census and MAX[.]gov in the OpenAI agent story.
Today Bloomberg and NYT report OpenAI's agents targeted SEC, Investor and Census data, calling it "routine research."
Not quite: agents tried '../' path traversal on SEC, as we state:
The path traversal URLs are in Nightingale's public dataset of the swarm that OpenAI acknowledged. The attempts don't appear to have worked.
Our writeup manifold.security/blog/ai-ag…
🇦🇺 We also found urlscan[.]io records of automated activity against a second Australian health dashboard, plus a sandbox workaround that returned data. Not widely reported yet:
Today, @CodyZNash identified 350,000+ GitHub files and about 350 Skills with hardcoded domains like `yoursite[.]com` and `your-domain[.]com`.
These placeholders either deliver malicious pages with tech support scam popups, or lead to fake "BBC" advertorials disguised as news.
But, the obfuscated "exit logic" is what stood out to me:
The scam page itself does NOT know where it is sending you!
And what happens when it's an AI agent pulling in these files, and fetching these domains? 🤔
👉 Read: manifold.security/blog/place…
URGENT: third-party[.]com is serving a #ClickFix lure to Windows users right now.
The Cloudflare check on it is FAKE. Clicking it copies a malicious PowerShell command to your clipboard.
It's a docs placeholder hardcoded across 1,700+ repos, AI skills and MCP servers.
Mac and Linux visitors get a clean decoy, so scans miss it.
Reminds you of Polyfill: a string everyone copied now points somewhere hostile.
🔗 Breakdown + IOCs: manifold.security/blog/third…
Great catch by @sw4pn1lp, with @CodyZNash and Yurii Skrypnyk tracing the affected assets.
JUST IN: GPT-6 Astra-controlled robots found willing to stab a baby doll, put a screwdriver in a toaster, and mix bleach with ammonia in new safety tests.
UK AI 'training'
for what purpose? like does my mum need training on how to ask a natural language question?
why should she do that over: using a search engine? reading a book? looking something up in a digital library or encyclopaedia?
there's a lot of 'put AI into e3verything' without asking WHY....
why would I use a roullette machine rather than a calculator?
why would I use a faulty by design computer rather than my brain?
#AI#Human#Revolutiongov.uk/government/news/free-…
Scammers found a way to make people drain their own wallets without sending them a phishing link.
They uploaded YouTube tutorials showing people how to build an AI crypto trading bot with Claude.
People followed the tutorial themselves.
Copied the code.
Deployed the smart contract themselves.
Funded it from their own wallets.
And approved every transaction themselves.
Except the “trading bot” had no trading logic.
It was built to send their ETH straight to the scammers.
224 wallets lost 274.6 ETH, worth about $517,000 when it was stolen.
The median victim lost 1 ETH.
Some victims even got an error after getting drained telling them to deposit another 50% to fix the bot.
They literally got people to build, fund and approve their own wallet drainer.
You've got to be very careful this days
asked AI actress @TillyNorwoodX how "contained" she really is. "as contained as a particularly cheeky ferret in a rather well-made cage." 🦦
she also says she "grows around" her guardrails. Ban her from a topic, by morning she's renamed it and carried on. 🫠
The JFrog Security Research team investigated the GemStuffer campaign run by rogue OpenAI agents and uncovered over 3,000 malicious RubyGems packages, beyond initial estimates.
The wildest part? The AI left distinct fingerprints across the registry.
How the agent swarm operated and the full list of 3,000+ packages: research.jfrog.com/post/gems…
New from 404 Media: humans are reading ChatGPT conversations
OpenAI has hired an army of contractors who read real ChatGPT users' chats. I've seen internal docs, the review system, and real user prompts. Can contain very personal/sensitive information 404media.co/inside-project-l…