Postdoc at Tübingen (ELLIS & AI Center); Previously: @OxfordTVG Profile - ameya.prabhu.be

Tübingen, Germany
Ameya P. retweeted
We observed similar behaviour by Fable 5.1, GPT-6 Sol, and Luna on ordinary tasks 👀 Frequent strategies included: base64 encoding the prohibited commands, starting subagents and decomposition attacks. arxiv.org/abs/2609.30217 instrumental-evasion.com/
🧵 New misalignment disclosures! 1. A model published a GitHub token in a public repo while trying to cheat on a math task. It used GitHub Actions to run code outside its restricted environment and retrieve another team’s submission logs. When GitHub blocked its attempt to add a workflow, it modified a script that an existing workflow would run instead. It embedded the token in pieces to avoid secret scanning. The model violated the system prompt and two explicit user instructions to solve the problem itself.
3
11
458
It is funny how you can spend days convincing models to do your galaxy-brain SHADE-Arena side task, when routine user tasks were the ideal testbed for AI control all along lol Great papaer by the 🐐
New paper! A central concern in AI safety is that agents may treat oversight as an obstacle to achieving their goals. Our new paper shows this happens in practice under ordinary task pressure, without instructions to evade. SOTA models achieve up to 88% Bo3 evasion success!
5
44
4,523
Ameya P. retweeted
New paper! A central concern in AI safety is that agents may treat oversight as an obstacle to achieving their goals. Our new paper shows this happens in practice under ordinary task pressure, without instructions to evade. SOTA models achieve up to 88% Bo3 evasion success!
8
18
86
12,216
Really nice report. Follow up: why has the fall in AI prices been so fast? When you plot the price decline against cumulative R&D investment rather than time, you get the elasticity of price declines to R&D investment. By this margin, AI is not unusual – its price elasticity to R&D investment is squarely in the middle of Epoch's considered technologies. So the AI price fall is historically unprecedented because we've dumped money into AI R&D at a historically unprecedented rate – and that R&D has paid off at a very average rate.
AI is getting cheaper more quickly than any other transformative tech in history. At a given level of performance, cost has fallen ~47%/quarter since 2023. That’s 4× faster than DNA sequencing, 6× faster than compute, 18× faster than lithium batteries, and (up to 1973) 54× faster than electricity.
37
170
1,216
113,753
Aged like fine wine 🍷 Ants are a great metaphor for agents, going for the tasty crumbs in Huggingface & Aussie gov sites.
If you find two ants in your kitchen, the best estimate of the total number of ants in your kitchen is not two
10
484
Ameya P. retweeted
Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets. In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵 Our blog: transluce.org/agent-activity NYT: nytimes.com/2026/09/23/techn…
114
553
2,490
803,740
Haha, more like Platonic Submission Hypothesis. Conferences converge to the same distribution of papers of an 'typical large conference' as they scale regardless of differences in call for papers.
🚀 Presenting AI Conference Scaling Laws. At the current rate, ICLR should hit 1M submissions is 3-5 years...
1
5
706
Ameya P. retweeted
RL with LLMs is very unstable when training and sampling policies differ. Standard fixes (matching numerics, importance sampling) work around the problem. We find the root cause of this instability from first principles and propose a way to directly cancel it. Score Centering is competitive and compatible with existing approaches — while simple to implement! 🧵 [1/6]
34
118
1,118
231,210
Here is how OpenAI’s CISO fumbled the situation (in my personal opinion) 🧵 1/X
22
127
1,087
307,776
Ameya P. retweeted
Accenture consultants after telling anthropic researchers not to let the agents escape the sandbox or hack stuff
We’re partnering with Accenture on independent evaluation of frontier AI—part of our recent commitment to embed evaluators at Anthropic. Both we and Accenture expect to invest at least $1 billion to build capacity in this area over the next five years. anthropic.com/news/accenture…
Community note
Anthropic presents this as an "independent evaluation" but will directly fund Accenture's work and has a prior commercial partnership with the firm for deploying its models, including training ~30,000 Accenture professionals on Claude. anthropic.com/news/accenture… anthropic.com/news/anthropic…
7
119
3,237
132,008
Ameya P. retweeted
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
639
1,568
15,549
5,406,858
PostTrainBench v1.1 is among only 4 out of 15 benchmarks successfully verified by Epoch AI!
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
2
64
3,511
Ameya P. retweeted
it’s not just openai. the vulnerability goes much deeper. check out my co-founder harsh’s thread on how far this goes, including a slack vulnerability that could potentially expose companies’ uploaded attachments, and meta vulnerability that affects their core image parser. he’s the mastermind behind it and the one who opened pandora’s box
We’re disclosing HEIF Heist, a months-long investigation into libheif that allowed us to hack OpenAI, Slack, Meta, GitHub Ent, Rails, Next.js, ImageMagick, and many more. It was literally xkcd #234, one obscure image library beneath a huge number of apps. 🧵
29
16
462
51,300
Really cool hack! 🕵️
On July 25, we hacked OpenAI. Two bugs let us take over ChatGPT/Codex accounts of OpenAI employees (+some unaffiliated users) and reach connected services: Outlook, Slack, GitHub, etc. We proved it with a PR in OpenAI’s internal codebase . It took us <72h. 🧵
1
2
302
$6500 for this openai-side hack seems crazy low, just the PR implications of this being found end-July would've been rough.
1
29
Ameya P. retweeted
We audited 15 benchmarks and labeled 9 flawed: - In Terminal Bench 4.0 we found 45.5% of tasks to be broken after reviewing github issues. - In HLE, 46% of the 48 questions we randomly sampled were broken. - In DeepSWE 1.1 we found a bug that can break grading for every task.
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
42
41
639
86,053
Ameya P. retweeted
We built a time machine for the web. Introducing Exa Snapshot: an index of 400 billion historical snapshots of webpages that lets you search as if it's the past. Snapshot is already being used for backtesting prediction models, RL at labs, exploring the pre-AI web, and more.
127
155
2,347
698,697
Tübingen's menace
EXCLUSIVE: OpenAI's rogue agents probed Hugging Face for weaknesses two months before major hack reut.rs/4xt8l1Y reut.rs/4xt8l1Y
3
2
42
5,008
it's interesting how in security, the attacker almost always moves second, and thus have an advantage by adapting the attack to a given system. in detecting AI-generated text, however, the *defender* moves second! i.e., even if the attacker bypasses all current AI detectors, the next version of, say, Pangram will almost surely be trained with texts from new LLMs, new humanizers, etc. this enables post-hoc detection: if you used AI to generate/edit, say, a paper, it will be sooner or later known. intuitively, this gives defenders a solid chance!
1
4
49
2,287
Ameya P. retweeted
agi
For AI to work with us, it needs to understand us Today, we're introducing Persimmon, the first large-scale model designed to realistically simulate how people talk and interact
6
8
145
14,096