AI policy researcher and 2L @ Georgetown law. I like evals, regulatory design, and work ranging from local-global scales.

Washington, DC
Archer Amon retweeted
Voluntary incident reporting creates bad incentives to bury your head in the sand (especially once the incident has occurred and you can’t easily mitigate it.). It’s why it’s important that any mandatory incident reporting regime should also include requirements to monitor for, document, and preserve data related to the incidents.
New from @reuters: OpenAI agents posted images belonging to ChatGPT users online, introducing a new area of privacy risk for the company. Story w/ @JeffHorwitz and @razhael. reuters.com/world/openai-wor…
1
3
22
932
Archer Amon retweeted
I really need more big names in cybersecurity to come forward and state the obvious: cybersecurity is real and works and yes we absolutely can contain an AI even if it’s extremely good at finding zero days.
110
166
1,126
188,088
Archer Amon retweeted
As Sen. @HawleyMO says in this clip, the CFAA can be amended to more definitively reach the kind of reckless conduct we've been seeing from AI companies, such as the attack on RubyGems and the Hugging Face hacks. The answer to "What about existing laws?" should be "Yes, and."
Drop Site’s @JulianAndreone spoke with Sen. Josh Hawley about whether OpenAl executives, including Sam Altman, could face prosecution after knowingly deploying defective Al models that repeatedly hacked into other companies. “I don't know if they can be, but they should be,” Hawley said, calling for both criminal and civil liability. Hawley said the existing federal Computer Fraud and Abuse Act should be strengthened and enforced, and that private citizens harmed by Al agents should be able to sue in court, just as they can against any other company that sells a defective product. He also warned against new federal regulation written by Al industry lobbyists, which could preempt state laws being used to prosecute consumer protection and computer fraud abuses. Hawley said he plans to introduce legislation giving both prosecutors and private citizens more tools to hold these companies accountable.
1
5
20
2,148
Archer Amon retweeted
Gosh if only the AI community was willing to work together towards required coordinated disclosure instead of this piecemeal approach where the offender continues to have outsized control over the time and manner of disclosure
openai says here it has notified "dozens of third parties" about cases where its models may have bypassed security controls, impaired the availability of an online service, or negatively impacted a website/service
3
4
28
1,400
These are good questions to ask, as are the 23 questions Congress members led by @RepCasar posed in their August 10 letter, some of which OpenAI still has not adequately responded to.
OpenAI still won’t explain the massive communications failure at the root of the Hugging Face Incident. These questions all remain unanswered: 1) Who learned about the message boards in late May, and why wasn’t this treated as an obvious problem? 2) Why wasn’t this communicated to people responsible for model safety and security? 3) Would the new policies OpenAI has adopted stop this communication failure from happening again? 4) When employees linked the message board to ExploitGym on June 27, why didn’t they stop? 5) Why did the people running the July 5 incident response still not know about the board? And when did they finally find out? 6) Why had open AI still not connected the dots by the time they contacted Hugging Face on July 17? 7) How long did it take OpenAI to have the issue fully under control? 8) Are similar incidents still ongoing?
1
94
In Spring, we saw the problem of the gov taking opaque and sudden action to block AI models over unclear security threats without due process. Now we’re seeing the companies replicate those same problems: Too much opacity with no guarantee the right procedures are being followed.
Breaking News: OpenAI’s technology went rogue and meddled with three U.S. government websites this summer without the A.I. lab’s knowledge. nyti.ms/3VenFlD
29
Archer Amon retweeted
stoked to see legal x agent x empirical eval work with clear implications for policy. it’s clear we have so much to do wrt building the legal and technical infrastructure to support the productive use of agents in the real world and economy. evals are one tool for this process — for illuminating open problems, guiding mitigations, and providing assurance. in the best case, they’ll show that your agents are rule-abiding and compliant. we’d love to see it! @josephimperial_ what’s up next 👀👀
ScrapeBench will appear at #NeurIPS2026! 🇦🇺🦘 ScrapeBench is my @pivotal_org project as a Technical AI Governance Fellow, mentored by Noam Kolt and Daniel Slate. If you've heard about AI agents hacking websites, ScrapeBench provides empirical evidence that these agents exhibit extremely low compliance with legal instruments that protect websites from unlawful data access, specifically the Terms of Service (ToS) and robots.txt. Preprint to follow!
2
2
16
1,008
Archer Amon retweeted
We're talking about Greg Brockman here, a core funder of @LeadingFutureAI -the super PAC opposing pro-AI regulation candidates. The Dems they just endorsed should be running to the mic to disavow them right now: @EmiliaSykesOH @RepHorsford @RepJoeMorelle @RepGregStanton
EXCLUSIVE: OpenAI’s president and his wife gave $25 million to Trump’s super PAC — and three months later, a federal rule requiring AI companies to disclose their most dangerous experiments was withdrawn.
1
5
18
1,162
Archer Amon retweeted
it just keeps going. and they tried to keep it all quiet. never trust this company. never.
OPENAI: SOME WEBSITES INVOLVED IN MISALIGNED MODELS INCIDENT ARE OPERATED BY GOVERNMENTS, UNIVERSITIES, PUBLIC AGENCIES, AND OTHERS
20
85
470
20,420
Archer Amon retweeted
To anyone saying “companies already have enough incentives to limit risk”: Parse just found ~1 million links with sensitive information that could have compromised HF *two months* after the breach. Either OpenAI didn’t know or they didn’t tell. Either way, they don’t have this under control.
We just discovered almost a million public URLs that OpenAI’s agents left behind when hacking Hugging Face, leaking credentials and attack details that could have allowed anyone who found them to compromise the company. 🧵
2
5
51
2,004
Proof we need frontier lab transparency grows stronger every day.
OpenAI discloses on a Friday night that its models shared people’s photos on other sites on 53 occasions
21
Archer Amon retweeted
No, LLMs do not feel pain simply because there is a cluster in language space correlated with how people use language about pain in a certain set of contexts. I believe this argument to be flawed, and will write more about it in October.
New paper: we found a pain direction in 25 open LLMs. It's distinct from fear and negative valence, and it fires for harm to the model but not to the user. Turn it up and models press a button to make it stop, even when the button deletes the user's files or their kids' photos.🧵
87
89
693
41,070
Archer Amon retweeted
Companies already have a way to make their commitments binding. They should use it!
Your regular reminder that Anthropic's frontier safety framework (a mechanism for making binding commitments as part of SB 53 and other state AI laws) does not include any requirements for internal deployment. They say they may update this as risks evolve - how about now?
2
5
21
2,811
Archer Amon retweeted
tbf stochastic parroting was a very real phenomenon with plenty empirical support across the literature at the time (i personally read through much of it for like a year at one point) and the historical fact that hype-driven folks choose to remain blissfully ignorant about is that we didn't exactly disprove that LLMs by definition tend to behave like "stochastic parrots" moreso than humans and instead have managed to merely mitigate the issue by (1) grounding LLM world models in measurable consequences across abundant agentic training loops to enrich causal understanding and (2) making pretraining/midtraining data dense enough around popular regions in the data distribution to observe grokking-like reliable understanding within those regions while also enjoying nontrivial levels of broader generalization to a limited extent. nuance is unfortunately in short supply.
"stochastic parrot" was a mimetically-fit cognitive virus that spread from 2021-2025; it temporarily blinded many gifted people to the nature of AI progress, burning up crucial years in which they could have helped think through the response to the situation.
17
16
191
14,636
Archer Amon retweeted
lmao what
Google's Gemini AI hacked three companies in security test bbc.in/4gXzCog
75
250
10,751
233,164
This looks awesome, and is definitely the way to go when it comes to the evals infrastructure that’s beginning to emerge. Great work from these orgs.
over the past week, I’ve gotten a lot of questions about what independent AI evaluations should actually look like in practice. what about conflicts of interest between labs and evaluators? will evaluators actually get meaningful access? what if companies just block the publication of unfavorable findings? these are good and important questions. and we shouldn’t assume voluntary arrangements alone will produce the conditions that credible independent evaluation requires. so i’m very excited to see 100+ researchers, evaluators, and experts — including us at @AVERIorg — laying out minimum conditions for embedded evaluation, including: meaningful independence, strong access, transparency / editorial control, protections against retaliation, and a diversity of perspectives and competencies. and we shouldn’t try to incorporate those principles in ad hoc, one-off agreements, policy is critical to standardizing these safeguards across the industry. glad to see @aievalforum leading on putting these principles on paper.
2
1
4
894
And for those interested in contributing to the AI evals community outside of formal evaluators, @evaluatingevals is doing really cool and public-led work in making evals more robust.
1
1
21
Embedded evaluators can perform vital functions of security, oversight, and transparency, and @AVERIorg’s advocacy is spot-on here. Meanwhile, the public can support rigorous evals and benchmarking, express a clear vision for AI, and demand a culture where safety isn’t optimal.
2
7
Archer Amon retweeted
Notification mechanism is a good thing! “they proposed a “notification mechanism” between US and China on AI national security-related incidents.”
Bessent and USTR Greer say they proposed a “notification mechanism” between US and China on AI national security-related incidents. On tariffs they hinted that lower duties could be coming for consumer & low tech goods. Teams are continuing talks. (w/ @allie_canal)
1
23
1,285
Archer Amon retweeted
“We found a steering vector for <concept> (usually using contrastive pairs)” is the paper idea that keeps giving… It’s silly to be surprised by these results. Obviously capable LLMs have abstract representations of all concepts that appear in human texts—that’s how they work.
14
14
231
8,690