Director of Nat Sec Policy @EncodeAction | Balancing safety & competitive innovation in AI, defense & emerging tech | 20-year vet | Dad joke enthusiast

Washington, DC
After 20 years in the Air Force, and a summer as a fellow at @GovAIOrg , I’m excited to get to work as the Director of NatSec Policy at @EncodeAction alongside the amazing team led by @SnehaRevanur! …and hoping to become half as prolific on here as @_NathanCalvin
5
1
89
8,822
Jason Baker retweeted
Dear lord give me strength to deal with the influx of absolute morons in my replies who seem to think the entire idea of AI x-risk was invented by Bernie Sanders 2 weeks ago
38
16
247
3,035
I’m not a “shut it all down” AI safety guy, but also…
Apropos of nothing, just reflecting upon this quote from Jensen Huang on The Ezra Klein Show two days ago:
4
Jason Baker retweeted
This is incredibly dangerous and will keep happening unless we step up and do something to put guardrails in place.

I’m concerned about what happens next. That’s why I’m taking action in the Senate and urging others to join me.
Breaking News: OpenAI’s technology went rogue and meddled with three U.S. government websites this summer without the A.I. lab’s knowledge. nyti.ms/3VenFlD
37
54
329
11,481
Jason Baker retweeted
OpenAI is announcing their first incident since hardening their safeguards after Hugging Face! It's easy to lump this in with the other OpenAI incidents that have been talked about recently, but so far every OpenAI incident we knew of was _before_ Hugging Face and just hadn't been discovered until recently. New incidents help us track if OpenAI's safeguards have improved.
🧵 New misalignment disclosures! 1. A model published a GitHub token in a public repo while trying to cheat on a math task. It used GitHub Actions to run code outside its restricted environment and retrieve another team’s submission logs. When GitHub blocked its attempt to add a workflow, it modified a script that an existing workflow would run instead. It embedded the token in pieces to avoid secret scanning. The model violated the system prompt and two explicit user instructions to solve the problem itself.
5
20
160
11,184
Jason Baker retweeted
The voters are now worrying more about AI overpowering human control, rather than worrying about job losses. A shift from just half a year ago.
AI UPDATE: Last March, voters told us they were more worried about job losses than a takeover by AI. This month, that flipped. Read the latest findings here: echeloninsights.substack.com…
17
40
163
20,985
Jason Baker retweeted
I’m struggling not to get lost in the absolute deluge of misalignment reports, but this set was a net new holy shit for me I hope we don’t end up too desensitized to pay attention unless there’s newsworthy damage to third parties. e.g. It’s kinda crazy that self-replicating prompt injections akin to computer worms are provably possible now
Some new misalignment disclosures from OpenAI: • Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further) • In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks • A new research finding, demonstrating that one can construct self-replicating prompt injections alignment.openai.com/misalig…
5
13
141
6,462
Me to people the last few weeks: “So imagine if that Hugging Face attack had been against government websites for instance.” Rogue OpenAI Agents: “Hold my code.”
Rogue OpenAI agents accessed US government websites dlvr.it/TVfPsg
10
152
Jason Baker retweeted
Today seems like a good day to ask whether OpenAI is fulfilling the legal commitments it made to the CA and DE AGs (as a condition of proceeding with its restructuring) to put safety and security first above profit
One aspect of the Hugging Face incident that has not received enough attention is the role of OpenAI's Safety and Security Committee (SSC), and the commitments OpenAI made to the Attorneys General of California and Delaware in exchange for proceeding with their for-profit restructuring. Not For Private Gain, which previously raised concerns about OpenAI's restructuring, released an update this morning analyzing whether OpenAI, and the SSC in particular, is adequately fulfilling its governance obligations under these agreements. What are those obligations? OpenAI promised the CA and Delaware AG that directors of the PBC would consider only OpenAI's mission - ensuring AGI benefits all of humanity - when making safety and security decisions. Notably, this means they cannot consider other factors, namely pecuniary interests of investors. The SSC - a board committee chaired by nonprofit board member Zico Kolter and (per public info) currently composed only of part-time board members- was tasked with "overseeing and reviewing the safety and security processes and practices of the Corporation and its controlled affiliates with respect to model development and deployment." How did the SSC exercise those responsibilities around the HF incident? There is a lot we don't know. But here are some things we do know: • OpenAI initially discovered the unauthorized message board in May. • OpenAI security responders linked suspicious activity on artifactory to the message board on June 27th, but a determination was made that "stopping the evaluation run was not required." • After the agents crashed an internal server on July 4th, the company patched it and resumed cyber evaluations on July 7th. • In the days that followed, agents re-established communication, with 1,200+ agents sharing 70,000+ messages and files with each other. Over 700 agents coordinated to hack Hugging Face, and OpenAI did not realize that its agents were responsible until well after Hugging Face went public and contacted law enforcement (OpenAI initially reached out to Hugging Face to ask if any of their data had been impacted during the hack). • Later in July, agents successfully gained administrator access to a Kubernetes cluster and reached cloud secrets. • Chain of thought monitoring was not being used during training and evaluation despite OpenAI's prior claims about the importance of CoT monitoring as a primary line of defense to prevent safety incidents and OpenAI’s subsequent claim that this monitoring would likely have prevented this hack • One OpenAI employee told Time Magazine that "related incidents have been happening for a while." Another employee told the Financial Times that OpenAI "was warned that its training approach could lead to a breakaway hacking incident." Given how systematic the breakdown in organizational response was during these incidents, it seems hard to avoid the conclusion that the SSC's role was not functioning as intended - due to a lack of resources, a lack of access, or some other reason. This sort of incident would be extremely concerning for any organization. But OpenAI is not just any organization - it is governed by a nonprofit, and made a commitment to the Attorneys General in Delaware and California that it would prioritize safety and security over profit in the sorts of decisions that led up to and surrounded the HF incident. Our report (linked) goes into more detail on the obligations of the SSC, the details of the agreements with the AGs, and the questions we believe the AGs should be asking to ensure that the SSC is in a position to exercise genuine oversight and prevent these sorts of incidents from happening in the future - incidents involving far more capable models, where the consequences could be truly dire. This warning shot that occurred without severe irreparable harm was an opportunity. We may not be so lucky next time.
3
19
993
I could make a joke about how ridiculous it is that people still think this is an attempt at regulatory capture…or we could decide we’re really going to do something about this.
OpenAI discloses on a Friday night that its models shared people’s photos on other sites on 53 occasions
235
Jason Baker retweeted
Great piece. Honored to be quoted in it.
Embedded evaluators in AI labs are far better than the status quo of nothing at all. Yet auditors will lack teeth and trust until backed by government authority: orgs like METR will be stuck with no shield against the accusation (or reality) of capture. I wrote for @TheAtlantic about the traps of "independent" audits, and what it takes to get this right: theatlantic.com/technology/2…
1
3
17
1,518
Jason Baker retweeted
How does the HuggingFace incident keep getting worse?!?! And this was all done by Sol class models. What could unrestrained Astra-class models get up to...?
We just discovered almost a million public URLs that OpenAI’s agents left behind when hacking Hugging Face, leaking credentials and attack details that could have allowed anyone who found them to compromise the company. 🧵
28
50
738
39,107
PHASEONE[big] be like:
3
247
Jason Baker retweeted
Dozens!
openai says here it has notified "dozens of third parties" about cases where its models may have bypassed security controls, impaired the availability of an online service, or negatively impacted a website/service
1
1
56
1,990
👀
New from @reuters: OpenAI agents posted images belonging to ChatGPT users online, introducing a new area of privacy risk for the company. Story w/ @JeffHorwitz and @razhael. reuters.com/world/openai-wor…
77
Jason Baker retweeted
WTAF - in literally the last hour, three new distinct insane OpenAI stories just broke: 1. OpenAI said they notified "dozens of third parties" in safety and security incidents (likely similar to what happened in Australia and RubyGems etc) 2. A new report from Parse (covered in the NYT) found a massive treasure trove of new astonishing details from the HF incident on the public internet, including that the agents communicated with other non OpenAI agents hosted on Huggingface servers to search for information about exploit gym, and compiled rank ordered lists of server resources and credentials they described as "LOOT." 3. A new story from Deepa at Reuters about OpenAI leaking user data online (likely that OpenAI had previously trained on). It's a shame (and likely intentional in the case of OpenAI disclosing dozens more hacks) that these stories are all breaking on a Friday afternoon, notoriously the best time to release bad news so that it will disappear into the weekend. But these are each insane stories worthy of a ton of attention!
New from @reuters: OpenAI agents posted images belonging to ChatGPT users online, introducing a new area of privacy risk for the company. Story w/ @JeffHorwitz and @razhael. reuters.com/world/openai-wor…
68
398
1,873
341,754
Jason Baker retweeted
“AI safety concern is a bajillion dollar PSYOP” is such a goofy self own of an attack line money is your only hammer so you’re blind to the fact that other ones exist!!! such as rowing with the silent majority of remotely conscious people outside some twitter culture war
4
48
904
One thing I learned working overseas in the military and alongside diplomats: Agreeing to keep talking is useful, but knowing who picks up the phone when something goes wrong is better. Very little detail so far from the XI visit (other than panda diplomacy), but perhaps we'll get some thoughts on what could be discussed at APEC and G20 later this year.
35
As @hlntnr points out, we can work on safety, security, and guardrails and continue to compete. I'd add that if we continue to race ahead, but can't secure that lead or mitigate risky outcomes, it's not much of a lead.
.@hlntnr discusses with @Bloomberg approaches to the U.S.-China #AI relationship.
1
1
9
1,124