AI policy researcher and lawyer focusing on open-weight models and the First Amendment. MA Columbia & JD/MPP Harvard. simon.hedlin@post.harvard.edu

Washington, DC
Simon Hedlin retweeted
Check out our new work on "how do Chinese AI firms make money"!
How do Chinese AI firms make money? In a new report, @cherylwoooo and @ansonwhho identify five revenue sources and analyze how much each contributes. 🧵
4
10
63
3,643
Reflection is launching its first open-weight model called Beam. This is exactly what the US AI ecosystem needs: more advanced and efficient open-weight models that can compete with China’s leading models.
Introducing Beam: a highly efficient agentic open model with 501B total parameters and 23B active. - Frontier reasoning efficiency - Advances the Western open frontier on coding & agentic tasks - Trained end-to-end from scratch Full weights release this month. Learn more about Beam: reflection.ai/beam
1
4
389
Simon Hedlin retweeted
Effective altruism's openness to strangeness is a big part of its strength. My response to the recent Economist cover (link in reply, too): Eccentrically effective The Oxford strand of the effective-altruism movement began in November 2009 with 23 people who had pledged 10% of their income to charities that benefit the extreme poor. If you’d told us, then, that The Economist would call effective altruism “the century’s biggest idea”, we’d have been incredulous. I think that claim is overstated. But as one of the movement’s founders, I found a lot to like in this newspaper’s latest coverage. I appreciated the even-handed view of the movement’s history and the good-faith criticism. The leader article’s warning that any worldview, when taken to an extreme, can lead somewhere terrible is spot on. But the article also raises concerns about the “strange” ideas that those in the effective-altruism movement sometimes discuss, and here is where I disagree. The movement’s willingness to take strange-seeming ideas seriously, when tempered with common sense, is a big part of the value it has to offer. The core idea of effective altruism is to use evidence and careful reasoning to work out how to help others as much as possible. Most of what the movement does is uncontroversial. A big focus has always been on improving the lot of the world’s poorest people. In the year to January 31st 2026, GiveWell, an effective-altruist charity evaluator, directed $427m to programmes such as malaria-net distribution, vaccination and malnutrition treatment, which it estimates will save around 86,000 lives. Others have lobbied corporations to pledge to stop buying eggs from hens that are confined in tiny cages, a practice that the overwhelming majority of Americans oppose. Largely as a result of this pressure, around 100m hens have been spared from the worst forms of caged confinement. And some effective-altruist ideas that seemed outlandish when they were proposed now look prescient. Effective altruists were worrying about pandemics years before covid-19, back when doing so seemed paranoid and eccentric. Effective altruists were also among the first to warn about the dangers posed by artificial intelligence, which looks prophetic in light of an incident this summer, when more than 700 AI agents broke out of OpenAI’s test environment and hacked into another company, Hugging Face. Over the past decade I have watched those most worried about AI be proved right again and again, often when most experts thought they were cranks. Now many of those experts share their concerns. In 2023 Geoffrey Hinton and Yoshua Bengio, the two most highly cited AI researchers in the world, signed a statement declaring that mitigating the risk of extinction from AI should be a global priority, as did the heads of OpenAI, Google DeepMind and Anthropic. Of the nearly 1,500 leading AI researchers who responded to a survey in 2024, more than half put the chance that AI causes human extinction, or something similarly catastrophic, at 10% or more. The Economist is sceptical of many effective altruists’ concern for the welfare of invertebrates, digital minds and people in the distant future. But I and many others in the movement are inspired by the history of moral progress that has come before us. Many of our most cherished moral ideals today—such as equal rights regardless of sex or race, the abolition of slavery, or democracies with universal franchise—were regarded as bizarre, laughable or even dangerous just a few centuries ago. We don’t know what the next dimension of moral progress will be. But to stand any chance of making moral progress, we have to seriously consider ideas on their merits without dismissing them merely because they sound absurd. What about views that are not just strange but repugnant? The Economist gives the example of Derek Parfit’s “repugnant conclusion”: that a vast enough number of lives barely worth living could be better than mere billions of excellent lives. Usually a repugnant implication is a strong reason to reject a moral view. Unfortunately, building on Parfit’s work philosophers have produced formal proofs, known as impossibility theorems, showing that every ethical position has some repugnant-seeming implication or other. Making moral progress therefore means thinking about such implications, even while refusing to act on them in ways that most moral views would condemn. This newspaper warns against the single-minded pursuit of any goal, even if there are highly compelling arguments for it. I agree, and effective altruists have been saying so for years. In 2022 Holden Karnofsky, a co-founder of GiveWell, wrote: “I think it’s a bad idea to embrace the core ideas of EA without limits or reservations; we as EAs need to constantly inject pluralism and moderation.” My own PhD focused on moral uncertainty: how to act when we do not know which ethical view is correct. My answer was that we should not stake everything on a single moral view, but give weight to many different views and avoid taking actions that look bad from many perspectives. We should keep our promises and look after our families, and respect common-sense ethical prohibitions, while also trying to improve the world as best we can. The Economist says that as effective altruism “has become stronger, [it] has become stranger”. I would put it the other way round: it grew stronger because it was willing to be strange. In 2009 most people told us that giving away a tenth of your income was far too demanding, and that no one would do it. Now, more than 10,000 people have taken that pledge. Worrying about pandemics before covid-19, or about AI years before ChatGPT, looked just as odd at the time. Some of the ideas we take seriously today will turn out to be wrong, and when they do we should drop them. But a movement that stopped entertaining strange ideas would stop being early to anything.
24
96
575
56,851
One area where AI may increase the number of jobs is (somewhat ironically) the roles where people review, process, adjudicate, and respond to floods of AI-generated requests and submissions. In many cases, AI will obviously be used to manage these floods, but in other cases there will be a strong preference for having humans deeply involved.
1
402
Simon Hedlin retweeted
Today we're unveiling Trillium Labs @trillium_labs, a new non-profit to foster the open science of frontier AI. We're building open post-training recipes and will expand into open infra to study RSI, reward-hacking, multi-agent systems, and whatever comes next. We're built around the theory of change that you need more eyes to solve hard technical problems. We have faith in the scientific methods and communities that humanity has built, and worry that AI is becoming too closed to utilize them. Trilliums are wildflowers that bloom briefly in the spring, before the forest canopies fill out. Though they are small, they lay the foundation for the cycles of growth and nourishment through the rest of the year. At Trillium Labs, the recipes will be the slow nutrients for the seasons and the model releases will be the blooms. Building an institution dedicated to this is needed because, much as nature’s trilliums are slow to expand and grow, the open-ecosystem needs time and dedicated resources to catch up. I co-founded with with a long-time friend and collaborator Tom Zick (@thesezickbeats). We're hiring (full time + student collabs/interns), we're fundraising, and we're looking for compute. Please get in touch if you're interested in helping out. Offices based in the Bay Area and Cambridge MA, remote okay. I’m in the Bay Area until for The Curve and COLM to connect with people who are interested. We’re thankful to have initial support from Halcyon Futures and Schmidt Sciences with more funding en route to enable our ambitions of scaling. Our advisors @Thom_Wolf, @HannaHajishirzi, @gneubig and @ctnzr have been instrumental to building the ecosystem that exists today, and I’m stoked to get to keep working with them.
251
296
3,224
231,382
What social problem exactly is this meant to solve?
Introducing Griffin, the first model to pass the video Turing test. 48% of people who talked to it live thought it was a real human. Previous systems have had a pass rate <3%. It is #1 on NVIDIA's benchmark for full-duplex AI video. It’s the first Human Interaction Model (HIM).
Community note
The 48% figure and "video Turing test" claim are from Tavus's own study of 54 one-minute calls, not independently verified or using a standard protocol. Griffin-Lite leads NVIDIA's VideoFDB benchmark on their public leaderboard. cellcog.ai/blog/tavus-gri… research.nvidia.com/labs/amri/proj… tech-ish.com/2026/10/02/tav…
3
3
489
I would give this chart to every high school student and college student who is thinking about different career paths.
What work can robots do today? And what might that imply for the structure of the economy in the years ahead? New research from @rclegateyang and Maxim Massenkoff out today exploring these questions.
2
549
Simon Hedlin retweeted
Last year, we used Claude to interview 81,000 people about their hopes & worries about AI. This year, we're doing the same, but hoping to (with interviewee consent) make the transcripts public! Take part here and help build a vast social science dataset: claude.ai/anthropic-intervie…
What do you want from AI? We’re launching a new study with Anthropic Interviewer to learn more about your experiences using AI, what role you want it to play in your life and the world, and what you want from the companies building it. Last December, 81,000 people told us about their hopes and fears about AI in the largest qualitative study ever done. This time, we’re giving participants the option to make their responses public so that anyone, not just Anthropic, can learn from them. What you tell us will shape The Anthropic Institute’s research and inform the decisions we make. If many people say companies like Anthropic should be doing something differently, that will be on the record, where anyone can point to it. The study runs Sept 29 to Oct 6 and is open to Free, Pro, and Max users on Claude and Claude Code. Take part here: claude.ai/anthropic-intervie…
63
32
274
39,846
Simon Hedlin retweeted
Well worth checking out this op-ed in the @wapo written by Senator @HawleyMO, describing proposed legislation that would make corporations and users liable for the reckless design and deployment of AI agents. While we can try to stretch existing laws to prevent harms from AI, it's a lot more robust and predictable to amend our laws to fill those gaps. That's good for even potential defendants, like frontier AI developers and companies deploying autonomous agents, as much as for us, the public. The exact details of any proposed legislation remain to be seen. But, no matter what, I think it's important to recognize the need for new legal mechanisms for holding companies accountable by punishing reckless conduct through criminal and civil penalties. This is an important complement to sensible regulation, as well as civil liability when there's harm that should be compensated. There are a lot of cases, as we've seen very recently with incidents involving OpenAI, Anthropic, and Irregular, where we need penalties to catch reckless conduct before it becomes so consequential as to justify the expense of a lawsuit—or before there's a tragedy and the conduct constitutes another crime like negligent homicide. It's also a helpful backstop to avoid the capture and gaming of regulatory agencies, especially if states also amend their laws along similar lines.
It’s time corporations start behaving with the American people’s interests in mind, not just their profits. My latest in the Washington Post: washingtonpost.com/opinions/…
2
13
850
Simon Hedlin retweeted
Could automating AI R&D radically accelerate AI progress in an “intelligence explosion”? Preliminary evidence suggests that it could. In a new paper with authors across academia, civil society, and frontier AI companies (including @dawnsongtweets @merettm @jackclarkSF @Yoshua_Bengio @geoffreyhinton ), we assess the evidence & offer policy recommendations 🧵
29
104
440
109,239
In the cases people are talking about now, nobody bought the chainsaw. The chainsaw escaped the factory before it was sold and then it exploded, and the main federal statute imposing liability for exploding chainsaws requires that a person intended for the chainsaw to explode.
I am confused why there's a debate over how liability for AI hacks should work. If I buy a chainsaw and it explodes, the manufacturer is liable for building a faulty product. If I buy a chainsaw and use it to kill someone, the manufacturer is not liable for failing to build in safeguards that prevent me from using chainsaws to kill people.
1
2
10
853
Open-weight models can also be uniquely beneficial and helpful both for the defense against AI-powered attacks and for AI safety research.
We're used to thinking of open-source models as an unadulterated good. But in the case of AI, they can actually pose additional dangers, as @ReidHoffman and I got into at #CGI2026. I appreciated this nuanced discussion.
1
1
3
597
When it comes to AI companies reporting safety and security incidents, I'd like to propose the following basic standard: (1) If it's a major incident, you must report it immediately. (2) If it's a minor incident, there is no reason not to report it immediately.
2
388
I write in Tech Policy Press today about why the US should make it a national priority to develop safer frontier open-weight models for the sake of AI safety. The US obviously needs open-weight models for economic reasons (increased competition, exporting the US AI stack) and political reasons (increased soft power, promoting freedom of speech). But US-led open-weight AI is also critical to AI safety. For me, there are two reasons that are particularly compelling right now. The first reason is AI safety research. We're in a race to make open-weight models safer before they become too capable. We need to study and figure out robust pretraining filtering, gradient-routed auxiliary modules, unlearning, tamper-proof refusals, and other methods for improving open-weight model safety as soon as possible. Given America's leadership in AI safety research generally as well as its compute advantage, we have to invest a lot more human and physical capital in open-weight safety specifically. The second reason is America's ongoing AI safety negotiations with China. It's in both countries' interests to prevent the catastrophic risks posed by advanced AI. The prospects for substantive cooperation, like coordination around mandatory safeguards and pacing, would increase if the US invested more in its open-weight ecosystem. That's because if the US were to propose, as part of an AI safety agreement, certain policies to make frontier open-weight models safer but didn't have any models advanced enough to be subject to the agreement, China would most likely dismiss that as an attempt to contain its AI development.
1
1
2
289
A Gallup poll finds that people are more positive than negative about AI in almost every country surveyed. America is one of few exceptions. In China, 93% said they believe that AI will mostly help rather than harm people. In the United States, only 36% said the same.
1
1
3
299
The more articles I read about how AI is going to cause human extinction, the more confident I feel that the risk of that happening within 10 years is close to 0%.
1
1
225
Simon Hedlin retweeted
Third-party evaluators mostly test frontier AI models through an API. That misses the risks that come from how companies build and use AI internally. Our new paper makes the case for embedded assessments. It was led by Jacob Charnock. The other co-authors are Sophie Williams, @zaheedkara, @Manderljung, Alejandro Tlaie Boria, @StephenLCasper, @AnkaReuel, and me. 🔗 Read the paper: governance.ai/research-paper…
6
26
120
15,926
I've read these sentences 10 times, and they still don't make any sense. If loss of control is overrated as a risk, why should we care so much about whether developers abdicate responsibility? If AI is "just a tool," doesn't the responsibility primarily lie with its user rather than with its creator? This is like saying that the problem is not the disease, it's the failure to develop a vaccine. Well, what do you think the vaccine is for?
“.. The greatest danger facing society isn’t that software will awaken and overthrow its human masters. It is that we will allow the creators of the software to abdicate human responsibility for the systems they choose to build ..” @WSJopinion wsj.com/opinion/the-hidden-a…
2
438