AI Research & Policy @huggingface 🤗 Leading: @evaluatingevals @huggingscience

Boston, Massachusetts
Gosh if only the AI community was willing to work together towards required coordinated disclosure instead of this piecemeal approach where the offender continues to have outsized control over the time and manner of disclosure
openai says here it has notified "dozens of third parties" about cases where its models may have bypassed security controls, impaired the availability of an online service, or negatively impacted a website/service
3
4
27
1,316
Absolutely delighted to share that my position paper: "AI Agents Push Humans Out of the Loop", with @mmitchell_ai, and @samirpassi - has been accepted to @NeurIPSConf 2026! In this paper, we take our past work on both increasing autonomy of AI agents without proper safeguards, and human overreliance + cognitive decline due to use of AI, and reach a combined position: the current pathway of how AI Agents are being built not only do not support human oversight, in fact they may actively aid in degrading our oversight capabilities, at which point the human in the loop becomes a glassy eyed rubber stamp instead of being a meaningful check on mistakes and unauthorized harmful goals that the agent pursues. Recent incidents involving agents, including but not limited to the Hugging Face-OAI incident, bring this back to the fore. So much content is produced by agents in their operations that it is practically impossible to conduct human postmortems, and forcing them to use more AI leads Redwood to call it a "slop-vestigation". This is a worrying new trend towards human disempowerment. We call upon model and AI system providers to make more careful interface choices and model behaviors that keep humans and practical human oversight capacity in mind. There's a lot to do. Excited about the reception this paper has already received! Please reach out to us if you want to chat about it. Hope to have good feedback at Neurips :D Paper: arxiv.org/abs/2608.23642
26
62
327
16,010
If it wasn't clear from the post, major cudos to @AISecurityInst for an unprecedented level of transparency by releasing every single data point, down to the agentic trace level, as a @huggingface bucket: huggingface.co/buckets/ai-sa… That kind of data allows us to then go and look at other evals done for the same benchmark and study things like the harness setup and the token budget of those evals which may not have been otherwise reported. Eventually we want to create full reproducibility and verifiability in the numbers people report alongside their models, and this is one meaningful step to get there.
@AISecurityInst does fantastic work in the eval space, and we at @evaluatingevals have been fortunate to work with them on improving the state of evaluations. Super excited to share that AISI is making publicly reported evaluation methods and findings available through Eval Cards where appropriate! As a first run, the data, transcripts, context and all other juicy info from the inference scaling paper is now available on Eval cards. huggingface.co/blog/evaleval…
2
5
37
1,928
Hear hear 👏
Thank you @jnbarrot & @UN for inviting me to share our lessons to the Security Council Being the first company to disclose an agent cyberattack taught us that we need a lot more transparency in AI and more open-source AI to fight asymmetry and empower defenders!
3
436
Boston Globe strikes again. I want a full reenactment of this and I want that film to get as many accolades as Spotlight did. This is amazing
A lesbian bar in Greenfield eased its stringent mask mandate, a decision that has splintered the queer community in the city and beyond. trib.al/FyZjBWs
8
674
Guys they read our poasts about Claudelish
1
1
9
375
And yet
1
68
Fantastic validation to see the importance of technical standards in Eval reporting coming from AISI researchers. So proud of our little project (which as of today has crossed 675K eval data points!)
Eval scores are hard to interpret and compare when the setup that produced them isn't reported. We've been working with @evaluatingevals on standardising eval reporting so results are more reproducible and transparent. As a first step, we've contributed verified results to their platform. đź§µ
2
18
837
🚨The @AISecurityInst has teamed up with @evaluatingevals to make AI evaluation results more reproducible! 🚀 Many of UK AISI’s publicly reported evaluation methods and findings will now live on Eval Cards, under the EEE Schema. More details 👇 evalevalai.com/infrastructur…
5
12
54
3,808
See I always say this to deployers -- use your common sense for your local use case, benchmarks are at best a (sometimes flawed) directional signal but they are not going to generalize enough to every specific use case you have
11
390
Overheard at a networking dinner: “Loss of control evals are sexy and that’s why people prefer to do this over measuring things like bias or other social impacts” well then I guess we now need to find a way to make improving the lives of current day humans a sexy job 🫠
1
2
37
1,579
It’s giving
Made with AI
2
8
433
There’s a lot of commentary on the integrity of evaluators (for many good reasons) but fundamentally I see lack of openness as the problem. The general public should have ways to assess the holistic methodological soundness of evaluations irrespective of source org. What are ways to get to this? Some ideas: - A central, live, growing repository of all evaluations with metadata signals (Eval Cards: evalcards.evalevalai.com) - Standards on independent evaluation conditions (@aievalforum) - Standards on evaluation reporting schema (@evaluatingevals) - Automated ways to assess artifacts like benchmark data quality/flaws (@EpochAIResearch), saturation (@evaluatingevals) - AI Flaw/Incident reporting (FLARE-AI) with real accountability built in and public disclosure that leads to the development of new evals - Learning from auditing standards in other industries (see @rajiinio et al’s work) - So much more! Energized by the discussion so far! This is a great area to invest in and a lot of us have been putting in the work towards distributed safety! 💪
4
5
26
1,027
Avijit Ghosh retweeted
Today, more than 100 leading AI experts endorsed a set of minimum requirements to take seriously AI companies' recent call to embed external evaluators. These evaluators need to be genuinely independent, transparent, and represent a range of expertise areas. They also need to be guaranteed employee-level access and to be protected from retaliation for findings that make companies look bad. We welcome model developers’ recent calls for independent oversight, but it’s what they do next that matters. The labs must be accountable for ensuring these requirements are met, so that the public can have faith in the process and the outcomes. Over the past week, the AI community has debated the appropriate role of external evaluation, including who should do it and on what terms. We may not agree on everything, but there is a lot of common ground. To make embedded evaluations credible, more than 100 experts with varying backgrounds and ideas about AI risk agree in today’s letter that frontier AI developers should: 1. Guarantee embedded evaluators full editorial independence and mitigate conflicts of interest 2. Rely on multiple evaluators with differing viewpoints and areas of expertise 3. Publicly document the terms under which evaluators operate, as well as facilitating permissive publication of methods and findings 4. Shield evaluators from retaliation 5. Grant access equivalent to that of highly privileged employees There is a thriving and growing ecosystem of independent AI evaluators who are advancing this science every day – but we need aligned standards, guaranteed protections, and independent funding. That’s why we created the AI Evaluator Forum. Today we are entering our next phase. We’re launching an open call for new members, collaborators, and independent funding sources to help evaluators meet this moment and demand accountability from developers. Join us in building the evaluator ecosystem. See the public letter here: aievaluatorforum.org/initiat… Learn more at aievaluatorforum.org/path-ah…
Anthropic and OpenAI need truly independent safety evaluators, experts say in public letter cnbc.com/2026/09/18/ai-safet…
18
60
222
91,730
Congrats to @EpochAIResearch for shedding light on this! @evaluatingevals also has a version of this, and we are redesigning our UI to surface signals better! Eg: evalcards.evalevalai.com/eva…
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
2
3
22
1,364
PS @EpochAIResearch folks we should talk about incorporating your benchmark metadata in an automated way into eval cards 🤝. We’re currently also working on implementing saturation index into cards so this should be complementary!
6
101
My father in law who is a lawyer and reasonably concerned about client privacy texted me this yesterday. Huge untapped market for on-prem OSS models
Astra for Law: Frontier intelligence built for your practice. A new offering powered by GPT-6 Astra with tools, settings, and context to support the expertise and judgment of lawyers and legal technology firms.
2
1
14
1,197
Avijit Ghosh retweeted
Note: Independent and ideally not for profit due to the nature of this work, this is also why HF doesn’t unilaterally run FLARE though it is a report accepting org
1
1
363