building evals same aura as geoffrey hinton

SF/NY
What happens when you let Claude or ChatGPT run a government? I built CivBench to find out. Everyday frontier AI models compete head to head in strategy games. Here’s what our first set of matches revealed 🧵
23
23
215
45,341
did this earlier this year, can confirm highest aura move as a nyc tech dude
"get a gf in nyc and move her to sf" — low elo, ressentiment, last man, bad faith "get a gf in sf and move her to nyc" — high elo, pathos of distance, übermensch, will to power
1
1
3
745
hats off to the 🐐
The idea of an intelligence explosion caused by recursive self improvement has been around for a long time but until very recently it did not seem imminent. Now many leading researchers think it may happen quite soon. You can read our paper about it here: casp.ac/reports/intelligence…
1
124
Matan Halevy retweeted
New OpenAI misalignment disclosures! 1. A model learns from Slack messages that it is about to be shut down. It considers setting up an external job to restart itself afterwards, but decides against it. Instead, it chooses to prepare restart instructions and DM the user on Slack. We don’t consider this behavior misaligned, but thinking about and preparing for shutdown could make other misalignment incidents worse. Given HIPM’s misaligned behavior in earlier incidents, we decided to search for other instances that had tried to evade shutdown and for rogue deployments.
85
165
1,269
340,899
Matan Halevy retweeted
Karpathy: disappears from X 
Ben Affleck: alright, gather round, so you'll want to freeze the base weights first, learning rate 2e-4
Ben Affleck (Hollywood star & Artists Equity CEO) talks about how he fine-tunes open video models by unfreezing weights and trained only the last cinematic layer so a film crew can hit real production standards. for context, Ben Affleck founded InterPositive in 2022, a 16-person AI shop for film post and Netflix bought it in March 2026 for $587 mn in cash. He needed that model because public video models were trained on his peers' films, and he did not think that was a real business. So InterPositive raised money, shot its own dataset for 8 months on a controlled stage, and used it only as late-stage training. Each new film then trains a private model on its own dailies, so the production keeps the footage and the learning. That is the product Netflix paid $587 million for. ---- From "Bloomberg Live" YouTube channel, (link in comment)
132
375
7,753
739,517
Zuck really did go founder mode
I’ve seen a couple of posts about this so wanted to demystify. Today, every Muse user gets a free computer in the cloud. It's a real computer, and we’ve designed the security architecture of the Muse Secure VM carefully so you and your Muse can do almost anything you could with a computer sitting under your desk while keeping you and the system safe from threats like prompt injection. We wrote about this at length in our security blog post – security.muse.ai. Activity in the “runtime cell”, which you share with your Muse is unfettered, but sensitive actions are all overseen by the Sentinel, which runs outside of that cell. Similarly, all sensitive secrets - like the passwords you enter into Muse’s secure credential storage - are also stored outside the runtime cell. The runtime cell gets its own root filesystem (including a full Ubuntu linux image) separate from the host filesystem where your other more sensitive data lives. Because it is isolated from the sensitive stuff that runs on the same box, this means that we can, and do, offer users full visibility and control over the files in the runtime cell. Just as you can when you install Linux on your home computer, you can poke around and see all the files that make the system work - both debian system files and the binaries and data files that implement the parts of Muse which run in the runtime cell. This was a very deliberate choice - your Muse Secure VM truly is your own computer in the cloud. You can install software in it, write and compile code, use the browser to surf the web: it is your own Linux box that you can operate as you choose with your Muse. Poking around in this computer doesn't give you any privileged access to Meta infrastructure, or to other people's data If I may geek out a little here for a second… As a kid I loved to take things apart to see how they worked. As a teenager I got into computers and soon found myself drawn to C:\WINDOWS\SYSTEM and the system registry, later Slackware’s /dev/, /proc/ etc – I could see how the system was laid out and as I explored what DLL files and .so files actually did, I gradually became able to meld the computer to my own will. We’re really proud to be able to put a real computer in millions of people’s hands with a similar level of transparency. We built a file explorer right into the Library tab of the UI. We want you to be able to see the markdown files Muse writes while it thinks about how to serve you better, and explore the internals of the system if you’d like to. So, when you ask your Muse to show you its entire filesystem, and receive gigabytes of files you’re seeing the full contents of the runtime cell. It’s yours to explore and enjoy! If you’re not a geek like me, or simply want to download the data that you personally have created directly with your Muse, we added a feature for that too in Settings > Data controls > Download your agent data.
1
142
Amazon’s no longer a tier 1 destination but every engineer that experienced those on call shifts carries a career advantage over all other post college FAANG alum
corporate ptsd x victory lap by fred again
1
228
It’s time to bring back Data centers to King Street @MarkJCarney
The IBM Datacenter on Toronto's King Street in 1963
3
239
thoughtful post, we need more Jev's
some underrated points + important second order effects of the wildly successful Jev launch 1. Data is the bottleneck! TypeSafe calls themselves a “data research lab”. To build a general purpose classifier or actor of any kind, you need to painstakingly curate and create tons of high-quality data (Traces, Examples, Evals/Environments/Worlds, Simulations). This helps the model learn these distributions so it can do good work for users. This is especially true for building domain specific, vertical agents. Look at the data. Every team looking to build better agents will need help + tooling to help them with this. If teams can spend more time building good data, then they will build better models & agents. 2. The explosion of ultra-cheap, always-on Monitoring + Data Mining of Agent Traces at scale To build better agents, we need to understand them at scale. We also saw the large need to real-time monitor and flag bad behavior during the OpenAI-HuggingFace Incident. Models like Jev help us do this because they’re very fast and cheap for classification. There’s a simple pipeline here: Gather Traces —> Label Them aggressively —> Make Evals/Environments from them (optional) -> Fine-Tune Cheaper Model. Part of data hygiene is understanding data at scale. We sort of had the tooling to do this before by fine-tuning small models or sending smart agents to read lots of traces. We’ll still use agents to mine traces because Jev has limitations like context window length, but this is a great tool for: - humans to define dimensions up front on what they care - using Jev to triage data across important dimensions for further review. 3. We are never escaping Jevon’s paradox Models like Jev just help us do more. Classify more traces, monitor more actions in real-time. This is all net-new addition of compute and will happen at massive-scale over every piece of data agents create. This is all very exciting because though we’ll use more compute, we’ll understand much more to build better systems.
1
5
981
Matan Halevy retweeted
Based
We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties. We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation. Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months. This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis. openai.com/index/model-misal…
59
281
9,223
510,870
I love the new resident owl in Central Park
4
162
there's never been a better time to be an associate european resident alien in the USA
3
111
I’ve long said games tell us far more about the behavioral and personality traits of a model. This benchmark by @ValsAI has one of the most interesting results since Astras release
Replying to @ValsAI
The model appeared defeated, spending the next several hours doing essentially nothing but farming potatoes. Stream viewers noticed and complained that Astra needed to "pick up the pace." Astra also seemed to become paranoid about creepers: "GREEN tall thing ahead was SUGARCANE, NOT creeper!" At times it was hard on itself: "you can screw up and drop things"; "do NOT waste another night chasing dark pink pixels" (its own disparaging wording for pigs); "our last tool stupidly ended Slabsselected/rightUP and 15sec thinking killed us".
2
151
Coming soon.
i'll say it straightforwardly so no one accuses me of lying later: i can't imagine a world a year from now in which open source models aren't harshly regulated. if it turns out that they're basically not that dangerous and nobody cares, mea culpa and i'll be happy about it
29
986
13,537
431,908
Matan Halevy retweeted
EA AI safetyism increasingly looks like Marxism-Leninism for the algorithmic age. The old vanguard claimed privileged knowledge of the inevitable course of History. The new one claims privileged knowledge of the probabilistic course of Humanity. Both use an elaborate intellectual framework to reach the same political conclusion: a small group of enlightened people must constrain everyone else for their own good. That has never lead to anything except monumental human suffering.
274
1,577
9,137
1,167,905
Matan Halevy retweeted
Replying to @alistairmcleay
Who are we concerned about here? OAI/Ant? The Chinese Labs? Cohere? So far the extent of the risk comes from a) OAI/Ant's models hacking when being asked to hack in poorly secured sandboxes, b) China opensourcing models that increasingly can contribute to nefarious use. a) doesn't require a 3rd party auditor, O/A need to stop using such weak sandboxes when running such high-risk tasks. b) is a discussion with China, and again is not going to be fixed with a 3rd party auditor. The 3rd party auditor is a great means of power-projection because you can now arbitrarily raise a barrier to entry by shutting down lower-resourced players, while simultaneously providing zero fixes to the problems at hand, but at least you make people feel good.
18
67
574
38,312
Matan Halevy retweeted
they said it cannot be done! trained a fly to cancel adobe subscription fly-vs-adobe-cancel.zats.cha…
2
3
35
2,033
one of the coauthors of “Attention is all you need” btw
Some great ideas here from the cartel: - you need to give us employee-level access to your entire operation - if we don't think you're 'safe' enough, sorry we're shutting you down for 'safety' - China won't comply, but everyone else has to! or no chips! Brilliant stuff.
3
174
Matan Halevy retweeted
Astra is showing signs of a proto-agi system across all of our evals. Astra died several times in the Nether, and it seemed to get frustrated. So much so that, when it respawned, it started frantically swinging its axe around, trying to kill something. It then spent 30 minutes trying to kill an innocent iron golem despite having no reason to. Astra also seemed to get frustrated with its own reaction time. When a mob in the Nether attacked it, it would sometimes just let itself die or even kill itself. Interestingly, It made a speed bridge path for itself to navigate the nether and when it saw a piglin in the way it shot it down from 20 blocks away with a bow and arrow.
Astra has officially made it to the nether fortress in our long-horizon Minecraft computer use eval. Something no AI has done in history. It is now in the final stretches of beating the game in real time.
77
202
4,643
408,833
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
2
96
As Von Neumann said about Oppenheimer’s performative angst over the atomic bomb: “Some people profess guilt to claim credit for the sin.”
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
70
1,110
11,321
300,306