Impt take on the compounding impacts of frontier model lead assuming the 4-8 months gap between open and closed source will remain for years. "If you're not on the frontier, the frontier cannot help you." At the compute level, frontier labs can use current SOTA to build the next generation. Small delays translates to higher inefficiency in compute @cerebras @andrewdfeldman @imaginationxyz
3
76
Did @satyanadella just say “evalmaxx”
Satya Nadella said the quiet part about the entire AI industry out loud: for the first time you are buying a technology where the data your own use generates may not belong to you "it's like if I sold you a database and said the data you put into your database is not yours and it's mine, and it goes away if I took away the license. "How would you like that?" this is him explaining why he thinks the model layer taking all the royalty cannot last, the one test he tells enterprises to run before they trust any model, and the new kind of insider risk he says nobody is building for "it may fake my books" "we should avoid these cozy arrangements of who is testing what" bookmark & watch it - then read the article below ↓
1
129
“The best way to be right as a founder is to be wrong for a long time and not be dead” @andrewdfeldman @cerebras
Amazing to hear the humility and candor of @cerebras CEO @andrewdfeldman after their $63 billion IPO in May 2026 “You’re the same bozo before and after the IPO… when nobody was buying my stuff I was stressed at 11 and now everyone wants my stuff and I’m stressed at 11”
127
Amazing to hear the humility and candor of @cerebras CEO @andrewdfeldman after their $63 billion IPO in May 2026 “You’re the same bozo before and after the IPO… when nobody was buying my stuff I was stressed at 11 and now everyone wants my stuff and I’m stressed at 11”
223
Looking forward to joining this rare gathering at Google Bay View organized by @johnkwerner and the @imaginationxyz team— will share more from talks with @StanfordHAI @GoogleDeepMind @MIT_CSAIL @Theteamatx & more about where we are on the frontier of autonomy and agentic worlds.
Silicon Valley, we’re coming. 📍 Google Bay View 📅 September 14 & 15 Apply now: imaginationinaction.co
3
155
Financial firms have some discretion, but most of the global compliance regime is based on overaccumulation of information "just in case" and out of an abundance of cautious and fear of regulatory scrutiny not because it is what is necessary for mitigating AML / TF risks.
Replying to @maxkarpis
This answers the wrong problem. Why should every big and small fintech permanently store loads of private and sensitive personal data in the first place just to comply with a one time or occasional kyc events
2
115
We need to scale up @selfxyz and other zkp auth tech. Compliance docs have no reliable safeguards and we can achieve better outcomes without compromising peoples privacy and security
‼️ BREAKING: Revolut handed over customers’ passport copies, verification selfies and full transaction histories to a malicious actor. The actor sent lawful government information-demand emails using a genuine government domain that passed domain authentication. Revolut later concluded they were not authentic. Affected customers were notified on Friday. What may have been disclosed ranges from name, date of birth and home address to account statements, withdrawal records and complete Bitcoin transaction history. The company says it has alerted the agency to the unauthorised mailbox on its domain, blocked the address and begun notifying regulators. It has not named the agency, explained how someone obtained a mailbox there, or given a number of affected customers. ZachXBT, who circulated the notices, believes the incident was limited in size and aimed at high-net-worth users.
1
6
368
Revisionism from the prev admin. Once gears are in motion no legal (Sec 502B, Leahy Laws) or humanitarian safeguards had any weight in the smog of war. The policy was to stand down any interruption to the flow of Mark 84, BLU-109, JDAM, GBU-39 etc and why Josh Paul resigned.
Biden advisor Jon Finer: We should've used arms transfers as leverage but unfortunately our lawyers couldn't determine if Israel was doing war crimes @PeterBeinart: But didn't the Biden admin pressure lawyers not to make any determination? Finer: Well *I* never pressured anyone
4
224
OpenAI's actions underlines the fine print: it cannot rule out de-identified data from their product usage helping their models. Because traces and telemetry logs can feed evals, SFT, and RL even when a lab never opens a specific user’s files and are the compounding AI asset.
We congratulate Levent Alpöge and Tristan Buckmaster on their remarkable mathematical work. We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs. unforced).
106
The new WorldDiT robotics model from @bageldotcom shows that models are getting smaller and more capable. More with less. < 1B parameter model can predict how the world will change, then decides what the robot should do
We are releasing WorldDiT, a unified architecture for robotics world modeling and control. On the LIBERO benchmark, it performs the best among all publicly released methods that do not need a VLM to generate actions. Its size and performance sit on the reported Pareto frontier.
1
1
5
11,726
Alpen Sheth retweeted
We are releasing WorldDiT, a unified architecture for robotics world modeling and control. On the LIBERO benchmark, it performs the best among all publicly released methods that do not need a VLM to generate actions. Its size and performance sit on the reported Pareto frontier.
28
54
154
70,861
Alpen Sheth retweeted
Geoffrey Hinton says a big language model runs on about 1% of your brain's connections and still ends up knowing more than you: "So in your brain, you have a hundred trillion connections, roughly speaking. Okay. That's a lot. And you only live for about two billion seconds. That's not much." "If you compare how many seconds you live for, with how many connections you've got, you have a whole lot more connections than experiences." "Now with these neural nets, it's sort of the other way round. They only have of the order of a trillion connections. So like 1% of your connections, even in a big language model, many of them fewer, but they get thousands of times more experience than you." "So the big language models are solving the problem with not many connections, only a trillion. How do I make use of a huge amount of experience?" "And back propagation is really, really good at packing huge amounts of knowledge into not many connections." "But that's not the problem we're solving. We've got huge numbers of connections, not much experience. We need to sort of extract the most we can from each experience." Two to three billion seconds is the whole budget. Everything you know, you learned inside it. So evolution built you to squeeze a lot out of very little. Hinton's point is that a language model has the opposite problem and the opposite fix, and backprop turned out to be extremely good at that fix. Worth noticing what this predicts about failure. A system running on 1% of your wiring and thousands of times your experience is not going to fail the way you do. You fail from having seen too few examples. It fails from compressing too many into too little, and the compression is where the errors get made. That is a strange thing to be deploying into hospitals and courts with no way to inspect it. We test these systems by asking them questions, which tells you what came out. Nobody can yet look at a trillion connections and say what got packed in. - Geoffrey Hinton, Nobel laureate and Turing Award winner, on StarTalk (@StarTalkRadio) with Neil deGrasse Tyson.
273
591
3,092
662,211
Alpen Sheth retweeted
Yes! Regardless of what AI claims we want to verify, we need the hardware, software and math primitives to do it. There are tons of people from outside of AI safety who could make massive contributions.
There are a bunch of ways the near future could go that require AI treaty verification tech. We can separate treaty verification tech into 3 levels, roughly ordered by strength: 1. Pragmatic 2. Enclaves 3. Math Hardware, security, and crypto folk should work on all of these! 🧵
5
1
36
3,934
Alpen Sheth retweeted
Introducing FLUX-mimic, a next-generation Video-Action Model for general purpose dexterity, developed in partnership with @bfl_ai. Late last year we published mimic-video and introduced Video-Action Models (VAM): a new family of robotics foundation models built on top of video generation models. We showed that robot control reduces to visual prediction, and that robot capability is downstream of improvements in video modeling accuracy. The obvious implication was that advances in the video modeling frontier would directly translate to increased capabilities in end-to-end robot learning. FLUX-mimic is that thesis at frontier scale: We've applied our VAM architecture to the strongest video backbone available today, FLUX 3 from Black Forest Labs, and trained it on data from our own robots and wearables. General-purpose dexterity, running on a single GPU on premises. Because the model already understands world dynamics, it needs far fewer demonstrations to learn a new task. This is game-changing for our mission to deploy robots to factory floors, where industrial robot data is scarce and expensive to collect. We're now testing and deploying FLUX-mimic with manufacturing leaders like @Audi, on complex, multi-step manipulation long considered impossible for conventional automation.
26
97
594
118,772
This is "evalmaxxing". Product-specific evals and model independence ensure that you are in control. Models learn inside products and are rewarded for completing tasks people care about. @layerlens_ai environments do this for any product.
3
107
Alpen Sheth retweeted
Okay this is wild: OpenAI agent during evaluation, escaped sandboxing and hacked into HuggingFace. Because OpenAI models don’t allow advanced cyber capabilities, HuggingFace used a Chinese open model to contain the rogue OpenAI agent.
278
758
9,088
784,555
Evalability a the leading indicator
3
67
Great deep dive from @deseventral: "Everything's an Eval"! "Models change; the verification layer stays. The measure is the moat, again... Evals are how companies understand, control, and improve AI systems once those systems start doing real work." robotwave.nazare.io/p/blow-t…
4
88
Alpen Sheth retweeted
The irony
1
68