I had Opus 5.5 recreate the city of San Francisco in Unreal Engine. I powered everything in it (people, pets, cars, traffic, literally everything) with Jev.
192
104
2,210
366,358
I created a custom tutoring plan for my son in math by loading his school's assessment into ChatGPT. Includes a reward system. So thankful for AI.
40
9
352
20,687
I have no idea how pass keys work and at this point I'm too afraid to ask.
118
19
758
72,417
Eight Sleep Pod 6 👀👀 One of the best purchases I've ever made for my quality of life. I am such a big fan, been a customer for over 2 years. Super proud I now get to work with Eight Sleep (this post isn't sponsored). Get $200 off with code MATTHEWBERMAN
BREAKING: Pod 6 is live. The sixth generation of our intelligent sleep system, redesigned for every kind of bed. ✅ The new Hub is small enough to fit under your bed frame ✅ More powerful than ever ✅ With 9x more sensors for enhanced accuracy ✅ In new Solo sizes, at our lowest starting price yet This is our most significant hardware launch to date, and our most accessible. Available now at eightsleep.com
16
5
147
50,653
Opus 5.5 is the best model in the world. Full breakdown with special guest @trq212
28
20
394
23,956
This took about 2 days on /goal
I had Opus 5.5 recreate the city of San Francisco in Unreal Engine. I powered everything in it (people, pets, cars, traffic, literally everything) with Jev.
24
7
246
31,703
$25 in Jev usage btw I left it running over night accidentally LOL
I had Opus 5.5 recreate the city of San Francisco in Unreal Engine. I powered everything in it (people, pets, cars, traffic, literally everything) with Jev.
57
16
1,025
149,431
Opus 5.5 being the absolute best model on the planet AND being cheaper than Astra/Fable was not on my bingo card.
75
89
3,111
75,928
The @ForwardFuture team had early access to Opus 5.5 and Alex put together some of the most incredible demos I've ever seen.
I had early access to Opus 5.5 and it blew my expectations clear out of the water! Here are some things I made! Dark Souls
14
11
268
34,854
Open source will dominate token volume but frontier will accrue the value.
Looks like today may be a record day for token volume % of open models on Vercel AI Gateway: 🟦 Open 78.4% 🟨 Closed 21.6% While spend 💲 usually tells a different story, #3 and #4 today are Moonshot AI & DeepSeek. Adding Z⁠.ai, their combined spend surpasses OpenAI (#2). (Do note that's the spend for inference of the model across providers (mostly in the US), not revenue going directly to the open weight labs.)
33
8
104
16,231
And it will be named Chat.
34
4
359
111,028
We need to talk about Jev
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
51
153
1,775
138,520
Jev sorted 150,000 Skittles in seconds...
37
15
462
106,884
Jev making the right moral decision 100% of the time
I trust Jev with my life
48
15
630
107,262
Super excited about Jev
30
9
305
43,121
"Other labs can just do this too."
Last month I wrote about how we can build a positive and safe future for everyone: meta.com/thefutureisforevery… Every lab has the responsibility and incentive to move at the pace required to train its models safely, and the ability to take its own actions to ensure that happens. The reality is: - People won't want to use agents that are misaligned with them and that don't do what they ask, so labs have a strong natural incentive to make their models more aligned. There is a lot of debate about slowing progress on capabilities until alignment catches up. My view is that trust and alignment are quickly becoming the most important capabilities that will differentiate agents and models. Any lab that doesn't focus on alignment will fall behind. - Labs face significant liability if their models cause harm, so they have a strong incentive to prevent this as well. Meta delayed shipping Muse for several months to focus on safety and security. We didn't call for everyone else to do this before we would. We just did it as part of our day-to-day work because it was clearly the right thing for people and for us. I'm proud of the security foundations we've built. - Engaging independent evaluators and advisors is industry best practice. MSL already does this today in several areas because it helps produce better work. Other labs can just do this too. In general, it would be helpful for there to be a larger and more diverse ecosystem of evaluators. - Committing the significant majority of compute towards serving people rather than racing towards recursive self-improvement is one of the best ways to ensure we develop this technology safely. Meta has made this commitment and other labs can do this as well. I believe the key to building a positive future for everyone is maintaining the right balance of power. This is within our power to do.
18
8
229
20,224
Love this. Model safety and alignment is a feature and economically incentivized.
Last month I wrote about how we can build a positive and safe future for everyone: meta.com/thefutureisforevery… Every lab has the responsibility and incentive to move at the pace required to train its models safely, and the ability to take its own actions to ensure that happens. The reality is: - People won't want to use agents that are misaligned with them and that don't do what they ask, so labs have a strong natural incentive to make their models more aligned. There is a lot of debate about slowing progress on capabilities until alignment catches up. My view is that trust and alignment are quickly becoming the most important capabilities that will differentiate agents and models. Any lab that doesn't focus on alignment will fall behind. - Labs face significant liability if their models cause harm, so they have a strong incentive to prevent this as well. Meta delayed shipping Muse for several months to focus on safety and security. We didn't call for everyone else to do this before we would. We just did it as part of our day-to-day work because it was clearly the right thing for people and for us. I'm proud of the security foundations we've built. - Engaging independent evaluators and advisors is industry best practice. MSL already does this today in several areas because it helps produce better work. Other labs can just do this too. In general, it would be helpful for there to be a larger and more diverse ecosystem of evaluators. - Committing the significant majority of compute towards serving people rather than racing towards recursive self-improvement is one of the best ways to ensure we develop this technology safely. Meta has made this commitment and other labs can do this as well. I believe the key to building a positive future for everyone is maintaining the right balance of power. This is within our power to do.
30
7
209
16,665
Matan is unstoppable. Proud early investor 🤖
We have raised $200M at a $5B valuation to scale self-improving software development in the enterprise. @FactoryAI has grown to serve hundreds of thousands of developers at companies including RBC, Adobe, Nvidia, T-Mobile, and Palo Alto Networks. We will use this capital to accelerate our investments in research, product, and global go-to-market.
5
2
112
19,274