@stanford Professor of Computer Science, @simile_ai co-founder, nationally bestselling author. I build interactive, social, and societal tech.

Stanford, CA
Minds and machines!
1/ Introducing PhilosophyBench from @StanfordAILab @StanfordHCI, the first independent, large-scale benchmark for evaluating AI’s philosophical capabilities. philosophybench.org
21
3,289
Now’s prob a good time to mention that I’ll start as an Assistant Professor with appointments at @Kennedy_School and @HarvardEngineer in 2027, working on… AI evals! My lab will also work on monitorability, incident analysis, and verification to advance technical AI governance 1/
51
50
712
140,670
Michael Bernstein retweeted
Itaú used Simile’s behavioral simulation to turn a five-week research cycle into four days and a two-week concept exploration into under three hours. The work also won an ESOMAR Award for AI and automation in market research. simile.com/blog/itau-at-esom…
2
9
39
6,005
Michael Bernstein retweeted
People have realized the influence of AI companionship on our mental well-being. But how this happens remains a mystery. In our new preprint, we followed users for about one year. Our data reveal that greater social engagement with AI companions consistently predicts lower well-being, which is driven in part by lower real-world social interaction with other people.
13
44
200
35,561
In which the lecturer gives a lecture:
What does it take to build an AI model that can accurately simulate human behavior? You need to understand both what people do and why they do it. On Sept 24, @msbernst is taking us under the hood of behavioral simulation. We’ll explore why neither behavioral data nor surveys and interviews are enough on their own, what each reveals about human behavior, and how Simile brings them together to model what people might do next. RSVP here simile.zoom.us/webinar/regis…
2
1
37
3,675
Michael Bernstein retweeted
Simile has been selected amongst this year's IA40, the top private companies shaping the future of applied AI. Proud to be building alongside this group.
The 2026 Intelligent Applications 40 is here. Now in its sixth year, the IA40 recognizes the top private companies shaping the future of applied AI, from foundational infrastructure to full-stack intelligent applications. This year's list reflects a market maturing fast: → 450+ companies nominated and voted on by 72 investors across 54 leading firms → $410B raised by this year's 45 winners since founding — more than 25x what the inaugural 2021 class had raised at the same point → 23 repeat winners and 22 new entrants, a 49% turnover rate, down from 68% last year, as some of 2025's breakout companies start to entrench The big pattern this year is how companies are “Harnessing the Value of AI.” As frontier models become more capable and commoditized, the winners are increasingly the companies proving they can turn AI capability into real ROI, whether they're building the application, aggregating the models, or supplying the infrastructure underneath. Databricks remains the only company to appear on every IA40 list since we launched it in 2021. Congratulations to all the 2026 IA40 winners. Your work is defining what's possible in this new era of intelligent applications. Explore all 45 winners and our analysis of the companies, categories, and funding patterns defining this year’s list → ordnl.link/RAFgnA7 Thank you to our great partners: @AWSstartups, @GoogleStartups, @Microsoft, @NYSE, @Delta, and @McKinsey.
3
8
48
6,414
Michael Bernstein retweeted
The "What-If Machine" is a vivid articulation of what Simile is building. It's not just about forecasting the future passively. It's about understanding how active interventions on the world change its course. It's the classic difference between causation and correlation.
Most big decisions are some version of a “what-if” question. What if we launched this? What if we raised the price? What if we didn’t? At Simile, we’re building a What-If Machine to simulate those decisions before making them for real. But getting the answer right requires understanding both what people do and why they do it. @msbernst explains what it takes to build one. simile.com/blog/what-if-mach…
5
6
98
21,117
I describe AI behavioral simulation as a “what-if” machine. I wrote about what a What-If Machine is, and how it works, at @simile_ai:
Most big decisions are some version of a “what-if” question. What if we launched this? What if we raised the price? What if we didn’t? At Simile, we’re building a What-If Machine to simulate those decisions before making them for real. But getting the answer right requires understanding both what people do and why they do it. @msbernst explains what it takes to build one. simile.com/blog/what-if-mach…
1
6
68
6,788
Michael Bernstein retweeted
For six months, Google Maps sent a slice of drivers in ten American cities down deliberately slower routes. About 30 seconds slower on average, and under 2% of trips were touched. The experiment ran on roughly 100 of the most congested road segments in each city, switching on and off day by day so each city acted as its own control. The results are out now in Nature Cities, from Google Research with collaborators at Berkeley and Stanford. The reasoning behind it is old and slightly counterintuitive. Every navigation app does the same thing: it finds you the fastest route right now, for you alone. When millions of phones do that at once they all pour onto the same handful of arteries, and those arteries stop being fast. Transport economists have been writing about the gap between what's good for one driver and what's good for the network since the 1950s. Nobody had tested a fix at city scale with real drivers, because you'd need to steer routing for a large share of a city's traffic, and about three companies on Earth are in that position. So Google put a penalty on the worst segments during their worst hours, which nudged the routing engine towards alternatives of similar road class and comparable travel time, then measured what happened across the whole network on weekdays between 7am and 8pm. On the targeted segments, traffic moved roughly 2% faster in the median city. Los Angeles got 4.56%, Atlanta 3.30%. Fuel burn on those stretches dropped between 0.5% and 1%. Across every road that saw a change in traffic, which covers around 80% of each city's driving, speeds rose 0.35%, reaching 0.5% during the morning and evening peaks. Total travel time on affected trips fell 0.69%. Their Bayesian model put the probability that the speed effect was genuinely positive at 99.8%. That adds up to more than 1,000 tonnes of CO2-equivalent saved per year, per city, in most of the places studied. Cars and vans account for around 10% of global carbon emissions, and the average driver spends about 2.6 years of their life behind the wheel, so a fraction of a percent across an entire road network is a serious amount of fuel. Per driver, the saving comes to about 0.25% of an average journey, roughly one fortieth of the normal day-to-day wobble in how long the same commute takes. Nobody in those ten cities noticed anything, in either direction, including the people sent the long way round. The authors are careful about what they haven't shown. They measured the immediate effect, not what happens once drivers work out the freed-up route is quicker and pile back onto it, which is the standard way road improvements get eaten. Their penalty scheme was deliberately crude, so the ceiling on the approach is unknown. And the underlying Google Maps data is commercially confidential, so nobody outside can check the numbers. The finding sitting underneath all this is that a private company's routing algorithm has become a piece of transport infrastructure, tunable in the way a traffic light or a congestion charge is tunable. The obvious follow-up experiments are the ones a city government would want to run, not the ones a mapping company would. link to full article: nature.com/articles/s44284-0…
116
404
4,672
1,702,650
🚨🚨🚨UIST Workshop Announcement: The Personalized Computer for the 21st Century Join us in Detroit on November 2nd for our morning workshop! 💻🤖☀️ personalized.computer/uist-2…
1
14
44
7,254
Michael Bernstein retweeted
Found in the @simile_ai office: @msbernst and @percyliang realizing, after over a decade, that they’ve been twins all along
5
10
323
32,824
Michael Bernstein retweeted
I just joined @simile_ai last week to help build out their confidence model! It's an exciting real world application of uncertainty quantification :) Check out our blog post to learn more about the progress that's been made so far, led by the amazing @wesel0
Anyone can simulate the future. But the simulation only matters if it’s trustworthy. At Simile, we train two types of models: simulation models and confidence models. Our first research blog post explores the origin of our proprietary confidence model, which predicts the accuracy of our population simulations and, in turn, makes them actionable. simile.com/blog/confidence?v…
7
2
127
30,322
Behavioral simulation requires confidence estimation. I'm 100% fine asking the system to simulate unlikely situations if it can signal lower confidence, so long as I can also get high confidence simulations and calibrate my trust.
Simile trains a simulation model. Also, we train another model that predicts the error of the simulation model (the “confidence model”). Tucked away in a secret dungeon below Mission Bay, we also have a “confidence confidence model” that predicts the error of the confidence model, but this one is saddled with a certain sinister statistical spirit that must be contained at all costs! To contain San Francisco’s crisis of confidence, Members of Technical Staff across the valley must come together and read Simile’s confidence blog post today!
1
2
42
6,266
Michael Bernstein retweeted
Simile trains a simulation model. Also, we train another model that predicts the error of the simulation model (the “confidence model”). Tucked away in a secret dungeon below Mission Bay, we also have a “confidence confidence model” that predicts the error of the confidence model, but this one is saddled with a certain sinister statistical spirit that must be contained at all costs! To contain San Francisco’s crisis of confidence, Members of Technical Staff across the valley must come together and read Simile’s confidence blog post today!
Anyone can simulate the future. But the simulation only matters if it’s trustworthy. At Simile, we train two types of models: simulation models and confidence models. Our first research blog post explores the origin of our proprietary confidence model, which predicts the accuracy of our population simulations and, in turn, makes them actionable. simile.com/blog/confidence?v…
6
13
87
15,390
Our new @PNASNews brief report on the value alignment of social media feeds is now live on @ConversationUS: theconversation.com/why-soci… Your For You algorithm is likely misaligned with your values, even when it's ranking accounts that you actively follow. Work led by @_ziv_e!
1
7
24
2,448
While the posts you follow on social media are value-aligned, those promoted by For You are misaligned. As a possible explanation: folks are especially to reply to value-misaligned content, even amongst accounts they follow. The algorithm misinterprets it as positive engagement.
1
1
383
Michael Bernstein retweeted
Excited to share that our paper is now published in @NatureHumBehav! How is AI companionship related to user well-being? Using survey data from 1,000+ users and 400,000+ chat messages, we find that using chatbots primarily for companionship is significantly associated with lower psychological well-being. nature.com/articles/s41562-0….
There’s a lot of speculation around how AI will change human relationships. To dig into this question, we collect surveys from 1000+ Character.AI users and 400,000+ messages to analyze the relationship between AI companionship and well-being. Preprint: arxiv.org/abs/2506.12605
2
10
53
8,456