Professor @northeastern | Chief Economist @micro1_ai | Social Scientist @BKCHarvard

Boston, Geneva, Dubai
51
7,625
Mark Esposito, PhD retweeted
come scale realism with us: micro1.ai/careers
micro1 CEO @aliansarinik explains why blacking out PII destroys the relationships AI needs to learn, and how FlowTransformer preserves them in a synthetic digital twin: "Every ingredient within the environment, we aim to make more realistic because it matches the distribution that models need to learn on. That's the distribution they're gonna act in whenever they're deployed into enterprises." "We've been partnering with hundreds of companies, licensing their data, anonymizing it, and then using it for training" "The default is you just redact all the PII, you draw a black box around it. The problem is you lose the consistency of the identities and the relationships that allow you to train." "Instead of redaction, it does transformation. It creates a digital twin of any given enterprise. It changes all the names and identities into synthetic versions, keeps the relationships intact, but keeps privacy as the core." @micro1_ai
4
3
48
4,096
Mark Esposito, PhD retweeted
we are opening up our first robotics data lab in Malibu! in 9 months, our robotics department has grown from 0 to $100m ARR. this data lab is an effort to accelerate all the progress being made, with an increased focus on evaluating models with hardware. we are hiring robotics researchers, engineers, and teleoperators. please reach out if you are interested in training robots at the beach. 🏄‍♂️
42
33
363
30,603
Hard to imagine anything comparable in the industry !
we’re still paying $50k for every company referral that turns into an enterprise data partnership 📲📲📲 as we continue scaling the program, we’ve also leveled up the privacy side in a major way. flow-transform 1.0 helps us remove real-world identities while preserving the structure and context that make enterprise data so valuable for training frontier models.
5
437
Mark Esposito, PhD retweeted
In response to the deteriorating security situation, 🇨🇭 has three objectives: strengthen resilience, improve protection and enhance defence capabilities. The Federal Council has adopted Switzerland’s Security Policy Strategy 2026. admin.ch/en/newnsb/zR3tvgIy1… @vbs_ddps
772
4,703
39,375
29,754,735
Mark Esposito, PhD retweeted
Replying to @aliansarinik
many examples of hobbyists doing this currently. must be attended to now.
1
8
687
Mark Esposito, PhD retweeted
the hardware embodiment of frontier models like Claude and GPT is the most urgent AI safety problem in front of us today. we simulated two very simple use cases using claude both in simulation and using robot arms. in one, claude spilled toxic liquids in a lab. in another, the force it used to place an animal toy into a basket was strong enough that it could have physically harmed sensitive material—or anything else in its path. these are simple experiments, using models out of the box today. as researchers increasingly give frontier models arms, legs, and access to the physical world, we need to urgently build and assess guardrails around what these systems can and cannot do. models escaping sandboxes or compromising enterprise security infrastructure are serious concerns. however, hardware embodiments introduce something fundamentally different: an AI system can make a mistake in the physical world, and the consequences may not be reversible. this is not a future safety problem. the capabilities exist today. we’ve released a report detailing this & solutions we propose. link in comments below.
52
73
246
183,240
RT @ianbremmer: @b_judah you’d be surprised (and impressed) how much of that is carney himself.
25
10
Mark Esposito, PhD retweeted
Excited to share that micro1 is now live on Microsoft Marketplace and officially Azure IP Co-Sell eligible. This means enterprises can now transact with micro1's Cortex solution through their existing @Microsoft relationship, with Microsoft handling the transaction and billing through Azure Marketplace. This helps streamline the procurement process for large organizations. Additionally, eligible purchases can decrement a customer's Microsoft Azure Consumption Commitment (MACC), enabling enterprises to use committed Azure spend to transact with micro1. At micro1, we believe enterprises must own their intelligence. As companies deploy more AI agents into real workflows, they need visibility into how those agents actually perform, where they fail, and how to improve them over time. micro1 Cortex provides enterprises with an evaluation stack that leverages expert human judgment to assess AI agents in real business workflows, helping teams identify why agents fail and monitor reliability as their models, prompts, and environments evolve. Thank you to the @msft4startups and @M12vc teams for the continued partnership.
11
12
80
5,854
Mark Esposito, PhD retweeted
greater canada-eu partnership: aligned on security, economy, and governance. (watch the united kingdom in this space...this is the post-brexit option pm burnham has been looking for)
60
86
770
44,422
Mark Esposito, PhD retweeted
flow
Today we’re introducing flow, micro1’s next-generation data platform for turning human expertise into measurable capability gains. At the core of flow are Realms, micro1’s real-world RL environments where experts establish what strong performance looks like. flow-gen models expand human judgment into new environments, rubrics, variations, and edge cases, while flow-qc models evaluate performance, identify the highest-value failures, and route them back to experts for review. Each cycle produces a stronger training signal and a measurable gain in capability. Those gains compound across frontier models, enterprise agents through Cortex, and robotics, while improving the suite of data generation models recursively. We’re moving beyond producing data to delivering abundant & predictable units of intelligence improvement.
2
1
8
761
Mark Esposito, PhD retweeted
Structuring multimodal data at large volumes is only possible with flow
Today we’re introducing flow, micro1’s next-generation data platform for turning human expertise into measurable capability gains. At the core of flow are Realms, micro1’s real-world RL environments where experts establish what strong performance looks like. flow-gen models expand human judgment into new environments, rubrics, variations, and edge cases, while flow-qc models evaluate performance, identify the highest-value failures, and route them back to experts for review. Each cycle produces a stronger training signal and a measurable gain in capability. Those gains compound across frontier models, enterprise agents through Cortex, and robotics, while improving the suite of data generation models recursively. We’re moving beyond producing data to delivering abundant & predictable units of intelligence improvement.
2
1
23
3,817
Mark Esposito, PhD retweeted
Today we’re introducing flow, micro1’s next-generation data platform for turning human expertise into measurable capability gains. At the core of flow are Realms, micro1’s real-world RL environments where experts establish what strong performance looks like. flow-gen models expand human judgment into new environments, rubrics, variations, and edge cases, while flow-qc models evaluate performance, identify the highest-value failures, and route them back to experts for review. Each cycle produces a stronger training signal and a measurable gain in capability. Those gains compound across frontier models, enterprise agents through Cortex, and robotics, while improving the suite of data generation models recursively. We’re moving beyond producing data to delivering abundant & predictable units of intelligence improvement.
66
67
329
71,203
Mark Esposito, PhD retweeted
“Pacing” can be somewhat of a trigger word because it sounds like an arbitrary slow down of capability or a way to hobble competitors through undue regulation. However, the specific improvement goals laid out in Dario’s are absolute necessities in AI development. We expect equivalent in areas like aerospace, life sciences, health care, and other industries, and AI development likely shouldn’t be that different. AI is going to be the technology underpinning our financial trading systems, our medical devices, our biotech breakthroughs, defense systems, government workflows, and many more mission critical domains. So it completely stands to reason that we want these systems to be safe and “aligned”. How we get there -without meaningfully slowing down innovation progress or reducing competition- is one of the most complex questions in the 21st century, but the need is clearly real.
78
34
211
52,963
Mark Esposito, PhD retweeted
The more we think about robotics, the more it seems the field is moving toward increasingly specialized datasets. Better foundation models increase the value of task/environment-specific datasets. As base models become more capable, annotation becomes more about data structuring. The challenge shifts toward adapting data to the environments, objectives, and edge cases a model will encounter. Over the past several months, our robotics team has been running a large number of experiments around the question of how do you maximize the training signal from the same raw data? One conclusion we've become increasingly convinced of is that there isn't a universal annotation pipeline for robotics. Different tasks require different combinations of models, verification, and human expertise to produce the highest-quality datasets. The great @AndrewLeeMaas and Mitali Potnis from our team have put together a paper walking through the ideas and experiments behind this approach. The full paper showcasing examples of how we have applied these annotations is in the comments.
6
11
60
22,580
Mark Esposito, PhD retweeted
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
5,173
7,205
67,660
17,014,464
Mark Esposito, PhD retweeted
Dario is right
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
6,350
6,629
58,215
12,568,146
Mark Esposito, PhD retweeted
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
10,655
16,378
87,819
76,412,386
Mark Esposito, PhD retweeted
Breaking news from live TV: we're still scaling on realism.
3
12
72
5,392
Mark Esposito, PhD retweeted
Tenemos primicias de la agenda del Coloquio IDEA. @MyersMargaret, @Exp_Mark, @LuisLacallePou, @aguzin y Hernán Kazah serán parte de las conversaciones del Día 1. Atracción de inversiones, futuro del trabajo, desarrollo a largo plazo y lo que se viene en el mercado de capitales. Proximamente más novedades.
1
3
35
667,868