Data and Simulation Infrastructure for embodied AI

San Francisco
Introducing Midcentury. We’re building the data and simulation infra for physical AI. Today, we’re coming out of stealth with a $15M Series Seed to scale robotics beyond polished demos. We’re already supporting frontier labs with: → The world’s largest egocentric dataset: 2M+ hours, 50+ environments, 20,000+ tasks → Matrix: a frontier simulation platform scaled with our real-world data to evaluate and post-train policies at scale
97
207
1,204
612,559
We then match this with high-resolution derived signals. L/R action annotations at milli-second level accuracy. Subtask and cycle annotations are frame-aligned with millisecond timestamps, while semantic labels identify actions, objects, tools, and context.
1
9
238
Finally, our dataset collection process is the most diverse possible, spanning hundreds of locations and channels. 20,000 tasks across 50+ environments: from datacenter buildout and jewelry manufacturing to sports, home cleaning, and agriculture. These include short-horizon demonstrations as well as long-horizon, multi-step workflows.
1
9
250
How do you scale 2 million hours of ego data without sacrificing quality? Earlier this week, we revealed the world’s largest egocentric dataset. Our frontier pipeline has also turned it into the most labeled, annotated, and human QA’ed robotics ego dataset in the world. We’ve focused on some of the hardest problems in computer vision and ego data to ensure every clip contains SOTA action supervision directly helpful for training dexterous manipulation.
4
11
80
5,957
1) Two hands and repetitive movement. Robot models need generalization and multi-step, compound actions to learn effectively. Clips full of repetitive movements or lacking two hands provide less useful training signals. We go through every frame in several passes with internally trained hand recognition models to guarantee hands in 85%+ of all frames.
1
10
600
2) PII, faces and written Finding moments with faces or unnecessary occlusions across millions of hours is a significant screening problem. Our pipeline can identify and flag these moments for further review. We also screen out obvious issues: strobing lights, graininess, user error.
1
8
304
Introducing Midcentury. We’re building the data and simulation infra for physical AI. Today, we’re coming out of stealth with a $15M Series Seed to scale robotics beyond polished demos. We’re already supporting frontier labs with: → The world’s largest egocentric dataset: 2M+ hours, 50+ environments, 20,000+ tasks → Matrix: a frontier simulation platform scaled with our real-world data to evaluate and post-train policies at scale
97
207
1,204
612,559
Thanks to our partners, collaborators, and investors on this journey, including: @buildpbc, @eddybuild, @cegapereira, @justswart, @jay_drainjr, @liz_harkavy, @Shaughnessy119.
2
35
8,956
The scaling era for robotics is here. If you’re training, evaluating, or deploying physical AI, we want to hear from you. Explore dataset samples and get early access to Matrix below. midcentury.xyz.
2
2
35
7,236
Midcentury retweeted
ego data is starting to see real evidence it helps scale robotic models! 20,000 hours used in total, one of the largest pre-training sets here.... what if I told you that certain players were already scaling to 1 million 👀
We trained a humanoid with 22-DoF dexterous hands to assemble model cars, operate syringes, sort poker cards, fold/roll shirts, all learned primarily from 20,000+ hours of egocentric human video with no robot in the loop. Humans are the most scalable embodiment on the planet. We discovered a near-perfect log-linear scaling law (R² = 0.998) between human video volume and action prediction loss, and this loss directly predicts real-robot success rate. Humanoid robots will be the end game, because they are the practical form factor with minimal embodiment gap from humans. Call it the Bitter Lesson of robot hardware: the kinematic similarity lets us simply retarget human finger motion onto dexterous robot hand joints. No learned embeddings, no fancy transfer algorithms needed. Relative wrist motion + retargeted 22-DoF finger actions serve as a unified action space that carries through from pre-training to robot execution. Our recipe is called "EgoScale": - Pre-train GR00T N1.5 on 20K hours of human video, mid-train with only 4 hours (!) of robot play data with Sharpa hands. 54% gains over training from scratch across 5 highly dexterous tasks. - Most surprising result: a *single* teleop demo is sufficient to learn a never-before-seen task. Our recipe enables extreme data efficiency. - Although we pre-train in 22-DoF hand joint space, the policy transfers to a Unitree G1 with 7-DoF tri-finger hands. 30%+ gains over training on G1 data alone. The scalable path to robot dexterity was never more robots. It was always us. Deep dives in thread:
2
23
13,931