Real-world data for physical AI

Pinned Tweet
Physical AI has a data problem. To operate in the real world, models need to learn from human experience. The scale of what that demands is increasingly clear: ↳ $7T potential humanoid market by 2050 ↳ $11.9B → $49.8B AI training dataset market by 2031 ↳ ~100K hours of training data today, with tens of millions being targeted Meeting that scale will require vastly more training data than exists today. It’s where we’re increasingly focused at Poseidon.
40
8
114
30,039
Data in the Wild, ep. 002. Task context •- Clean a small aquarium + replace water •- 10 minutes •- A home in the Philippines Capture •- Handheld monocular Samsung SM-A135F •- 1280×720, H.264 Embodied behavior •- Two-handed handling across glass, water, tools + small loose objects •- Actions: remove_decor, transfer_water, rinse_vessel, replace_decor, fit_lid, wipe_surface, return_container •- Challenges: transparent water, reflective glass, wet-tool handling + long-horizon sequencing Video collected from Numo, our data collection app that sources real-world data. Poseidon then refines these captures for model training, with a focus on long-tail tasks.
13
4
47
4,907
Data in the Wild, ep. 001. The data needed to train robotics is rare, unscripted, and full of long-tail behaviors. Here’s what that looks like in practice. Task context: •- Brush the pet, wipe its paws & body, collect loose fur •- 11 minutes •- A home in the Philippines Capture: •- Handheld monocular iPhone •- 1920×1080 H.264 Embodied behavior: •- One hand stabilizes or repositions the dog while the other brushes •- Actions: stabilize_pet, brush_coat, reposition_pet, set_down_brush •- Challenge: Moving occlusion, careful hand manipulation, unpredictable animal motion
10
5
62
4,604
Video sourced from Numo, Poseidon’s data collection app. Numo sources real-world data that existing datasets don’t capture. Poseidon then processes that data for model training. Pet brushing + paw wiping: ~19 hours in Numo vs. 0 across 12 public egocentric video datasets we analyzed.
5
8
856
Physical AI has a data problem. To operate in the real world, models need to learn from human experience. The scale of what that demands is increasingly clear: ↳ $7T potential humanoid market by 2050 ↳ $11.9B → $49.8B AI training dataset market by 2031 ↳ ~100K hours of training data today, with tens of millions being targeted Meeting that scale will require vastly more training data than exists today. It’s where we’re increasingly focused at Poseidon.
40
8
114
30,039
We work with AI teams to design and collect speech, spatial, sensor, and behavioral data in order to improve the models that power physical AI. More on what we’re building ↓ psdn.ai
2
10
819
Numo is now mobile-first. Contribute from anywhere with new video tasks, audio tasks in more languages, and faster withdrawals. Live today on iOS and Android ↓
56
12
144
13,855
Numo on web is winding down as we move everything to mobile. Download and keep contributing. iOS → apps.apple.com/app/id6783542…
2
4
1,309
Conversational voice data exposes a subtle failure mode in many ASR pipelines. The metric says the data is bad, but human reviewers say the audio is clear. Often the issue isn’t transcription quality itself, but that the evaluation stack wasn’t built for turn-taking, silence, and multi-speaker structure. Learn how to solve this in our blog:
21
6
63
6,843
A fascinating look at the challenge ahead for physical AI. Exactly what we're building: the pipeline to collect high-quality real-world data at the scale physical AI demands.
World Labs CEO Dr. Fei-Fei Li & SceniX Co-Founder Yunzhu Li on the data bottleneck in robotics and how world models help: "The lack of data in training, the lack of data in evaluation, this is very, very different from language models, where data is abundant on the internet." "And we know that in order for robotics to work, we have to somehow unlock the power of scaling law. But where does that come from?" "It's a profound problem that everybody's battling with in robotics." "We see a lot of unlock in being able to do this whole process through the modeling of the environments." "Being able to create these digital worlds... that's just going to unlock so much more potential for being able to replace all the costly and the unsafe data in the real environments with the data generated from the worlds for robots to be able to do scalable learning and evaluations." @drfeifei @YunzhuLiYZ @martin_casado
24
6
46
7,581
Poseidon retweeted
In today's AI landscape, every dataset must pass three critical questions: • Can you source it at scale? • Can you prove where it came from? • Can you trust its quality? DATA Network is built to answer all three.
41
17
113
22,189
Poseidon retweeted
For years, every voice AI conversation passed through text. Speech ▸ text ▸ model ▸ back to speech. Now, speech models cut the middleman, with far more realistic results. But while the architecture has made leaps, one bottleneck remains: data. Enter: @otodotearth x @psdnai
30
18
114
21,281
Quality audio is necessary for earning on Numo. For the best submissions possible, let's cover: 1. Recording environment 2. Voice clarity 3. Technical considerations ↓ ↓ ↓
44
11
106
467,014
Technical considerations: • Use your device's built-in mic or headset • Test your mic before recording • Make sure your mic is not covered by a phone case • Keep your device stable to avoid handling noise • Ensure you have a stable internet connection
1
11
1,547