When people talk about robotics, they usually talk about models, data, or hardware. Few people talk about the infrastructure that lets you iterate on all three quickly. Today we're publishing how we trained Dyna-2 on over 1,000,000 hours of egocentric video, repeatably. At this scale, most of what worked at ten thousand hours did not hold up:
• ingestion throughput was capped at 14,000 episode-hours per week — a million hours would have taken over a year
• building a training manifest took 48 hours before a run could even start
• reading a petabyte from cloud storage during training left GPUs exposed to latency and packet loss
🧵