@SemiAnalysis_ recently published a piece on where robotics companies are putting their inference.
Robot models are getting bigger, and on-robot compute isn't beefy enough to keep up. Jetson Thor is the top-of-market solution, but it has roughly a tenth of the FLOPs of a GB200. So companies are weaving between two approaches: either shrink the model to fit on the robot, or separate the "hard thinking" offboard.
They surveyed five different companies on how they did it:
𝟭)
@BostonDynamics
BD has split the stack. The visuomotor policy runs onboard on a Jetson Thor. The reasoning layer (which they say is similar to Gemini Robotics ER) runs in the cloud on Google TPUs. Their robot's tasks require sufficient generality, which means they need a frontier model. To make this split work, they've dedicated multiple teams entirely to networking.
𝟮)
@AgilityRobotics
Agility keeps inference on their humanoid, Digit, on Jetson-class compute. The only decision-making that goes into the cloud is orchestration between robots. Agility runs on-board inference to meet customer requirements. Overcoming networking/IT constraints is hard at brownfield sites; off-board inference also raises safety concerns.
𝟯)
@VerneRobotics
Verne runs everything locally on a Jetson or RTX 50-series. Verne bets that warehouse picking and packing doesn't need open-world reasoning, so a few-billion-parameter policy is sufficient.
𝟰)
@SundayRobotics
Sunday was originally built for cloud inference, but when they deployed ACT-2 into homes, they quickly moved everything onboard within a few days. For them, latency was fine, but jitter killed them as the robot moved through natural dead spots inside a house. To have 99% accuracy, the robot simply can't handle that much jitter.
𝟱)
@WeaveRobotics
Weave also runs inference on-board, but stays connected to the network anyway. Training data is constantly uploaded to the cloud, and teleoperators stand by to jump in if a robot fails a task.
The pattern I'm seeing, both in this post and across the industry, is that companies keep inference on-board, but the more general reasoning you need, the more impetus you have to split up your workload to use cloud compute.