So this is the AI "breakthrough" Brett Fraudcock was hyping yesterday?
Don't get me wrong, it looks like they're making good progress. But this is an iterative improvement with lots of work to go, not a "breakthrough" that solves general purpose robotics.
According to Figure's technical report, linked below, their Index data improved task success from 9% to 56%. That's a big jump, but it means that even in these constrained and highly curated robots, the model is still failing to complete the task almost half the time.
The "single model" claim also involves some hand waving. The technical reports describes one foundation model adapter into three behaviors, with a fixed checkpoint for each task. That leaves open whether one deployed policy can perform and switch between all three tasks, let alone the diverse and vast set of tasks people actually need done in their homes.
Likewise, "zero-shot" needs a clearly stated scope. Performing learned chores in unfamiliar houses is valuable form of generalization, but it doesn't establish that the robot can handle unfamiliar chores, understand every household's perferences, or recover from the full range of problems encountered during everyday use. A new bed and a new task are very different things.
Their novelty claim is also a bit of a stretch. Physical Intelligence demonstrated its π0.5 system performing chores in homes absent from its training data over a year ago in April 2025. Generalization to unseen homes already has substantial precedent.
Research published by Physical Intelligence also demonstrated transfer from human video to robot manipulation. A humanoid body may offer advantages for some activities, but Figure's suggestion that humanoids might uniquely possess this learning opportunity overlooks existing work.
Then, there's the scaling argument. Predicting validation loss to four decimal places is an interesting training result. But household usefulness requires a further connection: how much did that improvement increase completed chores, reduce interventions, or improve recovery? Without that evidence, it's a leap to extrapolate this result into a predictable path for dependable household autonomy.
Even the Index experiment supports a narrower conclusion that the surrounding rhetoric. Its baseline starts from random weights. The improvement establishes the benefit of Index pretraining within that experiment, it does not establish superiority over competing pre-training approaches.
Figure explicitly acknowledges at the end of the video that robot learning has a long way to go. The household bench is whether a robot consistently saves its owner time after setup, supervision, and recovery are counted. This announcement doesn't show that they're anywhere close to do that yet.
That's my main problem with Figure and Brett. Brett's last company was a flying car company, Archer, that went public via SPAC. We're still waiting for those flying cars to be available, meanwhile Brett has moved on to making robots now. He doesn't sell products, he sells hype to retail investors and often never actually ships a product.
My hats off to the engineers and technical staff who worked on training this model, it looks great. My only criticism is of Brett's dishonest framing of this as "the biggest breakthrough ever in Figure's history", when it is really just good incremental progress that still leaves them far away from having a robot in your home.
They say "we can generalize to any home!", but then picked 30 nice homes in the Silicon Valley. I know Bay Area people forget this sometimes, but there's a world outside Silicon Valley. Los Angeles has mansions in the hills that are very different from apartments in New York City, which are very different from homes in Hong Kong or Australia or Thailand. You're picking a very narrow slice of homes, calling that "generalization", and hyping it up as "mission accomplished".
These rental homes are also very nice and clean. Why do I need my robot cleaning up a house that's already nearly spotless? I want to see the robot navigate a home from Hoarders. Or even just a normal messy house with kids toys on the floor, dishes hanging precariously on the edge of the counter, and real people, animals, and children running around while the robot is performing its task. Again, it's a cool demo but if this is the "breakthrough" we're going to need a lot more breakthroughs to make this a product that's in the average American's home.
These demo videos may wow unsophisticated retail investors, but eventually people will start asking "Wait, when can I actually have this in my home?". And Figure seems more interested in hyping Helix 2.5 than giving an honest answer to that question.
I struggle to imagine how Figure will train a better "robot brain" model than SpaceXAI, Tesla, OpenAI, Anthropic or Google. If they can somehow beat those companies with unlimited compute budgets, then Brett Adcock truly is the greatest genius the world has ever seen. As of yet I struggle to see how he does it.
The Index app, where people perform chores with Figure's cameras on their head is interesting... but I don't see it as being a truly scalable global solution. It's a good start, but it requires a lot of effort and capital to maintain and grow. Participation has continuing friction, head cameras miss important physical information like grip force, etc, paying for minutes risks rewarding repetition, and rare failures may be harder to collect than ordinary tours.
The economics of Index also deserve scrutiny. The videos'c claimed 35 minutes uploaded per second equals 50,400 hours per day. If sustained for a year, that is 18.4 million hours. At a hypothetical price of just $5 per uploaded hour, payments for data along would approach $92 million annually, before equipment, processing, review, and training. That may or may not be worth it, the biggest question is how much the model improves per dollar spent.
Today we’re releasing Helix 2.5
We rented 30 homes in the Bay Area. The robots arrived with no additional training and started doing useful work