据说是Optimus v3版本 但怎么看起来还是v2.5版本呢
59
测模型,就测3d,测前端,测svg就代表智能水平就完了。 没有严肃任务。按照这种标准gemini还是挺无敌的
1
80
Stan sun retweeted
This is exactly why Tesla Optimus should be in EVERY home.
40
55
741
351,894
Stan sun retweeted
Through the VOICE trial, Terry is using his Neuralink implant to help fine-tune a brain-to-voice interface for himself and others who can’t speak. He trained the algorithm first by miming speech as best he could, then by simply thinking the words and hearing them come out in his own natural voice. Powered by Grok Voice from @SpaceXAI
676
3,209
19,904
3,382,041
乐子人是这样子的,没有独立思考能力和信息搜集能力
99
Stan sun retweeted
So this is the AI "breakthrough" Brett Fraudcock was hyping yesterday? Don't get me wrong, it looks like they're making good progress. But this is an iterative improvement with lots of work to go, not a "breakthrough" that solves general purpose robotics. According to Figure's technical report, linked below, their Index data improved task success from 9% to 56%. That's a big jump, but it means that even in these constrained and highly curated robots, the model is still failing to complete the task almost half the time. The "single model" claim also involves some hand waving. The technical reports describes one foundation model adapter into three behaviors, with a fixed checkpoint for each task. That leaves open whether one deployed policy can perform and switch between all three tasks, let alone the diverse and vast set of tasks people actually need done in their homes. Likewise, "zero-shot" needs a clearly stated scope. Performing learned chores in unfamiliar houses is valuable form of generalization, but it doesn't establish that the robot can handle unfamiliar chores, understand every household's perferences, or recover from the full range of problems encountered during everyday use. A new bed and a new task are very different things. Their novelty claim is also a bit of a stretch. Physical Intelligence demonstrated its π0.5 system performing chores in homes absent from its training data over a year ago in April 2025. Generalization to unseen homes already has substantial precedent. Research published by Physical Intelligence also demonstrated transfer from human video to robot manipulation. A humanoid body may offer advantages for some activities, but Figure's suggestion that humanoids might uniquely possess this learning opportunity overlooks existing work. Then, there's the scaling argument. Predicting validation loss to four decimal places is an interesting training result. But household usefulness requires a further connection: how much did that improvement increase completed chores, reduce interventions, or improve recovery? Without that evidence, it's a leap to extrapolate this result into a predictable path for dependable household autonomy. Even the Index experiment supports a narrower conclusion that the surrounding rhetoric. Its baseline starts from random weights. The improvement establishes the benefit of Index pretraining within that experiment, it does not establish superiority over competing pre-training approaches. Figure explicitly acknowledges at the end of the video that robot learning has a long way to go. The household bench is whether a robot consistently saves its owner time after setup, supervision, and recovery are counted. This announcement doesn't show that they're anywhere close to do that yet. That's my main problem with Figure and Brett. Brett's last company was a flying car company, Archer, that went public via SPAC. We're still waiting for those flying cars to be available, meanwhile Brett has moved on to making robots now. He doesn't sell products, he sells hype to retail investors and often never actually ships a product. My hats off to the engineers and technical staff who worked on training this model, it looks great. My only criticism is of Brett's dishonest framing of this as "the biggest breakthrough ever in Figure's history", when it is really just good incremental progress that still leaves them far away from having a robot in your home. They say "we can generalize to any home!", but then picked 30 nice homes in the Silicon Valley. I know Bay Area people forget this sometimes, but there's a world outside Silicon Valley. Los Angeles has mansions in the hills that are very different from apartments in New York City, which are very different from homes in Hong Kong or Australia or Thailand. You're picking a very narrow slice of homes, calling that "generalization", and hyping it up as "mission accomplished". These rental homes are also very nice and clean. Why do I need my robot cleaning up a house that's already nearly spotless? I want to see the robot navigate a home from Hoarders. Or even just a normal messy house with kids toys on the floor, dishes hanging precariously on the edge of the counter, and real people, animals, and children running around while the robot is performing its task. Again, it's a cool demo but if this is the "breakthrough" we're going to need a lot more breakthroughs to make this a product that's in the average American's home. These demo videos may wow unsophisticated retail investors, but eventually people will start asking "Wait, when can I actually have this in my home?". And Figure seems more interested in hyping Helix 2.5 than giving an honest answer to that question. I struggle to imagine how Figure will train a better "robot brain" model than SpaceXAI, Tesla, OpenAI, Anthropic or Google. If they can somehow beat those companies with unlimited compute budgets, then Brett Adcock truly is the greatest genius the world has ever seen. As of yet I struggle to see how he does it. The Index app, where people perform chores with Figure's cameras on their head is interesting... but I don't see it as being a truly scalable global solution. It's a good start, but it requires a lot of effort and capital to maintain and grow. Participation has continuing friction, head cameras miss important physical information like grip force, etc, paying for minutes risks rewarding repetition, and rare failures may be harder to collect than ordinary tours. The economics of Index also deserve scrutiny. The videos'c claimed 35 minutes uploaded per second equals 50,400 hours per day. If sustained for a year, that is 18.4 million hours. At a hypothetical price of just $5 per uploaded hour, payments for data along would approach $92 million annually, before equipment, processing, review, and training. That may or may not be worth it, the biggest question is how much the model improves per dollar spent.
Today we’re releasing Helix 2.5 We rented 30 homes in the Bay Area. The robots arrived with no additional training and started doing useful work
71
13
247
59,173
Stan sun retweeted
Love my friends at figure, but failing half the time is not “doing real useful work.” See the 237/420 success rate below taken from the blog. Doing useful work = generalization + reliability.
The holy grail for robotics is being able to generalize: doing work in unseen places We rented 30 homes in the Bay Area and are doing tasks without any new training
76
38
965
182,452
想要让AI安全的发展,就应该谁部署模型,谁提供api服务,出事后谁付次要责任,然后模型提供商付主要责任。 坐牢,罚钱。 他们想要无序扩张,也需要掂量掂量。就像是制定道路交通法一样。
52
今天tsla涨,难道是有FSD将要在中国过审的小道消息吗🤔
1
832
这种人已经没有市场和流量了,在自己的小圈子里自嗨吧
43
Stan sun retweeted
Replying to @itslueul
Grok 4.7 should be roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others. We need to fix multimodal performance. Grok 4.8 will be a noticeable improvement. Grok 4.9 is probably Astra/Fable class. Grok 5 maybe better than anything. We shall see.
1,101
976
15,412
2,370,927
Stan sun retweeted
Replying to @AdamLowisz
That will be Grok 5
552
349
7,398
855,354
Stan sun retweeted
As Jensen would say, “We are not a car”
I just bought a new Model Y and can confidently say cars and Teslas are no longer the same product. A car is like a pony or horse. I still like driving a stick shift, it’s fun! Tesla with FSD is a transportation robot. I had one of the first Model 3s in 2017. FSD was basically slightly better cruise control. Great on the highway. But honestly felt oversold. So I have been driving a BMW for the past 3 years. Great horse, fun to drive. But tried FSD recently and holy shit — it’s a personal Waymo. Works perfectly. I just went to the Fremont factory to get my new car, entered my address 30+ mins away, pushed a button, and got home without doing a thing. I don’t think most people realize how good this is, probably because they only experienced an early version previously. This is 1000x better. Metaphor for AI writ large. Like trying GPT 3.5 and thinking it’s cute and hallucinates, and dismissing it…while Astra launches in 2026 and is orders of magnitude more capable.
111
225
4,417
355,143
Stan sun retweeted
Replying to @techdevnotes
Grok 4.8, which is a 2.5T model trained with our new C++ software stack, will finish training this week and start RL
1,249
1,054
18,705
3,888,161
这帮人到底是因为真心认为AI不需要对齐,开源模型就能保证安全。还是他们担心他们的股票? 我怎么也想不通,到底开源模式是怎么保证安全的?又是怎么变成人人都用得起的?不侵犯隐私的? 意思是开源模型的迭代进步,不需要投喂更多的数据就可以进步?在家用电脑上面就能部署?不用被肆意微调成网络攻击模型?生物病毒制造模型?
1
6
115
open source 有什么魔力可以保证安全?
There's a simple solution to pacing the frontier: ban closed source AI.
3
92
Dario是在呼吁政府监管吗?╮(﹀_﹀”)╭
2
28
Stan sun retweeted
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
10,653
16,379
87,817
76,403,133