Pete Florence retweeted
We've reduced the time it takes to go from physical prompt → robot behavior. The faster anyone can teach a robot to do something new, the easier it becomes to scale physical work. Read more about GEN-1.5 in our blog post in the comments below.
8
27
236
26,735
Good question thanks Ani. Here's a video that gives a sense of the answer here about the effect of prompting: - in the first part, the model is just babbling with no prompt in context - at the end, we add a prompt for taking money out of a pouch (generalizing to the wallet)
Very cool work, Pete! I'm curious if you've tried an extreme version of this: no in-context example at all; just place objects in front of the robot and see if it can infer the task. I suspect this will have non-trivial success rates, which would also allow you to figure out how much of the benefit is actually coming from the prompt (rather than pure zero-shot task inference + capability). I remember Russ showed something like this in his Stanford talk on LBMs last year: piped.video/TN1M6vg4CsQ?si=-8J-…
17
28
270
32,910
Here's the prompt used at the end:
4
21
2,487
Last week @willknight came by and got a preview of one-shot prompting and we had some fun improvisational moments with the robot. Thanks Will for helping capture the moment.
I'm usually wary about talk of "the ChatGPT moment for robots" because language is less complex than the physical world. After visiting @GeneralistAI, though, where I saw their robots perform impressive manipulation with just a prompt, I think you can see it happening in a limited way. (Thanks to @peteflorence and co for having me!) wired.com/story/generalist-a…
2
3
42
9,165
Pete Florence retweeted
We started planning yesterday’s @generalistai GEN-1.5 announcement about 3 weeks ago. Two weeks ago, already feeling whiplash from new results, I asked ChatGPT to create this image to share in our internal meme channel.
Made with AI
3
6
41
3,845
Pete Florence retweeted
beautiful error correction behavior from a single human demonstration...
Replying to @GeneralistAI
GEN-1.5, our latest embodied foundation model, can learn new tasks prompted with 3 - 12 seconds of a single demonstration, no gradient updates or fine-tuning. It generalizes prompts to new situations, recovers from mistakes, and improvises new strategies to reach the same goal.
1
7
128
11,718
I rarely repost content, but this is a step-change for robotics research. As an embodied AI researcher, I find this to be a truly special moment for the field
Introducing GEN-1.5, a one-shot learner. It can learn new tasks in a few seconds. Show it what to do, and it generalizes. This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the physical world.
7
19
360
39,669
Pete Florence retweeted
At 10:06pm on Aug 3, I watched a robot do something I thought was years away. We were working on few-gradient learning with GEN-1.5, wondering how far we were from landing physical prompting: show the robot a task once, and it just does it. We YOLOed it, and the robot imitated exactly what we demonstrated, with 0 fine-tuning. We gave more prompts for different tasks. More successful rollouts landed. I could not believe my eyes. Looking back, this represents the culmination of a huge amount of work across the entire team, many failed experiments, and years of groundwork from the community. I just happened to be the one standing in front of the robot when it all clicked. @andyzengineer once asked me: "If Generalist failed tomorrow, what would we want to have left behind that could benefit humanity?" I believe we found one answer: an existence proof of physical prompting.
Introducing GEN-1.5, a one-shot learner. It can learn new tasks in a few seconds. Show it what to do, and it generalizes. This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the physical world.
43
90
1,158
108,253
Pete Florence retweeted
Introducing GEN-1.5, a one-shot learner. It can learn new tasks in a few seconds. Show it what to do, and it generalizes. This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the physical world.
320
1,688
12,166
3,417,184
Ever since I started working on robot foundation models a handful of years ago, the broad ability to one-shot in-context learn has been the single most vivid goal in my mind along the long road ahead. It’s something @andyzengineer and I have especially been thinking together on for several years now, and something shared by the whole team at Generalist as a major goal. But even before anybody started talking about robot foundation models, too, the capability this enables, i.e. generalized one-shot learning, has been both incredibly concrete and elusive. In terminology that I think many other people have come up with as well, in grad school we used to talk about a “Ctrl-C-Ctrl-V” type of capability. See something, and have the robot do it too. This same type of idea inspired the name of the 1970 MIT copy demo (people.csail.mit.edu/bkph/ph…). The thing is, it sounds simple, but is incredibly hard since the world is never quite the same when it got copied and where you want to paste it. The real world can be hard to predict and is full of variation. To do this, you need strong generalization, and it needs to be acquired in a single example. Doing this over a wide range of tasks, especially for dexterous tasks, is hard mode. Lots of the components of the idea of making this all happen have been there for a long time. As an example, this 2017 NeurIPS paper “one-shot imitation learning” arxiv.org/pdf/1703.07326 has excellent vision, with ambition well beyond what was achievable at the time, and although it’s not referred to as “in-context learning” since it was pre-Transformer, it actually uses attention to condition on a single demonstration. And now, many things have happened since early 2017, including the broadly celebrated arrival of one/few-shot in-context learning in language models in 2020. This new model GEN-1.5 takes in everything we have built and learned over the past couple years at Generalist. It has been training for 8 months. It has taken an incredible amount of commitment and grit from the whole team to get here. The level to which this model has survived many surgeries has continued to surprise me. And its capabilities have continued to surprise as well. We found compositional generalization on Friday. We found sim2real prompting earlier last week. We filmed the contiguous uncut videos of live prompting yesterday. To be clear, the success rates are modest, and there’s still a long way to go. But now I have definitely seen a ~decade-long imagination come into the real world.
Introducing GEN-1.5, a one-shot learner. It can learn new tasks in a few seconds. Show it what to do, and it generalizes. This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the physical world.
22
41
474
68,909
Pete Florence retweeted
Replying to @GeneralistAI
GPT3!
5
17
353
29,806
Butter, generalized
We've improved how GEN-1 learns to adapt to new actuators and new robots at the lowest level, with up to 10-20x gains on internal benchmarks. This significantly boosts performance on high-precision tasks like disassembling parts from a NIST board. Read more about GEN-1 in our blog posts in the comments below.
2
42
4,274
Yes, “all the hands”. Yesterday we showed a glimpse and announced we’re at over 9,000 variations so far available for pretraining. 2-finger, 5-finger, power tools, cooking tools, tape dispensers — it’s all good. Robot intelligence should be able to use them all. Amazing cooking across the full stack from the Generalist team!
Why create robot intelligence for just one hand, when we could have it learn from many? GEN-1, our latest embodied foundation model, now supports a broad range of end effectors from 5-finger hands, to specialized tools, and everything in between.
7
10
145
29,401
Pete Florence retweeted
Why create robot intelligence for just one hand, when we could have it learn from many? GEN-1, our latest embodied foundation model, now supports a broad range of end effectors from 5-finger hands, to specialized tools, and everything in between.
33
109
918
304,150
Pete Florence retweeted
Most robot demos are scripted. Generalist's GEN-1 is not. > GEN-1 was doing a task > Mid-task, an extra object was thrown into the bin > GEN-1 quickly adapted to the new environment > Still completed the task smoothly This is the difference between a scripted policy (seen in viral launch videos) and a real foundation model. GEN-1 was trained on 500,000 hours of proprietary, real-world dexterous manipulation data. It has NEVER seen that exact scenario before. This was the coolest demo at @AutomateShow @GeneralistAI @FlexivRobotics
9
6
79
11,697
Pete Florence retweeted
These are the kinds of tasks that felt impossible to automate just a few years ago. and here they are showing it live
That's one of the coolest robotics demos! 😮‍💨 @GeneralistAI showed GEN-1 handling box folding and screw packing during @AutomateShow. The boxes are cardboard with real variability: creasing, deformation, different configurations. GEN-1 retries when things go wrong. It adapts mid-task. This is a jump from GEN-0, shown 3 months ago during GTC. GEN-0 handled rigid, predictable boxes. GEN-1 handles deformable ones with chaotic behavior. The scaling is straightforward. GEN-1 trained on 500,000 hours of proprietary, real-world dexterous manipulation data. Scaling laws are working. More high-quality real-world data leads to better physics understanding, better retry behavior, better generalization. And the time to get a working demo is shrinking drastically. They're not even at scale yet. They're just starting to see what happens when you pre-train on massive amounts of actual robot data. ~~ ♻️ Join the weekly robotics newsletter, and never miss any news → ziegler.substack.com
12
13
166
31,763
Pete Florence retweeted
Replying to @GeneralistAI
@GeneralistAI mogging controls based robotics at Automate 2026 in Chicago
2
7
126
14,950
Pete Florence retweeted
Generalist is hiring our first product designer! Come help me build the Generalist design culture, and help shape the future of robotics! generalistai.com/careers/pro…
5
6
59
9,814