Co-founder @GeneralistAI general intelligence born from the physical world ✗DMs → email

Andy Zeng retweeted
A simple example of physical prompt steerability: same environment, different prompts, different behaviors. Thanks @JagdeepBhatia8 for the suggestion. Read more about GEN-1.5 in our blog post in the comments below.
Physical prompting is an elegant idea! One request for the @GeneralistAI team: can we see an example of *physical prompt steerability*, where different prompts induce distinct behavior in the *same* environment? I'm curious whether GEN-1.5 is listening to the physical prompt or simply executing the most likely action sequence for the scene under its pretraining prior.
7
10
112
21,681
Andy Zeng retweeted
We've reduced the time it takes to go from physical prompt → robot behavior. The faster anyone can teach a robot to do something new, the easier it becomes to scale physical work. Read more about GEN-1.5 in our blog post in the comments below.
8
27
236
26,678
Andy Zeng retweeted
I'm usually wary about talk of "the ChatGPT moment for robots" because language is less complex than the physical world. After visiting @GeneralistAI, though, where I saw their robots perform impressive manipulation with just a prompt, I think you can see it happening in a limited way. (Thanks to @peteflorence and co for having me!) wired.com/story/generalist-a…
5
9
79
26,860
Andy Zeng retweeted
Good question thanks Ani. Here's a video that gives a sense of the answer here about the effect of prompting: - in the first part, the model is just babbling with no prompt in context - at the end, we add a prompt for taking money out of a pouch (generalizing to the wallet)
Very cool work, Pete! I'm curious if you've tried an extreme version of this: no in-context example at all; just place objects in front of the robot and see if it can infer the task. I suspect this will have non-trivial success rates, which would also allow you to figure out how much of the benefit is actually coming from the prompt (rather than pure zero-shot task inference + capability). I remember Russ showed something like this in his Stanford talk on LBMs last year: piped.video/TN1M6vg4CsQ?si=-8J-…
17
28
270
32,865
Andy Zeng retweeted
We started planning yesterday’s @generalistai GEN-1.5 announcement about 3 weeks ago. Two weeks ago, already feeling whiplash from new results, I asked ChatGPT to create this image to share in our internal meme channel.
Made with AI
3
6
41
3,844
Andy Zeng retweeted
I remember having a conversation with @ilyasut in the early days of @OpenAI (~2018) about how the pace of progress in AI research can be surprisingly fast. Overnight, things can go from feeling nearly impossible to just working. I felt it then, and I feel it now.
At 10:06pm on Aug 3, I watched a robot do something I thought was years away. We were working on few-gradient learning with GEN-1.5, wondering how far we were from landing physical prompting: show the robot a task once, and it just does it. We YOLOed it, and the robot imitated exactly what we demonstrated, with 0 fine-tuning. We gave more prompts for different tasks. More successful rollouts landed. I could not believe my eyes. Looking back, this represents the culmination of a huge amount of work across the entire team, many failed experiments, and years of groundwork from the community. I just happened to be the one standing in front of the robot when it all clicked. @andyzengineer once asked me: "If Generalist failed tomorrow, what would we want to have left behind that could benefit humanity?" I believe we found one answer: an existence proof of physical prompting.
2
3
34
5,173
Andy Zeng retweeted
absolute GPT-3 moment for robotics holy moly
Introducing GEN-1.5, a one-shot learner. It can learn new tasks in a few seconds. Show it what to do, and it generalizes. This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the physical world.
32
218
4,261
479,069
Andy Zeng retweeted
Ever since I started working on robot foundation models a handful of years ago, the broad ability to one-shot in-context learn has been the single most vivid goal in my mind along the long road ahead. It’s something @andyzengineer and I have especially been thinking together on for several years now, and something shared by the whole team at Generalist as a major goal. But even before anybody started talking about robot foundation models, too, the capability this enables, i.e. generalized one-shot learning, has been both incredibly concrete and elusive. In terminology that I think many other people have come up with as well, in grad school we used to talk about a “Ctrl-C-Ctrl-V” type of capability. See something, and have the robot do it too. This same type of idea inspired the name of the 1970 MIT copy demo (people.csail.mit.edu/bkph/ph…). The thing is, it sounds simple, but is incredibly hard since the world is never quite the same when it got copied and where you want to paste it. The real world can be hard to predict and is full of variation. To do this, you need strong generalization, and it needs to be acquired in a single example. Doing this over a wide range of tasks, especially for dexterous tasks, is hard mode. Lots of the components of the idea of making this all happen have been there for a long time. As an example, this 2017 NeurIPS paper “one-shot imitation learning” arxiv.org/pdf/1703.07326 has excellent vision, with ambition well beyond what was achievable at the time, and although it’s not referred to as “in-context learning” since it was pre-Transformer, it actually uses attention to condition on a single demonstration. And now, many things have happened since early 2017, including the broadly celebrated arrival of one/few-shot in-context learning in language models in 2020. This new model GEN-1.5 takes in everything we have built and learned over the past couple years at Generalist. It has been training for 8 months. It has taken an incredible amount of commitment and grit from the whole team to get here. The level to which this model has survived many surgeries has continued to surprise me. And its capabilities have continued to surprise as well. We found compositional generalization on Friday. We found sim2real prompting earlier last week. We filmed the contiguous uncut videos of live prompting yesterday. To be clear, the success rates are modest, and there’s still a long way to go. But now I have definitely seen a ~decade-long imagination come into the real world.
Introducing GEN-1.5, a one-shot learner. It can learn new tasks in a few seconds. Show it what to do, and it generalizes. This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the physical world.
22
41
474
68,819
Andy Zeng retweeted
It feels like almost every day at this company there is another significant result! The speed of progress makes my job of documenting and communicating what we are doing a challenge, but it's also so much fun!
Introducing GEN-1.5, a one-shot learner. It can learn new tasks in a few seconds. Show it what to do, and it generalizes. This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the physical world.
1
1
36
3,629
Andy Zeng retweeted
Our latest pre-training comes with zero-gradient / in-context learning! So crazy to witness and glad we can talk about it publicly now.
Introducing GEN-1.5, a one-shot learner. It can learn new tasks in a few seconds. Show it what to do, and it generalizes. This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the physical world.
1
1
51
2,598
RT @felixwyw: At 10:06pm on Aug 3, I watched a robot do something I thought was years away. We were working on few-gradient learning with…
6
68
Andy Zeng retweeted
Replying to @GeneralistAI
GPT3!
5
17
352
29,796
Andy Zeng retweeted
Introducing GEN-1.5, a one-shot learner. It can learn new tasks in a few seconds. Show it what to do, and it generalizes. This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the physical world.
320
1,689
12,172
3,414,822
The models are getting better everyday. This is a pretty contact-rich task, and forces us to reckon with the fact that our policies not only need to learn the quirks of different robots, but also have to do it in a way that generalizes to the fine-grained subtleties of friction and forces e.g. so that objects don’t get wedged or stuck. Gave me goosebumps when I first saw these results internally. Something about it too, that just makes these models feel just a bit more “human” now in the way they move than they did before. Wild times.
We've improved how GEN-1 learns to adapt to new actuators and new robots at the lowest level, with up to 10-20x gains on internal benchmarks. This significantly boosts performance on high-precision tasks like disassembling parts from a NIST board. Read more about GEN-1 in our blog posts in the comments below.
20
25
318
61,802
Andy Zeng retweeted
What happens when you change the hand mid-task? We tested this by modifying the hands mid-rollout and letting the same model keep running. It perceives the new tool, and finds a new trajectory and contact strategy to complete the task. This works because training on mixed data forces the model to condition its behavior on the hand in front of it. This brings it closer to a general understanding of how shapes and contact surfaces interact with the physical world – and which actuation strategy should follow.
1
9
93
22,212
Andy Zeng retweeted
Each hand is a different sensorimotor interface by which GEN-1 experiences the physical world. Scaling pretraining across thousands of these interfaces teaches GEN-1 a universal physical commonsense that transfers to new hands and new ways to grasp, push, pull, twist, and more.
1
6
114
32,239
Andy Zeng retweeted
Sometime next year pairs of arms become reliable enough that they can do work on an assembly line or fulfillment center I think? In the same way that ~the most important trendline to have a view on is how many GW of accelerator capacity are deployed in 26/27/28/29, you should also have a view on how many pairs of arms are built and deployed
Why create robot intelligence for just one hand, when we could have it learn from many? GEN-1, our latest embodied foundation model, now supports a broad range of end effectors from 5-finger hands, to specialized tools, and everything in between.
17
20
447
96,197
Andy Zeng retweeted
A year ago, I thought world models would be disruptive by sometime in the 2030s. It’s now likely that these models will have a huge impact THIS decade, and massively accelerate innovation and disinflation even by 2030. Generalist is awesome / bullish 🇺🇸
Generalist
29
33
537
40,401
Andy Zeng retweeted
easy to miss detail and pretty neat: you can just modify the UMI style data collection devices to use different tools
Why create robot intelligence for just one hand, when we could have it learn from many? GEN-1, our latest embodied foundation model, now supports a broad range of end effectors from 5-finger hands, to specialized tools, and everything in between.
7
10
102
17,565