Building robots to build everything else @pantographPBC. Previously cofounder @sfcompute, @ExaAILabs

San Francisco
One of the dreams of RL is to train models that can quickly learn how to act in any RL environment, even ones they’ve never seen before. This new model is a step in that direction:
Introducing Pan-1, a highly capable Minecraft model trained with an RL-based pretraining technique that could unlock internet scale video for robotics models. It can achieve diverse goals—fight mobs, build structures, explore—without training on any of them specifically.
4
1
43
4,187
Maybe you do need giant models after all?
How do Opus 5.5 and Sol 6 perform on real robots? We find that both are much less performant than Astra and Fable on our hardware in early testing. On a simple task where they must put a ball in a bowl, Opus succeeds on only 1/10 attempts, and Sol on 0/10.
12
1,243
Early signs of GDP growth from AI:
What an insane summer !!! Several of the software libraries I maintain got massively faster! Some of this software is stuff that you probably use. For weeks I have been telling people that I was preparing an article... 'A summer of AI optimization'. Here it is. I have never seen anything like it. This was well optimized code in the first place. And we just made much of it much faster, much better, in a few short months. And it wasn't planned. It just happened.
11
791
Fixing all software vulnerabilities is very important for the coming cyber apocalypse!
Keep an eye out for: - plan for verifying all software coming very soon to @theoremlabs blog - verified (toy) sandbox demonstration in the coming weeks - verified production sandbox on Linux by end-of-year
9
939
Frontier language models are quite strong robot policies! There are many cool demos floating around twitter, but we wanted to run a larger systematic test on real robots. Some interesting results on the trade-offs between thinking harder and acting slower, recommend reading the post!
Astra outperforms Fable at robotic control! We ran a controlled benchmark of GPT-6 Astra and Claude Fable 5.1 on our Pandroid robots across eight manipulation tasks. Astra completed 35% of attempts, Fable 15% (p≈0.001). Human teleoperators completed every task.
1
30
1,523
Very cool!
on a quest to reproduce @pantographPBC 's research on goal-conditioned RL (implemented on DM control first for cheaper iterations). the goal is to scale this up to a tiny goal-conditioned minecraft agent. on DM Control, we beat plain IL by up to 3.75x.
2
15
1,696
Some of the largest-scale funding available for safety projects!
Today we’re launching Project Tailwind, a call for founders to start ambitious new AI safety initiatives: coefficientgiving.org/tailwi… We’re looking for great people to engage seriously with the risks of transformative AI, and to create the research, technologies, and institutions that will help humanity navigate them. There are critical, basic problems that no one owns. Over the last several months, my team at @coeff_giving brainstormed ideas we’d be excited to fund if we could find a promising founder, and the list quickly grew to more than 200 entries. We’ve narrowed it down to our favorites, but we also expect the best founders to bring their own ideas: coefficientgiving.org/tailwi… We’re providing funding at three levels: 1️⃣ Pre-seed: $200k to $2m to develop an idea and build a team 2️⃣ Seed: $2m to $20m to launch and scale 3️⃣ Scale: $20m to $200m+ for proven teams to scale, or world-class teams to start If you’re excited about something on our list, or you have another proposal for driving the field forward, you should get involved: coefficientgiving.org/tailwi… There are more good projects than there are people to work on them. Please help us make that stop being true! Also hi, I’m Emily! This is my first tweet. I lead AI and biosecurity grantmaking at @coeff_giving.
4
517
Exciting to see more companies building entirely independent pretraining stacks!
Frontier pretraining is said to be a big-lab-only game. We don’t have 100k chips yet, so there’s only one way: algorithmic efficiency. Our new recipe matches DeepSeek V4 Pro’s pretrain using 50x less compute – that’s roughly half the FLOPs used for GPT3, or ~$0.5M on GB200. magic.dev/blog/pretraining
30
2,585
Alex Gajewski retweeted
Two weeks ago, I resigned from OpenAI to join Conduit as a founding researcher, where we're training models to non-invasively read the human mind. I've written some thoughts about what telepathy could look like by 2035 and how to get there:
1,197
941
8,887
3,044,303
A year ago, @_kelsey_pool and I set out to make thousands of little robots to collect a dataset large enough to train a foundation model on. Today, we’re launching Pandroid! This is the kind of robot that I imagine diffusing out into our homes and lives. They feel like Ghibli tech to me! They have some kind of a spark to them. Especially when they’re running on their own, my brain just registers them as alive, like some kind of bird or chipmunk.
We set out to collect 10M hours of on-embodiment data as a starting point for robotics models, but couldn’t find a robot reliable enough at a price that made it possible. So, we built one. Meet Pandroid: an accessible hardware platform for AI. Small enough to safely experiment with, durable enough to run in fleets. Connect it to any model or coding agent over SSH to let it interact with the world. Join the waitlist:
5
6
71
6,040
Over time, this strategy of building smaller and easier to manufacture robots at really large scale will send prices towards zero. We’ll build larger and stronger robots later too, but we think it’s the right strategy to start with mass manufacturing for something very simple and make it more complex, rather than the other way around. We’ll be selling early Pandroids for about $4199, and these prices will fall over time, especially as actuators get cheaper.
1
10
190
We’re excited about helping to strengthen the US supply chain for robots. Actuators are the biggest bottleneck right now. If you’re working on American actuator manufacturing, we’d love to chat and see if there are ways we can help. DMs open!
2
10
173
What I think is most interesting about our model is how general the pretraining technique was (which we’re calling goal-conditioned pretraining). There wasn’t anything specific to Minecraft in it, or even about keyboard and mouse. Pretraining is where models get smartest. That’s where the most learning about the world happens, where we add the most information into the weights. We want a technique that allows models to learn how to accomplish goals through pretraining. You can think of video as an RL trajectory that has only observations (it doesn’t have actions or rewards). Video is a little bit special in this way (e.g. you can’t do this with text.) What to do about the missing actions and rewards? One way to deal with missing rewards is to reframe the problem. Rather than maximizing numerical rewards, we use goal-conditioned RL, where the goal is to learn how to reach any state. For video this means prompting the model with images. One of the things that’s great about goal conditioning is that we can do *hindsight-relabeling*, where we use frames from later in a video as goals for the earlier part. (It’s called that because we take a trajectory that may have been a failure for one goal and relabel it as a success for whatever actually happens.) Next, we’re planning on scaling this up to much more diverse videos, and we’re hoping to get a model that can achieve goals in any video game, as well as control robots in the real world. There is a wide space of ideas to experiment with here that have never been scaled up before. If you’re interested in working on these models, feel free to reach out :)
Introducing Pan-1, a highly capable Minecraft model trained with an RL-based pretraining technique that could unlock internet scale video for robotics models. It can achieve diverse goals—fight mobs, build structures, explore—without training on any of them specifically.
2
17
1,431
Alex Gajewski retweeted
pantograph has quietly been developing one of the most interesting approaches in robotics walking into their office half the team will be patiently studying minecraft evals while surrounded by tiny robots fiddling with wooden blocks and juggling balls
Introducing Pan-1, a highly capable Minecraft model trained with an RL-based pretraining technique that could unlock internet scale video for robotics models. It can achieve diverse goals—fight mobs, build structures, explore—without training on any of them specifically.
3
2
100
10,778
One of the dreams of RL is to train models that can quickly learn how to act in any RL environment, even ones they’ve never seen before. This new model is a step in that direction:
Introducing Pan-1, a highly capable Minecraft model trained with an RL-based pretraining technique that could unlock internet scale video for robotics models. It can achieve diverse goals—fight mobs, build structures, explore—without training on any of them specifically.
4
1
43
4,187
Next, we’re planning on scaling this up to much more diverse videos, and we’re hoping to get a model that can achieve goals in any video game, as well as controlling robots in the real world.
1
6
179
There is a wide space of ideas to experiment with here, which as far as we can tell haven’t been systematically explored yet at large scale. If you’re interested in these ideas, DMs are open!
6
164