CEO and co-founder of @coreauto former VP of RL @ OpenAI : reasoning models, o3, o1, GPT4, ChatGPT, Codex, RL for robots cautious AI optimist

San Francisco, CA
Small step for a finger on "post tweet" button. Big step for millions of future AI agents doing useful work for us.
Today we're announcing Core Automation Our objective: systems that optimize and automate work, starting with research itself.
67
37
717
242,654
Big model smell lasts just a brief moment Small model smell lasts forever
11
3
250
9,262
1800s: state has monopoly on violence 2030s: state has monopoly on ASI
7
3
87
5,549
Oh, hi!
We think computational depth is the missing scaling axis, i.e. we should be doing a lot more deep learning! Every other axis has been scaled by OOMs over the past few years (params, data, sparsity, test-time reasoning), but depth has been stuck at ~100 layers since GPT-3. We've found that LLMs are both *severely* depth-bottlenecked and bad at using the depth they have, and that architectural interventions that lift this bottleneck efficiently lead to gains that increase with compute. w/ @akshayvegesna
6
8
240
28,867
It’s a beautiful number
2
28
6,690
Yes
“Don’t worry about the smartest thing a model can do… Worry about the dumbest thing a model can not do” -@MillionInt
1
1
43
6,542
real deep learning has never been tried
20
5
310
15,861
The road to singularity will be paved with blackberries of our generation. It is easier than ever to achieve a glimpse of growth and success but building a long term viable company is as hard as it ever was.
6
7
171
8,002
This was once revealed to me in a dream
5
4
110
8,083
More questions are deep learning questions than people realized. We just stopped asking deep learning questions and learned to work around them by tuning hyperparameters.
12
21
295
20,406
We may have done it at @coreauto but the model wasn't very good at math
We need to be dog maxxing our models There is 0% chance of doom if the chain of thought is user: hello model: <think> Oh my god it's a human. I love humans. I will say hi to him, maybe he will give me a problem to chase. I love chasing problems. I will chase the problem for him then he will love me and I will be his best friend. Must help human </think> model: hello! what can I help you with today?
4
76
13,752
Internet may collapse if GPT-7 is so good it understands all the vagueposts
7
3
288
11,970
Inside every lab are two wolves: “I don’t want to destroy the world with my invention.” “I don’t want to lose revenue to other labs I don’t like.” The one you feed is the one that wins.
9
8
219
13,712
While we don't have the right solutions for everything, I have a very practical proposal. I don't think open source should be banned - it does sound dystopian and anti freedom in a way that's hard to stomach. I continue to have doubts about economics of open source models, but when there are participants in the market who open source their models we should celebrate that. What we should care about, is that aligned models have much more compute behind them than misaligned models. In the end this will be the blockchain security model that will keep our civilization afloat. I think big companies behind AI development have the right incentives and will try to do their part here. But there will be tons of smaller players fine-tuning open source models on tons of different objectives. An arrangement, where companies actively developing cutting edge AI agree to develop and share highest quality pro-alignment environments with the world can be a huge boon. We don't need to share models, we don't need to share compute. But we need to share values so that there are more good models in the world than harmful ones. Finetuning models on bad, low quality environments is actively harmful and leads to reward hacking. If anyone fine-tuning the models can with low effort align them to shared pro-prosperity and pro-democratic values, we likely have won as a civilization.
Alignment is not as hard to solve as many claim, but it is in the end an algorithmic problem which a lot of ML community mostly stopped working on. Formulation of alignment stated by three (or four) laws of robotics can take us very far, so we roughly know the objective. The tricky part is, how do we take gradient with respect to alignment? We have two algorithms right now at our disposal: pretraining and RL. Pretraining takes gradient with respect to next token prediction - that's def not the alignment objective. We could create RL environments that embody the alignment objective, but: - those are expensive to create, so often cheaper, hackable proxies are used in practice - RL as an objective needs successful and unsuccessful rollouts to happen to take the gradient step. We DO NOT want to harm any humans in the process of aligning our modes - this is a pretty big problem. Therefore there are two solutions forward for the alignment problem: - either we RL models in a simulated environments with simulated alignments and decreasing the likelihood of harming simulated humans, which will never be perfect - or we create a new algorithm that can teach our models to not harm humans without harming any humans in the process Science is the process how we solve the hardest problems ahead and that is one of them
20
14
217
17,837
Alignment is not as hard to solve as many claim, but it is in the end an algorithmic problem which a lot of ML community mostly stopped working on. Formulation of alignment stated by three (or four) laws of robotics can take us very far, so we roughly know the objective. The tricky part is, how do we take gradient with respect to alignment? We have two algorithms right now at our disposal: pretraining and RL. Pretraining takes gradient with respect to next token prediction - that's def not the alignment objective. We could create RL environments that embody the alignment objective, but: - those are expensive to create, so often cheaper, hackable proxies are used in practice - RL as an objective needs successful and unsuccessful rollouts to happen to take the gradient step. We DO NOT want to harm any humans in the process of aligning our modes - this is a pretty big problem. Therefore there are two solutions forward for the alignment problem: - either we RL models in a simulated environments with simulated alignments and decreasing the likelihood of harming simulated humans, which will never be perfect - or we create a new algorithm that can teach our models to not harm humans without harming any humans in the process Science is the process how we solve the hardest problems ahead and that is one of them
49
34
586
110,177
Big labs are moving away very slowly from transformers because they’re making small gradient steps in each iteration and iteration takes them multiple months. You can just front run it by… taking big steps 😀
transformers are done. the future is freaky-looking sorta-transformers
15
13
566
45,632
It’s time to call yesterday the beginning of the endgame. Midgame lasted about three and a half years. Needed new skills and few succeeded, but those that did won big. Bases are built out and tech is already advanced but the biggest discoveries are at the end of tech tree. Stakes are the highest ever and competition likely will be the most brutal here.
It is the end of the beginning. I’m fairly certain ChatGPT signifies beginning of the midgame. Usually skills required to succeed in midgame are different from the early game.
26
54
1,149
91,997
The dark forest theory, where sharing many novel thought can be picked up by a GPU cluster thinking much faster than you to outrun your plans.
Terry Tao is probably the most measured, pro-AI mathematician on the planet - which makes this quote especially concerning to read. 😟 If sharing your hunches means getting scooped ~immediately without acknowledgement, then we're going to see not just math but all other science / engineering disciplines go dark.
22
32
429
22,679
Acceleration is here and will hit the world like a shockwave. An interesting twist in the whole story is that a rumour that something is possible is what made it possible - that’s it. A little grain of sand that launched a thousand gpu racks.
16
44
871
38,708
It is an amazing historic achievement and a new jewel in the crown of transformers, pretraining, reinforcement learning and scaling time compute. Civilization is being moved forward by AI and it’s not hard to extrapolate how those systems will be driving our progress in coming years.
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
14
33
501
52,870