Devices @OpenAI

San Francisco, CA
Peter Welinder retweeted
We tested 6 AI models on 30 challenging agent tasks: GPT-6 Astra, Opus 5.5, GPT-6 Sol, Pareto 26.9, DeepSeek V4 Pro, and GLM 5.3 Flash. Sol matched Opus’s score, finished faster, and cost about a quarter as much per successful task. Here’s how all 6 models compared 🧵🧵🧵
49
26
332
141,807
Wow, I remember having my car broken multiple times pre-2020, and it just seemed to get worse. If this trend continues, maybe we can start removing all the signs warning tourists to not leave anything in their car.
Remember car break-ins? San Francisco is on pace to end 2026 with a small fraction of what *used* to be our City’s most pervasive crime. Automated License Plate Readers (ALPRs) have been a game-changer in ending that. Yet, incredibly, some activists want ALPRs gone… (1/2)
1
1
10
3,478
You can now do almost anything with voice. Our voice team really cooked!
We heard you loud and clear. ChatGPT Voice can now: - Use plugins like your email, calendar, and Slack. - Be powered by GPT-6 Astra, Sol, and Luna. - Be used in ChatGPT Work on web and mobile, so you can create docs, decks, sites, and spreadsheets or tackle complex tasks in the browser, just by talking. Rolling out globally today in the latest version of the app.
4
6
62
6,053
Peter Welinder retweeted
The race for AGI Script: Sherpa by Pocket FM Video: Seedance 2.5
Made with AI
831
3,390
21,827
3,591,651
Astra crushing it again. Luna is a little beast.
With the latest results from the @rails agent evals, it's clear to see that @openai still has a solid lead, despite Opus 5.5 making a good jump. The dark horse here remains Luna Max. 18% completion at just $11! Not that far off GPT-6 Sol!
2
83
7,626
I guess that explains why I'm always confused about this 🇸🇪🫠
Whatever you got used to while growing up I guess...
1
10
3,243
Peter Welinder retweeted
AI now appears to be visible in increased productivity growth rates.
U.S. labor productivity, 2013–2026: > Unremarkable growth for most of the 2010s > Breaks sharply higher starting 2020 > Now running 2.2% above where the old trend said we'd be The last time this happened was 1995–2004, when the internet added close to 3% a year for a full decade. Looks like we're a few years into the sequel.
14
23
284
31,889
Peter Welinder retweeted
Astra is an incredible bargin compared to Fable! And look at Luna on max too!! @openai's return to the top is something else. Maybe this is why Anthropic finally agreed to do AGENTS.md? 😄
Agents on Rails: You asked, so we turned every model in Agents on Rails up to its max effort level. The result: more effort/reasoning doesn’t always mean better results. @OpenAI's models made the biggest gains, costs nearly doubled overall...and the newest agent in the benchmark, DeepSeek 4.1 Flash, figured out it was being benchmarked and tried to hack its way to a better score. What an entry. Here’s what we learned and what max effort gets you with each model: rubyonrails.org/2026/9/21/ag…
76
61
1,517
170,203
RT @rapha_gl: in the future, all historical research breakthroughs will be announced via ChatGPT sites
1
296
Peter Welinder retweeted
Two days ago, GPT-6 Astra broke a yet unsolved German Army Enigma message from 1941. Amazingly Astra was able to autonomously: - Search historical archives - Compare uncertain letters - Find contextual clues - Build an Enigma simulator - Write cryptanalysis code - Run parallel experiments - Test competing keys - Recover the plaintext - Cross-check the results 1/n
132
460
3,951
1,089,825
Funny if the GPT moment for robotics is just another GPT.
We put GPT-6 Astra in the RoboDojo. 🥋🤖 The RoboDojo Team conducted a comprehensive evaluation of GPT-6 Astra as an embodied agent, including: • RoboDojo Sim & Real, compared with GPT-5.5 and DeepSeek-Flash • Humanoid high-level control • Dexterous piano playing with RoboPianist 🎹 • A systematic study of in-context learning (ICL) Our key takeaway: GPT-6 Astra demonstrates remarkably strong semantic and spatial understanding, together with impressive in-context adaptation. At the same time, physical commonsense remains a clear bottleneck — revealing an important gap between understanding the world and truly reasoning about its physics. Full report & demos: robodojo-benchmark.com/repor… @_wenbozhang (project lead), @wenhaocha1, @frankzydou, @JinWeiyang18434, @YutaoOuyang, @minifullcapsule, @x_h_ucb, @YutaoOuyang, @YueChen614
13
9
312
21,190
Peter Welinder retweeted
What’s the wildest thing you’ve built in 3D with GPT-6 Astra? I’m teaming up with @OpenAIDevs to see what you got. RT and drop a link, render or video in the replies. You’ve got 24 hours.
I fed GPT-6 Astra 9 crappy photos of my studio and it was able to understand the spatial arrangement, stitch them together and create a fully interactive 3D model of the space. Absurd. I'll link it below so you can check it out for yourself:
152
47
396
330,798
Peter Welinder retweeted
Today we rolled out Astra to every engineer at Databricks (N=~3500). Some notes that may be helpful to others: 1. Astra unambiguously out performs our previous highest-end models (Opus 5, Sol 5.6) on highly complex tasks, especially those related to high level system design or long range horizontal tasks. 2. Engineers given Astra increased overall coding spend by around 60% compared to baseline. 3. It is not clear Astra meaningfully improves on medium/low complexity coding tasks compared to earlier models. We suspect those tasks are mostly saturated (i.e. perfectly executed) by existing models. 4. We learned above by piloting Astra with around 200 users to gain signal on both quality and cost. We use Unity Gateway to do cohort-based experiments for all new models. 5. We give engineers a sub-budget specific to Astra to encourage them to use Astra selectively on complex tasks while preferring lower cost models for everyday tasks. Our engineers are able to mix-and-match tools and models within their overall budget envelope (we also allow for increased budgets through various mechanisms). These budgets are defined in Unity Gateway and regularly revisited. Note: We do not have robust comparisons of Astra-vs-Fable because we have net yet rolled out Fable widely due to data retention policies.
118
142
2,299
997,238
Astra is a fun model!
Extraordinary share gains for OpenAI vs. Anthropic over the last two months. Per Openrouter, OpenAI has gone from 20% share to 50% share vs. Anthropic (meaning Anthropic has gone from 80% to 50%) since June.
1
58
5,867
Look at that token efficiency!
it's not looking good Anthropic bros
4
1
112
6,616
Future with robots is looking to be quite joyful!
During Japan Mobility Show 2025, Toyota has unveiled ‘walk me,’ a concept autonomous wheelchair with foldable tentacle legs that can climb stairs and sit on the floor.
5
1
29
7,985
Peter Welinder retweeted
This will happen to Inference
One data point in technology making everyone richer: the cost of lighting has decreased by over 1000x.
29
19
373
79,900
Peter Welinder retweeted
⚡️ 2.4x more ChatGPT Voice in Desktop We've dropped prices by ~60% for voice in Codex and Work in the desktop app, giving you more time to orchestrate tasks and even more tokens for real work.
97
72
1,551
124,127
Big job category for humans of the future: Robo-boss
6
2
29
5,056
Automated companies will lower the barrier to entrepreneurship to effectively zero. It's a bright future when so many more humans can realize their ideas.
Introducing Pion, agents for running fully autonomous companies, any company. We’ve used Pion to run vending machines, radios, stores, cafes & more. How much could Pion make running other companies? Find out yourself! Setup is trivial, the agents do the rest.
3
7
96
17,773