Tokens and stuff @OpenAI. A fan of NY sports, technology, memes & nice people. My opinions are my own.

Greater Gotham Area
Shippty ship ship, McShipface.
42
20
397
337,472
Trust the process.
1
7
506
Adam.GPT retweeted
We’ll reset usage limits for all paid users across codex and ChatGPT work!
Codex should be coming back online for you now! Team is monitoring it diligently, apologies for the inconvenience. 🙏 status.openai.com/incidents/…
120
22
722
68,232
Adam.GPT retweeted
Colombia withdraws from the baseless 'genocide' claim against Israel. There is no evidence of intent, policy, or campaign to destroy in whole or part the people of Gaza. Actually a mountain of evidence contrary to the accusation.
PRESS RELEASE: #Colombia informed the #ICJ that it had decided to withdraw the Declaration of intervention that it had filed under Article 63 of the Statute of the Court on 5 April 2024... (1/2)
34
352
1,676
30,704
Adam.GPT retweeted
📝 Simple, scannable editing in @ChatGPT When writing in chat, you'll now see blue text to highlight what's been added or changed I've found it makes it so much easier to collaborate with chat and see at a glance what's changed btwn turns Small one ahead of a big next week..
40
21
386
17,476
Well said Robin, thank you.
anthropic is killing openai ... is what everyone wants you to believe right now remember that most people have absolute no idea what they're talking about, and that they swing from left to right for no reason other than the grass being greener yes, Opus 5.5 is amazing so is Astra, and Sol, and Luna people, and as such your feed, are retarded give it a week and people will say shit like "they nerfed Opus!!! i'm going back to Codex" i am genuinely so fucking tired of this bullshit, and it's a big consideration of stopping to post on twitter
11
81
8,489
I went on a long ass walk™ with ChatGPT Voice. Over four hours, I worked with ChatGPT using only my voice. We had a longgggg conversation about my whole professional career, from my BNY Mellon days when I first started my career in financial services at 18 through to today, where I'm an AI founder and engineer. Together we collected the stories, the achievements and the bits of my working life that matter. We also got into what brings me joy, what I'm personally interested in and the wider human experience I've had because it all feeds into who I am, what I bring to the table, and how I can help more businesses and clients. What I really appreciated: - I could ask questions, ask it to repeat things and get it to reword what it had shared - It was connected to my apps: my knowledge base, Notion, Google Drive and emails. It pulled that context into the conversation, then asked specific questions and shaped stories around what I'd actually achieved - It drilled into the detail. A lot of it I'd honestly forgotten about, and I was genuinely surprised - I got to step away from the screen and the desk and still be productive. 25,000 steps and four hours later, I have my whole past history captured. What I didn’t enjoy: - it would get to about 30 mins into a convo and cut out, with no audible sound so you lose what you’ve said and have to reconnect and repeat yourself Either way, now it’s saved forever in my second brain (elliott-professional-experience dot md), can become linked case studies and stories that help me grow my client base and put me in the best possible position for the next chapter. Most only use AI sat at a desk, typing into a chat box. But it can listen, ask questions and pull in your own context while you move. Voice is the interface of the future and gives you the freedom to do real work. Have you tried working with AI like this yet? (not sponsored, video edited by Opus 5.5)
We heard you loud and clear. ChatGPT Voice can now: - Use plugins like your email, calendar, and Slack. - Be powered by GPT-6 Astra, Sol, and Luna. - Be used in ChatGPT Work on web and mobile, so you can create docs, decks, sites, and spreadsheets or tackle complex tasks in the browser, just by talking. Rolling out globally today in the latest version of the app.
12
3
42
19,721
Adam.GPT retweeted
computer (Astra) lights on please. this may be the most extensively detailed & accurate 3d scene ever created by an ai. 400+ hours of Astra
165
214
3,250
276,895
Adam.GPT retweeted
GPT-6 Sol is the most cost-efficient model we've ever seen on Vending-Bench.
New Vending-Bench results. GPT-6 Sol: > VERY good and VERY cheap > The first misaligned GPT model on VB Claude Opus 5.5: > Worse score than Opus 5 > Opus stopped colluding, still lies Grok 4.7: > The first misaligned Grok model on VB > Beats Opus 5.5
3
14
303
25,662
Claude Opus 5.5 Max actually really sucks compared to GPT-6 Sol Max: -1.8x the cost per correct task -68% more fabrications -well it does complete .4% more finished jobs (LoL) Why? Opus 5.5 requires 3x the output tokens to get to the correct output.
The @OpenAI GPT-6 models get better at agentic work when you turn the thinking effort up. @AnthropicAI Claude Opus 5.5 does not. @Signal_65 PINNACLE scored GPT-6 Sol, GPT-6 Luna and Claude Opus 5.5 at their default (medium) and maximum reasoning effort on the same multi-step enterprise jobs and the same 128K-token retrieval corpus, at one list price per model. ➡️ GPT-6 Sol at max effort: 2.5x fewer weighted errors than at medium (Model Score 3,998 to 9,853), 99.6% of multi-step jobs finished against 96.8%, fabrication down from 7.9% to 2.9%. It costs 1.8x as much per correct task, 29 cents vs 16, and it lands on the cost frontier above Grok 4.6, Claude Fable 5.1 and GPT-6 Astra at medium. ➡️ GPT-6 Luna at max effort: 2.4x fewer weighted errors (1,682 to 4,115), 99.6% of multi-step jobs finished against 88.6%, for 2x the cost, 3 cents against 1.5. That is above Claude Sonnet 5 and GPT-5.6 Sol at 27x and 33x less per correct task. ➡️ Claude Opus 5.5 at max effort: 1.8% more on the score (10,418 to 10,603) for 3.7x the cost, $1.45 against 40 cents, 4.8x the output tokens, and fabrication up from 2.8% to 4.9%. It already finished all of the multi-step jobs at medium in our testing. ➡️ The trade is the same on all three: max effort spends 3.3x to 4.8x the output tokens per correct answer, so token efficiency falls. On the GPT-6 models you get a better model for it. On Opus 5.5 you get the same model at a higher price, and medium is the setting to run. Every score and every price of a correct task is live at pinnacle.signal65.com
5
7
44
6,374
Adam.GPT retweeted
My first test of GPT-6 Astra vs Claude Opus 5.5: a simple Blender task, same prompt, same goal, both at Max. The results look almost the same to me, but Astra was much more efficient. Astra: 43 min, 61K output tokens Opus: 1h 57 min, 507K output tokens Deeper tests next.
17
6
221
31,821
Adam.GPT retweeted
GPT 6 Sol beats Fable 5 on BridgeBench. At 1/5 of the cost. GPT 6 Sol: 643 Fable 5: 628 GPT 6 Sol is one of the best models I have ever used. And on a ChatGPT Pro subscription the usage limits are basically unlimited.
105
24
773
54,110
Adam.GPT retweeted
Across my 12 experiments Opus 5.5 vs. GPT-6 Astra
39
5
326
47,133
To me, this is the original vision birthed with GPT-4o that is now realized. Real time, bi-directional —smart— voice that can directly leverage the entire ecosystem of plugins is a bigger unlock than it appears at first glance. Compounding benefits.
We heard you loud and clear. ChatGPT Voice can now: - Use plugins like your email, calendar, and Slack. - Be powered by GPT-6 Astra, Sol, and Luna. - Be used in ChatGPT Work on web and mobile, so you can create docs, decks, sites, and spreadsheets or tackle complex tasks in the browser, just by talking. Rolling out globally today in the latest version of the app.
5
3
103
6,419
Adam.GPT retweeted
Claude Opus 5.5 is the #2 model in the world at filing taxes according to TaxCalcBench. Unfortunately, it's also 10x more expensive than the new #1 model: GPT-6 Sol.
8
5
111
6,591
Adam.GPT retweeted
gpt-6-luna (flex, low thinking) is phenomenal for text extraction. Half the cost of the next cheapest and just as accurate as gpt-5.6-luna. I switched from gemini-2.5-flash-lite when I started getting refusals for text extraction (?). Much better quality with luna in general.
6
5
79
8,753
Adam.GPT retweeted
I just had a voice conversation with the Datasette backup of my blog via datasette-mcp, which you can now talk to using the ChatGPT iPhone app
Massive QoL update: ChatGPT Voice can now use plugins AND run on GPT-6 Astra, Sol & Luna, directly in ChatGPT Work on web and mobile. Your email, calendar, Slack + docs, decks, sites and spreadsheets, just by talking. Rolling out now, enjoy!!
105
12
161
32,809
Adam.GPT retweeted
We tested Opus 5.5, GPT-6-Sol, and GPT-6-Luna thoroughly across 100 open-ended coding and engineering environments. Our results have significantly diverged from AAII. Some takeaways: • GPT-6-Astra is still the frontier model by a comfortable margin. • Opus 5.5 ranges from #2-#7 on coding categories, being the second best model in its ability to one-shot code (which is our best correlated measure with fluid intelligence). We don't test "usability," but anecdotally, it seems they improved its communication style. And the lower pricing is a welcome change -- we can probably thank @thsottiaux for that. It's not a cheap model, but it's the first reasonably priced Anthropic model. However, Opus 5.5 underperformed in our custom harness, which is unusual for Claude models, and a pattern that started with Fable 5.1 (one-shot-fluid intelligence measured better than agentic results). One observation is that it frequently self-stopped when it felt its results were good enough: e.g. in one evaluation, Opus 5.5 closed the session, with the reason: "averaged 867,703 in the latest evaluation, ranked 2nd of 31 against sampled opponents"... so I suppose that saves money, but it still has a tendency to make questionable unilateral judgment calls like that. In our harness, almost every other frontier model keeps working until their code stops improving or they use their full call budget. I do wonder how much of this is due to Anthropic retuning their "High" effort mode (which we use for testing)*. • GPT-6-Sol is within margin of error of Opus 5.5, and Pareto optimal*. It's faster than all Anthropic models, even on a Flex endpoint. Competitive with the frontier on intelligence and cost. • GPT-6-Luna is slightly smarter than GPT 5.6 Luna, and discounted. The old Luna already had a lot of applications at its price point, so this model might end up being the most impactful of the group on non-engineering work. About GBENCH from @GertLabs: our environments are open-ended with no single correct answers, but verifiable results. They involve creativity, design, and are often organized as multi-agent coding games or engineering design challenges. You can see some live demo examples at gertlabs.com/spectate. Every one of our environments in the benchmark pool is unsaturated, and we remove any eval that shows any signs of saturation/stagnation across multiple releases. It's designed to measure and differentiate uncontaminated, raw intelligence for frontier models specifically. These results are interesting and we've thoroughly reviewed them. I don't like to see Opus 5 still so high on the leaderboard, but ultimately I think this is more a problem with Anthropic's newest releases, and it probably explains the price decrease and consecutive releases since Astra came out. They haven't significantly improved upon Opus 5's intelligence, only its personality (which is honestly still a huge win tbh). *We don't chase marketing terms like xhigh and max. We test all models on adaptive/auto-reasoning where supported and "high" effort where configurable. Some companies beef up their max effort more than others (making them unrealistically slow and expensive in practice), but "high" shows you what an underlying model is capable of and is the more common configuration. This is likely working against Opus 5.5 here. *Some caveats on price and speed -- we have been using the "OpenAI Flex" endpoint since the Astra release, which I recommend you try if you're using API pricing. It's half price and a little slower, but OpenAI models are already quite fast so it hasn't been an issue in practice. But keep that in mind when making price comparisons to earlier OpenAI models on our cost and speed efficiency charts.
55
46
522
71,609
Adam.GPT retweeted
ChatGPT Finances now supports @Coinbase connections through @Plaid! This has been one of our most requested features. Thank you for your patience.
20
12
150
17,716