Tech investor. A skeptical empiricist surrendering to life; Studying fat tails, removing blind spots, seeking dynamic quality, and letting go. Views are my own.

New York, NY
Dickson Pau retweeted
One more thing: we’re halving the price of cache reads on Claude Sonnet 5.5, to $0.10 per million tokens. That makes Sonnet 5.5 around 20% cheaper to run on most long-running work.
131
222
5,256
297,569
OpenAI is on the ropes...
Replying to @claudeai
Haiku 5.5 is a significant step up over Haiku 4.5 across coding, computer use, and knowledge work.
4
242
Haiku 5.5
1
71
New Nvidia nemotron model coming next week!
3
186
So confused. So is all Grok Bot going to be based on Opus 5.5 or not? @poteto
Grok @Bot will use whatever achieves the best outcome for users. Simple questions will route to small, fast models. Questions with complex answers will route to large models.
2
109
No worries at all. SpaceXSI and Anthropic are still competitors at the end of the world.
It’s hard for OpenAI to compete when Anthropic and SpaceXAI are basically holding hands and backing each other while trying to crush it.
128
Dickson Pau retweeted
We shipped four things that were deemed good to great and some math proofs, but the vote is clear and the community demands a reset. I did calibrate it and it *seems* that the game is rigged in reset's favor, but such are the rules at the moment. Therefore ... the reset has been processed. Enjoy!
Roundup of Day 2/ 2.1/ Approve for me (auto-review) is now included and does not use usage. Can be between 2-10% of plan when used. Also better for you. 2.2/ Simplified API for builders. 2.3/ Meeting notes integrated. 2.4/ Decisions API live for builders. Will use in the app to improve the experience.
2,237
646
15,488
1,349,626
16GB is not enough!!
CA is in a bubble. Micro Center's Mac mini M6 $699 deal is in stock in every state except California.
1
1
192
😭😭😭
Grok 4.7 from @SpaceXAI on ARC-AGI (Verified): - ARC-AGI-3: 1.8%, $2.7k (standard harness), 10.0%, $4.8k (provider adapter harness) - ARC-AGI-2: 61.4%, $2.01/task - ARC-AGI-1: 90.2%, $0.64/task Grok 4.7 scores higher than Grok 4.6 on ARC-AGI-1, but lower on ARC-AGI-2 and 3.
1
2
224
Frontier scale 😭😭
Frontier scale, built in Europe. Nearly 4,000 NVIDIA Grace Blackwell Superchips powered the training of Mistral Large 4, now in public preview. Congrats to the @MistralAI team 🙌
363
Sad
Mistral has released Mistral Large 4, scoring 38 on the Artificial Analysis Intelligence Index; France is back to having the most intelligent model from outside the US and China @MistralAI has released Mistral Large 4 in Research Public Preview, with plans to release the weights of the 1T parameter (49B active) model at the end of October. It achieves 38 on the Artificial Analysis Intelligence Index, comparable to GPT-6 Luna (max, 38) and DeepSeek V4.1 Flash (max, 39). It also achieves 50% on the Artificial Analysis Cyber Index, level with GLM-5.3-Flash and ahead of models such as Kimi K3 and DeepSeek V4.1 Flash (max). Key benchmarking results for Mistral Large 4 Preview: ➤ Most intelligent model from outside the US and China: Mistral Large 4 Preview scores 38 on the Intelligence Index, comparable to DeepSeek V4.1 Flash (max, 39) and GPT-6 Luna (max, 38). This makes it the most intelligent model from outside the US and China, ahead of countries such as South Korea and the United Arab Emirates ➤ Level with GLM-5.3-Flash on cyber defense capability: Mistral Large 4 Preview scores 50 on the Artificial Analysis Cyber Index, level with GLM-5.3-Flash (50) and behind MiMo-V2.6-Pro (56). Once its weights are released, it will rank among the top three open weights models on the Cyber Index. Its strongest result is on CyberGym-E2E-AA, where it scores 82%, ahead of MiMo-V2.6-Pro (79%) and GPT-6 Luna (max, 78%) ➤ Over 4x the Cost per Task of similar-intelligence open weights models: Mistral Large 4 Preview costs $1.13 per Intelligence Index task with standard pricing of $1.36/$4.18 per 1M input/output tokens, with $0.14 per 1M cached input tokens. For the first two weeks, Mistral Large 4 Preview will be served at a 50% launch discount, bringing its Cost per Task down to $0.57. This is still more costly than GLM-5.3-Flash ($0.25) and DeepSeek V4.1 Flash (max, $0.27) ➤ Strong document and image reasoning: Mistral Large 4 Preview scores 19% on GDP.pdf, on par with MiMo-V2.6-Pro (19%) and behind Kimi K3 (22%). This is an 18-point improvement from Mistral Large 3, partly driven by improvements in their API, which now accepts 100 images per request, up from 8 for previous Mistral models Key model details: ➤ Context Window: 512k tokens ➤ Multimodality: Text and image input, with text output ➤ Pricing: $1.36/$4.18 per 1M input/output tokens ($0.14 per 1M cached input tokens), with 50% off for the first two weeks ($0.68/$2.09) ➤ Availability: Research Public Preview on Mistral's API, with open weights planned for the end of October
1
2
462
Dickson Pau retweeted
It's pretty funny that it seems the one thing Mistral Large 4 is actually leading on is a benchmark about regulation Another EU banger
99
255
6,957
165,173
Okay but do I get more jobs down or not?
For a better means of comparison, you get over 2.5X the Opus 5.5 tokens on Claude's Max 20x plan than Sol 6.1 on Pro 20x
555
cc @thsottiaux you might as well just give us 28 resets right now ha
this is brutal
366
Misleading headline
Reflection AI, an artificial intelligence startup from two former Google DeepMind researchers, has unveiled a new open-weight model that it says rivals leading options in the US and China bloomberg.com/news/articles/…
301
Reflection's first model Beam is only at the GLM5.2 level?????
2
1
18
5,025
Reflection's first model Beam is only at the GLM5.2 level?????
2
9
7,635
Dickson Pau retweeted
Day 1/ We have optimized the default speed to be ~50% faster across GPT-6 Astra and GPT-6.1 Sol through the subscription across all our products and partners using Sign in With ChatGPT (including OpenCode, Pi, Amp, Devin, ...). No changes needed on your end and this should be felt within the next two hours.
Over the next 28 days, each day we’ll either ship one thing that is a clear improvement and relevant for most codex/work users or ship a full reset. Let the improvements begin.
2,604
1,028
20,707
3,981,264
Dickson Pau retweeted
Over the next 28 days, each day we’ll either ship one thing that is a clear improvement and relevant for most codex/work users or ship a full reset. Let the improvements begin.
All right, we’re locking in. Only things being worked on are simplifications, more efficiency for more usage, groundbreaking features or new models. Sometimes you have to invest ahead of the curve, but feedback is clear that you all want things to get simpler. On it.
4,516
1,492
26,153
7,596,505
The best model that is also served cheaply wins. It's really as simple as that.
Today my ChatGPT Pro x20 sub ended, and im not renewing it. I used Codex since its release basically, back then with GPT-5. Its a great harness, GPT-6 Astra is a great model, but i cant justify buying it when Opus 5.5 is a better model, and the limits for it are better too. Byebye ChatGPT, thanks for last few months.
1
2
295