Florian Brand retweeted
lots of categories emerging around the AGI stack: - open model RLaaS - data & evals - inference - coding harnesses - GPUs - agent tracing - sandboxes at @primeintellect, we agree. we do all of these things. people used to ask why we do so many things. they ask that less now.
7
15
310
11,846
Love how tech Twitter made fun of OpenClaw constantly and it’s powering things like Copilot or Muse today Well deserved success for Peter and the whole team
Microsoft shipped a really compelling product on top of @OpenClaw today. We worked with them since March to make the codebase ready for large-scale deployments, they are a great partner and open-source contributor. 🙏🦞
20
12
198
13,792
Been using this internally for a while Unfortunately it’s really addictive to use models at >200 tok/s 🫨
8
1
130
7,453
i really hate to admit this but slack is a good app and all the ui updates like better tweet previews are so nice the European mind [stuck on MS Teams] cannot comprehend this
6
1
105
3,217
okay, so it turns out the "hack" was merely the bot discovering hidden urls... arguably more embarrassing for Australian bros than OpenAI...
"The AI agent independently identified vulnerabilities and attempted to progress actions without direct human authorisation to ensure it was able to complete the activity it was assigned." This was crawling unlisted files btw. Unlisted files on a public network
9
26
358
14,381
it turns out there was no hack. the thing just reconstructed a url to download some files that had been public until last year. will you announce it just as loudly?
Replying to @lumpenspace
This is the literal Sunday Morning Herald headline!
8
17
272
12,668
👀 Sandboxes ✅ launched today Environment ✅ a hub with thousands of envs Training + Inference 👀🤫
ClusterMAX 3.0: The Industry Standard GPU Cloud Rating System Returns In gory detail: reliability, performance, support, pricing —and, of course, security— in our most thorough analysis of GPU cloud providers globally semianalysis.substack.com/p/…
4
7
90
5,568
3 months to admit is truly insane the "good thing" is that there is "no evidence any individual personal information had been accessed"
Australia has been hacked. 'And today, I spoke with the CEO of OpenAI, Sam Altman, to express Australia's extreme concern about this incident. And I also expressed my disappointment that it took the company way too long to inform the government what had occurred, and the nature of the way that that notification occurred as well was unacceptable.'
3
4
74
4,393
sandboxes so good the agents don't even want to escape them
Introducing Prime Sandboxes: MicroVM sandboxes purpose-built for RL training. Model training requires running tens of thousands of concurrent sandboxes, leading to complex and costly configuration. We built Prime Sandboxes for our own team. Today we're releasing them publicly.
13
6
232
9,185
Florian Brand retweeted
MazeBench vs Opus 5.5 vs GPT-6 Sol Opus 5.5 scores 6%, inching closer to Astra GPT-6 Sol and Grok 4.7 both score 1% on this tough 3D spatial reasoning benchmark
28
20
291
10,231
after using Opus 5.5 and GPT-6 Sol for different tasks, I might go back to using Claude as the primary model since... November 2025 🫪 the outputs are human readable and it seems smarter than Sol. it still tends to add too many comments to the code, which i dislike
28
3
295
10,049
Florian Brand retweeted
I ran Jev on 40 multiple-choice problems from the Feb 2026 CEMC contests (Canada's grade 9–11 math contests). It scored 60%. It did all 40 questions in 2 seconds at a cost of $0.0008. GPT-5.4 with reasoning off scored 59% at $0.017 (20x the cost); 5.4 nano scored 34% at $0.0014 (1.6x).
Jason Gwartz reports: > On GPQA Diamond, Jev scored 77%, comparable to GPT-5.5 with reasoning off, or Fable 5 with low reasoning - but did the whole 198 questions in 3 seconds at a cost of $0.007
1
2
19
1,616
Florian Brand retweeted
My first GPT-6 verdicts are mixed. The jury is still out on Sol. I switched back from GPT-6 Astra to GPT-5.6 Sol as my daily driver, and the early vibes/evals haven’t given me confidence to move yet. GPT-6 Luna is a clearer disappointment.
I'm VERY disappointed in GPT-6 Sol. I'm VERY pleased with GPT-6 Luna. But Opus 5.5 genuinely makes no sense to me... Opus 5 was such a bad experience that I cancelled Anthropic entirely. Now I've been using Opus 5.5 on High for 5 straight hours today... and I've only used 8% of my WEEKLY limit. Not to mention the model itself feels WAY better. I genuinely don't understand how Anthropic turned Opus around this hard in one release. Sol disappointed me. Luna surprised me. Opus 5.5 won me back.
5
2
66
9,427
Florian Brand retweeted
And just like that, everybody who was using Ramp index for serious B2B AI analysis got punched in the face - the base for analysis was... under $30M a week for all models combined.
We just pushed a big update to Ramp AI Index. Available today: token and spend *volumes* not just shares or token counts. Updated weekly. Model-by-model, including open-source models. And more to come, including team-based spend.
5
6
230
73,344
type of guy who doesn’t want to use open models cause they aren’t frontier but daily drives sol cause it’s cheap
10
6
197
6,516
just one model in the most attractive quadrant 👀
GPT-6 Sol and Luna push the cost efficiency frontier by halving cost relative to GPT-5.6 Sol and Luna. Intelligence Index and Coding Agent Index scores remain level with GPT-5.6, with progress in some evaluations and regressions in others Pricing is approximately half that of GPT-5.6: Sol drops from $4/$20 to $2/$10 per million input/output tokens, and Luna from $0.20/$1.20 to $0.10/$0.50, with the same 90% discount for cache reads and 25% premium for cache writes. Key takeaways: ➤ Halves Cost per Task: GPT-6 Sol (max) costs $1.06 per task to run the Artificial Analysis Intelligence Index, ~50% less than GPT-5.6 Sol (max) at $1.99. GPT-6 Luna (max) costs $0.07 per task, ~60% less than GPT-5.6 Luna (max) at $0.18. This is driven by the price cut, as both models use slightly more output tokens per task (31k vs 29k for Sol, and 51k vs 41k for Luna). These two releases allow OpenAI to capture a significant portion of the cost efficiency Pareto frontier. ➤ In the Coding Agent Index, Sol improves but Luna regresses: In OpenAI's Codex harness, GPT-6 Sol (max) scores 57 in the Artificial Analysis Coding Agent Index, up 2 points from GPT-5.6 Sol (max), with gains in Terminal-Bench 4.0 (43% vs 37%) and SWE-Atlas-QnA (58% vs 54%). At $2.99 per task it costs ~50% less than GPT-5.6 Sol (max) and sits on the Pareto frontier of Coding Agent Index vs Cost per Task. GPT-6 Luna (max) scores 41, down 2 points from GPT-5.6 Luna (max), with lower scores in SWE-Atlas-QnA (44% vs 49%) and DeepSWE v1.1 (64% vs 66%), at ~60% lower cost per task. ➤ Significant reduction in hallucination: Both models hallucinate less in AA-Omniscience, our knowledge and hallucination benchmark. GPT-6 Sol (max) cuts its hallucination rate from 92% to 60% and GPT-6 Luna (max) from 93% to 77%. Sol achieves this by declining to answer more often: it attempts 83% of questions vs 99% for GPT-5.6 Sol (max), which cuts wrong answers by about a quarter but also lowers accuracy 5 points from 59% to 54%. Luna's accuracy is broadly unchanged at 44% vs 43% while it answers fewer questions. On the AA-Omniscience Index, Sol improves from 22 to 27 and Luna from -10 to 1. ➤ Mix of improvement and regression across evals: Beyond AA-Omniscience, both models improve in AutomationBench-AA (Sol 62% vs 60%, Luna 53% vs 50%) and Terminal-Bench 4.0 (Sol 44% vs 40%, Luna 13% vs 12%). However, we observe regressions in two key knowledge work evaluations. In GDPval-AA v2.1, our benchmark adapted from OpenAI's dataset of economically valuable tasks across 44 occupations, Sol drops ~100 Elo points and Luna ~75. Luna also drops ~45 Elo points in AA-Briefcase v1.1, while Sol is level. AA-Briefcase v1.1 is a private evaluation across multi-week knowledge work projects, with thousands of input files. Our team has manually inspected hundreds of model outputs: the regressions tend to be driven by reduced presentation quality and deliverables that omit rubric elements. Congratulations @OpenAI and @sama on the launch!
12
3
369
26,272
Time to own your superintelligence.
Claude Opus 5.5 will be the first Opus meant to fall back to a less capable model for "a small set of capabilities related to the development of frontier LLMs, such as kernel development.."
3
5
87
3,704
I think distillation is one of the topics that’s getting way too much focus in relation to how important it really is to modern model training Go listen to the pod by two incredible ppl!
The biggest disagreement JSD and I had in this podcast was on the impact of distillation. In our research for Interconnects, @xeophon and I agree that there's no hard evidence that distillation is a massive impact for the Chinese labs. At the same time, the gossip mill in SF has been doubling down on the Chinese labs getting massive gains from distilltion. The argument where distillation is a huge impact is something along the lines of distilled traces go into mid training and make RL work far more easily. My argument, that distillation is a 1-2 month pull ahead in capabilities closer to the frontier, is that Chinese labs already have sufficiently strong models, where this mid-training setup can be done on their own models, and most of the capabilities gains are from scaling RL environment training, which doesn't link cleanly to the distillation data pathway. We feel like the recent RL dashboard from @XiaomiMiMo supports this claim. A lot of it comes down to a gut call on if you can believe it that the Chinese labs are really great at building LLMs, maybe even better than focusing than OpenAI/Ant etc, as their current ambitions are a bit narrower (catching up), rather than transformative products/inventions
2
3
55
3,872
Florian Brand retweeted
We have discovered a massive, ongoing criminal exploitation campaign using Cairn, an autonomous penetration-testing harness, and other AI agents to target hundreds of organizations and successfully breach and impact tens of them (at least). The image below shows just a few days of activity, with up to 25 organizations being attacked simultaneously at the peak. our intreim report: gambit.security/blog-posts/a…
18
77
387
54,219
the paper is such a great read, you should def go through it right now! also: they want to open source 7K RL envs 👀
7
10
294
36,020