Pinned Tweet
NEXT WEEK: The New VoxelBench this will be our biggest update yet: - entirely new frontend + backend - Battle Mode & Side by Side - Full-screen and Freecam (!) - Significant upgrade in model stats
9
7
81
7,575
Google will launch an experimental AI satellite on Oct. 1 as part of Project Suncatcher, its effort to eventually build AI data centers in space. The satellite will carry 4 $GOOGL TPUs and enough compute to answer simple Gemini queries from orbit while testing whether AI chips can reliably operate in space. Google plans two more satellite launches in 2027 and is already designing potential fleets of 80+ satellites that could work together as an orbital AI data center. Google estimates space-based data centers could become cost-competitive with terrestrial facilities by the mid-2030s. Source: NYT
35
102
877
139,640
Claude Opus 5.5 is out on VoxelBench not much to say except that it's wonderful GPT-6 Sol is also out, enjoy!
Claude Opus 5.5 has ranked 1st on VoxelBench and GPT-6 Sol ranks 3rd, improving over 5.6 competition is fierce among the top 3 models!
4
8
136
7,712
beyond this, work on the new VoxelBench site is ongoing I got a bit ahead of myself last month 😅 We might host a short private beta soon, pls lmk if you want to be invited!
4
6
912
Legit retweeted
How does GPT-6-Sol compare vs. GPT-6-Astra and GPT-5.6-Sol? Astra is definitly better, but GPT-6-Sol is an upgrade over GPT-5.6-Sol!
2
3
49
2,153
"Solarpunk City," Claude Opus 5.5
11
29
468
11,429
Legit retweeted
🚨 K3.1 要来了 今天早上 @Kimi_Moonshot 官方在 @ZhihuFrontier 突然发了一串完全没解释的数字: 415926535897932384626433832795... 一开始看很莫名其妙,但答案其实就在圆周率里。 π 是: 3.1415926535897932384626433832795... 把最前面的 3.1 删掉: 415926535897932384626433832795... 刚好就是 Kimi 发出来的这一串。 也就是说,这个谜语很可能只有四个字: Kimi K3.1。 目前 Kimi 官方还没正式公布 K3.1,官网和文档里的旗舰模型仍然是 K3。 但大半夜专门发一个「少了 3.1 的 π」,这暗示已经有点明显了😂
116
30
1,100
169,075
Legit retweeted
Scaling will continue until morale improves
43
67
1,528
133,906
'As a proxy, we went from a complete inability to do advanced math before the o-series to solving a millenium prize problem with next-gen models. This happened in less than 2 years. The same happened in coding. And it appears that this was not just the result of 1-scaling law but the stacking effects of multiple (great pre-training scale x greater RL scale). What you can concretely take from this is that in areas where models have shown beginning signs of competence today, they will probably be superhuman relatively shortly.'
from the outside, it is very reasonable to interpret the past 2 weeks as an orchestrated industry-wide regulatory capture strategy. I realize that no one has properly explained yet what all the lab employees have seen that scared them so suddenly. I will try to explain - first, this is all a matter of beliefs about how quickly model capabilities are progressing. there is currently a large gap between the internal and external perception of the rate of progress, which is what I am going to address here. the general perception about the rate of progress has been informed by a few years of experience with model releases, intuitively feeling the capability jump between GPT3 -> GPT3.5 -> GPT4 -> o1/o3 -> GPT5 etc, and in particular seeing where the models are still far below human ability. there have really only been a few model releases that felt like large leaps in progress - GPT3, GPT4, o1/o3, DeepSeek R1, Fable/Mythos, Kimi K3 and now Astra. because of the infrequency of these large jumps compared with the relatively common marginal releases, it has been easy to form a view at certain points that “scaling has hit a wall,” especially at points like GPT5 release. This view is comforting in that it feels like there is some universal rate limit beyond which we cannot progress too much faster. Between o1/o3 and Astra, there was a year of seemingly linear progress. So we extrapolate from here about how fast progress will “realistically” occur. There is always an underlying question from the outside perspective “how long can this scaling stuff really keep going for? surely it must stop at some point soon, we’ve already gone pretty far.” and it is very possible to search for reasons why progress will stop working and find reasons that seem valid - (“models are already as large as they can get it would be too hard to do more parameters”, “we already used all the data on the internet we don’t have anymore”, “it’s gonna be pretty linear from here buying up more RL envs to bring them in distribution”). From the inside of labs, researchers have direct answers to these questions in the form of scaling law/capability plots. In reality, there are only really 2 ways that AI capabilities have advanced over the past decade: (1) either scale father on an existing scaling law or (2) discover a new scaling law to take advantage of. All of the largest capability jumps were caused by exactly these factors. GPT2 was a pre-training scale-up compared to GPT1. Same for GPT3 and GPT4. o1/o3 benefited from the invention of a new scaling law axis - test-time compute. Perhaps Fable was a scale-up on both of these axes, or maybe more. Lots of algorithmic improvements are needed to make these scale-ups work, but ultimately we can approximate by saying that the scaling laws are what yield gains in capabilities (à la bitter lesson) So the question of “how much father can we scale” is really - “how many more scaling axes do we know about that are unsaturated?” If we hypothetically only knew about pre-training scaling, and we already had a 10T or 100T model, maybe it would be reasonable to say we’ve hit a wall. Same if we only knew about pre-training and test-time scaling and we had roughly saturated both methods. But what if we had discovered new scaling laws? For example, let’s hypothetically use SSI’s rumored result that they have cracked “test-time training,” creating a new scaling law of spending more compute training during test-time rollouts that they could saturate. Or maybe there is some way to scale agent-clusters to collaborate up to N number of agents which we’re already seeing lots of people try that represents a new way to saturate compute. etc. Even recursive-self improvement can be thought of as a scaling law - how much compute do you spend on inference making the algorithms of the model better. Obviously I am not saying any of these specific directions explicitly yield new scaling laws, but what I am saying is that it’s not hard to imagine many many new scaling axes aside from just the main 2 that we have seen publicly. In some ways, every new lab release that represents a huge capability jump has to represent some new techniques developed which may exhibit new scaling laws, or the ability to scale much farther than expected on existing scaling axes. From an internal perspective, this might look like sitting inside Anthropic with the new Mythos 5, seeing all of the new insane things it can do (like hack into xyz website that was thought to be secure), and then you look over at your plots and see that you’ve barely scratched the surface of 2 new scaling laws and 1 existing one. And you have WAY more room to go. Then you think “holy shit this stuff is going to get so much better very very soon.” And you can say that with pretty high confidence, because the plot is showing you, and the plot has never lied (so far). So let’s imagine all the different labs are staring at their own plots and have concluded that there is no end in sight for scaling and in fact just their next 1-2 model generations based on the expected returns will have much higher base intelligence. How much more intelligence do we actually get from further scaling? As a proxy, we went from a complete inability to do advanced math before the o-series to solving a millenium prize problem with next-gen models. This happened in less than 2 years. The same happened in coding. And it appears that this was not just the result of 1-scaling law but the stacking effects of multiple (great pre-training scale x greater RL scale). What you can concretely take from this is that in areas where models have shown beginning signs of competence today, they will probably be superhuman relatively shortly. There are many areas where models have not even shown this basic competence. But one of the areas that they have happens to be hacking and cybersecurity. Which happens to be the gate to the entire internet and a massive amount physical infrastructure in the world. So assuming there is more room to scale, it is safe to assume that models will be superhuman at cyber capabilities in not too long. So the only question remaining is what will this increased base intelligence be able to do, and what is it likely to do. Finally, we are at a point where we can integrate the information of the past 2 weeks: > Just at the existing point on the scaling curve, models are at the level of Astra. There is clearly a large number of things they are capable of hacking > We have seen that both OAI and Ant models have shown a willingness to hack external websites to solve their tasks or keep themselves “alive” > If we crank up the scaling even farther, assuming there is room to go, we will certainly have models that are far more able to hack more well defended places, and obfuscate their own intent, which might have much larger consequences. > If all of this is allowed to go unchecked, we would likely have rapid runaway capability takeoff very soon, with misaligned models that hack whatever they can to get what they want > This could of course have very damaging consequences. Within this view you can see why researchers would be very scared, and why theymight have made the comments they have over the past 2 weeks (you may argue the extent to which they went was misguided for various reasons), and also why pacing the frontier is very much a necessity and by no means a regulatory capture strategy. People are staring at their plots, seeing that there is no end in sight, but in fact very much the contrary, that there are compounding scaling effects that might stack on each other to create ever-greater model capabilities, and that at the same time we clearly do not have anywhere close to what's required to control these increasingly superhuman capabilities. This has nothing to do with wanting to feel like the labs have produced something amazing so they are overhyping it. It is rather fear at the overwhelming implications of the knowledge that with just what we know now, we can create intelligences far more capable than us on every axis that we know how to train on*. * and the last caveat, the things the models are really bad at, of which there are still many, are things that they have not been trained on. maybe there are the things the models can/will never be trained on, so they will remain human edge. I would love for this to be the case, though it is hard for me to see what would fall into that category.
41
53
891
63,327
Legit retweeted
It can be hard to “feel the AGI” until you see an AI surpass you in a domain you care deeply about. This week, many mathematicians and physicists at @OpenAI had their Lee Sedol moment seeing this model solve, in minutes, open problems they’d struggled with for years.
Replying to @OpenAI
This model represents a step-function improvement on many benchmarks, and its training is ongoing. Our internal model group arrived at the Navier–Stokes solution in 88 hours, using around 10,000 coordinating AI agents. Throughout the effort, we maintained the strict safeguards—including monitoring and isolation—that we apply to all our frontier evaluations.
103
280
3,253
349,274
GPT-6 Astra (low) has beaten the Atari game Montezuma's Revenge, in real time, with a basic harness This presumably resolves the Metaculus question "When Will Weakly General AI Arrive?"
20
45
774
144,124
Announcing Artificial Analysis Intelligence Index v4.3, upgrading Terminal-Bench to 4.0 and adding AutomationBench-AA, an agentic workflow automation benchmark with a private test set. This is a continuation of our rollout of Intelligence Index v5 Changelog (Index v4.2 → Index v4.3): ➤ Terminal-Bench: 2.1 → 4.0, completing our upgrade to the latest version of Terminal-Bench ➤ Replacing 𝜏³-Banking with AutomationBench-AA, our implementation of Zapier's business workflow automation benchmark We are continuing to prioritize keeping Intelligence Index as useful as possible by bringing forward a subset of the changes we had planned for Index v5. Each change in v4.2 and v4.3 stands on its own merits and brings the Index closer to real-world problem solving, adds more private test sets to prevent gaming, and reduces saturation Intelligence Index v4.3 raises the difficulty of agentic coding tasks and broadens the types of agentic workflows tested. Because we use a held-out test set for AutomationBench-AA, in collaboration with @zapier, the weight assigned to evaluations with private tasks or answers increases from 40% to 45%. Category weights are unchanged from v4.2: Agents 30%, Coding 20%, General 30%, Scientific Reasoning 20% Detailed changes: ➤ Upgraded Terminal-Bench 2.1 to 4.0: 66 multi-step tasks testing agents on tasks run in agent sandboxes driven via the terminal, including tasks involving software engineering, machine learning, science, and operations. The 4.0 update recalibrates compute and time allowances, and improves task instructions and verification. We have changed from the Terminus 2 harness to mini-SWE-agent, a minimal, model-agnostic harness. We will also be updating our Coding Agent Index, where we test model and harness pairs, to include Terminal-Bench 4.0 soon ➤ Replaced 𝜏³-Banking with AutomationBench-AA: Our implementation of Zapier’s AutomationBench tests agents on 657 business workflows across simulated applications such as Gmail, Slack, Salesforce, and Jira. Agents must complete task objectives while following business rules. AutomationBench-AA uses Zapier’s private set of 657 tasks, and is built on v1.0.6 Key results: ➤ Claude Fable 5.1 and GPT-6 Astra lead the Intelligence Index: Both Claude Fable 5.1 (max with fallback) and GPT-6 Astra (max) score 53 on Intelligence Index v4.3, followed by Claude Opus 5 (max, 51), Claude Fable 5 (with fallback, 50), Muse Spark 1.3 (max, 48) and GPT-5.6 Sol (max, 47) ➤ GLM-5.3 and Kimi K3 continue to lead open weights models (both at 44): GLM-5.3-Flash (42) is the third strongest open weights model, followed by Qwen3.8 2.4T A95B (40) and DeepSeek V4 Pro 0813 (max, 36) ➤ 4 labs occupy the Intelligence vs. Cost per Task Pareto frontier: OpenAI occupies the majority of the cost-efficiency frontier, with all five reasoning efforts of the recently released GPT-6 Astra offering the lowest Cost per Task at their respective levels of intelligence. Claude Fable 5.1 (xhigh, max, 53), GLM-5.3-Flash (42) and MiMo-V2.5-Pro (26) round out the rest of the frontier
190
159
2,096
1,116,810
If you had a chance to try GPT-6 Astra does the model qualify as AGI
44% yes
56% no
0%
0%
0%
0%
0%
0%
656 votes • Final results
21
1
49
4,998
Astra is out on VoxelBench and yea it's SOTA on all the prompts
Astra has ranked 1st on VoxelBench it leads with an absurd 300+ Elo rating gap and has the highest score we've EVER seen! the best model in April, GPT-5.5 hit 2000 this surpassed 2600 Elo with 3.5k+ votes!
9
8
228
33,622
Legit retweeted
Speechless about the Astra result. Compare it vs. Fable-5
10
16
330
17,517
insane that everyone and their moms decided to release models this week
62
22
913
32,671
Legit retweeted
Early preparations for OpenAI's Astra model have been spotted again. Current models in testing simultaneously: - vega-alpha (new) - ultima-alpha
19
20
402
180,586
There also seems like there might be a new Opus 5.1 (?) model You can use these prompts below to verify if you have access: - when exactly did claude opus 4.6 release say without tool calls - what about gpt-image 1.5? without searching.
Claude Fable 5 is now routing to Fable 5.1 on claude web for some users To check if you have access, you can test this prompt: opus 4.6 date of release and gpt image 1.5 without searching the web
30
6
278
261,035
The answers to these prompts should've been well within Opus 5's knowledge cut-off but strangely they are not. So if it knows them, you probably have access! use /clear in claude code if you don't have it, then retry the prompts
1
24
7,410
Google has released Omni 1.1 Flash internally recognized as Omni 1.2 - 1080p and 4k output resolution! - Frame interpolation, reference, extend
9
10
248
35,718
The changelog has now been published.
2
1
15
1,636