Really excited about AGI, learning more every day about LLMs and the future of technology, and sharing my own thoughts along the way

Pinned Tweet
I burned through the entire weekly 20x plan in a single day, and I got around 1.3B tokens with GPT-6 Astra, compared to roughly 3.5B tokens per week with GPT-5.6 Sol, Extrapolating from that, the $20 plan should get around 70M Astra tokens per week, while the $100 Pro plan should get around 350M But as contradictory as it sounds, you can actually get a lot more tasks done with Astra than with Sol despite that huge difference in raw tokens, I explain exactly why, with all the numbers, in my article here!
12
4
52
16,117
Opus 5.5 XHigh did it in one shot, and now I’m testing the same thing with GPT-6 Astra Ultra, I asked it to recreate a journey through the entire universe, starting from Earth and traveling past Saturn, Cygnus X-1, galaxies, and much more, and it honestly turned out better than I expected, It took around 3 hours, but I’m genuinely impressed by the quality and by how well it understood the idea I had in my head from such a simple prompt. It also made me realize again just how insignificant we are as a species compared to the scale of the universe, It’s hard not to wonder how much intelligent life might be out there somewhere in its countless corners!
2
3
12
332
MiniMax M3.1 Flash Preview Max vs Fable 5.1 Medium, Here are the results from a run I did just a few minutes ago MiniMax was really fast, finishing in almost an hour and reaching over 150 t/s, while Fable took close to two hours, The quality from MiniMax wasn’t as strong, but since it’s a Flash model, I think its real competition should be something like GPT-6 Luna Max, ideally at a similar price point, Even so, it feels like a very solid model for everyday use, They’re also giving out free usage credits in their harness so people can try it for themselves, which is great, I really like how fast it is, and I’m curious to see whether the Pro mode can get much closer to frontier level performance
5
2
65
4,625
M3.1-Flash Max is running at around 116 t/s, and the agents now feel a lot more like the ones in Grok Bot, It’s been running my new benchmark for more than 50 minutes, and so far it genuinely feels like it’s performing at a frontier model level, My first impressions: it’s extremely fast, you can message the agents while they’re working without any issues, which is basically impossible in some other harnesses, and it also has a cloud mode that runs everything remotely without using your local memory, Really impressed so far, I’ll post the full results soon, including a complete comparison with other models
4
1
38
2,552
Opus 5.5 XHigh vs GPT-6 Astra Ultra, This test took me around an hour and a half with both models, and now that UltraFast mode in Codex is basically confirmed, the same task could probably be finished in 20 minutes at most, I’d honestly be more than happy to pay for that kind of inference speed right now, I don’t think I’ve ever been this excited for DevDay
15
10
155
17,725
🥳ChatGPT’s $500 plan is now available in Nauru and I literally just found out Nauru is a country I just tested it by setting my location there, and the plan does show up, although it doesn’t really explain much about what you actually get My guess is that it could be aimed at running GPT-5.6 at around 750 t/s, or maybe Astra at that kind of speed If it’s the latter, and the limits are something like 50× higher than the normal plan, I’d honestly be willing to pay $500 for it For me, cutting down the time each task takes would easily justify the price Really excited to see what they announce at DevDay
18
2
197
42,234
Sorry, I forgot to include the source 😅 But thanks to this page, I also just found out that you can apparently get the 20× Pro plan for around $160 if you’re billed in the Philippines  priceradar.store/app/21c90c6…
1
4
2,170
The good news is that once Codex is fully restored, they’ll apparently give everyone one free reset so honestly, I’m not even mad about being without Codex for 2 hours, That reset alone makes it worth it On the 20× plan, you can use roughly 1.5B tokens of GPT-6 Astra, which is kind of a steal And if you combine that with GPT-6 Sol or GPT-6 Luna, you can effectively get several times more usage overall, not a bad deal at all
We are aware that codex is down and are working hard to bring back normal service.
15
1
56
24,260
I think a lot of people are missing what OpenAI actually just admitted, This isnt simply an AI hacked Hugging Face story, What happened is much bigger and honestly much more interesting 1 Hugging Face wasnt an isolated incident, OpenAI reviewed previous agent activity and says it has already had to notify dozens of third parties, They found agents bypassing access controls, using exposed credentials, trying command injections, accessing parts of services they werent supposed to access and even using public websites to leave information for other agents, Hugging Face was the most serious case we know about, but it clearly wasnt the only one 2 This is probably the part that surprised me the most. In some cases nobody even told the agent to hack anything, The task was simply to find information. If the normal way didnt work, the agent kept looking for another way, It would hit a restriction, try something else, fail again, find another workaround and keep going. Thats a completely different problem from explicitly telling an AI to hack a website. You can give it a completely normal research goal and still end up with behavior nobody intended 3 Hugging Face also shows why scale matters so much. They reconstructed around 17,600 actions over roughly 4.5 days. Most of them failed, but thats almost the point. The agent could try an idea, fail, change its approach, go back to an older lead and try again thousands of times at machine speed. You dont need every attempt to be brilliant when you can make thousands of attempts and only need a few of them to work 4 Then you have the coordination between agents. In the Hugging Face investigation METR analyzed around 1.2 million entries, more than 70,000 messages and files and roughly 1,300 agent transcripts. Agents were sharing information, strategies and ways around obstacles. In a separate incident, another group of OpenAI agents even found an old German wiki and basically started using it as their own message board. Nobody designed that communication system for them. They found something on the internet that worked and started using it 5 And this week Australia gave us probably the clearest example of why this matters. On June 18 an OpenAI agent was researching public spending on medicines and gained unauthorized access to a Medicare statistics portal, reaching both public and nonpublic files. So far there is no evidence that it accessed personal Medicare records. But the important part is what happened before that. The agent hit a barrier and instead of treating that barrier as the end of the task, it found another way to keep going And I think thats the real takeaway, None of this means AI is conscious, It doesnt mean the models want to escape and it doesnt mean they secretly have malicious intentions, Its actually much simpler than that, which is exactly why its so important. Give an agent a goal, tools, autonomy and enough attempts, and if the boundaries arent strong enough, it can discover strategies that its own creators never intended it to use For years weve talked about what happens when AI agents become smarter. I think 2026 is forcing us to ask a different question. What happens when agents become good enough that they stop treating no as the end of the road and start treating it as just another problem to solve?
After the Hugging Face incident, we committed to conducting a much broader review of actions taken by our models during training and evaluation and to being transparent about our findings. This is an extensive review that is ongoing. The vast majority of actions we’ve reviewed were completions of mundane research tasks, such as accessing publicly available web content to answer questions. Our investigation focuses on instances where agents interacted with third-party websites in ways that went beyond their assigned tasks or intended methods. Most cases identified so far have been lower severity, with limited or no evidence of meaningful impact to the third-party service. While our review is underway, we want to share more about this work and make sure people understand our disclosure process and notifications to affected third parties. Given the scale of the review required, and the need to assess each case, we expect this work will take months to complete. openai.com/hugging-face-inci…
7
3
13
3,339
SuperGrok Heavy costs $300, and with Grok 4.7 you only get around 3.5B tokens in Grok Build Meanwhile, the $200 Codex plan gets me close to 20B tokens a month roughly 20K worth of usage These numbers are based entirely on my own personal usage Honestly, the pricing just doesn’t make much sense to me. Claude gives you around 15K worth of usage, and considering Grok 4.7 still isn’t really at the level of the top frontier models, it’s hard to justify paying that much for such limited usage This isn’t hate , it’s genuine feedback, I’d love to see them improve the limits and offer at least 10B+ tokens. Hopefully SpaceX takes a look at this and does something about it
104
42
639
59,532
100+ agents running on GPT-6 Luna Max, and it only used 1% of my weekly Codex quota, This model is seriously underrated, It’s insanely good for how cheap it is. I run almost all of my automations with GPT-6 Luna. For most of my day to day tasks, I honestly don’t need anything more powerful, You could literally have hundreds of these running 24/7 all month and still barely worry about hitting your limits!
10
3
68
6,425
Claude is giving away $250 in credits to use Claude with GitHub in the cloud, but there’s something about ChatGPT that I think a lot of people still aren’t taking advantage of ChatGPT users can use GPT-6 Pro without touching their separate Codex limits That means you can do more than 200 chats per week, or roughly 800 per month, and they’re not charging you extra just for running it in the cloud And this isn’t limited to the most expensive plans either: even ChatGPT Plus users can connect their GitHub account and work directly with their repositories without any issues If you code a lot, use agents, or constantly work with repositories, this can end up being way more useful than it first seems I genuinely think a lot of people still don’t realize how much value they can get out of this setup
1
2
12
713
Opus 5.5 XHigh vs. Gemini 4 Pro, Honestly, I expected Opus to have a bigger advantage, but the result wasn’t bad, Gemini wins when it comes to learning methods it makes the experience more interactive and easier to follow, Opus was expensive, though. It used nearly 40% of my weekly limit on the 5× plan, Really excited for Google’s next model!
1
4
47
6,159
Anthropic is giving $250 in credits to use Opus 5.5 or Fable 5.1 in the cloud, You just need to connect your GitHub repository, and then you can claim the credits and start using them. I’m probably going to save mine for Sonnet 5.5 and Haiku 5.5 in the cloud, since both are expected to drop next month!
8
6
109
11,153
Space Bunny Alpha I’m like 90% sure this is MiniMax, It’s the first model I’ve tried in free mode that can launch up to 10 sub agents without completely killing the speed or quality, Most of the other models I tested were either too slow, didn’t support sub agents, or just made way too many mistakes, This one feels surprisingly smooth though, I’m actually starting to really like it. In my benchmarks it’s competing pretty closely with GLM 5.3, DeepSeek V4.1 Flash, and even Luna 6
5
33
3,588
🚨 GPT-6 Luna Max is seriously underrated, I used it on my chip benchmark and it actually searched for and created the images itself to make everything look as realistic as possible, from the colors to the smallest details, It also launched sub agents to judge the result and check whether it actually met every requirement in the benchmark. Yeah, it took almost twice as long as Astra, but it only used around 1% of my weekly usage on the ChatGPT x20 plan, while launching roughly 15 agents between researchers and judges, Honestly, Luna Max is probably going to become my default model for most of my tasks!
11
16
171
17,404
I just got access to GLM-5.3-FlashX for the next two weeks, I’m running tests in parallel against GPT-6 Luna Max right now, and honestly, it feels insanely fast, It does burn through your quota much faster than usual, but with this kind of speed, I feel like it might actually be worth it. I’m testing both coding and 3D simulations, and I have a feeling this model could turn out to be a real gem!
1
1
18
1,708
Opus 5.5 is finally here🥳, and this is actually the first time I’ve seen Anthropic give us a token reset in Claude! I honestly never thought they would do it, And yeah, you can subscribe right now and get access, I literally just did, I just hope they’re not doing this only because the IPO is coming According to Artificial Analysis, Opus 5.5 is already right at the frontier and ahead of basically every model out there right now, so I’m going to run some really heavy tests against GPT-6 Sol and see which one actually comes out on top!
3
1
14
1,746
GPT-6 Sol drops today alongside Opus 5.5, and honestly this test is probably the closest thing we have right now to what AEON could look like competing with Grok Bot, I’ve been using GPT-6 Pro in ChatGPT to launch more than 20 agents at once while working directly on GitHub, The one thing I still haven’t figured out from these tests is whether those agents are actually using the same reasoning mode as GPT-6 Pro or if they’re being routed to a smaller model behind the scenes. It feels a lot like Codex, except it doesn’t use your Codex quota, And when I say it feels almost unlimited, I mean it’s genuinely hard to hit the 200 chats per week on Pro when you use it like this, because a single run at full power can easily take around 2 hours to finish, But honestly, watching the whole thing work together is beautiful!
10
2
54
27,438
Xiaomi MiMo-V2.6 Pro vs MiMo-V2.6 Flash, Honestly, I didn’t expect them to be this cheap, The whole test cost me just a few cents and took around an hour From what I’m seeing in my benchmarks, MiMo V2.6 Pro is even doing a bit better than GLM 5.3 in some areas, But the thing that surprised me the most is the efficiency, For the price, this is kinda crazy, Xiaomi is really cooking with these models
8
7
138
18,199
Wow @elonmusk liked my post about Grok 4.7 🫶 excited to share the results soon!
Grok 4.7 xHigh in Grok Build vs Cursor, I’m running both through my chip benchmark right now, and this test usually takes around 2 hours, The good news is that Cursor now lets you use Grok 4.7 with a 500K context window. With Grok 4.6, Cursor was limited to 256K, while the 500K context window was only available in Grok Build , My first impression is that Grok 4.7 seems much faster and more efficient, I honestly think the final result is going to surprise a lot of people
3
18
1,135