yells at @theo both on and off camera (co-host of the nerd snipe podcast)

San Francisco
Pinned Tweet
My vid on Grok 4.7: it's rough The model is really nice to interact with, it's very capable, but man the efficiency and price is just bad...
11
2
143
9,507
Ben Davis retweeted
I think I’m done mourning the death of programming. I’m ready to love again. I landed like 40 PRs while drinking with my friends in Scotland. This is the dream.
119
118
3,118
178,319
Ben Davis retweeted
SuperGrok Heavy costs $300, and with Grok 4.7 you only get around 3.5B tokens in Grok Build Meanwhile, the $200 Codex plan gets me close to 20B tokens a month roughly 20K worth of usage These numbers are based entirely on my own personal usage Honestly, the pricing just doesn’t make much sense to me. Claude gives you around 15K worth of usage, and considering Grok 4.7 still isn’t really at the level of the top frontier models, it’s hard to justify paying that much for such limited usage This isn’t hate , it’s genuine feedback, I’d love to see them improve the limits and offer at least 10B+ tokens. Hopefully SpaceX takes a look at this and does something about it
104
42
639
59,491
Ben Davis retweeted
The G in "AGI" stands for Jev
108
30
1,193
43,944
Getting bored of the standard one shot games/3d stuff, I think at least in the medium term the much more interesting side of things is the modding I had Opus 5.5 take the Majora's Mask recomp and build a full new area. It's got puzzles, enemies, rewards, and a mini dungeon...
36
11
426
29,362
This was ~15 prompts over a 4 hour period. Did a decent amount of steering, but most of it's just Opus 5.5 being insanely good at what it does
2
30
2,263
Ben Davis retweeted
You people are insane
I hate to break it to everyone, but Opus 5.5 has been nerfed starting the last few hours today. Autistic Claude speak is back, it's doing really dumb shit & this is not the same model I was using the last 2 days. Confirmed this on multiple accounts, it's turned into an idiot.
93
21
1,812
203,425
Opus 5.5 is the best model Anthropic's released since Opus 4.5 It's super natural and easy to interact with, writes great code, good at simplifying things, understands what I want, and most importantly the limits are absurd. It feels nearly unlimited...
31
6
690
16,641
I've had 5+ threads going all afternoon. Several of which are doing 5+ subagents each. Running for hours, all on xhigh reasoning
35
1,685
Anthropic won this week for sure The more I use it, the more impressed with Opus 5.5 I am - Great limits/price - It's writing/communication is great - Code quality is great And it's just super nice to work with
51
17
1,014
22,668
My vid on Grok 4.7: it's rough The model is really nice to interact with, it's very capable, but man the efficiency and price is just bad...
5
2,123
Grok 4.7 was a mess (especially the efficiency/price for it's capability level sucks) GPT-6 Sol/Luna are good, but kinda boring. They just work. They're better than last gen. They're efficient and cheap. Add the rest of the OpenAI points and that's pretty much it. Good models.
3
84
2,530
Success
Smuggling cheaper compute from the Ohio micro center back to SF
21
244
8,028
Goodbye Terra 🫡
GPT-6 Sol and Luna just landed in Astra’s orbit. Both launch today with API prices 50% lower than GPT-5.6. Build with Sol. Scale with Luna. To production and beyond.
54
62
2,880
102,655
Ben Davis retweeted
Periodic sadness reminder that the anthropic subscription still does not permit third party harnesses :(
93
71
1,398
63,589
Ben Davis retweeted
day 1 observations for grok 4.7 ignore the reports that say “it’s terrible” and the only thing they reference is a public benchmark. the same benchmarks told us opus 5 was better that fable - they are useless also ignore the reports that compare models with 3d games - that’s not real work. it's made for attention on social media i used grok 4.7 for a whole day as my firstmate, and it has been a really solid model with visible improvements over 4.5 (i'm ignoring 4.6 because 4.5 has been working better in my experience) key differences with 4.7 - 1. it follows system prompt very, very closely i noticed firstmate showing many new behaviors that i've never seen before, such as asking me to name specific red CI checks that i'm ok with bypassing, and refuse a simple "yolo" instruction i traced it and it's indeed how i instructed it in firstmate's system prompt, but none of the other models followed it closely enough to make this behavior visible - grok 4.7 is the first to pick that up there were a few other similar examples as well. so to me this is a clear behavioral difference 2. it's very "stable" if you've used astra then you know what a "spiky" model is. it can have some genius moments but you occasionally also wonder "how could it be so dumb and doesn't get me". grok 4.7 is the opposite of that throughout the whole day so far, i'll be honest i haven't get a "wow this is absolutely genius" moment yet, but grok 4.7 has been very steady with no big surprises. its behavior feels predictable, which does help it gain trust from me quickly 3. it's a conservative model it doesn't like to take actions without asking, and would explicitly say so this is a bit of a double edged sword, because it means i sometimes have to state the obvious "yes i do want that", but in hindsight a lot of those cases are indeed a bit ambiguous and i may not have preferred the model to just move forward without my confirmation 4. it's a bit slower and costs more than 4.5, visibly turns are taking a bit longer and my quota is draining at a visibly faster pace. i haven't quantified exactly where this is coming from yet so overall, i think it's showing some clearly different traits, and i mostly like the changes. i'm going to keep it as my primary firstmate and observe more if you've been using it, what qualitative insights have you gathered from real usage so far?
174
184
2,093
511,614
A story in two parts:
10
166
17,574
Cannot wait to use Gemini 3.8 Flash (74% on DeepSWE btw, higher than Fable 5) on this thing
2
61
3,742
i lied
Replying to @davis7
this is the last gemini gag i promise guys (grok 4.7 made the vid with @HyperFrames_, took like 5 mins with xhigh on fast mode)
23
3,579
Been using Grok 4.7 a bunch today and I really really like it. It's snappy, follows instructions well, and the "fast mode" is very fast But price is a huge issue... 1. I don't have good data for this yet, but it's pretty clearly not all that efficient (esp next to something like Astra) 2. The pricing really isn't the same as Grok 4.6. Over 200K tokens it doubles and with fast mode it increases even more (attached screenshot is from grok build analyzing it's own usage from today) 3. The usage limits at least in grok build are BAD. I'm on the $300/mo. supergrok heavy. I've used ~$40 and it's burned ~8% of my weekly limit. This is nothing compared to what u get from OpenAI or even Anthropic... I haven't been pushing it super hard in there yet b/c I'm on crappy airplane wifi, but I'm really worried I'm gonna rip through the limits at a brutal rate. One of the best parts of previous grok models was you basically got unlimited usage in Cursor/Grok Build on the big plan, that seems to be done. Which imo hurts this model way more than it hurts Fable b/c Fable (or Astra) is a true "premium token". It's so good I'll deal with the harder limits, if I'm getting limited on Grok I'll just go use Sol or Opus (well not opus but u get the idea) since this "tier" of model has tons of competition It sucks b/c I like the model a lot, and I really like the surfaces it's used in. Grok Build is the best CLI, Cursor's cloud + project system's really nice, and Grok Bot is a huge part of my main workflow at this point. They really need some of that OpenAI efficiency magic to stay competitive at this tier, or start making much bigger more powerful models
Looking like Grok 4.7 is not quite as cheap as before. Comes out more expensive than Astra on Artificial Analysis 🙃
32
12
345
33,864