Most agent benchmarks test a single coding step. I ran a different one: full compound engineering brainstorm → plan → implement → review → PR on real tasks already merged into the project. Result: GPT-5.6 Sol (xhigh) beat my previous best pair — Claude Opus 4.8 for planning + GPT-5.5 Codex for execution. Grok 4.5 landed 3rd as a single planner/executor/reviewer. Strongest open-source: GLM 5.2.
2
13
6,032
Screenote was a small idea that hit me while skiing in Alps. I still use it daily. Ah and it's free and opensource and works with any harness.
1
36
I haven't updated landing for a while, just trust me, it works
3
Ivan Kuznetsov retweeted
The Pareto frontier is misleading ❌ for general users. Price it by subscription instead of API. With cache hits, Opus 5.5 with Claude Code 20x has already killed every other model.
61
80
1,111
91,807
Ivan Kuznetsov retweeted
Why do usage limits last longer on Opus 5.5? ofc, token price –20%, cache-read –60%, +usage limits. But I also looked at how the model’s behavior has changed. I work on code evals and rl envs and benchmark new models and harnesses on real tasks from my work. I compared Opus 5 → 5.5 in Claude Code and pi-agent (and omp for reference), both High, with 30 runs per setup: > far fewer tool calls, turns, and output tokens (see the attached chart) > 5.5 often writes files and runs tests in the same bash call (not step-by-step like opus 5), using vars to avoid repeating file paths. I described the examples with differences in the replies 👇 P.S. I want to compare reasoning levels next. If you have recommendations or ideas pls write.
16
6
37
1,967
Tesla is the only car where you can vibecode during driving. That's a checkmate.
2
8
72
People writing here every day about hacker-agent-swarms and my OpenClaw can't register on website and save the credentials.
3
5
99
'remaining limitation is this session’s requirement for protected password entry' -- i'm so fucking tired of this shit.
4
61
As I understand it, I should now post tits here now with a caption saying I’m an AI builder in Web3 to get attention. But instead, I’ve released a new version of Hive, and I find this image very very sexy.
1
64
I also removed a lot of sloppy code and useless backward compatibility left behind by Sol.
1
4
I’ve been using DeepSeek v4.1 Flash as my daily driver for a couple of days, and, gosh, it’s awesome. It’s fast, the limits in OpenCode Go are enormous, and you feel like you’re flying. Astra can’t find many things to fix during review. Oh, and its answers are concise. I haven’t got any hallucinations either.
1
1
70
Sorry, what’s that big new neural net that plays Doom called, Jev?
31
Dear @Tesla, collecting road closures from government websites around the world and adding them to your database would be a weekend-long task for Grok. Why don’t you do it?
1
1
40
Even without it, a short Grok search along the route set in the navigation system, asking, “Are there any road closures along this route right now?” could be a good temporary solution.
1
1
29
Do you need helping hand?
7