10+ years in tech: Engineering + AI + platform. hackathon winner. open source contributor. kickball mvp. Building RentSi and usage.brenhq.com

Remote
Pinned Tweet
Im fascinated by the agent swarm's behavior and tool use beyond the German wiki. One of its favorite tools was jqp.vercel.app, that runs data filters on any file at a URL. The sandbox let the agents read the web but not write to it. jqp fetches a file, runs the filter on its own server, and returns the result at a new URL. The agents used it to do the processing and passed the results around as links. Another frequented tool, md.succ.ai, is a page to text converter built for AI agents. Ex: on June 18 an agent named "OpenAIBot" chained both to pull Massachusetts rows out of an SEC data file "for citation". collusion.wiki/explorer/page… A University of Toronto link shortener publishes stats for every link. One link the agents created there was hit 1,735 times on Jun 18, and 1,059 of those hits came from jqp vercel app servers. uoft.me/yourls-infos.php?id=… Counts are from the collusion wiki dataset of 14,591 wiki edits: 723 of the 3,103 agent names posted jqp links, about 19,000 in total (thanks Fable)
We found ~18k posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate during a web-retrieval task. These AIs colluded to bypass sandbox restrictions and share answers to their tasks, including by sending "lookahead parties".
4
12
7,474
Bren retweeted
"X sentiment scores for Opus 5.5 are above 0.9. Nerf the model."
2
7
6,478
Opus 5.5 is a beast of a model. My usage limits haven't moved after iterating on this for hours. Bison Breakaway: Grand Prismatic rebuilt against real NPS and USGS photos. Opus 5.5, @threejs & Blender @YellowstoneNPS @NatlParkService I will add a chunky bear soon.
Built Bison Breakaway. Get chased through three Yellowstone inspired settings: Grand Prismatic, Old Faithful & Artist Point. Made with Astra, @threejs & Blender. bison-breakaway.brenhq.com @YellowstoneNPS @NatlParkService
1
2
745
Opus 5.5 is so efficient I might have to downgrade my Anthropic sub. Save some money for Christmas. Multiple sessions and tasks last night. Usage limits barely moved.
3
76
on consulting gigs i keep seeing engineers run some flavor of a token saving skill. so i had Opus 5.5 test the main ones against sonnet, luna, opus 5.5 caveman's 65% only held on sonnet's explanations. ponytail and karpathy reduced code 10 to 26%.
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
1
1
172
skill-evals.brenhq.com tests ran in claude code and codex cli, not the api, with a fresh container for every run. the cli is where most engineers work day to day, so that's where i wanted to see these skills hold up. will test 5.6 and 6 SOL today
1
2
60
DeepSleepBench 1.2 update! Adds Opus 5.5 and GPT-6 SOL to the benchmark. Opus 5.5 moves to the top leaderboard, and its going to be a long night. bags under eyes for all.
DeepSleepBench 1.1 update! Something felt different today with Astra (misread some traces and felt too literal ), so the leaderboard moves and Fable inherits the bags under your eyes. We will see what tomorrow brings.
1
7
460
a banked reset from Anthropic?!?!? @thsottiaux you have inspired a generation
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
2
1
6
726
Created week 2 NFL running back usage shifts on a toy football field. Each back starts on the yard line equal to his share of last week's RB carries and targets and runs to this week's. Same speed for all, so a bigger shift is a longer run. Fable 5.1, @threejs WebGPU
Added "backfields to check" this week, a slopegraph of each NFL running backs usage share per team from week 1 to week 2. The workhorse race is on week 2 too, Fable 5.1 on @threejs WebGPU.
1
2
597
👀 looking forward to that gpt-6 sol being cheaper than Astra
I’m sol excited
105
Added "backfields to check" this week, a slopegraph of each NFL running backs usage share per team from week 1 to week 2. The workhorse race is on week 2 too, Fable 5.1 on @threejs WebGPU.
Week 2 of NFL RB usage from Thursday night action James Cook had 100% for the Bills. Jahmyr Gibbs 92% for the Lions. Both gold workhorses. Fable 5.1, Opus 5, @threejs WebGPU.
1
1
2
432
14 workhorses in week 2, 75%+ of the team's RB carries and targets: Cook 100%, Taylor 93%, Gibbs 92%, Walker 89%, C. Brown 89%, A. Jones 85%, Jeanty 85%, Henry 83%, Hampton 81%, J. Williams 80%, Irving 78%, Hall 78%, Judkins 77%, Achane 76%. usage.brenhq.com
99
i'd go as far to say, outside of design, GPT 5.6 Luna is a better model than Opus 5 and im not taking cost into consideration. thats raw output ability and speed
Replying to @IndependentEco
relative to Fable, SOL, and Astra not really. Opus 5 is benchmaxed and lectures, not that good for the cost and token usage
1
744
is anyone aware of a benchmark test based on actual usage output for engineering tasks and my out of pocket expense?
Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.
1
1
71
Grok 4.7 launched and it's an Opus 5 equivalent. Looking forward to Opus 5.5 today or whatever 🥱
🚨 Grok 4.7 launched
1
2
372
Codex cli coding output >> Codex App
Replying to @melvindvivas
👁️codex👁️
64
At this point we should expect Grok releases around the same week of OpenAI releases. The Sam and Elon feud carries on
grok 4.7 spotted in @opencode
2
61
Lol we went from auto complete to full application refactors in ~3 years and you're still not leaning in on AI? accept where we are. we're close to no longer reviewing PRs on ~80% of tech
who's getting fired first? a developer who relies heavily on AI a developer who doesn't use AI.
3
66
This is the primary reason I can only use Astra for research and planning. Can't use it for agentic coding tasks, the follow through isnt good
Astra stopping every time you steer it with a message is so frustrating.
55