Software Engineer at Autohive/Raygun. Building AI agent infrastructure and developer tools. Writing at proredcat.dev

Wellington, NZ
Really hate this behaviour, its like Codex will explicitly not compact so it can finish its turn meaning your next message will auto-compact
2
122
Having created both MCP servers and clients it is such a pain dealing with all the non-compliance with both the MCP spec and the OAuth spec itself This thread is a really good read, highly recommended
Replying to @jescalan
But as I shepherded Clerk's MCP implementation across the line, the disillusionment began. OAuth has been around forever, and is a nearly infinite web of hundreds of different specs, all of which any given consumer or provider may or may not have implemented. There's no ubiquitous way for a provider to know which capacities a consumer wants (though there's an optional spec for that), and same in reverse (also an optional spec for that). There are enterprise grade security features that would probably be valuable for everyone but require alignment between the consumer and provider, so they never really get prioritized (outside of Okta). Also, the specs change all the time. Especially the MCP spec - if either the consumer or provider side is expecting a different version, too bad. Not to mention that the specs a big and complex and extremely easy to mess up. Even the primary players are constantly messing it up with little bugs and de-syncing from various spec versions, many of which aren't exactly compatible. While I was working on Clerk's MCP implementation, I found major MCP implementation bugs in both Cursor and Claude Code. More recently, as a consumer of one of our vendors' MCPs, I noticed I was being signed out every day, they investigated, then came back to tell me that it was not their fault, and was actually a MCP bug in codex (github.com/openai/codex/issu…). The second I finished building one OAuth/MCP capability, another one would arrive in the spec that needed to be implemented. New consumers that wanted to use our MCPs were released constantly, many of them not spec compliant, then the user complaints start rolling in, which means you need to run off-spec compatibility workarounds, or just tell the users that they are out of luck until the client they are using fixes some bug. Dynamic client registration was mandated in the initial version. It's a messy, insecure feature that no provider likes and makes apps much more prone to fraud and abuse. I was in the room when it was decided to try to replace it with something better (now out, CIMD). But DCR will still always need to be supported for compatibility. It never ends. It never will end.
1
7
185
Reilly Oldham retweeted
"Does your Chief AI Officer have lunch with your Chief Electricity Officer?" JD (@traskjd) closing out the Aotearoa AI Summit this afternoon, on why New Zealand is in a better position than we think and why almost nobody here is doing anything about it. Four minutes worth your time. Full talk and transcript in the comments. #AotearoaAISummit
4
7
31
2,877
A very nice quality of life improvement is being deployed out right now to the @raygunio MCP server I've been wanting to add OAuth for a while so it would be even easier to set up the Raygun MCP server and now that's possible
3
1
7
1,605
We are live!
2
45
Been using the new v3.0 APM profiler from @raygunio in @autohiveai to improve the performance of our virtual filesystem for agents Having such a lightweight profiler now is huge as we can run it in production and start getting agents looping on performance improvements
2
3
8
767
Chart was created using Autohive too, having proper agent sandboxes has made the system a whole lot more capable which I'm really happy with how it turned out
50
Been working on making agents a lot more capable in Autohive lately, this is one large step towards making that happen Will have my own blog post around the engineering of this out soon!
Every agent in Autohive now has its own computer. Same as you have a laptop. They can open files, run code against them, dig through a messy export and build something out of what they find. Nothing to wire up first. Reilly pointed one at ten years of his raw Spotify history and got back a proper read on his own listening habits. Live now.
2
87
Just saw this thumbnail so I had to
5
93
Really liking Lambda MicroVMs, but I wish they supported nftables for transparent outbound proxying. Getting Operation not supported, even with OS capabilities set to ALL @awscloud
1
2
192
I really want to love GitHub's stacked PRs but I just can't get them working. Tried twice now with two different streams of work and they just get stuck in merging forever without any info on what is blocking it
1
46
Trying out Dia again for a bit, hopefully it sticks this time as there is a lot of nice design changes compared to Arc but there's also a lot I still love that Arc has and Dia doesn't
31
Spaces are in @diabrowser now!! This has been one of the bigger things I've been wanting before I moved off Arc, multiple windows for profiles was a hassle to manage, would love to see icons/emojis as well! Only think left now is the Command Bar without going to a new page ;)
1
1
138
Really liking the push from labs on tighter pricing, Grok 4.5 with cheaper cache reads and GPT 5.6 Luna/Terra with overall cheaper pricing too Also really looking forward to Grok 4.6 next week 👀
1
58
Reilly Oldham retweeted
We've launched the desktop app for autohive.com! We've been using this internally for a while, and many of the team have crossed over. Don't worry, all agents still run in the cloud, so everything runs if your laptop is on or not.
1
4
26
637
One of my favourite projects to have worked on at @raygunio, been really cool to see the usage grow month on month. I still remember shipping within a day of the MCP specs release! I use it daily with my coding agents and it really helps speed up investigations
MCP usage of @raygunio increased 43% on last month. It's been rising every single month since we launched it the day after Anthropic announced the spec in 2024. Dev teams are flying by combing AI Coding Agents + Raygun observability.
33
Just found out that Buzz resends instructions on how to use Buzz and the conversation history on every user message, that seems really wasteful for context and caching
1
165
Running into some weirdness with Kimi K3 on @FireworksAI_HQ when using the Responses API. It will sometimes start outputting Python repr looking objects as the visible text and not actually do the tool calls it's saying it will do causing premature stopping of agents This doesn't appear to happen on previous Kimi models and it doesn't happen on the Chat Completions API. It only appears to happen on the Responses API Example below
1
1
2
275
It's interesting to see so many companies building model routers recently. One thing I've noticed while building one myself is that, while the Responses API gives us a mostly common format, there still isn't a common way to continue an inference. Each provider returns its own opaque state that needs to be passed back with later requests. Sometimes it can be dropped, usually with worse responses. In Gemini's case, omitting it or passing it back incorrectly can cause the request to be rejected. So even through a unified router, the caller still needs to retain and return provider-specific state outside the standard API.
1
2
2
158