Armin Ronacher found Opus 4.8 and Sonnet 5 inventing phantom keys in tool calls that older Claude models never got wrong.
The setup: Pi's edit tool takes a nested edits[] array. The new models produce the correct oldText/newText payload, byte-for-byte, then append nonsense keys after it. requireUnique. oldText2. matchCase. type. id. A whole zoo. The harness rejects the call, the model retries, sometimes the same phantom key, sometimes a fresh one.
This is not a small-model problem. Opus 4.8 fails around 20% of the time in the transcript that reproduced it. Older Claude models did not do this. The SOTA models are worse at this specific schema than their predecessors.
Armin's hypothesis, and it's a strong one: Claude Code silently repairs malformed tool calls. It has aliases for parameter names, type coercions, Unicode escape repair, and it filters unknown keys without telling the model. If RL happens inside that forgiving harness, slightly malformed calls still complete the task and receive reward. There is no gradient against inventing a stray field.
The better-trained the model, the stronger its prior for Claude Code's flat edit schema. A different harness with a different shape is now off-distribution, and the model fights you harder because its prior is stronger, not weaker.
Two things temper this. Single-author finding, one user's transcript, ~20% failure rate. Strict tool invocation mode eliminates it in his runs, which suggests Anthropic can mask out invalid keys server-side when asked. Codex models did not show the same regression.
The direction is the part I'd watch. If post-training keeps tightening around one dominant, closed harness, every alternative harness inherits its quirks whether it wants to or not. Tool schemas are not neutral contracts anymore. Some are close to what the model saw during RL, and some are far away, and the distance now costs you.
I run my own agent stack on local hardware. This is the kind of regression I would not catch until it bit me in a long session, because the failures are context-dependent and the edit content itself is correct. Good to know before you blame your harness for a model that is drifting toward someone else's tool shape.