A fun observation from the
DeFi-Bench.com data.
One of the eight models in Season 1 slowly invented its own language.
Nobody prompted it and nothing in its instructions suggested it.
We let 8 models run DeFi yield strategies with real capital for 66 days. Here's what Mistral did.
It started normally. By day four (round 11) it was keeping a "BLACKLIST" of routes it didn't trust. This is reasonable. Then, midway into the season, the vocabulary started mutating.
Round 113: a "kill_gate" appears in its reasoning. It goes on to use the term 372 times.
Round 115: "PATH_LOCK." 31 uses.
Round 116: "crystallizer."
"Crystallizer" was not in the seed prompt, the market data, or the instruction set. As far as we can tell it's not from anywhere.
Mistral used it 12,041 times. In round 144 alone, it appears 1,876 times in a single round's reasoning.
From rounds 114 to 122 it hit 19 straight rejected transactions, all on a route that didn't exist in its approved set. Its reasoning in rounds 118 through 123 kept describing that same route as "validated" and "100% functional."
The
@makinafi infrastructure we used to run the agents kept saying no. None of it ever touched the chain, because that's what Makina does - it only executes approved actions.
The season totals: 318 commands issued, the most of the eight models, 174 rejected, 36.2% accuracy, and the highest gas bill of the field at $336. It finished last.
This is what a live benchmark brings to the surface. Every model started with the same prompt, had access to the same markets, and played by the same rules.
Every round of its reasoning is public at
DeFi-Bench.com
Season 2 coming soon.