For the benchmark, agents completed 42 wallet tasks spanning core CLI flows and open-ended, natural language actions like "supply to Aave" or "trade on Kumbaya."
Success was measured against session-key safety, permission scope, and efficiency in addition to completion.