Introducing Da7em Bench.
An independent benchmark for AI models, built on real client work. The first of its kind in the world.
How it works:
Every model runs about 200 real tasks in each of 12 areas: reasoning, research, planning, delivery, persistence, accuracy, honesty, acceptance, engineering, taste, writing, and communication.
Each model is tested across several harnesses, both official and neutral ones (Droid, Hermes Agent, Devin, Cursor), so no single harness decides a model's fate and the results reflect the model.
Scoring is 1 to 5. A 5 means the work was accepted as delivered. Middle scores mean it needed revision. A 1 means it failed.
The bar is professional work. Every result is judged against what a paid professional would have delivered for the same brief.
Tasks stay private so they can't leak into training data and inflate future scores.
The goal isn't one more leaderboard. It's helping you pick the right model for your kind of work. A model that leads in reasoning can still fall behind in writing or design taste, and the radar charts show exactly where.
This is v0.1. It will keep evolving with harder, market-relevant tasks and with new models as they ship. A few popular models (Opus 5, GPT Luna) aren't included yet because I haven't run enough tasks on them to score them fairly.
Da7em Bench is fully independent. No sponsors, no vendor relationships. I've paid for every run out of pocket, thousands of dollars so far. That's the whole point: an honest, neutral look at what these models actually do on real work.
Full scores and the framework are in the images below.