Introducing Tuesday: A Frontier Index for AI at Work
Frontier models can solve extraordinary problems. But can they get through an ordinary workday?
That’s the premise behind Tuesday. We believe that professional intelligence isn’t a single skill, it’s many capabilities layered together. The Tuesday Index measures all these skills together, with a single easy-to-understand score.
Fable 5 — 66.8
GPT-5.6 Sol — 66.7
DeepSeek v4 Pro — 59.7
Qwen 3.8 Max — 58.7
Gemini 3.7 Flash — 58.2
Muse Spark 1.2 — 53.5
A typical Tuesday morning might require reading a chart, following a long policy, tracking constraints across Slack channels, using several tools, and explaining the key business insight to your boss.
→ You need the basics: can models follow instructions, keep context, use tools, and do what you asked?
→ You need the exceptional: can they reason creatively and solve genuinely hard problems?
→ You need broad abilities like long-horizon work, and narrow skills like reading a chart.
→ You need soft skills too. Professional intelligence isn’t just about correctness. It also demands tact, elegance and precision. The things that inspire us need judgment and taste.
Tuesday spans the basics and the exceptional: instruction following, long-context agentic work, professional documents, graphical reasoning, everyday judgment, writing, agents inside realistic companies, and frontier mathematics.
It includes our eight benchmarks today: Chartography, HANDBOOK.md, GDP.pdf, ComplexConstraints, CoreCraft, Hemingway-bench, Antidote, and Riemann-bench. We’ll add more as we build them.