PolicyBench now scores 20 AI models on how accurately they compute US taxes and benefits — adding Claude Fable 5, Claude Sonnet 5, and five open-weight models: DeepSeek, Qwen, GLM, MiniMax, and Kimi.
GPT-5.5 leads, Fable ranks #2, and DeepSeek v4-pro tops the open-weight models.