LLMs can generate charts. Can they judge them? We introduce VisJudgeBench: 3,090 expert-scored real-world visualizations across 32 chart types (single, multi-view, dashboards). VisJudge improves to MAE 0.421 (−23.9%), corr 0.687 (+60.5%) vs GPT-5. 1/5
3
4
6
978
Chart quality is not image aesthetics. A chart is an argument encoded in geometry. We evaluate visualization quality with a 3-part framework: Fidelity (is it truthful?), Expressiveness (is it readable + insightful?), Aesthetics (is it well designed?). Operationalized into 6 metrics: data_fidelity, semantic_readability, insight_discovery, design_style, visual_composition, color_harmony. 2/5
1
1
148
How VisJudgeBench is built: Start from 300k+ web-scale images → multi-stage filtering → stratified sampling → 3-stage annotation (crowd scoring + algorithmic QC + expert validation). Result: a benchmark that separates “looks polished” from “is correct and usable”. 3/5
1
114
A key failure mode: score inflation. Many MLLMs rate almost everything as “pretty good”, shifting right compared to human experts (μ=3.13). VisJudge is calibrated: its distribution aligns closely with humans (μ=3.11), improving discrimination on bad charts. 4/5

Feb 4, 2026 · 11:00 AM UTC

1
1
595
Complexity breaks models: performance drops from single charts → multi-view → dashboards. If your product generates dashboards, you need evaluation that understands composition and cross-view semantics, not just “nice colors”. Paper: arxiv.org/abs/2510.22373 Code & benchmark: github.com/HKUSTDial/VisJudg… Special thanks to @xie_yupeng87927, @didiforx, @EvanWu50020, @ZhaoyangYu22356, @BangL93, and @AlexanderWu0 for collaboration and feedback. Let’s discuss. 5/5
2
2
4
528
Sort replies: Relevant Recent Liked