Measuring what it means to reason in diagrams
Most LLM benchmarks score whether an answer is correct. Very few ask whether a model can represent that answer visually, and faithfully. That gap is where our research lives.
The first benchmark for mathematical diagram generation
We created the first benchmark purpose-built to evaluate how well large language models generate mathematical diagrams — going beyond answer correctness to assess whether a model's diagram is geometrically and semantically faithful to the problem it's solving.
Math-Vision Diagrams: A Comprehensive Benchmark for Evaluating LLM Mathematical Diagram Generation Capabilities
Harish Kashyap, Kiran Byadarhaly, Sriram Chakaravarthy, Sanyukta Tuti, Aryan Mistry
The generation of mathematically precise diagrams from textual prompts is a critical, underexplored capability of large language models. This paper introduces the first benchmark purpose-built to evaluate LLMs on mathematical diagram generation, combining an ensemble of LLMs with subject-matter-expert curation, and proposes evaluation metrics spanning both text-to-code and text-to-image generation paradigms.
Want to see how we evaluate diagram generation in practice?
Talk to us