Diagram AI
Back
Research

Measuring what it means to reason in diagrams

Most LLM benchmarks score whether an answer is correct. Very few ask whether a model can represent that answer visually, and faithfully. That gap is where our research lives.

Benchmark

The first benchmark for mathematical diagram generation

We created the first benchmark purpose-built to evaluate how well large language models generate mathematical diagrams — going beyond answer correctness to assess whether a model's diagram is geometrically and semantically faithful to the problem it's solving.

Diagram fidelity score1.0
0.0Per-diagram evaluation grid

Math-Vision Diagrams: A Comprehensive Benchmark for Evaluating LLM Mathematical Diagram Generation Capabilities

Harish Kashyap, Kiran Byadarhaly, Sriram Chakaravarthy, Sanyukta Tuti, Aryan Mistry

The generation of mathematically precise diagrams from textual prompts is a critical, underexplored capability of large language models. This paper introduces the first benchmark purpose-built to evaluate LLMs on mathematical diagram generation, combining an ensemble of LLMs with subject-matter-expert curation, and proposes evaluation metrics spanning both text-to-code and text-to-image generation paradigms.

Want to see how we evaluate diagram generation in practice?

Talk to us