Diagram AI
Research

Measuring what it means to reason in diagrams

Most LLM benchmarks score whether an answer is correct. Very few ask whether a model can represent that answer visually, and faithfully. That gap is where our research lives.

Benchmark

The first benchmark for mathematical diagram generation

We created the first benchmark purpose-built to evaluate how well large language models generate mathematical diagrams — going beyond answer correctness to assess whether a model's diagram is geometrically and semantically faithful to the problem it's solving.

Diagram fidelity score1.0
0.0Per-diagram evaluation grid
Publication
Submitted — EACL

A metric for evaluating math diagrams — and training against it

Our submission proposes a metric for evaluating the quality of generated mathematical diagrams, and shows how that metric can be applied through reinforcement learning to fine-tune DiagramLLM 1.0's diagram-generation capability directly against it.

Diagram metric
Reinforcement learning
Fine-tuned DiagramLLM 1.0

Want to see how we evaluate diagram generation in practice?

Talk to us