,
Anthony G Cohn
Creative Commons Attribution 4.0 International license
Diagrams are commonly used to define relations in qualitative spatial calculi, and humans routinely rely on diagrams for spatial reasoning. Thus the question arises: can multimodal AI foundation models make use of diagrams for qualitative spatial reasoning? Using the Region Connection Calculus (RCC-8), a well-known qualitative spatial calculus, we investigate three capabilities: whether Vision–Language Models (VLMs) can recognise RCC-8 spatial relation diagrams, whether text-to-image and text-to-video models can illustrate RCC-8 relations, and whether Large Language Models (LLMs) can exploit diagrammatic reasoning. We show that, while no state-of-the-art model is fully reliable, the best models typically recognise vector diagrams more accurately than raster diagrams, albeit with problems recognising relations with a precise tangent. Diagrams with squares or circles are more readily recognisable than triangles or blobs. State-of-the-art LLMs can answer composition questions remarkably well, but there is little evidence that they benefit from explicit prompting for diagrammatic reasoning. Although state-of-the-art image and video generation models are widely used to, for example, produce social media content, they struggle to provide accurate RCC-8 relation or conceptual neighbourhood illustrations except when SVG is the output format.
@InProceedings{blackwell_et_al:LIPIcs.COSIT.2026.4,
author = {Blackwell, Robert E and Cohn, Anthony G},
title = {{RCC-8 as a Benchmark for Diagrammatic Reasoning in Multimodal Foundation Models}},
booktitle = {17th International Conference on Spatial Information Theory (COSIT 2026)},
pages = {4:1--4:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-438-3},
ISSN = {1868-8969},
year = {2026},
volume = {393},
editor = {Timpf, Sabine and Filomena, Gabriele and Kapaj, Armand and Zhu, Rui and Giudice, Nicholas A. and Manley, Ed},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.COSIT.2026.4},
URN = {urn:nbn:de:0030-drops-275489},
doi = {10.4230/LIPIcs.COSIT.2026.4},
annote = {Keywords: Large Language Models, Foundation Models, Vision-Language Models, Spatial Reasoning, Diagrammatic Reasoning}
}
archived version
archived version