<?xml version="1.0" encoding="UTF-8"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-09-10T09:44:57Z</responseDate>
  <request identifier="27561" metadataPrefix="oai_dc" verb="GetRecord">https://drops.dagstuhl.de/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:drops-oai.dagstuhl.de:27561</identifier>
        <datestamp>2026-09-10T05:38:43Z</datestamp>
        <setSpec>ddc:004</setSpec>
        <setSpec>open_access</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>Limits of Spatial Reasoning Using LLMs (Short Paper)</dc:title>
          <dc:creator>Ithivatana, Varis</dc:creator>
          <dc:creator>Bennett, Brandon</dc:creator>
          <dc:subject>Spatial Reasoning</dc:subject>
          <dc:subject>Large Language Models</dc:subject>
          <dc:subject>Frame of Reference</dc:subject>
          <dc:subject>Path Integration</dc:subject>
          <dc:subject>Planning</dc:subject>
          <dc:description>This paper investigates the limits of spatial reasoning of LLMs in a grid-based warehouse environment, with ground-truth answers computed by simulation. We test two kinds of task: executing given movement instructions, and generating shortest action sequences for planning problems. In the execution tasks, an agent must follow egocentric and allocentric actions and report its final absolute direction relative to a target object. In the planning tasks, the model must find a shortest route, sometimes with object carrying, allocentric descriptions, or ordered pallet-pushing goals. We evaluate 100 instances per experiment using GPT-5-mini, and compare a smaller planning set with ChatGPT-5.2 (extended thinking), Gemini-Pro-3.1, and Grok-Expert. The results show high accuracy on short and moderately long instruction-following tasks, but reduced accuracy on long trajectories with a much sharper drop on planning tasks. Static landmark labels in the form of coloured walls improve robustness, compared with problems described in terms of cardinal directions, while planning with ordered object-pushing remains difficult even for stronger models.</dc:description>
          <dc:publisher>Schloss Dagstuhl – Leibniz-Zentrum für Informatik</dc:publisher>
          <dc:contributor>Varis Ithivatana and Brandon Bennett</dc:contributor>
          <dc:date>2026</dc:date>
          <dc:relation>Is Part Of LIPIcs, Volume 393, 17th International Conference on Spatial Information Theory (COSIT 2026)</dc:relation>
          <dc:type>InProceedings</dc:type>
          <dc:type>Text</dc:type>
          <dc:type>doc-type:ResearchArticle</dc:type>
          <dc:type>publishedVersion</dc:type>
          <dc:format>application/pdf</dc:format>
          <dc:identifier>doi:10.4230/LIPIcs.COSIT.2026.17</dc:identifier>
          <dc:identifier>urn:nbn:de:0030-drops-275616</dc:identifier>
          <dc:identifier>https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.COSIT.2026.17</dc:identifier>
          <dc:language>eng</dc:language>
          <dc:rights>https://creativecommons.org/licenses/by/4.0/legalcode</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
