<?xml version="1.0" encoding="UTF-8"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-05T21:44:54Z</responseDate>
  <request identifier="27979" metadataPrefix="oai_dc" verb="GetRecord">https://drops.dagstuhl.de/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:drops-oai.dagstuhl.de:27979</identifier>
        <datestamp>2026-10-05T06:44:02Z</datestamp>
        <setSpec>ddc:004</setSpec>
        <setSpec>open_access</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>Don't Bother to Use a Second LLM and Write Tests Yourself! A Study on Elixir</dc:title>
          <dc:creator>Leu, Noah</dc:creator>
          <dc:creator>Oertel, Julian</dc:creator>
          <dc:creator>Hebig, Regina</dc:creator>
          <dc:subject>LLMs</dc:subject>
          <dc:subject>Code Generation</dc:subject>
          <dc:subject>Test Generation</dc:subject>
          <dc:subject>Rare Languages</dc:subject>
          <dc:description>Background. Generating software tests and implementations with Large Language Models (LLMs) is becoming more common. However, using the same LLM for the generation of implementation and tests might make both suffer from the same biases, potentially decreasing the tests' ability to detect faults compared to tests generated by a different LLM. &#13;
&#13;
Aims. In this paper, we want to investigate this notion by combining different LLMs for generating implementation and tests. &#13;
&#13;
Method. We generate Elixir code and test cases with 5 LLMs and compare them to a baseline of manually written code and tests. We wrote and generated 459 distinct implementations and 2709 test suites, manually analyzing 9417 test case failures. &#13;
Results. We find no difference between tests generated by the same or a different LLM as the code. However, of the 247 implementations flagged as faulty, 99 were only flagged by manually written tests. &#13;
&#13;
Conclusions. Our results indicate that LLM-generated tests alone are not yet sufficient.</dc:description>
          <dc:publisher>Schloss Dagstuhl – Leibniz-Zentrum für Informatik</dc:publisher>
          <dc:contributor>Noah Leu and Julian Oertel and Regina Hebig</dc:contributor>
          <dc:date>2026</dc:date>
          <dc:relation>Is Part Of LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)</dc:relation>
          <dc:type>InProceedings</dc:type>
          <dc:type>Text</dc:type>
          <dc:type>doc-type:ResearchArticle</dc:type>
          <dc:type>publishedVersion</dc:type>
          <dc:format>application/pdf</dc:format>
          <dc:identifier>doi:10.4230/LIPIcs.ESEM.2026.11</dc:identifier>
          <dc:identifier>urn:nbn:de:0030-drops-279793</dc:identifier>
          <dc:identifier>https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.11</dc:identifier>
          <dc:language>eng</dc:language>
          <dc:rights>https://creativecommons.org/licenses/by/4.0/legalcode</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
