<?xml version="1.0" encoding="UTF-8"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-05T21:32:58Z</responseDate>
  <request identifier="28010" metadataPrefix="oai_dc" verb="GetRecord">https://drops.dagstuhl.de/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:drops-oai.dagstuhl.de:28010</identifier>
        <datestamp>2026-10-05T06:44:04Z</datestamp>
        <setSpec>ddc:004</setSpec>
        <setSpec>open_access</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>Green AI: Cost of LLM-Based Code Completion</dc:title>
          <dc:creator>Alizadeh, Negar</dc:creator>
          <dc:creator>Saurabh, Nishant</dc:creator>
          <dc:creator>Castor, Fernando</dc:creator>
          <dc:subject>Green AI</dc:subject>
          <dc:subject>Green LLMs</dc:subject>
          <dc:subject>Code Completion</dc:subject>
          <dc:subject>Coding Assistant</dc:subject>
          <dc:subject>Energy Efficiency</dc:subject>
          <dc:subject>Trade-Offs</dc:subject>
          <dc:subject>Software Development</dc:subject>
          <dc:subject>Model Quantization</dc:subject>
          <dc:description>Background. Code completion is one of the most widely used applications of large language models (LLMs) in software development. In addition to proprietary coding assistants, powerful open-weight LLMs are increasingly adopted for locally deployed code completion systems, partly motivated by privacy concerns. Despite advances in LLM accuracy, the energy cost of inference in code completion tasks remains underexplored, particularly under large-context workloads and across programming languages. &#13;
&#13;
Aims. This study investigates the trade-off between accuracy and energy consumption in LLM-based code completion and analyzes how workload characteristics, context size, and model scale influence inference energy usage.&#13;
&#13;
Method. We evaluate 25 open-weight LLMs on two complementary code completion workloads: repository-level next-line completion (left-to-right prediction) with varying context sizes on the RepoBench dataset, and fill-in-the-middle (FIM) code completion across Python, Java, and Rust on the McEval dataset. We further analyze the influence of input tokens, output tokens, model size, and their interactions on energy consumption using correlation analysis and cluster-robust linear regression models.&#13;
&#13;
Results. Our findings show that the dominant drivers of energy consumption depend strongly on the structure of the completion task. In RepoBench, energy consumption is primarily influenced by input context size and its interaction with model scale, whereas in McEval, output generation and its interaction with active parameter count become the dominant factors. We further observe that output generation is substantially more energy-intensive per token than prompt processing. Across both benchmarks, smaller and heavily quantized models frequently achieve Pareto-optimal trade-offs, often providing accuracy comparable to larger FP16 models while consuming substantially less energy.&#13;
&#13;
Conclusions. The energy behavior of LLM inference in software engineering tasks depends strongly on workload structure, context length, and model scale. Our results suggest that increasing model size or context length does not necessarily lead to proportionally better completion quality, while quantization can substantially improve energy efficiency with limited accuracy degradation. These findings contribute toward more energy-aware deployment strategies for sustainable AI-assisted software development.</dc:description>
          <dc:publisher>Schloss Dagstuhl – Leibniz-Zentrum für Informatik</dc:publisher>
          <dc:contributor>Negar Alizadeh and Nishant Saurabh and Fernando Castor</dc:contributor>
          <dc:date>2026</dc:date>
          <dc:relation>Is Part Of LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)</dc:relation>
          <dc:type>InProceedings</dc:type>
          <dc:type>Text</dc:type>
          <dc:type>doc-type:ResearchArticle</dc:type>
          <dc:type>publishedVersion</dc:type>
          <dc:format>application/pdf</dc:format>
          <dc:identifier>doi:10.4230/LIPIcs.ESEM.2026.42</dc:identifier>
          <dc:identifier>urn:nbn:de:0030-drops-280104</dc:identifier>
          <dc:identifier>https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.42</dc:identifier>
          <dc:language>eng</dc:language>
          <dc:rights>https://creativecommons.org/licenses/by/4.0/legalcode</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
