<?xml version="1.0" encoding="UTF-8"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-07-28T23:36:30Z</responseDate>
  <request identifier="26700" metadataPrefix="oai_dc" verb="GetRecord">https://drops.dagstuhl.de/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:drops-oai.dagstuhl.de:26700</identifier>
        <datestamp>2026-07-28T09:32:41Z</datestamp>
        <setSpec>ddc:004</setSpec>
        <setSpec>open_access</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>Evaluating the Prosodic Diversity of TTS Models for L2 Prosody Assessment</dc:title>
          <dc:creator>Julião, Mariana</dc:creator>
          <dc:creator>Abad, Alberto</dc:creator>
          <dc:creator>Moniz, Helena</dc:creator>
          <dc:subject>Speech synthesis</dc:subject>
          <dc:subject>TTS</dc:subject>
          <dc:subject>prosody</dc:subject>
          <dc:subject>L1</dc:subject>
          <dc:subject>L2</dc:subject>
          <dc:description>Recent advances in text-to-speech (TTS) have led to synthetic speech that is often indistinguishable from natural speech at the level of individual utterances. However, it remains unclear whether such systems reproduce the prosodic variability observed in natural speech in a large native-speaker population. This question is particularly relevant for applications in computer-assisted language learning (CALL), as variability is a central property of prosody and a prerequisite for robust assessment.&#13;
In this work, we investigate whether TTS can approximate the distribution of prosodic patterns found in native speech, and whether it can serve as a reference for evaluating second language (L2) productions. To this end, we first compare different TTS models to native speakers, and then compare L2 speakers to synthetic speech by the model previously seen as the closest to native speakers.&#13;
Our results show that regardless of TTS achieving high perceptual quality, its prosodic variability substantially differs from that of native speakers. As a consequence, comparisons between L2 speech and TTS-based references reveal both the potential and the limitations of using synthetic speech for prosody assessment.</dc:description>
          <dc:publisher>Schloss Dagstuhl – Leibniz-Zentrum für Informatik</dc:publisher>
          <dc:contributor>Mariana Julião and Alberto Abad and Helena Moniz</dc:contributor>
          <dc:date>2026</dc:date>
          <dc:relation>Is Part Of OASIcs, Volume 144, 15th Symposium on Languages, Applications and Technologies (SLATE 2026)</dc:relation>
          <dc:type>InProceedings</dc:type>
          <dc:type>Text</dc:type>
          <dc:type>doc-type:ResearchArticle</dc:type>
          <dc:type>publishedVersion</dc:type>
          <dc:format>application/pdf</dc:format>
          <dc:identifier>doi:10.4230/OASIcs.SLATE.2026.2</dc:identifier>
          <dc:identifier>urn:nbn:de:0030-drops-267007</dc:identifier>
          <dc:identifier>https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026.2</dc:identifier>
          <dc:language>eng</dc:language>
          <dc:rights>https://creativecommons.org/licenses/by/4.0/legalcode</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
