<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD with OASIS Tables with MathML3 v1.1 20151215//EN" "JATS-journalpublishing-oasis-article1-mathml3.dtd">
<article article-type="research-article" dtd-version="1.1" xml:lang="en" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">
	<front>
		<journal-meta>
			<journal-id journal-id-type="publisher-id">Loquens</journal-id>
			<journal-title-group>
				<journal-title>Loquens</journal-title>
				<abbrev-journal-title abbrev-type="publisher">Loquens</abbrev-journal-title>
			</journal-title-group>
			<issn publication-format="electronic">2386-2637</issn>
			<issn-l>2386-2637</issn-l>
			<publisher>
				<publisher-name>Consejo Superior de Investigaciones Cient&#xed;ficas</publisher-name>
			</publisher>
		</journal-meta>
		<article-meta>
			<article-id pub-id-type="publisher-id">loquens.2022.e090</article-id>
			<article-id pub-id-type="doi">10.3989/loquens.2022.e090</article-id>
			<article-categories>
				<subj-group subj-group-type="heading">
					<subject>Articles</subject>
				</subj-group>
			</article-categories>
			<title-group>
				<article-title>Comparison of intensity-based methods for automatic speech rate computation</article-title>
				<trans-title-group xml:lang="es">
					<trans-title>Comparaci&#xf3;n de dos m&#xe9;todos basados en la intensidad para el c&#xe1;lculo autom&#xe1;tico de la velocidad de habla</trans-title>
				</trans-title-group>
			</title-group>
			<contrib-group>
				<contrib contrib-type="author" corresp="yes">
					<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0001-7002-9851</contrib-id>
					<name>
						<surname>Elvira-Garc&#xed;a</surname>
						<given-names>Wendy</given-names>
					</name>
					<email xlink:href="wendyelvira@ub.edu">wendyelvira@ub.edu</email>
					<aff id="aff1"><institution>Universitat de Barcelona</institution></aff>
				</contrib>
				<contrib contrib-type="author">
					<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-7160-9513</contrib-id>
					<name>
						<surname>Farr&#xfa;s</surname>
						<given-names>Mireia</given-names>
					</name>
					<email xlink:href="mfarrus@ub.edu">mfarrus@ub.edu</email>
					<aff id="aff2"><institution>Universitat de Barcelona</institution></aff>
				</contrib>
				<contrib contrib-type="author">
					<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-3310-8582</contrib-id>
					<name>
						<surname>Garrido Almi&#xf1;ana</surname>
						<given-names>Juan Mar&#xed;a</given-names>
					</name>
					<email xlink:href="jmgarrido@flog.uned.es">jmgarrido@flog.uned.es</email>
					<aff id="aff3"><institution>Universidad Nacional de Educaci&#xf3;n a Distancia (UNED)</institution></aff>
				</contrib>
			</contrib-group>
			<pub-date pub-type="epub">
				<day>01</day>
				<month>12</month>
				<year>2022</year>
			</pub-date>
			<pub-date pub-type="collection">
				<month>12</month>
				<year>2022</year>
			</pub-date>
			<volume>9</volume>
			<issue>1-2</issue>
			<elocation-id>e090</elocation-id>
			<history>
				<date date-type="received">
					<day>01</day>
					<month>05</month>
					<year>2022</year>
				</date>
				<date date-type="accepted">
					<day>05</day>
					<month>12</month>
					<year>2022</year>
				</date>
				<date date-type="pub">
					<day>09</day>
					<month>06</month>
					<year>2023</year>
				</date>
			</history>
			<permissions>
				<copyright-statement>&#xa9; 2022 CSIC</copyright-statement>
				<copyright-year>2022</copyright-year>
				<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by/4.0/">
					<license-p>This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International (CC BY 4.0) License.</license-p>
				</license>
			</permissions>
			<self-uri xlink:href="https://loquens.revistas.csic.es/index.php/loquens/article/view/XXXX/XXXX"/>
			<abstract>
				<title>Abstract</title>
				<p> Automatic computation of speech rate is a necessary task in a wide range of applications that require this prosodic feature, in which a manual transcription and time alignments are not available. Several tools have been developed to this end, but not enough research has been conducted yet to see to what extent they are scalable to other languages.</p>
				<p> In the present work, we take two off-the- shelf tools designed for automatic speech rate computation and already tested for Dutch and English (v1, which relies on intensity peaks preceded by an intensity dip to find syllable nuclei and v3, which relies on intensity peaks surrounded by dips) and we apply them to read and spontaneous Spanish speech. Then, we test which of them offers the best performance. The results obtained with precision and normalized mean squared error metrics showed that v3 performs better than v1. However, recall measurement shows a better performance of v1, which suggests that a more fine-grained analysis on sensitivity and specificity is needed to select the best option depending on the application we are dealing with.</p>
			</abstract>
			<trans-abstract xml:lang="es">
				<title>Resumen</title>
				<p> El c&#xe1;lculo autom&#xe1;tico de la velocidad de habla es una tarea fon&#xe9;tica &#xfa;til y que adem&#xe1;s se hace indispensable cuando no hay disponible una transcripci&#xf3;n manual a partir de la cual determinar una tasa de habla manual. Se han desarrollado varias herramientas para este fin, pero todav&#xed;a no se ha llevado a cabo suficiente investigaci&#xf3;n para ver hasta qu&#xe9; punto las herramientas son aplicables a lenguas distintas para las que fueron dise&#xf1;adas. En este art&#xed;culo probamos dos herramientas para el c&#xe1;lculo autom&#xe1;tico de la velocidad de habla ya evaluadas para el neerland&#xe9;s y el ingl&#xe9;s (v1, que se basa en la determinaci&#xf3;n de picos de intensidad precedidos de un valle para encontrar n&#xfa;cleos de s&#xed;laba, y v3, que se basa en picos de intensidad rodeados de valles) y las aplicamos a un corpus de habla le&#xed;da y espont&#xe1;nea del espa&#xf1;ol para analizar cu&#xe1;l ofrece mejores resultados en espa&#xf1;ol.</p>
				<p> Los resultados de precisi&#xf3;n y del error cuadr&#xe1;tico mediano normalizado obtenidos muestran que v3 funciona mejor que v1. No obstante, el <italic>recall</italic> muestra mejor rendimiento para la v1, lo que nos indica que se necesita un an&#xe1;lisis detallado de la sensibilidad y la especificidad para seleccionar la mejor opci&#xf3;n en funci&#xf3;n de los objetivos del an&#xe1;lisis posterior que se quiera hacer.</p>
			</trans-abstract>
			<kwd-group>
				<kwd>Prosody</kwd>
				<kwd>speech rate</kwd>
				<kwd>syllable count</kwd>
				<kwd>automatic assessment</kwd>
			</kwd-group>
			<kwd-group xml:lang="es">
				<kwd>Prosodia</kwd>
				<kwd>velocidad de habla</kwd>
				<kwd>evaluaci&#xf3;n autom&#xe1;tica</kwd>
			</kwd-group>
			<funding-group id="fw-01">
				<award-group id="aw1">
					<award-id>PGC2018-094233-B-C21</award-id>
				</award-group>
				<funding-statement> This work has been partially funded by the project &#x201c;M&#xe9;todos, modelos, m&#xe9;tricas y herramientas para la evaluaci&#xf3;n de la prosodia (ProA)&#x201d;, reference number PGC2018-094233-B-C21. The first author is a &#x201c;Serra H&#xfa;nter Fellow&#x201d;. The authors would like to thank Dr. Mar&#xed;a Machuca for providing the VILE corpus used in these experiments</funding-statement>
			</funding-group>
			<counts>
				<fig-count count="3"/>
				<table-count count="5"/>
				<equation-count count="7"/>
				<ref-count count="27"/>
				<page-count count="7"/>
			</counts>
		</article-meta>
	</front>
	<body>
		<sec id="sec1" sec-type="intro">
			<label>1.</label>
			<title>Introduction</title>
			<p> Automatic computation of speech rate has several applications in speech technologies, such as automatic evaluation of prosody. Several studies have explored, for example, its use for automatic evaluation of speech fluency (<xref ref-type="bibr" rid="B3">Cucchiarini, Strik, &amp; Boves, 1998</xref>, <xref ref-type="bibr" rid="B4">2000a</xref>, <xref ref-type="bibr" rid="B5">2000b</xref>, <xref ref-type="bibr" rid="B6">2002</xref>; <xref ref-type="bibr" rid="B19">Neumeyer, Franco, Digalakis, &amp; Weintraub, 2000</xref>; <xref ref-type="bibr" rid="B27">Zechner, Higgins, Xia, &amp; Williamson, 2009</xref>; <xref ref-type="bibr" rid="B15">Honig, Batliner, Weilhammer, &amp; N&#xf6;th, 2010</xref>, among others). Usual approaches to speech rate computation use a phonetic aligner to obtain the necessary phonetic segmentation. This is so because the performance of speech recognition systems, if available, is not good enough to guarantee that the obtained phonetic segmentation is reliable.</p>
			<p> Phonetic aligners appear then as an alternative to obtain a more accurate segmentation of the speech chain, but they need the orthographic transcription of the input discourse to be known. If the computation of the speech rate of unrestricted text -not previously known by the system- is attempted, there are some alternatives that do not require a full phonetic segmentation of the input speech to be available, such as the automatic detection of syllabic nuclei. With the aim of exploring this alternative, the current paper compares the performance of two different methods for the automatic computation of the number of syllabic nuclei using a similar technique based on intensity peak detection. The first one is a Praat script described in <xref ref-type="bibr" rid="B8">de Jong and Wempe (2009)</xref> and the second one is another Praat script developed by the same authors and other collaborators (<xref ref-type="bibr" rid="B7">de Jong, Pacilly, &amp; Wempe, 2021</xref>), in which a different approach to detect intensity peaks is applied. The final goal is to determine which of them would perform better in a task of syllable detection oriented to speech rate calculation and to establish if any of these two methods is adequate to be used in an automatic prosody evaluation system for Spanish.</p>
			<p> This paper is structured as follows: Section 2 briefly overviews the related work on this topic, Section 3 describes the experimental setup, Section 4 presents the assessment results, and finally, Sections 5 and 6 sketch the discussion and conclusions, respectively.</p>
		</sec>
		<sec id="sec2">
			<label>2.</label>
			<title>Related work</title>
			<p> Most of the studies that deal with automatic computation of speech rate are based on the transcriptions obtained -either manually or automatically- from speech material. They mainly differ in the units used to compute speech rate. Most of them are based on counting the number of syllables within a specific segment of speech, providing the speech rate computation as the number of syllables per second, while some other works also provide other measures. <xref ref-type="bibr" rid="B25">Verhasselt and Martens (1996)</xref>, for instance, defines speech rate as the number of phones per second and computes them over the sentences of the TIMIT corpus. <xref ref-type="bibr" rid="B23">Pfitzinger (1996)</xref> also used the number of phones per second as speech rate measure over a total of 240 sentences spoken by eight different speakers.</p>
			<p> The literature on automatic speech rate computation tools without transcriptions, which is the goal of the current paper, is scarce. One of the most relevant works in this respect is <xref ref-type="bibr" rid="B22">Pfau and Ruske (1998)</xref>, in which speech rate is computed by means of vowel detection, based on loudness in vowel regions, which tends to be higher than in consonant regions. Similarly, the method of <xref ref-type="bibr" rid="B21">Pellegrino, Farinas and Rouas (2004)</xref> is based on an unsupervised vowel detection algorithm scalable to any language. Validation was assessed on a spontaneous speech subset of the OGI Multilingual Telephone Speech Corpus. In <xref ref-type="bibr" rid="B18">Narayanan and Wang (2005)</xref> and <xref ref-type="bibr" rid="B26">Wang and Narayanan (2007)</xref>, the authors present novel methods for speech rate estimation, measured as the number of syllables per second, analyzing the segments contained between pauses in the Switchboard database (<xref ref-type="bibr" rid="B13">Godfrey &amp; Holliman, 1993</xref>). Both methods are based on an extension of signal correlation -essential for syllable detection- by including temporal correlation and prominent spectral sub-bands.</p>
			<p> The work described in <xref ref-type="bibr" rid="B10">Dekens, Demol, Verhelst and Verhoeve (2007)</xref> is also based on the number of syllables per second, and the authors evaluate the performance of several speech estimators on a multilingual database covering Dutch, English, French, Romanian and Spanish, by using sub-band and time correlation to detect the number of vowels and diphthongs.</p>
			<p> However, giving that speech rate can be computed using syllables or phones and total time of speech, any tool that identifies either syllable boundaries or vowels can be used for this task, for example, tools that syllabify conversational speech (Landsiedel et al., 2011; Mary et al., 2018) or tools that locate syllable nuclei (<xref ref-type="bibr" rid="B24">Sabu, Chaudhuri, Rao, &amp; Patil, 2021</xref>). Using this last method, <xref ref-type="bibr" rid="B9">de Jong et al. (2007)</xref> and de <xref ref-type="bibr" rid="B8">Jong &amp; Wempe (2009)</xref> compute speech rate over two corpora of spoken Dutch, by identifying peaks in intensity that are preceded by dips, which is then considered as a syllable nucleus. In <xref ref-type="bibr" rid="B24">Sabu et al. (2021)</xref>, the authors use the TIMIT dataset (<xref ref-type="bibr" rid="B12">Garofolo et al., 1993</xref>) and a children&#x2019;s oral reading corpus created ad hoc, for which they identify vowel sonority by means of local peak picking on a frequency-weighted energy contour.</p>
		</sec>
		<sec id="sec3">
			<label>3.</label>
			<title>Experimental setup</title>
			<sec id="sec3.1">
				<label>3.1.</label>
				<title>Evaluated tools for speech rate computation</title>
				<p> The present paper analyses the performance of two tools distributed under a GNU General Public License (<xref ref-type="bibr" rid="B8">de Jong &amp; Wempe, 2009</xref>; <xref ref-type="bibr" rid="B7">de Jong et al., 2021</xref>). Both of them are Praat-based scripts that use intensity in order to find syllable nuclei. More specifically, they extract an intensity object using the following parameters: &#x2018;minimum Pitch&#x2019; set to 50Hz and the autocorrelation method. After this point, their behavior differs.</p>
				<p>The first tool (v1), described in <xref ref-type="bibr" rid="B8">de Jong and Wempe (2009)</xref>, applies a predefined threshold (2dB above the median intensity of the total sound file) to find peaks preceded by a dip in intensity (see <xref ref-type="fig" rid="f1">Figure 1</xref>). Then, out of those peaks, it discards those that are unvoiced.</p>
				<fig id="f1">
					<label>Figure 1</label>
					<caption>
						<title>Intensity curve of the Spanish phrase &#x201c;La baba&#x201d; (the slime) with a syllable nucleus and its preceding and following dips highlighted.</title>
					</caption>
					<graphic id="gra-1" xlink:href="LOQUENS-9-1-2-e090-gf1.png"/>
				</fig>
				<p> The second tool (v3) relies on a different method (<xref ref-type="bibr" rid="B7">de Jong et al., 2021</xref>). It detects every intensity peak above 25 dB and below 95 % of the highest peak (in order to disregard loud bursts in the signal). Then, it measures the intensity surrounding the peak and if it is a dip of at least 2dB at both sides the peak is labelled as syllable nucleus (<xref ref-type="bibr" rid="B7">de Jong et al., 2021</xref>).</p>
				<p> Therefore, the main difference between the tools is that one (v1) considers as syllable nuclei those intensity peaks preceded by a dip, and the other (v3) considers as syllable nuclei those intensity peaks that are surrounded by intensity dips. This difference results in the same judgments most of the time, however in some cases it does not. Discrepancies between v1 and v3 are usually related to approximants, whose dip is short enough to be considered a whole with the next one, and laterals (and nasals to a lesser degree) in coda position (<xref ref-type="fig" rid="f2">Figure 2</xref>).</p>
				<fig id="f2">
					<label>Figure 2</label>
					<caption>
						<title>Waveform, spectrogram and intensity of the Spanish sentence &#x201c;Logra detener el paso del tiempo&#x201d; &#x2018;(It manages to stop time)&#x2019; depicting the vowel nuclei found by v1 (tier 4) and v3 (tier 5).</title>
					</caption>
					<graphic id="gra-2" xlink:href="LOQUENS-9-1-2-e090-gf2.png"/>
				</fig>
			</sec>
			<sec id="sec3.2">
				<label>3.2.</label>
				<title>Materials</title>
				<p> In order to test which method (preceding peak or surrounding peak) offers the best performance in Spanish, we used a subcorpus from the AHUMADA corpus (<xref ref-type="bibr" rid="B20">Ortega-Garcia, Gonzalez-Rodriguez, &amp; Marrero-Aguiar, 2000</xref>) selected for the VILE project (<xref ref-type="bibr" rid="B1">Albal&#xe1; et al., 2008</xref>; <xref ref-type="bibr" rid="B2">Battaner Moro et al., 2005</xref>), consisting of recordings of 30 male speakers, with a total of 3.5 hours of speech, recorded in three different sessions in different days (M1- M2-M3), and two different conditions: read speech (26984 vowels) and spontaneous speech (35366 vowels).</p>
				<p> The read subcorpus consists of the reading of a phonologically and syllabically balanced text of approximately one minute read at a normal speech rate. All speakers read the same text in the three sessions.</p>
				<p> The spontaneous subcorpus consists of at least one minute of speech describing a picture, explaining speakers&#x2019; last holidays, a well- known board game or simply something familiar to them.</p>
				<p> This material was manually annotated at the phoneme, syllable and word levels for the VILE project (<xref ref-type="bibr" rid="B1">Albal&#xe1; et al., 2008</xref>; <xref ref-type="bibr" rid="B2">Battaner Moro et al., 2005</xref>). The annotation procedure involved three steps: in the first one, a team of phoneticians orthographically transcribed intonational groups following the guidelines described in <xref ref-type="bibr" rid="B16">Llisterri, Machuca and R&#xed;os, (2017)</xref>; in the second one, EasyAlign (<xref ref-type="bibr" rid="B14">Goldman, 2011</xref>) was used to automatically align the annotation; finally, a human annotator revised the automatic segmentation.</p>
			</sec>
			<sec id="sec3.3">
				<label>3.3.</label>
				<title>Evaluation metrics</title>
				<p> One of the main challenges when assessing systems dealing with the automatic computation of speech rate is the diversity and sparseness of evaluation metrics. The metrics used in the literature to evaluate the speech rate estimators vary among the different works and include a wide range of metrics such as the relative prediction error, the correlation coefficient between the estimated and actual syllables, the syllable error rate, the vowel error rate, the linear regression coefficient, the mean error, the standard deviation error, and F-score, among others. Moreover, these metrics are computed either over the number of syllables (or phones) as units of measurement, or directly over the speech rate measurement.</p>
				<p> In the current paper, we present two different evaluations to compare the two tools addressed. Firstly, we show a performance analysis based on common metrics used for classification problems: accuracy, precision, recall, and F-score. For this assessment, we have considered the tier where vowel (syllable nuclei) and consonant intervals (non syllable nuclei) are labelled.</p>
				<p> Additionally, we provide the root mean square error (RMSE) and normalized root mean square error (NRMSE) for the assessment analysis, based on the syllable annotation tier and, more specifically, the number of syllables of each file in the VILE corpus.</p>
				<sec id="sec3.3.1">
					<label>3.3.1.</label>
					<title>Performance metrics</title>
					<p> For the first evaluation, we compare both tools using the standard performance metrics in classification problems:</p>
					<list list-type="bullet">
						<list-item>
							<p>Accuracy: defined as the number of cases of the correctly predicted class, that is:</p>
						</list-item>
					</list>
					<disp-formula id="e1">
						<mml:math id="mml-1">
							<mml:mi mathvariant="normal">a</mml:mi>
							<mml:mi mathvariant="normal">c</mml:mi>
							<mml:mi mathvariant="normal">c</mml:mi>
							<mml:mi mathvariant="normal">u</mml:mi>
							<mml:mi mathvariant="normal">r</mml:mi>
							<mml:mi mathvariant="normal">a</mml:mi>
							<mml:mi mathvariant="normal">c</mml:mi>
							<mml:mi mathvariant="normal">y</mml:mi>
							<mml:mo>=</mml:mo>
							<mml:mfrac>
								<mml:mrow>
									<mml:mi mathvariant="normal">T</mml:mi>
									<mml:mi mathvariant="normal">P</mml:mi>
									<mml:mo>+</mml:mo>
									<mml:mi mathvariant="normal">T</mml:mi>
									<mml:mi mathvariant="normal">N</mml:mi>
								</mml:mrow>
								<mml:mrow>
									<mml:mi mathvariant="normal">T</mml:mi>
									<mml:mi mathvariant="normal">P</mml:mi>
									<mml:mo>+</mml:mo>
									<mml:mi mathvariant="normal">T</mml:mi>
									<mml:mi mathvariant="normal">N</mml:mi>
									<mml:mo>+</mml:mo>
									<mml:mi mathvariant="normal">F</mml:mi>
									<mml:mi mathvariant="normal">P</mml:mi>
									<mml:mo>+</mml:mo>
									<mml:mi mathvariant="normal">F</mml:mi>
									<mml:mi mathvariant="normal">N</mml:mi>
								</mml:mrow>
							</mml:mfrac>
						</mml:math>
						<label>(1)</label>
					</disp-formula>
						<def-list id="d1">
							<title>where</title>
							<def-item>
								<term>
									<italic>TP</italic> =</term>
								<def>
									<p>True Positives (detected syllable nuclei)</p>
								</def>
							</def-item>
							<def-item>
								<term>
									<italic>TN</italic> =</term>
								<def>
									<p>True Negatives</p>
								</def>
							</def-item>
							<def-item>
								<term>
									<italic>FP</italic> =</term>
								<def>
									<p>False Positives</p>
								</def>
							</def-item>
							<def-item>
								<term>
									<italic>FN</italic> =</term>
								<def>
									<p>False Negatives</p>
								</def>
							</def-item>
						</def-list>
					<list list-type="bullet">
						<list-item>
							<p>Precision: defined as the number of correctly detected syllable nuclei over the actual cases, that is:</p>
						</list-item>
					</list>
					<disp-formula id="e2">
						<mml:math id="mml-2">
							<mml:mi mathvariant="normal">p</mml:mi>
							<mml:mi mathvariant="normal">r</mml:mi>
							<mml:mi mathvariant="normal">e</mml:mi>
							<mml:mi mathvariant="normal">c</mml:mi>
							<mml:mi mathvariant="normal">i</mml:mi>
							<mml:mi mathvariant="normal">s</mml:mi>
							<mml:mi mathvariant="normal">i</mml:mi>
							<mml:mi mathvariant="normal">o</mml:mi>
							<mml:mi mathvariant="normal">n</mml:mi>
							<mml:mo>=</mml:mo>
							<mml:mfrac>
								<mml:mrow>
									<mml:mi mathvariant="normal">T</mml:mi>
									<mml:mi mathvariant="normal">P</mml:mi>
								</mml:mrow>
								<mml:mrow>
									<mml:mi mathvariant="normal">T</mml:mi>
									<mml:mi mathvariant="normal">P</mml:mi>
									<mml:mo>+</mml:mo>
									<mml:mi mathvariant="normal">F</mml:mi>
									<mml:mi mathvariant="normal">P</mml:mi>
								</mml:mrow>
							</mml:mfrac>
						</mml:math>
						<label>(2)</label>
					</disp-formula>
					<list list-type="bullet">
						<list-item>
							<p>Recall: defined as the number of correctly detected syllable nuclei over the estimated cases, that is:</p>
						</list-item>
					</list>
					<disp-formula id="e3">
						<mml:math id="mml-3">
							<mml:mi mathvariant="normal">r</mml:mi>
							<mml:mi mathvariant="normal">e</mml:mi>
							<mml:mi mathvariant="normal">c</mml:mi>
							<mml:mi mathvariant="normal">a</mml:mi>
							<mml:mi mathvariant="normal">l</mml:mi>
							<mml:mi mathvariant="normal">l</mml:mi>
							<mml:mo>=</mml:mo>
							<mml:mfrac>
								<mml:mrow>
									<mml:mi mathvariant="normal">T</mml:mi>
									<mml:mi mathvariant="normal">P</mml:mi>
								</mml:mrow>
								<mml:mrow>
									<mml:mi mathvariant="normal">T</mml:mi>
									<mml:mi mathvariant="normal">P</mml:mi>
									<mml:mo>+</mml:mo>
									<mml:mi mathvariant="normal">F</mml:mi>
									<mml:mi mathvariant="normal">N</mml:mi>
								</mml:mrow>
							</mml:mfrac>
						</mml:math>
						<label>(3)</label>
					</disp-formula>
					<list list-type="bullet">
						<list-item>
							<p>F-score: defined as the combination of both precision and recall in the following form:</p>
						</list-item>
					</list>
					<disp-formula id="e4">
						<mml:math id="mml-4">
							<mml:mi>F</mml:mi>
							<mml:mo>-</mml:mo>
							<mml:mi>s</mml:mi>
							<mml:mi>c</mml:mi>
							<mml:mi>o</mml:mi>
							<mml:mi>r</mml:mi>
							<mml:mi>e</mml:mi>
							<mml:mo>=</mml:mo>
							<mml:mn>2</mml:mn>
							<mml:mi>&#xa0;</mml:mi>
							<mml:mfrac>
								<mml:mrow>
									<mml:mi>p</mml:mi>
									<mml:mi>r</mml:mi>
									<mml:mi>e</mml:mi>
									<mml:mi>c</mml:mi>
									<mml:mi>i</mml:mi>
									<mml:mi>s</mml:mi>
									<mml:mi>i</mml:mi>
									<mml:mi>o</mml:mi>
									<mml:mi>n</mml:mi>
									<mml:mo>&#x2219;</mml:mo>
									<mml:mi>r</mml:mi>
									<mml:mi>e</mml:mi>
									<mml:mi>c</mml:mi>
									<mml:mi>a</mml:mi>
									<mml:mi>l</mml:mi>
									<mml:mi>l</mml:mi>
								</mml:mrow>
								<mml:mrow>
									<mml:mi>p</mml:mi>
									<mml:mi>r</mml:mi>
									<mml:mi>e</mml:mi>
									<mml:mi>c</mml:mi>
									<mml:mi>i</mml:mi>
									<mml:mi>s</mml:mi>
									<mml:mi>i</mml:mi>
									<mml:mi>o</mml:mi>
									<mml:mi>n</mml:mi>
									<mml:mo>+</mml:mo>
									<mml:mi>r</mml:mi>
									<mml:mi>e</mml:mi>
									<mml:mi>c</mml:mi>
									<mml:mi>a</mml:mi>
									<mml:mi>l</mml:mi>
									<mml:mi>l</mml:mi>
								</mml:mrow>
							</mml:mfrac>
						</mml:math>
						<label>(4)</label>
					</disp-formula>
				</sec>
				<sec id="sec3.3.2">
					<label>3.3.2.</label>
					<title>Assessment metrics</title>
					<p>
						<xref ref-type="bibr" rid="B11">Farr&#xfa;s et al. (2021)</xref> explored the adequacy of several metrics commonly used in the literature, such as: (a) correlation coefficient between actual number of syllables -or speech rate- and estimated number of syllables -or speech rate measurement-, (b) mean error defined as the mean of the error in absolute values, (c) standard deviation error defined as the standard deviation of the previous mean, (d) coefficient of variation defined as (standard deviation error)/ (mean error), (e) mean square error (MSE), (f) root mean square error (RMSE), and (g) normalized root mean square error (NRMSE), by mean defined as:</p>
					<disp-formula id="e5">
						<mml:math id="mml-5">
							<mml:mi>R</mml:mi>
							<mml:mi>M</mml:mi>
							<mml:mi>S</mml:mi>
							<mml:mi>E</mml:mi>
							<mml:mo>=</mml:mo>
							<mml:msqrt>
								<mml:mfrac>
									<mml:mrow>
										<mml:mrow>
											<mml:msubsup>
												<mml:mo stretchy="false">&#x2211;</mml:mo>
												<mml:mrow>
													<mml:mi>i</mml:mi>
													<mml:mo>=</mml:mo>
													<mml:mn>1</mml:mn>
												</mml:mrow>
												<mml:mrow>
													<mml:mi>N</mml:mi>
												</mml:mrow>
											</mml:msubsup>
											<mml:mrow>
												<mml:msup>
													<mml:mrow>
														<mml:mfenced separators="|">
															<mml:mrow>
																<mml:msub>
																	<mml:mrow>
																		<mml:mover accent="true">
																			<mml:mrow>
																				<mml:mi>y</mml:mi>
																			</mml:mrow>
																			<mml:mo>^</mml:mo>
																		</mml:mover>
																	</mml:mrow>
																	<mml:mrow>
																		<mml:mi>i</mml:mi>
																	</mml:mrow>
																</mml:msub>
																<mml:mo>-</mml:mo>
																<mml:mover accent="true">
																	<mml:mrow>
																		<mml:mi>y</mml:mi>
																	</mml:mrow>
																	<mml:mo>^</mml:mo>
																</mml:mover>
															</mml:mrow>
														</mml:mfenced>
													</mml:mrow>
													<mml:mrow>
														<mml:mn>2</mml:mn>
													</mml:mrow>
												</mml:msup>
											</mml:mrow>
										</mml:mrow>
									</mml:mrow>
									<mml:mrow>
										<mml:mi>N</mml:mi>
									</mml:mrow>
								</mml:mfrac>
							</mml:msqrt>
						</mml:math>
						<label>(5)</label>
					</disp-formula>
					<disp-formula id="e6">
						<mml:math id="mml-6">
							<mml:mi mathvariant="normal">N</mml:mi>
							<mml:mi mathvariant="normal">R</mml:mi>
							<mml:mi mathvariant="normal">M</mml:mi>
							<mml:mi mathvariant="normal">S</mml:mi>
							<mml:mi mathvariant="normal">E</mml:mi>
							<mml:mo>=</mml:mo>
							<mml:mfrac>
								<mml:mrow>
									<mml:mi mathvariant="normal">R</mml:mi>
									<mml:mi mathvariant="normal">M</mml:mi>
									<mml:mi mathvariant="normal">S</mml:mi>
									<mml:mi mathvariant="normal">E</mml:mi>
								</mml:mrow>
								<mml:mrow>
									<mml:mover accent="true">
										<mml:mrow>
											<mml:mi mathvariant="normal">y</mml:mi>
										</mml:mrow>
										<mml:mo>-</mml:mo>
									</mml:mover>
								</mml:mrow>
							</mml:mfrac>
						</mml:math>
						<label>(6)</label>
					</disp-formula>
						<def-list id="d2">
							<title>where</title>
							<def-item>
								<term>
									<italic>N</italic>
								</term>
								<def>
									<p>is the number of observations</p>
								</def>
							</def-item>
							<def-item>
								<term>
									<italic>y</italic>
									<sub>
										<italic>i</italic>
									</sub>
								</term>
								<def>
									<p>is the <italic>i</italic>th reference (actual) value</p>
								</def>
							</def-item>
							<def-item>
								<term>
									<italic>&#x177;</italic>
									<sub>i</sub>
								</term>
								<def>
									<p>is its corresponding estimated value</p>
								</def>
							</def-item>
							<def-item>
								<term>
									<italic>&#x45e;</italic>
								</term>
								<def>
									<p>is the mean of the measured data</p>
								</def>
							</def-item>
						</def-list>
					<p>
						<xref ref-type="bibr" rid="B11">Farr&#xfa;s et al. (2021)</xref> concluded that correlation coefficients were not adequate for this kind of assessment, and that, instead, the use of the relative error as a unit for the different metrics should be encouraged, since it homogenizes the assessment based on the number of syllables and speech rate, apart from exhibiting consistent and coherent results. In the current paper, we evaluate the performance of both tools by computing the number of syllables, the speech rate, and the relative error. Moreover, as suggested in our previous study, we compare both tools by means of RMSE as an assessment metric, together with its normalized value (NRMSE) for a better comparison between models computed over different scales.</p>
				</sec>
			</sec>
		</sec>
		<sec id="sec4">
			<label>4.</label>
			<title>System comparison</title>
			<sec id="sec4.1">
				<label>4.1.</label>
				<title>Performance analysis</title>
				<p> The two tools analyzed provide the number of syllables detected via a TextGrid with a point tier (in which the syllable nuclei are indicated as points in time). The Spanish databases are labeled sound-by-sound using interval tiers. In order to make the results comparable, we have combined the automatic point tier with the manual interval tier. We considered that the system succeeded when either there is a point in the time range of a manual interval labelled as vowel (true positive, TP) or there is no point within the time range of a manual interval labelled as consonant (true negative, TN). The system fails when we have a point within a consonant time range (false positive, FP), we have more than a point within a vowel time range (as many false positives as surplus points) or there is no point within a vowel time range (false negative, FN).</p>
				<p> This comparison method is accurate for our purpose (computing the number of syllables detected by the script). However, it would not be accurate for tasks where the interest was the actual center of syllable nuclei, since the method counts as correct any point that falls within the vowel range without taking into account whether the script has placed the point in the vowel mid-point.</p>
				<p>
					<xref ref-type="table" rid="t1">Table 1</xref> shows the comparison of the results obtained by both tools in terms of performance analysis. It details the number of syllable nuclei (vowels) correctly detected (True Positives), wrongly detected (False Positives), missed (False Negatives) and correctly dismissed (True Negatives) by the two tools (v1 and v3) in the two analyzed conditions (read data and spontaneous data). The same results expressed in percentage are illustrated in <xref ref-type="fig" rid="f3">Figure 3</xref>.</p>
				<table-wrap id="t1">
					<label>Table 1</label>
					<caption>
						<title>Number of syllables detected (TP), false positives (FP) and missed (FN) and true negatives by the two scripts (v1 and v3 in the two situations (read corpus and spontaneous corpus))</title>
					</caption>
					<table>
						<colgroup>
							<col/>
							<col/>
							<col/>
							<col/>
							<col/>
						</colgroup>
						<thead>
							<tr>
								<th align="left" rowspan="2"> </th>
								<th align="center" colspan="2">read </th>
								<th align="center" colspan="2">spontaneous </th>
							</tr>
							<tr>
								<th align="center">V1</th>
								<th align="center">V3</th>
								<th align="center">V1</th>
								<th align="center">V3</th>
							</tr>
						</thead>
						<tbody>
							<tr>
								<td align="center">TP</td>
								<td align="center">18137</td>
								<td align="center">16989</td>
								<td align="center">23208</td>
								<td align="center">21671</td>
							</tr>
							<tr>
								<td align="center">FP</td>
								<td align="center">2034</td>
								<td align="center">1147</td>
								<td align="center">3655</td>
								<td align="center">1851</td>
							</tr>
							<tr>
								<td align="center">FN</td>
								<td align="center">8847</td>
								<td align="center">9995</td>
								<td align="center">12158</td>
								<td align="center">13695</td>
							</tr>
							<tr>
								<td align="center">TN</td>
								<td align="center">29740</td>
								<td align="center">30601</td>
								<td align="center">37990</td>
								<td align="center">39049</td>
							</tr>
						</tbody>
					</table>
				</table-wrap>
				<fig id="f3">
					<label>Figure 3</label>
					<caption>
						<title>Percentage of detected (TP), False Positive, Missed (FN), and True Negative syllables in v1 and v3.</title>
					</caption>
					<graphic id="gra-3" xlink:href="LOQUENS-9-1-2-e090-gf3.png"/>
				</fig>
				<p>
					<xref ref-type="fig" rid="f3">Figure 3</xref> shows that v1 results are better for detected syllables (higher value) and missed syllables (lower value), whereas v3 performs better for true negatives (a higher value) and false positives (a lower value), taking as a reference the number of manually annotated vowels (nucleus) and consonants (non-nucleus) in both subcorpora (26984 nuclei for read, 35366 for spontaneous speech).</p>
				<p>
					<xref ref-type="table" rid="t2">Table 2</xref> shows the main performance metrics for tool v1 and v3 for read and spontaneous speech with the best result highlighted in bold. Results show that, in general, v3 is the best tool when we consider precision and v1 shows a better performance in recall and F-score. Accuracy reveals contradictory results, having best results for tool v1 in read speech and the best result for tool v3 in spontaneous speech. However, accuracy is a discouraged metric in cases of heavily imbalanced cases (big difference between the number of false positives and false negatives or missed cases) (e.g. <xref ref-type="bibr" rid="B17">Mortaz, 2020</xref>) and in those cases performance analysis should rely on F-score metrics.</p>
				<table-wrap id="t2">
					<label>Table 2</label>
					<caption>
						<title>Performance metrics for tool v1 and v3 for read and spontaneous speech.</title>
					</caption>
					<table>
                <colgroup>
	                <col/>
	                <col/>
	                <col/>
	                <col/>
	                <col/>
                </colgroup>
                <thead>
                  <tr>
                    <th rowspan="2" align="center"> </th>
                    <th colspan="2" align="center">read </th>
                    <th colspan="2" align="center">spontaneous </th>
                  </tr>
                  <tr>
                    <th align="center">v1</th>
                    <th align="center">v3</th>
                    <th align="center">v1</th>
                    <th align="center">v3</th>
                  </tr>
                </thead>
                <tbody>
                  <tr>
                    <td align="left">accuracy</td>
                    <td align="center"><bold>0.815</bold></td>
                    <td align="center">0.810</td>
                    <td align="center">0.795</td>
                    <td align="center"><bold>0.796</bold></td>
                  </tr>
                  <tr>
                    <td align="left">precision</td>
                    <td align="center">0.899</td>
                    <td align="center"><bold>0.994</bold></td>
                    <td align="center">0.864</td>
                    <td align="center"><bold>0.921</bold></td>
                  </tr>
                  <tr>
                    <td align="left">recall</td>
                    <td align="center"><bold>0.672</bold></td>
                    <td align="center">0.629</td>
                    <td align="center"><bold>0.656</bold></td>
                    <td align="center">0.613</td>
                  </tr>
                  <tr>
                    <td align="left">F-score</td>
                    <td align="center"><bold>0.769</bold></td>
                    <td align="center">0.753</td>
                    <td align="center"><bold>0.746</bold></td>
                    <td align="center">0.733</td>
                  </tr>
                </tbody>
              </table>
				</table-wrap>
			</sec>
			<sec id="sec4.2">
				<label>4.2.</label>
				<title>Assessment metrics</title>
				<p> In this section, we present the assessment metrics obtained for the following units of analysis: number of syllables, speech rate, and relative error. Speech rate is defined as number of syllables per second, and the relative error is defined as:</p>
				<disp-formula id="e7">
					<mml:math id="mml-7">
						<mml:msub>
							<mml:mrow>
								<mml:mi>&#x3b5;</mml:mi>
							</mml:mrow>
							<mml:mrow>
								<mml:mi>r</mml:mi>
							</mml:mrow>
						</mml:msub>
						<mml:mo>=</mml:mo>
						<mml:mfrac>
							<mml:mrow>
								<mml:mfenced close="|" open="|" separators="|">
									<mml:mrow>
										<mml:msub>
											<mml:mrow>
												<mml:mo>#</mml:mo>
												<mml:mi mathvariant="normal">s</mml:mi>
												<mml:mi mathvariant="normal">y</mml:mi>
												<mml:mi mathvariant="normal">l</mml:mi>
												<mml:mi mathvariant="normal">l</mml:mi>
											</mml:mrow>
											<mml:mrow>
												<mml:mi>a</mml:mi>
											</mml:mrow>
										</mml:msub>
										<mml:mo>-</mml:mo>
										<mml:mo>#</mml:mo>
										<mml:msub>
											<mml:mrow>
												<mml:mi mathvariant="normal">s</mml:mi>
												<mml:mi mathvariant="normal">y</mml:mi>
												<mml:mi mathvariant="normal">l</mml:mi>
												<mml:mi mathvariant="normal">l</mml:mi>
											</mml:mrow>
											<mml:mrow>
												<mml:mi>m</mml:mi>
											</mml:mrow>
										</mml:msub>
									</mml:mrow>
								</mml:mfenced>
							</mml:mrow>
							<mml:mrow>
								<mml:mo>#</mml:mo>
								<mml:msub>
									<mml:mrow>
										<mml:mi mathvariant="normal">s</mml:mi>
										<mml:mi mathvariant="normal">y</mml:mi>
										<mml:mi mathvariant="normal">l</mml:mi>
										<mml:mi mathvariant="normal">l</mml:mi>
									</mml:mrow>
									<mml:mrow>
										<mml:mi>m</mml:mi>
									</mml:mrow>
								</mml:msub>
							</mml:mrow>
						</mml:mfrac>
					</mml:math>
					<label>(7)</label>
				</disp-formula>
				<p> where #<italic>sylla</italic> is the estimated (automatic) count of syllables, and #<italic>syllm</italic> in the actual (manual) count. Since speech rate is obtained using the number of syllables along the entire speech duration, and the length of the spurt analyzed is the same in both evaluations (automatic and manual), the relative error applied to the number of syllables and to speech rate coincides, making it a homogenized measurement.</p>
				<p> In <xref ref-type="table" rid="t3">Table 3</xref>, we show the total number of syllables obtained with the manual transcriptions in the entire corpus, as well as the number of syllables obtained in both automatic tools (v1 and v3) for read and spontaneous modalities. The results clearly show that v3 fails more than v1 when detecting syllable nuclei, although both tools underestimate the actual number of syllables.</p>
				<table-wrap id="t3">
					<label>Table 3</label>
					<caption>
						<title>Total number of syllables obtained with the manual transcriptions and the automatic tools for both read and spontaneous speech.</title>
					</caption>
					<table>
						<colgroup>
							<col/>
							<col/>
							<col/>
							<col/>
							<col/>
							<col/>
						</colgroup>
						<thead>
							<tr>
								<th align="center" colspan="3">read </th>
								<th align="center" colspan="3">spontaneous </th>
							</tr>
						</thead>
						<tbody>
							<tr>
								<th align="center">manual</th>
								<th align="center">v1</th>
								<th align="center">v3</th>
								<th align="center">manual</th>
								<th align="center">v1</th>
								<th align="center">v3</th>
							</tr>
							<tr>
								<td align="center">27005</td>
								<td align="center">20165</td>
								<td align="center">18155</td>
								<td align="center">35408</td>
								<td align="center">27491</td>
								<td align="center">24003</td>
							</tr>
						</tbody>
					</table>
				</table-wrap>
				<p>
					<xref ref-type="table" rid="t4">Tables 4</xref> and <xref ref-type="table" rid="t5">5</xref> show the root mean square error (RMSE) and normalized root mean square error (NRMSE) respectively, obtained for both tools, the different units of analysis (number of syllables, speech rate, and error rate), and both read and spontaneous modalities.</p>
				<table-wrap id="t4">
					<label>Table 4</label>
					<caption>
						<title>RMSE obtained for the different units of analysis</title>
					</caption>
					<table>
                <colgroup>
                <col/>
                <col/>
                <col/>
                <col/>
                <col/>
                </colgroup>
                <thead>
                  <tr>
                    <th rowspan="2" align="left"> </th>
                    <th colspan="2" align="center">read </th>
                    <th colspan="2" align="center">spontaneous </th>
                  </tr>
                  <tr>
                    <th align="center">v1</th>
                    <th align="center">v3</th>
                    <th align="center">v1</th>
                    <th align="center">v3</th>
                  </tr>
                </thead>
                <tbody>
                  <tr>
                    <td align="left">#syllables</td>
                    <td align="center"><bold>78.3</bold></td>
                    <td align="center">99.8</td>
                    <td align="center"><bold>99.2</bold></td>
                    <td align="center">135.3</td>
                  </tr>
                  <tr>
                    <td align="left">speech rate</td>
                    <td align="center"><bold>1.420</bold></td>
                    <td align="center">1.789</td>
                    <td align="center"><bold>3.604</bold></td>
                    <td align="center">4.032</td>
                  </tr>
                  <tr>
                    <td align="left">error rate</td>
                    <td align="center"><bold>0.261</bold></td>
                    <td align="center">0.333</td>
                    <td align="center"><bold>0.233</bold></td>
                    <td align="center">0.326</td>
                  </tr>
                </tbody>
              </table>
				</table-wrap>
				<table-wrap id="t5">
					<label>Table 5</label>
					<caption>
						<title>NRMSE obtained for the different units of analysis.</title>
					</caption>
					<table>
                <colgroup>
                <col/>
                <col/>
                <col/>
                <col/>
                <col/>
                </colgroup>
                <thead>
                  <tr>
                    <th rowspan="2" align="left"> </th>
                    <th colspan="2" align="center">read </th>
                    <th colspan="2" align="center">spontaneous </th>
                  </tr>
                  <tr>
                    <th align="center">v1</th>
                    <th align="center">v3</th>
                    <th align="center">v1</th>
                    <th align="center">v3</th>
                  </tr>
                </thead>
                <tbody>
                  <tr>
                    <td align="left">#syllables</td>
                    <td align="center">1.030</td>
                    <td align="center"><bold>1.015</bold></td>
                    <td align="center">1.128</td>
                    <td align="center"><bold>1.068</bold></td>
                  </tr>
                  <tr>
                    <td align="left">speech rate</td>
                    <td align="center">1.050</td>
                    <td align="center"><bold>1.033</bold></td>
                    <td align="center">1.136</td>
                    <td align="center"><bold>1.105</bold></td>
                  </tr>
                  <tr>
                    <td align="left">error rate</td>
                    <td align="center">1.031</td>
                    <td align="center"><bold>1.015</bold></td>
                    <td align="center">1.063</td>
                    <td align="center"><bold>1.030</bold></td>
                  </tr>
                </tbody>
              </table>
				</table-wrap>
				<p> The best result within both tools, with each assessment metric and for both read and spontaneous speech is highlighted in bold. The results mainly show that, while tool v1 performs better when it is evaluated by means of RMSE, tool v3 performs better if we consider NRMSE.</p>
			</sec>
		</sec>
		<sec id="sec5" sec-type="discussion">
			<label>5.</label>
			<title>Discussion</title>
			<p> The performance analysis (see 4.1) shows that both tools are reliable finding syllable nuclei (precision&gt; 0.8 and recall &gt; 0.5, in all cases). Also, both tools perform better with read speech than with spontaneous speech. However, they share a common problem in classification tasks: an imbalanced classification, with more false negatives than false positives, which complicates the assessment. For our data, this is a foreseeable result given that finding a syllable nucleus that is not preceded or followed by an intensity dip is more usual in speech than it is for a voiced intensity peak to be a consonant. This means that, if we want the tool to correctly disregard peaks that are not vowels, we need to use a more restrictive system -which v3 does by requiring a preceding and following intensity dip in order to consider an interval as a syllable nucleus- and that will give us a better result in accuracy and precision, which is exactly what is shown in <xref ref-type="table" rid="t2">Table 2</xref>, given that precision is as a measure of quality, meaning that the vowels that are marked as vowels with v3 are more likely to be real vowels. However, if we consider the global result of correctly identified and disregarded syllable nuclei (the quantity) a less restrictive rule (i.e., v1, which only considers previous intensity dips) has a better performance as illustrated by <xref ref-type="table" rid="t2">Table 2</xref> recall and F-score.</p>
			<p> For the aim of this paper, which is the automatic computation of speech rate, quantity measures can prove more relevant than quality measures given that, when computing speech rate, we are not interested in knowing whether the segment is a syllable nucleus but rather in getting a number of syllable nuclei as close as possible to the actual one. That is, if a false positive is later compensated by a missed nucleus the system is still accurate. This is the exact scenario when in a real syllable the automatic tool places the syllable nucleus within the onset instead of in the actual nucleus, but then does not label the vowel as nucleus.</p>
			<p> In <xref ref-type="table" rid="t3">Table 3</xref>, we can clearly see that the number of syllables counted by v1 is closer to the actual number of syllables counted by v3. In other words, v3 is missing a larger number of syllables, which results in larger values of RMSE for v3, both in the read and the spontaneous modalities (<xref ref-type="table" rid="t4">Table 4</xref>). These results are consistent with those shown in <xref ref-type="table" rid="t1">Table 1</xref>, also illustrated in <xref ref-type="fig" rid="f2">Figure 2</xref>: the number of detected (true positives) and false positive syllables is greater in v1. The number of missed syllables (false negatives) also contributes to enlarge the underestimation in the syllable counting.</p>
			<p> However, the NRMSE metric (<xref ref-type="table" rid="t5">Table 5</xref>) shows otherwise: the RMSE normalized values by the mean of the measured data in v3 outperform those obtained in v1. The fact is that, although v3 fails more in detecting syllables than v1, such failure is more stable. This is strengthened by the measurement of other metrics such as the standard error (standard deviation of the mean error) and the coefficient of variation -or relative standard deviation- defined as (standard deviation)/mean). For both measurements, v3 shows a better performance than v1 for both read and spontaneous modalities.</p>
			<p> On the one hand, this shows that, although v3 fails largely in missing syllables, such failure could be better compensated by a correction factor. On the other hand, and since v3 appears to be more restrictive in the detection conditions of syllables - we need an intensity dip in both side of the vowel and not only one as in v1-, but we can also ensure that the detected syllables come more often from actual syllable nuclei in v3 than in v1, in which the detected syllable could come more often from false nuclei. This is also strengthened by the larger number of true negatives in v3 encountered in <xref ref-type="table" rid="t1">Table 1</xref> for both modalities.</p>
		</sec>
		<sec id="sec6" sec-type="conclusions">
			<label>6.</label>
			<title>Conclusions</title>
			<p> The results presented and discussed in the previous sections indicate, on the one hand, that both methods of syllable detection are not fully reliable yet to face a speech rate analysis task: both detect a number of syllables which is remarkably lower than the number of syllables obtained from a manual annotation. However, v1 seems to offer a better performance for this task than v3, as the number of detected syllables is closer to the manually obtained value, which compensates the fact that it is less precise in the detection of actual syllables, a fact that is secondary in a speech rate calculation task if the number of detected syllables is close enough to the number of manually annotated ones.</p>
			<p> On the other hand, the results also show that, although v3 detects in general less true syllables than v1, it seems more adequate for tasks in which it is important that detected syllables correspond to actual syllables, such as automatic acoustic measurements of corpora involving the detection of syllabic nuclei.</p>
		</sec>
	</body>
	<back>
		<ack>
			<title>Acknowledgements</title>
			<p> This work has been partially funded by the project &#x201c;M&#xe9;todos, modelos, m&#xe9;tricas y herramientas para la evaluaci&#xf3;n de la prosodia (ProA)&#x201d;, reference number PGC2018-094233-B-C21. The first author is a &#x201c;Serra H&#xfa;nter Fellow&#x201d;. The authors would like to thank Dr. Mar&#xed;a Machuca for providing the VILE corpus used in these experiments</p>
		</ack>
		<ref-list>
			<label>7.</label>
			<title>References</title>
			<ref id="B1">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Albal&#xe1;</surname>
							<given-names>M. J.</given-names>
						</string-name>
						<string-name>
							<surname>Battaner</surname>
							<given-names>E.</given-names>
						</string-name>
						<string-name>
							<surname>Carranza</surname>
							<given-names>M.</given-names>
						</string-name>
						<string-name>
							<surname>Mota Gorriz</surname>
							<given-names>C. d. l.</given-names>
						</string-name>
						<string-name>
							<surname>Gil</surname>
							<given-names>J.</given-names>
						</string-name>
						<string-name>
							<surname>Llisterri</surname>
							<given-names>J.</given-names>
						</string-name>
						<etal/>
					</person-group>
					<year>2008</year>
					<article-title>VILE: An&#xe1;lisis estad&#xed;stico de los par&#xe1;metros relacionados con el grupo de entonaci&#xf3;n</article-title>
					<source>Language Design: Journal of Theoretical and Experimental Linguistics</source>
					<issue>Special Issue</issue>
					<fpage>15</fpage>
					<lpage>21</lpage>
				</mixed-citation>
			</ref>
			<ref id="B2">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Battaner Moro</surname>
							<given-names>E.</given-names>
						</string-name>
						<string-name>
							<surname>Gil Fern&#xe1;ndez</surname>
							<given-names>J.</given-names>
						</string-name>
						<string-name>
							<surname>Marrero Aguiar</surname>
							<given-names>V.</given-names>
						</string-name>
						<string-name>
							<surname>Carbo Marro</surname>
							<given-names>C.</given-names>
						</string-name>
						<string-name>
							<surname>Llisterri Boix</surname>
							<given-names>J.</given-names>
						</string-name>
						<string-name>
							<surname>Machuca Ayuso</surname>
							<given-names>J.</given-names>
						</string-name>
						<string-name>
							<surname>R&#xed;os Mestre</surname>
							<given-names>A.</given-names>
						</string-name>
						<etal/>
					</person-group>
					<year>2005</year>
					<article-title>VILE: estudio ac&#xfa;stico de la variaci&#xf3;n inter- e intralocutor en espa&#xf1;ol</article-title>
					<source>Procesamiento del Lenguaje Natural</source>
					<volume>35</volume>
					<fpage>435</fpage>
					<lpage>436</lpage>
				</mixed-citation>
			</ref>
			<ref id="B3">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Cucchiarini</surname>
							<given-names>C.</given-names>
						</string-name>
						<string-name>
							<surname>Strik</surname>
							<given-names>H.</given-names>
						</string-name>
						<string-name>
							<surname>Boves</surname>
							<given-names>L.</given-names>
						</string-name>
					</person-group>
					<year>1998</year>
					<conf-name>Quantitative assessment of second language learners&#x2019; fluency: An automatic approach</conf-name>
					<source>Proceedings of the 5th International Conference on Spoken Language Processing (ICSLP&#x2019;98)</source>
					<fpage>2619</fpage>
					<lpage>2622</lpage>
					<publisher-loc>Sydney, Australia</publisher-loc>
					<pub-id pub-id-type="doi">10.21437/ICSLP.1998-754</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B4">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Cucchiarini</surname>
							<given-names>C.</given-names>
						</string-name>
						<string-name>
							<surname>Strik</surname>
							<given-names>H.</given-names>
						</string-name>
						<string-name>
							<surname>Boves</surname>
							<given-names>L.</given-names>
						</string-name>
					</person-group>
					<year>2000a</year>
					<article-title>Different aspects of expert pronunciation quality ratings and their relation to scores produced by speech recognition algorithms</article-title>
					<source>Speech Communication</source>
					<volume>30</volume>
					<issue>2-3</issue>
					<fpage>109</fpage>
					<lpage>119</lpage>
					<pub-id pub-id-type="doi">10.1016/S0167-6393(99)00040-0</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B5">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Cucchiarini</surname>
							<given-names>C.</given-names>
						</string-name>
						<string-name>
							<surname>Strik</surname>
							<given-names>H.</given-names>
						</string-name>
						<string-name>
							<surname>Boves</surname>
							<given-names>L.</given-names>
						</string-name>
					</person-group>
					<year>2000b</year>
					<article-title>Quantitative assessment of second language learners&#x2019; fluency by means of automatic speech recognition technology</article-title>
					<source>Journal of the Acoustical Society of America</source>
					<volume>107</volume>
					<issue>2</issue>
					<fpage>989</fpage>
					<lpage>999</lpage>
					<pub-id pub-id-type="doi">10.1121/1.428279</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B6">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Cucchiarini</surname>
							<given-names>C.</given-names>
						</string-name>
						<string-name>
							<surname>Strik</surname>
							<given-names>H.</given-names>
						</string-name>
						<string-name>
							<surname>Boves</surname>
							<given-names>L.</given-names>
						</string-name>
					</person-group>
					<year>2002</year>
					<article-title>Quantitative assessment of second language learners&#x2019; fluency: comparisons between read and spontaneous speech</article-title>
					<source>Journal of the Acoustical Society of America</source>
					<volume>111</volume>
					<issue>6</issue>
					<fpage>2862</fpage>
					<lpage>2873</lpage>
					<pub-id pub-id-type="doi">10.1121/1.1471894</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B7">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>de Jong</surname>
							<given-names>N. H.</given-names>
						</string-name>
						<string-name>
							<surname>Pacilly</surname>
							<given-names>J.</given-names>
						</string-name>
						<string-name>
							<surname>Wempe</surname>
							<given-names>T.</given-names>
						</string-name>
					</person-group>
					<year>2021</year>
					<article-title>Praat scripts to measure speed fluency and breakdown fluency in speech automatically</article-title>
					<source>Assessment in Education: Principles, Policy and Practice</source>
					<volume>28</volume>
					<issue>4</issue>
					<fpage>456</fpage>
					<lpage>476</lpage>
					<pub-id pub-id-type="doi">10.1080/0969594X.2021.1951162</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B8">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>de Jong</surname>
							<given-names>N. H.</given-names>
						</string-name>
						<string-name>
							<surname>Wempe</surname>
							<given-names>T.</given-names>
						</string-name>
					</person-group>
					<year>2009</year>
					<article-title>Praat script to detect syllable nuclei and measure speech rate automatically</article-title>
					<source>Behavior Research Methods</source>
					<volume>41</volume>
					<issue>2</issue>
					<fpage>385</fpage>
					<lpage>390</lpage>
					<pub-id pub-id-type="doi">10.3758/BRM.41.2.385</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B9">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>de Jong</surname>
							<given-names>N. H.</given-names>
						</string-name>
						<string-name>
							<surname>Wempe</surname>
							<given-names>T.</given-names>
						</string-name>
						<etal/>
					</person-group>
					<year>2007</year>
					<article-title>Automatic measurement of speech rate in spoken Dutch</article-title>
					<source>ACLC Working Papers</source>
					<volume>2</volume>
					<fpage>51</fpage>
					<lpage>60</lpage>
				</mixed-citation>
			</ref>
			<ref id="B10">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Dekens</surname>
							<given-names>T.</given-names>
						</string-name>
						<string-name>
							<surname>Demol</surname>
							<given-names>M.</given-names>
						</string-name>
						<string-name>
							<surname>Verhelst</surname>
							<given-names>W.</given-names>
						</string-name>
						<string-name>
							<surname>Verhoeve</surname>
							<given-names>P.</given-names>
						</string-name>
					</person-group>
					<year>2007</year>
					<source>A comparative study of speech rate estimation techniques</source>
					<conf-name>Proceedings of the Eighth Annual Conference of the International Speech Communication Association (INTERSPEECH 2007)</conf-name>
					<fpage>510</fpage>
					<lpage>513</lpage>
					<pub-id pub-id-type="doi">10.21437/Interspeech.2007-237</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B11">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Farr&#xfa;s</surname>
							<given-names>M.</given-names>
						</string-name>
						<string-name>
							<surname>Elvira-Garc&#xed;a</surname>
							<given-names>W.</given-names>
						</string-name>
						<string-name>
							<surname>Garrido- Almi&#xf1;ana</surname>
							<given-names>J. M.</given-names>
						</string-name>
					</person-group>
					<year>2021</year>
					<source>On the need of standard assessment metrics for automatic speech rate computation tools</source>
					<conf-name>4th Phonetics and Phonology in Europe 2021 Conference (PAPE 2021)</conf-name>
				</mixed-citation>
			</ref>
			<ref id="B12">
				<mixed-citation publication-type="database">
					<person-group person-group-type="author">
						<string-name>
							<surname>Garofolo</surname>
							<given-names>J.-S.</given-names>
						</string-name>
						<etal/>
					</person-group>
					<year>1993</year>
					<source>TIMIT Acoustic-Phonetic Continuous Speech Corpus</source>
					<gov>LDC93S1</gov>
					<comment>Web Download</comment>
					<publisher-loc>Philadelphia</publisher-loc>
					<publisher-name>Linguistic Data Consortium</publisher-name>
				</mixed-citation>
			</ref>
			<ref id="B13">
				<mixed-citation publication-type="database">
					<person-group person-group-type="author">
						<string-name>
							<surname>Godfrey</surname>
							<given-names>J.-J.</given-names>
						</string-name>
						<string-name>
							<surname>Holliman</surname>
							<given-names>E.</given-names>
						</string-name>
					</person-group>
					<year>1993</year>
					<source>Switchboard-1 Release 2</source>
					<gov>LDC97S62</gov>
					<comment>Web Download</comment>
					<publisher-loc>Philadelphia</publisher-loc>
					<publisher-name>Linguistic Data Consortium</publisher-name>
				</mixed-citation>
			</ref>
			<ref id="B14">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Goldman</surname>
							<given-names>J.-P.</given-names>
						</string-name>
					</person-group>
					<year>2011</year>
					<source>Easyalign: an automatic phonetic alignment tool under Praat</source>
					<conf-name>Proceedings of the 12th Annual Conference of the International Speech Communication Association (INTERSPEECH 2011)</conf-name>
					<conf-date>28-21 August, 2011</conf-date>
					<conf-loc>Florence, Italy</conf-loc>
				</mixed-citation>
			</ref>
			<ref id="B15">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Honig</surname>
							<given-names>F.</given-names>
						</string-name>
						<string-name>
							<surname>Batliner</surname>
							<given-names>A.</given-names>
						</string-name>
						<string-name>
							<surname>Weilhammer</surname>
							<given-names>K.</given-names>
						</string-name>
						<string-name>
							<surname>N&#xf6;th</surname>
							<given-names>E.</given-names>
						</string-name>
					</person-group>
					<year>2010</year>
					<source>Automatic assessment of non-native prosody for English as L2</source>
					<conf-name>Speech Prosody 2010</conf-name>
					<conf-loc>Chicago, IL, USA</conf-loc>
				</mixed-citation>
			</ref>
			<ref id="B16">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Llisterri</surname>
							<given-names>J.</given-names>
						</string-name>
						<string-name>
							<surname>Machuca</surname>
							<given-names>M.</given-names>
						</string-name>
						<string-name>
							<surname>R&#xed;os</surname>
							<given-names>A.</given-names>
						</string-name>
					</person-group>
					<month>06</month>
					<year>2017</year>
					<source>VILE-P: un corpus para el estudio prosodico de la variaci&#xf3;n inter e intralocutor</source>
					<comment>Comunicaci&#xf3;n presentada en</comment>
					<conf-name>SUBSIDIA: Herramientas y recursos para las ciencias del habla</conf-name>
					<conf-loc>M&#xe1;laga, Spain</conf-loc>
				</mixed-citation>
			</ref>
			<ref id="B17">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Mortaz</surname>
							<given-names>E.</given-names>
						</string-name>
					</person-group>
					<year>2020</year>
					<article-title>Imbalance accuracy metric for model selection in multi-class imbalance classification problems</article-title>
					<source>Knowledge-Based Systems</source>
					<volume>210</volume>
					<elocation-id>106490</elocation-id>
					<pub-id pub-id-type="doi">10.1016/j.knosys.2020.106490</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B18">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Narayanan</surname>
							<given-names>S.</given-names>
						</string-name>
						<string-name>
							<surname>Wang</surname>
							<given-names>D.</given-names>
						</string-name>
					</person-group>
					<year>2005</year>
					<source>Speech rate estimation via temporal correlation and selected sub-band correlation</source>
					<conf-name>Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing</conf-name>
					<abbrev>ICASSP</abbrev>
					<volume>1</volume>
					<fpage>1</fpage>
					<lpage>413</lpage>
					<pub-id pub-id-type="doi">10.1109/ICASSP.2005.1415138</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B19">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Neumeyer</surname>
							<given-names>L.</given-names>
						</string-name>
						<string-name>
							<surname>Franco</surname>
							<given-names>H.</given-names>
						</string-name>
						<string-name>
							<surname>Digalakis</surname>
							<given-names>V.</given-names>
						</string-name>
						<string-name>
							<surname>Weintraub</surname>
							<given-names>M.</given-names>
						</string-name>
					</person-group>
					<year>2000</year>
					<article-title>Automatic scoring of pronunciation quality</article-title>
					<source>Speech Communication</source>
					<volume>30</volume>
					<issue>2-3</issue>
					<fpage>88</fpage>
					<lpage>93</lpage>
					<pub-id pub-id-type="doi">10.1016/S0167-6393(99)00046-1</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B20">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Ortega-Garc&#xed;a</surname>
							<given-names>J.</given-names>
						</string-name>
						<string-name>
							<surname>Gonz&#xe1;lez-Rodr&#xed;guez</surname>
							<given-names>J.</given-names>
						</string-name>
						<string-name>
							<surname>Marrero-Aguiar</surname>
							<given-names>V.</given-names>
						</string-name>
					</person-group>
					<year>2000</year>
					<article-title>Ahumada: A large speech corpus in Spanish for speaker characterization and identification</article-title>
					<source>Speech Communication</source>
					<volume>31</volume>
					<issue>2-3</issue>
					<fpage>255</fpage>
					<lpage>264</lpage>
					<pub-id pub-id-type="doi">10.1016/S0167-6393(99)00081-3</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B21">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Pellegrino</surname>
							<given-names>F.</given-names>
						</string-name>
						<string-name>
							<surname>Farinas</surname>
							<given-names>J.</given-names>
						</string-name>
						<string-name>
							<surname>Rouas</surname>
							<given-names>J.-L.</given-names>
						</string-name>
					</person-group>
					<year>2004</year>
					<source>Automatic estimation of speaking rate in multilingual spontaneous speech</source>
					<conf-name>Speech Prosody 2004</conf-name>
					<fpage>517</fpage>
					<lpage>520</lpage>
				</mixed-citation>
			</ref>
			<ref id="B22">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Pfau</surname>
							<given-names>T.</given-names>
						</string-name>
						<string-name>
							<surname>Ruske</surname>
							<given-names>G.</given-names>
						</string-name>
					</person-group>
					<year>1998</year>
					<source>Estimating the speaking rate by vowel detection</source>
					<conf-name>Proceedings of the 1998 IEEE International Conference on Acoustics, Speech and Signal Processing</conf-name>
					<abbrev>ICASSP &#x2018;98</abbrev>
					<volume>2</volume>
					<fpage>945</fpage>
					<lpage>948</lpage>
					<pub-id pub-id-type="doi">10.1109/ICASSP.1998.675422</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B23">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Pfitzinger</surname>
							<given-names>H. R.</given-names>
						</string-name>
					</person-group>
					<year>1996</year>
					<source>Two approaches to speech rate estimation</source>
					<conf-name>Proceedings of the 6th Australian International Conference on Speech Science and Technology</conf-name>
					<abbrev>SST, 96</abbrev>
					<volume>96</volume>
					<fpage>421</fpage>
					<lpage>426</lpage>
				</mixed-citation>
			</ref>
			<ref id="B24">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Sabu</surname>
							<given-names>K.</given-names>
						</string-name>
						<string-name>
							<surname>Chaudhuri</surname>
							<given-names>S.</given-names>
						</string-name>
						<string-name>
							<surname>Rao</surname>
							<given-names>P.</given-names>
						</string-name>
						<string-name>
							<surname>Patil</surname>
							<given-names>M.</given-names>
						</string-name>
					</person-group>
					<year>2021</year>
					<source>An optimized signal-processing pipeline for syllable detection and speech rate estimation</source>
					<conf-name>In National Conference on Communications</conf-name>
					<abbrev>NCC, 2020</abbrev>
					<pub-id pub-id-type="doi">10.48550/arXiv.2103.04346</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B25">
				<mixed-citation publication-type="confproc">
					<person-group person-group-type="author">
						<string-name>
							<surname>Verhasselt</surname>
							<given-names>J. P.</given-names>
						</string-name>
						<string-name>
							<surname>Martens</surname>
							<given-names>J.-P.</given-names>
						</string-name>
					</person-group>
					<year>1996</year>
					<source>A fast and reliable rate of speech detector</source>
					<conf-name>Proceedings of Fourth International Conference on Spoken Language Processing</conf-name>
					<abbrev>ICSLP&#x2019;96</abbrev>
					<volume>4</volume>
					<fpage>2258</fpage>
					<lpage>2261</lpage>
					<pub-id pub-id-type="doi">10.21437/ICSLP.1996-577</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B26">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Wang</surname>
							<given-names>D.</given-names>
						</string-name>
						<string-name>
							<surname>Narayanan</surname>
							<given-names>S. S.</given-names>
						</string-name>
					</person-group>
					<year>2007</year>
					<article-title>Robust speech rate estimation for spontaneous speech</article-title>
					<source>IEEE Transactions on Audio, Speech, and Language Processing</source>
					<volume>15</volume>
					<issue>8</issue>
					<fpage>2190</fpage>
					<lpage>2201</lpage>
					<pub-id pub-id-type="doi">10.1109/TASL.2007.905178</pub-id>
				</mixed-citation>
			</ref>
			<ref id="B27">
				<mixed-citation publication-type="journal">
					<person-group person-group-type="author">
						<string-name>
							<surname>Zechner</surname>
							<given-names>K.</given-names>
						</string-name>
						<string-name>
							<surname>Higgins</surname>
							<given-names>D.</given-names>
						</string-name>
						<string-name>
							<surname>Xia</surname>
							<given-names>X.</given-names>
						</string-name>
						<string-name>
							<surname>Williamson</surname>
							<given-names>D.</given-names>
						</string-name>
					</person-group>
					<year>2009</year>
					<article-title>Automatic scoring of non-native spontaneous speech in tests of spoken English</article-title>
					<source>Speech Communication</source>
					<volume>51</volume>
					<issue>10</issue>
					<pub-id pub-id-type="doi">10.1016/j.specom.2009.04.009</pub-id>
				</mixed-citation>
			</ref>
		</ref-list>
	</back>
</article>