<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "journalpublishing3.dtd">
<article article-type="research-article" dtd-version="3.0" xml:lang="en" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">LOQ</journal-id>
<journal-title-group>
<journal-title>Loquens</journal-title>
</journal-title-group>
<issn pub-type="epub">2386-2637</issn>
<publisher>
<publisher-name>Consejo Superior de Investigaciones Cientificas</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">loquens_e074</article-id>
<article-id pub-id-type="doi">10.3989/loquens.2020.074</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Articles</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Characterizing speech rhythm using spectral coherence between jaw displacement and speech temporal envelope</article-title>
<trans-title-group xml:lang="es">
<trans-title>Caracterizaci&#x00F3;n del ritmo del habla usando la coherencia espectral entre el desplazamiento de la mand&#x00ED;bula y la envolvente temporal del habla</trans-title>
</trans-title-group>
<alt-title alt-title-type="running-head">Characterizing speech rhythm using spectral coherence between jaw displacement and speech temporal envelope</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>He</surname>
<given-names>Lei</given-names>
</name>
<xref ref-type="aff" rid="aff0001">1</xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Zhang</surname>
<given-names>Yu</given-names>
</name>
<xref ref-type="aff" rid="aff0001">1</xref>
</contrib>
</contrib-group>
<aff id="aff0001"><label>1</label>Department of Computational Linguistics, University of Zurich, Switzerland</aff>
<author-notes>
<corresp id="cor1"><email xlink:href="lei.he@uzh.ch">lei.he@uzh.ch</email> ORCID: <ext-link ext-link-type="uri" xlink:href="https://orcid.org/0000-0002-9552-9075">https://orcid.org/0000-0002-9552-9075</ext-link>, <email xlink:href="yu.zhang@uzh.ch">yu.zhang@uzh.ch</email> ORCID: <ext-link ext-link-type="uri" xlink:href="https://orcid.org/0000-0002-0865-7897">https://orcid.org/0000-0002-0865-7897</ext-link></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>28</day>
<month>12</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<month>12</month>
<year>2020</year>
</pub-date>
<volume>7</volume>
<elocation-id>e074</elocation-id>
<history>
<date date-type="received">
<day>17</day>
<month>12</month>
<year>2020</year>
</date>
<date date-type="accepted">
<day>18</day>
<month>05</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2020 CSIC</copyright-statement>
<copyright-year>2020</copyright-year>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International (CC BY 4.0) License.</license-p>
</license>
</permissions>
<abstract>
<title>ABSTRACT</title>
<p>Lower modulation rates in the temporal envelope (ENV) of the acoustic signal are believed to be the rhythmic backbone in speech, facilitating speech comprehension in terms of neuronal entrainments at &#x03B4;- and &#x03B8;-rates (these rates are comparable to the foot- and syllable-rates phonetically). The jaw plays the role of a carrier articulator regulating mouth opening in a quasi-cyclical way, which correspond to the low-frequency modulations as a physical consequence. This paper describes a method to examine the joint roles of jaw oscillation and ENV in realizing speech rhythm using spectral coherence. Relative powers in the frequency bands corresponding to the &#x03B4;-and &#x03B8;-oscillations in the coherence (respectively notated as %&#x03B4; and %&#x03B8;) were quantified as one possible way of revealing the amount of concomitant foot- and syllable-level rhythmicities carried by both acoustic and articulatory domains. Two English corpora (mngu0 and MOCHA-TIMIT) were used for the proof of concept. %&#x03B4; and %&#x03B8; were regressed on utterance duration for an initial analysis. Results showed that the degrees of foot- and syllable-sized rhythmicities are different and are contingent upon the utterance length.</p>
</abstract>
<trans-abstract xml:lang="es">
<title>RESUMEN</title>
<p>Se piensa que las frecuencias de modulaci&#x00F3;n m&#x00E1;s bajas en la envolvente temporal (ENV) de la se&#x00F1;al ac&#x00FA;stica constituyen la columna vertebral r&#x00ED;tmica del habla, facilitando su comprensi&#x00F3;n a nivel de enlaces neuronales en t&#x00E9;rminos de los rangos &#x03B4; y &#x03B8; (estos rangos son comparables fon&#x00E9;ticamente a los rangos de pie m&#x00E9;trico y sil&#x00E1;bicos). La mand&#x00ED;bula funciona como un articulador que regula la abertura de la boca de una manera cuasi c&#x00ED;clica, lo que se corresponde, como una consecuencia f&#x00ED;sica, con las modulaciones de baja frecuencia. Este art&#x00ED;culo describe un m&#x00E9;todo para examinar el papel conjunto de la oscilaci&#x00F3;n de la mand&#x00ED;bula y de la envolvente ENV en la producci&#x00F3;n del ritmo del habla utilizando la coherencia espectral. Las potencias relativas en las bandas de frecuencia correspondientes a las oscilaciones &#x03B4; y &#x03B8; en la coherencia (indicadas respectivamente como %&#x03B4; y %&#x03B8;) se cuantificaron como un posible modo de revelar la cantidad de ritmicidad concomitante a nivel de pie m&#x00E9;trico y de s&#x00ED;laba que los dominios ac&#x00FA;sticos y articulatorios comportan. Para someter a prueba esta idea, en este estudio se analizaron dos corpus en ingl&#x00E9;s (mngu0 y MOCHA-TIMIT). Para un primer an&#x00E1;lisis, se realiz&#x00F3; una regresi&#x00F3;n de %&#x03B4; y %&#x03B8; en funci&#x00F3;n de la duraci&#x00F3;n del enunciado. Los resultados mostraron que los grados de ritmicidad del pie y de la s&#x00ED;laba son diferentes y dependen de la longitud del enunciado.</p>
</trans-abstract>
<kwd-group xml:lang="en">
<kwd>speech rhythm</kwd>
<kwd>spectral coherence</kwd>
<kwd>temporal envelope</kwd>
<kwd>jaw displacement</kwd>
</kwd-group>
<kwd-group xml:lang="es">
<kwd>ritmo del habla</kwd>
<kwd>coherencia espectral</kwd>
<kwd>envolvente temporal</kwd>
<kwd>desplazamiento de la mand&#x00ED;bula</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="sec1" sec-type="intro">
<title>1. INTRODUCTION</title>
<p>This paper characterizes speech rhythm in terms of the spectral coherence between jaw oscillations and speech temporal envelopes (ENV, henceforth). Two frequency bands in the coherence spectrum covering the neuronal &#x03B4;- and &#x03B8;-rates were particularly analyzed in terms of their relative contributions to the entire coherence power. These bands have been claimed to correspond to the foot- and syllable-timescales in speech and have been demonstrated to play a crucial role in neurological speech processing via brainwave-to-ENV entrainment (e.g. Doelling, Arnal, Ghitza, &#x0026; Poeppel, <xref ref-type="bibr" rid="cit0013">2014</xref>; Ghitza, <xref ref-type="bibr" rid="cit0019">2017</xref>; Poeppel &#x0026; Assaneo, <xref ref-type="bibr" rid="cit0039">2020</xref>). This paper reports an initial analysis on the relationships between relative powers of the &#x03B4;- and &#x03B8;-bands in their coherence and utterance length using two English corpora: mngu0 (Richmond, Hoole, &#x0026; King, <xref ref-type="bibr" rid="cit0042">2011</xref>) and MOCHA-TIMIT (Wrench, <xref ref-type="bibr" rid="cit0049">1999</xref>).</p>
<p>Historically, phoneticians described the rhythm of world languages in terms of intuitive isochronous units: stress-timed vs. syllable-timed rhythm<sup><xref ref-type="fn" rid="fn0001">1</xref></sup>
 (or metaphorically, Morse code vs. machine gun rhythm<sup><xref ref-type="fn" rid="fn0002">2</xref></sup>) (e.g. Abercrombie, <xref ref-type="bibr" rid="cit0001">1967</xref>; Jones, <xref ref-type="bibr" rid="cit0026">1922</xref>; Lloyd James, <xref ref-type="bibr" rid="cit0031">1940</xref>; Pike, <xref ref-type="bibr" rid="cit0038">1945</xref>). Failed attempts to corroborate strict isochrony instrumentally (e.g. Bertra&#x0301;n, <xref ref-type="bibr" rid="cit0004">1999</xref>; Dauer, <xref ref-type="bibr" rid="cit0010">1983</xref>; Roach, <xref ref-type="bibr" rid="cit0043">1982</xref>; Wenk &#x0026; Wioland, <xref ref-type="bibr" rid="cit0048">1982</xref>) motivated researchers to search for acoustic correlates of different rhythmicities with regard to durational variability of different phonetic intervals, i.e. the rhythm metrics (e.g. Dellwo, <xref ref-type="bibr" rid="cit0011">2006</xref>, <xref ref-type="bibr" rid="cit0012">2009</xref>; Grabe &#x0026; Low <xref ref-type="bibr" rid="cit0021">2002</xref>; Ramus, Nespor, &#x0026; Mehler, <xref ref-type="bibr" rid="cit0040">1999</xref>). Meanwhile, alternative approaches to speech rhythm have also been postulated focusing on different yet interrelated physical properties in the signal: (i) prominence (e.g. Cichocki, Selouani, &#x0026; Perreault, <xref ref-type="bibr" rid="cit0007">2014</xref>; Fuchs, <xref ref-type="bibr" rid="cit0017">2016</xref>; He, <xref ref-type="bibr" rid="cit0022">2012</xref>, <xref ref-type="bibr" rid="cit0023">2018</xref>; He &#x0026; Dellwo, <xref ref-type="bibr" rid="cit0024">2016</xref>; Lee &#x0026; Todd, <xref ref-type="bibr" rid="cit0028">2004</xref>), (ii) phase or power analyses of the modulation envelope (e.g. Lancia, Krasovitsky, &#x0026; Stuntebeck, <xref ref-type="bibr" rid="cit0027">2019</xref>; Leong, Stone, Turner, &#x0026; Goswami, <xref ref-type="bibr" rid="cit0029">2014</xref>; Tilsen &#x0026; Arvaniti, <xref ref-type="bibr" rid="cit0046">2013</xref>; Tilsen &#x0026; Johnson, <xref ref-type="bibr" rid="cit0047">2008</xref>) and (iii) coupling strength between feet and syllables (e.g. Barbosa, <xref ref-type="bibr" rid="cit0002">2002</xref>; Cummins &#x0026; Port, <xref ref-type="bibr" rid="cit0009">1998</xref>; Eriksson, <xref ref-type="bibr" rid="cit0016">1991</xref>; O&#x2019;Dell &#x0026; Nieminen, <xref ref-type="bibr" rid="cit0036">1999</xref>).<sup><xref ref-type="fn" rid="fn0003">3</xref></sup> In terms of phonological theorization, the metrical grid can be constructed based on intuitive assessment of prominent values, exhibiting the rhythmic skeleton of an utterance (e.g. Liberman &#x0026; Prince, <xref ref-type="bibr" rid="cit0030">1977</xref>; Nespor &#x0026; Vogel, <xref ref-type="bibr" rid="cit0035">1986</xref>; Selkirk, <xref ref-type="bibr" rid="cit0044">1980</xref>).</p>
<p>How did rhythm evolve in speech? From a Darwinian perspective, MacNeilage (<xref ref-type="bibr" rid="cit0033">1998</xref>) held that the rhythmicity in speech evolved from pre-existing cyclical jaw movements in ancestral primates. These movements were found to be important visuofacial gestures in extant non-human primate communications (Ghazanfar, Chandrasekaran, &#x0026; Morrill, <xref ref-type="bibr" rid="cit0018">2010</xref>). It is believed that the coupling between jaw cycles and vocalization arose in the course of human evolution: the sonority of speech typically waxes and wanes with mouth opening and closing gestures (Ghazanfar et al., <xref ref-type="bibr" rid="cit0018">2010</xref>; MacNeilage, <xref ref-type="bibr" rid="cit0033">1998</xref>; Morrill, Paukner, Ferrari, &#x0026; Ghazanfar, <xref ref-type="bibr" rid="cit0034">2012</xref>). Such opening-closing alternations are temporally organized into syllable-sized units corresponding to the ENV modulations, which constitute the rhythmic &#x201C;frames&#x201D;; the open and closed phases are filled with vocalic and consonantal &#x201C;contents&#x201D; &#x2014; the frame/content theory of speech evolution (MacNeilage, <xref ref-type="bibr" rid="cit0033">1998</xref>). By calculating ENV spectra (e.g. Tilsen &#x0026; Johnson <xref ref-type="bibr" rid="cit0047">2008</xref>) or the syllable intensity variability (e.g. He, <xref ref-type="bibr" rid="cit0023">2018</xref>), the characteristics related to the rhythmic &#x201C;frames&#x201D; can be revealed; by calculating the durational variability of vocalic and consonantal intervals (e.g. Grabe &#x0026; Low, <xref ref-type="bibr" rid="cit0021">2002</xref>; Ramus et al., <xref ref-type="bibr" rid="cit0040">1999</xref>), the characteristics related to the rhythmic &#x201C;contents&#x201D; may be evaluated.</p>
<p>Speech rhythm is not evolutionarily redundant; it is functional in the neurological processing of the speech signal. The recurring oscillations in the ENV &#x2013; which supposedly reflect the rhythmic frames &#x2013; facilitate the brain to parse the incoming speech signal for comprehension. It has been demonstrated that the &#x03B4;-oscillation (.5&#x2013;3 Hz, corresponding to foot/stress rates) and &#x03B8;-oscillation (3&#x2013;9 Hz, corresponding to syllable rates) in the auditory cortex entrain to the speech ENV at these modulation rates (Doelling et al., <xref ref-type="bibr" rid="cit0013">2014</xref>; Ghitza, <xref ref-type="bibr" rid="cit0019">2017</xref>; Giraud &#x0026; Poeppel, <xref ref-type="bibr" rid="cit0020">2012</xref>; Strau&#x00DF; &#x0026; Schwartz, <xref ref-type="bibr" rid="cit0045">2017</xref>). These slow neuronal oscillations formulated a temporal window structure whereby the auditory cortex tracks the speech signal at the foot and syllable rates. Within such longer temporal windows, information encoded in finer timescales (e.g. phonemes up to ~40 Hz, corresponding to the &#x03B3;-oscillation) can then be processed to achieve comprehension (Doelling et al., <xref ref-type="bibr" rid="cit0013">2014</xref>; Giraud &#x0026; Poeppel, <xref ref-type="bibr" rid="cit0020">2012</xref>).</p>
<p>The motor knowledge of speech production is arguably indispensable in the neurological processing of speech signals (Strau&#x00DF; &#x0026; Schwartz, <xref ref-type="bibr" rid="cit0045">2017</xref>). The jaw performs the role of a carrier articulator responsible for lower modulation frequencies that may correspond to the rhythmic frames (Strau&#x00DF; &#x0026; Schwartz, <xref ref-type="bibr" rid="cit0045">2017</xref>) to which the slower neuronal oscillations can be phase-locked, not only in the auditory cortex, but also in the visual cortex (Park, Kayser, Thut, &#x0026; Gross, <xref ref-type="bibr" rid="cit0037">2016</xref>). Seeing the speaker&#x2019;s mouth movements facilitates the listener to understand speech, particularly in adverse conditions with excessive noise (Park et al., <xref ref-type="bibr" rid="cit0037">2016</xref>
<sup><xref ref-type="fn" rid="fn0004">4</xref></sup>). The mouth movements help the listener visually access the rhythmic structure, like visual scaffolding. Therefore, the jaw as a carrier articulator plays an important role in both production and perception of speech rhythm; the temporal windows facilitating the neuronal entrainment to the speech ENV must be discoverable in the jaw oscillation as well. However, the roles of the jaw and ENV have been disjointedly studied: The jaw displacement has been shown to well explain the metrical structure of the utterance (Erickson, Suemitsu, Shibuya, &#x0026; Tiede, <xref ref-type="bibr" rid="cit0015">2012</xref>; Erickson &#x0026; Kawahara, <xref ref-type="bibr" rid="cit0014">2016</xref>; Huang &#x0026; Erickson, <xref ref-type="bibr" rid="cit0025">2019</xref>). The ENV has been extensively investigated in terms of its recurring patterns (He, <xref ref-type="bibr" rid="cit0023">2018</xref>; Tilsen &#x0026; Arvaniti, <xref ref-type="bibr" rid="cit0046">2013</xref>; Tilsen &#x0026; Johnson, <xref ref-type="bibr" rid="cit0047">2008</xref>) and synchronizations between different modulation rates (Cummins &#x0026; Port, <xref ref-type="bibr" rid="cit0009">1998</xref>; Lancia et al., <xref ref-type="bibr" rid="cit0027">2019</xref>; Leong et al., <xref ref-type="bibr" rid="cit0029">2014</xref>).</p>
<p>We thus propose to characterize speech rhythm using the spectral coherence between the jaw oscillation and speech ENV (hereinafter, <sc>jaw-env</sc> coherence). A spectral coherence is a Fourier transform-based representation that quantifies common periodicities in two signals. It evaluates the correlation between these two signals in the frequency domain, hence its advantage over assessing simple correlations in the time domain.<sup><xref ref-type="fn" rid="fn0005">5</xref></sup> A similar approach has been attempted, though, by calculating the coherence between the ENV and mouth opening size in terms of the number of pixels shrouded by the lip contour or the inter-lip distance (Chandrasekaran, Trubanova, Stillittano, Caplier, &#x0026; Ghazanfar, <xref ref-type="bibr" rid="cit0006">2009</xref>); however, the roles of the jaw elevation and depression and peripheral lip gestures could not be disentangled thereof (also see footnote 4). This study, instead, examined the sole role of the jaw movement.</p>
<p>Since the lower frequency components pertaining to the &#x03B4;- and &#x03B8;-oscillations are crucial for the neurological speech processing in the auditory, visual and motor cortices, the jaw oscillation and the ENV should be coherent in these frequency ranges. The degree of such coherence is measurable in terms of the percentage of the spectral integral bounded by the &#x03B4;- or &#x03B8;-band cutoffs out of the entire spectral integral of jaw-env coherence (notated as %&#x03B4; and %&#x03B8;, see Eq. (1) in &#x00A7;2.3). These two measures capture the relative amount of power shared by the jaw oscillation and ENV in terms of regularities at the frequency bands corresponding to the neuronal &#x03B4;- and &#x03B8;-samplings. Moreover, %&#x03B4; and %&#x03B8; are analyzed as a function of the utterance length (&#x00A7;3), because the rhythmic structure is more likely to evolve into a more complex pattern over time (a 5-sec utterance would intuitively have a more complex rhythmic structure than a 1-sec utterance &#x201C;Hello&#x0021;&#x201D; which contains a single iamb). It is expected that higher %&#x03B4; is associated with longer utterances, because more sizeable prosodic boundaries (including foot-sized timescales) may be included; for an utterance with higher %&#x03B4;, a smaller %&#x03B8; is expected because the total power of jaw-env coherence is fixed, and determined by the joint temporal amplitudes of both jaw oscillation and ENV (in reference to Parseval&#x2019;s theorem of energy conservation).</p>
</sec>
<sec id="sec2" sec-type="method">
<title>2. METHOD</title>
<sec id="sec2.1">
<title>2.1 The corpora</title>
<p>The mngu0 (Richmond et al., <xref ref-type="bibr" rid="cit0042">2011</xref>) contains one male English speaker producing over 1,000 utterances, amongst which 594 in the duration range of [2, 8] sec were chosen for the present study. The 2-sec cutoff allowed at least one cycle of the lowest &#x03B4; frequency (.5 Hz) to be included; the 8-sec cutoff excluded sentences with medial pauses. The MOCHA-TIMIT (Wrench, <xref ref-type="bibr" rid="cit0049">1999</xref>) contains three English speakers (1f, coded as &#x201C;fsew0&#x201D;; 2m, coded as &#x201C;maps0&#x201D; and &#x201C;msak0&#x201D;) producing the same set of 460 sentences. Altogether 5 sentences shorter than 2 sec were excluded. All utterances were shorter than 6 sec. For both corpora, the electromagnetic articulograph (Carstens AG500 for mngu0 and AG100 for MOCHA-TIMIT) was used to record the kinematic trajectories of various articulators (with 200 Hz or 500 Hz temporal resolutions) together with the audio speech signal (16-bit @ 16 kHz). All kinematic data were head-corrected and translated to a new Cartesian coordinate system in the midsagittal plane. Sensor histories data from the lower incisor were used for the jaw movements for the study.</p>
</sec>
<sec id="sec2.2">
<title>2.2. Calculating JAW-ENV coherence</title>
<p>Jaw-env coherences were calculated following three steps using Matlab<sup>&#x24C7;</sup> R2018b:</p>
<list list-type="roman-lower">
<list-item><p>Obtaining the spectra of the jaw oscillation functions (the matrix <bold>FFT</bold><sub>JAW</sub>) for each utterance. First, the jaw oscillation time series were estimated as the Euclidean distances of the lower incisor coordinates to zero (the vector <bold>d</bold><sub>JAW</sub>). To obtain <bold>FFT</bold><sub>JAW</sub>, a 512-point fast Fourier transform was applied to <bold>d</bold><sub>JAW</sub> which had been offset-removed, down-sampled to 80 Hz, cosine-tapered (&#x03B1; = .1), and zero-padded. The magnitude of <bold>FFT</bold><sub>JAW</sub> was then linearly normalized in 1 arbitrary unit (arb&#x2019;U, henceforth).</p></list-item>
<list-item><p>Obtaining the spectrum of the speech ENV (the matrix <bold>FFT</bold><sub>ENV</sub>). First, a &#x201C;beat&#x201D; detection filter (Cummins &#x0026; Port, <xref ref-type="bibr" rid="cit0009">1998</xref>; Tilsen &#x0026; Johnson, <xref ref-type="bibr" rid="cit0047">2008</xref>) was applied to the speech signal (first-order Butterworth, center frequency = 1,000 Hz, bandwidth = 300 Hz) to keep the vocalic energy while removing the glottal energy and obstruent noise. Then, the filtered signal was full-wave rectified and further bandpass filtered (fourth-order Butterworth, center frequency = 5 Hz, bandwidth = 10 Hz) to obtain the ENV. To obtain <bold>FFT</bold><sub>ENV</sub>, the ENV was offset-removed, down-sampled to 80 Hz, cosine-tapered (&#x03B1; = .1), zero-padded, and supplied to a 512-point fast Fourier transform. The magnitude of <bold>FFT</bold><sub>ENV</sub> was then linearly normalized in 1 arb&#x2019;U. The object obtained this way is called the beat histogram in music information retrieval (Lykartsis &#x0026; Lerch, <xref ref-type="bibr" rid="cit0032">2015</xref>).</p></list-item>
<list-item><p>The jaw-env coherence (the matrix <bold>COH</bold><sub>JAW-ENV</sub>) was calculated as the Hermitian inner product<sup><xref ref-type="fn" rid="fn0006">6</xref></sup> of the Fourier coefficients in <bold>FFT</bold><sub>JAW</sub> and <bold>FFT</bold><sub>ENV</sub> normalized to the individual power of <bold>FFT</bold><sub>JAW</sub> and <bold>FFT</bold><sub>ENV</sub> (a code snippet in Cohen, <xref ref-type="bibr" rid="cit0008">2017</xref> was applied); negative frequencies were neglected. <xref ref-type="fig" rid="f0001">Figure 1</xref> shows an example of calculating the jaw-env coherence from the spectra of jaw oscillation and the speech ENV. This process computes the common periodicities in two signals by evaluating the correlation between these two signals in the frequency domain.</p></list-item>
</list>
</sec>
<sec id="sec2.3">
<title>2.3. Calculating %&#x03B4; and %&#x03B8; in jaw-env coherence</title>
<p>Eq. (1) illustrates the conceptual calculations of %&#x03B4; and %&#x03B8; &#x2014; the percentage of the spectral integral bounded by the &#x03B4;-band cutoffs (f<sub>1</sub> = .5 Hz, f<sub>2</sub> = 3 Hz) or &#x03B8;-band cutoffs (f<sub>1</sub> = 3 Hz, f<sub>2</sub> = 9 Hz) over the entire spectral integral of the coherence function <italic>C</italic>(f) (f<sub>Nyq</sub> = 40 Hz). The Nyquist frequency (f<sub>Nyq</sub>) of 40 Hz was arbitrarily chosen at the upper &#x03B3;-band boundary responsible for processing phonemes and smaller features. Empirically, the frequency granularity (<italic>d</italic>f) is equal to 2 &#x00D7; f<sub>Nyq</sub> (40 Hz) &#x00F7; FFT points (512) = .16 Hz. Because of the frequency discretization, the coherence function <italic>C</italic>(f) is effectively the matrix <bold>COH</bold><sub>JAW-ENV</sub>. The integrals (approximated using Riemann sums) can be calculated through iterations at the step size of <italic>d</italic>f in <bold>COH</bold><sub>JAW-ENV.</sub></p>
<fig id="f0001">
<label>Figura 1</label>
<caption><p>The spectra of the jaw oscillation and the speech ENV (a); the JAW-ENV coherence calculated from the spectra of the jaw oscillation and the speech ENV (b).</p></caption>
<graphic xlink:href="loquens_074-g001.tif" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</fig>
<disp-formula id="eq1"><alternatives><mml:math id="M1"><mml:mrow><mml:mtext>&#x0025;</mml:mtext><mml:mo>&#x03B4;</mml:mo><mml:mtext>&#x2009;&#x2009;or&#x2009;&#x2009;&#x0025;</mml:mtext><mml:mo>&#x03B8;</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle='true'><mml:mrow><mml:msubsup><mml:mo>&#x222B;</mml:mo><mml:mrow><mml:msub><mml:mtext>f</mml:mtext><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mtext>f</mml:mtext><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:msubsup><mml:mrow><mml:mi>C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mtext>f</mml:mtext><mml:mo>)</mml:mo></mml:mrow><mml:mtext>df</mml:mtext></mml:mrow></mml:mrow></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle displaystyle='true'><mml:mrow><mml:msubsup><mml:mo>&#x222B;</mml:mo><mml:mn>0</mml:mn><mml:mrow><mml:msub><mml:mtext>f</mml:mtext><mml:mrow><mml:mtext>Nyq</mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:msubsup><mml:mrow><mml:mi>C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mtext>f</mml:mtext><mml:mo>)</mml:mo></mml:mrow><mml:mtext>df</mml:mtext></mml:mrow></mml:mrow></mml:mstyle></mml:mrow></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mn>100</mml:mn></mml:mrow></mml:math><graphic xlink:href="loquens_074-e001.tif" xmlns:xlink="http://www.w3.org/1999/xlink"/></alternatives></disp-formula></sec>
</sec>
<sec id="sec3">
<title>3. DATA ANALYSES AND RESULTS<sup><xref ref-type="fn" rid="fn0007">7</xref></sup></title>
<p>For the mngu0 data, simple linear regressions between utterance length and %&#x03B4; and %&#x03B8; were performed using R. The utterance duration was right skewed, hence was natural log transformed. <xref ref-type="table" rid="t0001">Table 1</xref> and <xref ref-type="fig" rid="f0002">Figure 2</xref> illustrate the results: %&#x03B4; increased as utterance duration increased, whereas %&#x03B8; decreased as utterance length increased, conforming to the expectation.</p>
<table-wrap id="t0001">
<label>Table 1</label>
<caption><p>Results of linear regression analyses for the mngu0 data.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" rowspan="2" align="left">Model (Y~X)</th>
<th colspan="2" align="center">F-test of overall significance<hr/></th>
<th colspan="3" align="center">t-test of estimated slope<hr/></th>
</tr>
<tr>
<th align="center">F(DoFs)</th>
<th align="center">p</th>
<th align="center">&#x03B2;</th>
<th align="center">99% CI</th>
<th align="center">|t|</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">%&#x03B4; ~ ln(utterance duration)</td>
<td align="center">881.5(1,592)</td>
<td align="center">&#x226A; .01</td>
<td align="center">30.66</td>
<td align="center">28.00, 33.32</td>
<td align="center">&#x003E; 2.576</td>
</tr>
<tr>
<td align="left">%&#x03B8; ~ ln(utterance duration)</td>
<td align="center">545.5(1,592)</td>
<td align="center">&#x226A; .01</td>
<td align="center">&#x2013;24.44</td>
<td align="center">&#x2013;27.14, &#x2013;21.75</td>
<td align="center">&#x003E; 2.576</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="f0002">
<label>Figura 2</label>
<caption><p>Regression lines and the 99% confidence intervals (shaded areas) superimposed over the scatterplots showing the relationships between %&#x03B4; and log utterance duration (a), and %&#x03B8; and log utterance duration (b) in the mngu0 corpus. Log durationvis-&#x00E0;-vis linear duration at abscissa tick marks in both subplots: .75 ln(sec) &#x21CC; 2.12 sec, 1.0 ln(sec) &#x21CC; 2.72 sec, 1.25 ln(sec) &#x21CC; 3.49 sec, 1.5 ln(sec) &#x21CC; 4.48 sec, 1.75 ln(sec) &#x21CC; 5.75 sec, and 2.0 ln(sec) &#x21CC; 7.39 sec.</p></caption>
<graphic xlink:href="loquens_074-g002.tif" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</fig>
<p>The MOCHA-TIMIT data were subsequently analyzed to examine whether consistent results would be obtained. Random-slope models were fitted by maximum likelihood (response variables: %&#x03B4; and %&#x03B8;; random effects: speaker and utterance; fixed effect: utterance length) using R{<italic>lme4</italic>, v1.1&#x2013;21} (Bates, M&#x00E4;chler, Bolker, &#x0026; Walker, <xref ref-type="bibr" rid="cit0003">2015</xref>). The significance of the slope estimate and between-speaker variability were tested in particular (see <xref ref-type="table" rid="t0002">Table 2</xref> and <xref ref-type="fig" rid="f0003">Figure 3</xref>): in general, a positive slope estimate was found significant between %&#x03B4; and utterance length, and a negative slope estimate was found significant between %&#x03B8; and utterance length. Moreover, individual differences were significant at the same time.</p>
<table-wrap id="t0002">
<label>Table 2</label>
<caption><p>Results of random-slope models for the MOCHA-TIMIT data.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" rowspan="2" align="left">Response variable</th>
<th colspan="3" align="center">Fixed effect: utterance length<hr/></th>
<th colspan="3" align="center">Random effect: speaker <xref ref-type="table-fn" rid="tf2-1"><sup>a</sup></xref><hr/></th>
<th align="center" rowspan="2">p</th>
</tr>
<tr>
<th align="center">&#x03B2;</th>
<th align="center">99% CI</th>
<th align="center">|t|</th>
<th align="center">AIC (full; reduced)</th>
<th align="center">&#x2013;LogLik (full; reduced)</th>
<th align="center">&#x03C7;2(DoF)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">%&#x03B4;</td>
<td align="center">8.38</td>
<td align="center">5.49, 11.27</td>
<td align="center">&#x003E; 2.576</td>
<td align="center">10121; 10510</td>
<td align="center">5248.8; 5051.5</td>
<td align="center">394.47(3)</td>
<td align="center">&#x226A; .01</td>
</tr>
<tr>
<td align="left">%&#x03B8;</td>
<td align="center">&#x2013;2.89</td>
<td align="center">&#x2013;5.16, &#x2013;.62</td>
<td align="center">&#x003E; 2.576</td>
<td align="center">10038; 10363</td>
<td align="center">5009.7; 5175.7</td>
<td align="center">331.92(3)</td>
<td align="center">&#x226A; .01</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="tf2-1"><label>a</label><p>Likelihood ratio test was used to test between-speaker variability between the full model and the speaker-reduced model.</p></fn>
<fn><p>The AICs of the full models were smaller than those of the reduced models, suggesting that the full models had better fits. The &#x03C7;2 values were calculated as the differences between twice the &#x2013;LogLik of the full and reduced models (the differences of the deviances).</p></fn>
</table-wrap-foot>
</table-wrap>
<fig id="f0003">
<label>Figura 3</label>
<caption><p>Regression lines and the 99% confidence intervals (shaded areas) superimposed over the scatterplots showing the relationships between %&#x03B4; and utterance duration (in sec) (a), and %&#x03B8; and utterance duration (b) for each of the three speakers in the MOCHA-TIMIT corpus.</p></caption>
<graphic xlink:href="loquens_074-g003.tif" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</fig>
</sec>
<sec id="sec4" sec-type="discussion">
<title>4. DISCUSSION</title>
<p>This paper introduced a method to characterize speech rhythm using spectral coherence between jaw oscillation and the speech ENV, i.e. the jaw-env coherence. It provides a spectro-temporal representation of the common periodicities in both signals. Two frequency bands corresponding to the brain &#x03B4;- and &#x03B8;-oscillations were analyzed in terms of the percentage of power accounted for by these two bands in jaw-env coherence, i.e. %&#x03B4; and %&#x03B8;. In general, utterance length was found to be a significant predictor of %&#x03B4; and %&#x03B8;, yet individual differences must not be neglected. The findings have several implications:</p>
<list list-type="roman-lower">
<list-item><p>The jaw oscillation and speech ENV possess strong spectral coherence in the low frequency bands of .5 &#x2013; 3 Hz and 3 &#x2013; 9 Hz. This upholds the role of jaw movement and speech ENV in speech rhythmicity. The semi-cyclical jaw movements constantly change the amount of radiated energy corresponding to the lower modulation frequencies in the speech signal, to which the auditory cortex of the listener entrains at the &#x03B4;- and &#x03B8;-rates. (Doelling et al., <xref ref-type="bibr" rid="cit0013">2014</xref>; Ghitza, <xref ref-type="bibr" rid="cit0019">2017</xref>; Giraud &#x0026; Poeppel, <xref ref-type="bibr" rid="cit0020">2012</xref>; Strau&#x00DF; &#x0026; Schwartz, <xref ref-type="bibr" rid="cit0045">2017</xref>). The jaw movements also invite neuronal entrainment in the listener&#x2019;s visual cortex (Park et al., <xref ref-type="bibr" rid="cit0037">2016</xref>). These entrainments play a useful role in speech processing and comprehension.</p></list-item>
<list-item><p>Both .5 &#x2013; 3 Hz and 3 &#x2013; 9 Hz bands (pertaining to the &#x03B4;- and &#x03B8;-rates) are represented in jaw-env coherence, but in different degrees as measured by %&#x03B4; and %&#x03B8;. This suggests that different levels of rhythmicities (including foot-sized and syllable-sized) are present simultaneously but differ in degrees. The amount of regularities at a larger timescale increases as the utterance length increases for all speakers from the two corpora (<xref ref-type="fig" rid="f0002">Figures 2</xref> and <xref ref-type="fig" rid="f0003">3</xref>). It is possible that longer utterances are more likely to contain larger prosodic boundaries or more extreme intonational accents, which would increase the power pertaining to the &#x03B4;-band. This may have a functional advantage: higher &#x03B4;-rate regularities would facilitate sensory chunking of a longer utterance under the neuronal &#x03B4;-sampling (an example of sensory chunking is using temporal groupings when memorizing a series of digits or syllables) (Boucher, Gilbert, &#x0026; Jemel, <xref ref-type="bibr" rid="cit0005">2019</xref>). Smaller units pertaining to faster rates (e.g., syllables, phonemes or even phonological features) could be processed within each &#x03B4;-window.</p></list-item>
<list-item><p>Individual differences are conspicuous in %&#x03B4; and %&#x03B8; as a function of utterance duration. The amounts of regularities in &#x03B4;- and &#x03B8;-bands are inversely proportional for the mngu0 speaker as well as speaker &#x201C;fsew&#x201D; in MOCHA-TIMIT (<xref ref-type="fig" rid="f0002">Figures 2</xref> and <xref ref-type="fig" rid="f0003">3</xref>), possibly because more &#x03B4; power has already taken up the majority of power in jaw-env coherence in longer utterances, leaving little power for the &#x03B8;-band regularity. For speaker &#x201C;msak&#x201D; in MOCHA-TIMIT, &#x03B4;-band regularities were already prominent even in short sentences (high intercept of &#x201C;msak&#x201D; in <xref ref-type="fig" rid="f0003">Figure 3a</xref>), leaving little power for syllable-sized frequencies regardless of the utterance length (low intercept and flat slope of &#x201C;msak&#x201D; in <xref ref-type="fig" rid="f0003">Figure 3b</xref>). Nevertheless, to investigate individual differences fully, it is mandatory to increase the sample size significantly.</p></list-item>
<list-item><p>The results may also explain why early phoneticians (e.g. Abercrombie, <xref ref-type="bibr" rid="cit0001">1967</xref>; Jones, <xref ref-type="bibr" rid="cit0026">1922</xref>; Lloyd James, <xref ref-type="bibr" rid="cit0031">1940</xref>; Pike, <xref ref-type="bibr" rid="cit0038">1945</xref>), despite having undergone rigorous ear training, would still inaccurately describe languages such as English as possessing isochronous feet. Higher %&#x03B4; may be a strong cue to foot-sized regularity in both jaw oscillation and speech temporal modulation. For all speakers analyzed in the study, a large amount of foot-sized regularity has been found. It is likely that early phoneticians have discerned such foot-sized regularity in English, yet unfortunately described it in absolute terms as &#x201C;stress-timed.&#x201D;</p></list-item>
</list>
<p>This study has limitations too:</p>
<list list-type="roman-lower">
<list-item><p>In terms of data variance, all speakers in the MOCHA-TIMIT corpus showed bigger variances than the mngu0 speaker (cf. <xref ref-type="fig" rid="f0002">Figures 2</xref> and <xref ref-type="fig" rid="f0003">3</xref>). This may be due to the data inconsistency issue of the MOCHA-TIMIT corpus. It has been demonstrated that even for a relatively stationary sensor at the velum, a tremendous amount of data inconsistency existed (Richmond, <xref ref-type="bibr" rid="cit0041">2009</xref>; Richmond et al., <xref ref-type="bibr" rid="cit0042">2011</xref>). Technical issues with respect to the early generation of the electromagnetic articulograph may be the culprit (Richmond, <xref ref-type="bibr" rid="cit0041">2009</xref>).</p></list-item>
<list-item><p>The two frequency bands analyzed in this study were informed by the low-neuronal oscillations that have been shown to play a key role in the rhythmic parsing in speech processing. Apart from considering these two bands as pertaining to the stress-rate or syllable-rate, further research still needs to be done to assess whether these frequency cutoffs are justifiable in linguistic/phonetic terms.</p></list-item>
<list-item><p>The corpora adopted in this study were small in terms of the number of speakers, and only English was analyzed. This reduced the generalizability of this study.</p></list-item>
</list>
<p>For future research, it is imperative to test the method using more speakers from different languages, including those traditionally labeled as &#x201C;syllable-timed.&#x201D; That they have been described as &#x201C;syllable-timed&#x201D; may be due to a high degree of syllable-sized cyclicity in jaw oscillations and speech temporal modulations (measurable as high %&#x03B8; in jaw-env coherence) even in longer sentences. So far, the coherence of the jaw oscillation and ENV has been investigated based on the power spectra. It will also be interesting to explore the coherence based on the phase spectra from multi-domain signals, including acoustic, articulatory and neurological, to further explore their temporal relationships in constituting speech rhythmicity both at the production and perception levels.</p>
</sec>
</body>
<back>
<ack>
<title>5. ACKNOWLEDGEMENTS</title>
<p>This study was supported by the Forschungskredit of the University of Zurich (Grant FK-19-069 to YZ and Grant FK-20-078 to LH). It was also benefited from a completed project from the Swiss National Science Foundation (Grant P2ZHP1_178109 to LH). We thank Alejandra Pesantez for her great help in the Spanish abstract.</p>
</ack>
<ref-list>
<title>REFERENCES</title>
<ref id="cit0001">
<mixed-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Abercrombie</surname>
<given-names>D.</given-names>
</name>
</person-group>
<year>1967</year>
<source>Elements of General Phonetics</source>
<publisher-loc>Edinburgh</publisher-loc>
<publisher-name>Edinburgh University Press</publisher-name>
</mixed-citation>
</ref>
<ref id="cit0002">
<mixed-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Barbosa</surname>
<given-names>P. A.</given-names>
</name>
</person-group>
<year>2002</year>
<chapter-title>Explaining cross-linguistic rhythmic variability via a coupled-oscillator model of rhythm production</chapter-title>
<person-group person-group-type="editor">
<name>
<surname>Bel</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Marlien</surname>
<given-names>I.</given-names>
</name>
</person-group>
<source>Proceedings of Speech Prosody 2002</source>
<fpage>163</fpage>
<lpage>166</lpage>
<publisher-loc>Aix-en-Provence, France</publisher-loc>
<publisher-name>Laboratoire Parole et Langage, SProSIG</publisher-name>
</mixed-citation>
</ref>
<ref id="cit0003">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bates</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>M&#x00E4;chler</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Bolker</surname>
<given-names>B. M.</given-names>
</name>
<name>
<surname>Walker</surname>
<given-names>S. C.</given-names>
</name>
</person-group>
<article-title>Fitting linear mixed-effects models using lme4</article-title>
<source>Journal of Statistical Software</source
>
<year>2015</year>
<volume>67</volume>
<issue>1</issue>
<fpage>1</fpage>
<lpage>48</lpage>
</mixed-citation>
</ref>
<ref id="cit0004">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bertra&#x0301;n</surname>
<given-names>A. P.</given-names>
</name>
</person-group>
<article-title>Prosodic typology: On the dichotomy between stress-timed and syllable-timed languages</article-title>
<source>Language Design</source>
<year>1999</year>
<volume>2</volume>
<fpage>103</fpage>
<lpage>131</lpage>
</mixed-citation>
</ref>
<ref id="cit0005">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Boucher</surname>
<given-names>V. J.</given-names>
</name>
<name>
<surname>Gilbert</surname>
<given-names>A. C.</given-names>
</name>
<name>
<surname>Jemel</surname>
<given-names>B.</given-names>
</name>
</person-group>
<article-title>The role of low-frequency neural oscillations in speech processing: Revising delta entrainment</article-title>
<source>Journal of Cognitive Neuroscience</source>
<year>2019</year>
<volume>31</volume>
<issue>8</issue>
<fpage>1205</fpage>
<lpage>1215</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1162/jocn_a_01410">http://dx.doi.org/10.1162/jocn_a_01410</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0006">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chandrasekaran</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Trubanova</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Stillittano</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Caplier</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Ghazanfar</surname>
<given-names>A. A.</given-names>
</name>
</person-group>
<article-title>The natural statistics of audiovisual speech</article-title>
<source>PLoS Computational Biology</source>
<year>2009</year>
<volume>5</volume>
<issue>7</issue>
<elocation-id>e1000436</elocation-id>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1371/journal.pcbi.1000436">http://dx.doi.org/10.1371/journal.pcbi.1000436</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0007">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cichocki</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Selouani</surname>
<given-names>S.-A.</given-names>
</name>
<name>
<surname>Perreault</surname>
<given-names>Y.</given-names>
</name>
</person-group>
<article-title>Measuring rhythm in dialects of New Brunswick French: Is there a role for intensity?</article-title>
<source>Canadian Acoustics &#x00B7; Acoustique Canadienne</source>
<year>2014</year>
<volume>42</volume>
<issue>3</issue>
<fpage>90</fpage>
<lpage>91</lpage>
</mixed-citation>
</ref>
<ref id="cit0008">
<mixed-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Cohen</surname>
<given-names>M. X.</given-names>
</name>
</person-group>
<year>2017</year>
<source>Matlab for Brain and Cognitive Scientists</source>
<publisher-loc>Cambridge, MA</publisher-loc>
<publisher-name>MIT Press</publisher-name>
</mixed-citation>
</ref>
<ref id="cit0009">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cummins</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Port</surname>
<given-names>R.</given-names>
</name>
</person-group>
<article-title>Rhythmic constraints on stress timing in English</article-title>
<source>Journal of Phonetics</source>
<year>1998</year>
<volume>26</volume>
<issue>2</issue>
<fpage>145</fpage>
<lpage>171</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1006/jpho.1998.0070">http://dx.doi.org/10.1006/jpho.1998.0070</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0010">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dauer</surname>
<given-names>R. M.</given-names>
</name>
</person-group>
<article-title>Stress-timing and syllable-timing reanalyzed</article-title>
<source>Journal of Phonetics</source>
<year>1983</year>
<volume>11</volume>
<fpage>51</fpage>
<lpage>62</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1016/S0095-4470(19)30776-4">http://dx.doi.org/10.1016/S0095-4470(19)30776-4</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0011">
<mixed-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Dellwo</surname>
<given-names>V.</given-names>
</name>
</person-group>
<year>2006</year>
<chapter-title>Rhythm and speech rate: A variation coefficient for &#x2206;C</chapter-title>
<person-group person-group-type="editor">
<name>
<surname>Karnowski</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Szigeti</surname>
<given-names>I.</given-names>
</name>
</person-group>
<source>Sprache und Sprachverarbeitung &#x2014; Language and language-processing</source>
<comment>Linguistik International 15</comment>
<fpage>231</fpage>
<lpage>241</lpage>
<publisher-loc>Frankfurt a/M</publisher-loc>
<publisher-name>Peter Lang</publisher-name>
</mixed-citation>
</ref>
<ref id="cit0012">
<mixed-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Dellwo</surname>
<given-names>V.</given-names>
</name>
</person-group>
<year>2009</year>
<chapter-title>Choosing the right rate normalization method for measurements of speech rhythm</chapter-title>
<person-group person-group-type="editor">
<name>
<surname>Schmid</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Schwarzenbach</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Studer</surname>
<given-names>D.</given-names>
</name>
</person-group>
<source>La dimensione temporale del parlato: Atti del 5&#x00B0; Convegno Nazionale AISV 2009</source>
<fpage>13</fpage>
<lpage>32</lpage>
<publisher-loc>Torriana</publisher-loc>
<publisher-name>EDK Editore</publisher-name>
</mixed-citation>
</ref>
<ref id="cit0013">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Doelling</surname>
<given-names>K. B.</given-names>
</name>
<name>
<surname>Arnal</surname>
<given-names>L. H.</given-names>
</name>
<name>
<surname>Ghitza</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Poeppel</surname>
<given-names>D.</given-names>
</name>
</person-group>
<article-title>Acoustic landmarks drive delta-theta oscillations to enable speech comprehension by facilitating perceptual parsing</article-title>
<source>NeuroImage</source>
<year>2014</year>
<volume>85</volume>
<issue>2</issue>
<fpage>761</fpage>
<lpage>768</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1016/j.neuroimage.2013.06.035">http://dx.doi.org/10.1016/j.neuroimage.2013.06.035</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0014">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Erickson</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Kawahara</surname>
<given-names>S.</given-names>
</name>
</person-group>
<article-title>Articulatory correlates of metrical structure: Studying jaw displacement patterns</article-title>
<source>Linguistics Vanguard</source>
<year>2016</year>
<volume>2</volume>
<issue>1</issue>
<elocation-id>20150025</elocation-id>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1515/lingvan-2015-0025">http://dx.doi.org/10.1515/lingvan-2015-0025</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0015">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Erickson</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Suemitsu</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Shibuya</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tiede</surname>
<given-names>M.</given-names>
</name>
</person-group>
<article-title>Metrical structure and production of English rhythm</article-title>
<source>Phonetica</source>
<year>2012</year>
<volume>69</volume>
<issue>3</issue>
<fpage>180</fpage>
<lpage>190</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1159/000342417">http://dx.doi.org/10.1159/000342417</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0016">
<mixed-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Eriksson</surname>
<given-names>A.</given-names>
</name>
</person-group>
<year>1991</year>
<source>Aspects of Swedish Speech Rhythm</source>
<publisher-loc>Gothenburg</publisher-loc>
<comment>(Gothenburg monographs in linguistics 9)</comment>
<publisher-name>University of Gothenburg Dissertation</publisher-name>
</mixed-citation>
</ref>
<ref id="cit0017">
<mixed-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Fuchs</surname>
<given-names>R.</given-names>
</name>
</person-group>
<year>2016</year>
<source>Speech Rhythm in Varieties of English</source>
<publisher-loc>Singapore</publisher-loc>
<publisher-name>Springer</publisher-name>
</mixed-citation>
</ref>
<ref id="cit0018">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ghazanfar</surname>
<given-names>A. A.</given-names>
</name>
<name>
<surname>Chandrasekaran</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Morrill</surname>
<given-names>R. J.</given-names>
</name>
</person-group>
<article-title>Dynamic, rhythmic facial expressions and the superior temporal sulcus of macaque monkeys: Implications for the evolution of audiovisual speech</article-title>
<source>European Journal of Neuroscience</source>
<year>2010</year>
<volume>31</volume>
<issue>10</issue>
<fpage>1807</fpage>
<lpage>1817</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1111/j.1460-9568.2010.07209.x">http://dx.doi.org/10.1111/j.1460-9568.2010.07209.x</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0019">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ghitza</surname>
<given-names>O.</given-names>
</name>
</person-group>
<article-title>Acoustic-driven delta rhythms as prosodic markers</article-title>
<source>Language, Cognition and Neuroscience</source>
<year>2017</year>
<volume>32</volume>
<issue>5</issue>
<fpage>545</fpage>
<lpage>561</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1080/23273798.2016.1232419">http://dx.doi.org/10.1080/23273798.2016.1232419</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0020">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Giraud</surname>
<given-names>A-L.</given-names>
</name>
<name>
<surname>Poeppel</surname>
<given-names>D.</given-names>
</name>
</person-group>
<article-title>Cortical oscillations and speech processing: Emerging computational principles and operations</article-title>
<source>Nature Neuroscience</source>
<year>2012</year>
<volume>15</volume>
<issue>4</issue>
<fpage>511</fpage>
<lpage>517</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1038/nn.3063">http://dx.doi.org/10.1038/nn.3063</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0021">
<mixed-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Grabe</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Low</surname>
<given-names>E. L.</given-names>
</name>
</person-group>
<year>2002</year>
<chapter-title>Durational variability in speech and rhythm class hypothesis</chapter-title>
<person-group person-group-type="editor">
<name>
<surname>Gussenhoven</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Warner</surname>
<given-names>N.</given-names>
</name>
</person-group>
<source>Laboratory Phonology 7</source>
<fpage>515</fpage>
<lpage>543</lpage>
<publisher-loc>Berlin &#x0026; New York</publisher-loc>
<publisher-name>Mouton de Gruyter</publisher-name>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1515/9783110197105.515">http://dx.doi.org/10.1515/9783110197105.515</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0022">
<mixed-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>He</surname>
<given-names>L.</given-names>
</name>
</person-group>
<year>2012</year>
<chapter-title>Syllabic intensity variations as quantification of speech rhythm: Evidence from both L1 and L2</chapter-title>
<person-group person-group-type="editor">
<name>
<surname>Ma</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Ding</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Hirst</surname>
<given-names>D.</given-names>
</name>
</person-group>
<source>Proceedings of Speech Prosody 2012</source>
<fpage>466</fpage>
<lpage>469</lpage>
<publisher-loc>Shanghai, China</publisher-loc>
</mixed-citation>
</ref>
<ref id="cit0023">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>He</surname>
<given-names>L.</given-names>
</name>
</person-group>
<article-title>Development of speech rhythm in first language: The role of syllable intensity variability</article-title>
<source>Journal of the Acoustical Society of America</source>
<year>2018</year>
<volume>143</volume>
<issue>6</issue>
<fpage>EL463</fpage>
<lpage>EL467</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1121/1.5042083">http://dx.doi.org/10.1121/1.5042083</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0024">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>He</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Dellwo</surname>
<given-names>V.</given-names>
</name>
</person-group>
<article-title>The role of syllable intensity in between-speaker rhythmic variability</article-title>
<source>International Journal of Speech, Language and the Law</source>
<year>2016</year>
<volume>23</volume>
<issue>2</issue>
<fpage>243</fpage>
<lpage>273</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1558/ijsll.v23i2.30345">http://dx.doi.org/10.1558/ijsll.v23i2.30345</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0025">
<mixed-citation publication-type="conf-proc">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Erickson</surname>
<given-names>D.</given-names>
</name>
</person-group>
<year>2019</year>
<chapter-title>Articulation of English &#x201C;prominence&#x201D; by L1 (English) and L2 (French) speaker</chapter-title>
<person-group person-group-type="editor">
<name>
<surname>Calhoun</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Escudero</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Tabain</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Warren</surname>
<given-names>P.</given-names>
</name>
</person-group>
<conf-name>Proceedings of the 19th International Congress of Phonetic Sciences (ICPhS-19), paper 134</conf-name>
<conf-loc>Melbourne, Australia</conf-loc>
</mixed-citation>
</ref>
<ref id="cit0026">
<mixed-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Jones</surname>
<given-names>D.</given-names>
</name>
</person-group>
<year>1922</year>
<source>An Outline of English Phonetics</source>
<publisher-loc>New York</publisher-loc>
<publisher-name>G. E. Stechert &#x0026; Co.</publisher-name>
</mixed-citation>
</ref>
<ref id="cit0027">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lancia</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Krasovitsky</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Stuntebeck</surname>
<given-names>F.</given-names>
</name>
</person-group>
<article-title>Coordinative patterns underlying cross-linguistic rhythmic differences</article-title>
<source>Journal of Phonetics</source>
<year>2019</year>
<volume>72</volume>
<fpage>66</fpage>
<lpage>80</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1016/j.wocn.2018.08.004">http://dx.doi.org/10.1016/j.wocn.2018.08.004</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0028">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lee</surname>
<given-names>C. S.</given-names>
</name>
<name>
<surname>Todd</surname>
<given-names>N. P. M.</given-names>
</name>
</person-group>
<article-title>Towards an auditory account of speech rhythm: Application of a model of the auditory &#x201C;primal sketch&#x201D; to two multi-language corpora</article-title>
<source>Cognition</source>
<year>2004</year>
<volume>93</volume>
<issue>3</issue>
<fpage>225</fpage>
<lpage>254</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1016/j.cognition.2003.10.012">http://dx.doi.org/10.1016/j.cognition.2003.10.012</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0029">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Leong</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Stone</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Turner</surname>
<given-names>R. E.</given-names>
</name>
<name>
<surname>Goswami</surname>
<given-names>U.</given-names>
</name>
</person-group>
<article-title>A role for amplitude modulation phase relationships in speech rhythm perception</article-title>
<source>Journal of the Acoustical Society of America</source>
<year>2014</year>
<volume>136</volume>
<issue>1</issue>
<fpage>366</fpage>
<lpage>381</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1121/1.4883366">http://dx.doi.org/10.1121/1.4883366</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0030">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liberman</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Prince</surname>
<given-names>A.</given-names>
</name>
</person-group>
<article-title>On stress and linguistic rhythm</article-title>
<source>Linguistic Inquiry</source>
<year>1977</year>
<volume>8</volume>
<issue>2</issue>
<fpage>249</fpage>
<lpage>336</lpage>
</mixed-citation>
</ref>
<ref id="cit0031">
<mixed-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Lloyd James</surname>
<given-names>A.</given-names>
</name>
</person-group>
<year>1940</year>
<source>Speech Signals in Telephony</source>
<publisher-loc>London</publisher-loc>
<publisher-name>Sir I. Pitman</publisher-name>
</mixed-citation>
</ref>
<ref id="cit0032">
<mixed-citation publication-type="conf-proc">
<person-group person-group-type="author">
<name>
<surname>Lykartsis</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Lerch</surname>
<given-names>A.</given-names>
</name>
</person-group>
<person-group person-group-type="editor">
<name>
<surname>Svensson</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Kristiansen</surname>
<given-names>U.</given-names>
</name>
</person-group>
<chapter-title>Beat histogram features for rhythm-based musical genre classification using multiple novelty functions</chapter-title>
<year>2015</year>
<conf-name>Proceedings of the 18th International Conference on Digital Audio Effects (DAFx), paper 42</conf-name>
<publisher-loc>Trondheim, Norway</publisher-loc>
<publisher-name>Department of Music and Department of Electronics and Telecommunications. Norwegian University of Science and Technology</publisher-name>
</mixed-citation>
</ref>
<ref id="cit0033">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>MacNeilage</surname>
<given-names>P. F.</given-names>
</name>
</person-group>
<article-title>The frame/content theory of evolution of speech production</article-title>
<source>Behavioral and Brain Sciences</source>
<year>1998</year>
<volume>21</volume>
<issue>4</issue>
<fpage>499</fpage>
<lpage>511</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1017/S0140525X98001265">http://dx.doi.org/10.1017/S0140525X98001265</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0034">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Morrill</surname>
<given-names>R. J.</given-names>
</name>
<name>
<surname>Paukner</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Ferrari</surname>
<given-names>P. F.</given-names>
</name>
<name>
<surname>Ghazanfar</surname>
<given-names>A. A.</given-names>
</name>
</person-group>
<article-title>Monkey lipsmacking develops like the human speech rhythm</article-title>
<source>Developmental Science,</source>
<year>2012</year>
<volume>15</volume>
<issue>4</issue>
<fpage>557</fpage>
<lpage>568</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1111/j.1467-7687.2012.01149.x">http://dx.doi.org/10.1111/j.1467-7687.2012.01149.x</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0035">
<mixed-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Nespor</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Vogel</surname>
<given-names>I.</given-names>
</name>
</person-group>
<year>1986</year>
<source>Prosodic Phonology</source>
<publisher-loc>Dordrecht</publisher-loc>
<publisher-name>Foris</publisher-name>
</mixed-citation>
</ref>
<ref id="cit0036">
<mixed-citation publication-type="conf-proc">
<person-group person-group-type="author">
<name>
<surname>O&#x2019;Dell</surname>
<given-names>M. L.</given-names>
</name>
<name>
<surname>Nieminen</surname>
<given-names>T.</given-names>
</name>
</person-group>
<year>1999</year>
<chapter-title>Coupled oscillator model of speech rhythm</chapter-title>
<person-group person-group-type="editor">
<name>
<surname>Ohala</surname>
<given-names>J. J.</given-names>
</name>
<name>
<surname>Hasegawa</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Ohala</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Granville</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Bailey</surname>
<given-names>A. C.</given-names>
</name>
</person-group>
<source>Proceedings of the 14th International Congress of Phonetic Sciences (ICPhS-14)</source>
<fpage>1075</fpage>
<lpage>1078</lpage>
<publisher-loc>San Francisco, California</publisher-loc>
<publisher-name>University of California</publisher-name>
</mixed-citation>
</ref>
<ref id="cit0037">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Park</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Kayser</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Thut</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Gross</surname>
<given-names>J.</given-names>
</name>
</person-group>
<article-title>Lip movements entrain the observers&#x2019; low-frequency brain oscillations to facilitate speech intelligibility</article-title>
<source>eLife</source>
<year>2016</year>
<volume>5</volume>
<elocation-id>e14521</elocation-id>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.7554/eLife.14521">http://dx.doi.org/10.7554/eLife.14521</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0038">
<mixed-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Pike</surname>
<given-names>K. L.</given-names>
</name>
</person-group>
<year>1945</year>
<source>The Intonation of American English</source>
<publisher-loc>Ann Arbor</publisher-loc>
<publisher-name>University of Michigan Press</publisher-name>
</mixed-citation>
</ref>
<ref id="cit0039">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Poeppel</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Assaneo</surname>
<given-names>M. F.</given-names>
</name>
</person-group>
<article-title>Speech rhythm and their neural foundations</article-title>
<source>Nature Reviews Neuroscience</source>
<year>2020</year>
<volume>21</volume>
<issue>6</issue>
<fpage>322</fpage>
<lpage>334</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1038/s41583-020-0304-4">http://dx.doi.org/10.1038/s41583-020-0304-4</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0040">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ramus</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Nespor</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Mehler</surname>
<given-names>J.</given-names>
</name>
</person-group>
<article-title>Correlates of linguistic rhythm in the speech signal</article-title>
<source>Cognition</source>
<year>1999</year>
<volume>73</volume>
<issue>3</issue>
<fpage>265</fpage>
<lpage>292</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1016/S0010-0277(99)00058-X">http://dx.doi.org/10.1016/S0010-0277(99)00058-X</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0041">
<mixed-citation publication-type="conf-proc">
<person-group person-group-type="author">
<name>
<surname>Richmond</surname>
<given-names>K.</given-names>
</name>
</person-group>
<article-title>Preliminary inversion mapping results with a new EMA corpus</article-title>
<year>2009</year>
<conf-name>Proceedings of the 10th Annual Conference of the International Speech Communication Association - INTERSPEECH 2009</conf-name>
<publisher-loc>Brighton, UK</publisher-loc>
<publisher-name>ISCA Archive</publisher-name>
<fpage>2835</fpage>
<lpage>2838</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://www.isca-speech.org/archive/interspeech_2009">http://www.isca-speech.org/archive/interspeech_2009</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0042">
<mixed-citation publication-type="conf-proc">
<person-group person-group-type="author">
<name>
<surname>Richmond</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Hoole</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>King</surname>
<given-names>S.</given-names>
</name>
</person-group>
<person-group person-group-type="editor">
<name>
<surname>Cosi</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>De Mori</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Di Fabbrizio</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Pieraccini</surname>
<given-names>R.</given-names>
</name>
</person-group>
<chapter-title>Announcing the electromagnetic articulography (Day 1) subset of the mngu0 articulatory corpus</chapter-title>
<year>2011</year>
<conf-name>Proceedings of the 12th Annual Conference of the International Speech Communication Association - INTERSPEECH 2011</conf-name>
<publisher-loc>Florence, Italy</publisher-loc>
<publisher-name>ISCA Archive</publisher-name>
<fpage>1505</fpage>
<lpage>1508</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://www.isca-speech.org/archive/interspeech_2011">http://www.isca-speech.org/archive/interspeech_2011</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0043">
<mixed-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Roach</surname>
<given-names>P.</given-names>
</name>
</person-group>
<year>1982</year>
<chapter-title>On the distinction between &#x201C;stress-timed&#x201D; and &#x201C;syllable-timed&#x201D; languages</chapter-title>
<person-group person-group-type="editor">
<name>
<surname>Crystal</surname>
<given-names>D.</given-names>
</name>
</person-group>
<source>Linguistic Controversies: Essays in Linguistic Theory and Practice in Honour of F. R. Palmer</source>
<fpage>73</fpage>
<lpage>79</lpage>
<publisher-loc>London</publisher-loc>
<publisher-name>Edwards Arnold</publisher-name>
</mixed-citation>
</ref>
<ref id="cit0044">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Selkirk</surname>
<given-names>E. O.</given-names>
</name>
</person-group>
<article-title>The role of prosodic categories in English word stress</article-title>
<source>Linguistic Inquiry</source>
<year>1980</year>
<volume>11</volume>
<issue>3</issue>
<fpage>563</fpage>
<lpage>605</lpage>
</mixed-citation>
</ref>
<ref id="cit0045">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Strau&#x00DF;</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Schwartz</surname>
<given-names>J-L.</given-names>
</name>
</person-group>
<article-title>The syllable in the light of motor skills and neural oscillations</article-title>
<source>Language, Cognition and Neuroscience</source>
<year>2017</year>
<volume>32</volume>
<issue>5</issue>
<fpage>562</fpage>
<lpage>569</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1080/23273798.2016.1253852">http://dx.doi.org/10.1080/23273798.2016.1253852</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0046">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tilsen</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Arvaniti</surname>
<given-names>A.</given-names>
</name>
</person-group>
<article-title>Speech rhythm analysis with decomposition of the amplitude envelope: Characterizing rhythmic patterns within and across languages</article-title>
<source>Journal of the Acoustical Society of America</source>
<year>2013</year>
<volume>134</volume>
<issue>1</issue>
<fpage>628</fpage>
<lpage>639</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1121/1.4807565">http://dx.doi.org/10.1121/1.4807565</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0047">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tilsen</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Johnson</surname>
<given-names>K.</given-names>
</name>
</person-group>
<article-title>Low-frequency Fourier analysis of speech rhythm</article-title>
<source>Journal of the Acoustical Society of America</source>
<year>2008</year>
<volume>124</volume>
<issue>2</issue>
<fpage>EL34</fpage>
<lpage>EL39</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1121/1.2947626">http://dx.doi.org/10.1121/1.2947626</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0048">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wenk</surname>
<given-names>B. J.</given-names>
</name>
<name>
<surname>Wioland</surname>
<given-names>F.</given-names>
</name>
</person-group>
<article-title>Is French really syllable-timed?</article-title>
<source>Journal of Phonetics</source>
<year>1982</year>
<volume>10</volume>
<issue>2</issue>
<fpage>193</fpage>
<lpage>216</lpage>
<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1016/S0095-4470(19)30957-X">http://dx.doi.org/10.1016/S0095-4470(19)30957-X</ext-link></comment>
</mixed-citation>
</ref>
<ref id="cit0049">
<mixed-citation publication-type="webpage">
<person-group person-group-type="author">
<name>
<surname>Wrench</surname>
<given-names>A.</given-names>
</name>
</person-group>
<year>1999</year>
<source>MOCHA MultiCHannel Articulatory database: English (MOCHA-TIMIT)</source>
<date-in-citation>accessed 25 December 2018</date-in-citation>
<comment><ext-link ext-link-type="uri" xlink:href="http://www.cstr.ed.ac.uk/research/projects/artic/mocha.html">http://www.cstr.ed.ac.uk/research/projects/artic/mocha.html</ext-link></comment>
</mixed-citation>
</ref>
</ref-list>
<fn-group>
<fn id="fn0001"><label>1</label><p>Quintessential &#x201C;stress-timed&#x201D; languages include the Germanic languages, and &#x201C;syllable-timed&#x201D; languages, the Romance languages.</p></fn>
<fn id="fn0002"><label>2</label><p>Arthur Lloyd James illustrated the &#x201C;Morse code&#x201D; rhythm of English to a foreign student whose native language was Sinhalese in a historical film <italic>48 Paddington Street</italic> archived by British Path&#x00E9; (URL: <ext-link ext-link-type="uri" xlink:href="https://www.britishpathe.com/video/48-paddington-street/">https://www.britishpathe.com/video/48-paddington-street/</ext-link>, accessed 2 January 2020). Lloyd James&#x2019; patronizing manner in the film is least appreciated though.</p></fn>
<fn id="fn0003"><label>3</label><p>The crux of all these approaches is the consensus of revealing different rhythmicities through different forms of variability in the speech signal. Variability can either be quantified via different physical quantities, such as duration (e.g. Dellwo, <xref ref-type="bibr" rid="cit0011">2006</xref>, <xref ref-type="bibr" rid="cit0012">2009</xref>; Grabe &#x0026; Low, <xref ref-type="bibr" rid="cit0021">2002</xref>; Ramus et al., <xref ref-type="bibr" rid="cit0040">1999</xref>), intensity (e.g. Cichocki et al., <xref ref-type="bibr" rid="cit0007">2014</xref>; Fuchs, <xref ref-type="bibr" rid="cit0017">2016</xref>; He, <xref ref-type="bibr" rid="cit0022">2012</xref>, <xref ref-type="bibr" rid="cit0023">2018</xref>; He &#x0026; Dellwo, <xref ref-type="bibr" rid="cit0024">2016</xref>), and a mixture of various parameters (e.g. Lee &#x0026; Todd, <xref ref-type="bibr" rid="cit0028">2004</xref>); or evaluated through the coordination between prosodic hierarchies, such as the phase difference between syllable- and stress/word-level timescales (e.g. Lancia et al., <xref ref-type="bibr" rid="cit0027">2019</xref>; Leong et al., <xref ref-type="bibr" rid="cit0029">2014</xref>), the power of recurring frequencies in the ENV (e.g. Tilsen &#x0026; Arvaniti, <xref ref-type="bibr" rid="cit0046">2013</xref>; Tilsen &#x0026; Johnson, <xref ref-type="bibr" rid="cit0047">2008</xref>), and a linear relationship between syllable and feet durations (e.g. Barbosa, <xref ref-type="bibr" rid="cit0002">2002</xref>; Eriksson, <xref ref-type="bibr" rid="cit0016">1991</xref>; O&#x2019;Dell &#x0026; Nieminen, <xref ref-type="bibr" rid="cit0036">1999</xref>).</p></fn>
<fn id="fn0004"><label>4</label><p>Park et al. (<xref ref-type="bibr" rid="cit0037">2016</xref>) did not examine the jaw movements <italic>per se</italic>, but the size of mouth opening. Although other factors such as the lip rounding or protrusion also affect the mouth aperture, the principal determinant of the mouth area is the jaw oscillation. It is sensible to include the whole mouth in the visual stimuli as it resembles personal communications more naturally. However, to characterize the spectro-temporal features of the jaw movements, it is more appropriate to measure the kinematics of the jaw free from the interference of other articulators.</p></fn>
<fn id="fn0005"><label>5</label><p>In fact, simple correlations for time-series data are problematic in general with spuriously high correlation coefficients.</p></fn>
<fn id="fn0006"><label>6</label><p>The Hermitian inner product of two signals (in this case, the jaw oscillation and speech ENV) is simply the multiplications of the Fourier coefficients of the first signal and the complex conjugates (sign change of the imaginary part) of the Fourier coefficients of the second one. It reveals the covariance between the two signals in the frequency domain.</p></fn>
<fn id="fn0007"><label>7</label><p>A more stringent &#x03B1;-level (= .01) was chosen in statistical analyses to reduce the chance of false positive findings.</p></fn>
</fn-group>
</back>
</article>
