<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "journalpublishing3.dtd">
<article article-type="research-article" dtd-version="3.0" xml:lang="en" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
	<front>
		<journal-meta>
			<journal-id journal-id-type="publisher-id">Loquens</journal-id>
			<journal-title-group>
				<journal-title>Loquens</journal-title>
				<abbrev-journal-title>Loquens</abbrev-journal-title>
			</journal-title-group>
			<issn pub-type="epub">2386-2637</issn>
			<publisher>
				<publisher-name>Consejo Superior de Investigaciones Científicas</publisher-name>
			</publisher>
		</journal-meta>
		<article-meta>
			<article-id pub-id-type="publisher-id">loquens.2015.021</article-id>
			<article-id pub-id-type="doi">10.3989/loquens.2015.021</article-id>
			<article-categories>
				<subj-group subj-group-type="heading">
					<subject>Articles</subject>
				</subj-group>
			</article-categories>
			<title-group>
				<article-title>Automatic speaker recognition of spanish siblings: (monozygotic and dizygotic) twins and non-twin brothers</article-title>
				<trans-title-group xml:lang="es">
					<trans-title>Reconocimiento automático de locutor con hermanos españoles: hermanos gemelos (monozigóticos y dizigóticos) y no gemelos</trans-title>
				</trans-title-group>
				<alt-title alt-title-type="running-head">Automatic speaker recognition of spanish siblings: (monozygotic and dizygotic) twins and non-twin brothers</alt-title>
			</title-group>
			<contrib-group>
				<contrib contrib-type="author" corresp="yes">
					<name>
						<surname>San Segundo</surname>
						<given-names>Eugenia</given-names>
					</name>
					<xref ref-type="corresp" rid="cor1"/>
				<aff>Department of Linguistics and Language Science, University of York</aff>
				</contrib>
				<contrib contrib-type="author" corresp="yes">
					<name>
						<surname>Künzel</surname>
						<given-names>Hermann </given-names>
					</name>
					<xref ref-type="corresp" rid="cor2"/>
				<aff>Department of Phonetics, University of Marburg</aff>
				</contrib>
			</contrib-group>
			<author-notes>
				<corresp id="cor1">e-mail:<email xlink:href="">eugenia.sansegundo@york.ac.uk</email></corresp>
				<corresp id="cor2">e-mail:<email xlink:href="">kuenzelh@uni-marburg.de</email></corresp>
			</author-notes>
			<pub-date pub-type="epub">
				<day>31</day>
				<month>July</month>
				<year>2015</year>
			</pub-date>
			<pub-date pub-type="collection">
				<year>2015</year>
			</pub-date>
			<volume>2</volume>
			<issue>2</issue>
			<elocation-id content-type="doi">10.3989/loquens.2015.021</elocation-id>
			<history>
				<date date-type="received">
					<day>13</day>
					<month>04</month>
					<year>2015</year>
				</date>
				<date date-type="accepted">
					<day>10</day>
					<month>07</month>
					<year>2015</year>
				</date> 
				<date date-type="available on line">
					<day>18</day>
					<month>05</month>
					<year>2016</year>
			</history>
			<permissions>
				<copyright-statement>© 2015 CSIC</copyright-statement>
				<copyright-year>2015</copyright-year>
				<license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by-nc/3.0/">
					<license-p>This is an open-access article distributed under the terms of the Creative Commons Attribution-Non Commercial (by-nc) Spain 3.0 License.</license-p>
				</license>
			</permissions>
			<abstract xml:lang="en" id="abstract01">
				<title>ABSTRACT</title>
				<p>The performance of the automatic speaker recognition (ASR) system Batvox<sup>TM</sup> (Version 4.1) has been tested with a male population of 24 monozygotic (MZ) twins, 10 dizygotic (DZ) twins, 8 non-twin siblings and 12 unrelated speakers (aged 18–52 with Standard Peninsular Spanish as their mother tongue). Since the cepstral features in which this ASR system is based depend largely on anatomical–physiological foundations, we hypothesized that such features ought to be gene-dependent. Therefore, higher similarity values should be found in MZ twins (100% shared genes) than in DZ twins, in brothers (B) or in a reference population of unrelated speakers (US).  Results corroborated the expected decreasing scale MZ &gt; DZ &gt; B &gt; US since the similarity coefficients yielded by the automatic system for these speakers decreased exactly in the same direction as the kinship degree of the four speaker groups diminishes. This suggests that the system features are to a great extent genetically conditioned and that they are hence useful and robust for comparing speech samples of known and unknown origin, as found in legal cases. Furthermore, the 9.9% EER (Equal Error Rate) obtained when testing MZ pairs lies around the same value (11% EER) found in <xref ref-type="bibr" rid="CIT28">Künzel (2010)</xref> with German twins.</p>
			</abstract>
			<trans-abstract xml:lang="es" id="abstract02">
				<title>RESUMEN</title>
				<p><italic>Reconocimiento automático de locutor con hermanos españoles: hermanos gemelos (monozigóticos y dizigóticos) y no gemelos.–</italic> Hemos utilizado el sistema de reconocimiento automático Batvox<sup>TM</sup> (versión 4.1) con una población de hablantes masculinos compuesta de 24 gemelos monocigóticos, 10 gemelos dicigóticos, 8 hermanos no gemelares y 12 hablantes no emparentados (edades comprendidas entre 18 y 52 años, con español centropeninsular como lengua materna). Puesto que los parámetros cepstrales en los que se basa Batvox<sup>TM</sup> dependen en gran medida de las bases anatómicas y fisiológicas del tracto vocal del hablante, se propuso que estos debían estar influenciados genéticamente.  Esta hipótesis se pudo corroborar, puesto que los coeficientes de similitud arrojados por el sistema automático decrecen exactamente en la misma dirección en la que disminuye el grado de parentesco de las parejas de hablantes, es decir: gemelos monocigóticos, dicigóticos, hermanos no gemelares y hablantes no emparentados. Esto es, los gemelos monocigóticos obtuvieron valores más altos que los dicigóticos; estos, a su vez, mayores que los hermanos no gemelares, y, finalmente, estos últimos mayores que los hablantes no emparentados.  Estos resultados sugieren que los parámetros en los que está basado este sistema de reconocimiento están condicionados en gran medida por aspectos genéticos y, por tanto, resultan útiles y robustos para la comparación de muestras de voz dubitadas e indubitadas que encontramos en un caso típicamente forense. Por otro lado, el EER (<italic>Equal Error Rate</italic>) del 9 % que se obtuvo en las comparaciones exclusivamente de gemelos monocigóticos supone un valor muy similar al hallado en estudios anteriores con gemelos monocigóticos alemanes, como <xref ref-type="bibr" rid="CIT28">Künzel (2010)</xref>: EER del 11 %. </p>
			</trans-abstract>
			<kwd-group xml:lang="en">
				<title>KEYWORDS</title>
				<kwd>forensic phonetics</kwd>
				<kwd>twins</kwd>
				<kwd>siblings</kwd>
				<kwd>automatic speaker recognition</kwd>
				<kwd>Spanish</kwd>
			</kwd-group>
			<kwd-group xml:lang="es">
				<title>PALABRAS CLAVE</title>
				<kwd>fonética judicial</kwd>
				<kwd>gemelos</kwd>
				<kwd>reconocimiento automático</kwd>
				<kwd>español</kwd>
			</kwd-group>
		</article-meta>
	</front>
	<body>
		<sec id="S1">
		<label>1.</label>
			<title>INTRODUCTION</title>
			<sec id="S1.1">
				<label>1.1.</label>
				<title>The forensic relevance of twins and non-twin siblings</title>
				<p>It is widely acknowledged that distinguishing twins poses a major challenge in the field of forensics because these individuals are physically very similar. For instance, biometrics such as fingerprints (<xref ref-type="bibr" rid="CIT21">Jain, Prabhakar &amp; Pankanti, 2002</xref>) or palmprints (<xref ref-type="bibr" rid="CIT26">Kong, Zhang &amp; Lu, 2006</xref>) have often been investigated in twins to study the subtle differences frequently observed between them. Similarly, researchers have investigated behavioral characteristics of twins such as handwriting (<xref ref-type="bibr" rid="CIT61">Srihari, Huang, &amp; Srinivasan, 2008</xref>). In the same way that handwriting depends on physiology as much as on behavioral factors like training and habits, the foundations of speaker recognition are largely grounded on the idea that a voice is determined not only by anatomical structure but also by nonbiological or behavioral factors. These factors include mainly social or dialectal aspects but other environmental influences are possible. <xref ref-type="bibr" rid="CIT38">Nolan and Oh (1996, p. 39)</xref> highlighted that aspects of personal voice quality are determined by anatomical inheritance, mimicking traits from other people, or else, they are arbitrarily chosen in order to mark someone’s personality. This <italic>organic-learned dichotomy </italic>(<xref ref-type="bibr" rid="CIT37">Nolan, 1997</xref>; <xref ref-type="bibr" rid="CIT38">Nolan &amp; Oh, 1996</xref>) may be a good translation in phonetic terms of the well-known <italic>nature–nurture dichotomy</italic>, first outlined by Sir Francis Galton in 1875 (<xref ref-type="bibr" rid="CIT15">Galton 1875</xref>, in <xref ref-type="bibr" rid="CIT60">Segal 1993, p. 45</xref>). </p>
		<p>This distinction, nature vs. nurture, has resulted in fruitful twin research in many disciplines, where heritability or concordance rates are calculated for certain traits in order to determine whether these could be genetically influenced. This happens when there is greater similarity on that trait between monozygotic (MZ) twin pairs than between dizygotic (DZ) twins. MZ twin pairs share 100% of their alleles and DZ twins, on average, share only half their genetic information, whereas both types of twin pairs share essentially the same prenatal and postnatal environments (<xref ref-type="bibr" rid="CIT62">Stromswold 2006, p. 334</xref>). This is the essence of the classical twin design, which requires that an important assumption be made: the <italic>equal environment assumption</italic> (EEA), i.e., it is assumed that the two twin types have similar environmental experience.<xref ref-type="fn" rid="NOTE1">1</xref> A number of studies have investigated the differences in MZ and DZ twins to assess the effect of genetic factors in voice (see Section 1.2), but—to the best of our knowledge—the joint consideration of MZ, DZ and non-twin siblings<xref ref-type="fn" rid="NOTE2">2</xref> has not been approached in phonetic studies before. </p>
		<p>Acknowledging the existence of these two “forces”, i.e., <italic>nature</italic> and <italic>nurture</italic> (alternatively also referred to as <italic>organic</italic> and <italic>learned factors</italic>, respectively) to explain the (dis)similarities between twins does not mean that their relative influence or importance can be clearly separated. Moreover, there is a third element, <italic>epigenetics</italic>, which is often neglected in twins’ studies even though it usually comes into play to explain how changes in gene expression caused by mechanisms other than changes in the underlying DNA sequence can cause divergence in twins, which may account for strikingly dissimilarities between MZ twins. See, for instance, how a particular epigenetic process called DNA methylation (<xref ref-type="bibr" rid="CIT32">Martino et al. 2013</xref>; <xref ref-type="bibr" rid="CIT40">Philips, 2008</xref>) is reported to make the expression of genes weaker or stronger. </p>
		<p>A recent study (<xref ref-type="bibr" rid="CIT12">Felson, 2014</xref>) has aimed at undertaking a comprehensive evaluation of the EEA, which has often caused some skepticism amongst researchers. Felson presents evidence that suggests that neither extreme of the opposing views is correct, and that the truth lies somewhere in the middle. In other words, it seems that although environmental similarity may not have been adequately measured in some sociology-related twin studies, “the resulting bias is likely modest” (p. 184). Therefore it could be argued that despite its limitations twin research is still greatly encouraged nowadays to shed light on the interplay of genetic and environmental factors. Particularly referring to the difficult task of searching for genetic influences of the voice, <xref ref-type="bibr" rid="CIT58">Sataloff (1995)</xref> pointed out that “the complexities of genetic research in humans have left most of the relevant questions unanswered” (p. 17). </p>
		<p>Studies on twins’ voices are undertaken for at least two main reasons. On the one hand, this type of studies can reveal—for the investigated voice characteristics—how the results of pairwise comparisons vary depending of the type of speaker considered. The comparison is usually between MZ twins and DZ twins; <xref ref-type="bibr" rid="CIT54">San Segundo (2014)</xref> also proposed drawing comparisons against non-twin siblings and unrelated speakers. The genetic influence of the analyzed voice characteristics is apparent when higher similarity is observed in MZ twins than in DZ twins, non-twin siblings or unrelated speakers. On the other hand, the relevance of twin studies to Forensic Phonetics<xref ref-type="fn" rid="NOTE3">3</xref> in particular lies in the search for robust<xref ref-type="fn" rid="NOTE4">4</xref> voice characteristics that could facilitate the discrimination of very similar speakers.<xref ref-type="fn" rid="NOTE5">5</xref> Hence, these four speaker groups (MZ twins, DZ twins, non-twin siblings and unrelated speakers) are proposed for testing the performance of a speaker-comparison system. As can be observed, the two highlighted aspects are strongly linked, since a set of characteristics may be robust for speaker comparison as far as they are maximally influenced by the speaker’s genetic endowment and minimally due to learned factors, the latter favoring voice disguise or imitation. The predominance of genes over environment is clearly related to the two most repeated (and probably important) criteria in the identification of characteristics for Forensic Speaker Comparison (FSC), namely that these characteristics should be as consistent as possible for each speaker (low within-speaker variability) and that they should exhibit large variation amongst speakers (high between-speaker variability). Among others, these criteria were already outlined by <xref ref-type="bibr" rid="CIT68">Wolf (1972)</xref> and <xref ref-type="bibr" rid="CIT36">Nolan (1983)</xref> in the phonetic realm, but they also appear in the literature specifically related to automatic speaker recognition (ASR). For example, <xref ref-type="bibr" rid="CIT25">Kinnunen and Li (2010)</xref> refer to the same characteristics for an ideal ASR system. </p>
			</sec>
			<sec id="S1.2">
				<label>1.2.</label>
				<title>Literature review: twins and ASR</title>
				<p>From a literature review of around 30 voice-related twin studies (<xref ref-type="bibr" rid="CIT54">San Segundo, 2014</xref>), we can draw some interesting conclusions. For instance, it seems that previous phonetic studies focusing on twins have aimed at basically one of the following objectives (see <xref ref-type="bibr" rid="CIT55">San Segundo, 2015</xref>): (a) trying to find a genetic component in the variation of certain voice characteristics by searching differences between MZ and DZ twin pairs (e.g., <xref ref-type="bibr" rid="CIT7">Debruyne, Decoster, Van Gijsel, &amp; Vercammen, 2002</xref>; <xref ref-type="bibr" rid="CIT43">Przybyla, Horii, &amp; Crawford, 1992</xref>) or else, in a forensic scenario, (b) creating a system capable of discriminating between MZ and DZ twins (e.g., <xref ref-type="bibr" rid="CIT13">Forrai &amp; Gordos, 1983</xref>) or, more frequently, testing whether it is possible to distinguish a speaker from his/her co-twin (e.g., <xref ref-type="bibr" rid="CIT2">Ariyaeeinia, Morrison, Malegaonkar, &amp; Black, 2008</xref>; <xref ref-type="bibr" rid="CIT20">Homayounpour &amp; Chollet, 1995</xref>; <xref ref-type="bibr" rid="CIT28">Künzel, 2010</xref>; <xref ref-type="bibr" rid="CIT31">Loakes, 2006</xref>; <xref ref-type="bibr" rid="CIT38">Nolan &amp; Oh, 1996</xref>; <xref ref-type="bibr" rid="CIT59">Scheffer, Bonastre, Ghio, &amp; Teston 2004</xref>). For a thorough discussion of the results derived from previous twin studies, see <xref ref-type="bibr" rid="CIT54">San Segundo (2014)</xref>, where previous works have been classified in four groups depending on whether they represent perceptual, acoustic, articulatory or automatic (ASR) approaches. </p>
		<p>While most of the studies undertaken from an acoustic perspective focus on <italic>traditional </italic>phonetic characteristics, as described in Künzel (2011) and Rose (2006)—for example, fundamental frequency (<italic>f</italic>
			<sub>0</sub>), formant patterns or temporal characteristics such as word duration, vowel duration or Voice Onset Time (VOT)—, research into laryngeal features and phonation characteristics derived from the glottal waveform has been very limited. Classical distortion characteristics such as jitter and shimmer have only occasionally been explored in twins (<xref ref-type="bibr" rid="CIT66">van Lierde, Vinck, De Ley, Clement, &amp; Van Cauwenberge 2005</xref>; <xref ref-type="bibr" rid="CIT67">Weirich &amp; Lancia, 2011</xref>). More recently, some investigations on twins’ voices (<xref ref-type="bibr" rid="CIT51">San Segundo, 2012</xref>; <xref ref-type="bibr" rid="CIT56">San Segundo &amp; Gómez-Vilda, 2013</xref>; <xref ref-type="bibr" rid="CIT54">San Segundo 2014</xref>; <xref ref-type="bibr" rid="CIT57">San Segundo &amp; Gómez-Vilda, 2015</xref>) have analyzed a considerably larger number of glottal features, on the basis of the voice analysis methodology described in <xref ref-type="bibr" rid="CIT17">Gómez-Vilda et al. (2007)</xref>, which relies on the decoupling of the vocal tract from the glottal source estimates. </p>
		<p>If we focus on ASR studies in particular, this approach to twins’ voices has not been extensively developed, in comparison with other acoustic studies investigating specific segmental features. The main objectives of the ASR studies reviewed in <xref ref-type="bibr" rid="CIT54">San Segundo (2014)</xref> are one of the following: (a) comparing the performance of ASR systems with the ability of familiar and non-familiar listeners to discriminate twins (<xref ref-type="bibr" rid="CIT20">Homayounpour &amp; Chollet, 1995</xref>); (b) testing if an ASR system is able to detect correctly the twin pair of a speaker (<xref ref-type="bibr" rid="CIT59">Scheffer et al. 2004</xref>), or (c) in general, testing the intra-speaker, inter-speaker and intra-pair similarity of twins, for example in terms of Likelihood Ratios (LRs) or similarity coefficients. In this last research line we find two recent studies, namely <xref ref-type="bibr" rid="CIT24">Kim (2010)</xref> and <xref ref-type="bibr" rid="CIT28">Künzel (2010)</xref>. Since both use the same ASR system that we are using in our study, we will devote an important part of this section to the description of their objectives and main findings. </p>
		<p><xref ref-type="bibr" rid="CIT24">Kim (2010)</xref> studied 22 Korean female twin pairs (17 MZ, including one triplet and five DZ) using <xref ref-type="bibr" rid="CIT1">Agnitio Voice Biometrics’ Batvox<sup>TM</sup> (Version 3.0)</xref>. Two different speaking styles—text reading and spontaneous interview—were used. The results of this investigation showed that every twin speaker was correctly identified in the same speaking style condition (when models and test files were <italic>read</italic> speech). According to the author, this would suggest that, at least in ASR, the same speaking style setting should be provided in order to get more confident results. Noteworthy of this study is also that in nine out of 22 pairs, intra-twin LRs in the same speaking style condition were higher than intra-speaker LRs in different speaking style condition. This situation is highly undesirable in a forensic context, where inter-speaker variation should be larger than intra-speaker variation (<xref ref-type="bibr" rid="CIT68">Wolf, 1972</xref>). </p>
		<p><xref ref-type="bibr" rid="CIT28">Künzel (2010)</xref> is the most recent study on automatic speaker recognition in which a Bayes-based system (Batvox<sup>TM</sup>, Version 3.1) was used to calculate LR distributions for inter-speaker, intra-pair and intra-speaker comparisons. A total of 35 German MZ pairs (26 female and nine male) participated in this study and two different tests were designed. In the first one, both target voices consisted of the same read text, while in the second one the speaker models were built from spontaneous speech samples but read speech samples were used as targets. The results showed that in the first experiment the automatic system allowed a perfect distinction of each member of a male twin pair (i.e., 0% of Equal Error Rate; EER) and 0.5% EER for female twin pairs. In the second experiment, the EER rose to 11% for male twin pairs and 4.4% for female twin pairs. These values represent the crossover point in the Tippett plot for the inter-speaker/intra-speaker LR distributions. However, the results for female twins are worse when considering intra-pair/intra-speaker distributions (19% EER in the first experiment and 48% EER for the second experiment). Therefore, the performance of the system was clearly superior for male than for female voices. The author’s explanation for this phenomenon is that “as a consequence of the higher fundamental frequency of female voices the spacing of the harmonics is less dense than for male voices, which in turn yields less speech sound- and speaker information in the spectrum” (<xref ref-type="bibr" rid="CIT28">Künzel, 2010, p. 270</xref>). This becomes clearer if we bear in mind that the spectrum is used for the extraction of the mel-frequency cepstrum coefficients (MFCCs), which are the features this automatic system used is based upon. </p>
		<p>Finally, references to siblings’ voices within an automatic approach are almost inexistent except for the study of <xref ref-type="bibr" rid="CIT6">Charlet and Lecha (2007)</xref>, which tested a text-dependent speaker recognition system with 33 families, finding that the son was highly confused with his brother. The implication is that someone could be a good impostor of his brother, making this type of speakers especially relevant in forensic studies and thus justifying not only the study of twins but also of non-twin siblings. </p>
			</sec>
		</sec>
		<sec id="S2">
			<label>2.</label>
			<title>DATA AND METHOD</title>
			<p>This section provides some details about the subjects recruited for this investigation, the data collection method and the characteristics of the speech samples analyzed. The methodology for carrying out the speaker comparison is also described, including the different stages of the automatic system Batvox<sup>TM </sup>as well as a description of the method for the measurement of system performance. </p>
			<sec id="S2.1">
				<label>2.1.</label>
				<title>Data collection</title>
				<p>This investigation is part of a larger research project (<xref ref-type="bibr" rid="CIT54">San Segundo, 2014</xref>); more details about the corpus of twin and non-twin subjects can be found in <xref ref-type="bibr" rid="CIT53">San Segundo (2013b</xref>, <xref ref-type="bibr" rid="CIT54">2014</xref>). The automatic analysis that we present here is based on speech samples extracted from the fifth corpus task: informal interview with the researcher. </p>
				<sec id="S2.1.1">
					<label>2.1.1.</label>
					<title>Subjects and recording characteristics</title>
					<p>Our corpus of speakers is made up of 24 MZ twins, 10 DZ twins, eight brothers and 12 unrelated speakers with no kinship relationship (friends or work colleagues). The importance of the first three speaker types has been explained in the introduction. The fourth group of speakers was recruited with the aim of creating a reference population, whose relevance for Likelihood-Ratio-based forensic studies has been acknowledged on numerous occasions in the literature (<xref ref-type="bibr" rid="CIT34">Morrison, 2010</xref>).<xref ref-type="fn" rid="NOTE6">6</xref> Friends or work colleagues were preferred instead of complete strangers in order to match as closely as possible the speaking style found in the conversations between brothers, characterized by their spontaneity due to a long-term relationship. The age of the speakers ranged between 18–52 years (mean: 28.96). The age difference between the brothers in each pair varied between four and 11 years. They were all male speakers of North-Central Peninsular Spanish with no speech pathologies or hearing difficulties. All speakers were recorded on two different occasions in order to account for intra-speaker variability. These two recording sessions were separated by 2–3 weeks, which served to obtain non-contemporaneous speech samples. </p>
		<p>Participants came in pairs (either with their twin or friend) to the recording sessions, which took place in the Phonetics Laboratory of the Spanish National Research Council. They were recorded with omnidirectional condenser microphones (head-mounted device) with flat frequency response. Recording specifications were: 44,1 kHz sample rate, 16-bit resolution and mono channel. Speakers were recorded in two different (acoustically isolated) rooms where they could communicate via landline telephone for certain cooperative tasks. Even though the recordings are high quality (telephone-degraded at a later stage), this set-up replicated forensic realistic conditions at the same time that it minimized the “observer’s paradox” (<xref ref-type="bibr" rid="CIT30">Labov, 1972</xref>) by avoiding the presence of the researcher at the place of the recording. </p>
				</sec>
				<sec id="S2.1.2">
					<label>2.1.2.</label>
					<title>Speech samples</title>
					<p>Speech samples were extracted from the fifth task of the corpus fully described in <xref ref-type="bibr" rid="CIT53">San Segundo (2013b</xref>, <xref ref-type="bibr" rid="CIT54">2014</xref>). In this speaking task (informal interview with the researcher), the researcher is at one end of the telephone and one member of each speaker pair at a time is at the other end of the telephone. In this interview, lasting around 10 minutes, the researcher asks each of the interviewees about any of the topics that they have been discussing with their twin/friend in the first task. Originally intended to elicit hesitation markers (i.e., vowel fillers) from the speakers, which could then facilitate glottal analyses (e.g., <xref ref-type="bibr" rid="CIT56">San Segundo &amp; Gómez-Vilda, 2013</xref>, <xref ref-type="bibr" rid="CIT57">2015</xref>), this corpus task was also considered the most appropriate for the ASR analysis. On the one hand, conversations here are long enough to allow the extraction of at least 120 seconds of net speech per speaker. According to <xref ref-type="bibr" rid="CIT28">Künzel (2010)</xref>, this is the recommended duration of a voice sample to be analyzed using the ASR system Batvox<sup>TM</sup>. On the other hand, this corpus task presents the advantage of having the same interlocutor in all conversations, i.e., the researcher. This leveled the speaking style of all speakers to the same degree of spontaneity/formality.<xref ref-type="fn" rid="NOTE7">7</xref> </p>
		<p>The speech fragments (120 s of duration on average) were extracted from the audio files belonging to the first and the second recording session of each speaker (average duration of 5 min). The speech material chosen for further analyses was selected from approximately the middle of the audio file, in order to avoid the beginning of the conversation, where the speaker has not already settled to his ordinary speaking style. Prior to the labeling and extraction using Praat (Version 5.3.79), the audio files were first aurally examined in order to remove extraneous noise, laughter, clicks, cough, etc., following the recommendations in <xref ref-type="bibr" rid="CIT28">Künzel (2010, p. 256)</xref>. </p>
				</sec>
			</sec>
			<sec id="S2.2">
				<label>2.2.</label>
				<title>Analysis tools and method </title>
				<sec id="S2.2.1">
					<label>2.2.1.</label>
					<title>ASR analysis</title>
					<p>For the ASR analysis, we have used the software Batvox<sup>TM</sup> (Version 4.1), which is based on parameters related to the resonances of the vocal tract, basically cepstral coefficients. One of the main assets of automatic systems is that between-sample differences in the speech content are not relevant because ASR systems exploit the voice itself and disregard the linguistic content of the utterances to a great extent. While this does not mean that Batvox<sup>TM</sup> is independent of the language mismatch between utterances to compare—which is not our case—, it still holds true that the relatively small influence of the linguistic content makes the extraction of speaker samples relatively easy, as there is no need for comparable phonetic units between speakers (in contrast with most traditional phonetic features). An overview of the first stages of a typical ASR system follows (see <xref ref-type="bibr" rid="CIT25">Kinnunen &amp; Li, 2010, pp. 2–3</xref>)<xref ref-type="fn" rid="NOTE8">8</xref>: </p>
					<list list-type="disc">
						<list-item>
							<p><italic>Parameter extraction</italic>: transformation of the raw signal into feature vectors in which speaker-specific properties are emphasized and statistical redundancies suppressed. </p>
						</list-item>
						<list-item>
							<p><italic>Speaker modeling</italic>: the feature vectors extracted from the training utterance of a speaker are used to train a speaker model, which is then stored in the system database. The Gaussian mixture model (GMM; <xref ref-type="bibr" rid="CIT45">Reynolds, Quatieri &amp; Dunn, 2000</xref>; <xref ref-type="bibr" rid="CIT46">Reynolds &amp; Rose, 1995</xref>) would be the most popular model for text-independent recognition, according to <xref ref-type="bibr" rid="CIT25">Kinnunen and Li (2010, p. 4)</xref>. </p>
						</list-item>
					</list>
					<p>Focusing on Batvox<sup>TM</sup> in particular, its main characteristics are, as explained in <xref ref-type="bibr" rid="CIT29">Künzel and Alexander (2014, p. 247)</xref>: a 38-dimensional feature vector consisting of 19 MFCCs plus their deltas, GMM-Channel-Factor analysis for the compensation of speaker models (<xref ref-type="bibr" rid="CIT23">Kenny, Boulianne, Ouellet, &amp; Dumouchel, 2005</xref>) and nuisance attribute projection (<xref ref-type="bibr" rid="CIT5">Campbell, Campbell, Reynolds, Singer, &amp; Torres-Carrasquillo, 2006</xref>) for the test files. </p>
		<p>A comparison between the statistical model for the reference speaker and the results for the target speaker’s model is carried out. The similarity score obtained after this procedure is then weighed using a reference population. For this study, the system was set to <italic>identification mode</italic>,<xref ref-type="fn" rid="NOTE9">9</xref> where results are indicated as normalized scores that can be used to calculate False Alarms (FA) and False Rejections (FR) rates, and eventually, EERs. This identification mode of operation was deemed the most appropriate for the purpose of this investigation (see <italic>Batvox 4.1 Basic User Manual</italic>, 2013). As reference population, a cohort of 31 Spanish male speakers was used (from Batvox<sup>TM</sup> databases), with characteristics matching those of the recordings in the twin corpus: male speakers, spontaneous conversations and high-quality recordings. </p>
		<p>The following tests were carried out: </p>
		<list list-type="disc">
			<list-item>
				<p><italic>Intra-speaker comparisons</italic> (<italic>matches</italic> or <italic>target trials</italic>): each speaker’s session one was compared with the same speaker’s session two.</p>
			</list-item>
			<list-item>
				<p><italic>Inter-speaker comparisons </italic>(<italic>non-matches </italic>or <italic>impostor trials</italic>): each speaker’s session one was compared with all other speakers’ session two. </p>
			</list-item>
			<list-item>
				<p><italic>Intra-pair comparisons: </italic>each speaker’s first session was compared with the first session of his sibling or conversation partner in the case of unrelated speakers. </p>
			</list-item>
		</list>
		<p>The first two types of comparisons served to test the general performance of the comparison system without taking into account the fact that some speakers are MZ, DZ or non-twin siblings. Yet, in order to investigate the magnitude of the <italic>sibling effect</italic>, the third type of test is also necessary. </p>
				</sec>
				<sec id="S2.2.2">
					<label>2.2.2.</label>
					<title>Performance measures</title>
					<p>Assessing the output accuracy of a forensic-comparison system is a very relevant aspect in forensic sciences. Several measures and graphical ways have therefore been developed to evaluate such accuracy: for instance, the <italic>log-likelihood-ratio cost </italic>(<italic>C<sub>llr</sub></italic>), originally envisaged for its use in ASR (<xref ref-type="bibr" rid="CIT4">Brümmer &amp; du Preez, 2006</xref>; <xref ref-type="bibr" rid="CIT65">van Leeuwen &amp; Brümmer, 2007</xref>) but also applied in forensic-comparison studies based in traditional acoustic parameters (e.g., <xref ref-type="bibr" rid="CIT19">Gonzalez-Rodriguez, Rose, Ramos, Toledano, &amp; Ortega-Garcia, 2007</xref>; <xref ref-type="bibr" rid="CIT35">Morrison &amp; Kinoshita, 2008</xref>). Besides, <italic>Tippett plots</italic> (<xref ref-type="bibr" rid="CIT33">Meuwly, 2001</xref>) have also been used as a graphical method to present the output of forensic systems and to assess its accuracy. In our study we have used EER, an accepted measure of the performance of an identification (also used in <xref ref-type="bibr" rid="CIT28">Künzel, 2010</xref>, or <xref ref-type="bibr" rid="CIT29">Künzel &amp; Alexander, 2014</xref>, for the performance testing of Batvox<sup>TM</sup>). The EER represents the point of intersection of matches and non-matches. Consequently, an EER of 0% indicates that there is no overlap of matches and non-matches, so neither FA nor FR occur. EERs were calculated using the Biometrics 1.2 software (Biometrics 1.2, 2012). </p>
				</sec>
			</sec>
		</sec>
		<sec id="S3">
			<label>3.</label>
			<title>RESULTS</title>
			<sec id="S3.1">
				<label>3.1.</label>
				<title>Overall system performance</title>
				<p>As explained above, we carried out three types of tests, which yielded results for intra-speaker, inter-speaker and intra-pair comparisons. If we first look at the results for intra-speaker and inter-speaker comparisons alone, we see that similarly high coefficients of recognition are obtained for all the pooled four speaker types (MZ, DZ, B and US). This can be observed in <xref ref-type="fig" rid="F1">Figure 1</xref>, which shows a 0% EER. The input values for the creation of this figure were of two types: </p>
				<list list-type="disc">
					<list-item>
						<p><italic>Matches </italic>(blue line): the values were obtained from the comparison of each speaker’s session one with his own session two. </p>
					</list-item>
					<list-item>
						<p><italic>Non-matches </italic>(red line): the values were obtained from the comparison of each speaker’s session one and all other speakers’ session two.<xref ref-type="fn" rid="NOTE10">10</xref></p>
					</list-item>
				</list>
				<fig id="F1">
					<label>Figure 1.</label>
					<caption>
						<title>Cumulative distribution of scores for same-speaker comparisons or matches (blue) and different-speaker comparisons or non-matches (red).</title>
					</caption>
					<graphic xlink:href="loquens021_f1.jpg" xmlns:xlink="http://www.w3.org/1999/xlink"/>
				</fig>
				<p>The 0% EER indicates that there is no overlap of matches and non-matches, so neither FA nor FR occur. This shows that the overall system performance with high-quality recordings and without taking into account the sibling effect (intra-pair comparisons) is perfect. </p>
			</sec>
			<sec id="S3.2">
				<label>3.2.</label>
				<title>Sibling effect</title>
				<p>When taking into account also intra-pair comparisons, in addition to matches and non-matches, the recognition coefficients are expected to be much lower, as the comparison is not between the same individuals. However, different patterns were observed depending on the type of speaker (MZ, DZ, B or US). This can be seen in <xref ref-type="table" rid="T1">Table 1</xref>, where the values obtained are classified per speaker (i.e., his intra-speaker coefficients) and per speaker pair (i.e., their intra-pair coefficients), depending on whether they are MZ, DZ, B or US. As it can be observed in this table, all intra-speaker comparisons yield similarly high coefficients of recognition. In relation to the intra-pair comparisons, <xref ref-type="table" rid="T1">Table 1</xref> is useful to observe the different values obtained by different speaker pairs, i.e., the performance of the system can be analyzed per speaker or per speaker pair. The fact that the speakers in this investigation are not very numerous is an advantage in order to carry out this kind of detailed examination. For instance, if we look at within-group differences, the value of MZ pair 39–40 (0.64) is very different from the other pairs’ coefficients (much higher in average). </p>
				<table-wrap id="T1">
		<label>Table 1:</label>
		<caption>
		<title>Summary of the results for the different comparison tests. MZ: Monozygotic twins; DZ: Dizygotic twins; B: Brothers; US: Unrelated Speakers. Divided columns are used in the intra-speaker scores for each pair member. Cases: xxvyy means speaker xx versus speaker yy. </title>
		</caption>
		<table frame="hsides" rules="groups">
			<thead>
				<tr>
					<th rowspan="2"></th>
					<th colspan="3">MZ</th>
					<th colspan="3">DZ</th>
					<th colspan="3">B</th>
					<th colspan="3">US</th>
				</tr>
				<tr>
					<th colspan="2">Intra-speaker</th>
					<th>Intra-pair</th>
					<th colspan="2">Intra-speaker</th>
					<th>Intra-pair</th>
					<th colspan="2">Intra-speaker</th>
					<th>Intra-pair</th>
					<th colspan="2">Intra-speaker</th>
					<th>Intra-pair</th>
				</tr>
			</thead>
			<tbody>
				<tr>
				<td>Cases</td>
				<td colspan="2">01v01/02v02</td>
				<td>01v02</td>
				<td colspan="2">13v13/14v14</td>
				<td>13v14</td>
				<td colspan="2">21v21/22v22</td>
				<td>21v22</td>
				<td colspan="2">25v25/26v26</td>
				<td>25v26</td>
				</tr>
		<tr>
				<td>Score</td>
				<td>4.22</td>
				<td>3.48</td>
				<td>3.79</td>
				<td>5.25</td>
				<td>6.17</td>
				<td>3.77</td>
				<td>4.51</td>
				<td>6.24</td>
				<td>0.64</td>
				<td>4.93</td>
				<td>4.47</td>
				<td>0.39</td>
				</tr>
		<tr>
				<td>Cases</td>
				<td colspan="2">03v03/04v04</td>
				<td>03v04</td>
				<td colspan="2">15v15/16v16</td>
				<td>15v16</td>
				<td colspan="2">23v23/24v24</td>
				<td>23v24</td>
				<td colspan="2">27v27/28v28</td>
				<td>27v28</td>
				</tr>
		<tr>
				<td>Score</td>
				<td>4.82</td>
				<td>4.79</td>
				<td>2.65</td>
				<td>4.27</td>
				<td>4.87</td>
				<td>2.53</td>
				<td>7.76</td>
				<td>5.27</td>
				<td>3.31</td>
				<td>3.99</td>
				<td>4.29</td>
				<td>0.64</td>
				</tr>
		<tr>
				<td>Cases</td>
				<td colspan="2">05v05/06v06</td>
				<td>05v06</td>
				<td colspan="2">17v17/18v18</td>
				<td>17v18</td>
				<td colspan="2">47v47/48v48</td>
				<td>47v48</td>
				<td colspan="2">29v29/30v30</td>
				<td>29v30</td>
				</tr>
		<tr>
				<td>Score</td>
				<td>4.29</td>
				<td>4.95</td>
				<td>3.45</td>
				<td>5.13</td>
				<td>6.35</td>
				<td>0.18</td>
				<td>5.53</td>
				<td>4.63</td>
				<td>0.79</td>
				<td>5.29</td>
				<td>5.42</td>
				<td>-0.66</td>
				</tr>
		<tr>
				<td>Cases</td>
				<td colspan="2">07v07/08v08</td>
				<td>07v08</td>
				<td colspan="2">19v19/20v20</td>
				<td>19v20</td>
				<td colspan="2">49v49/50v50</td>
				<td>49v50</td>
				<td colspan="2">31v31/32v32</td>
				<td>31v32</td>
				</tr>
		<tr>
				<td>Score</td>
				<td>4.23</td>
				<td>4.14</td>
				<td>2.31</td>
				<td>3.51</td>
				<td>5.46</td>
				<td>2.17</td>
				<td>2.78</td>
				<td>3.31</td>
				<td>0.36</td>
				<td>2.92</td>
				<td>4.67</td>
				<td>0.25</td>
				</tr>
		<tr>
				<td>Cases</td>
				<td colspan="2">09v09/10v10</td>
				<td>09v10</td>
				<td colspan="2">45v45/46v46</td>
				<td>45v46</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">51v51/52v52</td>
				<td>51v52</td>
				</tr>
		<tr>
				<td>Score</td>
				<td>3.64</td>
				<td>4.06</td>
				<td>2.66</td>
				<td>3.44</td>
				<td>3.83</td>
				<td>0.40</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
				<td>3.80</td>
				<td>3.52</td>
				<td>0.71</td>
				</tr>
		<tr>
				<td>Cases</td>
				<td colspan="2">11v11/12v12</td>
				<td>11v12</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
				<td colspan="2">53v53/54v54</td>
				<td>53v54</td>
				</tr>
		<tr>
				<td>Score</td>
				<td>3.24</td>
				<td>5.29</td>
				<td>1.34</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
				<td>4.03</td>
				<td>5.22</td>
				<td>0.22</td>
				</tr>
		<tr>
				<td>Cases</td>
				<td colspan="2">33v33/34v34</td>
				<td>33v34</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			</tr>
		<tr>
				<td>Score</td>
				<td>4.55</td>
				<td>6.06</td>
				<td>3.20</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			</tr>
		<tr>
				<td>Cases</td>
				<td colspan="2">35v35/36v36</td>
				<td>35v36</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			</tr>
		<tr>
				<td>Score</td>
				<td>6.44</td>
				<td>3.94</td>
				<td>4.93</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			</tr>
		<tr>
				<td>Cases</td>
				<td colspan="2">37v37/38v38</td>
				<td>37v38</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			</tr>
		<tr>
				<td>Score</td>
				<td>5.41</td>
				<td>4.52</td>
				<td>3.54</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			</tr>
		<tr>
				<td>Cases</td>
				<td colspan="2">39v39/40v40</td>
				<td>39v40</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			</tr>
		<tr>
				<td>Score</td>
				<td>6.05</td>
				<td>6.74</td>
				<td>0.64</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			</tr>
		<tr>
				<td>Cases</td>
				<td colspan="2">41v41/42v42</td>
				<td>41v42</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			</tr>
		<tr>
				<td>Score</td>
				<td>4.68</td>
				<td>5.9</td>
				<td>3.53</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			</tr>
		<tr>
				<td>Cases</td>
				<td colspan="2">43v43/44v44</td>
				<td>43v44</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			</tr>
		<tr>
				<td>Score</td>
				<td>4.43</td>
				<td>4.08</td>
				<td>2.59</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			<td colspan="2">&#160;</td>
			<td>&#160;</td>
			</tr>
			</tbody>
		</table>
	</table-wrap>
		<p>If we are interested in the behavior of the groups in general, and not specifically in each pair, <xref ref-type="table" rid="T2">Table 2</xref> and its corresponding figure (<xref ref-type="fig" rid="F2">Figure 2</xref>) are more insightful and probably more appropriate to assess the system performance depending on the speaker type. According to the information in <xref ref-type="table" rid="T2">Table 2</xref>, MZ intra-pair comparisons yield the highest values (i.e., the dissimilarity is the lowest, so they are the most similar speakers). From the average values obtained by the MZ pairs to the coefficient values yielded for US, we observe a gradation from largest to lowest, all through the average values of the DZ intra-pair comparisons and the B intra-pair comparisons. This trend is thus in agreement with our hypothesis, where we predicted the following scale (from more to less similar): MZ &gt; DZ &gt; B &gt; US. In other words, the coefficient grading goes in the same direction as the “magnitude” of kinship relationship. </p>
		<table-wrap id="T2">
		<label>Table 2.</label>
		<caption>
		<title>Average coefficients per speaker type and test type. All the intra-pair values per speaker type but also the intra-speaker values for MZ twins (last row) are shown, in order to highlight the grading in values (from lowest to largest), where the lowest means more dissimilar and the largest, more similar. </title>
		</caption>
		<table frame="hsides" rules="groups">
			<thead>
				<tr>
					<th>Speaker type</th>
					<th>Test type</th>
					<th>Average coefficient</th>
				</tr>
			</thead>
			<tbody>
				<tr>
					<td>Unrelated speakers (US)</td>
					<td rowspan="4">Intra-pair</td>
					<td>0.26</td>
				</tr>
				<tr>
					<td>Non-twin brothers (B)</td>
					<td>1.28</td>
				</tr>
				<tr>
					<td>Dyzigotic twins (DZ)</td>
					<td>1.81</td>
				</tr>
				<tr>
					<td>Monozygotic twins (MZ)</td>
					<td>2.89</td>
				</tr>
				<tr>
					<td>Monozygotic twins (MZ)</td>
					<td>Intra-speaker</td>
					<td>4.83</td>
				</tr>
			</tbody>
		</table>
	</table-wrap>
		<fig id="F2">
					<label>Figure 2.</label>
					<caption>
						<title>Grading of average coefficients from US (Unrelated Speakers) to MZ (Monozygotic) intra-speaker comparisons: the larger the value, the more similarity. Grey is used for intra-pair comparisons while black is used for intra-speaker comparisons; B: brothers; DZ: dizygotic.</title>
					</caption>
					<graphic xlink:href="loquens021_f2.jpg" xmlns:xlink="http://www.w3.org/1999/xlink"/>
				</fig>
		<p>We have added to <xref ref-type="table" rid="T2">Table 2</xref> the average coefficients obtained in (MZ) intra-speaker comparisons. As expected, these same-speaker comparisons yield the highest coefficients. The inclusion of these matches in the table is intended to serve as a baseline to which the rest of (intra-pair) coefficients can be compared, under the assumption that nobody could be more similar to anyone than to himself, although some exceptions may occur in the case of MZ twins, as we describe in Section 3.3. </p>
			</sec>
			<sec id="S3.3">
				<label>3.3.</label>
				<title>Special case study: MZ twins</title>
				<p>The MZ intra-pair comparisons deserve special consideration. As they represent the cases of highest similarity in human beings, they have been more often studied than the other types of kinship relationships considered in this investigation. In the case of FSC carried out using automatic recognition methods, the existence of previous studies that have also used Batvox<sup>TM</sup> for the voice comparison of MZ twins gives us the opportunity to compare our results with previous findings. </p>
		<p>For the MZ twins participating in our study, we have considered useful to compare the coefficients obtained by each speaker in the intra-speaker (IS) comparisons with the coefficients obtained by these same speakers in the intra-pair (IP) comparisons. <xref ref-type="table" rid="T3">Table 3</xref> contains this information, extracted from the general results shown in <xref ref-type="table" rid="T1">Table 1</xref>.</p>
		<table-wrap id="T3">
		<label>Table 3.</label>
		<caption>
		<title>For each of the MZ twin pairs, we show the IS–IP value, calculated as the difference between the intra-speaker (IS) comparison coefficient and the intra-pair (IP) comparison coefficient. Cases: xxvyy means speaker xx versus speaker yy. Only two out of 12 cases (greyshaded) show negative values.</title>
		</caption>
		<table frame="hsides" rules="groups">
			<thead>
				<tr>
					<th>MZ pair</th>
					<th>IS  comparison coefficient</th>
					<th>IP  comparison coefficient</th>
					<th>IS–IP difference</th>
				</tr>
			</thead>
			<tbody>
				<tr>
					<td>01v02</td>
					<td>3.48</td>
					<td>3.79 </td>
					<td style="color">–0.31 </td>
				</tr>
				<tr>
					<td>03v04</td>
					<td>4.79</td>
					<td>2.65 </td>
					<td>2.14 </td>
				</tr>
				<tr>
					<td>05v06</td>
					<td>4.95</td>
					<td>3.45 </td>
					<td>1.50 </td>
				</tr>
				<tr>
					<td>07v08</td>
					<td>4.14 </td>
					<td>2.31 </td>
					<td>1.83 </td>
				</tr>
				<tr>
					<td>09v10</td>
					<td>4.06 </td>
					<td>2.66 </td>
					<td>1.40 </td>
				</tr>
				<tr>
					<td>11v12</td>
					<td>5.29 </td>
					<td>1.34 </td>
					<td>3.95 </td>
				</tr>
				<tr>
					<td>33v34</td>
					<td>6.06 </td>
					<td>3.20 </td>
					<td>2.86 </td>
				</tr>
				<tr>
					<td>35v36</td>
					<td>3.94 </td>
					<td>4.93 </td>
					<td style="color">–0.99 </td>
				</tr>
				<tr>
					<td>37v38</td>
					<td>4.52 </td>
					<td>3.54 </td>
					<td>0.98 </td>
				</tr>
				<tr>
					<td>39v40</td>
					<td>6.74 </td>
					<td>0.64 </td>
					<td>6.10 </td>
				</tr>
				<tr>
					<td>41v42</td>
					<td>5.9 </td>
					<td>3.53 </td>
					<td>2.37 </td>
				</tr>
				<tr>
					<td>43v44</td>
					<td>4.08 </td>
					<td>2.59 </td>
					<td>1.49 </td>
				</tr>
			</tbody>
		</table>
	</table-wrap>
		<p>We have calculated an <italic>IS–IP value</italic> to measure the difference between the IS comparison coefficient and the IP comparison coefficient. This has been done per speaker and speaker pair. Note however that for the IS coefficients, we have only taken into account the values obtained by one member of the pair: the twin member with the even number in his pair (i.e., 02, 04, 06, 08, etc.). The selection of the IS coefficients of the odd pairs did not yield any negative value. That is the reason why we show the results of the even numbers; as explained above, the interest of this calculation lies in finding any possible speaker pair subject to discrimination errors by the system under test. </p>
		<p>As shown in <xref ref-type="table" rid="T3">Table 3</xref>, only two cases out of 12 MZ pairs show a negative value in their IS–IP value, meaning that the IP coefficient is larger than the IS coefficient. This implies that in these two cases the automatic system Batvox<sup>TM </sup>would not be able to discriminate between one twin and the other. In positive values, we can say that in 83.3% of the total MZ cases, the system identifies an identical twin without falsely accepting his co-twin. In <xref ref-type="fig" rid="F3">Figure 3</xref> we draw the IP and IS coefficient values per MZ twin pair, in IS-decreasing order to show how the trend “large IS–small IP” is followed in all cases except in the last two, corresponding to the MZ pairs 01v02 and 35v36, as we could also observe in <xref ref-type="table" rid="T3">Table 3</xref>. These two pairs account for the 16.7% not confirming the hypothesis that IS comparisons are always larger than MZ IP comparisons. However, as we will discuss below, this small percentage is in agreement with previous studies. </p>
		<fig id="F3">
					<label>Figure 3.</label>
					<caption>
						<title>IS–IP difference per speaker pair. We show in the x-axis the 12 MZ (monozygotic) pairs and in the y-axis the coefficient values for IS (intra-speaker) comparisons (grey) and IP (intra-pair) comparisons (black). Only the two last twin pairs would not be discriminated by the system.</title>
					</caption>
					<graphic xlink:href="loquens021_f3.jpg" xmlns:xlink="http://www.w3.org/1999/xlink"/>
				</fig>
		<p>The two specific cases of MZ twins that were not be recognized by the system explain the 9.9% EER obtained in <xref ref-type="fig" rid="F4">Figure 4</xref>, where the line for matches (right line) is used in this case for intra-pair comparisons (only MZ) and the line for non-matches (left curve) represents the inter-speaker comparisons. </p>
		<p>In <xref ref-type="fig" rid="F5">Figure 5</xref>, we have added the curves in <xref ref-type="fig" rid="F1">Figure 1</xref>, which showed the overall system performance. The line further to the right (black) is for IS comparisons of all the speakers in the corpus, and the other right line (blue) represents the IP comparisons, only for MZ. In this new figure, one can distinguish a left-shift from the general IS-curve to the MZ IP-curve, which indicates the performance deterioration from a situation where the system has to recognize same speakers to a situation where identical-twin recognition takes place. The lines for the non-matches in both cases (compare the two curves rising to the left) are practically identical. In both cases, they represent different-speaker comparisons, while in one case (yellow curve, i.e., non-matches in <xref ref-type="fig" rid="F1">Figure 1</xref>) these tests compared the first session of each speaker with the first session of all the other speakers in our corpus; and in the other case (red line, i.e, non-matches in <xref ref-type="fig" rid="F4">Figure 4</xref>), the different-speaker tests were obtained from comparing each speaker’s first session with all the other speakers’ second session. </p>
		<fig id="F4">
					<label>Figure 4.</label>
					<caption>
						<title>Cumulative distribution of scores for intra-pair (IP) comparisons or matches (blue) and inter-speaker (IS) comparisons or non-matches red). The EER obtained is 9.9%, indicating that some overlap between matches and non-matches exist.</title>
					</caption>
					<graphic xlink:href="loquens021_f4.jpg" xmlns:xlink="http://www.w3.org/1999/xlink"/>
				</fig>
		<fig id="F5">
					<label>Figure 5.</label>
					<caption>
						<title><italic>Lines rising to the right: </italic>cumulative distributions of scores for all-speakers intra-speaker (IS) comparisons (black line) and MZ intra-pair (IP) comparisons (blue line, crossing at the EER 9.9%). <italic>Curves rising to the left </italic>(yellow and red): both represent the cumulative distribution of IP comparisons or non-matches.<xref ref-type="fn" rid="NOTE11">11</xref></title>
					</caption>
					<graphic xlink:href="loquens021_f5.jpg" xmlns:xlink="http://www.w3.org/1999/xlink"/>
				</fig>
			</sec>
		</sec>
		<sec id="S4">
			<label>4.</label>
			<title>DISCUSSION</title>
			<p>Several aspects can be discussed in relation to the results obtained with the automatic system Batvox<sup>TM</sup>. On the one hand, we have tested the overall system performance with our speakers as <italic>tests</italic> and <italic>models</italic>, i.e., without taking into account the fact that part of these speakers are twins or siblings. This test has yielded intra- and inter-speaker comparisons. In other words: matches (for same-speaker comparisons) and non-matches (for different speaker comparisons). The 0% EER obtained for this first test shows that there were no FA or FR, which indicates a perfect performance of the system. </p>
		<p>On a second test, we introduced the concept of intra-pair (IP) comparison while taking into account the fact that out of the 54 speakers considered, 24 were MZ twins, 10 were DZ twins, eight were non-twin siblings and 12 were unrelated speakers. The results of comparing each speaker with his pair corroborated the hypothesis that higher similarity values would be found in MZ twins than in DZ twins, in siblings or in unrelated speakers. On average, higher coefficients were obtained by MZ IP-comparisons, followed by DZ twins, brothers and unrelated speakers, in that order. This is the scale that we expected taking into account the degree of shared genes and shared environmental factors by pairs in these four speaker types (see Section 1). </p>
		<p>Finally, when the IP comparison values only for the MZ twins were compared with the non-matches, we obtained a 9.9% EER, so a left-shift was observed in <xref ref-type="fig" rid="F5">Figure 5</xref> from the general IS-curve to the MZ IP-curve. This represents the deterioration in the system performance from a situation where the recognition is between same speakers to a situation where identical-twin recognition takes place. These results could be compared with the 11% EER obtained by <xref ref-type="bibr" rid="CIT28">Künzel (2010)</xref>, who also studied MZ twins. Although he studied both male and female twins, and two speaking styles (read speech and spontaneous speech) we have considered here only the results for male twins and spontaneous speech. The male participants in Künzel’s study were nine MZ pairs while in our investigation there are 12 pairs. Yet the EER percentages are very similar, indicating that the rate of false acceptance of other twin by this system is around 10%. Having a closer look at the data for the individual twin pairs (i.e., comparing the IP and the IS values), Künzel found that some speakers were more easily identified than others. Our study also points in this direction, as the coefficients in the IS and IP comparisons differ between pairs, sometimes considerably (see <xref ref-type="table" rid="T3">Table 3</xref> and <xref ref-type="fig" rid="F3">Figure 3</xref>). In fact, as it follows from the literature review carried out in <xref ref-type="bibr" rid="CIT54">San Segundo (2014)</xref>—and summarized in the introduction to this article—this heterogeneity appears as a common factor in most studies on twins’ voices. Previous analyses derived from the same corpus of Spanish twins showed the same phenomenon, namely that different twin pairs exhibit different results when an IP comparison is carried out, regardless of the type of phonetic-acoustic examination, be it formant trajectories (<xref ref-type="bibr" rid="CIT54">San Segundo, 2014</xref>) or glottal characteristics (<xref ref-type="bibr" rid="CIT56">San Segundo &amp; Gómez Vilda, 2013</xref>). Indeed, this need not be a characteristic exclusively linked to twins but common in speaker recognition. As <xref ref-type="bibr" rid="CIT9">Doddington, Liggett, Martin, Przybocki, and Reynolds (1998)</xref> explain, different speaker typologies could be established on the basis on how easily recognized/imitated speakers are. This implies that, in terms of FA and FR, “a considerable amount of the errors in an experiment, may be linked to only a few speakers” (<xref ref-type="bibr" rid="CIT28">Künzel, 2010, p. 264</xref>). </p>
		<p>Apart from <xref ref-type="bibr" rid="CIT28">Künzel (2010)</xref>, the other study that has analyzed twins’ voices using Batvox<sup>TM</sup> (Version 3.0) focused only on female voices (<xref ref-type="bibr" rid="CIT24">Kim, 2010</xref>), so the results in that study are not comparable with ours. From the investigation of <xref ref-type="bibr" rid="CIT28">Künzel (2010)</xref> we know that there is an important sex-related difference in the performance of the automatic system, this being superior for male as compared to female voices (see Section 1.2). Yet, it is worth-mentioning that <xref ref-type="bibr" rid="CIT24">Kim (2010)</xref> also found that in nine out of 22 cases, twins could be misidentified. She specifically refers to a situation where intra-twin LRs in the same speaking style condition were higher than intra-speaker LRs in different speaking style condition. </p>
		</sec>
		<sec id="S5">
			<label>5.</label>
			<title>CONCLUSIONS </title>
			<p>It is well known that the vocal tract is made up of different cavities (oral, nasal and pharyngeal). Each of these cavities has a resonance profile, which is supposed to be somehow typical and idiosyncratic for each speaker, at least similarly to what happens with other parts of the human anatomy, which are more or less individual (<xref ref-type="bibr" rid="CIT28">Künzel, 2010, p. 40</xref>). Automatic methods in general (as explained above), and Batvox<sup>TM</sup> specifically, extract a set of features representing the resonance profile of the vocal cavities of a speaker (MFCCs) and creates a multidimensional vector. These are the kind of parameters (<italic>low-level features</italic>) used in this type of analysis, in contrast with <italic>high-level features</italic>, which would refer to other linguistic aspects that also serve to characterize a speaker, such as intonation patterns, pausing behavior, jargon, sociolect, regional coloring, etc. (see <xref ref-type="bibr" rid="CIT25">Kinnunen &amp; Li, 2010</xref>; <xref ref-type="bibr" rid="CIT29">Künzel &amp; Alexander, 2014</xref>). No separation of linguistic or phonetic units is made, therefore, under the automatic approach. This is why <xref ref-type="bibr" rid="CIT22">Jessen (2008)</xref> classifies this type of automatic methods as <italic>holistic</italic>: “The distribution of the MFCCs over the entire course of the recording of a speaker is determined. (…) no segmentation of the speech stream into different linguistic categories, such as consonants, vowels or syllables is performed” (p. 699).<xref ref-type="fn" rid="NOTE12">12</xref> </p>
		<p>According to what has just been explained, we hypothesized that the cepstral features in which this ASR system is based would be strongly gene-dependent, as they depend largely on anatomical–physiological foundations. Therefore, higher similarity values should be found in MZ twins (100% shared genes) than in DZ twins, in brothers (B) or in a reference population of unrelated speakers (US). To the best of our knowledge, this represents the first investigation into the voice characteristic of Spanish twins and non-twin siblings from an ASR perspective. Previous studies (<xref ref-type="bibr" rid="CIT49">San Segundo, 2010a</xref>; <xref ref-type="bibr" rid="CIT50">San Segundo, 2010b</xref>; <xref ref-type="bibr" rid="CIT51">San Segundo, 2012</xref>; <xref ref-type="bibr" rid="CIT52">San Segundo, 2013a</xref>; and <xref ref-type="bibr" rid="CIT56">San Segundo &amp; Gómez-Vilda, 2013</xref>, <xref ref-type="bibr" rid="CIT57">2015</xref>) have tackled FSC of this set of twins and non-twin speakers from different points of view (mainly glottal analyses and formant trajectories).</p>
		<p>The most important conclusion that can be drawn from this analysis is that—as we have hypothesized—the similarity coefficients yielded by the automatic system Batvox<sup>TM</sup> decrease exactly as the kinship relationship of the speaker pairs decreases. In other words, the score sorting from largest to smallest resulted in the following scale of values: MZ &gt; DZ &gt; B &gt; US.</p>
		<p>In the introduction to this investigation we explained our reasons for sustaining the hypothesis that higher similarity values (hence worse recognition) would be found in MZ IP‑comparisons than in DZ IP‑comparisons. In turn, these speakers would be more similar than non-twin brothers (B) and the latter more similar than unrelated speakers (US). The justification for this lies in the fact that MZ twins share 100% of their genetic information and in general they also share educational and environmental backgrounds, while DZ twins share 50% of their genes but usually the same external influences as MZ twins. Sharing the same genetic information as DZ twins, brothers are supposed to share less environmental characteristics due to the age gap; and finally unrelated speakers share neither their genes nor their environmental background. This reasoning gives rise to the scale: MZ &gt; DZ &gt; B &gt; US, where “&gt;” means “more similar than”; for the aim of our investigation, at least in voice terms. To our knowledge, this is the first time that this hypothesis has been tested for an automatic system using the four types of speakers mentioned (MZ, DZ, B and US). The underlying idea behind this hypothesis is not foreign to phonetic studies, however. For instance, <xref ref-type="bibr" rid="CIT28">Künzel (2010, p. 251)</xref> sustains that “the more similar the geometry of two vocal tracts is, the more similar will be the respective similarity coefficients, or LRs” and that “this problem is particularly relevant to related speakers, most extremely for identical (MZ) twins” (<xref ref-type="bibr" rid="CIT28">Künzel, 2010, p. 251</xref>). As a matter of fact, the issue of how the comparison of very similar speakers can affect the recognition performance of an automatic system has been investigated before, albeit almost exclusively using MZ twins as participants. </p>
		<p>When comparing our results with previous findings by other authors who have tested the same automatic system with twins, we have been able to corroborate the widely reported finding in the ASR literature that some speakers are simply more easily identified than others. The 9.9% EER in our study corresponding to two out of 12 MZ twins who would be misidentified is comparable to the 11% EER in <xref ref-type="bibr" rid="CIT28">Künzel (2010)</xref>, indicating that confusion or non-distinction between twins occurred. The issue of the “striking performance inhomogeneities among speakers within a population” was already raised by <xref ref-type="bibr" rid="CIT9">Doddington et al. (1998)</xref> and we already referred to it in the glottal analysis described in <xref ref-type="bibr" rid="CIT54">San Segundo (2014)</xref>, where some cases (16.6%) were found of speakers exhibiting large self-unlikeness (i.e., they were very dissimilar when comparing their first and second recording session). </p>
		<p>To sum up, testing the performance of an ASR system using identical twins implies a strong reduction of inter-speaker variation and, as explained by <xref ref-type="bibr" rid="CIT28">Künzel (2010)</xref>, this is a most challenging task since “the <italic>a priori </italic>chances for a target voice to be very similar to the reference voice is much larger than within a set of unrelated speakers” (p. 269). We agree with him in considering that “a system that identifies an identical twin without falsely accepting the other twin is probably fit for use in the forensic environment” (<xref ref-type="bibr" rid="CIT28">Künzel, 2010, p. 274</xref>). The explanation for this seems logical: the system works even when it is being tested in a disadvantageous situation, which could be compared with a situation where there is channel distortion or cross-language samples to compare. All these are challenging situations. However, a real case where twins’ voices ought to be compared is not the most frequent situation in a forensic setting, basically because of the low incidence of twin births (rate of identical twins is four per thousand; fraternal rate is 22.8 per thousand). Yet, the importance of investigating twins’ voices goes beyond this pragmatic view, i.e., it is relevant per se, regardless of how many real cases involve the comparison of twins. First, the comparative study of MZ and DZ twins can reveal the genetic influence of the parameters under study (see EEA, Section 1.1). Hence the importance of carrying out studies with both types of twins, not only MZ twins. The finding that certain voice parameters are genetically marked entails a good performance of any system that would be based on such parameters because the typical speakers for comparison would be usually genetically unrelated, which means that the system would be good at separating them. Second, the consideration of further types of kinship relationships, apart from MZ and DZ twins, such as non-twin siblings can help clarify certain under-researched issues, such as the interplay between genetic and environmental influences in voice. </p>
		<p>From the results of our investigation, we suggest that the cepstral parameters on which the automatic system Batvox<sup>TM</sup> is based are genetically influenced. It is well known that these features relate to the geometry of the vocal tract, so some physical similarity between twins is expected to be encoded in DNA. Yet, the different use and configuration of the vocal apparatus could be exploited by twins in different ways, which could leave a generous margin for IP variation (<xref ref-type="bibr" rid="CIT31">Loakes, 2006</xref>; <xref ref-type="bibr" rid="CIT38">Nolan &amp; Oh, 1996</xref>). These different usage preferences—more related to learned aspects than to inborn characteristics—might be the key to explain the two out of 12 twin cases that were misidentified by the system, accounting for the 9.9% EER. </p>
		<p>All in all, as a direction for future work, it has not been mentioned so far that neither the group of MZ twins nor the DZ twin group are homogenous as far as their genes are concerned. MZ twins can be monochorionic or dichorionic, depending on whether they share the same placenta or have two different placentas instead; they can also be monoamniotic or diamnotic, depending on whether they share the same amniotic sac or not. How this can affect the differences found between one twin pair and another, as well as the influence of epigenetics in twin differences, has not been fully addressed in twin voice literature yet. For instance, the fact that spontaneous mutations tend to occur more often in dichorionic MZ twins makes them more likely to differ genetically than monochorionic MZ twins (see <xref ref-type="bibr" rid="CIT62">Stromswold, 2006</xref>). Whether the existence of different types of MZ twins affects their voice similarity or not is an open research question, which, in any case, would require specific DNA testing to obtain detailed information about the zygosity of the twin pairs. </p>
		<p>As regards epigenetics, future research focusing on twins’ voices should pay more attention to this concept, which we briefly introduced in Section 1.1. Although only two “forces” are typically mentioned in the twin literature to explain the (dis)similarities in twins voices, namely, genetic and environmental factors, the often-neglected third factor, i.e, epigenetics (which explains the alteration in the expression of specific genes caused by mechanisms other than changes in the underlying DNA sequence) may play an important role in our understanding of the striking dissimilarities found for some twin pairs. </p>
		</sec>
	</body>
	<back>
	<ack id="S6">
	<title>ACKNOWLEDGEMENTS</title>
	<p>This research has been possible thanks to a doctoral grant awarded to the first author by the Spanish Ministry of Education (<italic>Beca FPU-Programa Nacional de Formación de Profesorado Universitario, BOE 11-07-2009</italic>) and also thanks to a grant awarded by the IAFPA (International Association for Forensic Phonetics and Acoustics) to the project “Forensic comparison of Spanish twins and non-twin brothers”. </p>
	</ack>
		<fn-group id="S7">
		<title>NOTES</title>
			<fn id="NOTE1">
			<label>1</label>
				<p>From the EEA we can draw that the excess of similarity (for an investigated parameter) exhibited by MZ twins that is not present in DZ pairs must be due to genetic causes. Although we have taken advantage of this principle for our study, a strict application of the twin methodology would require the use of <italic>heritability estimates</italic> or <italic>concordance rates</italic>, in which the expected elevated similarity in MZs over DZs is often reported, depending on whether it is a continuous or a dichotomous trait (see <xref ref-type="bibr" rid="CIT63">Tomblin &amp; Buckwalter, 1998</xref>).</p>
			</fn>
			<fn id="NOTE2">
			<label>2</label>
				<p>Monozygotic twins (also called <italic>identical</italic>) develop from one zygote that splits and forms two embryos, while dizygotic (also called <italic>fraternal</italic>) develop from two separate eggs that are fertilized by two separate sperm cells (<xref ref-type="bibr" rid="CIT8">Del Abril Alonso et al., 2009, p. 90</xref>). Full brothers are male siblings with the same father and the same mother. </p>
			</fn>
			<fn id="NOTE3">
			<label>3</label>
				<p>A definition of Forensic Phonetics has been provided by different authors (e.g., <xref ref-type="bibr" rid="CIT22">Jessen, 2008</xref>; <xref ref-type="bibr" rid="CIT27">Künzel, 1994</xref>; <xref ref-type="bibr" rid="CIT37">Nolan, 1997</xref>; Rose, 2002). What all these definitions have in common is that they specify for the discipline of Phonetics the general definition of Forensics as the application of scientific knowledge to legal problems. <italic>Forensic Phonetics</italic> would then be the application of Phonetics aimed at solving any type of legal issue (see <xref ref-type="bibr" rid="CIT54">San Segundo, 2014</xref>). One of the most typical forensic cases where a phonetic expert is involved is one in which has to compare the voice of an offender (i.e., speech samples of an unknown speaker) with the voice of a suspect or several suspects (i.e., speech samples of known origin). It is widely accepted nowadays to refer to this kind of task as <italic>Forensic Speaker Comparison</italic> (FSC). Other possible tasks which a phonetician may be requested to perform for forensic purposes are described, for instance, in <xref ref-type="bibr" rid="CIT14">Foulkes and French (2012)</xref>.</p>
			</fn>
			<fn id="NOTE4">
			<label>4</label>
				<p><italic>Robustness</italic> is usually associated to a degradation factor, and could be defined as the reluctance of a system to lose performance when certain degradation factor is present. For our study, genetic similarity is seen as the degradation factor.</p>
			</fn>
			<fn id="NOTE5">
			<label>5</label>
				<p>While the typical question that a forensic phonetician has to answer in a FSC case is: “How much more likely the magnitude of the difference between samples is if they came from the same speaker than from different speakers?” (Rose, 2002, p. 89), in the case of siblings’ voices the question would have to be formulated in a slightly different way. For example, as pointed out by <xref ref-type="bibr" rid="CIT11">Feiser (2009)</xref>, “not uncommonly the question posed in court is whether a given unknown recording could have been spoken by the subject’s brother(s) instead of the subject himself. Other than being a possible legal strategy, this question suggests itself because siblings often have similar sounding voices” (2009, p. 1).</p>
			</fn>
			<fn id="NOTE6">
			<label>6</label>
				<p>Eventually, a cohort of 31 Spanish male speakers was used as background population (spontaneous conversation and high-quality recordings), coming from Batvox<sup>TM</sup> databases (see Section 2.2.1) because a minimum of 25 speakers is required using this system. However, the group of 12 unrelated speakers served to compare the matching scores of MZ, DZ and non-twin brothers with speakers without any type of genetic relationship.</p>
			</fn>
			<fn id="NOTE7">
			<label>7</label>
				<p>The importance of the same interlocutor is strongly linked to the theory of accommodation (<xref ref-type="bibr" rid="CIT16">Giles, Coupland &amp; Coupland, 1991</xref>). More recently, a fast-growing research line investigating convergence and imitation patterns in speech occurring between speakers in the course of conversational interactions (see e.g., <xref ref-type="bibr" rid="CIT39">Pardo, 2006</xref>; <xref ref-type="bibr" rid="CIT41">Pickering &amp; Garrod, 2004</xref>; <xref ref-type="bibr" rid="CIT64">Trouvain &amp; Truong, 2012</xref>), provides further evidence that speaker interlocutors actually converge in a number of phonetic features.</p>
			</fn>
			<fn id="NOTE8">
			<label>8</label>
				<p>More detailed information can be found in <xref ref-type="bibr" rid="CIT28">Künzel (2010, pp. 253–4)</xref> where he cites relevant bibliographic references in this field (<xref ref-type="bibr" rid="CIT10">Drygajlo, 2007</xref>; <xref ref-type="bibr" rid="CIT18">Gonzalez-Rodriguez, Fierrez-Aguilar, &amp; Ortega-Garcia, 2003</xref>; <xref ref-type="bibr" rid="CIT42">Przybocki, Martin, &amp; Le, 2007</xref>; <xref ref-type="bibr" rid="CIT44">Ramos, 2007</xref>). </p>
			</fn>
			<fn id="NOTE9">
			<label>9</label>
				<p>Note that, depending on the author followed, this type of recognition task could be named differently (e.g., <italic>verification task</italic>). Cf. <xref ref-type="bibr" rid="CIT3">Bimbot et al. (2004)</xref>.</p>
			</fn>
			<fn id="NOTE10">
			<label>10</label>
				<p>To avoid comparing a speaker with his sibling or conversational partner, at least in this first analysis which does not take into account the sibling effect, only the even members of each speaker pair were selected, both for the matches and non-matches. That is, only speakers 02, 04, 06, etc., were used in the analysis. Following the methodology described in <xref ref-type="bibr" rid="CIT28">Künzel (2010)</xref>, in order to facilitate this task, one member of the twin pairs was labeled <italic>red</italic> (the odd numbers) and the other member was labeled <italic>blue</italic> (the even numbers). <xref ref-type="fig" rid="F1">Figure 1</xref> shows the EER (0%) using the blue speakers. The same test was repeated using only the red speakers and a very similar EER was obtained (0.07%).</p>
			</fn>
			<fn id="NOTE11">
			<label>11</label>
				<p>The only difference between both curves rising to the left in <xref ref-type="fig" rid="F5">Figure 5</xref> is that one (yellow) compared first session of every speaker with first session of all other speakers, while the other (red) compared first session of every speaker with second session of all other speakers.</p>
			</fn>
			<fn id="NOTE12">
			<label>12</label>
				<p>As explained in <xref ref-type="bibr" rid="CIT22">Jessen (2008)</xref>, “as a means of smoothing the spectral shape and of making the outcome more realistic psycho-acoustically, the spectrum is then passed through a filterbank based on the non-linear Mel scale. The logarithms of the filter coefficients are transferred to the cepstrum by application of the Discrete Cosine Transform. The resulting vectors are now called cepstral coefficients” (p. 699).</p>
			</fn>
		</fn-group>
		<ref-list id="S8">
			<title>REFERENCES</title>
			<ref id="CIT1">
       <element-citation publication-type="software">
       <collab>Agnitio Voice trics</collab>
          <year>2013</year>
          <source>Batvox 4.1 Basic User Manual</source>
          <comment>Computer software</comment>
       </element-citation>
      </ref>
      <ref id="CIT2">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Ariyaeeinia</surname>
            <given-names>A.</given-names>
          </name>
          <name>
            <surname>Morrison</surname>
            <given-names>C.</given-names>
          </name>
          <name>
            <surname>Malegaonkar</surname>
            <given-names>A.</given-names>
          </name>
          <name>
            <surname>Black</surname>
            <given-names>S.</given-names>
          </name>
          </person-group>
          <year>2008</year>
          <article-title>A test of the effectiveness of speaker verification for differentiating between identical twins</article-title>
          <source>Science &amp; Justice</source>
          <volume>48</volume>
          <issue>4</issue>
          <fpage>182</fpage>
          <lpage>186</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1016/j.scijus.2008.02.002">http://dx.doi.org/10.1016/j.scijus.2008.02.002</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT3">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Bimbot</surname>
            <given-names>F.</given-names>
          </name>
          <name>
            <surname>Bonastre</surname>
            <given-names>J.-F.</given-names>
          </name>
          <name>
            <surname>Fredouille</surname>
            <given-names>C.</given-names>
          </name>
          <name>
            <surname>Gravier</surname>
            <given-names>G.</given-names>
          </name>
          <name>
            <surname>Magrin-Chagnolleau</surname>
            <given-names>I.</given-names>
          </name>
          <name>
            <surname>Meignier</surname>
            <given-names>S.</given-names>
          </name>
          <name>
            <surname>Reynolds</surname>
            <given-names>D. A.</given-names>
          </name>
          </person-group>
          <year>2004</year>
          <article-title>A tutorial on text-independent speaker verification</article-title>
          <source>EURASIP Journal on Advances in Signal Processing</source>
          <volume>4</volume>
          <fpage>1</fpage>
          <lpage>22</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi. org/10.1155/S1110865704310024">http://dx.doi. org/10.1155/S1110865704310024</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT4">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Brümmer</surname>
            <given-names>N.</given-names>
          </name>
          <name>
            <surname>du Preez</surname>
            <given-names>J.</given-names>
          </name>
          </person-group>
          <year>2006</year>
          <article-title>Application-independent evaluation of speaker detection</article-title>
          <source>Computer Speech &amp; Language</source>
          <volume>20</volume>
          <issue>2-3</issue>
          <fpage>230</fpage>
          <lpage>275</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1016/j.csl.2005.08.001">http://dx.doi.org/10.1016/j.csl.2005.08.001</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT5">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Campbell</surname>
            <given-names>W. M.</given-names>
          </name>
          <name>
            <surname>Campbell</surname>
            <given-names>J. P.</given-names>
          </name>
          <name>
            <surname>Reynolds</surname>
            <given-names>D. A</given-names>
          </name>
          <name>
            <surname>Singer</surname>
            <given-names>E.</given-names>
          </name>
          <name>
            <surname>Torres-Carrasquillo</surname>
            <given-names>P. A.</given-names>
          </name>
          </person-group>
          <year>2006</year>
          <article-title>Support vector machines for speaker and language recognition</article-title>
          <source>Computer Speech &amp; Language</source>
          <volume>20</volume>
          <issue>2-3</issue>
          <fpage>210</fpage>
          <lpage>229</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/doi:10.1016/j.csl.2005.06.003">http://dx.doi.org/doi:10.1016/j.csl.2005.06.003</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT6">
       <element-citation publication-type="book">
       <person-group person-group-type="author">
          <name>
            <surname>Charlet</surname>
            <given-names>D.</given-names>
          </name>
          <name>
            <surname>Lecha</surname>
            <given-names>V. P.</given-names>
          </name>
          </person-group>
          <person-group person-group-type="editor">
			  <name>
				  <surname>Filipe</surname>
				  <given-names>J.</given-names>
				</name>
				<name>
				  <surname>Coelhas</surname>
				  <given-names>H.</given-names>
				</name>
				<name>
				  <surname>Saramago</surname>
				  <given-names>M.</given-names>
				</name>
          </person-group>
          <year>2007</year>
          <chapter-title>Voice biometrics within the family: Trust, privacy and personalisation</chapter-title>
          <source>E-business and telecommunication networks: Second International Conference, ICETE 2005</source>
          <volume>3</volume>
          <fpage>93</fpage>
          <lpage>100</lpage>
          <publisher-name>Springer</publisher-name>
          <publisher-loc>Berlin</publisher-loc>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1007/978-3-540-75993-5_8">http://dx.doi.org/10.1007/978-3-540-75993-5_8</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT7">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Debruyne</surname>
            <given-names>F.</given-names>
          </name>
          <name>
            <surname>Decoster</surname>
            <given-names>W.</given-names>
          </name>
          <name>
            <surname>Van Gijsel</surname>
            <given-names>A.</given-names>
          </name>
          <name>
            <surname>Vercammen</surname>
            <given-names>J.</given-names>
          </name>
          </person-group>
          <year>2002</year>
          <article-title>Speaking fundamental frequency in monozygotic and dizygotic twins</article-title>
          <source>Journal of Voice</source>
          <volume>16</volume>
          <issue>4</issue>
          <fpage>466</fpage>
          <lpage>471</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1016/S0892-1997(02)00121-2">http://dx.doi.org/10.1016/S0892-1997(02)00121-2</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT8">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Del Abril Alonso</surname>
            <given-names>Á.</given-names>
          </name>
          <name>
            <surname>Ambrosio Flores</surname>
            <given-names>E.</given-names>
            </name>
            <name>
            <surname>Blas Calleja</surname>
            <given-names>M. d. R.</given-names>
          </name>
          <name>
            <surname>Caminero Gómez</surname>
            <given-names>Á.</given-names>
          </name>
          <name>
            <surname>García Lecumberri</surname>
            <given-names>C.</given-names>
          </name>
          <name>
            <surname>de Pablo González</surname>
            <given-names>J. M.</given-names>
          </name>
          </person-group>
          <year>2009</year>
          <source>Fundamentos de psicobiología</source>
          <publisher-name>Sanz y Torres</publisher-name>
          <publisher-loc>Madrid</publisher-loc>
       </element-citation>
      </ref>
      <ref id="CIT9">
       <element-citation publication-type="conf-proc">
       <person-group person-group-type="author">
          <name>
            <surname>Doddington</surname>
            <given-names>G.</given-names>
          </name>
          <name>
            <surname>Liggett</surname>
            <given-names>W.</given-names>
          </name>
          <name>
            <surname>Martin</surname>
            <given-names>A.</given-names>
          </name>
          <name>
            <surname>Przybocki</surname>
            <given-names>M.</given-names>
          </name>
          <name>
            <surname>Reynolds</surname>
            <given-names>D.</given-names>
          </name>
          </person-group>
          <year>1998</year>
          <article-title>SHEEP, GOATS, LAMBS and WOLVES: A statistical analysis of speaker performance in the NIST 1998 speaker recognition evaluation</article-title>
          <conf-name>Proceedings of the International Conference on Spoken Language (ICSLP '98)</conf-name>
          <comment>paper 0608</comment>
       </element-citation>
      </ref>
      <ref id="CIT10">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Drygajlo</surname>
            <given-names>A.</given-names>
          </name>
          </person-group>
          <year>2007</year>
          <article-title>Forensic automatic speaker recognition [Exploratory DSP]</article-title>
          <source>IEEE Signal Processing Magazine</source>
          <volume>24</volume>
          <issue>2</issue>
          <fpage>132</fpage>
          <lpage>135</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1109/MSP.2007.323278">http://dx.doi.org/10.1109/MSP.2007.323278</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT11">
       <element-citation publication-type="conf-proc">
       <person-group person-group-type="author">
          <name>
            <surname>Feiser</surname>
            <given-names>H. S.</given-names>
          </name>
          </person-group>
          <year>2009</year>
          <article-title>Acoustic similarities and differences in the voices of same-sex siblings</article-title>
          <conf-name>18th Annual Conference of the International Association for Forensic Phonetics and Acoustics (IAFPA)</conf-name>
          <conf-loc>Cambridge, UK</conf-loc>
          <comment>Paper presented</comment>
       </element-citation>
      </ref>
      <ref id="CIT12">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Felson</surname>
            <given-names>J.</given-names>
          </name>
          </person-group>
          <year>2014</year>
          <article-title>What can we learn from twin studies? A comprehensive evaluation of the equal environments assumption</article-title>
          <source>Social Science Research</source>
          <volume>43</volume>
          <fpage>184</fpage>
          <lpage>199</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/ doi:10.1016/j.ssresearch.2013.10.004">http://dx.doi.org/ doi:10.1016/j.ssresearch.2013.10.004</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT13">
       <element-citation publication-type="journal">
	   <person-group person-group-type="author">
          <name>
            <surname>Forrai</surname>
            <given-names>G.</given-names>
          </name>
          <name>
            <surname>Gordos</surname>
            <given-names>G.</given-names>
          </name>
          </person-group>
          <year>1983</year>
          <article-title>A new acoustic method for the discrimination of monozygotic and dizygotic twins</article-title>
          <source>Acta paediatrica Academiae Scientiarum Hungarica</source>
          <volume>24</volume>
          <issue>4</issue>
          <fpage>315</fpage>
          <lpage>322</lpage>
       </element-citation>
      </ref>
      <ref id="CIT14">
       <element-citation publication-type="book">
       <person-group person-group-type="author">
          <name>
            <surname>Foulkes</surname>
            <given-names>P.</given-names>
          </name>
          <name>
            <surname>French</surname>
            <given-names>J. P.</given-names>
          </name>
          </person-group>
          <person-group person-group-type="editor">
          <name>
				<surname>Tiersma</surname>
				<given-names>P.</given-names>
			</name>
			<name>
				<surname>Solan</surname>
				<given-names>L. M.</given-names>
			</name>
			</person-group>
          <year>2012</year>
          <chapter-title>Forensic speaker comparison: A linguistic-acoustic perspective</chapter-title>
          <source>Oxford handbook of language and law</source>
          <fpage>557</fpage>
          <lpage>572</lpage>
          <publisher-name>Oxford University Press</publisher-name>
			<publisher-loc>Oxford</publisher-loc>
			<comment><ext-link ext-link-type="uri" xlink:href="">http://dx.doi.org/10.1093/oxfordhb/9780199572120.013.0041</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT15">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Galton</surname>
            <given-names>F.</given-names>
          </name>
          </person-group>
          <year>1875</year>
          <article-title>The history of twins, as a criterion of the relative powers of nature and nurture (Rev. ed.)</article-title>
          <source>Journal of the Anthropological Institute of Great Britain and Ireland</source>
          <volume>5</volume>
          <fpage>391</fpage>
          <lpage>406</lpage>
       </element-citation>
      </ref>
      <ref id="CIT16">
       <element-citation publication-type="book">
       <person-group person-group-type="author">
          <name>
            <surname>Giles</surname>
            <given-names>H.</given-names>
          </name>
          <name>
            <surname>Coupland</surname>
            <given-names>J.</given-names>
          </name>
          <name>
            <surname>Coupland</surname>
            <given-names>N.</given-names>
          </name>
          </person-group>
          <year>1991</year>
          <source>Contexts of accommodation: Developments in applied sociolinguistics</source>
          <publisher-name>Cambridge University Press</publisher-name>
			<publisher-loc>Cambridge</publisher-loc>
       </element-citation>
      </ref>
      <ref id="CIT17">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Gómez-Vilda</surname>
            <given-names>P.</given-names>
          </name>
          <name>
            <surname>Fernández-Baillo</surname>
            <given-names>R.</given-names>
          </name>
          <name>
            <surname>Nieto</surname>
            <given-names>A.</given-names>
          </name>
          <name>
            <surname>Díaz</surname>
            <given-names>F.</given-names>
          </name>
          <name>
            <surname>Fernández-Camacho</surname>
            <given-names>F. J.</given-names>
          </name>
          <name>
            <surname>Rodellar</surname>
            <given-names>V.</given-names>
          </name>
          <name>
            <surname>Martínez</surname>
            <given-names>R.</given-names>
          </name>
          </person-group>
          <year>2007</year>
          <article-title>Evaluation of voice pathology based on the estimation of vocal fold biomechanical parameters</article-title>
          <source>Journal of Voice</source>
          <volume>21</volume>
          <issue>4</issue>
          <fpage>450</fpage>
          <lpage>476</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1016/j.jvoice.2006.01.008">http://dx.doi.org/10.1016/j.jvoice.2006.01.008</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT18">
       <element-citation publication-type="conf-proc">
       <person-group person-group-type="author">
          <name>
            <surname>Gonzalez-Rodriguez</surname>
            <given-names>J.</given-names>
          </name>
          <name>
            <surname>Fierrez-Aguilar</surname>
            <given-names>J.</given-names>
          </name>
          <name>
            <surname>Ortega-Garcia</surname>
            <given-names>J.</given-names>
          </name>
          </person-group>
          <year>2003</year>
          <article-title>Forensic identification reporting using automatic speaker recognition systems</article-title>
          <conf-name>Proceedings of the 2003 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP '03)</conf-name>
          <volume>2</volume>
          <fpage>93</fpage>
          <lpage>96</lpage>
       </element-citation>
      </ref>
      <ref id="CIT19">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Gonzalez-Rodriguez</surname>
            <given-names>J.</given-names>
          </name>
          <name>
            <surname>Rose</surname>
            <given-names>P.</given-names>
          </name>
          <name>
            <surname>Ramos</surname>
            <given-names>D.</given-names>
          </name>
          <name>
            <surname>Toledano</surname>
            <given-names>D. T.</given-names>
          </name>
          <name>
            <surname>Ortega-Garcia</surname>
            <given-names>J.</given-names>
          </name>
          </person-group>
          <year>2007</year>
          <article-title>Emulating DNA: Rigorous quantification of evidential weight in transparent and testable forensic speaker recognition</article-title>
          <source>IEEE Transactions on Audio, Speech, and Language Processing</source>
          <volume>15</volume>
          <issue>7</issue>
          <fpage>2104</fpage>
          <lpage>2115</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi. org/10.1109/TASL.2007.902747">http://dx.doi. org/10.1109/TASL.2007.902747</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT20">
       <element-citation publication-type="conf-proc">
       <person-group person-group-type="author">
          <name>
            <surname>Homayounpour</surname>
            <given-names>M. M.</given-names>
          </name>
          <name>
            <surname>Chollet</surname>
            <given-names>G.</given-names>
          </name>
          </person-group>
          <year>1995</year>
          <article-title>Discrimination of voices of twins and siblings for speaker verification</article-title>
          <conf-name>Proceedings of the 4th European Conference on Speech Communication and Technology (EUROSPEECH 1995)</conf-name>
          <fpage>345</fpage>
          <lpage>348</lpage>
       </element-citation>
      </ref>
      <ref id="CIT21">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Jain</surname>
            <given-names>A. K.</given-names>
          </name>
          <name>
            <surname>Prabhakar</surname>
            <given-names>S.</given-names>
          </name>
          <name>
            <surname>Pankanti</surname>
            <given-names>S.</given-names>
          </name>
          </person-group>
          <year>2002</year>
          <article-title>On the similarity of identical twin fingerprints</article-title>
          <source>Pattern Recognition</source>
          <volume>35</volume>
          <issue>11</issue>
          <fpage>2653</fpage>
          <lpage>2663</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1016/S0031-3203(01)00218-7">http://dx.doi.org/10.1016/S0031-3203(01)00218-7</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT22">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Jessen</surname>
            <given-names>M.</given-names>
          </name>
          </person-group>
          <year>2008</year>
          <article-title>Forensic phonetics</article-title>
          <source>Language and Linguistics Compass</source>
          <volume>2</volume>
          <issue>4</issue>
          <fpage>671</fpage>
          <lpage>711</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi. org/10.1111/j.1749-818X.2008.00066.x">http://dx.doi. org/10.1111/j.1749-818X.2008.00066.x</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT23">
       <element-citation publication-type="conf-proc">
       <person-group person-group-type="author">
          <name>
            <surname>Kenny</surname>
            <given-names>P.</given-names>
          </name>
          <name>
            <surname>Boulianne</surname>
            <given-names>G.</given-names>
          </name>
          <name>
            <surname>Ouellet</surname>
            <given-names>P.</given-names>
          </name>
          <name>
            <surname>Dumouchel</surname>
            <given-names>P.</given-names>
          </name>
          </person-group>
          <year>2005</year>
          <article-title>Factor analysis simplified</article-title>
          <conf-name>Proceedings of the 2005 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP '05)</conf-name>
          <volume>1</volume>
          <fpage>637</fpage>
          <lpage>640</lpage>
       </element-citation>
      </ref>
      <ref id="CIT24">
       <element-citation publication-type="conf-proc">
       <person-group person-group-type="author">
          <name>
            <surname>Kim</surname>
            <given-names>K.</given-names>
          </name>
          </person-group>
          <year>2010</year>
          <article-title>Automatic speaker identification of Korean male twins</article-title>
          <conf-name>19th Annual Conference of the International Association for Forensic Phonetics and Acoustics (IAFPA)</conf-name>
          <conf-loc>Trier</conf-loc>
          <comment>Paper presented</comment>
       </element-citation>
      </ref>
      <ref id="CIT25">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Kinnunen</surname>
            <given-names>T.</given-names>
          </name>
          <name>
            <surname>Li</surname>
            <given-names>H.</given-names>
          </name>
          </person-group>
          <year>2010</year>
          <article-title>An overview of text-independent speaker recognition: From features to supervectors</article-title>
          <source>Speech Communication</source>
          <volume>52</volume>
          <issue>1</issue>
          <fpage>12</fpage>
          <lpage>40</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1016/j. specom.2009.08.009">http://dx.doi.org/10.1016/j. specom.2009.08.009</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT26">
       <element-citation publication-type="journal">
	   <person-group person-group-type="author">
          <name>
            <surname>Kong</surname>
            <given-names>A. W.-K.</given-names>
          </name>
          <name>
            <surname>Zhang</surname>
            <given-names>D.</given-names>
          </name>
          <name>
            <surname>Lu</surname>
            <given-names>G.</given-names>
          </name>
          </person-group>
          <year>2006</year>
          <article-title>A study of identical twins' palmprints for personal verification</article-title>
          <source>Pattern Recognition</source>
          <volume>39</volume>
          <issue>11</issue>
          <fpage>2149</fpage>
          <lpage>2156</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/doi:10.1016/j. patcog.2006.04.035">http://dx.doi.org/doi:10.1016/j. patcog.2006.04.035</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT27">
       <element-citation publication-type="conf-proc">
       <person-group person-group-type="author">
          <name>
            <surname>Künzel</surname>
            <given-names>H. J.</given-names>
          </name>
          </person-group>
          <year>1994</year>
          <article-title>Current approaches to forensic speaker recognition</article-title>
          <conf-name>Proceedings of the ESCA Workshop on Automatic Speaker Recognition, Identification, and Verification</conf-name>
          <fpage>135</fpage>
          <lpage>141</lpage>
       </element-citation>
      </ref>
      <ref id="CIT28">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Künzel</surname>
            <given-names>H. J.</given-names>
          </name>
          </person-group>
          <year>2010</year>
          <article-title>Automatic speaker recognition of identical twins</article-title>
          <source>International Journal of Speech, Language and the Law</source>
          <volume>17</volume>
          <issue>2</issue>
          <fpage>251</fpage>
          <lpage>277</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1558/ijsll.v17i2.251">http://dx.doi.org/10.1558/ijsll.v17i2.251</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT29">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Künzel</surname>
            <given-names>H. J.</given-names>
          </name>
          <name>
            <surname>Alexander</surname>
            <given-names>P.</given-names>
          </name>
          </person-group>
          <year>2014</year>
          <article-title>Forensic automatic speaker recognition with degraded and enhanced speech</article-title>
          <source>Journal of the Audio Engineering Society</source>
          <volume>62</volume>
          <issue>4</issue>
          <fpage>244</fpage>
          <lpage>253</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi. org/10.17743/jaes.2014.0014">http://dx.doi. org/10.17743/jaes.2014.0014</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT30">
       <element-citation publication-type="book">
       <person-group person-group-type="author">
          <name>
            <surname>Labov</surname>
            <given-names>W.</given-names>
          </name>
          </person-group>
          <person-group person-group-type="editor">
          <name>
            <surname>Labov</surname>
            <given-names>W.</given-names>
          </name>
          </person-group>
          <year>1972</year>
          <chapter-title>The transformation of experience in the narrative syntax</chapter-title>
          <source>Language in the inner city: Studies in the Black English Vernacular</source>
          <fpage>354</fpage>
          <lpage>396</lpage>
          <publisher-name>University of Philadelphia Press</publisher-name>
		<publisher-loc>Philadelphia, PA</publisher-loc>
       </element-citation>
      </ref>
      <ref id="CIT31">
       <element-citation publication-type="thesis">
       <person-group person-group-type="author">
          <name>
            <surname>Loakes</surname>
            <given-names>D.</given-names>
          </name>
          </person-group>
          <year>2006</year>
          <publisher-name>University of Melbourne</publisher-name>
          <source>A forensic phonetic investigation into the speech patterns of identical and non-identical twins</source>
          <comment>Doctoral dissertation</comment>
       </element-citation>
      </ref>
      <ref id="CIT32">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Martino</surname>
            <given-names>D.</given-names>
          </name>
          <name>
            <surname>Loke</surname>
            <given-names>Y. J.</given-names>
          </name>
          <name>
            <surname>Gordon</surname>
            <given-names>L.</given-names>
          </name>
          <name>
            <surname>Ollikainen</surname>
            <given-names>M.</given-names>
          </name>
          <name>
            <surname>Cruickshank</surname>
            <given-names>M. N.</given-names>
          </name>
          <name>
            <surname>Saffery</surname>
            <given-names>R.</given-names>
          </name>
          <name>
            <surname>Craig</surname>
            <given-names>J. M.</given-names>
          </name>
          </person-group>
          <year>2013</year>
          <article-title>Longitudinal, genomescale analysis of DNA methylation in twins from birth to 18 months of age reveals rapid epigenetic change in early life and pair-specific effects of discordance</article-title>
          <source>Genome Biology</source>
          <volume>14</volume>
          <issue>5</issue>
          <fpage>R42</fpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1186/gb-2013-14-5-r42">http://dx.doi.org/10.1186/gb-2013-14-5-r42</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT33">
       <element-citation publication-type="thesis">
       <person-group person-group-type="author">
          <name>
            <surname>Meuwly</surname>
            <given-names>D.</given-names>
          </name>
          </person-group>
          <year>2001</year>
          <source>Reconnaissance de locuteurs en sciences forensiques: l'apport d'une approche automatique</source>
          <comment>PhD dissertation</comment>
          <publisher-name>University of Laussane</publisher-name>
       </element-citation>
      </ref>
      <ref id="CIT34">
       <element-citation publication-type="book">
       <person-group person-group-type="author">
          <name>
            <surname>Morrison</surname>
            <given-names>G. S.</given-names>
          </name>
          </person-group>
          <person-group person-group-type="editor">
			<name>
				<surname>Freckelton</surname>
				<given-names>I.</given-names>
			</name>
			<name>
				<surname>Selby</surname>
				<given-names>H.</given-names>
			</name>
			</person-group>
          <year>2010</year>
          <chapter-title>Forensic voice comparison</chapter-title>
          <source>Expert evidence</source>
          <comment>Chapter 99</comment>
          <publisher-name>Thomson Reuters</publisher-name>
			<publisher-loc>Sydney</publisher-loc>
       </element-citation>
      </ref>
      <ref id="CIT35">
       <element-citation publication-type="conf-proc">
       <person-group person-group-type="author">
          <name>
            <surname>Morrison</surname>
            <given-names>G. S.</given-names>
          </name>
          <name>
            <surname>Kinoshita</surname>
            <given-names>Y.</given-names>
          </name>
          </person-group>
          <year>2008</year>
          <article-title>Automatic-type calibration of traditionally derived likelihood ratios: Forensic analysis of Australian English /o/ formant trajectories</article-title>
          <conf-name>Proceedings of the 9th INTERSPEECH Conference</conf-name>
          <fpage>1501</fpage>
          <lpage>1504</lpage>
       </element-citation>
      </ref>
      <ref id="CIT36">
       <element-citation publication-type="book">
       <person-group person-group-type="author">
          <name>
            <surname>Nolan</surname>
            <given-names>F.</given-names>
          </name>
          </person-group>
          <year>1983</year>
          <source>The phonetic bases of speaker recognition</source>
          <publisher-name>Cambridge University Press</publisher-name>
			<publisher-loc>Cambridge</publisher-loc>
       </element-citation>
      </ref>
      <ref id="CIT37">
       <element-citation publication-type="book">
       <person-group person-group-type="author">
          <name>
            <surname>Nolan</surname>
            <given-names>F.</given-names>
          </name>
          </person-group>
           <person-group person-group-type="editor">
          <name>
            <surname>Hardcastle</surname>
            <given-names>W. J.</given-names>
          </name>
          <name>
            <surname>Laver</surname>
            <given-names>J.</given-names>
          </name>
          </person-group>
          <year>1997</year>
          <chapter-title>Speaker recognition and forensic phonetics</chapter-title>
          <source>The handbook of phonetic sciences</source>
          <fpage>744</fpage>
          <lpage>767</lpage>
          <publisher-name>Blackwell</publisher-name>
			<publisher-loc>Oxford</publisher-loc>
			<comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1111/b.9780631214786.1999.00025.x">http://dx.doi.org/10.1111/b.9780631214786.1999.00025.x</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT38">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Nolan</surname>
            <given-names>F.</given-names>
          </name>
          <name>
            <surname>Oh</surname>
            <given-names>T.</given-names>
          </name>
          </person-group>
          <year>1996</year>
          <article-title>Identical twins, different voices</article-title>
          <source>International Journal of Speech Language and the Law</source>
          <volume>3</volume>
          <issue>1</issue>
          <fpage>39</fpage>
          <lpage>49</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1558/ijsll.v3i1.39">http://dx.doi.org/10.1558/ijsll.v3i1.39</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT39">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Pardo</surname>
            <given-names>J. S.</given-names>
          </name>
          </person-group>
          <year>2006</year>
          <article-title>On phonetic convergence during conversational interaction</article-title>
          <source>The Journal of the Acoustical Society of America</source>
          <volume>119</volume>
          <issue>4</issue>
          <fpage>2382</fpage>
          <lpage>2393</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1121/1.2178720">http://dx.doi.org/10.1121/1.2178720</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT40">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Philips</surname>
            <given-names>T.</given-names>
          </name>
          </person-group>
          <year>2008</year>
          <article-title>The role of methylation in gene expression</article-title>
          <source>Nature Education</source>
          <volume>1</volume>
          <issue>1</issue>
          <fpage>116</fpage>
       </element-citation>
      </ref>
      <ref id="CIT41">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Pickering</surname>
            <given-names>M. J.</given-names>
          </name>
          <name>
            <surname>Garrod</surname>
            <given-names>S.</given-names>
          </name>
          </person-group>
          <year>2004</year>
          <article-title>Toward a mechanistic psychology of dialogue</article-title>
          <source>Behavioral and Brain Sciences</source>
          <volume>27</volume>
          <issue>2</issue>
          <fpage>169</fpage>
          <lpage>190</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1017/S0140525X04000056">http://dx.doi.org/10.1017/S0140525X04000056</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT42">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Przybocki</surname>
            <given-names>M. A.</given-names>
          </name>
          <name>
            <surname>Martin</surname>
            <given-names>A. F.</given-names>
          </name>
          <name>
            <surname>Le</surname>
            <given-names>A. N.</given-names>
          </name>
          </person-group>
          <year>2007</year>
          <article-title>NIST speaker recognition evaluations utilizing the Mixer corpora-2004, 2005, 2006</article-title>
          <source>IEEE Transactions on Audio, Speech, and Language Processing</source>
          <volume>15</volume>
          <issue>7</issue>
          <fpage>1951</fpage>
          <lpage>1959</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1109/ TASL.2007.902489">http://dx.doi.org/10.1109/ TASL.2007.902489</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT43">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Przybyla</surname>
            <given-names>B. D.</given-names>
          </name>
          <name>
            <surname>Horii</surname>
            <given-names>Y.</given-names>
          </name>
          <name>
            <surname>Crawford</surname>
            <given-names>M. H.</given-names>
          </name>
          </person-group>
          <year>1992</year>
          <article-title>Vocal fundamental frequency in a twin sample: Looking for a genetic effect</article-title>
          <source>Journal of Voice</source>
          <volume>6</volume>
          <issue>3</issue>
          <fpage>261</fpage>
          <lpage>266</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/ doi:10.1016/S0892-1997(05)80151-1">http://dx.doi.org/ doi:10.1016/S0892-1997(05)80151-1</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT44">
       <element-citation publication-type="thesis">
       <person-group person-group-type="author">
          <name>
            <surname>Ramos</surname>
            <given-names>D.</given-names>
          </name>
          </person-group>
          <year>2007</year>
          <source>Forensic evaluation of the evidence using automatic speaker recognition systems</source>
          <publisher-name>Universidad Autónoma de Madrid</publisher-name>
          <comment>Doctoral dissertation</comment>
          <comment>Retrieved from <ext-link ext-link-type="uri" xlink:href="http://hdl.handle.net/10486/1774">http://hdl.handle.net/10486/1774</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT45">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Reynolds</surname>
            <given-names>D. A.</given-names>
          </name>
          <name>
            <surname>Quatieri</surname>
            <given-names>T. F.</given-names>
          </name>
          <name>
            <surname>Dunn</surname>
            <given-names>R. B.</given-names>
          </name>
          </person-group>
          <year>2000</year>
          <article-title>Speaker verification using adapted Gaussian mixture models</article-title>
          <source>Digital Signal Processing</source>
          <volume>10</volume>
          <issue>1-3</issue>
          <fpage>19</fpage>
          <lpage>41</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1006/ dspr.1999.0361">http://dx.doi.org/10.1006/ dspr.1999.0361</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT46">
       <element-citation publication-type="journal">
	   <person-group person-group-type="author">
          <name>
            <surname>Reynolds</surname>
            <given-names>D. A.</given-names>
          </name>
          <name>
            <surname>Rose</surname>
            <given-names>R. C.</given-names>
          </name>
          </person-group>
          <year>1995</year>
          <article-title>Robust text-independent speaker identification using Gaussian mixture speaker models</article-title>
          <source>IEEE Transactions on Speech and Audio Processing</source>
          <volume>3</volume>
          <issue>1</issue>
          <fpage>72</fpage>
          <lpage>83</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1109/89.365379">http://dx.doi.org/10.1109/89.365379</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT47">
       <element-citation publication-type="book">
       <person-group person-group-type="author">
          <name>
            <surname>Rose</surname>
            <given-names>P.</given-names>
          </name>
          </person-group>
          <year>2002</year>
          <source>Forensic speaker identification</source>
          <publisher-name>Taylor &amp; Francis</publisher-name>
		<publisher-loc>London</publisher-loc>
       </element-citation>
      </ref>
      <ref id="CIT48">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Rose</surname>
            <given-names>P.</given-names>
          </name>
          </person-group>
          <year>2006</year>
          <article-title>Technical forensic speaker recognition: Evaluation, types and testing of evidence</article-title>
          <source>Computer Speech &amp; Language</source>
          <volume>20</volume>
          <issue>2-3</issue>
          <fpage>159</fpage>
          <lpage>191</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/doi:10.1016/j.csl.2005.07.003">http://dx.doi.org/doi:10.1016/j.csl.2005.07.003</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT49">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>San Segundo</surname>
            <given-names>E.</given-names>
          </name>
          </person-group>
          <year>2010</year>
          <article-title>Parametric representations of the formant trajectories of Spanish vocalic sequences for likelihood-ratio-based forensic voice comparison</article-title>
          <source>The Journal of the Acoustical Society of America</source>
          <volume>128</volume>
          <issue>4</issue>
          <fpage>2394</fpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1121/1.3508586">http://dx.doi.org/10.1121/1.3508586</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT50">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>San Segundo</surname>
            <given-names>E.</given-names>
          </name>
          </person-group>
          <year>2010b</year>
          <article-title>Variación inter e intralocutor: Parámetros acústicos segmentales que caracterizan fonéticamente a tres hermanos</article-title>
          <source>Interlingüística</source>
          <volume>21</volume>
          <fpage>352</fpage>
          <lpage>363</lpage>
       </element-citation>
      </ref>
      <ref id="CIT51">
       <element-citation publication-type="conf-proc">
       <person-group person-group-type="author">
         <name>
            <surname>San Segundo</surname>
            <given-names>E.</given-names>
          </name>
          </person-group>
          <year>2012</year>
          <article-title>Glottal source parameters for forensic voice comparison: An approach to voice quality in twins' voices</article-title>
          <conf-name>21st Annual Conference of the International Association for Forensic Phonetics and Acoustics (IAFPA)</conf-name>
          <conf-loc>Santander, Spain</conf-loc>
          <comment>Paper presented</comment>
       </element-citation>
      </ref>
      <ref id="CIT52">
       <element-citation publication-type="conf-proc">
       <person-group person-group-type="author">
          <name>
            <surname>San Segundo</surname>
            <given-names>E.</given-names>
          </name>
          </person-group>
          <year>2013</year>
          <article-title>Guess who is laughing: A perceptual experiment on twin and non-twin siblings' identification</article-title>
          <conf-name>31st International Conference AESLA (Asociación Española de Lingüística Aplicada</conf-name>
          <conf-loc>San Cristóbal de La Laguna</conf-loc>
          <publisher-name>Universidad de La Laguna</publisher-name>
           <comment>Paper presented</comment>
       </element-citation>
      </ref>
      <ref id="CIT53">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>San Segundo</surname>
            <given-names>E.</given-names>
          </name>
          </person-group>
          <year>2013</year>
          <article-title>A phonetic corpus of Spanish male twins and siblings: Corpus design and forensic application</article-title>
          <source>Procedia-Social and Behavioral Sciences</source>
          <volume>95</volume>
          <fpage>59</fpage>
          <lpage>67</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/ doi:10.1016/j.sbspro.2013.10.622">http://dx.doi.org/ doi:10.1016/j.sbspro.2013.10.622</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT54">
       <element-citation publication-type="thesis">
       <person-group person-group-type="author">
          <name>
            <surname>San Segundo</surname>
            <given-names>E.</given-names>
          </name>
          </person-group>
          <year>2014</year>
          <source>Forensic speaker comparison of Spanish twins and non-twin siblings: A phonetic-acoustic analysis of formant trajectories in vocalic sequences, glottal source parameters and cepstral characteristics</source>
          <publisher-name>Consejo Superior de Investigaciones Científicas-Universidad Internacional Menéndez Pelayo</publisher-name>
          <publisher-name>Universidad Internacional Menéndez Pelayo</publisher-name>
			<publisher-loc>Spain</publisher-loc>
          <comment>PhD thesis</comment>
       </element-citation>
      </ref>
      <ref id="CIT55">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
         <name>
            <surname>San Segundo</surname>
            <given-names>E.</given-names>
          </name>
          </person-group>
          <year>2015</year>
          <article-title>Forensic speaker comparison of Spanish twins and non-twin siblings: A phonetic-acoustic analysis of formant trajectories in vocalic sequences, glottal source parameters and cepstral characteristics (Thesis abstract)</article-title>
          <source>International Journal of Speech Language and the Law</source>
          <volume>22</volume>
          <issue>2</issue>
          <fpage>249</fpage>
          <lpage>253</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1558/ijsll.v22i2.28821">http://dx.doi.org/10.1558/ijsll.v22i2.28821</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT56">
       <element-citation publication-type="book">
       <person-group person-group-type="author">
          <name>
            <surname>San Segundo</surname>
            <given-names>E.</given-names>
          </name>
          <name>
            <surname>Gómez-Vilda</surname>
            <given-names>P.</given-names>
          </name>
          </person-group>
           <person-group person-group-type="editor">
          <name>
            <surname>Manfredi</surname>
            <given-names>C.</given-names>
          </name>
          </person-group>
          <year>2013</year>
          <chapter-title>Voice biometrical match of twin and non-twin siblings</chapter-title>
          <source>Models and analysis of vocal emissions for biomedical applications: 8th International Workshop, Firenze, Italy, 2013</source>
          <fpage>253</fpage>
          <lpage>256</lpage>
          <comment>Retrieved from <ext-link ext-link-type="uri" xlink:href="http://digital.casalini.it/9788866554707">http://digital.casalini.it/9788866554707</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT57">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>San Segundo</surname>
            <given-names>E.</given-names>
          </name>
          <name>
            <surname>Gómez-Vilda</surname>
            <given-names>P.</given-names>
          </name>
          </person-group>
          <year>2015</year>
          <article-title>Evaluating the forensic importance of glottal source features through the voice analysis of twins and non-twin siblings</article-title>
          <source>Language and Law/Linguagem e Direito</source>
          <volume>1</volume>
          <issue>2</issue>
          <fpage>22</fpage>
          <lpage>41</lpage>
       </element-citation>
      </ref>
      <ref id="CIT58">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Sataloff</surname>
            <given-names>R. T.</given-names>
          </name>
          </person-group>
          <year>1995</year>
          <article-title>Genetics of the voice</article-title>
          <source>Journal of Voice</source>
          <volume>9</volume>
          <issue>1</issue>
          <fpage>16</fpage>
          <lpage>19</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/doi:10.1016/S0892-1997(05)80218-8">http://dx.doi.org/doi:10.1016/S0892-1997(05)80218-8</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT59">
       <element-citation publication-type="conf-proc">
       <person-group person-group-type="author">
          <name>
            <surname>Scheffer</surname>
            <given-names>N.</given-names>
          </name>
          <name>
            <surname>Bonastre</surname>
            <given-names>J.-F.</given-names>
          </name>
          <name>
            <surname>Ghio</surname>
            <given-names>A.</given-names>
          </name>
          <name>
            <surname>Teston</surname>
            <given-names>B.</given-names>
          </name>
          </person-group>
          <year>2004</year>
          <article-title>Gémellité et reconnaissance automatique du locuteur</article-title>
          <conf-name>Actes des XXV Journées d'Étude sur la Parole (JEP)</conf-name>
          <fpage>445</fpage>
          <lpage>448</lpage>
       </element-citation>
      </ref>
      <ref id="CIT60">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Segal</surname>
            <given-names>N. L.</given-names>
          </name>
          </person-group>
          <year>1993</year>
          <article-title>Implications of twin research for legal issues involving young twins</article-title>
          <source>Law and Human Behavior</source>
          <volume>17</volume>
          <issue>1</issue>
          <fpage>43</fpage>
          <lpage>58</lpage>
       </element-citation>
      </ref>
      <ref id="CIT61">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Srihari</surname>
            <given-names>S.</given-names>
          </name>
          <name>
            <surname>Huang</surname>
            <given-names>C.</given-names>
          </name>
          <name>
            <surname>Srinivasan</surname>
            <given-names>H.</given-names>
          </name>
          </person-group>
          <year>2008</year>
          <article-title>On the discriminability of the handwriting of twins</article-title>
          <source>Journal of Forensic Sciences</source>
          <volume>53</volume>
          <issue>2</issue>
          <fpage>430</fpage>
          <lpage>446</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1111/j.1556-4029.2008.00682.x">http://dx.doi.org/10.1111/j.1556-4029.2008.00682.x</ext-link></comment>
          </element-citation>
      </ref>
      <ref id="CIT62">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Stromswold</surname>
            <given-names>K.</given-names>
          </name>
          </person-group>
          <year>2006</year>
          <article-title>Why aren't identical twins linguistically identical? Genetic, prenatal and postnatal factors</article-title>
          <source>Cognition</source>
          <volume>101</volume>
          <issue>2</issue>
          <fpage>333</fpage>
          <lpage>384</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/doi:10.1016/j.cognition.2006.04.007">http://dx.doi.org/doi:10.1016/j.cognition.2006.04.007</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT63">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Tomblin</surname>
            <given-names>J. B.</given-names>
          </name>
          <name>
            <surname>Buckwalter</surname>
            <given-names>P. P.</given-names>
          </name>
          </person-group>
          <year>1998</year>
          <article-title>Heritability of poor language achievement among twins</article-title>
          <source>Journal of Speech, Language, and Hearing Research</source>
          <volume>41</volume>
          <fpage>188</fpage>
          <lpage>189</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/ doi:10.1044/jslhr.4101.188">http://dx.doi.org/ doi:10.1044/jslhr.4101.188</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT64">
       <element-citation publication-type="conf-proc">
       <person-group person-group-type="author">
          <name>
            <surname>Trouvain</surname>
            <given-names>J.</given-names>
          </name>
          <name>
            <surname>Truong</surname>
            <given-names>K. P.</given-names>
          </name>
          </person-group>
          <year>2012</year>
          <article-title>Convergence of laughter in conversational speech: Effects of quantity, temporal alignment and imitation</article-title>
          <conf-name>International Symposium on Imitation and Convergence in Speech</conf-name>
          <conf-loc>Aix-en-Provence, France</conf-loc>
          <comment>Paper presented</comment>
       </element-citation>
      </ref>
      <ref id="CIT65">
       <element-citation publication-type="book">
       <person-group person-group-type="author">
          <name>
            <surname>van Leeuwen</surname>
            <given-names>D. A.</given-names>
          </name>
          <name>
            <surname>Brümmer</surname>
            <given-names>N.</given-names>
          </name>
          </person-group>
          <person-group person-group-type="editor">
          <name>
            <surname>Müller</surname>
            <given-names>C.</given-names>
          </name>
          </person-group>
          <year>2007</year>
          <chapter-title>An introduction to application-independent evaluation of speaker recognition systems</chapter-title>
          <source>Speaker classification I: Fundamentals, features, and methods</source>
          <fpage>330</fpage>
          <lpage>353</lpage>
          <publisher-name>Springer-Verlag</publisher-name>
			<publisher-loc>Heidelberg</publisher-loc>
			<comment><ext-link ext-link-type="uri" xlink:href="http:// dx.doi.org/10.1007/978-3-540-74200-5_19">http:// dx.doi.org/10.1007/978-3-540-74200-5_19</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT66">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Van Lierde</surname>
            <given-names>K. M.</given-names>
          </name>
          <name>
            <surname>Vinck</surname>
            <given-names>B.</given-names>
          </name>
          <name>
            <surname>De Ley</surname>
            <given-names>S.</given-names>
          </name>
          <name>
            <surname>Clement</surname>
            <given-names>G.</given-names>
          </name>
          <name>
            <surname>Van Cauwenberge</surname>
             <given-names>P.</given-names>
          </name>
          </person-group>
          <year>2005</year>
          <article-title>Genetics of vocal quality characteristics in monozygotic twins: A multiparameter approach</article-title>
          <source>Journal of Voice</source>
          <volume>19</volume>
          <issue>4</issue>
          <fpage>511</fpage>
          <lpage>518</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1016/j.jvoice.2004.10.005">http://dx.doi.org/10.1016/j.jvoice.2004.10.005</ext-link></comment>
       </element-citation>
      </ref>
      <ref id="CIT67">
       <element-citation publication-type="conf-proc">
       <person-group person-group-type="author">
          <name>
            <surname>Weirich</surname>
            <given-names>M.</given-names>
          </name>
          <name>
            <surname>Lancia</surname>
            <given-names>L.</given-names>
          </name>
          </person-group>
          <year>2011</year>
          <article-title>Perceived auditory similarity and its acoustic correlates in twins and unrelated speakers</article-title>
          <conf-name>Proceedings of the 17th International Congress of Phonetic Sciences (ICPhS 17-Hong Kong)</conf-name>
          <fpage>2118</fpage>
          <lpage>2121</lpage>
       </element-citation>
      </ref>
      <ref id="CIT68">
       <element-citation publication-type="journal">
       <person-group person-group-type="author">
          <name>
            <surname>Wolf</surname>
            <given-names>J. J.</given-names>
          </name>
          </person-group>
          <year>1972</year>
          <article-title>Efficient acoustic parameters for speaker recognition</article-title>
          <source>The Journal of the Acoustical Society of America</source>
          <volume>51</volume>
          <issue>6B</issue>
          <fpage>2044</fpage>
          <lpage>2056</lpage>
          <comment><ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1121/1.1913065">http://dx.doi.org/10.1121/1.1913065</ext-link></comment>
       </element-citation>
      </ref>
		</ref-list>
	</back>
</article>