Does vowel recognition relate to pitch?
Acoustics Virtually Everywhere (AVE), Maurer et al., 2020


> Close

Background


The spectral peak patterns, the estimated formant patterns and the spectral envelopes of vowel sounds are ambiguous, often representing two or three different vowel qualities if fo is varied.


Illustration part I, sound pairs of single speakers

Sound pairs in terms of comparisons of sounds of two adjacent vowels produced by single speakers are shown for which both the calculated vowel-related formant frequencies as well as the vowel-related spectral envelopes are similar. All sound series shown here represent a short extract of Maurer et al (2019). For reference, see below.

The below comparisons relate to formant frequency patterns F1–F2–F3 for sounds of front vowels and to F1–F2 for sounds of back vowels. In parallel, the spectral envelopes were compared for frequency ranges up to 3 kHz for adults and up to 3.5 kHz for children. – For the comparisons, the upper fo range of the sounds is limited to 400 Hz for men, 450 Hz for women, and 500 Hz for children, and the fo differences of a sound pair is limited to approximately 1 octave. These limitations are to reflect the range of the average, everyday speaking voice of women and children, and a range of chest and “mixed” voice for men.

Note: The ambiguity phenomenon can also be investigated in Klatt synthesis. Use the Klatt synthesiser given below the displayed sound spectra and perform synthesis with fo of the original natural sound as well as fo of the opposed sound. 


Series 1: Sounds of /e/ and /i/ produced by a single female speaker at fo of 167 and 310 Hz.

> Sounds

Series 2: Sounds of /ø/ and /y/ produced by a single female speaker at fo of 164 and 330 Hz.

> Sounds

Series 3: Sounds of /ɛ/ and /e/ produced by a single female speaker at fo of 181 and 412 Hz.

> Sounds

Series 4: Sounds of /ɛ/ and /ø/ produced by a single male speaker at fo of 127 and 240 Hz.

> Sounds

Series 5: Sounds of /a/ and /o/ produced by a single male speaker at fo of 161 and 317 Hz.

> Sounds

Series 6: Sounds of /o/ and /u/ produced by a single female speaker at fo of 215 and 401 Hz.

> Sounds



Illustration part II, sound triples of speakers of a given speaker group

Sound triples in terms of comparisons of sounds of adjacent and non-adjacent vowels produced by speakers of a given speaker group (men, women or children) are shown for which both the calculated vowel-related formant frequencies as well as the vowel-related spectral envelopes are similar. For details of comparison, see above. However, no limitations were set for fo variation.


Series 7: sounds of /ɛ/, /e/ and /i/ produced by a men at fo of 99, 263 and 539 Hz. (Note: For Klatt synthesis related to the sound of /ɛ/, use F1=500 Hz according to the peak in the vowel spectrum.)

> Sounds

Series 8: sounds of /a/, /o/ and /u/ produced by a men at fo of 107, 327 and 595 Hz.

> Sounds



Reference, details of vowel recognition test, and Klatt resynthesis


As mentioned, the series were previously shown in Maurer et al (2019). For details on a vowel recognition test of the natural sounds as well as of resynthesised sounds with a Klatt synthesiser (applying fo variation in resynthesis), see this publication. It also includes information how to interactively perform Klatt synthesis on the natural sounds for verification (see online presentation).


Maurer, D., Suter, H., d'Hereuse, C., & Dellwo, V. (2019). Formant Pattern and Spectral Shape Ambiguity of Vowel Sounds, and Related Phenomena of Vowel Acoustics-Exemplary Evidence. In INTERSPEECH (pp. 2368-2369).

> Interspeech paper
> Online presentation