MIT 6.300 — Lecture 16
Speech
Remark. During an actual class of 6.300, this lecture features videos and demonstrations that I do not have access to while writing these notes. Thus, these notes may exclude some 6.300 content.
§ Lecture: Speech as a System
Speech can be modeled as the output of an LTI system.

The input is roughly a periodic impulse train generated by the vocal cords. The behavior of the filter depends on the shape of the speaker's throat and nasal cavities.
§ Lecture: Vowels and Formants
Different vowels have different harmonic content.

Recall that the harmonic content pictured above is just the product . The source is easy to analyze: it is a train of impulses, where the spacing between impulses characterizes the perceived pitch of the vowel.
The filter has resonant frequencies known as formants. These resonant frequencies create peaks in the graph of at the position of each formant. It is the shape of —or, equivalently, the values of its formants—that gives a vowel sound its identity. Importantly,
Claim. A vowel's identity is not determined by its exact, local harmonic structure! Rather, it is determined by , which is represented by the “spectral envelope” of its rough, global harmonic structure.
You can play with the following (warning: AI-generated) demo to see how formant values (F1, F2, etc.) affect vowel quality.
There are more than two formants, but the first two are mostly sufficient for uniquely defining vowel quality.
F1 is roughly higher when the tongue is lower, rather than higher.
F2 is roughly lower when the tongue is retracted, rather than advanced.
In contrast, formants F4 and beyond control subtler qualities such as vocal timbre.
An interesting corollary of all of this is that higher-pitched vowels are harder to identify—they give a much blurrier picture of the spectral envelope because the train of impulses in is spaced much further apart. (Hear it for yourself in the demo!)
Remark. When whispering, vocal cords are no longer in use, so now represents a turbulent stream of air with no consistent, periodic harmonic structure at all. The result is still the same, though: the formants of reveal themselves in the output , determining vowel identity.