Two stages make a human voice. The useful thing is that they barely depend on each other. One makes a raw buzz. The other decides what the buzz turns into. Change either one and you can leave the other alone.
Acousticians call this the source–filter model. It is Gunnar Fant’s, laid out in Acoustic Theory of Speech Production in 1960.1 It is not a beginner’s shortcut. It is the model the field really uses. It holds for normal and disordered speech, and it works for any language.2 Everything in the practice half of this piece leans on it. So it is worth stating with care once.
The source: a valve that times itself
Sitting in your throat is the larynx (the voice box, a small cage of cartilage you can feel from the outside as the bump some people call an Adam’s apple). Stretched across the inside of it are the vocal folds. They are two small bands of tissue that meet in the middle. The gap between them is the glottis.
Here is the part almost everyone gets wrong. The folds are not muscles firing once per vibration. When you sing that A above middle C, they are opening and closing 440 times a second. No muscle in your body can fire 440 times a second on command. No nerve is sending 440 separate orders. Nothing is timing them at all.
What really happens is this. You close the folds. Your lungs and the muscles around them raise the subglottal pressure (the air pressure below the folds) until it wins. The folds are blown apart. Air rushes through the narrow gap. Two things then pull them shut again. One is their own elasticity, like a stretched rubber band pulling back. The other is a drop in pressure between them, caused by the air moving fast through the narrow gap. That is the Bernoulli effect (in a moving fluid, where the speed goes up, the sideways pressure goes down). The folds slam together. Flow stops. Pressure rebuilds underneath. They blow apart again.
That loop is the myoelastic-aerodynamic theory of voice production, named for its two parts: muscle-and-elastic on one side, moving air on the other. Janwillem van den Berg published it in the Journal of Speech and Hearing Research in 1958. It came from a paper he had given at the Chicago voice conference the year before. It has been the standard account ever since.3
One honest correction, because the Bernoulli story is told with more confidence than it has earned. Bernoulli suction is real and it helps. But modern models put most of the energy transfer somewhere else: in the changing shape of the gap through a cycle. The glottis is funnel-shaped one way while it is opening and the other way while it is closing. So the air pressure inside it is higher on the way open than on the way shut. That lopsidedness feeds energy into the tissue and keeps the oscillation alive against friction. Ingo Titze worked the physics out in 1988. The lopsided shape is doing more of the work than the suction is.4
Either way, the headline stands: the larynx is a self-oscillating valve. Steady flow in, chopped flow out. Direct current to alternating current. Nothing schedules the chopping.
The right thing to compare it to is not any part of a body. It is a reed. The reed in a clarinet, or the petal valve in a two-stroke engine, does exactly this. It is a bendy flap in a stream of air. It beats against its own seat because the air makes it. There is no camshaft and no timing chain anywhere in the system. Blow steadily and the flapping sorts itself out. The frequency is set by how long, how heavy, and how stiff the reed is, and by nothing else.
Two results follow directly.
A voice is cheap to keep going. Nothing has to fire per cycle. So the metabolic cost of phonating is close to the cost of breathing out slowly. This is why you can talk for hours. It is also why the tiredness from real singing shows up in the torso rather than the throat. The costly part is managing the pressure, not making the buzz.
Pitch belongs to the reed, not to effort. How fast the folds cycle depends on their length, mass and tension. Longer and heavier means slower. Shorter and tighter means faster. Adult male vocal folds are longer and heavier after puberty. That is most of why speaking fundamental frequency (the base rate of the cycle, written F0) sits near 116 Hz in adult men and near 205 Hz in adult women. That is roughly a 1.7-to-1 ratio.5 Treat those as centres of wide, overlapping ranges, not as facts about any one person.
On its own the source is not impressive. Pulses of air at 116 times a second sound like a kazoo. Buzzy, harmonically rich, and carrying almost no information about what you meant.
The filter: a tube you reshape with your tongue
Above the folds is a tube about 17 centimetres long in an adult. It is the pharynx, the mouth, the nose branching off, and the tongue sitting in the middle of all of it like a movable wall.
Any tube of air has frequencies it prefers. Blow across a bottle and it answers with one note. The note depends only on the shape of the bottle. Fill the bottle halfway with water and the note changes. You changed the shape of the air column, not the blowing. Your vocal tract does the same thing to the buzz coming from below. Some frequencies in that buzz get boosted. Others get damped down.
The boosted bands are called formants. They are numbered up from the lowest: F1, F2, F3. The first two are more or less the whole vowel system. Most often F1 and F2 alone are enough to tell which vowel you said.6 F1 tracks how open your jaw is. It climbs for an open vowel like “ah”. F2 tracks how far forward your tongue sits. It climbs for a front vowel like “ee”.6 Peterson and Barney measured this at Bell Labs in 1952, across 76 speakers and ten vowels. The plot of F1 against F2 they made is still the map the field uses.7
The engine comparison carries all the way through here. This is the half of it that is exactly right. A tuned exhaust pipe on a two-stroke does not make any pulses. It picks which ones resonate. Same engine, different pipe, different note. Your throat is the pipe. Unlike the engine’s, yours changes shape several times a second while you use it.
You do not make vowels
This is the reframe, and it is worth slowing down for.
It feels like you have a stock of vowel sounds and you make them one at a time. You do not. You make one buzz, a single, dumb, steady buzz. Then you change the shape of a tube while it passes through. “Ee” and “ah” are not two sounds you know how to make. They are one sound, filtered two ways. The difference between them lives only in where your tongue is.
Say “ee” and slide slowly to “ah” without stopping the sound. Nothing at the source changed. Your folds went on doing the same thing at the same rate. Everything you heard change was the tube.
There are two clean proofs of this, and you have already run one of them today.
A whisper has no source at all. When you whisper, your folds do not vibrate. There is no fundamental frequency, because there is no oscillation to have one. All you have is turbulent air hissing through a narrow glottis: noise, no pitch, no periodicity.8 By the substance model of voice, a whisper should be impossible to follow, because the voice has been switched off. Instead it is close to fully clear. The formants live on in the noise. The formants were carrying the words.8 You can hold a whole conversation with the source deleted. That is as direct a proof as physics ever gives you that the information was never in the source.
There are two whispers, and only one of them is that proof. A soft whisper is air through a shaped tube. The folds stay out of the way. A stage whisper is the opposite. You squeeze the larynx on purpose so the hiss will carry. That squeeze puts work back into the source, which is the thing the whisper was supposed to delete. Use the quiet one if you want to hear the filter. The loud one is acting.
Helium proves the other half. Everyone has heard the party trick. Almost everyone explains it wrong. Helium does not raise your pitch. Your folds are heavier and longer than the gas around them. They go on cycling at very nearly the rate they always did. What changes is the speed of sound. It is about 972 m/s in helium against 343 in air, because the gas is so much lighter.910 A tube’s resonances scale with the speed of sound inside it. So every formant jumps up by roughly the same factor. Your vocal tract sounds like a much smaller vocal tract. Measured in real conditions the shift is not perfectly even. It is non-linear. The resonance bands smear out badly, in the extreme case fourteen times wider than in air.11 But the direction is clear. Helium moves the filter and leaves the source where it was.
That is the whole machine. A reed that times itself, and a pipe you reshape with your tongue. Pitch belongs to the reed. Loudness tracks the pressure underneath it. Vowel and timbre belong to the pipe.
Those three controls come apart in the physics. They get badly tangled in practice. Untangling them is most of what voice training turns out to be.
References
-
Gunnar Fant, Acoustic Theory of Speech Production (The Hague: Mouton, 1960). Publisher preview · Semantic Scholar record ↩
-
Johan Sundberg and colleagues, The Gunnar Fant Legacy in the Study of Vocal Acoustics.
“It gives a highly accurate description of the speech signal and … explains how vowels and consonants get their acoustic properties. It is general, language-independent and valid for both normal and disordered speech.” ↩
-
Janwillem van den Berg, Myoelastic-Aerodynamic Theory of Voice Production, Journal of Speech and Hearing Research 1(3), 227–244 (1958). DOI 10.1044/jshr.0103.227. Based on a paper given at the Chicago International Voice Conference, May 1957. ↩
-
Ingo R. Titze, The physics of small-amplitude oscillation of the vocal folds, Journal of the Acoustical Society of America 83(4), 1536–1552 (1988). ↩
-
Holmberg, Hillman & Perkell (1988), as summarised in Average Speaking Frequencies: F0 Norms by Age, Sex, and Hormonal Status, Voice Science. Adult male mean about 116 Hz (range about 93–135 Hz); adult female mean about 205 Hz (range about 162–238 Hz). ↩
-
Formant, Wikipedia.
“Most often the two first formants, F1 and F2, are sufficient to identify the vowel.” ↩ ↩2
-
Gordon E. Peterson & Harold L. Barney, Control Methods Used in a Study of the Vowels, Journal of the Acoustical Society of America 24(2), 175–184 (1952). DOI 10.1121/1.1906875. Ten vowels, 76 speakers, at Bell Telephone Laboratories. ↩
-
Hamid Reza Sharifzadeh, Ian Vince McLoughlin & Martin J. Russell, A Comprehensive Vowel Space for Whispered Speech, Journal of Voice 26(2), e49–e56 (2012). ↩ ↩2
-
The Physics Classroom, The Speed of Sound; see also Speed of Sound in Air, The Physics Factbook. 343 m/s is the standard dry-air figure at 20 °C. ↩
-
E. O. Belcher & S. Hatlestad, Formant frequencies, bandwidths, and Qs in helium speech, Journal of the Acoustical Society of America 74(2), 428–432 (1983). DOI 10.1121/1.389758.
“Formant bandwidths in helium speech increased as much as 14 times their corresponding bandwidths in normal speech.” ↩