Hold one note. One vowel, steady, no vibrato. It sounds like one sound. It is not one sound.
Say your folds are opening and closing 220 times a second. They do not do it smoothly. They slam. Each cycle is a sharp puff of air with a hard edge on it. A sharp edge repeating 220 times a second is not one frequency. It is a whole ladder of them at once. There is 220. There is 440, twice as fast. There is 660, three times. Then 880, 1100, 1320, on up. Each rung is a whole-number multiple of the bottom rung. Each one is quieter than the last. That ladder is the harmonic series (the set of frequencies that are whole-number multiples of the lowest one).1
Nobody invented it. It falls out of the physics of anything that repeats a pattern over and over. Think of a plucked string, a column of air in a pipe, a pair of vocal folds. Repeating at a rate makes energy at that rate and at every multiple of it. You get the stack for free the moment you get the buzz.
You do not hear the rungs. Your brain fuses the whole stack into one sound sitting at one pitch. The only thing that survives the fusing is a flavour. That flavour has a name. It is timbre, and it is all about which rungs are loud and which are quiet.2 A flute and a violin can play the same A. Same bottom rung, same 440. You tell them apart at once, before you have finished thinking about it. The recipe above the bottom rung is different. The violin is loud in rungs the flute barely makes.
So singing is not really about making a note. The note is the easy part. Singing is about shaping a recipe.
Where resonance stops being mysticism
The buzz leaving the folds is already tilted. Left alone, each doubling of frequency arrives about 12 decibels quieter than the one below it. That is why the raw source sounds thin and buzzy rather than rich.3 If that tilted ladder were the whole story, everyone would sound like the same kazoo at different pitches.
It is not the whole story. The ladder has to get out through a tube, and tubes have opinions.
Push a child on a swing. The swing has one rate it wants to go at, set by the length of its chains. Nothing you do changes that rate. Push at the wrong moment and you fight it. Push in time with the rate it already has. Each small push adds to the last one, and the swing climbs. You did not choose the frequency. You gave energy at many timings. The swing took only one of them.
A tube of air does the same thing with sound. Air in a tube has a set of rates at which a push keeps adding to itself instead of cancelling itself. Those rates are set by the tube’s length and shape.4 Feed a whole ladder of frequencies into it. The tube boosts the ones near its preferred rates and lets the rest pass through quiet. That is all resonance is: a container that prefers some frequencies over others because of its shape. No energy is made. Nothing mystical is happening. The tube is a filter you can reshape with your tongue.
Now hold the two ideas side by side. The whole of vocal tone is the gap between them. The harmonic ladder is set by your pitch. It moves up and down as a whole when you change note. What the tube prefers is set by your shape. They stay put while you change note. Two separate grids. A harmonic is loud when a rung of the moving ladder lands under a peak of the fixed shape. Sing the same vowel up a scale. You can hear rungs light up and go dark as they pass under the peaks. It is like a train passing under streetlights.
Trained singers use this on purpose. Classical singers learn to narrow the space just above the larynx. That pulls several of the tube’s upper peaks into a cluster near 3 kHz. The result is a bump in the recipe called the singer’s formant. It sits in a band where an orchestra happens to be quiet.5 The singer is not louder than eighty instruments. The singer is loud in the one narrow place the eighty instruments left empty. That is a channel-allocation trick, solved by ear, centuries before anyone had the word for it.
Why two people singing the same note are never the same note
If timbre is the recipe, then a voice is a recipe that belongs to a body. Your tube has its own length, its own set of bends, its own soft-tissue lining. No other tube boosts the same rungs by the same amounts. Vocal tract length alone is one of the biggest sources of difference between speakers. The upper resonances are the ones above about 2.5 kHz. They are ruled by parts of your throat whose shape barely changes as you talk. That makes them close to a fixed signature.6
That is why two singers on the same pitch are two clearly different objects. They are not making different notes. They are making the same ladder through different tubes. You are hearing the tubes.
Here is the test that proves how much of the signal you can throw away and still have the person. A normal phone call carries about 300 Hz to 3,400 Hz and drops everything outside that band.7 Most adult speaking voices have a fundamental below 300 Hz. So when your mother calls, the bottom rung of her ladder never reaches you. That rung is the note she is on. It is not made quieter. It is gone.
She still sounds exactly like your mother, and you still hear her voice as low.
Two things do that. The first is that pitch does not live in the fundamental. Your hearing reads pitch from the spacing of the surviving rungs. So a ladder that goes 400, 500, 600, 700 is heard as a voice at 100. The 100 was never sent. That effect is called the missing fundamental. It is why a small phone speaker can give you a bass voice it cannot make.8 The second is that identity was never in the bottom rung anyway. It is in the pattern of relative loudness across the rungs that remain. The phone company kept exactly the band where that pattern is richest. The 300–3,400 Hz choice was not random. It was picked as the narrowest band that still lets a listener know who is talking.7
Which lands back on the idea this piece runs on. What crosses the wire is not her voice. What crossed the room was not air. Both times it is a pattern. A pattern survives losing most of itself, as long as the shape of what is left stays the same.
References
-
Harmonic series (music), Wikipedia. ↩
-
Christopher Dobrian, Harmonic/Overtone Series, Computer Music Pedagogy, University of California, Irvine.
“The relationship of the amplitude of the fundamental frequency component to the amplitude of the harmonics created by an instrument is referred to as the spectral envelope.” ↩
-
Spectral slope, Glottopedia. The −12 dB per octave figure is an idealisation derived from triangular source pulses; real voices vary around it, and breathy or falsetto phonation runs steeper, nearer −18 dB per octave. ↩
-
Acoustic resonance, Wikipedia. ↩
-
Johan Sundberg, Level and Center Frequency of the Singer’s Formant, Journal of Voice 15(2), 2001. The peak near 3 kHz arises from a clustering of the third, fourth and fifth resonances, produced by narrowing the epilaryngeal tube against a widened pharynx. ↩
-
Acoustic cues for the recognition of self-voice and other-voice, PubMed Central.
“Frequencies higher than 2500 Hz … are greatly related to the anatomy of one’s laryngeal cavity, whose anatomical configuration varies between speakers but virtually remains unchanged during articulation of different vowels, and therefore carry individual specificity.” ↩
-
ITU-T Recommendation G.712, Transmission performance characteristics of pulse code modulation channels — the standard that band-limits narrowband telephony to 300–3,400 Hz. The band was chosen as the minimum that preserves both intelligibility and recognition of the speaker. ↩ ↩2
-
Missing fundamental, Wikipedia. ↩