Your voice is not a thing you send. It is a pattern you press into air that was already there. You do it with a shape you control, using muscles you have never once felt. That is the whole subject. Almost everything odd about voice comes out of it. Why a deep supported breath leaves you winded in thirty seconds. Why your voice cracks in the same place every time. Why every language on earth carries about the same amount of information per second, no matter how fast it sounds. And why the first word almost every human says is the sound their mouth makes when it is doing nothing much.
What follows is in two halves. Theory is how a voice is made and what it carries. Practice is what to do with it. The first sound you made. The training that makes speaking and singing stop costing so much. And the reason a machine that listens is a bigger change than it looks.
Theory
Nothing actually travels
Start somewhere that looks like a detour and is not.
When you sing, and someone across the room hears you, it feels obvious what happened. Something went from you to them. Air, maybe. Your voice, carried over.
That is not what happens. Nothing goes from you to them.
The air molecules in front of your mouth get shoved forward a tiny bit. They bump the molecules next to them and bounce back to about where they started. Those molecules bump the next ones and bounce back. And so on, all the way across the room. Every molecule ends up roughly where it began. What crosses the room is not the air. It is the pattern of the bumping.
Physics has a name for this kind of wave. Sound is a longitudinal wave (one where the material moves back and forth along the same line the wave is travelling, rather than side to side across it).1 Shove one end of a stretched-out Slinky and you can watch the whole thing happen slowly enough to follow. A squeeze of tightly packed coils runs down the length of the toy. Every coil just jiggles a little and stays put. The squeeze is the wave. The coils are the room.
Where the molecules crowd together, the pressure goes slightly up. That is a compression. Where they thin out behind the crowd, the pressure goes slightly down. That is a rarefaction. A sound wave is nothing but compressions and rarefactions chasing each other outward. That is why it is also called a pressure wave.1 Your eardrum does not answer to wind. It answers to a tiny, fast wobble in pressure, riding on top of the plain weight of the air.
You have already seen this idea at a much bigger scale. A wave in the ocean travels for thousands of miles. The water does not. Each parcel of water goes up, forward, down, and back. It traces a little circle and ends up almost right where it started.2 That is why a duck sitting on the water bobs up and down as a wave passes instead of getting carried off to Portugal. The wave carries energy across an ocean. It does not carry the ocean.
Now take it somewhere with no water and no air in it at all. That is where the idea gets sharp.
Sit in stopped traffic on a highway. Up ahead, one driver taps the brakes. The car behind brakes a half-second later, then the next, then the next. A band of stopped cars forms. It travels backwards down the road toward you at maybe fifteen miles an hour. Every car in it is either stopped or moving forwards. The jam is real. You can measure how fast it moves. You can work out where it will be in ten minutes. And it is made of nothing. There is no object called a traffic jam. There is only a pattern in the spacing of cars, moving the other way from the cars themselves.
That is the whole trick. It is worth holding onto, because voice is that trick used on a body.
How fast, and what the numbers mean
Sound in air moves at about 343 metres per second at 20 °C. That is roughly a kilometre every three seconds.3 That is why you count between the lightning and the thunder. Light gets to you more or less at once. So the delay is pure sound-travel time. Three seconds of counting is about a kilometre of distance. Fifteen seconds and the storm is a long way off.
The speed depends on the temperature and on what the gas is made of. It does not depend on how loud the sound is. A shout and a whisper arrive at the same moment. This matters more than it sounds like it does. It comes back later. It is the reason a lungful of helium does something odd to your voice.
Two numbers describe any pattern like this. Neither of them is a substance.
Frequency is how many compressions pass a fixed point each second, measured in hertz. That is what you hear as pitch. When you sing an A above middle C, 440 compressions leave your mouth every second.
Amplitude is how big the pressure swings are: how far the crowding and thinning stray from the normal pressure of the room. That is what you hear as loudness.
Both belong to a pattern. Neither belongs to the air. You can take the same air and press a whole new pattern into it. You do it all the time, without adding or removing a thing.
Why this is the load-bearing idea
Almost everything people believe about voice quietly assumes that voice is a substance. That you produce it, that it comes out, that some people were handed more of it, that it can be used up. The language is built that way and it drags the gut sense along behind it.
The physics says something else. The thing that matters is usually a pattern, not a substance. Your voice is not a thing you make and then send. It is a pattern you press into something that was already there, filling the room, doing nothing.
Once you take that in, several things stop being a mystery.
It explains why a voice can be huge without being forceful. You are not throwing air. You are setting up a wobble in air that is already there. How well you set it up matters far more than the effort behind it. Trained singers are not blowing harder than you. Most of them are blowing less hard. They shape the pattern better.
It explains why sound goes around corners and through doors. A substance would need a path. A pattern just needs a medium that keeps touching itself. The air in your hallway does.
It explains why a recording works at all. A microphone does not capture your voice. It captures the pressure pattern and turns it into a matching pattern of voltage. Later a speaker pushes that pattern back into different air, in a different room, years on. Nothing of the original crossed over. The pattern was the whole content. Copying the pattern copied everything.
And it sets up the question the rest of this piece is about. If voice is a pattern rather than a substance, then getting better at voice cannot mean making more of something. It has to mean shaping the pattern more exactly. Which raises the next question: what, physically, does the shaping?
A shape, and air pushed through it
Two stages make a human voice. The useful thing is that they barely depend on each other. One makes a raw buzz. The other decides what the buzz turns into. Change either one and you can leave the other alone.
Acousticians call this the source–filter model. It is Gunnar Fant’s, laid out in Acoustic Theory of Speech Production in 1960.4 It is not a beginner’s shortcut. It is the model the field really uses. It holds for normal and disordered speech, and it works for any language.5 Everything in the practice half of this piece leans on it. So it is worth stating with care once.
The source: a valve that times itself
Sitting in your throat is the larynx (the voice box, a small cage of cartilage you can feel from the outside as the bump some people call an Adam’s apple). Stretched across the inside of it are the vocal folds. They are two small bands of tissue that meet in the middle. The gap between them is the glottis.
Here is the part almost everyone gets wrong. The folds are not muscles firing once per vibration. When you sing that A above middle C, they are opening and closing 440 times a second. No muscle in your body can fire 440 times a second on command. No nerve is sending 440 separate orders. Nothing is timing them at all.
What really happens is this. You close the folds. Your lungs and the muscles around them raise the subglottal pressure (the air pressure below the folds) until it wins. The folds are blown apart. Air rushes through the narrow gap. Two things then pull them shut again. One is their own elasticity, like a stretched rubber band pulling back. The other is a drop in pressure between them, caused by the air moving fast through the narrow gap. That is the Bernoulli effect (in a moving fluid, where the speed goes up, the sideways pressure goes down). The folds slam together. Flow stops. Pressure rebuilds underneath. They blow apart again.
That loop is the myoelastic-aerodynamic theory of voice production, named for its two parts: muscle-and-elastic on one side, moving air on the other. Janwillem van den Berg published it in the Journal of Speech and Hearing Research in 1958. It came from a paper he had given at the Chicago voice conference the year before. It has been the standard account ever since.6
One honest correction, because the Bernoulli story is told with more confidence than it has earned. Bernoulli suction is real and it helps. But modern models put most of the energy transfer somewhere else: in the changing shape of the gap through a cycle. The glottis is funnel-shaped one way while it is opening and the other way while it is closing. So the air pressure inside it is higher on the way open than on the way shut. That lopsidedness feeds energy into the tissue and keeps the oscillation alive against friction. Ingo Titze worked the physics out in 1988. The lopsided shape is doing more of the work than the suction is.7
Either way, the headline stands: the larynx is a self-oscillating valve. Steady flow in, chopped flow out. Direct current to alternating current. Nothing schedules the chopping.
The right thing to compare it to is not any part of a body. It is a reed. The reed in a clarinet, or the petal valve in a two-stroke engine, does exactly this. It is a bendy flap in a stream of air. It beats against its own seat because the air makes it. There is no camshaft and no timing chain anywhere in the system. Blow steadily and the flapping sorts itself out. The frequency is set by how long, how heavy, and how stiff the reed is, and by nothing else.
Two results follow directly.
A voice is cheap to keep going. Nothing has to fire per cycle. So the metabolic cost of phonating is close to the cost of breathing out slowly. This is why you can talk for hours. It is also why the tiredness from real singing shows up in the torso rather than the throat. The costly part is managing the pressure, not making the buzz.
Pitch belongs to the reed, not to effort. How fast the folds cycle depends on their length, mass and tension. Longer and heavier means slower. Shorter and tighter means faster. Adult male vocal folds are longer and heavier after puberty. That is most of why speaking fundamental frequency (the base rate of the cycle, written F0) sits near 116 Hz in adult men and near 205 Hz in adult women. That is roughly a 1.7-to-1 ratio.8 Treat those as centres of wide, overlapping ranges, not as facts about any one person.
On its own the source is not impressive. Pulses of air at 116 times a second sound like a kazoo. Buzzy, harmonically rich, and carrying almost no information about what you meant.
The filter: a tube you reshape with your tongue
Above the folds is a tube about 17 centimetres long in an adult. It is the pharynx, the mouth, the nose branching off, and the tongue sitting in the middle of all of it like a movable wall.
Any tube of air has frequencies it prefers. Blow across a bottle and it answers with one note. The note depends only on the shape of the bottle. Fill the bottle halfway with water and the note changes. You changed the shape of the air column, not the blowing. Your vocal tract does the same thing to the buzz coming from below. Some frequencies in that buzz get boosted. Others get damped down.
The boosted bands are called formants. They are numbered up from the lowest: F1, F2, F3. The first two are more or less the whole vowel system. Most often F1 and F2 alone are enough to tell which vowel you said.9 F1 tracks how open your jaw is. It climbs for an open vowel like “ah”. F2 tracks how far forward your tongue sits. It climbs for a front vowel like “ee”.9 Peterson and Barney measured this at Bell Labs in 1952, across 76 speakers and ten vowels. The plot of F1 against F2 they made is still the map the field uses.10
The engine comparison carries all the way through here. This is the half of it that is exactly right. A tuned exhaust pipe on a two-stroke does not make any pulses. It picks which ones resonate. Same engine, different pipe, different note. Your throat is the pipe. Unlike the engine’s, yours changes shape several times a second while you use it.
You do not make vowels
This is the reframe, and it is worth slowing down for.
It feels like you have a stock of vowel sounds and you make them one at a time. You do not. You make one buzz, a single, dumb, steady buzz. Then you change the shape of a tube while it passes through. “Ee” and “ah” are not two sounds you know how to make. They are one sound, filtered two ways. The difference between them lives only in where your tongue is.
Say “ee” and slide slowly to “ah” without stopping the sound. Nothing at the source changed. Your folds went on doing the same thing at the same rate. Everything you heard change was the tube.
There are two clean proofs of this, and you have already run one of them today.
A whisper has no source at all. When you whisper, your folds do not vibrate. There is no fundamental frequency, because there is no oscillation to have one. All you have is turbulent air hissing through a narrow glottis: noise, no pitch, no periodicity.11 By the substance model of voice, a whisper should be impossible to follow, because the voice has been switched off. Instead it is close to fully clear. The formants live on in the noise. The formants were carrying the words.11 You can hold a whole conversation with the source deleted. That is as direct a proof as physics ever gives you that the information was never in the source.
Helium proves the other half. Everyone has heard the party trick. Almost everyone explains it wrong. Helium does not raise your pitch. Your folds are heavier and longer than the gas around them. They go on cycling at very nearly the rate they always did. What changes is the speed of sound. It is about 972 m/s in helium against 343 in air, because the gas is so much lighter.312 A tube’s resonances scale with the speed of sound inside it. So every formant jumps up by roughly the same factor. Your vocal tract sounds like a much smaller vocal tract. Measured in real conditions the shift is not perfectly even. It is non-linear. The resonance bands smear out badly, in the extreme case fourteen times wider than in air.13 But the direction is clear. Helium moves the filter and leaves the source where it was.
That is the whole machine. A reed that times itself, and a pipe you reshape with your tongue. Pitch belongs to the reed. Loudness tracks the pressure underneath it. Vowel and timbre belong to the pipe.
Those three controls come apart in the physics. They get badly tangled in practice. Untangling them is most of what voice training turns out to be.
One note is a stack of notes
Hold one note. One vowel, steady, no vibrato. It sounds like one sound. It is not one sound.
Say your folds are opening and closing 220 times a second. They do not do it smoothly. They slam. Each cycle is a sharp puff of air with a hard edge on it. A sharp edge repeating 220 times a second is not one frequency. It is a whole ladder of them at once. There is 220. There is 440, twice as fast. There is 660, three times. Then 880, 1100, 1320, on up. Each rung is a whole-number multiple of the bottom rung. Each one is quieter than the last. That ladder is the harmonic series (the set of frequencies that are whole-number multiples of the lowest one).14
Nobody invented it. It falls out of the physics of anything that repeats a pattern over and over. Think of a plucked string, a column of air in a pipe, a pair of vocal folds. Repeating at a rate makes energy at that rate and at every multiple of it. You get the stack for free the moment you get the buzz.
You do not hear the rungs. Your brain fuses the whole stack into one sound sitting at one pitch. The only thing that survives the fusing is a flavour. That flavour has a name. It is timbre, and it is all about which rungs are loud and which are quiet.15 A flute and a violin can play the same A. Same bottom rung, same 440. You tell them apart at once, before you have finished thinking about it. The recipe above the bottom rung is different. The violin is loud in rungs the flute barely makes.
So singing is not really about making a note. The note is the easy part. Singing is about shaping a recipe.
Where resonance stops being mysticism
The buzz leaving the folds is already tilted. Left alone, each doubling of frequency arrives about 12 decibels quieter than the one below it. That is why the raw source sounds thin and buzzy rather than rich.16 If that tilted ladder were the whole story, everyone would sound like the same kazoo at different pitches.
It is not the whole story. The ladder has to get out through a tube, and tubes have opinions.
Push a child on a swing. The swing has one rate it wants to go at, set by the length of its chains. Nothing you do changes that rate. Push at the wrong moment and you fight it. Push in time with the rate it already has. Each small push adds to the last one, and the swing climbs. You did not choose the frequency. You gave energy at many timings. The swing took only one of them.
A tube of air does the same thing with sound. Air in a tube has a set of rates at which a push keeps adding to itself instead of cancelling itself. Those rates are set by the tube’s length and shape.17 Feed a whole ladder of frequencies into it. The tube boosts the ones near its preferred rates and lets the rest pass through quiet. That is all resonance is: a container that prefers some frequencies over others because of its shape. No energy is made. Nothing mystical is happening. The tube is a filter you can reshape with your tongue.
Now hold the two ideas side by side. The whole of vocal tone is the gap between them. The harmonic ladder is set by your pitch. It moves up and down as a whole when you change note. What the tube prefers is set by your shape. They stay put while you change note. Two separate grids. A harmonic is loud when a rung of the moving ladder lands under a peak of the fixed shape. Sing the same vowel up a scale. You can hear rungs light up and go dark as they pass under the peaks. It is like a train passing under streetlights.
Trained singers use this on purpose. Classical singers learn to narrow the space just above the larynx. That pulls several of the tube’s upper peaks into a cluster near 3 kHz. The result is a bump in the recipe called the singer’s formant. It sits in a band where an orchestra happens to be quiet.18 The singer is not louder than eighty instruments. The singer is loud in the one narrow place the eighty instruments left empty. That is a channel-allocation trick, solved by ear, centuries before anyone had the word for it.
Why two people singing the same note are never the same note
If timbre is the recipe, then a voice is a recipe that belongs to a body. Your tube has its own length, its own set of bends, its own soft-tissue lining. No other tube boosts the same rungs by the same amounts. Vocal tract length alone is one of the biggest sources of difference between speakers. The upper resonances are the ones above about 2.5 kHz. They are ruled by parts of your throat whose shape barely changes as you talk. That makes them close to a fixed signature.19
That is why two singers on the same pitch are two clearly different objects. They are not making different notes. They are making the same ladder through different tubes. You are hearing the tubes.
Here is the test that proves how much of the signal you can throw away and still have the person. A normal phone call carries about 300 Hz to 3,400 Hz and drops everything outside that band.20 Most adult speaking voices have a fundamental below 300 Hz. So when your mother calls, the bottom rung of her ladder never reaches you. That rung is the note she is on. It is not made quieter. It is gone.
She still sounds exactly like your mother, and you still hear her voice as low.
Two things do that. The first is that pitch does not live in the fundamental. Your hearing reads pitch from the spacing of the surviving rungs. So a ladder that goes 400, 500, 600, 700 is heard as a voice at 100. The 100 was never sent. That effect is called the missing fundamental. It is why a small phone speaker can give you a bass voice it cannot make.21 The second is that identity was never in the bottom rung anyway. It is in the pattern of relative loudness across the rungs that remain. The phone company kept exactly the band where that pattern is richest. The 300–3,400 Hz choice was not random. It was picked as the narrowest band that still lets a listener know who is talking.20
Which lands back on the idea this piece runs on. What crosses the wire is not her voice. What crossed the room was not air. Both times it is a pattern. A pattern survives losing most of itself, as long as the shape of what is left stays the same.
The instrument is made of arguing muscles
A valve that took a second job
Look at what the larynx is before you look at what it does for you.
Trace it back far enough and it is a sphincter. That is a ring of muscle around a hole, whose whole job is to close. The oldest version sits in air-breathing fish. There it is a simple valve of muscle. It guards the way into the swim bladder, so water cannot get in. Later animals add muscle that pulls the valve open on purpose. That way breathing can be timed. Making sound arrives last, and only in a serious way in mammals.22
So the order of business is this. Keep the lungs sealed against everything that is not air. Then let air in and out on schedule. Then, in the end, make noise. Speech is a late tenant in a building made for something else. The building was never rebuilt. Your larynx still slams shut the instant a crumb goes the wrong way. It does that rather than let you finish your sentence. The first tenant ranks first, and always will.
The parts follow from that first job. The cricoid (a full ring of cartilage at the base of the voice box) is the only complete cartilage ring anywhere in your airway. Every ring below it, all the way down the windpipe, is a C. It is open at the back, so your food pipe can bulge into the gap when you swallow.23 One place in the whole tube has to stay perfectly round at all times. That place is the mount that everything else pivots on.
On top of the ring sits the thyroid cartilage (the big shield in front, the bump some people call an Adam’s apple). At the back sit the arytenoids (two small pyramids that swivel to open and close the folds). Between the shield and the pyramids are the vocal folds themselves. Three pieces, one hinge, and a valve.
Two muscles pulling opposite ways
The folds are not a fixed reed. Their length, thickness and stiffness change all the time while you sing. Two muscles do most of that changing by pulling against each other.
The cricothyroid runs from the ring to the shield. When it tightens it tilts the shield forward and down against the ring. That opens up the distance between the front end of the folds and their back end. The folds get stretched: longer, thinner, tighter. Tighter and thinner means faster vibration, so pitch goes up. It is the main tensor of the folds and the main pitch-raiser. It is also the only intrinsic laryngeal muscle wired by a different nerve from all the others.24 Singers meet it as head voice: high, light, easy to hear and hard to keep full.
The thyroarytenoid runs from the shield back to the pyramids. It does the opposite. Tightening it pulls the pyramids forward toward the shield. That shortens and slackens the folds: shorter, thicker, heavier, lower.25 Singers meet it as chest voice: full, loud, and not willing to go high.
Now look at where that second muscle sits. Its deeper fibres run right alongside the vocal ligament. They are called the vocalis. They are part of the body of the vocal fold itself.25 This is not a muscle that pulls on the instrument from outside. It is a muscle that is a piece of the instrument. It changes its own mass and stiffness while it vibrates. Nothing in engineering works like that. A clarinet reed does not thicken itself halfway through a phrase.
Nothing about that setup is a bug. Two muscles pulling against each other is how you get fine control out of coarse parts. A hydraulic cylinder that only pushes gives you position control that is exactly as precise as your pump. Two cylinders pushing against each other give you position control that is as precise as the difference between them. They also give you a stiffness you can set on its own, by pressing them both harder. Your larynx runs that scheme. Pitch is the balance point between the two muscles. The fullness of the tone is how hard they are both working while they hold it.
Which means singing a scale is not a set of settings. It is a running argument between two muscles with opposite interests. It runs at a finer grain than either one has on its own.
A trick for reading muscle names
Those names look like noise until you know the rule. The rule is almost the whole of anatomy’s word list: muscles are named for the two things they connect.
Crico-thyroid runs from the cricoid to the thyroid cartilage. Thyro-arytenoid runs from the thyroid cartilage to the arytenoids. The name is a set of directions between two landmarks. Once you can name the landmarks you can read the muscle without being told what it does.
Try it on something that looks worse. Sternocleidomastoid is the thick rope you can see in the side of your neck when you turn your head. Sterno is breastbone. Cleido is collarbone. Mastoid is the lump of skull behind your ear. Three landmarks, three parts to the word. Now you know roughly what happens when it shortens. You just read a nine-syllable word by knowing where things are.
Why the crack happens in the same place
You have climbed a melody and had your voice flip. It goes thin, or it breaks outright. The maddening part is that it happens at almost the same pitch every time.
That pitch region is the passaggio (Italian for passage). It is fixed because it is where the two muscles trade the lead.26 Below it, thyroarytenoid leads and the folds are short and thick. Above it, cricothyroid has to lead and the folds must be long and thin. Somewhere in the middle the workload has to move from one to the other.
Done well it is a crossfade. One eases out at the rate the other eases in. Both stay partly active the whole way. That is all that trained “mixed voice” means. It is not a third register or a third muscle. It is the two you already have, refusing to fully hand over.26 Done badly, one of them quits all at once. The balance jumps instead of sliding, and you hear the jump.
There is a second change riding along with the first, and it is the more interesting one. Thick folds and thin folds do not just sound different. They vibrate in a different shape. A thick fold opens from its bottom edge first. The opening rolls upward through the tissue like a small wave, then the whole depth slams shut. A thin fold flutters shallowly, more like a ribbon, and may not fully close at all.
That slam is where your upper harmonics come from. Go back to the start of this half: the ladder exists because the puff has a sharp edge. A crisp, complete closure gives a sharp edge and so a tall ladder: brightness, ring, carrying power. A soft or partial closure rounds the edge off. The upper rungs collapse, and you get something flutey and pretty and small.
So the climb is two jobs at once. Hand the muscles over slowly, and thin the folds slowly without losing the closure. The short way to say it: you are shedding vibrating mass on a slope. The crack is what a step looks like when the slope was supposed to be smooth.
Getting better at it is not a strength problem. It is a coordination problem, in a place you cannot see.
You cannot feel any of it
This is the fact the rest of this piece leans on.
You cannot see your vocal folds. You cannot touch them either. Press anywhere on your throat and your fingers reach the outside of the shield and stop. The folds are sealed inside the cartilage box. And you cannot feel them the way you feel your hand. Everything you think you feel while singing is real, and is something else. There is vibration carried through cartilage and bone, air moving over your throat lining, the buzz in your face that voice teaching calls “the mask”. All of it is a side effect of the thing, not the thing.
The strange part is that this is not a shortage of sensors, or at least not on the face of it.
Small deep muscles across the body are fitted with muscle spindles (stretch-sensing receptors buried inside muscle) far more thickly than most. Measured per gram, the tiny muscles that aim your eye and hold your head carry densities an order of magnitude beyond big movers. One survey lists the inferior oblique of the eye at about 267 spindles per gram. It lists rectus capitis posterior at about 98. Large limb muscles in the same body are counted in single digits. The same survey is careful to say the field’s methods are inconsistent, and that the comparison should be handled gently.27
For the larynx itself, the evidence is genuinely contested. Studies from the 1950s through the 1980s reported spindles in the thyroarytenoid using traditional stains. Later immunohistochemical work argued that some of those were not spindles at all. And a recent animal study found canonical proprioceptors largely absent from intrinsic laryngeal muscle. A 2023 review of the whole question lands on unresolved. It suggests the laryngeal receptors may be odd enough in structure, with thinner capsules and fewer internal fibres, that they are simply hard to identify.28
What is not contested is the part that matters to you as a singer: whatever is down there does not report to you. Voice teachers work almost always in one way. They ask students to attend to sensations that are proxies. Direct position sense of the folds is not there to attend to.29
And notice that this is not a general failure of access. That is what makes it interesting. The command channel to your larynx stands out. Humans have a direct connection from motor cortex to the brainstem cell group that drives the laryngeal muscles. That link is absent in monkeys and weaker in apes. It is one of the leading candidates for what made voluntary speech possible at all.30 Your voluntary control of this equipment is finer than most. You just have no readout from it.
Think of a factory floor covered in sensors. All of them are wired into a control loop that works on its own and keeps the line running. None of them are wired to a screen in the control room. The instrumentation is excellent. The telemetry is fast and used all the time. It was just never routed to a display. For most of the machine’s life, no one stood in the control room who needed to look.
That is your larynx. Sensors, loops, reflexes, no dashboard.
Once you accept it, the whole odd dialect of voice teaching stops sounding like fortune telling. Spin it. Sing on the breath. Put it in the mask. Think the note before you sing it. None of those are orders to a muscle. An order to a muscle you cannot find is not a thing you can follow. They are descriptions of a result. They are handed to a nervous system. That system is very good at hunting for its own wiring, once it knows what it is hunting for. The teacher is not being vague. The teacher is aiming at the only input the system accepts.
Which raises the obvious question. It is the one the second half of this piece is about. If you cannot feel the instrument, how does anyone ever learn to play it?
Every language runs at the same speed
Play a Spanish news show, then play a Thai one. The Spanish comes at you like a machine gun. The Thai sounds like someone with all afternoon. We all notice this. And most of us draw the same lesson from it: some languages are faster, so their speakers must get more said per minute.
They do not. In 2019 a team of linguists measured it across seventeen languages. All of them send information at close to the same rate, near 39 bits per second.31 The fast languages are not saying more. They pay more syllables for the same content.
How it was measured
Christophe Coupé, Yoon Mi Oh, Dan Dediu and François Pellegrino taped speakers reading the same set of texts. The texts were put into each of the seventeen languages. Keeping the meaning the same is the whole trick. If everyone sends the same thing, then any gap in how long it takes is a gap in the wrapping, not in what is inside.
Then they measured two numbers for each language. The first is easy: syllable rate, syllables per second. That is just counting. It varies a lot. Japanese came in near 8.0 syllables per second and Spanish near 7.7. Thai came in near 4.7 and Vietnamese near 5.3. The fast ones are about 50 percent faster than the slow ones. That is just the gap your ear told you was there.
The second number is the good one: information density, in bits per syllable. Japanese sits near 5 bits per syllable. English is a bit over 7. Vietnamese tops the set at about 8.32
Now multiply. Eight syllables a second at five bits each is forty bits a second. Five and a bit syllables a second at eight bits each is about forty-two. Those are rough averages multiplied. So treat them as a neighbourhood, not a result. But the neighbourhood is the point. The study’s own figure across all seventeen languages was 39.15 bits per second. And information rate varied far less from language to language than either of the two numbers that make it.
What a bit actually is
The word “bit” is doing real work in that sentence. It does not mean what it means in day to day speech. It comes from Claude Shannon’s A Mathematical Theory of Communication, from 1948. That is the paper that made the field and coined the term.33 He is the same Shannon who runs the case in /on/entropy. This is the second deep dive in a row where he holds the whole thing up.
Shannon stopped asking what a message means. He asked how much it narrows down. Information, in his sense, is doubt cut away. A coin flip has two ways it can land. So learning one flip gives you one bit. Learning four flips gives you four bits. Four flips have sixteen ways to land, and sixteen is two times two times two times two. That is the whole formula. The number of bits is the number of times you halve the set of choices to get down to one.
So the worth of a symbol rests on how many other symbols could have shown up in its place. A letter drawn from a 26-letter alphabet tells you more than a digit drawn from ten. It rules out more of what the message might have been. This is why a password of random letters is harder to crack than a PIN of the same length. It is the same fact in a different mood.
Syllables work the same way. A language with a few hundred syllables gives each one a lot of work to do. Hearing it rules out a lot. A language with only a few dozen gives each one much less. In the study, the effective number of choices ran from about 32 syllables at the low end to about 256 at the high end. Effective is the right word there, and it matters. It is not the raw count in the dictionary. Some syllables are common and some are rare. And a syllable you were expecting anyway brings almost nothing when it lands.
Gross and net
This gives you a split worth keeping. It is your own hunch made exact.
Gross rate is syllables per second. It is what you hear. It is easy to count. And it differs wildly from language to language.
Net rate is bits per second. It is what gets across. It is hard to count. And it barely moves.
Your ear only ever reports the gross figure. That is why the trick of the ear holds so well. When Spanish sounds like it is outrunning you, you are hearing a language spend lots of cheap syllables. Each one has cut the field only a little, so it needs a lot of them. Vietnamese sounds slow because each syllable has already cut away more. There is less left to say.
The same trade, in a zip file
Take a text file and zip it. The zipped one is much smaller. And every byte in it is now much harder to guess. Before, the file was full of the letter e and the word the and long runs of the same sign. You could have guessed all of it. The whole job of a zip file is to delete what you could have guessed. What comes out the other side looks like noise. Noise is what a file looks like when nothing in it can be guessed. Which is the same as saying every byte is at full load.
The content did not change. The box got shorter and each unit of box got denser. The two changes cancel out. That is the trade. It is the same trade the seventeen languages are making. Vietnamese is the compressed file. Japanese is the uncompressed one. Neither is better. Neither says more. And if you measure the thing that counts — total content divided by total time — you cannot tell them apart.
Most people have handled the sound version of this. A song at 320 kilobits per second and the same song at 128 sound different. Their bitrates differ. And bitrate is bits per second: how much the encoder may spend on each moment of sound. Human speech runs at a fixed bitrate of about 39 bits per second. And every language on earth has landed on the same setting on its own.
Why 39 is not settled
The number looks like a limit set by the body. Something is capping it. Seventeen unrelated languages do not land in the same narrow band by chance. And the likely suspects all sit in the head, not in the mouth.
The best backed guess is on the listener’s side. Understanding speech seems to depend on brain rhythms locking onto the rhythm of the signal coming in. In particular, on activity in the theta band, which cycles at about 4 to 8 times per second. That is almost exactly the syllable rate of every language measured.34 If the decoder runs at a fixed rate, then the encoder gains nothing by going past it. Every language would be pushed toward the same ceiling by the plain fact that going faster gets you nothing. That is a good story, and the proof for the rhythm is real. But the last step is a guess, not a measure. Neural tracking of the speech envelope is needed for understanding, and it is not enough for it.
A rival account puts the limit on the speaker instead. Not how fast you can hear, but how fast you can put together what you are about to say. Listeners can follow taped speech played back well above normal speed without much trouble. That is awkward for a bottleneck made of hearing alone.
And the result itself carries a caveat worth saying plainly. What was measured is doubt over syllables, not doubt over meanings, and those are not the same thing.35 It also cannot be pulled fully apart from a simpler story. Languages with fewer syllables to draw on tend to be spoken faster. That alone might make the whole pattern, with nothing clever going on. The trade-off is solid. The claim that 39 bits per second is a best point, rather than a side effect, is not.
What else rides on the channel
The 39 bits is only the words. It is what a perfect transcript would catch. And a perfect transcript is famously not the talk itself.
Everything else your voice does rides a second channel that nobody plans for. Pitch shape, loudness, timbre, breathiness, the length of your pauses, where you speed up and where you stall. Linguists call this prosody, the tune and timing laid over the words rather than inside them. And it carries a live report on the state of the speaker.36
Some of that report is well proven. Listeners name emotions from voice alone at rates well above chance. And they do it across languages they do not speak. A large cross-cultural study found scores near 66 percent. They ran from about 74 percent in Germany down to about 52 percent in Indonesia. The pattern of mistakes was almost the same everywhere. The same emotions get mixed up with each other whatever the listener’s culture, and that is a stronger result than the score.37 A meta-analysis of 37 such studies backs the cross-cultural effect and adds an in-group edge. You read your own culture’s voices better, and the further apart two cultures sit, the wider the gap.38
The cleanest case is one you can test today, and it works by a route that leads straight back to the tube. Smiling is audible. Pull your lips back and you shorten the vocal tract. A shorter tube rings higher, so the formants rise. Vivien Tartter measured this in 1980. Smiling raised both the fundamental frequency and the formant frequencies for every speaker tested. And listeners picked the smiled tapes out of pairs at well above chance.39 Nobody is trying to send a smile. The change of shape does it for free. The filter cannot help but report its own shape.
Past emotion, we have to be less sure. Mental load slows speech rate and draws out pauses, and that is steady enough to measure in groups. The wider field of vocal biomarkers — using voice to spot disease, depression, fatigue, or being drunk — is real work with real signal. And right now it is oversold. The problems that keep coming up are the ones you would expect. Results that do not carry over between groups of people or recording setups. Small samples tied to one region. And models that read a voice with no knowledge of the person it belongs to.40 Voice clearly carries health information. Reading it well out of one tape, from a stranger, is not a solved problem. And products that claim it is are ahead of the proof.
This is also why a transcript of a good talk always reads thinner than the talk was. Nothing went missing from the words. The side channel just has no column in the file. The pause before someone answered. The drop in pitch that meant they had decided. The breath that meant they had not. All of it was information you used at the time, and none of it lives through the trip to text.
Which is the honest reason writing is hard. Text is voice with the side channel stripped out. Every trick writers reach for — italics, punctuation, a sentence cut short on purpose, a paragraph break set where a breath would go — is a fake limb for prosody, and a poor one. You are working with 39 bits per second and nothing else. And you are trying to rebuild a signal that first came in with a second track running under it.
Practice
A voice is a lossy squeeze of an inner state into a pattern in air. Everything you do with it is a fight over how much lives through the trip.
That sentence is the hinge, and the rest of this piece hangs on it. The first word you said to your mother was the crudest version there is. Barely any information at all, carried by the only noise your mouth could make. Singing is the same fight run at high fidelity and under load. Talking to a machine used to be the fight at its worst. The machine could not decode you, so you had to encode yourself into its language instead. That last one has just changed, and that is where this ends.
Our first sounds
Press your lips together and hum. That is /m/, a nasal consonant, a sound that leaves through the nose instead of the mouth. The soft palate, the fleshy back part of the roof of your mouth, drops away from the wall of your throat. That opens the door to your nasal cavity. Your lips never part. It is the only sound you can make with your mouth shut.
Now think what a baby’s mouth is doing for most of its first months. Lips sealed. Mouth full. The way out through the mouth is blocked. The way out through the nose is not.
Roman Jakobson saw this and built a case on it. “Often the sucking activities of a child are accompanied by a slight nasal murmur, the only phonation which can be produced when the lips are pressed to the mother’s breast or to the feeding bottle and the mouth is full.”41 That murmur is not speech. Nobody is being spoken to. It is the noise a working mouth makes when the mouth is busy. Jakobson’s next step is a proposal rather than a measured result, and worth flagging as one. He says the baby later makes the same sound when it wants the feed rather than has it, so the murmur moves from during to before. The body claim under it is not in doubt. If your lips are busy, /m/ is what is left.
The vowel that needs no shape
The other half of mama costs even less.
Go back to the source and the filter. The folds make one buzz. The tube above them shapes that buzz into a vowel. The vowel lives in the shape, not in the buzz. So ask the next question. What comes out when the tube is not shaped at all?
If the tube is a plain even pipe, about the same width from the voice box to the lips, you get a schwa, the flat, colourless uh in the second syllable of sofa. For an adult tract of about 17 cm, that shape puts the resonant peaks near 500, 1500 and 2500 Hz. That is the simplest number in all of speech acoustics.42 Now drop your jaw and leave the tongue lying where it lies. You have not aimed the tongue at all. You have not rounded, spread, bunched or curled a thing. What comes out is close to /a/, the open vowel in father.
So mama is not a word a baby has picked out of the sounds on offer. It is the two cheapest moves the vocal tract owns, taking turns. Shut the lips and let the voice out through the nose. Then open the jaw and let the voice out through the mouth. Repeat. Babies start making these same repeated consonant-and-vowel strings at about six months. The consonants that rule that stage are made at the front of the mouth with big, forgiving moves — b, d, m, n, p, t. Babies do this everywhere, before anyone has taught them a thing.43
The first word is not chosen. It is the sound left over.
Why the adults hear a name
George Murdock went through parent words in a large sample of the world’s languages. He found that four sound shapes — ma, na, pa, ta — turn up as words for mother and father far more often than chance allows. They turn up in languages with no shared parent and no contact with each other.44 That pattern is not handed down. You cannot inherit a word from a language you never touched. It is made fresh each time, over and over. Every new set of babies is issued the same mouth and finds the same easy noises in it.
What happens next is the good part, and it is done by the adults. A baby makes a nasal murmur that means nothing. The nearest adult hears it, decides it is a name, decides it is her name, and answers to it. Answering is what makes it a word. The baby seems to have said “mother”. What the baby did was breathe with its lips together.45
You have seen this happen with concrete. Walk across a college campus and you will find worn dirt lines cutting the corners off the lawn, right where the paved path was a nuisance. Nobody designed those lines. They are what thousands of people do when they take the shortest route. Which is to say they are the leftover of effort, not a choice. And then the grounds crew paves one, gives it a name, and puts up a sign. The path came from physics. The name came from a school looking at the path later and deciding what it meant.
Georgian pays it out the other way
The claim you often hear is that ma means mother in every language on earth. It does not. And the exception is the best part of the story.
Georgian is spoken in the Caucasus, and its language family has no proven relatives anywhere. It uses mama for father.46 Mother is deda.47 Same easy syllables, same babbling mouth, the jobs swapped. Papa in Georgian is grandfather.
That does not weaken the case. It is the case. The sounds are the same everywhere because mouths are the same everywhere. A Georgian baby’s lips and jaw do just what yours did. The meanings are not the same everywhere. Meanings are handed out by whoever is listening, and listeners belong to cultures rather than to anatomy.
Which is the source and the filter again, one level up. Down in your throat, the body gives a buzz and the shape gives the meaning. Out here, the body gives a small set of noises it is cheap to make. And the culture standing over the crib decides which noise is a mother and which is a father. In both cases the raw stuff carries no information at all. The shaping is where all the information lives.
It gets harder from here
Enjoy that first sound. It is the only free one you will ever get.
Mama is the one moment in your life when your voice does just what your body makes easiest, and the world answers as though you meant it. The rest of this half is the opposite of that. Speaking clearly means holding shapes your tongue does not fall into on its own. Singing means asking two muscles that disagree to hand a note back and forth without dropping it. Dictating to a machine means making a signal clean enough for a decoder that has no idea what you meant and no face to read.
None of that is left over. All of it is work.
Why a good session leaves you tired somewhere you cannot point to
You can walk for hours. You take the stairs without thinking about it. Then you sing for twenty minutes. Or you talk with real support for half an hour. And you are winded. Not sore. Not out of breath the way a hill leaves you out of breath. Heavy somewhere in the middle. Wrung out. And you cannot say where. You want to sit down. You cannot name the muscle that is asking.
Four things are going on at once, and they stack. That word matters. They are not four guesses fighting to be the right one. They are four costs, arriving at the same time, from four places. That is why the whole thing feels so odd and so hard to place.
A movement you have not learned costs more than the same movement once you have
Take a person. Sit them at a robot arm. Have them reach for a target. Now switch on a force field. It pushes the arm sideways as it moves. The old reaching program no longer works. A new one has to be built. Then measure the energy they burn while they do it. You do that by looking at the air they breathe out.
The energy cost of reaching went up by about 42%. That was the moment the movement became new. Then, as they learned it, the cost of that same movement fell by about 20%.48 Same arm. Same target. Same distance. The one thing that changed was how well the brain knew what it was doing.
That is what “muscle memory” really means. The name is a poor one. Nothing about the muscle changed in the half hour it took. What changed was the command sent to it.
Part of the reason is co-contraction (opposing muscles firing against each other at the same time). When you do not yet know what a movement needs, you brace. You switch on the muscle that makes the movement and the muscle that fights it. Bracing is what you do when you cannot guess what is coming. Both muscles burn fuel. Neither one gets anything done that the other is not cancelling. As the movement becomes known, that bracing drops away.
The same study has a wrinkle in it. It points somewhere odd. The bracing dropped off early. And the energy cost kept falling after it had already flattened out.48 So less bracing explains some of the saving and not all of it. Something else about how the nervous system sends the command keeps getting cheaper. That goes on for a good while after the part you can see has settled.
Move this out of the body for a second. The shape is easier to see somewhere else. Think of a person learning to drive a car with a manual gearbox. Twenty minutes of city traffic and they get out of the car truly tired. Jaw tight, shoulders up, worn out by a job that means pressing two pedals and moving a lever. A year later the same person drives the same route while holding a chat and eating a sandwich. The clutch did not get lighter. The hill did not get less steep. The cost fell because the program got written.
Your first year of singing well is that drive.
The breathing muscles are ordinary skeletal muscles, and they genuinely get tired
The diaphragm (the domed sheet of muscle under your lungs) is not a special organ. It is a skeletal muscle, like a bicep. It is made of the same stuff and has the same limits. It does about 70% of the work of a normal breath in.49 Around it and under it sit the deep intercostals (the muscles between the ribs) and the transversus abdominis (the deepest abdominal layer, wrapping the waist sideways like a belt). They manage the pressure in the trunk.
These muscles fatigue in the strict technical sense. Doctors who work on breathing use an exact meaning for the word. It is worth having. Fatigue is a reversible loss of the ability to produce force, caused by working under load, and recovered by rest.50 Reversible is the key word. It splits fatigue from weakness. It is why the answer to a session that wrecked you is a day off, not a diagnosis.
There is even a rough line for when a breathing pattern turns into a tiring one. It is called the tension–time index. It multiplies how hard the diaphragm is pulling by what share of each cycle it spends pulling. Above about 0.15, you cannot keep that pattern up for long.49 Read the second half of that. The fraction of time under tension counts as much as the force does.
Which is just what singing does to you on purpose. Breathing out is free, as a rule. The ribs drop, the belly comes in, air leaves. If you let that happen while singing, all your air escapes in a whoosh. The note falls apart. So a trained singer keeps the muscles that pull air in switched on. They stay on for the whole of the breath out. That holds the ribcage open and lets air leave in a trickle. Italians named it appoggio, from appoggiare, to lean. You lean on the breath.
That is co-contraction again. This time you are doing it on purpose, for the whole length of every phrase.
Here is the cheapest way to feel the cost. Press your palms together in front of your chest as hard as you can. Nothing moves. Nothing is lifted. Count to twenty and your arms are burning. Or throw a punch at full speed. One short burst of force. The arm flies on its own momentum. The opposing muscles catch it at the end. A few milliseconds of real work. Now throw the same punch in slow motion. There is no momentum to coast on. So every millimetre has to be pushed and held back at once. Ten of those and you need to sit down.
Singing well is the slow-motion punch. It runs without a break, inside your torso, for as long as the phrase lasts.
Two more details make it worse. The transversus abdominis is a feedforward muscle. It fires before the movement it is stabilising for, on prediction rather than on feedback. And it does so whichever way that movement goes.51 It is not waiting for your orders. And the diaphragm never gets a day off in your whole life. So its baseline is set by quiet breathing at rest. Ask it for something it does not normally do and, in that one pattern, it is untrained.
“Out of breath” is a sensation your brain builds, not a reading of your oxygen
This is the one that catches people out, and it is the most useful of the four.
The feeling of air hunger does not come from a sensor reading low oxygen. It comes from a comparison. Your brainstem sends out the command to breathe. It also sends a copy of that command upward: a corollary discharge (an internal copy of a motor command, forwarded so the brain knows what it just ordered). Meanwhile, stretch receptors in the lungs report what really happened. The command says more. The feedback says that was not enough. You feel that mismatch as air hunger.52
Corollary discharge is not a rare thing. The clearest case of it has nothing to do with breathing. It is why you cannot tickle yourself. Your brain forwards a copy of the command to your own hand. It guesses what that hand will feel like, and cancels it out. Someone else’s hand sends no copy. So nothing gets cancelled, and the same touch is too much to bear. Air hunger is that same wiring, with the cancelling gone wrong. The guess and the signal that turns up do not match. That mismatch is what you feel.
Two things drive the command: carbon dioxide, and the effort of breathing. In a lab you can pull them apart. Hold ventilation steady and change CO2, and air hunger moves sharply while the sense of effort barely does. Hold CO2 steady and change ventilation, and the sense of effort moves while air hunger stays put.53 They are two feelings with two causes. The one people describe as “I cannot get a satisfying breath” is the CO2 one.
Air hunger also lights up the insula (the fold of cortex that integrates the body’s internal states — hunger, pain, temperature, nausea) along with the parts that make anxiety.52 That is not a footnote. It is why breathlessness is scary in a way that a tired leg is not. It is the same routing described in /on/fascia for fascial sensation. That is an inner signal that turns up as a body feeling with an emotional colour, not as a reading at a point.
Here is what that means in practice, and it runs against the advice almost everyone gives. Taking huge, fast breaths on purpose washes carbon dioxide out of your blood. Your oxygen is fine. It was fine the whole time. But the CO2 that normally sets your breathing drive is now too low. Blood vessels in the brain narrow. You get light-headed, tingly, and gripped by the clear feeling that you cannot get a full breath.54 You have made the symptom by over-treating it. If a warm-up leaves you dizzy and gasping, the first thing to suspect is not weak lungs. It is too much air, too fast.
You have no internal map of any of it
The first half of this piece showed that the laryngeal muscles are packed with sensors. They report to almost nothing you can consciously read. The same problem runs down through the whole support system. And it makes the other three causes worse.
Interoception is the sense of your own inner state. Where it is good, control is cheap. You can find the position, hold it, and stop doing everything that is not the position. Where it is poor, the nervous system does the one safe thing it can. It switches on more than it needs. That is more co-contraction and more fuel burned. And the sense of respiratory effort is itself one of the inputs to how breathless you feel. So it is also more of the feeling of struggling. The poor map does not just make the work less efficient. It makes the same work feel harder, through a channel that is on the record.
So the tiredness is not proof that you are unfit. It is not proof that your breathing muscles are weak, though they may be. It is mostly the price of running an unfinished program.
The thing doing the learning, and the thing doing the tiring, is largely the brain. That is also why it gets better far faster than you expect.
Muscle takes months. A motor program takes weeks. And the measured energy cost of a movement starts falling inside one session.48 Say you have been treating this as a fitness problem and planning around that. Then you have the timescale wrong, and in the direction of good news.
Good tired and bad tired
One test makes all of the above easy to use. It is the thing to remember if you remember nothing else.
Good tired is deep, central, and late. It sits in the ribs, the back, and the belly. The trunk, not the neck. It tends to turn up after you finish rather than during. That is the way a hard set of anything shows up an hour later. That is the support system doing the job. It is the mark of a session that went right.
Bad tired is high, narrow, and right now. It is in the throat, the jaw, and the strap muscles down the front of the neck. And it shows up while you are working, not after. Scratchy, gripping, hot. That is substitution. The deep muscles you cannot feel were not carrying the load. So the outer muscles you can consciously aim at jumped in to cover. And they are truly bad at it. They are neck muscles. They have no business shaping a phrase.
Good tired is a training signal. You should go and get more of it. Bad tired is a technique signal and it means stop, not push. Telling the two apart in the moment, every time, is most of the skill.
Training a thing you cannot feel
Everything below shares one design idea. It makes feedback where the body gives none. Or it puts a load on a muscle that attention alone will not train. None of it is about trying harder. Trying harder is what causes substitution.
Building the map before building the strength
Sit or stand. Put one hand flat on your belly. Put the other hand on your lower ribs at the side, fingers spread. Breathe in slowly and quietly. The goal is that the belly moves and the lower ribs widen sideways under your hand. That is 360-degree breathing. It is so called because it opens out all the way around rather than only at the front. What should not happen is the chest rising and the shoulders lifting.
Then breathe out slowly, and make the exhale longer than the inhale.
This is not a strength exercise. Be clear about that or it will let you down. It is a mapping exercise. Your hands are doing the job your interoception cannot. They give you a clear report from outside on whether the pattern happened. It is the same trick as a mirror in a weights room. You cannot see your own back, so you borrow an outside channel.
The long exhale earns its place on its own. Exhale-weighted breathing lowers physiological arousal, and you can measure it. Take a controlled trial of five minutes a day for a month. In it, exhale-emphasised breathing beat both equal-ratio breathing and mindfulness meditation. It won on mood and on cutting respiratory rate.55 A calmer nervous system braces less. And bracing is what made the movement costly in the first place.
Three to five minutes. Not more. The moment it turns into a big effortful breath, it starts washing out CO2. And it makes the exact feeling you are trying to train away.
Semi-occluded vocal tract work, which is the best-evidenced exercise in the field
Hum. Do lip trills, the motorbike noise. Best of all, take a drinking straw. Put it in a glass with a few inches of water in it. Then phonate through it so the water bubbles.
SOVT stands for semi-occluded vocal tract. You narrow the exit of the tube while making sound. Do it at the lips, or by putting a straw in the way. That one change does something worth spelling out in full. It goes against what you would guess. And it is the mechanism that makes this the most-recommended exercise in voice science.
The vocal folds vibrate because air pressure from below blows them open. Then they snap shut again. What tires them and harms them is not vibrating. It is colliding: the force with which the two folds slam together, hundreds of times a second.
Narrow the exit and pressure builds up in the tube above the folds. The pressure below has not changed. So the difference across the folds, which is what drives them apart, gets smaller. Smaller driving pressure means smaller swings. Smaller swings mean the folds meet more gently. At the same time, the air column above them becomes a better acoustic match for the source. So more of the energy you put in comes out as sound instead of being wasted. Titze’s simulations, at the National Center for Voice and Speech, put it as impedance matching between the glottis and the vocal tract. The payoff is a voice that is more economic and makes lower collision forces in the tissue.56
Read that again in plain terms. You get more sound out for less impact on the folds. That is the whole reason this exercise wins. Almost every other way of making your voice easier makes it quieter or smaller. This one does not.
The straw in water adds a second thing on top. That is why it is the version to use. The bubbles are a picture of your airflow. Choppy bubbles mean choppy air. Steady bubbles mean steady air. You have just given yourself something to watch for a process you cannot see at all. That is the made-feedback idea in its purest form.
Here is the honest state of the evidence. The mechanism is well modelled. And the effect people report right away holds up. But short-term objective acoustic measures after one session are often unchanged.57 Three to five minutes before you sing, or before a day of heavy talking.
Loaded inspiratory work, because gentle breathing will not build strength
Say you want the breathing muscles stronger rather than better coordinated. Then you have to load them, the same way you would load any other muscle.
Inspiratory muscle training uses a handheld device. It has a spring-loaded valve you have to pull air through. A meta-analysis of pressure-threshold training found significant gains in maximal inspiratory pressure (MIP — the hardest suck you can generate, the standard measure of inspiratory strength). That held when the load was at least 15% of your current MIP. Real gains showed up inside four weeks. The largest gains, around 54%, came at twelve weeks.58 Another review found the feeling of breathlessness dropped by about 56–62%. That works through the respiratory metaboreflex: tiring breathing muscles set off a reflex that steals blood flow from your limbs.59
The caveat is the point of the paragraph. Gentle diaphragmatic breathing on its own is almost certainly not hard enough to build strength you can measure. It builds the pattern, which is the other job and a needed one. But the load is what builds the muscle. If you want both, you have to do both. The two do not stand in for each other.
About thirty breaths, twice a day, at a resistance you can only just finish, three or four days a week.
The version that runs all day
If you already walk and stretch through the day, this is the best change you can make. And it costs no extra time at all.
Hum for thirty seconds on a walk. Lip trill for twenty. Breathe out through a straw a few times while you are stretching. Do a physiological sigh whenever you notice your breathing has gone high and shallow. Take two inhales through the nose, the second one short and stacked on top of the first. Then a long slow exhale through the mouth.55 And speak on the breath. Start the sentence from an exhale that is already moving. Do not squeeze the first word out of a still chest.
The reason to break it up is not ease. Distributed practice beats massed practice for motor learning. A meta-analysis of 63 studies found the edge was large for simple motor tasks. And the effect holds most strongly for retention, not for how fast you improve inside one session.60 Massed practice also shows the mark of higher cognitive effort and attention demand. That is the fatigue described in the first half of this piece.
Six ninety-second doses across a day beat one nine-minute block. And they beat it on just the axis you care about: what is still there next week. The nervous system locks it in during the gaps. Give it more gaps.
Sirens, for the handover between registers
The register break was described earlier. It is the point where the two arguing muscles have to trade control. It is where the voice cracks if the handover is sudden. The exercise for it is a siren. Slide slowly from your lowest easy note up to your highest and back down, like a fire engine. Use an oo or a gentle ng.
Three rules make it work, and each one has a reason.
Go slowly, because the handover has to be gradual. Speed lets you skip the region rather than learn it. Go quietly, because volume is the single biggest driver of collision force. This exercise exists to train coordination, not to test strength. And if it cracks, do not push through it. Go slower and quieter and cross the same place again. A crack is the handover failing. And doing a failed handover louder just teaches the failure.
Do sirens on a lip trill or through the straw and you get both effects at once. The pressure above the folds is raised, so the crossing happens at lower collision force. The exercise most likely to cause harm becomes one of the safest in the set.56
Biofeedback, or how to measure a thing you cannot feel
A training plan is a loop. A loop needs a sensor. When a squat is too heavy you know it. The muscle files a report. The report comes to you as a feeling. The gear that makes your voice does not file reports. The cricothyroid and the thyroarytenoid are as thick with nerves as any part you own. None of that wiring reaches your mind. You can feel the work of singing. You cannot feel which muscle is doing it. You cannot feel if it is doing it well.
So you borrow a sensor from outside. That move is not special to voice. You have no detector at all for carbon monoxide. That is just why it kills people in their sleep. So we build a small box that beeps. You cannot feel your heart rate to within ten beats. So you strap a watch to your wrist. You cannot feel four pounds. So you stand on a scale. Each time, a number from outside stands in for a sense you were never issued. And the number is not a stand-in for the feeling. It is the feeling, coming in through a different door.
Biofeedback is not a technique. It is a prosthetic sense.
Voice is a very good fit for the prosthesis. A microphone is cheap. It is very exact. And the thing you train already leaves the body as a signal in air. A knee has to be guessed from video. A heart has to be read from electrodes on skin. But a voice comes to you pre-measured. Four numbers do most of the work. You can start taking all four today with a phone and a stopwatch.
Maximum phonation time
Take a full breath. Sing a steady /a/ at an easy pitch and hold it, evenly, until the sound stops. The seconds on the clock are your maximum phonation time, or MPT.
It is a mixed measure. That is both its weakness and its point. The number folds three things together. How much air you took in. How well you dole out the breath on the way out. And how well the folds turn that air into sound. Get better at any of the three and it climbs.
The target band is softer than it looks in print. Clinics often quote about 25 to 35 seconds for adult men, and 15 to 25 for women. But measured groups land lower. Maslan and his team timed 69 healthy older adults. They found means of 23.2 seconds for men and 21.0 for women. Age and sex made no real difference.61 Read the printed bands as a rough guide, not as a grade.
What is not soft is how repeatable it is. Speyer’s group found MPT reliable to an interclass correlation of 0.998 across raters. And it split dysphonic patients from healthy ones by about 6.6 seconds.62 That mix makes MPT a poor exam and a fine gauge. Your MPT against everyone else tells you very little. Your MPT against your own MPT from six weeks ago tells you almost all you wanted to know.
The s/z ratio
Hold /s/, the hiss, with no voice behind it, for as long as you can. Write down the seconds. Take a fresh full breath and hold /z/, the buzz, for as long as you can. Same tongue, same teeth, same lips. Write that down too. Then divide the first by the second.
The design here is lovely. It pays to see why before you see the numbers. The two sounds differ in just one thing. Everything above the larynx is the same. So the air is pushed through the same narrow gap both times. The one thing that differs is whether the folds are switched on. So /s/ measures your air supply and your breath control. And it does so with the larynx taken out of the picture. Then /z/ measures the same thing with the larynx put back in. Dividing one by the other cancels your lungs and leaves your larynx. It is a controlled experiment you can run standing in a hallway, with no gear at all.
Eckel and Boone tested it in 1981 across 150 people. Normal speakers sat near 1.0. So did dysphonic speakers whose vocal folds turned out to be sound. Speakers with nodules or polyps went past 1.4 ninety-five percent of the time.63 The telling part of that paper is the negative result. The /s/ hold did not differ between the groups at all. Only /z/ shortened. That is the control doing its job in public.
The reason is plain once you have the ratio. A lump on the edge of a fold stops it sealing cleanly. Glottal resistance falls. Air rushes past faster than it can be turned into sound. So the buzz dies while the hiss keeps going. A ratio well above 1.0 is literally the sound of leaking.
One caution. This is a screen, not a diagnosis. Say you sit well above 1.4 and you have been hoarse for weeks. That is a question for a laryngologist, not for an app.
Phrases per breath
Pick one set passage. Count how far you get before you have to take air. Same passage every time, same tempo.
This is the mixed, real-world measure. It is the one that changes how a Tuesday feels. It is also the noisiest of the four. It moves for reasons that have nothing to do with your breath. How you chose to phrase a line. How fast you took it. If you were on edge. Track it. But treat one reading as gossip, not as data.
Pitch accuracy, in cents
A cent is one twelve-hundredth of an octave, which is one hundredth of a semitone on a piano. Alexander Ellis brought in the unit in 1885. He split the octave into “1200 equal hundredths of an equal semitone, or cents as they may be briefly called”.64 It is logarithmic. So the same number of cents means the same felt distance. That holds whether you are up high or down low. It is just what a raw reading in hertz cannot give you.
The scale is easy to hold. People can hear a gap of about 5 to 6 cents in good conditions. And they spot 25 cents very reliably.64 So a tuner reading 18 cents flat is a real miss you can hear. And it is still a small one. The unit earns its keep. It turns “you were a bit under” into a number. You can average that across a take and plot it across a year.
The reference verse
One verse. Same microphone, same distance, same gain setting, recorded once a month into a folder named for the date.
Holding the setup fixed is not fussiness. It is the whole design. Move the mic six inches closer. Your tone gets warmer. Your levels rise. Those changes land in the recording as though your voice had done them. Freeze all you can control. Then your voice is the only thing left free to move. Then the takes match up. And takes that match up are a time series, not a set of impressions.
From each take, pull the four things. Pitch accuracy in cents. How steady the sustained notes are. Phrases per breath. And a plain listen-back. The listen-back matters as much as the numbers. Twelve months apart, the change is plain to anyone. One month apart, you cannot see it at all. That is just why the folder exists.
The room has already decided
There is one more reason to keep the folder. It is the honest one.
The people who have heard you sing for years made up their minds about your singing years ago. Views about someone you know well change very slowly. This is not them being unkind. It is how attention works. Once a thing is filed, it stops being looked at again. Your family stopped looking at your voice again around the time they stopped noticing your face. You can get much better. And you will still be met, warmly, with the same verdict they reached ten years ago.
You cannot argue with that. You should not try. What you can do is keep a dated folder of takes, recorded the same way every month. Against a room full of people who already decided, it is the only ground truth on hand. That is not a point about your family. It is a point about measuring. When the people watching do not update, the tool is the only thing left that will.
Voice is vernacular
In July I set a camera up to test a dictation tool and needed something to read into it. What I wrote turned out to be the argument itself, so here it is in the words I said:
For 2000 years, if you wanted to talk to power, you learned Latin. The scribes had the syntax. Everyone else just had a voice. Code was the same deal — a priestly language, gatekept by semicolons. But something shifted: the machine learned our tongue before most of us learned its. Voice is vernacular now. I don’t translate myself into the computer’s language anymore. It translates itself into mine.
The gate was always the syntax
Start with the Latin. The claim is not a figure of speech. In most of European history, three things could change your life. All three ran in a tongue almost no one spoke at home. Worship ran in Latin. Learning ran in Latin. Law ran in Latin.
The law case is the sharpest. You can date it. In England, court cases could be argued aloud in English from 1362. But the written record stayed in Latin and Law French for another three and a half centuries. That is the writs, the pleadings, the judgments, the part that really binds you. English became the required language of the courts only when the Proceedings in Courts of Justice Act took effect in 1733.65 For four hundred years, an English farmer could stand in an English court, speak English, and lose. He would never be able to read the sentence written against him.
The same law gives you the cleanest picture of what a syntax gate is worth. Under benefit of clergy, a man accused of a felony could escape the gallows. He had to prove he was a cleric. And over time the proof shrank into a reading test. Recite a passage of scripture, usually the opening of Psalm 51, in Latin.66 The Latin was not beside the point. It was the whole exam. Knowing one line of a dead language was, quite literally, the difference between hanging and not hanging. That is why it got the nickname the neck verse. People memorised it by sound without grasping a word. That tells you they understood the system perfectly.
Learning worked the same way. It held out longest. Newton published the Principia in 1687 in Latin. The first English translation did not appear until 1729, forty-two years later, two years after he was dead.67 The laws of motion existed in England for two generations before they existed in English.
Now put a number under it. In the diocese of Norwich in the late 1500s, about 61% of men could not write their own names. Among labourers, not being able to read stayed above 90% into the 1600s.68 That is plain English reading and writing, the low bar. Latin sat somewhere far below that. It lived inside a class of clerks, lawyers and priests. They read the official language on everyone else’s behalf.
That is the scribal class. And the key thing about it is that it was not a plot. Someone had to hold the syntax. The documents were real. The language was hard. Holding it was a full-time job. A scribal class forms wherever a needed language is costly to learn. It does not need bad intent to end up standing between you and the thing you need.
What “vernacular” actually meant
Today “vernacular” tends to mean casual, or slangy, or plain. In history it meant something much more exact. It meant the tongue you already had, when the official language shut you out. Not a cut-down Latin. Not a beginner’s version. The tongue you were already fluent in before anyone offered you a choice.
Dante is the founding case. And he is funny about it. Around 1304 he wrote De vulgari eloquentia, an argument that the Italian vernacular deserved the same standing as Latin. And he wrote that argument in Latin, because that was who he had to win over.69 Then he did the real thing. He wrote the Divine Comedy in Tuscan, the tongue spoken in the street. It was about God and hell and the shape of the universe. Until then those subjects belonged to the clerical language by default.
Two centuries later the same move came to scripture. This time the press was there. Luther’s German New Testament came out in September 1522, in a run of three to five thousand copies. It sold out fast enough that a second edition ran in December.70 His full German Bible came in 1534. By then more than two hundred thousand copies of the New Testament had gone out. In England the cost was higher. William Tyndale’s English New Testament was printed on the continent in 1526. It was smuggled in, hidden in bales of cloth. He was strangled and burned in 1536.71 Seventy-five years later, fifty-four scholars put together the King James Bible of 1611, with royal approval. And about 83% of its New Testament wording is Tyndale’s. The state killed the man and then took up his sentences.
The press is what made all of it cheap, not just brave. Book output in Europe had grown at about 1% a year for centuries as a hand-copying trade. After the middle of the 1400s it climbs steeply. And the cause historians point to is falling prices and rising literacy feeding each other.72 A vernacular Bible is worthless if a copy costs a year’s wages. Luther’s 1522 Testament cost one guilder. That is about two months’ pay for a schoolmaster. Brutal, and still perhaps a hundred times cheaper than what a scriptorium would have charged.
None of this killed Latin quickly. The Catholic mass was still said in Latin worldwide until 1963. That is when the Second Vatican Council’s constitution on the liturgy opened the door to vernacular languages.73 That is four hundred and forty years after Luther. Gates close slowly.
The translation now runs the other way
Code is the same shape, squeezed into one lifetime. The ideas in most programs are not hard. Add these up. If this, then that. Send it there. What is hard is the syntax. The exact spelling, the brackets, the semicolon that must be there and the one that must not. Miss it and the machine does not do a worse job. It refuses. That is a gate made only of form. And it made a scribal class right on cue. There are about 47 million developers in the world,74 against roughly eight billion people. Under one percent of us hold the syntax.
People have been trying to widen that gate for seventy years. And they all pushed the same way. Grace Hopper’s FLOW-MATIC in the late 1950s, and COBOL after it, swapped maths symbols for English words on purpose. Business customers were not at ease with notation.75 It made code look like English. It did not stop being a syntax you had to learn exactly. Every step since has moved the gate a bit closer to you. And each one left it a gate. Friendlier languages. Better error messages. Visual editors.
What changed is direction, not distance. Natural-language models did not make code easier to learn. They made learning it a choice. The translation runs the other way now. The machine learned our tongue before most of us learned its.
What happens to the scribes
The prediction follows. It is worth saying flat out, not hinting at it. The scribal class loses its monopoly. Not its skill. Not its use. Its monopoly. That is the spot of being the only route between an ordinary person and a machine that does what they meant.
But be honest about the analogy. It cuts both ways. The printing press did not get rid of scribes. It moved them. In 1492 Johannes Trithemius, abbot of Sponheim, wrote De laude scriptorum, “In praise of scribes”. He defended hand-copying as holy work that printing could not replace. He had it printed, in 1494, because that was how you got read now.76 He also grew his monastery’s library from about 40 volumes to 2,000, most of them printed. He was not a fool fighting the future. He could see the craft being redrawn around him. He was arguing about which part of it was the point. The copying went away. But the editing stayed. So did the correcting, the choosing of what was worth copying at all, and the huge new trade of running presses. Those took in the people. What happened to any single scribe is a truly open question, and always was. Nobody gets to promise you that the new jobs land on the same people as the old ones.
The receiving end does the decoding
Here is why this belongs beside the anatomy and not in an essay of its own about computers.
Everything above in this piece has been about a channel. A voice is a pattern pressed into air by a shape you control. Every language, fast-sounding or slow, pushes about 39 bits per second through that channel.31 And prosody — pitch, timing, breath, hesitation — carries a second stream of information that text just drops on the floor.
For the whole history of computing, you did the encoding work. You took an inner state. You turned it into the machine’s syntax by hand. Then you typed it in a form the machine could parse with no effort at all. All the loss happened on your side of the wire.
Think of what became of travel adapters. Mains electricity comes in two rough families. One is around 100–127 volts. The other is around 220–240 volts. For decades, crossing between them meant carrying a converter. It was a heavy brick whose only job was to make you fit the wall.77 Then power supplies became universal input. The label on your charger says 100–240 V. The conversion happens inside the device. The mismatch did not go away. The work of settling it moved from the traveller to the receiver.
That is just what a voice interface is. The first time in this whole story that the receiving end does the decoding. That is what makes it vernacular in the precise historical sense. Not a simpler language offered to you. It is the one you already had, finally accepted.
The first sound you ever made was not chosen. It was the one your anatomy made easiest, a closed mouth and an open jaw, and the adults around you decided it meant them. Everything after that has been learning to shape air on purpose — vowels, then words, then a sung phrase held longer than felt possible, then, for a while, a language of brackets and semicolons that your body had no natural way to produce at all. The machine meeting you where your voice already was is not a new capability so much as the end of a long detour.
References
-
Longitudinal wave, Wikipedia.
“Mechanical longitudinal waves are also called compressional or compression waves, because they produce compression and rarefaction when travelling through a medium, and pressure waves, because they produce increases and decreases in pressure.” ↩ ↩2
-
Wind wave, Wikipedia.
“Parcels near the surface move not plainly up and down but in circular orbits: forward above and backward below (compared to the wave propagation direction).” ↩
-
The Physics Classroom, The Speed of Sound; see also Speed of Sound in Air, The Physics Factbook. 343 m/s is the standard dry-air figure at 20 °C. ↩ ↩2
-
Gunnar Fant, Acoustic Theory of Speech Production (The Hague: Mouton, 1960). Publisher preview · Semantic Scholar record ↩
-
Johan Sundberg and colleagues, The Gunnar Fant Legacy in the Study of Vocal Acoustics.
“It gives a highly accurate description of the speech signal and … explains how vowels and consonants get their acoustic properties. It is general, language-independent and valid for both normal and disordered speech.” ↩
-
Janwillem van den Berg, Myoelastic-Aerodynamic Theory of Voice Production, Journal of Speech and Hearing Research 1(3), 227–244 (1958). DOI 10.1044/jshr.0103.227. Based on a paper given at the Chicago International Voice Conference, May 1957. ↩
-
Ingo R. Titze, The physics of small-amplitude oscillation of the vocal folds, Journal of the Acoustical Society of America 83(4), 1536–1552 (1988). ↩
-
Holmberg, Hillman & Perkell (1988), as summarised in Average Speaking Frequencies: F0 Norms by Age, Sex, and Hormonal Status, Voice Science. Adult male mean about 116 Hz (range about 93–135 Hz); adult female mean about 205 Hz (range about 162–238 Hz). ↩
-
Formant, Wikipedia.
“Most often the two first formants, F1 and F2, are sufficient to identify the vowel.” ↩ ↩2
-
Gordon E. Peterson & Harold L. Barney, Control Methods Used in a Study of the Vowels, Journal of the Acoustical Society of America 24(2), 175–184 (1952). DOI 10.1121/1.1906875. Ten vowels, 76 speakers, at Bell Telephone Laboratories. ↩
-
Hamid Reza Sharifzadeh, Ian Vince McLoughlin & Martin J. Russell, A Comprehensive Vowel Space for Whispered Speech, Journal of Voice 26(2), e49–e56 (2012). ↩ ↩2
-
E. O. Belcher & S. Hatlestad, Formant frequencies, bandwidths, and Qs in helium speech, Journal of the Acoustical Society of America 74(2), 428–432 (1983). DOI 10.1121/1.389758.
“Formant bandwidths in helium speech increased as much as 14 times their corresponding bandwidths in normal speech.” ↩
-
Harmonic series (music), Wikipedia. ↩
-
Christopher Dobrian, Harmonic/Overtone Series, Computer Music Pedagogy, University of California, Irvine.
“The relationship of the amplitude of the fundamental frequency component to the amplitude of the harmonics created by an instrument is referred to as the spectral envelope.” ↩
-
Spectral slope, Glottopedia. The −12 dB per octave figure is an idealisation derived from triangular source pulses; real voices vary around it, and breathy or falsetto phonation runs steeper, nearer −18 dB per octave. ↩
-
Acoustic resonance, Wikipedia. ↩
-
Johan Sundberg, Level and Center Frequency of the Singer’s Formant, Journal of Voice 15(2), 2001. The peak near 3 kHz arises from a clustering of the third, fourth and fifth resonances, produced by narrowing the epilaryngeal tube against a widened pharynx. ↩
-
Acoustic cues for the recognition of self-voice and other-voice, PubMed Central.
“Frequencies higher than 2500 Hz … are greatly related to the anatomy of one’s laryngeal cavity, whose anatomical configuration varies between speakers but virtually remains unchanged during articulation of different vowels, and therefore carry individual specificity.” ↩
-
ITU-T Recommendation G.712, Transmission performance characteristics of pulse code modulation channels — the standard that band-limits narrowband telephony to 300–3,400 Hz. The band was chosen as the minimum that preserves both intelligibility and recognition of the speaker. ↩ ↩2
-
Missing fundamental, Wikipedia. ↩
-
Clarence Sasaki, Anatomy and development and physiology of the larynx, GI Motility online (2006).
“Viewed phylogenetically, the primary function of the larynx is its use as a sphincter, protecting the lower airway from the intrusion of liquids and food … The third function of the larynx, phonation … appears to be a late phylogenetic acquisition.” ↩
-
Anatomy, Head and Neck: Cricoid Cartilage, StatPearls, NCBI Bookshelf. The cricoid is the only complete cartilage ring encircling any part of the airway; the tracheal cartilages below it are C-shaped and open posteriorly, where the trachea abuts the oesophagus. ↩
-
Cricothyroid muscle, Wikipedia. Contraction rotates the thyroid cartilage at the cricothyroid joint, stretching, tensing and thinning the vocal folds; it is the principal tensor and pitch-raiser, and the only intrinsic laryngeal muscle not supplied by the recurrent laryngeal nerve. ↩
-
Thyroarytenoid muscle, Wikipedia.
“Its main use is to draw the arytenoid cartilages forward toward the thyroid, thus relaxing and shortening the vocal folds.” Its deeper fibres form the vocalis, a band lying against and adherent to the vocal ligament. ↩ ↩2
-
Passaggio, Wikipedia; see also Passaggio: the register transition zone in singing, Voice Science. The passaggio is a zone of several adjacent pitches rather than a single threshold. ↩ ↩2
-
Sun et al., Quantity and Distribution of Muscle Spindles in Animal and Human Muscles, International Journal of Molecular Sciences, 2024. Reported densities include inferior oblique at 266.67 spindles per gram, superior oblique at 189.47, and rectus capitis posterior at 98.31.
“We have reservations” — the authors’ own caution about concluding that fine-motor muscles necessarily carry higher spindle densities, given inconsistent counting methods across the literature. ↩
-
Hernández-Morato, Yu & Pitman, A review of the peripheral proprioceptive apparatus in the larynx, Frontiers in Neuroanatomy 17:1114817, 2023. See also Canonical Proprioceptors Are Largely Absent in the Intrinsic Laryngeal Muscles of the Rat Larynx, Journal of Comparative Neurology.
“Between 1950 and 1987, multiple groups observed MuSp in the TA using various traditional stains… Others have found MuSp to be absent in the TA.” ↩
-
“I ask them what they can feel”: proprioception and the voice teacher’s approach, James Cook University research repository. ↩
-
Kristina Simonyan & Barry Horwitz, Laryngeal Motor Cortex and Control of Speech in Humans, The Neuroscientist, 2011. In humans the laryngeal motor cortex sits in primary motor cortex with direct projections to the brainstem nucleus ambiguus; in non-human primates it sits in premotor cortex with only indirect connections. ↩
-
Christophe Coupé, Yoon Mi Oh, Dan Dediu & François Pellegrino, Different languages, similar encoding efficiency: Comparable information rates across the human communicative niche, Science Advances 5(9), 2019. DOI 10.1126/sciadv.aaw2594
“We show here, using quantitative methods on a large cross-linguistic corpus of 17 languages, that the coupling between language-level (information per syllable) and speaker-level (speech rate) properties results in languages encoding similar information rates (~39 bits/s) despite wide differences in each property individually.” ↩ ↩2
-
CNRS, Similar information rates across languages, despite divergent speech rates (2019).
“The 17 languages studied have information densities ranging from 5 (i.e. choice of 2^5 = 32 possible syllables) to 8 (2^8 = 256 syllables) bits per syllable.” ↩
-
Claude E. Shannon, A Mathematical Theory of Communication, Bell System Technical Journal 27, 379–423 and 623–656 (1948). Full text at the Internet Archive · DOI 10.1002/j.1538-7305.1948.tb01338.x. The paper that introduced the bit. ↩
-
Kösem et al., Neural speech tracking in the theta and in the delta frequency band differentially encode clarity and comprehension of speech in noise, Journal of Neuroscience 39(29), 2019; and Effects of syllable rate on neuro-behavioral synchronization across modalities, Neurobiology of Language 4(2), 2023. The framework is Giraud and Poeppel’s. ↩
-
Sean Trott, Do different languages really convey information at the same rate? — a research review setting out the interpretive limits of the 39 bits/s result, in particular that uncertainty over signals is not uncertainty over meanings. ↩
-
Pauline Larrouy-Maestri, David Poeppel & Marc D. Pell, The sound of emotional prosody: nearly 3 decades of research and future directions, Perspectives on Psychological Science, 2025. ↩
-
Klaus R. Scherer, Rainer Banse & Harald G. Wallbott, Emotion inferences from vocal expression correlate across languages and cultures, Journal of Cross-Cultural Psychology 32(1), 2001. ↩
-
Petri Laukka & Hillary Anger Elfenbein, Cross-cultural emotion recognition and in-group advantage in vocal expression: a meta-analysis, Emotion Review 13(1), 2021. Thirty-seven studies, expressers from 26 cultural groups, perceivers from 44. ↩
-
Vivien C. Tartter, Happy talk: perceptual and acoustic effects of smiling on speech, Perception & Psychophysics 27(1), 24–27 (1980). ↩
-
Master protocols in vocal biomarker development to reduce variability and advance clinical precision: a narrative review, Frontiers in Digital Health, 2025; and Using voice and speech data in healthcare: a scoping review of the ethical, legal and social implications, 2025. ↩
-
Roman Jakobson, “Why ‘Mama’ and ‘Papa’?” — first published in Bernard Kaplan and Seymour Wapner (eds.), Perspectives in Psychological Theory: Essays in Honor of Heinz Werner (International Universities Press, 1960), reprinted in Selected Writings I: Phonological Studies (Mouton, 1962). Record
“Often the sucking activities of a child are accompanied by a slight nasal murmur, the only phonation which can be produced when the lips are pressed to the mother’s breast or to the feeding bottle and the mouth is full.” ↩
-
Macquarie University Department of Linguistics, Vocal Tract Resonance; see also National Center for Voice and Speech, How the Vocal Tract Filters Sound.
A vocal tract of uniform cross-section, about 17 cm long, resonates at approximately 500, 1500, 2500 and 3500 Hz — the acoustic definition of a neutral vowel. ↩
-
Babbling, Wikipedia — canonical (reduplicated) babbling begins around six months; the early consonant set is dominated by p, b, t, d, k, g, m, n, and infants worldwide follow the same broad tendencies before native-language influence appears. ↩
-
George Peter Murdock, “Cross-Language Parallels in Parental Kin Terms”, Anthropological Linguistics 1 (1959), 1–5. JSTOR · HRAF record
The survey found “the universal tendency for languages, regardless of their historical relationships, to develop similar words for mother and father on the basis of nursery forms”, with Ma, Na, Pa and Ta significantly over-represented. ↩
-
Mama and papa, Wikipedia — the cross-linguistic pattern, Jakobson’s account, and the counter-examples including Georgian. ↩
-
Wiktionary, მამა (mama) — Georgian for “father”, from Old Georgian მამაჲ, from Proto-Kartvelian. ↩
-
Wiktionary, დედა (deda) — Georgian for “mother”. ↩
-
Huang, Kram & Ahmed, Reduction of Metabolic Cost during Motor Learning of Arm Reaching Dynamics, Journal of Neuroscience 32(6), 2182–2190 (2012). Journal of Neuroscience · PMC full text
“Interestingly, distinct and significant reductions in metabolic power occurred even after muscle activity and coactivation had stabilized.” ↩ ↩2 ↩3
-
Reviews of respiratory muscle function and the tension–time index, following Roussos and Macklem. Respiratory muscle function · Respiratory muscle fatigue and breathing pattern, PubMed. ↩ ↩2
-
American Thoracic Society / European Respiratory Society, ATS/ERS Statement on Respiratory Muscle Testing, American Journal of Respiratory and Critical Care Medicine 166(4), 518–624 (2002). Full statement (PDF)
Muscle fatigue is defined as a reduced force-generating capacity of the muscle at a given level of recruitment, resulting from activity under load, and reversible by rest. ↩
-
Hodges & Richardson, Feedforward contraction of transversus abdominis is not influenced by the direction of arm movement, Experimental Brain Research 114, 362–370 (1997). PubMed ↩
-
Banzett, Lansing & Binks, Air Hunger: A Primal Sensation and a Primary Element of Dyspnea, Comprehensive Physiology 11(2), 1449–1483 (2021). DOI 10.1002/cphy.c200001 · PubMed
Functional neuroimaging shows air hunger activating the insular cortex — an integration centre for homeostatic perceptions including pain and hunger — together with limbic structures associated with anxiety. ↩ ↩2
-
Lansing, Im, Thwing, Legedza & Banzett, The perception of respiratory work and effort can be independent of the perception of air hunger, American Journal of Respiratory and Critical Care Medicine 162(5) (2000). DOI 10.1164/ajrccm.162.5.9907096 · PubMed
“Air hunger ratings changed more steeply when PCO2 was altered and ventilation was constant; work or effort ratings changed more steeply when ventilation was altered and PCO2 was constant.” ↩
-
Hyperventilation syndrome: hypocapnia, respiratory alkalosis, cerebral vasoconstriction and the resulting paradoxical sensation of breathlessness. Medscape · Hyperventilation, Wikipedia. ↩
-
Balban, Neri, Kogon et al., Brief structured respiration practices enhance mood and reduce physiological arousal, Cell Reports Medicine 4(1), 100895 (2023). DOI 10.1016/j.xcrm.2022.100895 · PubMed ↩ ↩2
-
Ingo R. Titze, Voice Training and Therapy With a Semi-Occluded Vocal Tract: Rationale and Scientific Underpinnings, Journal of Speech, Language, and Hearing Research 49(2), 448–459 (2006). Publisher page
“Benefits to the voice are derived from a heightened interaction of the vibration source (vocal folds) with the vocal tract to make the voice more economic and to reduce collision forces of the vocal fold tissues.” ↩ ↩2
-
Costa, Costa, Oliveira & Behlau, Immediate effects of the phonation into a straw exercise, Brazilian Journal of Otorhinolaryngology 77(4) (2011). DOI 10.1590/S1808-86942011000400009 · PubMed ↩
-
Fernández-Lázaro et al., Inspiratory Muscle Training Program Using the PowerBreath: Does It Have Ergogenic Potential for Respiratory and/or Athletic Performance? A Systematic Review with Meta-Analysis, International Journal of Environmental Research and Public Health 18(13), 6703 (2021). DOI 10.3390/ijerph18136703 · PMC full text ↩
-
Lorca-Santiago, Jiménez, Pareja-Galeano & Lorenzo, Inspiratory Muscle Training in Intermittent Sports Modalities: A Systematic Review, International Journal of Environmental Research and Public Health 17(12), 4448 (2020). DOI 10.3390/ijerph17124448 · PMC full text ↩
-
Kwon, Kwon & Lee, Effectiveness of motor sequential learning according to practice schedules in healthy adults; distributed practice versus massed practice, Journal of Physical Therapy Science 27(3) (2015). PubMed · concept: Spacing effect, Wikipedia. ↩
-
Maslan, Leng, Rees, Blalock & Butler, Maximum Phonation Time in Healthy Older Adults, Journal of Voice 25(6), 709–713 (2011). Journal of Voice
“Females and males had mean MPTs of 20.96 (SE = 0.92) and 23.23 (SE = 0.96) seconds, respectively. MPTs did not vary significantly with age or gender.” ↩
-
Speyer et al., Maximum phonation time: variability and reliability, Journal of Voice 24(3), 281–284 (2010). DOI 10.1016/j.jvoice.2008.10.004
“Patients showed significantly shorter maximum phonation times compared with healthy controls (on average, 6.6 seconds shorter).” ↩
-
Eckel & Boone, The S/Z ratio as an indicator of laryngeal pathology, Journal of Speech and Hearing Disorders 46(2), 147–149 (1981). DOI 10.1044/jshd.4602.147
“While no statistical difference was found between the three groups in their ability to sustain /s/, the subjects with laryngeal pathology had significantly lower duration times for /z/ than subjects in the other two groups… The dysphonic subjects with laryngeal pathology produced s/z ratios in excess of 1.4 ninety-five percent of the time.” ↩
-
Cent (music), Wikipedia — the unit, its introduction by Alexander John Ellis in 1885, and the perceptual thresholds.
“99 other notes were interposed, making exactly equal intervals with each other, we should divide the octave into 1200 equal hundredths of an equal semitone, or cents as they may be briefly called.” The article also notes that humans can distinguish a difference in pitch of about 5–6 cents, and recognise 25 cents very reliably. ↩ ↩2
-
Proceedings in Courts of Justice Act 1730 (in force 25 March 1733), and the earlier Pleading in English Act 1362. Wikipedia
“All writs, process, pleadings, rules, orders, indictments, records, judgments, and all proceedings whatsoever in any courts of justice within England… shall be in the English tongue and language only, and not in Latin or French.” ↩
-
Benefit of clergy, Wikipedia — the Latin reading test, usually Psalm 51, that transferred a felony case from the secular courts to the ecclesiastical ones. The literacy test was abolished in 1706; the privilege itself survived until 1827. ↩
-
Isaac Newton, Philosophiæ Naturalis Principia Mathematica (1687), written in Latin; first English translation by Andrew Motte, 1729. Wikipedia ↩
-
Historical literacy is usually measured by whether people could sign their names — a very low bar. In the diocese of Norwich in the late sixteenth century about 61% of men could not; among day labourers in northern England illiteracy remained above 90% around 1600. Our World in Data: Literacy ↩
-
Dante Alighieri, De vulgari eloquentia (c. 1304–1307), written in Latin, arguing that the Italian vernacular deserved the dignity given to Latin. Left unfinished at one and a half books. Wikipedia ↩
-
Luther Bible, Wikipedia. The September Testament of 1522 ran to roughly 3,000–5,000 copies at one guilder each, about two months’ salary for a schoolmaster; a second edition followed in December. The complete Bible was printed by Hans Lufft at Wittenberg in 1534, by which time over 200,000 copies of the New Testament had sold. ↩
-
William Tyndale (c. 1494–1536), Wikipedia — first English New Testament translated from the Greek, printed on the continent in 1526 and smuggled into England; strangled and burned as a heretic on 6 October 1536. Estimates put roughly 83% of the King James New Testament (1611) as Tyndale’s wording. ↩
-
Eltjo Buringh & Jan Luiten van Zanden, Charting the “Rise of the West”: Manuscripts and Printed Books in Europe, A Long-Term Perspective from the Sixth through Eighteenth Centuries, The Journal of Economic History 69(2), 2009. Cambridge Core
Book production grew at roughly one percent a year across the medieval period; the sharp rise after the mid-fifteenth century is attributed to lower book prices and rising literacy. ↩
-
Sacrosanctum Concilium, the Second Vatican Council’s Constitution on the Sacred Liturgy, promulgated 4 December 1963, which extended the use of vernacular languages in Catholic worship. Wikipedia ↩
-
SlashData, Global Developer Population Trends 2025 — an estimated 47.2 million developers worldwide, of whom about 36.5 million are professionals. SlashData ↩
-
FLOW-MATIC, Wikipedia — developed under Grace Hopper at Remington Rand for the UNIVAC I between 1955 and 1959, the first programming language to express operations in English-like statements rather than mathematical symbols, and a direct ancestor of COBOL. ↩
-
Johannes Trithemius, De laude scriptorum manualium (“In Praise of Scribes”), written 1492, printed 1494. As abbot of Sponheim he grew the monastery library from about 40 volumes to 2,000, many of them printed, and described printing itself as a marvellous art. Wikipedia · Internet Archive ↩
-
Mains electricity by country, Wikipedia. Mains supply splits into roughly 100–127 V and 220–240 V systems; modern switch-mode power supplies are rated for universal input across 100–240 V, so the conversion happens in the device rather than in a separate converter. ↩