Play a Spanish news show, then play a Thai one. The Spanish comes at you like a machine gun. The Thai sounds like someone with all afternoon. We all notice this. And most of us draw the same lesson from it: some languages are faster, so their speakers must get more said per minute.
They do not. In 2019 a team of linguists measured it across seventeen languages. All of them send information at close to the same rate, near 39 bits per second.1 The fast languages are not saying more. They pay more syllables for the same content.
How it was measured
Christophe Coupé, Yoon Mi Oh, Dan Dediu and François Pellegrino taped speakers reading the same set of texts. The texts were put into each of the seventeen languages. Keeping the meaning the same is the whole trick. If everyone sends the same thing, then any gap in how long it takes is a gap in the wrapping, not in what is inside.
Then they measured two numbers for each language. The first is easy: syllable rate, syllables per second. That is just counting. It varies a lot. Japanese came in near 8.0 syllables per second and Spanish near 7.7. Thai came in near 4.7 and Vietnamese near 5.3. The fast ones are about 50 percent faster than the slow ones. That is just the gap your ear told you was there.
The second number is the good one: information density, in bits per syllable. Japanese sits near 5 bits per syllable. English is a bit over 7. Vietnamese tops the set at about 8.2
Now multiply. Eight syllables a second at five bits each is forty bits a second. Five and a bit syllables a second at eight bits each is about forty-two. Those are rough averages multiplied. So treat them as a neighbourhood, not a result. But the neighbourhood is the point. The study’s own figure across all seventeen languages was 39.15 bits per second. And information rate varied far less from language to language than either of the two numbers that make it.
What a bit actually is
The word “bit” is doing real work in that sentence. It does not mean what it means in day to day speech. It comes from Claude Shannon’s A Mathematical Theory of Communication, from 1948. That is the paper that made the field and coined the term.3 He is the same Shannon who runs the case in /on/entropy. This is the second deep dive in a row where he holds the whole thing up.
Shannon stopped asking what a message means. He asked how much it narrows down. Information, in his sense, is doubt cut away. A coin flip has two ways it can land. So learning one flip gives you one bit. Learning four flips gives you four bits. Four flips have sixteen ways to land, and sixteen is two times two times two times two. That is the whole formula. The number of bits is the number of times you halve the set of choices to get down to one.
So the worth of a symbol rests on how many other symbols could have shown up in its place. A letter drawn from a 26-letter alphabet tells you more than a digit drawn from ten. It rules out more of what the message might have been. This is why a password of random letters is harder to crack than a PIN of the same length. It is the same fact in a different mood.
Syllables work the same way. A language with a few hundred syllables gives each one a lot of work to do. Hearing it rules out a lot. A language with only a few dozen gives each one much less. In the study, the effective number of choices ran from about 32 syllables at the low end to about 256 at the high end. Effective is the right word there, and it matters. It is not the raw count in the dictionary. Some syllables are common and some are rare. And a syllable you were expecting anyway brings almost nothing when it lands.
Gross and net
This gives you a split worth keeping. It is your own hunch made exact.
Gross rate is syllables per second. It is what you hear. It is easy to count. And it differs wildly from language to language.
Net rate is bits per second. It is what gets across. It is hard to count. And it barely moves.
Your ear only ever reports the gross figure. That is why the trick of the ear holds so well. When Spanish sounds like it is outrunning you, you are hearing a language spend lots of cheap syllables. Each one has cut the field only a little, so it needs a lot of them. Vietnamese sounds slow because each syllable has already cut away more. There is less left to say.
The same trade, in a zip file
Take a text file and zip it. The zipped one is much smaller. And every byte in it is now much harder to guess. Before, the file was full of the letter e and the word the and long runs of the same sign. You could have guessed all of it. The whole job of a zip file is to delete what you could have guessed. What comes out the other side looks like noise. Noise is what a file looks like when nothing in it can be guessed. Which is the same as saying every byte is at full load.
The content did not change. The box got shorter and each unit of box got denser. The two changes cancel out. That is the trade. It is the same trade the seventeen languages are making. Vietnamese is the compressed file. Japanese is the uncompressed one. Neither is better. Neither says more. And if you measure the thing that counts — total content divided by total time — you cannot tell them apart.
Most people have handled the sound version of this. A song at 320 kilobits per second and the same song at 128 sound different. Their bitrates differ. And bitrate is bits per second: how much the encoder may spend on each moment of sound. Human speech runs at a fixed bitrate of about 39 bits per second. And every language on earth has landed on the same setting on its own.
Why 39 is not settled
The number looks like a limit set by the body. Something is capping it. Seventeen unrelated languages do not land in the same narrow band by chance. And the likely suspects all sit in the head, not in the mouth.
The best backed guess is on the listener’s side. Understanding speech seems to depend on brain rhythms locking onto the rhythm of the signal coming in. In particular, on activity in the theta band, which cycles at about 4 to 8 times per second. That is almost exactly the syllable rate of every language measured.4 If the decoder runs at a fixed rate, then the encoder gains nothing by going past it. Every language would be pushed toward the same ceiling by the plain fact that going faster gets you nothing. That is a good story, and the proof for the rhythm is real. But the last step is a guess, not a measure. Neural tracking of the speech envelope is needed for understanding, and it is not enough for it.
A rival account puts the limit on the speaker instead. Not how fast you can hear, but how fast you can put together what you are about to say. Listeners can follow taped speech played back well above normal speed without much trouble. That is awkward for a bottleneck made of hearing alone.
And the result itself carries a caveat worth saying plainly. What was measured is doubt over syllables, not doubt over meanings, and those are not the same thing.5 It also cannot be pulled fully apart from a simpler story. Languages with fewer syllables to draw on tend to be spoken faster. That alone might make the whole pattern, with nothing clever going on. The trade-off is solid. The claim that 39 bits per second is a best point, rather than a side effect, is not.
References
-
Christophe Coupé, Yoon Mi Oh, Dan Dediu & François Pellegrino, Different languages, similar encoding efficiency: Comparable information rates across the human communicative niche, Science Advances 5(9), 2019. DOI 10.1126/sciadv.aaw2594
“We show here, using quantitative methods on a large cross-linguistic corpus of 17 languages, that the coupling between language-level (information per syllable) and speaker-level (speech rate) properties results in languages encoding similar information rates (~39 bits/s) despite wide differences in each property individually.” ↩
-
CNRS, Similar information rates across languages, despite divergent speech rates (2019).
“The 17 languages studied have information densities ranging from 5 (i.e. choice of 2^5 = 32 possible syllables) to 8 (2^8 = 256 syllables) bits per syllable.” ↩
-
Claude E. Shannon, A Mathematical Theory of Communication, Bell System Technical Journal 27, 379–423 and 623–656 (1948). Full text at the Internet Archive · DOI 10.1002/j.1538-7305.1948.tb01338.x. The paper that introduced the bit. ↩
-
Kösem et al., Neural speech tracking in the theta and in the delta frequency band differentially encode clarity and comprehension of speech in noise, Journal of Neuroscience 39(29), 2019; and Effects of syllable rate on neuro-behavioral synchronization across modalities, Neurobiology of Language 4(2), 2023. The framework is Giraud and Poeppel’s. ↩
-
Sean Trott, Do different languages really convey information at the same rate? — a research review setting out the interpretive limits of the 39 bits/s result, in particular that uncertainty over signals is not uncertainty over meanings. ↩