The 39 bits is only the words. It is what a perfect transcript would catch. And a perfect transcript is famously not the talk itself.
Everything else your voice does rides a second channel that nobody plans for. Pitch shape, loudness, timbre, breathiness, the length of your pauses, where you speed up and where you stall. Linguists call this prosody, the tune and timing laid over the words rather than inside them. And it carries a live report on the state of the speaker.1
Some of that report is well proven. Listeners name emotions from voice alone at rates well above chance. And they do it across languages they do not speak. A large cross-cultural study found scores near 66 percent. They ran from about 74 percent in Germany down to about 52 percent in Indonesia. The pattern of mistakes was almost the same everywhere. The same emotions get mixed up with each other whatever the listener’s culture, and that is a stronger result than the score.2 A meta-analysis of 37 such studies backs the cross-cultural effect and adds an in-group edge. You read your own culture’s voices better, and the further apart two cultures sit, the wider the gap.3
The cleanest case is one you can test today, and it works by a route that leads straight back to the tube. Smiling is audible. Pull your lips back and you shorten the vocal tract. A shorter tube rings higher, so the formants rise. Vivien Tartter measured this in 1980. Smiling raised both the fundamental frequency and the formant frequencies for every speaker tested. And listeners picked the smiled tapes out of pairs at well above chance.4 Nobody is trying to send a smile. The change of shape does it for free. The filter cannot help but report its own shape.
Past emotion, we have to be less sure. Mental load slows speech rate and draws out pauses, and that is steady enough to measure in groups. The wider field of vocal biomarkers — using voice to spot disease, depression, fatigue, or being drunk — is real work with real signal. And right now it is oversold. The problems that keep coming up are the ones you would expect. Results that do not carry over between groups of people or recording setups. Small samples tied to one region. And models that read a voice with no knowledge of the person it belongs to.5 Voice clearly carries health information. Reading it well out of one tape, from a stranger, is not a solved problem. And products that claim it is are ahead of the proof.
This is also why a transcript of a good talk always reads thinner than the talk was. Nothing went missing from the words. The side channel just has no column in the file. The pause before someone answered. The drop in pitch that meant they had decided. The breath that meant they had not. All of it was information you used at the time, and none of it lives through the trip to text.
Which is the honest reason writing is hard. Text is voice with the side channel stripped out. Every trick writers reach for — italics, punctuation, a sentence cut short on purpose, a paragraph break set where a breath would go — is a fake limb for prosody, and a poor one. You are working with 39 bits per second and nothing else. And you are trying to rebuild a signal that first came in with a second track running under it.
References
-
Pauline Larrouy-Maestri, David Poeppel & Marc D. Pell, The sound of emotional prosody: nearly 3 decades of research and future directions, Perspectives on Psychological Science, 2025. ↩
-
Klaus R. Scherer, Rainer Banse & Harald G. Wallbott, Emotion inferences from vocal expression correlate across languages and cultures, Journal of Cross-Cultural Psychology 32(1), 2001. ↩
-
Petri Laukka & Hillary Anger Elfenbein, Cross-cultural emotion recognition and in-group advantage in vocal expression: a meta-analysis, Emotion Review 13(1), 2021. Thirty-seven studies, expressers from 26 cultural groups, perceivers from 44. ↩
-
Vivien C. Tartter, Happy talk: perceptual and acoustic effects of smiling on speech, Perception & Psychophysics 27(1), 24–27 (1980). ↩
-
Master protocols in vocal biomarker development to reduce variability and advance clinical precision: a narrative review, Frontiers in Digital Health, 2025; and Using voice and speech data in healthcare: a scoping review of the ethical, legal and social implications, 2025. ↩