1.1 Sound is a pattern

  1. Longitudinal wave, Wikipedia.

    “Mechanical longitudinal waves are also called compressional or compression waves, because they produce compression and rarefaction when travelling through a medium, and pressure waves, because they produce increases and decreases in pressure.”

  2. Wind wave, Wikipedia.

    “Parcels near the surface move not plainly up and down but in circular orbits: forward above and backward below (compared to the wave propagation direction).”

  3. The Physics Classroom, The Speed of Sound; see also Speed of Sound in Air, The Physics Factbook. 343 m/s is the standard dry-air figure at 20 °C.

1.2 Voice is source + filter

  1. Gunnar Fant, Acoustic Theory of Speech Production (The Hague: Mouton, 1960). Publisher preview · Semantic Scholar record

  2. Johan Sundberg and colleagues, The Gunnar Fant Legacy in the Study of Vocal Acoustics.

    “It gives a highly accurate description of the speech signal and … explains how vowels and consonants get their acoustic properties. It is general, language-independent and valid for both normal and disordered speech.”

  3. Janwillem van den Berg, Myoelastic-Aerodynamic Theory of Voice Production, Journal of Speech and Hearing Research 1(3), 227–244 (1958). DOI 10.1044/jshr.0103.227. Based on a paper given at the Chicago International Voice Conference, May 1957.

  4. Ingo R. Titze, The physics of small-amplitude oscillation of the vocal folds, Journal of the Acoustical Society of America 83(4), 1536–1552 (1988).

  5. Holmberg, Hillman & Perkell (1988), as summarised in Average Speaking Frequencies: F0 Norms by Age, Sex, and Hormonal Status, Voice Science. Adult male mean about 116 Hz (range about 93–135 Hz); adult female mean about 205 Hz (range about 162–238 Hz).

  6. Formant, Wikipedia.

    “Most often the two first formants, F1 and F2, are sufficient to identify the vowel.”

  7. Gordon E. Peterson & Harold L. Barney, Control Methods Used in a Study of the Vowels, Journal of the Acoustical Society of America 24(2), 175–184 (1952). DOI 10.1121/1.1906875. Ten vowels, 76 speakers, at Bell Telephone Laboratories.

  8. Hamid Reza Sharifzadeh, Ian Vince McLoughlin & Martin J. Russell, A Comprehensive Vowel Space for Whispered Speech, Journal of Voice 26(2), e49–e56 (2012).

  9. The Physics Classroom, The Speed of Sound; see also Speed of Sound in Air, The Physics Factbook. 343 m/s is the standard dry-air figure at 20 °C.

  10. Helium, Wikipedia — speed of sound listed as 972 m/s.

  11. E. O. Belcher & S. Hatlestad, Formant frequencies, bandwidths, and Qs in helium speech, Journal of the Acoustical Society of America 74(2), 428–432 (1983). DOI 10.1121/1.389758.

    “Formant bandwidths in helium speech increased as much as 14 times their corresponding bandwidths in normal speech.”

1.3 A note is a stack

  1. Harmonic series (music), Wikipedia.

  2. Christopher Dobrian, Harmonic/Overtone Series, Computer Music Pedagogy, University of California, Irvine.

    “The relationship of the amplitude of the fundamental frequency component to the amplitude of the harmonics created by an instrument is referred to as the spectral envelope.”

  3. Spectral slope, Glottopedia. The −12 dB per octave figure is an idealisation derived from triangular source pulses; real voices vary around it, and breathy or falsetto phonation runs steeper, nearer −18 dB per octave.

  4. Acoustic resonance, Wikipedia.

  5. Johan Sundberg, Level and Center Frequency of the Singer’s Formant, Journal of Voice 15(2), 2001. The peak near 3 kHz arises from a clustering of the third, fourth and fifth resonances, produced by narrowing the epilaryngeal tube against a widened pharynx.

  6. Acoustic cues for the recognition of self-voice and other-voice, PubMed Central.

    “Frequencies higher than 2500 Hz … are greatly related to the anatomy of one’s laryngeal cavity, whose anatomical configuration varies between speakers but virtually remains unchanged during articulation of different vowels, and therefore carry individual specificity.”

  7. ITU-T Recommendation G.712, Transmission performance characteristics of pulse code modulation channels — the standard that band-limits narrowband telephony to 300–3,400 Hz. The band was chosen as the minimum that preserves both intelligibility and recognition of the speaker.

  8. Missing fundamental, Wikipedia.

1.4 How singing works

  1. Clarence Sasaki, Anatomy and development and physiology of the larynx, GI Motility online (2006).

    “Viewed phylogenetically, the primary function of the larynx is its use as a sphincter, protecting the lower airway from the intrusion of liquids and food … The third function of the larynx, phonation … appears to be a late phylogenetic acquisition.”

  2. Anatomy, Head and Neck: Cricoid Cartilage, StatPearls, NCBI Bookshelf. The cricoid is the only complete cartilage ring encircling any part of the airway; the tracheal cartilages below it are C-shaped and open posteriorly, where the trachea abuts the oesophagus.

  3. Cricothyroid muscle, Wikipedia. Contraction rotates the thyroid cartilage at the cricothyroid joint, stretching, tensing and thinning the vocal folds; it is the principal tensor and pitch-raiser, and the only intrinsic laryngeal muscle not supplied by the recurrent laryngeal nerve.

  4. Thyroarytenoid muscle, Wikipedia.

    “Its main use is to draw the arytenoid cartilages forward toward the thyroid, thus relaxing and shortening the vocal folds.” Its deeper fibres form the vocalis, a band lying against and adherent to the vocal ligament.

  5. Passaggio, Wikipedia; see also Passaggio: the register transition zone in singing, Voice Science. The passaggio is a zone of several adjacent pitches rather than a single threshold.

  6. Sun et al., Quantity and Distribution of Muscle Spindles in Animal and Human Muscles, International Journal of Molecular Sciences, 2024. Reported densities include inferior oblique at 266.67 spindles per gram, superior oblique at 189.47, and rectus capitis posterior at 98.31.

    “We have reservations” — the authors’ own caution about concluding that fine-motor muscles necessarily carry higher spindle densities, given inconsistent counting methods across the literature.

  7. Uwe Proske & Simon C. Gandevia, The proprioceptive senses: their roles in signaling body shape, body position and movement, and muscle force, Physiological Reviews 92(4), 1651–1697 (2012). DOI 10.1152/physrev.00048.2011. Gamma motor neurons reset intrafusal fibre length so the spindle stays sensitive as the parent muscle shortens.

  8. J. E. Macintosh & Nikolai Bogduk, The biomechanics of the lumbar multifidus, Clinical Biomechanics 1(4), 205–213 (1986). Segmental fibres of multifidus span one vertebral pair and can report local orientation that a long multi-level muscle would sum and lose.

  9. Peck, Buxton & Nitz, A comparison of spindle concentrations in large and small muscles acting in parallel combinations, Journal of Morphology 180(3), 243–252 (1984); and Hallgren, Rowan, Ackermann et al., Implied Evidence of the Functional Role of the Rectus Capitis Posterior Muscles, Journal of the American Osteopathic Association 120(6), 395–403 (2020). The muscle is too small to move the head in a useful way; the spindle density is the point.

  10. Hernández-Morato, Yu & Pitman, A review of the peripheral proprioceptive apparatus in the larynx, Frontiers in Neuroanatomy 17:1114817, 2023. See also Canonical Proprioceptors Are Largely Absent in the Intrinsic Laryngeal Muscles of the Rat Larynx, Journal of Comparative Neurology.

    “Between 1950 and 1987, multiple groups observed MuSp in the TA using various traditional stains… Others have found MuSp to be absent in the TA.”

  11. “I ask them what they can feel”: proprioception and the voice teacher’s approach, James Cook University research repository.

  12. Kristina Simonyan & Barry Horwitz, Laryngeal Motor Cortex and Control of Speech in Humans, The Neuroscientist, 2011. In humans the laryngeal motor cortex sits in primary motor cortex with direct projections to the brainstem nucleus ambiguus; in non-human primates it sits in premotor cortex with only indirect connections.

1.5 The bandwidth of voice

  1. Christophe Coupé, Yoon Mi Oh, Dan Dediu & François Pellegrino, Different languages, similar encoding efficiency: Comparable information rates across the human communicative niche, Science Advances 5(9), 2019. DOI 10.1126/sciadv.aaw2594

    “We show here, using quantitative methods on a large cross-linguistic corpus of 17 languages, that the coupling between language-level (information per syllable) and speaker-level (speech rate) properties results in languages encoding similar information rates (~39 bits/s) despite wide differences in each property individually.”

  2. CNRS, Similar information rates across languages, despite divergent speech rates (2019).

    “The 17 languages studied have information densities ranging from 5 (i.e. choice of 2^5 = 32 possible syllables) to 8 (2^8 = 256 syllables) bits per syllable.”

  3. Claude E. Shannon, A Mathematical Theory of Communication, Bell System Technical Journal 27, 379–423 and 623–656 (1948). Full text at the Internet Archive · DOI 10.1002/j.1538-7305.1948.tb01338.x. The paper that introduced the bit.

  4. Kösem et al., Neural speech tracking in the theta and in the delta frequency band differentially encode clarity and comprehension of speech in noise, Journal of Neuroscience 39(29), 2019; and Effects of syllable rate on neuro-behavioral synchronization across modalities, Neurobiology of Language 4(2), 2023. The framework is Giraud and Poeppel’s.

  5. Sean Trott, Do different languages really convey information at the same rate? — a research review setting out the interpretive limits of the 39 bits/s result, in particular that uncertainty over signals is not uncertainty over meanings.

1.6 Voice beyond words

  1. Pauline Larrouy-Maestri, David Poeppel & Marc D. Pell, The sound of emotional prosody: nearly 3 decades of research and future directions, Perspectives on Psychological Science, 2025.

  2. Klaus R. Scherer, Rainer Banse & Harald G. Wallbott, Emotion inferences from vocal expression correlate across languages and cultures, Journal of Cross-Cultural Psychology 32(1), 2001.

  3. Petri Laukka & Hillary Anger Elfenbein, Cross-cultural emotion recognition and in-group advantage in vocal expression: a meta-analysis, Emotion Review 13(1), 2021. Thirty-seven studies, expressers from 26 cultural groups, perceivers from 44.

  4. Vivien C. Tartter, Happy talk: perceptual and acoustic effects of smiling on speech, Perception & Psychophysics 27(1), 24–27 (1980).

  5. Master protocols in vocal biomarker development to reduce variability and advance clinical precision: a narrative review, Frontiers in Digital Health, 2025; and Using voice and speech data in healthcare: a scoping review of the ethical, legal and social implications, 2025.

2.1 Our first sounds

  1. Roman Jakobson, “Why ‘Mama’ and ‘Papa’?” — first published in Bernard Kaplan and Seymour Wapner (eds.), Perspectives in Psychological Theory: Essays in Honor of Heinz Werner (International Universities Press, 1960), reprinted in Selected Writings I: Phonological Studies (Mouton, 1962). Record

    “Often the sucking activities of a child are accompanied by a slight nasal murmur, the only phonation which can be produced when the lips are pressed to the mother’s breast or to the feeding bottle and the mouth is full.”

  2. Macquarie University Department of Linguistics, Vocal Tract Resonance; see also National Center for Voice and Speech, How the Vocal Tract Filters Sound.

    A vocal tract of uniform cross-section, about 17 cm long, resonates at approximately 500, 1500, 2500 and 3500 Hz — the acoustic definition of a neutral vowel.

  3. Babbling, Wikipedia — canonical (reduplicated) babbling begins around six months; the early consonant set is dominated by p, b, t, d, k, g, m, n, and infants worldwide follow the same broad tendencies before native-language influence appears.

  4. George Peter Murdock, “Cross-Language Parallels in Parental Kin Terms”, Anthropological Linguistics 1 (1959), 1–5. JSTOR · HRAF record

    The survey found “the universal tendency for languages, regardless of their historical relationships, to develop similar words for mother and father on the basis of nursery forms”, with Ma, Na, Pa and Ta significantly over-represented.

  5. Mama and papa, Wikipedia — the cross-linguistic pattern, Jakobson’s account, and the counter-examples including Georgian.

  6. Wiktionary, მამა (mama) — Georgian for “father”, from Old Georgian მამაჲ, from Proto-Kartvelian.

  7. Wiktionary, დედა (deda) — Georgian for “mother”.

2.2 Why singing tires you

  1. Huang, Kram & Ahmed, Reduction of Metabolic Cost during Motor Learning of Arm Reaching Dynamics, Journal of Neuroscience 32(6), 2182–2190 (2012). Journal of Neuroscience · PMC full text

    “Interestingly, distinct and significant reductions in metabolic power occurred even after muscle activity and coactivation had stabilized.”

  2. Reviews of respiratory muscle function and the tension–time index, following Roussos and Macklem. Respiratory muscle function · Respiratory muscle fatigue and breathing pattern, PubMed.

  3. American Thoracic Society / European Respiratory Society, ATS/ERS Statement on Respiratory Muscle Testing, American Journal of Respiratory and Critical Care Medicine 166(4), 518–624 (2002). Full statement (PDF)

    Muscle fatigue is defined as a reduced force-generating capacity of the muscle at a given level of recruitment, resulting from activity under load, and reversible by rest.

  4. Hodges & Richardson, Feedforward contraction of transversus abdominis is not influenced by the direction of arm movement, Experimental Brain Research 114, 362–370 (1997). PubMed

  5. Banzett, Lansing & Binks, Air Hunger: A Primal Sensation and a Primary Element of Dyspnea, Comprehensive Physiology 11(2), 1449–1483 (2021). DOI 10.1002/cphy.c200001 · PubMed

    Functional neuroimaging shows air hunger activating the insular cortex — an integration centre for homeostatic perceptions including pain and hunger — together with limbic structures associated with anxiety.

  6. Lansing, Im, Thwing, Legedza & Banzett, The perception of respiratory work and effort can be independent of the perception of air hunger, American Journal of Respiratory and Critical Care Medicine 162(5) (2000). DOI 10.1164/ajrccm.162.5.9907096 · PubMed

    “Air hunger ratings changed more steeply when PCO2 was altered and ventilation was constant; work or effort ratings changed more steeply when ventilation was altered and PCO2 was constant.”

  7. Hyperventilation syndrome: hypocapnia, respiratory alkalosis, cerebral vasoconstriction and the resulting paradoxical sensation of breathlessness. Medscape · Hyperventilation, Wikipedia.

  8. Ingo R. Titze, Sheila S. Schmidt & Michael Titze, Phonation threshold pressure in a physical model of the vocal fold mucosa, Journal of the Acoustical Society of America 97(5), 3080–3084 (1995). DOI 10.1121/1.411870. “There was a consistent hysteresis effect; that is, phonation threshold pressure was always lower for oscillation offset than onset.” See also Jorge C. Lucero, The minimum lung pressure to sustain vocal fold oscillation, JASA 98(2), 779–784 (1995).

  9. Thomas J. Hixon, Respiratory Function in Speech and Song (San Diego: Singular, 1991); and the earlier Hixon, Siebens & Ewanowski report, Respiratory Mechanics during Speech Production, JASA 44(1), 376 (1968). At high lung volume the passive recoil often exceeds the pressure you want, so the inhale muscles brake the spring. See also Ladefoged & Loeb on checking recoil.

  10. Elliot, Sundberg & Gramming, What happens during vocal warm-up?, Journal of Voice 9(1), 37–44 (1995). Warm-up lowers the pressure needed to phonate; Titze’s 1995 mucosa model (above) ties that drop to lower viscosity.

2.3 Training what you cannot feel

  1. Balban, Neri, Kogon et al., Brief structured respiration practices enhance mood and reduce physiological arousal, Cell Reports Medicine 4(1), 100895 (2023). DOI 10.1016/j.xcrm.2022.100895 · PubMed

  2. Ingo R. Titze, Voice Training and Therapy With a Semi-Occluded Vocal Tract: Rationale and Scientific Underpinnings, Journal of Speech, Language, and Hearing Research 49(2), 448–459 (2006). Publisher page

    “Benefits to the voice are derived from a heightened interaction of the vibration source (vocal folds) with the vocal tract to make the voice more economic and to reduce collision forces of the vocal fold tissues.”

  3. Costa, Costa, Oliveira & Behlau, Immediate effects of the phonation into a straw exercise, Brazilian Journal of Otorhinolaryngology 77(4) (2011). DOI 10.1590/S1808-86942011000400009 · PubMed

  4. Fernández-Lázaro et al., Inspiratory Muscle Training Program Using the PowerBreath: Does It Have Ergogenic Potential for Respiratory and/or Athletic Performance? A Systematic Review with Meta-Analysis, International Journal of Environmental Research and Public Health 18(13), 6703 (2021). DOI 10.3390/ijerph18136703 · PMC full text

  5. Lorca-Santiago, Jiménez, Pareja-Galeano & Lorenzo, Inspiratory Muscle Training in Intermittent Sports Modalities: A Systematic Review, International Journal of Environmental Research and Public Health 17(12), 4448 (2020). DOI 10.3390/ijerph17124448 · PMC full text

  6. Kwon, Kwon & Lee, Effectiveness of motor sequential learning according to practice schedules in healthy adults; distributed practice versus massed practice, Journal of Physical Therapy Science 27(3) (2015). PubMed · concept: Spacing effect, Wikipedia.

2.4 The feedback loop

  1. Maslan, Leng, Rees, Blalock & Butler, Maximum Phonation Time in Healthy Older Adults, Journal of Voice 25(6), 709–713 (2011). Journal of Voice

    “Females and males had mean MPTs of 20.96 (SE = 0.92) and 23.23 (SE = 0.96) seconds, respectively. MPTs did not vary significantly with age or gender.”

  2. Speyer et al., Maximum phonation time: variability and reliability, Journal of Voice 24(3), 281–284 (2010). DOI 10.1016/j.jvoice.2008.10.004

    “Patients showed significantly shorter maximum phonation times compared with healthy controls (on average, 6.6 seconds shorter).”

  3. Eckel & Boone, The S/Z ratio as an indicator of laryngeal pathology, Journal of Speech and Hearing Disorders 46(2), 147–149 (1981). DOI 10.1044/jshd.4602.147

    “While no statistical difference was found between the three groups in their ability to sustain /s/, the subjects with laryngeal pathology had significantly lower duration times for /z/ than subjects in the other two groups… The dysphonic subjects with laryngeal pathology produced s/z ratios in excess of 1.4 ninety-five percent of the time.”

  4. Cent (music), Wikipedia — the unit, its introduction by Alexander John Ellis in 1885, and the perceptual thresholds.

  5. Sonic Visualiser, Centre for Digital Music, Queen Mary University of London. The layer-property labels and zoom-wheel controls above match the 4.3 reference manual.

    “99 other notes were interposed, making exactly equal intervals with each other, we should divide the octave into 1200 equal hundredths of an equal semitone, or cents as they may be briefly called.” The article also notes that humans can distinguish a difference in pitch of about 5–6 cents, and recognise 25 cents very reliably.

2.5 Voice as interface

  1. Proceedings in Courts of Justice Act 1730 (in force 25 March 1733), and the earlier Pleading in English Act 1362. Wikipedia

    “All writs, process, pleadings, rules, orders, indictments, records, judgments, and all proceedings whatsoever in any courts of justice within England… shall be in the English tongue and language only, and not in Latin or French.”

  2. Benefit of clergy, Wikipedia — the Latin reading test, usually Psalm 51, that transferred a felony case from the secular courts to the ecclesiastical ones. The literacy test was abolished in 1706; the privilege itself survived until 1827.

  3. Isaac Newton, Philosophiæ Naturalis Principia Mathematica (1687), written in Latin; first English translation by Andrew Motte, 1729. Wikipedia

  4. Historical literacy is usually measured by whether people could sign their names — a very low bar. In the diocese of Norwich in the late sixteenth century about 61% of men could not; among day labourers in northern England illiteracy remained above 90% around 1600. Our World in Data: Literacy

  5. Dante Alighieri, De vulgari eloquentia (c. 1304–1307), written in Latin, arguing that the Italian vernacular deserved the dignity given to Latin. Left unfinished at one and a half books. Wikipedia

  6. Luther Bible, Wikipedia. The September Testament of 1522 ran to roughly 3,000–5,000 copies at one guilder each, about two months’ salary for a schoolmaster; a second edition followed in December. The complete Bible was printed by Hans Lufft at Wittenberg in 1534, by which time over 200,000 copies of the New Testament had sold.

  7. William Tyndale (c. 1494–1536), Wikipedia — first English New Testament translated from the Greek, printed on the continent in 1526 and smuggled into England; strangled and burned as a heretic on 6 October 1536. Estimates put roughly 83% of the King James New Testament (1611) as Tyndale’s wording.

  8. Eltjo Buringh & Jan Luiten van Zanden, Charting the “Rise of the West”: Manuscripts and Printed Books in Europe, A Long-Term Perspective from the Sixth through Eighteenth Centuries, The Journal of Economic History 69(2), 2009. Cambridge Core

    Book production grew at roughly one percent a year across the medieval period; the sharp rise after the mid-fifteenth century is attributed to lower book prices and rising literacy.

  9. Sacrosanctum Concilium, the Second Vatican Council’s Constitution on the Sacred Liturgy, promulgated 4 December 1963, which extended the use of vernacular languages in Catholic worship. Wikipedia

  10. SlashData, Global Developer Population Trends 2025 — an estimated 47.2 million developers worldwide, of whom about 36.5 million are professionals. SlashData

  11. FLOW-MATIC, Wikipedia — developed under Grace Hopper at Remington Rand for the UNIVAC I between 1955 and 1959, the first programming language to express operations in English-like statements rather than mathematical symbols, and a direct ancestor of COBOL.

  12. Johannes Trithemius, De laude scriptorum manualium (“In Praise of Scribes”), written 1492, printed 1494. As abbot of Sponheim he grew the monastery library from about 40 volumes to 2,000, many of them printed, and described printing itself as a marvellous art. Wikipedia · Internet Archive

  13. Christophe Coupé, Yoon Mi Oh, Dan Dediu & François Pellegrino, Different languages, similar encoding efficiency: Comparable information rates across the human communicative niche, Science Advances 5(9), 2019. DOI 10.1126/sciadv.aaw2594

    “We show here, using quantitative methods on a large cross-linguistic corpus of 17 languages, that the coupling between language-level (information per syllable) and speaker-level (speech rate) properties results in languages encoding similar information rates (~39 bits/s) despite wide differences in each property individually.”

  14. Mains electricity by country, Wikipedia. Mains supply splits into roughly 100–127 V and 220–240 V systems; modern switch-mode power supplies are rated for universal input across 100–240 V, so the conversion happens in the device rather than in a separate converter.