A training plan is a loop. A loop needs a sensor. When a squat is too heavy you know it. The muscle files a report. The report comes to you as a feeling. The gear that makes your voice does not file reports. The cricothyroid and the thyroarytenoid are as thick with nerves as any part you own. None of that wiring reaches your mind. You can feel the work of singing. You cannot feel which muscle is doing it. You cannot feel if it is doing it well.

So you borrow a sensor from outside. That move is not special to voice. You have no detector at all for carbon monoxide. That is just why it kills people in their sleep. So we build a small box that beeps. You cannot feel your heart rate to within ten beats. So you strap a watch to your wrist. You cannot feel four pounds. So you stand on a scale. Each time, a number from outside stands in for a sense you were never issued. And the number is not a stand-in for the feeling. It is the feeling, coming in through a different door.

Biofeedback is not a technique. It is a prosthetic sense.

Voice is a very good fit for the prosthesis. A microphone is cheap. It is very exact. And the thing you train already leaves the body as a signal in air. A knee has to be guessed from video. A heart has to be read from electrodes on skin. But a voice comes to you pre-measured. Four numbers do most of the work. You can start taking all four today with a phone and a stopwatch.

The voice in your head is a lie

Before any of the numbers, one correction that decides how you read all of them.

The voice you hear while you sing is not the voice in the room. Some of it reaches your ears through the air, the ordinary way. Much of it arrives through the bones of your skull, straight from the larynx to the inner ear, skipping the room entirely. Bone carries low tones far better than high ones. So the inside version is warmer and deeper than the one anyone else gets. This is why nearly every person is mildly let down the first time they hear a tape of themselves. The tape is not unkind. The inside version was flattering.

The rule that falls out of it is short. Trust the recording over the feeling, every time. If it felt wonderful and sounds thin, the recording is right. If it felt like nothing and sounds good, the recording is still right. Your sense of your own voice is a rough guess made through the wrong channel. Playback is the report.

Which makes a pair of headphones training gear rather than a nicety.

Maximum phonation time

Take a full breath. Sing a steady /a/ at an easy pitch and hold it, evenly, until the sound stops. The seconds on the clock are your maximum phonation time, or MPT.

It is a mixed measure. That is both its weakness and its point. The number folds three things together. How much air you took in. How well you dole out the breath on the way out. And how well the folds turn that air into sound. Get better at any of the three and it climbs.

The target band is softer than it looks in print. Clinics often quote about 25 to 35 seconds for adult men, and 15 to 25 for women. But measured groups land lower. Maslan and his team timed 69 healthy older adults. They found means of 23.2 seconds for men and 21.0 for women. Age and sex made no real difference.1 Read the printed bands as a rough guide, not as a grade.

What is not soft is how repeatable it is. Speyer’s group found MPT reliable to an interclass correlation of 0.998 across raters. And it split dysphonic patients from healthy ones by about 6.6 seconds.2 That mix makes MPT a poor exam and a fine gauge. Your MPT against everyone else tells you very little. Your MPT against your own MPT from six weeks ago tells you almost all you wanted to know.

The s/z ratio

Hold /s/, the hiss, with no voice behind it, for as long as you can. Write down the seconds. Take a fresh full breath and hold /z/, the buzz, for as long as you can. Same tongue, same teeth, same lips. Write that down too. Then divide the first by the second.

The design here is lovely. It pays to see why before you see the numbers. The two sounds differ in just one thing. Everything above the larynx is the same. So the air is pushed through the same narrow gap both times. The one thing that differs is whether the folds are switched on. So /s/ measures your air supply and your breath control. And it does so with the larynx taken out of the picture. Then /z/ measures the same thing with the larynx put back in. Dividing one by the other cancels your lungs and leaves your larynx. It is a controlled experiment you can run standing in a hallway, with no gear at all.

Eckel and Boone tested it in 1981 across 150 people. Normal speakers sat near 1.0. So did dysphonic speakers whose vocal folds turned out to be sound. Speakers with nodules or polyps went past 1.4 ninety-five percent of the time.3 The telling part of that paper is the negative result. The /s/ hold did not differ between the groups at all. Only /z/ shortened. That is the control doing its job in public.

The reason is plain once you have the ratio. A lump on the edge of a fold stops it sealing cleanly. Glottal resistance falls. Air rushes past faster than it can be turned into sound. So the buzz dies while the hiss keeps going. A ratio well above 1.0 is literally the sound of leaking.

The s/z ratio
The s/z ratio
Two holds that differ in exactly one thing. Divide one by the other and your lungs cancel out, leaving your larynx.
Two holds that differ in exactly one thing. Divide one by the other and your lungs cancel out, leaving your larynx.
Hold “sssss” as long as you can
Hold “sssss” as long as you can
Folds open. Air is spent on the hiss alone.
Folds open. Air is spent on the hiss alone.
about 20 seconds
about 20 seconds
Hold “zzzzz” as long as you can
Hold “zzzzz” as long as you can
Same breath, same mouth, folds now closed and buzzing.
Same breath, same mouth, folds now closed and buzzing.
about 20 seconds
about 20 seconds
s ÷ z
s ÷ z
Lung size, posture, effort and practice all sit in both numbers, so dividing throws them away.
Lung size, posture, effort and practice all sit in both numbers,so dividing throws them away.
0.8
0.8
1.0
1.0
1.2
1.2
1.4
1.4
1.6
1.6
sealing cleanly
sealing cleanly
worth a question
worth a question
healthy voices sit near 1.0
healthy voices sit near 1.0
Why a high ratio means leaking
Why a high ratio means leaking
Anything that stops the folds sealing — swelling, a lump on an edge — lets air rush past faster than it can be turned into sound. So the buzz runs out while the hiss keeps going, and the ratio climbs. It is a screen and not a diagnosis: a ratio well above 1.4 with weeks of hoarseness is a question for a laryngologist, not for an app.
Anything that stops the folds sealing — swelling, a lump on an edge — lets air rush past faster than it can be turned into sound. So the buzz runs out while the hisskeeps going, and the ratio climbs. It is a screen and not a diagnosis: a ratio well above 1.4 with weeks of hoarseness is a question for a laryngologist, not for anapp.
Text is not SVG - cannot display
Two holds that differ in exactly one thing. Divide them and your lungs cancel out, leaving your larynx. Ariel Diaz · CC BY-SA 4.0

One caution. This is a screen, not a diagnosis. Say you sit well above 1.4 and you have been hoarse for weeks. That is a question for a laryngologist, not for an app.

Phrases per breath

Pick one set passage. Count how far you get before you have to take air. Same passage every time, same tempo.

This is the mixed, real-world measure. It is the one that changes how a Tuesday feels. It is also the noisiest of the four. It moves for reasons that have nothing to do with your breath. How you chose to phrase a line. How fast you took it. If you were on edge. Track it. But treat one reading as gossip, not as data.

Pitch accuracy, in cents

A cent is one twelve-hundredth of an octave, which is one hundredth of a semitone on a piano. Alexander Ellis brought in the unit in 1885. He split the octave into “1200 equal hundredths of an equal semitone, or cents as they may be briefly called”.4 It is logarithmic. So the same number of cents means the same felt distance. That holds whether you are up high or down low. It is just what a raw reading in hertz cannot give you.

The scale is easy to hold. People can hear a gap of about 5 to 6 cents in good conditions. And they spot 25 cents very reliably.4 So a tuner reading 18 cents flat is a real miss you can hear. And it is still a small one. The unit earns its keep. It turns “you were a bit under” into a number. You can average that across a take and plot it across a year.

The reference verse

One verse. Same microphone, same distance, same gain setting, recorded once a month into a folder named for the date.

Holding the setup fixed is not fussiness. It is the whole design. Move the mic six inches closer. Your tone gets warmer. Your levels rise. Those changes land in the recording as though your voice had done them. Freeze all you can control. Then your voice is the only thing left free to move. Then the takes match up. And takes that match up are a time series, not a set of impressions.

From each take, pull the four things. Pitch accuracy in cents. How steady the sustained notes are. Phrases per breath. And a plain listen-back. The listen-back matters as much as the numbers. Twelve months apart, the change is plain to anyone. One month apart, you cannot see it at all. That is just why the folder exists.

A spectrogram, so you can see the stack

The four numbers tell you what happened. A spectrogram shows you why. It is a picture of a recording. Time runs left to right. Frequency runs bottom to top. Loudness is brightness. A sung note is never one line. It is a stack. The bottom line is the fundamental, the pitch you would name. Above it sit the harmonics, exact multiples. Tone is which of those upper lines are bright.

Sonic Visualiser is a free desktop app that draws this.5 Its defaults are built for general audio, not for voice. They smear the harmonics together and squash the low end. Three purpose-built panes, saved once as a session template, give you the same picture every month. That is the only way comparison means anything.

Install it. Open a take. Then:

Pane 1 stays the waveform. You want both.

Pane 2 is harmonics and vibrato. Fastest path: Layer → Add Melodic Range Spectrogram, which presets a log scale over the singing range. Then in the layer properties column on the right:

  • Window: 4096. Overlap: 75%.
  • Frequency scale: Log. Min Frequency: 50 Hz. Max Frequency: 2000 Hz.
  • Colour scale: dB. Colour: Green or Sunset.
  • Gain: up until the harmonic ladder is clear and the background is not solid.

A strong note shows five or more clean stacked bands with dark gaps between them. Vibrato is a gentle waviness in every band at once.

Pane 3 is breath and the two tone breaks. Add a second spectrogram.

  • Window: 1024. Overlap: 50%.
  • Frequency scale: Log. Min Frequency: 100 Hz. Max Frequency: 10000 Hz.
  • Colour scale: dB. Gain: moderate.

When the note goes hissy, the gaps between the bands fill with haze in the 2–10 kHz region while the bands themselves fade. That haze is leaked air. When the note goes empty, the bands fade and the haze does not arrive. The filter moved.

Pane 4 is a spectrum plot. Layer → Add Spectrum. Log frequency, dB vertical, peak-hold on. Read it paused on one sustained vowel, never in a gap. A gap measures the room.

The vertical frequency scale belongs to the pane’s current layer. Click inside a pane to make it current. Give the pane enough height or the scale is suppressed. Min Frequency and Max Frequency in the property box are what zoom the view. The mouse wheel only zooms time. If you cannot see the zoom wheels in the bottom-right corner of a pane: View → Show Zoom Wheels, or press Z. The vertical wheel sets frequency zoom. The panner next to it scrolls the range. Double-click a wheel to type a value. The small empty box at the very bottom-right corner resets both zooms.

Same microphone, same distance, same room, same gain, every time. One file, in this order:

  1. Sustained /a/ at a comfortable pitch, about 10 seconds.
  2. The same vowel on a five-note scale up and down.
  3. About 20 seconds of one set phrase.

Save it as a dated wav, reopen it through the saved session template so the settings never drift, and screenshot pane 2 at a fixed zoom. Monthly is enough. Change shows over weeks, not days.

This is the same folder as the reference verse. The spectrogram is how you read it. A later tool can run the setup automatically. Until then, the app and the saved template are the instrument.

Guess, then check

There is one drill that trains the sensor rather than the voice, and it costs nothing.

Sing a phrase into the headphones. Before you press play, say out loud what you think it sounded like. Be specific. Flat on the top note. Ran out of air in the last bar. Then listen.

What you are training is not the singing. It is the accuracy of your own report. Every good singer has an internal read that matches the tape, and nobody is born with it, because the channel it would come through does not exist. You build it by guessing and being corrected, a few hundred times. It is also the cheapest way to tell whether a session went well when there is no gear in the room.

The room has already decided

There is one more reason to keep the folder. It is the honest one.

The people who have heard you sing for years made up their minds about your singing years ago. Views about someone you know well change very slowly. This is not them being unkind. It is how attention works. Once a thing is filed, it stops being looked at again. Your family stopped looking at your voice again around the time they stopped noticing your face. You can get much better. And you will still be met, warmly, with the same verdict they reached ten years ago.

You cannot argue with that. You should not try. What you can do is keep a dated folder of takes, recorded the same way every month. Against a room full of people who already decided, it is the only ground truth on hand. That is not a point about your family. It is a point about measuring. When the people watching do not update, the tool is the only thing left that will.

References

  1. Maslan, Leng, Rees, Blalock & Butler, Maximum Phonation Time in Healthy Older Adults, Journal of Voice 25(6), 709–713 (2011). Journal of Voice

    “Females and males had mean MPTs of 20.96 (SE = 0.92) and 23.23 (SE = 0.96) seconds, respectively. MPTs did not vary significantly with age or gender.”

  2. Speyer et al., Maximum phonation time: variability and reliability, Journal of Voice 24(3), 281–284 (2010). DOI 10.1016/j.jvoice.2008.10.004

    “Patients showed significantly shorter maximum phonation times compared with healthy controls (on average, 6.6 seconds shorter).”

  3. Eckel & Boone, The S/Z ratio as an indicator of laryngeal pathology, Journal of Speech and Hearing Disorders 46(2), 147–149 (1981). DOI 10.1044/jshd.4602.147

    “While no statistical difference was found between the three groups in their ability to sustain /s/, the subjects with laryngeal pathology had significantly lower duration times for /z/ than subjects in the other two groups… The dysphonic subjects with laryngeal pathology produced s/z ratios in excess of 1.4 ninety-five percent of the time.”

  4. Cent (music), Wikipedia — the unit, its introduction by Alexander John Ellis in 1885, and the perceptual thresholds. 2

  5. Sonic Visualiser, Centre for Digital Music, Queen Mary University of London. The layer-property labels and zoom-wheel controls above match the 4.3 reference manual.

    “99 other notes were interposed, making exactly equal intervals with each other, we should divide the octave into 1200 equal hundredths of an equal semitone, or cents as they may be briefly called.” The article also notes that humans can distinguish a difference in pitch of about 5–6 cents, and recognise 25 cents very reliably.