An intuition kit · resonance, poles & the voice

Why is a formant
a pole?

A guided walk from the ringing of a tube to the dot on a circle that a computer hunts for in your voice — built to be re-told at a dinner table, with the real reasons tucked just underneath.

When you say “ah”, the air in your throat and mouth rings — the way a bottle hums when you blow across its lip. The frequencies it likes to ring at are the formants, and they shift as you move your tongue and jaw. That shifting is, quite literally, the difference between one vowel and the next.

The strange-sounding claim we'll earn is this: each of those ringing frequencies shows up, in the math, as a place where a fraction blows up — where the bottom of a fraction heads toward zero. We call those bottom-spots poles. By the end you'll see why that isn't a coincidence or a trick, but the same fact said in two languages. And you'll see how a machine, handed nothing but the raw sound, walks its way to those poles.

Every panel below is an instrument you can touch. Drag things. The numbers are real.

01The tube that rings

Resonance is cheap input, expensive output

Push a child on a swing. One badly-timed shove does nothing. But push gently in time with the swing's own rhythm, again and again, and the arc grows huge. You barely spent any effort; the swing is now carrying a lot of energy. That is the whole idea of resonance: at one special frequency, tiny repeated nudges accumulate into a large oscillation. More resonance means more amplitude means more stored energy — for almost no input.

Drag the drive frequency and watch the swing find — or miss — its favourite rhythm.

Resonance · driven oscillator gain ×1.0
The peak sits at the tube's natural rhythm. Off it, the pushes fight each other and nothing builds.

A vocal tract is just a tube of air — from the vocal folds up to the lips. Like any tube, it has a set of favourite ringing frequencies, set by its length and shape. Those are your formants. Change the shape (move your tongue), and the favourite frequencies slide. Watch three of them move as the tube lengthens and shortens:

Vocal tract · favourite frequencies F1 500 · F2 1500 · F3 2500 Hz
Same tube, three resonances. They all slide together because they all come from one length.
Dinner-table version: a formant is a frequency the mouth-tube loves to ring at. Tiny effort in, big ring out. Move your tongue, move the ring.
02Why a denominator

The pole isn't a metaphor — it's the equation of motion, upside down

Here's the honest reason the bottom of a fraction enters the story. A ringing tube is a damped oscillator — the same object as a mass on a spring, or a swing with friction. It has inertia (the air's mass), a springiness (the air pushing back when compressed), and a little loss (friction, sound leaking out). You already know its behaviour: nudge it and it rings while slowly fading.

Ask: how big is its response if you drive it steadily at frequency f? Near its favourite frequency, the answer is enormous — that's resonance from §1. “Enormous response” means a number that is shooting toward infinity. And the only way a clean ratio shoots to infinity is if its bottom is going to zero. So the resonance must live in a denominator. It has nowhere else to live.

That denominator, where it touches zero, defines a point in the plane of frequencies. That point is the pole. Drag it. As you push it toward the “no loss” line, the ring lasts longer and the spectral peak grows taller and sharper — the bottom of the fraction is getting closer to zero.

The pole & what it controls f 700 Hz · ring 12 ms
Drag the dot. Left–right = how fast the ring dies. Up–down = the ringing frequency.
pole the ring it makes the spectral peak

So a single pole carries two readable facts: how high it sits tells you the ringing frequency (the formant), and how close it sits to the no-loss line tells you how sharp and long-lived that ring is. That's the entire soul of the topic in one dot.

The actual derivation, if you want it (no transforms required)

Write the damped oscillator the way a physicist would — position x driven by a force:

ẍ  +  2σ·ẋ  +  ω₀²·x  =  drive(t)

where ω₀ is the natural frequency and σ is the loss. Now the one and only trick: feed it a steady oscillation and notice that differentiating just multiplies by a number we'll call s. (That swap — “take a derivative” becomes “multiply by s” — is all we borrow from the Laplace transform. Treat it as a black box that turns calculus into algebra. You never have to compute one.)

Every derivative becomes a factor of s, so the equation collapses into plain algebra:

(s² + 2σ·s + ω₀²)·X  =  drive

and the response — output per unit drive — is simply one over that polynomial:

H(s)  =  1s² + 2σ·s + ω₀²

Look at what the denominator is: it's the equation of motion itself. The physics — the inertia, the springiness, the loss — didn't go anywhere; the transform just dropped it straight into the bottom of the fraction. So “where does the response blow up?” becomes “where does this polynomial hit zero?”, and its roots are

s  =  −σ  ±  i·ωd

These roots are the poles. The imaginary part ωd is the pitch of the ring; the real part −σ is how fast it fades. A sharp formant is a slowly-fading ring is a pole sitting close to the no-loss line. That's the panel above, made of symbols.

Dinner-table version: the response is a fraction. Resonance is the response going huge. A fraction only goes huge when its bottom goes to zero — and that zero-of-the-bottom is the pole.
03Poles on the circle

A tract is a handful of poles, and the circle is their home

Recorded sound is sampled — a stream of numbers, tens of thousands per second. When you move from continuous air to sampled numbers, the flat plane from §2 curls up into a circle (the famous “unit circle”). Nothing about the idea changes, only the costume:

Angle around the circle = frequency.   Distance out to the rim = sharpness (a pole hugging the rim is a long, ringing, narrow resonance).

A real vocal tract isn't one resonator but a chain of them, so its response is several of these poles at once — and because they multiply together, the whole thing is one big fraction with only a denominator. That “all-pole” shape isn't a simplification we impose; it's the honest form of a tube with resonances and no side-branches. Drag the poles; build a vowel.

Signature instrument · drag the poles F1 700 · F2 1220 · F3 2600 Hz
Drag any coloured dot. Toward the rim → taller, sharper peak. Around the circle → the peak slides in frequency.
poles (on the circle) resulting spectrum
Dinner-table version: a vowel is a few dots pinned near the edge of a circle. Where they sit is which vowel you're saying.
04Source meets filter

Two independent things: a buzz, and the tube that shapes it

Voiced speech has two parts that don't know about each other. The source is the buzz of the vocal folds — a dense comb of evenly-spaced harmonics, and its spacing is your pitch. The filter is the tube from §3 — the formants. The sound that comes out is the source seen through the filter: the comb, with its teeth scaled up near the formants and down elsewhere.

The punchline for detection: the formants are in the filter, not the source. Slide the pitch and the comb's teeth move, but the formant humps stay put. Change the vowel and the humps move, but the comb spacing doesn't. They're separable — which is exactly what lets us pull the formants out later.

Source × filter = speech pitch 120 Hz
Move pitch: the teeth slide, the humps don't. Move vowel: the humps slide, the teeth don't.
05The vowel space

Two formants are enough to name a vowel

Phoneticians draw vowels on a trapezoid — close/open up the side, front/back across. It turns out that map is just the first two formants in disguise. F1 tracks how open your mouth is (jaw down → F1 up); F2 tracks how far forward your tongue is (tongue front → F2 up). Pin those two numbers and you've essentially pinned the vowel.

Drag the tongue around its space. The label is the vowel you'd be making; press play to hear it. (This is the same pair of poles from §3, now wearing their phonetic names.)

Vowel space · F1 × F2 “ah” · F1 730 · F2 1090
Left ↔ right is F2 (tongue front/back). Up ↕ down is F1 (mouth open/closed). The grey dots are real vowels.
Dinner-table version: the vowel chart you saw in linguistics class is the formant map. Two numbers, one vowel.
06Prediction finds the tract

How a machine recovers the poles from raw sound

Now the algorithm — the standard one, called linear prediction (LPC). It rests on a single idea you can feel. Try to guess each new sample of the sound from the few samples just before it. The part that's predictable is the tube's ring — a resonance is exactly “memory”, it carries the wave forward smoothly. The part that's unpredictable is the source — the sudden kicks of the vocal folds, which arrive without warning.

So the best-predicting filter is the vocal tract, and the leftover error is the source. Two birds. Watch the predictor glide along the ring and stumble only at the glottal kicks — and the error (the residual) collapse to near-silence in between.

Predict the next sample predictable = tract · surprise = source
the sound predicted from the past residual (the surprise)

Why does this find the answer cleanly, with no guessing or getting stuck? Because the all-pole choice from §3 makes the “how wrong am I” landscape a single smooth bowl — one bottom, no false valleys to fall into. The method (a tidy recursion named Levinson–Durbin) just rolls straight downhill to the one true minimum, adding one pole at a time. That is the convergence: not a hopeful search, but a ball in a bowl.

The error landscape · one bowl, one bottom least-squares = guaranteed minimum
Every starting point rolls to the same bottom. Zeros (anti-resonances) would dent this bowl with traps — all-pole keeps it clean.
Dinner-table version: guess the next bit of sound from the last few. What you can guess is the tube; what surprises you is the voice. Best guesser = the tube.
07Putting it together

Detection, start to finish

Here is the whole pipeline in one picture. The jagged grey line is the raw measured spectrum of a vowel — bumpy because of the source's comb of harmonics. Linear prediction fits the smooth all-pole envelope (amber) that hugs the tops of those bumps, ignoring the comb. Then we read the envelope's peaks — or equivalently, find the poles on the circle — and those are the formants. Angle gives frequency; the labels drop out.

Drag the tongue once more: the measured spectrum, the fitted envelope, and the poles all move together. You're now watching exactly what a formant tracker does, thirty times a second, to follow your speech.

Formant detection · the full loop F1 730 · F2 1090 · F3 2600 Hz
measured spectrum all-pole envelope poles = formants

The four sentences to carry home

  1. A formant is a frequency the mouth-tube loves to ring at — small effort in, big oscillation out. Vowels are made of them.
  2. A response is a fraction; resonance is that fraction blowing up; a fraction can only blow up where its bottom hits zero — and that bottom-zero is the pole. (The denominator is the equation of motion in disguise.)
  3. On the circle, a pole's angle is its frequency and its nearness to the rim is its sharpness. A vowel is a few poles near the rim.
  4. To find them, a machine predicts each sound-sample from the last few: what it can predict is the tube, what surprises it is the voice — and the best predictor's poles are the formants.