0% READ
● A Visual Essay

The Window at
Three Kilohertz

Why Stevie Wonder sits on top of a dense arrangement without ever sounding pushed.

2026.09.05·14 min·Sound
↓ Scroll

Put on Superstition and try to lose the vocal. You cannot.

There is a clavinet riff doing most of the harmonic work, a horn section stabbing through the gaps, drums Stevie played himself, and a Moog holding the bottom. It is a busy record. And the voice is never once in trouble. It is not louder than the band in any way you would notice, and it never sounds like it is being pushed to stay on top. It simply sits there, above everything, entirely at ease.

Engineers usually explain this with a compliment: he has a great voice, it just cuts. That is true and it explains nothing. Plenty of great voices get buried by three guitars. What I want to know is the physical account. What is actually happening in the air, and in the ear, that makes this particular voice survive that particular arrangement.

There are four mechanisms, and they compound. One is in his throat, one is in your head, one is in the arrangement, and one is in the way the brain decides what counts as a thing.

A buzz and a tube

A voice is two separate machines, and it helps enormously to keep them apart.**

The first is the source. Air from the lungs forces the vocal folds to open and slam shut, over and over, a hundred-odd times a second for a man. That is all the larynx does. It makes a buzz. On its own it sounds like a duck call: harsh, dense, and completely unlike speech. Crucially, that buzz already contains a huge stack of harmonics, evenly spaced, running from your pitch up past ten kilohertz.

The second is the filter. Above the folds sits a bent tube about seventeen centimetres long: throat, mouth, and optionally the nose. Like any tube it has resonances, and by moving your tongue and jaw you change its shape and therefore where those resonances sit. Each resonance is a formant. The formants do not add anything. They select. They amplify whichever harmonics of the buzz happen to fall near them and let the rest through quietly.

The first two formants are the vowel. That is genuinely all a vowel is: two bumps in a filter. Move the first to 730 Hz and the second to 1090 and you get ah. Move them to 270 and 2290 and the same buzz becomes ee. Nothing about the source changed.

Everything above the second formant is not the vowel. It is the voice. And it is where this essay happens.

The ring

In the 1970s Johan Sundberg went looking for why an opera singer with no microphone can be heard over eighty players who are, collectively, far louder.* The answer was not volume. Measured at the singer, the orchestra usually wins.

What he found instead was a shape. Trained singers cluster their third, fourth and fifth formants tightly together, around 2.8 kHz, instead of leaving them spread out. Several modest resonances stacked in one place become one tall narrow peak, and that peak is what the ear hears as ring. Sundberg called it the singer's formant.

The elegant part is that it costs nothing. It is a change in the shape of the tube, chiefly a lowered larynx and a widened throat, not a change in effort at the folds. The singer is not pushing harder. They are re-aiming the same energy into a narrower band.

Play the vowel below and then switch the ring on. The pitch does not change, the vowel does not change, and the loudness barely changes. What changes is whether the thing sounds like it is in the room with you.

Fig. 01 · Source and filter
A sawtooth buzz through a bank of resonances. The first two formants pick the vowel; the cluster near 3 kHz is the ring. Same source either way.

Stevie is not an opera singer and this is not a trained classical placement. But the same physics is available to anyone with a bright, forward, slightly nasal production, and that is a fair description of what he does. The nasal coupling in particular tends to put a strong component right in this region.

Your ear is not flat

Here is the part that turns a modest advantage into a decisive one. The region a singer's formant lands in happens to be the region where human hearing is most sensitive, by a wide margin.

Some of this is plumbing. Your ear canal is a tube about two and a half centimetres long, closed at one end by the eardrum, and it resonates like any such tube, adding a broad boost somewhere between two and four kilohertz before the sound reaches anything neural. The rest is the machinery behind it. Put together, the difference in sensitivity between 100 Hz and 3 kHz is roughly twenty decibels.

Twenty decibels is not a subtlety. It means a sound at 3 kHz is heard as dramatically louder than a sound at 100 Hz carrying the same physical energy. Put energy at three kilohertz and you are given loudness you never paid for.

Drag the slider. The amplitude sent to your speakers never changes.

Fig. 02 · Free loudness
Sensitivity by frequency. The tone is generated at a constant level throughout, so everything you hear change is your own hearing.
"The ring is not loud. It is simply aimed at the part of you that is listening hardest."

What actually buries a voice

Masking is not a matter of overall level. A sound hides another sound mainly when the two are close in frequency, because the inner ear analyses sound in overlapping bands rather than as one continuum. Two things in the same band fight. Two things in different bands mostly do not.§

This has an asymmetry that surprises people the first time they meet it. Masking spreads upward far more readily than downward. A loud low tone will smear over the region above it; a high tone does very little to what sits below.

Which means the intuition that a big bass sound will swamp a vocal is mostly wrong. Bass and voice are far enough apart that the bass can be enormous and the voice remains legible. What kills a vocal is something with real energy in the same couple of octaves as the voice: a strummed acoustic guitar, a bright pad, a doubled synth line, a second vocal in the same register.

So the practical question is never how loud the band is. It is what the band is doing between roughly two and four kilohertz.

The band nobody plays in

Now look at what Stevie actually plays. On the run of records from Music of My Mind through Innervisions he played most of the instruments himself, which means the arrangement and the vocal were decided by the same person, often on the same day.

The Moog bass is filtered low and stays there. The drums put their weight in the kick and the snare, with the cymbals living far above the voice rather than across it. The Rhodes has a bell-like attack but the body of it sits under a kilohertz. None of these are anywhere near the window.

The clavinet is the interesting one, because it genuinely is bright enough to compete. And listen to what it does: it plays in the gaps. The riff answers the vocal line rather than running underneath it. The horns do the same. The arrangement is full of bright, aggressive sounds that almost never occupy the same instant as the lead vocal.

Toggle the parts below and watch the window fill up.

Fig. 03 · Where everything sits
Indicative spectral shapes for the palette, not measurements. The outlined curve is the voice. Add the bright parts and the window it depends on starts closing.

This is the part that is craft rather than anatomy. A singer with a beautiful ring will still lose to an arrangement that parks a bright rhythm part directly underneath them. Stevie is not just a voice that cuts. He is a voice that cuts arranged by someone who knew not to fill the space it needed.

Why a moving voice separates

There is one more mechanism, and it is the one I find genuinely beautiful, because it is not about frequency content at all.

Your auditory system is not handed separate sounds. It is handed one pressure wave containing everything at once, and it has to decide which parts belong together. Albert Bregman's term for the problem is auditory scene analysis. One of the strongest cues the brain uses is common fate: components that change together are probably one object.

Vibrato is common fate made audible. When a singer's pitch moves, every harmonic moves with it, by the same proportion, at the same instant. Nothing else in a mix does that. A sustained organ chord sits perfectly still. A pad does not breathe. So the voice becomes the one thing in the arrangement whose parts are visibly, coherently moving together, and the brain lifts it out as a single object almost effortlessly.

Stevie's voice is never still. The pitch is constantly inflected, scooped, shaken, bent into and out of notes. That reads as expression, and it is. It is also, in the plainest engineering terms, a continuous stream of grouping cues.

The demonstration below makes the point better than a paragraph can. The pad never changes. Only the behaviour of the voice does.

Fig. 04 · Common fate
A harmonic stack over a static pad. Steady, it blends. Moving together, it detaches. Moving separately, it stops being a voice at all.

The third setting is the one worth sitting with. The amount of movement is identical to the second. All that changed is whether the harmonics agree with each other. Coherence, not motion, is what makes an object.

Four things at once

None of these mechanisms is remarkable alone. Stacked, they explain the whole effect:

  • A production that concentrates energy near three kilohertz, at no cost in effort.
  • An ear that happens to be about twenty decibels more sensitive there than down at the bass.
  • An arrangement that keeps that same region unusually empty, because the arranger was also the singer.
  • A delivery that never holds still, handing the listener a constant supply of cues for treating the voice as one separate thing.

What you experience as effortlessness is four independent advantages pointing the same direction. And the reason it sounds natural rather than engineered is that none of it is a boost. Nothing was made louder. The energy was aimed, and then not competed with.

If you are mixing or singing

The practical version, in rough order of how much difference it makes:

  • Subtract, do not add. If the vocal is buried, the reflex is to boost it around 3 kHz. It usually works better to find whatever else is sitting there and take it out of the way. Boosting the voice raises the whole mix's harshness; carving the competition costs nothing.
  • Suspect the mid-bright things, not the bass. Masking spreads upward, so the sub is rarely your problem. The acoustic guitar, the bright pad and the doubled synth are.
  • Arrange in time, not just in frequency. The cheapest way to give a vocal room is for the bright parts to play between the lines. This is the single biggest lesson from those records.
  • Let the voice move. Heavy pitch correction and hard compression both flatten exactly the cues the brain uses to separate a voice from its background. A perfectly steady, perfectly level vocal is easier to lose than an imperfect one.
  • If you sing: find the ring before you find the volume. More air is the expensive way to be heard. Changing the shape of the tube is the cheap one, and it is what actually carries.

Not louder

I like this as an explanation because it takes nothing away. Knowing the mechanism does not make Superstition less astonishing. If anything it makes the arrangement more so, because you start hearing all the places something bright could have been and deliberately is not.

The thing that stays with me is that none of the four mechanisms is about force. A voice does not win a mix by being louder than the mix. It wins by being aimed where the ear is listening, in a band nobody else is using, moving in a way that marks it out as one thing. That is a description of a voice, and it is also a fairly good description of how to be heard in general.

(If the source-and-filter idea was new here, there is a lot more on how waveform and envelope shape a sound in The Shape of a Sound.)

END · 14 MINReply by email
◆ Newsletter

New essays. New tracks. One email a month, max.

Reply to anything I send and it goes straight to me.

NEXT →

What Music Knows About Emotion

Why some sounds make us feel happy and others make us feel sad.

Notes
  • *Johan Sundberg, The Science of the Singing Voice (1987), and the papers behind it. The singer’s formant is the clustering of F3–F5 near 2.8 kHz, and the orchestral spectrum it exploits falls away above roughly 500 Hz.
  • The formant values used in Fig. 01 are the classic measured averages for adult male vowels from Peterson & Barney (1952). They are a reference set, not a description of any particular person.
  • Fig. 02 plots the standard A-weighting curve, the usual engineering model of frequency-dependent hearing sensitivity. It peaks a little under 3 kHz. Part of that is ear-canal resonance, part is the rest of the auditory system.
  • §Critical bands, and the upward spread of masking, are core psychoacoustics: a masker is most effective on frequencies near it and above it, much less so below. This is also the principle every perceptual audio codec is built on.
  • Albert Bregman, Auditory Scene Analysis (1990). Common fate, harmonicity and onset asynchrony are the main cues the brain uses to decide which parts of a pressure wave belong to the same source.
  • An honest caveat: I have no spectrographic analysis of Stevie Wonder’s voice, and I am not aware of a published one. The mechanisms here are well established; the claim that his particular production exploits them is my reading as a listener, not a measurement.
  • **Source-filter theory is Gunnar Fant’s, from Acoustic Theory of Speech Production (1960). It is the idea that makes the whole of speech acoustics tractable: the larynx sets pitch, the tract sets timbre, and the two are largely independent.