Ducking music under voice:a starting recipe
Ducking means lowering the music while someone speaks and bringing it back up when they stop. Done well, the viewer never notices. Done badly, the music pumps up and down, or the voice is hard to follow.
The starting recipe
These are starting points, not rules. Every voice and track is different.
- Set the music at a comfortable level on its own, then lower it by around 15 to 20 dB while the voice is speaking.
- Make the dip start just before the first word, not on it. Around a quarter of a second ahead works well.
- Let the music come back up slowly, over half a second to a second, after the last word.
- Do not bring it back up for short pauses of a second or two between sentences. It sounds like pumping.
Automatic ducking
Most editors now have an automatic ducking feature that lowers the music when it detects speech. It is a good first pass. Then listen through and fix the places where it reacts too quickly or too slowly.
Choose music that ducks well
Some music is easier to talk over than others. Tracks with a lot happening in the middle of the frequency range, where voices sit, have to be ducked much further. Ambient, low and airy tracks can sit higher under speech.
- Evidence Room Ambience on the Mystery & True Crime shelf
- Morning in the Kitchen on the Kitchen & Lifestyle shelf
Both sit well under a voice: one patient and tense, one warm and homely.
Check on a phone
Most of your audience will listen on a phone speaker or cheap earbuds, where music can mask speech more than on studio monitors. Play the finished video on your phone before publishing.
For podcast-specific levels, read Mixing podcast music to loudness targets.