Start a project →
How-to · Audio

How loud should background music be under a YouTube voiceover?

Most advice on music levels is a rule of thumb nobody measured. So I measured three narrated documentary channels with a loudness meter: one client and two reference channels. They land on three different answers, and all three are deliberate. Here are the numbers, why they work, and how to set the same thing in your editor.

By Kevin Tabares · Sep 26, 2026 · Updated · 8 min read

How loud should background music be under a voiceover?

For a narrated YouTube video, keep the music bed about 24 to 31 dB quieter than the voice, measured as loudness (LUFS), not peaks. Two measured documentary channels that keep music running sit at roughly 24 and 31 dB under. Short sound hits can rise to about 10 dB under.

That is much quieter than most people set music by ear. It sounds almost too low while you edit, and exactly right to a viewer who is following a story for twenty minutes. The rest of this guide shows where those numbers come from and how to hit them.

What do narrated documentary channels actually do with music?

These are public videos run through an EBU R128 loudness meter. The Interpreter is a client of mine (faceless aviation, maritime and rail disaster documentaries narrated with stickman characters); Philip Thompson and Librarian are reference channels I analyzed, not clients.

Channel Where music plays Music level vs. voice Whole mix Sound effects
The Interpreter (client) The whole video, as a constant bed. No absolute silence anywhere (longest gap in the voice about 1.5 s) About 24 dB under the narration, no ducking About −22.5 LUFS integrated, true peak −5.5 dBFS, loudness range 2.5–2.8 LU Short hits (impacts, engines) around 10 dB under the voice
Librarian, “The Real James Bond” The whole video; the music changes character as the story moves About 31 dB under (−46.8 LUFS in the pauses vs. −15.8 LUFS for the whole program) About −15.8 LUFS integrated; voice about 157 words a minute None
Philip Thompson, “The Unbreakable Spy of WW2” Only the cold open (about 0:10–1:30) and the ending No bed in the body: dry voice, digital silence in the pauses About −13.7 LUFS integrated, true peak −1.2 dBFS, loudness range 3.0 LU None

Three things stand out. First, none of them puts music close to the voice. Second, the loudness range is tiny on all three (2.5 to 3.0 LU): the voice is kept very even, and that evenness is what lets a quiet bed stay audible without fighting it. Third, "no music at all" is a real option. Philip Thompson's body is just a voice, and the music at the open and close feels bigger because of it.

How do you measure "dB under the voice"?

This is where most arguments about music levels go wrong. "Music 12 dB below the dialogue" and "music 24 dB under the narration" can describe mixes that sound surprisingly close, because one number compares peaks on a waveform and the other compares loudness. Peaks tell you how tall the spikes are; loudness (LUFS) tells you how loud something sounds over time.

The numbers in the table compare loudness. The simplest version you can do yourself:

  1. Solo the voice track, play a typical 30 seconds and read the short-term loudness on a LUFS meter.
  2. Solo the music track over the same 30 seconds and read it the same way.
  3. Subtract. If the voice reads −16 LUFS and the music reads −40 LUFS, the bed is 24 dB under.

On a finished video you cannot solo tracks, so the measurement above works the other way round: the meter reads the loudness in the gaps between sentences (music only) and compares it with the loudness of the whole program. That is how the Librarian figure was taken.

Should you use a fixed music bed or ducking?

Ducking means the music drops automatically whenever someone speaks and comes back up in the pauses. It is useful when the music is a feature: an intro with no voice, a montage, a talking-head video with long stretches of B-roll and no speech.

Under continuous narration it tends to backfire. A documentary voiceover pauses for a breath every few seconds, so a ducked bed rises and falls all the time. Viewers hear that as pumping, and it draws attention to the one layer that should stay invisible. The Interpreter's bed does not move with the voice at all. At 24 dB under, there is nothing to duck: the music is already below the narration everywhere, so it can simply stay put.

What does change the feel is the other layer. The short hits (an impact, an engine, a squeal) come in much closer to the voice, around 10 dB under, and then get out of the way. The bed stays a floor; emphasis comes from hits and from changing the track, not from pushing the bed up.

Rule of thumb from the measurements: bed low and constant, hits short and closer, never both loud at the same time. If a moment needs more weight, change the music cue or add a hit instead of raising the bed.

How do you set it in Premiere Pro?

Menu names shift a little between versions; the logic does not.

  1. Put the narration on its own audio track and the music on another. Keep sound hits on a third track so you can level them separately.
  2. Add the Loudness Meter effect to the mix (older versions call it Loudness Radar). Solo the voice, play a typical section and note the short-term reading.
  3. Solo the music. Select the music clips, open Audio Gain (shortcut G) and lower the gain until the music reads about 24 dB below the voice reading. Go toward 31 dB if the music is busy or has melody in the voice's range.
  4. In the Essential Sound panel, tag the music as Music and leave Ducking off for a narrated video.
  5. Set the hits on their track around 10 dB under the voice and listen to the transitions in and out of each one.
  6. When you export, use the Loudness Normalization option in the export Effects settings to bring the whole mix to your target without changing the balance you just built.

How do you set it in DaVinci Resolve?

  1. On the Fairlight page, give voice, music and hits their own tracks.
  2. Set the loudness standard for the meter in Project Settings → Fairlight, then open the loudness section of the meters panel. It shows short-term and integrated loudness.
  3. Solo the voice and note the short-term reading. If the narration was recorded in several sessions, right-click the voice clips and use Normalize Audio Levels in the ITU-R BS.1770 (loudness) mode so every clip starts from the same level.
  4. Solo the music and lower its clip volume in the Inspector (or the track fader) until it reads about 24 to 31 dB below the voice.
  5. Do not put a sidechain compressor keyed from the voice on the music bus. That is ducking, and the point here is a bed that stays still.
  6. Play the full timeline once with the integrated meter reset, and check the result before you render.

What loudness should the finished video be for YouTube?

This is a separate decision from the bed level. The balance between voice and music is set inside the mix; the master level only decides how loud the whole thing plays.

YouTube normalizes playback. Mastering engineer Ian Shepherd documents that YouTube turns down videos louder than its reference playback level and does not turn quieter ones up. The reference most often cited is −14 LUFS, for example on Wikipedia's audio normalization article, citing Pro Video Coalition. You can see what YouTube does to any video: right-click the player, open Stats for nerds and read the "content loudness" figure, which shows the difference between YouTube's estimate of that video and its reference.

Read the table with that in mind. Philip Thompson's −13.7 LUFS sits almost exactly at the commonly cited reference. Librarian's −15.8 is a little under it. The Interpreter's −22.5 LUFS is a deliberate, quieter house level with plenty of headroom (true peak −5.5 dBFS); YouTube plays it as uploaded. A quiet master is fine, as long as it is a choice and you know YouTube will not raise it for you.

What are the most common mistakes with music under voiceover?

If you run a narrated channel, the documentary YouTube editing page shows how these audio measurements sit next to pacing and visuals, and the faceless channel page covers voiceover-led formats more broadly. When you hand this over to an editor, put the target in writing (for example "bed 24 dB under the voice, no ducking, hits around 10 dB under"); this briefing guide shows where it goes. For the picture side of the same videos, see how to use archival footage without copyright strikes, and for mixing as a service, YouTube sound design.

Sources and how this was measured

Related guides

By niche
Documentary YouTube editor for history, disaster and true crime deep dives
Service
YouTube sound design and audio mix
How-to
How to brief your video editor (so they don't waste cycles guessing)