Tyler Pratt← All articles
Content

How Faceless Creators Stop The Bed Eating The Voice Before You Export — Without Guessing The Faders Again

Why does the music eat your voiceover - or disappear?

2026-08-21 · 7 min read
Quick answer

Faceless videos fail the mix in two ways. The bed is too loud and the line gets buried. Or the bed is so quiet the video feels empty. Creators then raise the voice in Audacity and the phone still loses it, because the bed is the mask. YouTube speech wants about 8 to 12 dB of air under the voice. A bed left at the same level as the VO has almost none. The fix is a map: duck on talk, extra duck on lists, mute for a second before a reveal, and a phone pass at arm's length. Do not pick a new song. Move the fader.

Key points

You paste a finished narration and a note about the bed you already picked. The tool scores hearability, flags every buried line, and writes the duck, lift, and mute points. You move faders. You do not hunt a new song.

A creator on r/NewTubers said it this week: "I have background music but it's hard to find the perfect balance between it being too invasive or being washed out."

Same week, a faceless Shorts builder named the other failure. His AI voice was too low one foot from a phone, headphones off, after Audacity, CapCut, and Auphonic.

Those two quotes are one job. The bed and the voice are fighting. This article is the map.

Why can't you hear the line on a phone?

Because the bed is sitting in the same loudness as the voice, and a phone speaker cannot split them.

Headphones lie. They have stereo, bass, and a sealed cup. A phone is a tiny mono driver one foot from a noisy room. When a piano chord and a quiet sentence hit at once, the chord wins.

That is why raising the clip in Audacity failed. The gap between voice and bed stayed at about 2 dB. Both got louder. The mask stayed. The phone still lost the line.

Faceless videos get this worse than talking-head videos. There is no mouth to watch. If the words vanish, the video is a slideshow with a song.

What is the balance, in numbers?

It is a gap, not a vibe. YouTube's home is about -14 LUFS for the whole video. Speech should live there. The bed should live under it.

  1. Voice. About -14 LUFS integrated. True peak under -1.5 dB. Fold to mono so a phone does not lose a side.
  2. Bed under talk. About -24 to -28 LUFS. That is 10 to 14 dB of air. The ear still knows a song is there. The words win.
  3. Bed in a pause. Lift to about -20 so the video does not feel empty. This is how you avoid 'washed out' without eating the next line.
  4. Lists and reveals. Extra duck, or a mute. Short words like 'Eggs. Gas. Rent.' need more air than a long sentence.

Under 6 dB of gap is buried. Over 18 dB for a long stretch is washed. The 'perfect balance' is not a middle slider. It is two sliders that move.

Why does one bed under the whole video fail?

Because your story changes and the fader does not.

A faceless money video opens curious, lists three prices, then lands a reveal. If the piano stays at one level, the list is noisy and the reveal is covered. That is 'too invasive' and 'the quiet line got lost' in the same track.

Sound-design tools tell you which song to search for. Mix is the next job. You already have the piano. The piano is fine. The piano never sitting down is the problem.

The cheapest retention win on the channel is a mute. One second of no bed before 'Here is the part the ads never said' does more than a new library.

What should you actually move in CapCut?

Keyframes on the bed. Not the VO gain.

  1. Open. Let the bed play under a second alone, then duck as the first line enters. If both hit at once, neither wins.
  2. Talk. Hold the duck. Do not ride every word. A flat -26 under speech is cleaner than a jumpy fader.
  3. List. Extra 2 dB down. Short words die first on a phone.
  4. Reveal. Mute 0.8 to 1.5 seconds, then come back under the payoff, still ducked.
  5. Close. Do not swell the song over the last line. That is how a good mix gets washed at the end.

Then do the phone pass. Play the reveal at arm's length with no headphones. If you lean in, the mute is too short. Lengthen the hole. Do not add gain.

Why is 'just normalize' not enough?

Normalize sets a peak. Mix sets a relationship.

A VO normalized to -1 dB still loses to a bed at -18 if the voice is thin and the piano is dense. Peak is not loudness. Loudness is not gap. Gap is what a phone hears.

Compression helps the phone pass because it brings quiet syllables up without lifting the bed. A slow 3:1 on the VO keeps t, k, p, and d popping. That was the other clue in the too-low thread: plosives that do not pop in a noisy room.

Do both. Compress the voice a little. Duck the bed a lot. If you only normalize, you will keep turning the TV up.

How do you run this on every video?

You paste one prompt into Claude Code and it builds the tool. It comes pre-filled with a savings-account script under an un-ducked piano, so the first run already shows hearability in the 30s, a phone fail, and a mute on the reveal.

Then you paste your own narration, the bed you actually used, and the spot that already feels wrong. It maps the beats and writes the keyframes.

Re-run it every time you drop a loop under a new VO. The song can stay. The map cannot be last week's map. Lists, reveals, and quiet lines move.

Grab it below - drop your email and the prompt is on the very next page.

Can you turn this into a side hustle?

Yes — think of it as a skill you just acquired in one paste. Skills can be sold, and this one sells by the deliverable.

The play is simple. Faceless documentary and money channels that lose watch time to a bed that is too loud or a voice that dies on a phone already want a mix map for a finished narration - every section scored, duck/lift points, YouTube LUFS targets so the bed never eats the line — they just dread producing it. You produce it in minutes, deliver a clean result, and bill for the outcome. Going rates run $40 to $120 per video, or $350 a month per channel.

The best part is the cost to start: a free prompt - one prompt that mixes every video you make. The tool does the heavy lifting in minutes, so your margin is high and you can take on more clients without more hours. To get your first client, reach out to a few faceless documentary and money channels that lose watch time to a bed that is too loud or a voice that dies on a phone you already know. Do one for free, show them the result, and ask who else needs it.

FAQ

Is this just a loudness calculator?

No. A calculator gives one number for the whole file. This reads the script and changes the bed by beat - talk, list, pause, reveal. The number is the gap. The job is the fader list.

Do I need a new music track?

Almost never. If the bed fits the mood, keep it. The failure is a bed that never ducks. Replace the song only if it has vocals or a drop that fights the line no matter the fader.

Will this make my video loud enough for YouTube?

It targets about -14 LUFS for the voice and a real gap under it, which is what YouTube and a phone both want. It will not master the file for you. It tells you what to do in CapCut or Resolve so the export is hearable.

What if I do not use music?

Then you do not need this tool. If you drop any bed, loop, or ambience under an AI voice, you need it every upload.

Written alongside the Voice Bed Mixer · More AI tools & articles