43 min read

> "Films are 50 percent visual and 50 percent sound. Sometimes sound even overplays the visual."

Prerequisites

  • 15
  • 29

Learning Objectives

  • Sequence the audio-post workflow — organize, clean, level, score, design, and master — in the order that puts dialogue first.
  • Clean location dialogue by removing rumble, hum, clicks, and broadband noise without introducing artifacts.
  • Shape and level dialogue with EQ and compression so every word is clear, present, and consistent.
  • Select, license, and place a music bed that supports the story without burying the voice.
  • Design ambience and spot sound effects that build a believable, immersive space.
  • Master a finished mix to a target loudness (about -14 LUFS for streaming) with a safe true-peak ceiling.

Chapter 33: Audio Post-Production

"Films are 50 percent visual and 50 percent sound. Sometimes sound even overplays the visual." — David Lynch

Overview

Take a locked cut that you have already colored and graded — the picture is beautiful, the story is tight, every shot is where it belongs — and play it for someone with the sound turned down. They will say it looks professional. Now turn the sound up on the raw audio, straight off the timeline: the dialogue lurches from loud to nearly inaudible between shots, a low hum sits under every word, the room the interview was shot in rings with echo, a piece of music you dropped in fights the voice for the same space, and the whole thing is either so quiet the viewer reaches for the volume or so hot it distorts on a phone speaker. The picture is finished. The video is not — and everyone can hear it, even the people who could never tell you why.

This is the chapter where "sounds professional" is finally won or lost. Back in Chapter 14 we made a promise as a threshold concept — that sound is half the picture, that audiences forgive a soft image and abandon bad audio within seconds. On set, in Chapters 14 and 15, you kept the first half of that promise: you captured clean dialogue, protected your levels, and recorded room tone at every location. Audio post is where the second half comes due. It is the craft of turning a pile of recorded audio and picture into a single, balanced, delivery-ready soundtrack — cleaning the dialogue, leveling it, scoring it, designing its world, and mastering it to a loudness standard so it plays back correctly everywhere.

It is also, quietly, the most rescuable stage of the whole pipeline and the least forgiving. Rescuable, because a great deal of what makes amateur audio sound amateur — the hum, the unevenness, the harshness, the dead silence between words — is fixable at the timeline with tools that are free. Unforgiving, because you cannot un-clip a distorted peak, cannot remove the echo of a bad room after the fact, and cannot invent a music license you never bought. Post can polish what the shoot captured; it cannot resurrect what the shoot destroyed. That is why this chapter leans, again and again, on the work you already did on set.

By the end you will be able to sit down at a locked cut, open its audio, and walk it through a repeatable order of operations — clean, level, score, design, master — that produces a mix a client would pay for and a platform will play back at the right volume. And you will finally hear the Café Scene the way it was always meant to sound.

In this chapter you will learn to:

  • Run the audio-post workflow in the right order, with the mix built dialogue-first.
  • Perform dialogue cleanup and noise reduction — killing rumble, hum, clicks, and hiss without wrecking the voice.
  • Use EQ and compression to make dialogue clear, present, and even.
  • Choose and place a properly licensed music bed that lifts the story instead of smothering it.
  • Design ambience and sound effects (SFX) that make an ordinary scene feel like a real place.
  • Master the finished mix to a target loudness in LUFS with a safe true-peak ceiling.

Learning Paths

Every reader needs a clean, level, correctly-loud mix — that is not optional. But weight your attention by your goal:

  • 📱 Phone-first: §33.2 and §33.6 are your highest-return sections. A two-minute noise-reduction pass and a correct loudness export will do more for your videos than any camera upgrade — and both are free in DaVinci Resolve or Audacity.
  • 🎥 Creator: §33.4 (music and licensing) and §33.6 (loudness) decide whether your channel sounds "produced" and whether it gets normalized down to mush. Master to spec; use music you are actually allowed to use.
  • 💼 Pro-track: all six sections. Clients hear the mix before they see the grade. §33.3 (leveling) and §33.6 (delivery loudness) are where you meet a spec and keep the account.
  • 🎓 Student: follow the order of operations in §33.1 like a checklist; it is the single most transferable thing in the chapter. Then do the Production Checkpoint on Project 3.

33.1 The audio-post workflow and the mix

Let us name the destination before we start walking toward it. Audio post (audio post-production) is the whole set of tasks that turn the raw sound on your timeline into a finished soundtrack: organizing and syncing the audio, cleaning and leveling the dialogue, adding music and sound effects and ambience, balancing all of it together, and mastering the result to a delivery loudness. The mix is the balanced blend that comes out the other end — the single act of setting every element's level, tone, and placement so they combine into one soundtrack in which nothing is fighting and nothing is lost, and above all the dialogue is always clear.

Two ideas make everything else in this chapter make sense, so we install them first.

You do not start until the picture is locked. Every move in audio post is keyed to a specific frame — a noise reduction learned from a specific pause, a music cue that lands on a specific cut, a sound effect synced to a specific door opening. If you re-edit the picture afterward, every one of those alignments slides out of place and you do the work twice. So audio post is a finishing stage: the cut is done, the color is done (Chapters 31 and 32), and now the sound gets its turn. This is the same discipline as the rest of the book — fix it in pre, not in post — pointed at your own workflow: lock the cut, then finish the sound.

The dialogue is the boss of the mix. Everything in a mix exists in a priority order, and dialogue sits at the top of it, always. Music serves the dialogue. Sound effects serve the dialogue. Ambience serves the dialogue. If any of them ever makes a word harder to understand, that element is too loud, full stop — no matter how much you love the track. A viewer will forgive a quiet music bed; they will not forgive missing a line. Build the mix in that order of importance and you cannot go far wrong; build it by dropping in whatever is fun (music first, usually) and you will spend the rest of the session fighting your own soundtrack.

Those two ideas give us the order of operations — the spine of the entire chapter:

FIGURE 33.1 — The audio-post workflow: dialogue first, master last

  LOCKED         ORGANIZE &        DIALOGUE          EQ &            MUSIC          SFX &          FINAL MIX        MASTER &
  PICTURE   →    SYNC        →     CLEANUP     →     LEVEL      →    BED       →    AMBIENCE   →    (balance)   →    DELIVER
  ─────────      ──────────        ──────────        ────────        ──────         ─────────       ──────────       ────────
  the cut is     import audio      remove noise,     high-pass,      choose &       room tone       set the         to ~-14 LUFS
  finished;      to its tracks;    hum, clicks;      carve EQ,       license;       bed + spot      dialogue-first  (streaming);
  no more        sync double-      lower the         even out        duck under     effects,        balance;        peak ≤ -1 dBFS;
  edits          system sound      noise floor       the levels      dialogue       sync to pix     automate rides   QC on 4 systems
                 (Ch.15/27)                                          (Ch.29/38)     (Ch.15 tone)
   ↑ you do NOT begin until picture is locked — every fix below is keyed to a frame that will not move.
   Read left to right: this is the order you work in, and the order of the six sections of this chapter.

Notice the shape of it. The first third is repair (organize, clean, level) — unglamorous, essential, and where most of the professional-versus-amateur difference actually lives. The middle third is building the world (music, effects, ambience) — the creative, fun part everyone wants to jump to. The last third is finishing (balance and master) — the technical gate that decides whether all your work survives contact with a real playback system. Skip the first third and no amount of the second will save you; skip the last and a beautiful mix arrives at the viewer too quiet, too loud, or distorted.

The mix is organized on the timeline into layers called stems — grouped tracks, one group per kind of sound: dialogue, music, effects, and ambience. Keeping them separated is not tidiness for its own sake; it is what lets you set each group's level independently, and it is what a client or a broadcaster will sometimes ask you to deliver (a "music-and-effects" or M&E version with the dialogue removed, so the piece can be re-voiced in another language). Here is the whole mix, laid out as stems, which is also our map for the rest of the chapter:

FIGURE 33.2 — The mix, laid out as stems (the Café Scene soundscape)   [constructed teaching example]

  TRACK            CONTENT ─ left→right is time →                          TYPICAL LEVEL      METER (peak, dBFS)
  ───────────────  ─────────────────────────────────────────────────     ─────────────      ──────────────────────────
  A1 DIALOGUE  ●   [ "one flat white" ][  "for here?"  ][ "thanks" ]       peaks ~ -12 dBFS   -60 -48 -36 -24 -12  0
                   the spine — always on top, always clear                 (the star)         |░░░░░░░░░░░░░░░█▓·|  ◄peak
  A2 DIALOGUE  ○   [ ...second speaker / safety lav... ]                   matched to A1      |░░░░░░░░░░░░░░░█▓·|
  A3 MUSIC     ♪   [ music bed ....................................... ]   -24 to -30 dBFS    |░░░░░░░░░█▓·······|  (under)
                   ducks down under every line, lifts in the gaps          under dialogue
  A4 SFX       ✦   [ door♪ ]         [ espresso ]      [ cup ]             accents ~ -18 dBFS  |░░░░░░░░░░░░█▓····|
                   spot effects, synced to the frame                       (below dialogue)
  A5 AMBIENCE  ≈   [ café room tone / low murmur ..................... ]   -28 to -34 dBFS    |░░░░░░█▓·········|  (bed)
                   the continuous "you are here" bed (from Ch.15 tone)     quietest layer
  ══════════════   ═════════ MIX BUS — everything sums here ═════════      MASTER             INTEGRATED: -14 LUFS
                                                                           limiter @ -1 dBFS   TRUE PEAK: -1.0 dBFS ✓
   Read top to bottom = the priority order. Dialogue is loudest and clearest; every other layer is set in
   reference to it. The MIX BUS sums all five tracks; the master limiter guards the -1 dBFS peak ceiling, and
   the whole program is metered to -14 LUFS integrated for streaming. This is the whole chapter in one picture.

Study that figure, because we will build it piece by piece — dialogue in §§33.2–33.3, music in §33.4, effects and ambience in §33.5, the bus and the loudness meter in §33.6. The levels shown are starting points, not laws; the ratios between them are the real lesson. Dialogue on top. Music and ambience well beneath it. Effects punctuating in between.

One practical note before we dig in, because it decides whether your judgments are trustworthy: you must monitor honestly. Mix on the best headphones or speakers you have, at a consistent, moderate volume (loud mixing tricks you into pulling levels down; quiet mixing tricks you into pushing them up). Then — and this is the step beginners skip — check the mix on the worst speaker your audience will actually use, which is a phone speaker. A mix that is perfect on studio headphones and unintelligible on a phone is a failed mix, because most of your viewers are holding the phone. Reading a meter (§33.3, §33.6) protects you from your own ears; listening on real-world systems protects you from your own studio.

🚪 Threshold Concept: the mix is the moment "sound is half the picture" comes true. In Chapter 14 that phrase was a promise about the shoot — capture clean sound or lose the audience. Here it becomes a promise about the finish. The same footage, cut identically and graded identically, can feel like a home video or a film depending entirely on what you do in the next few hours of audio post: whether the dialogue is even, whether the room hum is gone, whether the music lifts without smothering, whether the whole thing arrives at the right loudness. Nothing on screen changes. The audience will call the difference "production value" and never once say the word sound — but sound is exactly, and entirely, what they are hearing.

🔄 Check Your Eye. 1. Why must the picture be locked before you start audio post? 2. In the mix's priority order, what always sits on top, and what does that imply about a too-loud music track? 3. What are the four kinds of stem, and why keep them separated?

Check yourself

  1. Every audio move is keyed to a specific frame (a noise print from a pause, a cue on a cut, an effect on an action). Re-editing the picture slides all those alignments out of place and forces you to redo the work.
  2. Dialogue always sits on top. If a music track ever makes a word harder to understand, the music is too loud — clarity of dialogue beats everything, always.
  3. Dialogue, music, effects, and ambience. Separated stems let you set each group's level independently and let you deliver an M&E (music-and-effects) version for re-voicing or translation.

33.2 Dialogue cleanup and noise reduction

The first real work is on the dialogue, because the dialogue is the mix, and because it usually arrives with problems the shoot could not fully prevent. Dialogue cleanup is the process of removing everything from the recorded voice that is not the voice — hums, hiss, clicks, thumps, rustle, and room ring — so that what remains is intelligible and clean. Noise reduction is the specific tool at the center of it: a process that identifies steady, unwanted background sound (an air-conditioner drone, electrical hum, tape-style hiss, the constant character of a room) and lowers it while leaving the voice as intact as possible.

Here is the honest framing, and it is the same one this book has used about every stage: cleanup reduces, it does not resurrect. If the shoot gave you dialogue riding well above the noise floor — the quiet inherent hiss of the recording system you met in Chapter 15 — cleanup will make it sing. If the shoot buried the voice down in the noise, or let the room echo swallow it, cleanup can only trade one ugliness for another. This is the payoff of every discipline from Chapters 14 and 15: close mic, good gain, quiet room, room tone recorded. You are now cashing those in. When cleanup goes badly, the fix is almost never a better plugin; it is a better recording next time.

Work the cleanup in a fixed order, gentlest tools first, because each step makes the next one's job smaller:

  1. Spot-fix the obvious transients. Before any broadband work, hunt down the isolated events: a mouth click, a lip smack, a chair creak, a bump when someone knocked the table, a single loud plosive "p." These are handled surgically — a tiny volume dip on that one spot, a de-click tool, or simply cutting the offending few frames and patching the gap with room tone. Do these first so they do not confuse the noise-reduction step.
  2. High-pass to kill the rumble. Roll off the very low frequencies — below roughly 80 Hz — where traffic rumble, HVAC thrum, handling thumps, and wind energy live but the human voice essentially does not. You met the high-pass filter (low-cut) in Chapter 15 as a switch on the mic; here it is the first band of your EQ, and it instantly removes a layer of muddy weight from almost any location recording. (Back it off if the voice is very deep and starts to sound thin.)
  3. De-hum the electrical buzz. A steady tonal hum at 50 or 60 Hz (and its harmonics) comes from mains electricity — bad cables, dimmers, fluorescent ballasts, ground loops. A de-hum tool notches out that exact frequency and its overtones, removing the buzz with almost no cost to the voice because it is so narrow and so specific.
  4. Broadband noise reduction — and here is where Chapter 15 pays off. This is the big one: the tool that pulls down steady hiss and drone across the whole spectrum. It works by learning a noise print — a short sample of the noise by itself, with no voice over it — and then subtracting that fingerprint from the whole clip. And what is the cleanest possible sample of a location's noise, recorded at the same mic, position, and levels as the dialogue? It is the room tone you were told, in Chapter 15, to record thirty seconds of at every single setup. That habit, which felt like a chore on set, is now the exact key that makes noise reduction work. Learn the print from your room tone; apply a modest reduction — 8 to 10 dB is plenty for most dialogue — and stop.
FIGURE 33.3 — Broadband noise reduction: lowering the noise floor

  BEFORE (hiss + drone baked under the voice)          AFTER (noise print learned, reduced ~8-10 dB)
   0 dBFS ┤                                             0 dBFS ┤
  -12     ┤     ╱‖╲    ╱‖╲   ╱‖╲                        -12     ┤     ╱‖╲    ╱‖╲   ╱‖╲     ◄ voice: untouched
  -24     ┤    ╱ ‖ ╲  ╱ ‖ ╲ ╱ ‖ ╲                       -24     ┤    ╱ ‖ ╲  ╱ ‖ ╲ ╱ ‖ ╲
  -40     ┤▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓  ◄ noise floor       -40     ┤
  -55     ┤                                             -55     ┤·············· ◄ floor pushed down, clean
  -inf    ┤                                             -inf    ┤   quiet parts now near-silent
          └────────────────────────────►time                    └────────────────────────────►time
   The drone/hiss (▓) sits around -40 dBFS under and     Learn the "noise print" from a bar of pure room
   between the words — audible in every pause, and       tone (recorded on set in Ch.15!), reduce 8-10 dB,
   the thing that most says "amateur."                   and STOP. Overdo it and the voice turns watery.

Once the floor is down and the tone is clean, most dialogue is ready to level and EQ (the next section). Two special problems deserve a mention because they are common and because their fixes are limited:

  • Reverb / room echo (the recording sounds "roomy," like it was shot in an empty box) can be reduced by a de-reverb tool but never truly removed. Echo is the voice smeared across time, tangled into itself; software can lessen it but not un-tangle it. This is why Chapter 15 spent so long on close-micing and taming the room: room echo is the one location problem post is worst at fixing.
  • Sibilance (harsh, spitty "s" and "sh" sounds, often worsened by a presence boost) is tamed with a de-esser, a compressor that clamps down only on the narrow high-frequency band where sibilance lives, and only when it spikes.

⚠️ Common Mistake: over-processing until the voice sounds underwater. Every cleanup tool trades noise for artifacts, and past a certain point the cure is worse than the disease. Crank noise reduction too hard and the voice develops a watery, swirly, robotic quality — "chirping" or "underwater" — as the tool chews into the sound of the voice itself. The professional standard is just enough: reduce the noise until it stops being distracting, not until it is gone. A little honest room sound under a clear voice is invisible to an audience; an over-cleaned voice full of digital artifacts is not. When in doubt, do less, and bring the ambience bed (§33.5) up slightly to mask what remains.

✂️ In the Edit: the room tone you grabbed on set is now three tools at once. Recording thirty seconds of room tone felt like the most pointless thirty seconds of the shoot. In post it is doing three jobs. First, it is the noise print that makes broadband noise reduction work. Second, it is the patch you lay under edit points and cut-out breaths so the background does not drop to dead silence and "blink" at every cut. Third, it is the seed of your ambience bed (§33.5). One habit on set; three rescues in post. This is what "you shoot for the edit" sounds like.

🔄 Check Your Eye. 1. Cleanup can reduce but not do what? Name the one location problem post is worst at fixing. 2. What on-set recording makes broadband noise reduction actually work, and how? 3. You reduced noise hard and the voice now sounds watery. What went wrong, and what is the fix?

Check yourself

  1. It can reduce noise but not resurrect a buried or echo-swamped voice. Room echo/reverb is the problem post handles worst — software can lessen it but never truly remove it.
  2. Room tone (Ch.15): a clean sample of the location's noise by itself, at the same mic/position/levels, becomes the "noise print" the tool subtracts from the whole clip.
  3. Over-processing — you reduced too far and the tool chewed into the voice, creating artifacts. Back the reduction off to "just enough" (often ~8-10 dB) and mask any residual with a slightly louder ambience bed.

33.3 EQ and leveling dialogue

Clean dialogue is not yet finished dialogue. It may be quiet in one shot and loud in the next, dull and boxy from a lav under a shirt, or thin and harsh from a shotgun in a hard room. Two tools fix this: EQ shapes its tone, and leveling controls its loudness over time. Together they turn cleaned dialogue into dialogue that is present, consistent, and effortless to listen to.

EQ (short for equalization) is the tool that raises or lowers specific frequency bands of a sound — turning down the "mud" in the low-mids, turning up the "presence" that makes a voice intelligible, rolling off a harsh spike. If cleanup decides what is in the recording, EQ decides the balance of tone within it. Think of it as a set of volume knobs, one for each region of the frequency spectrum from the deepest lows (measured in Hz) up to the airy highs (kHz).

For the human voice, a handful of regions matter, and the professional habit is subtractive first: cut the problems before you boost the good parts, because cutting is cleaner and does not add loudness or harshness.

FIGURE 33.4 — Dialogue EQ, before and after (the frequency picture of one voice)

  BEFORE (raw location dialogue)                        AFTER (EQ'd for clarity)
  level                                                 level
   │        ██                     ██                    │                             ▁▄█▄▁
   │    ██  ██  ██             ██  ██                    │                        ▁▄███████▄▁
   │  ████████████▄▄     ▄▄▄███████████▄▄                │         ▁▂▄▄▄▄▄▄▄▄███████████████████▄▄
   └─────────────────────────────────────►freq           └─────────────────────────────────────►freq
     20   80  200  400   1k   3k   8k  16kHz               20   80  200  400   1k   3k   8k  16kHz
     ▲rumble ▲"mud"          ▲boxy   ▲harsh                ✂cut  ✂dip            ▲+presence ▲air
     (nothing rolled off; boomy, cloudy, dull)            (HP@80 · -3dB@300 · +3dB@3-5k · +2dB@10k)
   Problems: the low rumble adds nothing but weight;      Fixes: high-pass rolls off the rumble; a gentle
   a 200-400 Hz build makes the voice "muddy"; a          300 Hz dip clears the mud; a 3-5 kHz lift adds
   3 kHz spike is harsh; the top is dull, so the          presence and intelligibility; a touch of 10 kHz
   voice sits "behind glass."                             adds "air." Rule: CUT problems, THEN boost.

Read the figure as a recipe you will adapt to every voice:

  • High-pass at ~80 Hz — the same low-cut from cleanup, removing sub-bass rumble that carries no voice.
  • Cut the mud, ~200-400 Hz — a gentle dip (a couple of dB) clears the boomy, cloudy build-up that makes a voice sound like it is talking into a barrel. Lavs under clothing especially need this.
  • Watch the boxiness, ~500 Hz-1 kHz — small cuts here open up a "boxy," small-room quality.
  • Boost presence, ~2-5 kHz — a modest lift adds the consonant clarity that makes speech intelligible. This is the single most useful move for making a voice "cut through" a mix. Overdo it and you get harshness and sibilance (reach for the de-esser).
  • Add air, ~10-12 kHz — a small high-shelf boost adds an open, expensive-sounding sheen. A little goes a long way.

Every voice is different — a deep voice needs less mud-cutting, a thin voice needs the presence boost more — so use the figure as a starting map and listen, ideally comparing to the raw version by toggling the EQ on and off.

With the tone fixed, we fix the dynamics — the swing between loud and quiet. Location dialogue is uneven: a subject leans in and booms, then sits back and murmurs; one shot was recorded a little hotter than the next. Two tools even it out. Riding the levels (also called automation, or clip gain) means manually raising the quiet bits and lowering the loud bits, moment to moment — the most transparent method and worth doing on the biggest swings. Compression does it automatically: a compressor turns down whatever crosses a threshold, narrowing the gap between loud and quiet so the whole performance sits in a tighter, more consistent band. A gentle compressor on dialogue (a low ratio, a few dB of reduction on peaks) is standard; a heavy one pumps and sounds unnatural.

The target you are leveling toward is a consistent working level, read on the same dBFS meter you used to set gain in Chapter 15. Aim dialogue to peak around -12 dBFS and to sit at a steady average, with the loudest moments still leaving headroom below the 0 dBFS ceiling. (dBFS governs peaks and clipping; the overall program loudness in LUFS — how loud it feels — is the separate, final measurement we master to in §33.6.) The point of leveling is not a number; it is that the viewer never once reaches for the volume knob during your piece.

⚙️ Settings Box: a starting dialogue chain (adjust by ear, do not worship). Order matters — cleanup, then tone, then dynamics.

text STAGE TOOL / SETTING (starting point) WHY ─────────────── ─────────────────────────────────────── ───────────────────────────── 1 spot fixes de-click / manual dips on transients kill isolated clicks & bumps first 2 high-pass low-cut @ ~80 Hz remove rumble that carries no voice 3 de-hum notch @ 50/60 Hz + harmonics (if buzzing) surgically remove electrical hum 4 noise reduce learn print from room tone; reduce 8-10 dB lower the noise floor, no artifacts 5 EQ (subtract) -2 to -3 dB @ 250-400 Hz (mud) clear the boom/box 6 EQ (boost) +2 to +4 dB @ 3-5 kHz (presence) intelligibility / "cut through" 7 EQ (air) +1 to +2 dB high-shelf @ 10 kHz open, polished top end 8 de-ess tame 5-8 kHz only when "s"/"sh" spike stop harsh sibilance 9 compress ratio ~2:1-3:1, 3-6 dB gain reduction even out the loud/quiet swing 10 level/ride automate the big swings by hand transparency the compressor can't give TARGET: dialogue peaks ~ -12 dBFS, steady average, always below 0 dBFS. Bypass often to compare to raw.

🔬 The Tech: dBFS versus LUFS — peak versus loudness (skippable). These two scales measure different things and beginners conflate them. dBFS (decibels relative to full scale) measures the instantaneous peak level of the signal; 0 dBFS is the hard digital ceiling, and crossing it is clipping — permanent distortion. It answers "am I about to distort?" LUFS (Loudness Units relative to Full Scale) measures perceived loudness averaged over time, weighted to match how human hearing responds across frequencies. It answers "how loud does this feel to a listener?" A gunshot and a sustained pad can hit the same dBFS peak yet feel wildly different in loudness — LUFS captures that difference; dBFS cannot. You level dialogue by peaks (dBFS) during the mix, and you master the finished program by loudness (LUFS) at the end (§33.6). You can skip the math entirely and still get both right by watching the two meters.

🎬 On Set (well — at the timeline): clean and level one minute of dialogue. Take one continuous minute of your own recorded dialogue (Project 1's talking-head is ideal). Run the chain above in order: spot-fix, high-pass, de-hum if needed, a modest noise reduction learned from your room tone, subtractive then additive EQ, and a gentle compressor. Constraint: bypass the whole chain every thirty seconds and listen to the raw version, so you never drift into over-processing. Self-review question: play the finished minute on your phone speaker at half volume — is every single word clear without you leaning in? If not, the fix is almost always more presence (3-5 kHz) and more consistent leveling, not more volume.

🔄 Check Your Eye. 1. Why is EQ done "subtractive first"? 2. Which frequency region, boosted modestly, most improves a voice's intelligibility — and what is the risk of overdoing it? 3. What is the difference between what dBFS measures and what LUFS measures?

Check yourself

  1. Cutting a problem frequency is cleaner than boosting around it: it adds no loudness or harshness and avoids piling energy into an already-busy band. Cut the mud/harshness first, then boost the good parts.
  2. The presence range, roughly 2-5 kHz — it adds consonant clarity and makes the voice "cut through." Overdone, it turns harsh and spitty (sibilance), which you then have to tame with a de-esser.
  3. dBFS measures instantaneous peak level (are you about to clip at 0 dBFS?); LUFS measures perceived loudness averaged over time (how loud does the whole program feel?). You level by peaks and master by loudness.

33.4 Music: selection, licensing, and the bed

With the dialogue clean and level, we start building the world around it — and the first layer most videos reach for is music. Used well, music is the fastest emotional lever you have; used badly, it is the fastest way to smother a good piece or to get it taken down for copyright. This section is about both halves: choosing and placing music that lifts the story, and doing it legally.

A music bed is a piece of music laid underneath the dialogue and picture to set mood, energy, and pace — sitting low in the mix, supporting the voice rather than competing with it. The word bed is the whole idea: the music is the surface the scene rests on, not the thing on stage. It fills the emotional space, carries momentum through quiet stretches, and tells the viewer how to feel — all from beneath the dialogue, never on top of it.

Choosing the right bed is a storytelling decision, and it is the same skill you began building in Chapter 29, where music and the emotional edit shape a piece's feeling. A few principles:

  • Match the energy and arc, not just the genre. The music should track the shape of your piece — building where it builds, pulling back where it breathes. A three-minute doc that stays on one unchanging loop feels flat; one whose music swells and recedes with the story feels alive.
  • Prefer one evolving theme to a playlist of songs. This is the lesson of the Up "Married Life" montage from Chapter 1: a single melody, re-orchestrated as the story changes, binds a piece into a whole, while a pile of different tracks fragments it. When you can, pick one cue and let it carry.
  • Leave room for the voice. The best dialogue-under-music tracks are sparse in the exact frequency range where the voice lives (the presence range you just boosted). Dense, vocal-heavy, or bright-and-busy music fights your dialogue for the same space; simple, warm, instrumental beds get out of its way.
  • Cut music to picture, and picture to music. Land musical changes on your edit points, and — a trick from Chapter 29 — sometimes nudge a cut to land on a musical beat. When music and image change together, the edit feels intentional.

Now the part that is not optional. You may not use commercial or popular music you have not licensed. Dropping a song you love off a streaming service into your video is copyright infringement, whether or not you make money from it, and on most platforms it will get your video muted, blocked, demonetized, or taken down by an automated content-matching system. This is real, it is enforced, and "but I credited them" is not a license. The legal, professional paths are:

  • Royalty-free / production-library music — you pay once (or a subscription) for a license to use tracks in your projects. This is the workhorse for creators and corporate work.
  • Creative Commons music — free to use if you follow the specific license, which usually requires attribution and sometimes forbids commercial use. Read the exact terms; "Creative Commons" is not one rule but several.
  • Commissioned/original music — you pay a composer to write it, and you own or license the result cleanly. Best for high-end work.
  • Platform-provided libraries — some editing and hosting platforms include cleared music for use on that platform.

The one rule that keeps you safe: know why you are allowed to use every piece of music in your project, and be able to prove it. The full legal treatment — how licensing actually works, sync rights, the difference between the composition and the recording, and what a release covers — is the business of Chapter 38. Here, just build the habit: licensed music only, terms read, receipt kept.

Placing the bed is a level-and-automation job. The bed sits well under the dialogue — often in the -24 to -30 dBFS range against dialogue peaking near -12 dBFS — and it ducks: it drops a few more dB under every line and lifts back up in the gaps between speech. You can do this by hand (drawing volume automation down under each line) or automatically with a sidechain (the dialogue signal tells a compressor on the music to pull the music down whenever anyone speaks). Either way, the effect is the same and it is the mark of a professional mix: the music breathes around the voice, loud in the silences, respectfully quiet under the words.

⚠️ Common Mistake: music that fights the dialogue (or the copyright system). Two errors, both fatal. The first is level: the music feels great while you are choosing it, so you set it too loud, and now the viewer strains to hear the words over it — remember the priority order, dialogue is the boss, and duck the music. The second is rights: using a famous track "just for now" and forgetting to replace it, then publishing and getting the video blocked. Never build a mix on music you are not licensed to use, because you will fall in love with the timing and the swap will hurt. Pick a legal track first.

♿ Accessibility & Inclusion: loud music makes your video unusable for many. Music mixed too high under dialogue is not just an aesthetic problem — for viewers who are hard of hearing, for non-native speakers, and for anyone watching in a noisy place, an over-loud bed can push the dialogue below the threshold of intelligibility entirely. Keeping music respectfully under the voice is an accessibility practice, not only a taste one. And it does not replace captions: however clean your mix, accurate captions (Chapter 36 covers them on export) are what make the piece reachable by deaf viewers and by the enormous audience watching with the sound off.

🔗 Connection. Music's role in shaping emotion and rhythm is Chapter 29 (§29.5, music and the emotional edit); the "one evolving theme" idea traces back to the Up montage in Chapter 1's case study; and the legal machinery of licensing — sync rights, releases, copyright — is Chapter 38 (§38.5). This section is where those threads meet the timeline.

🔄 Check Your Eye. 1. What does it mean for a music bed to "duck," and why does it matter? 2. Name two legal ways to score a video and the one habit that keeps you out of copyright trouble. 3. Why is "I credited the artist" not a defense for using an unlicensed commercial song?

Check yourself

  1. Ducking means the music drops a few dB under every line of dialogue and lifts back in the gaps — done by hand or by sidechain. It keeps the dialogue (the boss) clear while still letting the music support the scene.
  2. Royalty-free/production-library, Creative Commons (per its exact terms), commissioned original, or a platform's cleared library. The habit: use only music you can prove you are licensed to use, terms read and receipt kept.
  3. Credit is not a license. Copyright gives the owner control over use regardless of attribution; unlicensed use is infringement and automated systems will block/mute/demonetize it. (Full treatment: Ch.38.)

33.5 Sound effects and ambience design

Clean dialogue and a good music bed can still sound strangely empty — like the scene is happening in a soundproof void. The missing layer is the world: the continuous ambience of the place and the specific sounds of things happening in it. This is sound design, and it is where an ordinary scene becomes a place you believe you are sitting in.

Sound effects (SFX) are added, non-musical sounds — a door's bell, an espresso machine, a cup on a saucer, footsteps, a phone buzz — that either sync to an on-screen action or fill the world off-screen. They come in two broad kinds, and a good soundscape uses both:

  • Ambience (the bed of the place). A continuous, low background that establishes where we are: the murmur-and-clink of a café, the hum of an office, the wind-and-birds of a park. This is the direct descendant of the room tone you recorded in Chapter 15 — often it is that room tone, looped and extended, sometimes layered with a richer library ambience to add life. The ambience bed runs under the entire scene at a low level (around -28 to -34 dBFS), and it does a quiet miracle: it makes the whole scene cohere, and it hides your dialogue edits, because the background no longer drops out and "blinks" at every cut.
  • Spot effects (the specific events). Discrete sounds tied to a moment: the door bell as someone enters, the hiss-and-clunk of the espresso machine as the order is made, the cup set down, a chair scraping. These are placed frame-accurately against the picture — the sound must land on the same frame as the action, or the eye and ear disagree and the illusion breaks. Small human sounds like footsteps and cloth movement are their own craft, traditionally called Foley, performed and recorded in sync to picture; you can approximate a lot of it by recording your own.

Sourcing effects follows the same rules as everything else: record your own whenever you can (it is free, exclusive, and exactly right), draw from properly licensed effects libraries otherwise, and mind the license on anything you download. A door you record yourself will always fit your door better than a stranger's.

Here is the anchor of the whole book arriving at its sonic payoff. We built the Café Scene across thirty-two chapters — framed it, covered it, moved the camera, lit it with the window, recorded its dialogue and room tone, cut it, found its story, corrected and graded it to a warm café look. The picture has not changed since Chapter 32. Watch what sound alone now does to it:

FIGURE 33.5 — Building the Café Scene soundscape, one layer at a time   [constructed teaching example]

  LAYER ADDED                     WHAT IT DOES                              WITHOUT IT
  ──────────────────────────────  ───────────────────────────────────      ──────────────────────────
  0. cleaned dialogue only        every word clear — but it sounds like     "recorded in a closet";
     (NR + EQ + level)            the café is empty and dead                no place, no life
  1. + café ambience bed          the room comes alive: a low murmur,       dialogue floats in a void
     (the Ch.15 room tone,        a distant fridge, the sense of walls
      looped, ~-30 dBFS)          — now we are somewhere
  2. + door bell (spot SFX)       the "a story is starting" sound from      the entrance means nothing
     synced to the door open      Chapter 1 — punctuates the walk-in
  3. + espresso machine (SFX)     the counter gets a life and a rhythm;     the café has no craft, no
     under the order              also masks a dialogue edit beneath it     texture, no "coffee" in it
  4. + cup / saucer / chair       small human sounds sell the reality       the world feels unfinished
     (Foley, synced)
  5. + music bed, ducked          emotion and momentum; lifts in the        the scene is real but flat —
     under the dialogue           gaps, bows under every line               no feeling steering us
  ─────────────────────────────────────────────────────────────────────────────────────────────────
   Six layers, no new picture. The ordinary scene became a place you could believe you were sitting in.
   The scene never changed — your command of it did.

Read down that figure and you can hear the scene assemble itself. Layer 0 is intelligible and dead. By layer 5 it is alive, located, and steering your feeling — and every layer beneath the dialogue is quieter than the dialogue, exactly as the priority order demands. Notice layer 3's hidden gift: the espresso machine, added for realism, also masks an edit in the dialogue underneath it, because a continuous louder sound covers the small discontinuity of a cut. Sound design is not only decoration; it is repair.

💡 Why It Works: the audience believes what it hears. There is an old truth in this craft: the eye is easy to fool, but a scene sounds real before it looks real. Add the right ambience and spot effects and the viewer's mind stitches the world together and stops questioning it — a flat image gains depth, a cut becomes invisible, an empty set becomes a busy café. That is why sound designers say sound "sells" the picture. The audience is not consciously listening to the espresso machine; they are simply, quietly, believing they are in a café. Take it away and they cannot say what is wrong — only that something is.

🎬 On Set (at the timeline): build a three-layer soundscape. Take any 20-30 second clip with an action in it (someone entering a room, making food, walking outside). Build three layers under the dialogue: (1) a continuous ambience bed appropriate to the place — your own recorded room tone is perfect; (2) at least two spot effects synced frame-accurately to on-screen actions; (3) one small Foley element (footsteps, a cup, a click). Constraint: every added layer must sit below the dialogue level. Self-review question: mute all three layers, then unmute them — does the scene go from "a void with talking" to "a real place"? If not, your ambience is probably too quiet or your spot effects are off-sync.

🔄 Check Your Eye. 1. What is the difference between an ambience bed and a spot effect? 2. Name two jobs the ambience bed does beyond "adding realism." 3. Why must spot effects be placed frame-accurately?

Check yourself

  1. An ambience bed is a continuous, low background establishing where we are (often looped room tone). A spot effect is a discrete sound tied to a specific on-screen or off-screen event (a door, a cup, footsteps).
  2. It makes the scene cohere into one place, and it hides dialogue edits by keeping the background continuous so it never drops to dead silence and "blinks" at a cut. (A loud spot effect can also mask an edit beneath it.)
  3. Because the eye and ear must agree — a sound that lands on a different frame than its action breaks the illusion; the viewer notices the mismatch even if they cannot name it.

33.6 The final mix and mastering to LUFS

Every layer now exists: clean, level dialogue; a licensed, ducked music bed; ambience and spot effects. The mix is the act of balancing them into one soundtrack, and mastering is the final step that sets the whole program to the correct loudness and a safe peak for delivery. This is the gate every piece of audio must pass through, and it is where a lot of otherwise-good work quietly fails.

Balancing the mix. With the stems built, you set their relative levels against the one fixed reference — dialogue — and you do it while watching the picture and listening on honest monitors. Dialogue on top and consistent. Ambience a soft, continuous floor beneath everything. Music ducked under the voice, rising in the gaps. Effects punctuating at their moments, below the dialogue. Most of this you have already dialed in section by section; the mix is where you audition it all together and adjust, because layers interact — the music that sounded fine alone may crowd the dialogue once ambience is added. Three more moves finish the balance:

  • Automate the rides. Real mixes are not static. You draw level automation across the timeline — bringing music up in a montage and down under a key line, lifting a quiet speaker, dipping an effect that competes with a word. A mix that never moves sounds flat; the movement is invisible and it is the craft.
  • Keep dialogue centered; place the world. Pan dialogue to the center (it should feel like it comes from the screen, from the person). Ambience and some effects can spread wider (stereo) to create space, but a voice panned off-center is disorienting — center it.
  • Place sounds in one space with reverb. A tiny, matched reverb can glue a clean, close-mic'd voice into the same room as the ambience — so the dialogue and the world sound like they share a space rather than like the voice was pasted on top. A little; too much and you have recreated the room echo you fought to remove.

Mastering to loudness. Here is the term the whole delivery world turns on. Loudness is how loud a program feels to a listener over its whole duration — not its momentary peaks, but its sustained perceived level — and we measure it in LUFS (Loudness Units relative to Full Scale). LUFS mastering is the final process of adjusting the entire mix so its overall loudness hits a target number that the destination platform expects, while keeping its true peaks below a safe ceiling.

Why does this matter so much? Because every streaming platform normalizes loudness. When you upload, the platform measures your program's integrated (whole-program) loudness and turns it up or down to hit their house target — around -14 LUFS for most streaming and social video. This has two consequences that beginners learn the hard way:

  • If you master too quiet, the platform turns you up — and there is usually no penalty, though you may raise your noise floor.
  • If you master too loud (crushing the mix with a limiter to be "competitive"), the platform simply turns you down to its target, and now your over-compressed, lifeless, dynamically-flat mix plays back at the same loudness as everyone else's — you destroyed your dynamics for nothing. Loudness normalization means the "loudness war" is unwinnable on these platforms; the only thing over-limiting buys you is a worse-sounding mix at the same volume.

So you master to the target, not past it. You put a loudness meter on the master bus, you set a true-peak limiter with its ceiling at -1 dBFS (so nothing clips, with a little margin for the distortion that lossy delivery codecs can add), and you adjust the overall level until the integrated reading sits at your target.

FIGURE 33.6 — The master bus: from mixed stems to a delivered loudness

  DIALOGUE bus ┐
  MUSIC bus    ├──►  MIX BUS  ──►  gentle bus  ──►  TRUE-PEAK   ──►  LOUDNESS       ──►  EXPORT
  SFX bus      ┘     (sum)         glue comp        LIMITER          METER               to spec
  AMBIENCE bus       balance       (optional,       ceiling          reads the WHOLE
                     set here      1-2 dB)          -1 dBFS          program, not peaks
                                                                     ▼
                          ┌──────────────── LOUDNESS METER ──────────────────┐
                          │ INTEGRATED (whole program):  -14.0 LUFS   ◄ target │
                          │ SHORT-TERM (last 3 s):       -13.2 LUFS           │
                          │ TRUE PEAK (max):              -1.0 dBTP  ✓ safe    │
                          │ LOUDNESS RANGE (LRA):          6 LU                │
                          └────────────────────────────────────────────────────┘
   Peaks (dBFS) tell you about clipping; LOUDNESS (LUFS) tells you how loud it FEELS over time.
   You master to the INTEGRATED LUFS number, and you never let the TRUE PEAK cross -1 dBFS.

Different destinations publish different targets, and you should match the one you are delivering to. The common ones:

Destination Integrated loudness target Peak ceiling Notes
Streaming / social video (YouTube, most platforms) about -14 LUFS -1 dBFS (true peak) platforms normalize to roughly this; the book's Project 3 target
Podcast / spoken-word about -16 LUFS -1 dBFS some services target -16; mono deliverables common
Broadcast TV — Europe (EBU R128) -23 LUFS -1 dBTP a strict, enforced standard; ±1 LU tolerance
Broadcast TV — North America (ATSC A/85) -24 LKFS -2 dBTP LKFS is the same measurement as LUFS under another name

Treat the exact numbers as current, revisable targets — platforms publish and occasionally change them, so confirm the spec for your destination before you deliver (Chapter 36 covers matching delivery specs in full). The principle is stable even when a number shifts: measure your integrated loudness, master to the destination's target, and never cross the peak ceiling.

Quality-control the finished mix. Before you call it done, run the checks that catch the failures a quiet studio hides:

  • Play it on four systems: good headphones, laptop/computer speakers, a phone speaker, and — if you can — a car or a TV. A mix that survives all four is a real mix. The phone-speaker pass is non-negotiable, because that speaker cannot reproduce the low end and will expose any dialogue that was leaning on it.
  • Check it in mono. Sum the mix to mono and listen: some stereo effects and doubled sounds partly cancel in mono (phase problems), and a surprising number of viewers hear your video through a single phone speaker. If the dialogue survives mono, it survives anywhere.
  • Watch the whole thing at delivery loudness, once, silently taking notes, then fix and re-check. The final pass is a viewing, not a listening — problems you tune out when hunting for them jump out when you just watch.

⚙️ Settings Box: the master bus and delivery targets (a starting point).

text MASTER-BUS CHAIN (in order) DELIVER TO MASTER AT ───────────────────────────────────────────── ──────────────────────── ───────────────── 1 loudness meter (monitor integrated + true peak) streaming / social ~ -14 LUFS, -1 dBFS 2 optional glue/bus compressor: 1-2 dB, slow podcast / voice ~ -16 LUFS, -1 dBFS 3 true-peak limiter: ceiling -1 dBFS broadcast (EBU R128) -23 LUFS, -1 dBTP 4 adjust master level until INTEGRATED = target broadcast (ATSC A/85) -24 LKFS, -2 dBTP QC: phone speaker + mono check + one silent watch-through at final loudness before you export. Do NOT crush with the limiter to be "loud" — platforms normalize; over-limiting only costs you dynamics.

🔬 The Tech: true peak, integrated vs short-term, and loudness range (skippable). A loudness meter shows several numbers. Integrated loudness is the average over the entire program — the one you master to. Momentary and short-term are averages over very short windows (roughly 0.4 s and 3 s), useful for watching moment-to-moment loudness while you work. Loudness Range (LRA) describes how much the loudness varies across the program — a big LRA means big dynamics (a dramatic piece), a small LRA means consistent loudness (a talky podcast). True peak (dBTP) is subtler than sample peak: when a digital signal is converted back to analog, the reconstructed waveform can overshoot the highest sample, so a signal reading 0 dBFS in samples might actually peak above 0 in the analog output and distort. A true-peak limiter predicts and prevents those inter-sample overs — which is why the ceiling is set at -1 dBFS rather than exactly 0. You can master perfectly well by watching just two numbers (integrated LUFS and true peak); the rest is diagnosis.

🔄 Check Your Eye. 1. If you crush your master with a limiter to make it as loud as possible, what does a streaming platform do to it — and what have you lost? 2. What two numbers do you actually have to watch to master a mix correctly, and what does each control? 3. Why check the mix on a phone speaker and in mono?

Check yourself

  1. The platform normalizes it down to its house target (around -14 LUFS), so it plays back no louder than anyone else's — but now it is over-compressed and dynamically flat. You destroyed your dynamics for no gain in loudness.
  2. Integrated loudness in LUFS (set it to the destination's target — about -14 for streaming) and true peak in dBFS/dBTP (keep it below -1 to avoid clipping). Loudness governs how loud it feels; true peak guards against distortion.
  3. The phone speaker is what most of your audience uses and it cannot reproduce low end, so it exposes any weak or bass-dependent dialogue; the mono check reveals phase cancellation and matches single-speaker playback. If dialogue survives both, it survives everywhere.

Production Checkpoint

Your task: mix Project 3. Your 5-minute branded piece is edited, corrected, and graded. Now give it a finished soundtrack, working the order of operations from this chapter:

  1. Lock and organize. Confirm the picture is locked. Lay your audio out in stems — dialogue, music, effects, ambience — on separate tracks.
  2. Clean and level the dialogue. Spot-fix transients, high-pass the rumble, de-hum if needed, and run a modest noise reduction learned from your room tone. Then EQ for clarity (cut the mud, add presence) and level it so dialogue peaks around -12 dBFS and never makes the viewer reach for the volume.
  3. Add a licensed music bed. Choose one properly licensed track (royalty-free, Creative Commons per its terms, or original — you must be able to prove your right to use it), place it under the dialogue, and duck it so it lifts in the gaps and bows under every line.
  4. Design the ambience and effects. Lay a continuous ambience bed under the piece and add the spot effects the story needs, synced to picture.
  5. Master to loudness. Put a loudness meter and a true-peak limiter (ceiling -1 dBFS) on the master bus and master the whole program to about -14 LUFS integrated. QC it on a phone speaker and in mono.

Why this matters: this is the step where Project 3 stops being "a nicely edited video" and becomes a piece that sounds like professional work — the half of "production value" that no color grade can supply. It is also the payoff of every audio discipline you have practiced since Chapter 14. Keep your stems; Chapter 36 will export the master to delivery specs, and Chapter 39 will fold this piece into your reel. The scene never changed — your command of it did, and now that command includes its sound.

Summary

  • Audio post turns raw timeline sound into a finished soundtrack; the mix is the balanced blend, built dialogue-first. You do not start until the picture is locked.
  • The order of operations (memorize it):
Step What you do The rule
Organize & sync audio into stems (dialogue/music/SFX/ambience) separate so you can control each
Dialogue cleanup de-click, high-pass, de-hum, noise-reduce reduce, don't resurrect; less is more
EQ & level cut mud, add presence; compress/ride cut before you boost; peaks ~-12 dBFS
Music bed choose, license, place, duck dialogue is the boss; music serves
SFX & ambience ambience bed + synced spot effects ambience hides edits; sync to frame
Final mix & master balance, automate, master to loudness ~-14 LUFS, true peak ≤ -1 dBFS
  • Dialogue cleanup guide: spot-fix transients → high-pass ~80 Hz → de-hum 50/60 Hz → broadband noise reduction (print from room tone, ~8-10 dB) → de-reverb (limited) / de-ess as needed. Over-processing = "underwater" voice; do less.
  • Dialogue EQ, subtractive first: high-pass 80 Hz; cut mud 200-400 Hz; boost presence 3-5 kHz (intelligibility); a touch of air at 10 kHz. Then compress gently and ride the big swings.
  • Music: match the arc, prefer one evolving theme, leave room for the voice, and only ever use music you are licensed to use (royalty-free, Creative Commons per terms, original, or platform-cleared). Full legal treatment: Chapter 38.
  • Sound design: an ambience bed (often your looped room tone) makes the scene a real place and hides edits; spot effects and Foley, synced frame-accurately, give it life. The audience believes what it hears.
  • Loudness targets: streaming/social ~-14 LUFS; podcast ~-16 LUFS; EBU R128 broadcast -23 LUFS; ATSC A/85 -24 LKFS. True peak ≤ -1 dBFS. Platforms normalize loudness, so do not over-limit — master to the target, not past it.
  • QC: listen on headphones, computer, phone speaker, and a car/TV; check in mono; watch it through once at final loudness.
  • dBFS vs LUFS: dBFS = instantaneous peak (clipping at 0); LUFS = perceived loudness over time. Level by peaks, master by loudness.

Spaced Review

Bring back three earlier chapters this one depends on:

  1. (Chapter 14) In one sentence, why is "sound is half the picture," and how does audio post make good on that promise that the shoot began?
  2. (Chapter 15) You are about to run broadband noise reduction on a noisy clip. What did you record on set that makes it work, and exactly what does the tool do with it?
  3. (Chapter 15) Dialogue was aimed to peak around -12 dBFS on set. Why leave headroom below 0 dBFS, and what permanently happens if a peak crosses 0?
  4. (Chapter 29) Chapter 29 shaped the emotional edit with music. Give one reason a single evolving theme beats a playlist of different songs under a piece.
Check yourself 1. Audiences forgive a soft image but abandon bad audio within seconds, so half of "professional" is *sounding* professional. The shoot captured clean dialogue; audio post cleans, levels, scores, designs, and masters it into a balanced soundtrack — delivering the second half of the promise. 2. Room tone — 30+ seconds of the location's noise by itself, recorded at the same mic/position/levels. Noise reduction learns a "noise print" from it and subtracts that fingerprint from the whole clip, lowering the steady background while leaving the voice. 3. Headroom is the safety gap so an unexpected loud moment does not exceed the 0 dBFS ceiling. Crossing 0 dBFS is clipping — the peak is chopped flat into permanent, uncorrectable distortion you cannot fix in post. 4. One transforming theme binds the piece into a single emotional arc (as in *Up*'s "Married Life"); a pile of different songs fragments it. The recurring melody also lets the audience build meaning through repetition.

What's Next

Your video now sounds finished — clean, balanced, and correctly loud. The next layer of the finish is visual information: the title that names the piece, the lower third that names a speaker, the animated element that guides the eye. Chapter 34 takes on motion graphics — titles, lower thirds, keyframes, and the restraint that makes on-screen graphics feel designed rather than decorated. As with sound, the discipline is the same one that has run through the whole book: every graphic, like every cut, light, and cue, has to be motivated — there to serve the story, or not there at all.