Case Study 2: Building the Café Scene Soundscape — A Full Audio-Post Session

This is a production walkthrough, not a film analysis. We sit down at a locked, graded cut of the Café Scene and take its sound from "raw and dead" to "finished and alive," in the exact order of operations from Chapter 33. Follow along in DaVinci Resolve's Fairlight page or in any editor with tracks, an EQ, and a loudness meter (Appendix E). Every diagram is a real session layout you can copy. The scene is a [constructed teaching example].

The brief and the constraints

We have carried the Café Scene through the entire book. We framed it (Chapter 6), covered it with shot sizes and the 180° line on the order exchange (Chapter 7), gave it a slow push-in on the reaction (Chapter 8), lit it with the window as the key (Chapters 11, 13), recorded its counter dialogue and room tone (Chapter 15), assembled the coverage (Chapter 26), cut it with J- and L-cuts (Chapter 28), found its story (Chapter 29), and corrected and graded it to a warm café look (Chapters 31-32). The picture is finished and will not change.

What we have not done is finish its sound. Right now the timeline audio is:

  • Counter dialogue on two tracks — a lav on the customer and a lav/boom on the barista (the two-mic setup from Chapter 15's FIGURE 15.8), plus the camera's scratch track for reference. The order exchange: "One flat white." — "For here?" — "Yeah, thanks."
  • 45 seconds of room tone recorded at the counter, same mics and levels, per the Chapter 15 discipline.
  • Nothing else. No ambience laid, no effects, no music. Played back, the scene is a strange thing: the dialogue is intelligible, but the café sounds empty — a dead room with two disembodied voices, a low fridge hum swelling and vanishing at every cut, and a couple of clicks where someone bumped the counter.

The brief: turn this into a soundscape a client would accept — clean, level dialogue; a living café ambience; the spot effects the story needs (the door bell, the espresso machine, a cup); a licensed music bed under a 90-second scene; mastered to about -14 LUFS for streaming. Free tools only. One evening's work.

The constraints: the picture is locked (nothing we do can require a re-edit); every sound must sit below the dialogue; and every piece of music must be one we can prove we are licensed to use.

Gear and session setup

You need almost nothing to do this well, which is the point.

⚙️ Settings Box: the audio-post session (a starting point).

text ITEM CHOICE WHY ─────────────────── ──────────────────────────────────── ─────────────────────────────── software Resolve (Fairlight page) or Audacity free; full cleanup + meters + LUFS monitoring closed-back headphones + check on phone honest judgment + real-world test sample rate 48 kHz (match the project) video-standard; no resampling project frame rate matches the picture (e.g. 24 fps) effects must sync to the frame DIALOGUE chain HP 80 Hz → de-hum → NR (from room tone) clean before you shape → EQ (cut 300 Hz, +3 dB @ 3-5 kHz) clarity and presence → gentle comp (2.5:1, 3-5 dB) → ride even levels DIALOGUE target peaks ~ -12 dBFS, steady headroom below 0 dBFS MUSIC bed -24 to -30 dBFS, ducked under dialogue dialogue is the boss AMBIENCE bed -28 to -34 dBFS, continuous "you are here" without crowding SFX ~ -18 dBFS accents, synced to frame punctuate below dialogue MASTER loudness meter + true-peak limiter @ -1 to ~-14 LUFS integrated QC phone speaker + mono + one silent watch catch the failures a studio hides

The session, laid out

First we build the stems. A tidy session is a fast session, and separated stems (Chapter 33 §33.1) are what let us control each layer and, if the client ever asks, deliver a music-and-effects version. Here is the empty session we are about to fill:

FIGURE CS2.1 — The Café Scene session, stems laid out (empty, ready to build)

  TRACK            ROLE                              STARTING CONTENT
  ───────────────  ────────────────────────────      ─────────────────────────────────────
  A1 DIALOGUE  ●   customer (lav)                    "one flat white" / "yeah, thanks"
  A2 DIALOGUE  ●   barista (lav/boom)                "for here?"
  A3 MUSIC     ♪   the licensed bed                  (empty — added in Phase 5)
  A4 SFX       ✦   spot effects                      (empty — added in Phase 4)
  A5 AMBIENCE  ≈   café room tone / life             (empty — filled in Phase 3)
  ── scratch ──    camera reference audio            muted; kept only for sync safety
  ══ MIX BUS ══    everything sums here → master     (built in Phase 6)

  We will fill this top-down in priority order: dialogue first (Phases 1-2), then the world
  (Phases 3-5), then the balance and master (Phase 6). This figure is the whole session in one glance.

Phase 1 — Clean the dialogue

We start where the mix always starts: the dialogue, because it is the mix. We work the §33.2 order, gentlest tools first.

The thinking. Before touching a plugin, we listen through the two dialogue tracks once, just noting problems: two clicks (a counter bump on the customer lav, a lip smack before "thanks"), a steady low fridge hum and HVAC drone under everything, and a slightly boomy, "under-the-shirt" quality on the customer's lav. No clipping, thankfully — the gain discipline from Chapter 15 held. Nothing here is fatal; it is all cleanup, not rescue.

The moves, in order:

  1. Spot-fix the two clicks. We zoom into the counter-bump click, select the few frames, and either de-click it or cut it and patch the tiny gap with a sliver of room tone. Same for the lip smack. Two seconds of work each, done first so they do not confuse the noise reducer.
  2. High-pass at 80 Hz on both dialogue tracks — the fridge/HVAC rumble and the boomy weight roll away instantly, and the voices lose nothing. (The customer's deep-ish voice we check at 70 Hz so we do not thin it.)
  3. De-hum. The fridge contributes a tonal buzz; a notch at its fundamental and harmonics removes the buzz cleanly.
  4. Broadband noise reduction — the Chapter 15 payoff. We drop the 45 seconds of room tone onto a scratch track, select a clean bar of it (no voice, no bump), and teach the noise reducer that print. Then we apply it to both dialogue tracks at about 9 dB of reduction. The HVAC drone under the words falls away; the voices stay natural. We deliberately stop there — a test at 18 dB turned the customer's voice watery, exactly the artifact §33.2 warns about.
FIGURE CS2.2 — Dialogue cleanup, before and after (the fridge-and-HVAC floor)

  BEFORE (A1 customer lav, raw)                      AFTER (cleaned)
   0 dBFS ┤                                          0 dBFS ┤
  -12     ┤   ╱‖╲  · ╱‖╲   ╱‖╲                       -12     ┤   ╱‖╲    ╱‖╲   ╱‖╲    ← voice intact
  -24     ┤  ╱ ‖ ╲   ‖ ╲ ╱ ‖ ╲   ← "·" = click       -24     ┤  ╱ ‖ ╲  ╱ ‖ ╲ ╱ ‖ ╲   (clicks gone)
  -38     ┤▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓  ← hum + HVAC floor    -40     ┤
  -55     ┤                                          -55     ┤················· ← floor pushed down
          └──────────────────────────►time                   └──────────────────────────►time
   Boomy, clicky, a droning floor at ~-38 dBFS.      HP 80 Hz + de-hum + 9 dB NR (print from room tone).
                                                      Voice untouched; the café's dead hum is gone.

What we adjusted. After the NR we found the customer's lav still a touch boomy (that under-the-shirt quality). We noted it for the EQ in Phase 2 rather than reducing noise harder — a reminder that not every problem is a noise problem. The clean dialogue now sits over a near-silent floor, which is too clean for a café — but that is correct: we want the dialogue clean and the café to come from a designed ambience (Phase 3), not from leftover recording noise we could not control.

Phase 2 — EQ and level the dialogue

Clean is not finished. Now we shape tone and even the levels, per §33.3.

The thinking. The customer's lav is boomy and slightly dull (typical of a lav under clothing); the barista's boom is thinner and brighter (typical of a boom in a hard-surfaced counter area). Left alone, cutting between them would sound like two different rooms. Our job is (a) make each voice clear and (b) make them match, the audio cousin of the shot-matching from Chapter 31.

The moves:

  • Customer (A1): subtractive first — a 3 dB dip at 300 Hz kills the boom, a small cut at 500 Hz opens the boxiness. Then a +3 dB presence lift at 4 kHz for clarity and a whisper of 10 kHz air. It now sounds like a person, not a shirt.
  • Barista (A2): less mud to cut; mostly a gentle presence match and a touch of low-mid added (the boom was thin) so the two voices share a tonal family.
  • De-ess both lightly after the presence boost, so the added clarity does not become spitty.
  • Compress each gently (about 2.5:1, 3-4 dB reduction) to tame the swing where the customer leans in on "one flat white."
  • Ride the levels by hand on the one big swing: the customer's "yeah, thanks" trailed off quiet, so we draw its clip gain up a few dB. We level both tracks to peak around -12 dBFS.
FIGURE CS2.3 — Matching two mics into one conversation (EQ moves)

  CUSTOMER lav (boomy, dull)                 BARISTA boom (thin, bright)
   cut  -3 dB @ 300 Hz  (kill boom)           add  +2 dB @ 250 Hz  (add body)
   cut  -2 dB @ 500 Hz  (open box)            boost +2 dB @ 4 kHz  (match presence)
   boost +3 dB @ 4 kHz  (presence)            de-ess lightly
   boost +1 dB @ 10 kHz (air)                 gentle comp 2.5:1
   ─────────────────────────────────         ─────────────────────────────────
   GOAL: after EQ, cutting A1↔A2 sounds like ONE room and one conversation, not two recordings.
   Toggle each EQ on/off against the raw to be sure you improved it — cut before you boost.

What we adjusted. After matching, the barista's line "for here?" felt slightly detached — recorded a hair farther back. A tiny, matched room reverb (added later in the mix, Phase 6) will seat both voices in the same space; we make a note. The dialogue is now clean, clear, matched, and even. Play it alone and it is broadcast-usable — and completely, unnaturally empty. Time to build the café.

Phase 3 — Lay the ambience bed

Here the room comes alive, per §33.5. This is the single biggest transformation in the whole session.

The thinking. A café is never silent. It has a low, continuous life — a distant murmur of other customers, the hum of the fridge and machines, the soft clink of crockery somewhere off. That bed does three jobs at once: it establishes place, it fills the unnatural void under our over-clean dialogue, and it hides our dialogue edits by keeping the background continuous so it never "blinks" at a cut.

The moves:

  • Loop the recorded room tone across the entire scene on A5. Our 45 seconds of counter room tone, looped (with crossfades so the loop point is inaudible), is the foundation — it is the actual sound of this café, which no library can match.
  • Layer a richer café ambience underneath, from a properly licensed effects library — a low, sparse "coffee shop murmur" — to add the sense of other people and a larger room than 45 seconds of tone alone implies. We keep it low and roll off its highs so it does not add hiss or compete with the dialogue's presence range.
  • Set the bed level to about -30 dBFS — present enough to feel, low enough that it never touches the dialogue.
FIGURE CS2.4 — Ambience: from a dead room to a place

  BEFORE (A5 empty)                          AFTER (A5 ambience bed)
  A1 DIALOGUE  ● [ flat white ][ thanks ]    A1 DIALOGUE  ● [ flat white ][ thanks ]
  A5 AMBIENCE  ≈ [ ................ empty ]   A5 AMBIENCE  ≈ [ room tone loop + café murmur ....... ]
                                                             └ continuous, ~-30 dBFS, under everything
  Cuts "blink" to silence; the café is a     Background is continuous; cuts vanish; we are unmistakably
  soundproof void with two floating voices.  IN a café before anyone has said a word or seen a wide shot.

What we adjusted. Our first ambience pass was too loud (-24 dBFS) and started to crowd the dialogue's clarity — a small version of the §33.4 "music too loud" mistake, but with ambience. We pulled it to -30 dBFS. Instantly better: you feel the café without noticing the bed. We also confirmed the loop's crossfade point landed under a line of dialogue, hiding it further.

Phase 4 — Add the spot effects

Now the specific events. The ambience is the room; the spot effects are the things that happen in it, synced frame-accurately per §33.5.

The thinking. The scene has a handful of actions that want a sound: the door opening as the customer enters (the bell — canon since Chapter 1's first look at the scene, "the sound that says a story is starting"), the espresso machine working as the order is made, the cup and saucer set down on the counter, and the chair as the customer sits by the window. Each must land on the exact frame of its action.

The moves:

  • The door bell on A4, synced to the frame the door opens in the establishing wide. We record our own — a small bell, one clean ring — because a self-recorded sound always fits better than a stranger's.
  • The espresso machine — a hiss-and-clunk cycle — laid under the order exchange. This one does double duty: it gives the counter a craft and a rhythm, and it masks a dialogue edit between the customer's and barista's lines, exactly the trick from §33.5 (and the "masking" principle dramatized by the waterfall in Case Study 1).
  • The cup and saucer on the frame the barista sets the drink down; the chair as the customer sits. Small Foley, recorded at home in a minute.
  • Level the effects to about -18 dBFS — present and punchy, but clearly below the dialogue.
FIGURE CS2.5 — Spot effects, synced to the frame

  PICTURE   [ wide: door opens ][ counter: order + drink made ][ sits by window: phone ]
  A1 DIAL   ●         [ "one flat white" ][ "for here?" ][ "thanks" ]
  A4 SFX    ✦   [door♪]            [ espresso hiss-clunk ][cup]        [chair]
                  ▲ on the frame     ▲ under + masking a       ▲ on set-  ▲ on sit
                    the door opens     dialogue edit             down frame
  A5 AMB    ≈   [ café bed ........................................................ ]
   Every ✦ lands on its action's exact frame. The espresso SFX also hides the A1 edit beneath it.

What we adjusted. The door bell, on first placement, was a frame late — the ring came after the door visibly opened, and the eye-ear mismatch was obvious even at a glance. We nudged it one frame earlier to land exactly on the door's movement. We also found our espresso SFX ran slightly long past the "thanks," poking out into the sit-down; we trimmed its tail so it settled back into the ambience as the customer moved to the window. Sync is not "close enough"; it is frame-exact or it reads as wrong.

Phase 5 — Choose and place the music bed

The world is built and real. Now we add feeling, and we do it legally, per §33.4.

The thinking. The Café Scene's story is small and warm — an ordinary ritual, a person checking their phone, a reaction to a message. The music should be gentle, warm, and unobtrusive, with room in the presence range for the voices. And it must be one theme, deployed simply, not a montage of songs (the Up lesson from Chapter 1).

The moves:

  • License first. We pick a warm, sparse instrumental track from a royalty-free library we subscribe to, confirm the license covers this use, and keep the receipt. (Had we wanted a specific commercial song, we would need a sync license — the business of Chapter 38 — and for this scene it is not worth it.) We do not build the mix on a "temporary" famous track we would have to swap.
  • Place it on A3 under the whole scene, starting a beat before the door opens so it establishes mood, and let a gentle swell land as the customer reads the phone message (the emotional beat, Chapter 29).
  • Duck it. The bed sits at about -26 dBFS in the gaps and ducks to about -32 dBFS under every line — done here with volume automation drawn by hand under the dialogue, so the music lifts in the silences and bows under the words.
FIGURE CS2.6 — The ducked music bed (automation follows the dialogue)

  A1 DIAL   ●      [ "one flat white" ]      [ "for here?" ]   [ "thanks" ]
  A3 MUSIC  ♪  ────╲___________________╱─────╲____________╱────╲__________╱──── swell ──►
               up   ▼ ducked under line  up    ▼ ducked      up  ▼ ducked   (lifts on
               (-26)  (-32 dBFS)        (-26)   (-32)        (-26) (-32)      the reaction)
   The music is loud in the gaps and quiet under every word. Dialogue is the boss; the bed serves it.

What we adjusted. Our first track was lovely but bright and slightly busy in the 3-5 kHz range — it fought the dialogue's presence even when ducked. We swapped it for a warmer, sparser cue with its energy lower down, and the dialogue immediately sat clear on top. This is the §33.4 lesson in practice: a bed's frequency content, not just its level, decides whether it crowds the voice.

Phase 6 — The final mix and master

Every layer exists. Now we balance them together, automate the moves, and master to loudness, per §33.6.

The thinking. We have set each layer's level in isolation, but layers interact — with the ambience and music both in, does the dialogue still win everywhere? We audition the whole scene, watching the picture, on honest headphones, and adjust.

The moves:

  • Balance and automate. Dialogue on top and consistent; ambience a soft floor; music ducked; effects punctuating. We ride a couple of spots: the music swell under the phone reaction comes up a touch; the espresso SFX dips a hair where it briefly competed with "for here?"
  • Center the dialogue; place the world. Both voices panned center. The café ambience spread wide in stereo for space; the door bell panned slightly to screen-left where the door is.
  • Seat the voices in one room. A tiny, matched reverb on both dialogue tracks glues them into the same café space and fixes the barista's slightly-detached "for here?" from Phase 2.
  • Master. On the mix bus: a loudness meter and a true-peak limiter with its ceiling at -1 dBFS. We play the whole scene and read the integrated loudness, then raise or lower the master until it sits at -14 LUFS. The true peak reads -1.0 dBTP — safe.
FIGURE CS2.7 — The finished Café Scene mix (the whole session, one picture)

  TRACK            CONTENT ─ time →                              LEVEL           METER (peak, dBFS)
  ───────────────  ─────────────────────────────────────       ─────────       ──────────────────────
  A1 DIALOGUE  ●   [ flat white ][ for here? ][ thanks ]        ~ -12 dBFS      |░░░░░░░░░░░░░░░█▓·|
  A2 DIALOGUE  ●   [ ...matched, seated in reverb... ]          matched         |░░░░░░░░░░░░░░░█▓·|
  A3 MUSIC     ♪   [ warm bed, ducked, swells on reaction ]     -26/-32 dBFS    |░░░░░░░░░█▓·······|
  A4 SFX       ✦   [bell]     [ espresso ][cup]      [chair]    ~ -18 dBFS      |░░░░░░░░░░░░█▓····|
  A5 AMBIENCE  ≈   [ room tone + café murmur ............. ]    ~ -30 dBFS      |░░░░░░█▓·········|
  ══ MIX BUS ══    sum → glue → true-peak limiter @ -1 dBFS     MASTER          INTEGRATED: -14.0 LUFS
                                                                                 TRUE PEAK:  -1.0 dBTP ✓
   Compare to FIGURE CS2.1: the same session, now full. Dialogue on top, the world beneath it, mastered.

Quality control. We do not trust the headphones alone. We export a check file and play it (1) on a phone speaker — every word clear, ambience and music present but not muddy, good; (2) in mono — the wide café ambience narrows but the dialogue survives cleanly, no phase cancellation; (3) one silent watch-through at final loudness, just watching, notes in hand. One catch: on the phone, the low end of the music briefly rumbled; we high-passed the music bed a little to clean it. Re-check, done.

Here is what we made, as a Described Shot — the same opening the book first showed you in Chapter 1, now heard the way it was always meant to sound:

FIGURE CS2.8 — "The Café Scene, finished sound"        [constructed teaching example]
  THE FRAME    The wide establishing shot: a person pushes through the glass door, the counter waits, the
               window glows behind — graded warm (Ch.32). The picture is unchanged from Chapter 1.
  THE MOVE     Locked off, as it has always been; later, the slow push-in on the reaction (Ch.8).
  THE LIGHT    The window key, warm counter practicals — the look built across Part III and graded in Ch.32.
  THE SOUND    THIS is what changed. A low, living café bed — murmur, fridge, distant crockery. The door's
               bell rings exactly as it opens. At the counter, clean, warm, matched dialogue rides over a
               hiss-and-clunk of espresso. A cup settles. Under it all, a warm music bed breathes — quiet
               under the words, lifting as the customer reads the message by the window. Every word is clear;
               the room is unmistakably real; mastered to -14 LUFS, it sounds right on a phone and in a car.
  THE CUT      J- and L-cuts (Ch.28) let the ambience and dialogue flow across the picture cuts; the
               continuous bed hides every seam.
  THE EFFECT   You believe you are sitting in the café. The scene has depth, place, and feeling it never had
               as picture alone — and you cannot point to a single sound and call it "the effect."
  THE LESSON   Sound is half the picture, finished. The scene never changed — your command of it did.

Discussion questions

  1. The dialogue was cleaned until the café sounded empty, and then a designed ambience bed was added back. Why is that better than simply leaving the original recording noise in? What does it let you control?
  2. The espresso-machine effect does two jobs at once. Name them, and connect the second (masking an edit) to the waterfall in Case Study 1.
  3. The first music track was rejected not for its level but for its frequency content. Explain, using the dialogue's presence range, why a bright, busy track fights a voice even when it is quiet.
  4. Every spot effect had to be nudged to land on the exact frame of its action. Why is "close enough" not good enough, and how did the door bell reveal this?
  5. The final mix was checked on a phone and in mono, and both revealed problems the headphones hid. Why are these two checks non-negotiable, and which of your own past videos would have failed them?
  6. Trace one thread from set to this session: pick a decision made during the Chapter 15 shoot (mic choice, room tone, gain) and show exactly how it paid off — or would have cost you — in this mix.

Your turn: mix your own scene

Take a 60-90 second scene of your own — ideally your Café Scene equivalent, or Project 2/3 — and run this exact session, in order:

  1. Build the stems: dialogue, music, effects, ambience on separate tracks.
  2. Clean the dialogue (Phase 1): spot-fix, high-pass, de-hum, a modest noise reduction learned from your room tone. Stop before it sounds watery.
  3. EQ and level (Phase 2): cut mud, add presence, match any mismatched shots, compress gently, ride the big swings to ~-12 dBFS peaks.
  4. Lay the ambience (Phase 3): loop your room tone, layer a licensed ambience, set it low (~-30 dBFS).
  5. Add spot effects (Phase 4): synced frame-accurately, below the dialogue; record your own where you can.
  6. Place a licensed music bed (Phase 5): one theme, room for the voice, ducked under every line.
  7. Mix and master (Phase 6): balance, center the dialogue, master to about -14 LUFS with a -1 dBFS ceiling, and QC on a phone and in mono.

Then do the honest test: play someone the before (raw timeline audio) and the after. Ask what changed. They will struggle to name it — and that struggle is the whole point of this chapter. The picture never changed. Your command of its sound did.

Key takeaways

  • Audio post follows a fixed order: organize into stems → clean dialogue → EQ and level → ambience → spot effects → music → mix and master. Work top-down in priority order and the session stays calm.
  • Clean the dialogue past natural, then design the room back in. Over-clean dialogue plus a controlled ambience bed gives you a café you control, instead of leftover recording noise you do not.
  • Match your mics with EQ so cutting between a lav and a boom sounds like one room, not two — the audio cousin of shot-matching.
  • The ambience bed is the biggest single transformation: it establishes place, fills the void under clean dialogue, and hides your edits. Often it is your looped room tone.
  • Spot effects must be frame-exact, and a loud effect (the espresso machine) can mask a dialogue edit beneath it — sound design as repair as well as decoration.
  • License the music first, and choose it by frequency content, not just level — a bright, busy bed fights the voice even when quiet; a warm, sparse one gets out of its way and ducks under every line.
  • Master to the target (~-14 LUFS), ceiling -1 dBFS, then QC on a phone and in mono. The failures that matter are the ones a quiet studio hides.
  • The scene never changed — your command of it did. That is the payoff of "sound is half the picture," and it is now yours to run on every project.