Appendix D: Audio Reference
The screenshot-and-scan card for production sound — Chapters 14, 15, and 33 compressed into tables you can check on set and at the timeline. It assumes those chapters taught you the why; this is the what and the number. The single idea beneath all of it: sound is half the picture (Chapter 14) — get the right mic close to the source, aim it, reject the rest, and give sound equal weight to light.
Notation used throughout: dBFS (decibels relative to full scale; 0 dBFS = the digital ceiling, where clipping happens), LUFS (Loudness Units relative to Full Scale; a program's perceived loudness over time), dB (a relative change in level), Hz / kHz (frequency). In the diagrams: ((• = microphone, ● = mic capsule, ↑ = the subject in front of the mic, → = sound/signal travelling, ))) = a wireless (radio) link.
The whole appendix in one line. Capture dialogue close, on-axis, out of frame at ~-12 dBFS with headroom below
0 dBFS; monitor on headphones and record room tone every time; in post, clean → EQ → level → music → SFX/ambience → master to loudness. Everything below is detail on those six moves.
D.1 Microphone selection — match the mic to the job
Choose by the job, not the brand. Four families cover almost everything you will shoot.
| Mic | What it is / how it sounds | Best for | Watch out for |
|---|---|---|---|
| Lavalier (lav) | tiny clip-on worn on the person ~6–8" below the chin; usually omni; rides close, so it's present and consistent | interviews, testimonials, presenters, moving subjects; the phone-first hero | clothing rustle, cable thumps, off-center dropout on head-turns; it must be hidden |
| Shotgun | long, narrow, directional (interference tube); "reach with rejection" | dialogue at a distance, scripted scenes, run-and-gun, news — usually on a boom | roomy/hollow indoors; still must be close; aim it true; directional is not telephoto for sound |
| Handheld ("stick") | rugged mic held right at the mouth; usually cardioid | reporting, hosting, stages, music, vox pops — genres where seeing a mic is fine | it's visible; occupies a hand; handling noise if gripped carelessly |
| Built-in / on-camera | the camera's own mic (or a shotgun in the hotshoe); stuck at the camera's distance | an arm's-length vlog only; otherwise a scratch track to sync better audio to | almost always too far → roomy and distant; never the final sound unless the camera is already close |
Lav — wired vs wireless.
| Wired lav | Wireless lav | |
|---|---|---|
| Buys you | cheapest, reliable, nothing to run out | freedom to move; no cable back to the recorder |
| Costs you | tethers the subject to the recorder | batteries, a little setup, occasional radio dropouts |
| Reach for it when | the subject is seated / static | the subject walks, or sits far from the camera |
The boom is a delivery system, not a mic. A boom is the pole; you mount a shotgun or supercardioid on it to get a directional mic close and out of frame — the backbone of film and TV dialogue. Full booming technique with a partner is Chapter 15. Buy exactly one audio thing for this whole book? Buy a lav — the highest-return purchase in video.
D.2 Polar patterns — where each mic listens
A polar pattern is the map of a mic's directional sensitivity — the invisible shape of what it hears.
FIGURE D.1 — Four polar patterns, seen from above
(↑ = the subject, in front of the mic; ● = the capsule)
OMNI CARDIOID SUPERCARDIOID SHOTGUN (lobar)
(most lavs) (many handhelds) (interior dialogue) (film / outdoors)
↑ ↑ ↑ ↑
,-----. ,-----. .-` `-. |''|
/ ● \ / ● \ / ● \ |● | long,
( ) ( front ) ( front ) | | narrow
\ / \ / \ / | | reach
`-----` `--.--` `-.o.-` `o `
hears every front + sides, tight front, tiny REAR lobe
direction; no rejects the small REAR lobe (o); rejects the
rejection REAR (behind ●) (o) to watch sides / the room
| Pattern | Hears | Rejects | When to use | Note |
|---|---|---|---|---|
| Omnidirectional | all directions equally | nothing | most lavs | fine because it's close; forgives head-turns |
| Cardioid | front + sides | the rear | many handhelds | point the back at the loudest noise |
| Super / hypercardioid | tight front | the sides (small rear lobe) | interior dialogue — the small-room champ | keep noise out of the zone directly behind it |
| Shotgun (lobar) | narrow front lobe, long reach | the sides | film, outdoors, on a boom | roomy indoors; directional ≠ telephoto for sound |
Off-axis rejection (the whole point): point the sensitive (on-axis) direction at the mouth, and the deaf (off-axis) zone at the noise. For a directional mic, aim is the recording (Chapter 14).
D.3 Mic placement — close, on-axis, out of frame
Almost every dialogue-audio problem is a placement problem, and almost every fix is one of three rules.
| Rule | What it means | For a lav | For a boom |
|---|---|---|---|
| Close | near the mouth = the voice beats the room | 6–8" below the chin (built in) | 1–2 ft, a hand's width out of frame — as near as the shot allows |
| On-axis | the sensitive direction points at the source | center it on the sternum (omni forgives turns) | aim the barrel down at the mouth |
| Out of frame | never seen (unless it's a handheld) | hidden under a stable layer; cable dressed down | just above (or below) the frame line |
FIGURE D.2 — Why closer wins: the direct-to-reverberant ratio
Halving the mouth-to-mic distance dramatically lifts the direct voice over the room.
MIC FAR (across the room) MIC CLOSE (6–8", or a boom a hand off-frame)
( S )↑ · · · · · · · · · ((• ( S )↑((•
voice + FULL room echo the VOICE dominates; the room falls back
+ noise, roughly equal → clean, present, "talking TO you"
= "a recording OF a room" = "professional" — and nothing else changed
Two-mic safety: on important dialogue, record a lav and a boom at once — lav = the consistent backup, boom = the natural preferred sound. It covers a rustle, a dead battery, or a drifted aim, and gives the editor a choice (Chapter 14; drilled in the interview, Chapter 19).
On- vs off-camera: get the mic off the camera and onto the subject. When the off-camera mic records to its own device, that is double-system sound — sync it in the edit (Chapter 15; the edit-side sync is Chapter 27).
D.4 Levels and loudness reference
Two different jobs, two different scales. Capture levels (dBFS, on set) protect against clipping; delivery loudness (LUFS, in post) sets how loud the finished program plays back. Do not confuse them.
Capture — set gain by these (Chapters 14–15):
| Target | Value | Why |
|---|---|---|
| Dialogue average | ~-12 dBFS | well above the noise floor, comfortably below the ceiling |
| Loudest peak (laugh / shout) | under ~-6 dBFS, always below 0 |
0 dBFS = clipping = destroyed forever |
| Headroom to keep | 6–12 dB | the safety margin for the unexpected loud word |
| Room tone / quiet | ~ -40 to -50 dBFS | near-silence should read low; if it's high, the room is too noisy |
| Noise floor (system hiss) | ~ -60 dBFS | keep dialogue far above it |
| Auto gain (AGC) | OFF | it "pumps" — raises the noise floor in the pauses |
| Safety / backup track | on, ~-12 dB lower | insurance against a surprise clip on unrepeatable moments |
| Bit depth | 24-bit if available | forgiving of a conservative (quiet) level |
| Set the level using | the subject's real, loudest delivery | people get louder once "record" is rolling |
FIGURE D.3 — The dBFS capture meter: aim for -12, leave headroom
0 dBFS ┤▓▓ ← THE CEILING — never touch it (clipping is permanent)
-6 ┤██ ← the loudest peaks may flick up here, no higher
-12 ┤██████ ← AIM: normal dialogue lives here
-24 ┤███████
-45 ┤████████ ← room tone / quiet moments
-60 ┤█ ← the noise floor (system hiss)
-inf ┴──────────
The gap from your peaks up to 0 dBFS is HEADROOM. Too quiet is recoverable; too loud is not.
The asymmetry to memorize: a little too quiet is a nudge in the edit (you pay only in some hiss); a clipped peak is a re-shoot. When in doubt, record a touch low and protect the headroom.
Delivery — master by these (Chapter 33). Streaming platforms normalize loudness, so you master to a target, not past it. LUFS measures how loud the whole program feels; the peak ceiling (dBFS / dBTP true peak) guards against distortion.
| Destination | Integrated loudness | Peak ceiling |
|---|---|---|
| Streaming / social video | about -14 LUFS | -1 dBFS (true peak) |
| Podcast / spoken-word | about -16 LUFS | -1 dBFS |
| Broadcast — Europe (EBU R128) | -23 LUFS | -1 dBTP |
| Broadcast — North America (ATSC A/85) | -24 LKFS | -2 dBTP |
Hedge these numbers. The targets above are typical values at the time of writing. Platforms publish and periodically revise their loudness normalization, and targets differ by destination — always confirm the current number at the source before you deliver. The full delivery/codec specs and how to match them live in Appendix H (and Chapter 36). (
LKFSis the same measurement asLUFSunder another name.)
dBFS vs LUFS in one line: level dialogue by its peaks (dBFS) during the mix; master the finished program by its loudness (LUFS) at the end.
D.5 Location-sound checklist — run it every time
The discipline of Chapter 15, as a checklist. Run it the way a pilot runs a preflight.
Before you roll:
[ ] Headphones ON — closed-back, wired. Monitor the RECORDER'S output, not the room.
[ ] Kill every motor — HVAC, fans, fridge, dimmers, appliances. (Note anything to switch BACK on!)
[ ] Mic close, on-axis, out of frame (D.3). Lav centered & hidden, or boom a hand off the frame.
[ ] Set gain: dialogue ~-12 dBFS; test the LOUDEST line stays under ~-6 dBFS. AGC off.
[ ] Tame the room if it rings: get closer + hang blankets/coats behind the mic AND the subject.
[ ] Wind? foam windscreen (indoor / breath) · furry cover (outdoor) · your own body · a sock.
[ ] Rumble / handling? shock mount + high-pass (low-cut) at ~80 Hz.
[ ] Double-system? camera mic ON as a SCRATCH track; CLAP once, in frame, per take (→ sync, Ch.27).
Monitor for these — the meter shows level, not quality, so your ears are the only judge:
FIGURE D.4 — What to LISTEN for on headphones
[ ] Distortion / clipping ..... crackle or harsh edge on loud words → lower gain
[ ] Hum / buzz ................ steady electrical drone (mains, 50/60 Hz) → move cables, kill dimmers
[ ] HVAC / fridge / fans ...... a low whoosh or motor → turn it OFF
[ ] Wind / rumble ............. low roar or thumps → wind protection, high-pass
[ ] Clothing rustle (lav) ..... scratchy friction on movement → re-mount the lav
[ ] Plosives ................. "p"/"b" pops thumping the mic → angle it off-axis, add foam
[ ] Echo / room .............. hollow "bathroom" sound → get closer, deaden the room
[ ] Off-mic / swimming voice .. level rises & falls as they turn → re-aim, brief the subject
[ ] Intermittent hits ........ a plane, a truck, a phone buzz → pause, re-take
[ ] Dropouts (wireless) ...... momentary silence / static → batteries, range, interference
Before you wrap — the one habit beginners always skip:
[ ] ROOM TONE: record 30–60 s at EVERY setup — same mic, position, and levels; everyone still & silent.
It fills gaps, smooths cuts, patches edits, and becomes the noise print + ambience bed in post.
A borrowed room's tone will NOT match — every location gets its own.
Fix-on-set vs fixable-in-post: reverb and wind roar you must fix here — post cannot cleanly remove them; a steady hum you may leave for noise reduction. But a clipped peak, an echoey room, and un-recorded room tone can never be repaired — fix it in pre and on set, not in post.
D.6 Audio-post order of operations
Lock the picture first — every fix below is keyed to a frame that must not move (Chapter 33).
FIGURE D.5 — The audio-post pipeline: dialogue first, master last
LOCKED ORGANIZE DIALOGUE EQ & MUSIC SFX & FINAL MIX MASTER &
PICTURE → & SYNC → CLEANUP → LEVEL → BED → AMBIENCE → (balance) → DELIVER
stems: de-noise, cut mud, license room-tone dialogue ~-14 LUFS,
DIA/MUS/ de-hum, add & duck bed + on top; peak ≤ -1 dBFS;
SFX/AMB de-click presence under DIA spot SFX automate QC on 4 systems
Dialogue is the boss of the mix: if any layer makes a word harder to hear, that layer is too loud.
The dialogue chain (in order — cleanup, then tone, then dynamics):
| # | Move | Starting point | Why |
|---|---|---|---|
| 1 | Spot-fix transients | de-click / manual dips | kill isolated clicks, bumps, one-off plosives first |
| 2 | High-pass (low-cut) | ~80 Hz | remove rumble that carries no voice |
| 3 | De-hum | notch 50 / 60 Hz + harmonics | surgically remove electrical buzz |
| 4 | Broadband noise reduction | learn the print from room tone, reduce 8–10 dB | lower the noise floor without artifacts |
| 5 | EQ — subtract | -2 to -3 dB @ 200–400 Hz | clear the "mud" / boominess |
| 6 | EQ — boost presence | +2 to +4 dB @ 3–5 kHz | intelligibility; make the voice "cut through" |
| 7 | EQ — air | +1 to +2 dB @ ~10 kHz | an open, polished top end |
| 8 | De-ess | tame 5–8 kHz, only when "s"/"sh" spike | stop harsh sibilance |
| 9 | Compress | ratio ~2:1–3:1, 3–6 dB of reduction | even out the loud/quiet swing |
| 10 | Level / ride | automate the big swings by hand | transparency a compressor can't give |
Rule: cut problems before you boost, and reduce noise to just enough — overdone, the voice goes "underwater." Bypass the chain often to compare against the raw take.
Stem levels (starting points; the ratios between them are the real lesson):
| Stem | Typical level | Its job |
|---|---|---|
| Dialogue | peaks ~-12 dBFS | the star — always on top, always clear |
| SFX (spot) | accents ~-18 dBFS | punctuate on the exact frame of the action |
| Music bed | -24 to -30 dBFS | ducks under every line, lifts in the gaps |
| Ambience | -28 to -34 dBFS | the quietest layer — the "you are here" bed |
| Master bus | -14 LUFS integrated | true-peak limiter ceiling at -1 dBFS |
Music, briefly (Chapter 33; licensing in Chapter 38): match the arc, not just the genre; prefer one evolving theme to a playlist; leave room for the voice; and only ever use music you can prove you're licensed to use (royalty-free / production library, Creative Commons per its exact terms, commissioned original, or a platform's cleared library). Then duck it under the dialogue.
QC before you call it done: play it on four systems (good headphones, computer speakers, a phone speaker, and a car or TV), check it in mono, and watch it through once at final loudness. The phone-speaker pass is non-negotiable. And captions are not optional — they don't replace a clean mix, but a clean mix makes accurate captions (Chapter 36).
D.7 Troubleshooting — symptom → cause → fix
The field-and-timeline diagnostic. Set = fix it on the shoot; Post = fixable at the timeline.
| Symptom | Likely cause | Fix | Where |
|---|---|---|---|
| Hollow / echoey / "bathroom" | mic too far in a hard, reflective room | get closer; add absorption (blankets, coats, rugs); indoors use a supercardioid, not a shotgun | Set (post can't cleanly un-echo) |
| Wind roar / low rumble | moving air striking the diaphragm | foam windscreen (indoor) / furry cover (outdoor) / your body / a sock; high-pass ~80 Hz | Set |
| Steady hum / buzz | mains electricity — dimmers, ballasts, ground loop, bad cable | move cables, kill dimmers; in post, de-hum notch at 50/60 Hz + harmonics | Set + Post |
| Distortion / crackle on peaks | gain too hot — the signal clipped at 0 dBFS |
prevent: set to the loudest line, keep peaks under -6; a clipped take is a re-shoot | Set (no post fix) |
| Harsh, spitty "s" / "sh" | sibilance, often worsened by a presence boost | a de-esser on the 5–8 kHz band | Post |
| Muddy / boomy / "in a barrel" | low-mid build-up (common on a lav under clothing) | cut -2 to -3 dB @ 200–400 Hz | Post |
| Thin / dull / "behind glass" | no presence or air; or high-pass set too aggressively | boost 3–5 kHz (presence) and a touch of 10 kHz (air); ease the low-cut | Post |
| Clothing rustle (lav) | fabric rubbing the capsule; loose cable | re-mount away from fabric; add a strain-relief cable loop; tape it down | Set (a rustle sits inside the word) |
| "Pumping" / breathing noise floor | auto gain (AGC) riding levels in the pauses | switch to manual gain | Set |
| Level swims up and down | subject drifting off-axis on a directional mic as they turn | re-aim / center the mic; brief the subject; or use an omni lav | Set |
| Wireless dropouts | dead batteries, out of range, radio interference | fresh batteries, closer receiver, change the channel | Set |
| Voice too quiet, buried in hiss | gain set too low, then boosted in post | next time aim for -12; some rescue via careful gain + noise reduction | Set (mostly) |
| Background "blinks" at every cut | no room tone laid under the edit | lay a room-tone bed across the whole scene | Post — only if you recorded tone (Ch.15) |
| Lips out of sync with the voice | double-system drift / no sync reference | line up the clap spike; keep a scratch track rolling | Post (Ch.27) |
| Whole video too quiet / loud on the platform | not mastered to the platform's loudness | master integrated to target (~-14 LUFS streaming), peak ≤ -1 dBFS | Post (Ch.33 / Appendix H) |
D.8 Numbers to memorize
The short list, in one place.
| Number | Meaning |
|---|---|
| -12 dBFS | dialogue average on capture (Chapters 14–15) |
| -6 dBFS | keep the loudest peaks under this |
| 0 dBFS | the ceiling — clipping; never touch it |
| 6–12 dB | headroom to protect |
| ~80 Hz | high-pass / low-cut for rumble |
| 50 / 60 Hz | mains hum to notch out (de-hum) |
| 3–5 kHz | presence boost for intelligibility |
| 8–10 dB | a modest, safe noise-reduction amount |
| 30–60 s | room tone at every setup |
| ~-14 LUFS | streaming master target (confirm at source → Appendix H) |
| -1 dBFS | true-peak ceiling for delivery |
The one sentence: capture clean dialogue close, on-axis, and out of frame at ~-12 dBFS with headroom, monitor on headphones, and record room tone every time — then clean, EQ, level, score, design, and master to your platform's loudness — because sound is half the picture.