49 min read

> "We don't hear a scene. We hear a hundred decisions about a scene."

Prerequisites

  • 14

Learning Objectives

  • Set up a double-system recording and mark it with a sync clap the editor can align.
  • Operate a boom overhead and from below, alone or with a partner, without dipping into frame or casting a shadow.
  • Monitor a recording on closed-back headphones and name the specific faults you are listening for.
  • Set audio levels to average around -12 dBFS while protecting at least 6 dB of headroom below clipping.
  • Tame reverb, wind, and handling noise on location using proximity, absorption, and wind protection.
  • Record 30–60 seconds of matched room tone at every setup and explain why the edit cannot live without it.

Chapter 15: Audio for Video II — Location Recording, Monitoring, Booming, and Capturing Clean Dialogue

"We don't hear a scene. We hear a hundred decisions about a scene." — a working sound recordist's maxim, widely repeated on set

Overview

Here is a shoot that looks like a success and is a disaster. The light is beautiful — a soft window key, a warm practical glowing behind the subject, everything you learned in Chapters 11 through 13 put to work. The framing is clean, the focus is sharp, the performance is real. You wrap, you drive home happy, you drop the card into your editor that night. And then you put on headphones for the first time all day, and you hear it: a refrigerator hum that never stops, a hollow bathroom-tile echo around every word, one clipped syllable where the subject laughed too loud, and — worst of all — a hard silence between every take, because you never recorded the sound of the room. The picture is gorgeous. The audio is unusable. And there is no version of "we'll fix it in post" that saves it, because you cannot un-clip a peak, you cannot un-echo a room, and you cannot invent the ambient sound you never captured.

That gap — between a shoot that looked professional and a recording that is professional — is the entire subject of this chapter. In Chapter 14 you learned what a microphone is, how a lav and a shotgun differ, what a polar pattern does, why placement close and on-axis matters, and how to set gain for a clean signal. You have the tools. This chapter is about the discipline of using them in the real world, where the world does not cooperate: where air conditioners drone, where tile bounces sound like a racquetball court, where wind turns a lovely shotgun into a bag of thumps, and where the loudest word of the interview arrives without warning. Location sound is not a gadget you buy. It is a set of habits you run every single time, the way a pilot runs a checklist — and the habits are the whole difference between sound an audience forgets (which is the goal) and sound an audience flees.

This is where the book's second throughline stops being a slogan and becomes a daily practice: sound is half the picture. We gave audio two full chapters, equal weight with light, because audiences abandon bad sound faster than anything else you can do to them — and because so few beginners take it seriously that simply doing so vaults you past people who have been shooting for years. By the end of this chapter your Project 1 talking-head will be fully in the can — clean dialogue, protected headroom, and thirty seconds of room tone — ready to assemble the moment you reach the edit.

In this chapter you will learn to:

  • Record double-system sound — picture on the camera, audio on a separate device — and sync the two with a simple clap.
  • Boom a scene from above and below, solo or with a partner, keeping the mic just out of frame and out of the light.
  • Practice real monitoring: listening through closed headphones to what the mic is actually capturing, not what the room sounds like to your ear.
  • Set levels and protect headroom so nothing clips and nothing drowns in the noise floor.
  • Apply wind protection and tame a bad room, and handle the hums, rustles, and rumbles that ruin location dialogue.
  • Record room tone at every setup — the single most-skipped, most-needed thirty seconds on any shoot.

Learning Paths

This chapter is a discipline, and every reader needs the discipline — but here is where your attention pays off most.

  • 📱 Phone-first: §15.1 shows you that your phone is a second recorder; §15.5 gives you the sock-over-the-mic wind fix and the blanket-fort room; §15.6 costs nothing and will save every edit you ever make. These three sections are your whole toolkit.
  • 🎥 Creator: §15.3 (monitoring) and §15.4 (levels) are the reason your voice will suddenly sound "pro" — most creator audio dies on these two. §15.6 is your secret weapon for clean-sounding cuts.
  • 💼 Pro-track: §15.1 (double-system + sync) and §15.2 (booming with a partner) are the craft that separates a videographer from a filmmaker. Learn the boom as an athletic skill.
  • 🎓 Student: read all six. §15.6 (room tone) is the discipline instructors test for and beginners always forget; §15.4 is where the numbers (dBFS, headroom) get graded.

15.1 Double-system sound and syncing

Start with a choice you make before you point anything: where does the sound get recorded? You have two options, and their names are worth knowing because everyone on a set uses them.

Single-system sound records the audio inside the camera, married to the picture in the same file, the way your phone does by default. It is simple: one device, one file, nothing to line up later. Double-system sound records the picture on the camera and the sound on a separate device — a dedicated audio recorder, a mixer, or, yes, a second phone — to be married to the picture later in the edit. That is the definition, and the "double" is the point: two independent machines, two files, deliberately reunited in post.

Why would anyone choose the more complicated path? Because separating the sound from the camera buys you three things you cannot get otherwise. First, better sound electronics: a dedicated recorder has cleaner preamps (the circuit that boosts a mic's tiny signal to a usable level — see Chapter 14, §14.5) than most cameras, so the same mic sounds warmer and quieter. Second, independence: the mic no longer has to live near the camera. A recorder in a sound person's bag can follow a boom anywhere in the room while the camera stays on its tripod across the space. Third, redundancy: two machines mean two chances, and a dead camera battery no longer costs you the audio.

The cost of that freedom is one job you now owe the editor: the two files have to be lined up. Aligning separately-recorded audio to its picture so that lips and sound match again is called sync — the single most important word in this section. Get sync wrong and the whole illusion collapses; a mouth a few frames out of step with its voice reads as wrong to a viewer even when they cannot say why. So the entire discipline of double-system sound comes down to making sync easy, and you make it easy on set, in one second, with a clap.

Here is the trick, and it is beautiful in its simplicity. A sharp hand clap (or a clapperboard — the sound-and-picture marker you will formalize as the slate in Chapter 18) makes a sudden, loud transient — a tall, narrow spike — that lands at the same instant in both recordings: the camera's reference audio and the recorder's clean audio. In the edit you slide one track until the two spikes stack on top of each other, and from that frame forward, everything is in sync. That is the whole method.

FIGURE 15.1 — Two systems, married in the edit (signal flow)

  SOUND SYSTEM                                   PICTURE SYSTEM
  ( S ) mouth                                    ( S ) framed in the shot
     |                                              |
     v                                              v
   ((•  mic (boom or lav, Ch.14)                 [ CAM ] records picture
     |   balanced cable / wireless                  |   + a SCRATCH track from its
     v                                              v    own mic (reference only)
  [ RECORDER ]  clean preamp, gain set          [ picture file + scratch audio ]
     |   monitored on headphones
     v
  [ sound file: 24-bit WAV ]
              \                                    /
               \____________ the CLAP ___________/
                    one sharp spike in BOTH files
                                |
                                v
             EDIT: line up the spikes  ->  SYNC  (Chapter 27, §27.4)

Notice the small, non-negotiable detail in that diagram: the camera still records a scratch track — a rough reference recording from its own built-in mic. You are not using that audio in the final piece; it sounds thin and distant, exactly as Chapter 14 warned. You are recording it as reference, for two reasons. It carries the clap spike the editor lines up against, and it is your backup if the good recorder fails. The rule is iron: never shoot double-system without a scratch track. A double-system shoot with no reference audio and no clap is a pile of picture and a pile of sound with no way to marry them — a genuine post-production nightmare that hours of manual nudging may never fully fix.

FIGURE 15.2 — The clap: one spike in two files, then lined up

  CAMERA scratch   ~~~~~~~ |# ~~~~~~~~~~~~~~~    <- the clap = a tall, narrow spike
  RECORDER sound   ~~~~~~~~~~~~ |# ~~~~~~~~~~~    <- same clap, offset in time
                                |
                                |  slide the recorder track left until the spikes stack:
                                v
  CAMERA scratch   ~~~~~~~ |# ~~~~~~~~~~~~~~~
  RECORDER sound   ~~~~~~~ |# ~~~~~~~~~~~~~~~     <- now sound and picture are in SYNC

The workflow on set is a small ritual you run before every take, and it takes three seconds. Point the camera, roll the recorder, and — with both running — hold a hand (or the clapper) up in frame, say the take number out loud so the audio names itself ("scene two, take three"), and clap once, sharply, where the camera can see it. Then act. Do it even when you are shooting alone: reach into your own frame and clap. Future-you, at the timeline at midnight, will thank present-you every time.

🎒 Gear Note: your phone is a second recorder. You do not need to buy a dedicated recorder to shoot double-system. Clip a lav to your subject, run it into a second phone in their pocket set to record a voice memo, and let your camera (or first phone) shoot the picture with its own mic as the scratch. Clap once at the start. You now have a clean, close, independent audio track and a reference track, which is exactly what a $600 recorder gives you — the principle is identical, the price is zero. The only rule is the same as the pros': always get the clap.

🔬 The Tech: timecode, and why big productions don't clap. On larger sets, cameras and recorders share timecode — a running clock (HH:MM:SS:FF, hours:minutes:seconds:frames) stamped identically into every file so software can align them automatically, no clap needed. Devices are "jam-synced" from one master clock at the start of the day and drift is periodically re-synced. It is faster across dozens of takes and multiple cameras, and it is completely unnecessary for you right now: a clap and a waveform do the same job for free. Skip this box entirely if you like — the practitioner path never depends on timecode. File it under "things that exist for the day you shoot a hundred takes."

✂️ In the Edit. This whole section is a favor you do the editor (usually you). In Chapter 27 (§27.4) you will drop the picture and the recorder audio onto a timeline, and either line up the two clap spikes by eye or let the software auto-sync them by matching the waveforms — a one-click operation if you gave it a scratch track and a clean transient to match. Skip the clap and you sentence yourself to nudging clips a frame at a time until the lips look right. Ten seconds of discipline on set buys an hour of ease in post. That trade — cheap on set, expensive later — is the whole logic of you shoot for the edit.

⚠️ Common Mistake: the silent double-system shoot. The classic beginner failure is to turn off the camera's mic ("I have a good recorder, I don't need it") and skip the clap ("I'll just eyeball it later"). Now the camera files have no audio at all, so there is nothing to sync against and no backup — and every clip has to be aligned by watching lips, one frame at a time, with no guarantee it will ever sit perfectly. Always leave the camera mic on as a scratch. Always clap. There are no exceptions worth the risk.

🔄 Check Your Eye. 1. In one sentence each, distinguish single-system from double-system sound. 2. Why do you still record the camera's own (bad-sounding) mic on a double-system shoot? 3. A clap works for sync because it creates what, exactly, in both recordings?

Check yourself

  1. Single-system records audio inside the camera, in the same file as the picture; double-system records audio on a separate device, to be re-married to the picture in the edit.
  2. As a scratch/reference track — it carries the sync clap the editor lines up against, and it is a backup if the good recorder fails. You never use it in the final mix.
  3. A sharp, sudden transient — a tall, narrow spike — that lands at the same instant in both files, so you can slide one until the spikes stack and the two are in sync.

15.2 Booming technique and working with a partner

Chapter 14 told you what a boom is — a microphone (usually a shotgun) on the end of a long pole, held out of frame to get the mic close to the subject without seeing it. This section is about operating it, because a boom is not a mount, it is an instrument, and it is played by a person. Boom operating is genuinely athletic — arms up, still, for minutes at a time — and it is a skill you can start building today with a broom handle and a rubber band.

The default and best position is overhead: the mic hangs from above, just out of the top of the frame, angled down at the subject's mouth. There are two reasons this is the go-to. First, aiming down at the mouth puts the voice on-axis (Chapter 14, §14.3) while the mic's rejection points at the floor, so foot shuffles, chair creaks, and floor bounce fall off the back of the pattern. Second, coming from above captures the voice from over the chest, which sounds natural — it is roughly how we're used to hearing people, and it avoids the chesty boom of a mic pointed up from below.

FIGURE 15.3 — The overhead boom, just out of frame (side view)

     boom op: arms up, elbows in, braced
          \
           \______ pole ______
                              \
                               ((•  <- mic angled DOWN at the mouth, on-axis
   top of frame - - - - - - - - -\- - - - - - - - - - - - - - - - -
   (the frame edge)               \   ~one hand's width above the frame line
                                   (o)  ( S ) speaking
                                   /|\
   ------------------------------ floor ------------------------------
   Aim the capsule at the mouth from above; keep it a hand out of the top of frame.
   Overhead points the mic's DEAD side at the floor -> foot & floor noise rejected.

You go from below — mic pointing up at the mouth from just under the frame — only when overhead is impossible: a very wide shot where an overhead boom cannot get close enough without dropping into frame, a low ceiling that a top boom would bounce off, or a lighting setup where an overhead boom would throw a shadow across the subject's face. Booming from below is a legitimate tool, but it comes with costs: it tends to pick up more chest resonance and less crisp diction, and now your mic's rejection points at the ceiling, so if there is an air vent or a ceiling fan up there, you will hear it. Choose overhead unless something forces you low; then commit to low cleanly.

Wherever you boom from, three fundamentals hold. Aim at the mouth, not the chest, not the forehead — the capsule's on-axis line should point right at the source of the words. Get as close as the frame allows — closer is always better for the ratio of voice to room (that is the inverse-square law from Chapter 14 doing its work), so you ride the mic right at the edge of the frame, a hand's width out, no more. And hold it steady — a wandering mic makes the voice swim in level and tone as it moves on- and off-axis. The grip that makes steadiness possible: hands apart along the pole for leverage, arms up and braced (elbows tucked toward your ribs, not winged out), and, when you can, a hip or a wall to plant against. It burns. That is normal. Boom operators train for it like an isometric hold, and so should you.

Booming a two-person exchange adds one move: you favor whoever is speaking. Rather than parking the mic dead-center between two people (which leaves both of them slightly off and slightly distant), you gently swing or rotate the pole so the on-axis line lands on the current talker, then lead the next line by anticipating who speaks next. Good boom ops read the conversation a beat ahead, the way a good camera operator anticipates a move. You are, in effect, editing with the microphone in real time — pointing the audience's ear at the voice that matters.

FIGURE 15.4 — Booming a two-person exchange: favor the talker (top-down)

     ( A ) speaking now  ●                    ●  ( B ) speaks next
                          \                  /
                           \    ((•        /
                            \   / |  \_____/  <- swing the capsule to lead
                             \ /  |   the incoming line
                          [ CAM ] frame edge
   Aim on-axis at the current talker; anticipate and swing to B a beat before B speaks.
   Dead-center between them leaves BOTH off-axis and distant — favor, don't split.

Now the partnership, because a boomed scene is usually two people cooperating: whoever holds the boom and whoever runs the camera. The two of you need a tiny shared language, spoken quietly between takes and sometimes during them. The boom op cannot see the frame; the camera op can. So the camera op is the boom op's eyes: "you're clear," "a hair lower," "you dipped in on that last line," "watch the shadow." Before a take, the boom op dips the mic slowly into frame until the camera op says "there" — that is the floor, the lowest the mic can go — and then holds a hand's width above it. During a take, the camera op watches the top of the frame like a hawk, because a boom that drifts down into the shot ruins the take as surely as an actor forgetting a line.

And that word shadow is not incidental — it is where this chapter shakes hands with Part III. A boom is a physical object between your lights and your subject, and if it crosses a hard key light (Chapter 11) it will throw a boom shadow — a telltale bar sliding across the wall or the subject's face that instantly reads as amateur. The fixes are all spatial: keep the boom out of the key's throw, work it from the shadow side of the subject, angle the pole so its shadow falls out of frame, or soften the key until the shadow melts. This is why sound and light are taught in the same part of this book: on a real set, the person booming and the person lighting are constantly negotiating the same few feet of air above the subject's head.

💡 Why It Works: proximity beats everything. A boom sounds good for the same physical reason a lav does — it gets close. Every time you halve the distance from mouth to mic, the direct voice gets dramatically louder relative to the room's reverberation and noise, because the voice follows the inverse-square law while the room's reflected sound is roughly even everywhere. A shotgun's tight pattern (Chapter 14, §14.3) then rejects whatever is off to the sides. Close plus directional is why a boom a foot above someone's head, in a noisy room, can sound like a quiet studio. The mic did not silence the room; the geometry did.

🎬 On Set: boom thirty seconds of dialogue. Grab a partner and a boom (a real pole, or a broom handle with your shotgun or lav rubber-banded to the end, or even a phone-plus-lav gaffer-taped on). Shoot a 30-second two-person exchange with one of you booming overhead and one running the camera. Constraints: mic a hand's width out of the top of frame the whole time, favor whoever is speaking, and no boom shadow. Self-review question: play it back on headphones — does the voice stay even in level and tone as you swing between speakers, or does it swim and drift? If it swims, you were moving the capsule off the mouths; do it again and aim more precisely. If you are solo, clamp the boom into a light stand angled over the subject and treat the stand as your partner.

🔄 Check Your Eye. 1. Why is overhead the default boom position — what does aiming down at the mouth reject? 2. Name two situations that force you to boom from below instead. 3. What is a boom shadow, and give two ways to kill it without moving the subject.

Check yourself

  1. Aiming down puts the voice on-axis while pointing the mic's dead side at the floor, rejecting foot noise, chair creaks, and floor bounce; it also captures the voice naturally from above the chest.
  2. Any two of: a very wide shot where overhead can't get close without entering frame; a low ceiling a top boom would bounce off; a hard key light where an overhead boom would throw a shadow on the face.
  3. A boom shadow is a bar of shadow the pole/mic casts when it crosses a (usually hard) light. Kill it by working from the subject's shadow side, keeping the boom out of the key's throw, angling the pole so its shadow falls out of frame, or softening the key.

15.3 Monitoring: trust your ears, not the meters alone

Here is a scene that plays out on beginner shoots constantly. The meters looked perfect all day — bars bouncing in a healthy range, never pinning, never dead. Everyone felt good. Then, in the edit, the recording turns out to be riddled with a buzz, or an off-camera phone's vibration, or the sub-audible thud of someone bumping the mic stand, or the sizzle of a wireless dropout — none of which the meters flagged, because a meter measures loudness, not quality. A signal can be perfectly, healthily loud and completely ruined. The only instrument that catches ruin is a human ear, actively listening. That active listening is called monitoring, and it is the habit that most separates people whose location sound is reliable from people who get unlucky.

Let us define it precisely. Monitoring is listening to the recorded signal — the actual output of your recorder or camera, through closed-back headphones, while you record — so that you hear what is being captured to the file, not what the room sounds like to your naked ear. Every word of that matters. Closed-back headphones, because they seal out the room and force you to hear only the mic's world. The recorder's output, not the mic directly, so you are checking the entire signal chain all the way to the file (this is sometimes called "confidence monitoring" — you are confirming the recording is real and clean, not just that the mic is alive). And while you record, because a problem heard live can be fixed in the next take; a problem discovered in the edit cannot be fixed at all.

The reason monitoring feels like a superpower is that your brain is an extraordinary noise filter and the microphone is not. Sitting in a room, you effortlessly ignore the air conditioner, the traffic, the fluorescent hum — your auditory system tunes them out so you can focus on a voice. The mic has no such filter. It captures the drone at full strength, sitting right on top of the dialogue. Put on sealed headphones and you become the microphone: suddenly you hear the fridge you had stopped noticing, the echo you were ignoring, the highway two blocks away. That is not the headphones exaggerating. That is the truth of the recording, which your unaided ears were politely hiding from you. The job of monitoring is to hear like the mic hears, so nothing surprises you later.

So what are you actually listening for? A named checklist, because "listen carefully" is useless advice and "listen for these nine things" is a skill. Run your ear down this list on every setup:

FIGURE 15.5 — The monitoring checklist: what to listen FOR (not just watch on the meter)

  [ ] Distortion / clipping ...... crackle or harsh edge on loud words -> lower gain (§15.4)
  [ ] Hum / buzz ................. steady electrical drone (mains at 50/60 Hz) -> move cables, kill dimmers
  [ ] HVAC / fridge / fans ....... a low continuous whoosh or motor -> turn it OFF (note to turn back on!)
  [ ] Wind / rumble .............. low roar or thumps -> wind protection, high-pass filter (§15.5)
  [ ] Clothing rustle (lav) ...... scratchy friction on movement -> re-mount the lav (§15.5)
  [ ] Plosives .................. "p"/"b" pops thumping the mic -> angle mic off-axis, add foam
  [ ] Echo / room ............... hollow, distant, "bathroom" sound -> get closer, deaden the room (§15.5)
  [ ] Off-mic / swimming voice ... level rises & falls as they turn -> re-aim mic, brief the subject
  [ ] Intermittent hits ......... a plane, a truck, a phone buzz, a fridge kicking on -> pause, re-take
  [ ] Dropouts (wireless) ....... momentary silence/static -> check batteries, range, interference

Two operational habits make monitoring real. First, listen-back before you commit. Before the take that counts, roll twenty seconds, have the subject talk, stop, and play it back on the headphones. Live monitoring catches most things; a playback catches the rest, and it confirms the file actually recorded (the number of "great takes" that turn out to be silent because Record was never armed is not zero). Second, keep the cans on during the take. The temptation is to set levels, feel satisfied, and take the headphones off to do something else. Don't. The plane that flies over on take four, the phone that buzzes in someone's pocket, the moment the subject turns away and goes off-mic — you only catch those live if you are listening live.

🎒 Gear Note: closed-back headphones are the cheapest upgrade you own. The tool here is not exotic: any closed-back headphones (the kind that seal around or onto your ears and block outside sound) will do the job, and even wired closed earbuds shoved firmly in beat nothing at all. What you are avoiding is open headphones and loose earbuds that let the room leak in — because then you are hearing the room plus the recording and cannot tell what is on the file. Wired beats wireless for monitoring: Bluetooth adds a delay that makes lips and sound feel out of step and can mask problems. The whole principle: seal out the room so you hear only what the mic hears. A $15 pair used religiously outperforms a $300 pair left in the bag.

♿ Accessibility & Inclusion. Clean, intelligible dialogue is an accessibility issue before it is an aesthetic one. Muddy, echoey, noise-buried audio is hardest on exactly the people who most need clarity — viewers who are hard of hearing, viewers watching in a second language, viewers in a noisy environment. Monitoring for intelligibility is therefore part of making the video reachable, not just polished. And it does not replace captions: even flawless audio needs accurate captions and subtitles for the roughly one in five viewers watching with sound off and for every deaf or hard-of-hearing viewer. Capture clean sound and caption it — the two serve overlapping but different audiences.

⚠️ Common Mistake: trusting the meter, skipping the ears. Beginners watch the bars and feel safe. But the meter is blind to buzz, echo, rustle, plosives, off-mic drift, and intermittent noise — all of which can sit inside a perfectly "healthy" level. The meter tells you how loud; only your ears tell you how clean. Use both, and when they disagree, believe your ears.

🔄 Check Your Eye. 1. What does a meter measure, and name three problems it cannot catch that your ears can. 2. Why do sealed headphones reveal noise you did not notice while standing in the same room? 3. What is a "listen-back," and what two things does it confirm?

Check yourself

  1. A meter measures loudness/level, not quality. It cannot catch (any three) hum/buzz, HVAC or fridge noise, room echo, clothing rustle, plosives, off-mic drift, wireless dropouts, or intermittent hits like a passing plane.
  2. Because your brain filters ambient noise so you can focus on a voice, but the mic does not filter — sealed headphones make you hear like the mic, exposing the drone/echo your ears were tuning out.
  3. Recording ~20 seconds and playing it back on headphones before the real take. It confirms the sound is clean (no hidden faults) and that the file actually recorded (Record was armed and the signal is there).

15.4 Setting levels and protecting headroom

Chapter 14 taught you to set gain for a clean signal — enough boost to lift the voice well above the electronic hiss, not so much that it distorts. This section sharpens that into a specific target and a specific safety margin, because location dialogue lives or dies on two numbers. First, the vocabulary, measured the way this book measures all digital audio: in dBFS, decibels relative to full scale, where 0 dBFS is the absolute ceiling — the single loudest value a digital recorder can encode.

Levels are how loud your recorded signal is, read on the meter in dBFS. Because 0 dBFS is a hard wall, everything you record sits below it, as a negative number: normal speech might average -12 dBFS, a breath might sit at -40 dBFS, the noise floor might hiss down at -60 dBFS. The higher the number toward zero, the louder. Headroom is the gap, measured in dB, between your loudest peaks and that 0 dBFS ceiling — your safety margin against the sudden loud moment you did not see coming. If your peaks hit -6 dBFS, you have 6 dB of headroom. If they hit 0, you have none, and the next surprise clips.

And clipping is the thing you are protecting against, because it is the one audio disaster with no cure. When a signal tries to go above 0 dBFS, the recorder cannot represent it, so it simply flattens the top off the wave — squares it into a harsh, crackling buzz. That distortion is baked into the file permanently. Chapter 33's cleanup tools can reduce steady hiss and hum; nothing can rebuild a peak that was chopped off at the ceiling. This is why, in digital audio, you always err toward more headroom: a clean voice recorded a little quiet at -18 dBFS is trivially raised in post with no penalty, but a voice that clipped at 0 is dead on arrival. Quiet is a nudge in the edit; clipped is a re-shoot.

So here is the target, and it is a range, not a magic number. Set your levels so normal dialogue averages around -12 dBFS, with the loudest peaks staying under about -6 dBFS. That keeps the voice comfortably above the noise floor (so it is clean when you bring it up) while leaving roughly 6 to 12 dB of headroom (so a laugh, a raised voice, or an emphatic word does not hit the ceiling). Picture the meter as a road with a wall at the top:

FIGURE 15.6 — The dBFS meter: aim low, leave headroom (one channel)

    0 dBFS |##|   <- THE CEILING. Touch it and you CLIP. No fix in post.        ^
   -3      |##|      _                                                          |  DANGER
   -6      |##|     | peaks may flick up here on the loudest word — no higher   |
   -9      |######| _|                                                        --+-- HEADROOM
  -12      |##########|  <- TARGET: normal dialogue averages here               |  (~6–12 dB)
  -18      |##########|                                                         |
  -24      |######|                                                             v
  -40      |###|   <- breaths, quiet moments, room tone live down here
  -60      |#|     <- the NOISE FLOOR (system hiss). Keep dialogue FAR above it.
           +-------------------------------------------------
   Set gain so speech lives around -12 dBFS and the loudest peak stays under -6.
   The gap from your peaks up to 0 dBFS is HEADROOM — your safety margin.

The method to hit that target has one step beginners skip: test the loudest thing. Set gain so ordinary talking sits near -12 dBFS, then ask the subject to deliver their loudest likely line — the laugh, the punchline, the emphatic sentence — at full volume, and watch that it peaks under -6 without kissing the ceiling. If it clips, back the gain off and re-check. You are not setting levels for the average; you are setting them so the peaks survive. The average takes care of itself.

Two more tools protect you when the world is unpredictable. A safety track (or backup track) is a second recording of the same mic set a good deal quieter — often around -12 dB lower than the main. If a surprise shout clips the main track, the safety track, recording lower, captured it clean, and the editor swaps to it for that moment. Many dedicated recorders can do this dual-record automatically; it is the professional insurance policy for unrepeatable moments like events and live speeches. And you must turn off automatic gain (often labeled AGC, automatic gain control). Auto-gain constantly re-adjusts loudness on its own, which sounds convenient and is a trap: in the pauses between sentences it cranks the gain up hunting for signal, dragging the room noise and hiss up into an audible "breathing" swell, then slams it back down when the voice returns. Set your gain manually, deliberately, and leave it.

Finally, a judgment call: set-and-forget versus riding the fader. For a controlled situation — a seated interview, a talking-head — you set levels once for that subject and that mic, and leave them alone; consistency is the goal and fiddling only introduces jumps. For an unpredictable situation — an event, a moving subject, a speaker who whispers then shouts — you may need to gently ride the level, easing gain down as they get loud and up as they get quiet, like a hand on a dimmer. Riding well is a subtle skill; when in doubt on a controlled shoot, set conservatively (a touch quiet, plenty of headroom) and don't touch it.

⚙️ Settings Box — location dialogue levels (a starting point to adjust, not a recipe).

Setting Starting point Why
Dialogue average ~-12 dBFS Well above the noise floor, comfortably below clipping.
Peak ceiling under -6 dBFS Leaves ~6 dB of headroom for the unexpected loud word.
Headroom to keep 6–12 dB Digital 0 dBFS is a hard wall; quiet is fixable, clipped is not.
Auto gain (AGC) OFF Prevents the noise floor "pumping" up in the pauses.
Safety / backup track ON, ~-12 dB lower Insurance against a surprise clip on unrepeatable moments.
Bit depth 24-bit if available Deeper quiet detail and more forgiving of a conservative level.
Riding vs set-and-forget set-and-forget for interviews; ride for events Consistency where you can control it; adapt where you can't.

🔬 The Tech: bit depth, and the 32-bit-float promise. The reason a quiet-but-clean recording rescues so easily is bit depth — the number of loudness steps the system can store (Chapter 3 introduced this for picture; audio is the same idea). Recording in 24-bit gives you a vast quiet range, so a signal sitting at -18 dBFS still holds plenty of detail when you raise it. A newer format, 32-bit float, is widely marketed as making clipping nearly impossible — its enormous numeric range means even a signal that appears to hit the ceiling can often be pulled back down intact in post. It is a genuine advance where it is available, but treat it as a safety net, not a license to stop caring: good levels, good headroom, and monitoring remain the discipline. Skip this box freely; the -12 dBFS/headroom habit works on any recorder ever made.

⚠️ Common Mistake: recording "hot to be safe." Beginners often push levels high — up near -3 dBFS — reasoning that a loud, strong signal is a good signal. It is the opposite of safe. A hot level has almost no headroom, so the first unexpected laugh or raised voice clips and is destroyed forever. The safe direction in digital is down: aim for -12, keep your 6+ dB of headroom, and remember that a quiet clean take is a five-second fix while a clipped take is a re-shoot. When unsure, record a little too quiet, never a little too loud.

🔄 Check Your Eye. 1. What is 0 dBFS, and what happens to a signal that tries to exceed it? 2. State the target average and peak ceiling for location dialogue, and how much headroom that leaves. 3. Why should automatic gain (AGC) be turned off for dialogue?

Check yourself

  1. 0 dBFS is the absolute ceiling — the loudest value a digital recorder can encode. A signal that exceeds it clips: the top of the wave is chopped off into permanent, uncorrectable distortion.
  2. Average around -12 dBFS, peaks under about -6 dBFS, leaving roughly 6–12 dB of headroom.
  3. Because AGC cranks the gain up in the silent pauses hunting for signal, pulling the room noise/hiss up into an audible "pumping" swell, then drops it when the voice returns — inconsistent and noisy. Set gain manually and leave it.

15.5 Taming the room, wind, and handling noise

Two rooms, one microphone, one voice. Room one is a kitchen: tile floor, bare walls, a window, hard cabinets — say a sentence and it rings, hollow and distant, like you are talking in a stairwell. Room two is a carpeted bedroom with a bed, curtains, and a closet full of clothes — the same sentence, same mic, comes back close, warm, and dry, like the voice is right next to you. Nothing changed but the surfaces. That difference is the whole subject of this section, and learning to hear it and fix it is what turns "I recorded some audio" into "I recorded this room's dialogue on purpose."

Give the discipline its name. Location sound is the craft of capturing usable audio in real, uncontrolled spaces — the kitchen, the café, the office, the street — as opposed to a treated studio built to sound neutral. On location you do not get to silence the world or rebuild the walls. You get to manage what is there: choose the least-bad spot, kill the noise you can, get close enough that the voice wins, and shape the surfaces you can reach. It is problem-solving under constraint, and the constraints are exactly why it is a skill.

The first enemy is reverberation — the echo of a voice bouncing off hard, parallel surfaces (tile, glass, drywall, bare floors) and arriving at the mic a fraction of a second after the direct voice, smeared and hollow. Reverb is the single most common thing that makes amateur audio sound amateur, and — crucially — it is the thing post-production is worst at removing. You can reduce steady hum in the edit; you cannot cleanly un-echo a voice, because the echo is tangled into every word. So you fix reverb on set, two ways at once. Get closer: proximity is your strongest weapon, because the direct voice obeys the inverse-square law and gets much louder as the mic approaches, while the room's reflected sound stays roughly even — so closing from three feet to one foot can bury the reverb under the voice. And add absorption: drape soft, thick, non-parallel things into the space. Blankets, moving pads, coats, a mattress against a wall, cushions, curtains drawn, a rug on a hard floor — even bodies. The move that pays off most is hanging a blanket on the hard wall behind the mic and another behind the subject, breaking the parallel bounce that makes a room ring.

FIGURE 15.7 — Deadening a hard room with what's already in it (top-down)

     ######## blanket / moving pad on the hard wall behind subject ########
                              ( S ) ---> eyeline
                               ((•  mic CLOSE (one foot) — the room can't compete
     # blanket #                                       # open closet of coats #
     # on the  #            [ CAM ]                     #  (a free bass trap)  #
     # parallel#                                        #                      #
     #  wall   #     -- rug on the hard floor --  -- curtains drawn --
     ###########       -- cushions, a mattress, bodies: all absorb --
   Soft + non-parallel + CLOSE. Absorption plus proximity beats reverb every time.
   You cannot cleanly un-echo a voice in post — so kill the echo here, on set.

The second enemy is wind, and its fix has its own name. Wind protection is shielding the microphone from moving air — the low-frequency roar and the sudden thumps that air makes when it strikes the mic's diaphragm directly. Even a light indoor draft or the puff of your own breath (a plosive "p") can overload a bare mic. The tools scale with the threat: a foam windscreen (the little sponge cap) handles indoor air, breath, and light breeze; a furry cover — the "dead cat" or windjammer — is for real outdoor wind, its fibers breaking up the airflow before it reaches the capsule. And the free tool is position: put your own body between the wind and the mic, angle the mic away from the gust, tuck into the lee of a wall or a car. On a phone with no accessories, a literal sock or a bit of foam pulled over the mic will take the edge off a breeze — inelegant, effective, and always available.

The third enemy is the family of self-inflicted low-frequency noises: handling noise (the thumps and rumbles that travel up the pole or cable when you grip, shift, or knock the mic) and rumble (traffic, HVAC, footfalls transmitted through the floor). The fixes: a shock mount — the elastic suspension that floats the mic in its cradle so vibration cannot travel into it — and simple discipline (don't touch the mic during a take, give the cable a little slack so tugs don't transmit, brace against something solid). And there is one filter worth knowing on set: the high-pass filter (also called a low-cut), a switch on many mics and recorders that rolls off the very low frequencies — typically below around 80 Hz — where rumble and handling noise live but the human voice mostly does not. Flip it on for a boom outdoors or near traffic and the rumble drops away while the voice stays intact. Flip it off if you ever hear it thinning a deep voice; it is a scalpel, not a default.

There is a fourth enemy specific to the lav (Chapter 14): clothing rustle — the scratchy friction noise of fabric rubbing across the tiny mic as the subject moves and breathes. The cause is almost always a mounting problem, so the fix is in the mount: secure the capsule so nothing brushes it, hide it just under a stable collar or placket rather than against a moving, noisy synthetic fabric, form a small loop of cable for strain relief, and tape down anything that can flap. Nylon jackets and stiff synthetics are the worst offenders; when you can, mount to a calmer, softer piece of the wardrobe. Then monitor (§15.3) while the subject moves and listen for the rustle before you roll for real.

All of this comes to a head in the noisiest, most honest location-sound challenge in the book: the Café Scene. A working café is a wall of sound — the espresso machine's hiss and grind, the fridge, the murmur of other tables, the clatter of cups, music on the house system. You cannot make it quiet; it is supposed to sound like a café. So you do not fight the room, you win within it: you mic the order-counter exchange close on two lavs and cover it with a boom, you time your takes around the loudest machine bursts, you position to keep the espresso machine off the back of the shotgun's pattern, and — because the background is part of the scene — you make sure to capture it clean and separately as room tone (which is the whole of §15.6). The window key you set in Chapter 13 stays exactly where it was; now you add the sound plan under it.

FIGURE 15.8 — Miking the Café order counter: boom + lav safety (top-down)

     [ back bar / espresso machine — the loud one ]  ((≈ hiss & grind
     ================== COUNTER ==================
   barista ( B ) ●                              ● ( C ) customer ordering
     (lav under the apron)                        (lav under the collar)
                        ((•  boom favors the talker, from above & between them
                            \  swing to ( C ) on the order, back to ( B ) on the reply
     window key (Ch.13) ☀                   [ CAM ]  medium of the exchange
   Two close lavs = a clean, present track on EACH voice through the noise;
   the boom is the natural-sounding master. Then grab the café's ROOM TONE (§15.6).

🎒 Gear Note: wind and room tools, as principle. You need three cheap things and a habit. For wind: a foam windscreen for indoor air and breath, a furry cover for real outdoor wind — or, on a phone, a sock. For rumble/handling: a shock mount and the high-pass filter switch. For reverb: nothing to buy at all — blankets, coats, curtains, rugs, and closeness are the whole kit, and a carpeted room with soft furniture is a better "studio" than a bare, beautiful, ringing one. The habit that ties it together: when you scout a location (a pre-production skill you will formalize in Chapter 16, §16.6), listen before you look — clap your hands once in the space and hear how long it rings.

✂️ In the Edit. Post can help with some location noise and not others, and knowing the difference tells you what to fix on set. Steady, unchanging sound — a constant hum, an air-conditioner drone — can often be reduced in the edit with noise reduction (Chapter 33, §33.2), because its consistency is exactly what the tool learns and subtracts. But reverb and wind roar resist cleanup: reverb is the voice itself, delayed and smeared, and wind is broadband chaos, so neither subtracts cleanly, and aggressive attempts leave the voice sounding underwater. The rule that follows: fix reverb and wind on set (get close, absorb, shield); you may leave a steady background hum for post if you must — but only because it is steady. As always, the cleaner the capture, the less the edit has to rescue.

⚠️ Common Mistake: leaving the HVAC on (and forgetting the fridge). The most common ruined-audio story is a continuous background noise nobody switched off — the air conditioning, a ceiling fan, a mini-fridge, a computer's fan — humming under every take because it was so constant that everyone stopped hearing it. Before you roll, walk the room and kill every motor you can: HVAC, fans, fridge, buzzing dimmers, appliances. One warning: if you unplug or switch off a fridge or an aquarium, put a big visible reminder somewhere and turn it back on when you wrap — recordists have quietly cost people a fridge of groceries. Monitoring (§15.3) is what catches the drone you have gone deaf to; killing it at the source is what fixes it.

🔄 Check Your Eye. 1. Why is reverb the location problem you most need to fix on set rather than in post? 2. Name the two simultaneous moves that beat reverb, and the physics behind the first one. 3. Match the tool to the noise: foam windscreen, furry cover, shock mount, high-pass filter — wind roar outdoors, indoor breath/plosives, handling thumps, low rumble.

Check yourself

  1. Because reverb is the voice itself, delayed and smeared into every word; post cannot cleanly separate it back out, so cleanup leaves the voice sounding underwater. Steady hum can be reduced later; echo cannot.
  2. Get closer (proximity — the direct voice obeys the inverse-square law and gets much louder as the mic approaches, while the reflected room sound stays roughly even, so the voice buries the reverb) and add absorption (blankets, coats, curtains, rugs, cushions on hard, parallel surfaces).
  3. Furry cover → wind roar outdoors; foam windscreen → indoor breath/plosives; shock mount → handling thumps; high-pass filter → low rumble.

15.6 Always record room tone

You have done everything right. The dialogue is close, clean, monitored, well within headroom, in a room you tamed. You wrap and go home. And in the edit, you cut two good takes together — and at the exact frame of the cut, the background drops out. One take had the faint hum of the space under it; the next was recorded a moment later when the hum happened to be quieter, and the join lands as a tiny, jarring hole of dead silence that a viewer feels even if they cannot name it. The fix costs thirty seconds and you had to record it on set, because it cannot be made afterward. It is room tone, and this section is the single most important habit in the chapter.

Chapter 14 gave you the term: room tone is the unique ambient sound of a space with no one talking and nothing happening — not true silence (there is no such thing indoors) but the specific, quiet signature of that room: the breath of its air handling, the hum of its lights, the distant character of its walls. This chapter gives you the practice, and the practice is a rule with no exceptions: at every location, at every setup, record 30 to 60 seconds of room tone. Same microphone, same position, same levels as the dialogue you just shot — because the whole point is that it matches the dialogue's acoustic fingerprint and can be laid invisibly underneath it.

Why does the editor need it so badly? Four jobs, all invisible when done right. It fills gaps — laid under a scene, room tone means there is never a hole of dead silence; the background is continuous even where the dialogue is not. It smooths cuts — a bed of matching tone across a cut hides the little shifts in background between takes, so the join disappears (the problem in the story above). It patches trouble — when the editor cuts out a cough, a stumble, or a bumped word, the little gap gets filled with room tone so the deletion is seamless. And it provides sound continuity — the aural cousin of the visual continuity you learned in Chapter 9. Just as a prop must not jump between shots, the background must not jump; room tone is what keeps the sonic world continuous while the picture cuts around inside it.

The protocol is a small ceremony, and doing it out loud makes it stick. At the end of every setup — before you move the mic, strike the light, or let anyone relax — you call it: "Room tone, everybody, thirty seconds — hold still and quiet." Everyone freezes in place (their bodies are part of the room's sound), and you record a clean minute of the space doing nothing, with the exact same mic, placement, and levels. Then you wrap. Every location gets its own — a room tone from the kitchen will not match the living room, and the café's tone will not match the office — because each space has a different fingerprint. It feels faintly absurd the first time, a crew standing in silence staring at each other. Do it anyway. It is the cheapest insurance in production, and the editors who receive your footage will quietly conclude you know what you are doing.

For the Café Scene, room tone is not a background afterthought — it is half the recording plan, because the café's ambient wall of sound is the scene's atmosphere. After you shoot the order-counter dialogue (§15.5), you grab a long, clean take of the café's own voice: the espresso machine idling, the fridge, the murmur, the music bed of the place, with no one delivering lines over it. Handed both — the close, clean dialogue and the rich café tone — the editor can build a bed of that ambience under the entire scene, cut the dialogue freely on top of it, and the café will sound continuous and alive from first frame to last, even though the coverage was shot in a dozen pieces (Chapter 33, §33.5, is where that ambience bed gets designed). This is the anchor's whole lesson made audible: the scene never changed — your command of it did. You framed it, covered it, moved on it, lit it, shot it by the window — and now you have captured its sound like a professional, clean voice over honest atmosphere.

FIGURE 15.9 — Room tone is the glue under a cut (edit timeline preview)

  V1  [ dialogue: take A ###### | dialogue: take B ############### ]
  A1  [ take A sync audio ###### | take B sync audio ############# ]
  A2  [ ROOM TONE bed .......................................... ]  <- one continuous layer
                                 ^ the picture & dialogue cut here
   Without A2, the background lurches (or drops to silence) at the cut and the seam shows.
   With a matching room-tone bed spanning the whole scene, the join is inaudible.

🚪 Threshold Concept: silence is never silent, and you have to record it. The idea that permanently changes how you hear once you truly get it: every room has a sound, even an empty, quiet one, and that sound is a thing you must capture deliberately. Beginners think of "silence" as nothing — as the absence of sound they can ignore. Professionals know that the quiet background of a space is a real, specific, recordable signal that the edit cannot function without, and that it will not exist unless someone stood still and recorded it. Once you believe this, you will never wrap a setup without grabbing room tone again — and you will start noticing, in every film and show you watch, the seamless continuous atmosphere you were never meant to notice, which is the whole point.

🎬 On Set: hear how different "silence" is. Record 30 seconds of room tone in three different spaces — a bathroom, a bedroom, and outdoors, say — with the same mic and level. Then listen to all three back-to-back on headphones. Constraint: no talking, hold completely still, same gain. Self-review question: can you tell, with your eyes closed, which is which? (You can, easily — the bathroom rings, the bedroom is dry and close, outdoors has a wide airy hiss.) That obvious difference is exactly why you can never borrow one room's tone for another, and why every setup needs its own thirty seconds.

🔗 Connection. Room tone is where three chapters shake hands. It is the aural version of the visual continuity you learned to protect in Chapter 9 — the background must not jump any more than a prop should. It is a gift to the assembly you will build in Chapter 26 (§26.6), when your Project 1 talking-head gets cut together and the tone bed keeps the cuts invisible. And it is the raw material for the ambience and soundscape design of Chapter 33 (§33.5), where the café's captured tone becomes a designed bed under the whole scene. Capture it now; it pays off three times downstream.

🔄 Check Your Eye. 1. Define room tone, and say why it is not the same as "silence." 2. Give two of the four jobs room tone does for the editor. 3. Why must every location and setup have its own room tone, recorded with the same mic and levels?

Check yourself

  1. Room tone is the unique ambient sound of a space with no one talking — the specific quiet signature of that room (air handling, lights, walls). It is not "silence" because there is no true silence indoors; every room has a real, recordable background.
  2. Any two of: fills gaps (no holes of dead silence), smooths cuts (hides background shifts between takes), patches edits (fills where a cough/word was removed), provides sound continuity (keeps the background from jumping as the picture cuts).
  3. Because each space has a different acoustic fingerprint — a borrowed tone won't match and the mismatch is audible. Same mic/placement/levels ensures it matches the dialogue it will sit under.

Production Checkpoint

Project 1 locks here. Your task: record your final talking-head audio, plus 30 seconds of room tone. Everything in this chapter and the last converges on one clean recording of the person you are putting on camera.

Run the full discipline, in order:

  1. Choose your system. Single-system (a good mic straight into the camera) or double-system (mic into a separate recorder or a second phone). If double, leave the camera mic on as a scratch track and clap once at the top of each take so the editor can sync it in Chapter 27.
  2. Place and monitor. Get the mic close and off-frame (lav or boom, from Chapter 14). Put on closed-back headphones and listen — not the meters alone — for hum, HVAC, echo, rustle, and plosives. Kill every motor you can (and note anything to switch back on).
  3. Set levels with headroom. Gain so normal speech averages ~-12 dBFS; have your subject deliver their loudest line and confirm it peaks under ~-6 dBFS. Turn AGC off.
  4. Tame the room. If it rings, get closer and hang a blanket behind the mic and the subject. If there is any draft, use a windscreen (or a sock).
  5. Record the takes — several, relaxed, per your directing from Chapter 10.
  6. Grab room tone. Before you move anything: "Room tone, thirty seconds, hold still." Record a clean 30–60 seconds with the same mic, position, and levels.

Why this matters: with this recording, Project 1 is fully shot. Trace it: you chose the concept (Chapter 1), set the camera and exposure (Chapters 2–5), framed and covered it (Chapters 6–7), lit it (Chapters 11–13), and now recorded it cleanly (Chapters 14–15). Every element of your 60-second talking-head is on a card — picture, light, clean dialogue, and matched room tone. You will not touch it again until you assemble it in Chapter 26, where all of this capture becomes a cut. Sound was given equal weight with light across this whole part for one reason, and you have now proven it on your own project: sound is half the picture, and you captured your half like a professional.

Summary

  • Double-system sound records picture on the camera and audio on a separate device for cleaner electronics, mic independence, and redundancy; single-system records both in the camera. Double-system's one obligation is sync.
  • Make sync free: always keep a scratch track on the camera and clap once at the top of each take. The clap's spike lines up in both files (auto-synced or by eye) in Chapter 27.
  • Booming is an athletic skill. Default overhead (aim down at the mouth; reject the floor); go from below only when a wide shot, low ceiling, or boom shadow forces it. Get close, aim at the mouth, hold steady, and favor the talker in a two-person scene. The camera op is the boom op's eyes.
  • Monitoring — active listening on closed-back headphones to the recorder's output — is the only tool that catches quality faults a meter misses. Listen for the nine faults; do a listen-back; keep the cans on during the take.
  • Levels and headroom, in one table:
Number Target Rule
Dialogue average ~-12 dBFS above the noise floor, below the ceiling
Peak ceiling under ~-6 dBFS test the loudest line, not the average
Headroom 6–12 dB digital 0 dBFS is a hard wall — quiet is fixable, clipped is not
AGC OFF stops the noise floor "pumping" in pauses
  • Location sound is managing real spaces, not silencing them. Beat reverb with proximity + absorption (fix it on set — post can't). Beat wind with wind protection (foam indoors, furry cover outdoors, your body, a sock). Beat rumble/handling with a shock mount and the high-pass filter; beat lav clothing rustle with a better mount.
  • Room tone is non-negotiable: 30–60 seconds at every setup, same mic/position/levels. It fills gaps, smooths cuts, patches edits, and keeps sound continuous — the aural version of continuity (Chapter 9), the glue under the assembly (Chapter 26), the raw ambience for the mix (Chapter 33).
  • Fix it in pre and on set, not in post: you cannot un-clip a peak, un-echo a room, or invent room tone you never recorded. The discipline is the whole craft.

Spaced Review

Bring back Chapter 14 — the microphone foundations this chapter is built on. Answer from memory before checking.

  1. Chapter 14 named three microphone types you'd reach for on a shoot. What are they, and which one rides on the end of a boom?
  2. What is a polar pattern, and why does a shotgun's tight pattern help you in a noisy room (the principle you leaned on in §15.2 and §15.5)?
  3. Chapter 14's rule for placement was three words about where to put the mic. What were they — and how does §15.5's "get closer" extend that rule?
  4. What is gain, and how does §15.4's -12 dBFS target turn Chapter 14's "clean signal" into a specific number?
  5. Chapter 14 flagged a threshold concept that this whole part is built on. State it — and say how recording room tone (§15.6) is one more way you honor it.
Check yourself 1. **Lav** (lavalier), **shotgun**, and **handheld** mics. The **shotgun** typically rides on the boom. 2. A **polar pattern** is the map of directions a mic is sensitive to. A shotgun is highly directional (narrow pickup, strong off-axis rejection), so pointed at a voice it hears the voice and rejects sound arriving from the sides — which lets a close boom win over a noisy room. 3. **Close, on-axis, out of frame.** §15.5 extends "close" with the physics: halving the mic-to-mouth distance dramatically raises the direct voice over the room's reverb and noise (inverse-square law). 4. **Gain** is the amount a preamp boosts the mic's weak signal. Chapter 14 said "enough to lift the voice above the hiss, not so much it distorts"; §15.4 makes that concrete — average `~-12 dBFS`, peaks under `~-6 dBFS`, `6–12 dB` of headroom. 5. **Sound is half the picture** — audiences forgive a soft image but leave over bad audio. Recording room tone honors it by protecting the *other* half's continuity, so the finished piece sounds as intentional as it looks.

What's Next

Project 1 is shot, and with it you close Part III: you can now light a scene and record it cleanly, the two skills that most reliably make video look and sound professional — and you will never again point a camera without first asking where is my light, and how is my sound? But everything so far has assumed a small, controllable shoot. Real productions have moving parts: a plan, a schedule, a crew, a call sheet, a location scouted in advance so its light and its sound are known before you arrive. Part IV takes you onto a real set. Chapter 16 begins where every good shoot begins and where this book keeps insisting the cheapest fixes live — in pre-production: the treatment, the script, the storyboard, the shot list, the schedule, and the scout. We are about to make the day itself run on purpose.