48 min read

Play someone a beautifully shot video with hollow, echoing sound and watch their face. Within a few seconds something goes slightly wrong behind their eyes — a small discomfort they usually cannot name — and their thumb starts to itch toward the...

Prerequisites

  • 2

Learning Objectives

  • Explain why audiences abandon bad sound faster than a soft image, and defend the principle that a mic's placement matters more than the recorder behind it.
  • Distinguish the three workhorse microphones — lavalier, shotgun, and handheld — and choose the right one for a given subject and location.
  • Read a microphone's polar pattern and use its off-axis rejection to capture a subject while rejecting the room.
  • Place a mic close, on-axis, and out of frame to maximize the ratio of clean dialogue to noise and reverberation.
  • Set gain for a clean signal — dialogue peaking around -12 dBFS with headroom below 0 dBFS — and recognize clipping on a meter.
  • Choose between on-camera and off-camera sound, and record your talking-head with a proper mic, level-checked and monitored on headphones.

Chapter 14: Audio for Video I — Microphones and Signal Flow

Overview

Play someone a beautifully shot video with hollow, echoing sound and watch their face. Within a few seconds something goes slightly wrong behind their eyes — a small discomfort they usually cannot name — and their thumb starts to itch toward the scrubber or the back button. Now play them a plainly shot video, even a shaky phone clip, where the voice is close, warm, and clear. They settle in. They stay. They believe it. Nothing about the picture explained the difference. The difference was sound, and it is the single most reliable way to make a video feel professional or feel like an accident.

This is the chapter where the book's second promise stops being a slogan and becomes a skill. You have spent three chapters learning to put light on a face (Chapters 11–13). Now you learn to put a microphone near a mouth — which, dollar for dollar and minute for minute, will do even more for your videos than the light did. The astonishing part, and the part beginners refuse to believe until they hear it themselves, is that the recorder barely matters. A twenty-dollar microphone six inches from a mouth beats a two-thousand-dollar camera's built-in mic across a room, every single time, without exception. The craft of production audio is not about owning expensive gear. It is about getting the right kind of microphone close to the sound and rejecting everything else — and that is a skill, not a purchase.

We will keep it concrete and practitioner-simple. You will meet the three microphones that cover almost every job you will ever shoot, learn to read the invisible shape of what a mic hears so you can point it at your subject and away from the air conditioner, place it where it captures clean speech without wandering into frame, and set your levels so the sound is loud enough to use but never distorted past saving. By the end you will mic your talking-head properly and hear — on headphones, which is the only honest way to judge — the exact moment it starts to sound like something people will actually watch.

In this chapter you will learn to:

  • Explain, and prove with a ten-minute test, why sound is half the picture — and why the mic matters more than the recorder.
  • Tell a lavalier, a shotgun, and a handheld mic apart, and know which job each is for.
  • Read a mic's polar pattern and use its off-axis rejection to keep the subject and lose the room.
  • Place a mic close, on-axis, and out of frame — the three rules that fix most bad audio before it happens.
  • Set gain for a clean signal, read a dBFS meter, and recognize clipping before it ruins a take.
  • Choose on-camera vs off-camera sound and record your Project 1 talking-head with a real mic.

Learning Paths

Everyone in this book needs this chapter — bad audio is the great equalizer of failure, and it strikes the phone shooter and the cinema owner alike. Weight your attention like this:

  • 📱 Phone-first: §14.2 (the lav) and §14.4 (placement) are your whole world. A cheap wired or wireless lav plugged into your phone — or simply getting the phone close — is the biggest single upgrade available to you, bigger than any lens or light. Read these two sections twice.
  • 🎥 Creator: §14.5 (gain and levels) and §14.6 (on- vs off-camera) are where your talking-to-camera stops sounding like a bedroom and starts sounding like a channel. Learn to set a level and never clip.
  • 💼 Pro-track: §14.3 (polar patterns, off-axis rejection) and §14.4 (booming close and out of frame) are the fundamentals clients pay you to get right in rooms you did not choose. This is craft you can charge for.
  • 🎓 Student: read in order. This chapter is the theory (how a mic hears) behind the discipline of Chapter 15 (how you run a location sound recording). The Production Checkpoint records the audio for the talking-head you have been lighting since Chapter 11.

14.1 Why sound is half the picture

Let us start by being precise about what a microphone even is, because the definition contains the whole reason placement beats price. A microphone is a device — a transducer — that converts the tiny pressure changes of a sound wave in the air into a matching electrical signal a camera or recorder can store. That is all it does: it listens at one point in space and turns "the air is vibrating like this" into "the voltage is wiggling like that." And here is the consequence that governs everything else in this chapter: a microphone can only capture the sound that reaches it. It cannot lean toward the person. It cannot ignore the fridge. It hears whatever arrives at its little diaphragm, in whatever proportion arrives — the voice and the room echo and the traffic and the air conditioner — and bakes them together into one signal that can never again be pulled apart. Which means the entire game is decided by where you put the mic and which way it faces, long before the sound ever reaches the recorder.

Now the two things a microphone is usually capturing on a shoot, both of which get their names here. Dialogue is the recorded speech of the people in your video — the interview answer, the piece to camera, the line an actor delivers, the vendor explaining their craft. It is almost always the most important sound you record, because it carries the literal meaning, and a viewer's tolerance for unclear dialogue is close to zero: they will strain for a sentence or two and then leave. Room tone is the opposite and its necessary partner — the ambient sound of a location when nobody is talking and nothing is "happening": the hum of the refrigerator, the hiss of the ventilation, the faint traffic, the character of the air in that specific space. Every room has one, even a "silent" room; true digital silence sounds unnervingly dead to us. You will use room tone constantly in the edit to smooth over cuts and hide seams, and we treat recording it as an unbreakable habit in the next chapter — but you meet the term here, because the whole reason you fight for clean dialogue is to keep it separate from, and cleaner than, the room tone underneath it.

Why does sound punish us so much faster than picture? Because of how human perception is wired. We evolved to trust our ears as an early-warning system — a snapping twig in the dark mattered more than a blurry shape — and bad audio trips ancient alarm bells that a soft image does not. A viewer will happily accept a slightly out-of-focus shot, a bit of noise, an imperfect grade; the brain smooths those over. But hollow, distant, echoing, or distorted sound reads as wrong at a level below conscious thought, and it produces an itch to escape. There is an idea, widely credited to filmmakers including George Lucas, that sound is half — or more than half — of the experience of a film. Whether or not the fraction is exact, the working truth for you is blunt: half of "looking professional" is sounding professional, and it is the half almost every beginner neglects, which is exactly why taking it seriously puts you instantly ahead.

🚪 Threshold Concept: sound is half the picture. This is the idea this whole book has been building toward since Chapter 1, and it is the chapter where you own it. Audiences forgive a soft image for a long time; they abandon bad audio in seconds. Once you truly believe this, your whole approach to a shoot changes. You stop treating sound as something the camera happens to also record, and you start treating the microphone as a second camera that you aim, place, and check with the same care. You budget time for it. You listen on headphones the way you look at a monitor. You never again say "we'll get the audio somehow" — because you now know that a video which looks like a million dollars and sounds like a phone in a bucket is a bad video, and the fix costs twenty dollars and thirty seconds of thought. This single reflex — sound gets equal weight — will do more for your work over the next year than any camera upgrade you could make.

To feel this rather than just read it, run the experiment — it takes ten minutes and you will never forget it. Sit a person down and have them say the same three sentences twice, on camera. The first time, put the camera across the room and record with its built-in microphone. The second time, get a cheap microphone (or simply the phone itself) six to eight inches from their mouth, just out of frame, and record the identical lines. Do not change the light. Do not change the framing. Change only the microphone's distance. Then listen to both on headphones.

🎞️ Read This Sequence. The most important comparison in Part III is one you hear, not see. Below is the same talking-head shot rendered twice — every visual field identical, only THE SOUND changed. Read the two SOUND fields against each other; that gap is the entire chapter.

FIGURE 14.1 — "One shot, two soundtracks"        [constructed teaching example]

  Legend for this chapter's audio diagrams:
   ((•  = microphone     ● = mic capsule     ↑ = the subject/source in front of the mic
   → = signal or sound travelling      ))) = a wireless (radio) link

  TAKE A — built-in mic, camera across the room
  THE FRAME    Medium close-up, subject on the right third, soft window key camera-left (your Ch.13 setup).
  THE LIGHT    Unchanged — the good window light you already built.
  THE SOUND    Thin and distant. The voice sits *behind* a wash of room echo; the refrigerator hum and a
               far-off street are as present as the words. It sounds like a recording OF a room that has a
               person in it. You lean in and still lose the ends of sentences.
  THE EFFECT   The viewer's brain flags "amateur / far away / not for me" within a second or two and disengages.
  THE LESSON   A good picture cannot rescue distant sound. The camera heard the room; it barely heard the person.

  TAKE B — cheap mic six inches from the mouth
  THE FRAME    Identical. Same lens, same light, same third.
  THE SOUND    Close, warm, and present. The voice is *in front of* the room; you hear breath, the soft edges
               of consonants, intimacy. The fridge is still there but now it is a distant background, not a
               competitor. The person sounds like they are talking TO you.
  THE EFFECT   The viewer settles in and believes it. It reads as "professional" and they cannot say why.
  THE LESSON   Nothing changed but the mic's distance — and that alone crossed the line from amateur to
               professional. Proximity is the whole game. This is *sound is half the picture*, made audible.

💡 Why It Works: the direct-to-reverberant ratio (the reason close wins). When a person speaks, some sound reaches the mic directly — a straight line from mouth to diaphragm — and much more reaches it indirectly, after bouncing off walls, ceiling, and floor (that bounced, smeared sound is what we hear as "echo" or "roominess"). The direct sound is loudest right at the source and falls off fast with distance; the reflected sound is roughly the same everywhere in the room. So the ratio of clean direct sound to muddy reflected sound is fantastic up close and terrible far away. Halving the distance from mouth to mic hugely improves that ratio — the voice leaps forward and the room falls back. You are not making the mic "better." You are changing the proportion of good sound to bad, and distance is the dial. This is why a cheap mic close beats an expensive mic far, and why the first question of production audio is never "what mic?" but "how do I get it close?"

That is the thesis of the chapter, and honestly of the whole discipline: get the right microphone close to the source and reject everything else. Everything that follows — the mic types, the polar patterns, the placement rules, the levels — is in service of those two moves. Keep the thesis in your head as we go, because every technical detail is just a different way of accomplishing it.

🔄 Check Your Eye. 1. In one sentence, what does a microphone actually do, and why does that make placement matter more than price? 2. A viewer will tolerate a slightly soft, noisy image far longer than bad audio. Why — what is it about sound? 3. You record the same line from across the room and again from six inches away, changing nothing else. Why does the close version sound so much cleaner?

Check yourself

  1. A microphone converts the sound that reaches it into an electrical signal — it captures whatever arrives at one point in space, voice and room and noise all baked together. Because it can only record what reaches it, where you put it and which way it faces decide the result far more than how expensive it is.
  2. Human perception treats sound as an early-warning system, so "wrong" sound (hollow, distant, distorted) registers as alarming below conscious thought and drives the viewer away; the brain smooths over an imperfect image but not bad audio. Half of "sounding professional" is what people mean by "looking professional."
  3. Close up, the direct sound from the mouth dominates the reflected/echoed sound from the room (a great direct-to-reverberant ratio); far away, the reflections and room noise are as loud as the voice. Distance is the dial that sets the proportion of clean sound to muddy sound.

14.2 Microphone types: lav, shotgun, and handheld

There are hundreds of microphones on the market, but for the video you will actually shoot, three families cover almost everything. Learn what each is for — the problem it solves — and you will choose correctly without memorizing model numbers, exactly as you learned lights by their job (key, fill, rim) rather than by brand in Chapter 11.

The lavalier — the clip-on. A lavalier (universally shortened to lav, also called a lapel mic) is a tiny microphone designed to be clipped to a person's clothing near the chest, typically six to eight inches below the chin, so it stays close to the mouth no matter where the person moves or where the camera is. That closeness is its superpower: because it rides on the talent, it is always near the source, so it delivers consistent, present, intelligible dialogue with very little room in it — the direct-to-reverberant ratio (§14.1) is excellent by design. Lavs come in two flavors. A wired lav runs a thin cable from the person to the recorder or phone — cheap, reliable, nothing to run out of battery, but it tethers the subject. A wireless lav clips a small transmitter to the person that beams the signal by radio to a receiver on the camera — freeing them to move, at the cost of batteries, a bit of setup, and occasional radio interference. For a seated talking-head, a five-dollar wired lav into a phone is genuinely all you need. For a walk-and-talk or a moving subject, wireless earns its keep. Lavs are the workhorse of interviews, testimonials, presenters, and reality TV, precisely because one person, one mic, always close, is such a reliable formula.

The shotgun — the reach. A shotgun mic is a long, narrow, highly directional microphone that hears a tight zone straight in front of it and rejects most of what is to its sides and rear — so you can aim it at a subject from outside the frame and capture their voice while ignoring much of the room. The name comes from its shape (a long "barrel," often with an interference tube — more on the physics in §14.3), and its gift is reach with rejection: it lets the mic be a couple of feet away instead of touching the person, which matters enormously when you cannot or do not want to put a mic on the talent. Shotguns are the standard for dialogue where a lav would be seen or is impractical, for run-and-gun documentary and news, and for anything scripted where a mic on a pole hovers just above the frame. One important honesty, because beginners buy a shotgun expecting magic: a shotgun is directional, not telephoto for sound — it does not "zoom in" and pluck a whisper from across a stadium. It narrows what it hears, but it still obeys the close-is-king rule. A shotgun three feet away, aimed right, beats a shotgun ten feet away every time. And indoors, in a small reflective room, a shotgun's long tube can actually pick up smeared reflections and sound roomy — which is why on interiors many pros reach for a shorter, tighter mic (a supercardioid; §14.3) instead. Outdoors and at a respectful distance, the shotgun is superb.

The boom — not a mic, a delivery system. The word boom names the long pole (a boompole) that a mic is mounted on so an operator can hold it just outside the frame and position it inches from the action — and, by extension, the whole technique of "booming" a scene that way. A shotgun (or a supercardioid) on a boom, held over the actors and angled down at their mouths, is how the majority of film and television dialogue is captured, because it combines the two things this chapter worships: a directional mic (rejects the room) placed close (great ratio) and out of frame (invisible). The boom is a person's job — someone holds it, aims it, and follows the dialogue — and we give the technique its own full treatment, with a partner, in Chapter 15. For now, know the vocabulary: the boom gets the shotgun close and out of frame; that pairing is the backbone of professional dialogue.

The handheld — the reporter's stick. A handheld mic is the rugged, familiar "stick" a reporter or singer holds and speaks directly into, usually with a rounded foam or metal ball on the end. It is built to be held right at the mouth (great proximity), to survive abuse, and often to reject handling noise and wind. You will use one for on-camera interviews where it is acceptable to see the mic (person-on-the-street, event hosting, music, vox pops, a stage), and it is the honest tool when hiding a mic is not worth the trouble. Its whole trade is that it is visible and it occupies a hand — which is fine, and even expected, in the genres that use it.

Whichever mic you choose, the signal it produces has to travel a path from mouth to recorded file — and knowing that path tells you where things can go wrong (a dead battery, a bad plug, gain set at the wrong stage). Read it left to right, the way the sound actually flows.

FIGURE 14.2 — Signal flow: how a voice becomes a recording (left to right)

  Legend:  ((•  = microphone     → = sound / signal travelling     ))) = a wireless (radio) link

  WIRED (a lav or shotgun cabled into a recorder or camera):

   ( speaker )        ((•              [ preamp + GAIN ]         [ recorder / camera ]
     mouth  ──sound──►  mic  ──signal──►  boosts the weak    ──►  turns it into a digital
   (the source)        converts air      mic signal to a          file, stored to a card
                       pressure → volts  usable level
                       (a transducer)    (§14.5)

  WIRELESS (a lav on a body-pack transmitter):

   ((•  ──►  [ TX transmitter ]  ))) radio ))) [ RX receiver ]  ──►  [ camera / recorder ]
   lav       clipped to the subject             on/near the camera     records the audio
   (on the                                      (re-check GAIN here)
    person)

The figure names the links in the chain, and every one is a place a recording can die: the mic can be pointed wrong or too far (the transducer only captures what reaches it, §14.1); the gain stage can be set too low or too high (§14.5); a wireless link can drop out or run out of battery; a plug can be the wrong type. The lesson for now is simply that a microphone is never "just a mic" — it is the front of a chain, and you are responsible for every link between the mouth and the file. Test the whole path on headphones before you trust it.

🎒 Gear Note: the classes of mic, and the phone-only path. Do not buy a microphone before you understand the class you need — the job, not the brand. Lav = a mic on the person, for consistent close dialogue (interviews, presenters, moving subjects). Shotgun = a directional mic at a distance, for dialogue you can't or won't put a mic on (film scenes, run-and-gun, on a boom). Handheld = a visible stick at the mouth, for reporting, hosting, and stages. 📱 Phone-first path: you have two excellent free-to-cheap options. First, get the phone close — a phone's own mic six inches from a mouth in a quiet room is dramatically better than most people expect (it's the proximity, not the mic). Second, a cheap wired lav with a TRRS plug (or a lightning/USB-C adapter) clipped to your subject turns your phone into a genuinely professional-sounding dialogue rig for the price of a sandwich; a budget wireless lav does the same for a moving subject. If you buy exactly one audio thing for this whole book, buy a lav. It is the highest-return purchase in video, full stop.

⚠️ Common Mistake: buying a shotgun to use indoors, up on the camera. The classic beginner error is to mount a shotgun on top of the camera and expect it to fix everything. Two problems compound. First, on the camera it is as far from the subject as the camera is — usually far too far — so you get reach but not closeness, and the room floods in. Second, indoors the shotgun's long interference tube smears in reflections and sounds hollow and roomy, the opposite of what you wanted. The fix is almost always the same: for a seated interior subject, use a lav (on the person, always close) or a shotgun/supercardioid on a boom held close and just out of frame — not a directional mic bolted to a camera across the room. The mic wants to be near the mouth, and the camera is not near the mouth.

🔄 Check Your Eye. 1. Match the mic to the job: (a) a seated testimonial where you don't want to see a mic; (b) a scripted two-person scene shot from a few feet back; (c) a reporter doing person-on-the-street interviews at a festival. 2. Why is a lav so reliable for dialogue, in terms of the §14.1 principle? 3. A friend says their new shotgun will "zoom in on sound from across the room." Correct them in one sentence.

Check yourself

  1. (a) a lav — on the person, hidden, always close and consistent; (b) a shotgun (or supercardioid) on a boom — directional, close, and out of frame; (c) a handheld stick — rugged, visible (which is fine for reporting), held right at the mouth.
  2. Because it rides on the talent, six to eight inches from the mouth, it is always close to the source no matter where they or the camera move — so the direct-to-reverberant ratio stays excellent and the dialogue stays present and clean.
  3. A shotgun is directional, not telephoto for sound — it narrows what it hears and rejects the sides, but it still obeys close-is-king; aimed from across the room it just captures a narrow slice of a distant, roomy sound.

14.3 Polar patterns and off-axis rejection

Every microphone has an invisible shape to what it hears — it is more sensitive in some directions than others — and that shape is the second half of "get close and reject the rest." A polar pattern is the map of a microphone's directional sensitivity: which directions it hears loudly, which it hears quietly, and which it rejects almost entirely. Learn to picture the pattern of the mic in your hand and you can aim it — pointing its sensitive zone at your subject and its deaf zones at the air conditioner, the road, or the reflective wall. This aiming is called using the mic's off-axis rejection: "on-axis" is the direction the mic points and hears best; "off-axis" is everything else, which it hears less (or not at all). Off-axis rejection is how a directional mic keeps the subject and loses the room.

There are four patterns worth knowing, from least to most directional:

  • Omnidirectional ("omni") hears all directions equally — a sphere of sensitivity with no rejection anywhere. It has no "back" to point away from noise, which sounds like a weakness but is often a strength: omnis are natural, low-distortion, unfussy about exact aim, and immune to a quirk called the proximity effect. Most lavs are omnidirectional, and that is deliberate — a lav is already so close to the mouth that it doesn't need to reject the room (closeness did that job), and an omni forgives the subject turning their head. The trade is that an omni cannot help you at a distance; its lack of rejection only works because it is close.
  • Cardioid hears the front and sides and rejects the rear — its sensitivity is a heart shape (hence the name) pointing forward. Point a cardioid at your subject and the wall or noise directly behind the mic largely disappears. Handheld mics are often cardioid, which is why a singer's mic rejects the monitor speaker behind it.
  • Supercardioid / hypercardioid is a tighter front lobe with stronger side rejection — but with a small rear lobe, a pocket of sensitivity directly behind the mic. These are the interior dialogue champions: tight enough to reject a reflective room's side and floor bounce, without a shotgun's roomy tube artifacts. The catch is that little rear lobe — you must keep noise sources out of the zone directly behind the mic, not just in front of it.
  • Shotgun (lobar) is the narrowest pattern of all: a long, thin lobe reaching forward with heavy rejection of the sides. It achieves this with an interference tube — a slotted barrel in front of the capsule that causes off-axis sound (arriving from the sides) to partly cancel itself out, while on-axis sound passes straight through. That is the "reach with rejection" of §14.2, and its limits live here too: the rejection is strong for high frequencies but weaker for lows, and in a small room the tube also gathers reflections, which is the roominess problem. Outdoors, aimed true, a shotgun is a scalpel; indoors, it can be a blunt instrument.
FIGURE 14.3 — Four polar patterns from above (● = mic capsule; ↑ = the subject, in front)

   OMNIDIRECTIONAL        CARDIOID            SUPER/HYPERCARDIOID       SHOTGUN (lobar)
   (typical lav)          (many handhelds)    (interior dialogue)       (film / outdoors)

          ↑                    ↑                     ↑                       ↑
       ,------.             ,------.              .-`  `-.                  |''|
      /        \           /   ●    \            /   ●    \                 |● |
     (    ●     )         (  front   )          (  front   )                |  |   long
      \        /           \        /            \        /                 |  |   narrow
       `------`             `--..--`              `-.oo.-`                   |  |   reach
     hears every         heart shape:          tight front,               ` oo `
     direction the       front + sides,        rejects sides,             tiny rear
     same; no rear       rejects the           small REAR lobe (oo)       lobe (oo);
     to aim away         rear (behind ●)       to watch behind ●          rejects the room

   on-axis  = ↑ toward the subject (loudest)     off-axis = the sides/rear the mic rejects

Read the figure as a set of aiming instructions. The pattern tells you where to point the mic's front (at the mouth) and — just as important — where to point its rejection (at the noise). With a cardioid handheld, put the back of the mic toward the loudest interference. With a supercardioid indoors, keep both the front on the subject and the rear lobe pointed away from, say, the espresso machine. With a shotgun on a boom, aim the barrel straight down the axis to the mouth so the voice is dead-on and the room is off to the sides where the tube cancels it. The lav gets a pass on all of this — being omni, it just needs to be close — which is one more reason it is the beginner's best friend.

⚠️ Common Mistake: aiming a directional mic at the room instead of the mouth. A directional mic only rejects noise if it is pointed right. The frequent errors: booming a shotgun so it points at the subject's chest or the wall behind them instead of straight at the mouth (you lose the crispness and gain the room); forgetting a supercardioid's rear lobe and standing the noisy fridge directly behind the mic; or clipping a directional lav and letting the subject turn their head off-axis so the level drops every time they look away. The rule: point the mic's most sensitive axis at the sound you want, and its deaf zone at the sound you don't. Aim is not a detail; for a directional mic, aim is the recording.

🔬 The Tech: condenser vs dynamic, and where the power comes from. Skippable, but it demystifies the spec sheet. Microphones convert sound to signal two main ways. A dynamic mic uses a coil and magnet (like a tiny speaker in reverse); it is rugged, needs no power, handles very loud sources, and is common in handhelds. A condenser mic uses an electrically charged diaphragm and a backplate (a capacitor); it is far more sensitive and detailed — the choice for quiet, nuanced dialogue — but it needs power to work. Shotguns and studio mics are usually condensers powered by phantom power (+48V sent up the cable from a pro recorder or mixer, labeled "48V"). Most lavs are electret condensers that run on a small battery or "plug-in power" from the device. The practical upshots: if a pro mic seems dead, check whether it needs phantom power switched on; and the reason a good condenser lav sounds so present is its sensitivity, which is another argument for keeping loud room noise away from it. None of this changes the core lesson — close and aimed — it just explains what the buttons do.

🔄 Check Your Eye. 1. What is a polar pattern, and why does knowing a mic's pattern let you improve a recording without moving the subject? 2. Why are most lavs omnidirectional, when "rejects the room" sounds so desirable? 3. You're recording dialogue in a small, echoey room and your shotgun sounds hollow and roomy. What kind of mic pattern might serve you better, and why?

Check yourself

  1. A polar pattern is the map of a mic's directional sensitivity — where it hears loudly and where it rejects sound. Knowing it lets you aim: point the sensitive on-axis zone at the mouth and the deaf off-axis zones at the noise, improving the ratio of subject to room without touching the subject.
  2. Because a lav is already extremely close to the mouth, closeness has already won the direct-to-reverberant battle, so it doesn't need directional rejection — and an omni pattern is more natural, less fussy about exact position, and forgives the subject turning their head (which would drop the level on a directional mic).
  3. A supercardioid / hypercardioid — its tighter pattern rejects the room's side and floor reflections without a shotgun's long interference tube, which indoors gathers smeared reflections and sounds roomy. Tight front, controlled rejection, no tube artifacts.

14.4 Placement: close, on-axis, out of frame

Now we put the two ideas together — proximity (§14.1) and pattern (§14.3) — into the physical act of placing the mic, which is where clean audio is actually won or lost. Almost every dialogue-audio problem a beginner has is a placement problem, and almost every fix is one of three rules: close, on-axis, and out of frame. Get all three and mediocre gear sounds great; miss one and great gear sounds mediocre.

Close you already believe: the nearer the mic to the mouth, the more the voice dominates the room and the noise (the ratio again). For a lav, "close" is built in — six to eight inches below the chin. For a boom, "close" means the operator holds the mic as near the mouth as they dare without dipping into frame — often just above the top of the shot, a foot or two away, not up near the ceiling. The instinct to keep the mic "safely" far away is the enemy; safe and far sounds like a hostage video. Push it as close as the frame allows.

On-axis means the mic's most sensitive direction points at the source — the mouth — so the voice hits it dead-on and the rejection lands on the room. A boomed shotgun is aimed straight down the barrel at the mouth. A lav, being omni, is forgiving, but you still center it on the sternum so it hears both sides of the mouth evenly as the head turns. Off-axis, a directional mic sounds duller and quieter, so a mispointed boom robs the voice of its crispness even at the right distance.

Out of frame is the constraint that makes production audio an art: the mic must not be seen (unless it's a handheld, where seeing it is the convention). This is the tension every boom operator lives in — get the mic as close as possible while keeping it a hair above (or below) the frame line. For a lav, "out of frame" usually means hidden: tucked under a layer of clothing or clipped low and unobtrusively, cable dressed down out of sight. Hiding a lav well is a small craft of its own (and a source of the noises we'll warn about in a second), but the principle is simple — the audience should hear the mic, never see it.

FIGURE 14.4 — Mic the talking-head two ways: close, on-axis, out of frame

  A) LAV, clipped and hidden                 B) BOOM (shotgun/supercardioid) from front-top
                                                          ((•  ← mic just ABOVE the frame line,
      ________________________ top of frame                \        aimed DOWN at the mouth (on-axis)
     |                        |                    _________ \_____________ top of frame
     |       ( face )         |                   |            ▼            |
     |          |             |                   |        ( face )         |
     |        ((•  lav on the |                   |           |             |
     |         sternum, 6–8"  |                   |     the mic is 1–2 ft    |
     |         below the chin,|                   |     from the mouth, just |
     |         hidden, cable  |                   |     out of the top edge  |
     |         dressed DOWN   |                   |_________________________|
     |________________________|                     Boom from the FRONT-TOP so the voice is on-axis
        centered so head-turns                       and the mic points down past the face into
        don't change the level                       the floor's dead zone, not back at the room.

Read the two options as the same goal reached two ways. The lav gets close by riding the person; the boom gets close by reaching in from just outside the frame. Both keep the mic near the mouth and both keep it invisible. Which you choose is a story-and-logistics decision: a lav is fast, consistent, and survives the subject moving, but it's a single fixed point on the chest (and it must be hidden); a boom sounds more open and natural and can serve two people, but it needs an operator and constant attention to the frame line. Many professionals record both at once on important dialogue — lav as the safe, consistent backup and boom as the preferred, natural sound — a "two-mic safety" habit we return to when we build the interview in Chapter 19.

⚠️ Common Mistake: the lav that rustles, the cable that thumps. A hidden lav lives against clothing, and clothing is a noise machine. The three classic failures: (1) the mic head rubs a collar, scarf, or jacket every time the person moves, adding a maddening scritch to every gesture — fix it by mounting the mic where fabric won't touch it and taping down anything that can rub; (2) the cable, left loose, transmits thumps up to the capsule as it swings — fix it with a small strain-relief loop taped just below the mic so cable movement never reaches the head; (3) the mic is placed too low or off-center, so the voice is dull and drops out when the head turns — fix it by centering it high on the sternum. None of these is fixable in post; a rustle sits inside the voice. The lav is a five-minute job done right and an unusable take done carelessly.

✂️ In the Edit. Placement is the purest example of you shoot for the edit, because clean, close, consistent dialogue is what makes an edit possible. An editor can cut freely between two interview answers only if they match — same presence, same room level, same tone. If your mic distance wandered (a boom that crept closer and farther, a subject who leaned in and out), the level and the roominess jump at every cut and the piece sounds broken, no matter how good the picture. Worse, some damage cannot be repaired at all: a clipped peak (§14.5) is gone, a rustle lives inside the word, an echoey wide-room recording can never be "de-roomed" back to closeness. The editor and the sound mixer (Chapter 33) can polish clean audio — EQ it, level it, gently reduce steady noise — but they cannot manufacture presence that the mic never captured. Consistent placement on set is the difference between an edit that assembles smoothly and one that fights you every cut. And this is exactly why you also grab thirty seconds of room tone: it is the connective tissue that lets the editor smooth the small mismatches that remain.

🎬 On Set: mic the talking-head, three placements. Sit your subject in your Chapter 13 window light. Record the same three sentences three ways, changing only the mic placement: (1) camera's built-in mic, camera at its normal framing distance; (2) a lav clipped centered on the sternum, hidden under a layer if you can; (3) the phone or a shotgun held just out of the top of frame, a foot or two from the mouth, aimed at the mouth. Constraint: identical framing, light, and lines; headphones on for all three. Self-review: line up the three recordings and rank them; write one sentence naming what closeness and aim each added. Keep the best — it is a direct rehearsal for this chapter's Production Checkpoint.

🔄 Check Your Eye. 1. Name the three placement rules, and say which one a lav gives you "for free." 2. Why do many pros record a lav and a boom on the same important dialogue? 3. A hidden lav adds a scratchy sound whenever the subject moves. What are the two likely causes and their fixes?

Check yourself

  1. Close, on-axis, and out of frame. A lav gives you close for free — it rides the person six to eight inches from the mouth, so proximity is built in (you still handle on-axis by centering it, and out-of-frame by hiding it).
  2. Two-mic safety: the lav is the consistent, always-close backup that survives movement, while the boom gives a more open, natural sound; recording both means one covers the other if a cable rustles, a battery dies, or an aim drifts — and gives the editor a choice.
  3. Clothing rubbing the mic head (fix: mount it where fabric won't touch, tape down rub sources) and cable movement transmitting thumps to the capsule (fix: a small taped strain-relief loop just below the mic). Both are placement problems, not post problems — a rustle sits inside the voice and can't be removed cleanly.

14.5 Gain, levels, and clean signal

You have the mic close, aimed, and out of frame. The last job before you press record is to set how loud the signal is — the difference between a recording you can use and one that is either buried in hiss or shattered by distortion. This is the domain of gain, and it is where a surprising number of otherwise-good recordings die.

Gain is the amount by which your recorder or camera amplifies the weak incoming microphone signal to bring it up to a useful recording level — set with a knob or menu marked "gain," "mic level," "input level," or "volume." A microphone puts out a very small signal; gain boosts it. Set the gain too low and the recorded waveform is tiny — you'll have to raise it later in the edit, and raising it also raises the mic and circuit hiss (the noise floor) that was hiding underneath, so a quiet-but-clean-looking take turns hissy and nasty when you bring it up. Set the gain too high and the signal slams into the digital ceiling and clips — the peaks are chopped flat and reproduce as a horrible crackling distortion that no software on earth can undo. Gain is a Goldilocks control: not too low, not too high, and — crucially — set before you record, because it is baked into the file.

To set it, you have to read a meter, and the scale you read is dBFS — decibels relative to full scale. On this scale, 0 dBFS is the absolute maximum, the ceiling; every usable level is a negative number below it, and the quieter the sound, the more negative (further below zero). This is the notation this book uses for levels: you want your dialogue to peak around -12 dBFS, loud and healthy, while making sure the very loudest peaks (a laugh, an emphatic word) still stay comfortably below 0 dBFS so they never clip. That gap between your normal peaks and the 0 ceiling is your headroom — the safety margin that catches the unexpectedly loud moment. Aim for dialogue peaking near -12, and you have plenty of headroom and plenty of level.

FIGURE 14.5 — Reading an audio meter in dBFS (set gain so dialogue peaks near -12)

    0 dBFS ┤▓▓  ← THE CEILING. Clipping. Touch this and the peak is destroyed FOREVER.
   -3      ┤     ← danger zone: too hot, no headroom left
   -6      ┤██
   -9      ┤████
  -12      ┤██████   ← AIM your DIALOGUE PEAKS here — loud, clean, room to spare above
  -18      ┤██████
  -24      ┤███████
  -40      ┤████████  ← room tone / near-silence lives down here (that's fine)
  -inf     ┴──────────
            quiet  ......  normal speech  ......  a loud laugh / shout
    Too low (peaks at -40): you'll boost it in post and boost the HISS with it → noisy.
    Too hot (peaks at 0):   clipped, cracked, unusable → nothing recovers it. Leave headroom.

Read the meter the way you read the waveform for exposure in Chapter 5 — as an instrument that tells you the truth your ears alone might miss. You are watching for two things: the peaks living up around -12 dBFS during normal speech (loud enough), and never seeing them pin against 0 (safe). If the peaks are crawling along near the bottom at -40, raise the gain until speech reaches roughly -12. If they're kissing the top, pull the gain down. And set it with the real sound: have the person talk at their actual performance volume — including their loudest likely moment — while you set the level, because people always get louder once recording starts. There is a deep asymmetry here worth burning in: being a little too quiet is a recoverable mistake; being too loud is not. A quiet-but-clean take can be lifted (at the cost of some hiss); a clipped take is broken glass. When in genuine doubt, err a touch low and protect your headroom.

⚙️ Settings Box — gain and levels, a starting point (adjust to your gear and scene).

Control Starting point Why
Dialogue peak level around -12 dBFS Loud and clean with generous headroom below the 0 dBFS ceiling.
Absolute peak (loudest laugh/shout) stays below 0 dBFS 0 dBFS is clipping — the one number you may never hit. Headroom catches surprises.
Room tone / quiet down around -40 to -50 dBFS Near-silence should read low; if room tone is high, the room is too noisy — quiet it (Ch.15).
Auto gain / AGC OFF (manual) Auto gain "pumps" — it raises the noise floor in pauses and rides levels unpredictably. Set it yourself.
Monitoring headphones, always The meter shows level, not quality — you only catch rustle, hum, and distortion by listening.
Set level using the subject's real, loudest voice People get louder on "action"; set headroom for the peak, not the rehearsal murmur.
If you can, record a safety track a few dB lower (dual-record) Some recorders capture a second, quieter copy as insurance against a surprise clip (more in Ch.15).

⚠️ Common Mistake: trusting auto gain and never wearing headphones. Two habits wreck more audio than bad mics ever do. The first is leaving the camera on automatic gain control (AGC), which constantly re-rides the level: in a pause it cranks the gain and the room hiss swells up, then it ducks when speech returns, producing a queasy breathing/pumping sound under everything. Switch to manual and set the level yourself. The second is not listening — judging audio by the meter or, worse, by the little on-camera speaker. The meter tells you the level, not the quality; it will happily show a perfect -12 while the take is full of clothing rustle, a ground-loop hum, wind, or a distant leaf-blower you never noticed. Headphones are to sound what the monitor is to picture — the only honest judge. Put them on before the first take and keep them on. We build monitoring into a discipline in Chapter 15; start the habit now.

🔄 Check Your Eye. 1. What is gain, and what goes wrong if you set it too low? Too high? 2. On a dBFS meter, what is 0, and where should your dialogue peaks sit? 3. Why is "a little too quiet" a recoverable mistake but "too loud" usually fatal?

Check yourself

  1. Gain is how much the recorder amplifies the weak mic signal to a usable level. Too low: you must boost it in post, which raises the hidden hiss (noise floor) and makes it noisy. Too high: the signal clips against the digital ceiling and distorts permanently.
  2. 0 dBFS is the absolute maximum — the ceiling and the clipping point. Dialogue peaks should sit around -12 dBFS, loud and clean, with the loudest peaks still safely below 0 (that gap is headroom).
  3. A too-quiet but clean take can be raised later (you pay only in some added hiss); a clipped (too-loud) take has had its peaks chopped flat and destroyed — no software can rebuild the lost waveform, so it's usually unusable.

14.6 On-camera vs off-camera sound

The final decision of this chapter is where the microphone lives relative to the camera — and it is really a restatement of everything above, because it is another way of asking "how do I get the mic close?" There are two families.

On-camera sound means the microphone is on or in the camera: the built-in mic, or an external shotgun mounted in the camera's hotshoe. Its virtue is simplicity — one device, nothing to sync, nothing extra to carry — and its curse is fixed distance: the mic is exactly as far from the subject as the camera is. Since the camera is usually where the frame wants it (a flattering distance for the face, a wide enough shot), and that is almost never as close as the sound wants it, on-camera audio is compromised by default. It gets worse as you shoot wider or longer-lensed: a nice tight portrait shot on a long lens might put the camera fifteen feet back, and now your on-camera mic is fifteen feet from the mouth, drowning in room. On-camera sound is acceptable only when the camera is already close (a handheld vlog at arm's length), or as a scratch track — a rough reference recording used only to sync better audio in the edit, never as the final sound.

Off-camera sound means the microphone is placed independently, near the subject, regardless of where the camera sits: a lav on the person, a boom over them, a handheld in their hand. This is how you honor close, on-axis, out of frame no matter how the shot is framed — the camera goes where the picture needs it, and the mic goes where the sound needs it, and the two are decoupled. Off-camera is the professional default for a reason: it is the only way to keep the mic close while the camera is free to be wherever the composition demands.

FIGURE 14.6 — On-camera vs off-camera: why the frame and the mic want different places

  ON-CAMERA (mic tied to the camera)          OFF-CAMERA (mic goes to the subject)

    ( S ) ↑                                       ( S ) ↑
      :                                          ((•  ← lav on the person (6–8")
      :  the mic is stuck                          :        OR a boom just off-frame
      :  way back HERE with                        :
      :  the camera →                              :   camera is free to sit
      :                                            :   wherever the FRAME wants
   [CAM]((•   far from the mouth =             [CAM]      it — near or far, wide or long,
      room floods in, voice recedes              without ever moving the MIC off the mouth

When the mic is placed off-camera on a separate recorder (rather than cabled back into the camera), you have entered what the next chapter will call double-system sound — camera records picture, a second device records sound, and you sync them in the edit. That is a powerful, standard workflow with its own technique (and its own way of syncing, from a clap to timecode), and it is the opening subject of Chapter 15. For this chapter, the decision is simpler and it is the one that matters most: get the mic off the camera and onto the subject. Whether the audio then rides back into the camera on a cable or lands on a separate recorder, the win is the same — the mic is close because it went to the sound instead of staying with the picture.

🔗 Connection. This chapter is the theory and gear of production audio — what a mic is, how it hears, where to place it, how to set a level. The discipline of running it on a real location — monitoring on headphones as a rule, protecting against wind, taming a bad room, syncing double-system sound, and always recording room tone as an unbreakable habit — is Chapter 15, the second half of this pair. And what happens to your recordings after the shoot — cleaning up steady noise, EQ-ing and leveling the dialogue, laying music and effects, and mastering the whole mix to a loudness standard — is audio post-production, Chapter 33. Two truths bridge all three: the cleaner your capture here, the less those later chapters have to rescue; and no amount of post can add presence a mic never captured. Fix it at the mic first.

♿ Accessibility & Inclusion. Sound and access are deeply linked, in two directions. First, because many viewers watch with the sound off or cannot hear it at all, captions are not optional — but captions are only as good as the dialogue behind them: clean, close audio produces accurate captions (including from automatic transcription), while a muddy room recording produces garbled ones that fail exactly the viewers who depend on them. Recording clean dialogue is an accessibility act. Second, when you put a microphone on a person — clipping a lav under their clothing, running a transmitter at the small of their back — you are entering their personal space, so ask first, explain what you're doing, and let them place or clip the mic themselves when that's more comfortable. Consent to be mic'd is part of consent to be filmed (releases are Chapter 38). Good audio practice and respectful practice are the same practice.

🔄 Check Your Eye. 1. Why is on-camera sound compromised by default, and when is it actually acceptable? 2. What does "off-camera sound" let you do that on-camera sound cannot? 3. What is the one word that separates plain off-camera sound from "double-system sound," and where do you learn that workflow?

Check yourself

  1. Because the mic is stuck exactly as far from the subject as the camera is, and the camera is placed for the frame, which is almost never as close as the sound needs — so the room floods in, worse the wider or longer you shoot. It's acceptable only when the camera is already close (an arm's-length vlog) or as a rough scratch track used only to sync better audio.
  2. It decouples mic from camera: the camera goes where the picture needs it while the mic stays close, on-axis, and out of frame on the subject — honoring all three placement rules regardless of framing.
  3. A separate recorder — when the off-camera mic records to its own device (not back into the camera), it's double-system sound, synced in the edit; you learn that workflow in Chapter 15.

Production Checkpoint

Your task: record your Project 1 talking-head with a proper microphone — a lav or a shotgun, placed close and out of frame — set your levels, and listen back on headphones.

Return to the window-lit talking-head you have been building since Chapter 11. This time the job is entirely the sound. Choose your mic by the job (§14.2): a lav clipped centered on the sternum, six to eight inches below the chin and hidden if you can, is the fast, reliable choice for a seated subject; a shotgun or supercardioid on a boom (or simply held just out of the top of the frame, aimed at the mouth) is the alternative. Whichever you pick, honor the three rules: close, on-axis, out of frame (§14.4). Set your gain manually so the subject's dialogue peaks around -12 dBFS with the loudest moments safely below 0 (§14.5) — turn off any automatic level control. Then, the non-negotiable step: put on headphones and listen to a test take before you shoot for real. Listen past the words for the enemies — clothing rustle, a cable thump, a hum, the room's echo, an air conditioner you had tuned out. Fix what you hear at the mic; do not "leave it for post."

Shoot the same thirty seconds to camera you have been refining, now clean. If you can, record a lav and a boom (or a lav and the phone close) at once — two-mic safety, and a preview of Chapter 19's interview discipline.

Why this matters: your talking-head now looks good and, for the first time, sounds good — and you have felt the truth of this chapter in your own ears, which is the only way it ever really lands. You have also just done the thing that most separates amateur from professional video, and it cost you a cheap mic and thirty seconds of listening. Project 1's picture and sound are nearly complete; in Chapter 15 you add the location-sound discipline — monitoring, taming the room, and recording room tone — that locks it.

Summary

  • Sound is half the picture. Audiences forgive a soft image but abandon bad audio in seconds, because perception treats wrong sound as alarming. Give sound equal weight to picture from the first shoot.
  • A microphone captures only the sound that reaches it, so placement beats price. Getting a cheap mic close to the mouth beats an expensive mic far away, every time — because closeness wins the direct-to-reverberant ratio (voice over room echo and noise). The whole craft: get the right mic close, and reject the rest.
  • The three workhorse mics, by job:
Mic What it is Best for Watch out for
Lavalier (lav) tiny clip-on, on the person, usually omni interviews, presenters, moving subjects; the phone-first hero clothing rustle, cable thumps; hide it well, center it
Shotgun long, narrow, directional (interference tube) dialogue at a distance, film scenes, run-and-gun, on a boom roomy indoors; still needs to be close; aim it true
Handheld rugged visible "stick," at the mouth reporting, hosting, stages, vox pops it's visible (fine for those genres) and uses a hand
Boom a shotgun/supercardioid on a pole, off-frame most film/TV dialogue — close + directional + invisible needs an operator; mind the frame line (Ch.15)
  • Polar pattern = the map of where a mic hears. Omni (most lavs) hears all directions — fine because it's close. Cardioid rejects the rear. Supercardioid is tighter (interior dialogue champ) with a small rear lobe to watch. Shotgun/lobar is narrowest (great outdoors, roomy indoors). Use off-axis rejection: point the sensitive axis at the mouth, the deaf zone at the noise.
  • Placement in three rules: close, on-axis, out of frame. Lav gives "close" for free (center it, hide it); boom reaches in from just off-frame, aimed at the mouth. Record both on important dialogue (two-mic safety).
  • Gain and levels: set gain manually so dialogue peaks around -12 dBFS, peaks below 0 dBFS (the clipping ceiling). Too low → boost adds hiss; too high → clipping is permanent. Too quiet is recoverable; too loud is not. Turn AGC off; judge on headphones, not the meter or the camera speaker.
  • On-camera sound is stuck at the camera's distance (usually too far) — a scratch track at best. Off-camera puts the mic on the subject regardless of framing — the professional default. A mic on its own recorder is double-system sound (defined and drilled in Chapter 15).
Decision Choose… Because
Seated interview, hide the mic Lav (or lav + boom) on the person = always close and consistent
Two people, mic out of shot Boom (shotgun/supercardioid) directional + close + invisible
Reporting / hosting / stage Handheld visible is fine; rugged; right at the mouth
Small echoey room Supercardioid over a shotgun rejects reflections without a tube's roominess
Any shot where camera is far Off-camera mic decouple the mic from the frame; go to the sound

Spaced Review

Retrieval keeps skills from fading. Answer these from earlier chapters before you check — then read on.

  1. (Ch.2) A microphone converts sound to an electrical signal that gets recorded; a camera sensor does the parallel job for light. In one sentence each, what does the sensor do, and why does "resolution" describe the picture but tell you nothing about the sound?
  2. (Ch.2) You shoot a portrait on a long lens with the camera fifteen feet back. What does that framing distance do to your on-camera audio, and which chapter-14 idea explains it?
  3. (Ch.11) In three-point lighting, the key is the main shaping source you place with intention. Draw the analogy: what is the "key" of your audio, and where do you place it?
  4. (Ch.11) A motivated light is justified by a believable source in the scene. How is choosing a mic also a "motivated" decision — matched to the subject and location rather than to a spec sheet?
  5. (Ch.2 + Ch.5) You judge exposure with a waveform/zebras rather than your eye. What is the audio equivalent you set your gain against, and what is the one number on it you may never hit?
Check yourself 1. The **sensor** converts light into an electrical signal that becomes the image (Chapter 2). **Resolution** counts the picture's pixels — a purely visual measure — and audio is a completely separate signal captured by a completely separate device (the mic), so a "4K" spec says nothing about whether the sound is any good. (This is *sound is half the picture* in a sentence: the two halves are captured independently.) 2. It puts the on-camera mic *fifteen feet from the mouth*, so the room floods in and the voice recedes — because **the on-camera mic is exactly as far from the subject as the camera is**, and the frame (not the sound) chose that distance. Off-camera mic'ing is the fix. 3. Your audio "key" is the **microphone** — your main, intentional capture of the most important sound (the dialogue) — and you place it *close, on-axis, and out of frame*, near the mouth, just as you place a key light with purpose rather than leaving it to chance. 4. Choosing a mic is motivated when it's driven by the *situation* — a lav because the subject moves and you can't see a mic; a shotgun on a boom because it's a two-person scene; a supercardioid because the room is small and reflective — rather than by which mic is most expensive. The job justifies the tool, exactly as a believable source justifies a light. 5. The audio meter in **dBFS** — you set gain so dialogue peaks near **-12 dBFS**, and the number you may never hit is **0 dBFS**, the clipping ceiling (the ears-on equivalent of judging by headphones, since the meter shows level but not quality).

What's Next

You can now choose a microphone, read how it hears, place it close and out of frame, and set a clean level — which is most of what separates audio that works from audio that drives viewers away. But knowing which mic and where is only half of production sound. The other half is the discipline of running it on a real, uncooperative location: trusting your ears over the meter, taming a room that echoes and a wind that roars, syncing a separate recorder to the picture, and the habit that will save more edits than any other — recording room tone every single time. In Chapter 15 we take the microphone into the field, boom a scene with a partner, and turn "I have a mic" into "I reliably capture clean dialogue anywhere." Bring your talking-head; it is about to become bulletproof.