Here is a hard truth about nonfiction video: the interview is usually the whole show. The B-roll is beautiful, the music is moving, the graphics are clean — but the reason the piece exists, the thing a viewer actually leans in for, is a real person...
Prerequisites
- 11
- 14
Learning Objectives
- Build a repeatable interview setup with correct framing, looking room, and an off-axis eyeline.
- Light a sit-down interview by applying three-point and window-key principles from Part III to a real face.
- Record a two-mic-safe interview with a lav and a boom, monitored on headphones and protected from clipping.
- Write and sequence interview questions that are open, specific, and quiet enough to produce full-sentence answers.
- Listen actively — using silence and follow-ups — to get the answer in a self-contained sentence the editor can actually cut.
- Choose between a single-camera and a two-camera interview, and shoot each so the edit will hold together.
In This Chapter
- Overview
- Learning Paths
- 19.1 The interview setup: framing and off-axis eyeline
- 19.2 Lighting the interview (applying Part III)
- 19.3 Audio for the interview: two-mic safety
- 19.4 Question design: open, specific, quiet
- 19.5 Listening and getting the answer in a full sentence
- 19.6 The two-camera and single-camera interview
- Production Checkpoint
- Summary
- Spaced Review
- What's Next
Chapter 19: The Interview
Overview
Here is a hard truth about nonfiction video: the interview is usually the whole show. The B-roll is beautiful, the music is moving, the graphics are clean — but the reason the piece exists, the thing a viewer actually leans in for, is a real person saying something true on camera. Documentaries, testimonials, brand films, YouTube explainers, oral histories, news packages: strip away the polish and most of them are an interview with pictures laid over it. Get the interview right and you have gold to cut. Get it wrong and no amount of B-roll, color, or sound design will save you, because you cannot grade or mix your way to an answer that was never spoken.
And "getting it wrong" almost never means the camera failed. It means one of four ordinary, avoidable things: the subject was framed and lit so they looked uncomfortable and untrustworthy; the audio had a single point of failure and it failed; the questions were built to produce one-word answers; or the interviewer talked so much, and listened so little, that the subject never relaxed into the real story. Every one of those is a decision made before or during the interview — which is to say, every one of them is fixable, for free, by you. This chapter is where the whole book so far comes together in one chair.
That is what makes this a synthesis chapter. You already know how to frame a face on a third (Chapter 6), build coverage that cuts (Chapter 7), direct a real person to relax (Chapter 10), light with a key and a window (Chapters 11–13), and record clean dialogue with a lav or a boom (Chapters 14–15). We are not going to re-teach any of that. We are going to point all of it at a single, high-value target — a person in a chair, telling you something worth keeping — and add the specific craft that only the interview needs: the off-axis eyeline, two-mic safety, question design, and the discipline of listening for the full-sentence answer.
In this chapter you will learn to:
- Build the interview setup — the repeatable arrangement of subject, camera, interviewer, light, and mic — with correct framing and looking room.
- Place an off-axis eyeline so your subject looks like they are in a real conversation, not lost or staring past the viewer.
- Light the interview by applying Part III, not reinventing it: a motivated key, a controlled ratio, separation from the background.
- Record audio with two-mic safety so a single failure never costs you the answer.
- Practice question design — open, specific, quiet — and recognize the non-answer before it wastes your day.
- Use active listening and silence to pull the answer out as a self-contained sentence the editor can cut.
- Decide between the single-camera and two-camera interview, and shoot either one so the cut holds.
Learning Paths
Every reader needs this chapter — the interview is the one setup that shows up in every kind of paid and personal work. Weight your attention like this:
- 📱 Phone-first: §19.1, §19.3, and §19.4 are your core. A phone on a tripod, one window, and one clip-on mic is a completely legitimate interview rig — the craft is in the eyeline, the questions, and the listening, none of which cost money.
- 🎥 Creator: §19.4 and §19.5 will change your channel. Most "talking to camera" creators never learn to interview others, and it is the fastest way to add credibility and variety to a feed.
- 💼 Pro-track: all of it, but especially §19.3 (two-mic safety is non-negotiable on a paid job) and §19.6 (the two-camera setup you will be expected to run). This is bread-and-butter billable work.
- 🎓 Student: §19.1 and §19.2 connect your Part II and Part III skills into one deliverable; §19.5 is the craft that separates a passing interview from a memorable one.
19.1 The interview setup: framing and off-axis eyeline
Before a single question is asked, you make a set of physical decisions that will shape every second of footage: where the subject sits, how tightly you frame them, where they look, and where you — the interviewer — sit relative to the lens. Together, that arrangement is the interview setup: the repeatable configuration of subject, camera, interviewer, light, and microphone that you build before the conversation starts. It is the interview's equivalent of blocking (Chapter 10) — staging worked out in advance so that, once you roll, you can forget the mechanics entirely and just listen. A good setup is invisible in the finished piece; a bad one announces itself in the first two seconds and never lets the viewer settle.
Start with the frame. A sit-down interview almost always lives in the medium-to-medium-close-up range you learned in Chapter 7 — roughly from mid-chest to just above the head. Wider than that and the face, where all the meaning is, gets small; tighter than a comfortable close-up and you lose the shoulders and hands that carry so much of how a person reads. Seat the subject on a third, not dead center, and leave looking room (the lead room from Chapter 6, applied to a face): more space on the side they are facing than behind their head. A person pushed to the edge they are looking toward feels cramped and about to leave the frame; a person centered feels like a mugshot. On a third, with air in front of their gaze, they feel composed and in conversation with someone.
FIGURE 19.1 — Interview framing: subject on a third, with looking room (16:9)
+---------------------------------------------------+
| . . |
| O·················|··········|·············| <- upper thirds line
| ( head ) |
| ( S = subject ) |
| LOOKING ROOM --> ( shoulders ) <-- less |
| (more space on | | space |
| the side they O O behind |
| face) . . the head |
+---------------------------------------------------+
subject sits on the RIGHT third, facing LEFT;
eyes land on the upper-third line; the gaze crosses
the open left side of the frame toward the interviewer.
The eyes should land on or just below the upper-third line — that is where a viewer looks first, so that is where you put the most important feature. Give a hand's width of headroom above the hair (Chapter 6): enough that they are not crammed under the top edge, not so much that a lake of empty ceiling floats over them.
Now the decision that defines the interview look: where does the subject look? Not into the lens. In a standard nonfiction interview, the subject looks at you, the interviewer, who sits just to one side of the camera — a placement we call the off-axis eyeline: the subject's gaze is directed slightly off the lens axis, at a real person beside the camera, rather than straight down the barrel. This one choice carries a surprising amount of meaning. When someone looks into the lens, they are talking to the viewer — direct address, the register of a host, a spokesperson, a piece to camera. When they look just off the lens at an interviewer, they are talking to someone else, and the viewer becomes a privileged observer of a real conversation. That second mode is the grammar of documentary and testimonial: it feels overheard, candid, true. (There is a famous, deliberate exception — the direct-to-lens interview — and we analyze it in this chapter's Case Study 1. Know the rule first; then you can break it on purpose.)
How far off-axis? This is a dial, not a switch, and it is worth setting deliberately. The closer you sit to the lens, the more of the subject's eyes the camera sees and the more intimate and direct the gaze feels — almost, but not quite, talking to us. The farther you sit from the lens, the more the subject turns into profile, the more we see the side and back of the eye, and the more detached and "watching from a distance" it reads. The tightest, warmest documentary look usually comes from sitting close to the lens — your head nearly touching the matte box or phone — so the subject's near eye still reads full and the gaze skims just past the camera. A common beginner error is to sit a meter to the side, which throws the subject into a lost-looking three-quarter-profile that feels evasive on screen even when the person is being completely honest.
⚠️ Common Mistake: the interviewer parked too far from the lens. Set the camera up, then sit down to run the interview from wherever there happens to be a chair — often well off to the side. The subject dutifully turns to face you, and now their eyeline is so wide that the camera is looking at their cheek. On screen they seem to be avoiding us. The fix: sit as close to the lens as you physically can without being in the shot — directly beside it, at the subject's eye height — so the off-axis angle is small and the gaze feels like near-eye-contact. If you are solo and can't sit by a camera you're also operating, put a small object or your own face right next to the lens and have them talk to that fixed point.
Two more setup decisions finish the arrangement. First, eye height: put the lens at roughly the subject's eye level. Below eye level and you shoot up their nose and hand them unearned power and a double chin; well above and you look down on them, literally and figuratively. Eye level is the neutral, respectful default — deviate only when the story wants the subtext that a high or low angle carries (Chapter 7). Second, the space behind them: an interview background should have some depth and separation (we will light for that in §19.2) and nothing that grows out of their head or distracts. A wall pushed right up behind the subject is the flattest, most hostage-video-looking choice there is; give them a few feet of room so the background can fall soft.
Here is the whole thing as a top-down map — the single most useful diagram in this chapter, because it is the setup you will build again and again for the rest of your career.
FIGURE 19.2 — The interview setup (top-down): subject, off-axis eyeline, key, and mic
☀ / ▢ KEY (soft, on the side the subject faces,
\ just above eye level, feathered down)
\
↘
( S ) ──────► eyeline to interviewer
/ | \ |
(fill) ○ | ▣ rim/back |
camera-near | (separates | ((• boom just
shadow | hair from | out of top frame
side | background) | OR ⟂ lav on chest
| ▼
((• lav clipped, hidden ( I ) = interviewer,
| seated a hair to the
[ CAM ] ~50mm-equiv, eye height, LENS's side, as CLOSE
| subject on a third to [CAM] as possible
▼
small off-axis angle ⟵ I sits here, beside the lens
Result: subject looks *just past* the lens at a real person → the "candid conversation" look.
Walk the map. The subject sits at eye height to a camera on a third. The interviewer — you — sits as close to the lens as possible on one side, so the subject's eyeline crosses only a few degrees off-axis. The key light comes from the side the subject faces (motivated, as if by the interviewer's presence or a window; see §19.2), placed just above eye level. A fill lifts the near, camera-side shadow to taste; a rim separates them from the background. Audio is doubled — a hidden lav on the chest and a boom just out of the top of frame (§19.3). Build this once and it becomes muscle memory; then every interview is just this map, adjusted to the room.
🎬 On Set: build the setup, no questions yet. Seat a willing friend and build the FIGURE 19.2 arrangement with whatever you own: subject on a third, eye-height camera, you sitting right beside the lens, one soft light or window as key. Shoot 30 seconds of them just describing their morning. Constraint: do not touch a single question-writing or lighting-refinement idea yet — this drill is only framing and eyeline. Self-review: play it back on mute. Does it look like they are talking to a person just off-screen, or staring lost into space? If it's the second, you sat too far from the lens. Move in and reshoot.
Finally, a Described Shot of the target — what a strong, ordinary, achievable interview frame actually looks like, so you have a picture to aim at. Nothing here requires more than a window, a chair, and a phone.
FIGURE 19.3 — "A strong interview frame" [constructed teaching example]
THE FRAME Medium close-up. The subject sits on the right third, framed from mid-chest up, turned
a few degrees to look camera-left at an interviewer just beside the lens. A hand's width of
headroom; open looking room across the left of the frame. Background: a room falling soft
a few feet behind, one out-of-focus lamp glowing in the corner for depth.
THE MOVE Locked off on a tripod at the subject's eye height. Absolute stillness — the interview is
about the words and the face; a moving camera would only compete with them.
THE LIGHT Soft key from camera-left (the side they face), just above eye level, wrapping the front of
the face; the camera-near cheek sits about a stop and a half darker for gentle modeling; a
faint rim skims the far shoulder and hair, lifting them off the background.
THE SOUND Their voice, close and warm, from a lav under the collar with a boom backing it up; a low
bed of room tone underneath. No music yet — in an interview, the voice is the event.
THE CUT This is the "spine" shot the whole piece is built on. It will hold on the strong sentences
and cut away to B-roll over the weaker connective bits, the interview audio continuing (an
L-cut) beneath the pictures.
THE EFFECT Because the eyeline is only slightly off-lens and the light leads into their gaze, the viewer
feels spoken-*near*, not spoken-*at* — a candid, trustworthy, someone-is-confiding register.
THE LESSON The interview look is made of small, cheap decisions — a third, a few degrees of eyeline,
a key on the looking side — stacked. No single one is expensive; together they read as "pro."
🔄 Check Your Eye. 1. What is the difference in meaning between a subject looking into the lens versus just off it? 2. On which third do you place a subject who is facing camera-left — and where does the looking room go? 3. Your subject looks like they're avoiding the camera even though they're being sincere. What's the most likely setup cause?
Check yourself
- Into the lens = direct address, talking to the viewer (host/spokesperson). Just off the lens (an off-axis eyeline) = talking to an interviewer, and the viewer becomes an observer of a real conversation — the documentary/testimonial register.
- Place them on the right third so their gaze crosses the open left side of the frame; the looking room is the extra space on the side they face (camera-left).
- You (the interviewer) are sitting too far from the lens, so the eyeline is too wide and the camera sees too much profile. Move in beside the lens.
19.2 Lighting the interview (applying Part III)
You already know how to light. Chapter 11 gave you the key, fill, and rim of three-point lighting and the rule that light should be motivated; Chapter 12 gave you soft versus hard, diffusion, and white balance; Chapter 13 gave you the window as a free softbox. This section does not add a new lighting system — it aims the one you have at a face in a chair and makes the handful of decisions the interview specifically needs.
Key from the side the subject faces. The single most reliable, most flattering interview key sits on the looking side — the same side as the interviewer, the direction the subject's gaze travels. Put it there, soft, just above eye level, and feather it down across the face. Two things happen. First, the light "leads" the eye line, so the bright side of the face is the side opening toward the conversation and the shadow falls away on the camera-near side, giving you natural modeling instead of a flat mask. Second, it reads as motivated — the subject looks lit by the person they are talking to, or by the window they are sitting near, exactly the intuition Chapter 11 drilled. Keying from the opposite side (shadow on the looking side) is the "sinister" choice; it's a legitimate mood tool, but it is a deliberate departure, not a default.
FIGURE 19.4 — Interview lighting map (top-down): a motivated key + controlled ratio
[ WINDOW or soft KEY ] ☀/▢
\ (side the subject faces; ~45° round, just above eye level)
↘
( S ) ───────► eyeline / looking side (lit)
shadow | ↖
side ○ | ▣ rim/back light (opposite/high-behind,
(fill, | a stop or so over key, edges the hair
~1.5 | and far shoulder off the background)
stops |
down) ▽ practical lamp in the soft background (depth, not exposure)
|
[ CAM ]
Key : fill ≈ 2:1–3:1 (gentle) for a warm doc look; open it up (near 1:1) for corporate/clean;
crush the fill (4:1+) only when the story wants weight. Rim optional but cheap separation.
Set the ratio to the tone. The relationship between your key and your fill — the contrast ratio you met in Chapter 11 — is a storytelling dial, not a technical constant. A gentle 2:1 or 3:1 (fill about one to one-and-a-half stops under the key) gives the warm, dimensional, slightly moody look most documentaries want. Opening the fill up toward 1:1 gives the bright, even, trustworthy look most corporate and testimonial work wants (the brand does not want their customer looking noir). Crushing the fill to 4:1 or beyond hands you weight and unease — save it for the story that has earned it. You are not choosing a "correct" ratio; you are choosing what the piece is trying to make the viewer feel, which is Theme 1 doing its job.
Separate them from the background. A face on a flat, evenly lit wall looks pasted on. Two cheap moves fix it, both from Part III: add a rim or hair light (Chapter 11) from behind to edge the subject off the background, and light — or simply place — the background so it is darker, softer, and deeper than the subject (a practical lamp glowing out of focus, per FIGURE 19.4, does this for free). Separation is what gives an interview that three-dimensional, "there's a world behind them" feel instead of the passport-photo flatness of a wall pressed to the back of their head.
The window-key interview. For the phone-first and no-budget reader, the entire section above is available with zero lights. Seat the subject so a large window is on their looking side, a few feet away — that is your soft key (Chapter 13, the window as a free softbox). Bounce a little light back into the camera-near shadow with anything white (a foam board, a bedsheet, a reflector) for your fill. Let a lamp or a bright bit of the room behind them provide separation. Motivated, soft, flattering, and free. The window is the professional interview light; the only thing money buys you is control over when the sun goes behind a cloud.
⚠️ Common Mistake: a window or lamp behind the subject. The most common interview lighting disaster is seating someone with the brightest thing in the room behind them — a window at their back, a lamp over their shoulder. The camera exposes for the bright background and the face goes to silhouette, or you expose for the face and blow the window to white. The fix, decided in pre (Theme 5): put the brightest source in front of the subject, on their looking side, and make sure the background is darker than the face. If you're stuck with a window behind them, either close the blind and light the face yourself, or reframe the whole setup 90–180° so the window becomes the key.
Now, the anchor duty for this chapter, and a look you have been promised since Chapter 11: the classic three-point interview look, rendered as a "Watch This" Described Shot. This is the grammar you have seen ten thousand times without naming it — the documentary talking head, lit so cleanly it disappears.
🎞️ Read This Sequence: the three-point interview look (the "Watch This" entry Chapter 11 previewed). This is not one film; it is the shared visual language of broadcast and documentary interviews for decades — a Tier-1 convention, described here as a teaching case, never a specific reproduced clip.
FIGURE 19.5 — "The classic three-point interview look" [constructed teaching example — a genre convention]
THE FRAME Medium close-up, subject on a third, off-axis eyeline to an unseen interviewer. Clean, calm,
symmetrical enough to feel composed, asymmetrical enough (the third) to feel alive.
THE MOVE Locked off. The stillness is the format — your attention goes entirely to the face and words.
THE LIGHT A soft KEY on the looking side, just above eye level; a gentle FILL from near camera lifting
the shadow to a ~2:1 ratio; a RIM from high-behind on the opposite side edging the hair and
shoulder off a deep, soft background with one out-of-focus practical glowing in it.
THE SOUND Close, warm dialogue (lav + boom); a bed of room tone; music, if any, kept low and out of the
way of the voice — because in this look the voice is the entire point.
THE CUT Built to be cut *away from*: the setup is so consistent that B-roll can drop over any weak
moment and the interview audio carries underneath (the L-cut), then return to this same frame.
THE EFFECT Invisibility. Done right, the viewer never thinks about the lighting at all; they think only
about the person and what they're saying. That disappearance IS the craft.
THE LESSON Three-point lighting isn't a look you notice — it's a look that gets *out of the way* so the
story can land. When people say an interview looks "professional," this quiet setup is usually why.
✂️ In the Edit. A production choice in this section pays a direct dividend in post: consistency lets you cut invisibly. If every interview is lit the same way — same key side, same ratio, same background separation — then B-roll can drop over any moment and the returns to the face all match, and a two-camera cut (§19.6) won't jar. But if the sun moves during a long window-key interview and the exposure drifts, the editor is stuck matching shots that don't match (a job you'll do the hard way in Chapter 31). The cheap fix is on set: watch your exposure across the whole interview, and if you're on a window, either shoot fast or supplement it so the light doesn't wander.
🔄 Check Your Eye. 1. Which side of the face do you usually key in an interview, and why does it look motivated? 2. What two cheap moves separate a subject from a flat background? 3. A subject is a dark silhouette against a bright window behind them. Name the fix and the stage it should have been solved in.
Check yourself
- The looking side (the side they face / the interviewer's side). The light leads the eyeline and reads as coming from the conversation or the window — motivated, and it drops the shadow onto the camera-near side for natural modeling.
- A rim/hair light from behind, and making the background darker/softer/deeper than the subject (e.g., an out-of-focus practical).
- Put the brightest source in front of them on their looking side and keep the background darker than the face; ideally solved in pre-production by choosing where they sit — fix it in pre, not in post.
19.3 Audio for the interview: two-mic safety
Chapters 14 and 15 taught you that sound is half the picture, how a lav and a shotgun differ, how polar patterns reject off-axis noise, how to set gain without clipping, and why you always record room tone. The interview asks you to apply all of it under one unforgiving condition: you usually get one take of the truth. A performer can do a line again; a person telling you about the worst day of their life, or a CEO giving you eight minutes between meetings, often cannot re-say it the same way, and would lose all their spontaneity if they tried. So the interview's audio rule is not just "record it well." It is: record it so that no single failure can lose it. That principle has a name.
Two-mic safety is the practice of recording the subject on two independent microphones at once — a lav and a boom/shotgun — onto two separate channels, so that if either one fails you still have a clean, usable track. It is the audio equivalent of a spare parachute, and on any interview that matters it is not optional. Lavs and booms fail in different, uncorrelated ways, which is exactly why running both is safe: a lav can catch clothing rustle, a scarf, a seatbelt-style cable snag, or a radio dropout; a boom can drift off-axis, catch a room reflection, or throw a shadow that forces a reframe. The chance that both fail on the same sentence, in the same way, at the same moment is vanishingly small. Run both and you have turned a single point of failure into a redundant system.
FIGURE 19.6 — Two-mic safety: signal flow for a redundant interview recording
(( lav on subject's chest )) ──► [ wireless TX ] ··· [ wireless RX ] ──► CH 1 ─┐
├─► [ RECORDER ]
(( boom / shotgun overhead )) ──────────────────────► [ preamp/gain ] ──► CH 2 ─┘ │
▼
monitor BOTH channels on 🎧 headphones two clean,
set each level so peaks sit around -12 dBFS, independent
well below 0 dBFS (no clipping) takes of the
same words
If CH 1 (lav) catches a rustle on one sentence → the editor uses CH 2 (boom) for that line, and back.
Record 30 s of ROOM TONE at the end (Ch.15) so gaps and patches sit in the same silence.
Follow the flow. The lav — clipped under the collar, hidden, close to the mouth — is your primary in most interviews because it stays consistently close no matter how the subject moves, and it rejects room reflections by sheer proximity. The boom or shotgun — just out of the top of the frame, angled down at the mouth, on its own channel — is your safety and often your better-sounding track, because a good boom in a good room has a fuller, more natural voice than a tiny lav capsule. Both land on separate channels of your recorder (or a camera with two inputs, or a second recorder run in double-system, Chapter 15). And you monitor both on headphones the entire time — not the meters alone, your actual ears (Chapter 15) — because a meter cannot hear a lav rustling against a collar or a fridge kicking on in the next room, and you can.
⚙️ Settings Box: interview audio — a starting point to adjust, not a recipe.
Setting Starting point Why Primary mic Lav, hidden under collar, ~a hand's width below the chin Consistent close pickup; immune to subject movement Safety mic Shotgun/boom just above frame, angled at the mouth Fuller natural tone; independent failure mode Channels Lav → CH1, boom → CH2 (separate) Two independent takes; pick the best per line in the edit Levels Peaks around -12 dBFS, never touching 0 dBFS Headroom protects against a sudden loud laugh or emphasis Monitoring Closed-back 🎧 headphones, both channels, whole take Your ears catch what meters can't (rustle, hum, room) Room tone 30 seconds, everyone silent, at the same setup Patches edits and covers gaps invisibly (Ch.15) Sample/format 48 kHz, 24-bit if available Standard for video; 24-bit gives extra headroom
Two failure modes deserve a specific warning, because they cost real footage. First, clipping: if you ride your levels too hot and the subject laughs, cries, or slaps the table for emphasis, the peak hits 0 dBFS and distorts — and distortion, unlike a quiet recording, cannot be repaired (Chapter 5's lesson about blown highlights applies to sound; a clipped peak is information that was never recorded). Leave headroom: aim dialogue around -12 dBFS so the loud moments have room to breathe below 0. Second, the unmonitored lav: a lav clipped once and never listened to is a bet that nothing rubs against it for the whole interview — and something always does. Monitor it, and re-dress or re-clip it the instant you hear friction.
⚠️ Common Mistake: one mic, no backup, no headphones. The single most expensive interview mistake is trusting one microphone and not listening to it. The camera's meters look fine; you find out in the edit that the lav was buzzing, or the on-camera mic captured a room full of echo, or a truck idled outside for your subject's best answer. There is no fix — the words are gone or ruined. The fix is the whole section: two mics on two channels, both in your headphones, levels with headroom, and 30 seconds of room tone. It is fifteen extra minutes of setup that saves the entire shoot.
🎒 Gear Note: the phone-and-cheap-mic path is real. You do not need a field recorder and two wireless kits to be safe. A workable two-mic-ish approach on no budget: run a wired lav into your phone or camera as your primary, and let a second phone (or the camera's own mic) capture the room as a rough safety and sync reference. It isn't a broadcast rig, but it gives you a second, independent recording of the same words — which is the entire point of two-mic safety. As you take paid work, a small two-channel recorder and a lav are the highest-value audio purchase you will make; until then, redundancy matters more than price.
🔄 Check Your Eye. 1. What is two-mic safety, and why do a lav and a boom make a safe pair rather than a redundant-but-useless one? 2. Where should dialogue peaks sit, and what happens if you clip? 3. Why is monitoring on headphones not the same as watching the meters?
Check yourself
- Recording the subject on two independent mics (lav + boom) on two channels so either can fail without losing the take. They're a safe pair because they fail in different, uncorrelated ways (rustle/dropout vs. off-axis drift/reflection), so both rarely fail on the same line.
- Around -12 dBFS with headroom below 0 dBFS. Clipping (hitting 0) distorts irreparably — the information was never recorded.
- Meters show level, not content. Your ears catch rustle, hum, echo, and background noise that a level meter reads as perfectly fine.
19.4 Question design: open, specific, quiet
Everything so far gets you a subject who is framed, lit, and recorded beautifully. None of it matters if they say nothing worth keeping. What they say is governed almost entirely by what you ask and how you ask it — a craft we call question design: writing and sequencing your interview questions so they produce usable, story-carrying answers instead of dead ends. Question design is where interviews are actually won or lost, and it is invisible to beginners, who think the interview is about the subject's eloquence. It isn't. A dull answer is almost always the fault of a dull question.
The first principle is open, not closed. A closed question can be answered with one word — "Did you like working there?" invites "Yes." An open question demands a built-out reply — "What was it like working there?" invites a story. Closed questions produce the non-answer: a reply that cannot stand on its own in the edit — a bare "yes" or "no," a fragment, an "it depends," an echo of your question, or an evasion. The non-answer is the raw material of mush; you can shoot for two hours and cut nothing. Train yourself to hear a closed question forming in your mouth and open it up before it lands. "Were you nervous?" becomes "Walk me through what was going through your head." "Is the process hard?" becomes "Take me through the process — where does it get hard?"
The second principle is specific, not vague. Big abstract questions produce big abstract answers, and abstraction is death on screen. "What does community mean to you?" gets you a paragraph of nothing. "Tell me about a time this neighborhood showed up for you" gets you a scene — a day, a person, a moment, the concrete stuff that cuts. The most powerful phrasing in all of interviewing is some version of "tell me about the day..." or "walk me through..." — because it asks for a specific narrated event, and events have detail, emotion, and a beginning-middle-end you can build a story from. When an answer goes abstract, pull it back to ground: "Can you give me an example?" is the follow-up that turns a platitude into gold.
The third principle is quiet. Your job during the answer is to disappear. That means the question itself is short — one question, not three stacked into a paragraph the subject can't hold in their head. It means you do not step on the beginning or end of their answer, because the editor needs clean air on both sides to cut (Theme 3 — you're shooting for the edit even while you talk). And crucially, it means you kill your verbal encouragements — the "mm-hmm," "right," "yeah" you'd normally say to show you're listening — because every one of them lands on top of the subject's audio and can make an otherwise perfect soundbite uncuttable. Nod instead. Smile. Lean in. Show you're listening with your face, silently, so the recording stays clean.
FIGURE 19.7 — Rewriting closed/vague questions into open, specific, cuttable ones
CLOSED / VAGUE (produces the non-answer) → OPEN / SPECIFIC (produces a story)
------------------------------------------ ---------------------------------------------
"Did the fire scare you?" → "Take me back to that night — what happened?"
"Is your job stressful?" → "Walk me through your hardest shift."
"Do you love what you do?" → "Tell me about a moment you knew this was it."
"What does resilience mean to you?" → "Tell me about a time you almost quit — and didn't."
"Was the community supportive?" → "Give me an example of the neighborhood showing up."
"So you started in 2010, right?" (leading) → "How did it begin?" (let them tell it, and date it)
There is a fourth trap to name: the leading question, which hands the subject your answer and gets it parroted back. "You must have been devastated, right?" gets you "Yes, I was devastated" — an echo, not a discovery, and it makes the piece feel coached. Ask the neutral, open version — "How did you feel?" — and let them find the word themselves; the word they choose is always more true, and more usable, than the one you fed them. (This is the same discipline as directing non-actors from Chapter 10: you create the conditions for a real response; you do not perform it for them.)
Sequence matters too. Open with easy, warming questions — name, what they do, an unthreatening bit of background — so the subject relaxes and forgets the camera before you get to the questions that matter. Put the hard or emotional questions in the middle, once trust is built and before fatigue sets in. And end soft, on something they're glad to talk about, plus the single most valuable closing question in interviewing: "Is there anything I didn't ask that you think I should have?" It routinely produces the best answer of the day, because it hands the subject the floor and they say the thing they came to say.
🎬 On Set: write ten, and the two that matter. For your Project 2 interview subject, write ten questions using FIGURE 19.7's open/specific/quiet rules. Then mark the two you most need answered for the story to work — your must-gets. Constraint: no closed questions, no leading questions, and at least three that start with "tell me about the time..." or "walk me through...". Self-review: read them aloud. Can any be answered "yes"? Rewrite it. Do any contain the answer? Neutralize it. Keep this list; you'll run it in the Production Checkpoint.
Interviewing a real person is not only a craft problem; it is an ethical one, and the standard is informed consent. Before you roll, the subject should understand what the piece is, where it will appear, and that they can pause, skip a question, or stop. For anything you intend to publish or be paid for, that understanding is formalized in a signed release — the model release that grants you the right to use their likeness and words — which we cover fully in Chapter 38; on a documentary or testimonial, get it signed before they leave. Handle sensitive subjects with care: a person reliving trauma is doing you a favor, not performing for you, and "getting the shot" never outranks their wellbeing. If a subject asks to stop, you stop.
♿ Accessibility & Inclusion. Two habits, woven in from the interview stage, not bolted on later. First, caption every interview — the roughly one in five viewers watching sound-off, and every viewer who is deaf or hard of hearing, cannot follow a talking head without captions, and a clean two-mic recording (§19.3) is what makes accurate captions possible (Chapter 36). Second, make the room accessible and the questions humane: choose a location your subject can physically reach and be comfortable in, offer questions in advance if that helps a nervous or neurodivergent subject give their best answers, and never treat someone's accent, disability, or way of speaking as something to "fix." The goal is their real voice, captured so everyone in your audience can receive it.
19.5 Listening and getting the answer in a full sentence
Question design gets the answer started. Active listening gets the real answer out — and gets it in a shape you can actually cut. Active listening is the interviewer's discipline of genuinely attending to what the subject is saying right now — not mentally rehearsing your next question — so you can follow the thread, ask the follow-up the moment invites, and use silence to let them go deeper. It is the difference between an interviewer running down a list and an interviewer having a conversation, and viewers can feel which one they're watching even when they can't say why.
The mechanics are simple and hard. Follow the thread, not the list. When a subject says something surprising, your written next question is now wrong — the right next question is "wait, say more about that." A list is a safety net, not a script; the best material almost always comes from the follow-up you didn't plan, prompted by something you'd have missed if you were reading ahead instead of listening. Echo to open, not to lead. Repeating a subject's own last few words back as a gentle question — "...and that changed everything?" — invites them to expand without putting words in their mouth. And above all, use the silence.
The pause is the interviewer's most powerful and most counterintuitive tool. When a subject finishes an answer, the amateur instinct — and it is a strong one, because silence feels socially unbearable — is to jump in with the next question. Resist it. Hold two or three seconds of silence, keep your eyes on them, and something remarkable happens: they fill it, and what they add is often the truest, most vulnerable, least rehearsed thing they say all day. The first answer is the prepared one; the thing that comes after the pause is the real one. Learning to sit in silence — to let it get slightly uncomfortable — is a craft skill you can practice, and it will get you material no question can.
FIGURE 19.8 — "The answer after the pause" [constructed teaching example]
THE FRAME Medium close-up, subject on a third, off-axis eyeline. They've just finished a composed,
slightly rehearsed answer and glance down. The frame holds — nobody speaks.
THE MOVE Locked off. The stillness of the shot mirrors the stillness of the room; the camera waits.
THE LIGHT The same soft key on the looking side; as they look down into the pause, the modeling deepens
and the moment turns inward.
THE SOUND Two seconds of near-silence — just room tone and a breath. Then, quietly, the real sentence:
the one they didn't plan, lower and slower than the answer before it.
THE CUT The editor will almost certainly cut the composed first answer and keep *this* — so it must be
a full, self-contained sentence. This is why the pause matters: it produces the keeper.
THE EFFECT The held silence makes the viewer lean in; the unrehearsed line lands as truth precisely
because it arrived slowly, after the performance dropped away.
THE LESSON Don't fill the silence — the subject will, and with the better answer. The interviewer's
hardest skill is saying nothing for three seconds.
Now the technical demand that shapes everything about how you listen: the answer must be a full, self-contained sentence, because in the edit your question disappears. When the piece is cut, the audience hears the subject, not you (§19.6 and Chapter 30). So an answer that depends on the question is useless. Ask "How long have you done this?" and if they say "Twenty years," you have nothing — cut into the piece, "Twenty years" answers a question no one heard. You need them to say, "I've been doing this for twenty years." That is a soundbite; "twenty years" is a fragment. This single fact reorganizes your whole technique, and it earns a threshold:
🚪 Threshold Concept: in the edit, the question disappears — so the answer must stand alone. Once you truly internalize that the viewer will never hear your question, you start engineering every answer to be self-contained. You ask the subject, up front, to put the question into the answer: "Instead of 'yes,' could you answer in a full sentence — 'I did feel that way because...'?" You re-ask, warmly, when they give you a fragment: "That's great — can you say the whole thing as one sentence for me?" You stop worrying about your own eloquence and start listening for whether each answer could be lifted out and dropped into the film whole. This is the mental shift from having a conversation to harvesting soundbites — and the best interviewers do both at once.
So, on set, do three things with every important answer. One: brief them at the top. Before you roll, tell the subject the one rule that will save your edit: "You'll never hear my questions in the final piece, so try to answer in full sentences that include a bit of the question — instead of 'yes,' say 'I do think so, because...'." Most people get it instantly and it transforms your footage. Two: re-ask for the clean version. When a great thought comes out tangled or as a fragment, don't move on — say, "That was perfect; give it to me one more time as a single sentence," and they'll hand you a cuttable gem. Three: get the emphasis and the pause. Leave air after their answers (both for the silence-technique above and so the editor has clean audio to cut on), and if a line is the heart of the piece, it's fine to say, "Can we do that one again? I want to make sure I got it" — you're not faking anything; you're getting a usable take of a true thing.
✂️ In the Edit. Everything in this section is you doing the editor's job in advance (usually the editor is you). A full-sentence answer is a clip you can drop straight onto the timeline; a fragment is a clip you have to nurse, or throw away. Clean air after answers is where your cuts live; talk over the tail and you've welded that soundbite to your own voice forever. The interview you conduct with the edit in mind is the interview that cuts itself — which is exactly the payoff we cash in Chapter 30, when you build a whole piece's spine out of nothing but these self-contained answers.
🔄 Check Your Eye. 1. Why must an interview answer be a full, self-contained sentence? 2. A subject just gave a composed answer and stopped. What should you do, and why? 3. What's the difference between echoing to open and a leading question?
Check yourself
- Because in the edit the question is cut out — the viewer only hears the subject. A fragment like "twenty years" answers an unheard question and can't stand alone; "I've been doing this for twenty years" can.
- Hold two or three seconds of silence rather than jumping to the next question — the subject usually fills it with the truer, unrehearsed answer.
- Echoing repeats their words back to invite expansion without adding content ("...and that changed everything?"); a leading question inserts your answer for them to parrot ("You were devastated, right?"), which reads as coached and less true.
19.6 The two-camera and single-camera interview
You have a framed, lit, well-recorded subject giving you full-sentence answers. The last setup decision is how many cameras cover them — and it changes how you shoot and how you'll cut. The core problem both approaches solve is the same one: an interview edit is full of cuts within a single continuous answer — you trim out an "um," tighten a rambling middle, join the front of one sentence to the end of another. Cut like that on one unbroken shot and you get a jump cut (Chapter 9): the subject's head visibly leaps as two nearly-identical frames butt together. The whole game of shooting an interview is giving yourself a way to hide those jumps.
The single-camera interview covers the subject from one angle. It is the phone-first, solo, run-and-gun reality, and it's completely viable — but you must plan for the jump cuts. Three ways to hide them, all of which you already know or will meet next chapter: (1) lay B-roll over the cut — the single most common solution, and the entire reason B-roll exists (Chapter 20); when you cut within an answer, you drop a cutaway over the seam and the interview audio continues underneath (an L-cut). (2) Punch in. If you're shooting in 4K and delivering in 1080p, you can crop one continuous take into a "wide" and a "tight" version in post, and cut between the two sizes to hide a trim — a free second angle from a single camera. (3) Reframe between questions. Stop between topics, physically change your shot size or angle (respecting the 30° rule so the change reads as a new shot, not a jump — Chapter 9), then resume; now you have genuinely different sizes to cut between. The cost of single-camera is simply that you must have enough B-roll (Chapter 20) — without it, a single-camera interview with lots of internal cuts becomes a "hostage video" of jumps.
The two-camera interview covers the subject from two angles at once, and it is the professional standard for a reason: it lets you cut within an answer invisibly, with no B-roll required, by cutting from Camera A to Camera B on the trim. Set it up right and the two cameras obey the same rules that make any coverage cut (Chapter 7): put both cameras on the same side of the line (the 180-degree rule — the subject's eyeline stays consistent) and offset them by at least 30° in angle and ideally a step in size (say, A is a medium, B is a close-up), so cutting between them reads as a deliberate shot change, not a jump. The classic arrangement: A-cam is your safe, wider "spine" shot straight-ish on; B-cam is a tighter, more side-on angle for emphasis and for hiding edits.
FIGURE 19.9 — Two-camera interview (top-down): same side of the line, 30°+ apart
( S ) ──────► eyeline (off-axis, to interviewer)
/ \
/ \
≥30° / \
/ \
[ CAM B ] [ CAM A ] ( I ) interviewer, beside CAM A,
tighter, wider "spine", near the lens the subject looks past
more side-on straighter on
\ /
\___________/
BOTH cameras on the SAME side of the subject's eyeline (don't cross the line);
A and B differ by ≥30° and by a shot size → cutting A↔B hides internal edits.
Match white balance, exposure, and height between cameras so the cut doesn't jar.
Two things make or break the two-camera setup, both learnable. First, match the cameras: same white balance, same exposure, same eye height, ideally similar picture profile, so cutting from A to B doesn't jump in brightness or color (the alternative is a painful matching job in Chapter 31). Second, sync them: run a clap or use timecode so the editor can line the two angles up frame-accurate (Chapter 15's double-system discipline, paid off in Chapter 27). Put the interviewer beside Camera A (usually the wider one) so the eyeline stays consistent and the subject isn't ping-ponging between two operators.
✂️ In the Edit. This is the cleanest set-to-timeline trade in the book. Two cameras buy you invisible cuts; one camera buys you a B-roll obligation. With two matched, synced, 30°-offset angles, you can tighten any answer by cutting A↔B and the audience never sees the seam — the interview practically edits itself. With one camera, every internal trim is a debt you pay in Chapter 20 by shooting enough B-roll to cover it. Neither is wrong; you're choosing where to spend the effort — on set with a second camera, or later with a cutaway. Decide before you shoot, because if you plan single-camera and skimp on B-roll, the edit has no way out.
🎬 On Set: shoot the same answer two ways. Record one subject answering one good question, first single-camera (one locked medium), then re-ask it two-camera if you have a second body/phone (a medium and a close-up, same side of the line, 30° apart). In the edit, try to tighten each version by 20%. Self-review: which internal cuts could you hide, and how? The single-camera version will need B-roll or a punch-in; the two-camera version will hide its cuts on the A↔B switch. That contrast is the lesson of this section.
One honesty note carried over from §19.4: the classic single-camera trick of shooting "noddies" — cutaways of the interviewer nodding, filmed after the subject leaves — is a legitimate editing tool, but it has an ethics line. Using a noddy to bridge a cut is fine; using one to fabricate a reaction that implies something false — a skeptical nod stitched onto an answer to editorialize — is the kind of dishonest edit documentary ethics (Chapter 21) exists to rule out. The tools are neutral; what you imply with them is not.
Production Checkpoint
Project 2 — shoot the interview for your documentary short. This is the centerpiece of your 3-minute documentary, and everything in this chapter converges on it. Using your pre-production plan (the treatment and shot list from Chapter 16, the story arc from Chapter 17, the call sheet from Chapter 18), shoot a real interview with your subject that clears four bars:
- Framing: subject on a third, at eye height, with correct headroom and looking room, and an off-axis eyeline — you sitting right beside the lens so the gaze reads as candid conversation (§19.1, FIGURE 19.2).
- Motivated light: a soft key on the looking side (a window or one lamp is plenty), a controlled ratio for the tone you want, and some separation from the background (§19.2).
- Two-mic-safe audio: a lav and a boom (or your best two-independent-recordings approximation), on separate channels, monitored on headphones, peaking around -12 dBFS, plus 30 seconds of room tone (§19.3).
- Full-sentence answers: a designed question list — open, specific, quiet — run with active listening and the silence technique, with the subject briefed to answer in self-contained sentences (§19.4–19.5).
Why this matters: this single sit-down is the spine your entire documentary will hang on. In Chapter 20 you'll shoot the B-roll that covers and illustrates it; in Chapter 30 you'll build the whole piece out of the self-contained answers you capture today. If you get gold here — a subject who was comfortable, well-lit, cleanly recorded, and talking in cuttable sentences — the edit will feel like discovery. If you get mush, no later chapter can rescue it. Also: get your subject's release signed before they leave (Chapter 38), and note whether you shot single- or two-camera, because that decides how much B-roll you owe yourself next.
Summary
Reference-grade recap. Screenshot this before an interview shoot.
The interview setup (build this every time):
| Element | Default | The reason |
|---|---|---|
| Frame | Medium / medium-close-up, subject on a third | Face is where the meaning is; a third feels alive |
| Eye height | Lens at subject's eye level | Neutral, respectful; avoids up-the-nose or looking-down |
| Eyeline | Off-axis — subject looks at interviewer just beside the lens | The candid "conversation" register of documentary |
| Interviewer position | As close to the lens as possible | Small off-axis angle = near-eye-contact, not lost profile |
| Looking room | More space on the side they face | Composed, not cramped or mugshot-centered |
| Background | Deep, soft, darker than the subject, nothing growing from their head | Separation and dimension, not a flat wall |
Lighting (apply Part III): key on the looking side, soft, just above eye level, motivated; fill to a ratio that matches tone (2:1–3:1 doc-warm; ~1:1 corporate-clean; 4:1+ only for weight); rim + a deeper background for separation. The window is a free, professional interview key.
Audio — two-mic safety: lav (primary) + boom (safety) on two separate channels; monitor both on headphones; peaks ~-12 dBFS with headroom below 0; 30 s room tone. No single failure should cost the take.
Question design — open, specific, quiet:
| Do | Instead of |
|---|---|
| "Walk me through..." / "Tell me about the day..." | "Did you...?" / yes-no closed questions |
| A specific example | A vague abstraction ("what does X mean to you?") |
| One short question, then silence | Three stacked questions in one breath |
| Neutral phrasing; let them find the word | Leading ("You were devastated, right?") |
| Nod silently; kill the "mm-hmm"s | Verbal encouragements over their audio |
| Close with "Anything I didn't ask?" | Ending abruptly on your last listed question |
Getting the answer: the question disappears in the edit, so the answer must be a full self-contained sentence — brief the subject up front, re-ask for the clean version, and use silence to get the truer answer after the pause. Watch for the non-answer (yes/no, fragment, echo, evasion) and open it up on the spot.
One vs. two cameras:
| Single-camera | Two-camera | |
|---|---|---|
| Hides internal cuts with | B-roll / punch-in (4K→1080p) / reframe between Qs | Cutting A↔B (invisible, no B-roll needed) |
| Requires | Enough B-roll (Ch.20); the 30° rule when reframing | Same side of line, ≥30° apart + a size step; matched WB/exposure; sync |
| Best for | Solo, run-and-gun, phone-first | Paid work, long interviews, minimal B-roll |
Themes surfaced: story is the boss (ratio and eyeline chosen for feeling); sound is half the picture (two-mic safety); you shoot for the edit (full-sentence answers, two-camera cuts, clean air); fix it in pre, not in post (light in front, not a window behind); motivate every choice (key on the looking side).
Spaced Review
Bringing back the two chapters this one is built on — retrieve before you read the answers.
- (Ch.11) Name the three lights of three-point lighting and the job of each — and which one you'd sacrifice first if you only had one light.
- (Ch.11) What does it mean for a light to be motivated, and why does keying an interview from the subject's looking side read as motivated?
- (Ch.14) Why is a lav clipped close to the mouth so good at rejecting room echo, and how does that differ from what a shotgun's polar pattern does?
- (Ch.14) State the threshold idea from Chapter 14 in one sentence — and give one reason an interview proves it.
- (Ch.11 & 14 together) A subject is beautifully lit but the audio echoes badly. Which problem loses you the viewer faster, and why?
Check yourself
1. **Key** (primary shaping light, sets exposure and modeling), **fill** (lifts the shadow side, sets the contrast ratio), **rim/back** (separates subject from background). Sacrifice the *fill* first — you can live with a deeper shadow, and even bounce a little light back for free; losing the key leaves you with no shape, losing the rim only costs separation. 2. Motivated light appears to come from a source that makes sense in the scene (a window, a lamp, the conversation). Keying from the looking side reads as motivated because the light seems to come from where the subject's attention and the interviewer are — the eye accepts it as natural. 3. Proximity: the lav is inches from the mouth, so the direct voice is far louder than reflected room sound (inverse-square law working for you). A shotgun instead uses a tight polar pattern to *reject* off-axis room sound from a distance — same goal, different mechanism. 4. "Sound is half the picture" — audiences forgive a soft image but leave over bad audio in seconds. An interview proves it: a perfectly lit talking head with echoey, buzzing audio is unwatchable, while a plain frame with clean, close sound feels professional. 5. The audio. Viewers tolerate an imperfect image far longer than bad sound; echoey dialogue reads as "amateur" and drives them off within seconds, no matter how good the lighting is.What's Next
You now have the backbone of your documentary in the can: a real person, framed and lit and cleanly recorded, telling you true things in sentences you can cut. But an interview alone — even a great one — is a "hostage video": a head talking, with no way to breathe, illustrate, or hide a single edit. The thing that turns those answers into a film is the footage you lay over them. In Chapter 20 we shoot B-roll and coverage — the cutaways, inserts, and sequences that let you cut invisibly, show what the subject is describing, and give the viewer's eye somewhere to go. You'll leave Part IV with everything a real documentary needs: the interview from this chapter, and the pictures that set it free.