Case Study 2: Shoot-Along — A Talking-Head Interview for a Documentary Short
This is a production walkthrough, not an analysis. You are going to follow one interview from an empty room to a rough-cut timeline, watching every decision get made and adjusted in real time. The setup is the book's anchor talking-head at a table (Chapter 1's four setups), and everything in it is doable by one person with a phone, one light or a window, and a clip-on mic. Mirror it for your own Project 2 and you will have shot the centerpiece of your documentary by the end.
Our subject for the walkthrough is a neighborhood bicycle-repair shop's owner-mechanic — a generic, unnamed real person, the kind of craftsperson a 3-minute documentary short is built around. Everything below applies unchanged whether your subject is a baker, a nurse, a luthier, or your grandmother. The subject is not the variable. The craft is.
The brief and the constraints
The brief: capture a 6–8 minute interview that will cut down to roughly 90 seconds of self-contained soundbites carrying a 3-minute documentary about a small shop that has fixed the neighborhood's bikes for decades. We need: the origin story, one specific vivid memory, the emotional "why," and a strong closing line. We are shooting single-location, in the shop's back office, in a two-hour window.
The constraints — and they are typical:
- One shooter (you), maybe one helper. No crew.
- A real working room, not a studio — a cramped office with a small window, fluorescent overheads, and street noise through a thin wall.
- A subject who is not a performer, is a little camera-shy, and has a customer coming in two hours.
- Modest kit: one camera (or a phone), one light or the window, a lav, and a shotgun/recorder. That is enough. It is, in fact, plenty.
Notice how much of this brief is pre-production (Theme 5). We know the four things we need before we arrive, which means we can build a question list aimed at exactly those, and we know the room's problems (fluorescents, street noise, small window) before we walk in, which means we can plan around them. A vague brief — "go interview the bike guy" — is how you come home with two hours of mush. A specific one is how you come home with gold.
Gear and the settings
Here is the whole kit, and the starting settings we will dial in on the day. Treat the box as a launch point, not a recipe — we will adjust several of these once we see the room.
⚙️ Settings Box: single-shooter documentary interview (starting point).
Domain Setting Why / adjust-when Camera 4K, 24 fps, 180° shutter (1/48–1/50 s) 4K lets us punch in for a "second angle" (§19.6); 24 fps for the doc look Lens / focal length ~50–85mm equivalent, aperture ~f/2.8 Flattering compression + soft background separation Framing Medium close-up, subject on a third, eye-height The interview default (§19.1) White balance Set manually to the key (e.g. ~5600K daylight if window-keyed) Avoid auto-WB drifting mid-answer; match to dominant source Exposure Expose for the face; protect highlights on the window Use zebras/waveform (Ch.5); face is the priority Primary mic Lav, hidden under collar, ~a hand's width below the chin → CH1 Consistent close pickup regardless of movement (§19.3) Safety mic Shotgun on a stand/boom, just above frame, angled at mouth → CH2 Independent failure mode; often the better tone Levels Peaks ~-12 dBFS, never at 0; monitor both on 🎧 Headroom for laughs/emphasis; ears beat meters (Ch.15) Room tone 30 s at wrap, everyone still Patches every edit invisibly (Ch.15, Ch.30) Second angle Punch-in from 4K, or a second phone as B-cam Hides internal cuts without B-roll (§19.6)
The single most important line in that box is the audio one: two mics, two channels, both in the headphones. If you forget everything else, remember that. A perfectly framed, beautifully lit interview with one failed mic is worthless; a plain frame with clean, safe sound is usable. Sound is half the picture (Chapter 14), and on a one-take interview it is the half most likely to sink you.
The setup: building the room
We arrive, and the office is exactly the problem we expected: a desk against a wall, one smallish window on the left wall, harsh fluorescent tubes overhead, and a thin wall to a noisy corridor on the right. Before touching a camera, we solve the room, in this order: sound, light, then frame.
Sound first, because it's the hardest to fix later. We kill the fluorescents (they can buzz, and they're ugly light anyway) and the HVAC if we can. We move the subject's chair away from the noisy corridor wall and toward the quieter interior, so the street noise is behind the mic's rejection, not in front of it. We drape a coat over a hard shelf to cut one slap-back reflection. None of this is gear; it's just choosing where in the room to sit, which is a §19.3 decision made with our ears.
Light second. The window becomes our key (Chapter 13, the window as a free softbox). We seat the subject so the window is on their looking side — they'll face toward it and toward us (§19.2). If the day is dim, we add one soft LED just above and beside the lens on the same side, matched to daylight, as the real key and let the window fill. We bounce a little light back into the camera-near shadow with a white board. We do not light the wall behind them — we let it fall dark for separation, and we place a small desk lamp in the background, out of focus, for a warm point of depth.
Frame third, because now the room supports it. Camera on a tripod at the subject's eye height, subject on the right third facing left, looking room across the open frame, medium close-up.
Here is the room as a top-down map — the exact FIGURE 19.2 arrangement, adapted to this office.
FIGURE CS2.1 — The shoot-along setup (top-down): window-key talking head in a small office
[ WINDOW ] ☀ (looking side; soft key)
\
↘ ▽ desk lamp (background, out of focus, warm depth)
( S ) ──────► eyeline
/ | \
bounce ○ | · (dark wall behind — separation, no light on it)
(fill, near |
shadow) ((• shotgun on a stand, just out of top of frame → CH2
|
((• lav clipped under collar → CH1
|
[ CAM ] 4K/24p, ~f/2.8, eye height, subject on right third
|
▼
( YOU ) seated right beside the lens ← quiet corridor behind camera,
noisy wall now BEHIND the subject's
mic rejection
Two hours, one shooter, one window, one light, two mics. This is a complete documentary interview rig.
Now the framing, checked against the overlay we learned in §19.1:
FIGURE CS2.2 — Framing check for the shoot-along (16:9)
+---------------------------------------------------+
| . . |
| | window glow | |
| O·······|····( eyes on )····|···················| <- upper third: eyes land here
| | ( head ) | |
| LOOKING | ( shoulders ) | dark wall + |
| ROOM → | | | | soft desk lamp |
| O | | O (separation) |
| . | | . |
+---------------------------------------------------+
subject on RIGHT third, facing LEFT into the looking room; a hand of headroom;
background dark and deep on the right → dimensional, not flat.
We frame it, set exposure for the face (letting the window sit a touch hot but not fully blown, watching zebras from Chapter 5), lock white balance to daylight so it can't drift mid-answer, and clip the lav. Then — the step beginners skip — we sit down beside the lens and put on the headphones, and we listen. We hear the lav is a hair muffled under a thick collar, so we re-dress it higher and more exposed. We hear a faint hum, hunt it down to a mini-fridge, and unplug it. Now both channels are clean in our ears. Only now are we ready to talk.
The thinking: we spent our first thirty minutes and did not roll a frame of interview. That is correct. On a two-hour window, the setup is not stolen from the interview time — it is the interview, because a badly-set-up interview is unusable no matter how good the answers. Fix it in pre and at setup, not in post.
The shoot, phase by phase
Phase 1 — settle the subject (the first five minutes are not for keeping)
We roll, but the first questions are throwaways by design (Chapter 10, getting a relaxed performance; §19.4's warming sequence). "Tell me your name and what you do here." "How long has the shop been open?" We are not fishing for gold yet; we are letting the subject get used to the lens, the light, and the sound of their own voice, and we are checking our levels on real speech. We also deliver the single most important brief of the day:
"You'll never hear my questions in the finished film — just you. So try to answer in complete sentences that fold my question in. If I ask how long you've been open, instead of 'thirty years,' say 'The shop's been open for thirty years.' Don't worry if you forget; I'll just ask you to say it again."
Most subjects get this instantly, and it transforms the footage from fragments into cuttable soundbites (§19.5). We watch our headphones and our waveform through these warm-ups; the subject laughs at one point and the peak jumps to -6 dBFS — good, our -12 target left the headroom to catch it clean.
FIGURE CS2.3 — "The warm-up take" [constructed teaching example]
THE FRAME Medium close-up, subject on the right third, window glow on the left of their face, a hand of
headroom, dark shop wall falling soft behind with a warm lamp bokeh in the corner.
THE MOVE Locked off, eye height. Still and calm — the calm is partly to settle the nervous subject.
THE LIGHT Window key from the looking side, soft and slightly cool; white bounce lifting the near cheek;
the background a stop and a half down for separation.
THE SOUND Two clean channels — a close lav and a fuller shotgun — the voice a little stiff still, warming.
THE CUT Won't be used, and that's the point; it's the runway. The keeper takes come once they forget the camera.
THE EFFECT We can see the shoulders drop over the first two minutes; the setup is doing its job of making
an anxious person comfortable enough to be honest.
THE LESSON Spend the first five minutes making the subject comfortable, not making the film. Comfort is
what buys you the real answers later.
Phase 2 — the origin story (open, specific, quiet)
Now we go for our first must-get. Not "why did you open the shop?" (abstract, invites a shrug) but the specific, narrated version:
"Walk me back to the day you decided to open this place. Where were you, what was going on in your life?"
This is question design (§19.4) doing its work: it asks for a scene, not an opinion, so we get a day, a place, a feeling. The subject starts talking, and we do the hard part — we shut up. We nod, we hold eye contact beside the lens, we do not say "mm-hmm," and when they reach what feels like the end, we wait three seconds. Into that silence, they add the real line: "...honestly, I was terrified I'd fail. I just didn't want to work for anyone else anymore." That unrehearsed sentence is worth more than the whole prepared answer before it, and it only exists because we left the silence open (§19.5).
FIGURE CS2.4 — "The answer after the pause" [constructed teaching example]
THE FRAME Same setup; the subject looks down slightly as they drop from the rehearsed story into the
honest afterthought — the eyeline breaks toward the desk, then back up to us.
THE MOVE Locked off. We do not push in or adjust; stillness lets the small human moment carry itself.
THE LIGHT Unchanged; as they look down, the window key models the face a touch more, the moment turns inward.
THE SOUND A breath, two seconds of clean room tone, then the quiet line — lower and slower than the story
before it. Both channels clean; the lav catches the intimacy, the shotgun the room's air.
THE CUT This is a keeper. It's a full, self-contained sentence ("I was terrified I'd fail...") that can
drop straight onto the timeline with no question needed.
THE EFFECT The confession lands as truth precisely because it arrived after the performance dropped away.
THE LESSON The best soundbite of the interview came from silence, not from a question. Don't fill the gap.
Phase 3 — the vivid memory and the emotional why
We keep climbing the ladder into the material that matters, now that trust is built. "Tell me about a repair you'll never forget." (A specific request; it produces a story about a kid's first bike, or a race saved, or a regular who became a friend.) When an answer goes abstract — "it's just about community, really" — we pull it back to ground with the most reliable follow-up in interviewing: "Can you give me an example?" And the platitude turns into a scene.
One answer comes out brilliant but tangled — the subject stumbles over the middle of a beautiful thought. We don't move on. We say, warmly: "That was perfect — can you give me that one more time, as a single sentence?" And they hand us a clean, cuttable version of a true thing (§19.5). We are not faking; we are getting a usable take of something real.
The thinking: every one of these is a §19.4/§19.5 move — open question, specific request, the "give me an example" rescue, the "say it as one sentence" re-ask, and above all the discipline of listening to this answer instead of reading ahead to the next question. The list is a safety net, not a script; two of our best exchanges came from following a thread we didn't plan.
Phase 4 — the close, and the magic question
We wind down with easy, warm questions so the subject leaves feeling good, and then we ask the single most valuable closing question in interviewing (§19.4):
"Is there anything I didn't ask that you think I should have?"
And — as it does more often than not — it produces the best line of the day. The subject pauses, then says the thing they actually came to say: something about hoping the shop outlasts them, about the neighborhood changing. It's the closing line of our documentary, and we would never have gotten it from our planned list. We hand the subject the floor, and they gave us our ending.
Phase 5 — the second angle and the wrap
Because we're single-camera, we protect the edit two ways (§19.6). First, we shot in 4K, so we can punch in to a tighter "B-angle" from the same footage in post to hide internal cuts. Second, for two or three key answers, we quietly ask the subject to say them once more and we reframe tighter between takes (past 30°, per Chapter 9) so we have a genuinely different size to cut to. If we'd had a second phone, we'd have run it as a B-cam the whole time.
Then the two things every professional does at wrap and every amateur forgets:
- Room tone. "I need thirty seconds of everyone completely silent — this is the most important quiet of the day." We record 30 seconds of the room doing nothing (Chapter 15), which will patch every edit invisibly.
- The release. Before the subject returns to their customer, we get the model release signed (Chapter 38) — the piece of paper that lets us actually use everything we just shot. No release, no film.
The thinking: the shoot took ninety minutes of our two-hour window, and we leave with clean, safe, well-lit, full-sentence answers, a magic closing line, a second angle, room tone, and a signed release. Nothing here required a crew or a studio. It required a plan, a window, two mics, and the discipline to listen and wait.
The edit pass
We are not cutting the whole documentary here (that's Chapter 30) — but let's do the first, revealing pass that proves the shoot worked: pulling the self-contained soundbites and laying the spine. In the edit, our question disappears, so every clip must stand alone — and because we briefed the subject and re-asked for clean versions, they do.
FIGURE CS2.5 — Rough assembly: the interview spine before B-roll (edit timeline)
V1 [ MCU: origin ###### ][ punch-in: "terrified" ##### ][ MCU: memory ####### ][ tight: close line #### ]
A1 [ lav (CH1) ################################################################################## ]
A2 [ shotgun (CH2) safety .......... used here → ][ ............................................. ]
A3 [ room tone bed ............................................................................... ]
^cut within answer: ^cut within answer: ^ends on the
hidden by punch-in hidden by size change "magic question" line
The spine is ~85 seconds of full-sentence soundbites. Every internal cut is hidden by a punch-in or a
reframe (§19.6); the lav is primary, the shotgun patches one rustled line; room tone fills the gaps.
NEXT (Ch.20): lay B-roll of the shop, hands, bikes over every cut → the documentary breathes.
Read the timeline. The interview audio (the lav on A1) is the spine; nothing else is decided yet. Where we cut within an answer to tighten it — trimming a stumble out of the origin story — the picture jump is hidden by cutting to the 4K punch-in or the reframed tighter take (§19.6), so there's no jump cut (Chapter 9). On one line, the lav caught a faint clothing rustle, so we drop to the shotgun safety (A2) for that sentence — two-mic safety paying off exactly as designed (§19.3). The room tone (A3) runs underneath so the silences between soundbites are the same silence, not a series of clicks. And the whole thing ends on the "magic question" line, which will be our documentary's closing beat.
Before we could build that spine, we did one unglamorous step that makes everything faster: we transcribed the interview. With two clean audio channels, an auto-transcription (Resolve, a platform tool, or a dedicated app) turned six minutes of talk into a searchable page of text. We read it, not the footage, and highlighted the self-contained sentences — the "terrified I'd fail" line, the vivid-memory story, the closing beat — then built the spine from those highlights. This is the paper edit we formalize in Chapter 30, and it is only possible because the audio is clean and the answers are full sentences. It is worth noticing how the whole chain pays off here: the full-sentence brief (§19.5) made the transcript's highlights liftable; two-mic safety (§19.3) made the transcript accurate; and the designed questions (§19.4) meant the highlights were worth lifting. A sloppy interview produces a transcript full of fragments and "um"s that no highlighting can rescue. The edit was easy because the shoot was disciplined — which is the entire argument of this book in one timeline.
The key edit decisions, named:
- Cut for meaning, not chronology. We led with the "terrified I'd fail" confession, not the literal first thing the subject said. The paper edit (Chapter 30) arranges soundbites by story, not by shoot order.
- Every internal cut is hidden — by a punch-in, a size change, or (next chapter) B-roll. No naked jump cuts.
- The safety mic saved one line. Without the second channel, that sentence — a good one — would have been lost to a rustle. The fifteen extra minutes of two-mic setup paid for itself in one cut.
- Room tone made the gaps invisible. The spine breathes because the silence is continuous.
This rough spine is exactly what you'll hand to Chapter 30 to finish, and exactly what Chapter 20's B-roll will cover. If your own Project 2 interview produces a timeline that looks like FIGURE CS2.5 — full-sentence soundbites, hidden internal cuts, a safety track you can fall back on, room tone underneath — you have shot a real documentary interview.
Discussion questions
- We solved the room in the order sound, light, frame. Why that order? What goes wrong if you frame first and deal with audio last?
- The first five minutes were deliberately throwaway. Defend spending shooting time on takes you know you won't use.
- On one line the lav failed and we used the shotgun. Walk through what the edit would have looked like without two-mic safety. Was the redundancy worth the extra setup?
- We got our best two lines from silence and from the "anything I didn't ask?" question — neither from our planned list. What does that say about how tightly you should script an interview?
- We're single-camera and hid our cuts with a 4K punch-in and reframes. Compare the effort/cost of that against simply running a second camera. When would you choose each?
- The release got signed before the subject left. What exactly is lost if you skip it and "get it later"?
Your turn: run this shoot for Project 2
This walkthrough is your Production Checkpoint. Run it for real with your own documentary subject.
The brief: shoot a 6–8 minute interview that will cut to ~90 seconds of self-contained soundbites, in one location, in a two-hour window, single-shooter.
Hit every mark from this case study:
- Solve the room in the order sound → light → frame, using the FIGURE CS2.1 setup adapted to your space.
- Build the frame and eyeline to FIGURE CS2.2: subject on a third, eye height, off-axis eyeline with you right beside the lens.
- Run two-mic safety (lav + shotgun, two channels, both in headphones, ~-12 dBFS) and get 30 seconds of room tone.
- Brief the full-sentence rule, warm up for five minutes, then run a designed question list of open, specific, quiet questions.
- Listen and wait. Use the three-second silence on every important answer. Ask "can you give me an example?" and "say it as one sentence" when you need to.
- Close with the magic question, protect the edit (punch-in or reframe), and get the release signed before they leave.
Then do the first edit pass: pull your self-contained soundbites into a spine like FIGURE CS2.5, hide the internal cuts, and confirm your safety track and room tone are there. If you can build that spine, you're ready for Chapter 20's B-roll — and your documentary is real.
Key takeaways
- Solve the room before the camera, in the order sound → light → frame. On a real location, where you sit the subject solves more problems than any setting.
- Two mics, two channels, both in the headphones. It is the one non-negotiable of the interview shoot, and in this walkthrough it saved a line the lav lost to a rustle.
- Spend the first five minutes on comfort, not footage. A settled subject gives the real answers; the warm-up is the runway that earns them.
- The full-sentence brief transforms your raw footage from fragments into cuttable soundbites — because the question disappears in the edit.
- The best material comes from silence and from the "anything I didn't ask?" close — not from the planned list. The list is a safety net; listening is the craft.
- Protect the edit on set (punch-in, reframe, or a second camera) so internal cuts have somewhere to hide, and record room tone and get the release before you wrap.
- Every bit of this is doable with a phone, a window, and one clip-on mic. The setup, the questions, and the listening are the whole job — and none of them is for sale.