You have twelve minutes of a real person telling you true things. Somewhere in that twelve minutes is a three-minute film — a tight, moving, watchable story. But it is not sitting there waiting to be trimmed down to size, the way a sculptor's statue...
Prerequisites
- 19
- 28
Learning Objectives
- Build a radio edit — cut the interview words first, from the transcript, before touching a frame of picture.
- Pull selects from raw interview footage and lay them into a string-out that becomes the story spine.
- Assemble a soundbite spine that tells the whole story in audio, ordered for the arc rather than for chronology.
- Layer B-roll over every interview seam so the cuts vanish and the picture proves what the voice claims.
- Pace an interview piece by alternating the face and the B-roll, holding the face for the emotional lines.
- Take an interview-driven piece to a fine cut and lock it, running the mute and eyes-closed quality passes.
In This Chapter
- Overview
- Learning Paths
- 30.1 The paper/radio edit: story before picture
- 30.2 Pulling selects and the string-out
- 30.3 Building the spine from soundbites
- 30.4 Layering B-roll to hide cuts and show
- 30.5 Pacing an interview piece
- 30.6 The fine cut and lock
- Production Checkpoint
- Summary
- Spaced Review
- What's Next
Chapter 30: Editing the Interview-Driven Piece
Overview
You have twelve minutes of a real person telling you true things. Somewhere in that twelve minutes is a three-minute film — a tight, moving, watchable story. But it is not sitting there waiting to be trimmed down to size, the way a sculptor's statue supposedly waits inside the block. It is scattered. The best sentence of the whole interview is nine minutes in, right after a rambling non-answer. The line that should open the film is the last thing your subject said, almost as an afterthought, when they thought the camera was off. The middle of the story is spread across four different answers to four different questions, and none of them, on its own, makes sense. The raw interview is not a rough draft of your film. It is the quarry. The film is something you will build out of it, one soundbite at a time, in an order the subject never spoke and could not have.
That building is the craft of this chapter, and it is a craft of its own — different in kind from cutting a scripted scene. When you cut the Café Scene in Chapters 28 and 29, you had a shot list and coverage: a plan on the page that told you what every shot was and where it went. An interview-driven piece has no such plan, because the story does not exist yet. It has to be found in what a nervous human being happened to say, and then constructed from fragments, and then hidden — because the moment you start joining those fragments, the picture jumps and betrays every join. This is the hardest and most common edit in all of nonfiction video, and almost every documentary, testimonial, brand film, explainer, and oral history you have ever watched was made this way: a spine of talking-head audio, with pictures laid over the seams.
The method has four moves, and they are the four things this chapter teaches. First you find the story in the words — a radio edit, cutting the audio first with your eyes closed to the picture. Then you pull your keepers into selects and lay them into a string-out. Then you order those keepers into a spine that carries the whole story in sound. And then — only then — you lay B-roll over every seam to make a B-roll-driven narrative: a film whose voice leads and whose pictures prove. Get this sequence right and a pile of rambling answers becomes a story that holds a stranger to the end. Get it wrong — reach for the pretty pictures before the words work — and you get the thing everyone recognizes and no one watches: a handsome, hollow video that never quite says anything.
This is also the chapter where Project 2 locks. You shot the interview in Chapter 19 and the B-roll in Chapter 20; you ingested, synced, and strung it out in Chapter 27; you learned the grammar of the cut in Chapter 28 and how to cut for feeling in Chapter 29. Now you finish it — build the spine, layer the B-roll, and freeze the fine cut. By the end of this chapter your documentary short is a film.
In this chapter you will learn to:
- Run a radio edit — cut the interview's words first, from a transcript, before you touch a frame of picture, and prove the story works for the ear alone.
- Pull selects from raw interview footage and lay them into a string-out that becomes your first watchable pass.
- Build the spine — the ordered sequence of soundbites that carries the whole story — for the arc, not for chronology.
- Layer B-roll over every interview seam to create a B-roll-driven narrative where the cut vanishes and the picture proves the point.
- Pace an interview piece by alternating the face and the B-roll — and knowing when to hold on the face.
- Take the piece to a fine cut and lock it, running the two quality passes that catch what your tired eyes miss.
Learning Paths
This is the capstone edit of Part VI — everyone who wants to finish a documentary needs the whole method. Weight your attention like this:
- 📱 Phone-first: §30.1 (the radio edit) and §30.4 (layering B-roll) are your core, and they cost nothing — a text document and a free editor. A single-camera phone interview depends on B-roll to be cuttable at all, so §30.4 is not optional for you; it is survival.
- 🎥 Creator: §30.3 (the spine) and §30.5 (pacing) are what keep a talking-head video from losing viewers. If you interview guests or cut your own long-form to camera, this is your retention chapter.
- 💼 Pro-track: all of it, and especially §30.1 (the radio edit is how you cut a client's rambling CEO into ninety persuasive seconds) and §30.6 (lock and hand-off). This is the single most billable edit in the business.
- 🎓 Student: the method in order — radio edit → selects → spine → B-roll → pace → lock. The Production Checkpoint runs all of it on your documentary, and it is the exam.
30.1 The paper/radio edit: story before picture
Open your interview in the timeline, start scrubbing, and you will make a specific, seductive mistake within about ninety seconds: you will find a beautiful shot — a moment where the light catches your subject just right, or a gesture you love — and you will start building the film around it. You are now editing pictures. And you are lost, because you have let the footage decide the story instead of deciding it yourself. The single most important discipline in interview editing is to refuse that temptation entirely at the start — to not look at the pictures at all until the words already tell the story.
In Chapter 29 you met the tool for this in its general form: the paper edit, working out the shape of a piece on paper before you touch the timeline. The interview-driven piece has a specialized, more radical version of it, and it is the beating heart of this whole chapter. A radio edit (also called a radio cut) is a paper edit built entirely from the interview's audio — you find and order the story using only the spoken words, with your eyes closed to the picture, so that the piece works as a piece of radio before a single frame of image is laid over it. The name is the whole philosophy: if the story holds up with nothing but the voice — the way a great radio documentary holds you with no picture at all — then the pictures you add later can only make it better. If it does not hold up as audio, no amount of gorgeous B-roll will save it; you will have built a beautiful slideshow that says nothing.
Here is why this works, and why it is the opposite of how beginners edit. Interview-driven video is carried by its A-roll — the interview audio (Chapter 20). The words are the spine; the pictures are the flesh. So the story lives in the audio, and the audio is a thing you can edit with your ears alone, faster and more honestly than you can edit picture. When you cut the words first, you are forced to make the story work on its actual load-bearing element. When you cut pictures first, you are decorating a spine you have not built yet.
The raw material of a radio edit is not the footage — it is a transcript. You take your synced interview (Chapter 27) and get a text of everything the subject said, timecoded so each line points back to where it lives in the footage. Modern editors will transcribe an interview automatically in a few minutes, and — as Chapter 27's accessibility note promised — that same transcript is the raw material for your captions later, so you are doing double-duty work. Now the interview is a document you can read, and reading is far faster than scrubbing. You can see the whole twelve minutes on a few pages, mark the strong lines with a highlighter, cross out the mush, and rearrange paragraphs — the entire story-finding job, done at reading speed, away from the seductive time-sink of the clips.
Here is the method as a picture, because seeing it is the fastest way to understand it. This is the layout you will build — on paper, in a document, or on index cards — every time you edit an interview for the rest of your life.
FIGURE 30.1 — The radio edit: from raw transcript to a soundbite spine (read, mark, reorder)
STEP 1: THE MARKED TRANSCRIPT STEP 2: SOUNDBITE CARDS STEP 3: THE SPINE (ordered)
(read it; highlight keepers, cut mush) (one keeper per card) (cards in STORY order, not
--------------------------------------- ------------------------ chronological order)
00:14 "Um, so, I guess I started..." ✗ ┌──────────────────────┐ ┌──────────────────────┐
00:41 "My father made chairs, and ✓ │ CARD A @ 09:12 │ ───► │ 1 (HOOK) CARD C │
his father made chairs..." │ "I still love it." │ │ "I still love it" │
01:50 "It's...complicated." (vague) ✗ │ (6s) — the HOOK │ └──────────────────────┘
02:30 "Every joint is cut by hand, ✓ └──────────────────────┘ ┌──────────────────────┐
no screws, no glue-gun..." ┌──────────────────────┐ ───► │ 2 CARD A (origin) │
04:05 [rambles about tool brands] ✗ │ CARD B @ 02:30 │ │ "my father made..." │
06:20 "Nobody makes them like this ✓ │ "every joint by hand"│ └──────────────────────┘
anymore. It's a dying craft." │ (5s) — the CRAFT │ ┌──────────────────────┐
09:12 "I still love it. Fifty years ✓ └──────────────────────┘ ───► │ 3 CARD B (the how) │
in, I still love it." ┌──────────────────────┐ │ "every joint..." │
... │ CARD D @ 06:20 │ └──────────────────────┘
│ "a dying craft" │ ───► │ 4 CARD D (the why) │
The transcript is the QUARRY. You │ (7s) — the STAKES │ │ 5 (RESOLVE) ... │
read it, not scrub it — reading is └──────────────────────┘ └──────────────────────┘
10x faster than watching. Each card = one self- The story the SUBJECT
contained sentence. never spoke in this order.
Read the figure left to right, because it is the whole chapter in miniature. On the left, the raw transcript — most of it crossed out, because most of any interview is throat-clearing, tangents, and non-answers (Chapter 19). The few lines that survive are the keepers: self-contained sentences that say something. In the middle, each keeper becomes a card — a soundbite you can pick up and move. On the right, you deal the cards into story order: the hook first (Chapter 17), then origin, then the craft, then the stakes, then the resolution — an order the subject never spoke, because no one tells their own story in perfect dramatic sequence. That right-hand column is your spine, and you built it without watching a single frame. The pictures come later. The story came first.
🚪 Threshold Concept: cut the words first, and the film builds itself. The instinct of every beginner is to edit an interview by watching it and trimming. The professional does the opposite: they read it and arrange it, as text, until the story works on the page — and only then open the timeline to execute what the page already decided. Once you internalize that the story of an interview-driven piece lives in the words, and that words are faster and more honest to edit than pictures, your whole workflow inverts. You stop scrubbing hopefully through footage looking for a film, and start constructing a film from a document you can hold in your hand. The radio edit is where the hours you would have lost nudging clips get compressed into twenty minutes with a highlighter. Story before picture — always, and especially here.
There is a subtle, hard skill hiding in the marking pass: telling a keeper from a non-answer on the page. A keeper is a sentence that stands on its own — it states its own subject, it does not depend on the question you asked (Chapter 19's full-sentence rule), and it says one clear thing. A non-answer is a fragment ("twenty years"), an echo of your question, a hedge ("it's complicated"), or a lovely-sounding line that, read cold, means nothing. The transcript is merciless about this in a way the footage is not: on screen, a warm delivery can disguise an empty sentence, but on the page, stripped of the charming voice, an empty sentence is obviously empty. That is another reason to work from text — it strips away the performance so you can judge the substance.
⚠️ Common Mistake: falling in love with a shot before the story works. The classic interview-editing failure is to open the footage, find a gorgeous moment, build a beautiful opening around it, and then spend two days trying to make a story reach that shot. You have let one pretty picture take the wheel, and the film will forever feel like it is straining toward a moment instead of telling a story. The fix is the radio edit: decide the story in words first, with the pictures hidden, so that no single shot — however lovely — can hijack the structure. When the spine is built and the story holds as audio, then you go find the pictures, and the gorgeous shot either serves the story you built (great — use it) or it doesn't (save it for the reel, Chapter 39). The story decides; the shot auditions.
🔗 Connection. The radio edit is the paper edit (Chapter 29, §29.2) pointed specifically at dialogue, and it rests on the story tools of Chapter 17 — the arc, the hook, and the story beat. When you deal your soundbite cards into order, you are building the three-act shape of Chapter 17 out of real, unscripted sentences. The transcript itself is the payoff of Chapter 27's sync-and-transcribe step (§27.4), and it will pay off again as your captions at delivery (Chapter 36). One transcript, three jobs: find the story, cut the words, caption the film.
🔄 Check Your Eye. 1. What is a radio edit, and why do you cut the words before the picture? 2. Why work from a transcript instead of scrubbing the footage? 3. On the page, how do you tell a keeper soundbite from a non-answer?
Check yourself
- A radio edit is a paper edit built from the interview audio — you find and order the story using only the spoken words, eyes closed to the picture, so the piece works as radio before any image is added. You cut words first because the story of an interview-driven piece lives in the words (the A-roll is the spine); if it works for the ear, the pictures can only improve it, and no B-roll can rescue a spine that doesn't.
- Reading is roughly ten times faster than scrubbing, and text strips away the charming delivery so you judge substance, not performance — you can see the whole interview on a few pages, mark keepers, and reorder at reading speed.
- A keeper is a self-contained sentence that states its own subject, doesn't depend on your question, and says one clear thing; a non-answer is a fragment, an echo of the question, a hedge, or a nice-sounding line that means nothing when read cold.
30.2 Pulling selects and the string-out
The radio edit told you what the story is and what order it goes in. Now you cross from the page back to the timeline and gather the actual footage of those keeper lines. This is the culling job you previewed in Chapter 27, and now you do it in full and for real, with the interview specifically in mind.
A select is a clip — or, far more often in an interview, the good part of a clip — that you have marked as a keeper and set aside to use. For an interview, a select is almost always a single soundbite: one self-contained sentence or thought, carved out of a long, rambling take with an in-point right before the first word and an out-point right after the last, cleanly, on the breath. You are not keeping "answer to question four" (two minutes of talking); you are keeping the seven seconds inside it where the person said the true, cuttable thing. Everything else — the wind-up, the tangent, the second attempt, the "does that make sense?" at the end — stays in the footage but gets left behind, out of your way. Your radio edit already told you which soundbites to pull; the select pass is just going and getting them.
You gather those selects into the Selects bin you built in Chapter 27, and — this is the interview-specific refinement — you label each one with what it says, not what number take it was. A clip called INT_A_Take04 tells you nothing; a select called SB_love-it-50-years tells you exactly which soundbite it is, so that when you are ordering the spine you can find "the fifty-years line" in a second. You are turning your footage into the deck of cards from FIGURE 30.1, made real on the timeline.
FIGURE 30.2 — The soundbite log (your selects) → the radio string-out on the timeline
THE SOUNDBITE LOG (selects, labeled by CONTENT — this is your card deck, made real)
------------------------------------------------------------------------------------------
LABEL | SOURCE TC | LEN | STRONG? | JOB IN THE STORY
-------------------------+-------------+-----+---------+------------------------------------
SB_still-love-it | 09:12–09:18 | 6s | ★★★ | THE HOOK — opens the film
SB_father-made-chairs | 00:41–00:52 | 11s | ★★ | origin / who they are
SB_every-joint-by-hand | 02:30–02:38 | 8s | ★★★ | the craft (the "how")
SB_dying-craft | 06:20–06:27 | 7s | ★★★ | the stakes (the "why")
SB_keep-going | 10:40–10:46 | 6s | ★★ | the resolution / button
------------------------------------------------------------------------------------------
THE RADIO STRING-OUT (the selects laid end to end, IN SPINE ORDER, audio only for now)
------------------------------------------------------------------------------------------
V1 [ head:love-it ][ head:father ][ head:by-hand ][ head:dying ][ head:keep-going ]
A1 [ "still love it" ][ "father made..." ][ "every joint" ][ "dying craft" ][ "keep going" ]
└── each is ONE self-contained soundbite, trimmed on the breath, in story order ──┘
Watch it with your EYES CLOSED. Does the story hold as pure audio? If yes → you have a spine.
Look at what the string-out is. A string-out is your selects laid end to end on the timeline in a rough, first-pass order — not trimmed to the frame, not covered with B-roll, not scored, just the keepers in a line. In Chapter 27 you built a string-out to prep any edit; here it is the specific, powerful thing the radio edit was aiming at all along — a radio string-out, the soundbite spine laid out as actual audio you can play. It is the moment your paper plan becomes a thing you can listen to. And you test it exactly the way its name demands: you press play, you close your eyes, and you listen to the whole thing through as if it were radio. Does the story hold? Does it hook you, take you somewhere, and land? If yes, you have found your film, and everything from here is making it look as good as it already sounds. If no — if it drags, or the order is wrong, or a beat is missing — you fix it now, in the audio, where fixing is cheap, before you spend a single hour on picture.
🎬 On Set (at the desk): pull selects and build a radio string-out. Take your Project 2 interview (or any interview you have) and its transcript. From the transcript, mark your keeper soundbites, then go into the timeline and pull each one as a select — in and out on the breath, labeled by content. Lay them end to end in the story order your radio edit chose. Constraint: no B-roll, no music, no frame-trimming — just the soundbites, in order, as raw audio. Self-review: play it with your eyes closed, start to finish, and answer one question — does it tell a story a stranger would follow? If it drags or confuses, reorder the selects (not the words within them) until the audio holds. You are not allowed to add a single picture until it does.
⚠️ Common Mistake: keeping too much. The beginner's selects bin is enormous, because everything the subject said feels precious and cutting hurts. So they keep forty soundbites for a three-minute film, string them all out, and end up with a twelve-minute lecture that says the same thing six times. The fix is ruthlessness at the select stage: for every beat of your story, keep the one best soundbite that carries it, and leave the rest — not deleted, just not selected. If two soundbites say the same thing, the weaker one is not a "backup," it is a delay. A tight film is not a long film trimmed; it is a short film built from only the strongest pieces. When in doubt, cut the bite — the story is almost always clearer with fewer, better lines. (This is Chapter 29's "kill your darlings," moved to the front of the process, where it is cheapest.)
✂️ In the Edit. The quality of your selects is set entirely by the quality of your interview, which was set entirely in Chapter 19. Every full-sentence answer you fought for — briefing the subject to answer in complete sentences, re-asking for the clean version, holding the silence for the truer line — becomes a pullable select now. A subject who answered in fragments gives you a bin full of clips you cannot use whole; a subject you coached into self-contained sentences gives you a deck of cards ready to deal. This is the exact moment the interviewing discipline of Chapter 19 either pays out or comes due. Note it honestly for your next shoot: the selects bin is the most brutally accurate report card an interview ever gets.
🔄 Check Your Eye. 1. In an interview, what exactly is a select — the whole answer, or something smaller? 2. What is a radio string-out, and how do you test it? 3. Why label selects by their content rather than by take number?
Check yourself
- Something smaller — a single self-contained soundbite, one sentence or thought, carved out of a long take with clean in/out points on the breath. You keep the seven good seconds, not the two-minute answer they lived inside.
- A radio string-out is your soundbite selects laid end to end on the timeline in story order, as audio you can play. You test it by closing your eyes and listening to the whole thing through — if the story holds as pure radio, you have a spine.
- So you can find any bite instantly by what it says when you're ordering the spine — "the fifty-years line" is findable; "Take 04" is not. Content labels turn your footage into a deck of cards you can deal.
30.3 Building the spine from soundbites
You have a radio string-out that holds up with your eyes closed. Now you turn that loose line of selects into a real, tight spine — the ordered, trimmed sequence of interview soundbites that carries the entire story on the audio track, from hook to resolution. The string-out was the rough draft of the spine; building the spine is where you commit to the order, tighten every join, and make the words themselves as lean and strong as they can be. This is still audio-first work. You are not adding pictures. You are perfecting the thing the pictures will hang on.
Three jobs build the spine. First, order for the arc, not for chronology. Your subject did not answer in story order, and you must not cut in question order. Deal the soundbites into the three-act shape from Chapter 17: a hook that earns the first ten seconds, a middle that develops, and a resolution that lands. Often the best hook is something the subject said late, once they had relaxed and forgotten the camera — the throwaway line that is secretly the thesis. Put it first. The chronology of the interview is irrelevant; the drama of the story is everything.
Second, cut the questions out — the subject stands alone. As Chapter 19 drilled, in the finished piece the viewer never hears you. So your questions come out entirely, and each answer must stand on its own (which is exactly why you engineered full-sentence answers on the shoot). Where a great thought lives inside a tangle, you carve the clean sentence out of it. Where the subject started an answer twice, you take the better start. Where an "um" or a false start sits in the middle of a keeper, you snip it — and you have just created a tiny jump cut on the picture, which is fine, because the whole next section is about hiding it.
Third, tighten the words themselves — but honestly. You may trim within a soundbite, join the front of one sentence to the end of another, and remove an "um" or a redundant clause. What you may never do is build a frankenbite that makes the subject appear to say something they did not mean. This is the ethical line of interview editing, and it is bright: you are allowed to make a person more concise than they were; you are never allowed to make them say something false. Joining "I was angry" from one answer to "...at myself" from another, when they actually said they were angry at their boss, is not editing — it is fabrication, and it is the kind of dishonest cut documentary ethics (Chapter 21) exists to forbid. The test is simple: would the subject, watching the cut, agree that is what they meant? If yes, tighten freely. If no, you have crossed the line.
FIGURE 30.3 — The spine on the timeline: soundbites joined, and the jump-cut problem it creates
THE SPINE — five soundbites, trimmed and joined in STORY order (audio carries the whole story)
V1 [ head: love-it ]|[ head: father ]|[ head: by-hand ]|[ head: dying ]|[ head: keep-going ]
A1 [ "still love it" ]|[ "my father..." ]|[ "every joint" ]|[ "dying craft" ]|[ "and I keep going" ]
^ ^ ^ ^
every | is a splice between two different moments of the interview —
the AUDIO flows as a story, but the PICTURE JUMPS at every one of them.
WITHIN a single soundbite, too:
A1 [ "every joint is cut by hand // [um] // no screws, no glue" ]
^ snip the "um" → the head JUMPS on that frame as well
The spine is DONE as audio and BROKEN as picture. Every splice — between bites AND inside them —
is a visible jump cut on the talking head. That is not a failure. That is the SETUP for §30.4:
the spine is deliberately full of seams, because B-roll is about to cover every one of them.
Study the figure, because it names the exact problem that makes interview editing its own craft. The audio is a finished story — five soundbites flowing hook-to-resolution, plus internal tightening. But every one of those splices, between bites and inside them, makes the talking head jump: the subject's head snaps, their hand teleports, because you have joined moments filmed seconds or minutes apart (Chapter 9's jump cut, Chapter 19's warning about cutting within one continuous shot). A spine, by design, is riddled with these seams. If you had shot the interview with two cameras (Chapter 19, §19.6), you could hide many of them by cutting between angles A and B. But most of the time — and always for the phone-first, single-camera shooter — you hide them with the thing you shot Chapter 20 for: B-roll. The broken picture is not a mistake. It is the reason the next section exists.
💡 Why It Works: the ear tolerates what the eye will not. You can splice audio far more aggressively than picture and get away with it, because the ear forgives a join that the eye catches instantly. A tiny gap or overlap between two soundbites, a breath trimmed, a room-tone patch under a splice — the ear glides over all of it if the sense flows. But the eye is unforgiving: the moment a face jumps, the viewer sees the edit and the spell breaks. This asymmetry is the whole physics of interview editing. It is why you build the story on the audio (which bends) and repair the picture with B-roll (which broke). Work with the grain of human perception: cut the flexible thing freely, and cover the brittle thing completely.
⚠️ Common Mistake: the frankenbite that lies. The most serious error in this entire chapter is not technical — it is ethical. Under deadline pressure, or to make a subject sound better than they were, an editor stitches words the subject never joined into a sentence they never meant, or drops the "not" that reverses their meaning, or cuts a "yes" onto a question they were actually answering "no." The audience cannot tell; the subject can, and so can your conscience, and so, eventually, can everyone when it comes out. The rule is absolute: tighten what a person said, reorder their thoughts, cut their fillers — but never manufacture a meaning they did not intend. If you would be ashamed to show the subject the cut, do not make it. (The releases and consent you handle in Chapter 38 assume you edited them honestly; a frankenbite betrays that assumption.)
🔄 Check Your Eye. 1. When you build the spine, do you order the soundbites by chronology or by something else? What? 2. Why is a finished spine "done as audio and broken as picture"? 3. Where is the ethical line between tightening a soundbite and building a frankenbite?
Check yourself
- By the arc (Chapter 17) — hook, development, resolution — not by the order the subject spoke or the order you asked. The best hook is often a line said late, once they relaxed; chronology is irrelevant, drama is everything.
- Because joining soundbites from different moments (and snipping fillers inside them) makes the audio flow as a story while the talking-head picture jumps at every splice — the ear accepts the joins, the eye catches them. The seams are deliberate; B-roll will cover them.
- You may tighten, reorder, and cut fillers to make a person more concise than they were; you may never stitch words to make them say something false. The test: would the subject, watching, agree that is what they meant?
30.4 Layering B-roll to hide cuts and show
Here is the payoff of the entire book's central promise — you shoot for the edit — arriving exactly on schedule. You have a spine that sounds like a finished story and looks like a broken one, jumping at every seam. Now you lay the B-roll you gathered in Chapter 20 over those seams, and two things happen at once: every jump vanishes, and every claim gets proof. This is the move that turns a talking head into a film. When you do it well and completely, you have made a B-roll-driven narrative: a piece whose story is carried by the interview audio and told by the B-roll laid over it — the voice leads, unbroken, while the picture roams over the work, the place, and the person, hiding every cut and showing what the words describe.
Recall from Chapter 20 that B-roll does two jobs at once, and now you cash in both. It covers the cut: laid over a splice in the spine, the viewer is looking at the subject's hands, not their jumping head, at the exact frame the audio joins — the seam disappears. And it shows what is described: over the line "every joint is cut by hand," you lay the shot of the hands cutting the joint, and the picture proves what the voice claims (show-don't-tell, Chapter 17, made literal). One shot, both jobs — which is why the matched B-roll you planned against your beats (Chapter 20's coverage plan, FIGURE 20.7) is worth so much more than pretty footage that covers a cut while saying nothing.
The mechanism is the L-cut you learned in Chapter 28, now used not once but as the engine of the whole piece: the interview audio runs continuously on its track while the picture cuts to B-roll above it. In your editor, the interview lives on the base video track with its audio; the B-roll drops onto the track above it, covering the picture while the interview sound plays on underneath. (In DaVinci Resolve — the free suite this book anchors to — you place the B-roll clip on the video track above the interview and trim its handles to time the cover; Appendix E maps the same move in other software. The concept is identical everywhere: picture on top, voice underneath, sound never stops.)
FIGURE 30.4 — B-roll-driven narrative: the same spine, now covered (the voice leads, the picture proves)
V2 ........[ B-roll: hands cutting the joint ][ B-roll: the shelf ][ ....][ B-roll: the sign ]
V1 [ FACE ]|________(picture cut away)________|[FACE]|____________|[FACE]|__(covered)__|[ FACE ]
A1 [ interview spine (UNBROKEN): "...still love it // my father // every joint by hand // dying craft // keep going" ]
A2 [ wild sound: the plane, the mallet, the shop hum, birds outside ............................... ]
^open on ^the splices from FIG 30.3 now live UNDER B-roll — every jump is hidden
the face | | |
(hook) proves "by hand" proves "dying craft" returns to the face for the button
THE RULE FOR THE LOCK: every seam in the spine (§30.3) must have B-roll over it OR a two-camera cut.
Open on the face (meet the person), roam over the B-roll (hide cuts + prove claims), and RETURN to
the face for the lines that must be SEEN said. The audio never stops. The story is carried by the
voice; the film is TOLD by the pictures laid over it. That is a B-roll-driven narrative.
This figure is the destination of the whole chapter, so read it slowly. On A1, the interview spine runs unbroken — the story, in the subject's voice, exactly as your radio edit built it. On V1, the face appears at the start (we meet the person) and returns at key moments. On V2, B-roll covers the picture over every splice from FIGURE 30.3 — and each piece of B-roll is matched: the hands over "by hand," the empty shelf over "dying craft." On A2, wild sound (Chapter 20) gives the B-roll its own texture so it feels alive, not like a silent slideshow. Nothing about the audio story changed from the broken spine of §30.3 — the exact same soundbites, in the exact same order. The B-roll simply gave the picture somewhere to be at every seam. The interview, un-cuttable a section ago, is now free to be as tight as your radio edit made it, because the jumps are all hidden. That freedom is what B-roll buys, and it is why an interview with no B-roll is a hostage video and an interview with rich, matched B-roll can be superb.
Now let us render the single most important cut in the whole method as a Described Shot — the moment a soundbite hands off to its B-roll — because the timing of that handoff is a craft in itself, and getting it right is what separates an edit that breathes from one that feels like a machine gun of cutaways.
🎞️ Read This Sequence. This is the signature cut of interview editing: the held face, the pause, and the reach for B-roll on exactly the right frame. Read it as a decision about when, not just what.
FIGURE 30.5 — "The pause before the answer" [constructed teaching example]
THE FRAME Medium close-up. The subject sits right-of-center on the thirds, looking camera-left to an
off-screen interviewer; the background falls soft at f/2.8, a window bokeh'd behind them.
THE MOVE Locked off on a tripod. Stillness is the point — the subject is about to say something hard,
and a moving camera would compete with the moment.
THE LIGHT Soft key from the window camera-left, just forward of the subject; the far cheek a stop
darker; a gentle rim from a practical lamp lifts their shoulder off the background.
THE SOUND We stay on their sync audio, but the interviewer's last question has already ended (an
L-cut): room tone, a breath, then the answer begins two full seconds in.
THE CUT We cut IN from a wider two-shot on the interviewer's final word; we HOLD through the silence
and on the face as the first strong sentence lands — then, as the sentence completes, we cut
AWAY to the B-roll that proves it, the interview audio continuing beneath (the L-cut).
THE EFFECT The held silence makes the viewer lean in. Because the picture is still and the sound is
bare, the pause becomes the loudest thing in the piece — and because we held the FACE for the
hard sentence and only left for the B-roll once it was said, the emotion registers before the
picture wanders.
THE LESSON Silence is a shot, and timing is everything: hold the face through the pause and the key
sentence so the viewer SEES it said, then hand off to B-roll on the tail. Cut away too early
and you rob the line of its face; cut away too late and the talking head goes stale.
The lesson in that box is the whole rhythm of §30.5, previewed here in one cut: you do not simply bury the interview under wall-to-wall pictures. You choose which lines the viewer must see said on the face — the emotional ones, the vulnerable pause, the line where the person's expression is the point — and you hold the face for those, handing off to B-roll only on the connective tissue between them. B-roll hides the cuts you want hidden; it must never hide the face at the moment the face is the story.
There is one more piece of craft to name, and it connects straight back to the scene you have been building all book. The Café Scene is not an interview — it is a scripted narrative scene with planned coverage — but the B-roll-over-voice move you are using here is exactly the one you first performed on it. In Chapter 20 you shot café cutaways (the steam rim-lit by the window, the barista's hands, the sign on the door, FIGURE 20.6), and in Chapter 28 you laid them over the order exchange as L-cuts to hide the joins between takes. That was this same technique in miniature: sound continues, picture roams, the cut vanishes. What is new here is scale and purpose — you are covering not one dialogue cut but an entire story's worth of them, and your B-roll is not just texture but proof. The scene never changed; the move you learned on it is now the engine of a whole film.
✂️ In the Edit: every seam needs a cover, and you decided that on the shoot. This is the callout the whole book has been building toward, so let it land. In the edit, you cover every splice in the spine with B-roll — and if you did not shoot that B-roll, the seam stays naked and the film stays broken. There is no button that generates the hands, the shop, the sign. The overlap between "seams I need to cover" and "cover I actually have" was decided entirely on your Chapter 20 B-roll day. The editor who shot a rich, matched coverage plan cuts freely and hides every jump; the editor who shot thin is trapped, forced to leave the talking head naked or hold shots too long. Right now, editing your own footage, you are meeting the version of you who held the camera — and either thanking or cursing them. Whichever it is, remember it on your next shoot, and over-shoot the B-roll.
⚙️ Settings Box: laying B-roll over an interview (a starting point, not a recipe).
Decision Starting point Why Where the B-roll goes On the video track above the interview (V2 over V1) Covers the picture while the interview audio plays on underneath (the L-cut) When to cut away On the tail of a key sentence, after the face has said it Hold the face for the line; leave for the B-roll once the point is made When to return to the face For emotional lines, the pause, the button The viewer needs to see the person for the moments that matter most How long each B-roll shot holds Long enough to read, short enough to stay alive (often 2–4 s) A cutaway that overstays becomes its own dead spot Wild sound under B-roll Low, on its own track (A2), beneath the voice Makes B-roll feel live, not like a silent slideshow (Chapter 20) Interview audio level Always clear and on top; dip music/wild sound under it The voice is the spine — nothing may fight it (sound is half the picture) Exact loudness targets in LUFS, music choice, and the real mix are audio post (Chapter 33); here you are timing the cover to serve the story, not finishing the sound.
♿ Accessibility & Inclusion. A B-roll-driven narrative has a built-in accessibility trap and a built-in accessibility win. The trap: because the story is on the audio and the pictures are illustrative, a blind or low-vision viewer who cannot hear will be fine — but a deaf or hard-of-hearing viewer watching sound-off gets your beautiful B-roll and no story at all, because the story lives in the voice they cannot hear. The fix is captions (Chapter 36), and the good news is you already have them: your radio-edit transcript, trimmed to match the spine, is your caption file. The win: because you shot and laid in matched B-roll, the muted picture track still shows the craft and the beats — so a sound-off viewer following captions sees a real film, not a static head. Build the piece so both halves of your audience receive the story: the voice for those who hear it, the captions and the matched pictures for those who don't.
🔄 Check Your Eye. 1. What is a B-roll-driven narrative — what carries the story, and what tells it? 2. What two jobs does a single matched B-roll shot do when you lay it over an interview seam? 3. When should you not cut away to B-roll — when must the viewer stay on the face?
Check yourself
- A piece whose story is carried by the unbroken interview audio (the spine) and told by the B-roll laid over it — the voice leads and never stops; the picture roams over the work, place, and person, hiding cuts and proving claims.
- It covers the cut (the viewer sees B-roll, not the jumping head, at the splice — the seam vanishes) and it shows what the voice describes (proving the claim — show-don't-tell). One shot, both jobs.
- On the emotional lines, the vulnerable pause, and the moments where the person's expression is the point — hold the face so the viewer sees it said. B-roll hides the cuts you want hidden; it must never hide the face when the face is the story.
30.5 Pacing an interview piece
An interview-driven piece can fail in two opposite ways, and both are about pacing. The first is the hostage video: too much face, too little B-roll, a head talking for ninety unbroken seconds until the viewer's attention drains no matter how good the words are. The second is its mirror image, and it is the one that catches people after they learn to love B-roll: wall-to-wall pictures, the face buried under a relentless stream of cutaways so that we never actually see the person, never connect with them, and the piece feels like a corporate sizzle reel with a voiceover — slick, busy, and strangely cold. The craft of pacing an interview is finding the living rhythm between those two failures: showing the person enough to bond with them, and the B-roll enough to prove and breathe.
The governing principle is the reveal-and-return rhythm. Reveal the person at the top — open on the face (as FIGURE 30.4 does), so the viewer meets a human being before the pictures start roaming. Then roam over the B-roll to prove claims, hide seams, and give the eye variety. Then return to the face for the lines that must be seen said — the emotional beat, the pause, the confession, the button at the end. A well-paced interview piece breathes in and out of the face this way: we are with the person, then out in their world, then back with them for the moment that matters. The return to the face is not a default you fall back to when you run out of B-roll; it is a choice you make for the specific lines whose power is in the expression.
Everything you learned about rhythm in Chapter 29 applies directly here, with an interview-specific accent. Vary your shot lengths so the piece has a pulse, not a metronome. Let it accelerate through a build and then exhale on a long hold — and in an interview, the most powerful exhale is often a held shot of the face saying nothing, or a slow, quiet piece of B-roll under a pause. Cut on the thought, not on a fixed count (Chapter 29's blink): let a soundbite finish its idea before you move, and let the B-roll change when the point changes. And protect the living silences — the two-second pause before the hard answer (FIGURE 30.5) is not dead air to be trimmed for pace; it is the loudest moment in the piece, and pace is what earns it.
Here is a thirty-second stretch of a real interview piece, rendered as a Described Sequence, so you can see the reveal-and-return rhythm as an actual cut — the alternation of face and B-roll, and the deliberate choice of which line stays on the face.
FIGURE 30.6 — Thirty seconds of a paced interview piece: face → B-roll → face [constructed teaching example]
# | Picture (V1 face / V2 B-roll) | Dur | Audio (A1 interview spine + A2 wild) | Why here
---+--------------------------------------+------+-------------------------------------+------------------------
1 | FACE, medium CU (the hook) | 5 s | "Fifty years in, I still love it." | meet the person first
2 | B-ROLL seq: hands cutting the joint | 4 s | (spine continues) "...every joint | prove "by hand"; hide the
| (wide→detail) | | is cut by hand..." + plane hiss | splice under it
3 | B-ROLL: ECU the chisel in the wood | 3 s | "...no screws, no glue." + mallet | the detail that lands
4 | B-ROLL: slow pan, shelf of finished | 4 s | "Nobody makes them like this | prove "dying craft";
| work; an empty apprentice's chair | | anymore." + shop hum | the B-story (the "why")
5 | FACE, medium CU (HOLD) | 6 s | "It's a dying craft." [2s pause] | RETURN to the face — this
| | | "...and that scares me." | line must be SEEN said
6 | B-ROLL: dust in the window light | 4 s | (room tone; music bed swells low) | the exhale — let it land
7 | FACE (the button) | 4 s | "But I keep going." | end on the person
---+--------------------------------------+------+-------------------------------------+------------------------
Face → world → face. We bond (1), we see the proof (2–4), we return for the hard line (5), we
breathe (6), we close on the human (7). B-roll hides every seam; the FACE owns the emotional beats.
Read the right-hand column, because it is the pacing logic laid bare. We open on the face to bond (shot 1). We roam the B-roll to prove the craft and the stakes, hiding the spine's splices under every cutaway (shots 2–4). Then — and this is the decision that makes it a film instead of a reel — we return to the face for the vulnerable line, "and that scares me," holding through the two-second pause so the viewer sees it said (shot 5). We exhale on a quiet piece of atmosphere (shot 6), and we close on the person (shot 7). Notice that the two most important moments — the hard confession and the final button — are both on the face, not on B-roll. That is not an accident of coverage; it is the pacing choice at the center of the craft. The B-roll serves the story; the face is the story at the moments that matter most.
⚠️ Common Mistake: wall-to-wall B-roll (burying the person). Once you learn that B-roll hides cuts, the temptation is to cover everything — to never let the face breathe, because every cutaway hides another seam and looks "cinematic." The result is a piece where we never really meet the person, and it feels slick and hollow: a voice narrating a montage instead of a human being telling a story. The fix is the reveal-and-return rhythm: open on the face, return to it for the emotional lines and the button, and reserve wall-to-wall B-roll for the connective, informational stretches. If you cannot remember what your subject looked like after watching your own cut, you buried them. A documentary is about a person; let us see them.
🎬 On Set (at the desk): the reveal-and-return pass. Take your covered spine and do one dedicated pass on pacing alone. Mark the two or three lines whose power is in the face — the emotional beat, the pause, the closing button — and make sure each of those plays on the face, held, not under B-roll. Then make sure you open on the face and return to it at least twice. Constraint: somewhere in the piece, hold one silent or near-silent beat (a pause, a quiet piece of B-roll) for at least two full seconds — the exhale. Self-review: watch it and ask two questions — did I meet the person, and did I stay with them for the moment that mattered? If the face never gets a held emotional beat, you are pacing a sizzle reel, not cutting a story.
🔗 Connection. Pacing here is Chapter 29's rhythm (§29.3) — accelerate, hold, exhale; cut on the thought; protect the living silence — applied to the specific in-and-out of face and B-roll. And when you lay a music bed under the piece to feel its emotional shape, run Chapter 29's mute test (§29.5): mute the music and the story must still hold on the spine and the pictures. Music that carries a piece the spine cannot is a crutch, not a bed. (Choosing and licensing that music is Chapter 38's territory — 🔗 see Chapter 38 before you publish with any track; the technical mix is Chapter 33.)
🔄 Check Your Eye. 1. Name the two opposite ways an interview piece fails on pacing. 2. What is the "reveal-and-return" rhythm, and which moments must play on the face? 3. In an interview piece, what makes the most powerful "exhale"?
Check yourself
- The hostage video (too much face, no B-roll, the head talks unbroken until attention drains) and wall-to-wall B-roll (the face buried under relentless cutaways so we never meet the person and it feels slick and cold).
- Reveal the person at the top (open on the face), roam over B-roll to prove and breathe, and return to the face for the lines whose power is in the expression — the emotional beat, the pause, the closing button. Those must-be-seen lines play on the face, not under B-roll.
- A held shot of the face saying nothing, or a slow quiet piece of B-roll under a pause — the two-second silence before or after a hard line, which pace has earned. Sustained speed exhausts; the breath is what makes the peak land.
30.6 The fine cut and lock
You have a paced, covered, honest cut. The last stage is to make it finished and then to stop — two disciplines that sound easy and are not. The fine cut (Chapter 29) is where you tighten every frame; the lock is where you declare the picture done and hand it off. Both have interview-specific forms, and both are where good editors separate from people who fiddle forever.
Take the fine cut first. Everything from Chapter 29 applies — enter late, leave early; trim the fat; protect the living silences — with one addition specific to this kind of edit: you now tighten in two directions at once, the audio spine and the picture cover, and they must stay married. When you trim a soundbite tighter, the B-roll over it may now be too long and need pulling in; when you shorten a B-roll shot, you may expose a seam that needs new cover. The fine cut of an interview piece is a pass down the whole timeline asking, of every second: is this word earning its place, and is this picture earning its place? Almost every interview piece is fifteen or twenty percent too long in its first covered version — a soundbite that says the same thing as the one before, a B-roll shot held two beats past its point, a pause that is dead rather than living. Find them and cut them, and protect the handful of moments that must be long.
Then you run the two quality passes that catch what a tired editor's eyes miss — and they are the two "close a channel" tests this chapter has been building toward:
- The eyes-closed pass (the radio test, again). Play the whole finished piece and listen only. Does the story still hold as pure audio, hook to resolution? If a stretch drags or confuses with your eyes shut, no picture will save it — the spine has a hole, and you fix the words. This is the radio edit returning as final quality control.
- The sound-off pass (the mute test, for picture). Play it again and watch only, muted. Does the picture track still show a film — the person, the craft, the place, the beats — or does it collapse into a static head and a few random cutaways? If it collapses, your B-roll is not matched to the story, and a sound-off viewer (and a caption reader) gets nothing.
A piece that passes both — a story that holds for the ear alone and a film that shows for the eye alone — is genuinely finished, because it serves both halves of your audience and both halves of the craft. That double test is the highest standard an interview-driven piece can meet, and it is the target of your fine cut.
Now the hard part: stopping. Picture lock is the moment you declare that no more picture changes will be made — the cut is frozen — and hand the film off to the finishing stages: color (Chapters 31–32), audio post (Chapter 33), titles and lower thirds (Chapter 34), and captions and export (Chapter 36). Lock exists because those downstream stages build on the exact timing of your cut: a colorist grades specific frames, a sound mixer balances specific moments, a motion-graphics artist times a lower third to a specific soundbite. If you keep nudging the picture after they start, their work breaks. Lock is a promise — this timing is final — that lets the finish begin. On a solo project where you do everything yourself, lock is just as important, because it is the decision that saves you from re-cutting forever; it is the line you draw that says the storytelling is done; now I make it shine.
FIGURE 30.7 — The interview-edit pipeline: from transcript to lock (run it in order)
1 RADIO EDIT transcript ─► mark keepers ─► order the cards ─► the story, on paper (§30.1)
│ (eyes closed to picture — the story must work as WORDS first)
▼
2 SELECTS pull each keeper soundbite (in/out on the breath), label by CONTENT (§30.2)
│
▼
3 RADIO STRING-OUT selects end to end, in story order ─► CLOSE YOUR EYES: does it hold? (§30.2)
│
▼
4 SPINE order for the ARC; cut the questions; tighten honestly (no frankenbite) (§30.3)
│ (done as audio, deliberately broken as picture — seams everywhere)
▼
5 B-ROLL LAYER cover EVERY seam; match B-roll to the beats; voice leads, picture proves (§30.4)
│
▼
6 PACE reveal → roam → return; hold the face for the emotional lines; exhale (§30.5)
│
▼
7 FINE CUT enter late/leave early on words AND pictures; protect living silences (§30.6)
│ ─► run BOTH passes: eyes-closed (radio) + sound-off (mute)
▼
8 LOCK freeze the picture ─► hand off to color / sound / titles / captions (§30.6)
───────────────────────────────────────────────────────────────────────────────────────────────►
Project 2 is now LOCKED. Everything after this makes the finished film shine (Part VII).
That ladder is the whole chapter on one card, and it is the method you will run on every interview-driven piece you ever cut. Note the shape of it: the first four rungs are audio and story, done before you lay a single frame of picture; picture only enters at rung five. That is the radical, counterintuitive core of interview editing, and the ladder exists to keep you honest about it. When a piece is not working and you cannot say why, walk back down the ladder — is the spine broken (rung 4), or just the cover (rung 5)? Nine times out of ten a "boring" interview piece has a fine cover and a broken spine, and the fix is not more B-roll — it is back to the words.
✂️ In the Edit: lock is the payoff of every promise you made the editor. From Chapter 1 you were told you shoot for the edit, and every chapter since has been you making promises to this moment — the full-sentence answers (Chapter 19), the matched B-roll and wild sound (Chapter 20), the clean sync and safe offload (Chapter 27). Lock is where those promises come due and get paid. An editor who was shot for — deep coverage, clean audio, honest interviews — reaches lock with a tight, moving film and time to spare. An editor who was not is still, at lock, papering over holes that a shot never filled. Reaching a clean lock is the quiet proof that the whole chain worked. Savor it: your documentary short is a film.
🎒 Gear Note: none of this needs a fast machine or a paid tool. The entire method in this chapter — radio edit, selects, spine, B-roll layer, pace, lock — runs on a free editor and a modest computer. The transcript is text; the spine is audio; the B-roll layer is dragging a clip onto the track above. If your interview and B-roll stutter, the proxies from Chapter 27 make even a phone-shot 4K documentary cut smoothly on a laptop. The one thing that would actually help — and it is free — is a second pair of eyes: show your locked cut to one person who has not seen the footage, and watch their face for where attention drifts. The most valuable edit tool in this chapter costs nothing and lives in someone else's reactions.
🔄 Check Your Eye. 1. What are the two quality passes you run on an interview fine cut, and what does each one catch? 2. What is picture lock, and why does it matter even on a solo project? 3. A locked interview piece feels boring and you can't say why. Where do you look first — the spine or the B-roll?
Check yourself
- The eyes-closed pass (listen only — does the story hold as pure audio? catches a broken spine) and the sound-off pass (watch muted — does the picture still show a film? catches unmatched B-roll and fails a sound-off/caption viewer). A finished piece passes both.
- Picture lock is declaring no more picture changes and handing off to color, sound, titles, and captions — it matters because those stages build on your exact timing, and on a solo project it is the decision that stops you re-cutting forever and lets you finish.
- The spine (the words), first. A "boring" interview piece usually has fine cover and a broken spine; the fix is back to the radio edit and the order of the soundbites, not more B-roll.
Production Checkpoint
Project 2 — LOCK. Finish your documentary short: build the spine from your interview selects, layer B-roll over every cut, and lock the fine cut. This is the assignment the last fifteen chapters have been building toward. You shot the interview (Chapter 19) and the B-roll (Chapter 20); you ingested, synced, and strung out the media (Chapter 27); you learned the grammar and the feeling of the cut (Chapters 28–29). Now you finish the film.
Run the whole pipeline, in order:
- Radio edit. From your interview transcript, mark the keeper soundbites, cross out the mush, and deal them into story order on paper — hook, middle, resolution (Chapter 17). Decide the story in words before you touch the picture.
- Selects and radio string-out. Pull each keeper as a select (in and out on the breath, labeled by content) and lay them end to end in story order. Close your eyes and listen: does it hold as radio? Reorder until it does.
- Build the spine. Commit the order, cut your questions out, tighten every soundbite honestly (never a frankenbite), and accept that the picture now jumps at every seam.
- Layer the B-roll. Cover every seam with matched B-roll on the track above — the hands over "by hand," the place over the establishing lines — with wild sound underneath. Voice leads; picture proves.
- Pace it. Reveal the person, roam the B-roll, and return to the face for the emotional lines and the button. Give it one held exhale.
- Fine cut and lock. Tighten words and pictures together; run the eyes-closed pass and the sound-off pass; then freeze the picture.
Why this matters: this is the checkpoint where Project 2 stops being footage and becomes a film — a real, finished, three-minute documentary short you can show a stranger. It is also the single most transferable edit in the book: every testimonial, brand film, explainer, and documentary you will ever be paid to cut is this exact method. Do not chase perfect color or a finished mix here — that is Part VII. Chase a story that holds with the sound off and with the eyes closed, then lock it. When you do, you have crossed the finish line the whole book pointed at: you have turned a real person's words into a film.
Summary
Editing the interview-driven piece is a four-part construction — find the story in the words, build the spine, cover it with B-roll, and lock — done in a strict order that puts audio and story before a single frame of picture.
The interview-edit pipeline (run it in order):
| Step | Move | The key idea |
|---|---|---|
| 1. Radio edit | Mark keepers on the transcript; order the cards on paper | Cut the words first, eyes closed to picture — story before picture |
| 2. Selects | Pull each keeper soundbite, in/out on the breath, labeled by content | A select is the good seven seconds, not the whole answer |
| 3. String-out | Selects end to end in story order; listen with eyes closed | If it holds as radio, you have a spine |
| 4. Spine | Order for the arc; cut the questions; tighten honestly | Done as audio, deliberately broken as picture |
| 5. B-roll layer | Cover every seam; match B-roll to the beats | Voice leads, picture proves — a B-roll-driven narrative |
| 6. Pace | Reveal → roam → return; hold the face for emotional lines | Never bury the person; the face owns the key beats |
| 7. Fine cut | Enter late/leave early on words and pictures | Run the eyes-closed pass and the sound-off pass |
| 8. Lock | Freeze the picture; hand off to finishing | The storytelling is done; now it shines (Part VII) |
The four owned terms:
| Term | What it is |
|---|---|
| Radio edit | A paper edit built from interview audio alone — cut and order the words with eyes closed to picture, so the piece works as radio first |
| Select | A clip, or the good part of one, marked as a keeper — in an interview, a single self-contained soundbite |
| String-out | Selects laid end to end in rough story order — the spine, before frame-trimming or B-roll |
| B-roll-driven narrative | A piece whose story is carried by unbroken interview audio and told by B-roll laid over it — the voice leads, the picture proves |
Decision rules to keep:
- Cut the words first. If the story doesn't hold with your eyes closed, no B-roll will save it.
- Order for the arc, never for chronology or question order. The best hook is often said last.
- The subject stands alone: cut the questions out; every soundbite must be a full sentence.
- Tighten honestly. Make a person more concise, never make them say something false — no frankenbites.
- Cover every seam with matched B-roll (or a two-camera cut). Voice on top, unbroken; wild sound beneath.
- Hold the face for the lines that must be seen said; reserve wall-to-wall B-roll for connective stretches.
- A finished piece passes both tests: it holds for the ear alone (eyes closed) and shows for the eye alone (sound off).
- Lock the picture before you finish. The storytelling ends so the shining can begin.
Themes surfaced: story is the boss (the radio edit puts story before every pretty picture); sound is half the picture (the spine is audio; the voice leads everything); you shoot for the edit (every seam you cover was a B-roll shot you did or didn't grab in Chapter 20); motivate every choice (every cutaway hides a cut and proves a beat).
Spaced Review
Reach back to the two chapters this edit stands on — the interview you shot and the grammar you cut with.
- (Chapter 19) Why must an interview answer be a full, self-contained sentence, and how does that rule make the selects and spine of this chapter possible?
- (Chapter 19) You shot a single-camera interview. Name the internal cuts it will create when you build the spine, and the two Chapter 19 ways (besides B-roll) to hide them.
- (Chapter 28) What is an L-cut, and why is it the single most important cut in a B-roll-driven narrative?
- (Chapter 28) How does cutting on action in a B-roll sequence (Chapter 20) help a cutaway feel like part of the film rather than an interruption?
Check yourself
1. Because in the edit the question is cut out — the viewer only hears the subject — so a fragment ("twenty years") answers an unheard question and can't stand alone, while "I've been doing this twenty years" can. Full sentences are what make each soundbite a *pullable select* and a *movable card* in the spine; without them you have clips you can't lift, reorder, or join. 2. Trimming within one continuous take creates *jump cuts* (the head snaps at every splice, between bites and inside them). Besides laying B-roll over the seam, Chapter 19's two ways are: *punch in* (crop a 4K take into a "wide" and a "tight" and cut between the sizes) and *reframe between questions* (change size/angle on the 30° rule so a resumed shot cuts cleanly) — and, on the shoot, a *two-camera* setup lets you hide cuts by switching A↔B. 3. An L-cut is where the audio of the outgoing shot leads into the incoming picture — the sound continues across the picture cut. It is the engine of a B-roll-driven narrative because the interview audio runs *unbroken* while the picture cuts to B-roll, which is exactly what hides the spine's seams and lets the voice carry the story over the roaming pictures. 4. Cutting on the movement (Chapter 9's matching action) carries the motion across the cut so the seam disappears, making a B-roll sequence read as one fluid action; a cutaway timed to a motion or a beat feels *motivated* and woven in, rather than a random pretty shot dropped on top of the voice.What's Next
Project 2 is locked — a real documentary short, built from a real person's words, with B-roll over every seam, holding up with the eyes closed and the sound off. That is a genuine milestone: you have proven you can do the hardest, most common edit in nonfiction video, the one every client and every story will ask of you. Part VI, the edit, is complete.
Now the film gets to shine. Chapter 31 opens Part VII — Finishing, beginning with Color Correction: balancing every shot to a neutral, consistent baseline with the scopes, so the interview you lit in Chapter 19 and the B-roll you grabbed across a changing afternoon all match, shot to shot. Your cut told the story; the finish makes it look and sound like the story deserves. You built the film. Next you make it beautiful.