Case Study 2: From Cards to String-Out — Prepping a Documentary Edit, Step by Step
This is a from-scratch, production-side walkthrough — the opposite of Case Study 1's analysis. We take a real-shaped documentary shoot (a constructed teaching example, no real people named) all the way through the edit-prep pipeline: the media comes off the cards, gets organized, gets synced, gets proxied, gets culled to selects, and ends as a string-out ready to cut. Every phase shows the thinking, a diagram, and what got adjusted when reality intervened. Do this once on your own footage and the whole chapter becomes muscle memory.
The brief and the constraints
The assignment is a Project-2-style piece: a 3-minute documentary short profiling the owner of a small neighborhood bicycle-repair shop — a talking-head interview cut against B-roll of the work, in the "talking head at a table" and "product/detail" setups you've used all book. One-person crew, one shoot day, modest laptop for the edit. The goal of this case study is not the finished film (that's Chapters 28–30); it's getting from "a bag of cards at the end of the day" to "a synced, organized, strung-out project I can actually cut."
Here's what came back from the shoot day — the pile we have to tame:
FIGURE CS2.1 — What came back from the shoot (the raw pile, one bag of media)
CAMERA A (main interview + some B-roll) 4K/25, 10-bit, Long-GOP — one 128 GB card
• the interview: ~40 min across 6 questions, several takes
• hands-on B-roll: wheel truing, tools, the workbench
CAMERA B (second interview angle) 4K/25 — one 64 GB card
• locked-off wider two-shot, rolling the whole interview
AUDIO RECORDER (double-system) 48 kHz, 24-bit — one SD card
• lav on subject + boom overhead, both channels
• 90 sec of ROOM TONE recorded before we wrapped
PHONE (grab shots) 4K/HEVC
• the shop sign, the street, a customer arriving
MUSIC one licensed track (a folder on the desktop)
Total: ~120 GB across four devices, three codecs, two frame-rate-safe sources,
and audio living entirely apart from the picture. Nothing is synced. Nothing is named.
This is a completely normal end-of-day mess. The pipeline turns it into an edit.
Notice the shape of the problem, because it's the shape of almost every shoot: media scattered across several devices, in different codecs, with the good sound living separately from the picture (that's double-system sound from Chapter 15 doing its job), and not one file yet named or synced. This is not a failure. This is exactly what a well-run shoot looks like at wrap. The rest of this case study is the pipeline from FIGURE 27.6, run on this pile.
Scale this up or down and the shape holds. A one-person phone shoot has fewer devices but the same steps (offload, organize, maybe proxy, select, string-out). A three-camera event has more of everything but the same logic. Learn the pipeline on this mid-sized documentary and you can run it on anything — the number of cards changes, the discipline doesn't. So take these six phases not as "how this one shoot went" but as the template you'll adapt for the rest of your working life.
⚙️ Settings Box: the ingest decisions for this shoot. Made once, at the start, and written down so the whole project is consistent.
Decision Choice for this shoot Why Offload method Verified (checksum) copy via the editor's clone tool Catches a bad copy while the cards still exist Copies before wiping Two drives (working SSD + backup HDD) The two-copy floor; full 3-2-1 comes at archive (Ch.37) Project frame rate 25 fps (matches both cameras) Set the timeline to the footage; no conforming headaches Proxies? Yes — 4K Long-GOP on a laptop will stutter Half-res, easy codec; edit light, deliver heavy Sync method Waveform auto-sync, clap as backup We slated each interview take; scratch audio is clean Naming BIKE_prefix + setup + subject/shot + takeConsistent and sortable; full scheme deferred to Ch.37
Phase 1 — Offload: get it off the cards, twice
Before anything creative, the media has to survive. Each card gets a verified copy to two separate drives; no card is touched again until both copies are confirmed.
FIGURE CS2.2 — The offload, four cards to two drives (verified)
CAM-A card ─┐
CAM-B card ─┤ ┌─► SSD (working / edit drive) ── "01_FOOTAGE/..."
REC card ─┼─► verified copy ──►─┤
PHONE ─┘ └─► HDD (backup drive) ── mirror of the same
Confirmed: SSD ✓ HDD ✓ → cards are now safe to reuse (but we don't reuse them today —
no reason to, and a wrapped card in the bag is a free third copy until we need it).
The thinking here is pure paranoia, and it's correct. The one-line offload log gets filled in as we go — card ID, contents, copied-to-SSD ✓, copied-to-HDD ✓, verified ✓ — so "did I copy that card?" is never a question we have to answer from memory. What got adjusted: the verified copy of Camera A's 128 GB card flagged one file that failed verification on the first pass. On a naked drag-and-drop we'd never have known until that clip refused to play mid-edit; instead we re-copied it from the still-intact card and moved on. That single caught error is the entire argument for verification, delivered on day one.
Phase 2 — Organize: folders on disk, bins in the project
Now the pile gets homes. We build the FIGURE 27.2 structure on the SSD and mirror it as bins after importing.
FIGURE CS2.3 — This project's folders (disk) and bins (project), mirrored
ON THE SSD IN THE PROJECT
BIKE-DOC/ 📁 BIKE-DOC
├── 01_FOOTAGE/ ├── 01 FOOTAGE
│ ├── CAM-A/ │ ├── Cam A — interview
│ ├── CAM-B/ │ ├── Cam B — wide angle
│ ├── B-ROLL/ │ ├── B-roll — the work
│ └── PHONE/ │ └── Phone — grabs
├── 02_AUDIO/ ├── 02 AUDIO
│ ├── RECORDER/ │ ├── Recorder — lav+boom
│ └── ROOM-TONE/ │ └── Room tone
├── 03_MUSIC-SFX/ ├── 03 Music & SFX
├── 06_PROXIES/ ├── 04 SELECTS ◄ project-only
├── 07_EXPORTS/ │ ├── Interview selects
└── 08_DOCS/ (brief, release, questions) │ └── B-roll selects
└── 05 TIMELINES ◄ project-only
├── String-out
└── Rough cut
The thinking: every kind of asset now has exactly one obvious home, numbered to sort in workflow order. Room tone gets its own folder and bin, because in three weeks when a cut needs to breathe we'll want it in one reach, not buried among the recorder files. The Selects and Timelines bins have no disk folder — they'll hold pointers, not new media. What got adjusted: the phone's HEVC grabs were 4K but a different frame rate than the 25 fps project; we noted it in the DOCS folder so future-us isn't surprised when they need a slight conform, and filed them anyway. Organization isn't just tidiness — it's where you notice problems early, while they're cheap.
✂️ In the Edit. This is the payoff of every "boring" thing the shoot did right. The slate on each interview take (Chapter 18) means the takes are identifiable. The 90 seconds of room tone (Chapter 15) means the edit can breathe and hide cuts. The extra B-roll sequence of the wheel-truing (Chapter 20) means we'll have somewhere to cut when we trim the interview. None of it helped until this moment — and now, filed and labeled, all of it is at our fingertips. A shoot with this coverage and no organization would be a warehouse with no shelves; the coverage would exist but we couldn't use it.
Phase 3 — Sync: marry the double-system audio to the picture
The good sound lives on the recorder; the picture lives on the cameras. Time to make them one. Because we slated every interview take and kept both cameras' scratch mics running, this is fast — we feed the cameras and the recorder clips to the waveform auto-sync, and it matches them. For any take the software struggles with, we fall back to the clap by hand.
FIGURE CS2.4 — Syncing interview take 3 at the clap (what auto-sync is matching)
BEFORE (three separate recordings, started at different instants)
CAM-A scratch ─────╮_______________/\_________________ ← clap spike
CAM-B scratch ──────────╮__________/\_________________ ← same clap, offset
REC lav+boom ───╮__________________/\________________ ← same clap, clean audio
AFTER (all three aligned on the shared spike → one synced clip)
V Cam A ____________________________/\________________ picture (main angle)
V Cam B ____________________________/\________________ picture (wide angle)
A Recorder (lav+boom) _____________/\________________ the clean sound we'll actually use
^ clap frame: three spikes stacked = the take is locked
The thinking: the clap gives every source one shared transient, so aligning that single frame syncs the entire take across both cameras and the recorder at once. The camera scratch audio — which we'll never put in the final mix — earns its whole existence right here as the reference the sync matches against. Once synced, each interview take becomes a single tidy clip carrying good picture and good sound, and it goes into the Cam A / Cam B bins ready to select from. What got adjusted: auto-sync nailed five of six questions instantly; on question four, someone had bumped the camera mic and the scratch was too quiet to match, so we synced that one by hand off the clap spike — thirty seconds of manual work, entirely because a clap existed. Question four is the reason you always slate.
⚠️ What would have happened with no slate. Imagine this same shoot with the camera mics off and no clap. Every one of the six questions, several takes each, would have to be nudged into sync by ear, guessing at the frame where lips and voice meet, with nothing to check against. A ten-minute job becomes an afternoon of tedium and doubt. The 90 seconds of clapping on set bought back an afternoon in post — "fix it in pre, not in post" written in a single hand-slap.
Phase 4 — Proxies: make it play on a real laptop
4K Long-GOP from Camera A stutters on the edit laptop the instant we scrub — expected. So we generate proxies: half-resolution, in an easy all-intra editing codec, parked in the 06_PROXIES/ folder. The computer grinds through the render while we get coffee; when it's done, the timeline plays like glass.
FIGURE CS2.5 — The proxy round-trip for this project
4K/25 Long-GOP ORIGINALS ──┐ (stutter on the laptop; scrubbing lurches)
(masters — never deleted) │
│ generate proxies once (half-res, all-intra)
▼
1080p PROXIES ◄────────────── edit on THESE ── smooth scrub, instant response
(in 06_PROXIES/) │
│ ...cut the whole documentary on proxies...
│
│ at export (Chapter 36): relink to the ORIGINALS
▼
4K ORIGINALS ────────────────► final 4K render ── the audience gets the real image
The thinking: we trade a little disk space and one up-front render for a comfortable edit — the best deal in post. The golden rule stays taped to the wall: proxies are disposable stand-ins; the originals are the master. We will not delete the originals, and the very first item on our eventual export checklist (Chapter 36) will be "confirm the project is on ORIGINALS." What got adjusted: the phone's HEVC grabs actually stuttered worse than Camera A's footage — highly compressed phone codecs are hard to decode — so we proxied those too, even though they're only a few seconds each. Heavy is heavy, regardless of the camera's price.
Phase 5 — Selects: watch everything, keep the gold
Only now, with everything synced and playing smoothly, do we look at the footage as content. We watch every clip — all six questions, every take, all the B-roll, even the parts we assume are junk (the Murch immersion move from Case Study 1). As we go, we flag keepers and, for long answers, range-select just the good ten or twenty seconds. The keepers gather into the Selects bins.
FIGURE CS2.6 — The selects pass: from everything to the keepers
ALL FOOTAGE (~45 min of interview + B-roll)
│ watch everything; flag keepers; carve the good part of long takes
▼
INTERVIEW SELECTS B-ROLL SELECTS
• Q1 "how I started" take 3 • wheel truing, close (sharp, has life)
• Q2 "the hardest repair" take 2 • hands on the wrench (great detail)
• Q3 "why bikes matter" take 4 ★★★ • the shop sign, morning (an opener?)
• Q5 "a customer story" take 1 • customer arriving (a closer?)
│
▼ ~4 min of genuine keepers — the raw material of a 3-min film
(Q4 and several weak takes stay filed as outtakes — not deleted, just out of the way.)
The thinking: selecting is ruthless subtraction. Out of 45 minutes we keep maybe four, and the discipline is to resist flagging everything — if half your footage is "selects," you haven't selected. We also mark a standout: the answer to Q3 ("why bikes matter") is the emotional heart, flagged three stars, and it will probably anchor the film. What got adjusted: Q1's best-delivered take (take 5) had a leaf-blower whining outside; take 3 was slightly less polished but clean, and clean beats polished when you can't fix the noise — so take 3 became the select. That's a selection call the sound forced, exactly the kind you can only make once you've watched (and listened to) everything.
Phase 6 — The string-out: selects end to end, then a first shape
Finally, the selects go onto the timeline, end to end, in a rough order — the string-out. No frame-trimming, no music, no transitions. Just every good piece in one watchable strip, so we can see the whole thing for the first time.
FIGURE CS2.7 — The string-out (top) becoming a first shape (bottom)
STRING-OUT (selects in the order we happened to shoot them — long and loose)
V [ Q1 ][ Q2 ][ Q3★ ][ Q5 ][ truing ][ hands ][ sign ][ customer ]
A [ ---- interview sync audio across the answers ---- ]
~6 min. Watchable, but shapeless. Now we ask: what's the STORY order?
FIRST SHAPE (reordered toward a beginning–middle–end; obvious fat trimmed)
V [ sign ][ Q1 start ][ truing ][ Q2 hardest ][ hands ][ Q3★ heart ][ customer ][ Q5 close ]
A [ room tone ][ ------------- interview + B-roll audio ------------- ][ room tone ]
^ open on place ^ B-roll over answers hides cuts ^ end on a human note
~4 min. Not trimmed to the frame, not scored — but it has a shape you could show someone.
The thinking: the string-out's only job is to get everything into one strip; the first shape is where we start ordering for story (Chapter 17's arc) — open on the shop sign to establish place, let the heart answer (Q3) land near the end, close on the human moment of a customer. We lay the room tone under the head and tail so the piece can breathe. Crucially, we do not perfect any single scene yet: we get the whole thing watchable top to bottom first, because you can't judge the pace of the Q3 answer until you've felt what comes before and after it. This is a rough cut in its earliest form — every beat present, roughly the right length, deeply unpolished, and exactly what Chapters 28–30 need as raw clay.
A few of the ordering calls are worth naming, because they show how a string-out thinks even before the real cutting starts. We moved Q3 — "why bikes matter," our three-star heart — from where it was shot (third) to near the end, because an emotional high lands harder as a destination than as a middle stop. We pushed Q2 ("the hardest repair") up front-ish, because it's concrete and visual and pairs naturally with the wheel-truing B-roll, giving the viewer something to watch while they settle in. And we ended on the customer arriving rather than on a talking-head answer, because a documentary about a person and a place wants to close on life continuing, not on a sentence. None of these is final — the whole point of a rough cut is that it's re-orderable — but making them now, roughly, means the next pass (Chapter 28's real cutting) starts from a shape rather than from a shapeless pile. That is the difference between editing forward and editing in circles.
What got adjusted, one last time: watching the first shape through, the middle sagged — two interview answers back to back with no B-roll between them felt like a lull. We didn't fix it (that's the fine-cut work of Chapter 29); we just noted it, in the DOCS folder, as the film's known weak spot. A rough cut's job isn't to be good. It's to make the film's problems visible so the later passes know where to aim. Ours just told us exactly where it needs help — which is a rough cut doing its job perfectly.
✂️ In the Edit. Look at the bottom timeline and notice how much of it is only possible because of decisions made on the shoot. B-roll laid over the interview audio (hiding the cuts between takes) exists because Chapter 20's coverage got shot. The room tone under the open and close exists because Chapter 15's discipline recorded it. The clean Q1 take exists because we had options to choose from. The string-out doesn't just reveal the story — it reveals, with total clarity, whether you shot for the edit. Here, we did, and the edit has somewhere to go.
What the whole thing cost — and bought
Beginners resist edit-prep because it feels like time stolen from "real" editing. So let's look at what this pipeline actually cost on this shoot, honestly, and what it bought back.
FIGURE CS2.8 — Edit-prep time on this shoot (rough, one-person crew, ~120 GB)
PHASE TIME MOSTLY WAITING? THE PAYOFF
─────────────────────────────────────────────────────────────────────────────────
1 Offload (verified) ~35 min yes (copying) 2 safe copies; 1 corrupt file caught
2 Organize (folders/bins) ~20 min no everything findable in seconds
3 Sync ~15 min no 6 questions married to clean sound
4 Proxies ~40 min yes (rendering) a timeline that plays like glass
5 Selects ~55 min no (watching) 4 min of gold out of 45
6 String-out + shape ~30 min no a watchable rough shape, top to bottom
─────────────────────────────────────────────────────────────────────────────────
TOTAL "hands-on" ~2.5 hrs, and a big chunk of it is the computer working while you don't.
Two things stand out. First, a large share of the clock — the offload and the proxy render — is the machine working while you make coffee or watch footage; the true hands-on cost is smaller than it looks. Second, and more important: nearly all of this is time that would otherwise be spent during the edit, except worse. The clip you didn't organize you will hunt for mid-cut, breaking your flow; the audio you didn't sync you will fight one take at a time; the footage you didn't proxy will stutter under you for a week. Edit-prep doesn't add time to a project. It moves time to the front, batches it, does most of it while the computer grinds, and converts a hundred small future interruptions into one calm afternoon. That's the trade, and once you've felt it you never go back.
There's a psychological payoff too, easy to underrate. When you finally open the timeline to cut, the project gets out of your way. Nothing is offline, nothing stutters, the good takes are flagged, the sound is married to the picture. Your entire attention is free for the only question that's actually art — what goes next, and why? That freedom is what the two-and-a-half hours really bought. It is the same freedom, at a smaller scale, that let the editor in Case Study 1 build the impossible.
Discussion questions
- In FIGURE CS2.1, the good audio lives entirely apart from the picture. Why is that deliberately true of a well-run shoot, and what chapter's discipline put it there?
- Verification caught one corrupt file on day one (Phase 1). Walk through what would have happened to that clip if the offload had been a plain drag-and-drop instead.
- The proxy round-trip (Phase 4) proxied even the tiny phone grabs. Explain, using Chapter 3's ideas, why a few seconds of phone HEVC can stutter worse than 4K from a "real" camera.
- In the selects pass (Phase 5), a cleaner take beat a better-delivered one because of background noise. Why does sound so often win these calls, and what does that say about the "sound is half the picture" idea?
- Compare the two timelines in FIGURE CS2.7. Name three specific decisions that turned the shapeless string-out into a first shape, and which earlier chapter each one draws on.
- Why is it a mistake to perfect the Q3 answer scene before the rest of the film is roughed out? What can you only judge once the whole piece exists?
Your turn: run the pipeline on your own shoot
Take your Project 2 documentary media — or any real multi-source shoot you have — and run this exact walkthrough end to end.
- Inventory it like FIGURE CS2.1: list every source, its codec, and where the audio lives. Seeing the mess plainly is the first step to taming it.
- Offload verified to two drives and fill in a one-line log per card. Note whether verification catches anything.
- Build the folders and mirror them as bins, giving room tone and selects their own homes.
- Sync your double-system audio — by waveform if you can, by the clap if you must — and turn each take into one synced clip.
- Proxy anything that stutters, phone footage included, and confirm the timeline now plays smoothly.
- Do a ruthless selects pass — watch everything, keep maybe a tenth — then lay a string-out and nudge it into a first shape with a beginning, middle, and end.
Stop there. Do not trim to the frame or add music — that's Chapters 28–30. When you can watch your rough shape top to bottom and say the one sentence of story it tells, you've done exactly what a professional does before the "real" editing starts. You'll have turned a bag of cards into a film waiting to be cut.
Key takeaways
- A normal shoot returns a mess — media scattered across devices and codecs, audio living apart from picture, nothing named or synced. The pipeline, not talent, is what turns it into an edit.
- Verify at offload. A checksum copy caught a corrupt file on day one that a drag-and-drop would have hidden until it broke the edit weeks later.
- Organize before you sync, sync before you select. Each phase makes the next one possible: homes first, then married sound, then chosen keepers, then a string-out.
- Proxy by the footage, not the camera's price. Heavy Long-GOP stutters whether it's from a cinema camera or a phone; proxies fix both, and originals stay sacred.
- Selecting is ruthless subtraction, and sound often decides the call — a clean take beats a better-delivered noisy one when you can't fix the noise.
- Get the whole piece to a rough shape before perfecting any scene. The string-out reveals the story and reveals, unmistakably, whether you shot for the edit.
- Every powerful move in the final timeline — B-roll over the cuts, room tone under the ends, a clean take to choose — was earned on the shoot. Edit-prep is where production and post finally shake hands.