Case Study 2: Shoot-Along — Directing a Real Person to a Relaxed Take

This is a from-scratch production walkthrough in the book's talking-head setup — one real, non-actor subject at a table, directed to a natural performance. Where Case Study 1 analyzed a legendary scene, this one puts you in the director's chair with a nervous human and a phone. It is deliberately doable by one person today: a friend, a window, and something to mount the camera on. We'll shoot the same short answer several times and watch it go from frozen to true — because that transformation, not any setting, is the whole job. Read it with your phone in one hand and a willing person in mind.

The brief and the constraints

The brief: capture a 45-second, natural-sounding answer from a real person who knows something — for our walkthrough, a ceramicist talking about why they make bowls by hand. The subject is a genuine expert and a genuine non-actor: comfortable at the wheel, terrified in front of a lens. Our entire job is to get them to forget the camera and simply talk the way they'd talk to a curious friend. This is exactly Project 1's challenge, and exactly the daily reality of paid work — the client is almost always a nervous real person, not an actor.

The constraints (chosen so anyone can do this today):

  • One subject, one non-actor. A friend, a family member, a coworker with a craft or a passion. The more they know their subject and the less they've been on camera, the better this teaches.
  • One camera — a phone is perfect and, as we'll see, actively helps.
  • A table by a window — the talking-head's natural home. The window is our light (we're not lighting yet — that's Part III; here we just use what's free and keep it consistent).
  • One microphone if you have it — a lav or the phone held close; if not, get the camera near and record in a quiet room. Sound is half the picture (Chapters 14–15), and a nervous voice recorded badly is a double loss.
  • 30–40 minutes. Most of that is not rolling. Most of directing a real person is the talking you do around the takes.

The single discipline that makes or breaks the whole shoot: you are directing a nervous system, not a face. Every choice below serves one goal — turning the subject's adrenaline down until their real self can come through the lens.

Gear and settings

You need almost nothing: a phone or any camera, a way to mount it at eye height (a small tripod, a stack of books on the table, a shelf), a chair for the subject with a defined spot, and a window. If you have a lav mic or can get the phone close, capture clean audio; if not, the directing lessons land silently.

⚙️ Settings Box — the directed talking-head (a starting point, not a recipe)

Setting Choice Why
Resolution 1080p or 4K 4K lets you punch in for a "second angle" in post to hide trims; 1080p is plenty otherwise.
Frame rate 24 or 30 fps Match the rest of Project 1 — frame-rate continuity across a piece (Ch.2).
Shutter ~1/50 s (24 fps) / 1/60 s (30 fps) The 180° motion look (Ch.2); lock it so it never shifts.
Aperture Wide-ish (e.g. f/2.8–f/4 equivalent) Soft background separates the subject and quiets a busy room (Ch.4 §4.3) — but not so shallow they drift out of focus if they move.
ISO As low as the window allows a clean exposure Noise is ugly on faces; let the window do the work (Ch.5).
White balance Set manually to the window Lock it — auto WB will shift between takes and break continuity (Ch.9).
Focus Locked on the subject's eyes at their mark Hunting autofocus on a face is fatal; lock it where their eyes will be.
Lens / distance ~50–85mm-equiv, framed medium/medium-CU A slightly longer lens flatters a face and lets you sit back, less "in their space" — which relaxes people.
Camera height The subject's eye level Eye level reads as equal and honest (§10.4); a low table angle quietly makes them look distant.
Eyeline Off-lens, to you beside the camera The "overheard" interview look (§10.4) is far more relaxing for a non-actor than the bare lens.
Mic Lav or close phone mic, out of frame Close, clean sound; and react silently so your "mm-hm" doesn't land on the track (Ch.14).

The through-line of the whole box: set every technical thing once, lock it, and then forget it — so that when the subject sits down you are 100% a director and 0% an operator. You cannot relax a person while you're fiddling with focus. Get the gear out of the way first.

The setup: block for a person, not a pose

Before the subject arrives, we block the talking-head — and yes, even a person sitting at a table is blocked (§10.2). We decide where the chair goes (their mark), where the camera sits (eye level, on a third), where we sit (just beside the lens, at the subject's eye height, so the off-lens eyeline is a hand's width off the camera), and how far the subject is from the background (far enough for soft separation). We put their hands where they'll have something to do — here, a half-finished bowl on the table in front of them.

FIGURE CS2.1 — Blocking the directed talking-head (top-down view)

        [ WINDOW / soft key ] ☀
        =======================================
                 \  daylight across the subject's face (camera-left)
                  ↘
                 ( S )  seated at the table, framed on the right third (Ch.6)
                  |\        a half-finished bowl on the table -> the hands' task
                  | \
       eyeline    |  ` ` ` ` ` ` ` `> looks JUST OFF the lens at YOU
       (off-lens) |                    (a hand's width off-axis, at eye height)
                  |                   ((•  lav / close mic, out of frame
                  |
               [ CAM ]  ~50-85mm-equiv, at the subject's EYE LEVEL, on a third
                  ·
                 (you)  seated right beside the lens, at eye height, asking questions

   MARKS:   the chair is the subject's mark; the bowl marks the hands.
   DEPTH:   subject well forward of the back wall -> soft, calm background.
   The person is blocked for comfort and connection, not frozen into a pose.

Read it as a floor plan. The window keys the subject's face from camera-left (the free, motivated light we'll formalize in Chapter 11). The camera sits at the subject's eye level — not on the table looking up, the classic amateur error that makes a subject read as distant. You sit right beside the lens at the same height, so when the subject talks to you, their eyeline is a whisker off the lens: the warm, "overheard" interview look. The bowl on the table gives their hands a real task. And the subject is placed well forward of the back wall, so the background falls soft and calm behind them. Everything about the blocking is designed to do one thing: make a frightened person feel like they're having a chat, not facing a firing squad.

The shoot, phase by phase

Phase 0 — before you roll (which is when the directing actually starts)

The subject arrives nervous — they've been dreading this. We do not sit them down and start. We get them a glass of water, talk about their drive, ask how the studio's going. We show them the setup casually so the camera stops being a mystery. And — crucially — we start recording during this chat, quietly, with no announcement. There is no "Action." There is no bright line to seize up at. We're just two people talking, and the camera happens to be on.

The thinking: the most valuable minutes of a real-person shoot are the ones the subject doesn't think count. Adrenaline is highest at the start and drains as the person habituates to the room. Our job in Phase 0 is to let it drain — and to be rolling while it does, because the relaxed gold sometimes shows up before the "real" interview ever begins.

Phase 1 — the frozen first take

We slide from the chat into the first question — "So tell me why you make your bowls by hand" — and watch the freeze arrive on schedule. The moment it becomes "the interview," the subject's whole body changes: they sit up too straight, the voice climbs and goes formal, the eyes flick to the lens, and out comes a stiff, generic sentence that sounds nothing like the person we were just chatting with.

FIGURE CS2.2 — "Take 1: the freeze"        [constructed teaching example — shoot-along]
  THE FRAME    Medium of the subject at the table. Correct: on the third, eye-level, soft background, window
               key camera-left. Everything technical is right.
  THE MOVE     Locked off. Still.
  THE LIGHT    Soft window key, gentle falloff. Flattering. Not the problem.
  THE SOUND    A voice gone thin and high — "radio voice." Fast, careful, over-articulated. The words are
               generic: "I'm very passionate about the craft of ceramics and I really value handmade..."
  THE CUT      Would cut cleanly. Technically usable. Humanly dead.
  THE EFFECT   A person performing "being interviewed." The eyes keep darting to the lens. We don't believe
               a word, not because it's false, but because it's *frightened*.
  THE LESSON   A technically perfect take of a frozen subject is a failed take. Everything the camera can
               control is right, and it doesn't matter. The problem is the nervous system, and it's ours to fix.

We do not panic and we do not say "relax" (the word makes it worse). We do not say "great, can you do that again with more energy" (directing a result, §10.3, which would only pump up the fakeness). We keep our own tone easy — the set takes its temperature from us (§10.1) — and we change the circumstance.

Phase 2 — the adjustment: give the hands a task, and ask instead of instruct

Two moves, together. First, we point at the half-finished bowl: "Actually — show me what you were doing with this one. Walk me through it." The subject's hands go to the clay, their attention snaps to the task, and their eyes come off the lens. Second, we stop instructing and start asking — real, specific, memory questions: "What was the first bowl you ever made that you were proud of?" A generic prompt gets a generic answer; a specific memory gets a specific human.

The thinking: every relaxation tool shares one mechanism — redirect attention outward (§10.5). The bowl pulls their focus onto a task they love and know; the memory question pulls it onto a real moment. Self-consciousness is attention pointed inward, and a person genuinely absorbed in describing their first proud bowl simply has no attention left over to watch themselves. We didn't calm them down with soothing words. We gave their mind somewhere better to be.

FIGURE CS2.3 — "Take 3: loosening"        [constructed teaching example — shoot-along]
  THE FRAME    Same clean framing — but the body has changed. Shoulders down, a slight lean over the bowl,
               hands moving as they talk.
  THE MOVE     Still locked off; all the new life is in the subject.
  THE LIGHT    Unchanged — same locked window key. (We never touched the gear; we changed the person.)
  THE SOUND    The voice has dropped into its real register — lower, warmer, with actual rhythm and the odd
               unguarded laugh. Real words now: "Honestly the first one that didn't collapse, I kept it, it's
               terrible, it's on my shelf..."
  THE CUT      Cuts beautifully, and now there's something worth cutting *to* — the hands on the clay.
  THE EFFECT   We start to believe them. The darting eyes have settled onto the bowl and onto us. A person is
               emerging from behind the performance.
  THE LESSON   You don't fix a frozen take at the camera; you fix it by changing the subject's attention.
               A task for the hands plus a real question does more than any amount of "just relax."

Better — much better. But "loosening" is not yet "true." The subject is aware they're loosening, which is its own small self-consciousness. For the last step we need them to forget the shoot entirely.

Phase 3 — the eyeline, the silence, and reacting with the face

Three small directing moves refine the take. We make sure the eyeline is landing just off the lens, on us — we keep our own face at lens height and engaged, because the subject mirrors us (§10.5). We start letting silences sit: when they finish an answer, instead of jumping in, we hold a beat, and — reliably — they fill it with something they hadn't planned, which is usually the truest line of the day. And we do all our reacting with our face, not our voice — nodding, smiling, leaning in — so our encouragement fuels them without landing on the audio track.

The thinking: these are the finishing tools. The eyeline keeps the connection warm; the held silence harvests the unguarded material; the silent reaction gives the subject the engaged listener a bare lens can't. Notice none of them is a note to the subject — we're not directing them at all anymore. We're managing the room so the performance happens on its own. That's the deepest version of §10.1: sometimes directing is entirely about what you do off-camera.

Phase 4 — the steal

We have plenty of good material. Now we go for the best of all. We say, warmly, "That's perfect — I think we've got it, thank you so much." The subject exhales, sits back, the last of the performance drops away — and, still rolling, we ask, almost as an afterthought: "Just curious — what do you love most about it, really?" And now, believing the pressure is off, they give us the answer we came for: unguarded, specific, and completely true.

FIGURE CS2.4 — "Take 7: the steal (after 'that's a wrap')"        [constructed teaching example — shoot-along]
  THE FRAME    Same framing, but the whole person has softened — sitting back, half-smiling, hands loose.
  THE MOVE     Locked off, still rolling past the "end."
  THE LIGHT    The same locked window key it's been the whole time.
  THE SOUND    Their real voice, all the way home — slower, warmer, a small laugh, a pause, then: "...I like
               that someone eats out of something my hands made. That's the whole thing, really."
  THE CUT      The take. This is the spine of the 45 seconds; everything else supports it.
  THE EFFECT   We completely believe them, because they've stopped performing and are simply telling the
               truth to a person they've forgotten is being recorded.
  THE LESSON   The relaxed take is the only take that matters, and it often arrives *after* the subject thinks
               you're done. Keep rolling. The steal is the payoff of every relaxation move that came before.

Put the four figures side by side and you have the whole chapter in one shoot. Take 1 (frozen) and Take 7 (true) are identically framed, lit, and exposed — the camera did nothing different. Everything that changed happened in the person, and everything that changed the person was directing: an easy tone, a task for the hands, real questions, a held silence, an engaged silent face, and the steal after the "wrap." That progression — frozen to loosening to true — is the arc you are learning to produce on demand, with anyone, on any shoot.

The edit pass

Drop the takes into any timeline. The directing did the hard part; the edit just protects it.

FIGURE CS2.5 — The talking-head assembly (tracks: V=video, A=audio; cuts marked | )

  V1  [ TAKE 7 main (the steal) ####|#### CUTAWAY: hands on the bowl #### |## TAKE 7 continues ##### ]
  A1  [ TAKE 7 sync audio — the real answer, continuous ################################################ ]
  A2  [ room tone bed ................................................................................. ]
                          ^ trim a stumble here, HIDDEN under the cutaway     ^ back to the face to finish

   The spine is one continuous relaxed answer (A1). Where we trim a stumble or an "um," we cover the
   video join with the hands-on-the-bowl cutaway (shot in Ch.9's Production Checkpoint) so the picture
   never jumps. Sound leads; picture hides the seams. That's the whole edit.

Decision 1 — build on the relaxed take. We choose Take 7 as the spine, because a relaxed take with a small flaw beats a stiff take that's technically clean, every time (§10.5). The whole edit is in service of the one true answer.

Decision 2 — trim the stumbles, hide them under the cutaway. Real speech has "um"s and false starts. We trim them out of the audio — which creates a jump in the picture (§9.4) — and we hide each jump under the cutaway of the hands on the bowl, which is exactly the matching insert Chapter 9's Production Checkpoint had us shoot. The task we gave the subject to relax them in Phase 2 turns out to also be the shot that saves the edit. Directing and shooting-for-the-edit were the same act.

Decision 3 — let the sound lead. We keep the real answer running continuously on the audio track and cut the picture around it — to the hands, back to the face — so the seams disappear behind a voice that never stops. (This is a preview of the J- and L-cuts of Chapter 28; here you just feel why they exist.)

✂️ In the Edit: the directing is the footage. Notice that the edit was easy — and it was easy because the directing was good. A relaxed spine take, a cutaway shot on set, clean sync sound: the timeline practically assembled itself. Had we shot only the frozen Take 1, no editing skill on earth would have rescued it, because the raw material would have been dead. The lesson of the whole chapter, felt at the timeline: the performance you direct on set is the only raw material the edit has to work with. Direct a true take, and the edit is a formality. Direct a frozen one, and there is nothing to cut.

What we'd adjust next time

Every shoot teaches its refinements. On this one, the notes that recur:

  • Roll even earlier. We nearly missed a lovely, unguarded line in the Phase 0 chat because we started recording a beat too late. On a real-person shoot, when in doubt, you're already rolling.
  • Prepare more real questions than you think you need. We ran low on specific, memory-triggering questions and started drifting toward generic ones — and the generic ones got generic answers. Come with a deep list of specific prompts.
  • Watch your own audio discipline. Twice, our "mm-hm" landed under the subject's best line. React with the face, keep the voice silent — write it on your hand if you have to.
  • Give a second task in reserve. When the bowl stopped being novel, the hands went still and a little of the tension crept back. Have a second thing for the hands (a tool, a second piece) ready.

Troubleshooting: when the subject won't thaw

You did everything above and the subject is still frozen. It happens — some people are genuinely camera-shy, and thawing is a skill that sharpens with reps. Here's the diagnostic table, most to least common cause.

Symptom Likely cause Fix
Formal "radio voice," generic words They're performing "being interviewed" Change the circumstance: a task for the hands + a specific memory question. Stop the "interview."
Eyes darting to the lens No safe place to look Sit at the lens, at eye height, and give them you to talk to (off-lens eyeline). Say "just talk to me."
Answers are one stiff sentence, then silence Yes/no or abstract questions Ask "tell me about the first time..." or "walk me through..." — open, specific, concrete.
They get more tense with each take You're directing results ("more energy!") Stop giving notes. Change the circumstance, roll long, and let a take be a "rehearsal that doesn't count."
Great in the chat, frozen "on camera" The bright line of "Action" Never announce it. Roll through the chat and slide into questions with no starting gun.
Good, but never truly natural They know they're "doing well" The steal: say "that's a wrap," keep rolling, ask one more real question.
Voice fine, but the face is dead No one to react to; a bored room Give them an engaged, warm, silent face at the lens. The subject mirrors the director.

Notice how few rows mention the camera. Almost every fix is a directing move — a change to the circumstance, the eyeline, the question, or the room's tone. That is the entire thesis of the chapter, proved in a troubleshooting table: when a performance is dead, the answer is almost never in the settings.

Discussion questions

  1. Take 1 and Take 7 are identically framed, lit, and exposed. Everything that changed was directing. What does that tell you about where a beginner's effort is best spent once the basic technical craft is in place?
  2. We started recording during the "warm-up chat," before the subject thought it counted. Is there an ethical line here? How do you reconcile "roll early to catch the relaxed take" with the duty to let a real person know they're on camera (§10.3's ♿ callout)?
  3. The bowl did double duty — it relaxed the subject and became the cutaway that saved the edit. Find another example of a single on-set choice that pays off twice, once for performance and once for the edit.
  4. We refused to say "relax," "be natural," or "more energy." Why are those three specifically counterproductive, and what did we say instead each time?
  5. The "steal after the wrap" works because the subject believes the pressure is off. Does knowing this trick change how you'll behave the next time someone tells you a shoot is over?
  6. Which single item in the Settings Box, if we'd gotten it wrong, would most have sabotaged the performance (as opposed to the image)? Argue for one.

Your turn: the direct-address version (extension)

Level up by switching the eyeline from off-lens to straight down the barrel.

The brief: direct the same kind of subject to deliver a 30-second direct-address piece — talking to the lens, as a host or a testimonial-to-camera would (§10.4).

  • Give the lens a face. Direct address is harder for a non-actor because the lens is a blank eye. Put a small photo or a pair of drawn eyes right at the lens, or sit at the lens yourself and have them "talk to you, but at the camera." Direct them to imagine one specific person.
  • Keep the height honest. Camera at their eye level, so direct address reads as connection, not confrontation from below.
  • Relax first, aim second. Do all the Phase 0–4 relaxation work before you ask for the harder eyeline. A relaxed person can handle the lens; a frozen one can't.
  • Compare the two. Put the off-lens interview take and the direct-address take side by side. Which suits this subject and this message? Testimonials often work either way; hosting wants the lens. Directing is choosing.
  • Show a friend both versions and ask which one they trust more. Their answer teaches you what the eyeline choice actually does to an audience.

Key takeaways

  • You direct a nervous system, not a face. Every move on a real-person shoot serves one goal: turning the subject's adrenaline down until their true self comes through the lens.
  • Set the gear once, lock it, and forget it. You cannot relax a person while you're fiddling with focus — be 100% director the moment they sit down.
  • A technically perfect frozen take is a failed take. Take 1 was flawless and dead; the problem was never in the camera, and neither was the fix.
  • Every relaxation tool redirects attention outward — a task for the hands, a specific memory question, an engaged face to talk to. Self-consciousness can't survive genuine outward attention.
  • Never say "relax," "be natural," or "more energy." Change the circumstance instead; direct causes, not results.
  • Roll early, let silences sit, react with your face, and steal the take after "that's a wrap." The truest answer often arrives once the subject believes the pressure is off.
  • The relaxed take is the only take that matters — build the edit on it, and hide its small flaws under a cutaway.
  • Directing and shooting-for-the-edit are one act: the bowl that relaxed the subject was the same shot that saved the cut. Good directing is good footage; there is nothing else for the edit to work with.