Open any short-video feed and start scrolling. Watch what your thumb does. It flicks past a video in under a second, past another, past another — and then, on maybe the fourth or the tenth, it stops. Something in the first moment reached out of the...
Prerequisites
- 6
Learning Objectives
- Compose for a 9:16 vertical frame, keeping faces and key information inside the platform's safe zones.
- Write and build a social hook that earns the first two seconds and buys the next few.
- Cut for pace and design burned-in captions for a viewer watching with the sound off.
- Match the typical platform specs — aspect ratio, resolution, length, loudness — for Reels, TikTok, and Shorts, and know why they change.
- Repurpose a horizontal video into a vertical one by reframing, restacking, and re-hooking it.
- Read a retention curve and use it to diagnose and fix a short-form video.
In This Chapter
- Overview
- Learning Paths
- 23.1 Vertical framing and safe zones
- 23.2 The hook: earning the first two seconds
- 23.3 Pace, captions, and sound-off viewing
- 23.4 Platform specs and aspect ratios
- 23.5 Repurposing horizontal into vertical
- 23.6 Retention and what the analytics tell you
- Production Checkpoint
- Summary
- Spaced Review
- What's Next
Chapter 23: Social and Vertical Video
"A wealth of information creates a poverty of attention." — Herbert A. Simon, 1971
Overview
Open any short-video feed and start scrolling. Watch what your thumb does. It flicks past a video in under a second, past another, past another — and then, on maybe the fourth or the tenth, it stops. Something in the first moment reached out of the screen and caught you: a face already mid-sentence, a bold line of text, an action you had to see the end of. You did not decide to watch that video. It took you, before you had a conscious thought, in less time than it takes to read this sentence.
That thumb is the whole ballgame. On a social feed there is no title card you chose to click, no lean-back audience that already committed to a film, no patient viewer who will "give it a minute." There is a person half-paying-attention with autoplay running and the sound off, one flick away from something else forever. The video that survives that person is not usually the one with the best camera or the prettiest light. It is the one built, from its first frame, to stop the scroll — and then to earn every half-second after that, because the same thumb that stopped is still hovering, ready to leave the instant you get boring.
Vertical, short-form video is not "regular video, cropped." It is its own craft, with its own grammar, and this chapter teaches it. The frame is a tall column, because the phone is held upright. The opening is measured in single seconds, not minutes. The captions carry the message, because most of the audience never turns the sound on. And the reward — whether the platform shows your video to ten more people or ten million — is decided almost entirely by one thing you can now learn to design: how much of it people actually watch.
Everything you have built so far still applies. Story is still the boss (§1.3). Sound is still half the picture — even when it starts muted. You still shoot for the edit, still motivate every choice. But those principles now bend around a new set of constraints, and learning where they bend is what separates a video made for a phone from a horizontal video shoved onto one.
In this chapter you will learn to:
- Compose for vertical video — the tall 9:16 frame — and keep your subject inside the platform's safe zones, clear of the interface that covers the edges.
- Build the social hook — the first two seconds engineered to stop a scrolling thumb — and choose from a toolbox of hook types.
- Cut for pace and design captions for the majority who watch with the sound off, so your work is understood and reaches more people.
- Match the typical platform specs — aspect ratio, resolution, length, loudness — while knowing why any spec you memorize will change.
- Repurpose a horizontal video into a vertical one without simply chopping the sides off.
- Read retention and watch time — the analytics that actually matter — and use the curve to fix a video that isn't working.
Learning Paths
This is the most phone-native chapter in the book — everyone should read it, but here is where your attention pays off most:
- 📱 Phone-first: this entire chapter is home turf. Your phone is not a compromise here; it is the native instrument of vertical video, held exactly the way the format wants. §23.1 and §23.2 are your core.
- 🎥 Creator: all six sections are your livelihood. §23.2 (the hook) and §23.6 (retention) are the difference between a video that reaches your existing followers and one that reaches a million strangers. Read them twice.
- 💼 Pro-track: clients now ask for vertical cut-downs of everything. §23.4 (specs) and §23.5 (repurposing) are the deliverables you will actually be paid to produce — often from footage you shot horizontal.
- 🎓 Student: §23.1 (framing) and §23.3 (captions, accessibility) connect straight back to composition (Chapter 6) and forward to delivery (Chapter 36). This is applied composition under real-world constraints.
23.1 Vertical framing and safe zones
Hold your phone the way you actually hold it to text someone — upright, tall. That shape, roughly twice as tall as it is wide, is the native frame of social video, and it changes composition from the ground up. Vertical video is video composed for a tall frame — almost always the 9:16 aspect ratio — designed to fill a phone held upright, edge to edge, with no black bars. It is not a horizontal video turned sideways or cropped. It is a frame you compose for, from the moment you decide where to stand.
In Chapter 6 (§6.5) you met this frame briefly and saw the core problem in FIGURE 6.6: a 16:9 shot plays a lateral game — subject on a left or right third, nose room, lead room across the width. The 9:16 frame has almost no width to play with. The game goes vertical. You get closer to your subject, and you stack the elements of the shot top-to-bottom instead of spreading them left-to-right. A wide landscape becomes a tall canyon; a two-shot becomes almost impossible; a single face fills the frame comfortably and reads clearly even on a small screen. The tall frame wants a single subject, close, centered or on a vertical third.
Why does vertical work at all, when a horizontal frame suits the side-to-side sweep of the human gaze? Because the viewing situation changed. Nobody watches a feed on a wall-mounted television; they watch it on a device held in one hand, filling their whole field of view, one video at a time. The vertical frame fills that device completely, so your video occupies 100% of the viewer's attention with no letterbox bars and no competing pixels. The format is "correct" not in the abstract but for the exact way a phone is held and a feed is consumed — motivate the choice by the viewer (Chapter 6's fifth throughline), and vertical stops feeling like a compromise and starts feeling like the right tool.
The interface eats your frame
Here is the trap that catches everyone the first time. You compose a beautiful vertical shot on your camera, your subject perfectly placed. You upload it. And now, on the actual platform, a username sits across the lower-left, a caption runs along the bottom, a stack of like / comment / share / save buttons climbs the right edge, a progress bar pins the very bottom, and the app's own chrome nibbles the top. Half of what you carefully composed is now hidden behind the platform's own graphics — and you cannot move them.
The region that survives all of that is the safe zone: the part of the vertical frame that stays reliably visible and uncovered by platform interface and cropping, no matter which app plays it. (You met this idea in passing in Chapter 6; here is its formal home.) Everything that must be seen — your subject's face, the key action, your own on-screen text — belongs inside the safe zone. Everything outside it is a gamble you will usually lose.
FIGURE 23.1 — The 9:16 vertical frame and its safe zones (schematic overlay)
┌───────────────────────┐ ← TOP ~10–12%: app header, profile icon, "Following/For You"
│ x x x x x x x x x x x x│ tabs may sit here. Keep titles OUT of the top strip.
│·······················│ ┐
│· ·│ │
│· ·│ │
│· ( FACE ) ·│ │ CENTER COLUMN = ACTION-SAFE.
│· ·│ │ Keep faces, key motion, and
│· eyes on the ·│ │ anything that MUST be seen
│· upper third ·│ ├─ inside this central ~80% width
│· ·│ │ and clear of the bottom fifth.
│· ·│ │
│· ·│ │ ▓ ← RIGHT EDGE ~10%:
│· [ your caption ] ·│ │ ▓ like / comment /
│·······················│ ┘ ▓ share / save buttons
│ @username · ♪ sound │ ← BOTTOM ~15–20%: username, caption,
│ ▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬ │ audio label, and the progress bar.
└───────────────────────┘ Platform UI lives here. Assume it's COVERED.
Legend: · = safe-zone boundary x/▓/▬ = zones the platform's own UI will cover.
The proportions above are a practical working guide, not an official spec — they
vary by app and update often. When in doubt, keep it centered and give the edges away.
Read the overlay as a set of instructions you can shoot to. The center column, roughly the middle 80% of the width and clear of the bottom fifth, is action-safe: put the face there, put the motion there, put anything the video dies without there. The bottom strip belongs to the platform's username, caption, and progress bar — and, crucially, to your own burned-in captions, which we design in §23.3. The right edge belongs to the button stack. The top belongs to the app's header. Give those edges away on purpose and compose as if they are already gone, because on most of the apps, most of the time, they are.
The practical habit is simple: frame a little looser than feels right, and keep the important thing centered. A face that sits dead-center with breathing room survives every app's interface. A face jammed into the bottom-left — exactly where a beginner puts it, because that felt "balanced" on the bare frame — ends up behind the username. When in doubt, protect the center and sacrifice the edges.
🎒 Gear Note: your phone already shoots this natively — turn on the grid. The single most phone-native fact in this book: held upright, your phone shoots 9:16 vertical with no adapters, no rigs, nothing. It is the correct tool, not a downgrade. Do two things to shoot vertical well. First, turn on your camera app's grid (usually in the camera settings) so you can place a face on the upper-third line and keep it centered horizontally. Second, if your app offers it, enable grid/level guides to keep verticals actually vertical — a tilted upright frame reads as "amateur" instantly. Everything else in this chapter you can do on the device in your hand. (For a locked-off vertical talking-head, a five-dollar phone clamp on any tripod beats holding it — steady still beats shaky, exactly as in Chapter 1.)
⚠️ Common Mistake: composing edge-to-edge on the bare frame. You compose in your camera app or editor, where the frame is clean and the whole 9:16 rectangle is yours. It looks perfect. Then the platform lays its interface over the bottom and right, and your subtitle is half-buried under the username while your subject's chin is behind a progress bar. The fix is to always preview with the interface in mind — mentally (or with a safe-zone overlay template, see Appendix I) block out the bottom fifth and right tenth before you finalize a frame. Compose for the frame the viewer sees, not the frame your editor shows you.
🔄 Check Your Eye. 1. Why does a 9:16 frame push you toward a single, closer subject compared to 16:9? 2. Name three things the platform's own interface will typically cover on a vertical video, and where. 3. A creator frames a talking-head with the face in the lower-left third (as they would for 16:9). What goes wrong on the platform, and what is the fix?
Check yourself
- The tall, narrow frame has almost no left/right room to spread elements across, so you compose vertically and get closer; a single face fills it comfortably while a wide two-shot barely fits.
- Username and caption (bottom-left), the like/comment/share/save button stack (right edge), and the progress bar (very bottom) — plus app header/tabs at the top. Assume all are covered.
- The face lands in the bottom-left where the username and caption sit, so it gets partly hidden behind the interface. Fix: recompose with the face centered and higher (eyes near the upper-third line), keeping it inside the action-safe center column.
23.2 The hook: earning the first two seconds
In Chapter 17 (§17.3) you met the hook — the opening that earns the first ten seconds of any story, the reason a viewer keeps watching instead of drifting off. Short-form has a crueler, faster cousin, and it deserves its own name. The social hook is the first roughly two seconds of a vertical video, engineered to stop a scrolling thumb and buy the next few seconds. It is a hook not against boredom but against departure: the viewer is not settling in to watch you, they are actively scrolling away, and your opening has to physically arrest that motion before a conscious decision is even made.
The difference in scale is everything. A narrative hook has ten seconds to raise a question. A social hook has about two, and it is competing not with the other videos in your genre but with the scroll itself — with the near-infinite supply of something-else one flick away. This is why "start with a slow logo animation" or "hey guys, welcome back to my channel" is fatal here: every second of throat-clearing is a second the thumb uses to leave. The first frame must already be the interesting part.
🚪 Threshold Concept: you are not competing with other videos — you are competing with the scroll. The mental shift that changes everything about short-form is this. In a cinema or on a chosen YouTube video, the audience has committed; your job is to reward the commitment. On a feed, no one has committed to anything; your job is to win the commitment, frame by frame, starting at zero. That reframes every opening decision. You do not "set up" a short-form video; you detonate it. You do not ease the viewer in; you grab them mid-motion. Once you truly internalize that the default action is leaving, you stop making openings that assume the viewer will wait — because they won't.
What a social hook is made of
A strong hook usually works on three channels at once, because the viewer might be perceiving any one of them first — and remember, the sound is probably off.
FIGURE 23.2 — The anatomy of a two-second social hook (first ~0–2s)
CHANNEL IN THE FIRST TWO SECONDS… EXAMPLE
─────────────── ─────────────────────────────────────────── ─────────────────────────
VISUAL hook motion, a face mid-action, a striking image, a hand already cracking an
(the eye) a pattern-interrupt — something happening. egg into a shoe (?!)
VERBAL hook the first spoken words land ON the payoff, "This is the mistake that
(the ear, IF not on setup. No "so today I want to…" cost me the whole shoot."
sound is on)
TEXT hook an on-screen line the muted viewer reads "I edited for 6 years.
(the caption) instantly — the promise or the question. Nobody told me THIS."
─────────────── ─────────────────────────────────────────── ─────────────────────────
THE OPEN LOOP All three point at a question the viewer NEEDS answered — a gap they
(the engine) must stay to close. "Wait for it." "Here's what happened." The loop
is what converts a stopped thumb into a watched video.
The open loop is the engine underneath all three channels. You open a question, a curiosity gap, a "wait — how does that end?" — and you do not immediately close it. The viewer stays because leaving means never knowing. "I tried the cheapest and the most expensive version of this — one of them shocked me." That sentence, on screen in the first second, opens a loop (which one? why?) that the viewer has to stay to close. Master short-form creators are, above all, masters of the open loop: they are always making you a small promise that the rest of the video keeps.
Here is a toolbox of hook types you will see again and again. None is a gimmick; each is a different way to open a loop or stop a scroll:
- The result-first / payoff-first hook. Show the finished thing, the transformation, or the punchline in the first frame, then rewind to how you got there. ("Here's the final shot. Now here's how three lamps and a bedsheet made it.") The viewer stays to close the gap between the impressive result and the ordinary start.
- The bold claim. State something confident, surprising, or mildly contrarian. ("Your ring light is ruining your videos.") The viewer stays to see if you can back it up — or to argue.
- The direct question. Ask the exact thing your target viewer is quietly wondering. ("Why do your phone videos look cheap and TikToks look expensive?") It self-selects the right audience and opens the loop by definition.
- In medias res. Start in the middle of the action, already moving — no setup, no context. The viewer arrives mid-story and stays to understand what they walked into.
- The pattern interrupt. Do something visually unexpected in frame one — an odd object, a fast move, a jarring cut — that breaks the scroll's rhythm. Novelty stops the thumb; then your loop keeps it.
- The stakes / "wait for it." Name what is at risk or promise a payoff worth waiting for. ("I have ten seconds to save this shot before the light's gone.") An explicit promise that the ending is worth the middle.
🎞️ Read This Sequence. Here is a constructed hook, built to show all three channels firing at once in the first two seconds. Picture it as a muted autoplay — then imagine the sound on.
FIGURE 23.3 — "The two-second hook" (a constructed vertical opening) [constructed teaching example]
THE FRAME 9:16 vertical, tight. A person's hands, already in motion, hold a phone up to a window;
the face is half in frame at the top, mid-word. Nothing is being "set up" — we've walked
in on something happening. Big bold text sits in the center-safe zone: "STOP filming
people like THIS."
THE MOVE Handheld, a hair of energy — not locked off. The slight motion reads as "live, now, real,"
which the feed rewards; a dead-still tripod frame can read as an ad and get scrolled.
THE LIGHT Window light on the hands and face — free, soft, motivated (Chapter 13). It's flattering
enough to look intentional but natural enough to not look like a commercial.
THE SOUND IF the sound is on: the first spoken words are the payoff — "…because this one setting is
why your face looks flat." No greeting. The sentence lands on the hook, not before it.
THE CUT This IS the first shot, and it was almost certainly placed first in the EDIT, not shot
first. It cuts hard — on the word "THIS" — to the demonstration. No breathing room.
THE EFFECT Three channels hit at once: a striking image (hands + text), a promise (the bold text),
and, for the unmuted, a payoff sentence. All three open the same loop: which setting?
THE LESSON A social hook works the eye, the ear, and the reader simultaneously, and it opens a loop
in under two seconds. Setup is the enemy; the interesting part must be frame one.
Notice the field that does the quiet heavy lifting: THE CUT. The hook is not something you shoot so much as something you build in the edit. Your best two seconds — the most surprising line, the most striking image, the clearest promise — almost never happen at the literal start of your recording. They happen in the middle. The craft is to find that moment in your footage and move it to the front, then cut hard into the payoff.
✂️ In the Edit: the hook is assembled, not filmed. This is the chapter's clearest example of "you shoot for the edit," turned inside out. On the shoot, you capture the whole thing — the intro, the demo, the result. In the edit, you throw the intro away and lead with the result or the best line, because that is your hook. The single highest-leverage edit in all of short-form is choosing which two seconds go first. When a video isn't performing, the first thing a pro does is not reshoot — it is re-order: pull a stronger moment to the front and cut everything before it. (You will do this deliberately with J-cuts, hard cuts, and re-ordering in Chapter 28.)
🎬 On Set: shoot three hooks for one video. Take any 20–30 second idea you can film on your phone (a tip, a demonstration, a reveal). Shoot the same video, but record three different openings for it: (1) a bold-claim hook, (2) a result-first hook, and (3) an in-medias-res hook. Constraint: each opening must land its promise in under two seconds, with a legible line of on-screen text. Self-review: show all three to one person, muted, and ask which made them want to keep watching — and why. Keep the winner; you just A/B-tested a hook the way every serious creator does. This feeds your Production Checkpoint teaser.
🔄 Check Your Eye. 1. In one sentence, how does the social hook differ from the narrative hook you learned in Chapter 17? 2. What is an "open loop," and why does it keep a viewer watching? 3. Why is the most important hook decision usually an editing decision, not a shooting one?
Check yourself
- The narrative hook has about ten seconds to raise a question for a viewer who has already committed; the social hook has about two seconds to win commitment from a viewer actively scrolling away.
- A curiosity gap — a question or promise you open but don't immediately close; the viewer stays because leaving means never learning the answer.
- Because your most compelling two seconds almost always occur in the middle of the footage, not at the start; the craft is finding that moment and re-ordering it to the front in the edit.
23.3 Pace, captions, and sound-off viewing
Now a fact that reshapes every short-form video you will ever make: most feed video is watched with the sound off. Autoplay in a social feed starts muted by default, and a large share of viewers — very often the majority, especially on the first pass through a feed — never turn it on. This is sound-off viewing, and designing for it is not optional. If your video only makes sense with audio, most of your audience will scroll past it never knowing what it was about.
This does not contradict our second throughline — sound is half the picture (Chapter 14). It sharpens it. Sound is still half the picture; it is just the half that has to be redundant with the other half in short-form. The unmuted viewer should be rewarded richly — a good voice, a clean mix, trending audio that fits — because they watch longer and the platform notices. But the muted viewer must be able to follow the entire video from the picture and the text alone. You build for both at once.
Captions carry the message
The tool that makes a muted video legible is captions, and in short-form they are usually burned-in — rendered permanently into the picture as on-screen text, not toggled on by the viewer. There are two kinds, and the distinction matters:
- Open captions are burned into the video itself; everyone sees them, always, in the style you designed. This is the short-form default, because it guarantees the muted majority reads your words and because you control the look.
- Closed captions are a separate text track the viewer (or the platform) can switch on or off. Platforms increasingly auto-generate these, which is a genuine accessibility win — but auto-captions are often wrong, mistiming words and mangling names, so you cannot simply trust them.
♿ Accessibility & Inclusion: captions are the point here, not a bolt-on. Captions serve the huge muted-autoplay audience and every viewer who is deaf or hard of hearing — the same design decision does both jobs, which is why accessibility and reach point the same direction in short-form. Do it right: (1) Correct the auto-captions. Platform auto-captioning is a fast start, but proofread it — fix names, jargon, and mistimed words, because an uncorrected caption is worse than none. (2) Design captions to be read on a phone at arm's length: large type, high contrast, a stroke or a background plate behind the text so it survives a busy background. (3) Keep captions inside the safe zone (§23.1) — up out of the bottom strip where the platform's own username and caption live, or they collide and neither is readable. (4) Don't strobe or hard-flash text — photosensitivity-safe motion matters. The styling and animation of captions and titles is craft you will build in Chapter 34, and burning captions cleanly onto your export is covered in Chapter 36 (§36.5); the exact button in each app is in Appendix E. Here, the rule is simply: caption everything, place it safe, and check it.
FIGURE 23.4 — Caption placement in the vertical safe zone
┌───────────────────────┐
│ │ Captions go in the CENTER band, not
│ ( FACE ) │ the very bottom (that's the platform's
│ │ username/progress bar) and not the very
│ ┌─────────────────┐ │ top (app header).
│ │ BIG, HIGH- │ │
│ │ CONTRAST TEXT │ │ • 1–6 words on screen at a time, synced
│ │ WITH A STROKE │ │ to the speech (or the beat).
│ └─────────────────┘ │ • A stroke or plate so text survives any
│ │ background.
│·······················│ • Kept clear of the bottom ~18% and the
│ @username ▬▬▬▬▬▬▬▬▬▬ │ right ~10% button/UI zone.
└───────────────────────┘
Pace: cut the dead air
Short-form is faster than the video you have edited so far, and for a concrete reason: every dull half-second is an exit ramp. The discipline is to remove dead air ruthlessly — the breaths, the "um"s, the pause before the point, the wind-up before the demonstration. Where a documentary might hold a silence for weight (you will do exactly that in Chapter 30), a short-form cut keeps the information density high and the pauses out.
But — and this is the throughline talking — motivate the pace. Fast is not automatically good. Cutting for the sake of frantic cutting is as empty as any other unmotivated choice (Chapter 6's fifth throughline: motivate every choice). You cut fast in short-form because the format punishes dead air, not because speed is a virtue. The goal is not "as fast as possible"; it is "no wasted moment," which is different. A held beat that lands a joke or a reveal earns its place; a held beat that is just a person inhaling does not. Cut the second kind; keep the first.
A few concrete pace tools, most of which you will formalize in Part VI:
- Lead with motion, sustain with change. Something should be visibly changing at all times — a cut, a caption appearing, a zoom, a new shot. Static + silent = scrolled.
- Cut on the sentence, not the breath. Trim the gaps between spoken points so the ideas stack without air (a first taste of the tight interview edit in Chapter 30).
- Ride a rhythm. Trending audio and music give you a beat to cut on; matching your cuts to the beat makes even simple footage feel intentional (you will formalize cutting to music in Chapter 29 and 33).
- The loop. Design the last frame to flow back into the first, so the video plays seamlessly on repeat. A clean loop quietly multiplies watch time — the platform counts the replay.
⚠️ Common Mistake: uncorrected auto-captions and captions in the danger zone. Two failures, one fix. First, creators turn on the platform's automatic captions and never check them, so "DaVinci Resolve" becomes "the vinci resolved" and a name is mangled on screen for everyone — worse than no caption. Second, they place captions at the very bottom of the frame (where text "belongs" in a movie), so the platform's username and progress bar sit right on top of them. Fix both: proofread every auto-caption, and place your text in the center-safe band (FIGURE 23.4), never the bottom strip.
🔄 Check Your Eye. 1. Why must a short-form video be fully understandable with the sound off? 2. What is the difference between open (burned-in) and closed captions, and why is burned-in the short-form default? 3. "Cut as fast as possible" — what's wrong with this as a rule, and what's the better version?
Check yourself
- Feed video autoplays muted and a large share of viewers — often the majority on the first pass — never unmute; a video that only makes sense with audio loses most of its audience.
- Open captions are burned permanently into the picture (everyone sees them, you control the style); closed captions are a separate toggleable/auto-generated track. Burned-in is the default because it guarantees the muted majority reads your words and you control the look and placement.
- Speed isn't a virtue in itself; unmotivated fast cutting is just noise. The better rule is "no wasted moment" — cut dead air ruthlessly, but keep a held beat that lands a joke or a reveal.
23.4 Platform specs and aspect ratios
Every platform imposes and prefers a set of technical requirements: the aspect ratio it wants, the resolution and frame rate it accepts, the length it allows, the loudness it normalizes to, the caption formats it supports. Call these the platform specs — the delivery requirements a given platform expects for a given kind of video. Match them and your video looks and sounds as intended; miss them and the platform crops, squishes, re-compresses, or turns your work down until it looks amateur through no fault of your craft.
The single most important spec is aspect ratio, and short-form has settled on a small family:
FIGURE 23.5 — The aspect-ratio family (and what each is for)
16:9 9:16 4:5 1:1
┌──────────┐ ┌────┐ ┌──────┐ ┌────────┐
│ │ │ │ │ │ │ │
│ │ │ │ │ │ │ │
└──────────┘ │ │ │ │ │ │
│ │ │ │ └────────┘
horizontal │ │ └──────┘
(landscape) │ │ "tall-ish" square
└────┘ portrait
YouTube (main), Reels, TikTok, In-feed posts Older feed
TV, most Shorts — the (esp. Instagram standard;
corporate/ SHORT-FORM feed): a still safe,
broadcast work STANDARD. vertical-ish rarely ideal
compromise that now.
also reads on a
grid/profile.
Three of these you will use constantly. 9:16 is the full-screen short-form standard — Reels, TikTok, and Shorts all live here, and it is what you should shoot and cut for when the destination is a short-form feed. 4:5 is a useful in-feed portrait ratio, especially for Instagram's main feed: it is taller than 16:9 (so it commands more screen as someone scrolls a feed) but not as tall as 9:16 (so it still displays cleanly in a profile grid and doesn't get cropped). 1:1 square is the older feed standard — still safe, rarely the best choice now that full-screen vertical exists. 16:9 remains the home of long-form and most client/broadcast work, and it is the ratio you will most often repurpose from (§23.5).
⚙️ Settings Box: typical short-form vertical specs (a starting point — verify against Appendix H).
Spec Typical value (at time of writing) Notes Aspect ratio 9:16 (vertical) The full-screen short-form standard. Resolution 1080 × 1920 (1080p vertical) Upload at least this; higher is fine and re-scaled. Frame rate 30 fps (24 or 60 also accepted) 60 fps for smooth motion/slow-mo; 24/30 for a filmic look (Chapter 2). Length ~15–60 s sweet spot Maximums keep rising (minutes, not seconds) but short still wins retention. Loudness around −14 LUFS Platforms normalize loudness anyway; mix clean (Chapter 33). Codec / file H.264/H.265 .mp4, high bitrate Upload quality high; the platform re-compresses. See Appendix H. Captions burned-in and/or an uploaded caption file Do both where you can (§23.3, ♿). These numbers will be out of date the moment they're useful. Treat this box as a starting instinct, then confirm the current spec for your target platform in Appendix H (Delivery & Codec Reference) and Appendix J (Resources & Communities). The craft lesson underneath does not change: shoot and deliver at the platform's native vertical resolution, keep the loudness sane, and caption it.
The one spec everyone gets wrong is length, because it is a moving target. Platforms keep raising their maximum durations — what began as fifteen-second clips now stretches to minutes. But the maximum a platform allows and the length that holds attention are different questions. Retention (§23.6) almost always favors the shortest version that fully tells the story. Longer is permission, not advice. Make it exactly as long as the idea needs and not one second longer — a discipline, not a limit.
🔬 The Tech: what the platform does to your upload (optional). When you upload, the platform does not serve your file as-is. It transcodes it — re-encodes your video to its own codec, resolution ladder, and (usually much lower) bitrate so it streams fast to phones on cell networks. This is why a crisp export can look soft or blocky after upload, especially on high-motion or high-detail footage: the platform's compression crushed data your export had. You cannot beat the transcode, but you can feed it well: upload at the platform's native vertical resolution or higher, use a high bitrate on your own export (give the transcoder clean data to work from), and avoid extremely noisy or fast-flickering footage that compresses badly. The platform also normalizes loudness — turning loud uploads down toward a target so no one video blasts the feed — which is why chasing maximum loudness is pointless; mix for clarity, not volume. All of this is delivery craft you will formalize in Chapter 36 and reference in Appendix H. Skip this box freely; the practitioner rule is just "upload high-quality, native-resolution vertical."
🔗 Connection. Platform delivery is a whole craft of its own, and it has a dedicated home: Chapter 36 (§36.3, Platform delivery specs) covers exporting to spec for every destination, and §36.5 covers burning in and attaching captions on export. The living, always-current spec tables are in Appendix H. When a client asks for "a version for Instagram and a version for TikTok," that is a delivery task (Chapter 36) built on the framing and hook craft you are learning here.
🔄 Check Your Eye. 1. What are the four common social aspect ratios, and which is the full-screen short-form standard? 2. Why is a platform's maximum length not a target? 3. Why is it pointless to chase maximum loudness on a social upload?
Check yourself
- 16:9 (horizontal/long-form), 9:16 (vertical — the short-form standard), 4:5 (in-feed portrait), and 1:1 (square). 9:16 is the full-screen short-form standard.
- The maximum is only permission; retention favors the shortest version that fully tells the story, so length should match the idea, not the limit.
- Platforms normalize loudness — they turn loud uploads down toward a target — so extra loudness is discarded. Mix for clarity around −14 LUFS instead.
23.5 Repurposing horizontal into vertical
Most of the video in the world — including your Project 1 talking-head and your Project 2 documentary — was shot 16:9. And most of the time, the request that pays is "give me a vertical cut of that for socials." So the everyday craft is repurposing: turning a horizontal video into a vertical one that feels made-for-the-format, not chopped down to it.
The core problem is unavoidable. A 16:9 frame is wide; a 9:16 frame is tall. Fit one inside the other and you must lose most of the width — you keep only a tall central slice. If your subject was composed on the left third with nice horizontal nose room (exactly what Chapter 6 taught), a naive crop to vertical either cuts them out of frame or strands them awkwardly against a wall of empty background.
FIGURE 23.6 — Repurposing 16:9 → 9:16: three honest methods
ORIGINAL 16:9 You must lose the sides. Three ways to handle it:
┌───────────────────┐
│ ┌─────┐ │ (A) CROP & REFRAME (B) STACKED / letterbox (C) BLUR-FILL
│ │ S │ │ ┌───┐ ┌─────────┐ ┌─────────┐
│ └─────┘ │ │┌─┐│ keep the │▓▓▓▓▓▓▓▓▓│ title/space │░░░░░░░░░│ blurred
└───────────────────┘ ││S││ subject in the ├─────────┤ │░┌─────┐░│ copy of
│└─┘│ tall slice; if │ ┌─────┐ │ the 16:9 │░│ S │░│ the frame
whole width visible │ │ they move, KEYFRAME│ │ S │ │ clip sits │░└─────┘░│ fills the
└───┘ the crop to follow.│ └─────┘ │ in a band, │░░░░░░░░░│ top/bottom
Best when there's ├─────────┤ captions/ └─────────┘ bars.
room around S. │▓caption▓│ graphics fill Use sparingly —
└─────────┘ the rest. it's a fallback.
Three honest approaches, in rough order of quality:
- Crop and reframe (pan-and-scan). Slide a 9:16 window over the 16:9 footage to keep the subject in the tall slice. If the subject moves, or if two people talk, keyframe the crop so it follows the action (or cuts between two crop positions) — the same way a camera operator would pan. This is the most "native" result and the most work. It is best when you shot with room around the subject.
- Stacked / letterbox layout. Place the 16:9 clip in a horizontal band (often the upper or middle third) and use the remaining vertical space for a title, captions, or a second element. This preserves the entire original frame (nothing is cropped out) and gives you a built-in place for text. It is the workhorse for talking-heads, tutorials, and anything where losing the sides would lose information.
- Blur-fill. Put the 16:9 clip centered and fill the empty top and bottom with a blurred, enlarged copy of the same frame. It fills the screen without cropping, but it reads as a repurpose (everyone recognizes it) and wastes screen on blur. Use it as a fallback, not a plan.
The professional move is not any one of these — it is to shoot with the repurpose in mind so you have the choice. This is "shoot for the edit" (and Chapter 6's "compose for the crop") applied directly: if you know a horizontal shoot will also live vertically, frame the subject centered and a little loose, keep the essential action inside a central vertical column, and avoid putting anything you'll need out at the far edges. Then a vertical crop survives cleanly. A shot composed only for 16:9, with the subject hard on a side third, gives the vertical editor nothing but bad options.
✂️ In the Edit: the reframe is a real edit, not a checkbox. Repurposing tempts people into a one-click "auto-reframe" and a shrug. Auto-reframe tools (which track a subject and move the crop automatically) are a genuine time-saver and a fine first pass — but they guess, and they guess wrong on cuts, on two-person scenes, and on fast motion. Treat the automatic result as a rough assembly to fix, exactly like any first cut: check every shot, correct the crop where the tool lost the subject, and re-time captions to the new frame. The vertical cut is a new edit of old footage, and it deserves the same eye. (Where to find these tools in each editor: Appendix E.)
🎬 On Set: reframe a horizontal clip to vertical. Take one horizontal clip you already have — ideally from your Project 1 talking-head or Project 2 footage — and produce a 9:16 vertical version of a 15-second chunk of it. Constraint: use method (A) crop-and-reframe if there's room around the subject, or method (B) stacked layout if cropping would lose something important; then add burned-in captions in the safe zone. Self-review: play your vertical cut on your actual phone, in a feed if you can, and ask — does the subject stay inside the safe zone the whole time, and can I follow it muted? This is a direct rehearsal of your Production Checkpoint.
🔄 Check Your Eye. 1. Why can't you simply "crop" most 16:9 shots to 9:16 and keep the composition? 2. When would you choose a stacked/letterbox layout over a crop-and-reframe? 3. What one habit on a horizontal shoot makes a later vertical crop easy?
Check yourself
- A 16:9 frame is wide and a 9:16 frame is tall, so you must discard most of the width; if the subject was composed on a side third (correct for 16:9), the tall crop strands or cuts them.
- When cropping would lose essential information (e.g., a wide demonstration, a two-shot, on-screen text), the stacked layout preserves the whole original frame and gives you space for captions/titles.
- Frame the subject centered and a little loose, keeping essential action in a central vertical column — "compose for the crop" — so a vertical window survives cleanly.
23.6 Retention and what the analytics tell you
You have built the frame, the hook, the captions, and the cut. Now the video is live, and the platform makes a decision about it — and it makes that decision using one signal above all others. Retention (or watch time) is the measure of how much of your video people actually watch: the average view duration, the percentage of the video the typical viewer sees, and the shape of the retention curve — the graph of how many viewers are still watching at each second. It is the closest thing short-form has to a single scoreboard, because it is the clearest evidence that your video is worth showing to more people.
Here is the general mechanism, described the way platforms themselves describe it (the exact math is proprietary and changes, so hold the specifics loosely): a platform shows your new video to a small sample of viewers, watches how they respond — how long they stay, whether they rewatch, share, save, or follow — and then either expands distribution to more people or quietly stops. Retention is the through-line of every one of those signals. A video people watch to the end, rewatch, and share is a video the platform will keep pushing, because keeping people watching is the platform's entire business. This is why the hook (§23.2) matters so much: it protects the most fragile part of the curve.
FIGURE 23.7 — Reading a retention curve (viewers still watching vs. time)
100% ┤██▓▒░ ← the HOOK CLIFF: the steep drop
│ ░░░ in the first 1–3 seconds. A weak
75% ┤ ░░░░░ hook = a cliff here. Fix the first
│ ░░░░░░░ two seconds first.
50% ┤ ░░░░░░░░░____ ← the SAG: a slump in the middle
│ ▒▒▒▒ ▒▒ means a dull stretch. Tighten or
25% ┤ ▒▒ ▒▒▒ cut the boring part.
│ ▒▒▒▒▒
0% ┼───┬───┬───┬───┬───┬───┬───┬───┬───┬───┬──▒▒▒→ ↑ a small BUMP back UP near the end
0s 2 4 6 8 10 12 14 16 18 20 = REWATCHES / a good loop. Prized.
Diagnose by shape: cliff at the start → the hook. Sag in the middle → the pacing/content.
A late bump → people are looping it (good). Where the line falls is where the video fails.
The retention curve is the most honest teacher you have, because it tells you not just whether a video failed but where. Learn to read the shape:
- A cliff in the first one to three seconds means the hook failed — people arrived and immediately left. Do not touch the middle; fix the first two seconds (§23.2). This is the most common and most fixable problem.
- A steady, gentle decline is normal and healthy; some falloff always happens as a video runs.
- A sag or sudden drop in the middle points to a specific dull stretch — a moment where you lost them. Find that timestamp, watch what happens there, and cut or tighten it.
- A bump back up near the end means people are rewatching or the video is looping cleanly — a strong signal the platform loves. A clean loop (§23.3) manufactures this.
Beyond the curve, a handful of metrics actually matter, and it is worth knowing which to chase:
- Average % viewed / completion rate — the core retention number. The percentage of the video the typical viewer sees. This, more than raw views, tells you if the video works.
- Shares and saves — the strongest positive signals a viewer can give. A share means "this is worth someone else's time"; a save means "I want this again." Platforms weight these heavily, and they are the metric most correlated with a video reaching new people.
- Rewatches / loops — watch time above 100% means people looped it. Prized.
- Follows from a video — the sign a video didn't just entertain but converted a stranger into an audience.
- Views, likes — the vanity metrics. Pleasant, but they lag and they mislead. A video with many likes and terrible retention is a video the platform will stop showing. Do not steer by likes.
This is the single most common analytics mistake: the beginner opens their stats and reads views and likes — the numbers that feel like a report card — while the professional opens the retention curve, because it is the only metric that says what to change. Likes tell you a video was seen; the curve tells you where it lost people, which is the only information you can act on. If you check one number after posting, make it average percentage viewed; if you study one picture, make it the retention graph.
💡 Why It Works: retention is just "story is the boss," measured. It can feel like short-form runs on a mysterious algorithm, but strip away the machinery and the platform is measuring one thing: did this hold a human's attention? That is the exact question every principle in this book already serves. A strong hook, a clear story, a tight pace, sound that rewards the unmuted, a clean loop — these are not algorithm hacks, they are the old craft of holding attention, now with a scoreboard attached. The retention curve is story is the boss (§1.3) rendered as a graph. Serve the viewer's attention honestly and the metrics follow; try to trick the metric and the viewer feels it and leaves, which the metric then reports. Make it good; the numbers are downstream.
🔗 Connection: extend your Frame Log — analyze a short. Add a standing prompt to the Frame Log you started in Chapter 1: whenever a short-form video stops your scroll, pause and dissect its first two seconds — what was the visual hook, the verbal hook, and the on-screen text; what open loop did it create; and where would you guess its retention curve dips? Then, on one video you scrolled past, name the specific reason its opening lost you. Do this for one short a day. You are training the exact eye that §23.2 and §23.6 depend on, building a private swipe-file of hooks that work — and feeding the taste that becomes your reel in Chapter 39.
🔄 Check Your Eye. 1. What does a steep cliff in the first two seconds of a retention curve tell you to fix? 2. Why are shares and saves stronger signals than likes? 3. Restate the relationship between "the algorithm" and "story is the boss" in one sentence.
Check yourself
- The hook failed — viewers arrived and left immediately; fix the first two seconds, not the middle.
- A share means a viewer thinks the video is worth someone else's time and a save means they want it again — both predict a video reaching new people, whereas a like is a passive, lagging vanity metric.
- The platform is measuring whether your video holds human attention, which is exactly what "story is the boss" has always been about — retention is that principle with a scoreboard attached.
Production Checkpoint
Your task (Projects 1–2): cut a vertical social teaser. Take your finished (or in-progress) Project 1 talking-head or a piece of your Project 2 documentary footage, and produce a 15–30 second vertical teaser for a social feed. It must:
- Be reframed to 9:16 vertical, with your subject and key action inside the safe zone the whole time (§23.1, FIGURE 23.1) — crop-and-reframe or stacked layout, whichever protects what matters (§23.5).
- Hook in the first two seconds (§23.2). Lead with your single most compelling moment or line — pulled to the front in the edit, even if it happened in the middle of the footage. Open a loop.
- Be fully legible with the sound off (§23.3): burned-in captions, large and high-contrast, proofread, placed in the center-safe band; nothing essential in the platform's UI zones.
- Match the typical vertical spec (§23.4): 1080 × 1920, sane loudness, exported clean (confirm current specs against Appendix H).
Why this matters: this is the first time you deliver the same story to a different audience and a different frame — the exact skill clients pay for, because almost every finished piece now needs a vertical cut-down. It proves you can carry your craft across formats instead of starting over, and it turns work you already own into reach you didn't have. Watch it back on your actual phone, muted, and ask the only question that counts: would this stop my own thumb?
Summary
Vertical short-form is its own craft: the frame stands upright, the opening is measured in seconds, the captions carry the message, and retention is the scoreboard.
The vertical frame and safe zones:
| Rule | What to do |
|---|---|
| The frame is a column | Compose vertically; get closer; favor a single subject. |
| Give the edges away | Keep faces/action in the center; assume the bottom fifth and right tenth are covered by UI. |
| Frame a little loose | Center-safe survives every app's interface and any later re-crop. |
| Turn on the grid | Place eyes on the upper third; keep verticals actually vertical. |
The social hook (first ~2 seconds): hits the eye (a striking image/motion), the ear (the first words land on the payoff, not setup), and the reader (an on-screen promise) at once, and opens a loop the viewer must stay to close. Hook types: result-first, bold claim, direct question, in medias res, pattern interrupt, stakes. The hook is assembled in the edit — lead with your best moment, wherever it was shot.
Sound-off, captions, pace: most feed video plays muted, so the video must be understandable from picture + text alone. Burn in captions (large, high-contrast, safe-zone, proofread — ♿). Cut dead air ruthlessly, but motivate the pace — "no wasted moment," not "as fast as possible." Design a clean loop.
Platform specs (starting point; verify in Appendix H):
| Spec | Typical |
|---|---|
| Aspect ratio | 9:16 (also 4:5 in-feed, 1:1 legacy, 16:9 long-form) |
| Resolution | 1080 × 1920 |
| Frame rate | 30 fps (24/60 accepted) |
| Length | ~15–60 s sweet spot; maximums keep rising |
| Loudness | ~ −14 LUFS (platform normalizes anyway) |
Repurposing 16:9 → 9:16: crop-and-reframe (keyframe to follow action), stacked/letterbox (keeps the whole frame + room for text), or blur-fill (fallback). Best of all: shoot centered and loose so a vertical crop survives — compose for the crop.
Retention / watch time — reading the curve:
| Curve shape | Diagnosis | Fix |
|---|---|---|
| Cliff at 0–3 s | Weak hook | Rebuild the first two seconds |
| Gentle decline | Normal | Nothing |
| Mid-video sag | A dull stretch | Cut/tighten that timestamp |
| Late bump | Rewatches / good loop | Keep doing it |
Chase average % viewed, shares, and saves — not likes. The "algorithm" is just story is the boss with a scoreboard.
Spaced Review
Retrieval practice from earlier chapters — answer before checking.
- (Chapter 6) The rule of thirds, headroom, and nose room were built for a 16:9 frame. Which of these still applies in 9:16, and which becomes almost irrelevant — and why?
- (Chapter 6) Chapter 6 warned against "crop to vertical later." Using this chapter, give the two-part fix for shooting one subject that must live in both 16:9 and 9:16.
- (Chapter 7) Why does the tall 9:16 frame push you toward tighter shot sizes and away from the wide two-shot you learned to cover a scene with?
- (Chapter 7) You're building a vertical teaser from coverage you shot for a horizontal edit. Why does having shot proper coverage (wide/medium/close-up) make the vertical re-cut easier?
Check yourself
1. **Headroom** and placing the **eyes on the upper third** still apply (vertical placement matters as much as ever). **Nose room / lead room across the width** becomes nearly irrelevant, because the narrow frame has almost no horizontal space to open in front of a gaze — the game moves to vertical placement and getting closer. 2. (a) Compose the *subject centered and a little loose*, keeping essential action in a central vertical column; (b) then crop-and-reframe (or stack) for the vertical version. Shoot for the crop so both ratios survive. 3. The narrow frame can't fit a wide two-shot or a broad landscape without shrinking the subject to nothing; a single subject, closer, reads clearly on a small muted screen — the format wants intimacy, not breadth. 4. A re-cut lets you *choose* the shot that crops best to vertical for each beat (a tight close-up often survives a 9:16 crop where a wide doesn't), and gives you cutaways to hide reframes — coverage is what makes any re-edit, including a format change, possible.What's Next
You have now taken the whole production kit and bent it to the most constrained, most attention-hungry format there is — a phone held upright, a muted feed, and two seconds to win. The discipline it teaches (earn every moment, caption for everyone, deliver to spec) will sharpen everything else you make. Next, we go to the opposite extreme of control. Chapter 24 puts you into live, event, and multi-camera production — weddings, conferences, and streams, where there is no second take, no re-ordering the hook in the edit, and no fixing it later. If short-form is about the ruthless edit, live is about ruthless preparation: getting it right in the moment, because the moment is all you get.