skills/ad-creative/references/short-form-video-specs.md
The platform-craft layer beneath any 9:16 video for TikTok, Reels, or Shorts — the constraints that decide whether a good idea survives contact with the feed — plus a tiered library of creator/UGC and founder formats that consistently perform for growth and paid.
Part 1 (the spec) applies to every vertical video this skill produces — the iMessage reveals in imessage-video-ads.md, the motion ads in motion-video-ads.md, and the creator formats below. Part 2 is the format library.
object-fit: cover) — mixed source resolutions are fine.Platform UI covers the frame edges — the action rail, caption stack, music button, and account row all sit on top of your video. Text or key visuals in those bands get covered. Keep everything inside the cross-platform safe band — the worst case of TikTok and IG Reels margins on a 1080×1920 canvas:
| Edge | Keep clear | Why |
|---|---|---|
| Top | 220px | TikTok tabs + IG account row |
| Bottom | 500px | Caption / music / CTA stack (both platforms) |
| Left | 180px | Symmetry with right |
| Right | 180px | Action rail (like/comment/share/music) |
Result: a 720×1200 centered text band, from y=220 to y=1420. Compose all captions and load-bearing visuals inside it. Preview against a safe-zone overlay before a big push. (These numbers drift with app updates — re-verify occasionally; they're a well-sourced worst-case, not a permanent law.)
White fill, black outline, no background pill — the native look that reads as organic, not as an ad:
color: #fff;
font-family: "TikTok Sans", sans-serif; /* or a close variable sans; embed it, don't assume it's installed */
font-weight: 700;
paint-order: stroke fill; /* stroke behind fill — keeps glyphs crisp */
-webkit-text-stroke: 8px #000;
text-shadow: 0 2px 10px rgba(0, 0, 0, 0.35);
document.fonts.ready) so sizing uses the real face, not a fallback.Renders must be reproducible: no clocks (Date.now()), no Math.random(), no network fetches at render time. Same inputs → same MP4, every time. (Applies whether you're on Remotion, HyperFrames, or an ffmpeg pipeline — see the video skill for framework choice.)
UGC- and creator-driven short-form formats that reliably perform for growth and paid. Each is a structure, not a script — feed it your own footage and hook. All obey Part 1.
Tiers rank a format on one axis: does it scale a cold ad into net-new audiences (a "unicorn scaling" format), or does it just convert people already in mid/low funnel (a "supporting cast" format)? S = the rare formats that both scale cold and carry heavy education. A = scales up well. B = solid supporting cast under the right conditions. C = situational or operationally complex (rights, specific talent, or better faked than captured). Build a portfolio across tiers — don't expect every format to scale. Meta's persona-based delivery is why creator-fronted formats (Yapper, Investigation, Authority, VSL) rank so high: they reach personas natively through the creators those personas already follow. For the full 51-format taxonomy and where each sits, see meta-creative-formats.md (companion reference) and the tier/portfolio logic in ads/references/meta-decision-system.md.
Shape: creator reaction clip with a hook caption → hard cut to an app/product demo screen recording. ~9–12s total.
[ reaction · ~3s · hook caption ] → [ demo · full length · optional payoff caption ]
Shape: silent, fast tutorial. Fullscreen intro → 50/50 split (typing/action on one half, live result on the other), step captions at the seam. The "…but no yapping" promise = pure value, no talking.
[ intro · fullscreen · hook ] → [ split: input | output · ordered step captions at the seam ]
Shape: one video plays fullscreen; the creator is cut out of their background (greenscreen/segmentation) and composited on top — reacting to or narrating over the underlying content. Optionally start centered, then shrink/drag into a corner so the underlying video takes over.
[ fullscreen video (e.g. a screen recording / another post) + creator cutout overlay · optional hook text ]
Shape: one creator talks straight to camera, telling a personal story that lands on your product. No cuts required — the story is the ad. ~20–60s.
[ creator talking to camera · hook line first · personal story → product as the resolution ]
Shape: a creator "investigates" your product, niche, or a question on the viewer's behalf — visiting places, comparing options, testing claims. The discovery arc is the retention engine.
[ creator sets up the question · goes and investigates (real footage) · lands on your product as the finding ]
Shape: position the brand as the underdog (David) against a big industry, incumbent, or broken status quo (Goliath). Root-for-you storytelling.
[ name the Goliath (the villain / broken norm) · the brand's fight against it · why you win / how you're different ]
Shape: a credentialed expert — doctor, dermatologist, engineer, practitioner — presents or endorses the product on the strength of their expertise.
[ expert on camera (credentials clear) · the problem in their domain · why this product is the right answer ]
Shape: long-form (60s to several minutes) direct-response video that educates before it sells — problem → mechanism → proof → offer.
[ hook + problem · why it happens (the mechanism) · the solution + proof · the offer + CTA ]
Shape: the creator talks over full-frame imagery — screenshots, product shots, charts, a competitor's page — pairing an educational take with the visual it references. (Distinct from Format 3's reaction: this is a teaching overlay, not a reaction to a post.)
[ creator cutout + full-frame reference imagery behind them · educational narration keyed to what's on screen ]
Shape: two people in a real exchange — interview, dialogue, back-and-forth — where the product surfaces naturally in the conversation.
[ two people talking · a real question/answer exchange · product enters as part of the dialogue ]
Shape: react to, duet, or stitch another creator's video — your commentary alongside or after their clip.
[ original creator's clip · your reaction / duet / stitch responding to it ]
Shape: sensory-forward, sound-led video — tapping, unboxing, application, texture — with the product as the sensory object.
[ close-up sensory action · product-forward · ASMR audio carries (no VO) ]
Shape: person-on-the-street questions — real or recreated — capturing candid reactions to your product or category question.
[ on-the-street setup · question to passersby · candid answers → your angle ]
For founder-led video ads and organic-native brand content, four narrative structures (Oren John) give a founder something to say, and a shooting + edit system makes it fast to produce. These aren't a separate tier — they're the story arc inside a Yapper, Investigation, or vlog. Founder's content is typically a brand's first top performer: telling the story of why you built the brand auto-connects with same-problem buyers.
The four structures (pick the arc, then shoot to it):
The three-capture shooting system (makes any of the above fast):
The 0.5–1s cut formula (the edit): every shot is 0.5–1 second — a 45-second voiceover becomes ~45 one-second shots. Record the voiceover/talk track first, lay clips under it, reorder, trim. Cut in CapCut or Instagram's Edits app — don't reach for Premiere/DaVinci. This cut cadence is the vlog-speed cousin of Format 1's hard cut, and it's what makes the footage read as energetic rather than slow.
Vertical-video spec (safe-zone band, caption recipe, auto-sizing, organic-vs-baked audio) and the first three creator formats are distilled from Daniel Hangan's reelclaw-templates (built on HeyGen's HyperFrames; TikTok Sans redistributed under SIL OFL 1.1) — patterns credited, no code vendored. The tiered format library (Yapper, Investigation, David & Goliath, Authority, VSL, and the tier logic) is adapted from Dara Denney's Meta creative-type tier list; the founder / organic-vlog structures, three-capture shooting system, and 0.5–1s cut formula are adapted from Oren John's vlog + yapping playbooks — sources credited, expressed originally. Safe-zone numbers are a cross-platform worst case; re-verify against current app UI. For framework/tooling choices to actually render these, see the video skill.