Talking-Head Multicam Director

Make one-camera talking heads feel multicam: punch-ins, caption beats and B-roll plans

$19.90

For creators, founders, educators and editors with single-camera talking-head videos.

One locked-off camera can look like a three-camera shoot if you plan it. Give your agent a transcript (SRT, VTT or rough timings) or a script, and this skill returns an edit plan you can build in any editor:

  • Angle map: wide, medium, close, reframes and slow pushes, with the crop math that tells you how far you can punch in before the picture goes soft (for example 200% from 4K to 1080, but only about 112% when cutting a horizontal 4K frame to vertical).
  • Edit table: in/out times, angle, caption treatment, insert, sound and the reason for each change, with angle changes on new ideas, emphasis words and every jump cut.
  • Caption beats: chunking rules, three style presets, emphasis words and beat cards for numbers and list items.
  • Inserts: where B-roll earns its place, and Seedance 2.5/2.0 or MiniMax H3 prompts for faceless cutaways (hands, desk, objects, the room) that match your room's light. The skill keeps generated shots faceless, because these models restrict real-face references; extra angles of you come from punch-ins.

Included: SKILL.md, an edit-grammar reference (crop table, cut rules, caption presets, insert decision table, 10 faceless insert prompts) and a complete worked example. No script and no dependencies.

Highlights

  • Punch-in angle map with crop math for your source and delivery size
  • Cut points on ideas, emphasis words and every jump cut
  • Kinetic caption beats: base captions, emphasis words, beat cards
  • B-roll placement plus faceless Seedance/H3 inserts that match your room
  • Full worked plan for a 42-second vertical founder clip

Example

Angle map from the source
You

Shot on a 4K horizontal camera but I need a 9:16 Reel. How far can I punch in?

Agent + skill
Max sharp punch-in = source pixels ÷ delivery pixels on the limiting axis. A 9:16 crop of a 3840×2160 frame uses the full 2160 px height for a 1920 px tall delivery: 2160 ÷ 1920 ≈ 112%. So: A 100%, A-push 100→106%, and one reframe at 112% with the subject shifted toward a third. Get the rest of the variety from caption beats and inserts. Next time, shoot vertical 4K (200% headroom) or frame wider.
Edit plan rows
You

Transcript: 'Mistake two: you ask for thirty minutes. Um, that's a huge ask from a stranger. Ask for something tiny. A yes-or-no question.'

Agent + skill
| In | Out | Line | Angle | Captions | Insert | Why |
| 14.80 | 17.20 | Mistake two: you ask for thirty minutes. | A 100% | beat card #2 THE 30-MINUTE ASK | — | list item |
| 17.20 | 19.10 | That's a huge ask from a stranger. | C 185% | a HUGE ask / from a stranger | — | hides the jump cut (um removed) |
| 19.10 | 20.80 | Ask for something tiny. | B 140% | ask for something TINY | — | new idea |
| 20.80 | 22.60 | A yes-or-no question. | insert | beat card YES / NO? | generated: a finger taps the trackpad once | makes 'tiny' visual |
Faceless insert that matches the room
You

Write a Seedance 2.5 insert of my desk for the line about follow-ups. I have a frame of the empty room.

Agent + skill
Bindings: the room = @Image 1 (layout, furniture, window light, colours; ignore any people).
Talking-head B-roll insert, photoreal, 9:16, 5 s, matching @Image 1: soft daylight from frame-left, shallow depth of field.
0-5s: [Close-up, static] A phone lies face-up on the oak desk beside the white mug; its screen lights up with a soft blank glow and it buzzes once.
No people's faces in frame. No subtitles. Audio: none needed.

Use the best 1.6 s; add the notification UI as a graphic in post.

What's inside

talking-head-multicam-director/
├── agents/
│   └── openai.yaml
├── references/
│   ├── edit-grammar.md
│   └── worked-example.md
├── LICENSE.txt
├── README.md
└── SKILL.md

Install by unzipping into your agent's skills folder. Install guide →