Studio Edition

AI Music Video Director

Song to beat map, concepts, synced shot plan, Seedance/H3 prompts and edit-on-beat guide

$199.00

For artists, directors and editors making music videos with AI video models.

Music videos work when the picture follows the song's structure and the editor has enough coverage to cut on the beat. AI generation adds hard limits: short clips, lip sync that isn't guaranteed, and looks that drift between generations. This skill has your agent plan around all of that, from song facts to the final cut.

  • Beat and section map: the included stdlib Python script turns BPM, meter, first-downbeat offset and a section list (bar counts or timestamps) into bar/beat timecodes at your frame rate, a suggested cut interval per section from its energy score, every cut point labelled section/phrase/bar/beat, and phrase-aligned generation windows that fit Seedance 2.5 (30 s) or Seedance 2.0/H3 (15 s), each with the exact audio excerpt to cut. It flags timestamps that drift off the grid and suggests half-time for fast tracks.
  • Concepts: three distinct treatments with motif, chorus evolution, money moment, a 14-point AI feasibility score and a generation estimate.
  • Shot plans: coverage passes (wide, mid, close sync, detail, narrative, texture) planned per section and energy.
  • Prompts: patterns for Seedance 2.5/2.0 and MiniMax H3 lip-sync, wide, narrative and texture passes, plus image-model keyframes.
  • Edit on beat: frame math, cutting density by section, sync slipping, accents, ramps, cutdowns and delivery.

Included: SKILL.md, 5 playbooks, 9 templates, the script and 2 worked examples (a 108 BPM narrative video and a 140 BPM half-time vertical teaser). Lip sync to an existing vocal is not guaranteed by any model; the skill plans coverage and edit fixes for that.

Highlights

  • Beat-map script: BPM + sections to bar/beat timecodes and cut points
  • Phrase-aligned generation windows with exact audio excerpts for lip sync
  • Three concept treatments with an AI feasibility score and generation estimate
  • Shot plans per section: performance, narrative and texture coverage passes
  • 9 templates, 5 playbooks, edit-on-beat guide and 2 complete worked examples

Example

Beat map from BPM and sections
You

python scripts/beat_map.py --bpm 108 --offset 0.28 --max-clip 30 --sections "Intro:4=1,Verse 1:16=2,Pre 1:8=3,Chorus 1:16=5" --format md

Agent + skill
**Song** 108.0 BPM, 4 beats/bar, beat 0.556 s, bar 2.222 s, 44 bars, ends 1:38.058 (24.0 fps)

| Section | Energy | Bars | Start | End | Start TC | Cut every | Avg shot | Cuts |
| Intro | 1 | 1-4 | 0:00.280 | 0:09.169 | 00:00:00:07 | 8 beats | 4.44 s | 2 |
| Verse 1 | 2 | 5-20 | 0:09.169 | 0:44.724 | 00:00:09:04 | 8 beats | 4.44 s | 8 |
| Pre 1 | 3 | 21-28 | 0:44.724 | 1:02.502 | 00:00:44:17 | 4 beats | 2.22 s | 8 |
| Chorus 1 | 5 | 29-44 | 1:02.502 | 1:38.058 | 00:01:02:12 | 2 beats | 1.11 s | 32 |

| Segment | Bars | Start TC | Length | Generate | Audio excerpt |
| Chorus 1/1 | 29-36 | 00:01:02:12 | 17.78 s | 18 s | 1:02.502-1:20.280 |
| Chorus 1/2 | 37-44 | 00:01:20:07 | 17.78 s | 18 s | 1:20.280-1:38.058 |
Three concepts, one recommendation
You

Indie-pop song about keeping in touch across a city, 'Paper Planets'. I don't want to appear on camera. Give me concepts.

Agent + skill
A: Phone Booth (performance, literal). The singer alone in a rain-streaked glass booth; more paper planes press on the glass each chorus; final chorus they rise as stars. Faces 1. Feasibility 4/14. ~105 generations.

B: Paper Route (hybrid, parallel). A boy throws paper-plane letters from his rooftop; they cross the city to the booth where the singer catches one each chorus; the bridge is the night they stop; the final chorus fills the sky with gold. Faces 2. Feasibility 5/14. ~95 generations.

C: Fold (conceptual). The world is folded paper that unfolds into colour each chorus. No faces. Feasibility 3/14. ~56 generations.

Recommendation: B. It gives the song a story and keeps a performer at the centre of every chorus at a manageable risk. If the budget is under ~60 generations, C delivers the same payoff with no faces to keep consistent.
H3 lip-sync pass for a vertical teaser
You

140 BPM trap-pop, hook 0:54.86-1:22.29, AI performer, Hailuo app, 9:16 teaser.

Agent + skill
Half-time check: snare on beat 3, re-mapped at 70 BPM. Windows: Hook 1/1 bars 17-20 (13.71 s, excerpt 0:54.857-1:08.571), Hook 1/2 bars 21-24.

14s, 9:16. Night-time vertical music video; rooftop above a glowing city, cold blue moonlight with hot red aviation-light accents.
References: Image 1 = Juno's face; Image 2 = Juno's outfit; Image 3 = the rooftop with the satellite dish; Audio 1 = the song excerpt for this shot, performed by Juno.
Shot 1 (0-14s): close-up from a low angle, mostly static, one small push-in from 7s. Juno performs along to Audio 1, lips following the vocal, looking down the lens on each title line.
Sound: Audio 1 is the full soundtrack.

Edit: place at teaser 00:00:02:14, slip until the 'b' in the title line closes on the beat.

What's inside

ai-music-video-director/
├── agents/
│   └── openai.yaml
├── examples/
│   ├── example-01-indie-pop-narrative.md
│   └── example-02-trap-pop-vertical-teaser.md
├── references/
│   ├── 01-song-analysis.md
│   ├── 02-concepts-and-treatments.md
│   ├── 03-shot-design-by-section.md
│   ├── 04-ai-generation-playbook.md
│   └── 05-edit-on-beat.md
├── scripts/
│   └── beat_map.py
├── templates/
│   ├── 01-song-brief.md
│   ├── 02-beat-section-map.csv
│   ├── 03-concept-treatment.md
│   ├── 04-visual-world-bible.md
│   ├── 05-shot-plan.csv
│   ├── 06-prompt-pack.md
│   ├── 07-coverage-tracker.md
│   ├── 08-edit-on-beat-plan.md
│   └── 09-delivery-checklist.md
├── LICENSE.txt
├── README.md
└── SKILL.md

Install by unzipping into your agent's skills folder. Install guide →