lip-sync-spokesperson-videolisted
Install: claude install-skill apimageorg/apimage-skills
# Lip-Sync Spokesperson Video
`generate_lip_sync` turns a face image plus audio into a talking avatar. It's the most constrained tool in APImage and the most expensive per second, so the craft is mostly in what you decide *before* generating.
## The constraints, up front
| Constraint | Value | Consequence |
|---|---|---|
| Model | seedance 2 0 only | No draft path. No cheap iteration |
| Cost | **3 credits per second** | 30s = 90 credits |
| Max duration | **30 seconds** | Longer scripts must be split |
| Input | Face image + audio | Both have to be right first |
**90 credits for a 30-second clip** is a third of a Starter plan's monthly allowance. That reframes the work: you're not iterating your way to a good clip, you're getting the inputs right and generating once.
## Script to the second, first
At 3 credits/second, every unnecessary word is money. Write the script, read it aloud with a timer, and cut.
| Duration | Realistic word count | Use for |
|---|---|---|
| 8s | ~20 words | A single claim or hook |
| 12s | ~30 words | Hook plus one supporting point |
| 15s | ~38 words | Problem, then solution |
| 20s | ~50 words | Three beats, tightly |
| 30s | ~75 words | A full short ad. The cap |
Cut before generating: greetings, "in this video", self-introduction, brand name repetition, and any sentence that doesn't advance the point. A 30-second script that says what a 15-second one could costs double for nothing.
See talking head scripting for structure.
## The face i