ai-voiceover

Solid

The AI narration / voiceover mini-skill (ElevenLabs-led). Use when someone wants an "AI voiceover," "narration," "text-to-speech for a video," "voice for my Reel/Short/explainer," "clone my voice," or to "dub a video into other languages." Picks the voice and model, writes for the ear, and directs the delivery; ElevenLabs generates the audio, the human mixes/reviews, WoopSocial schedules/publishes. Sits below the ai-video router, sibling to veo-3 and heygen. Consented voices only; disclose AI voice in ads/political.

AI & Automation 76 stars 16 forks Updated 1 weeks ago MIT

Install

View on GitHub

Quality Score: 85/100

Stars 20%
63
Recency 20%
90
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
50
License 10%
100
Description 5%
100

Skill Content

# ai-voiceover The **audio** producer of the video cluster — the counterpart to **veo-3** (scenes) and **heygen** (avatars) under the **ai-video** router. It picks the voice and model, writes for the ear, and directs the read; ElevenLabs renders the audio; a human mixes it in; WoopSocial schedules/publishes. ## The POV: 80% script + direction, 20% tool Most AI VO sounds robotic because people feed it **eye-written copy** and accept the **default read**. A great voiceover is mostly the script-for-the-ear and the direction. Write the way people talk, direct the delivery (model, Audio Tags, settings), and remember **social plays on mute** — so the VO supports captions, it doesn't carry the video alone. ## Read these first 1. **brand-profile** — audience, platform, non-negotiables. 2. **voice-builder** — the brand's **written** voice. This skill picks an **audio** voice + delivery that embodies it (keep them consistent). ## The framework: VOICE (Depth: `references/the-voice-framework.md`.) - **V — Voice match:** library / Voice Design / consented clone; fit brand + platform. - **O — Own the script for the ear:** spoken cadence, contractions, short sentences; read it aloud. - **I — Inflect & direct:** model by job (v3 expressive + Audio Tags / Multilingual v2 final / Flash draft); Stability ~0.3–0.5 expressive vs ~0.7–1.0 consistent; Similarity ~0.75–0.85; pronunciation. - **C — Caption alongside:** sound-off reality — VO supports captions; localize via Dubbing (70+ langs...

Details

Author
social-media-skills
Repository
social-media-skills/skills
Created
1 months ago
Last Updated
1 weeks ago
Language
Shell
License
MIT

Similar Skills

Semantically similar based on skill content — not just same category

AI & Automation Solid

heygen

The AI avatar / talking-head mini-skill (HeyGen). Use when someone wants an "AI avatar video," "talking-head video," "digital twin / clone of myself on camera," "faceless presenter video," "spokesperson video," or to "translate/localize a video into many languages with lip-sync." The creator/social avatar lane; enterprise L&D/training/SCORM avatar work routes to synthesia, and a real human on camera (founder/trust content) routes to talking-head-and-piece-to-camera. Scripts and sets up the video; HeyGen renders; the human reviews/edits; WoopSocial schedules/ publishes. Sits below the ai-video router, sibling to veo-3. Consented avatars only; AI disclosure mandatory.

76 Updated 1 weeks ago
social-media-skills
AI & Automation Solid

ai-video

The model-agnostic AI-video router and brief — the counterpart to image-prompt. Use when someone asks "which AI video tool should I use," "make an AI video," "generate B-roll / a talking-head / a voiceover," "turn this long video into Shorts," or needs a video brief. Routes the job to the right tool by fit and writes a portable brief; the tool generates, the human assembles, WoopSocial schedules/publishes. Sits above the tool skills: veo-3, kling, luma (generative scenes), heygen, synthesia (avatars), ai-voiceover, captions-and-clipping. A real human on camera routes to talking-head-and-piece-to-camera. Never routes to discontinued tools.

76 Updated 1 weeks ago
social-media-skills
AI & Automation Listed

voiceover-production

The authority on AI-video audio + voice DURING GENERATION. Native audio ON by default (generate_audio=true); multi-clip voice consistency carried via Seedance reference-to-video (@Audio1), NEVER by muting clips + layering a TTS voiceover. Replacing or changing the voice on an EXISTING video is a separate, user-triggered flow — see the `voice-replacement` skill. Loaded BY orchestration skills — not a standalone entry point.

2 Updated 5 days ago
Nagellabs