generating-audio

Solid

ALWAYS read this skill before generating spoken audio or calling audio_generate — a voiceover, narration, an ad read, a character line, or any script read aloud. Turns a script into speech — picks the model and voice, prepares the text for reading, and splits a long script into clips. Use whenever the user asks for text-to-speech, a voiceover, narration, or to have something read or spoken aloud.

AI & Automation 24 stars 3 forks Updated 4 days ago Apache-2.0

Install

View on GitHub

Quality Score: 81/100

Stars 20%
47
Recency 20%
100
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
50
License 10%
100
Description 5%
100

Skill Content

# Audio Generation Turn a script into spoken audio via the `audio_generate` tool. Two decisions drive quality: **which voice** reads it, and **how the script is written for the ear**. **Scope.** Speech — a voice reading written words, delivered as its own audio file. Nothing here makes non-speech sound, alters audio that already exists, or combines two tracks into one. Lip-synced dialogue spoken by a character *inside* a video clip belongs to `generating-videos`. ## Workflow ### Step 1: Settle the script Generate only from the exact words that will be spoken. - **Supplied** → use them verbatim. - **Enough to write them** — the product, audience, platform, and length are known → draft the script and show it before generating. - **Not enough** → ask. Never invent a tagline, product claim, or brand name to fill the gap. **Write to a duration.** Speech runs about two to three words a second, so a fifteen-second read is thirty to forty-five words. Set the word count before writing, and trim words rather than speeding up the delivery. ### Step 2: Pick the model **`eleven-v3` is the default** — the most expressive read, and right for anything heard as a performance. Reach for another only on a clear signal: | Reach for another model when the script… | Model | | --- | --- | | Is a long, even read — an explainer, documentary narration, an audiobook chapter — where the voice must not drift | `eleven-multilingual-v2` | | Is high-volume, a throwaway draft, or cost-sensitive,...

Details

Author
SupercmoHQ
Repository
SupercmoHQ/superCMO-skills
Created
3 weeks ago
Last Updated
4 days ago
Language
Python
License
Apache-2.0

Bundled in these plugins

Similar Skills

Semantically similar based on skill content — not just same category

AI & Automation Solid

generating-audio

ALWAYS read this skill before generating spoken audio or calling audio_generate — a voiceover, narration, an ad read, a character line, or any script read aloud. Turns a script into speech — picks the model and voice, prepares the text for reading, and splits a long script into clips. Use whenever the user asks for text-to-speech, a voiceover, narration, or to have something read or spoken aloud.

413 Updated today
aiskillstore
AI & Automation Solid

generating-videos

ALWAYS read this skill before generating or animating any video, or calling video_generate — text-to-video, image-to-video, a start→end transition, or a reference / motion / audio-guided clip. Generates a video from a brief — one clip, or several clips joined into a single file at any length. Use when the user wants to make or generate a video, animate a photo, bring an image to life, produce b-roll, film a described scene, or move from one held frame to another. This skill should also be used when the video takes its motion or style from an existing video, or has to run to an existing audio track.

24 Updated 4 days ago
SupercmoHQ
AI & Automation Solid

ai-voiceover

The AI narration / voiceover mini-skill (ElevenLabs-led). Use when someone wants an "AI voiceover," "narration," "text-to-speech for a video," "voice for my Reel/Short/explainer," "clone my voice," or to "dub a video into other languages." Picks the voice and model, writes for the ear, and directs the delivery; ElevenLabs generates the audio, the human mixes/reviews, WoopSocial schedules/publishes. Sits below the ai-video router, sibling to veo-3 and heygen. Consented voices only; disclose AI voice in ads/political.

32 Updated today
social-media-skills