← ClaudeAtlas

audio_generationlisted

Generate audio in two ways: text-to-speech, or custom sound effects. ### Speech (text-to-speech): - High-quality text-to-speech conversion using a selected voice - Several pre-built Mandarin voices with different characteristics - Use clear, well-formatted text (punctuation helps natural speech) ### Sound effects: - AI-powered sound effect generation from an English description - Customizable duration (0.5 to 22 seconds) - Covers ambient, action, musical, foley, emotional, and abstract sounds - The description MUST be in English ### Usage Guidelines: - For speech: pick a voice ID and provide the text to read - For sound effects: give a detailed English description and a duration - Specify an output path with a .mp3 extension - Output is saved locally as mp3
serejaris/kimi-skills · ★ 4 · Code & Development · score 75
Install: claude install-skill serejaris/kimi-skills
# Audio Generation Use this skill to generate audio. There are two distinct flows — pick the one that matches the user's intent: - **Generate speech** (text-to-speech): the user wants spoken audio of some text. Use the `speech` flow. - **Generate sound effects**: the user wants a sound effect / ambience / SFX described in words. Use the `sound-effects` flow. ## Setup Before the first use, ensure the agent-gw Python SDK (version 0.2.6 or newer) is installed. This checks the current environment and installs or upgrades it only when needed: ```bash python3 scripts/audio_generation_tool.py ensure-deps ``` The SDK needs an API key from `api_key=...`, `KIMI_API_KEY`, or `~/.kimi/agent-gw.json`. ## Choosing the flow 1. If the user wants their **text read aloud / a voiceover / TTS** → **speech**. 2. If the user wants a **sound effect, ambience, music bed, or SFX described in words** → **sound-effects**. Then build the parameters for that flow, run the matching command, and on success surface the saved mp3 to the user. On failure, explain the error from the script; do not invent audio or a local path. ## Flow A — Generate speech (text-to-speech) Parameters: - `text` (required): the text to convert to speech. - `voice_id` (required): one of the supported voices below. Default is `05Cdh2gw2NMzDvykn1nm`. - `output` (required): local output path ending in `.mp3`. Supported voice IDs: - `05Cdh2gw2NMzDvykn1nm`: calm middle-aged Mandarin male (default) - `Q63G7WZ5riIGb