audio_generationlisted
Install: claude install-skill serejaris/kimi-skills
# Audio Generation
Use this skill to generate audio. There are two distinct flows — pick the one
that matches the user's intent:
- **Generate speech** (text-to-speech): the user wants spoken audio of some
text. Use the `speech` flow.
- **Generate sound effects**: the user wants a sound effect / ambience / SFX
described in words. Use the `sound-effects` flow.
## Setup
Before the first use, ensure the agent-gw Python SDK (version 0.2.6 or newer) is installed. This checks the current environment and installs or upgrades it only when needed:
```bash
python3 scripts/audio_generation_tool.py ensure-deps
```
The SDK needs an API key from `api_key=...`, `KIMI_API_KEY`, or
`~/.kimi/agent-gw.json`.
## Choosing the flow
1. If the user wants their **text read aloud / a voiceover / TTS** → **speech**.
2. If the user wants a **sound effect, ambience, music bed, or SFX described in
words** → **sound-effects**.
Then build the parameters for that flow, run the matching command, and on
success surface the saved mp3 to the user. On failure, explain the error from
the script; do not invent audio or a local path.
## Flow A — Generate speech (text-to-speech)
Parameters:
- `text` (required): the text to convert to speech.
- `voice_id` (required): one of the supported voices below. Default is
`05Cdh2gw2NMzDvykn1nm`.
- `output` (required): local output path ending in `.mp3`.
Supported voice IDs:
- `05Cdh2gw2NMzDvykn1nm`: calm middle-aged Mandarin male (default)
- `Q63G7WZ5riIGb