ai-media-generation-expert

Solid

Expert guide for AI image generation (Flux, DALL-E, Stable Diffusion), video generation (Sora, Runway), voice synthesis (ElevenLabs TTS), and speech recognition (Whisper STT) integration / Panduan ahli integrasi AI generasi gambar, video, suara (TTS), dan pengenalan suara (STT).

AI & Automation 46 stars 9 forks Updated 3 days ago MIT

Install

View on GitHub

Quality Score: 87/100

Stars 20%
56
Recency 20%
100
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
50
License 10%
100
Description 5%
100

Skill Content

# AI Media Generation Expert (2026 Edition) [English](#english) | [Bahasa Indonesia](#bahasa-indonesia) --- <a name="english"></a> ## English ### Orchestration & Integration - **`ai-llm-integration-expert`**: Core LLM patterns and model selection. - **`file-upload-media-expert`**: Storage, CDN, and media pipeline for generated assets. - **`sse-websocket-streaming-expert`**: Real-time streaming for progressive image/audio generation. - **`async-queue-temporal-expert`**: Background job processing for long-running generation tasks. - **`zero-trust-secret-vault`**: Secure API key management for Replicate, OpenAI, ElevenLabs. ### Description Production-grade guide for integrating AI-powered media generation into web and mobile applications. Covers image generation (Flux 1.1 Pro, SDXL, DALL-E 3), video generation (Sora, Runway Gen-3), text-to-speech (ElevenLabs v3, OpenAI TTS), speech-to-text (Whisper large-v3, Deepgram Nova-3), and voice cloning. Includes async processing patterns, cost optimization, and content safety filtering. ### Trigger Conditions - Integrating AI image generation (Flux, DALL-E, Stable Diffusion, Midjourney API). - Building text-to-speech or speech-to-text features. - Implementing voice cloning or AI avatar generation. - Adding AI video generation or editing capabilities. - Building media processing pipelines with AI models. --- ### Core Architecture #### 1. Image Generation **Provider Selection Matrix:** | Provider | Model | Speed | Quality | Cost...

Details

Author
roedyrustam
Repository
roedyrustam/vibes-plug
Created
3 months ago
Last Updated
3 days ago
Language
Python
License
MIT

Integrates with

Similar Skills

Semantically similar based on skill content — not just same category

AI & Automation Solid

ai-llm-integration-expert

Expert guide for integrating Large Language Models (LLMs), Model Context Protocol (MCP), RAG architecture, vector databases, and AI agents / Panduan ahli untuk integrasi LLM, Model Context Protocol (MCP), arsitektur RAG, vector database, dan agen AI.

46 Updated 3 days ago
roedyrustam
AI & Automation Listed

super-claudiomedia-content-creation

Media content creation skill. Use when the user wants to create, generate, or produce any kind of video, audio, or image. This is the main skill for all media generation tasks. Trigger on video: "I want to make a video", "create a TikTok video", "generate a realistic video", "make a promo video", "animate my photo", "create a video ad", "Remotion", "Higgsfield", "Kling", "Seedance", "Weavy AI", "Hailuo". Trigger on audio: "read this article aloud", "create a voiceover", "text to speech", "TTS", "generate narration in Portuguese", "background music", "create a jingle", "ElevenLabs", "Francisca Neural", "Suno", "Udio", "audio summary". Trigger on image: "generate an image", "create a graphic", "make a diagram", "draw X", "generate a photo of Y", "make an infographic", "Midjourney", "DALL-E", "Flux", "Napkin.ai", "Nano Banana 2", "animate a static image". Also triggers for: marketing creatives, social media visuals, product photos, content creator tools.

4 Updated today
toolbox-playground
AI & Automation Listed

model-selector

Recommend the best AI model for image, video, and audio generation tasks. Triggers on "which model should I use", "recommend a model", "best model for", "what model for", "compare models", "fastest model for", "cheapest model for".

1 Updated 4 weeks ago
genfeedai