Nagellabs
OrganizationAI video studio — a desktop editor that connects CLI coding agents (Claude Code, Codex) over ACP and lets them compose, edit and export video.
Categories
Indexed Skills (37)
text-to-speech
Convert text to speech using ElevenLabs voice AI. Use when generating audio from text, creating voiceovers, building voice apps, or synthesizing speech in 70+ languages.
shadcn
Manages shadcn components and projects — adding, searching, fixing, debugging, styling, and composing UI. Provides project context, component docs, and usage examples. Applies when working with shadcn/ui, component registries, presets, --preset codes, or any project with a components.json file. Also triggers for "shadcn init", "create an app with --preset", or "switch to --preset".
ai-asset-generation
Produce ONE AI asset (image / video / audio / 3D) via a generation MCP (fal.ai, etc.) — the call + save layer: discover provider, pick model, read schema, disclose cost, build the prompt, run + poll, and import the file with provenance. Also owns the two universal video invariants: no in-video text, native audio on. The realism-image craft and the physical-action/FLF craft live in their own skills (realistic-image-generation, physical-action-video); the keyframe→clip WORKFLOW belongs to the Storyboard.
animated-text-overlays
Add or FIX an ANIMATED text overlay — text that reveals or moves over time (typewriter / letter-by-letter, word-by-word fade, slide-up, pop, gradient shine, lower-third, kinetic hook / title / caption). Load this BEFORE hand-writing any code overlay (add_overlay kind code) that animates text — it owns the element-local timing contract that prevents the "only the first few letters show" reveal bug. Triggers — "make the caption type out", "animate the caption", "letter-by-letter typewriter", "kinetic hook", "the animated caption is broken / cut off". NOT for subtitles synced to speech (use speech-captions); NOT for plain static text (use add_overlay kind text).
animating-overlays
Animate an overlay's MOTION — move / slide / zoom / spin / fade an overlay's position, scale, rotation, or opacity from one value to another over time. Load this whenever the user asks to animate a transform or opacity transition on a text/image/video/code/three overlay. It owns the keyframes-first rule — use `libi.add_keyframe` (visible, user-editable timeline diamonds), NEVER bake the motion into a code overlay's draw function. Triggers — "make the title slide up", "fade this in", "zoom the logo in", "spin it as it enters", "animate it moving across". NOT for text-reveal typewriter/word-by-word (use animated-text-overlays) and NOT for looping/parametric motion like bob/shake/pulse (use effects).
audio-analysis
Transcribe a video or audio file. Default is local Whisper (faster-whisper, free, on-device). ElevenLabs is opt-in for speaker diarization / audio events or on explicit request. Triggers on "transcribe", "captions", "speech-to-text", or any request to extract spoken text from a media file.
generic-video
Create a genre-agnostic AI video — either recreating a source (handed over by mimic-video) or from a fresh brief. Owns the creative intake (fidelity, theme/style, pacing, duration, stitch-vs-fully-AI, model, voice) and the build flow. References ugc-craft for prompt craft, ai-video-models for per-engine rules, ai-asset-generation for mechanics.
guiding-manual-edits
Point the user at the exact overlay inspector control instead of editing it for them — use when the user asks how to change something themselves, rejects your edit and wants to hand-tweak it, OR when your own automated refinements keep missing and the user would get there faster by hand. Drives `libi.highlight_property` (flash a field + reveal its tab) and `libi.set_complexity_mode` (switch one overlay's tab). Load when guiding a manual edit, NOT when the user wants you to make the change.
mimic-video-captions
Reproduce / mimic the ON-SCREEN CAPTIONS of an existing video — lyric typography, kinetic text, road/perspective captions, glowing animated subtitles. Owns the caption-mimic flow — the authoritative words+timing come from the Whisper transcript, the visual treatment + motion from a caption-focused paid analysis, routing each caption to the right renderer (3D/perspective → three-overlays, flat kinetic 2D → animated-text-overlays, plain subtitle → speech-captions), then RENDER and self-correct via the verify loop. Load this when the user wants the source's captions reproduced faithfully — NOT for recreating the video's content (that is mimic-video).
mimic-video
Top-level dispatcher for recreating / mimicking an existing video. Analyzes the source (via video-analysis), classifies it, and routes to the right creation skill — ugc-product-video for product ads, music-video-creation for music videos, generic-video otherwise. Generates nothing itself; it picks the path and hands off.
music-creation
Interview-style music generation. Asks the user about genre, vocals, lyrics, length, optional reference track. Dispatches via ai-asset-generation skill which routes to local-music (default), or elevenlabs / fal-ai on explicit request.
music-video-creation
Build a music video — generate a track, attach it under the visuals, optionally render synced lyrics on screen, and iterate cleanly when the user wants a different song. Encodes the composition rules that prevent stale text, double-rendered lyrics, off-by-200ms caption sync, and "I don't see anything" surprises.
onboarding-libi-explainer-short
Run ONLY during first-run onboarding when the user clicks "show me how it works" (or asks for the libi intro/demo). Builds libi's own 52-second explainer film into a real piece with a single call to libi.build_onboarding_piece (~15 MB download, no generation), reveals it, then tells the user honestly that the film is pre-made but was itself built in libi and is fully editable — and asks what they want to make.
physical-action-video
Make hard physical-action / manipulation video beats survive generation — applying, peeling, pressing, pouring, gripping-and-releasing, twisting, writing, cutting. Owns the FLF-first (first-last-frame) approach, prompt decomposition (3–5 one-verb sub-steps, object anchoring, affordance pre-conditions, frame-relative direction), the model-escalation ladder, the editorial before/after fallback, and the levers that keep isolated clips looking like ONE video. Loaded BY ugc-product-video / generic-video / production-routes when a beat manipulates an object. NOT a standalone entry point.
realistic-image-generation
Generate realistic AI images — especially photoreal people / creator portraits and video KEYFRAMES (start/end frames). Owns the realism model picker (gpt-image-2 default, never let recommend_model downgrade it), the anti-'AI-look' banned tokens + Flux negative prompts, the UGC selfie + demographic templates, the prompt-plausibility (anatomy) pre-check, and the post-generation image-validation rubric. Loaded BY the Storyboard keyframe step / ugc-product-video / generic-video — it produces ONE good image; the board sequences keyframe→clip. NOT a standalone entry point.
removing-and-replacing-backgrounds
Remove a video's or photo's background into a reusable alpha "cutout" asset (subject isolated, background transparent), then compose it over any new background or transplant it into another video — local free MatAnyone matting for video, paid fal fallback (bria video / birefnet photos). Triggers — "remove the background", "put her on a beach", "green screen this", "cut out the product", "transparent background", "place him in the other video".
speech-captions
Add subtitles/captions synced to spoken audio — one call to libi.generate_captions builds a styled, time-synced caption track from the file's word-level timings (local Whisper STT), in a chosen STYLE (cumulative / word-by-word / karaoke / letter-by-letter). State the style in the result. Use for "add captions", "subtitle this", "sync the caption to her speech". For decorative (non-speech) animated text use animated-text-overlays.
stitching-multi-clip
Build a multi-clip timeline from N independently-generated short video clips as SEPARATE per-beat video overlays. The editor's playback engine smooths the clip-boundary seams; concatenation into one file is a FINAL-EXPORT concern, not a way to fix preview jumps.
three-overlays
Add a real 3D / WebGL (three.js) overlay — perspective captions (text laid on a ground plane, floating billboard text that moves with the camera) and simple animated 3D objects, composited over the layers beneath. Load this BEFORE calling `libi.add_overlay` with kind three. Use when a caption needs DEPTH/PERSPECTIVE that flat Canvas2D code overlays cannot express (text mapped onto a road/floor, 3D-positioned lyrics that play with the footage), or for a simple rotating/animated 3D object. NOT for flat animated text (typewriter, word reveal, kinetic 2D caption → animated-text-overlays); NOT for speech subtitles (→ speech-captions); NOT for character rigs, physics, or imported heavy 3D models.
ugc-craft
Internal craft reference for UGC video generation: the 9-layer UGC formula, clip-duration methodology, pacing / natural-motion / skin-realism cue banks, character-consistency phrasing, and negative-prompt + forbidden-word lists. Loaded BY `ugc-product-video` and `stitching-multi-clip` — it is NOT a standalone entry point. If the user wants a UGC / product / demo / social video, start from `ugc-product-video` (or `stitching-multi-clip` for a source+AI stitch); do not begin a build from here.
ugc-product-video
Walks the user through creating a UGC-style AI product video — brief, ad format + production route, character + product references, a real scripted ad, per-clip generation on Seedance 2.0 (default), validation, audio, captions, end card. Thin router over the prompt files in `prompts/`. Use when the user wants to create a product ad, demo video, or social UGC. Default to ONE full-length multi-beat clip (15s with in-prompt jump-cut beats), not one short clip per beat.
using-asset-folders
Use when generating or uploading MULTIPLE related assets — an extend chain, several concept/style variations of one image or video, or a batch of takes — and you want them grouped instead of flooding the piece with loose files. Teaches the one-file-one-asset model plus asset folders. There are no "options" or a "default file" anymore.
using-character-library
Catalog and reuse recurring characters and items across pieces — when to suggest saving, how to disambiguate, and how to link assets.
using-effects
Apply tasteful in/out/loop animation effects to any layer — captions, stickers, logos, backgrounds, audio. Fade a caption in, pop a logo, gently pulse an emphasis, Ken-Burns a photo, fade audio. "make the title bounce in", "add a subtle float", "fade this out".
using-object-tracking
Pin any visual element (emoji, text, image, video, JS draw fn, blur/pixelate/mask) to a moving subject across a video — face or object tracking, censor a face, follow a product. "smiley on her face", "blur the license plate", "logo on his shirt", "name tag follows him", "pixelate the plate".
using-storyboard
Use when building or planning a multi-scene video via the Storyboard (the editor's Storyboard tab) — including mimic/recreate flows and any time you'd otherwise write a piece script. Teaches the free schematic tier, the per-endpoint generation spec + model-schema cache workflow (cache-gate → populate → set spec → validate → fix), keyframing/reference/audio params, live continuity references between scenes, and versioned takes.
video-analysis
Analyze a video's visual content. Default to the free agent-driven flow (extract keyframes, describe each, produce VideoSummary) — this covers most tasks. Only mention the paid Gemini-via-fal.ai script flow when the task genuinely needs audio/music understanding or the user explicitly asks for it.
video-planning
Think like a senior video editor BEFORE generating — reverse-engineer a target video into its building blocks and build algorithm, turn it into an explicit reviewable plan, then direct the build through the Storyboard, loading the right craft specialist per block. Resolves three entry modes (extract a plan from a demo video · reuse a saved recipe skill · create a fresh plan) and offers to capture a liked plan as a reusable skill. Loaded BY mimic-video and every creation skill (generic-video, ugc-product-video, music-video-creation) — the planning/director layer ABOVE the Storyboard. Not a standalone entry point.
ai-video-models
Per-engine prompting guides for AI video models (Seedance 2.0, Veo 3.1, Kling). Genre-neutral — how to write a good prompt for each engine (reference-image token, prompt order, length, motion language, FLF, duration caps, style whitelist). Loaded BY creation skills (ugc-product-video, generic-video) once a model is chosen — NOT a standalone entry point.
skill-eval
Run libi's agent-driven skill-eval scenarios. Use after editing a bundled skill, MCP wiring, or agent instructions to verify the inner libi agent still behaves correctly (e.g. picks gpt-image-2, keeps native audio). Heavy + token-costly — run manually, only the scenarios a change warrants.
feature-testing
Verify libi features end-to-end via the actual libi agent in test mode (LIBI_TEST_MODE=1 npx libi — fal-ai is swapped for a fake that mirrors the real tool surface with placeholder outputs) before declaring agent-facing work complete. Use after changes to MCP tools, skills, agent instructions, or any libi.* route.
windows-qa
Build, install and verify a libi desktop build on the Windows QA VM. Use before shipping anything Windows-facing, or whenever a build must be put in front of a real Windows user. Encodes the traps that have each cost hours — npm ci vs npm install, CI=1, processes dying with the SSH session, what "installed" actually means, and the big one — the VM is an elevated administrator, so an elevated pass proves nothing about a real user.
installing-mcps
Use when the user asks you to install, set up, configure, repair, or fix an MCP server. Drives the get_install_plan → follow plan → update_dep_status → verify_install flow with appropriate progress updates.
using-piece-duplication
Use when the user wants multiple versions of a video, several videos about one subject, or a safe place to try a fundamentally different creative direction.
using-snapshot-draft
Use whenever you mutate a piece's composition or when the user asks to "go back," "undo," "save," "discard," or talks about prior versions. Teaches the snapshot/draft mental model: every edit lands in the draft; commit promotes to snapshot; discard reverts; restore_snapshot recovers a prior committed state.
voice-replacement
Use when the user asks to CHANGE, REPLACE, RE-VOICE, or DUB the voice on one or more EXISTING videos/scenes in a piece — a deliberate, user-triggered step AFTER the video exists ("change the voice", "give it a different voiceover", "redo the voice", "dub this", "clone my voice over it", "new narrator"). This is NOT initial generation (that keeps native audio via voiceover-production). It transcribes the target scenes, asks whether to clone the existing voice or pick a new one, lip-syncs the sections where a character speaks on camera, and mutes + re-voices the rest. A standalone entry point with its own trigger.
voiceover-production
The authority on AI-video audio + voice DURING GENERATION. Native audio ON by default (generate_audio=true); multi-clip voice consistency carried via Seedance reference-to-video (@Audio1), NEVER by muting clips + layering a TTS voiceover. Replacing or changing the voice on an EXISTING video is a separate, user-triggered flow — see the `voice-replacement` skill. Loaded BY orchestration skills — not a standalone entry point.
Bio shown is the top-scored skill's repo description as a fallback — real GitHub bios land in a future update.