Video Generation
bytedance/deer-flow
Generates short videos from a structured JSON prompt, optionally guided by a reference image used as the first or last frame.
Production guidance for Kling AI Open Platform Advanced Lip-Sync.
$ npx skills add calesthio/generative-media-skills --skill kling-advanced-lip-sync -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install calesthio/generative-media-skills kling-advanced-lip-sync --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/providers/lip-sync/kling-advanced-lip-sync .claude/skills/kling-advanced-lip-sync && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "kling-advanced-lip-sync" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/lip-sync/kling-advanced-lip-sync into .claude/skills/kling-advanced-lip-sync/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kling-advanced-lip-sync", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/lip-sync/kling-advanced-lip-syncType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add calesthio/generative-media-skills --skill kling-advanced-lip-sync -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install calesthio/generative-media-skills kling-advanced-lip-sync --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/providers/lip-sync/kling-advanced-lip-sync .agents/skills/kling-advanced-lip-sync && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "kling-advanced-lip-sync" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/lip-sync/kling-advanced-lip-sync into .agents/skills/kling-advanced-lip-sync/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kling-advanced-lip-sync", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add calesthio/generative-media-skills --skill kling-advanced-lip-sync -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install calesthio/generative-media-skills kling-advanced-lip-sync --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/providers/lip-sync/kling-advanced-lip-sync .cursor/skills/kling-advanced-lip-sync && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "kling-advanced-lip-sync" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/lip-sync/kling-advanced-lip-sync into .cursor/skills/kling-advanced-lip-sync/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kling-advanced-lip-sync", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/calesthio/generative-media-skills.git --path skills/providers/lip-sync/kling-advanced-lip-sync--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add calesthio/generative-media-skills --skill kling-advanced-lip-sync -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install calesthio/generative-media-skills kling-advanced-lip-sync --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/providers/lip-sync/kling-advanced-lip-sync .gemini/skills/kling-advanced-lip-sync && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "kling-advanced-lip-sync" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/lip-sync/kling-advanced-lip-sync into .gemini/skills/kling-advanced-lip-sync/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kling-advanced-lip-sync", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install calesthio/generative-media-skills kling-advanced-lip-syncInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add calesthio/generative-media-skills --skill kling-advanced-lip-sync -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/providers/lip-sync/kling-advanced-lip-sync .github/skills/kling-advanced-lip-sync && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "kling-advanced-lip-sync" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/lip-sync/kling-advanced-lip-sync into .github/skills/kling-advanced-lip-sync/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kling-advanced-lip-sync", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add calesthio/generative-media-skills --skill kling-advanced-lip-sync -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install calesthio/generative-media-skills kling-advanced-lip-sync --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/providers/lip-sync/kling-advanced-lip-sync .opencode/skills/kling-advanced-lip-sync && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "kling-advanced-lip-sync" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/lip-sync/kling-advanced-lip-sync into .opencode/skills/kling-advanced-lip-sync/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kling-advanced-lip-sync", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
kling-advanced-lip-syncProduction guidance for Kling AI Open Platform Advanced Lip-Sync.
Kling Advanced Lip Sync is an agent skill from calesthio/generative-media-skills. Production guidance for Kling AI Open Platform Advanced Lip-Sync. Use for identifying/selecting one face in an existing video, assigning URL/Base64 or Kling TTS audio, cropping and inserting audio in milliseconds, mixing original sound, creating and monitoring tasks, downloading expiring results, consent/privacy review, sync QA, and repair. Do not use for still-image avatar generation or ordinary Kling video generation.
Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `EVAL.md`).
It sits in Media & Creative, covering AI video generation. The repository describes itself as: Research-backed agent skills and tools for premium image, video, audio, voice, and generative media production across AI coding assistants. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 8c85352. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are http and json).
From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
api-singapore.klingai.comAlso links to:
kling.aiklingai.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Kling Advanced Lip Sync loads about 2.9k tokens when it runs. Until then it costs about 112 tokens; SKILL.md has 1,306 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from calesthio/generative-media-skills at commit 8c85352, republished under its MIT licence (© calesthio). 1,306 words, ~2,859 tokens.
.claude/skills/kling-advanced-lip-sync/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Use this skill for Kling AI Open Platform's documented Advanced Lip-Sync workflow on an existing video. It is distinct from Avatar, which generates a new talking performance from a still image and audio.
Facts were verified 2026-07-12. Endpoint visibility, entitlement, pricing, callback behavior, retention, limits, and policy can change. Re-check official documentation, authenticated console, account package, and applicable agreement before paid use.
Base URL in current examples:
https://api-singapore.klingai.comOperations:
POST /v1/videos/identify-facePOST /v1/videos/advanced-lip-syncGET /v1/videos/advanced-lip-sync/{id}GET /v1/videos/advanced-lip-syncAlthough the navigation may call a tab Face Recognition, the operation is Identify Face. Describe it as face identification/detection for selecting a visible face; do not claim general identity matching.
Official docs use JWT Bearer authentication derived from Kling Open Platform access/secret credentials:
Authorization: Bearer ${KLING_JWT}
Content-Type: application/jsonFollow the current authentication page for token construction and lifetime. Generate server-side, use least privilege, and never place keys/secrets/JWTs in browser code, logs, callbacks, or fixtures.
Identify Face accepts exactly one:
video_id: Kling-generated video from the documented recent window;video_url: accessible external video.They are mutually exclusive.
Current external-video constraints include .mp4 or .mov, 2-60 seconds, at most 100 MB, 720p or 1080p, and width/height each 512-2160 pixels. Current docs also limit Kling video_id recency to 30 days. Kling performs content validation.
Preflight codec/container, duration, dimensions, file size, URL reachability, audio presence, face visibility, cuts, occlusion, and consent. Do not upload a private URL that Kling cannot access or a mutable URL that can change between operations.
Example:
POST /v1/videos/identify-face
{"video_url":"https://example.com/source.mp4"}The response includes session_id and face_data[]; each face can include face_id, face_image, start_time, and end_time in milliseconds.
Present detected face crops/intervals to the user or authorized operator. If several faces exist, require an explicit face_id unless a documented, approved selection rule exists. Record selected face, reason, and interval.
face_choose is an array, but current docs explicitly say only one-person lip-sync is supported. Do not claim simultaneous multi-speaker replacement. Re-identify after any source-video edit; do not reuse a stale session against changed media.
Kling does not publicly establish session expiry/reuse semantics in the verified docs. Keep identify/create close together and handle invalid sessions as a new identification job.
For the selected face, provide exactly one:
audio_id: recent Kling TTS audio;sound_file: accessible URL or Base64 audio.They are mutually exclusive.
Current external-audio constraints include .mp3, .wav, .m4a, or .aac, 2-60 seconds, and at most 5 MB. Current docs limit audio_id recency to 30 days.
Timing fields are milliseconds:
sound_start_time: crop start in source audio;sound_end_time: crop end;sound_insert_time: insertion point in source video.Validate:
$$ cropDuration = soundEnd - soundStart \ge 2000\text{ ms} $$
The crop must fit inside the source audio, the inserted range must fit inside the video, and the inserted range must overlap the selected face's detected appearance interval by at least two seconds.
Use explicit millisecond arithmetic. Never pass seconds by accident.
Mix controls:
sound_volume: current documented range 0-2, default 1;original_audio_volume: range 0-2, default 1; no effect for silent source.Preserve ambience when it matters, but ensure source dialogue does not create double speech. Review the final mix separately from lip motion.
Illustrative request:
{
"session_id": "<identify-face-session>",
"face_choose": [
{
"face_id": "<approved-face-id>",
"sound_file": "https://example.com/dialogue.wav",
"sound_start_time": 0,
"sound_end_time": 3000,
"sound_insert_time": 1000,
"sound_volume": 1,
"original_audio_volume": 0.25
}
],
"external_task_id": "scene-014-take-03",
"callback_url": "https://example.com/hooks/kling",
"watermark_info": {"enabled": false}
}This assumes a three-second crop, sufficient video duration, and at least two seconds of overlap with the selected face interval.
external_task_id uniqueness is caller-owned. Store it with the system task_id, source/audio hashes, session/face IDs, timings, volumes, attempt, consent record, and cost preflight.
Task statuses:
submittedprocessingsucceedfailedUse the single-task endpoint as source of truth. Callbacks can reduce polling latency, but production handling should be idempotent and accept duplicates/out-of-order events. Unless the current callback page documents a signature mechanism, do not invent one. Retrieve the task before trusting or publishing callback data.
List queries use current pagination ranges; do not scan the list when a task ID is known.
Retry transient rate/concurrency/server errors with bounded jitter. Do not retry auth, permission, parameter, policy, or invalid-media failures unchanged. Preserve provider code/message/request ID.
The public pricing page did not clearly establish a dedicated Advanced Lip-Sync list price during verification. Do not assume Avatar pricing applies. Check authenticated console/package entitlement and create a dated budget approval.
Successful task responses expose final deduction fields; reconcile actual units/balance against the estimate.
Generated media is documented as deleted after 30 days. Download promptly into controlled storage, checksum, probe, and preserve provenance. Do not use provider URLs as permanent delivery storage.
Endpoint entitlement and concurrency are account/package dependent. Errors such as rate, concurrency, package exhaustion, or permission denial should be surfaced precisely.
This workflow processes recognizable face video and voice/audio. Require:
Kling's technical face_id does not prove identity or consent. Review current API privacy/terms for hosting, storage, transfer, processing, model improvement, and enterprise exceptions. Do not make unqualified training, residency, or deletion claims.
Kling does not publish a formal sync-quality metric or repair API. The following are production heuristics:
Do not claim a rerun will necessarily improve output. Change one variable, record it, and compare.
This is a complete example, not a mandatory formula.
Intent: replace a three-second line in an authorized interview while keeping room ambience.
Preflight a 12-second 1080p MP4 under 100 MB, confirm one visible consented speaker, and provide a 3.2-second WAV under 5 MB. Identify faces, select the approved face whose interval covers seconds 2-10, then crop audio 100-3100 ms and insert at 3500 ms. This yields a 3000 ms insertion within video and at least two seconds inside the face interval. Set source dialogue low enough to avoid doubling while retaining ambience.
Monitor by task ID, download/checksum immediately, then review consonants, start/end motion, identity, occlusion, and ambience. If sync starts late, adjust insertion/crop before changing volume or regenerating audio.
This is a complete example, not a mandatory formula.
Intent: replace one host line in a two-person panel clip.
Identify Face returns two faces with distinct crops and intervals. Present both; the authorized editor selects the host. Because only one-person lip-sync is documented, create one task for that face and do not promise simultaneous guest replacement. Choose a segment where the host remains visible and the guest does not occlude the mouth. Keep original audio if only ambience remains; otherwise use a prepared mix.
If the requirement becomes replacing both speakers, stop and choose another supported workflow rather than chaining undocumented multi-face tasks.
Official sources verified 2026-07-12:
© calesthio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in skills/providers/lip-sync/kling-advanced-lip-sync of calesthio/generative-media-skills.
Open the folder on GitHubat commit 8c85352
Kling Advanced Lip Sync next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Kling Advanced Lip Sync this skillcalesthio/generative-media-skills | 197 | — | ~2.9k | Automated safety check: Pass | MIT | |
| Video Generationbytedance/deer-flow | 84k | 3 repos | ~1.4k | Automated safety check: Pass | MIT | |
| Video Cover Imageitwanger/toBeBetterJavaer | 18k | — | ~3.3k | Automated safety check: Pass | None | |
| Seedancesongguoxs/seedance-prompt-skill | 2.9k | 1 repos | ~2.5k | Automated safety check: Pass | None | |
| HyperFrames Video Entry Pointheygen-com/hyperframes | 60k | 3 repos | ~5.2k | Automated safety check: Pass | Apache-2.0 | |
| Lanshu Create AI Presenter Videocclank/lanshu-create-ai-presenter-video | 2.6k | — | ~3.6k | Automated safety check: Pass | MIT |
bytedance/deer-flow
Generates short videos from a structured JSON prompt, optionally guided by a reference image used as the first or last frame.
itwanger/toBeBetterJavaer
Generate matched 3:4, 16:9, and 4:3 short-video cover images from toBeBetterJavaer video scripts or AI/Java technical topics.
songguoxs/seedance-prompt-skill
This skill should be used when the user asks to "generate video prompts", "create Seedance prompts", "write video descriptions", mentions "Seedance", "seedance", "即梦", "即梦平台", "视频提示词", "视频生成"…
heygen-com/hyperframes
Entry point for making, editing and rendering videos from HTML compositions with HyperFrames, routing each request to the right workflow.
cclank/lanshu-create-ai-presenter-video
Turn a topic or finished script into a complete, publish-ready explainer video — led by an AI presenter from an authorized adult presenter image, or performed in one of nine visual explainer styles…
eternityspring/reelbench-skills
拉片:把一条成片拆成逐镜头的分析表——每个镜头的时长、景别、类别、运镜、画面. An agent skill from eternityspring/reelbench-skills.
calesthio/generative-media-skills
A skill your agent uses to turn generated, captured, scanned, or modeled 3D output into production-ready standalone assets for DCC, real-time engine, web, or interchange delivery.
calesthio/generative-media-skills
Provider-independent audio mixing and mastering direction for AI agents finishing generated videos, ads, trailers, explainers, podcasts, recuts, avatar clips, music videos, documentaries, and social…
calesthio/generative-media-skills
Provider-independent captions and media accessibility direction for AI agents producing or finishing generated videos, ads, social clips, explainers, avatar videos, documentaries, podcasts/video…
calesthio/generative-media-skills
Provider-independent production workflow for agents assembling, auditing, executing, and handing off ComfyUI node-graph workflows for image, video, upscale, inpaint, conditioning, and batch media…
calesthio/generative-media-skills
Provider-independent FFmpeg finishing workflow for AI agents preparing generated or edited media deliverables.
calesthio/generative-media-skills
Provider-independent quality assurance for AI-generated and AI-assisted media.
Categories
Production guidance for Kling AI Open Platform Advanced Lip-Sync. Kling Advanced Lip Sync is an agent skill from calesthio/generative-media-skills. Production guidance for Kling AI Open Platform Advanced Lip-Sync.
Kling Advanced Lip Sync fits situations like: identifying/selecting one face in an existing video; assigning URL/Base64; kling TTS audio; cropping and inserting audio in milliseconds.
Run `npx skills add calesthio/generative-media-skills --skill kling-advanced-lip-sync -a claude-code`. Or copy the skill folder (skills/providers/lip-sync/kling-advanced-lip-sync in calesthio/generative-media-skills) into .claude/skills/kling-advanced-lip-sync in your project. Claude Code loads it when a task matches its description.
Run `npx skills add calesthio/generative-media-skills --skill kling-advanced-lip-sync -a codex`. Or copy the skill folder (skills/providers/lip-sync/kling-advanced-lip-sync in calesthio/generative-media-skills) into .agents/skills/kling-advanced-lip-sync in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calesthio/generative-media-skills --skill kling-advanced-lip-sync -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kling-advanced-lip-sync, .gemini/skills/kling-advanced-lip-sync, .github/skills/kling-advanced-lip-sync and .opencode/skills/kling-advanced-lip-sync in your project.
SKILL.md names no scripts, command-line tools or credentials: Kling Advanced Lip Sync is instructions for the agent only.
SKILL.md names 3 domains. In commands or code: api-singapore.klingai.com; the agent is likely to contact it when it follows the instructions. As links in the text: kling.ai and klingai.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Kling Advanced Lip Sync is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Kling Advanced Lip Sync: Video Generation (bytedance/deer-flow, 84k stars), Video Cover Image (itwanger/toBeBetterJavaer, 18k stars), Seedance (songguoxs/seedance-prompt-skill, 2.9k stars) and HyperFrames Video Entry Point (heygen-com/hyperframes, 60k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
calesthio (a GitHub user) maintains it in calesthio/generative-media-skills, which has 197 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on July 14, 2026.
Source: calesthio/generative-media-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.