Agent skill

Kling Advanced Lip Sync

by calesthio in calesthio/generative-media-skills

Production guidance for Kling AI Open Platform Advanced Lip-Sync.

MITAuto-check passedMedia & Creative

Install Kling Advanced Lip Sync

skills CLI
$ npx skills add calesthio/generative-media-skills --skill kling-advanced-lip-sync -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install calesthio/generative-media-skills kling-advanced-lip-sync --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/providers/lip-sync/kling-advanced-lip-sync .claude/skills/kling-advanced-lip-sync && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
kling-advanced-lip-sync
GitHub stars
197
Token cost
~2.9k tokens
SKILL.md length
1,306 words
Files
2
Skills in repo
26
Repo updated
First seen
Licence
MIT

At a glance

Production guidance for Kling AI Open Platform Advanced Lip-Sync.

  • Works in 4 steps: POST /v1/videos/identify-face → POST /v1/videos/advanced-lip-sync → GET /v1/videos/advanced-lip-sync/{id} → …
  • Identifying/selecting one face in an existing video
  • SKILL.md covers Verification and evidence, Official workflow, Authentication and Validate the video source, plus 10 more sections
  • Reaches api-singapore.klingai.com

What it does

Kling Advanced Lip Sync is an agent skill from calesthio/generative-media-skills. Production guidance for Kling AI Open Platform Advanced Lip-Sync. Use for identifying/selecting one face in an existing video, assigning URL/Base64 or Kling TTS audio, cropping and inserting audio in milliseconds, mixing original sound, creating and monitoring tasks, downloading expiring results, consent/privacy review, sync QA, and repair. Do not use for still-image avatar generation or ordinary Kling video generation.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `EVAL.md`).

It sits in Media & Creative, covering AI video generation. The repository describes itself as: Research-backed agent skills and tools for premium image, video, audio, voice, and generative media production across AI coding assistants. The licence is MIT.

When your agent uses it

  • Identifying/selecting one face in an existing video
  • Assigning URL/Base64
  • Kling TTS audio
  • Cropping and inserting audio in milliseconds

Example prompts

  • “/kling-advanced-lip-sync”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. POST /v1/videos/identify-face
  2. POST /v1/videos/advanced-lip-sync
  3. GET /v1/videos/advanced-lip-sync/{id}
  4. GET /v1/videos/advanced-lip-sync

What it can do on your machine

Read from SKILL.md and the folder at commit 8c85352. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are http and json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api-singapore.klingai.com

    Also links to:

    • kling.ai
    • klingai.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Kling Advanced Lip Sync loads about 2.9k tokens when it runs. Until then it costs about 112 tokens; SKILL.md has 1,306 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~112
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from calesthio/generative-media-skills at commit 8c85352, republished under its MIT licence (© calesthio). 1,306 words, ~2,859 tokens.

Download SKILL.mdSave it as .claude/skills/kling-advanced-lip-sync/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
kling-advanced-lip-sync
description
Production guidance for Kling AI Open Platform Advanced Lip-Sync. Use for identifying/selecting one face in an existing video, assigning URL/Base64 or Kling TTS audio, cropping and inserting audio in milliseconds, mixing original sound, creating and monitoring tasks, downloading expiring results, consent/privacy review, sync QA, and repair. Do not use for still-image avatar generation or ordinary Kling video generation.

Kling Advanced Lip-Sync

Use this skill for Kling AI Open Platform's documented Advanced Lip-Sync workflow on an existing video. It is distinct from Avatar, which generates a new talking performance from a still image and audio.

Verification and evidence

  • Documented fact: behavior stated in first-party Kling Open Platform pages.
  • Production heuristic: a practical sync/repair recommendation not guaranteed by Kling.
  • Empirical observation: a result measured from an actual task/output.

Facts were verified 2026-07-12. Endpoint visibility, entitlement, pricing, callback behavior, retention, limits, and policy can change. Re-check official documentation, authenticated console, account package, and applicable agreement before paid use.

Official workflow

Base URL in current examples:

text
https://api-singapore.klingai.com

Operations:

  1. POST /v1/videos/identify-face
  2. POST /v1/videos/advanced-lip-sync
  3. GET /v1/videos/advanced-lip-sync/{id}
  4. GET /v1/videos/advanced-lip-sync

Although the navigation may call a tab Face Recognition, the operation is Identify Face. Describe it as face identification/detection for selecting a visible face; do not claim general identity matching.

Authentication

Official docs use JWT Bearer authentication derived from Kling Open Platform access/secret credentials:

http
Authorization: Bearer ${KLING_JWT}
Content-Type: application/json

Follow the current authentication page for token construction and lifetime. Generate server-side, use least privilege, and never place keys/secrets/JWTs in browser code, logs, callbacks, or fixtures.

Validate the video source

Identify Face accepts exactly one:

  • video_id: Kling-generated video from the documented recent window;
  • video_url: accessible external video.

They are mutually exclusive.

Current external-video constraints include .mp4 or .mov, 2-60 seconds, at most 100 MB, 720p or 1080p, and width/height each 512-2160 pixels. Current docs also limit Kling video_id recency to 30 days. Kling performs content validation.

Preflight codec/container, duration, dimensions, file size, URL reachability, audio presence, face visibility, cuts, occlusion, and consent. Do not upload a private URL that Kling cannot access or a mutable URL that can change between operations.

Identify and select a face

Example:

http
POST /v1/videos/identify-face

{"video_url":"https://example.com/source.mp4"}

The response includes session_id and face_data[]; each face can include face_id, face_image, start_time, and end_time in milliseconds.

Present detected face crops/intervals to the user or authorized operator. If several faces exist, require an explicit face_id unless a documented, approved selection rule exists. Record selected face, reason, and interval.

face_choose is an array, but current docs explicitly say only one-person lip-sync is supported. Do not claim simultaneous multi-speaker replacement. Re-identify after any source-video edit; do not reuse a stale session against changed media.

Kling does not publicly establish session expiry/reuse semantics in the verified docs. Keep identify/create close together and handle invalid sessions as a new identification job.

Prepare audio and timing

For the selected face, provide exactly one:

  • audio_id: recent Kling TTS audio;
  • sound_file: accessible URL or Base64 audio.

They are mutually exclusive.

Current external-audio constraints include .mp3, .wav, .m4a, or .aac, 2-60 seconds, and at most 5 MB. Current docs limit audio_id recency to 30 days.

Timing fields are milliseconds:

  • sound_start_time: crop start in source audio;
  • sound_end_time: crop end;
  • sound_insert_time: insertion point in source video.

Validate:

$$ cropDuration = soundEnd - soundStart \ge 2000\text{ ms} $$

The crop must fit inside the source audio, the inserted range must fit inside the video, and the inserted range must overlap the selected face's detected appearance interval by at least two seconds.

Use explicit millisecond arithmetic. Never pass seconds by accident.

Mix controls:

  • sound_volume: current documented range 0-2, default 1;
  • original_audio_volume: range 0-2, default 1; no effect for silent source.

Preserve ambience when it matters, but ensure source dialogue does not create double speech. Review the final mix separately from lip motion.

Create the task

Illustrative request:

json
{
  "session_id": "<identify-face-session>",
  "face_choose": [
    {
      "face_id": "<approved-face-id>",
      "sound_file": "https://example.com/dialogue.wav",
      "sound_start_time": 0,
      "sound_end_time": 3000,
      "sound_insert_time": 1000,
      "sound_volume": 1,
      "original_audio_volume": 0.25
    }
  ],
  "external_task_id": "scene-014-take-03",
  "callback_url": "https://example.com/hooks/kling",
  "watermark_info": {"enabled": false}
}

This assumes a three-second crop, sufficient video duration, and at least two seconds of overlap with the selected face interval.

external_task_id uniqueness is caller-owned. Store it with the system task_id, source/audio hashes, session/face IDs, timings, volumes, attempt, consent record, and cost preflight.

Monitor and reconcile

Task statuses:

  • submitted
  • processing
  • succeed
  • failed

Use the single-task endpoint as source of truth. Callbacks can reduce polling latency, but production handling should be idempotent and accept duplicates/out-of-order events. Unless the current callback page documents a signature mechanism, do not invent one. Retrieve the task before trusting or publishing callback data.

List queries use current pagination ranges; do not scan the list when a task ID is known.

Retry transient rate/concurrency/server errors with bounded jitter. Do not retry auth, permission, parameter, policy, or invalid-media failures unchanged. Preserve provider code/message/request ID.

Cost, access, and result custody

The public pricing page did not clearly establish a dedicated Advanced Lip-Sync list price during verification. Do not assume Avatar pricing applies. Check authenticated console/package entitlement and create a dated budget approval.

Successful task responses expose final deduction fields; reconcile actual units/balance against the estimate.

Generated media is documented as deleted after 30 days. Download promptly into controlled storage, checksum, probe, and preserve provenance. Do not use provider URLs as permanent delivery storage.

Endpoint entitlement and concurrency are account/package dependent. Errors such as rate, concurrency, package exhaustion, or permission denial should be surfaced precisely.

Show full SKILL.md (513 more words)Show less

This workflow processes recognizable face video and voice/audio. Require:

  • explicit recorded authority for likeness and audio/voice;
  • intended use, audience, territory, term, modification scope, revocation, and retention;
  • additional review for minors, employees, public figures, political use, deceased persons, and sensitive contexts;
  • no deceptive impersonation or undisclosed material alteration;
  • access controls and deletion procedure for source video, face crops, audio, sessions, and output;
  • platform/legal disclosure and provenance where required.

Kling's technical face_id does not prove identity or consent. Review current API privacy/terms for hosting, storage, transfer, processing, model improvement, and enterprise exceptions. Do not make unqualified training, residency, or deletion claims.

Sync QA and repair

Kling does not publish a formal sync-quality metric or repair API. The following are production heuristics:

  1. Review normal speed and frame-by-frame around plosives, labials, word starts, and phrase ends.
  2. Check mouth motion begins/ends with speech rather than room tone.
  3. Verify the selected face stays visible across the insertion.
  4. Inspect profiles, cuts, occlusion, hand-over-mouth, facial hair, and motion blur.
  5. Check jaw, teeth/tongue, frozen expression, identity drift, and surrounding pixels.
  6. Check ambience continuity and double speech after mix changes.
  7. Repair timing first by recropping or changing insertion time.
  8. Shorten difficult segments rather than repeatedly processing a full clip.
  9. Preserve an untouched master and compare outside the altered range.

Do not claim a rerun will necessarily improve output. Change one variable, record it, and compare.

Example 1: single-speaker localization insert

This is a complete example, not a mandatory formula.

Intent: replace a three-second line in an authorized interview while keeping room ambience.

Preflight a 12-second 1080p MP4 under 100 MB, confirm one visible consented speaker, and provide a 3.2-second WAV under 5 MB. Identify faces, select the approved face whose interval covers seconds 2-10, then crop audio 100-3100 ms and insert at 3500 ms. This yields a 3000 ms insertion within video and at least two seconds inside the face interval. Set source dialogue low enough to avoid doubling while retaining ambience.

Monitor by task ID, download/checksum immediately, then review consonants, start/end motion, identity, occlusion, and ambience. If sync starts late, adjust insertion/crop before changing volume or regenerating audio.

Example 2: multiple detected faces

This is a complete example, not a mandatory formula.

Intent: replace one host line in a two-person panel clip.

Identify Face returns two faces with distinct crops and intervals. Present both; the authorized editor selects the host. Because only one-person lip-sync is documented, create one task for that face and do not promise simultaneous guest replacement. Choose a segment where the host remains visible and the guest does not occlude the mouth. Keep original audio if only ambience remains; otherwise use a prepared mix.

If the requirement becomes replacing both speakers, stop and choose another supported workflow rather than chaining undocumented multi-face tasks.

Sources

Official sources verified 2026-07-12:

© calesthio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/providers/lip-sync/kling-advanced-lip-sync of calesthio/generative-media-skills.

  • SKILL.md
  • EVAL.md

Open the folder on GitHubat commit 8c85352

Compare with similar skills

Kling Advanced Lip Sync next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Kling Advanced Lip Sync compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Kling Advanced Lip Sync this skillcalesthio/generative-media-skills197—~2.9kAutomated safety check: PassMIT
Video Generationbytedance/deer-flow84k3 repos~1.4kAutomated safety check: PassMIT
Video Cover Imageitwanger/toBeBetterJavaer18k—~3.3kAutomated safety check: PassNone
Seedancesongguoxs/seedance-prompt-skill2.9k1 repos~2.5kAutomated safety check: PassNone
HyperFrames Video Entry Pointheygen-com/hyperframes60k3 repos~5.2kAutomated safety check: PassApache-2.0
Lanshu Create AI Presenter Videocclank/lanshu-create-ai-presenter-video2.6k—~3.6kAutomated safety check: PassMIT

Similar skills

  • Video Generation

    bytedance/deer-flow

    Generates short videos from a structured JSON prompt, optionally guided by a reference image used as the first or last frame.

    84k GitHub starsUsed in 3 repos~1.4k tokens
    Media & CreativeAuto-check passed
  • Video Cover Image

    itwanger/toBeBetterJavaer

    Generate matched 3:4, 16:9, and 4:3 short-video cover images from toBeBetterJavaer video scripts or AI/Java technical topics.

    18k GitHub stars~3.3k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Seedance

    songguoxs/seedance-prompt-skill

    This skill should be used when the user asks to "generate video prompts", "create Seedance prompts", "write video descriptions", mentions "Seedance", "seedance", "即梦", "即梦平台", "视频提示词", "视频生成"…

    2.9k GitHub starsUsed in 1 repo~2.5k tokens
    Media & CreativeAuto-check passed
  • HyperFrames Video Entry Point

    heygen-com/hyperframes

    Entry point for making, editing and rendering videos from HTML compositions with HyperFrames, routing each request to the right workflow.

    60k GitHub starsUsed in 3 repos~5.2k tokens
    Media & CreativeAuto-check passed
  • Lanshu Create AI Presenter Video

    cclank/lanshu-create-ai-presenter-video

    Turn a topic or finished script into a complete, publish-ready explainer video — led by an AI presenter from an authorized adult presenter image, or performed in one of nine visual explainer styles…

    2.6k GitHub stars~3.6k tokensUpdated today
    Media & CreativeAuto-check passed
  • Video Shots

    eternityspring/reelbench-skills

    拉片:把一条成片拆成逐镜头的分析表——每个镜头的时长、景别、类别、运镜、画面. An agent skill from eternityspring/reelbench-skills.

    878 GitHub starsUsed in 1 repo~1.8k tokens
    Media & CreativeAuto-check: notes

More from calesthio/generative-media-skills

All 26 skills in this repo
  • 3D Asset Production

    calesthio/generative-media-skills

    A skill your agent uses to turn generated, captured, scanned, or modeled 3D output into production-ready standalone assets for DCC, real-time engine, web, or interchange delivery.

    197 GitHub stars~9.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Audio Mixing Mastering

    calesthio/generative-media-skills

    Provider-independent audio mixing and mastering direction for AI agents finishing generated videos, ads, trailers, explainers, podcasts, recuts, avatar clips, music videos, documentaries, and social…

    197 GitHub stars~7k tokensUpdated 2 mo ago
    Auto-check passed
  • Captions Media Accessibility

    calesthio/generative-media-skills

    Provider-independent captions and media accessibility direction for AI agents producing or finishing generated videos, ads, social clips, explainers, avatar videos, documentaries, podcasts/video…

    197 GitHub stars~6.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Comfyui Media Workflows

    calesthio/generative-media-skills

    Provider-independent production workflow for agents assembling, auditing, executing, and handing off ComfyUI node-graph workflows for image, video, upscale, inpaint, conditioning, and batch media…

    197 GitHub stars~8.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Ffmpeg Media Finishing

    calesthio/generative-media-skills

    Provider-independent FFmpeg finishing workflow for AI agents preparing generated or edited media deliverables.

    197 GitHub stars~8.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Generated Media QA

    calesthio/generative-media-skills

    Provider-independent quality assurance for AI-generated and AI-assisted media.

    197 GitHub stars~8k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Kling Advanced Lip Sync

What does Kling Advanced Lip Sync do?

Production guidance for Kling AI Open Platform Advanced Lip-Sync. Kling Advanced Lip Sync is an agent skill from calesthio/generative-media-skills. Production guidance for Kling AI Open Platform Advanced Lip-Sync.

When should I use Kling Advanced Lip Sync?

Kling Advanced Lip Sync fits situations like: identifying/selecting one face in an existing video; assigning URL/Base64; kling TTS audio; cropping and inserting audio in milliseconds.

How do I install Kling Advanced Lip Sync in Claude Code?

Run `npx skills add calesthio/generative-media-skills --skill kling-advanced-lip-sync -a claude-code`. Or copy the skill folder (skills/providers/lip-sync/kling-advanced-lip-sync in calesthio/generative-media-skills) into .claude/skills/kling-advanced-lip-sync in your project. Claude Code loads it when a task matches its description.

How do I install Kling Advanced Lip Sync in Codex?

Run `npx skills add calesthio/generative-media-skills --skill kling-advanced-lip-sync -a codex`. Or copy the skill folder (skills/providers/lip-sync/kling-advanced-lip-sync in calesthio/generative-media-skills) into .agents/skills/kling-advanced-lip-sync in your project. Codex loads it when a task matches its description.

Can I use Kling Advanced Lip Sync in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calesthio/generative-media-skills --skill kling-advanced-lip-sync -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kling-advanced-lip-sync, .gemini/skills/kling-advanced-lip-sync, .github/skills/kling-advanced-lip-sync and .opencode/skills/kling-advanced-lip-sync in your project.

What does Kling Advanced Lip Sync need to run?

SKILL.md names no scripts, command-line tools or credentials: Kling Advanced Lip Sync is instructions for the agent only.

Does Kling Advanced Lip Sync access the network?

SKILL.md names 3 domains. In commands or code: api-singapore.klingai.com; the agent is likely to contact it when it follows the instructions. As links in the text: kling.ai and klingai.com. This is read from the text; nothing was executed.

Is Kling Advanced Lip Sync safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Kling Advanced Lip Sync use?

Kling Advanced Lip Sync is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Kling Advanced Lip Sync use?

About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Kling Advanced Lip Sync?

Skills that share tags, products or a category with Kling Advanced Lip Sync: Video Generation (bytedance/deer-flow, 84k stars), Video Cover Image (itwanger/toBeBetterJavaer, 18k stars), Seedance (songguoxs/seedance-prompt-skill, 2.9k stars) and HyperFrames Video Entry Point (heygen-com/hyperframes, 60k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Kling Advanced Lip Sync?

calesthio (a GitHub user) maintains it in calesthio/generative-media-skills, which has 197 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on July 14, 2026.

Source: calesthio/generative-media-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.