Agent skill

Aliyun Vidu Video

by cinience in cinience/alicloud-skills

A skill your agent uses when generating videos with DashScope Vidu models.

MITAuto-check passedMedia & Creative

Install Aliyun Vidu Video

skills CLI
$ npx skills add cinience/alicloud-skills --skill aliyun-vidu-video -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cinience/alicloud-skills aliyun-vidu-video --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cinience/alicloud-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai/video/aliyun-vidu-video .claude/skills/aliyun-vidu-video && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
aliyun-vidu-video
GitHub stars
397
Token cost
~2.7k tokens
SKILL.md length
785 words
Files
4 (incl. scripts, references)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when generating videos with DashScope Vidu models.

  • Works in 6 steps: Confirm user intent: text-to-video,… → Select the appropriate model name based… → Prepare input: prompt and/or media array… → …
  • Generating videos with DashScope Vidu models
  • SKILL.md covers Validation, Output And Evidence, Prerequisites and Critical model names, plus 10 more sections
  • Runs Python scripts from its folder; calls python; reaches dashscope.aliyuncs.com; needs DASHSCOPE_API_KEY and API_KEY

What it does

Aliyun Vidu Video is an agent skill from cinience/alicloud-skills. Use when generating videos with DashScope Vidu models. Use when implementing text-to-video, image-to-video (first frame), keyframe-to-video (first+last frame), or reference-to-video generation via the video-synthesis async API.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/api_reference.md`, `references/sources.md` and `scripts/generate_vidu_video.py`).

It sits in Media & Creative, covering AI video generation. It works with Alibaba Cloud. The repository describes itself as: alibaba cloud skills,qwen ,wan and all skills. The licence is MIT.

When your agent uses it

  • Generating videos with DashScope Vidu models
  • Implementing text-to-video
  • Image-to-video (first frame)
  • Keyframe-to-video (first+last frame)

Example prompts

  • “/aliyun-vidu-video”

Requirements

  • Python 3
  • A credential in DASHSCOPE_API_KEY
  • A credential in API_KEY

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Confirm user intent: text-to-video, image-to-video, keyframe, or reference-to-video.
  2. Select the appropriate model name based on capability and quality tier (Q3 pro/turbo or Q2).
  3. Prepare input: prompt and/or media array with correct types and valid public URLs.
  4. Set resolution, size, and duration parameters.
  5. Create async task and poll for results (15s interval recommended).
  6. Download and save generated video before URL expiration (24 hours).

What it can do on your machine

Read from SKILL.md and the folder at commit 1818263. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • dashscope.aliyuncs.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DASHSCOPE_API_KEY
    • API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Aliyun Vidu Video loads about 2.7k tokens when it runs, and up to ~4.7k if it reads all its reference files. Until then it costs about 61 tokens; SKILL.md has 785 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~61
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from cinience/alicloud-skills at commit 1818263, republished under its MIT licence (© cinience). 785 words, ~2,696 tokens.

Download SKILL.mdSave it as .claude/skills/aliyun-vidu-video/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
aliyun-vidu-video
description
Use when generating videos with DashScope Vidu models. Use when implementing text-to-video, image-to-video (first frame), keyframe-to-video (first+last frame), or reference-to-video generation via the video-synthesis async API.

Vidu Video Generation

Validation

bash
mkdir -p output/aliyun-vidu-video
python -m py_compile skills/ai/video/aliyun-vidu-video/scripts/generate_vidu_video.py && echo "py_compile_ok" > output/aliyun-vidu-video/validate.txt

Pass criteria: command exits 0 and output/aliyun-vidu-video/validate.txt is generated.

Output And Evidence

  • Save task IDs, polling responses, and final video URLs to output/aliyun-vidu-video/.
  • Keep at least one end-to-end run log for troubleshooting.

Prerequisites

  • Set DASHSCOPE_API_KEY in your environment (Beijing region key required).
  • Region: China Mainland (Beijing) only. Model, Endpoint URL, and API Key must belong to the same region.
  • Enable Vidu models in the Alibaba Cloud Model Studio console before first use.

Critical model names

Text-to-video
  • vidu/viduq3-pro_text2video
  • vidu/viduq3-turbo_text2video
  • vidu/viduq2_text2video
Image-to-video (first frame)
  • vidu/viduq3-pro_img2video
  • vidu/viduq3-turbo_img2video
  • vidu/viduq2-pro_img2video
  • vidu/viduq2-turbo_img2video
Keyframe-to-video (first+last frame)
  • vidu/viduq3-pro_start-end2video
  • vidu/viduq3-turbo_start-end2video
  • vidu/viduq2-pro_start-end2video
  • vidu/viduq2-turbo_start-end2video
Reference-to-video
  • vidu/viduq2_reference2video
  • vidu/viduq2-pro_reference2video

Capabilities

CapabilityDescriptionModel suffixRequired input
Text-to-videoGenerate video from text prompt only_text2videoprompt
Image-to-videoGenerate video from a single image + optional prompt_img2videomedia[image]
Keyframe-to-videoInterpolate video between first and last frame images_start-end2videomedia[image x2] + prompt
Reference-to-videoEmbed reference subject(s) into prompted scene_reference2videomedia[image 1-7] + prompt

API endpoint (async only)

POST https://dashscope.aliyuncs.com/api/v1/services/aigc/video-generation/video-synthesis

Required headers:

  • Authorization: Bearer $DASHSCOPE_API_KEY
  • Content-Type: application/json
  • X-DashScope-Async: enable

Normalized interface

Request
  • model (string, required) -- one of the model names listed above
  • input.prompt (string) -- up to 5000 characters, describes desired video content
    • Required for text-to-video, keyframe, and reference modes
    • Optional for image-to-video
  • input.media (array) -- media objects with type and url fields (not used for text-to-video)
    • type: image or video
    • url: public URL (HTTP/HTTPS)
  • parameters.resolution (string, optional) -- 540P, 720P (default), or 1080P
  • parameters.size (string, optional) -- pixel dimensions width*height (e.g., 1280*720). Values depend on resolution tier. For text-to-video and reference-to-video, explicit size values are supported.
  • parameters.duration (integer, optional) -- video length in seconds
    • Q3 models: [1, 16], default 5
    • Q2 models: [1, 10], default 5
  • parameters.audio (boolean, optional) -- generate audio track (Q3 models only, default false)
  • parameters.watermark (boolean, optional) -- add "AI generated" watermark (default false)
  • parameters.seed (integer, optional) -- range [0, 2147483647]
Size values by resolution tier (text-to-video)
ResolutionAspect ratioSize (width*height)
540P16:9960*528
540P9:16528*960
540P1:1720*720
540P4:3816*608
540P3:4608*816
720P16:91280*720
720P9:16720*1280
720P1:1960*960
720P4:31104*816
720P3:4816*1104
1080P16:91920*1080
1080P9:161080*1920
1080P1:11440*1440
1080P4:31674*1238
1080P3:41238*1674
Size values by resolution tier (reference-to-video)
ResolutionAspect ratioSize (width*height)
540P16:9960*540
540P9:16540*960
540P1:1540*540
540P4:3720*540
540P3:4540*720
720P16:91280*720
720P9:16720*1280
720P1:1720*720
720P4:3960*720
720P3:4720*960
1080P16:91920*1080
1080P9:161080*1920
1080P1:11080*1080
1080P4:31440*1080
1080P3:41080*1440
Media input limits

Images (type=image):

  • Formats: JPG, PNG, WEBP
  • Aspect ratio: 1:4 to 4:1
  • Max size: 50MB

Videos (type=video, reference-to-video only):

  • Formats: mp4, avi, mov
  • Resolution: min 128x128 pixels
  • Aspect ratio: 1:4 to 4:1
  • Duration: 1-5s
  • Max size: 50MB
Response (task creation)
  • output.task_id (string) -- use for polling, valid 24 hours
  • output.task_status (string) -- PENDING | RUNNING | SUCCEEDED | FAILED | CANCELED | UNKNOWN
  • request_id (string)
Show full SKILL.md (320 more words)Show less
Response (task result)
  • output.video_url (string) -- generated video URL (MP4, H.264), valid 24 hours
  • output.orig_prompt (string) -- original prompt
  • usage.duration (integer) -- billable video duration in seconds
  • usage.output_video_duration (integer) -- actual output duration
  • usage.size (string) -- output resolution
  • usage.fps (integer) -- frame rate (24)
  • usage.audio (boolean) -- whether audio was generated
  • usage.SR (string) -- resolution tier

Quick start (Python + HTTP)

python
import os
import json
import time
import requests

API_KEY = os.getenv("DASHSCOPE_API_KEY")
BASE_URL = "https://dashscope.aliyuncs.com/api/v1"

def create_vidu_task(req: dict) -> str:
    """Create a Vidu video generation task and return task_id."""
    payload = {
        "model": req["model"],
        "input": {},
        "parameters": {
            "resolution": req.get("resolution", "720P"),
            "duration": req.get("duration", 5),
        },
    }
    if req.get("prompt"):
        payload["input"]["prompt"] = req["prompt"]
    if req.get("media"):
        payload["input"]["media"] = req["media"]
    if req.get("size"):
        payload["parameters"]["size"] = req["size"]
    if req.get("audio") is not None:
        payload["parameters"]["audio"] = req["audio"]
    if req.get("watermark") is not None:
        payload["parameters"]["watermark"] = req["watermark"]
    if req.get("seed") is not None:
        payload["parameters"]["seed"] = req["seed"]

    resp = requests.post(
        f"{BASE_URL}/services/aigc/video-generation/video-synthesis",
        headers={
            "Authorization": f"Bearer {API_KEY}",
            "Content-Type": "application/json",
            "X-DashScope-Async": "enable",
        },
        json=payload,
    )
    resp.raise_for_status()
    data = resp.json()
    return data["output"]["task_id"]


def poll_task(task_id: str, interval: int = 15) -> dict:
    """Poll until task completes. Returns final response."""
    while True:
        resp = requests.get(
            f"{BASE_URL}/tasks/{task_id}",
            headers={"Authorization": f"Bearer {API_KEY}"},
        )
        resp.raise_for_status()
        data = resp.json()
        status = data["output"]["task_status"]
        if status in ("SUCCEEDED", "FAILED", "CANCELED"):
            return data
        time.sleep(interval)

Mode-specific examples

python
# Text-to-video
task_id = create_vidu_task({
    "model": "vidu/viduq3-turbo_text2video",
    "prompt": "A cat running under moonlight",
    "resolution": "540P",
    "size": "960*528",
    "duration": 5,
})

# Image-to-video (first frame)
task_id = create_vidu_task({
    "model": "vidu/viduq3-pro_img2video",
    "prompt": "Camera slowly pans upward",
    "media": [{"type": "image", "url": "https://example.com/image.jpg"}],
    "resolution": "720P",
    "duration": 5,
})

# Keyframe-to-video (first + last frame)
task_id = create_vidu_task({
    "model": "vidu/viduq3-turbo_start-end2video",
    "prompt": "A cat jumps from windowsill to sofa",
    "media": [
        {"type": "image", "url": "https://example.com/first.png"},
        {"type": "image", "url": "https://example.com/last.png"},
    ],
    "resolution": "540P",
    "duration": 5,
})

# Reference-to-video
task_id = create_vidu_task({
    "model": "vidu/viduq2_reference2video",
    "prompt": "Man playing guitar in a cafe",
    "media": [
        {"type": "image", "url": "https://example.com/ref1.jpg"},
        {"type": "image", "url": "https://example.com/ref2.jpg"},
    ],
    "resolution": "720P",
    "size": "1280*720",
    "duration": 5,
})

Error handling

ErrorLikely causeAction
401/403Missing or invalid DASHSCOPE_API_KEYCheck env var; ensure Beijing region key
400 InvalidParameterUnsupported resolution/size combo, bad duration, missing mediaValidate parameters against size tables
"does not support synchronous calls"Missing X-DashScope-Async: enable headerAdd required header
429Rate limit or quotaRetry with backoff
Cross-region errorModel and API Key from different regionsEnsure all are Beijing region

Output location

  • Default output: output/aliyun-vidu-video/videos/
  • Override base dir with OUTPUT_DIR.

Anti-patterns

  • Do not use model names not listed in "Critical model names" above.
  • Do not call this API synchronously -- async header is required.
  • Do not omit size when using reference-to-video -- it is required for that mode.
  • Do not pass audio=true with Q2 models -- only Q3 models support audio generation.
  • Video URLs expire after 24 hours; download and persist immediately.
  • For image-to-video, supply exactly 1 image. For keyframe, supply exactly 2 images (first, then last).
  • For reference-to-video with viduq2_reference2video, only images are accepted (1-7). For viduq2-pro_reference2video, images (1-4) plus optional videos (1-2) are accepted.
  • Keyframe mode: first and last frame pixel count ratio must be between 0.8 and 1.25.

Workflow

  1. Confirm user intent: text-to-video, image-to-video, keyframe, or reference-to-video.
  2. Select the appropriate model name based on capability and quality tier (Q3 pro/turbo or Q2).
  3. Prepare input: prompt and/or media array with correct types and valid public URLs.
  4. Set resolution, size, and duration parameters.
  5. Create async task and poll for results (15s interval recommended).
  6. Download and save generated video before URL expiration (24 hours).

References

  • See references/api_reference.md for full HTTP API details.
  • See references/sources.md for source links.

© cinience, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in skills/ai/video/aliyun-vidu-video of cinience/alicloud-skills.

  • SKILL.md
  • references/api_reference.md
  • references/sources.md
  • scripts/generate_vidu_video.py

Open the folder on GitHubat commit 1818263

Compare with similar skills

Aliyun Vidu Video next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Aliyun Vidu Video compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Aliyun Vidu Video this skillcinience/alicloud-skills397—~2.7kAutomated safety check: PassMIT
Bailian Media Generationmodelstudioai/cli542—~2kAutomated safety check: PassApache-2.0
Video Generationbytedance/deer-flow84k3 repos~1.4kAutomated safety check: PassMIT
Video Cover Imageitwanger/toBeBetterJavaer18k—~3.3kAutomated safety check: PassNone
Seedancesongguoxs/seedance-prompt-skill2.9k1 repos~2.5kAutomated safety check: PassNone
HyperFrames Video Entry Pointheygen-com/hyperframes60k3 repos~5.2kAutomated safety check: PassApache-2.0

Similar skills

  • Bailian Media Generation

    modelstudioai/cli

    Chinese-language entry point into Alibaba Cloud Bailian's image, video and speech generation and understanding, routed through separate image, video, speech and vision commands.

    542 GitHub stars~2k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Video Generation

    bytedance/deer-flow

    Generates short videos from a structured JSON prompt, optionally guided by a reference image used as the first or last frame.

    84k GitHub starsUsed in 3 repos~1.4k tokens
    Media & CreativeAuto-check passed
  • Video Cover Image

    itwanger/toBeBetterJavaer

    Generate matched 3:4, 16:9, and 4:3 short-video cover images from toBeBetterJavaer video scripts or AI/Java technical topics.

    18k GitHub stars~3.3k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Seedance

    songguoxs/seedance-prompt-skill

    This skill should be used when the user asks to "generate video prompts", "create Seedance prompts", "write video descriptions", mentions "Seedance", "seedance", "即梦", "即梦平台", "视频提示词", "视频生成"…

    2.9k GitHub starsUsed in 1 repo~2.5k tokens
    Media & CreativeAuto-check passed
  • HyperFrames Video Entry Point

    heygen-com/hyperframes

    Entry point for making, editing and rendering videos from HTML compositions with HyperFrames, routing each request to the right workflow.

    60k GitHub starsUsed in 3 repos~5.2k tokens
    Media & CreativeAuto-check passed
  • Lanshu Create AI Presenter Video

    cclank/lanshu-create-ai-presenter-video

    Turn a topic or finished script into a complete, publish-ready explainer video — led by an AI presenter from an authorized adult presenter image, or performed in one of nine visual explainer styles…

    2.6k GitHub stars~3.6k tokensUpdated today
    Media & CreativeAuto-check passed

More from cinience/alicloud-skills

All 96 skills in this repo
  • Aliyun Skill Creator

    cinience/alicloud-skills

    A skill your agent uses when creating, migrating, or optimizing skills for this alicloud-skills repository.

    397 GitHub stars~2.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Alicloud Acs Agent Sandbox

    cinience/alicloud-skills

    Bootstrap, create, connect to, operate, secure, scale, upgrade, troubleshoot, inspect, and tear down Alibaba Cloud Container Compute Service (ACS) Agent Sandbox environments.

    397 GitHub stars~2.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Alicloud Acs Cluster

    cinience/alicloud-skills

    Create, inspect, connect to, inventory, and delete Alibaba Cloud Container Compute Service (ACS) clusters through the official CS OpenAPI.

    397 GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Aliyun Adb Mysql

    cinience/alicloud-skills

    A skill your agent uses when managing Alibaba Cloud AnalyticDB for MySQL (ADB) via OpenAPI/SDK, including the user needs AnalyticDB resource lifecycle and configuration operations, status checks, or…

    397 GitHub stars~708 tokensUpdated 2 mo ago
    Auto-check passed
  • Aliyun Aicontent Generate

    cinience/alicloud-skills

    A skill your agent uses when managing Alibaba Cloud AIContent (AiContent) via OpenAPI/SDK, including the user needs AI content generation or content workflow operations in Alibaba Cloud, including…

    397 GitHub stars~734 tokensUpdated 2 mo ago
    Auto-check passed
  • Aliyun Aimiaobi Generate

    cinience/alicloud-skills

    A skill your agent uses when managing Alibaba Cloud Quan Miao (AiMiaoBi) via OpenAPI/SDK, including the user asks for Alibaba Cloud MiaoBi content operations, including listing resources…

    397 GitHub stars~724 tokensUpdated 2 mo ago
    Auto-check passed

Works with

Questions about Aliyun Vidu Video

What does Aliyun Vidu Video do?

A skill your agent uses when generating videos with DashScope Vidu models. Aliyun Vidu Video is an agent skill from cinience/alicloud-skills. Use when generating videos with DashScope Vidu models.

When should I use Aliyun Vidu Video?

Aliyun Vidu Video fits situations like: generating videos with DashScope Vidu models; implementing text-to-video; image-to-video (first frame); keyframe-to-video (first+last frame).

How do I install Aliyun Vidu Video in Claude Code?

Run `npx skills add cinience/alicloud-skills --skill aliyun-vidu-video -a claude-code`. Or copy the skill folder (skills/ai/video/aliyun-vidu-video in cinience/alicloud-skills) into .claude/skills/aliyun-vidu-video in your project. Claude Code loads it when a task matches its description.

How do I install Aliyun Vidu Video in Codex?

Run `npx skills add cinience/alicloud-skills --skill aliyun-vidu-video -a codex`. Or copy the skill folder (skills/ai/video/aliyun-vidu-video in cinience/alicloud-skills) into .agents/skills/aliyun-vidu-video in your project. Codex loads it when a task matches its description.

Can I use Aliyun Vidu Video in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cinience/alicloud-skills --skill aliyun-vidu-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/aliyun-vidu-video, .gemini/skills/aliyun-vidu-video, .github/skills/aliyun-vidu-video and .opencode/skills/aliyun-vidu-video in your project.

What does Aliyun Vidu Video need to run?

Going by SKILL.md and its folder, Aliyun Vidu Video needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named DASHSCOPE_API_KEY and API_KEY. Our summary lists: Python 3; A credential in DASHSCOPE_API_KEY; A credential in API_KEY.

Does Aliyun Vidu Video access the network?

SKILL.md names 1 domain. In commands or code: dashscope.aliyuncs.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Aliyun Vidu Video safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Aliyun Vidu Video use?

Aliyun Vidu Video is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Aliyun Vidu Video use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2k tokens, read only when the agent opens those files.

What are the alternatives to Aliyun Vidu Video?

Skills that share tags, products or a category with Aliyun Vidu Video: Bailian Media Generation (modelstudioai/cli, 542 stars), Video Generation (bytedance/deer-flow, 84k stars), Video Cover Image (itwanger/toBeBetterJavaer, 18k stars) and Seedance (songguoxs/seedance-prompt-skill, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Aliyun Vidu Video?

cinience (a GitHub user) maintains it in cinience/alicloud-skills, which has 397 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on August 11, 2026.

Source: cinience/alicloud-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.