Bailian Media Generation
modelstudioai/cli
Chinese-language entry point into Alibaba Cloud Bailian's image, video and speech generation and understanding, routed through separate image, video, speech and vision commands.
A skill your agent uses when authoring or validating PAIDF augmentation YAML configs, or running remote Cosmos Transfer (including Cosmos3 WSM controls), Cosmos Predict, image-edit, or…
$ npx skills add NVIDIA/skills --skill paidf-augmentation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills paidf-augmentation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/paidf-augmentation .claude/skills/paidf-augmentation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "paidf-augmentation" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/paidf-augmentation into .claude/skills/paidf-augmentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "paidf-augmentation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/paidf-augmentationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill paidf-augmentation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills paidf-augmentation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/paidf-augmentation .agents/skills/paidf-augmentation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "paidf-augmentation" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/paidf-augmentation into .agents/skills/paidf-augmentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "paidf-augmentation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill paidf-augmentation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills paidf-augmentation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/paidf-augmentation .cursor/skills/paidf-augmentation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "paidf-augmentation" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/paidf-augmentation into .cursor/skills/paidf-augmentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "paidf-augmentation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/paidf-augmentation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill paidf-augmentation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills paidf-augmentation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/paidf-augmentation .gemini/skills/paidf-augmentation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "paidf-augmentation" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/paidf-augmentation into .gemini/skills/paidf-augmentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "paidf-augmentation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills paidf-augmentationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill paidf-augmentation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/paidf-augmentation .github/skills/paidf-augmentation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "paidf-augmentation" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/paidf-augmentation into .github/skills/paidf-augmentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "paidf-augmentation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill paidf-augmentation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills paidf-augmentation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/paidf-augmentation .opencode/skills/paidf-augmentation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "paidf-augmentation" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/paidf-augmentation into .opencode/skills/paidf-augmentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "paidf-augmentation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
paidf-augmentationA skill your agent uses when authoring or validating PAIDF augmentation YAML configs, or running remote Cosmos Transfer (including Cosmos3 WSM controls), Cosmos Predict, image-edit, or…
Paidf Augmentation is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Use when authoring or validating PAIDF augmentation YAML configs, or running remote Cosmos Transfer (including Cosmos3 WSM controls), Cosmos Predict, image-edit, or image-to-video inference.
Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including reference files (for example `BENCHMARK.md`, `evals/evals.json` and `references/captioning-strategy-guide.md`).
It sits in Media & Creative, covering AI video generation. It works with NVIDIA AI Platform, CUDA and Qwen. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 67a13c0. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
dockergituvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use docker, git and uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HF_TOKENVLM_API_KEYLLM_API_KEYVEO_API_KEYAWS_ACCESS_KEY_IDAWS_SECRET_ACCESS_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Paidf Augmentation loads about 5.3k tokens when it runs, and up to ~36k if it reads all its reference files. Until then it costs about 52 tokens; SKILL.md has 2,268 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit 67a13c0, republished under its Apache-2.0 licence (© NVIDIA). 2,268 words, ~5,309 tokens.
.claude/skills/paidf-augmentation/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.Remote-API pipeline for captioning, generating, and evaluating augmented camera data. Models are configured as HTTP endpoints; no local model weights are included.
Use it to select a supported generation mode, author and validate
PipelineConfig YAML, configure captioning/evaluators, and run the
paidf-augmentation:1.2.0 container. It covers Cosmos Transfer/Predict,
Cosmos3 WSM controls, image editing, image-to-video, BYOM endpoints, and
quality gates.
Do not use this skill for training or fine-tuning models, deploying clusters or NIM endpoints, or unrelated application/database development.
| Requirement | Detail |
|---|---|
| Docker | docker --version. The image is remote-API only — it bundles no Cosmos/torch weights, so plain remote inference needs no GPU and no HF_TOKEN. |
| NVIDIA GPU (conditional) | Only when a configured local stage requires CUDA or decodes H.264 in the augmentation container. See Limitations. |
| Endpoint URLs | One reachable URL per role used by the selected config. The examples name local Qwen services (Qwen/Qwen3.6-27B-FP8 on vlm, Qwen/Qwen2.5-14B-Instruct on llm), but those are not guaranteed to be running. Resolve every required role independently: keep each reachable configured endpoint and ask only for the missing role URLs. |
| API keys (conditional) | Only for endpoints requiring authentication. Pass the environment variable named by each endpoint's api_key_env; never hardcode values in YAML. Local unauthenticated endpoints need none. |
| Input media | A video/image reachable by multistorageclient (s3://, msc://, gs://, az://, or HTTPS). Local paths are allowed only when no external Cosmos evaluator is enabled; external evaluation requires checker-readable remote URIs. |
| Approved release commit (execution only) | A 40-character commit hash confirmed by the user as release-owner-approved or retrieved from the verified release manifest. Never invent it or rely on an unverified environment value. |
Resolve each value in this precedence order: state file → explicit prompt arguments → agent context → user prompt. Ask the user only for what remains unresolved.
| Input | Required | Description |
|---|---|---|
config_path | Yes | Path to the pipeline YAML, e.g. configs/cookbook/video-data-augmentation/config_video_transfer_CT3_omni.yaml. If absent, pick a starting config from Supported Models and confirm with the user. |
input_media | Conditional | Source video/image → data[].inputs.rgb. Required for every mode except Cosmos Predict inference_type: text2world, where inputs may be null or rgb omitted. Overridable at run time via data.0.inputs.rgb=.... |
control_media | For precomputed control | Control video → data[].inputs.controls.<type>. For Cosmos3 WSM transfer, set controls.wsm; the client uploads it rather than exposing its local path to the server. |
output_paths | Yes | data[].output.{video,caption,metadata}; evaluation optional. |
model_name | Yes | augmentation.model.name — an endpoint id, a role, or a known model name. Free-form string, not an enum. |
endpoint_urls | Yes | One endpoints[] entry per role in use. |
api_key_env | If auth | Env-var name per endpoint; the value comes from the environment. |
target_attributes | No | captioning.llm.variables (e.g. weather_condition, lighting_condition). |
generation_params | No | augmentation.parameters — pass-through; only set knobs are sent. |
seed | No | Under augmentation.parameters. Before the first candidate, null resolves once to the generator/executor seed when available, otherwise the current Unix time. Retry n uses that resolved base seed plus n. pipeline.retry defaults to 1, so evaluation permits at most 2 candidates by default. |
Each endpoints: list entry declares role, url, wire model, and optional
id, adapter, api_key_env, and timeout. Roles are vlm, llm,
image_edit, video_transfer, video_predict, image2video, and evaluator.
augmentation.model.name resolves by endpoint id, then role, then the known
model-name mapping. Adapter contracts and full fields are in
configuration-schema.md.
When the user hasn't specified a model, choose from their input type and goal:
| Input Type → Goal | model.name | Role / default adapter | Input → Output |
|---|---|---|---|
| Video — change scene attributes (weather, lighting, style) | cosmos-transfer2.5 | video_transfer / nim | Video (+ controls) → Video |
| Video + precomputed WSM — follow world-state geometry/motion | A video_transfer endpoint id, e.g. cosmos3-transfer-wsm | video_transfer / openai.video.async | RGB video + WSM control + prompt → Video |
| Video + text — extend or predict continuation | cosmos-predict | video_predict / nim | Video+Text → Video |
| Text only — generate video from scratch | cosmos-predict (inference_type: text2world) | video_predict / nim | Text → Video |
| Image — edit specific attributes | image-edit | image_edit / nim (or openai.chat.completions, openai.images.edits) | Image → Image |
| Image — animate a first frame | cosmos3-image2video (or your Veo endpoint id) | image2video / openai.video.sync (Veo: openai.video.async) | Image + prompt → Video |
Key rule: use Cosmos Transfer for video scene changes, Cosmos Predict for
new/continued video, image edit for still edits, and image-to-video to animate a
frame. For Cosmos3 WSM, set data[].inputs.controls.wsm. Config selection
details are in config-decision-tree.md.
All models run via remote HTTP through one BaseExecutor; there is no local torchrun and no executor_type field.
Follow this ordered decision tree; open the linked references only for the selected mode's details.
config_attempts counts YAML versions, not validation or preflight calls. Keep
it across the whole workflow and never reset it.
1 after initial authoring.3, STOP and report the error.
Otherwise increment it once and make one correction pass containing all
currently reported YAML fixes.Execute this gate before every other step:
data, endpoints, captioning,
augmentation, evaluators, pipeline, then data_processing. Omit
optional sections rather than creating placeholders. Set
config_attempts = 1 (config attempt 1 of 3).EXPECTED_RELEASE_REF from the user as a release-owner-approved
full commit or retrieve it from the verified release manifest.
b. Verify the Git working tree is clean.
c. Verify HEAD matches EXPECTED_RELEASE_REF. If a release tag selected
the revision, verify the tag first, then compare its resolved full commit.
d. Build the image and capture its immutable sha256: image ID, or obtain a
verified registry digest.
e. Create the paidf Docker bridge if needed and attach local services.
f. Launch the container using the recorded image ID, the paidf network,
required volume mounts, --entrypoint /bin/bash, and only the endpoint
key variables named by api_key_env plus the scoped storage credential
variables required for the configured remote media.
If any sub-step fails, STOP and report the exact failure and relevant
remediation: clean a dirty tree, supply the approved revision, resolve a
revision mismatch, check the Docker daemon/build, fix network setup, or fix
the container launch. Do not continue with a partial runtime.config_attempts (current config
attempt N of 3):config_attempts
(current config attempt N of 3) and run its hard preflight gate before
captioning/generation by querying /health, /checkers, and
/dependency-graph:/health.status == "healthy"; require every configured
checks[].name to exactly equal one string in the top-level checkers
array returned by /checkers and one key in the top-level checkers
object returned by /dependency-graph. Do not use substring, alias, or
nested-field matching. Then proceed to step 7./checkers or
/dependency-graph whose body explicitly identifies a configured checker
name or invalid config field qualifies. Apply the Config Correction
Budget rule, then return to step 5; revalidation does not increment the
counter./health failure, unhealthy service, or missing deployed checker stops the
workflow before generation. Create no candidate and report the root cause.pipeline.retry + 1 candidates (retry defaults to 1, so the default
maximum is 2). Resolve the base seed once before the first candidate as
defined in Inputs; retry n uses base_seed + n:base_seed + next_retry_number. If regenerate_caption_on_retry: true,
regenerate the caption next; then generate the replacement. If false,
generate the replacement immediately.passed. With pipeline.evaluation.strict: true (the
default), apply retain_failures: its default true retains video, caption,
metadata, and evaluation outputs for inspection; explicit false deletes
them. With strict: false, retain the files and report the evaluator
failure without failing the sample.error_code: ambiguous_post, preserve the
retained candidate, do not resubmit automatically, STOP, and defer
to operator recovery. Follow the
ambiguous-POST recovery flow.Set EXPECTED_RELEASE_REF to a release-owner-reviewed full 40-character commit,
never a tag or branch. If a signed release tag selects the revision, verify the
tag first and record its resolved full commit as EXPECTED_RELEASE_REF. Build
only from a clean matching checkout, then record the resulting immutable image
ID. For a prebuilt release, use its verified digest.
set -e
EXPECTED_RELEASE_REF="${EXPECTED_RELEASE_REF:?set a reviewed full commit ID}"
test "${#EXPECTED_RELEASE_REF}" -eq 40
case "$EXPECTED_RELEASE_REF" in *[!0-9a-fA-F]*) exit 1 ;; esac
test -z "$(git status --porcelain)"
test "$(git rev-parse HEAD)" = "$(git rev-parse "${EXPECTED_RELEASE_REF}^{commit}")"
DOCKER_BUILDKIT=0 docker build \
-t paidf-augmentation:1.2.0 \
-f docker/Dockerfile .
PAIDF_IMAGE_ID="$(docker image inspect --format '{{.Id}}' paidf-augmentation:1.2.0)"
case "$PAIDF_IMAGE_ID" in sha256:*) ;; *) exit 1 ;; esac
test "$(docker image inspect --format '{{.Id}}' paidf-augmentation:1.2.0)" = "$PAIDF_IMAGE_ID"
docker network inspect paidf >/dev/null 2>&1 || \
docker network create paidf
# Attach each local model container once, for example:
# docker network connect paidf vlm
# Example only: replace these with exactly the api_key_env names declared by
# the selected config; use an empty array when every endpoint is unauthenticated.
PAIDF_ENDPOINT_KEY_ARGS=()
# Authenticated example:
# PAIDF_ENDPOINT_KEY_ARGS=(-e VLM_API_KEY -e LLM_API_KEY -e VEO_API_KEY)
# Forward only credential names required by the configured remote media. Give
# checker containers their own scoped credentials in their deployment.
PAIDF_STORAGE_CREDENTIAL_ARGS=()
# S3-compatible example (include only the names your storage config requires):
# PAIDF_STORAGE_CREDENTIAL_ARGS=(-e AWS_ACCESS_KEY_ID -e AWS_SECRET_ACCESS_KEY \
# -e AWS_DEFAULT_REGION -e AWS_ENDPOINT_URL)
docker run -it --rm \
--network paidf \
"${PAIDF_ENDPOINT_KEY_ARGS[@]}" \
"${PAIDF_STORAGE_CREDENTIAL_ARGS[@]}" \
-v "$(pwd)/modules:/workspace/modules" \
-v "$(pwd)/configs:/workspace/configs" \
-v "$(pwd)/data:/workspace/data" \
--entrypoint /bin/bash \
"$PAIDF_IMAGE_ID"Persist the captured sha256: ID and use it directly for every later launch;
never re-resolve the mutable tag as the launch target. For registry releases,
verify the signed manifest and use name:tag@sha256:<manifest-digest>.
paidf bridge for the augmentation container and
every local model, Arbitrator, and checker container. cosmos-evaluator may
be a container/DNS hostname on that bridge; it is not a second bridge name.
Change endpoint URLs from host localhost addresses to container DNS names,
for example http://vlm:8000/v1. The default bridge is sufficient when every
endpoint is remote.-e VAR_NAME; never mount or load a broad credential file.--gpus for data_processing.alignment and any H.264 decode; pick a GPU not shared with a busy model server. Container runs as uid 10000; ensure data/ is writable (or --user "$(id -u):$(id -g)").Warning: Never enable Docker host-network mode on shared, multi-tenant, or production hosts. It is permitted only when all three conditions hold: (1) a required host-local service cannot be moved onto the
paidfbridge, (2) the host is single-tenant and isolated with no other workloads, and (3) the user sends an explicit message in the current conversation choosing host networking after receiving a warning that it removes network namespace isolation. Silence, prior documentation, or an earlier unrelated approval is not acceptance. Do not select it automatically. Review pipeline-operations.md.
# Schema and cross-section preflight only: no endpoint calls or inference.
uv run --no-sync modules/cli.py --config configs/<config_file>.yaml --validate-only
# Run only after validation succeeds.
uv run --no-sync modules/cli.py --config configs/<config_file>.yaml
# With OmegaConf CLI overrides (dot-list syntax)
uv run --no-sync modules/cli.py --config configs/<config_file>.yaml \
data.0.inputs.rgb=/workspace/data/input.mp4 \
augmentation.parameters.seed=42Environment variables: generation/captioning keys resolve as the api_key_env
var → the role's default env var. External Cosmos evaluation reads only the
variable explicitly named by its endpoint and fails initialization when that
variable is unset. Leave api_key_env off for unauthenticated endpoints.
LOG_LEVEL sets logging.
Configs are validated against PipelineConfig (modules/aug_utils/schema/) and have seven top-level sections. Author them in the canonical order above; YAML key order does not change runtime semantics. Full per-section YAML is in configuration-schema.md; runtime flow and common editing tasks are in pipeline-operations.md.
Run all inference and schema validation inside the Docker container for a consistent environment. For config-validation errors, runtime/endpoint errors, and typical per-stage timings, see troubleshooting.md.
torchrun, no executor_type, no Gradio executor.data_processing.alignment (cupy) and by any local stage decoding H.264 because the image ships only the hardware h264_cuvid decoder (software AVC decode is off for licensing). VP9 decodes in software. Video output is VP9-only.data[].inputs.rgb, control media, generated data[].output.video, and configured companion media. Use checker-readable s3://, msc://, gs://, az://, or HTTP(S) URIs; the pipeline does not upload local files solely for evaluation.api_key_env; local endpoints (e.g. vLLM) need none.© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 12 other files (references) in skills/paidf-augmentation of NVIDIA/skills.
Open the folder on GitHubat commit 67a13c0
Paidf Augmentation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Paidf Augmentation this skillNVIDIA/skills | 3.5k | — | ~5.3k | Automated safety check: Pass | Apache-2.0 | |
| Bailian Media Generationmodelstudioai/cli | 542 | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| Aliyun Modelstudio Entrycinience/alicloud-skills | 397 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Qianwen Video GenerationQianWen-AI/qianwen-ai | 104 | — | ~5k | Automated safety check: Notes | Apache-2.0 | |
| Qianwen VisionQianWen-AI/qianwen-ai | 104 | — | ~4.9k | Automated safety check: Notes | Apache-2.0 | |
| Hunyuan VideoVectorSpaceLab/AREX-Skill | 328 | — | ~1.1k | Automated safety check: Pass | Custom licence |
modelstudioai/cli
Chinese-language entry point into Alibaba Cloud Bailian's image, video and speech generation and understanding, routed through separate image, video, speech and vision commands.
cinience/alicloud-skills
A skill your agent uses when routing Alibaba Cloud Model Studio requests to the right local skill (Qwen text, coder, deep research, image, video, audio, search and multimodal skills).
QianWen-AI/qianwen-ai
Generate videos using Wan and HappyHorse models. An agent skill from QianWen-AI/qianwen-ai.
QianWen-AI/qianwen-ai
Understand images and videos with Qwen vision models. An agent skill from QianWen-AI/qianwen-ai.
VectorSpaceLab/AREX-Skill
Use this operating skill for Tencent-Hunyuan/HunyuanVideo text-to-video setup, checkpoint layout, inference commands, Gradio launch, and CUDA/FP8/xDiT troubleshooting.
vllm-project/vllm-omni
Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
A skill your agent uses when authoring or validating PAIDF augmentation YAML configs, or running remote Cosmos Transfer (including Cosmos3 WSM controls), Cosmos Predict, image-edit, or…. Paidf Augmentation is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Use when authoring or validating PAIDF augmentation YAML configs, or running remote Cosmos Transfer (including Cosmos3 WSM controls), Cosmos Predict, image-edit, or image-to-video inference.
Paidf Augmentation fits situations like: validating PAIDF augmentation YAML configs; running remote Cosmos Transfer (including Cosmos3 WSM controls); image-to-video inference.
Run `npx skills add NVIDIA/skills --skill paidf-augmentation -a claude-code`. Or copy the skill folder (skills/paidf-augmentation in NVIDIA/skills) into .claude/skills/paidf-augmentation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill paidf-augmentation -a codex`. Or copy the skill folder (skills/paidf-augmentation in NVIDIA/skills) into .agents/skills/paidf-augmentation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill paidf-augmentation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/paidf-augmentation, .gemini/skills/paidf-augmentation, .github/skills/paidf-augmentation and .opencode/skills/paidf-augmentation in your project.
Going by SKILL.md and its folder, Paidf Augmentation needs the command-line tools its instructions call (docker, git and uv) and credentials named HF_TOKEN, VLM_API_KEY, LLM_API_KEY and VEO_API_KEY. Our summary lists: Docker.
SKILL.md contains no URLs. Its commands use docker, git and uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Paidf Augmentation is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 31k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Paidf Augmentation: Bailian Media Generation (modelstudioai/cli, 542 stars), Aliyun Modelstudio Entry (cinience/alicloud-skills, 397 stars), Qianwen Video Generation (QianWen-AI/qianwen-ai, 104 stars) and Qianwen Vision (QianWen-AI/qianwen-ai, 104 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,539 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.