Agent skill

Dashscope

by calesthio in calesthio/OpenMontage

DashScope (Alibaba Cloud Bailian / 阿里云百炼) integration — image generation (qwen-image-2.0-pro), text-to-speech (qwen3-tts-flash), and ASR with word-level timestamps (qwen3-asr-flash-filetrans).

AGPL-3.0Auto-check: notesMedia & Creative

Install Dashscope

skills CLI
$ npx skills add calesthio/OpenMontage --skill dashscope -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install calesthio/OpenMontage dashscope --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/calesthio/OpenMontage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/dashscope .claude/skills/dashscope && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dashscope
GitHub stars
65k
Token cost
~1.5k tokens
SKILL.md length
496 words
Files
1
Skills in repo
41
Repo updated
First seen
Licence
AGPL-3.0

At a glance

DashScope (Alibaba Cloud Bailian / 阿里云百炼) integration — image generation (qwen-image-2.0-pro), text-to-speech (qwen3-tts-flash), and ASR with word-level timestamps (qwen3-asr-flash-filetrans).

  • Works in 4 steps: Image: Generate a sample first. Check… → TTS: Generate a 10-15 second sample… → ASR: Audio must be at a publicly… → …
  • Generating images via Qwen-Image
  • SKILL.md covers Current API, OpenMontage Usage, Recommended Workflow and Parameters, plus 2 more sections
  • Reaches dashscope.aliyuncs.com; needs DASHSCOPE_API_KEY

What it does

Dashscope is an agent skill from calesthio/OpenMontage. DashScope (Alibaba Cloud Bailian / 阿里云百炼) integration — image generation (qwen-image-2.0-pro), text-to-speech (qwen3-tts-flash), and ASR with word-level timestamps (qwen3-asr-flash-filetrans). Use when generating images via Qwen-Image, narrating via Qwen-TTS, or transcribing with word-level timestamps via Qwen-ASR.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice, Speech recognition and synthesis and Image generation. It works with Qwen and Alibaba Cloud. The repository describes itself as: World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant… The licence is AGPL-3.0.

When your agent uses it

  • Generating images via Qwen-Image
  • Narrating via Qwen-TTS
  • Transcribing with word-level timestamps via Qwen-ASR

Example prompts

  • “/dashscope”

Requirements

  • Python 3
  • A credential in DASHSCOPE_API_KEY

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Image: Generate a sample first. Check prompt_extend: true (default) — DashScope rewrites your prompt for better results. Disable if you…
  2. TTS: Generate a 10-15 second sample before full narration. Approve voice and pacing before committing to full generation.
  3. ASR: Audio must be at a publicly accessible URL. Upload to any public host (S3, etc.) first. Local paths are rejected with a clear error.
  4. Subtitles: Build from result.data["words"] — each word has begin_time_seconds and end_time_seconds. Group words into caption phrases by…

What it can do on your machine

Read from SKILL.md and the folder at commit 9327439. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • dashscope.aliyuncs.com

    Also links to:

    • dashscope.aliyun.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DASHSCOPE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dashscope loads about 1.5k tokens when it runs. Until then it costs about 82 tokens; SKILL.md has 496 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:8
    Requires `DASHSCOPE_API_KEY` in `.env`. Get one at https://dashscope.aliyun.com/.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from calesthio/OpenMontage at commit 9327439, republished under its AGPL-3.0 licence (© calesthio). 496 words, ~1,507 tokens.

Download SKILL.mdSave it as .claude/skills/dashscope/SKILL.md (or your agent's skills folder).
name
dashscope
description
DashScope (Alibaba Cloud Bailian / 阿里云百炼) integration — image generation (qwen-image-2.0-pro), text-to-speech (qwen3-tts-flash), and ASR with word-level timestamps (qwen3-asr-flash-filetrans). Use when generating images via Qwen-Image, narrating via Qwen-TTS, or transcribing with word-level timestamps via Qwen-ASR.

DashScope

Requires DASHSCOPE_API_KEY in .env. Get one at https://dashscope.aliyun.com/.

Current API

CRITICAL: DashScope's /compatible-mode/v1/ only supports /chat/completions and /embeddings. Image generation, TTS, and ASR all use DashScope-native endpoints — not OpenAI-compatible paths.

All three tools use Authorization: Bearer $DASHSCOPE_API_KEY.

Image Generation
text
POST https://dashscope.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation
  • Model: qwen-image-2.0-pro (default), qwen-image-max, wan2.7-image, z-image-turbo
  • Body: {model, input: {messages: [{role: "user", content: [{text: "prompt"}]}]}, parameters: {size: "W*H", n, prompt_extend, watermark}}
  • Size format uses asterisk: "1024*1024" not "1024x1024"
  • Response: output.choices[0].message.content[0].image (URL, valid ~24h) — must download separately
Text-to-Speech
text
POST https://dashscope.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation

Same endpoint as image gen, different body.

  • Model: qwen3-tts-flash (default), qwen3-tts-instruct-flash, qwen-tts-2025-05-22
  • Body: {model, input: {text, voice: "Cherry", language_type: "Auto"}}
  • Response: output.audio.url (WAV, valid ~24h) — must download separately
ASR with Word-Level Timestamps
text
POST https://dashscope.aliyuncs.com/api/v1/services/audio/asr/transcription
Header: X-DashScope-Async: enable
  • Model: qwen3-asr-flash-filetrans (NOT qwen3-asr-flash — the sync version has no word timestamps)
  • Body: {model, input: {file_url: "https://public-url/audio.mp3"}, parameters: {enable_words: true, language_hints: ["zh","en"]}}
  • Returns task_id → poll GET /api/v1/tasks/{task_id} until SUCCEEDED → download output.result.transcription_url → JSON with transcripts[].sentences[].words[]
  • Timestamps in begin_time/end_time are in milliseconds — the tool normalizes to seconds

OpenMontage Usage

Image via selector
python
from tools.graphics.image_selector import ImageSelector

result = ImageSelector().execute({
    "preferred_provider": "dashscope",
    "prompt": "一只猫坐在沙发上",
    "output_path": "projects/my-video/assets/images/cat.png",
})
TTS via selector
python
from tools.audio.tts_selector import TTSSelector

result = TTSSelector().execute({
    "preferred_provider": "dashscope",
    "text": "如果 AI 真的会改变未来,普通人到底该怎么参与?",
    "voice": "Cherry",
    "output_path": "projects/my-video/assets/audio/narration.wav",
})
ASR directly (word timestamps for subtitles)
python
from tools.analysis.dashscope_asr import DashscopeAsr

result = DashscopeAsr().execute({
    "audio_url": "https://example.com/narration.wav",
    "output_path": "projects/my-video/assets/audio/transcription.json",
})

# result.data["words"] is a flat list of {text, begin_time_seconds, end_time_seconds}
  1. Image: Generate a sample first. Check prompt_extend: true (default) — DashScope rewrites your prompt for better results. Disable if you need literal prompt adherence.
  2. TTS: Generate a 10-15 second sample before full narration. Approve voice and pacing before committing to full generation.
  3. ASR: Audio must be at a publicly accessible URL. Upload to any public host (S3, etc.) first. Local paths are rejected with a clear error.
  4. Subtitles: Build from result.data["words"] — each word has begin_time_seconds and end_time_seconds. Group words into caption phrases by language semantics, not fixed character count.

Parameters

Image (dashscope_image)
  • prompt (required): text prompt
  • model: default qwen-image-2.0-pro
  • size: default "1024*1024" — asterisk separator, not "x"
  • n: 1-6 images
  • negative_prompt: things to avoid (max 500 chars)
  • prompt_extend: default true — auto-rewrite prompt for better results
  • watermark: default false
  • seed: for reproducibility
Show full SKILL.md (194 more words)Show less
TTS (dashscope_tts)
  • text (required): text to synthesize (max 600 chars for qwen3-tts-flash)
  • model: default qwen3-tts-flash
  • voice: default "Cherry" — other voices: "Ethan", "Chelsie", etc.
  • language_type: default "Auto" — "Chinese", "English", "Japanese", "Korean"
  • instructions: natural language delivery instructions (only for qwen3-tts-instruct-flash)
ASR (dashscope_asr)
  • audio_url (required): must be publicly accessible URL
  • model: qwen3-asr-flash-filetrans (only model that supports word timestamps)
  • language_hints: default ["zh", "en"]
  • enable_words: default true — required for word-level timestamps
  • poll_interval_seconds: default 5.0
  • timeout_seconds: default 300

Troubleshooting

  • Image size error: Use "W*H" with asterisk, not "WxH". Example: "2048*2048".
  • TTS no audio URL: Check output.audio.url — if empty, the model name or voice may be wrong.
  • ASR "file not accessible": audio_url must be publicly reachable. DashScope servers fetch the file; local paths and auth-gated URLs don't work.
  • ASR poll timeout: Increase timeout_seconds (default 300). Long audio files take longer to transcribe.
  • ASR no word timestamps: Ensure enable_words: true and model is qwen3-asr-flash-filetrans (not the sync qwen3-asr-flash).
  • Auth error (401): Verify DASHSCOPE_API_KEY is set. Use Authorization: Bearer $KEY header.

Safety

Never print or write the API key to logs, metadata, patches, or project artifacts. .env.example should contain only empty variable names. The tool's _safe_error() method redacts the key from error messages.

© calesthio, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/dashscope of calesthio/OpenMontage.

Open the folder on GitHubat commit 9327439

Compare with similar skills

Dashscope next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dashscope compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dashscope this skillcalesthio/OpenMontage65k—~1.5kAutomated safety check: NotesAGPL-3.0
Bailian Media Generationmodelstudioai/cli542—~2kAutomated safety check: PassApache-2.0
Video Productionspeechlab0210/video-production-skill105—~4.1kAutomated safety check: NotesMIT
Local AI Useamd/skills398—~5kAutomated safety check: NotesMIT
Local AI App Integrationamd/skills398—~6kAutomated safety check: PassMIT
Aliyun Qwen Imagecinience/alicloud-skills397—~1.8kAutomated safety check: PassMIT

Similar skills

  • Bailian Media Generation

    modelstudioai/cli

    Chinese-language entry point into Alibaba Cloud Bailian's image, video and speech generation and understanding, routed through separate image, video, speech and vision commands.

    542 GitHub stars~2k tokensUpdated 8 days ago
    Media & CreativeAuto-check passed
  • Video Production

    speechlab0210/video-production-skill

    AI educational video production pipeline. An agent skill from speechlab0210/video-production-skill.

    105 GitHub stars~4.1k tokensUpdated 3 mo ago
    Media & CreativeAuto-check: notes
  • Local AI Use

    amd/skills

    Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.

    398 GitHub stars~5k tokensUpdated today
    Media & CreativeAuto-check: notes
  • Integrates local AI capabilities into applications using Embeddable Lemonade.

    398 GitHub stars~6k tokensUpdated today
    Media & CreativeAuto-check passed
  • Aliyun Qwen Image

    cinience/alicloud-skills

    A skill your agent uses when generating images with Model Studio DashScope SDK using Qwen Image generation models (qwen-image, qwen-image-plus, qwen-image-max, qwen-image-2.0 series and snapshots).

    397 GitHub stars~1.8k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Aliyun Qwen Tts

    cinience/alicloud-skills

    A skill your agent uses when generating human-like speech audio with Model Studio DashScope Qwen TTS models (qwen3-tts-flash, qwen3-tts-instruct-flash).

    397 GitHub stars~951 tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed

More from calesthio/OpenMontage

All 41 skills in this repo
  • Video Understand

    calesthio/OpenMontage

    Understand video content locally using ffmpeg frame extraction and Whisper transcription.

    65k GitHub stars~841 tokensUpdated 4 days ago
    Auto-check passed
  • Avatar Video

    calesthio/OpenMontage

    Create AI avatar videos with precise control over avatars, voices, scripts, scenes, and backgrounds using HeyGen's v2 API.

    65k GitHub stars~1.6k tokensUpdated 4 days ago
    Auto-check passed
  • D3 Viz

    calesthio/OpenMontage

    Creating interactive data visualisations using d3.js. An agent skill from calesthio/OpenMontage.

    65k GitHub starsUsed in 3 repos~5.4k tokens
    Auto-check passed
  • Create Video

    calesthio/OpenMontage

    Create videos from a text prompt using HeyGen's Video Agent.

    65k GitHub stars~1.3k tokensUpdated 4 days ago
    Auto-check passed
  • Threejs World Generation

    calesthio/OpenMontage

    Build deterministic, editable, free-viewpoint Three.js worlds from text or structured briefs.

    65k GitHub stars~2k tokensUpdated 4 days ago
    Auto-check passed
  • Video Edit

    calesthio/OpenMontage

    Edit videos locally using ffmpeg. An agent skill from calesthio/OpenMontage.

    65k GitHub stars~855 tokensUpdated 4 days ago
    Auto-check: notes

Questions about Dashscope

What does Dashscope do?

DashScope (Alibaba Cloud Bailian / 阿里云百炼) integration — image generation (qwen-image-2.0-pro), text-to-speech (qwen3-tts-flash), and ASR with word-level timestamps (qwen3-asr-flash-filetrans). Dashscope is an agent skill from calesthio/OpenMontage.0-pro), text-to-speech (qwen3-tts-flash), and ASR with word-level timestamps (qwen3-asr-flash-filetrans).

When should I use Dashscope?

Dashscope fits situations like: generating images via Qwen-Image; narrating via Qwen-TTS; transcribing with word-level timestamps via Qwen-ASR.

How do I install Dashscope in Claude Code?

Run `npx skills add calesthio/OpenMontage --skill dashscope -a claude-code`. Or copy the skill folder (.agents/skills/dashscope in calesthio/OpenMontage) into .claude/skills/dashscope in your project. Claude Code loads it when a task matches its description.

How do I install Dashscope in Codex?

Run `npx skills add calesthio/OpenMontage --skill dashscope -a codex`. Or copy the skill folder (.agents/skills/dashscope in calesthio/OpenMontage) into .agents/skills/dashscope in your project. Codex loads it when a task matches its description.

Can I use Dashscope in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calesthio/OpenMontage --skill dashscope -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dashscope, .gemini/skills/dashscope, .github/skills/dashscope and .opencode/skills/dashscope in your project.

What does Dashscope need to run?

Going by SKILL.md and its folder, Dashscope needs credentials named DASHSCOPE_API_KEY. Our summary lists: Python 3; A credential in DASHSCOPE_API_KEY.

Does Dashscope access the network?

SKILL.md names 2 domains. In commands or code: dashscope.aliyuncs.com; the agent is likely to contact it when it follows the instructions. As links in the text: dashscope.aliyun.com. This is read from the text; nothing was executed.

Is Dashscope safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Dashscope use?

Dashscope is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dashscope use?

About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Dashscope?

Skills that share tags, products or a category with Dashscope: Bailian Media Generation (modelstudioai/cli, 542 stars), Video Production (speechlab0210/video-production-skill, 105 stars), Local AI Use (amd/skills, 398 stars) and Local AI App Integration (amd/skills, 398 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dashscope?

calesthio (a GitHub user) maintains it in calesthio/OpenMontage, which has 65,101 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 3, 2026.

Source: calesthio/OpenMontage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.