Agent skill

Aliyun Qwen Multimodal Embedding

by cinience in cinience/alicloud-skills

A skill your agent uses when multimodal embeddings are needed from Alibaba Cloud Model Studio models such as qwen3-vl-embedding for image, video, and text retrieval, cross-modal search, clustering…

MITAuto-check passedAI & LLM Engineering

Install Aliyun Qwen Multimodal Embedding

skills CLI
$ npx skills add cinience/alicloud-skills --skill aliyun-qwen-multimodal-embedding -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cinience/alicloud-skills aliyun-qwen-multimodal-embedding --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cinience/alicloud-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai/search/aliyun-qwen-multimodal-embedding .claude/skills/aliyun-qwen-multimodal-embedding && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
aliyun-qwen-multimodal-embedding
GitHub stars
397
Token cost
~751 tokens
SKILL.md length
249 words
Files
4 (incl. scripts, references)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when multimodal embeddings are needed from Alibaba Cloud Model Studio models such as qwen3-vl-embedding for image, video, and text retrieval, cross-modal search, clustering…

  • Multimodal embeddings are needed from Alibaba Cloud Model Studio models such as qwen3-vl-embedding for image
  • SKILL.md covers Validation, Output And Evidence, Critical model names and Prerequisites, plus 5 more sections
  • Runs Python scripts from its folder; calls python; needs DASHSCOPE_API_KEY
  • Cross-modal search

What it does

Aliyun Qwen Multimodal Embedding is an agent skill from cinience/alicloud-skills. Use when multimodal embeddings are needed from Alibaba Cloud Model Studio models such as qwen3-vl-embedding for image, video, and text retrieval, cross-modal search, clustering, or offline vectorization pipelines.

Its SKILL.md is about 750 tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `agents/openai.yaml`, `references/sources.md` and `scripts/prepare_multimodal_embedding_request.py`).

It sits in AI & LLM Engineering, covering Embeddings. It works with Qwen and Alibaba Cloud. The repository describes itself as: alibaba cloud skills,qwen ,wan and all skills. The licence is MIT.

When your agent uses it

  • Multimodal embeddings are needed from Alibaba Cloud Model Studio models such as qwen3-vl-embedding for image
  • Cross-modal search
  • Offline vectorization pipelines

Example prompts

  • “/aliyun-qwen-multimodal-embedding”

Requirements

  • Python 3
  • A credential in DASHSCOPE_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 1818263. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DASHSCOPE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Aliyun Qwen Multimodal Embedding loads about 751 tokens when it runs, and up to ~817 if it reads all its reference files. Until then it costs about 62 tokens; SKILL.md has 249 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~62
When it runs · the whole SKILL.md, loaded when a task matches
~751
With references · SKILL.md plus every file in references/, read only if the agent opens them
~817

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from cinience/alicloud-skills at commit 1818263, republished under its MIT licence (© cinience). 249 words, ~751 tokens.

Download SKILL.mdSave it as .claude/skills/aliyun-qwen-multimodal-embedding/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
aliyun-qwen-multimodal-embedding
description
Use when multimodal embeddings are needed from Alibaba Cloud Model Studio models such as `qwen3-vl-embedding` for image, video, and text retrieval, cross-modal search, clustering, or offline vectorization pipelines.
version
1.0.0

Category: provider

Model Studio Multimodal Embedding

Validation

bash
mkdir -p output/aliyun-qwen-multimodal-embedding
python -m py_compile skills/ai/search/aliyun-qwen-multimodal-embedding/scripts/prepare_multimodal_embedding_request.py && echo "py_compile_ok" > output/aliyun-qwen-multimodal-embedding/validate.txt

Pass criteria: command exits 0 and output/aliyun-qwen-multimodal-embedding/validate.txt is generated.

Output And Evidence

  • Save normalized request payloads, selected dimensions, and sample input references under output/aliyun-qwen-multimodal-embedding/.
  • Record the exact model, modality mix, and output vector dimension for reproducibility.

Use this skill when the task needs text, image, or video embeddings from Model Studio for retrieval or similarity workflows.

Critical model names

Use one of these exact model strings as needed:

  • qwen3-vl-embedding
  • qwen2.5-vl-embedding
  • tongyi-embedding-vision-plus-2026-03-06

Selection guidance:

  • Prefer qwen3-vl-embedding for the newest multimodal embedding path.
  • Use qwen2.5-vl-embedding when you need compatibility with an older deployed pipeline.

Prerequisites

  • Set DASHSCOPE_API_KEY in your environment, or add dashscope_api_key to ~/.alibabacloud/credentials.
  • Pair this skill with a vector store such as DashVector, OpenSearch, or Milvus when building retrieval systems.

Normalized interface (embedding.multimodal)

Request
  • model (string, optional): default qwen3-vl-embedding
  • texts (array<string>, optional)
  • images (array<string>, optional): public URLs or local paths uploaded by your client layer
  • videos (array<string>, optional): public URLs where supported
  • dimension (int, optional): e.g. 2560, 2048, 1536, 1024, 768, 512, 256 for qwen3-vl-embedding
Response
  • embeddings (array<object>)
  • dimension (int)
  • usage (object, optional)

Quick start

bash
python skills/ai/search/aliyun-qwen-multimodal-embedding/scripts/prepare_multimodal_embedding_request.py \
  --text "A cat sitting on a red chair" \
  --image "https://example.com/cat.jpg" \
  --dimension 1024

Operational guidance

  • Keep input.contents as an array; malformed shapes are a common 400 cause.
  • Pin the output dimension to match your index schema before writing vectors.
  • Use the same model and dimension across one vector index to avoid mixed-vector incompatibility.
  • For large image or video batches, stage files in object storage and reference stable URLs.

Output location

  • Default output: output/aliyun-qwen-multimodal-embedding/request.json
  • Override base dir with OUTPUT_DIR.

References

  • references/sources.md

© cinience, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in skills/ai/search/aliyun-qwen-multimodal-embedding of cinience/alicloud-skills.

  • SKILL.md
  • agents/openai.yaml
  • references/sources.md
  • scripts/prepare_multimodal_embedding_request.py

Open the folder on GitHubat commit 1818263

Compare with similar skills

Aliyun Qwen Multimodal Embedding next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Aliyun Qwen Multimodal Embedding compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Aliyun Qwen Multimodal Embedding this skillcinience/alicloud-skills397—~751Automated safety check: PassMIT
Openrouter EmbeddingsQinghongLin/data2story-skill155—~499Automated safety check: PassMIT
Bailian Media Generationmodelstudioai/cli542—~2kAutomated safety check: PassApache-2.0
Dashscopecalesthio/OpenMontage66k—~1.5kAutomated safety check: NotesAGPL-3.0
AI SDKvercel-labs/ai-facts16820 repos~1.2kAutomated safety check: PassNone
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k7 repos~2.3kAutomated safety check: PassMIT

Similar skills

  • Openrouter Embeddings

    QinghongLin/data2story-skill

    Generate text embeddings via OpenRouter using Qwen3-Embedding-8B.

    155 GitHub stars~499 tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Bailian Media Generation

    modelstudioai/cli

    Chinese-language entry point into Alibaba Cloud Bailian's image, video and speech generation and understanding, routed through separate image, video, speech and vision commands.

    542 GitHub stars~2k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Dashscope

    calesthio/OpenMontage

    DashScope (Alibaba Cloud Bailian / 阿里云百炼) integration — image generation (qwen-image-2.0-pro), text-to-speech (qwen3-tts-flash), and ASR with word-level timestamps (qwen3-asr-flash-filetrans).

    66k GitHub stars~1.5k tokensUpdated 7 days ago
    Media & CreativeAuto-check: notes
  • AI SDK

    vercel-labs/ai-facts

    Official

    Answer questions about the AI SDK and help build AI-powered features.

    168 GitHub starsUsed in 20 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed

More from cinience/alicloud-skills

All 96 skills in this repo
  • Aliyun Skill Creator

    cinience/alicloud-skills

    A skill your agent uses when creating, migrating, or optimizing skills for this alicloud-skills repository.

    397 GitHub stars~2.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Alicloud Acs Agent Sandbox

    cinience/alicloud-skills

    Bootstrap, create, connect to, operate, secure, scale, upgrade, troubleshoot, inspect, and tear down Alibaba Cloud Container Compute Service (ACS) Agent Sandbox environments.

    397 GitHub stars~2.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Alicloud Acs Cluster

    cinience/alicloud-skills

    Create, inspect, connect to, inventory, and delete Alibaba Cloud Container Compute Service (ACS) clusters through the official CS OpenAPI.

    397 GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Aliyun Adb Mysql

    cinience/alicloud-skills

    A skill your agent uses when managing Alibaba Cloud AnalyticDB for MySQL (ADB) via OpenAPI/SDK, including the user needs AnalyticDB resource lifecycle and configuration operations, status checks, or…

    397 GitHub stars~708 tokensUpdated 2 mo ago
    Auto-check passed
  • Aliyun Aicontent Generate

    cinience/alicloud-skills

    A skill your agent uses when managing Alibaba Cloud AIContent (AiContent) via OpenAPI/SDK, including the user needs AI content generation or content workflow operations in Alibaba Cloud, including…

    397 GitHub stars~734 tokensUpdated 2 mo ago
    Auto-check passed
  • Aliyun Aimiaobi Generate

    cinience/alicloud-skills

    A skill your agent uses when managing Alibaba Cloud Quan Miao (AiMiaoBi) via OpenAPI/SDK, including the user asks for Alibaba Cloud MiaoBi content operations, including listing resources…

    397 GitHub stars~724 tokensUpdated 2 mo ago
    Auto-check passed

Questions about Aliyun Qwen Multimodal Embedding

What does Aliyun Qwen Multimodal Embedding do?

A skill your agent uses when multimodal embeddings are needed from Alibaba Cloud Model Studio models such as qwen3-vl-embedding for image, video, and text retrieval, cross-modal search, clustering…. Aliyun Qwen Multimodal Embedding is an agent skill from cinience/alicloud-skills. Use when multimodal embeddings are needed from Alibaba Cloud Model Studio models such as qwen3-vl-embedding for image, video, and text retrieval, cross-modal search, clustering, or offline vectorization pipelines.

When should I use Aliyun Qwen Multimodal Embedding?

Aliyun Qwen Multimodal Embedding fits situations like: multimodal embeddings are needed from Alibaba Cloud Model Studio models such as qwen3-vl-embedding for image; cross-modal search; offline vectorization pipelines.

How do I install Aliyun Qwen Multimodal Embedding in Claude Code?

Run `npx skills add cinience/alicloud-skills --skill aliyun-qwen-multimodal-embedding -a claude-code`. Or copy the skill folder (skills/ai/search/aliyun-qwen-multimodal-embedding in cinience/alicloud-skills) into .claude/skills/aliyun-qwen-multimodal-embedding in your project. Claude Code loads it when a task matches its description.

How do I install Aliyun Qwen Multimodal Embedding in Codex?

Run `npx skills add cinience/alicloud-skills --skill aliyun-qwen-multimodal-embedding -a codex`. Or copy the skill folder (skills/ai/search/aliyun-qwen-multimodal-embedding in cinience/alicloud-skills) into .agents/skills/aliyun-qwen-multimodal-embedding in your project. Codex loads it when a task matches its description.

Can I use Aliyun Qwen Multimodal Embedding in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cinience/alicloud-skills --skill aliyun-qwen-multimodal-embedding -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/aliyun-qwen-multimodal-embedding, .gemini/skills/aliyun-qwen-multimodal-embedding, .github/skills/aliyun-qwen-multimodal-embedding and .opencode/skills/aliyun-qwen-multimodal-embedding in your project.

What does Aliyun Qwen Multimodal Embedding need to run?

Going by SKILL.md and its folder, Aliyun Qwen Multimodal Embedding needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named DASHSCOPE_API_KEY. Our summary lists: Python 3; A credential in DASHSCOPE_API_KEY.

Does Aliyun Qwen Multimodal Embedding access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Aliyun Qwen Multimodal Embedding safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Aliyun Qwen Multimodal Embedding use?

Aliyun Qwen Multimodal Embedding is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Aliyun Qwen Multimodal Embedding use?

About 751 tokens (SKILL.md is roughly 3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 66 tokens, read only when the agent opens those files.

What are the alternatives to Aliyun Qwen Multimodal Embedding?

Skills that share tags, products or a category with Aliyun Qwen Multimodal Embedding: Openrouter Embeddings (QinghongLin/data2story-skill, 155 stars), Bailian Media Generation (modelstudioai/cli, 542 stars), Dashscope (calesthio/OpenMontage, 66k stars) and AI SDK (vercel-labs/ai-facts, 168 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Aliyun Qwen Multimodal Embedding?

cinience (a GitHub user) maintains it in cinience/alicloud-skills, which has 397 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on August 11, 2026.

Source: cinience/alicloud-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.