Agent skill

Aliyun Qwen OCR

by cinience in cinience/alicloud-skills

A skill your agent uses when OCR-specialized extraction is needed with Alibaba Cloud Model Studio Qwen OCR models (qwen-vl-ocr, qwen-vl-ocr-latest, and snapshots), including document parsing, table…

MITAuto-check passedDocuments & Office

Install Aliyun Qwen OCR

skills CLI
$ npx skills add cinience/alicloud-skills --skill aliyun-qwen-ocr -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cinience/alicloud-skills aliyun-qwen-ocr --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cinience/alicloud-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai/multimodal/aliyun-qwen-ocr .claude/skills/aliyun-qwen-ocr && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
aliyun-qwen-ocr
GitHub stars
397
Token cost
~953 tokens
SKILL.md length
309 words
Files
5 (incl. scripts, references)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when OCR-specialized extraction is needed with Alibaba Cloud Model Studio Qwen OCR models (qwen-vl-ocr, qwen-vl-ocr-latest, and snapshots), including document parsing, table…

  • OCR-specialized extraction is needed with Alibaba Cloud Model Studio Qwen OCR models (qwen-vl-ocr
  • SKILL.md covers Validation, Output And Evidence, Critical model names and Prerequisites, plus 6 more sections
  • Runs Python scripts from its folder; calls python and python3; needs DASHSCOPE_API_KEY
  • Qwen-vl-ocr-latest

What it does

Aliyun Qwen OCR is an agent skill from cinience/alicloud-skills. Use when OCR-specialized extraction is needed with Alibaba Cloud Model Studio Qwen OCR models (qwen-vl-ocr, qwen-vl-ocr-latest, and snapshots), including document parsing, table parsing, multilingual OCR, formula recognition, and key information extraction.

Its SKILL.md is about 950 tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `agents/openai.yaml`, `references/api_reference.md` and `references/sources.md`).

It sits in Documents & Office, covering Document parsing. It works with Qwen and Alibaba Cloud. The repository describes itself as: alibaba cloud skills,qwen ,wan and all skills. The licence is MIT.

When your agent uses it

  • OCR-specialized extraction is needed with Alibaba Cloud Model Studio Qwen OCR models (qwen-vl-ocr
  • Qwen-vl-ocr-latest
  • Including document parsing
  • Multilingual OCR

Example prompts

  • “/aliyun-qwen-ocr”

Requirements

  • Python 3
  • A credential in DASHSCOPE_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 1818263. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DASHSCOPE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Aliyun Qwen OCR loads about 953 tokens when it runs, and up to ~1.4k if it reads all its reference files. Until then it costs about 69 tokens; SKILL.md has 309 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~69
When it runs · the whole SKILL.md, loaded when a task matches
~953
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from cinience/alicloud-skills at commit 1818263, republished under its MIT licence (© cinience). 309 words, ~953 tokens.

Download SKILL.mdSave it as .claude/skills/aliyun-qwen-ocr/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
aliyun-qwen-ocr
description
Use when OCR-specialized extraction is needed with Alibaba Cloud Model Studio Qwen OCR models (`qwen-vl-ocr`, `qwen-vl-ocr-latest`, and snapshots), including document parsing, table parsing, multilingual OCR, formula recognition, and key information extraction.
version
1.0.0

Category: provider

Model Studio Qwen OCR

Validation

bash
mkdir -p output/aliyun-qwen-ocr
python -m py_compile skills/ai/multimodal/aliyun-qwen-ocr/scripts/prepare_ocr_request.py && echo "py_compile_ok" > output/aliyun-qwen-ocr/validate.txt

Pass criteria: command exits 0 and output/aliyun-qwen-ocr/validate.txt is generated.

Output And Evidence

  • Save request payloads, selected OCR task name, and normalized output expectations under output/aliyun-qwen-ocr/.
  • Keep the exact model, image source, and task configuration with each saved run.

Use Qwen OCR when the task is primarily text extraction or document structure parsing rather than broad visual reasoning.

Critical model names

Use one of these exact model strings:

  • qwen-vl-ocr
  • qwen-vl-ocr-latest
  • qwen-vl-ocr-2025-11-20
  • qwen-vl-ocr-2025-08-28
  • qwen-vl-ocr-2025-04-13
  • qwen-vl-ocr-2024-10-28

Selection guidance:

  • Use qwen-vl-ocr for the stable channel.
  • Use qwen-vl-ocr-latest only when you explicitly want the newest OCR behavior.
  • Pin qwen-vl-ocr-2025-11-20 when you need reproducible document parsing based on the Qwen3-VL OCR upgrade.

Prerequisites

  • Install dependencies (recommended in a venv):
bash
python3 -m venv .venv
. .venv/bin/activate
python -m pip install requests
  • Set DASHSCOPE_API_KEY in environment, or add dashscope_api_key to ~/.alibabacloud/credentials.

Normalized interface (ocr.extract)

Request
  • image (string, required): HTTPS URL, local path, or data: URL.
  • model (string, optional): default qwen-vl-ocr.
  • prompt (string, optional): use when you want custom extraction instructions.
  • task (string, optional): built-in OCR task.
  • task_config (object, optional): configuration for built-in task such as extraction fields.
  • enable_rotate (bool, optional): default false.
  • min_pixels (int, optional)
  • max_pixels (int, optional)
  • max_tokens (int, optional)
  • temperature (float, optional): recommended to keep near default/low values.
Response
  • text (string): extracted text or structured markdown/html-style output.
  • model (string)
  • usage (object, optional)

Built-in OCR tasks

Use one of these values in task:

  • text_recognition
  • key_information_extraction
  • document_parsing
  • table_parsing
  • formula_recognition
  • multi_lan
  • advanced_recognition

Quick start

Custom prompt:

bash
python skills/ai/multimodal/aliyun-qwen-ocr/scripts/prepare_ocr_request.py \
  --image "https://example.com/invoice.png" \
  --prompt "Extract seller name, invoice date, amount, and tax number in JSON."

Built-in task:

bash
python skills/ai/multimodal/aliyun-qwen-ocr/scripts/prepare_ocr_request.py \
  --image "https://example.com/table.png" \
  --task table_parsing \
  --model qwen-vl-ocr-2025-11-20

Operational guidance

  • Prefer built-in OCR tasks for standard parsing jobs because they use official task prompts.
  • For critical business fields, add downstream validation rules after OCR.
  • qwen-vl-ocr and older snapshots default to 4096 max output tokens unless higher limits are approved by Alibaba Cloud; qwen-vl-ocr-2025-11-20 follows the model maximum.
  • Increase max_pixels only when small text is missed; this raises token cost.

Output location

  • Default output: output/aliyun-qwen-ocr/request.json
  • Override base dir with OUTPUT_DIR.

References

  • references/api_reference.md
  • references/sources.md

© cinience, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in skills/ai/multimodal/aliyun-qwen-ocr of cinience/alicloud-skills.

  • SKILL.md
  • agents/openai.yaml
  • references/api_reference.md
  • references/sources.md
  • scripts/prepare_ocr_request.py

Open the folder on GitHubat commit 1818263

Compare with similar skills

Aliyun Qwen OCR next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Aliyun Qwen OCR compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Aliyun Qwen OCR this skillcinience/alicloud-skills397—~953Automated safety check: PassMIT
Bailian Media Generationmodelstudioai/cli542—~2kAutomated safety check: PassApache-2.0
Dashscopecalesthio/OpenMontage66k—~1.5kAutomated safety check: NotesAGPL-3.0
DOCX ToolkitXiaomiMiMo/MiMo-Code14k—~2.4kAutomated safety check: PassApache-2.0
Office File Processxstongxue/best-skills3k—~1.8kAutomated safety check: PassProprietary
PDF ToolkitXiaomiMiMo/MiMo-Code14k—~1.7kAutomated safety check: PassApache-2.0

Similar skills

  • Bailian Media Generation

    modelstudioai/cli

    Chinese-language entry point into Alibaba Cloud Bailian's image, video and speech generation and understanding, routed through separate image, video, speech and vision commands.

    542 GitHub stars~2k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Dashscope

    calesthio/OpenMontage

    DashScope (Alibaba Cloud Bailian / 阿里云百炼) integration — image generation (qwen-image-2.0-pro), text-to-speech (qwen3-tts-flash), and ASR with word-level timestamps (qwen3-asr-flash-filetrans).

    66k GitHub stars~1.5k tokensUpdated 7 days ago
    Media & CreativeAuto-check: notes
  • DOCX Toolkit

    XiaomiMiMo/MiMo-Code

    Produces, edits and reads Microsoft Word files through python-docx and lxml, with a decision table for picking the lightest workflow for a given task.

    14k GitHub stars~2.4k tokensUpdated 2 days ago
    Documents & OfficeAuto-check passed
  • Office File Process

    xstongxue/best-skills

    处理 Office 文档的一站式 skill:Word(.doc/.docx/.dotx)、Excel(.xls/.xlsx/.xlsm/.csv)、PowerPoint(.ppt/.pptx/.potx) 的创建、读取、编辑、提取、转换、校验。触发:『读取 word 文档』『提取 excel 内容』『看 ppt 讲了什么』、.doc 老格式打不开、生成/编辑 Word…

    3k GitHub stars~1.8k tokensUpdated 28 days ago
    Documents & OfficeAuto-check passed
  • PDF Toolkit

    XiaomiMiMo/MiMo-Code

    Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling.

    14k GitHub stars~1.7k tokensUpdated 2 days ago
    Documents & OfficeAuto-check passed
  • Pullmd

    AeternaLabsHQ/pullmd

    Read any web page, document, or YouTube video as clean Markdown using PullMD.

    486 GitHub stars~2.6k tokensUpdated yesterday
    Documents & OfficeAuto-check passed

More from cinience/alicloud-skills

All 96 skills in this repo
  • Aliyun Skill Creator

    cinience/alicloud-skills

    A skill your agent uses when creating, migrating, or optimizing skills for this alicloud-skills repository.

    397 GitHub stars~2.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Alicloud Acs Agent Sandbox

    cinience/alicloud-skills

    Bootstrap, create, connect to, operate, secure, scale, upgrade, troubleshoot, inspect, and tear down Alibaba Cloud Container Compute Service (ACS) Agent Sandbox environments.

    397 GitHub stars~2.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Alicloud Acs Cluster

    cinience/alicloud-skills

    Create, inspect, connect to, inventory, and delete Alibaba Cloud Container Compute Service (ACS) clusters through the official CS OpenAPI.

    397 GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Aliyun Adb Mysql

    cinience/alicloud-skills

    A skill your agent uses when managing Alibaba Cloud AnalyticDB for MySQL (ADB) via OpenAPI/SDK, including the user needs AnalyticDB resource lifecycle and configuration operations, status checks, or…

    397 GitHub stars~708 tokensUpdated 2 mo ago
    Auto-check passed
  • Aliyun Aicontent Generate

    cinience/alicloud-skills

    A skill your agent uses when managing Alibaba Cloud AIContent (AiContent) via OpenAPI/SDK, including the user needs AI content generation or content workflow operations in Alibaba Cloud, including…

    397 GitHub stars~734 tokensUpdated 2 mo ago
    Auto-check passed
  • Aliyun Aimiaobi Generate

    cinience/alicloud-skills

    A skill your agent uses when managing Alibaba Cloud Quan Miao (AiMiaoBi) via OpenAPI/SDK, including the user asks for Alibaba Cloud MiaoBi content operations, including listing resources…

    397 GitHub stars~724 tokensUpdated 2 mo ago
    Auto-check passed

Questions about Aliyun Qwen OCR

What does Aliyun Qwen OCR do?

A skill your agent uses when OCR-specialized extraction is needed with Alibaba Cloud Model Studio Qwen OCR models (qwen-vl-ocr, qwen-vl-ocr-latest, and snapshots), including document parsing, table…. Aliyun Qwen OCR is an agent skill from cinience/alicloud-skills. Use when OCR-specialized extraction is needed with Alibaba Cloud Model Studio Qwen OCR models (qwen-vl-ocr, qwen-vl-ocr-latest, and snapshots), including document parsing, table parsing, multilingual OCR, formula recognition, and key information extraction.

When should I use Aliyun Qwen OCR?

Aliyun Qwen OCR fits situations like: OCR-specialized extraction is needed with Alibaba Cloud Model Studio Qwen OCR models (qwen-vl-ocr; qwen-vl-ocr-latest; including document parsing; multilingual OCR.

How do I install Aliyun Qwen OCR in Claude Code?

Run `npx skills add cinience/alicloud-skills --skill aliyun-qwen-ocr -a claude-code`. Or copy the skill folder (skills/ai/multimodal/aliyun-qwen-ocr in cinience/alicloud-skills) into .claude/skills/aliyun-qwen-ocr in your project. Claude Code loads it when a task matches its description.

How do I install Aliyun Qwen OCR in Codex?

Run `npx skills add cinience/alicloud-skills --skill aliyun-qwen-ocr -a codex`. Or copy the skill folder (skills/ai/multimodal/aliyun-qwen-ocr in cinience/alicloud-skills) into .agents/skills/aliyun-qwen-ocr in your project. Codex loads it when a task matches its description.

Can I use Aliyun Qwen OCR in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cinience/alicloud-skills --skill aliyun-qwen-ocr -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/aliyun-qwen-ocr, .gemini/skills/aliyun-qwen-ocr, .github/skills/aliyun-qwen-ocr and .opencode/skills/aliyun-qwen-ocr in your project.

What does Aliyun Qwen OCR need to run?

Going by SKILL.md and its folder, Aliyun Qwen OCR needs Python for the scripts in its folder, the command-line tools its instructions call (python and python3) and credentials named DASHSCOPE_API_KEY. Our summary lists: Python 3; A credential in DASHSCOPE_API_KEY.

Does Aliyun Qwen OCR access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Aliyun Qwen OCR safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Aliyun Qwen OCR use?

Aliyun Qwen OCR is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Aliyun Qwen OCR use?

About 953 tokens (SKILL.md is roughly 3.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 435 tokens, read only when the agent opens those files.

What are the alternatives to Aliyun Qwen OCR?

Skills that share tags, products or a category with Aliyun Qwen OCR: Bailian Media Generation (modelstudioai/cli, 542 stars), Dashscope (calesthio/OpenMontage, 66k stars), DOCX Toolkit (XiaomiMiMo/MiMo-Code, 14k stars) and Office File Process (xstongxue/best-skills, 3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Aliyun Qwen OCR?

cinience (a GitHub user) maintains it in cinience/alicloud-skills, which has 397 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on August 11, 2026.

Source: cinience/alicloud-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.