Agent skill

Wps Doc Scraper

by daymade in daymade/claude-code-skills

Faithfully archive public WPS/KDocs/金山文档 links, especially embedded ProcessOn .pof mind maps and canvases, as raw source data, original SVG/PNG, and Markdown.

MITAuto-check passedData & Analytics

Install Wps Doc Scraper

skills CLI
$ npx skills add daymade/claude-code-skills --skill wps-doc-scraper -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install daymade/claude-code-skills wps-doc-scraper --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/daymade/claude-code-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/wps-doc-scraper .claude/skills/wps-doc-scraper && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
wps-doc-scraper
GitHub stars
1.4k
Token cost
~1.1k tokens
SKILL.md length
480 words
Files
8 (incl. scripts, references)
Skills in repo
103
Repo updated
First seen
Licence
MIT

At a glance

Faithfully archive public WPS/KDocs/金山文档 links, especially embedded ProcessOn .pof mind maps and canvases, as raw source data, original SVG/PNG, and Markdown.

  • Works in 4 steps: Identify the URL type. → For ProcessOn .pof mind maps or… → Capture the original image. → …
  • A user gives a kdocs.cn
  • SKILL.md covers Overview, Workflow Decision Tree, Hard Rules and Bundled Resources
  • Runs Python scripts from its folder; calls python3; reaches kdocs.cn

What it does

Wps Doc Scraper is an agent skill from daymade/claude-code-skills. Faithfully archive public WPS/KDocs/金山文档 links, especially embedded ProcessOn .pof mind maps and canvases, as raw source data, original SVG/PNG, and Markdown. Use when a user gives a kdocs.cn or wps.processon.com link and asks to scrape, save, download, 扒下来, 归档, or 转 Markdown without logging in or saving the document to an account.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `agents/openai.yaml`, `references/capture-manifest.md` and `references/permission-and-failure-boundaries.md`).

It sits in Data & Analytics, covering Web scraping. The repository describes itself as: Professional Claude Code skills marketplace featuring production-ready skills for enhanced development workflows. The licence is MIT.

When your agent uses it

  • A user gives a kdocs.cn
  • Wps.processon.com link and asks to scrape
  • 转 Markdown without logging in
  • Saving the document to an account

Example prompts

  • “/wps-doc-scraper”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Identify the URL type.
  2. For ProcessOn .pof mind maps or canvases, run the API extractor
  3. Capture the original image.
  4. Validate the archive.

What it can do on your machine

Read from SKILL.md and the folder at commit 91bed2b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • kdocs.cn

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Wps Doc Scraper loads about 1.1k tokens when it runs, and up to ~2.7k if it reads all its reference files. Until then it costs about 87 tokens; SKILL.md has 480 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~87
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from daymade/claude-code-skills at commit 91bed2b, republished under its MIT licence (© daymade). 480 words, ~1,093 tokens.

Download SKILL.mdSave it as .claude/skills/wps-doc-scraper/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
wps-doc-scraper
description
Faithfully archive public WPS/KDocs/金山文档 links, especially embedded ProcessOn .pof mind maps and canvases, as raw source data, original SVG/PNG, and Markdown. Use when a user gives a kdocs.cn or wps.processon.com link and asks to scrape, save, download, 扒下来, 归档, or 转 Markdown without logging in or saving the document to an account.

WPS Doc Scraper

Overview

Archive WPS/KDocs public documents with source fidelity. Prefer unauthenticated data APIs, keep the raw payloads, capture the original visual artifact when the document is a canvas or mind map, and generate Markdown only as a structured representation of the source.

This is an extraction skill, not a writing skill. Do not use an LLM to rewrite, summarize, smooth, or infer missing document content unless the user explicitly asks for a separate analysis after the archive is complete.

Workflow Decision Tree

  1. Identify the URL type.

    • kdocs.cn/view/l/<share_id> or kdocs.cn/l/<share_id>: read public link metadata first.
    • wps.processon.com/diagrams/view?...: extract file_id and group_id, then use the ProcessOn data API.
    • wps.processon.com/wpsapi/diagrams/view/api?...: treat as the source API directly.
    • Other WPS/KDocs pages: collect public metadata, try official public export/download only when available, then use browser DOM capture as a fallback.
  2. For ProcessOn .pof mind maps or canvases, run the API extractor:

    bash
    python3 /path/to/wps-doc-scraper/scripts/wps_processon_extract.py \
      --url "https://www.kdocs.cn/view/l/..." \
      --output-dir "/path/to/archive-dir"

    Expected outputs: processon-api.json, processon-definition.json, capture-manifest.json, and <title>.md.

  3. Capture the original image.

    • If the page exposes a rendered SVG, save the serialized full SVG as <title>-全画布.svg.
    • If the SVG uses foreignObject, do not trust ImageMagick alone for text rendering. Use scripts/render_svg_tiles.py to make the PNG through macOS Quick Look square tiles.
    bash
    python3 /path/to/wps-doc-scraper/scripts/render_svg_tiles.py \
      --svg "/path/to/<title>-全画布.svg" \
      --output "/path/to/<title>-全画布.png"
  4. Validate the archive.

    • Markdown is a structural conversion, not an editorial rewrite.
    • The original image exists for visual documents.
    • Raw JSON and manifest exist.
    • No output text contains the Unicode replacement character �.
    • Any failed or permission-limited path is recorded in the manifest or final response.
Show full SKILL.md (235 more words)Show less

Hard Rules

  • Never log in, save to an account, copy into the user's cloud drive, or mutate the remote document unless the user explicitly requests it.
  • Never bypass CAPTCHAs, paywalls, tenant restrictions, or permission walls. A public link that requires login is a hard boundary unless the user provides authorized access.
  • Do not infer hidden nodes or missing text. Preserve what the API or rendered DOM actually exposes.
  • HTTP 200 is not enough. Detect login-wall payloads such as 用户未登录, empty definitions, and placeholder shells.
  • Keep raw source artifacts before transformation: API JSON, parsed definition JSON, serialized SVG, screenshots or PNGs, and a capture manifest.
  • For browser fallback, capture DOM/source data before screenshots when possible; screenshots alone are not a faithful text archive.
  • Report extraction gaps plainly. Do not silently produce a polished Markdown file from partial data.

Bundled Resources

scripts/
  • wps_processon_extract.py: deterministic extractor for public WPS/KDocs ProcessOn .pof mind maps. It resolves KDocs share metadata, downloads the ProcessOn data API JSON, parses the embedded definition, and writes Markdown plus a manifest.
  • render_svg_tiles.py: macOS Quick Look based SVG-to-PNG renderer for full-canvas SVGs with foreignObject text. It renders square vertical tiles and stitches them into one PNG.
references/
  • processon-mindmap-api.md: endpoint pattern and payload shape for WPS-hosted ProcessOn files.
  • rendered-svg-capture.md: browser-side process for extracting the rendered full-canvas SVG and producing a PNG.
  • capture-manifest.md: minimum manifest fields and acceptance checks.
  • permission-and-failure-boundaries.md: login walls, forbidden escalation, and failure reporting rules.

© daymade, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references) in wps-doc-scraper of daymade/claude-code-skills.

  • SKILL.md
  • agents/openai.yaml
  • references/capture-manifest.md
  • references/permission-and-failure-boundaries.md
  • references/processon-mindmap-api.md
  • references/rendered-svg-capture.md
  • scripts/render_svg_tiles.py
  • scripts/wps_processon_extract.py

Open the folder on GitHubat commit 91bed2b

Compare with similar skills

Wps Doc Scraper next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Wps Doc Scraper compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Wps Doc Scraper this skilldaymade/claude-code-skills1.4k—~1.1kAutomated safety check: PassMIT
Tmuxtrpc-group/trpc-agent-go1.8k24 repos~868Automated safety check: PassApache-2.0
Ketch1broseidon/ketch6971 repos~3.9kAutomated safety check: PassMIT
Crawl4AI Web Scrapingsmallnest/goclaw5981 repos~2.5kAutomated safety check: PassMIT
Boss Zhipin Scrapereatmoreduck/boss-zhipin-scraper1.5k—~2.6kAutomated safety check: PassMIT
Axyusukebe/ax7191 repos~918Automated safety check: PassMIT

Similar skills

  • Tmux

    trpc-group/trpc-agent-go

    Remote-control tmux sessions for interactive CLIs by sending keystrokes and scraping pane output.

    1.8k GitHub starsUsed in 24 repos~868 tokens
    Data & AnalyticsAuto-check passed
  • Ketch

    1broseidon/ketch

    Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but…

    697 GitHub starsUsed in 1 repo~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Crawl4AI Web Scraping

    smallnest/goclaw

    Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.

    598 GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed
  • Boss Zhipin Scraper

    eatmoreduck/boss-zhipin-scraper

    Scrape BOSS直聘 (job listing site) via Chrome CDP. An agent skill from eatmoreduck/boss-zhipin-scraper.

    1.5k GitHub stars~2.6k tokensUpdated 9 days ago
    Data & AnalyticsAuto-check passed
  • Ax

    yusukebe/ax

    Use the ax CLI instead of curl + throwaway parsing scripts whenever you fetch a URL, explore an unknown web page, or extract structured data from HTML.

    719 GitHub starsUsed in 1 repo~918 tokens
    Data & AnalyticsAuto-check passed
  • Anakinscraper

    Anakin-Inc/anakin

    Scrape any website into clean markdown or structured JSON. An agent skill from Anakin-Inc/anakin.

    4.5k GitHub stars~859 tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed

More from daymade/claude-code-skills

All 103 skills in this repo
  • Video Comparer

    daymade/claude-code-skills

    This skill should be used when comparing two videos to analyze compression results or quality differences.

    1.4k GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check: notes
  • CLI Demo Generator

    daymade/claude-code-skills

    Generates professional animated CLI demos as GIFs using VHS terminal recordings.

    1.4k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Doc To Markdown

    daymade/claude-code-skills

    Converts DOCX/PDF/PPTX and saved HTML/HTM to high-quality Markdown with automatic post-processing.

    1.4k GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Interaction Design Board

    daymade/claude-code-skills

    Generates several distinct, clickable HTML interaction prototypes for one product surface into a Design Board and collects selection/remix feedback before implementation.

    1.4k GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Auto Repo Setup

    daymade/claude-code-skills

    Diagnoses and repairs repository setup and guarded Git workflows for Claude Code or Codex — environment repair, startup sync, hook auditing, collaborator handoff.

    1.4k GitHub stars~2.6k tokensUpdated today
    Auto-check: notes
  • Bigdata Skill

    daymade/claude-code-skills

    Pulls Bigdata.com (RavenPack) financial and news data via the official bigdata-client SDK and /v1/ REST endpoints — structured financials, prices, analyst estimates, entity-sentiment series…

    1.4k GitHub stars~3.7k tokensUpdated today
    Auto-check passed

Questions about Wps Doc Scraper

What does Wps Doc Scraper do?

Faithfully archive public WPS/KDocs/金山文档 links, especially embedded ProcessOn .pof mind maps and canvases, as raw source data, original SVG/PNG, and Markdown. Wps Doc Scraper is an agent skill from daymade/claude-code-skills.pof mind maps and canvases, as raw source data, original SVG/PNG, and Markdown.

When should I use Wps Doc Scraper?

Wps Doc Scraper fits situations like: A user gives a kdocs.cn; wps.processon.com link and asks to scrape; 转 Markdown without logging in; saving the document to an account.

How do I install Wps Doc Scraper in Claude Code?

Run `npx skills add daymade/claude-code-skills --skill wps-doc-scraper -a claude-code`. Or copy the skill folder (wps-doc-scraper in daymade/claude-code-skills) into .claude/skills/wps-doc-scraper in your project. Claude Code loads it when a task matches its description.

How do I install Wps Doc Scraper in Codex?

Run `npx skills add daymade/claude-code-skills --skill wps-doc-scraper -a codex`. Or copy the skill folder (wps-doc-scraper in daymade/claude-code-skills) into .agents/skills/wps-doc-scraper in your project. Codex loads it when a task matches its description.

Can I use Wps Doc Scraper in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add daymade/claude-code-skills --skill wps-doc-scraper -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/wps-doc-scraper, .gemini/skills/wps-doc-scraper, .github/skills/wps-doc-scraper and .opencode/skills/wps-doc-scraper in your project.

What does Wps Doc Scraper need to run?

Going by SKILL.md and its folder, Wps Doc Scraper needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Wps Doc Scraper access the network?

SKILL.md names 1 domain. In commands or code: kdocs.cn; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Wps Doc Scraper safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Wps Doc Scraper use?

Wps Doc Scraper is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Wps Doc Scraper use?

About 1.1k tokens (SKILL.md is roughly 4.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to Wps Doc Scraper?

Skills that share tags, products or a category with Wps Doc Scraper: Tmux (trpc-group/trpc-agent-go, 1.8k stars), Ketch (1broseidon/ketch, 697 stars), Crawl4AI Web Scraping (smallnest/goclaw, 598 stars) and Boss Zhipin Scraper (eatmoreduck/boss-zhipin-scraper, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Wps Doc Scraper?

daymade (a GitHub user) maintains it in daymade/claude-code-skills, which has 1,444 GitHub stars. The repository holds 103 skills in this directory. The repository was last updated on October 8, 2026.

Source: daymade/claude-code-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.