Agent skill

Clean Content Fetch

by LeoYeAI in LeoYeAI/openclaw-master-skills

获取干净、可读的网页正文内容,适合现代网页、博客、新闻、公告和微信公众号文章抓取;支持网页正文提取、内容清洗、去噪、Markdown 输出,适用于普通 fetch 效果不佳、页面噪音较多或动态渲染干扰的场景。Clean content fetch for modern web pages, article extraction, WeChat article capture, content…

MITAuto-check passedDocuments & Office

Install Clean Content Fetch

skills CLI
$ npx skills add LeoYeAI/openclaw-master-skills --skill clean-content-fetch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install LeoYeAI/openclaw-master-skills clean-content-fetch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/clean-content-fetch .claude/skills/clean-content-fetch && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
clean-content-fetch
GitHub stars
2.2k
Token cost
~574 tokens
SKILL.md length
85 words
Files
1
Skills in repo
1,235
Repo updated
First seen
Licence
MIT

At a glance

获取干净、可读的网页正文内容,适合现代网页、博客、新闻、公告和微信公众号文章抓取;支持网页正文提取、内容清洗、去噪、Markdown 输出,适用于普通 fetch 效果不佳、页面噪音较多或动态渲染干扰的场景。Clean content fetch for modern web pages, article extraction, WeChat article capture, content…

  • Works in 5 steps: 使用 python3 scripts/scrapling_fetch.py → 默认正文选择器优先级 → 命中正文后,使用 html2text 转 Markdown → …
  • Tasks that involve Messaging and chat bots
  • SKILL.md covers 默认流程, 用法, 依赖 and 输出约定, plus 4 more sections
  • Calls python3, python and pip; reaches xhslink.com

What it does

Clean Content Fetch is an agent skill from LeoYeAI/openclaw-master-skills. 获取干净、可读的网页正文内容,适合现代网页、博客、新闻、公告和微信公众号文章抓取;支持网页正文提取、内容清洗、去噪、Markdown 输出,适用于普通 fetch 效果不佳、页面噪音较多或动态渲染干扰的场景。Clean content fetch for modern web pages, article extraction, WeChat article capture, content cleanup, noise reduction, and markdown output when ordinary fetch is not clean enough.

Its SKILL.md is about 570 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering Messaging and chat bots, Web clipping and read-later and Markdown. It works with WeChat and Python. The repository describes itself as: 🧠 Curated collection of 1209+ best OpenClaw skills — weekly updated by MyClaw.ai. The licence is MIT.

When your agent uses it

  • Tasks that involve Messaging and chat bots
  • Tasks that involve Web clipping and read-later
  • Tasks that involve Markdown

Example prompts

  • “/clean-content-fetch”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. 使用 python3 scripts/scrapling_fetch.py
  2. 默认正文选择器优先级
  3. 命中正文后,使用 html2text 转 Markdown
  4. 若都未命中,回退到 body
  5. 最终按 max_chars 截断输出

What it can do on your machine

Read from SKILL.md and the folder at commit e5199b5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • xhslink.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Clean Content Fetch loads about 574 tokens when it runs. Until then it costs about 76 tokens; SKILL.md has 85 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~76
When it runs · the whole SKILL.md, loaded when a task matches
~574

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from LeoYeAI/openclaw-master-skills at commit e5199b5, republished under its MIT licence (© LeoYeAI). 85 words, ~574 tokens.

Download SKILL.mdSave it as .claude/skills/clean-content-fetch/SKILL.md (or your agent's skills folder).
name
clean-content-fetch
description
获取干净、可读的网页正文内容,适合现代网页、博客、新闻、公告和微信公众号文章抓取;支持网页正文提取、内容清洗、去噪、Markdown 输出,适用于普通 fetch 效果不佳、页面噪音较多或动态渲染干扰的场景。Clean content fetch for modern web pages, article extraction, WeChat article capture, content cleanup, noise reduction, and markdown output when ordinary fetch is not clean enough.

Scrapling Web Fetch

当用户要获取网页内容、正文提取、把网页转成 markdown/text、抓取文章主体时,优先使用此技能。

默认流程

  1. 使用 python3 scripts/scrapling_fetch.py <url> <max_chars>
  2. 默认正文选择器优先级:
    • article
    • main
    • .post-content
    • [class*="body"]
  3. 命中正文后,使用 html2text 转 Markdown
  4. 若都未命中,回退到 body
  5. 最终按 max_chars 截断输出

用法

bash
python3 /Users/zzd/.openclaw/workspace/skills/scrapling-web-fetch/scripts/scrapling_fetch.py <url> 30000

依赖

优先检查:

  • scrapling
  • html2text
  • curl_cffi
  • playwright
  • browserforge

推荐使用独立虚拟环境,避免系统 Python 的 PEP 668 限制:

bash
python3 -m venv /Users/zzd/.openclaw/workspace/.venvs/clean-content-fetch
/Users/zzd/.openclaw/workspace/.venvs/clean-content-fetch/bin/pip install scrapling html2text curl_cffi playwright browserforge
/Users/zzd/.openclaw/workspace/.venvs/clean-content-fetch/bin/python -m playwright install chromium

如直接运行脚本,优先使用该虚拟环境中的 Python:

bash
/Users/zzd/.openclaw/workspace/.venvs/clean-content-fetch/bin/python /Users/zzd/.openclaw/workspace/skills/scrapling-web-fetch/scripts/scrapling_fetch.py <url> 30000

输出约定

脚本默认输出 Markdown 正文内容。 如需结构化输出,可追加 --json。 如需调试提取命中了哪个 selector,可查看 stderr 输出。

附加资源

  • 用法参考:/Users/zzd/.openclaw/workspace/skills/scrapling-web-fetch/references/usage.md
  • 选择器策略:/Users/zzd/.openclaw/workspace/skills/scrapling-web-fetch/references/selectors.md
  • 统一入口:/Users/zzd/.openclaw/workspace/skills/scrapling-web-fetch/scripts/fetch-web-content

何时用这个技能

  • 获取文章正文
  • 抓博客/新闻/公告正文
  • 将网页转成 Markdown 供后续总结
  • 常规 fetch 效果差,希望提升现代网页抓取稳定性
  • 抓小红书分享短链或笔记落地页正文

小红书抓取方法

对于 xhslink.com 短链或小红书笔记页,推荐直接使用虚拟环境中的脚本运行:

bash
/Users/zzd/.openclaw/workspace/.venvs/clean-content-fetch/bin/python /Users/zzd/.openclaw/workspace/skills/scrapling-web-fetch/scripts/scrapling_fetch.py 'http://xhslink.com/o/9745hugimlD' 30000

说明:

  • 脚本会先解析短链并抓取落地页正文
  • 适合提取小红书笔记文案、标题和主体内容
  • 若页面需要更复杂交互,再切到浏览器自动化

何时不用

  • 需要完整浏览器交互、点击、登录、翻页时:改用浏览器自动化
  • 只是简单获取 API JSON:直接请求 API 更合适

© LeoYeAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/clean-content-fetch of LeoYeAI/openclaw-master-skills.

Open the folder on GitHubat commit e5199b5

Compare with similar skills

Clean Content Fetch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Clean Content Fetch compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Clean Content Fetch this skillLeoYeAI/openclaw-master-skills2.2k—~574Automated safety check: PassMIT
Web To Markdownrookie-ricardo/erduo-skills935—~894Automated safety check: PassMIT
Wechat Article To Markdownjackwener/wechat-article-to-markdown1.1k—~352Automated safety check: PassNone
Hailaobao Gzh Designwwwzhouhui/skills_collection283—~1.1kAutomated safety check: PassNone
Gzh Designisjiamu/gzh-design-skill4k—~2.2kAutomated safety check: PassAGPL-3.0
URL to Markdown Fetcherjoeseesun/qiaomu-markdown-proxy509—~1.4kAutomated safety check: PassMIT

Similar skills

  • Web To Markdown

    rookie-ricardo/erduo-skills

    Convert a web URL into cleaned Markdown with deterministic routing.

    935 GitHub stars~894 tokensUpdated 2 mo ago
    Productivity & AutomationAuto-check passed
  • Wechat Article To Markdown

    jackwener/wechat-article-to-markdown

    Fetch WeChat Official Account (微信公众号) articles from mp.weixin.qq.com and convert to Markdown.

    1.1k GitHub stars~352 tokensUpdated 6 mo ago
    Documents & OfficeAuto-check passed
  • Hailaobao Gzh Design

    wwwzhouhui/skills_collection

    把 Markdown 文章排版成微信公众号可直接粘贴的 HTML,内置 7 套排版风格。触发词:公众号排版、微信排版、gzh 排版、typeset for wechat、Markdown 转公众号。不触发:只想改文字不改格式、要生成小红书图文。

    283 GitHub stars~1.1k tokensUpdated 4 days ago
    Documents & OfficeAuto-check passed
  • Gzh Design

    isjiamu/gzh-design-skill

    微信公众号文章排版引擎,将 Markdown 转换为可直接粘贴到公众号编辑器的 HTML。主题风格从 references/theme-index.md 注册的自定义主题库中选取,自动章节编号、关键词下划线标记、引言卡片、目录导航、代码块、图片/GIF、作者签名。支持 Markdown / Word(.docx) / PDF / 纯文本输入(非 Markdown…

    4k GitHub stars~2.2k tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • URL to Markdown Fetcher

    joeseesun/qiaomu-markdown-proxy

    Converts a URL or PDF into clean Markdown, with dedicated routes for WeChat articles, Feishu docs, arXiv papers and login-gated pages, before any summary or rewrite.

    509 GitHub stars~1.4k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Qiaomu Epub Book Generator

    joeseesun/qiaomu-epub-book-generator

    Generate EPUB ebooks from Markdown files with SVG→PNG conversion, remote/local image embedding, compressed illustrations, and WeChat Reading compatibility.

    134 GitHub stars~1.1k tokensUpdated 6 mo ago
    Documents & OfficeAuto-check passed

More from LeoYeAI/openclaw-master-skills

All 1,235 skills in this repo
  • DevOps Pipeline Management

    LeoYeAI/openclaw-master-skills

    Manages pipelines on a DevOps quality and efficiency platform through its OpenAPI: list workspaces and templates, create, update, run and cancel pipelines, and read run records.

    2.2k GitHub stars~4.2k tokensUpdated 2 mo ago
    Auto-check: notes
  • Feishu Document Collaboration

    LeoYeAI/openclaw-master-skills

    Patches OpenClaw's Feishu extension so an edited document triggers an isolated agent session that reads the doc and replies inline, turning it into a live chat space.

    2.2k GitHub stars~2k tokensUpdated 2 mo ago
    Auto-check passed
  • Files Memory System

    LeoYeAI/openclaw-master-skills

    Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.

    2.2k GitHub stars~3.8k tokensUpdated 2 mo ago
    Auto-check passed
  • GEO-Claw AI Visibility Agent

    LeoYeAI/openclaw-master-skills

    Runs a brand's AI-search visibility work end to end: diagnosing how AI platforms represent it, repositioning it, producing AI-optimized content and monitoring ongoing mentions.

    2.2k GitHub stars~4.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Google Workspace CLI

    LeoYeAI/openclaw-master-skills

    Installs and authenticates the gws CLI, then automates Gmail, Drive, Sheets, Calendar, Docs, Chat and Tasks with ready-made recipes, persona bundles and security audits.

    2.2k GitHub stars~2.6k tokensUpdated 2 mo ago
    Auto-check: notes
  • HealthFit Health Advisors

    LeoYeAI/openclaw-master-skills

    Runs four advisor roles, a fitness coach, nutritionist, data analyst and TCM practitioner, to build a health profile and track workouts, diet and wellness over time.

    2.2k GitHub stars~4.4k tokensUpdated 2 mo ago
    Auto-check passed

Works with

Questions about Clean Content Fetch

What does Clean Content Fetch do?

获取干净、可读的网页正文内容,适合现代网页、博客、新闻、公告和微信公众号文章抓取;支持网页正文提取、内容清洗、去噪、Markdown 输出,适用于普通 fetch 效果不佳、页面噪音较多或动态渲染干扰的场景。Clean content fetch for modern web pages, article extraction, WeChat article capture, content…. Clean Content Fetch is an agent skill from LeoYeAI/openclaw-master-skills. 获取干净、可读的网页正文内容,适合现代网页、博客、新闻、公告和微信公众号文章抓取;支持网页正文提取、内容清洗、去噪、Markdown 输出,适用于普通 fetch 效果不佳、页面噪音较多或动态渲染干扰的场景。Clean content fetch for modern web pages, article extraction, WeChat article capture, content cleanup, noise reduction, and markdown output when ordinary fetch is not clean enough.

When should I use Clean Content Fetch?

Clean Content Fetch fits situations like: tasks that involve Messaging and chat bots; tasks that involve Web clipping and read-later; tasks that involve Markdown.

How do I install Clean Content Fetch in Claude Code?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill clean-content-fetch -a claude-code`. Or copy the skill folder (skills/clean-content-fetch in LeoYeAI/openclaw-master-skills) into .claude/skills/clean-content-fetch in your project. Claude Code loads it when a task matches its description.

How do I install Clean Content Fetch in Codex?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill clean-content-fetch -a codex`. Or copy the skill folder (skills/clean-content-fetch in LeoYeAI/openclaw-master-skills) into .agents/skills/clean-content-fetch in your project. Codex loads it when a task matches its description.

Can I use Clean Content Fetch in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LeoYeAI/openclaw-master-skills --skill clean-content-fetch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/clean-content-fetch, .gemini/skills/clean-content-fetch, .github/skills/clean-content-fetch and .opencode/skills/clean-content-fetch in your project.

What does Clean Content Fetch need to run?

Going by SKILL.md and its folder, Clean Content Fetch needs the command-line tools its instructions call (python3, python and pip). Our summary lists: Python 3.

Does Clean Content Fetch access the network?

SKILL.md names 1 domain. In commands or code: xhslink.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Clean Content Fetch safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Clean Content Fetch use?

Clean Content Fetch is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Clean Content Fetch use?

About 574 tokens (SKILL.md is roughly 2.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Clean Content Fetch?

Skills that share tags, products or a category with Clean Content Fetch: Web To Markdown (rookie-ricardo/erduo-skills, 935 stars), Wechat Article To Markdown (jackwener/wechat-article-to-markdown, 1.1k stars), Hailaobao Gzh Design (wwwzhouhui/skills_collection, 283 stars) and Gzh Design (isjiamu/gzh-design-skill, 4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Clean Content Fetch?

LeoYeAI (a GitHub user) maintains it in LeoYeAI/openclaw-master-skills, which has 2,161 GitHub stars. The repository holds 1,235 skills in this directory. The repository was last updated on July 20, 2026.

Source: LeoYeAI/openclaw-master-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.