Agent skill

Content Extract

by aAAaqwq in aAAaqwq/AGI-Super-Team

Robust URL-to-Markdown extraction for OpenClaw workflows. An agent skill from aAAaqwq/AGI-Super-Team.

MITAuto-check passedKnowledge Management

Install Content Extract

skills CLI
$ npx skills add aAAaqwq/AGI-Super-Team --skill content-extract -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aAAaqwq/AGI-Super-Team content-extract --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aAAaqwq/AGI-Super-Team.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/content-extract .claude/skills/content-extract && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
content-extract
GitHub stars
105
Used in
1 other repo
Token cost
~650 tokens
SKILL.md length
132 words
Files
4 (incl. scripts, references)
Skills in repo
161
Repo updated
First seen
Licence
MIT

At a glance

Robust URL-to-Markdown extraction for OpenClaw workflows. An agent skill from aAAaqwq/AGI-Super-Team.

  • The user wants to extract/summarize/convert a webpage to markdown (especially WeChat mp.weixin.qq.com) and webfetch/browser is blocked
  • SKILL.md covers 工作流(Decision Tree), MinerU 调用(给 agent 的确定性脚本), 交付规范(强制) and 本 skill 自身不做什么, plus 1 more section
  • Runs Python scripts from its folder; calls python3
  • Tasks that involve Messaging and chat bots

What it does

Content Extract is an agent skill from aAAaqwq/AGI-Super-Team. Robust URL-to-Markdown extraction for OpenClaw workflows. Use when the user wants to "extract/summarize/convert a webpage to markdown" (especially WeChat mp.weixin.qq.com) and webfetch/browser is blocked or messy. Uses a cheap probe via webfetch first, then falls back to the official MinerU API (via the local mineru-extract skill) and returns a traceable result contract with source links.

Its SKILL.md is about 650 tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/domain-whitelist.md`, `references/heuristics.md` and `scripts/content_extract.py`).

It sits in Knowledge Management, covering Messaging and chat bots and Web clipping and read-later. It works with WeChat and Model Context Protocol. The repository describes itself as: An installable, cross-framework AI organization: C-suite agents, expert subagents, curated skills, independent review, and one-command setup across 18 AI client/runtime adapters. The licence is MIT.

When your agent uses it

  • The user wants to extract/summarize/convert a webpage to markdown (especially WeChat mp.weixin.qq.com) and webfetch/browser is blocked
  • Tasks that involve Messaging and chat bots
  • Tasks that involve Web clipping and read-later

Example prompts

  • “extract/summarize/convert a webpage to markdown”
  • “/content-extract”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 331ecd3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • mineru.net
    • opendatalab.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Content Extract loads about 650 tokens when it runs, and up to ~1.1k if it reads all its reference files. Until then it costs about 102 tokens; SKILL.md has 132 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~102
When it runs · the whole SKILL.md, loaded when a task matches
~650
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from aAAaqwq/AGI-Super-Team at commit 331ecd3, republished under its MIT licence (© aAAaqwq). 132 words, ~650 tokens.

Download SKILL.mdSave it as .claude/skills/content-extract/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
content-extract
description
Robust URL-to-Markdown extraction for OpenClaw workflows. Use when the user wants to "extract/summarize/convert a webpage to markdown" (especially WeChat mp.weixin.qq.com) and web_fetch/browser is blocked or messy. Uses a cheap probe via web_fetch first, then falls back to the official MinerU API (via the local mineru-extract skill) and returns a traceable result contract with source links.
author
Daniel Li

content-extract — 上层内容解析入口(MCP 语义对齐,但不跑 MCP Server)

  • Author: Daniel Li
  • Copyright © Daniel Li. All rights reserved.

目标:把“给我一个 URL → 产出可读 Markdown + 可追溯入口”变成一个统一入口,供后续所有业务 skill(github-explorer、写作类 skills、日报等)复用。

核心原则(来自你发的 Excel Skill 拆解文章的启发):

  • 行为规约层:永远给出可追溯入口(原文 URL + 解析产物路径/链接),绝不编造来源。
  • Token 探针:先用低成本 probe 判断可不可以直接抓;不行再走重解析(MinerU)。
  • 反弹机制:失败时返回“下一步动作建议”,而不是一堆异常栈。

工作流(Decision Tree)

输入:url

  1. Domain Whitelist(跳过 probe):若 URL 属于高概率反爬/动态站点(微信/知乎等),直接走 MinerU
  • 白名单文件:references/domain-whitelist.md
  • 对命中白名单的 URL:强制 model_version=MinerU-HTML
  1. Probe(低成本):优先用 web_fetch(url)
  • 目标:拿到正文 markdown(便宜、快)
  • 判断“失败/不合格”条件(见 references/heuristics.md)包括:
    • 403/401/反爬
    • 只有“环境异常/验证码/请在微信打开”等提示
    • 内容极短/明显导航页/丢正文
  1. Fallback(高保真):走 MinerU 官方 API
  • 调用下游 driver:skills/mineru-extract/scripts/mineru_parse_documents.py
  • 对 HTML 页面(微信等):强制 model_version=MinerU-HTML
  1. 输出统一结果合同(Result Contract)

无论用 probe 还是 MinerU,都返回同一套结构:

json
{
  "ok": true,
  "source_url": "...",
  "engine": "web_fetch" ,
  "markdown": "...",
  "artifacts": {
    "out_dir": "...",
    "markdown_path": "...",
    "zip_path": "..."
  },
  "sources": [
    "原文URL",
    "(如使用MinerU)MinerU full_zip_url",
    "(如使用MinerU)本地markdown_path"
  ],
  "notes": ["任何重要限制/失败原因/下一步建议"]
}

注意:engine 可能是 web_fetch 或 mineru。

MinerU 调用(给 agent 的确定性脚本)

当需要 MinerU 时,用这个命令(返回 JSON,且可把 markdown 内联进 JSON,便于下游总结):

bash
python3 mineru-extract/scripts/mineru_parse_documents.py \
  --file-sources "<URL>" \
  --model-version MinerU-HTML \
  --emit-markdown --max-chars 20000

路径说明: 上述命令假设你在 skills 安装根目录下执行。如果 mineru-extract 安装在其他位置,请替换为实际路径。

交付规范(强制)

  • 输出必须包含 sources(原文入口 + 解析产物入口)。
  • 如果 MinerU 成功:必须把 markdown_path(本地路径)写进 sources,方便复查。
  • 如果两条链路都失败:必须明确失败原因,并给出下一步(例如:让 Boss 提供可访问镜像链接 / 允许我用浏览器 relay 导出 HTML / 走上传 HTML 文件解析的兜底方案)。

本 skill 自身不做什么

  • 不跑 MCP Server(避免常驻服务与运维负担)
  • 不试图绕过登录/验证码(这属于访问层问题;我们只做解析层和工作流路由)

References

© aAAaqwq, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in skills/content-extract of aAAaqwq/AGI-Super-Team.

  • SKILL.md
  • references/domain-whitelist.md
  • references/heuristics.md
  • scripts/content_extract.py

Open the folder on GitHubat commit 331ecd3

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in aAAaqwq/AGI-Super-Team, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Content Extract next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Content Extract compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Content Extract this skillaAAaqwq/AGI-Super-Team1051 repos~650Automated safety check: PassMIT
Web To Markdownrookie-ricardo/erduo-skills935—~894Automated safety check: PassMIT
Clean Content FetchLeoYeAI/openclaw-master-skills2.2k—~574Automated safety check: PassMIT
Md2wechatgeekjourneyx/md2wechat-skill3.7k—~3.8kAutomated safety check: PassCustom licence
Burp MCP Vuln Checklangbyyi/CyberStrikeAI-SRC133—~3.1kAutomated safety check: PassApache-2.0
Wechat To Mdbzd6661/wechat-article-for-ai101—~654Automated safety check: PassNone

Similar skills

  • Web To Markdown

    rookie-ricardo/erduo-skills

    Convert a web URL into cleaned Markdown with deterministic routing.

    935 GitHub stars~894 tokensUpdated 2 mo ago
    Productivity & AutomationAuto-check passed
  • Clean Content Fetch

    LeoYeAI/openclaw-master-skills

    获取干净、可读的网页正文内容,适合现代网页、博客、新闻、公告和微信公众号文章抓取;支持网页正文提取、内容清洗、去噪、Markdown 输出,适用于普通 fetch 效果不佳、页面噪音较多或动态渲染干扰的场景。Clean content fetch for modern web pages, article extraction, WeChat article capture, content…

    2.2k GitHub stars~574 tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Md2wechat

    geekjourneyx/md2wechat-skill

    Convert Markdown to WeChat Official Account HTML. An agent skill from geekjourneyx/md2wechat-skill.

    3.7k GitHub stars~3.8k tokensUpdated 13 days ago
    Media & CreativeAuto-check passed
  • Burp MCP Vuln Check

    langbyyi/CyberStrikeAI-SRC

    Automate low-impact web vulnerability verification through Burp MCP.

    133 GitHub stars~3.1k tokensUpdated 10 days ago
    SecurityAuto-check passed
  • Wechat To Md

    bzd6661/wechat-article-for-ai

    Convert WeChat Official Account (微信公众号) articles to clean Markdown files with locally downloaded images.

    101 GitHub stars~654 tokensUpdated 7 mo ago
    Productivity & AutomationAuto-check passed
  • Wiki Ingest

    zhuzhaoyun/Molio

    将源文件/资料增量导入(入库)到现有 wiki,使知识持续积累和演进。读取源文件,生成或更新 source 摘要页与实体/概念/对比等页面,建立交叉链接,检测矛盾,更新分层 INDEX/log/hot(旧单索引库首次入库时自动升级为分层索引)。支持显式文件路径、URL、或无显式目标时自动找最近 raw/wechat 暂存资料。Triggers on: 入库, 导入, 整理进知识库…

    431 GitHub stars~1.5k tokensUpdated today
    Knowledge ManagementAuto-check passed

More from aAAaqwq/AGI-Super-Team

All 161 skills in this repo
  • Content Creator

    aAAaqwq/AGI-Super-Team

    Create SEO-optimized marketing content with consistent brand voice.

    105 GitHub starsUsed in 3 repos~1.9k tokens
    Auto-check passed
  • Financial Calculator

    aAAaqwq/AGI-Super-Team

    Advanced financial calculator with future value tables, present value, discount calculations, markup pricing, and compound interest.

    105 GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Performing Security Code Review

    aAAaqwq/AGI-Super-Team

    This skill provides automated assistance for security agent tasks Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    105 GitHub starsUsed in 1 repo~920 tokens
    Auto-check passed
  • Frontend Design Ultimate

    aAAaqwq/AGI-Super-Team

    Create distinctive, production-grade static sites with React, Tailwind CSS, and shadcn/ui — no mockups needed.

    105 GitHub starsUsed in 2 repos~2.7k tokens
    Auto-check passed
  • Sysadmin Toolbox

    aAAaqwq/AGI-Super-Team

    Tool discovery and shell one-liner reference for sysadmin, DevOps, and security tasks.

    105 GitHub starsUsed in 2 repos~775 tokens
    Auto-check passed
  • Zsxq Smart Publish

    aAAaqwq/AGI-Super-Team

    Publish and manage content on 知识星球 (zsxq.com). An agent skill from aAAaqwq/AGI-Super-Team.

    105 GitHub stars~1.5k tokensUpdated 10 days ago
    Auto-check passed

Questions about Content Extract

What does Content Extract do?

Robust URL-to-Markdown extraction for OpenClaw workflows. An agent skill from aAAaqwq/AGI-Super-Team. Content Extract is an agent skill from aAAaqwq/AGI-Super-Team. Robust URL-to-Markdown extraction for OpenClaw workflows.

When should I use Content Extract?

Content Extract fits situations like: the user wants to extract/summarize/convert a webpage to markdown (especially WeChat mp.weixin.qq.com) and webfetch/browser is blocked; tasks that involve Messaging and chat bots; tasks that involve Web clipping and read-later.

How do I install Content Extract in Claude Code?

Run `npx skills add aAAaqwq/AGI-Super-Team --skill content-extract -a claude-code`. Or copy the skill folder (skills/content-extract in aAAaqwq/AGI-Super-Team) into .claude/skills/content-extract in your project. Claude Code loads it when a task matches its description.

How do I install Content Extract in Codex?

Run `npx skills add aAAaqwq/AGI-Super-Team --skill content-extract -a codex`. Or copy the skill folder (skills/content-extract in aAAaqwq/AGI-Super-Team) into .agents/skills/content-extract in your project. Codex loads it when a task matches its description.

Can I use Content Extract in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aAAaqwq/AGI-Super-Team --skill content-extract -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/content-extract, .gemini/skills/content-extract, .github/skills/content-extract and .opencode/skills/content-extract in your project.

What does Content Extract need to run?

Going by SKILL.md and its folder, Content Extract needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Content Extract access the network?

SKILL.md names 2 domains. As links in the text: mineru.net and opendatalab.github.io. This is read from the text; nothing was executed.

Is Content Extract safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Content Extract use?

Content Extract is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Content Extract use?

About 650 tokens (SKILL.md is roughly 2.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 479 tokens, read only when the agent opens those files.

What are the alternatives to Content Extract?

Skills that share tags, products or a category with Content Extract: Web To Markdown (rookie-ricardo/erduo-skills, 935 stars), Clean Content Fetch (LeoYeAI/openclaw-master-skills, 2.2k stars), Md2wechat (geekjourneyx/md2wechat-skill, 3.7k stars) and Burp MCP Vuln Check (langbyyi/CyberStrikeAI-SRC, 133 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Content Extract?

aAAaqwq (a GitHub user) maintains it in aAAaqwq/AGI-Super-Team, which has 105 GitHub stars. The repository holds 161 skills in this directory. The repository was last updated on September 27, 2026.

Source: aAAaqwq/AGI-Super-Team on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.