Clean Content Fetch
LeoYeAI/openclaw-master-skills
获取干净、可读的网页正文内容,适合现代网页、博客、新闻、公告和微信公众号文章抓取;支持网页正文提取、内容清洗、去噪、Markdown 输出,适用于普通 fetch 效果不佳、页面噪音较多或动态渲染干扰的场景。Clean content fetch for modern web pages, article extraction, WeChat article capture, content…
Web取得の標準前処理レイヤー。URL・公式ドキュメント・ブログ・ニュース・OSSページを読むときは、原則として毎回このSkillでDefuddleを使い本文をMarkdown/JSON化してから読む。生HTMLのまま要約・分析・比較・レビューしない。トリガー例: URLを読む/Webページ要約/OSS調査/公式Doc確認/記事解析/競合サイト確認/Web一次情報確認。
$ npx skills add cloudnative-co/claude-code-starter-kit --skill web-content-extraction -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install cloudnative-co/claude-code-starter-kit web-content-extraction --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/cloudnative-co/claude-code-starter-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/web-content-extraction .claude/skills/web-content-extraction && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "web-content-extraction" agent skill from https://github.com/cloudnative-co/claude-code-starter-kit/tree/main/skills/web-content-extraction into .claude/skills/web-content-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-content-extraction", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/cloudnative-co/claude-code-starter-kit/tree/main/skills/web-content-extractionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add cloudnative-co/claude-code-starter-kit --skill web-content-extraction -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install cloudnative-co/claude-code-starter-kit web-content-extraction --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cloudnative-co/claude-code-starter-kit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/web-content-extraction .agents/skills/web-content-extraction && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "web-content-extraction" agent skill from https://github.com/cloudnative-co/claude-code-starter-kit/tree/main/skills/web-content-extraction into .agents/skills/web-content-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-content-extraction", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cloudnative-co/claude-code-starter-kit --skill web-content-extraction -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install cloudnative-co/claude-code-starter-kit web-content-extraction --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cloudnative-co/claude-code-starter-kit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/web-content-extraction .cursor/skills/web-content-extraction && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "web-content-extraction" agent skill from https://github.com/cloudnative-co/claude-code-starter-kit/tree/main/skills/web-content-extraction into .cursor/skills/web-content-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-content-extraction", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/cloudnative-co/claude-code-starter-kit.git --path skills/web-content-extraction--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add cloudnative-co/claude-code-starter-kit --skill web-content-extraction -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install cloudnative-co/claude-code-starter-kit web-content-extraction --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cloudnative-co/claude-code-starter-kit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/web-content-extraction .gemini/skills/web-content-extraction && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "web-content-extraction" agent skill from https://github.com/cloudnative-co/claude-code-starter-kit/tree/main/skills/web-content-extraction into .gemini/skills/web-content-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-content-extraction", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install cloudnative-co/claude-code-starter-kit web-content-extractionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add cloudnative-co/claude-code-starter-kit --skill web-content-extraction -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/cloudnative-co/claude-code-starter-kit.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/web-content-extraction .github/skills/web-content-extraction && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "web-content-extraction" agent skill from https://github.com/cloudnative-co/claude-code-starter-kit/tree/main/skills/web-content-extraction into .github/skills/web-content-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-content-extraction", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cloudnative-co/claude-code-starter-kit --skill web-content-extraction -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install cloudnative-co/claude-code-starter-kit web-content-extraction --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cloudnative-co/claude-code-starter-kit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/web-content-extraction .opencode/skills/web-content-extraction && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "web-content-extraction" agent skill from https://github.com/cloudnative-co/claude-code-starter-kit/tree/main/skills/web-content-extraction into .opencode/skills/web-content-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-content-extraction", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
web-content-extractionWeb取得の標準前処理レイヤー。URL・公式ドキュメント・ブログ・ニュース・OSSページを読むときは、原則として毎回このSkillでDefuddleを使い本文をMarkdown/JSON化してから読む。生HTMLのまま要約・分析・比較・レビューしない。トリガー例: URLを読む/Webページ要約/OSS調査/公式Doc確認/記事解析/競合サイト確認/Web一次情報確認。
Web Content Extraction is an agent skill from cloudnative-co/claude-code-starter-kit. Web取得の標準前処理レイヤー。URL・公式ドキュメント・ブログ・ニュース・OSSページを読むときは、原則として毎回このSkillでDefuddleを使い本文をMarkdown/JSON化してから読む。生HTMLのまま要約・分析・比較・レビューしない。トリガー例: URLを読む/Webページ要約/OSS調査/公式Doc確認/記事解析/競合サイト確認/Web一次情報確認。
Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 20 other files, including scripts (for example `README.md`, `package-lock.json` and `package.json`).
It sits in Documents & Office, covering Web clipping and read-later. The repository describes itself as: One-command setup of a complete Claude Code development environment with interactive wizard. The licence is MIT.
Read from SKILL.md and the folder at commit f00e7ce. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 7 files in scripts/ (JavaScript and Shell), which the agent can run.
Shell commands in SKILL.md call:
npmFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Web Content Extraction loads about 1.2k tokens when it runs. Until then it costs about 52 tokens; SKILL.md has 264 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from cloudnative-co/claude-code-starter-kit at commit f00e7ce, republished under its MIT licence (© cloudnative-co). 264 words, ~1,219 tokens.
.claude/skills/web-content-extraction/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.Webページ本文を抽出し、LLMが読みやすい Markdown/JSON に整形する。 Claude Code が Webページ、URL、公式ドキュメント、ブログ記事、ニュース記事、OSSページを読む場合は、原則として毎回このSkillを使う。
When reading any public web URL, use Defuddle first.
Do not summarize, analyze, compare, or review a web page from raw HTML unless Defuddle extraction fails or the page type is explicitly unsupported.
# 公開URLを取得して本文をMarkdown/JSON化(SSRFガードあり)
~/.claude/skills/web-content-extraction/scripts/run-node.sh \
~/.claude/skills/web-content-extraction/scripts/defuddle-url.mjs <url># ローカルHTMLファイルを本文抽出(外部通信なし)
~/.claude/skills/web-content-extraction/scripts/run-node.sh \
~/.claude/skills/web-content-extraction/scripts/defuddle-file.mjs <file>出力は JSON(stdout)。最低限 success, url, fetchedAt/parsedAt, title, author,
site, domain, published, description, wordCount, content(Markdown) を含む。
warnings / fetchWarnings がある場合は抽出の信頼性に注意する。
| フィールド | 意味 |
|---|---|
success | 本文抽出に成功したか(false は抽出失敗/空) |
warnings | 本文が短い・空など低信頼の警告 |
url / requestedUrl / finalUrl | 対象URL(リダイレクト後の最終URL含む) |
fetchedAt / parsedAt | 取得・解析時刻(ISO8601, 監査用に必ず保持) |
title author site domain published description | メタデータ |
wordCount | 語数(空白区切り。日本語は極端に小さく出る) |
charCount | 非空白の文字数(日本語の実分量はこちらで判断) |
cjkCharCount | CJK文字数(日本語/中国語/韓国語の量の目安) |
content | 本文(HTMLはMarkdown、PDFはプレーンテキスト) |
extractorType | サイト固有抽出器が使われた場合の種別 |
extractorEngine | PDF抽出時のみ "pdf"。pageCount も付く |
useAsync はupstreamに存在しないため意図を構造で担保)。localhost / プライベートIP(10/8,172.16/12,192.168/16,127/8,169.254/16,100.64/10 等) /
.local/.internal 等 / 単一ラベルの内部ホスト名 / 非http(s) / 認証情報付きURL は標準で拒否。2000::/3 以外は全拒否。Teredo/site-local/documentation/NAT64/IPv4-mapped/6to4(private埋め込み)等を含む)。ALLOW_PRIVATE_URLS=true(バイパスは stderr に監査記録)。success:false や warnings あり)で断定しない。公開URLが PDF(content-type: application/pdf / .pdf / 先頭 %PDF-)の場合、defuddle-url.mjs
は自動で pdfjs-dist によるテキスト抽出にフォールバックする(extractorEngine:"pdf", pageCount 付き)。
content はMarkdownでなくプレーンテキスト。charCount:0+警告 → OCRが必要(本Skillの対象外)。defuddle-file.mjs はHTML専用)。cd ~/.claude/skills/web-content-extraction && ./scripts/run-node.sh --testtest/url-guard.test.mjs(SSRFガード)/ test/defuddle-core.test.mjs(charCount)/
test/extract-smoke.test.mjs(実抽出スモーク: HTML+PDF)/
test/defuddle-url.test.mjs(URL CLI exit code契約)を実行。DNS非依存・オフラインの
決定的テストのみ。CIは .github/workflows/skill-web-content-extraction.yml(Node 22/24 マトリクス)。
web-content-update feature が有効な場合のみ、SessionStart フックで依存(defuddle/jsdom/pdfjs-dist/undici)の更新を確認する。npm test 通過時のみ採用、失敗時は自動ロールバックする。手動実行は npm run update:deps。詳細は README 参照。
PDFやGitHubリポジトリなど、Defuddleだけでは不十分な対象では適切な専用手段を併用する。
本Skillは defuddle の実API検証に基づく(0.6.x で検証、0.18.x で再確認。依存は自動更新)。当初設計(linkedom /
useAsync:false)からの逸脱:
defuddle/node は jsdom 専用(peerDependency)。linkedom では
getComputedStyle/メディアクエリ評価が未実装で例外となり、ノイズ除去に失敗する(実証済み)。
→ jsdom を採用。useAsync は存在しない(0.6.x–0.18.x の dist 全体に出現なし)。意図(非同期外部取得をしない)は
「同期 parse() + resources:'usable' を付けないDOM + スクリプト非実行」で構造的に担保。defuddle/node に文字列を渡すと内部JSDOMが resources:'usable' で外部フェッチするため、
本Skillは自前で安全オプションのJSDOMを構築して渡し、外部取得を防いでいる。© cloudnative-co, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 17 other files (scripts) in skills/web-content-extraction of cloudnative-co/claude-code-starter-kit.
Open the folder on GitHubat commit f00e7ce
Web Content Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Web Content Extraction this skillcloudnative-co/claude-code-starter-kit | 151 | — | ~1.2k | Automated safety check: Pass | MIT | |
| Clean Content FetchLeoYeAI/openclaw-master-skills | 2.2k | — | ~574 | Automated safety check: Pass | MIT | |
| X to Markdown ConverterJimLiu/baoyu-skills | 26k | 4 repos | ~1.8k | Automated safety check: Warn | MIT | |
| Read URLs and PDFstw93/Waza | 7.2k | — | ~1.8k | Automated safety check: Pass | MIT | |
| Canghe URL To Markdownfreestylefly/canghe-skills | 461 | 4 repos | ~1.1k | Automated safety check: Pass | None | |
| URL to Markdown Fetcherjoeseesun/qiaomu-markdown-proxy | 508 | — | ~1.4k | Automated safety check: Pass | MIT |
LeoYeAI/openclaw-master-skills
获取干净、可读的网页正文内容,适合现代网页、博客、新闻、公告和微信公众号文章抓取;支持网页正文提取、内容清洗、去噪、Markdown 输出,适用于普通 fetch 效果不佳、页面噪音较多或动态渲染干扰的场景。Clean content fetch for modern web pages, article extraction, WeChat article capture, content…
JimLiu/baoyu-skills
Saves tweets, threads and X Articles as Markdown files with YAML front matter, using an unofficial API that asks for your consent first.
tw93/Waza
Fetches web pages and PDFs and returns a source-grounded summary, clean Markdown, quotes or citations, routing each kind of link to a suitable fetch method.
freestylefly/canghe-skills
Fetch any URL and convert to markdown using Chrome CDP. An agent skill from freestylefly/canghe-skills.
joeseesun/qiaomu-markdown-proxy
Converts a URL or PDF into clean Markdown, with dedicated routes for WeChat articles, Feishu docs, arXiv papers and login-gated pages, before any summary or rewrite.
rookie-ricardo/erduo-skills
Convert a web URL into cleaned Markdown with deterministic routing.
cloudnative-co/claude-code-starter-kit
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.
cloudnative-co/claude-code-starter-kit
A skill your agent uses when the user explicitly asks for TDD or a tests-first workflow, or when developing a new feature with a test coverage requirement.
cloudnative-co/claude-code-starter-kit
Example project-specific skill template. An agent skill from cloudnative-co/claude-code-starter-kit.
cloudnative-co/claude-code-starter-kit
Practical prompt patterns and techniques for effective Claude Code usage.
cloudnative-co/claude-code-starter-kit
Comprehensive verification system for Claude Code sessions. An agent skill from cloudnative-co/claude-code-starter-kit.
cloudnative-co/claude-code-starter-kit
日本語の業務文書を作成・修正するときに使用する共通品質基準。事実性、確度、論理、簡潔さ、自然な日本語を守る。提案書、報告、技術説明、議事録、メール、Slack、要約、レビューに適用する。創作、広告コピー、コードや構造化データだけの生成には使用しない。
Categories
Web取得の標準前処理レイヤー。URL・公式ドキュメント・ブログ・ニュース・OSSページを読むときは、原則として毎回このSkillでDefuddleを使い本文をMarkdown/JSON化してから読む。生HTMLのまま要約・分析・比較・レビューしない。トリガー例: URLを読む/Webページ要約/OSS調査/公式Doc確認/記事解析/競合サイト確認/Web一次情報確認。. Web Content Extraction is an agent skill from cloudnative-co/claude-code-starter-kit.
Web Content Extraction fits situations like: tasks that involve Web clipping and read-later.
Run `npx skills add cloudnative-co/claude-code-starter-kit --skill web-content-extraction -a claude-code`. Or copy the skill folder (skills/web-content-extraction in cloudnative-co/claude-code-starter-kit) into .claude/skills/web-content-extraction in your project. Claude Code loads it when a task matches its description.
Run `npx skills add cloudnative-co/claude-code-starter-kit --skill web-content-extraction -a codex`. Or copy the skill folder (skills/web-content-extraction in cloudnative-co/claude-code-starter-kit) into .agents/skills/web-content-extraction in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cloudnative-co/claude-code-starter-kit --skill web-content-extraction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/web-content-extraction, .gemini/skills/web-content-extraction, .github/skills/web-content-extraction and .opencode/skills/web-content-extraction in your project.
Going by SKILL.md and its folder, Web Content Extraction needs JavaScript and a shell for the scripts in its folder and the command-line tools its instructions call (npm). Our summary lists: Node.js; A Bash shell.
SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Web Content Extraction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.2k tokens (SKILL.md is roughly 4.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Web Content Extraction: Clean Content Fetch (LeoYeAI/openclaw-master-skills, 2.2k stars), X to Markdown Converter (JimLiu/baoyu-skills, 26k stars), Read URLs and PDFs (tw93/Waza, 7.2k stars) and Canghe URL To Markdown (freestylefly/canghe-skills, 461 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
cloudnative-co (a GitHub organization) maintains it in cloudnative-co/claude-code-starter-kit, which has 151 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 7, 2026.
Source: cloudnative-co/claude-code-starter-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.