Agent skill

Web Content Extraction

by cloudnative-co in cloudnative-co/claude-code-starter-kit

Web取得の標準前処理レイヤー。URL・公式ドキュメント・ブログ・ニュース・OSSページを読むときは、原則として毎回このSkillでDefuddleを使い本文をMarkdown/JSON化してから読む。生HTMLのまま要約・分析・比較・レビューしない。トリガー例: URLを読む/Webページ要約/OSS調査/公式Doc確認/記事解析/競合サイト確認/Web一次情報確認。

MITAuto-check passedDocuments & Office

Install Web Content Extraction

skills CLI
$ npx skills add cloudnative-co/claude-code-starter-kit --skill web-content-extraction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cloudnative-co/claude-code-starter-kit web-content-extraction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cloudnative-co/claude-code-starter-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/web-content-extraction .claude/skills/web-content-extraction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
web-content-extraction
GitHub stars
151
Token cost
~1.2k tokens
SKILL.md length
264 words
Files
18 (incl. scripts)
Skills in repo
8
Repo updated
First seen
Licence
MIT

At a glance

Web取得の標準前処理レイヤー。URL・公式ドキュメント・ブログ・ニュース・OSSページを読むときは、原則として毎回このSkillでDefuddleを使い本文をMarkdown/JSON化してから読む。生HTMLのまま要約・分析・比較・レビューしない。トリガー例: URLを読む/Webページ要約/OSS調査/公式Doc確認/記事解析/競合サイト確認/Web一次情報確認。

  • Tasks that involve Web clipping and read-later
  • SKILL.md covers Purpose, Mandatory Rule, Use Cases and Standard Commands, plus 8 more sections
  • Runs JavaScript and Shell scripts from its folder; calls npm

What it does

Web Content Extraction is an agent skill from cloudnative-co/claude-code-starter-kit. Web取得の標準前処理レイヤー。URL・公式ドキュメント・ブログ・ニュース・OSSページを読むときは、原則として毎回このSkillでDefuddleを使い本文をMarkdown/JSON化してから読む。生HTMLのまま要約・分析・比較・レビューしない。トリガー例: URLを読む/Webページ要約/OSS調査/公式Doc確認/記事解析/競合サイト確認/Web一次情報確認。

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 20 other files, including scripts (for example `README.md`, `package-lock.json` and `package.json`).

It sits in Documents & Office, covering Web clipping and read-later. The repository describes itself as: One-command setup of a complete Claude Code development environment with interactive wizard. The licence is MIT.

When your agent uses it

  • Tasks that involve Web clipping and read-later

Example prompts

  • “/web-content-extraction”

Requirements

  • Node.js
  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit f00e7ce. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 7 files in scripts/ (JavaScript and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Web Content Extraction loads about 1.2k tokens when it runs. Until then it costs about 52 tokens; SKILL.md has 264 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from cloudnative-co/claude-code-starter-kit at commit f00e7ce, republished under its MIT licence (© cloudnative-co). 264 words, ~1,219 tokens.

Download SKILL.mdSave it as .claude/skills/web-content-extraction/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.
name
web-content-extraction
description
Web取得の標準前処理レイヤー。URL・公式ドキュメント・ブログ・ニュース・OSSページを読むときは、原則として毎回このSkillでDefuddleを使い本文をMarkdown/JSON化してから読む。生HTMLのまま要約・分析・比較・レビューしない。トリガー例: URLを読む/Webページ要約/OSS調査/公式Doc確認/記事解析/競合サイト確認/Web一次情報確認。
when_to_use
URL・Webページ・公式ドキュメント・ブログ・ニュース・OSS/GitHubページを読む前に毎回使う。生HTMLを直接読まず、まず Defuddle で本文を Markdown/JSON 化してから読む。URL要約・OSS調査・記事解析・競合確認・Web一次情報確認に対応。

Web Content Extraction

Purpose

Webページ本文を抽出し、LLMが読みやすい Markdown/JSON に整形する。 Claude Code が Webページ、URL、公式ドキュメント、ブログ記事、ニュース記事、OSSページを読む場合は、原則として毎回このSkillを使う。

Mandatory Rule

When reading any public web URL, use Defuddle first.

Do not summarize, analyze, compare, or review a web page from raw HTML unless Defuddle extraction fails or the page type is explicitly unsupported.

Use Cases

  • URL要約 / 公式ドキュメント確認 / OSS調査
  • ブログ記事分析 / ニュース記事分析 / 競合サイト分析
  • ベンダー公式ブログの調査 / Web上の一次情報確認
  • 技術記事の読み取り / LLM・RAG向けのWeb本文抽出

Standard Commands

bash
# 公開URLを取得して本文をMarkdown/JSON化(SSRFガードあり)
~/.claude/skills/web-content-extraction/scripts/run-node.sh \
  ~/.claude/skills/web-content-extraction/scripts/defuddle-url.mjs <url>
bash
# ローカルHTMLファイルを本文抽出(外部通信なし)
~/.claude/skills/web-content-extraction/scripts/run-node.sh \
  ~/.claude/skills/web-content-extraction/scripts/defuddle-file.mjs <file>

出力は JSON(stdout)。最低限 success, url, fetchedAt/parsedAt, title, author, site, domain, published, description, wordCount, content(Markdown) を含む。 warnings / fetchWarnings がある場合は抽出の信頼性に注意する。

Output Fields

フィールド意味
success本文抽出に成功したか(false は抽出失敗/空)
warnings本文が短い・空など低信頼の警告
url / requestedUrl / finalUrl対象URL(リダイレクト後の最終URL含む)
fetchedAt / parsedAt取得・解析時刻(ISO8601, 監査用に必ず保持)
title author site domain published descriptionメタデータ
wordCount語数(空白区切り。日本語は極端に小さく出る)
charCount非空白の文字数(日本語の実分量はこちらで判断)
cjkCharCountCJK文字数(日本語/中国語/韓国語の量の目安)
content本文(HTMLはMarkdown、PDFはプレーンテキスト)
extractorTypeサイト固有抽出器が使われた場合の種別
extractorEnginePDF抽出時のみ "pdf"。pageCount も付く

Security Rules

  • 同期コア + 非フェッチDOM で動作する(useAsync はupstreamに存在しないため意図を構造で担保)。
  • 外部フォールバック・サブリソース外部取得・ページ内スクリプト実行は行わない。
  • 社内URL、顧客URL、認証付きURL、個人情報・機密を含むページを外部送信しない。
  • localhost / プライベートIP(10/8,172.16/12,192.168/16,127/8,169.254/16,100.64/10 等) / .local/.internal 等 / 単一ラベルの内部ホスト名 / 非http(s) / 認証情報付きURL は標準で拒否。
  • IP判定はバイト単位。10進/8進/16進 IPv4 も拒否。IPv6はdefault-deny(グローバルユニキャスト 2000::/3 以外は全拒否。Teredo/site-local/documentation/NAT64/IPv4-mapped/6to4(private埋め込み)等を含む)。
  • 接続IPをpin(guarded undici dispatcher)して DNSリバインディング/TOCTOU を封じる。
  • リダイレクトは手動追従し各ホップを送信前に再検査。本文はストリームでサイズ上限(メモリDoS対策)。
  • 開発用途で明示的に許可する場合のみ ALLOW_PRIVATE_URLS=true(バイパスは stderr に監査記録)。
  • 抽出結果には必ず URL と取得日時を残す。
  • 抽出結果だけを唯一の真実として扱わない。重要な事実は一次情報で再確認する。

Failure Handling

  • Defuddle抽出に失敗した場合は、必ず「Defuddle抽出失敗」と明示する。
  • 不完全な抽出結果(success:false や warnings あり)で断定しない。
  • 必要なら以下の代替手段を検討する:
    • raw HTML取得 / GitHub raw / 公式API / Playwright(MCP) / PDF専用抽出 / 手動確認
  • 代替手段を使った場合は、Defuddleではなく代替手段を使ったことを明示する。

PDF対応

公開URLが PDF(content-type: application/pdf / .pdf / 先頭 %PDF-)の場合、defuddle-url.mjs は自動で pdfjs-dist によるテキスト抽出にフォールバックする(extractorEngine:"pdf", pageCount 付き)。

  • 日本語/CJK PDF も文字化けしない(fsベースのCMapReaderFactoryで packed CMap を解決)。
  • content はMarkdownでなくプレーンテキスト。
  • スキャン画像のみのPDFは charCount:0+警告 → OCRが必要(本Skillの対象外)。
  • ローカルPDFファイルは現状未対応(defuddle-file.mjs はHTML専用)。

Tests

bash
cd ~/.claude/skills/web-content-extraction && ./scripts/run-node.sh --test

test/url-guard.test.mjs(SSRFガード)/ test/defuddle-core.test.mjs(charCount)/ test/extract-smoke.test.mjs(実抽出スモーク: HTML+PDF)/ test/defuddle-url.test.mjs(URL CLI exit code契約)を実行。DNS非依存・オフラインの 決定的テストのみ。CIは .github/workflows/skill-web-content-extraction.yml(Node 22/24 マトリクス)。

自動アップデート

web-content-update feature が有効な場合のみ、SessionStart フックで依存(defuddle/jsdom/pdfjs-dist/undici)の更新を確認する。npm test 通過時のみ採用、失敗時は自動ロールバックする。手動実行は npm run update:deps。詳細は README 参照。

Unsupported or Caution Cases

  • 認証が必要なページ / 社内・顧客・非公開ページ / 個人情報・機密を含むページ
  • JavaScript必須のSPA(初期HTMLに本文がない) / PDF
  • robots.txt や利用規約に反する大量取得

PDFやGitHubリポジトリなど、Defuddleだけでは不十分な対象では適切な専用手段を併用する。

Implementation Notes (重要・現実との差分)

本Skillは defuddle の実API検証に基づく(0.6.x で検証、0.18.x で再確認。依存は自動更新)。当初設計(linkedom / useAsync:false)からの逸脱:

  • linkedom は使用不可: defuddle/node は jsdom 専用(peerDependency)。linkedom では getComputedStyle/メディアクエリ評価が未実装で例外となり、ノイズ除去に失敗する(実証済み)。 → jsdom を採用。
  • useAsync は存在しない(0.6.x–0.18.x の dist 全体に出現なし)。意図(非同期外部取得をしない)は 「同期 parse() + resources:'usable' を付けないDOM + スクリプト非実行」で構造的に担保。
  • defuddle/node に文字列を渡すと内部JSDOMが resources:'usable' で外部フェッチするため、 本Skillは自前で安全オプションのJSDOMを構築して渡し、外部取得を防いでいる。

© cloudnative-co, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 17 other files (scripts) in skills/web-content-extraction of cloudnative-co/claude-code-starter-kit.

  • SKILL.md
  • .gitignore
  • README.md
  • package-lock.json
  • package.json
  • scripts/defuddle-file.mjs
  • scripts/defuddle-url.mjs
  • scripts/lib/defuddle-core.mjs
  • scripts/lib/pdf-extract.mjs
  • scripts/lib/url-guard.mjs
  • scripts/run-node.sh
  • scripts/update-deps.mjs
  • test/defuddle-core.test.mjs
  • test/defuddle-url.test.mjs
  • test/extract-smoke.test.mjs
  • test/update-deps-lock.test.mjs
  • test/update-deps-run.test.mjs
  • test/url-guard.test.mjs

Open the folder on GitHubat commit f00e7ce

Compare with similar skills

Web Content Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Web Content Extraction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Web Content Extraction this skillcloudnative-co/claude-code-starter-kit151—~1.2kAutomated safety check: PassMIT
Clean Content FetchLeoYeAI/openclaw-master-skills2.2k—~574Automated safety check: PassMIT
X to Markdown ConverterJimLiu/baoyu-skills26k4 repos~1.8kAutomated safety check: WarnMIT
Read URLs and PDFstw93/Waza7.2k—~1.8kAutomated safety check: PassMIT
Canghe URL To Markdownfreestylefly/canghe-skills4614 repos~1.1kAutomated safety check: PassNone
URL to Markdown Fetcherjoeseesun/qiaomu-markdown-proxy508—~1.4kAutomated safety check: PassMIT

Similar skills

  • Clean Content Fetch

    LeoYeAI/openclaw-master-skills

    获取干净、可读的网页正文内容,适合现代网页、博客、新闻、公告和微信公众号文章抓取;支持网页正文提取、内容清洗、去噪、Markdown 输出,适用于普通 fetch 效果不佳、页面噪音较多或动态渲染干扰的场景。Clean content fetch for modern web pages, article extraction, WeChat article capture, content…

    2.2k GitHub stars~574 tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • X to Markdown Converter

    JimLiu/baoyu-skills

    Saves tweets, threads and X Articles as Markdown files with YAML front matter, using an unofficial API that asks for your consent first.

    26k GitHub starsUsed in 4 repos~1.8k tokens
    Knowledge ManagementAuto-check: warnings
  • Fetches web pages and PDFs and returns a source-grounded summary, clean Markdown, quotes or citations, routing each kind of link to a suitable fetch method.

    7.2k GitHub stars~1.8k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Canghe URL To Markdown

    freestylefly/canghe-skills

    Fetch any URL and convert to markdown using Chrome CDP. An agent skill from freestylefly/canghe-skills.

    461 GitHub starsUsed in 4 repos~1.1k tokens
    Knowledge ManagementAuto-check passed
  • URL to Markdown Fetcher

    joeseesun/qiaomu-markdown-proxy

    Converts a URL or PDF into clean Markdown, with dedicated routes for WeChat articles, Feishu docs, arXiv papers and login-gated pages, before any summary or rewrite.

    508 GitHub stars~1.4k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Web To Markdown

    rookie-ricardo/erduo-skills

    Convert a web URL into cleaned Markdown with deterministic routing.

    935 GitHub stars~894 tokensUpdated 2 mo ago
    Productivity & AutomationAuto-check passed

More from cloudnative-co/claude-code-starter-kit

All 8 skills in this repo
  • Eval Harness

    cloudnative-co/claude-code-starter-kit

    Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.

    151 GitHub starsUsed in 10 repos~1.3k tokens
    Auto-check passed
  • TDD Workflow

    cloudnative-co/claude-code-starter-kit

    A skill your agent uses when the user explicitly asks for TDD or a tests-first workflow, or when developing a new feature with a test coverage requirement.

    151 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Project Guidelines Example

    cloudnative-co/claude-code-starter-kit

    Example project-specific skill template. An agent skill from cloudnative-co/claude-code-starter-kit.

    151 GitHub stars~446 tokensUpdated today
    Auto-check passed
  • Prompt Patterns

    cloudnative-co/claude-code-starter-kit

    Practical prompt patterns and techniques for effective Claude Code usage.

    151 GitHub stars~887 tokensUpdated today
    Auto-check passed
  • Verification Loop

    cloudnative-co/claude-code-starter-kit

    Comprehensive verification system for Claude Code sessions. An agent skill from cloudnative-co/claude-code-starter-kit.

    151 GitHub stars~733 tokensUpdated today
    Auto-check passed
  • Cloudnative Writing Baseline

    cloudnative-co/claude-code-starter-kit

    日本語の業務文書を作成・修正するときに使用する共通品質基準。事実性、確度、論理、簡潔さ、自然な日本語を守る。提案書、報告、技術説明、議事録、メール、Slack、要約、レビューに適用する。創作、広告コピー、コードや構造化データだけの生成には使用しない。

    151 GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Questions about Web Content Extraction

What does Web Content Extraction do?

Web取得の標準前処理レイヤー。URL・公式ドキュメント・ブログ・ニュース・OSSページを読むときは、原則として毎回このSkillでDefuddleを使い本文をMarkdown/JSON化してから読む。生HTMLのまま要約・分析・比較・レビューしない。トリガー例: URLを読む/Webページ要約/OSS調査/公式Doc確認/記事解析/競合サイト確認/Web一次情報確認。. Web Content Extraction is an agent skill from cloudnative-co/claude-code-starter-kit.

When should I use Web Content Extraction?

Web Content Extraction fits situations like: tasks that involve Web clipping and read-later.

How do I install Web Content Extraction in Claude Code?

Run `npx skills add cloudnative-co/claude-code-starter-kit --skill web-content-extraction -a claude-code`. Or copy the skill folder (skills/web-content-extraction in cloudnative-co/claude-code-starter-kit) into .claude/skills/web-content-extraction in your project. Claude Code loads it when a task matches its description.

How do I install Web Content Extraction in Codex?

Run `npx skills add cloudnative-co/claude-code-starter-kit --skill web-content-extraction -a codex`. Or copy the skill folder (skills/web-content-extraction in cloudnative-co/claude-code-starter-kit) into .agents/skills/web-content-extraction in your project. Codex loads it when a task matches its description.

Can I use Web Content Extraction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cloudnative-co/claude-code-starter-kit --skill web-content-extraction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/web-content-extraction, .gemini/skills/web-content-extraction, .github/skills/web-content-extraction and .opencode/skills/web-content-extraction in your project.

What does Web Content Extraction need to run?

Going by SKILL.md and its folder, Web Content Extraction needs JavaScript and a shell for the scripts in its folder and the command-line tools its instructions call (npm). Our summary lists: Node.js; A Bash shell.

Does Web Content Extraction access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Web Content Extraction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Web Content Extraction use?

Web Content Extraction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Web Content Extraction use?

About 1.2k tokens (SKILL.md is roughly 4.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Web Content Extraction?

Skills that share tags, products or a category with Web Content Extraction: Clean Content Fetch (LeoYeAI/openclaw-master-skills, 2.2k stars), X to Markdown Converter (JimLiu/baoyu-skills, 26k stars), Read URLs and PDFs (tw93/Waza, 7.2k stars) and Canghe URL To Markdown (freestylefly/canghe-skills, 461 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Web Content Extraction?

cloudnative-co (a GitHub organization) maintains it in cloudnative-co/claude-code-starter-kit, which has 151 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 7, 2026.

Source: cloudnative-co/claude-code-starter-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.