Agent skill

Web To Markdown

by rookie-ricardo in rookie-ricardo/erduo-skills

Convert a web URL into cleaned Markdown with deterministic routing.

MITAuto-check passedProductivity & Automation

Install Web To Markdown

skills CLI
$ npx skills add rookie-ricardo/erduo-skills --skill web-to-markdown -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install rookie-ricardo/erduo-skills web-to-markdown --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/rookie-ricardo/erduo-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/web-to-markdown .claude/skills/web-to-markdown && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
web-to-markdown
GitHub stars
935
Token cost
~894 tokens
SKILL.md length
315 words
Files
9 (incl. scripts, references)
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Convert a web URL into cleaned Markdown with deterministic routing.

  • Works in 2 steps: Normalize and validate the input URL. → Select route
  • Use defuddle.md for YouTube links
  • SKILL.md covers Quick Workflow, Commands, Routing Policy and Special-Site Extraction Behavior, plus 2 more sections
  • Runs JavaScript scripts from its folder; calls node and npm; reaches r.jina.ai and defuddle.md

What it does

Web To Markdown is an agent skill from rookie-ricardo/erduo-skills. Convert a web URL into cleaned Markdown with deterministic routing. Use when Codex needs to read article-like content from links and should apply source-aware fetch strategies: default to r.jina.ai for general pages (including X/Twitter), use defuddle.md for YouTube links, and use browser-impersonated extraction for WeChat/Zhihu/Feishu pages with Mozilla Readability cleanup.

Its SKILL.md is about 890 tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including scripts and reference files (for example `agents/openai.yaml`, `package-lock.json` and `package.json`).

It sits in Productivity & Automation, covering Messaging and chat bots, Markdown and Plain language and style rules. It works with X (Twitter), YouTube, Feishu (Lark) and WeChat. The licence is MIT.

When your agent uses it

  • Use defuddle.md for YouTube links
  • Use browser-impersonated extraction for WeChat/Zhihu/Feishu pages with Mozilla Readability cleanup

Example prompts

  • “/web-to-markdown”

Requirements

  • Node.js

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Normalize and validate the input URL.
  2. Select route

What it can do on your machine

Read from SKILL.md and the folder at commit 1fb6ec5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • node
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • r.jina.ai
    • defuddle.md

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Web To Markdown loads about 894 tokens when it runs, and up to ~1.4k if it reads all its reference files. Until then it costs about 98 tokens; SKILL.md has 315 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~98
When it runs · the whole SKILL.md, loaded when a task matches
~894
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from rookie-ricardo/erduo-skills at commit 1fb6ec5, republished under its MIT licence (© rookie-ricardo). 315 words, ~894 tokens.

Download SKILL.mdSave it as .claude/skills/web-to-markdown/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
web-to-markdown
description
Convert a web URL into cleaned Markdown with deterministic routing. Use when Codex needs to read article-like content from links and should apply source-aware fetch strategies: default to r.jina.ai for general pages (including X/Twitter), use defuddle.md for YouTube links, and use browser-impersonated extraction for WeChat/Zhihu/Feishu pages with Mozilla Readability cleanup.

Web To Markdown

Convert URLs into usable Markdown by applying domain-aware fetching routes, then return the cleaned content directly.

Quick Workflow

  1. Normalize and validate the input URL.
  2. Select route:
  • r.jina.ai: general web + X/Twitter.
  • defuddle.md: YouTube transcript/content extraction.
  • special-browser-fetch: WeChat/Zhihu/Feishu.
  1. Return markdown text (or JSON metadata if needed).

For generic URLs (non-YouTube, non-WeChat/Zhihu/Feishu), use this fallback chain:

  • try r.jina.ai first,
  • if it fails, fallback to direct HTTP fetch + Readability,
  • if direct fetch still fails or returns shell-like content, fallback to browser extraction.

Commands

Run from this skill directory (skills/web-to-markdown):

bash
npm install
node scripts/url_to_markdown.mjs <url>

Return metadata with markdown:

bash
node scripts/url_to_markdown.mjs <url> --json

Force special-site browser extraction:

bash
node scripts/fetch_special_sites.mjs <url> --json

Routing Policy

  • Default route: https://r.jina.ai/<url>.
  • YouTube (youtube.com, youtu.be): https://defuddle.md/<url>.
  • X/Twitter (x.com, twitter.com): https://r.jina.ai/<url>.
  • WeChat/Zhihu/Feishu: run scripts/fetch_special_sites.mjs.
  • If input is already proxy-formatted (https://defuddle.md/https://... or https://r.jina.ai/https://...), normalize back to the original URL and re-apply routing.

Special-Site Extraction Behavior

Use a two-stage strategy for WeChat/Zhihu/Feishu:

  1. Try cuimp HTTP/TLS impersonation first, then clean HTML with Mozilla Readability.
  2. If stage 1 fails or returns blocked/shell content, fallback to puppeteer-extra browser impersonation.
  • HTTP stage impersonates modern Chrome TLS/HTTP profile via cuimp.
  • Browser stage impersonates a modern Chrome user agent and standard sec-ch-ua headers.
  • Remove known login modals and backdrop overlays (best effort).
  • Scroll the page to trigger lazy-loaded article blocks.
  • Parse cleaned document with Mozilla Readability.
  • Convert extracted HTML body to Markdown via Turndown.
  • Resolve browser executable from CHROME_PATH first, then system Chrome/Chromium/Edge paths.

If special-site extraction fails due to anti-bot checks, account-only pages, or network limits, report failure clearly and ask for fallback input (for example raw page text).

Output Contract

For normal usage, output markdown only.

When --json is used, return:

  • source: backend source (r.jina.ai, defuddle, cuimp, browser-readability).
  • strategy: selected route (r-jina, defuddle, special-http-fetch, special-browser-fetch-fallback).
  • requestedUrl: original input.
  • resolvedUrl: normalized/final URL.
  • markdown: extracted markdown body.

Resources

  • references/routing-and-notes.md: domain routing rules and operational caveats.
  • scripts/url_to_markdown.mjs: primary entrypoint.
  • scripts/fetch_special_sites_http.mjs: WeChat/Zhihu/Feishu HTTP impersonation fetcher (cuimp JS).
  • scripts/fetch_special_sites.mjs: two-stage extractor (HTTP-first, browser-fallback).

© rookie-ricardo, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references) in skills/web-to-markdown of rookie-ricardo/erduo-skills.

  • SKILL.md
  • agents/openai.yaml
  • package-lock.json
  • package.json
  • references/routing-and-notes.md
  • scripts/fetch_generic_fallback.mjs
  • scripts/fetch_special_sites.mjs
  • scripts/fetch_special_sites_http.mjs
  • scripts/url_to_markdown.mjs

Open the folder on GitHubat commit 1fb6ec5

Compare with similar skills

Web To Markdown next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Web To Markdown compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Web To Markdown this skillrookie-ricardo/erduo-skills935—~894Automated safety check: PassMIT
Feedgrab BatchiBigQiang/feedgrab614—~1.8kAutomated safety check: PassMIT
URL to Markdown Fetcherjoeseesun/qiaomu-markdown-proxy509—~1.4kAutomated safety check: PassMIT
Clean Content FetchLeoYeAI/openclaw-master-skills2.2k—~574Automated safety check: PassMIT
Content CollectorvigorX777/content-collector-skill240—~2kAutomated safety check: PassNone
Agent ReachEdisonChenAI/agent-reach1151 repos~1.3kAutomated safety check: PassMIT

Similar skills

  • Feedgrab Batch

    iBigQiang/feedgrab

    Batch content grabber — bulk fetch bookmarks, user tweets, search results, author notes, wiki pages, subreddit posts, and more from X/Twitter, Xiaohongshu, Reddit, WeChat, YouTube, Feishu, Zhihu.

    614 GitHub stars~1.8k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check passed
  • URL to Markdown Fetcher

    joeseesun/qiaomu-markdown-proxy

    Converts a URL or PDF into clean Markdown, with dedicated routes for WeChat articles, Feishu docs, arXiv papers and login-gated pages, before any summary or rewrite.

    509 GitHub stars~1.4k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Clean Content Fetch

    LeoYeAI/openclaw-master-skills

    获取干净、可读的网页正文内容,适合现代网页、博客、新闻、公告和微信公众号文章抓取;支持网页正文提取、内容清洗、去噪、Markdown 输出,适用于普通 fetch 效果不佳、页面噪音较多或动态渲染干扰的场景。Clean content fetch for modern web pages, article extraction, WeChat article capture, content…

    2.2k GitHub stars~574 tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Content Collector

    vigorX777/content-collector-skill

    Collect social media content (X/Twitter, WeChat, Jike, Reddit, etc.) into Feishu bitable.

    240 GitHub stars~2k tokensUpdated 6 mo ago
    Productivity & AutomationAuto-check passed
  • Agent Reach

    EdisonChenAI/agent-reach

    Use the internet: search, read, and interact with 13+ platforms including Twitter/X, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu (小红书), Douyin (抖音), WeChat Articles (微信公众号), LinkedIn, Boss直聘…

    115 GitHub starsUsed in 1 repo~1.3k tokens
    Productivity & AutomationAuto-check passed
  • Viral Title

    kangarooking/kangarooking-skills

    Generate high-potential viral title candidates for content across WeChat public account articles, X/Twitter posts, YouTube videos, Bilibili videos, and similar content platforms.

    662 GitHub stars~1.7k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check passed

More from rookie-ricardo/erduo-skills

  • Anthropic Style Diagram

    rookie-ricardo/erduo-skills

    Draw architecture, flow and structural diagrams in the Anthropic/Claude visual language as SVG, then render and save them as PNG.

    935 GitHub stars~3k tokensUpdated 2 mo ago
    Auto-check passed
  • Ak Rss Digest

    rookie-ricardo/erduo-skills

    Curate a Chinese reading digest from a fixed bundle of RSS and Atom feeds, with a strong preference for AI agent thinking, frontier AI commentary, deep interviews, and non-boring high-signal essays.

    935 GitHub stars~1.4k tokensUpdated 2 mo ago
    Auto-check passed
  • Gemini Watermark Remover

    rookie-ricardo/erduo-skills

    Remove the visible Gemini AI watermark from images using reverse alpha blending.

    935 GitHub stars~359 tokensUpdated 2 mo ago
    Auto-check passed
  • Translate Polisher

    rookie-ricardo/erduo-skills

    高质量文章翻译技能,采用"分析→初译→审校→终稿"四步精翻工作流。仅支持中文↔英文、中文↔日文翻译。当用户明确提出"翻译"、"translate"、"精翻"、"翻訳"、"翻译文章"、"translate to Chinese/English/Japanese"、"改成中文"、"改成英文"、"改成日文"、"翻成中文"、"翻成日文"、"翻成英文"、"英译中"、"中译英"、"中译日"、"日译中"、"日…

    935 GitHub stars~3.6k tokensUpdated 2 mo ago
    Auto-check passed
  • Transcript Polisher

    rookie-ricardo/erduo-skills

    将语音转录文本(访谈、演讲、播客、会议)精修为可读性更高的文章段落。当用户提到"字幕精修"、"transcript polish"、"润色字幕"、"把视频字幕整理成文章"、"访谈文字整理"、处理访谈记录、转录文本优化、语音转文字整理、或者需要将大段对话/演讲文本整理成可读文章时触发。适用于单人演说或多人对谈的转录文本整理,要求保留原句原词、拒绝高度概括。即使用户只是说"帮我整理一下这段文字"并附…

    935 GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Daily News Report

    rookie-ricardo/erduo-skills

    基于预设 URL 列表抓取内容,筛选高质量技术信息并生成每日 Markdown 报告. An agent skill from rookie-ricardo/erduo-skills.

    935 GitHub stars~2.1k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Web To Markdown

What does Web To Markdown do?

Convert a web URL into cleaned Markdown with deterministic routing. Web To Markdown is an agent skill from rookie-ricardo/erduo-skills. Convert a web URL into cleaned Markdown with deterministic routing.

When should I use Web To Markdown?

Web To Markdown fits situations like: use defuddle.md for YouTube links; use browser-impersonated extraction for WeChat/Zhihu/Feishu pages with Mozilla Readability cleanup.

How do I install Web To Markdown in Claude Code?

Run `npx skills add rookie-ricardo/erduo-skills --skill web-to-markdown -a claude-code`. Or copy the skill folder (skills/web-to-markdown in rookie-ricardo/erduo-skills) into .claude/skills/web-to-markdown in your project. Claude Code loads it when a task matches its description.

How do I install Web To Markdown in Codex?

Run `npx skills add rookie-ricardo/erduo-skills --skill web-to-markdown -a codex`. Or copy the skill folder (skills/web-to-markdown in rookie-ricardo/erduo-skills) into .agents/skills/web-to-markdown in your project. Codex loads it when a task matches its description.

Can I use Web To Markdown in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rookie-ricardo/erduo-skills --skill web-to-markdown -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/web-to-markdown, .gemini/skills/web-to-markdown, .github/skills/web-to-markdown and .opencode/skills/web-to-markdown in your project.

What does Web To Markdown need to run?

Going by SKILL.md and its folder, Web To Markdown needs JavaScript for the scripts in its folder and the command-line tools its instructions call (node and npm). Our summary lists: Node.js.

Does Web To Markdown access the network?

SKILL.md names 2 domains. In commands or code: r.jina.ai and defuddle.md; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Web To Markdown safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Web To Markdown use?

Web To Markdown is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Web To Markdown use?

About 894 tokens (SKILL.md is roughly 3.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 510 tokens, read only when the agent opens those files.

What are the alternatives to Web To Markdown?

Skills that share tags, products or a category with Web To Markdown: Feedgrab Batch (iBigQiang/feedgrab, 614 stars), URL to Markdown Fetcher (joeseesun/qiaomu-markdown-proxy, 509 stars), Clean Content Fetch (LeoYeAI/openclaw-master-skills, 2.2k stars) and Content Collector (vigorX777/content-collector-skill, 240 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Web To Markdown?

rookie-ricardo (a GitHub user) maintains it in rookie-ricardo/erduo-skills, which has 935 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on July 25, 2026.

Source: rookie-ricardo/erduo-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.