Agent skill

URL to Markdown Fetcher

by JimLiu in JimLiu/baoyu-skills

Fetches a web page, X post, YouTube transcript or Hacker News thread through a Chrome-driven CLI and saves it as clean markdown.

MITAuto-check passedKnowledge Management

Install URL to Markdown Fetcher

skills CLI
$ npx skills add JimLiu/baoyu-skills --skill baoyu-url-to-markdown -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install JimLiu/baoyu-skills baoyu-url-to-markdown --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/JimLiu/baoyu-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/baoyu-url-to-markdown .claude/skills/baoyu-url-to-markdown && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
baoyu-url-to-markdown
GitHub stars
27k
Used in
1 other repo
Token cost
~2.1k tokens
SKILL.md length
814 words
Files
55 (incl. scripts, references)
Skills in repo
22
Repo updated
First seen
Licence
MIT

At a glance

Fetches a web page, X post, YouTube transcript or Hacker News thread through a Chrome-driven CLI and saves it as clean markdown.

  • Works in 3 steps: Prefer built-in user-input tools exposed… → Fallback: if no such tool exists, emit a… → Batching: if the tool supports multiple…
  • Saving a web article as a markdown file
  • SKILL.md covers User Input Tools, CLI Setup, Preferences (EXTEND.md) and Usage, plus 6 more sections
  • Runs TypeScript scripts from its folder

What it does

The skill runs the vendored `baoyu-fetch` command-line tool, which drives Chrome over the DevTools protocol and applies site-specific adapters. Built-in adapters cover X (Twitter), YouTube transcripts and Hacker News threads, and everything else goes through a generic adapter built on Defuddle. Pages that need a login or CAPTCHA can be handled with interaction wait modes.

The CLI needs Bun; when `scripts/node_modules` is missing, the agent installs dependencies with `bun install`. Preferences live in an `EXTEND.md` file, looked up in the project folder, then the XDG config folder, then the home folder, and they cover whether to download media and the default output directory. If none exists, a blocking first-time setup asks you questions through the runtime's user-input tool before anything is saved. Reference files describe the adapters and a quality gate.

When your agent uses it

  • Saving a web article as a markdown file
  • Capturing a YouTube transcript or a Hacker News discussion as notes
  • Archiving an X post or article thread in markdown

Example prompts

  • “Save this article as markdown and download its images into my notes folder.”
  • “Get the transcript of this YouTube video as a markdown file.”
  • “Convert the Hacker News thread I just pasted into a markdown file.”

Requirements

  • Bun
  • Chrome
  • Network access to the pages you want to fetch

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Prefer built-in user-input tools exposed by the current agent runtime — e.g., AskUserQuestion, request_user_input, clarify, ask_user, or…
  2. Fallback: if no such tool exists, emit a numbered plain-text message and ask the user to reply with the chosen number/answer for each…
  3. Batching: if the tool supports multiple questions per call, combine all applicable questions into a single call; if only single-question…

What it can do on your machine

Read from SKILL.md and the folder at commit 1567581. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 9 files in scripts/ (TypeScript, from the files we listed), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

URL to Markdown Fetcher loads about 2.1k tokens when it runs, and up to ~4.2k if it reads all its reference files. Until then it costs about 83 tokens; SKILL.md has 814 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from JimLiu/baoyu-skills at commit 1567581, republished under its MIT licence (© JimLiu). 814 words, ~2,117 tokens.

Download SKILL.mdSave it as .claude/skills/baoyu-url-to-markdown/SKILL.md (or your agent's skills folder). This skill also uses 54 other files; get the full folder from GitHub.
name
baoyu-url-to-markdown
description
Fetch any URL and convert to markdown using baoyu-fetch CLI (Chrome CDP with site-specific adapters). Built-in adapters for X/Twitter, YouTube transcripts, Hacker News threads, and generic pages via Defuddle. Handles login/CAPTCHA via interaction wait modes. Use when user wants to save a webpage as markdown.
version
1.61.0

URL to Markdown

Fetches any URL via baoyu-fetch CLI (Chrome CDP + site-specific adapters) and converts it to clean markdown.

User Input Tools

When this skill prompts the user, follow this tool-selection rule (priority order):

  1. Prefer built-in user-input tools exposed by the current agent runtime — e.g., AskUserQuestion, request_user_input, clarify, ask_user, or any equivalent.
  2. Fallback: if no such tool exists, emit a numbered plain-text message and ask the user to reply with the chosen number/answer for each question.
  3. Batching: if the tool supports multiple questions per call, combine all applicable questions into a single call; if only single-question, ask them one at a time in priority order.

Concrete AskUserQuestion references below are examples — substitute the local equivalent in other runtimes.

CLI Setup

Important: The CLI source is vendored in {baseDir}/scripts/lib. scripts/package.json installs only third-party runtime dependencies.

Agent Execution Instructions:

  1. Determine this SKILL.md file's directory path as {baseDir}
  2. Resolve ${BUN} runtime: if bun installed → bun; else suggest installing Bun
  3. If {baseDir}/scripts/node_modules does not exist, run ${BUN} install --cwd {baseDir}/scripts
  4. ${READER} = {baseDir}/scripts/baoyu-fetch
  5. Replace all ${READER} in this document with the resolved value

Preferences (EXTEND.md)

Check EXTEND.md in priority order — the first one found wins:

PriorityPathScope
1.baoyu-skills/baoyu-url-to-markdown/EXTEND.mdProject
2${XDG_CONFIG_HOME:-$HOME/.config}/baoyu-skills/baoyu-url-to-markdown/EXTEND.mdXDG
3$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.mdUser home
ResultAction
FoundRead, parse, apply settings
Not foundMUST run first-time setup (see below) — do NOT silently create defaults

EXTEND.md supports: download media by default, default output directory.

First-Time Setup ⛔ BLOCKING

When EXTEND.md is not found, you MUST use AskUserQuestion to gather preferences before creating EXTEND.md. NEVER create EXTEND.md with silent defaults. Generation is BLOCKED until setup completes. Batch all three questions into a single call:

  • Q1 — Media (header "Media"): "How to handle images and videos in pages?"
    • "Ask each time (Recommended)" — Prompt after each save
    • "Always download" — Download to local imgs/ and videos/
    • "Never download" — Keep remote URLs
  • Q2 — Output (header "Output"): "Default output directory?"
    • "url-to-markdown (Recommended)" — Save to ./url-to-markdown/{domain}/{slug}.md
    • User may pick "Other" and type a custom path
  • Q3 — Save (header "Save"): "Where to save preferences?"
    • "User (Recommended)" — ~/.baoyu-skills/ (all projects)
    • "Project" — .baoyu-skills/ (this project only)

After answers, write EXTEND.md, confirm "Preferences saved to [path]", then continue.

Full template: references/config/first-time-setup.md.

Supported Keys
KeyDefaultValuesDescription
download_mediaaskask / 1 / 0ask = prompt each time, 1 = always, 0 = never
default_output_diremptypath or emptyDefault output directory (empty = ./url-to-markdown/)

EXTEND.md → CLI mapping:

EXTEND.md keyCLI argumentNotes
download_media: 1--download-mediaRequires --output to be set
default_output_dir: ./posts/Agent constructs --output ./posts/{domain}/{slug}.mdAgent generates path, not a direct flag

Value priority: CLI arguments → EXTEND.md → skill defaults.

Usage

bash
# Default: headless capture, markdown to stdout
${READER} <url>

# Save to file
${READER} <url> --output article.md

# Save with media download
${READER} <url> --output article.md --download-media

# Wait for interaction (login/CAPTCHA) — auto-detect and continue
${READER} <url> --wait-for interaction --output article.md

# Wait for interaction — manual control (Enter to continue)
${READER} <url> --wait-for force --output article.md

# JSON output
${READER} <url> --format json --output article.json

# Force specific adapter
${READER} <url> --adapter youtube --output transcript.md
Show full SKILL.md (384 more words)Show less

Options

OptionDescription
<url>URL to fetch
--output <path>Output file path (default: stdout)
--format <type>Output format: markdown (default) or json
--jsonShorthand for --format json
--adapter <name>Force adapter: x, youtube, hn, or generic (default: auto-detect)
--headlessForce headless Chrome (no visible window)
--wait-for <mode>Interaction wait mode: none (default), interaction, or force
--wait-for-interactionAlias for --wait-for interaction
--wait-for-loginAlias for --wait-for interaction
--timeout <ms>Page load timeout (default: 30000)
--interaction-timeout <ms>Login/CAPTCHA wait timeout (default: 600000 = 10 min)
--interaction-poll-interval <ms>Poll interval for interaction checks (default: 1500)
--download-mediaDownload images/videos to local imgs/ and videos/, rewrite markdown links. Requires --output
--media-dir <dir>Base directory for downloaded media (default: same as --output directory)
--cdp-url <url>Reuse existing Chrome DevTools Protocol endpoint
--browser-path <path>Custom Chrome/Chromium binary path
--chrome-profile-dir <path>Chrome user data directory (default: BAOYU_CHROME_PROFILE_DIR env or ./baoyu-skills/chrome-profile)
--debug-dir <dir>Write debug artifacts (document.json, markdown.md, page.html, network.json)

Agent Quality Gate

CRITICAL: treat default headless capture as provisional. Some sites render differently in headless mode and can silently return low-quality content without failing the CLI.

After every headless run, inspect the saved markdown. See references/quality-gate.md for the full checklist, recovery workflow, and capture-mode table. Read it whenever a run looks suspicious or the user asks about login/CAPTCHA handling.

Output Path Generation

The agent must construct the output file path — baoyu-fetch does not auto-generate paths.

Algorithm:

  1. Determine base directory from EXTEND.md default_output_dir or default ./url-to-markdown/
  2. Extract domain from URL (e.g., example.com)
  3. Generate slug from URL path or page title (kebab-case, 2-6 words)
  4. Construct: {base_dir}/{domain}/{slug}/{slug}.md — each URL gets its own directory so media files stay isolated
  5. Conflict resolution: append timestamp {slug}-YYYYMMDD-HHMMSS/{slug}-YYYYMMDD-HHMMSS.md

Pass the constructed path to --output. Media files (--download-media) are saved into subdirectories next to the markdown file, keeping each URL's assets self-contained.

Adapters & Media

See references/adapters.md for the adapter catalog (X, YouTube, Hacker News, generic), per-adapter notes, the media download flow (ask / always / never), and the JSON output schema. Read it before answering adapter-specific questions or handling media prompts.

Environment Variables

VariableDescription
BAOYU_CHROME_PROFILE_DIRChrome user data directory (can also use --chrome-profile-dir)

Troubleshooting: Chrome not found → use --browser-path. Timeout → increase --timeout. Login/CAPTCHA → --wait-for interaction. Debug → --debug-dir to inspect captured HTML and network logs.

Extension Support

Custom configurations via EXTEND.md. See Preferences section above for paths and supported keys.

© JimLiu, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 54 other files (scripts, references) in skills/baoyu-url-to-markdown of JimLiu/baoyu-skills.

  • SKILL.md
  • references/adapters.md
  • references/config/first-time-setup.md
  • references/quality-gate.md
  • scripts/baoyu-fetch
  • scripts/bun.lock
  • scripts/lib/adapters/generic/index.ts
  • scripts/lib/adapters/hn/index.ts
  • scripts/lib/adapters/index.ts
  • scripts/lib/adapters/types.ts
  • scripts/lib/adapters/x/article.ts
  • scripts/lib/adapters/x/index.ts
  • scripts/lib/adapters/x/login.ts
  • … and 42 more

Open the folder on GitHubat commit 1567581

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in JimLiu/baoyu-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

URL to Markdown Fetcher next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

URL to Markdown Fetcher compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
URL to Markdown Fetcher this skillJimLiu/baoyu-skills27k1 repos~2.1kAutomated safety check: PassMIT
Web To Markdownrookie-ricardo/erduo-skills935—~894Automated safety check: PassMIT
Baoyu URL To Markdownsdyckjq-lab/llm-wiki-skill2.5k2 repos~3.2kAutomated safety check: PassNone
Multi-Source to NotebookLM Processorjoeseesun/qiaomu-anything-to-notebooklm6.2k—~3.6kAutomated safety check: PassMIT
Youtube Transcriptbrowser-act/skills6.1k—~2.1kAutomated safety check: PassMIT
YouTube Talk Notetakerdair-ai/dair-academy-plugins614—~2.3kAutomated safety check: PassMIT

Similar skills

  • Web To Markdown

    rookie-ricardo/erduo-skills

    Convert a web URL into cleaned Markdown with deterministic routing.

    935 GitHub stars~894 tokensUpdated 2 mo ago
    Productivity & AutomationAuto-check passed
  • Baoyu URL To Markdown

    sdyckjq-lab/llm-wiki-skill

    Fetch any URL and convert to markdown using Chrome CDP. An agent skill from sdyckjq-lab/llm-wiki-skill.

    2.5k GitHub starsUsed in 2 repos~3.2k tokens
    Knowledge ManagementAuto-check passed
  • Multi-Source to NotebookLM Processor

    joeseesun/qiaomu-anything-to-notebooklm

    Collects content from WeChat articles, web pages, YouTube, podcasts, documents and more, uploads it to NotebookLM and generates podcasts, slides or mind maps.

    6.2k GitHub stars~3.6k tokensUpdated 6 days ago
    Knowledge ManagementAuto-check passed
  • Youtube Transcript

    browser-act/skills

    YouTube transcript extraction and content reformatting: given a YouTube video URL, opens the video's transcript panel, extracts all timestamped segments, and transforms the raw transcript into…

    6.1k GitHub stars~2.1k tokensUpdated 1 mo ago
    Knowledge ManagementAuto-check passed
  • YouTube Talk Notetaker

    dair-ai/dair-academy-plugins

    Converts a YouTube talk into a markdown study note with slide images, a timestamped transcript and editable notes, browsable through a small local server.

    614 GitHub stars~2.3k tokensUpdated 2 mo ago
    Knowledge ManagementAuto-check passed
  • URL to Markdown Fetcher

    joeseesun/qiaomu-markdown-proxy

    Converts a URL or PDF into clean Markdown, with dedicated routes for WeChat articles, Feishu docs, arXiv papers and login-gated pages, before any summary or rewrite.

    509 GitHub stars~1.4k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed

More from JimLiu/baoyu-skills

All 22 skills in this repo
  • Markdown Article Formatter

    JimLiu/baoyu-skills

    Reformats plain text or Markdown articles with frontmatter, a title, a summary, headings, bold, lists and code blocks, and saves a separate formatted copy.

    27k GitHub starsUsed in 6 repos~3.5k tokens
    Auto-check passed
  • X to Markdown Converter

    JimLiu/baoyu-skills

    Saves tweets, threads and X Articles as Markdown files with YAML front matter, using an unofficial API that asks for your consent first.

    27k GitHub starsUsed in 4 repos~1.8k tokens
    Auto-check: warnings
  • SVG Diagram Generator

    JimLiu/baoyu-skills

    Creates standalone dark-themed SVG diagrams, including architecture, flowchart, sequence, structural, mind map, timeline and state machine types.

    27k GitHub starsUsed in 1 repo~3.1k tokens
    Auto-check passed
  • Publishes articles and image-text posts to a WeChat Official Account through the API or Chrome CDP, converting markdown to WeChat-ready HTML with link citations.

    27k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check: warnings
  • Image Compressor

    JimLiu/baoyu-skills

    Compresses images to WebP by default, or to PNG or JPEG, picking the best available tool on the machine and optionally processing whole folders.

    27k GitHub starsUsed in 5 repos~598 tokens
    Auto-check passed
  • Generates text and images through an unofficial, reverse-engineered Gemini Web API, supporting reference images and multi-turn conversations.

    27k GitHub starsUsed in 5 repos~1.6k tokens
    Auto-check passed

Questions about URL to Markdown Fetcher

What does URL to Markdown Fetcher do?

Fetches a web page, X post, YouTube transcript or Hacker News thread through a Chrome-driven CLI and saves it as clean markdown. The skill runs the vendored `baoyu-fetch` command-line tool, which drives Chrome over the DevTools protocol and applies site-specific adapters. Built-in adapters cover X (Twitter), YouTube transcripts and Hacker News threads, and everything else goes through a generic adapter built on Defuddle.

When should I use URL to Markdown Fetcher?

URL to Markdown Fetcher fits situations like: saving a web article as a markdown file; capturing a YouTube transcript or a Hacker News discussion as notes; archiving an X post or article thread in markdown.

How do I install URL to Markdown Fetcher in Claude Code?

Run `npx skills add JimLiu/baoyu-skills --skill baoyu-url-to-markdown -a claude-code`. Or copy the skill folder (skills/baoyu-url-to-markdown in JimLiu/baoyu-skills) into .claude/skills/baoyu-url-to-markdown in your project. Claude Code loads it when a task matches its description.

How do I install URL to Markdown Fetcher in Codex?

Run `npx skills add JimLiu/baoyu-skills --skill baoyu-url-to-markdown -a codex`. Or copy the skill folder (skills/baoyu-url-to-markdown in JimLiu/baoyu-skills) into .agents/skills/baoyu-url-to-markdown in your project. Codex loads it when a task matches its description.

Can I use URL to Markdown Fetcher in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add JimLiu/baoyu-skills --skill baoyu-url-to-markdown -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/baoyu-url-to-markdown, .gemini/skills/baoyu-url-to-markdown, .github/skills/baoyu-url-to-markdown and .opencode/skills/baoyu-url-to-markdown in your project.

What does URL to Markdown Fetcher need to run?

Going by SKILL.md and its folder, URL to Markdown Fetcher needs TypeScript for the scripts in its folder. Our summary lists: Bun; Chrome; Network access to the pages you want to fetch.

Does URL to Markdown Fetcher access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is URL to Markdown Fetcher safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does URL to Markdown Fetcher use?

URL to Markdown Fetcher is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does URL to Markdown Fetcher use?

About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2k tokens, read only when the agent opens those files.

What are the alternatives to URL to Markdown Fetcher?

Skills that share tags, products or a category with URL to Markdown Fetcher: Web To Markdown (rookie-ricardo/erduo-skills, 935 stars), Baoyu URL To Markdown (sdyckjq-lab/llm-wiki-skill, 2.5k stars), Multi-Source to NotebookLM Processor (joeseesun/qiaomu-anything-to-notebooklm, 6.2k stars) and Youtube Transcript (browser-act/skills, 6.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains URL to Markdown Fetcher?

JimLiu (a GitHub user) maintains it in JimLiu/baoyu-skills, which has 26,507 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on September 10, 2026.

Source: JimLiu/baoyu-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.