Agent skill

Fetch URL As Markdown

by CodeAlive-AI in CodeAlive-AI/ai-driven-development

Fetch a web page (URL) and return clean Markdown via local trafilatura, with Exa MCP as a fallback for JS-rendered or anti-bot pages.

MITAuto-check passedDocuments & Office

Install Fetch URL As Markdown

skills CLI
$ npx skills add CodeAlive-AI/ai-driven-development --skill fetch-url-as-markdown -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install CodeAlive-AI/ai-driven-development fetch-url-as-markdown --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/CodeAlive-AI/ai-driven-development.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/fetch-url-as-markdown .claude/skills/fetch-url-as-markdown && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fetch-url-as-markdown
GitHub stars
155
Token cost
~906 tokens
SKILL.md length
308 words
Files
4 (incl. scripts)
Skills in repo
22
Repo updated
First seen
Licence
MIT

At a glance

Fetch a web page (URL) and return clean Markdown via local trafilatura, with Exa MCP as a fallback for JS-rendered or anti-bot pages.

  • Works in 3 steps: Try trafilatura first → If exit code is 1 or 2 → fall back to… → Exit code 3 means trafilatura is not…
  • The user asks to read
  • SKILL.md covers Workflow (the only thing the…, Exit codes (what they mean for…, Defaults baked into the script and Useful flags, plus 1 more section
  • Runs Python scripts from its folder; calls python3

What it does

Fetch URL As Markdown is an agent skill from CodeAlive-AI/ai-driven-development. Fetch a web page (URL) and return clean Markdown via local trafilatura, with Exa MCP as a fallback for JS-rendered or anti-bot pages. Use when the user asks to read, fetch, scrape, summarize, or quote a URL — prefer this over the built-in WebFetch tool. Don't use for binary files (PDFs, images, archives) or for fetching API/JSON endpoints.

Its SKILL.md is about 910 tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `README.md` and `scripts/fetch_url.py`).

It sits in Documents & Office, covering Markdown and Web scraping. It works with Model Context Protocol. The repository describes itself as: Practices, protocols, and skills for AI-driven software development. Skills and safety hooks for Claude Code, Codex, OpenCode, Cursor, Antigravity, and any agent supporting the… The licence is MIT.

When your agent uses it

  • The user asks to read
  • Quote a URL — prefer this over the built-in WebFetch tool
  • Binary files (PDFs
  • For fetching API/JSON endpoints

Example prompts

  • “/fetch-url-as-markdown”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Try trafilatura first
  2. If exit code is 1 or 2 → fall back to Exa MCP with the same URL
  3. Exit code 3 means trafilatura is not installed — install once

What it can do on your machine

Read from SKILL.md and the folder at commit 25b7b1d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fetch URL As Markdown loads about 906 tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 308 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~91
When it runs · the whole SKILL.md, loaded when a task matches
~906

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from CodeAlive-AI/ai-driven-development at commit 25b7b1d, republished under its MIT licence (© CodeAlive-AI). 308 words, ~906 tokens.

Download SKILL.mdSave it as .claude/skills/fetch-url-as-markdown/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
fetch-url-as-markdown
description
Fetch a web page (URL) and return clean Markdown via local trafilatura, with Exa MCP as a fallback for JS-rendered or anti-bot pages. Use when the user asks to read, fetch, scrape, summarize, or quote a URL — prefer this over the built-in WebFetch tool. Don't use for binary files (PDFs, images, archives) or for fetching API/JSON endpoints.

URL to Markdown

Fetch any web URL and get clean, readable Markdown — main content only, no navigation/footer/ads. Local + free by default; smart fallback to Exa MCP when the page can't be extracted locally.

Workflow (the only thing the agent needs to remember)

  1. Try trafilatura first:

    bash
    python3 ~/.claude/skills/fetch-url-as-markdown/scripts/fetch_url.py "<URL>"
  2. If exit code is 1 or 2 → fall back to Exa MCP with the same URL:

    mcp__exa__web_search_advanced_exa(
        query="<URL>",
        includeDomains=["<host of URL>"],
        numResults=1,
        textMaxCharacters=50000,
        type="auto"
    )

    (mcp__exa__crawling works too if the server exposes it; the web_search_advanced_exa call above is the always-available variant — pin the host with includeDomains and use the URL itself as the query.)

  3. Exit code 3 means trafilatura is not installed — install once:

    bash
    python3 -m pip install --break-system-packages trafilatura

Exit codes (what they mean for the fallback decision)

CodeMeaningAction
0Markdown printed to stdoutdone
1DownloadError — network/HTTP/timeout/anti-bot block at fetchfall back to Exa
2ExtractionError — empty extract, JS/Cloudflare wall, or stub body (<200 chars)fall back to Exa
3trafilatura missinginstall (see above), then retry
4UnsupportedContentTypeError — URL is binary (PDF, image, archive)don't fall back to Exa; use the right specialized skill (e.g. pdf for PDFs)

Defaults baked into the script

  • output_format="markdown", include_formatting=True — keeps headings/lists/code structure where the source HTML uses real <h1..h6> etc.
  • include_links=True, include_tables=True
  • with_metadata=True → emits a YAML frontmatter (title, author, date, url, hostname)
  • favor_recall=True, deduplicate=True — readable but trims duplicates
  • Real-browser User-Agent + 30s timeout configured in scripts/settings.cfg
  • Anti-stub guards (built into the script):
    • rejects Content-Type other than text/html|application/xhtml+xml|text/plain|application/xml|text/xml → exit 4
    • sniffs raw HTML for Cloudflare / "Please enable JavaScript" / Imperva / DataDome wall markers → exit 2
    • rejects extracted bodies under 50 chars (configurable via --min-body N, 0 to disable) → exit 2

Useful flags

bash
... fetch_url.py "<URL>" --no-links     # strip hyperlinks
... fetch_url.py "<URL>" --no-tables    # strip tables
... fetch_url.py "<URL>" --no-metadata  # omit YAML header
... fetch_url.py "<URL>" --comments     # include user comments (off by default — usually noise)
... fetch_url.py "<URL>" --images       # include image refs (experimental)
... fetch_url.py "<URL>" --precision    # terser output, drops borderline content

When to choose what

SituationTool
Article, blog post, docs, README, wikitrafilatura (default) — local, free
JS-heavy SPA, login-walled, CloudflareExa fallback (the script will signal exit 2)
Bulk / many URLstrafilatura — no quota, no API key
Already failed twice on a domainExa directly

© CodeAlive-AI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in skills/fetch-url-as-markdown of CodeAlive-AI/ai-driven-development.

  • SKILL.md
  • README.md
  • scripts/fetch_url.py
  • scripts/settings.cfg

Open the folder on GitHubat commit 25b7b1d

Compare with similar skills

Fetch URL As Markdown next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fetch URL As Markdown compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fetch URL As Markdown this skillCodeAlive-AI/ai-driven-development155—~906Automated safety check: PassMIT
Open Knowledgeojowwalker77/BonsAI145—~5.1kAutomated safety check: PassCustom licence
Web Article Extractordongbeixiaohuo/writing-agent435—~740Automated safety check: PassMIT
Gdoc To Markdowniurykrieger/claude-bedrock1051 repos~3.8kAutomated safety check: NotesMIT
Adk Docs WriterBrainDAO/adk-ts119—~4.9kAutomated safety check: PassMIT
Email Formattingprovos/ironcurtain613—~518Automated safety check: PassApache-2.0

Similar skills

  • Open Knowledge

    ojowwalker77/BonsAI

    Authoritative agent-runtime contract for working inside an OpenKnowledge project — a markdown-CRDT knowledge base exposed over MCP.

    145 GitHub stars~5.1k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Web Article Extractor

    dongbeixiaohuo/writing-agent

    使用隔离的 Chrome DevTools MCP 从博客、新闻、公众号等网页提取正文,返回结构化内容,或保存 Markdown 和远程图片。用户要求提取文章、抓取网页正文、保存为 Markdown、下载文章图片或排查正文选择器时调用。

    435 GitHub stars~740 tokensUpdated 3 days ago
    Documents & OfficeAuto-check passed
  • Gdoc To Markdown

    iurykrieger/claude-bedrock

    Internal fetcher module for Google Docs and Sheets. An agent skill from iurykrieger/claude-bedrock.

    105 GitHub starsUsed in 1 repo~3.8k tokens
    Documents & OfficeAuto-check: notes
  • Adk Docs Writer

    BrainDAO/adk-ts

    ADK-TS documentation specialist for Fumadocs MDX pages. An agent skill from BrainDAO/adk-ts.

    119 GitHub stars~4.9k tokensUpdated 3 mo ago
    Documents & OfficeAuto-check passed
  • Email Formatting

    provos/ironcurtain

    Markdown formatting conventions for email summary documents — heading depth, list style, line length, emoji policy, and a mandatory provenance footer.

    613 GitHub stars~518 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Official

    Build a marketing agency database from Clutch.co with the Clutch.co Agency API Actor (johnvc/clutch-agency-api).

    262 GitHub stars~2.5k tokensUpdated 14 days ago
    Documents & OfficeAuto-check passed

More from CodeAlive-AI/ai-driven-development

All 22 skills in this repo
  • Investigating Repository History

    CodeAlive-AI/ai-driven-development

    Investigate GitHub repository history before risky code changes using git blame/log, GitHub PRs, review comments, squash/rebase/cherry-pick/rename heuristics, and cited evidence.

    155 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Plugins Management

    CodeAlive-AI/ai-driven-development

    Create, publish, delete, and submit plugins for coding agents (Claude Code, OpenCode, Devin CLI/Desktop).

    155 GitHub stars~3.1k tokensUpdated yesterday
    Auto-check: notes
  • Semantic Scholar Deep

    CodeAlive-AI/ai-driven-development

    Deep research over the Semantic Scholar Graph API. An agent skill from CodeAlive-AI/ai-driven-development.

    155 GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed
  • Windows QA Engineer

    CodeAlive-AI/ai-driven-development

    A skill your agent uses when testing Windows 11 desktop apps (WinForms/WPF/UWP) via UFO UIA/Win32 automation MCP.

    155 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Agentic Readiness

    CodeAlive-AI/ai-driven-development

    Audit and improve repositories for reliable agentic work across Codex and Codex App, Claude Code, and OpenCode.

    155 GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Hooks Management

    CodeAlive-AI/ai-driven-development

    Manage hooks and automation for coding agents (Claude Code, Codex CLI, OpenCode, Devin CLI/Desktop).

    155 GitHub stars~4.8k tokensUpdated yesterday
    Auto-check: notes

Questions about Fetch URL As Markdown

What does Fetch URL As Markdown do?

Fetch a web page (URL) and return clean Markdown via local trafilatura, with Exa MCP as a fallback for JS-rendered or anti-bot pages. Fetch URL As Markdown is an agent skill from CodeAlive-AI/ai-driven-development. Fetch a web page (URL) and return clean Markdown via local trafilatura, with Exa MCP as a fallback for JS-rendered or anti-bot pages.

When should I use Fetch URL As Markdown?

Fetch URL As Markdown fits situations like: the user asks to read; quote a URL — prefer this over the built-in WebFetch tool; binary files (PDFs; for fetching API/JSON endpoints.

How do I install Fetch URL As Markdown in Claude Code?

Run `npx skills add CodeAlive-AI/ai-driven-development --skill fetch-url-as-markdown -a claude-code`. Or copy the skill folder (skills/fetch-url-as-markdown in CodeAlive-AI/ai-driven-development) into .claude/skills/fetch-url-as-markdown in your project. Claude Code loads it when a task matches its description.

How do I install Fetch URL As Markdown in Codex?

Run `npx skills add CodeAlive-AI/ai-driven-development --skill fetch-url-as-markdown -a codex`. Or copy the skill folder (skills/fetch-url-as-markdown in CodeAlive-AI/ai-driven-development) into .agents/skills/fetch-url-as-markdown in your project. Codex loads it when a task matches its description.

Can I use Fetch URL As Markdown in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add CodeAlive-AI/ai-driven-development --skill fetch-url-as-markdown -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fetch-url-as-markdown, .gemini/skills/fetch-url-as-markdown, .github/skills/fetch-url-as-markdown and .opencode/skills/fetch-url-as-markdown in your project.

What does Fetch URL As Markdown need to run?

Going by SKILL.md and its folder, Fetch URL As Markdown needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Fetch URL As Markdown access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Fetch URL As Markdown safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Fetch URL As Markdown use?

Fetch URL As Markdown is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fetch URL As Markdown use?

About 906 tokens (SKILL.md is roughly 3.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Fetch URL As Markdown?

Skills that share tags, products or a category with Fetch URL As Markdown: Open Knowledge (ojowwalker77/BonsAI, 145 stars), Web Article Extractor (dongbeixiaohuo/writing-agent, 435 stars), Gdoc To Markdown (iurykrieger/claude-bedrock, 105 stars) and Adk Docs Writer (BrainDAO/adk-ts, 119 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fetch URL As Markdown?

CodeAlive-AI (a GitHub organization) maintains it in CodeAlive-AI/ai-driven-development, which has 155 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 6, 2026.

Source: CodeAlive-AI/ai-driven-development on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.