Agent skill

URL to Markdown Fetcher

by joeseesun in joeseesun/qiaomu-markdown-proxy

Converts a URL or PDF into clean Markdown, with dedicated routes for WeChat articles, Feishu docs, arXiv papers and login-gated pages, before any summary or rewrite.

MITAuto-check passedDocuments & Office

Install URL to Markdown Fetcher

skills CLI
$ npx skills add joeseesun/qiaomu-markdown-proxy --skill qiaomu-markdown-proxy -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install joeseesun/qiaomu-markdown-proxy qiaomu-markdown-proxy --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
qiaomu-markdown-proxy
GitHub stars
509
Token cost
~1.4k tokens
SKILL.md length
280 words
Files
26 (incl. scripts, references, assets)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Converts a URL or PDF into clean Markdown, with dedicated routes for WeChat articles, Feishu docs, arXiv papers and login-gated pages, before any summary or rewrite.

  • Works in 3 steps: Route by URL Type → Display Content → Continue or Save
  • Reading a WeChat public-account article or a Feishu document from its link
  • SKILL.md covers Trigger Priority(先触发,再分流), URL Routing (先判断再执行), Workflow and Examples, plus 1 more section
  • Calls bash, python3 and curl; reaches r.jina.ai and x.com; needs FEISHU_APP_SECRET

What it does

Whenever a user gives a link to read, this skill goes first and hands the Markdown to whatever comes next, such as a summary, translation, article or podcast script. It sorts each URL by type: WeChat article links use `scripts/fetch_weixin.sh`, which tries a proxy and falls back to Playwright on a verification page; Feishu and Lark docs use `scripts/fetch_feishu.py`, which needs Feishu API authentication; arXiv links use `scripts/extract_tex.py` to pull sections, figures and formulas from the LaTeX source; PDFs use `scripts/extract_pdf.sh`; everything else goes through `scripts/fetch.sh` with a cascade of proxy services.

It is not used when you already pasted the full text, for YouTube links (a separate download skill handles those) or for ordinary web search. After fetching, it shows the title, author and source platform. A request that only asks for extraction saves the page to `~/Downloads` as a Markdown file with YAML front matter (title, author, date, URL, source), unless you say preview only; a read-then-write request continues straight into the downstream task.

When your agent uses it

  • Reading a WeChat public-account article or a Feishu document from its link
  • Converting a webpage or PDF into Markdown before summarizing or translating it
  • Pulling the structured content of an arXiv paper from its LaTeX source
  • Fetching a post from a site that requires login, such as X

Example prompts

  • “Read https://mp.weixin.qq.com/s/abc123 and summarize the main argument.”
  • “Convert this arXiv paper into Markdown and keep the equations.”
  • “Fetch the article at example.com/blog/launch and save it as Markdown in my Downloads folder.”

Requirements

  • Bash, to run the fetch scripts
  • Feishu API credentials, for Feishu and Lark documents
  • Playwright, for the WeChat fallback

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Route by URL Type
  2. Display Content
  3. Continue or Save

What it can do on your machine

Read from SKILL.md and the folder at commit ab56e0a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • bash
    • python3
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • r.jina.ai
    • x.com
    • mp.weixin.qq.com
    • xxx.feishu.cn
    • arxiv.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • FEISHU_APP_SECRET

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

URL to Markdown Fetcher loads about 1.4k tokens when it runs, and up to ~2.3k if it reads all its reference files. Until then it costs about 131 tokens; SKILL.md has 280 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~131
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from joeseesun/qiaomu-markdown-proxy at commit ab56e0a, republished under its MIT licence (© joeseesun). 280 words, ~1,363 tokens.

Download SKILL.mdSave it as .claude/skills/qiaomu-markdown-proxy/SKILL.md (or your agent's skills folder). This skill also uses 25 other files; get the full folder from GitHub.
name
qiaomu-markdown-proxy
description
Read, fetch, extract, parse, or convert a URL or link into clean Markdown. Use FIRST whenever a user asks to read source content from a webpage, especially WeChat/微信公众号 mp.weixin.qq.com, Feishu/Lark docs, X/Twitter, PDFs, arXiv papers, or a URL that will feed a later summary, rewrite, article, podcast, or analysis. For WeChat and Feishu, prefer this specialist extractor over generic web open or generic content parsers. Exclude YouTube, pure web search, already-pasted text, and local non-PDF files.
version
2.1.0

Markdown Proxy - URL to Markdown

将任意 URL 转为干净的 Markdown。支持需要登录的页面、PDF、专有平台。

Trigger Priority(先触发,再分流)

  • 用户给出 URL 并要求“读取、抓取、提取、解析、转 Markdown”时,优先使用本 Skill。
  • 用户要求“读取这篇链接后再写稿、总结、翻译、做播客或分析”时,先用本 Skill 取回原文,再把 Markdown 交给下游 Skill;不要因为最终任务是写作而跳过抓取。
  • mp.weixin.qq.com 和飞书文档属于专用路由。不要先试普通网页打开或通用内容解析器。
  • 用户已经贴出全文时不触发。YouTube 交给 qiaomu-youtube-download,普通搜索问题交给搜索工具。

URL Routing (先判断再执行)

收到 URL 后,先判断类型,不同类型走不同通道:

URL PatternRoute ToReason
mp.weixin.qq.comscripts/fetch_weixin.sh先代理,验证码页自动回退 Playwright
feishu.cn/docx/ feishu.cn/wiki/ larksuite.com/docx/scripts/fetch_feishu.py需飞书 API 认证
youtube.com youtu.beqiaomu-youtube-download skillYouTube 有专用工具链
huggingface.co/papers/提取 arXiv ID → scripts/extract_tex.pyHuggingFace 论文页实际是 arXiv 镜像,先找到 arXiv 链接再走 LaTeX 提取
arxiv.org/abs/ arxiv.org/pdf/scripts/extract_tex.py从 LaTeX 源码提取结构化内容 (章节/图表/公式)
.pdf (URL or local path)scripts/extract_pdf.shPDF 专用提取
All other URLsscripts/fetch.sh代理级联自动 fallback

Workflow

Step 1: Route by URL Type
if URL contains "mp.weixin.qq.com":
    → bash ~/.agents/skills/qiaomu-markdown-proxy/scripts/fetch_weixin.sh "URL"
    → Done

if URL contains "feishu.cn/docx/" or "feishu.cn/wiki/" or "larksuite.com/docx/":
    → python3 ~/.agents/skills/qiaomu-markdown-proxy/scripts/fetch_feishu.py "URL"
    → Done

if URL contains "huggingface.co/papers/":
    → First fetch the page (WebFetch) to find the arXiv URL
    → Then python3 ~/.agents/skills/qiaomu-markdown-proxy/scripts/extract_tex.py "{arxiv_url}"
    → Done

if URL contains "arxiv.org/abs/" or "arxiv.org/pdf/":
    → python3 ~/.agents/skills/qiaomu-markdown-proxy/scripts/extract_tex.py "URL"
    → Done

if URL contains "youtube.com" or "youtu.be":
    → Call qiaomu-youtube-download skill
    → Done

if URL ends with ".pdf" or is local PDF path:
    if remote URL:
        → Try: curl -sL "https://r.jina.ai/{url}"
        → If fails: download + extract_pdf.sh
    if local path:
        → bash ~/.agents/skills/qiaomu-markdown-proxy/scripts/extract_pdf.sh "PATH"
    → Done

else:
    → bash ~/.agents/skills/qiaomu-markdown-proxy/scripts/fetch.sh "URL"
    → Done
Step 2: Display Content

After fetching, show to user:

Title:  {title}
Author: {author} (if available)
Source: {platform} (公众号 / 飞书文档 / 网页 / PDF)
URL:    {original_url}

Summary
{3-5 sentence summary}

Content
{full Markdown, truncated at 200 lines if long}
Step 3: Continue or Save
  • Composite request(“读取后写稿/总结/分析”):把提取结果直接交给下游任务,在同一轮继续;除非用户要求,不必额外保存源文件。

  • Extraction-only request(“只读取/转 Markdown”):保存到 ~/Downloads/{title}.md,使用 YAML frontmatter。

  • Filename: use article title, remove special characters.

  • Format: YAML frontmatter (title, author, date, url, source) + Markdown body.

  • Tell the user the saved path.

  • Skip saving if the user says “just preview” or “don't save”.

Only stop after extraction when extraction was the complete request. If the user asked for a downstream deliverable, continue to that deliverable.

Examples

General URL
bash
bash ~/.agents/skills/qiaomu-markdown-proxy/scripts/fetch.sh "https://example.com/article"
X/Twitter Post
bash
bash ~/.agents/skills/qiaomu-markdown-proxy/scripts/fetch.sh "https://x.com/username/status/1234567890"
WeChat Article
bash
bash ~/.agents/skills/qiaomu-markdown-proxy/scripts/fetch_weixin.sh "https://mp.weixin.qq.com/s/abc123"
Feishu Document
bash
python3 ~/.agents/skills/qiaomu-markdown-proxy/scripts/fetch_feishu.py "https://xxx.feishu.cn/docx/xxxxxxxx"
arXiv LaTeX Source
bash
python3 ~/.agents/skills/qiaomu-markdown-proxy/scripts/extract_tex.py "https://arxiv.org/abs/1706.03762"
PDF (Remote)
bash
curl -sL "https://r.jina.ai/https://example.com/paper.pdf"
PDF (Local)
bash
bash ~/.agents/skills/qiaomu-markdown-proxy/scripts/extract_pdf.sh "/path/to/paper.pdf"
With Custom Proxy
bash
bash ~/.agents/skills/qiaomu-markdown-proxy/scripts/fetch.sh "https://example.com" "http://127.0.0.1:7890"

Notes

  • r.jina.ai and defuddle.md require no API key
  • fetch.sh handles proxy cascade with automatic fallback
  • Content validation: filters error, login-wall, and WeChat verification pages; requires >5 lines
  • WeChat wrapper prefers dependency-free proxies, then uses local Playwright; when Python packages are missing and uv is available, it runs them in an isolated environment
  • Playwright fallback requires a Chromium runtime; install once with python3 -m playwright install chromium if absent
  • Feishu script requires: FEISHU_APP_ID + FEISHU_APP_SECRET env vars
  • PDF extraction tries: marker-pdf → pdftotext → pypdf
  • For detailed method documentation, see references/methods.md

© joeseesun, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 25 other files (scripts, references, assets) in the repository root of joeseesun/qiaomu-markdown-proxy.

  • SKILL.md
  • .gitignore
  • LICENSE
  • README.md
  • agents/interface.yaml
  • assets/qiaomu-profile/qiaomu_avatar.jpeg
  • assets/qiaomu-profile/qiaomu_reward_qr.png
  • assets/qiaomu-profile/qiaomu_wechat_public_account_qr.jpg
  • evals/trigger_cases.json
  • manifest.json
  • references/methods.md
  • reports/creation-handoff.md
  • reports/output-evidence.json
  • reports/prior-art-research.md
  • reports/release-local.json
  • … and 11 more

Open the folder on GitHubat commit ab56e0a

Compare with similar skills

URL to Markdown Fetcher next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

URL to Markdown Fetcher compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
URL to Markdown Fetcher this skilljoeseesun/qiaomu-markdown-proxy509—~1.4kAutomated safety check: PassMIT
Read URLs and PDFstw93/Waza7.2k—~1.8kAutomated safety check: PassMIT
Web To Markdownrookie-ricardo/erduo-skills935—~894Automated safety check: PassMIT
Clean Content FetchLeoYeAI/openclaw-master-skills2.2k—~574Automated safety check: PassMIT
Huashu Markdown Publishing Pipelinealchaincyf/huashu-md-html907—~4.8kAutomated safety check: PassMIT
Feishu Docs Exporter and Writerleemysw/feishu-docx268—~1.8kAutomated safety check: PassMIT

Similar skills

  • Fetches web pages and PDFs and returns a source-grounded summary, clean Markdown, quotes or citations, routing each kind of link to a suitable fetch method.

    7.2k GitHub stars~1.8k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Web To Markdown

    rookie-ricardo/erduo-skills

    Convert a web URL into cleaned Markdown with deterministic routing.

    935 GitHub stars~894 tokensUpdated 2 mo ago
    Productivity & AutomationAuto-check passed
  • Clean Content Fetch

    LeoYeAI/openclaw-master-skills

    获取干净、可读的网页正文内容,适合现代网页、博客、新闻、公告和微信公众号文章抓取;支持网页正文提取、内容清洗、去噪、Markdown 输出,适用于普通 fetch 效果不佳、页面噪音较多或动态渲染干扰的场景。Clean content fetch for modern web pages, article extraction, WeChat article capture, content…

    2.2k GitHub stars~574 tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Huashu Markdown Publishing Pipeline

    alchaincyf/huashu-md-html

    Converts files and web pages into clean Markdown, then turns Markdown into polished HTML, Word, PDF and EPUB using four templates.

    907 GitHub stars~4.8k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Exports Feishu and Lark documents, sheets, bitables and wikis to Markdown, and creates, appends to or manages cloud docs from the command line.

    268 GitHub stars~1.8k tokensUpdated 16 days ago
    Documents & OfficeAuto-check passed
  • Qiaomu Epub Book Generator

    joeseesun/qiaomu-epub-book-generator

    Generate EPUB ebooks from Markdown files with SVG→PNG conversion, remote/local image embedding, compressed illustrations, and WeChat Reading compatibility.

    134 GitHub stars~1.1k tokensUpdated 6 mo ago
    Documents & OfficeAuto-check passed

Questions about URL to Markdown Fetcher

What does URL to Markdown Fetcher do?

Converts a URL or PDF into clean Markdown, with dedicated routes for WeChat articles, Feishu docs, arXiv papers and login-gated pages, before any summary or rewrite. Whenever a user gives a link to read, this skill goes first and hands the Markdown to whatever comes next, such as a summary, translation, article or podcast script.sh` with a cascade of proxy services.

When should I use URL to Markdown Fetcher?

URL to Markdown Fetcher fits situations like: reading a WeChat public-account article or a Feishu document from its link; converting a webpage or PDF into Markdown before summarizing or translating it; pulling the structured content of an arXiv paper from its LaTeX source; fetching a post from a site that requires login, such as X.

How do I install URL to Markdown Fetcher in Claude Code?

Run `npx skills add joeseesun/qiaomu-markdown-proxy --skill qiaomu-markdown-proxy -a claude-code`. Or copy the skill folder (the joeseesun/qiaomu-markdown-proxy repository) into .claude/skills/qiaomu-markdown-proxy in your project. Claude Code loads it when a task matches its description.

How do I install URL to Markdown Fetcher in Codex?

Run `npx skills add joeseesun/qiaomu-markdown-proxy --skill qiaomu-markdown-proxy -a codex`. Or copy the skill folder (the joeseesun/qiaomu-markdown-proxy repository) into .agents/skills/qiaomu-markdown-proxy in your project. Codex loads it when a task matches its description.

Can I use URL to Markdown Fetcher in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add joeseesun/qiaomu-markdown-proxy --skill qiaomu-markdown-proxy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/qiaomu-markdown-proxy, .gemini/skills/qiaomu-markdown-proxy, .github/skills/qiaomu-markdown-proxy and .opencode/skills/qiaomu-markdown-proxy in your project.

What does URL to Markdown Fetcher need to run?

Going by SKILL.md and its folder, URL to Markdown Fetcher needs the command-line tools its instructions call (bash, python3 and curl) and credentials named FEISHU_APP_SECRET. Our summary lists: Bash, to run the fetch scripts; Feishu API credentials, for Feishu and Lark documents; Playwright, for the WeChat fallback.

Does URL to Markdown Fetcher access the network?

SKILL.md names 5 domains. In commands or code: r.jina.ai, x.com, mp.weixin.qq.com, xxx.feishu.cn and arxiv.org; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is URL to Markdown Fetcher safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does URL to Markdown Fetcher use?

URL to Markdown Fetcher is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does URL to Markdown Fetcher use?

About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 914 tokens, read only when the agent opens those files.

What are the alternatives to URL to Markdown Fetcher?

Skills that share tags, products or a category with URL to Markdown Fetcher: Read URLs and PDFs (tw93/Waza, 7.2k stars), Web To Markdown (rookie-ricardo/erduo-skills, 935 stars), Clean Content Fetch (LeoYeAI/openclaw-master-skills, 2.2k stars) and Huashu Markdown Publishing Pipeline (alchaincyf/huashu-md-html, 907 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains URL to Markdown Fetcher?

joeseesun (a GitHub user) maintains it in joeseesun/qiaomu-markdown-proxy, which has 509 GitHub stars. The repository was last updated on August 5, 2026.

Source: joeseesun/qiaomu-markdown-proxy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.