Agent skill

News Extractor

by NanmiCoder in NanmiCoder/NewsCrawler

新闻站点内容提取。支持 12 个平台:微信公众号、今日头条、网易新闻、搜狐新闻、腾讯新闻、BBC News、CNN News、Twitter/X、Lenny's Newsletter、Naver Blog、Detik News、Quora。当用户需要提取新闻内容、抓取公众号文章、爬取新闻、或获取新闻JSON/Markdown时激活。

GPL-3.0Auto-check passedWriting & Content

Install News Extractor

skills CLI
$ npx skills add NanmiCoder/NewsCrawler --skill news-extractor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NanmiCoder/NewsCrawler news-extractor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NanmiCoder/NewsCrawler.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/news-extractor .claude/skills/news-extractor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
news-extractor
GitHub stars
622
Token cost
~1.3k tokens
SKILL.md length
198 words
Files
25 (incl. scripts, references)
Skills in repo
1
Repo updated
First seen
Licence
GPL-3.0

At a glance

新闻站点内容提取。支持 12 个平台:微信公众号、今日头条、网易新闻、搜狐新闻、腾讯新闻、BBC News、CNN News、Twitter/X、Lenny's Newsletter、Naver Blog、Detik News、Quora。当用户需要提取新闻内容、抓取公众号文章、爬取新闻、或获取新闻JSON/Markdown时激活。

  • Works in 5 steps: 接收 URL - 用户提供新闻链接 → 平台检测 - 自动识别平台类型 → 内容提取 - 调用对应爬虫获取并解析内容 → …
  • Tasks that involve Newsletters
  • SKILL.md covers 支持平台 (12), 依赖安装, 使用方式 and 工作流程, plus 6 more sections
  • Runs Python scripts from its folder; calls uv; reaches x.com and mp.weixin.qq.com

What it does

News Extractor is an agent skill from NanmiCoder/NewsCrawler. 新闻站点内容提取。支持 12 个平台:微信公众号、今日头条、网易新闻、搜狐新闻、腾讯新闻、BBC News、CNN News、Twitter/X、Lenny's Newsletter、Naver Blog、Detik News、Quora。当用户需要提取新闻内容、抓取公众号文章、爬取新闻、或获取新闻JSON/Markdown时激活。

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 27 other files, including scripts and reference files (for example `references/platform-patterns.md`, `scripts/crawlers/__init__.py` and `scripts/crawlers/base.py`).

It sits in Writing & Content, covering Newsletters and Markdown. It works with X (Twitter). The repository describes itself as: 多平台新闻 & 内容爬虫集合 | Multi-platform News & Content Crawler Suite 支持:微信公众号、Twiiter、今日头条、网易新闻、搜狐新闻、腾讯新闻、Naver、Detik、Quora 等主流平台 Supports crawling news & content from WeChat, Toutiao…. The licence is GPL-3.0.

When your agent uses it

  • Tasks that involve Newsletters
  • Tasks that involve Markdown

Example prompts

  • “/news-extractor”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. 接收 URL - 用户提供新闻链接
  2. 平台检测 - 自动识别平台类型
  3. 内容提取 - 调用对应爬虫获取并解析内容
  4. 格式转换 - 生成 JSON 和 Markdown
  5. 输出文件 - 保存到指定目录

What it can do on your machine

Read from SKILL.md and the folder at commit 25fb3b4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 15 files in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • x.com
    • mp.weixin.qq.com
    • bbc.com
    • toutiao.com
    • 163.com
    • sohu.com
    • news.qq.com
    • edition.cnn.com
    • lennysnewsletter.com
    • blog.naver.com
    • news.detik.com
    • quora.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

News Extractor loads about 1.3k tokens when it runs, and up to ~2.6k if it reads all its reference files. Until then it costs about 46 tokens; SKILL.md has 198 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NanmiCoder/NewsCrawler at commit 25fb3b4, republished under its GPL-3.0 licence (© NanmiCoder). 198 words, ~1,310 tokens.

Download SKILL.mdSave it as .claude/skills/news-extractor/SKILL.md (or your agent's skills folder). This skill also uses 24 other files; get the full folder from GitHub.
name
news-extractor
description
新闻站点内容提取。支持 12 个平台:微信公众号、今日头条、网易新闻、搜狐新闻、腾讯新闻、BBC News、CNN News、Twitter/X、Lenny's Newsletter、Naver Blog、Detik News、Quora。当用户需要提取新闻内容、抓取公众号文章、爬取新闻、或获取新闻JSON/Markdown时激活。
license
GPL-3.0

News Extractor Skill

从主流新闻平台提取文章内容,输出 JSON 和 Markdown 格式。

标准兼容、独立可迁移:本目录遵循 Agent Skills 规范,包含完整的 SKILL.md、脚本与依赖,不依赖 NewsCrawler 的其他模块。通过兼容的 Skills 安装器安装后,运行一次 uv sync 即可使用。

运行环境需要 Python 3.9+、uv 和目标站点的网络访问。

支持平台 (12)

中文平台
平台IDURL 示例
微信公众号wechathttps://mp.weixin.qq.com/s/xxxxx
今日头条toutiaohttps://www.toutiao.com/article/123456/
网易新闻neteasehttps://www.163.com/news/article/ABC123.html
搜狐新闻sohuhttps://www.sohu.com/a/123456_789
腾讯新闻tencenthttps://news.qq.com/rain/a/20251016A07W8J00
国际平台
平台IDURL 示例
BBC Newsbbchttps://www.bbc.com/news/articles/c797qlx93j0o
CNN Newscnnhttps://edition.cnn.com/2025/10/27/uk/article-slug
Twitter/Xtwitterhttps://x.com/user/status/123456789
Lenny's Newsletterlennyhttps://www.lennysnewsletter.com/p/article-slug
Naver Blognaverhttps://blog.naver.com/username/123456
Detik Newsdetikhttps://news.detik.com/internasional/d-123456/slug
Quoraquorahttps://www.quora.com/question/answers/123456

依赖安装

本 skill 使用 uv 管理依赖。首次使用前需要安装:

bash
cd <Agent Skills 目录>/news-extractor
uv sync

重要: 所有脚本必须使用 uv run 执行,不要直接用 python 运行。

依赖列表
包名用途
pydantic数据模型验证
requestsHTTP 请求
curl_cffi浏览器模拟抓取
tenacity重试机制
parselHTML/XPath 解析
demjson3非标准 JSON 解析

使用方式

基本用法
bash
# 以下命令均在 news-extractor Skill 目录内执行

# 提取新闻,自动检测平台,输出 JSON + Markdown
uv run scripts/extract_news.py "URL"

# 指定输出目录
uv run scripts/extract_news.py "URL" --output ./output

# 仅输出 JSON
uv run scripts/extract_news.py "URL" --format json

# 仅输出 Markdown
uv run scripts/extract_news.py "URL" --format markdown

# Twitter 受保护推文 (需要 Cookie)
uv run scripts/extract_news.py "URL" --cookie "auth_token=xxx; ct0=yyy"

# 列出支持的平台
uv run scripts/extract_news.py --list-platforms
输出文件

脚本默认输出两种格式到指定目录(默认 ./output):

  • {news_id}.json - 结构化 JSON 数据
  • {news_id}.md - Markdown 格式文章

工作流程

  1. 接收 URL - 用户提供新闻链接
  2. 平台检测 - 自动识别平台类型
  3. 内容提取 - 调用对应爬虫获取并解析内容
  4. 格式转换 - 生成 JSON 和 Markdown
  5. 输出文件 - 保存到指定目录

输出格式

JSON 结构
json
{
  "title": "文章标题",
  "news_url": "原始链接",
  "news_id": "文章ID",
  "meta_info": {
    "author_name": "作者/来源",
    "author_url": "",
    "publish_time": "2024-01-01 12:00"
  },
  "contents": [
    {"type": "text", "content": "段落文本", "desc": ""},
    {"type": "image", "content": "https://...", "desc": ""},
    {"type": "video", "content": "https://...", "desc": ""}
  ],
  "texts": ["段落1", "段落2"],
  "images": ["图片URL1", "图片URL2"],
  "videos": []
}
Markdown 结构
markdown
# 文章标题

## 文章信息
**作者**: xxx
**发布时间**: 2024-01-01 12:00
**原文链接**: [链接](URL)

---

## 正文内容

段落内容...

![图片](URL)

---

## 媒体资源
### 图片 (N)
1. URL1
2. URL2

使用示例

提取微信公众号文章
bash
uv run scripts/extract_news.py \
  "https://mp.weixin.qq.com/s/ebMzDPu2zMT_mRgYgtL6eQ"
提取 BBC 新闻
bash
uv run scripts/extract_news.py \
  "https://www.bbc.com/news/articles/c797qlx93j0o"
提取 Twitter 推文
bash
# 公开推文 (无需认证)
uv run scripts/extract_news.py \
  "https://x.com/BarackObama/status/896523232098078720"

# 受保护推文 (需要 Cookie)
uv run scripts/extract_news.py \
  "https://x.com/user/status/123456" --cookie "auth_token=xxx; ct0=yyy"

错误处理

错误类型说明解决方案
无法识别该平台URL 不匹配任何支持的平台检查 URL 是否正确
平台不支持非支持的站点本 Skill 仅支持列出的 12 个平台
提取失败网络错误或页面结构变化重试或检查 URL 有效性
认证失败Twitter Cookie 无效重新获取 Cookie

注意事项

  • 仅用于教育和研究目的
  • 不要进行大规模爬取
  • 尊重目标网站的 robots.txt 和服务条款
  • 微信公众号可能需要有效的 Cookie(当前默认配置通常可用)
  • Twitter 公开推文无需认证,受保护推文需要 Cookie

目录结构

news-extractor/
├── SKILL.md                      # [必需] Skill 定义文件
├── pyproject.toml                # 依赖管理
├── references/
│   └── platform-patterns.md      # 平台 URL 模式说明
└── scripts/
    ├── extract_news.py           # CLI 入口脚本
    ├── models.py                 # 数据模型
    ├── detector.py               # 平台检测
    ├── formatter.py              # Markdown 格式化
    └── crawlers/                 # 爬虫模块
        ├── __init__.py
        ├── base.py               # BaseNewsCrawler 基类
        ├── fetchers.py           # HTTP 获取策略
        ├── wechat.py             # 微信公众号
        ├── toutiao.py            # 今日头条
        ├── netease.py            # 网易新闻
        ├── sohu.py               # 搜狐新闻
        ├── tencent.py            # 腾讯新闻
        ├── bbc.py                # BBC News
        ├── cnn.py                # CNN News
        ├── twitter.py            # Twitter/X
        ├── twitter_client.py     # Twitter API 客户端
        ├── twitter_types.py      # Twitter 数据类型
        ├── lenny.py              # Lenny's Newsletter
        ├── naver.py              # Naver Blog
        ├── detik.py              # Detik News
        └── quora.py              # Quora

参考

© NanmiCoder, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 24 other files (scripts, references) in .claude/skills/news-extractor of NanmiCoder/NewsCrawler.

  • SKILL.md
  • pyproject.toml
  • references/platform-patterns.md
  • scripts/crawlers/__init__.py
  • scripts/crawlers/base.py
  • scripts/crawlers/bbc.py
  • scripts/crawlers/cnn.py
  • scripts/crawlers/detik.py
  • scripts/crawlers/fetchers.py
  • scripts/crawlers/lenny.py
  • scripts/crawlers/naver.py
  • scripts/crawlers/netease.py
  • scripts/crawlers/quora.py
  • scripts/crawlers/sohu.py
  • scripts/crawlers/tencent.py
  • scripts/crawlers/toutiao.py
  • scripts/crawlers/twitter.py
  • scripts/crawlers/twitter_client.py
  • … and 7 more

Open the folder on GitHubat commit 25fb3b4

Compare with similar skills

News Extractor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

News Extractor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
News Extractor this skillNanmiCoder/NewsCrawler622—~1.3kAutomated safety check: PassGPL-3.0
Changelog Social RecapFlorianBruniaux/claude-code-ultimate-guide6.1k—~1.8kAutomated safety check: NotesCC-BY-SA-4.0
Voice Buildercharlie947/social-media-skills3.8k—~3.3kAutomated safety check: PassMIT
AI News Digestlselector/seminar164—~1.6kAutomated safety check: PassNone
ShareSerhiiKorniienko/bullshit-detector154—~857Automated safety check: PassMIT
Content Repurposer Smsblacktwist/social-media-skills560—~3.7kAutomated safety check: PassMIT

Similar skills

  • Changelog Social Recap

    FlorianBruniaux/claude-code-ultimate-guide

    Turns CHANGELOG.md entries for a release or a week into LinkedIn, Twitter/X, newsletter and Slack posts in French and English.

    6.1k GitHub stars~1.8k tokensUpdated yesterday
    Writing & ContentAuto-check: notes
  • Voice Builder

    charlie947/social-media-skills

    Build a personalised voice profile inside a Codex or Claude project from a short interview plus 3 to 5 sample pieces of writing.

    3.8k GitHub stars~3.3k tokensUpdated 23 days ago
    Writing & ContentAuto-check passed
  • AI News Digest

    lselector/seminar

    Extract the last N days of AI/tech newsletters from the local Thunderbird mail folder and write a deduplicated markdown digest of the most important events, each with a short description and 1-2…

    164 GitHub stars~1.6k tokensUpdated today
    Writing & ContentAuto-check passed
  • Share

    SerhiiKorniienko/bullshit-detector

    Turn a BS report (or any analysis result) into ready-to-paste posts for X/Twitter, LinkedIn, Facebook, Reddit, Hacker News, or a newsletter issue — plus a branded image carousel (PNGs + PDF) for…

    154 GitHub stars~857 tokensUpdated 9 days ago
    Writing & ContentAuto-check passed
  • Content Repurposer Sms

    blacktwist/social-media-skills

    When the user wants to turn one piece of content into multiple formats or adapt content across text-first and visual-first platforms (LinkedIn, Twitter/X, Threads, Bluesky, Facebook, Instagram…

    560 GitHub stars~3.7k tokensUpdated 5 mo ago
    Writing & ContentAuto-check passed
  • Creator Research Tracker

    AlphaMao1/AlphaMao_Skills

    把 X/Twitter 博主、YouTube 博主、newsletter、播客或类似个人信息源的持续更新转成 progressive-investment-research 研究工作区的可审计增量。适用于:跟踪博主更新、生成每日报告、通过 Chrome 登录态或本地 browser bridge 抓 X、用 yt-dlp 抓 YouTube 字幕、把观点拆成 source…

    130 GitHub stars~1.6k tokensUpdated 15 days ago
    Writing & ContentAuto-check passed

Works with

Questions about News Extractor

What does News Extractor do?

新闻站点内容提取。支持 12 个平台:微信公众号、今日头条、网易新闻、搜狐新闻、腾讯新闻、BBC News、CNN News、Twitter/X、Lenny's Newsletter、Naver Blog、Detik News、Quora。当用户需要提取新闻内容、抓取公众号文章、爬取新闻、或获取新闻JSON/Markdown时激活。. News Extractor is an agent skill from NanmiCoder/NewsCrawler.

When should I use News Extractor?

News Extractor fits situations like: tasks that involve Newsletters; tasks that involve Markdown.

How do I install News Extractor in Claude Code?

Run `npx skills add NanmiCoder/NewsCrawler --skill news-extractor -a claude-code`. Or copy the skill folder (.claude/skills/news-extractor in NanmiCoder/NewsCrawler) into .claude/skills/news-extractor in your project. Claude Code loads it when a task matches its description.

How do I install News Extractor in Codex?

Run `npx skills add NanmiCoder/NewsCrawler --skill news-extractor -a codex`. Or copy the skill folder (.claude/skills/news-extractor in NanmiCoder/NewsCrawler) into .agents/skills/news-extractor in your project. Codex loads it when a task matches its description.

Can I use News Extractor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NanmiCoder/NewsCrawler --skill news-extractor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/news-extractor, .gemini/skills/news-extractor, .github/skills/news-extractor and .opencode/skills/news-extractor in your project.

What does News Extractor need to run?

Going by SKILL.md and its folder, News Extractor needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does News Extractor access the network?

SKILL.md names 12 domains. In commands or code: x.com, mp.weixin.qq.com, bbc.com, toutiao.com, 163.com, sohu.com, news.qq.com, edition.cnn.com, lennysnewsletter.com, blog.naver.com, news.detik.com and quora.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is News Extractor safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does News Extractor use?

News Extractor is published under the GPL-3.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does News Extractor use?

About 1.3k tokens (SKILL.md is roughly 5.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.3k tokens, read only when the agent opens those files.

What are the alternatives to News Extractor?

Skills that share tags, products or a category with News Extractor: Changelog Social Recap (FlorianBruniaux/claude-code-ultimate-guide, 6.1k stars), Voice Builder (charlie947/social-media-skills, 3.8k stars), AI News Digest (lselector/seminar, 164 stars) and Share (SerhiiKorniienko/bullshit-detector, 154 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains News Extractor?

NanmiCoder (a GitHub user) maintains it in NanmiCoder/NewsCrawler, which has 622 GitHub stars. The repository was last updated on August 8, 2026.

Source: NanmiCoder/NewsCrawler on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.