Agent skill

Xiaohongshu Post Extractor

by chenxiachan in chenxiachan/xhs-claude-skills

Pulls a Xiaohongshu post's text, image text and video subtitles or transcript into a Markdown note saved in an Obsidian folder, in Chinese.

MITAuto-check: warningsKnowledge Management

SKILL.md written in Chinese; this summary is our English description.

Install Xiaohongshu Post Extractor

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add chenxiachan/xhs-claude-skills --skill xhs -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install chenxiachan/xhs-claude-skills xhs --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/chenxiachan/xhs-claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/xhs .claude/skills/xhs && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
xhs
GitHub stars
421
Token cost
~1.5k tokens
SKILL.md length
215 words
Files
1
Skills in repo
4
Repo updated
First seen
Licence
MIT

At a glance

Pulls a Xiaohongshu post's text, image text and video subtitles or transcript into a Markdown note saved in an Obsidian folder, in Chinese.

  • Works in 2 steps: 检查 ~/cookies.json 是否存在 → 如果不存在,告知用户需要从 Chrome 导出 cookies
  • Saving a Xiaohongshu article post as a Markdown note
  • SKILL.md covers 常量定义, 输入 and 提取流程
  • Calls curl and ffmpeg; reaches xiaohongshu.com

What it does

You give the agent a Xiaohongshu link. It first looks for a cookies file at `~/cookies.json` exported from a logged-in Chrome session and, if it is missing, explains how to export one and stops. It then reads the post ID and `xsec_token` from the URL, requests the page with those cookies and parses the post data from `window.__INITIAL_STATE__`. If the page redirects to an error, the cookies have expired and you are asked to export them again.

For video posts it prefers subtitles embedded by the platform and only falls back to downloading the video, extracting 16000 Hz mono audio with ffmpeg and transcribing it with a Whisper large-v3-turbo model through `mlx_whisper`, then cleans the text and removes temporary files. For image posts it runs OCR when the description is short and there are several images. The result is saved as Markdown under `~/Documents/Obsidian Vault/xhs`. The excerpt is truncated in the OCR step.

When your agent uses it

  • Saving a Xiaohongshu article post as a Markdown note
  • Getting a transcript or subtitles from a Xiaohongshu video post
  • Extracting text from image-only posts such as screenshots of long articles

Example prompts

  • “提取这个小红书帖子的内容并保存到我的 Obsidian 笔记里。”
  • “Get the transcript of this Xiaohongshu video and save it as a Markdown note.”
  • “这篇图文帖子的文字都在图片里,帮我做 OCR 整理成笔记。”

Requirements

  • Cookies exported from a logged-in Chrome session and saved as `~/cookies.json`
  • Python and `curl`
  • `ffmpeg` and `mlx_whisper` for video transcription
  • An Obsidian vault folder at `~/Documents/Obsidian Vault`
  • Pre-approved tools (allowed-tools): Bash, Read, Write, Edit, Glob, Grep

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. 检查 ~/cookies.json 是否存在
  2. 如果不存在,告知用户需要从 Chrome 导出 cookies

What it can do on your machine

Read from SKILL.md and the folder at commit e140727. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write
    • Edit
    • Glob
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • ffmpeg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • xiaohongshu.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Xiaohongshu Post Extractor loads about 1.5k tokens when it runs. Until then it costs about 12 tokens; SKILL.md has 215 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~12
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningMentions a credentials file (SSH keys, cloud or package-manager tokens)SKILL.md:12
    - Cookies 文件: `~/cookies.json`(从 Chrome 导出的小红书 cookies)
  • WarningMentions a credentials file (SSH keys, cloud or package-manager tokens)SKILL.md:23
    2. 如果不存在,告知用户需要从 Chrome 导出 cookies:
  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write, Edit, Glob, Grep

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from chenxiachan/xhs-claude-skills at commit e140727, republished under its MIT licence (© chenxiachan). 215 words, ~1,490 tokens.

Download SKILL.mdSave it as .claude/skills/xhs/SKILL.md (or your agent's skills folder).
name
xhs
description
提取小红书帖子内容(文字、图片 OCR、视频字幕/转录),整理为 Markdown 并保存
allowed-tools
Bash, Read, Write, Edit, Glob, Grep
user-invocable
true
argument-hint
<小红书链接>

用户希望提取小红书帖子内容。请按以下步骤处理:

常量定义

  • Cookies 文件: ~/cookies.json(从 Chrome 导出的小红书 cookies)
  • Obsidian 保存目录: ~/Documents/Obsidian Vault/xhs
  • Whisper 模型: mlx-community/whisper-large-v3-turbo

输入

用户提供的小红书链接: $ARGUMENTS

提取流程

步骤 0:检查 Cookies
  1. 检查 ~/cookies.json 是否存在
  2. 如果不存在,告知用户需要从 Chrome 导出 cookies:
    • 在 Chrome 打开 xiaohongshu.com 并确认已登录
    • 打开 DevTools Console,运行以下代码将 cookies 复制到剪贴板:
    javascript
    copy(JSON.stringify(document.cookie.split('; ').map(c => {
      const [name, ...rest] = c.split('=');
      return { name, value: rest.join('='), domain: '.xiaohongshu.com', path: '/',
        expires: Date.now()/1000 + 86400*30, size: name.length + rest.join('=').length,
        httpOnly: false, secure: false, session: false, priority: 'Medium',
        sameParty: false, sourceScheme: 'Secure', sourcePort: 443 };
    })))
    • 将剪贴板内容保存到 ~/cookies.json
    • 然后终止流程,等用户完成后重新运行
步骤 1:解析链接

从 URL 中提取帖子 ID(24 位十六进制字符串)和 xsec_token 参数。

步骤 2:获取帖子内容

使用 Python 脚本,通过 Cookies 请求帖子页面 HTML,从 window.__INITIAL_STATE__ 解析全部帖子数据:

python
import json, urllib.request, ssl, re

with open('<Cookies 文件>') as f:
    cookies = json.load(f)
cookie_str = '; '.join(f"{c['name']}={c['value']}" for c in cookies)

ctx = ssl.create_default_context()
ctx.check_hostname = False
ctx.verify_mode = ssl.CERT_NONE

req = urllib.request.Request('<帖子URL>')
req.add_header('Cookie', cookie_str)
req.add_header('User-Agent', 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36')

resp = urllib.request.urlopen(req, timeout=15, context=ctx)
html = resp.read().decode('utf-8', errors='ignore')

m = re.search(r'window\.__INITIAL_STATE__\s*=\s*(\{.+?\})\s*</script>', html, re.DOTALL)
raw = m.group(1).replace('undefined', 'null')
data = json.loads(raw)

# 帖子数据在: data['note']['noteDetailMap'][<key>]['note']
# 包含: title, desc, type, time, user, imageList, video, interactInfo, ipLocation

如果请求失败(被重定向到 404/错误页),说明 cookies 过期,提示用户按步骤 0 重新导出。

步骤 3:视频内容提取(仅视频帖子)

如果帖子 type 为 video,优先使用平台内嵌字幕,仅在无字幕时回退到本地 Whisper 转录。

3a. 检查平台字幕(优先)

从步骤 2 获取的视频数据中检查是否有内嵌字幕:

note['video']['media'] 或 note['video']['mediaV2'](JSON 字符串,需二次解析)
-> 查找 subtitles 字段
-> 优先级:source > zh-CN > en-US
-> 取对应语言的 SRT URL

如果找到字幕 URL:

bash
# 注意:字幕 CDN 域名必须使用 HTTPS(HTTP 可能超时)
curl -sL --connect-timeout 10 -o /tmp/xhs_{post_id}.srt \
  -H "User-Agent: Mozilla/5.0" \
  -H "Referer: https://www.xiaohongshu.com/" \
  "<字幕URL(确保 https://)>"

解析 SRT 文件,合并为连续文本(去除时间戳和序号),按语义断句重新组织段落。 字幕比 Whisper 转录更准确,且无需下载视频,应优先使用。

3b. Whisper 转录(回退方案)

仅当步骤 3a 未找到字幕时,执行以下子步骤:

提取视频 URL:

note['video']['media']['stream'] -> 按 h264 > h265 > av1 优先级取第一个的 masterUrl

下载视频并提取音频:

bash
curl -L -o /tmp/xhs_{post_id}.mp4 -H "Referer: https://www.xiaohongshu.com/" <视频URL>
ffmpeg -y -i /tmp/xhs_{post_id}.mp4 -vn -acodec pcm_s16le -ar 16000 -ac 1 /tmp/xhs_{post_id}.wav

语音转录:

python
import mlx_whisper
result = mlx_whisper.transcribe("/tmp/xhs_{post_id}.wav",
    path_or_hf_repo="mlx-community/whisper-large-v3-turbo", language="zh", verbose=False)
3c. 清理转录/字幕文本
  • 去除尾部重复字符(背景音乐噪音)
  • 按语义断句,添加标点和段落
  • 如有步骤/要点结构,用 Markdown 格式化
3d. 清理临时文件
bash
rm -f /tmp/xhs_{post_id}.mp4 /tmp/xhs_{post_id}.wav /tmp/xhs_{post_id}.srt
步骤 3B:图片文字识别(仅图文帖子)

如果帖子 type 为 normal(图文帖子),且图片中可能包含大量文字内容(如长文截图、PPT 翻拍、信息图表等),执行以下子步骤进行 OCR 识别。

判断是否需要 OCR: 如果帖子 desc 已经包含完整的文章内容(超过 500 字),通常不需要 OCR。但如果 desc 较短(如仅有标题或几句引言),而图片数量较多(≥3 张),则图片很可能是文章的载体,需要 OCR 提取。

3B-a. 下载图片

从步骤 2 获取的 imageList 中提取每张图片的 urlDefault URL。

关键:必须将 HTTP URL 改为 HTTPS(HTTP 连接小红书图片 CDN 可能超时)。

使用 curl 批量下载:

bash
# 单张下载
curl -sL --connect-timeout 10 -o /tmp/xhs_{post_id}_img_{序号}.jpg \
  -H "Referer: https://www.xiaohongshu.com/" \
  -H "User-Agent: Mozilla/5.0" \
  "<图片URL(http:// 替换为 https://)>"

# 批量下载(curl 多输出模式,一条命令下载所有图片)
curl -sL --connect-timeout 10 \
  -H "Referer: https://www.xiaohongshu.com/" \
  -H "User-Agent: Mozilla/5.0" \
  -o /tmp/xhs_{post_id}_img_00.jpg "<URL_0>" \
  -o /tmp/xhs_{post_id}_img_01.jpg "<URL_1>" \
  ...
3B-b. 读取图片文字

使用 Claude Code 的 Read 工具读取每张图片(多模态能力,直接识别图中文字)。

注意多图限制: Claude 的多图上下文限制为每张图片最长边 ≤ 2000px。每次最多同时读取 4 张图片,超过 4 张需分批读取。

# 分批读取,每批最多 4 张
Read /tmp/xhs_{post_id}_img_00.jpg
Read /tmp/xhs_{post_id}_img_01.jpg
Read /tmp/xhs_{post_id}_img_02.jpg
Read /tmp/xhs_{post_id}_img_03.jpg
# (下一批)
Read /tmp/xhs_{post_id}_img_04.jpg
...

从每张图片中提取所有中文/英文文字内容,按图片顺序拼接为完整文章。

3B-c. 整理 OCR 文本
  • 合并所有图片的文字为连续文章
  • 修复跨图片的断句(上一张图最后一行可能和下一张图第一行是同一句话)
  • 按逻辑结构分节,添加小标题
  • 保留关键数据、引用和结论
3B-d. 清理临时文件
bash
rm -f /tmp/xhs_{post_id}_img_*.jpg
步骤 4:整理输出并保存

将内容整理为 Markdown 文件,保存到 <Obsidian 保存目录>/{YYYY-MM-DD} {短标题}.md。

  • 文件名格式:{发布日期} {短标题}.md,短标题不超过15个字,是核心洞察的极简概括
  • 日期前缀确保按时间排序
  • 不创建子目录,所有帖子 md 直接放在 xhs 文件夹下
  • 媒体文件统一放在 <Obsidian 保存目录>/img/ 或 <Obsidian 保存目录>/video/

写作风格:Peter Thiel 式——直接、反直觉、一句话给判断。笔记是决策工具,不是知识库。用户扫一眼就能决定:深挖还是跳过。

文件结构(无 YAML frontmatter):

markdown
# 一句话核心洞察(反直觉的判断,不是描述性标题)

核心论点,2-3句话。直接给出"大多数人觉得X,但其实Y"的判断。
不废话,不铺垫,像 Thiel 在董事会上说话。

**与我的关联:** 一句话。读取用户的 memory(~/.claude/projects/*/memory/ 下的
user 和 project 类型记忆)了解用户背景、研究方向和当前工作,据此说清楚
这个内容跟用户有什么关系。如果 memory 不可用,从通用的个人发展/工具/方法论角度切入。

**值得深挖吗:** 是/否。一句话理由。

> [!tip]- 详情
> 帖子核心内容的结构化整理(折叠状态,点开才看到):
> - 从 desc、视频字幕/转录、图片 OCR 文字中提炼,清理 `#xxx[话题]#` 标记
> - 按逻辑结构分节,保留关键数据和结论
> - 纯装饰性图片用 `![图N](urlDefault)` 嵌入
> - 含大量文字的图片:嵌入 OCR 提取的结构化文本(不嵌入图片 URL)
> - 视频帖子在此处放整理后的字幕/转录内容

> [!info]- 笔记属性
> - **来源**: 小红书 · 作者名
> - **帖子ID**: xxx
> - **链接**: 原始链接
> - **日期**: YYYY-MM-DD
> - **类型**: image/video
> - **互动**: N赞 / N收藏 / N评论
> - **标签**: 标签1, 标签2, ...

关键约束:

  • 折叠区域外的可见内容不超过 6 行
  • 标题必须是洞察/判断,不是"XX帖子的总结"
  • 图片使用 urlDefault 字段的 URL

© chenxiachan, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/xhs of chenxiachan/xhs-claude-skills.

Open the folder on GitHubat commit e140727

Compare with similar skills

Xiaohongshu Post Extractor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Xiaohongshu Post Extractor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Xiaohongshu Post Extractor this skillchenxiachan/xhs-claude-skills421—~1.5kAutomated safety check: WarnMIT
Link Curator Obsidian Vaultdodo-reach/hermes-link-curator142—~2.8kAutomated safety check: PassMIT
Obsidian Web Clipper Templatesdavila7/claude-code-templates32k7 repos~632Automated safety check: PassMIT
Obsidian CLIpablo-mano/Obsidian-CLI-skill450—~3.2kAutomated safety check: PassNone
Daily Notes Workflowballred/obsidian-claude-pkm1.9k—~2kAutomated safety check: PassMIT
Summarize Callreysu/ai-life-skills269—~3.8kAutomated safety check: NotesMIT

Similar skills

  • Link Curator Obsidian Vault

    dodo-reach/hermes-link-curator

    Archives links a user sends into a profile's Obsidian-style vault with a save script and replies with a one-word confirmation instead of a summary.

    142 GitHub stars~2.8k tokensUpdated 4 mo ago
    Knowledge ManagementAuto-check passed
  • Obsidian Web Clipper Templates

    davila7/claude-code-templates

    Builds importable JSON templates for the Obsidian Web Clipper by checking your Bases, analyzing a sample page and drafting the template.

    32k GitHub starsUsed in 7 repos~632 tokens
    Knowledge ManagementAuto-check passed
  • Obsidian CLI

    pablo-mano/Obsidian-CLI-skill

    Lets the agent read, write, search and tidy an Obsidian vault directly through the official Obsidian command-line interface.

    450 GitHub stars~3.2k tokensUpdated 7 mo ago
    Knowledge ManagementAuto-check passed
  • Daily Notes Workflow

    ballred/obsidian-claude-pkm

    Creates today's daily note in an Obsidian vault and guides morning, midday and evening routines for planning, task review and reflection.

    1.9k GitHub stars~2k tokensUpdated 7 mo ago
    Knowledge ManagementAuto-check passed
  • Summarize Call

    reysu/ai-life-skills

    Transcribe a call recording with speaker diarization, summarize it, and create Obsidian vault notes (call note, transcript, person notes for participants).

    269 GitHub stars~3.8k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Obsidian Canvas Boards

    AgriciDaniel/claude-obsidian

    Creates, inspects and updates Obsidian JSON Canvas boards in a vault, with text, file, link, group and edge nodes, using safe recoverable edits.

    15k GitHub stars~1.4k tokensUpdated 28 days ago
    Knowledge ManagementAuto-check passed

More from chenxiachan/xhs-claude-skills

  • Xiaohongshu Cover Generator

    chenxiachan/xhs-claude-skills

    Creates a Xiaohongshu-style 3:4 cover image from a title by building an HTML page in one of six visual styles and capturing it with Playwright.

    421 GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check: notes
  • Xhs Analyze

    chenxiachan/xhs-claude-skills

    分析已提取的小红书收藏内容,做 AI 总结、对比、提炼

    421 GitHub stars~153 tokensUpdated 1 mo ago
    Auto-check: notes
  • Xhs Batch

    chenxiachan/xhs-claude-skills

    批量提取小红书帖子,整理为本地 Markdown 知识库

    421 GitHub stars~187 tokensUpdated 1 mo ago
    Auto-check: notes

Questions about Xiaohongshu Post Extractor

What does Xiaohongshu Post Extractor do?

Pulls a Xiaohongshu post's text, image text and video subtitles or transcript into a Markdown note saved in an Obsidian folder, in Chinese. You give the agent a Xiaohongshu link.json` exported from a logged-in Chrome session and, if it is missing, explains how to export one and stops.

When should I use Xiaohongshu Post Extractor?

Xiaohongshu Post Extractor fits situations like: saving a Xiaohongshu article post as a Markdown note; getting a transcript or subtitles from a Xiaohongshu video post; extracting text from image-only posts such as screenshots of long articles.

How do I install Xiaohongshu Post Extractor in Claude Code?

Run `npx skills add chenxiachan/xhs-claude-skills --skill xhs -a claude-code`. Or copy the skill folder (skills/xhs in chenxiachan/xhs-claude-skills) into .claude/skills/xhs in your project. Claude Code loads it when a task matches its description.

How do I install Xiaohongshu Post Extractor in Codex?

Run `npx skills add chenxiachan/xhs-claude-skills --skill xhs -a codex`. Or copy the skill folder (skills/xhs in chenxiachan/xhs-claude-skills) into .agents/skills/xhs in your project. Codex loads it when a task matches its description.

Can I use Xiaohongshu Post Extractor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add chenxiachan/xhs-claude-skills --skill xhs -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/xhs, .gemini/skills/xhs, .github/skills/xhs and .opencode/skills/xhs in your project.

What does Xiaohongshu Post Extractor need to run?

Going by SKILL.md and its folder, Xiaohongshu Post Extractor needs the command-line tools its instructions call (curl and ffmpeg). Our summary lists: Cookies exported from a logged-in Chrome session and saved as `~/cookies.json`; Python and `curl`; `ffmpeg` and `mlx_whisper` for video transcription; An Obsidian vault folder at `~/Documents/Obsidian Vault`. Its frontmatter pre-approves these tools: Bash, Read, Write, Edit, Glob, Grep.

Does Xiaohongshu Post Extractor access the network?

SKILL.md names 1 domain. In commands or code: xiaohongshu.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Xiaohongshu Post Extractor safe to install?

Our automated static check of SKILL.md flagged 2 warning(s): mentions a credentials file (ssh keys, cloud or package-manager tokens). Read the flagged lines before installing; the check is not a guarantee either way.

What licence does Xiaohongshu Post Extractor use?

Xiaohongshu Post Extractor is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Xiaohongshu Post Extractor use?

About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Xiaohongshu Post Extractor?

Skills that share tags, products or a category with Xiaohongshu Post Extractor: Link Curator Obsidian Vault (dodo-reach/hermes-link-curator, 142 stars), Obsidian Web Clipper Templates (davila7/claude-code-templates, 32k stars), Obsidian CLI (pablo-mano/Obsidian-CLI-skill, 450 stars) and Daily Notes Workflow (ballred/obsidian-claude-pkm, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Xiaohongshu Post Extractor?

chenxiachan (a GitHub user) maintains it in chenxiachan/xhs-claude-skills, which has 421 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on August 14, 2026.

Source: chenxiachan/xhs-claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.