Agent skill

Media Crawler

by tsingyuai in tsingyuai/growth-lab

Install, authenticate, configure, operate, and troubleshoot the external MediaCrawler client shared by Douyin, Kuaishou, Bilibili, Weibo, Tieba, and Zhihu collectors.

Apache-2.0Auto-check passedData & Analytics

Install Media Crawler

skills CLI
$ npx skills add tsingyuai/growth-lab --skill media-crawler -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tsingyuai/growth-lab media-crawler --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tsingyuai/growth-lab.git skills-src && mkdir -p .claude/skills && cp -r skills-src/collectors/media-crawler .claude/skills/media-crawler && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
media-crawler
GitHub stars
2k
Token cost
~731 tokens
SKILL.md length
261 words
Files
3 (incl. references)
Skills in repo
22
Repo updated
First seen
Licence
Apache-2.0

At a glance

Install, authenticate, configure, operate, and troubleshoot the external MediaCrawler client shared by Douyin, Kuaishou, Bilibili, Weibo, Tieba, and Zhihu collectors.

  • Works in 7 steps: If install or authentication is missing,… → Treat… → Before a run, record upstream commit,… → …
  • Auditing this client
  • SKILL.md covers Contract, Standard invocation and Completion report
  • Calls uv

What it does

Media Crawler is an agent skill from tsingyuai/growth-lab. Install, authenticate, configure, operate, and troubleshoot the external MediaCrawler client shared by Douyin, Kuaishou, Bilibili, Weibo, Tieba, and Zhihu collectors. Xiaohongshu uses the separate browser-first xiaohongshu-mcp Collector. Use when auditing this client, onboarding a supported platform account, selecting search/detail/creator modes, enabling comments or media, locating outputs, or diagnosing crawler failures.

Its SKILL.md is about 730 tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `agents/openai.yaml` and `references/operations.md`).

It sits in Data & Analytics, covering Web scraping. It works with Bilibili, Douyin, Xiaohongshu and Model Context Protocol. The repository describes itself as: An end-to-end growth tool that understands the product, fetch the data it needs, researches the market, executes campaigns, and reviews results to improve the next round of… The licence is Apache-2.0.

When your agent uses it

  • Auditing this client
  • Onboarding a supported platform account
  • Selecting search/detail/creator modes
  • Enabling comments

Example prompts

  • “/media-crawler”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. If install or authentication is missing, invoke onboard-growth-lab. Do not duplicate the global audit here.
  2. Treat ${MEDIACRAWLER_DIR:-${GROWTHLAB_CLIENT_ROOT:-$HOME/.growth-lab/clients}/MediaCrawler} as an external checkout. Never vendor it or…
  3. Before a run, record upstream commit, platform, crawl type, keywords/IDs, config changes, login type, comment/media flags, and destination.
  4. Modify only the documented platform config and config/base_config.py; show the diff before running. Restore unrelated example values.
  5. Run serially and conservatively. Never silently retry risk-control or authentication errors.
  6. Copy the required output into the invoking Model's memory//...; leave source provenance beside it. A Collector does not invent a new…
  7. Apply the upstream non-commercial learning license and each target platform's terms.

What it can do on your machine

Read from SKILL.md and the folder at commit 2d0807c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Media Crawler loads about 731 tokens when it runs, and up to ~2.2k if it reads all its reference files. Until then it costs about 110 tokens; SKILL.md has 261 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~110
When it runs · the whole SKILL.md, loaded when a task matches
~731
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from tsingyuai/growth-lab at commit 2d0807c, republished under its Apache-2.0 licence (© tsingyuai). 261 words, ~731 tokens.

Download SKILL.mdSave it as .claude/skills/media-crawler/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
media-crawler
description
Install, authenticate, configure, operate, and troubleshoot the external MediaCrawler client shared by Douyin, Kuaishou, Bilibili, Weibo, Tieba, and Zhihu collectors. Xiaohongshu uses the separate browser-first xiaohongshu-mcp Collector. Use when auditing this client, onboarding a supported platform account, selecting search/detail/creator modes, enabling comments or media, locating outputs, or diagnosing crawler failures.

MediaCrawler

This is the shared tool layer. Read operations.md before changing the external checkout. Then invoke exactly one platform Skill:

MediaCrawler does not support Twitter/X or Reddit. Do not imply otherwise.

Contract

  1. If install or authentication is missing, invoke onboard-growth-lab. Do not duplicate the global audit here. Onboarding must obtain the user's explicit ban-risk acknowledgement before login or crawling, require existing-Chrome CDP with no browser or Cookie fallback, and verify each enabled platform with a non-empty minimal real read. Installation, a persisted profile, or a visible login alone is not readiness.
  2. Treat ${MEDIACRAWLER_DIR:-${GROWTHLAB_CLIENT_ROOT:-$HOME/.growth-lab/clients}/MediaCrawler} as an external checkout. Never vendor it or commit its browser profile, cookies, databases, or downloaded data.
  3. Before a run, record upstream commit, platform, crawl type, keywords/IDs, config changes, login type, comment/media flags, and destination.
  4. Modify only the documented platform config and config/base_config.py; show the diff before running. Restore unrelated example values.
  5. Run serially and conservatively. Never silently retry risk-control or authentication errors.
  6. Copy the required output into the invoking Model's memory/<model>/...; leave source provenance beside it. A Collector does not invent a new Memory owner.
  7. Apply the upstream non-commercial learning license and each target platform's terms.

Standard invocation

bash
cd "${MEDIACRAWLER_DIR:-${GROWTHLAB_CLIENT_ROOT:-$HOME/.growth-lab/clients}/MediaCrawler}"
uv run main.py --platform <dy|ks|bili|wb|tieba|zhihu> --lt qrcode --type <search|detail|creator>

Use --lt qrcode with CDP and an existing Chrome session. Do not fall back to standard Playwright, a newly launched clean browser, or Cookie injection.

Completion report

Return: exact source query/URLs, run time, upstream commit, raw and copied paths, record/media/comment counts, filters, partial failures, and any risk-control signal. Never report a search-card excerpt as full detail.

© tsingyuai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in collectors/media-crawler of tsingyuai/growth-lab.

  • SKILL.md
  • agents/openai.yaml
  • references/operations.md

Open the folder on GitHubat commit 2d0807c

Compare with similar skills

Media Crawler next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Media Crawler compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Media Crawler this skilltsingyuai/growth-lab2k—~731Automated safety check: PassApache-2.0
Chrome Daily Update Checkdragon-hh/ai-boshu-crawler106—~1.3kAutomated safety check: PassNone
Catbuscv-cat/catbus280—~986Automated safety check: PassMIT
Opinions Crawlerinfometa/workbuddyskills348—~2.2kAutomated safety check: PassNone
Agent ReachPanniantong/Agent-Reach95k—~1.4kAutomated safety check: PassMIT
bb-browser Site Commands for OpenClawepiral/bb-browser6.2k—~1kAutomated safety check: PassMIT

Similar skills

  • Chrome Daily Update Check

    dragon-hh/ai-boshu-crawler

    Run the AI blogger crawler daily update check for Bilibili, Douyin, Xiaohongshu, and YouTube with Chrome producers only where required, the merged Xiaohongshu profile-video skill, and the yt-dlp…

    106 GitHub stars~1.3k tokensUpdated 2 mo ago
    Productivity & AutomationAuto-check passed
  • Catbus

    cv-cat/catbus

    用 catbus CLI 读取或操作小红书(RedNote)、抖音、TikTok、B 站(Bilibili)、快手、微博、闲鱼、淘宝、京东、X(Twitter)这 10 个平台:搜索、详情、评论、用户、推荐流、下载、发布、点赞关注、私信、直播弹幕,输出统一的 JSON。用户要查询、采集、导出、监听或操作这些平台上的内容和账号时使用。

    280 GitHub stars~986 tokensUpdated 3 days ago
    Data & AnalyticsAuto-check passed
  • Opinions Crawler

    infometa/workbuddyskills

    基于 OpenCLI 的舆情数据抓取技能。覆盖国内外主流社媒(搜索、用户信息、帖子/视频列表、视频详情及互动量、评论、弹幕等)和商店平台数据爬取。当用户需要抓取舆情数据、采集社媒内容、获取商店评分评论、或者需要安装和配置 OpenCLI 时使用。

    348 GitHub stars~2.2k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Agent Reach

    Panniantong/Agent-Reach

    Routes web research and platform lookups across 16 sites, including Twitter, Reddit, YouTube, Bilibili, Xiaohongshu and GitHub, through one command-line tool.

    95k GitHub stars~1.4k tokensUpdated 2 days ago
    Productivity & AutomationAuto-check passed
  • Runs structured data commands against sites such as Twitter, Reddit, GitHub, YouTube and Zhihu through OpenClaw's browser, reusing your existing login state.

    6.2k GitHub stars~1k tokensUpdated 4 mo ago
    Productivity & AutomationAuto-check passed
  • Ketch

    1broseidon/ketch

    Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but…

    702 GitHub starsUsed in 1 repo~3.9k tokens
    Data & AnalyticsAuto-check passed

More from tsingyuai/growth-lab

All 22 skills in this repo
  • Screenshot Assets

    tsingyuai/growth-lab

    Capture authenticated product screenshots through the repository-owned Playwright CDP script and archive them in the invoking loop's Memory.

    2k GitHub stars~583 tokensUpdated 12 days ago
    Auto-check passed
  • Run SEO Page Loop

    tsingyuai/growth-lab

    Run an SEO page observation-action-review loop with persistent Memory by coordinating demand research, page creation, adversarial review, image generation, IndexNow submission, and performance review.

    2k GitHub stars~1.2k tokensUpdated 12 days ago
    Auto-check passed
  • Wechat Article Compose

    tsingyuai/growth-lab

    把一个已确认的选题写成微信公众号长文。锁定复刻锚与用户价值、写出 article.md / wechat.yml / review.md,并通过公众号通用合规检查和人工预览前自检。准备、改写或核验公众号文章时使用;不负责发布,也不产出小红书卡片。

    2k GitHub stars~860 tokensUpdated 12 days ago
    Auto-check passed
  • Wechat Mp Publish

    tsingyuai/growth-lab

    通过微信公众号官方 API 把公众号文章生产单元渲染为微信排版 HTML、生成阅读页预览、上传封面与正文图片并创建草稿;在三重确认下提交发布并查询状态。支持本机直连与固定 IP 远程发布服务两种模式。用户要求预览公众号、同步微信草稿、发布公众号或查询发布状态时使用。

    2k GitHub stars~797 tokensUpdated 12 days ago
    Auto-check: notes
  • Xhs Render Cards

    tsingyuai/growth-lab

    把已批准的小红书草稿、单一分析参考和真实产品素材变成可审查卡片:先完成 DAI 与 image plan,再按确定性、完整效果或可分离图层模式制作,机械验证 PNG 与清单并运行合规检查。精确文字和真实 UI 不交给模型猜测;AI 生图需要单独配置和授权。

    2k GitHub stars~1k tokensUpdated 12 days ago
    Auto-check passed
  • Xiaohongshu MCP

    tsingyuai/growth-lab

    使用本机 browser-first xiaohongshu-mcp 只读搜索小红书、下载候选首图、补全用户选择的笔记详情,并把脱敏证据写入调用方 Memory。用于 xhs-replicate 的选题和视觉参考研究,取代小红书 MediaCrawler 路径。

    2k GitHub stars~822 tokensUpdated 12 days ago
    Auto-check passed

Questions about Media Crawler

What does Media Crawler do?

Install, authenticate, configure, operate, and troubleshoot the external MediaCrawler client shared by Douyin, Kuaishou, Bilibili, Weibo, Tieba, and Zhihu collectors. Media Crawler is an agent skill from tsingyuai/growth-lab. Install, authenticate, configure, operate, and troubleshoot the external MediaCrawler client shared by Douyin, Kuaishou, Bilibili, Weibo, Tieba, and Zhihu collectors.

When should I use Media Crawler?

Media Crawler fits situations like: auditing this client; onboarding a supported platform account; selecting search/detail/creator modes; enabling comments.

How do I install Media Crawler in Claude Code?

Run `npx skills add tsingyuai/growth-lab --skill media-crawler -a claude-code`. Or copy the skill folder (collectors/media-crawler in tsingyuai/growth-lab) into .claude/skills/media-crawler in your project. Claude Code loads it when a task matches its description.

How do I install Media Crawler in Codex?

Run `npx skills add tsingyuai/growth-lab --skill media-crawler -a codex`. Or copy the skill folder (collectors/media-crawler in tsingyuai/growth-lab) into .agents/skills/media-crawler in your project. Codex loads it when a task matches its description.

Can I use Media Crawler in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tsingyuai/growth-lab --skill media-crawler -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/media-crawler, .gemini/skills/media-crawler, .github/skills/media-crawler and .opencode/skills/media-crawler in your project.

What does Media Crawler need to run?

Going by SKILL.md and its folder, Media Crawler needs the command-line tools its instructions call (uv).

Does Media Crawler access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Media Crawler safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Media Crawler use?

Media Crawler is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Media Crawler use?

About 731 tokens (SKILL.md is roughly 2.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.5k tokens, read only when the agent opens those files.

What are the alternatives to Media Crawler?

Skills that share tags, products or a category with Media Crawler: Chrome Daily Update Check (dragon-hh/ai-boshu-crawler, 106 stars), Catbus (cv-cat/catbus, 280 stars), Opinions Crawler (infometa/workbuddyskills, 348 stars) and Agent Reach (Panniantong/Agent-Reach, 95k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Media Crawler?

tsingyuai (a GitHub organization) maintains it in tsingyuai/growth-lab, which has 2,000 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on September 28, 2026.

Source: tsingyuai/growth-lab on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.