Agent skill

Bggg Data Reddit

by binggandata in binggandata/bggg-skills

Collect auditable Reddit search results and full comment trees at scale, preserve the source JSON, and normalize posts and comments into analysis-ready JSONL.

MITAuto-check passedMarketing & SEO

Install Bggg Data Reddit

skills CLI
$ npx skills add binggandata/bggg-skills --skill bggg-data-reddit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install binggandata/bggg-skills bggg-data-reddit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/binggandata/bggg-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/bggg-data-reddit .claude/skills/bggg-data-reddit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bggg-data-reddit
GitHub stars
605
Token cost
~1.2k tokens
SKILL.md length
430 words
Files
6 (incl. scripts, references)
Skills in repo
13
Repo updated
First seen
Licence
MIT

At a glance

Collect auditable Reddit search results and full comment trees at scale, preserve the source JSON, and normalize posts and comments into analysis-ready JSONL.

  • Works in 5 steps: Open https://www.reddit.com/ in Chrome.… → Save each unmodified response under… → Rank directly relevant parent posts by… → …
  • Market research
  • SKILL.md covers VOC Project Layout(bggg 系列共用), Workflow, Quality Rules and Degradation, plus 1 more section
  • Runs Python scripts from its folder; calls python3; reaches reddit.com

What it does

Bggg Data Reddit is an agent skill from binggandata/bggg-skills. Collect auditable Reddit search results and full comment trees at scale, preserve the source JSON, and normalize posts and comments into analysis-ready JSONL. Use for VOC, market research, community discovery, keyword snowballing, or any task that needs Reddit post and comment text without inventing unavailable fields. Reddit data layer of the bggg VOC suite (shared project folder; orchestrated by industry-orchestrator, reported by bggg-voc-report).

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `agents/openai.yaml`, `references/chrome_same_origin.md` and `references/schema.md`).

It sits in Marketing & SEO, covering Market research. It works with Reddit. The repository describes itself as: Open-source Codex skills from BGGG. The licence is MIT.

When your agent uses it

  • Market research
  • Community discovery
  • Keyword snowballing
  • Any task that needs Reddit post and comment text without inventing unavailable fields

Example prompts

  • “/bggg-data-reddit”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Open https://www.reddit.com/ in Chrome. Fetch each planned URL from that Reddit page with same-origin fetch, including pagination through…
  2. Save each unmodified response under data/raw/source_json/ and append one manifest row per saved file. Follow…
  3. Rank directly relevant parent posts by engagement and topical fit. Download complete comment trees for the highest-value parents with…
  4. Normalize and validate
  5. Report query count, parent-post count, comment count, unique rows, deleted/removed rows skipped, unresolved more nodes, HTTP failures, and…

What it can do on your machine

Read from SKILL.md and the folder at commit 1034ee5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • reddit.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bggg Data Reddit loads about 1.2k tokens when it runs, and up to ~1.8k if it reads all its reference files. Until then it costs about 118 tokens; SKILL.md has 430 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~118
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from binggandata/bggg-skills at commit 1034ee5, republished under its MIT licence (© binggandata). 430 words, ~1,194 tokens.

Download SKILL.mdSave it as .claude/skills/bggg-data-reddit/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
bggg-data-reddit
description
Collect auditable Reddit search results and full comment trees at scale, preserve the source JSON, and normalize posts and comments into analysis-ready JSONL. Use for VOC, market research, community discovery, keyword snowballing, or any task that needs Reddit post and comment text without inventing unavailable fields. Reddit data layer of the bggg VOC suite (shared project folder; orchestrated by industry-orchestrator, reported by bggg-voc-report).

BGGG Reddit Data

Use Reddit's public JSON listings through a logged-in or public Chrome tab when direct shell access is blocked. Save every response before transforming it, then normalize locally with the bundled script.

VOC Project Layout(bggg 系列共用)

bggg VOC 系列 skill(bggg-data-amazon / bggg-data-reddit / bggg-data-x / bggg-voc-report / industry-orchestrator)共用一个项目文件夹,让多平台数据规整到同一处、下游分析零改路径。开工先确定项目根目录 <project>(用户指定,或新建 voc-<产品或主题slug>/),并从 <project> 根目录执行本 skill 的全部命令(下文相对路径都基于它):

text
<project>/
  PROJECT.md            # 研究简报 + 决策日志(编排 skill 维护;单独使用可省)
  config/               # 采集目标:amazon_targets.tsv / reddit_queries.tsv / x_queries.tsv / keywords.txt
  work/<platform>/…     # 各平台原始证据、attempt 日志、request plan、manifest
  data/raw/             # 各平台规范化 JSONL(统一行契约,分析共用层)
  data/clean|coded/     # 下游清洗与编码(industry-orchestrator 维护)
  output/               # 报告与交付物(bggg-voc-report 写 output/report/)

本 skill 的落点:config/reddit_queries.tsv → work/reddit/(request plan 与 manifest)+ data/raw/source_json/(未修改源响应)→ data/raw/reddit_<lang>_<date>.jsonl。

Workflow

  1. Define a small query file and build a request plan:
bash
python3 scripts/build_request_plan.py \
  --queries config/reddit_queries.tsv \
  --output work/reddit/request_plan.json

Use tab-separated columns query, lang, and round. lang and round are optional.

  1. Open https://www.reddit.com/ in Chrome. Fetch each planned URL from that Reddit page with same-origin fetch, including pagination through data.after. Do not read or export browser cookies, local storage, profile data, or credentials.

  2. Save each unmodified response under data/raw/source_json/ and append one manifest row per saved file. Follow references/chrome_same_origin.md for the browser loop and manifest format.

  3. Rank directly relevant parent posts by engagement and topical fit. Download complete comment trees for the highest-value parents with /comments/{post_id}.json?limit=500&depth=10&sort=top&raw_json=1. Treat replies as clustered under their parent thread; do not count thousands of replies from one viral thread as independent market prevalence.

  4. Normalize and validate:

bash
python3 scripts/normalize_reddit.py \
  --manifest work/reddit/source_manifest.jsonl \
  --output data/raw/reddit_EN_2026-07-25.jsonl \
  --summary work/reddit/normalize_summary.json
  1. Report query count, parent-post count, comment count, unique rows, deleted/removed rows skipped, unresolved more nodes, HTTP failures, and parent-thread concentration.
Show full SKILL.md (220 more words)Show less

Quality Rules

  • Keep raw source JSON immutable. Write retries to new files rather than overwriting evidence.
  • Use the Reddit fullname (t3_… or t1_…) as native_id; derive source_id with SHA-256.
  • Preserve exact text_raw. Put translations or labels in downstream clean/coded files.
  • Skip [deleted], [removed], empty bodies, ads, and promoted posts; count every exclusion.
  • Search listings are discovery samples, not complete census data. Record sort, time window, query, page cursor, and collection time.
  • Stop pagination only when after is null, the requested cap is reached, or repeated cursors/no-new-ID protection fires.
  • Respect login walls, rate limits, and challenge pages. Back off and record the failure; never bypass authentication or change the user's account.
  • Deduplicate exact native IDs first. Use similarity dedupe only in a later cleaning stage so raw evidence remains auditable.
  • For prevalence statistics, cap or weight rows per parent thread and disclose the rule.

Degradation

If Chrome same-origin fetching is unavailable, try a normal public Reddit JSON request with a conservative user agent. If it is blocked, deliver the request plan and collection manifest with the exact failure; do not substitute search-engine snippets for full Reddit comments.

Resources

  • scripts/build_request_plan.py: create deterministic search request and manifest templates.
  • scripts/normalize_reddit.py: parse saved search listings and nested comment trees into JSONL.
  • references/chrome_same_origin.md: Chrome collection loop, checkpointing, and manifest contract.
  • references/schema.md: normalized row and audit requirements.

© binggandata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in bggg-data-reddit of binggandata/bggg-skills.

  • SKILL.md
  • agents/openai.yaml
  • references/chrome_same_origin.md
  • references/schema.md
  • scripts/build_request_plan.py
  • scripts/normalize_reddit.py

Open the folder on GitHubat commit 1034ee5

Compare with similar skills

Bggg Data Reddit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bggg Data Reddit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bggg Data Reddit this skillbinggandata/bggg-skills605—~1.2kAutomated safety check: PassMIT
Customer ResearchNexus-JPF/note-companion8706 repos~3.2kAutomated safety check: PassMIT
Comment MiningScrapeCreators/social-media-research-skills3.4k—~1kAutomated safety check: NotesMIT
Reddit InsightsBrianRWagner/ai-marketing-claude-code-skills4411 repos~3.1kAutomated safety check: PassNone
Customer Researchunifapi-agent/agents589—~2.1kAutomated safety check: PassMIT
Blog DiscourseAgriciDaniel/claude-blog2.3k1 repos~3.4kAutomated safety check: WarnMIT

Similar skills

  • Customer Research

    Nexus-JPF/note-companion

    When the user wants to conduct, analyze, or synthesize customer research.

    870 GitHub starsUsed in 6 repos~3.2k tokens
    Marketing & SEOAuto-check passed
  • Comment Mining

    ScrapeCreators/social-media-research-skills

    A skill your agent uses when the user wants to mine comments and replies for audience reactions, customer language, questions, objections, complaints, product ideas, buying intent, sentiment, or…

    3.4k GitHub stars~1k tokensUpdated 1 mo ago
    Marketing & SEOAuto-check: notes
  • Reddit Insights

    BrianRWagner/ai-marketing-claude-code-skills

    Search and analyze Reddit content using semantic AI search via reddit-insights.com MCP server.

    441 GitHub starsUsed in 1 repo~3.1k tokens
    Marketing & SEOAuto-check passed
  • Customer Research

    unifapi-agent/agents

    When the user wants to research customers from public communities, or synthesize customer language, pains, and objections.

    589 GitHub stars~2.1k tokensUpdated 1 mo ago
    Marketing & SEOAuto-check passed
  • Blog Discourse

    AgriciDaniel/claude-blog

    Research what people are actually saying about a topic in the last 30 days across Reddit, X / Twitter, YouTube, Hacker News, dev.to, Medium, and other public discourse platforms.

    2.3k GitHub starsUsed in 1 repo~3.4k tokens
    Marketing & SEOAuto-check: warnings
  • Product Demand Research

    ScrapeCreators/social-media-research-skills

    A skill your agent uses when the user wants to validate a product idea, find pain points, mine demand signals, discover objections, or gather voice-of-customer language from Reddit, social posts…

    3.4k GitHub stars~636 tokensUpdated 1 mo ago
    Marketing & SEOAuto-check: notes

More from binggandata/bggg-skills

All 13 skills in this repo
  • Bggg Data Amazon

    binggandata/bggg-skills

    Collect Amazon.com written product reviews at scale through Woot's public review AJAX route, retain every attempt and error log, reconcile partial runs, and normalize exact review text into…

    605 GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Bggg Data X

    binggandata/bggg-skills

    Collect auditable public X/Twitter posts by controlling the user's already logged-in Chrome, searching X's rendered web interface, scrolling visible results, and extracting original post text and…

    605 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Ngs Amazon Image Studio

    binggandata/bggg-skills

    Generate Amazon and cross-border ecommerce product images from a product business card.

    605 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Bggg Tiktok Capcut

    binggandata/bggg-skills

    通用 CapCut 草稿生成与 AI 视频检查 skill。用于把本地 AI 视频套用现有 CapCut 草稿模板, 生成可在 CapCut 首页显示并可编辑的新草稿;也用于提取模板样式、验证草稿结构、抽帧检查 AI 痕迹、规划修复窗口和做本地 RIFE 补帧。

    605 GitHub stars~854 tokensUpdated 1 mo ago
    Auto-check passed
  • Bggg Tiktok Cut

    binggandata/bggg-skills

    用于把 AI 生成的视频、本地素材、口播素材或产品短片剪成可发布到 TikTok 的竖屏成片. An agent skill from binggandata/bggg-skills.

    605 GitHub stars~954 tokensUpdated 1 mo ago
    Auto-check passed
  • Bggg Tiktok Downloader

    binggandata/bggg-skills

    下载 TikTok 视频到本地。当用户给出 TikTok 单个视频链接、分享链接、视频链接文本、 TikTok 博主主页链接、@handle,并要求下载、保存、抓取、批量下载、下载博主作品、 下载指定数量或下载全部作品时,使用此 skill。优先用 yt-dlp,单视频下载失败时用 tikwm 兜底; 可选引用本地 TikTokDownloader 作为链接类型识别辅助。

    605 GitHub stars~494 tokensUpdated 1 mo ago
    Auto-check passed

Works with

Categories

Questions about Bggg Data Reddit

What does Bggg Data Reddit do?

Collect auditable Reddit search results and full comment trees at scale, preserve the source JSON, and normalize posts and comments into analysis-ready JSONL. Bggg Data Reddit is an agent skill from binggandata/bggg-skills. Collect auditable Reddit search results and full comment trees at scale, preserve the source JSON, and normalize posts and comments into analysis-ready JSONL.

When should I use Bggg Data Reddit?

Bggg Data Reddit fits situations like: market research; community discovery; keyword snowballing; any task that needs Reddit post and comment text without inventing unavailable fields.

How do I install Bggg Data Reddit in Claude Code?

Run `npx skills add binggandata/bggg-skills --skill bggg-data-reddit -a claude-code`. Or copy the skill folder (bggg-data-reddit in binggandata/bggg-skills) into .claude/skills/bggg-data-reddit in your project. Claude Code loads it when a task matches its description.

How do I install Bggg Data Reddit in Codex?

Run `npx skills add binggandata/bggg-skills --skill bggg-data-reddit -a codex`. Or copy the skill folder (bggg-data-reddit in binggandata/bggg-skills) into .agents/skills/bggg-data-reddit in your project. Codex loads it when a task matches its description.

Can I use Bggg Data Reddit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add binggandata/bggg-skills --skill bggg-data-reddit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bggg-data-reddit, .gemini/skills/bggg-data-reddit, .github/skills/bggg-data-reddit and .opencode/skills/bggg-data-reddit in your project.

What does Bggg Data Reddit need to run?

Going by SKILL.md and its folder, Bggg Data Reddit needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Bggg Data Reddit access the network?

SKILL.md names 1 domain. In commands or code: reddit.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Bggg Data Reddit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Bggg Data Reddit use?

Bggg Data Reddit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bggg Data Reddit use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 614 tokens, read only when the agent opens those files.

What are the alternatives to Bggg Data Reddit?

Skills that share tags, products or a category with Bggg Data Reddit: Customer Research (Nexus-JPF/note-companion, 870 stars), Comment Mining (ScrapeCreators/social-media-research-skills, 3.4k stars), Reddit Insights (BrianRWagner/ai-marketing-claude-code-skills, 441 stars) and Customer Research (unifapi-agent/agents, 589 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bggg Data Reddit?

binggandata (a GitHub user) maintains it in binggandata/bggg-skills, which has 605 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on August 13, 2026.

Source: binggandata/bggg-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.