Blog
AgriciDaniel/claude-blog
Full-lifecycle blog engine with 31 sub-skills, 12 templates, 100-point scoring, and 5 agents.
Find a dataset for a Data2Story blog. An agent skill from QinghongLin/data2story-skill.
$ npx skills add QinghongLin/data2story-skill --skill find-data -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install QinghongLin/data2story-skill find-data --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/QinghongLin/data2story-skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/find-data .claude/skills/find-data && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "find-data" agent skill from https://github.com/QinghongLin/data2story-skill/tree/main/skills/find-data into .claude/skills/find-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "find-data", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/QinghongLin/data2story-skill/tree/main/skills/find-dataType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add QinghongLin/data2story-skill --skill find-data -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install QinghongLin/data2story-skill find-data --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/QinghongLin/data2story-skill.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/find-data .agents/skills/find-data && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "find-data" agent skill from https://github.com/QinghongLin/data2story-skill/tree/main/skills/find-data into .agents/skills/find-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "find-data", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add QinghongLin/data2story-skill --skill find-data -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install QinghongLin/data2story-skill find-data --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/QinghongLin/data2story-skill.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/find-data .cursor/skills/find-data && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "find-data" agent skill from https://github.com/QinghongLin/data2story-skill/tree/main/skills/find-data into .cursor/skills/find-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "find-data", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/QinghongLin/data2story-skill.git --path skills/find-data--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add QinghongLin/data2story-skill --skill find-data -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install QinghongLin/data2story-skill find-data --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/QinghongLin/data2story-skill.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/find-data .gemini/skills/find-data && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "find-data" agent skill from https://github.com/QinghongLin/data2story-skill/tree/main/skills/find-data into .gemini/skills/find-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "find-data", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install QinghongLin/data2story-skill find-dataInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add QinghongLin/data2story-skill --skill find-data -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/QinghongLin/data2story-skill.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/find-data .github/skills/find-data && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "find-data" agent skill from https://github.com/QinghongLin/data2story-skill/tree/main/skills/find-data into .github/skills/find-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "find-data", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add QinghongLin/data2story-skill --skill find-data -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install QinghongLin/data2story-skill find-data --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/QinghongLin/data2story-skill.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/find-data .opencode/skills/find-data && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "find-data" agent skill from https://github.com/QinghongLin/data2story-skill/tree/main/skills/find-data into .opencode/skills/find-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "find-data", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
find-dataFind a dataset for a Data2Story blog. An agent skill from QinghongLin/data2story-skill.
Find Data is an agent skill from QinghongLin/data2story-skill. Find a dataset for a Data2Story blog. Accepts a topic, a URL, or a DIP-style category. Downloads + validates against 4 completeness gates before handing off to /data2story-pro. Local-first: searches Economist/Pudding/TidyTuesday clones before going online. Supports --validate-only to audit a folder you already have.
Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including reference files (for example `references/completeness_gates.md`, `references/examples/good_economist.md` and `references/examples/good_theme.md`).
The repository describes itself as: Data Journalist Agent: Transforming Data into Verifiable Multimodal Story. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 63a55c1. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
Bash(python:*)ReadWriteGlobGrepWebSearchWebFetchAskUserQuestionFrom allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonpipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Find Data loads about 3.3k tokens when it runs, and up to ~7k if it reads all its reference files. Until then it costs about 82 tokens; SKILL.md has 1,510 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from QinghongLin/data2story-skill at commit 63a55c1, republished under its MIT licence (© QinghongLin). 1,510 words, ~3,261 tokens.
.claude/skills/find-data/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.Turn an idea, URL, or category into a phase2/datasets/<name>/ folder
that the /data2story-pro pipeline can run on without crashing.
You are the gatekeeper before the 7-agent newsroom. Detective, Analyst, Editor, Designer, Programmer, Auditor, Inspector all assume the data is already there, parseable, and provenanced. Your job is to make sure that's true before they start.
Refuse to mark a folder ready until it passes 4 gates. See
references/completeness_gates.md for the criteria. The gates are codified
in tools/audit.py, not in this prose.
Python deps for the tools/:
pandas — required (audit.py reads/inspects CSV/JSON).openpyxl — required only when a source is .xlsx (audit.py's
pd.ExcelFile); a mid-run missing-openpyxl is the usual cause of an XLSX
audit error.urllib — stdlib, no install (fetch.py downloads, audit.py HEAD-checks).One install line:
pip install pandas openpyxlAfter editing anything under tools/, run the no-network regression suite:
py tools/selftest.py (exit 0 = pass).
SKILL_DIR = directory containing this SKILL.mdWORKSPACE = ancestor that contains phase2/datasets/ (the parent repo root), if one existsDATASETS_ROOT = WORKSPACE/phase2/datasets when WORKSPACE exists; otherwise ./datasets
(clone-relative, under the current working dir) — the open-source clone has no phase2/datasetsBLOGS_ROOT = WORKSPACE/phase2/blogs when WORKSPACE existsOUT_DIR default = DATASETS_ROOT/<name> (so ./datasets/<name> on an OSS clone). Callers may
override with an explicit --out, which always wins.INPUT = first positional argument from $ARGUMENTS$ARGUMENTS: --mode, --source, --out, --validate-onlyNever hard-code machine paths. Resolve at runtime by walking up from SKILL_DIR.
Decision tree (the FIRST condition that matches wins):
--validate-only flag present → MODE = validate-only. INPUT must be an existing folder path.INPUT is an existing folder path (and --validate-only absent) → still ask the user before re-fetching; default to validate-only behaviour with a confirmation.INPUT starts with http:// or https:// → MODE = url.INPUT (case-insensitive) exactly matches a DIP category from phase2/datasets/data_is_plural/dip_category_summary.md → MODE = category.topic.State your classification out loud in one line before doing anything else.
No local corpora (open-source clone): on a machine without the Economist/, Pudding/,
tidytuesday/ clones and without the Data-is-Plural CSV under phase2/datasets/, browse_local.py
and dip_query.py both return empty (no crash). In that case there is nothing to match locally, so
topic/category discovery degrades straight to WebSearch (Step 1). No flag is needed — detect it
from the empty browse_local.py + dip_query.py results. Keep the local-first path intact for
machines that DO have the corpora.
Just audit. Skip discovery, skip fetch.
python "SKILL_DIR/tools/audit.py" "<folder>"This writes <folder>/validate.json and prints the gate summary. Read it back
with Read, surface the verdict, and stop.
If overall.ready_for_data2story is true → print the /data2story-pro <folder>
command for the user to copy.
If it's false → list the specific gate failures and recommend remediations.
Determine OUT_DIR (use --out if given, else derive a slug from the URL path's
last meaningful segment, store under DATASETS_ROOT/<slug>/ — i.e. ./datasets/<slug>/
on an open-source clone with no phase2/datasets).
Branch on URL shape:
| URL pattern | Tool |
|---|---|
github.com/{owner}/{repo}/tree/{branch}/{path} | python fetch.py github-folder <url> <out_dir> |
github.com/{owner}/{repo}/blob/{branch}/{path} | Same (fetch.py auto-detects blob) |
github.com/{owner}/{repo} with no path | Refuse — tell user to narrow to a folder |
| Anything else | python fetch.py url <url> <out_dir> |
fetch.py is source-aware: if the URL points to TheEconomist/graphic-detail-data,
the-pudding/data, or rfordatascience/tidytuesday AND the local clone already
contains that path, it copies from the local clone (no network). Otherwise it
fetches via the GitHub API. This happens automatically — do not handle it yourself.
Several hosts serve a human-facing page at the obvious URL, not the raw bytes,
so a naive download returns HTML instead of a .csv/.json/.xlsx. fetch.py
auto-rewrites GitHub /blob/ URLs to raw.githubusercontent.com and, as a
backstop, warns on stderr if a data-extension download comes back as an HTML
page (<!doctype html> / <html). The other common share-hosts are not
auto-rewritten — pass the direct-download form yourself:
| Host | Page URL (returns HTML) | Direct-download form to use |
|---|---|---|
| GitHub | github.com/<o>/<r>/blob/<branch>/<path> | auto-handled → raw.githubusercontent.com/<o>/<r>/<branch>/<path> |
| Google Drive | drive.google.com/file/d/<id>/view | drive.google.com/uc?export=download&id=<id> |
| Dropbox | ...?dl=0 | swap to ...?dl=1 (or dl.dropboxusercontent.com) |
| OneDrive | share link (1drv.ms/...) | append &download=1 to the direct link |
| Kaggle | dataset page (kaggle.com/datasets/...) | the file API / CLI download — the page is not a file |
If a download trips the HTML warning, you almost certainly handed fetch.py a
share/page URL — fix the URL to its direct-download form and re-fetch.
After fetch: check whether the folder has a README.md.
tools/README_template.md. Auto-fill what you can
(title from folder name, column headers from the first CSV, source URL from
the manifest). Leave the codebook definition cells as {TODO} for the user
to fill in. Tell the user explicitly that README is a stub.INPUT is one of the 14 DIP categories. Run both browse and DIP query in parallel:
python tools/browse_local.py "<INPUT>" --top 10
python tools/dip_query.py --category "<INPUT>" --top 10Both emit JSON. Merge:
Present top 5 to the user as a numbered list via AskUserQuestion:
Then route the pick:
links array. If multiple URLs,
pick the one with a data-file extension (csv/xlsx/json) or the GitHub one.
Route to url mode for that URL.Same as category mode but use full-text query:
python tools/browse_local.py "<INPUT>"
python tools/dip_query.py "<INPUT>" --top 10If BOTH return empty (no local hits, no DIP hits), only THEN do WebSearch:
<INPUT> open dataset CSV site:ourworldindata.org OR site:github.com OR site:data.gov.
Cap web results at 5. Surface them with a "no local matches found, querying web" disclaimer.
From the WebSearch results, pick the one whose URL ends in a data-file extension
(.csv / .xlsx / .json / .tsv) or is a github.com/{owner}/{repo}/tree/{branch}/{path}
folder. Route that URL through url mode: python tools/fetch.py url <url> <OUT_DIR> for a direct
data file, or python tools/fetch.py github-folder <url> <OUT_DIR> for a GitHub folder. Then
continue to the README step (Step 2) and the audit step (Step 3) exactly as in url mode — this
closes the topic → web → fetch → audit loop. If no result has a usable data URL, report that
honestly and do not fabricate one.
After fetch, the OUT_DIR may or may not have a README. Decide:
tools/README_template.mdWhen generating, fill in:
{TITLE}: human-readable from folder slug (snake_case → "snake case", title-cased){PRIMARY SOURCE URL}: from manifest.json items' urls{filename.csv} rows: scan each CSV's columns{Codebook}: list each column with {TODO: definition} for unknown ones; you
may guess from column names (launch_year → "year of launch (integer)") but
flag guesses with a {?} prefix so the user can correct.python "SKILL_DIR/tools/audit.py" "<OUT_DIR>"This writes <OUT_DIR>/validate.json. Read it back.
For Gate 4 (multimodal — advisory only), audit.py cannot judge by itself.
Do your own quick check after reading validate.json:
From the CSV columns and entity samples, classify subjects:
For concrete subjects, do ONE WebSearch with site:commons.wikimedia.org
to verify at least 1 reference photo exists for a representative subject.
Do NOT download photos — just verify they exist. That's Detective's job.
Write your Gate 4 finding back into validate.json under
gates.multimodal.info.
Read the final validate.json. Print a concise summary:
=== find-data verdict ===
folder: <OUT_DIR>
files: <N>
[OK ] Technical
[WARN] Story material
- <filename.csv>: only 47 rows (< 50)
[OK ] Provenance
[INFO] Multimodal: concrete subjects (agencies, rocket types)
overall: READY
next: /data2story-pro <OUT_DIR>If overall is BLOCKED, do NOT print the /data2story-pro command. Instead print:
overall: BLOCKED
to fix: <specific remediations from the fails list>Suggest concrete fixes:
<folder> after adding README"<path> and re-run")Append a one-line entry to DATASETS_ROOT/_find-data-log.md (create if not
present) so the user has a running history:
- 2026-05-25T14:00Z · single · datasets/<name> · READY · invoked via URLDATASETS_ROOT/<OUT_DIR>/) and the digest log.{PRIMARY SOURCE URL} placeholder and warn the user.references/completeness_gates.md — exact pass/warn/fail criteria for each gatereferences/input_mode_dispatch.md — full URL classification tablereferences/examples/good_economist.md — what single mode output should look likereferences/examples/good_theme.md — what theme mode output should look liketools/audit.py — deterministic gate evaluatortools/fetch.py — source-aware downloader (manifest / github-folder / single-url modes)tools/browse_local.py — index Economist/Pudding/TidyTuesday for candidate matchingtools/dip_query.py — filter dip_categorized.csv by category or topictools/README_template.md — graphic-detail-style skeleton for auto-generation© QinghongLin, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 10 other files (references) in skills/find-data of QinghongLin/data2story-skill.
Open the folder on GitHubat commit 63a55c1
Find Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Find Data this skillQinghongLin/data2story-skill | 156 | — | ~3.3k | Automated safety check: Pass | MIT | |
| BlogAgriciDaniel/claude-blog | 2.3k | — | ~6.2k | Automated safety check: Pass | MIT | |
| BlogAgriciDaniel/claude-blog | 2.3k | 1 repos | ~8.6k | Automated safety check: Warn | MIT | |
| Notion To Blogwasp-lang/wasp | 19k | — | ~922 | Automated safety check: Pass | MIT | |
| Blog Writing Guidesickn33/agentic-awesome-skills | 47k | 2 repos | ~2.2k | Automated safety check: Pass | MIT | |
| Blog Topic Researchjeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~2k | Automated safety check: Pass | MIT-0 |
AgriciDaniel/claude-blog
Full-lifecycle blog engine with 31 sub-skills, 12 templates, 100-point scoring, and 5 agents.
AgriciDaniel/claude-blog
Full-lifecycle blog engine with 31 sub-skills, 12 content templates, 5-category 100-point scoring, and 5 specialized agents.
wasp-lang/wasp
Transfer a blog post from Notion to the Wasp blog. An agent skill from wasp-lang/wasp.
sickn33/agentic-awesome-skills
This skill enforces Sentry's blog writing standards across every post — whether you're helping an engineer write their first blog post or a marketer draft a product announcement.
jeremylongshore/tons-of-skills-marketplace
Build a deduplicated editorial backlog from current, traceable demand and authoritative product evidence.
luongnv89/claude-howto
Guides a blog post from idea to draft in stages: research notes from your sources, brainstorming and clarifying questions, then outlining and versioned drafting.
QinghongLin/data2story-skill
Run sentence-level traceability verification on a Data2Story blog (verify.py - verifier.json), then emit the in-page Inspector panel (the reader-facing runnable verifier) + the verify/ artifacts…
QinghongLin/data2story-skill
Audit a generated Data2Story blog for build correctness across ALL modalities by ACTUALLY RENDERING it in a real headless browser (when available) — catching blank/0-width charts, broken/oversized…
QinghongLin/data2story-skill
Review a finished Data2Story blog against the 5 quality rubric dimensions (visualdesign, narrativepacing, datamethodtransparency, claimdataalignment, insightvalue), score each 1-7 with on-page…
QinghongLin/data2story-skill
Research external context for a dataset — domain background, history, related studies, and why this data matters.
QinghongLin/data2story-skill
Generate music (NOT speech) via OpenRouter using Google Lyria 3 Pro.
QinghongLin/data2story-skill
Run sentence-level traceability verification on a blog, then generate viewer.html with interactive evidence panel.
Find a dataset for a Data2Story blog. An agent skill from QinghongLin/data2story-skill. Find Data is an agent skill from QinghongLin/data2story-skill. Find a dataset for a Data2Story blog.
Run `npx skills add QinghongLin/data2story-skill --skill find-data -a claude-code`. Or copy the skill folder (skills/find-data in QinghongLin/data2story-skill) into .claude/skills/find-data in your project. Claude Code loads it when a task matches its description.
Run `npx skills add QinghongLin/data2story-skill --skill find-data -a codex`. Or copy the skill folder (skills/find-data in QinghongLin/data2story-skill) into .agents/skills/find-data in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add QinghongLin/data2story-skill --skill find-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/find-data, .gemini/skills/find-data, .github/skills/find-data and .opencode/skills/find-data in your project.
Going by SKILL.md and its folder, Find Data needs Python for the scripts in its folder and the command-line tools its instructions call (python and pip). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash(python:*), Read, Write, Glob, Grep, WebSearch, WebFetch, AskUserQuestion.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Find Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.8k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Find Data: Blog (AgriciDaniel/claude-blog, 2.3k stars), Blog (AgriciDaniel/claude-blog, 2.3k stars), Notion To Blog (wasp-lang/wasp, 19k stars) and Blog Writing Guide (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
QinghongLin (a GitHub user) maintains it in QinghongLin/data2story-skill, which has 156 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on July 5, 2026.
Source: QinghongLin/data2story-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.