Agent skill

Find Data

by QinghongLin in QinghongLin/data2story-skill

Find a dataset for a Data2Story blog. An agent skill from QinghongLin/data2story-skill.

MITAuto-check passed

Install Find Data

skills CLI
$ npx skills add QinghongLin/data2story-skill --skill find-data -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install QinghongLin/data2story-skill find-data --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/QinghongLin/data2story-skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/find-data .claude/skills/find-data && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
find-data
GitHub stars
156
Token cost
~3.3k tokens
SKILL.md length
1,510 words
Files
11 (incl. references)
Skills in repo
31
Repo updated
First seen
Licence
MIT

At a glance

Find a dataset for a Data2Story blog. An agent skill from QinghongLin/data2story-skill.

  • Works in 6 steps: Classify the input → Branch on mode → Optional: generate / repair the README → …
  • SKILL.md covers Prerequisites, Resolve paths first, Step 0 — Classify the input and Step 1 — Branch on mode, plus 6 more sections
  • Runs Python scripts from its folder; calls python and pip

What it does

Find Data is an agent skill from QinghongLin/data2story-skill. Find a dataset for a Data2Story blog. Accepts a topic, a URL, or a DIP-style category. Downloads + validates against 4 completeness gates before handing off to /data2story-pro. Local-first: searches Economist/Pudding/TidyTuesday clones before going online. Supports --validate-only to audit a folder you already have.

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including reference files (for example `references/completeness_gates.md`, `references/examples/good_economist.md` and `references/examples/good_theme.md`).

The repository describes itself as: Data Journalist Agent: Transforming Data into Verifiable Multimodal Story. The licence is MIT.

Example prompts

  • “/find-data”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Bash(python:*), Read, Write, Glob, Grep, WebSearch, WebFetch, AskUserQuestion

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Classify the input
  2. Branch on mode
  3. Optional: generate / repair the README
  4. Audit (the 4 gates)
  5. Verdict
  6. Save a digest

What it can do on your machine

Read from SKILL.md and the folder at commit 63a55c1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(python:*)
    • Read
    • Write
    • Glob
    • Grep
    • WebSearch
    • WebFetch
    • AskUserQuestion

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Find Data loads about 3.3k tokens when it runs, and up to ~7k if it reads all its reference files. Until then it costs about 82 tokens; SKILL.md has 1,510 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from QinghongLin/data2story-skill at commit 63a55c1, republished under its MIT licence (© QinghongLin). 1,510 words, ~3,261 tokens.

Download SKILL.mdSave it as .claude/skills/find-data/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
find-data
description
Find a dataset for a Data2Story blog. Accepts a topic, a URL, or a DIP-style category. Downloads + validates against 4 completeness gates before handing off to /data2story-pro. Local-first: searches Economist/Pudding/TidyTuesday clones before going online. Supports --validate-only to audit a folder you already have.
allowed-tools
Bash(python:*), Read, Write, Glob, Grep, WebSearch, WebFetch, AskUserQuestion
argument-hint
<topic | URL | category | folder-path> [--mode single|theme] [--source Economist|Pudding|tidytuesday] [--out ./datasets/<name>] [--validate-only]

find-data

Turn an idea, URL, or category into a phase2/datasets/<name>/ folder that the /data2story-pro pipeline can run on without crashing.

You are the gatekeeper before the 7-agent newsroom. Detective, Analyst, Editor, Designer, Programmer, Auditor, Inspector all assume the data is already there, parseable, and provenanced. Your job is to make sure that's true before they start.

Refuse to mark a folder ready until it passes 4 gates. See references/completeness_gates.md for the criteria. The gates are codified in tools/audit.py, not in this prose.

Prerequisites

Python deps for the tools/:

  • pandas — required (audit.py reads/inspects CSV/JSON).
  • openpyxl — required only when a source is .xlsx (audit.py's pd.ExcelFile); a mid-run missing-openpyxl is the usual cause of an XLSX audit error.
  • urllib — stdlib, no install (fetch.py downloads, audit.py HEAD-checks).

One install line:

bash
pip install pandas openpyxl

After editing anything under tools/, run the no-network regression suite: py tools/selftest.py (exit 0 = pass).

Resolve paths first

  • SKILL_DIR = directory containing this SKILL.md
  • WORKSPACE = ancestor that contains phase2/datasets/ (the parent repo root), if one exists
  • DATASETS_ROOT = WORKSPACE/phase2/datasets when WORKSPACE exists; otherwise ./datasets (clone-relative, under the current working dir) — the open-source clone has no phase2/datasets
  • BLOGS_ROOT = WORKSPACE/phase2/blogs when WORKSPACE exists
  • OUT_DIR default = DATASETS_ROOT/<name> (so ./datasets/<name> on an OSS clone). Callers may override with an explicit --out, which always wins.
  • INPUT = first positional argument from $ARGUMENTS
  • Parse flags from $ARGUMENTS: --mode, --source, --out, --validate-only

Never hard-code machine paths. Resolve at runtime by walking up from SKILL_DIR.

Step 0 — Classify the input

Decision tree (the FIRST condition that matches wins):

  1. --validate-only flag present → MODE = validate-only. INPUT must be an existing folder path.
  2. INPUT is an existing folder path (and --validate-only absent) → still ask the user before re-fetching; default to validate-only behaviour with a confirmation.
  3. INPUT starts with http:// or https:// → MODE = url.
  4. INPUT (case-insensitive) exactly matches a DIP category from phase2/datasets/data_is_plural/dip_category_summary.md → MODE = category.
  5. Otherwise → MODE = topic.

State your classification out loud in one line before doing anything else.

No local corpora (open-source clone): on a machine without the Economist/, Pudding/, tidytuesday/ clones and without the Data-is-Plural CSV under phase2/datasets/, browse_local.py and dip_query.py both return empty (no crash). In that case there is nothing to match locally, so topic/category discovery degrades straight to WebSearch (Step 1). No flag is needed — detect it from the empty browse_local.py + dip_query.py results. Keep the local-first path intact for machines that DO have the corpora.

Step 1 — Branch on mode

validate-only mode

Just audit. Skip discovery, skip fetch.

bash
python "SKILL_DIR/tools/audit.py" "<folder>"

This writes <folder>/validate.json and prints the gate summary. Read it back with Read, surface the verdict, and stop.

If overall.ready_for_data2story is true → print the /data2story-pro <folder> command for the user to copy.

If it's false → list the specific gate failures and recommend remediations.

url mode

Determine OUT_DIR (use --out if given, else derive a slug from the URL path's last meaningful segment, store under DATASETS_ROOT/<slug>/ — i.e. ./datasets/<slug>/ on an open-source clone with no phase2/datasets).

Branch on URL shape:

URL patternTool
github.com/{owner}/{repo}/tree/{branch}/{path}python fetch.py github-folder <url> <out_dir>
github.com/{owner}/{repo}/blob/{branch}/{path}Same (fetch.py auto-detects blob)
github.com/{owner}/{repo} with no pathRefuse — tell user to narrow to a folder
Anything elsepython fetch.py url <url> <out_dir>

fetch.py is source-aware: if the URL points to TheEconomist/graphic-detail-data, the-pudding/data, or rfordatascience/tidytuesday AND the local clone already contains that path, it copies from the local clone (no network). Otherwise it fetches via the GitHub API. This happens automatically — do not handle it yourself.

URL normalization (share-page vs direct-download)

Several hosts serve a human-facing page at the obvious URL, not the raw bytes, so a naive download returns HTML instead of a .csv/.json/.xlsx. fetch.py auto-rewrites GitHub /blob/ URLs to raw.githubusercontent.com and, as a backstop, warns on stderr if a data-extension download comes back as an HTML page (<!doctype html> / <html). The other common share-hosts are not auto-rewritten — pass the direct-download form yourself:

HostPage URL (returns HTML)Direct-download form to use
GitHubgithub.com/<o>/<r>/blob/<branch>/<path>auto-handled → raw.githubusercontent.com/<o>/<r>/<branch>/<path>
Google Drivedrive.google.com/file/d/<id>/viewdrive.google.com/uc?export=download&id=<id>
Dropbox...?dl=0swap to ...?dl=1 (or dl.dropboxusercontent.com)
OneDriveshare link (1drv.ms/...)append &download=1 to the direct link
Kaggledataset page (kaggle.com/datasets/...)the file API / CLI download — the page is not a file

If a download trips the HTML warning, you almost certainly handed fetch.py a share/page URL — fix the URL to its direct-download form and re-fetch.

After fetch: check whether the folder has a README.md.

  • If yes → leave it alone, jump to Step 2.
  • If no → generate one from tools/README_template.md. Auto-fill what you can (title from folder name, column headers from the first CSV, source URL from the manifest). Leave the codebook definition cells as {TODO} for the user to fill in. Tell the user explicitly that README is a stub.
category mode

INPUT is one of the 14 DIP categories. Run both browse and DIP query in parallel:

bash
python tools/browse_local.py "<INPUT>" --top 10
python tools/dip_query.py --category "<INPUT>" --top 10

Both emit JSON. Merge:

  • Local candidates (already-cloned folders) score higher than DIP leads (which still need fetching).
  • For each local candidate, run a cheap pre-score (does Gate 1 trivially pass — is there ≥ 1 CSV in the folder?).

Present top 5 to the user as a numbered list via AskUserQuestion:

  • For each candidate, show: source · title · date · top-3 files · 1-line readme excerpt
  • Options: pick 1–5, or "all" (theme mode), or "skip" (abort)

Then route the pick:

  • Local candidate → copy folder into OUT_DIR (or in-place audit if user accepts).
  • DIP lead → extract URL from the candidate's links array. If multiple URLs, pick the one with a data-file extension (csv/xlsx/json) or the GitHub one. Route to url mode for that URL.
Show full SKILL.md (614 more words)Show less
topic mode

Same as category mode but use full-text query:

bash
python tools/browse_local.py "<INPUT>"
python tools/dip_query.py "<INPUT>" --top 10

If BOTH return empty (no local hits, no DIP hits), only THEN do WebSearch: <INPUT> open dataset CSV site:ourworldindata.org OR site:github.com OR site:data.gov. Cap web results at 5. Surface them with a "no local matches found, querying web" disclaimer.

From the WebSearch results, pick the one whose URL ends in a data-file extension (.csv / .xlsx / .json / .tsv) or is a github.com/{owner}/{repo}/tree/{branch}/{path} folder. Route that URL through url mode: python tools/fetch.py url <url> <OUT_DIR> for a direct data file, or python tools/fetch.py github-folder <url> <OUT_DIR> for a GitHub folder. Then continue to the README step (Step 2) and the audit step (Step 3) exactly as in url mode — this closes the topic → web → fetch → audit loop. If no result has a usable data URL, report that honestly and do not fabricate one.

Step 2 — Optional: generate / repair the README

After fetch, the OUT_DIR may or may not have a README. Decide:

  • README present + has codebook table + has source mention → leave alone
  • README missing → generate from tools/README_template.md
  • README present but no codebook table → ADD a codebook section, do not overwrite the existing prose

When generating, fill in:

  • {TITLE}: human-readable from folder slug (snake_case → "snake case", title-cased)
  • {PRIMARY SOURCE URL}: from manifest.json items' urls
  • {filename.csv} rows: scan each CSV's columns
  • {Codebook}: list each column with {TODO: definition} for unknown ones; you may guess from column names (launch_year → "year of launch (integer)") but flag guesses with a {?} prefix so the user can correct.

Step 3 — Audit (the 4 gates)

bash
python "SKILL_DIR/tools/audit.py" "<OUT_DIR>"

This writes <OUT_DIR>/validate.json. Read it back.

For Gate 4 (multimodal — advisory only), audit.py cannot judge by itself. Do your own quick check after reading validate.json:

  • From the CSV columns and entity samples, classify subjects:

    • Concrete (people / places / species / events / specific objects) — e.g., "launches.csv has agency names like SpaceX, NASA, Roscosmos → concrete"
    • Abstract (rates, indices, generic counts) — e.g., "GDP per capita time series → abstract"
  • For concrete subjects, do ONE WebSearch with site:commons.wikimedia.org to verify at least 1 reference photo exists for a representative subject. Do NOT download photos — just verify they exist. That's Detective's job.

  • Write your Gate 4 finding back into validate.json under gates.multimodal.info.

Step 4 — Verdict

Read the final validate.json. Print a concise summary:

=== find-data verdict ===
folder: <OUT_DIR>
files: <N>
[OK ] Technical
[WARN] Story material
       - <filename.csv>: only 47 rows (< 50)
[OK ] Provenance
[INFO] Multimodal: concrete subjects (agencies, rocket types)

overall: READY
next:    /data2story-pro <OUT_DIR>

If overall is BLOCKED, do NOT print the /data2story-pro command. Instead print:

overall: BLOCKED
to fix: <specific remediations from the fails list>

Suggest concrete fixes:

  • "README missing → run /find-data --validate-only <folder> after adding README"
  • "Primary URL dead → find an alternate source via /find-data --source-finder"
  • (or open-ended: "drop a corrected file into <path> and re-run")

Step 5 — Save a digest

Append a one-line entry to DATASETS_ROOT/_find-data-log.md (create if not present) so the user has a running history:

- 2026-05-25T14:00Z · single · datasets/<name> · READY · invoked via URL

Constraints

  • Never overwrite an existing README.md without asking. Generate-from-template only fires when README is absent. If README exists but is thin, ADD sections, don't replace.
  • Never modify the skill repo itself. Your only writes are inside the dataset output directory (DATASETS_ROOT/<OUT_DIR>/) and the digest log.
  • Never fabricate a source URL or license. If you can't find a real source for a generated README, leave the {PRIMARY SOURCE URL} placeholder and warn the user.
  • Cap WebSearch at 5 results, cap WebFetch at 3 calls per invocation. This is a gating skill, not a research tool — depth is Detective's job.
  • Do not run /data2story-pro yourself. Your last word is the verdict + the command for the user to copy. Hand-off is manual.

Reference files

  • references/completeness_gates.md — exact pass/warn/fail criteria for each gate
  • references/input_mode_dispatch.md — full URL classification table
  • references/examples/good_economist.md — what single mode output should look like
  • references/examples/good_theme.md — what theme mode output should look like
  • tools/audit.py — deterministic gate evaluator
  • tools/fetch.py — source-aware downloader (manifest / github-folder / single-url modes)
  • tools/browse_local.py — index Economist/Pudding/TidyTuesday for candidate matching
  • tools/dip_query.py — filter dip_categorized.csv by category or topic
  • tools/README_template.md — graphic-detail-style skeleton for auto-generation

© QinghongLin, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (references) in skills/find-data of QinghongLin/data2story-skill.

  • SKILL.md
  • references/completeness_gates.md
  • references/examples/good_economist.md
  • references/examples/good_theme.md
  • references/input_mode_dispatch.md
  • tools/README_template.md
  • tools/audit.py
  • tools/browse_local.py
  • tools/dip_query.py
  • tools/fetch.py
  • tools/selftest.py

Open the folder on GitHubat commit 63a55c1

Compare with similar skills

Find Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Find Data compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Find Data this skillQinghongLin/data2story-skill156—~3.3kAutomated safety check: PassMIT
BlogAgriciDaniel/claude-blog2.3k—~6.2kAutomated safety check: PassMIT
BlogAgriciDaniel/claude-blog2.3k1 repos~8.6kAutomated safety check: WarnMIT
Notion To Blogwasp-lang/wasp19k—~922Automated safety check: PassMIT
Blog Writing Guidesickn33/agentic-awesome-skills47k2 repos~2.2kAutomated safety check: PassMIT
Blog Topic Researchjeremylongshore/tons-of-skills-marketplace2.8k—~2kAutomated safety check: PassMIT-0

Similar skills

  • Blog

    AgriciDaniel/claude-blog

    Full-lifecycle blog engine with 31 sub-skills, 12 templates, 100-point scoring, and 5 agents.

    2.3k GitHub stars~6.2k tokensUpdated yesterday
    Writing & ContentAuto-check passed
  • Blog

    AgriciDaniel/claude-blog

    Full-lifecycle blog engine with 31 sub-skills, 12 content templates, 5-category 100-point scoring, and 5 specialized agents.

    2.3k GitHub starsUsed in 1 repo~8.6k tokens
    Writing & ContentAuto-check: warnings
  • Notion To Blog

    wasp-lang/wasp

    Transfer a blog post from Notion to the Wasp blog. An agent skill from wasp-lang/wasp.

    19k GitHub stars~922 tokensUpdated yesterday
    Writing & ContentAuto-check passed
  • Blog Writing Guide

    sickn33/agentic-awesome-skills

    This skill enforces Sentry's blog writing standards across every post — whether you're helping an engineer write their first blog post or a marketer draft a product announcement.

    47k GitHub starsUsed in 2 repos~2.2k tokens
    Writing & ContentAuto-check passed
  • Blog Topic Research

    jeremylongshore/tons-of-skills-marketplace

    Build a deduplicated editorial backlog from current, traceable demand and authoritative product evidence.

    2.8k GitHub stars~2k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Blog Post Drafter

    luongnv89/claude-howto

    Guides a blog post from idea to draft in stages: research notes from your sources, brainstorming and clarifying questions, then outlining and versioned drafting.

    42k GitHub stars~2.2k tokensUpdated 8 days ago
    Writing & ContentAuto-check passed

More from QinghongLin/data2story-skill

All 31 skills in this repo
  • Inspector

    QinghongLin/data2story-skill

    Run sentence-level traceability verification on a Data2Story blog (verify.py - verifier.json), then emit the in-page Inspector panel (the reader-facing runnable verifier) + the verify/ artifacts…

    156 GitHub starsUsed in 1 repo~3.1k tokens
    Auto-check: notes
  • Auditor

    QinghongLin/data2story-skill

    Audit a generated Data2Story blog for build correctness across ALL modalities by ACTUALLY RENDERING it in a real headless browser (when available) — catching blank/0-width charts, broken/oversized…

    156 GitHub stars~6.4k tokensUpdated 3 mo ago
    Auto-check: notes
  • Critic

    QinghongLin/data2story-skill

    Review a finished Data2Story blog against the 5 quality rubric dimensions (visualdesign, narrativepacing, datamethodtransparency, claimdataalignment, insightvalue), score each 1-7 with on-page…

    156 GitHub stars~4.7k tokensUpdated 3 mo ago
    Auto-check: notes
  • Detective

    QinghongLin/data2story-skill

    Research external context for a dataset — domain background, history, related studies, and why this data matters.

    156 GitHub stars~2.4k tokensUpdated 3 mo ago
    Auto-check: notes
  • Openrouter Text2music

    QinghongLin/data2story-skill

    Generate music (NOT speech) via OpenRouter using Google Lyria 3 Pro.

    156 GitHub starsUsed in 1 repo~407 tokens
    Auto-check passed
  • Inspector

    QinghongLin/data2story-skill

    Run sentence-level traceability verification on a blog, then generate viewer.html with interactive evidence panel.

    156 GitHub stars~697 tokensUpdated 3 mo ago
    Auto-check: notes

Questions about Find Data

What does Find Data do?

Find a dataset for a Data2Story blog. An agent skill from QinghongLin/data2story-skill. Find Data is an agent skill from QinghongLin/data2story-skill. Find a dataset for a Data2Story blog.

How do I install Find Data in Claude Code?

Run `npx skills add QinghongLin/data2story-skill --skill find-data -a claude-code`. Or copy the skill folder (skills/find-data in QinghongLin/data2story-skill) into .claude/skills/find-data in your project. Claude Code loads it when a task matches its description.

How do I install Find Data in Codex?

Run `npx skills add QinghongLin/data2story-skill --skill find-data -a codex`. Or copy the skill folder (skills/find-data in QinghongLin/data2story-skill) into .agents/skills/find-data in your project. Codex loads it when a task matches its description.

Can I use Find Data in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add QinghongLin/data2story-skill --skill find-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/find-data, .gemini/skills/find-data, .github/skills/find-data and .opencode/skills/find-data in your project.

What does Find Data need to run?

Going by SKILL.md and its folder, Find Data needs Python for the scripts in its folder and the command-line tools its instructions call (python and pip). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash(python:*), Read, Write, Glob, Grep, WebSearch, WebFetch, AskUserQuestion.

Does Find Data access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Find Data safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Find Data use?

Find Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Find Data use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.8k tokens, read only when the agent opens those files.

What are the alternatives to Find Data?

Skills that share tags, products or a category with Find Data: Blog (AgriciDaniel/claude-blog, 2.3k stars), Blog (AgriciDaniel/claude-blog, 2.3k stars), Notion To Blog (wasp-lang/wasp, 19k stars) and Blog Writing Guide (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Find Data?

QinghongLin (a GitHub user) maintains it in QinghongLin/data2story-skill, which has 156 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on July 5, 2026.

Source: QinghongLin/data2story-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.