Agent skill

Authoritative Data Harvester

by yushui2022 in yushui2022/MathModel-Skill

Finds authoritative public data sources for modeling tasks, prefers official APIs and bulk downloads, and outputs a reproducible fetch and cleaning plan with citations.

MITAuto-check passedData & Analytics

SKILL.md written in Chinese; this summary is our English description.

Install Authoritative Data Harvester

skills CLI
$ npx skills add yushui2022/MathModel-Skill --skill authoritative-data-harvester -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install yushui2022/MathModel-Skill authoritative-data-harvester --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/yushui2022/MathModel-Skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/trae/.trae/skills/authoritative-data-harvester .claude/skills/authoritative-data-harvester && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
authoritative-data-harvester
GitHub stars
454
Used in
1 other repo
Token cost
~1.1k tokens
SKILL.md length
231 words
Files
2 (incl. scripts)
Skills in repo
10
Repo updated
First seen
Licence
MIT

At a glance

Finds authoritative public data sources for modeling tasks, prefers official APIs and bulk downloads, and outputs a reproducible fetch and cleaning plan with citations.

  • Works in 6 steps: 需求归一化 → 选源策略(先权威后便利) → 访问方式决策树 → …
  • Finding official statistics or open government data for a modeling task
  • SKILL.md covers 全局流程协作约束(长对话防漂移), 执行契约, 目标 and 何时调用, plus 8 more sections
  • Runs Python scripts from its folder; calls python

What it does

This skill is one stage of a mathematical modeling paper workflow and is not meant to run alone. For a full paper or an unclear stage it returns to the workflow orchestrator, runs a workflow guard script before formal work and stops if the report fails, writes only its own artifacts and records progress through a context memory skill. Its input is the variable needs you give or the data needs found in earlier analysis files.

Required outputs are a reproducible description of the data sources, a fetch or download plan, and raw and processed data saved with source details under `crawled_data/`, preferably with a `sources.json`, plus a data dictionary and citation information such as source, update time, licence and access date. Official APIs and bulk downloads come first and HTML parsing last, robots.txt, terms and rate limits are respected, and logins, paywalls and CAPTCHAs are never bypassed. If automatic retrieval fails, it offers equivalent authoritative alternatives, notes definition differences and gives a manual download path.

When your agent uses it

  • Finding official statistics or open government data for a modeling task
  • Building a repeatable data download with fixed API parameters
  • Filling gaps in a variable list with documented substitute indicators

Example prompts

  • “Find authoritative sources for national GDP and population data and write a reproducible download script.”
  • “Fetch the World Bank indicators for my variable list and save sources.json with citations.”
  • “We lack data for one variable; propose an official alternative indicator and explain the definition difference.”

Requirements

  • Python to run scripts/run.py and the workflow guard script
  • Network access to official data sources

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. 需求归一化
  2. 选源策略(先权威后便利)
  3. 访问方式决策树
  4. 抓取实现规范
  5. 清洗与校验(最低要求)
  6. 交付物清单(必须输出)

What it can do on your machine

Read from SKILL.md and the folder at commit 7712876. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Authoritative Data Harvester loads about 1.1k tokens when it runs. Until then it costs about 27 tokens; SKILL.md has 231 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~27
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from yushui2022/MathModel-Skill at commit 7712876, republished under its MIT licence (© yushui2022). 231 words, ~1,096 tokens.

Download SKILL.mdSave it as .claude/skills/authoritative-data-harvester/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
authoritative-data-harvester
description
自动定位并获取权威公开数据(优先API/官方批量下载),输出可复现抓取与清洗方案。Invoke when用户需要权威数据、官方统计、API下载或数据源爬取。

权威数据自动获取(Authoritative Data Harvester)

全局流程协作约束(长对话防漂移)

  • 本 skill 不得作为孤立入口。用户要求完整论文、生成 Word、继续流程或不确定阶段时,先回到 paper-workflow-orchestrator 判断当前 S0-S8 阶段。
  • 启动或继续本 skill 的正式任务前,必须运行:
    bash
    python .trae/skills/paper-workflow-orchestrator/scripts/workflow_guard.py --skill authoritative-data-harvester
  • 如果输出 [WORKFLOW FAIL] 或报告 status != "PASS",停止本 skill,按 paper_output/qa/workflow_guard_report.json 的失败项回补前置阶段,不得凭记忆继续。
  • 本 skill 只写入自己契约范围内的 paper_output/ 产物;完成后必须回到 paper-workflow-orchestrator 判断下一步,并用 context-memory-keeper 记录已完成产物、阻塞项和下一步。
  • 长对话中如果上下文变长、阶段不确定或用户分开调用 skill,先运行:
    bash
    python .trae/skills/paper-workflow-orchestrator/scripts/workflow_guard.py --status
    再读取 paper_output/qa/workflow_guard_report.json、paper_output/preflight_report.json、paper_output/input_manifest.json、paper_output/results/run_manifest.json 和本 skill 的上游 JSON 契约,按报告里的 recommended_skill 与 next_action 继续。
  • 继续流程前,必须把 paper_output/context/workflow_memory.json 视为长期断点记录;若其中的 current_step、next_step、recommended_skill 与 workflow_guard.py --status 不一致,以 guard 报告为准。
  • 每次完成本 skill 的产物后,先回到 paper-workflow-orchestrator 或运行 workflow_guard.py --status,再更新 workflow memory:
    bash
    python .trae/skills/context-memory-keeper/scripts/update_workflow_memory.py
    更新后读取 paper_output/context/workflow_memory.json / .md,确认下一步和推荐 skill 已记录。

执行契约

  • 上游输入:用户给出的变量需求,或 paper_output/step1/problem_analysis.json、paper_output/plan/model_route.json 中识别出的外部数据需求。
  • 必须输出:可复现的数据源说明、抓取或下载方案,并将原始/处理后数据与来源信息保存到 crawled_data/,优先包含 crawled_data/sources.json。
  • 下游交接:data-cleaning-and-visualization 读取 crawled_data/ 做统一清洗、图表计划和论文级配图。
  • 推荐下一步:数据落盘后进入 data-cleaning-and-visualization;完整论文目标应回到 paper-workflow-orchestrator 判断后续阶段。
  • 失败回退:若无法自动获取,应给出同级权威替代源、口径差异和人工下载路径;不得使用无来源或不可引用的数据冒充权威数据。

目标

在数学建模任务中,快速找到“权威、可引用、可复现”的公开数据源,并以尽量不爬网页、优先 API/批量下载的方式获取数据,最终输出:

  • 数据获取脚本/方案(含链接、参数、时间范围、字段解释)
  • 原始数据与清洗后的数据(CSV/Parquet)
  • 数据字典与引用信息(来源、更新时间、许可证/条款、访问日期)

何时调用

  • 需要权威/官方数据源(统计局、国际组织、政府开放数据)
  • 需要可复现的数据抓取流程(接口参数固定、可重复运行)
  • 已有变量清单但缺数据,或需要补充替代指标

总原则(必须遵守)

  1. 优先使用官方 API 或 Bulk Download,最后才做 HTML 解析。
  2. 遵守 robots.txt 与服务条款,尊重速率限制;不得绕过登录/付费/验证码。
  3. 全流程可复现:记录来源 URL、接口参数、访问日期、版本/更新时间、字段含义与单位。
  4. 数据质量优先:对齐口径、单位、频率、地理范围;明确缺失与异常处理策略。

标准工作流(每次执行都按此输出)

1) 需求归一化

输出“数据需求表”,至少包含:

  • 变量名(中英)、单位、期望频率(日/月/年)、时间范围
  • 地区粒度(国家/省/市/网格)与口径说明
  • 允许替代指标(主指标不可得时)
2) 选源策略(先权威后便利)

按任务类型优先匹配:

  • 宏观/发展:World Bank、IMF、OECD、UNData
  • 人口/社会:UN DESA、World Bank、各国统计局/人口普查
  • 公共卫生:WHO(必要时补充二次聚合源并标注来源链)
  • 气象/气候:NOAA、NASA(按开放政策选择)
  • 欧盟统计:Eurostat
  • 中国:国家统计局、部委开放数据、地方统计局(优先可下载/接口)
3) 访问方式决策树
  • 有官方 API:直接 API
  • 无 API 但有批量下载(CSV/Excel/ZIP):直链下载
  • 仅网页表格:优先找页面背后的 XHR/JSON;仍不行再做 HTML 解析
4) 抓取实现规范

交付脚本必须具备:

  • 参数化:start/end、region、indicator、output_dir
  • 稳健性:重试(指数退避)、超时、速率限制、缓存/断点
  • 落盘:raw/ 与 processed/ 分目录;保留原始响应或原始文件
  • 日志:只记录必要信息,不输出敏感信息
5) 清洗与校验(最低要求)
  • 字段:统一命名(snake_case)、类型转换、单位换算、时间索引对齐
  • 缺失:说明缺失原因(不可得/断档/口径变化),给出处理策略
  • 异常:基本规则校验(范围、同比/环比跳变阈值)
  • 抽检:与来源页面/元数据对照样本行
6) 交付物清单(必须输出)
  • 数据文件:raw.*、processed.csv(或 parquet)
  • 元数据:sources.json(name/url/access_method/params/license/updated_at/accessed_at)
  • 数据字典:字段含义、单位、频率、地区粒度、缺失策略
  • 引用格式:可直接用于论文/报告的参考条目

常用权威数据源(可扩展)

  • World Bank Data API
  • IMF Data
  • UNData / UN agencies
  • OECD Data
  • Eurostat
  • WHO
  • NOAA / NASA
  • 各国统计局与政府开放数据平台

用户输入模板(用于快速启动)

  • “我要做【主题】建模,变量有【A,B,C】;时间【YYYY-YYYY】;地区【国家/省/市】;请给权威来源与可复现的数据获取脚本+清洗结果。”
  • “我需要【某指标】的官方数据,优先 API,没有就批量下载;请输出 sources.json + processed.csv。”

失败回退策略

当某源不可用/受限:

  1. 换同级权威源(例如 UN ↔ World Bank ↔ OECD)
  2. 换替代指标并明确口径差异
  3. 仅交付最权威的可下载版本,并说明无法自动化获取原因

目录约定(与项目全局对齐)

  • 原始下载与接口响应建议保存到:crawled_data/raw/。
  • 清洗后的结构化数据建议保存到:crawled_data/processed/。
  • 来源与可复现信息建议保存为:crawled_data/sources.json(包含 url、参数、访问日期、许可证/条款)。

前后衔接

  • 后续通常先做:data-cleaning-and-visualization(把 crawled_data/ 里的数据统一清洗并出图)。
  • 若要继续到论文草稿:回到 paper-workflow-orchestrator。

约束(必须遵守)

  • Memory Interaction (必做):
    • 获取数据后:必须调用 context-memory-keeper,记录“新增数据源名称”与“存放路径”到 Short-term Workbench。
  • 本技能只负责“把权威数据拿到手并保证可复现”,不负责直接写论文正文;产物必须落盘到 crawled_data/。
  • 若数据要进入论文,必须同时满足两点:
    • crawled_data/sources.json 中记录可引用来源信息。
    • 后续调用 data-cleaning-and-visualization,把数据清洗并生成 paper_output/figures/ 的证据图表。

© yushui2022, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in packages/trae/.trae/skills/authoritative-data-harvester of yushui2022/MathModel-Skill.

  • SKILL.md
  • scripts/run.py

Open the folder on GitHubat commit 7712876

Used in 1 other repository

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in yushui2022/MathModel-Skill, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Authoritative Data Harvester next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Authoritative Data Harvester compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Authoritative Data Harvester this skillyushui2022/MathModel-Skill4541 repos~1.1kAutomated safety check: PassMIT
Crawl4AI Web Scrapingsmallnest/goclaw5991 repos~2.5kAutomated safety check: PassMIT
Monitor With HaolemeHaolemeApp/Haoleme157—~1.3kAutomated safety check: PassAGPL-3.0
Credit Risk Data Cleaninggithub/awesome-copilot40k1 repos~1.5kAutomated safety check: PassMIT
Data Quality Frameworkswshobson/agents40k11 repos~1.1kAutomated safety check: PassMIT
Bio Batch ProcessingGPTomics/bioSkills1.2k1 repos~3kAutomated safety check: PassMIT

Similar skills

  • Crawl4AI Web Scraping

    smallnest/goclaw

    Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.

    599 GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed
  • Monitor With Haoleme

    HaolemeApp/Haoleme

    Selectively monitor important long-running or resource-intensive commands with Haoleme by prefixing them with hao, so status, output, and completion notifications sync to the mobile app.

    157 GitHub stars~1.3k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Credit Risk Data Cleaning

    github/awesome-copilot

    Official

    Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.

    40k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Sets up data quality checks with Great Expectations, dbt tests and data contracts, with checkpoints and pass-fail reports for pipelines.

    40k GitHub starsUsed in 11 repos~1.1k tokens
    Data & AnalyticsAuto-check passed
  • Bio Batch Processing

    GPTomics/bioSkills

    Process many sequence files in batch (count, merge, split, convert, summarize) with memory-safe streaming and on-disk indexing using Biopython, pysam, or pyfastx.

    1.2k GitHub starsUsed in 1 repo~3k tokens
    Data & AnalyticsAuto-check passed
  • Preprocessing Data With Automated Pipelines

    jeremylongshore/tons-of-skills-marketplace

    Process automate data cleaning, transformation, and validation for ML tasks.

    2.8k GitHub stars~1k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from yushui2022/MathModel-Skill

All 10 skills in this repo
  • Builds a scoring-aligned outline for a mathematical modeling paper and a model selection plan with baseline, improvement and validation experiments.

    454 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Modeling Code and Result Contracts

    yushui2022/MathModel-Skill

    Generates result-evidence contracts, tables and runnable q1 to q3 modeling code scaffolds for a math modeling paper from a model route, a data plan and cleaned data.

    454 GitHub stars~1.4k tokensUpdated 4 days ago
    Auto-check passed
  • Formal Modeling Paper Writer

    yushui2022/MathModel-Skill

    Plans, drafts, audits, formats and verifies a formal mathematical-modeling paper from an evidence chain, delivering audited Markdown and a Word file with native equations.

    454 GitHub stars~1.6k tokensUpdated 4 days ago
    Auto-check passed
  • Paper Micro-Unit Generator

    yushui2022/MathModel-Skill

    Repairs one failing section of a mathematical modeling paper from the repair queue, or builds a legacy or quickstart scaffold when you ask for one by name.

    454 GitHub stars~1.1k tokensUpdated 4 days ago
    Auto-check passed
  • Context Memory Keeper

    yushui2022/MathModel-Skill

    Maintains a two-layer persistent memory for a math-modeling paper workflow: long-term rules plus a short-term workbench, with finished tasks archived.

    454 GitHub starsUsed in 1 repo~893 tokens
    Auto-check passed
  • Math Modeling Data Cleaning and Charts

    yushui2022/MathModel-Skill

    Cleans raw or scraped competition data and produces exploratory charts and a figure plan as one stage of a mathematical modeling paper workflow.

    454 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed

Works with

Questions about Authoritative Data Harvester

What does Authoritative Data Harvester do?

Finds authoritative public data sources for modeling tasks, prefers official APIs and bulk downloads, and outputs a reproducible fetch and cleaning plan with citations. This skill is one stage of a mathematical modeling paper workflow and is not meant to run alone. For a full paper or an unclear stage it returns to the workflow orchestrator, runs a workflow guard script before formal work and stops if the report fails, writes only its own artifacts and records progress through a context memory skill.

When should I use Authoritative Data Harvester?

Authoritative Data Harvester fits situations like: finding official statistics or open government data for a modeling task; building a repeatable data download with fixed API parameters; filling gaps in a variable list with documented substitute indicators.

How do I install Authoritative Data Harvester in Claude Code?

Run `npx skills add yushui2022/MathModel-Skill --skill authoritative-data-harvester -a claude-code`. Or copy the skill folder (packages/trae/.trae/skills/authoritative-data-harvester in yushui2022/MathModel-Skill) into .claude/skills/authoritative-data-harvester in your project. Claude Code loads it when a task matches its description.

How do I install Authoritative Data Harvester in Codex?

Run `npx skills add yushui2022/MathModel-Skill --skill authoritative-data-harvester -a codex`. Or copy the skill folder (packages/trae/.trae/skills/authoritative-data-harvester in yushui2022/MathModel-Skill) into .agents/skills/authoritative-data-harvester in your project. Codex loads it when a task matches its description.

Can I use Authoritative Data Harvester in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yushui2022/MathModel-Skill --skill authoritative-data-harvester -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/authoritative-data-harvester, .gemini/skills/authoritative-data-harvester, .github/skills/authoritative-data-harvester and .opencode/skills/authoritative-data-harvester in your project.

What does Authoritative Data Harvester need to run?

Going by SKILL.md and its folder, Authoritative Data Harvester needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python to run scripts/run.py and the workflow guard script; Network access to official data sources.

Does Authoritative Data Harvester access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Authoritative Data Harvester safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Authoritative Data Harvester use?

Authoritative Data Harvester is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Authoritative Data Harvester use?

About 1.1k tokens (SKILL.md is roughly 4.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Authoritative Data Harvester?

Skills that share tags, products or a category with Authoritative Data Harvester: Crawl4AI Web Scraping (smallnest/goclaw, 599 stars), Monitor With Haoleme (HaolemeApp/Haoleme, 157 stars), Credit Risk Data Cleaning (github/awesome-copilot, 40k stars) and Data Quality Frameworks (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Authoritative Data Harvester?

yushui2022 (a GitHub user) maintains it in yushui2022/MathModel-Skill, which has 454 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 7, 2026.

Source: yushui2022/MathModel-Skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.