Crawl4AI Web Scraping
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
Finds authoritative public data sources for modeling tasks, prefers official APIs and bulk downloads, and outputs a reproducible fetch and cleaning plan with citations.
SKILL.md written in Chinese; this summary is our English description.
$ npx skills add yushui2022/MathModel-Skill --skill authoritative-data-harvester -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install yushui2022/MathModel-Skill authoritative-data-harvester --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/yushui2022/MathModel-Skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/trae/.trae/skills/authoritative-data-harvester .claude/skills/authoritative-data-harvester && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "authoritative-data-harvester" agent skill from https://github.com/yushui2022/MathModel-Skill/tree/standard/packages/trae/.trae/skills/authoritative-data-harvester into .claude/skills/authoritative-data-harvester/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "authoritative-data-harvester", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/yushui2022/MathModel-Skill/tree/standard/packages/trae/.trae/skills/authoritative-data-harvesterType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add yushui2022/MathModel-Skill --skill authoritative-data-harvester -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install yushui2022/MathModel-Skill authoritative-data-harvester --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yushui2022/MathModel-Skill.git skills-src && mkdir -p .agents/skills && cp -r skills-src/packages/trae/.trae/skills/authoritative-data-harvester .agents/skills/authoritative-data-harvester && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "authoritative-data-harvester" agent skill from https://github.com/yushui2022/MathModel-Skill/tree/standard/packages/trae/.trae/skills/authoritative-data-harvester into .agents/skills/authoritative-data-harvester/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "authoritative-data-harvester", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add yushui2022/MathModel-Skill --skill authoritative-data-harvester -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install yushui2022/MathModel-Skill authoritative-data-harvester --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yushui2022/MathModel-Skill.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/packages/trae/.trae/skills/authoritative-data-harvester .cursor/skills/authoritative-data-harvester && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "authoritative-data-harvester" agent skill from https://github.com/yushui2022/MathModel-Skill/tree/standard/packages/trae/.trae/skills/authoritative-data-harvester into .cursor/skills/authoritative-data-harvester/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "authoritative-data-harvester", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/yushui2022/MathModel-Skill.git --path packages/trae/.trae/skills/authoritative-data-harvester--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add yushui2022/MathModel-Skill --skill authoritative-data-harvester -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install yushui2022/MathModel-Skill authoritative-data-harvester --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yushui2022/MathModel-Skill.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/packages/trae/.trae/skills/authoritative-data-harvester .gemini/skills/authoritative-data-harvester && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "authoritative-data-harvester" agent skill from https://github.com/yushui2022/MathModel-Skill/tree/standard/packages/trae/.trae/skills/authoritative-data-harvester into .gemini/skills/authoritative-data-harvester/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "authoritative-data-harvester", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install yushui2022/MathModel-Skill authoritative-data-harvesterInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add yushui2022/MathModel-Skill --skill authoritative-data-harvester -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/yushui2022/MathModel-Skill.git skills-src && mkdir -p .github/skills && cp -r skills-src/packages/trae/.trae/skills/authoritative-data-harvester .github/skills/authoritative-data-harvester && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "authoritative-data-harvester" agent skill from https://github.com/yushui2022/MathModel-Skill/tree/standard/packages/trae/.trae/skills/authoritative-data-harvester into .github/skills/authoritative-data-harvester/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "authoritative-data-harvester", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add yushui2022/MathModel-Skill --skill authoritative-data-harvester -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install yushui2022/MathModel-Skill authoritative-data-harvester --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yushui2022/MathModel-Skill.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/packages/trae/.trae/skills/authoritative-data-harvester .opencode/skills/authoritative-data-harvester && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "authoritative-data-harvester" agent skill from https://github.com/yushui2022/MathModel-Skill/tree/standard/packages/trae/.trae/skills/authoritative-data-harvester into .opencode/skills/authoritative-data-harvester/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "authoritative-data-harvester", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
authoritative-data-harvesterFinds authoritative public data sources for modeling tasks, prefers official APIs and bulk downloads, and outputs a reproducible fetch and cleaning plan with citations.
This skill is one stage of a mathematical modeling paper workflow and is not meant to run alone. For a full paper or an unclear stage it returns to the workflow orchestrator, runs a workflow guard script before formal work and stops if the report fails, writes only its own artifacts and records progress through a context memory skill. Its input is the variable needs you give or the data needs found in earlier analysis files.
Required outputs are a reproducible description of the data sources, a fetch or download plan, and raw and processed data saved with source details under `crawled_data/`, preferably with a `sources.json`, plus a data dictionary and citation information such as source, update time, licence and access date. Official APIs and bulk downloads come first and HTML parsing last, robots.txt, terms and rate limits are respected, and logins, paywalls and CAPTCHAs are never bypassed. If automatic retrieval fails, it offers equivalent authoritative alternatives, notes definition differences and gives a manual download path.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 7712876. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Authoritative Data Harvester loads about 1.1k tokens when it runs. Until then it costs about 27 tokens; SKILL.md has 231 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from yushui2022/MathModel-Skill at commit 7712876, republished under its MIT licence (© yushui2022). 231 words, ~1,096 tokens.
.claude/skills/authoritative-data-harvester/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.paper-workflow-orchestrator 判断当前 S0-S8 阶段。python .trae/skills/paper-workflow-orchestrator/scripts/workflow_guard.py --skill authoritative-data-harvester[WORKFLOW FAIL] 或报告 status != "PASS",停止本 skill,按 paper_output/qa/workflow_guard_report.json 的失败项回补前置阶段,不得凭记忆继续。paper_output/ 产物;完成后必须回到 paper-workflow-orchestrator 判断下一步,并用 context-memory-keeper 记录已完成产物、阻塞项和下一步。python .trae/skills/paper-workflow-orchestrator/scripts/workflow_guard.py --statuspaper_output/qa/workflow_guard_report.json、paper_output/preflight_report.json、paper_output/input_manifest.json、paper_output/results/run_manifest.json 和本 skill 的上游 JSON 契约,按报告里的 recommended_skill 与 next_action 继续。paper_output/context/workflow_memory.json 视为长期断点记录;若其中的 current_step、next_step、recommended_skill 与 workflow_guard.py --status 不一致,以 guard 报告为准。paper-workflow-orchestrator 或运行 workflow_guard.py --status,再更新 workflow memory:python .trae/skills/context-memory-keeper/scripts/update_workflow_memory.pypaper_output/context/workflow_memory.json / .md,确认下一步和推荐 skill 已记录。paper_output/step1/problem_analysis.json、paper_output/plan/model_route.json 中识别出的外部数据需求。crawled_data/,优先包含 crawled_data/sources.json。data-cleaning-and-visualization 读取 crawled_data/ 做统一清洗、图表计划和论文级配图。data-cleaning-and-visualization;完整论文目标应回到 paper-workflow-orchestrator 判断后续阶段。在数学建模任务中,快速找到“权威、可引用、可复现”的公开数据源,并以尽量不爬网页、优先 API/批量下载的方式获取数据,最终输出:
输出“数据需求表”,至少包含:
按任务类型优先匹配:
交付脚本必须具备:
当某源不可用/受限:
crawled_data/raw/。crawled_data/processed/。crawled_data/sources.json(包含 url、参数、访问日期、许可证/条款)。data-cleaning-and-visualization(把 crawled_data/ 里的数据统一清洗并出图)。paper-workflow-orchestrator。context-memory-keeper,记录“新增数据源名称”与“存放路径”到 Short-term Workbench。crawled_data/。crawled_data/sources.json 中记录可引用来源信息。data-cleaning-and-visualization,把数据清洗并生成 paper_output/figures/ 的证据图表。© yushui2022, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (scripts) in packages/trae/.trae/skills/authoritative-data-harvester of yushui2022/MathModel-Skill.
Open the folder on GitHubat commit 7712876
We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in yushui2022/MathModel-Skill, which our catalogue first saw on October 7, 2026.
Authoritative Data Harvester next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Authoritative Data Harvester this skillyushui2022/MathModel-Skill | 454 | 1 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Crawl4AI Web Scrapingsmallnest/goclaw | 599 | 1 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Monitor With HaolemeHaolemeApp/Haoleme | 157 | — | ~1.3k | Automated safety check: Pass | AGPL-3.0 | |
| Credit Risk Data Cleaninggithub/awesome-copilot | 40k | 1 repos | ~1.5k | Automated safety check: Pass | MIT | |
| Data Quality Frameworkswshobson/agents | 40k | 11 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Bio Batch ProcessingGPTomics/bioSkills | 1.2k | 1 repos | ~3k | Automated safety check: Pass | MIT |
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
HaolemeApp/Haoleme
Selectively monitor important long-running or resource-intensive commands with Haoleme by prefixing them with hao, so status, output, and completion notifications sync to the mobile app.
github/awesome-copilot
Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.
wshobson/agents
Sets up data quality checks with Great Expectations, dbt tests and data contracts, with checkpoints and pass-fail reports for pipelines.
GPTomics/bioSkills
Process many sequence files in batch (count, merge, split, convert, summarize) with memory-safe streaming and on-disk indexing using Biopython, pysam, or pyfastx.
jeremylongshore/tons-of-skills-marketplace
Process automate data cleaning, transformation, and validation for ML tasks.
yushui2022/MathModel-Skill
Builds a scoring-aligned outline for a mathematical modeling paper and a model selection plan with baseline, improvement and validation experiments.
yushui2022/MathModel-Skill
Generates result-evidence contracts, tables and runnable q1 to q3 modeling code scaffolds for a math modeling paper from a model route, a data plan and cleaned data.
yushui2022/MathModel-Skill
Plans, drafts, audits, formats and verifies a formal mathematical-modeling paper from an evidence chain, delivering audited Markdown and a Word file with native equations.
yushui2022/MathModel-Skill
Repairs one failing section of a mathematical modeling paper from the repair queue, or builds a legacy or quickstart scaffold when you ask for one by name.
yushui2022/MathModel-Skill
Maintains a two-layer persistent memory for a math-modeling paper workflow: long-term rules plus a short-term workbench, with finished tasks archived.
yushui2022/MathModel-Skill
Cleans raw or scraped competition data and produces exploratory charts and a figure plan as one stage of a mathematical modeling paper workflow.
Works with
Categories
Finds authoritative public data sources for modeling tasks, prefers official APIs and bulk downloads, and outputs a reproducible fetch and cleaning plan with citations. This skill is one stage of a mathematical modeling paper workflow and is not meant to run alone. For a full paper or an unclear stage it returns to the workflow orchestrator, runs a workflow guard script before formal work and stops if the report fails, writes only its own artifacts and records progress through a context memory skill.
Authoritative Data Harvester fits situations like: finding official statistics or open government data for a modeling task; building a repeatable data download with fixed API parameters; filling gaps in a variable list with documented substitute indicators.
Run `npx skills add yushui2022/MathModel-Skill --skill authoritative-data-harvester -a claude-code`. Or copy the skill folder (packages/trae/.trae/skills/authoritative-data-harvester in yushui2022/MathModel-Skill) into .claude/skills/authoritative-data-harvester in your project. Claude Code loads it when a task matches its description.
Run `npx skills add yushui2022/MathModel-Skill --skill authoritative-data-harvester -a codex`. Or copy the skill folder (packages/trae/.trae/skills/authoritative-data-harvester in yushui2022/MathModel-Skill) into .agents/skills/authoritative-data-harvester in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yushui2022/MathModel-Skill --skill authoritative-data-harvester -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/authoritative-data-harvester, .gemini/skills/authoritative-data-harvester, .github/skills/authoritative-data-harvester and .opencode/skills/authoritative-data-harvester in your project.
Going by SKILL.md and its folder, Authoritative Data Harvester needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python to run scripts/run.py and the workflow guard script; Network access to official data sources.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Authoritative Data Harvester is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.1k tokens (SKILL.md is roughly 4.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Authoritative Data Harvester: Crawl4AI Web Scraping (smallnest/goclaw, 599 stars), Monitor With Haoleme (HaolemeApp/Haoleme, 157 stars), Credit Risk Data Cleaning (github/awesome-copilot, 40k stars) and Data Quality Frameworks (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
yushui2022 (a GitHub user) maintains it in yushui2022/MathModel-Skill, which has 454 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 7, 2026.
Source: yushui2022/MathModel-Skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.