Authoritative Data Harvester
yushui2022/MathModel-Skill
Finds authoritative public data sources for modeling tasks, prefers official APIs and bulk downloads, and outputs a reproducible fetch and cleaning plan with citations.
Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.
$ npx skills add github/awesome-copilot --skill datanalysis-credit-risk -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install github/awesome-copilot datanalysis-credit-risk --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/datanalysis-credit-risk .claude/skills/datanalysis-credit-risk && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "datanalysis-credit-risk" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/datanalysis-credit-risk into .claude/skills/datanalysis-credit-risk/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datanalysis-credit-risk", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/github/awesome-copilot/tree/main/skills/datanalysis-credit-riskType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add github/awesome-copilot --skill datanalysis-credit-risk -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install github/awesome-copilot datanalysis-credit-risk --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/datanalysis-credit-risk .agents/skills/datanalysis-credit-risk && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "datanalysis-credit-risk" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/datanalysis-credit-risk into .agents/skills/datanalysis-credit-risk/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datanalysis-credit-risk", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add github/awesome-copilot --skill datanalysis-credit-risk -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install github/awesome-copilot datanalysis-credit-risk --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/datanalysis-credit-risk .cursor/skills/datanalysis-credit-risk && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "datanalysis-credit-risk" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/datanalysis-credit-risk into .cursor/skills/datanalysis-credit-risk/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datanalysis-credit-risk", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/github/awesome-copilot.git --path skills/datanalysis-credit-risk--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add github/awesome-copilot --skill datanalysis-credit-risk -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install github/awesome-copilot datanalysis-credit-risk --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/datanalysis-credit-risk .gemini/skills/datanalysis-credit-risk && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "datanalysis-credit-risk" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/datanalysis-credit-risk into .gemini/skills/datanalysis-credit-risk/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datanalysis-credit-risk", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install github/awesome-copilot datanalysis-credit-riskInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add github/awesome-copilot --skill datanalysis-credit-risk -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/datanalysis-credit-risk .github/skills/datanalysis-credit-risk && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "datanalysis-credit-risk" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/datanalysis-credit-risk into .github/skills/datanalysis-credit-risk/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datanalysis-credit-risk", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add github/awesome-copilot --skill datanalysis-credit-risk -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install github/awesome-copilot datanalysis-credit-risk --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/datanalysis-credit-risk .opencode/skills/datanalysis-credit-risk && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "datanalysis-credit-risk" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/datanalysis-credit-risk into .opencode/skills/datanalysis-credit-risk/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datanalysis-credit-risk", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
datanalysis-credit-riskCleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.
The pipeline runs eleven steps in order and never deletes the original data. It loads and formats the raw records, analyzes samples by organization, separates out-of-sample records, filters months with too few samples, computes missing rates, and then removes features that fail each check.
Feature filters cover high missing rate, low information value (IV), unstable PSI, noise found by shuffling labels in a Null Importance test, and high correlation. The last step exports an Excel report with details and statistics for every stage. A quick-start script, scripts/example.py, runs the whole flow, and reusable functions such as get_dataset, missing_check and drop_lowiv_features sit in the references folder.
11 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 727ff2e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Credit Risk Data Cleaning loads about 1.5k tokens when it runs, and up to ~17k if it reads all its reference files. Until then it costs about 153 tokens; SKILL.md has 620 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from github/awesome-copilot at commit 727ff2e, republished under its MIT licence (© github). 620 words, ~1,546 tokens.
.claude/skills/datanalysis-credit-risk/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.# Run the complete data cleaning pipeline
python ".github/skills/datanalysis-credit-risk/scripts/example.py"The data cleaning pipeline consists of the following 11 steps, each executed independently without deleting the original data:
| Function | Purpose | Module |
|---|---|---|
get_dataset() | Load and format data | references.func |
org_analysis() | Organization sample analysis | references.func |
missing_check() | Calculate missing rate | references.func |
drop_abnormal_ym() | Filter abnormal months | references.analysis |
drop_highmiss_features() | Drop high missing rate features | references.analysis |
drop_lowiv_features() | Drop low IV features | references.analysis |
drop_highpsi_features() | Drop high PSI features | references.analysis |
drop_highnoise_features() | Null Importance denoising | references.analysis |
drop_highcorr_features() | Drop high correlation features | references.analysis |
iv_distribution_by_org() | IV distribution statistics | references.analysis |
psi_distribution_by_org() | PSI distribution statistics | references.analysis |
value_ratio_distribution_by_org() | Value ratio distribution statistics | references.analysis |
export_cleaning_report() | Export cleaning report | references.analysis |
DATA_PATH: Data file path (best are parquet format)DATE_COL: Date column nameY_COL: Label column nameORG_COL: Organization column nameKEY_COLS: Primary key column name listOOS_ORGS: Out-of-sample organization listmin_ym_bad_sample: Minimum bad sample count per month (default 10)min_ym_sample: Minimum total sample count per month (default 500)missing_ratio: Overall missing rate threshold (default 0.6)overall_iv_threshold: Overall IV threshold (default 0.1)org_iv_threshold: Single organization IV threshold (default 0.1)max_org_threshold: Maximum tolerated low IV organization count (default 2)psi_threshold: PSI threshold (default 0.1)max_months_ratio: Maximum unstable month ratio (default 1/3)max_orgs: Maximum unstable organization count (default 6)n_estimators: Number of trees (default 100)max_depth: Maximum tree depth (default 5)gain_threshold: Gain difference threshold (default 50)max_corr: Correlation threshold (default 0.9)top_n_keep: Keep top N features by original gain ranking (default 20)The generated Excel report contains the following sheets:
© github, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (scripts, references) in skills/datanalysis-credit-risk of github/awesome-copilot.
Open the folder on GitHubat commit 727ff2e
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in github/awesome-copilot, which our catalogue first saw on October 7, 2026.
Credit Risk Data Cleaning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Credit Risk Data Cleaning this skillgithub/awesome-copilot | 40k | 1 repos | ~1.5k | Automated safety check: Pass | MIT | |
| Authoritative Data Harvesteryushui2022/MathModel-Skill | 452 | 1 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Data Quality Frameworkswshobson/agents | 40k | 10 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Ingesting Dataancoleman/ai-design-components | 526 | — | ~1.9k | Automated safety check: Pass | MIT | |
| Bio Batch ProcessingGPTomics/bioSkills | 1.2k | 1 repos | ~3k | Automated safety check: Pass | MIT | |
| Data Cleaningmagnus919/agent-skills | 111 | — | ~2.1k | Automated safety check: Pass | MIT |
yushui2022/MathModel-Skill
Finds authoritative public data sources for modeling tasks, prefers official APIs and bulk downloads, and outputs a reproducible fetch and cleaning plan with citations.
wshobson/agents
Sets up data quality checks with Great Expectations, dbt tests and data contracts, with checkpoints and pass-fail reports for pipelines.
ancoleman/ai-design-components
Data ingestion patterns for loading data from cloud storage, APIs, files, and streaming sources into databases.
GPTomics/bioSkills
Process many sequence files in batch (count, merge, split, convert, summarize) with memory-safe streaming and on-disk indexing using Biopython, pysam, or pyfastx.
magnus919/agent-skills
Clean, profile, validate, reshape, and document messy tabular, text, JSON, and relational data through an evidence-first, reproducible workflow.
XiaomiMiMo/MiMo-Code
Builds, edits, cleans, recalculates and reads Excel workbooks and CSV files with openpyxl and pandas, plus LibreOffice for recalculation and PDF export.
github/awesome-copilot
Maps an unfamiliar codebase into seven evidence-backed documents in docs/codebase/, using a scan script and templates, for onboarding or architecture write-ups.
github/awesome-copilot
Designs Azure infrastructure from a natural-language description, or diagrams an existing resource group, then refines the design through conversation and deploys it with Bicep.
github/awesome-copilot
Generates, edits and validates draw.io files with correct mxGraph XML, covering flowcharts, architecture, sequence, ER and UML class diagrams.
github/awesome-copilot
Builds a warm, browser-based daily focus board the user updates by talking to their agent, with Eisenhower priorities, a brain-dump box and kind not-today carryover.
github/awesome-copilot
End-to-end skill for building, testing, linting, versioning, and publishing a production-grade Python library to PyPI.
github/awesome-copilot
Analyze Terraform plan JSON output for AzureRM Provider to distinguish between false-positive diffs (order-only changes in Set-type attributes) and actual resource changes.
Works with
Categories
Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step. The pipeline runs eleven steps in order and never deletes the original data. It loads and formats the raw records, analyzes samples by organization, separates out-of-sample records, filters months with too few samples, computes missing rates, and then removes features that fail each check.
Credit Risk Data Cleaning fits situations like: preparing raw credit data before building a pre-loan model; screening candidate variables by missing rate, IV and PSI stability; removing noisy or highly correlated features before modeling; producing a cleaning report that lists what was dropped at each step.
Run `npx skills add github/awesome-copilot --skill datanalysis-credit-risk -a claude-code`. Or copy the skill folder (skills/datanalysis-credit-risk in github/awesome-copilot) into .claude/skills/datanalysis-credit-risk in your project. Claude Code loads it when a task matches its description.
Run `npx skills add github/awesome-copilot --skill datanalysis-credit-risk -a codex`. Or copy the skill folder (skills/datanalysis-credit-risk in github/awesome-copilot) into .agents/skills/datanalysis-credit-risk in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add github/awesome-copilot --skill datanalysis-credit-risk -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/datanalysis-credit-risk, .gemini/skills/datanalysis-credit-risk, .github/skills/datanalysis-credit-risk and .opencode/skills/datanalysis-credit-risk in your project.
Going by SKILL.md and its folder, Credit Risk Data Cleaning needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python, to run the bundled scripts.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Credit Risk Data Cleaning is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.5k tokens (SKILL.md is roughly 6.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 16k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Credit Risk Data Cleaning: Authoritative Data Harvester (yushui2022/MathModel-Skill, 452 stars), Data Quality Frameworks (wshobson/agents, 40k stars), Ingesting Data (ancoleman/ai-design-components, 526 stars) and Bio Batch Processing (GPTomics/bioSkills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
github (a GitHub organization, an official publisher) maintains it in github/awesome-copilot, which has 39,748 GitHub stars. The repository holds 417 skills in this directory. The repository was last updated on October 7, 2026.
Source: github/awesome-copilot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.