Official agent skill

Credit Risk Data Cleaning

by github in github/awesome-copilot

Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.

OfficialMITAuto-check passedData & Analytics

Install Credit Risk Data Cleaning

skills CLI
$ npx skills add github/awesome-copilot --skill datanalysis-credit-risk -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install github/awesome-copilot datanalysis-credit-risk --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/datanalysis-credit-risk .claude/skills/datanalysis-credit-risk && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
datanalysis-credit-risk
GitHub stars
40k
Used in
1 other repo
Token cost
~1.5k tokens
SKILL.md length
620 words
Files
4 (incl. scripts, references)
Skills in repo
417
Repo updated
First seen
Licence
MIT

At a glance

Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.

  • Works in 11 steps: Get Data - Load and format raw data → Organization Sample Analysis -… → Separate OOS Data - Separate… → …
  • Preparing raw credit data before building a pre-loan model
  • SKILL.md covers Quick Start, Complete Process Description, Core Functions and Parameter Description, plus 2 more sections
  • Runs Python scripts from its folder; calls python

What it does

The pipeline runs eleven steps in order and never deletes the original data. It loads and formats the raw records, analyzes samples by organization, separates out-of-sample records, filters months with too few samples, computes missing rates, and then removes features that fail each check.

Feature filters cover high missing rate, low information value (IV), unstable PSI, noise found by shuffling labels in a Null Importance test, and high correlation. The last step exports an Excel report with details and statistics for every stage. A quick-start script, scripts/example.py, runs the whole flow, and reusable functions such as get_dataset, missing_check and drop_lowiv_features sit in the references folder.

When your agent uses it

  • Preparing raw credit data before building a pre-loan model
  • Screening candidate variables by missing rate, IV and PSI stability
  • Removing noisy or highly correlated features before modeling
  • Producing a cleaning report that lists what was dropped at each step

Example prompts

  • “Run the credit risk cleaning pipeline on the loan applications dataset and give me the Excel report.”
  • “Which features in our pre-loan data have too many missing values or unstable PSI?”
  • “Drop low-IV and highly correlated variables before I fit the scorecard.”

Requirements

  • Python, to run the bundled scripts

Workflow steps

11 steps, taken from the first numbered list in SKILL.md.

  1. Get Data - Load and format raw data
  2. Organization Sample Analysis - Statistics of sample count and bad sample rate for each organization
  3. Separate OOS Data - Separate out-of-sample (OOS) samples from modeling samples
  4. Filter Abnormal Months - Remove months with insufficient bad sample count or total sample count
  5. Calculate Missing Rate - Calculate overall and organization-level missing rates for each feature
  6. Drop High Missing Rate Features - Remove features with overall missing rate exceeding threshold
  7. Drop Low IV Features - Remove features with overall IV too low or IV too low in too many organizations
  8. Drop High PSI Features - Remove features with unstable PSI
  9. Null Importance Denoising - Remove noise features using label permutation method
  10. Drop High Correlation Features - Remove high correlation features based on original gain
  11. Export Report - Generate Excel report containing details and statistics of all steps

What it can do on your machine

Read from SKILL.md and the folder at commit 727ff2e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Credit Risk Data Cleaning loads about 1.5k tokens when it runs, and up to ~17k if it reads all its reference files. Until then it costs about 153 tokens; SKILL.md has 620 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~153
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~17k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from github/awesome-copilot at commit 727ff2e, republished under its MIT licence (© github). 620 words, ~1,546 tokens.

Download SKILL.mdSave it as .claude/skills/datanalysis-credit-risk/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
datanalysis-credit-risk
description
Credit risk data cleaning and variable screening pipeline for pre-loan modeling. Use when working with raw credit data that needs quality assessment, missing value analysis, or variable selection before modeling. it covers data loading and formatting, abnormal period filtering, missing rate calculation, high-missing variable removal,low-IV variable filtering, high-PSI variable removal, Null Importance denoising, high-correlation variable removal, and cleaning report generation. Applicable scenarios arecredit risk data cleaning, variable screening, pre-loan modeling preprocessing.

Data Cleaning and Variable Screening

Quick Start

bash
# Run the complete data cleaning pipeline
python ".github/skills/datanalysis-credit-risk/scripts/example.py"

Complete Process Description

The data cleaning pipeline consists of the following 11 steps, each executed independently without deleting the original data:

  1. Get Data - Load and format raw data
  2. Organization Sample Analysis - Statistics of sample count and bad sample rate for each organization
  3. Separate OOS Data - Separate out-of-sample (OOS) samples from modeling samples
  4. Filter Abnormal Months - Remove months with insufficient bad sample count or total sample count
  5. Calculate Missing Rate - Calculate overall and organization-level missing rates for each feature
  6. Drop High Missing Rate Features - Remove features with overall missing rate exceeding threshold
  7. Drop Low IV Features - Remove features with overall IV too low or IV too low in too many organizations
  8. Drop High PSI Features - Remove features with unstable PSI
  9. Null Importance Denoising - Remove noise features using label permutation method
  10. Drop High Correlation Features - Remove high correlation features based on original gain
  11. Export Report - Generate Excel report containing details and statistics of all steps

Core Functions

FunctionPurposeModule
get_dataset()Load and format datareferences.func
org_analysis()Organization sample analysisreferences.func
missing_check()Calculate missing ratereferences.func
drop_abnormal_ym()Filter abnormal monthsreferences.analysis
drop_highmiss_features()Drop high missing rate featuresreferences.analysis
drop_lowiv_features()Drop low IV featuresreferences.analysis
drop_highpsi_features()Drop high PSI featuresreferences.analysis
drop_highnoise_features()Null Importance denoisingreferences.analysis
drop_highcorr_features()Drop high correlation featuresreferences.analysis
iv_distribution_by_org()IV distribution statisticsreferences.analysis
psi_distribution_by_org()PSI distribution statisticsreferences.analysis
value_ratio_distribution_by_org()Value ratio distribution statisticsreferences.analysis
export_cleaning_report()Export cleaning reportreferences.analysis

Parameter Description

Data Loading Parameters
  • DATA_PATH: Data file path (best are parquet format)
  • DATE_COL: Date column name
  • Y_COL: Label column name
  • ORG_COL: Organization column name
  • KEY_COLS: Primary key column name list
OOS Organization Configuration
  • OOS_ORGS: Out-of-sample organization list
Abnormal Month Filtering Parameters
  • min_ym_bad_sample: Minimum bad sample count per month (default 10)
  • min_ym_sample: Minimum total sample count per month (default 500)
Missing Rate Parameters
  • missing_ratio: Overall missing rate threshold (default 0.6)
IV Parameters
  • overall_iv_threshold: Overall IV threshold (default 0.1)
  • org_iv_threshold: Single organization IV threshold (default 0.1)
  • max_org_threshold: Maximum tolerated low IV organization count (default 2)
PSI Parameters
  • psi_threshold: PSI threshold (default 0.1)
  • max_months_ratio: Maximum unstable month ratio (default 1/3)
  • max_orgs: Maximum unstable organization count (default 6)
Show full SKILL.md (257 more words)Show less
Null Importance Parameters
  • n_estimators: Number of trees (default 100)
  • max_depth: Maximum tree depth (default 5)
  • gain_threshold: Gain difference threshold (default 50)
High Correlation Parameters
  • max_corr: Correlation threshold (default 0.9)
  • top_n_keep: Keep top N features by original gain ranking (default 20)

Output Report

The generated Excel report contains the following sheets:

  1. 汇总 - Summary information of all steps, including operation results and conditions
  2. 机构样本统计 - Sample count and bad sample rate for each organization
  3. 分离OOS数据 - OOS sample and modeling sample counts
  4. Step4-异常月份处理 - Abnormal months that were removed
  5. 缺失率明细 - Overall and organization-level missing rates for each feature
  6. Step5-有值率分布统计 - Distribution of features in different value ratio ranges
  7. Step6-高缺失率处理 - High missing rate features that were removed
  8. Step7-IV明细 - IV values of each feature in each organization and overall
  9. Step7-IV处理 - Features that do not meet IV conditions and low IV organizations
  10. Step7-IV分布统计 - Distribution of features in different IV ranges
  11. Step8-PSI明细 - PSI values of each feature in each organization each month
  12. Step8-PSI处理 - Features that do not meet PSI conditions and unstable organizations
  13. Step8-PSI分布统计 - Distribution of features in different PSI ranges
  14. Step9-null importance处理 - Noise features that were removed
  15. Step10-高相关性剔除 - High correlation features that were removed

Features

  • Interactive Input: Parameters can be input before each step execution, with default values supported
  • Independent Execution: Each step is executed independently without deleting original data, facilitating comparative analysis
  • Complete Report: Generate complete Excel report containing details, statistics, and distributions
  • Multi-process Support: IV and PSI calculations support multi-process acceleration
  • Organization-level Analysis: Support organization-level statistics and modeling/OOS distinction

© github, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in skills/datanalysis-credit-risk of github/awesome-copilot.

  • SKILL.md
  • references/analysis.py
  • references/func.py
  • scripts/example.py

Open the folder on GitHubat commit 727ff2e

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in github/awesome-copilot, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Credit Risk Data Cleaning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Credit Risk Data Cleaning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Credit Risk Data Cleaning this skillgithub/awesome-copilot40k1 repos~1.5kAutomated safety check: PassMIT
Authoritative Data Harvesteryushui2022/MathModel-Skill4521 repos~1.1kAutomated safety check: PassMIT
Data Quality Frameworkswshobson/agents40k10 repos~1.1kAutomated safety check: PassMIT
Ingesting Dataancoleman/ai-design-components526—~1.9kAutomated safety check: PassMIT
Bio Batch ProcessingGPTomics/bioSkills1.2k1 repos~3kAutomated safety check: PassMIT
Data Cleaningmagnus919/agent-skills111—~2.1kAutomated safety check: PassMIT

Similar skills

  • Authoritative Data Harvester

    yushui2022/MathModel-Skill

    Finds authoritative public data sources for modeling tasks, prefers official APIs and bulk downloads, and outputs a reproducible fetch and cleaning plan with citations.

    452 GitHub starsUsed in 1 repo~1.1k tokens
    Data & AnalyticsAuto-check passed
  • Sets up data quality checks with Great Expectations, dbt tests and data contracts, with checkpoints and pass-fail reports for pipelines.

    40k GitHub starsUsed in 10 repos~1.1k tokens
    Data & AnalyticsAuto-check passed
  • Ingesting Data

    ancoleman/ai-design-components

    Data ingestion patterns for loading data from cloud storage, APIs, files, and streaming sources into databases.

    526 GitHub stars~1.9k tokensUpdated 10 mo ago
    Data & AnalyticsAuto-check passed
  • Bio Batch Processing

    GPTomics/bioSkills

    Process many sequence files in batch (count, merge, split, convert, summarize) with memory-safe streaming and on-disk indexing using Biopython, pysam, or pyfastx.

    1.2k GitHub starsUsed in 1 repo~3k tokens
    Data & AnalyticsAuto-check passed
  • Data Cleaning

    magnus919/agent-skills

    Clean, profile, validate, reshape, and document messy tabular, text, JSON, and relational data through an evidence-first, reproducible workflow.

    111 GitHub stars~2.1k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • XLSX Spreadsheet Toolkit

    XiaomiMiMo/MiMo-Code

    Builds, edits, cleans, recalculates and reads Excel workbooks and CSV files with openpyxl and pandas, plus LibreOffice for recalculation and PDF export.

    14k GitHub stars~2.9k tokensUpdated 4 days ago
    Documents & OfficeAuto-check passed

More from github/awesome-copilot

All 417 skills in this repo
  • Acquire Codebase Knowledge

    github/awesome-copilot

    Official

    Maps an unfamiliar codebase into seven evidence-backed documents in docs/codebase/, using a scan script and templates, for onboarding or architecture write-ups.

    40k GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Azure Architecture Autopilot

    github/awesome-copilot

    Official

    Designs Azure infrastructure from a natural-language description, or diagrams an existing resource group, then refines the design through conversation and deploys it with Bicep.

    40k GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Draw.io Diagram Generator

    github/awesome-copilot

    Official

    Generates, edits and validates draw.io files with correct mxGraph XML, covering flowcharts, architecture, sequence, ER and UML class diagrams.

    40k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Daily Focus Board

    github/awesome-copilot

    Official

    Builds a warm, browser-based daily focus board the user updates by talking to their agent, with Eisenhower priorities, a brain-dump box and kind not-today carryover.

    40k GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Python Pypi Package Builder

    github/awesome-copilot

    Official

    End-to-end skill for building, testing, linting, versioning, and publishing a production-grade Python library to PyPI.

    40k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed
  • Official

    Analyze Terraform plan JSON output for AzureRM Provider to distinguish between false-positive diffs (order-only changes in Set-type attributes) and actual resource changes.

    40k GitHub starsUsed in 1 repo~547 tokens
    Auto-check passed

Questions about Credit Risk Data Cleaning

What does Credit Risk Data Cleaning do?

Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step. The pipeline runs eleven steps in order and never deletes the original data. It loads and formats the raw records, analyzes samples by organization, separates out-of-sample records, filters months with too few samples, computes missing rates, and then removes features that fail each check.

When should I use Credit Risk Data Cleaning?

Credit Risk Data Cleaning fits situations like: preparing raw credit data before building a pre-loan model; screening candidate variables by missing rate, IV and PSI stability; removing noisy or highly correlated features before modeling; producing a cleaning report that lists what was dropped at each step.

How do I install Credit Risk Data Cleaning in Claude Code?

Run `npx skills add github/awesome-copilot --skill datanalysis-credit-risk -a claude-code`. Or copy the skill folder (skills/datanalysis-credit-risk in github/awesome-copilot) into .claude/skills/datanalysis-credit-risk in your project. Claude Code loads it when a task matches its description.

How do I install Credit Risk Data Cleaning in Codex?

Run `npx skills add github/awesome-copilot --skill datanalysis-credit-risk -a codex`. Or copy the skill folder (skills/datanalysis-credit-risk in github/awesome-copilot) into .agents/skills/datanalysis-credit-risk in your project. Codex loads it when a task matches its description.

Can I use Credit Risk Data Cleaning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add github/awesome-copilot --skill datanalysis-credit-risk -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/datanalysis-credit-risk, .gemini/skills/datanalysis-credit-risk, .github/skills/datanalysis-credit-risk and .opencode/skills/datanalysis-credit-risk in your project.

What does Credit Risk Data Cleaning need to run?

Going by SKILL.md and its folder, Credit Risk Data Cleaning needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python, to run the bundled scripts.

Does Credit Risk Data Cleaning access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Credit Risk Data Cleaning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Credit Risk Data Cleaning use?

Credit Risk Data Cleaning is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Credit Risk Data Cleaning use?

About 1.5k tokens (SKILL.md is roughly 6.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 16k tokens, read only when the agent opens those files.

What are the alternatives to Credit Risk Data Cleaning?

Skills that share tags, products or a category with Credit Risk Data Cleaning: Authoritative Data Harvester (yushui2022/MathModel-Skill, 452 stars), Data Quality Frameworks (wshobson/agents, 40k stars), Ingesting Data (ancoleman/ai-design-components, 526 stars) and Bio Batch Processing (GPTomics/bioSkills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Credit Risk Data Cleaning?

github (a GitHub organization, an official publisher) maintains it in github/awesome-copilot, which has 39,748 GitHub stars. The repository holds 417 skills in this directory. The repository was last updated on October 7, 2026.

Source: github/awesome-copilot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.