Credit Risk Data Cleaning
github/awesome-copilot
Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.
Clean, profile, validate, reshape, and document messy tabular, text, JSON, and relational data through an evidence-first, reproducible workflow.
$ npx skills add magnus919/agent-skills --skill data-cleaning -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install magnus919/agent-skills data-cleaning --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/data-cleaning .claude/skills/data-cleaning && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "data-cleaning" agent skill from https://github.com/magnus919/agent-skills/tree/main/data-cleaning into .claude/skills/data-cleaning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/magnus919/agent-skills/tree/main/data-cleaningType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add magnus919/agent-skills --skill data-cleaning -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install magnus919/agent-skills data-cleaning --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/data-cleaning .agents/skills/data-cleaning && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "data-cleaning" agent skill from https://github.com/magnus919/agent-skills/tree/main/data-cleaning into .agents/skills/data-cleaning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add magnus919/agent-skills --skill data-cleaning -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install magnus919/agent-skills data-cleaning --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/data-cleaning .cursor/skills/data-cleaning && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "data-cleaning" agent skill from https://github.com/magnus919/agent-skills/tree/main/data-cleaning into .cursor/skills/data-cleaning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/magnus919/agent-skills.git --path data-cleaning--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add magnus919/agent-skills --skill data-cleaning -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install magnus919/agent-skills data-cleaning --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/data-cleaning .gemini/skills/data-cleaning && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "data-cleaning" agent skill from https://github.com/magnus919/agent-skills/tree/main/data-cleaning into .gemini/skills/data-cleaning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install magnus919/agent-skills data-cleaningInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add magnus919/agent-skills --skill data-cleaning -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/data-cleaning .github/skills/data-cleaning && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "data-cleaning" agent skill from https://github.com/magnus919/agent-skills/tree/main/data-cleaning into .github/skills/data-cleaning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add magnus919/agent-skills --skill data-cleaning -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install magnus919/agent-skills data-cleaning --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/data-cleaning .opencode/skills/data-cleaning && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "data-cleaning" agent skill from https://github.com/magnus919/agent-skills/tree/main/data-cleaning into .opencode/skills/data-cleaning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
data-cleaningClean, profile, validate, reshape, and document messy tabular, text, JSON, and relational data through an evidence-first, reproducible workflow.
Data Cleaning is an agent skill from magnus919/agent-skills. Clean, profile, validate, reshape, and document messy tabular, text, JSON, and relational data through an evidence-first, reproducible workflow. Use when preparing data for analysis, reporting, modeling, ingestion, migration, or matching, including AI-suggested repair, entity-match review, score calibration limits, and reversible repair ledgers. Do not use for statistical modeling, dashboard design, or operating a named data platform; route those tasks to data-scientist, data-engineering, or the relevant tool…
Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 26 other files, including scripts and reference files (for example `README.md`, `evals/evals.json` and `evals/trigger-cases.json`). Compatibility notes: Works with any Agent Skills client. The bundled profiler requires Python 3.9+ and the standard library; ecosystem tools are optional.
It sits in Data & Analytics, covering Data cleaning, Data pipelines and ETL and Performance reviews. It works with Python. The repository describes itself as: Curated collection of AI agent skills for Hermes and other agent frameworks. The licence is MIT.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit c545c2b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 4 files in scripts/ (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Works with any Agent Skills client. The bundled profiler requires Python 3.9+ and the standard library; ecosystem tools are optional.
From compatibility in the SKILL.md frontmatter.
Data Cleaning loads about 2.1k tokens when it runs, and up to ~6k if it reads all its reference files. Until then it costs about 134 tokens; SKILL.md has 888 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from magnus919/agent-skills at commit c545c2b, republished under its MIT licence (© magnus919). 888 words, ~2,066 tokens.
.claude/skills/data-cleaning/SKILL.md (or your agent's skills folder). This skill also uses 22 other files; get the full folder from GitHub.Treat cleaning as a controlled transformation of an observed dataset, not cosmetic editing. Preserve raw input, state the target use and grain, make every lossy decision explicit, and prove that the cleaned output satisfies a contract.
When a model proposes data transformations, load the AI boundary workflow and use the companion record.
| Need | Read next |
|---|---|
| End-to-end method, scope, and stopping rules | references/methodology.md |
| Choose a library or platform | references/tool-selection.md |
| Missingness, duplicates, types, ranges, categories, dates, joins | references/operations.md |
| Text, identifiers, Unicode, and entity resolution | references/text-and-entity.md |
| Schemas, contracts, validation, drift, scale | references/validation-and-scale.md |
| CLI, OpenRefine, monitoring, and interactive remediation | references/cli-and-interactive-tools.md |
| Source claims and version-sensitive caveats | references/sources.md |
| Plan, logs, exceptions, contracts, or reports | templates/cleaning-plan.md, templates/transformation-log.jsonl, templates/exception-register.csv, templates/schema-contract.yml, templates/quality-report.md |
| Lightweight profile or reconciliation | Run python3 scripts/profile_dataset.py --help or python3 scripts/reconcile_dataset.py --help |
| Script | Purpose | Invocation |
|---|---|---|
scripts/profile_dataset.py | Dependency-free first-pass profiling of a CSV, TSV, or JSONL input without modifying it: missingness, cardinality, type candidates, duplicates, ranges, and value anomalies. Run it at workflow step 3 (Profile before changing) as the evidence-gathering pass before designing any cleaning decision. | python3 scripts/profile_dataset.py data.csv --output profile.json |
scripts/reconcile_dataset.py | Reconciliation between a before and after delimited dataset: row counts, key uniqueness/overlap, and per-column sums (--sum), keyed by --key, writing a machine-readable report. Run it during Validate twice / Review to prove grain preservation and quantify exactly what a transformation changed. | python3 scripts/reconcile_dataset.py raw.csv cleaned.csv --key id --sum amount --output reconciliation.json |
scripts/test_profile_dataset.py | Pytest suite covering the profiler's behavior on representative inputs. Run it after modifying the profiler or when auditing its output; CI discovers it automatically. | python3 -m pytest scripts/test_profile_dataset.py |
scripts/test_reconcile_dataset.py | Pytest suite covering the reconciler's keying, summing, and reporting behavior. Run it after modifying the reconciler or when auditing its output; CI discovers it automatically. | python3 -m pytest scripts/test_reconcile_dataset.py |
A cleaning task is complete only when the output, transformation/decision log, validation evidence, provenance, and unresolved issues exist; raw data remains intact; acceptance checks pass; and a reviewer can reproduce or audit the result. If semantic ambiguity remains, stop at quarantine or escalation rather than inventing a value.
Do not use this skill for inferential statistics or model selection, which belong to data-scientist; for ETL orchestration, storage, or production data-quality operations, route to data-engineering; or for operating a named validation or database platform, route to that tool's skill. This skill supplies cleaning judgment and artifacts those workflows consume.
compatibility); ecosystem tools (OpenRefine, pandas-backed tooling) are optional accelerators covered in references/cli-and-interactive-tools.md.pytest only when running the bundled test suites.© magnus919, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 22 other files (scripts, references) in data-cleaning of magnus919/agent-skills.
Open the folder on GitHubat commit c545c2b
Data Cleaning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Data Cleaning this skillmagnus919/agent-skills | 116 | — | ~2.1k | Automated safety check: Pass | MIT | |
| Credit Risk Data Cleaninggithub/awesome-copilot | 40k | 1 repos | ~1.5k | Automated safety check: Pass | MIT | |
| Authoritative Data Harvesteryushui2022/MathModel-Skill | 453 | 1 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Data Quality Frameworkswshobson/agents | 40k | 11 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Bio Batch ProcessingGPTomics/bioSkills | 1.2k | 1 repos | ~3k | Automated safety check: Pass | MIT | |
| Preprocessing Data With Automated Pipelinesjeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~1k | Automated safety check: Pass | MIT |
github/awesome-copilot
Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.
yushui2022/MathModel-Skill
Finds authoritative public data sources for modeling tasks, prefers official APIs and bulk downloads, and outputs a reproducible fetch and cleaning plan with citations.
wshobson/agents
Sets up data quality checks with Great Expectations, dbt tests and data contracts, with checkpoints and pass-fail reports for pipelines.
GPTomics/bioSkills
Process many sequence files in batch (count, merge, split, convert, summarize) with memory-safe streaming and on-disk indexing using Biopython, pysam, or pyfastx.
jeremylongshore/tons-of-skills-marketplace
Process automate data cleaning, transformation, and validation for ML tasks.
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
magnus919/agent-skills
Organize durable agent research outputs as summaries, analysis, and evidence dossiers.
magnus919/agent-skills
Build portable, first-person colored ASCII city engines and small GIS-derived city packs.
magnus919/agent-skills
Manage color workflows with ICC profiles, working spaces, gamut mapping, and color science.
magnus919/agent-skills
A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…
magnus919/agent-skills
Use Docker Compose to define, run, debug, and harden multi-container applications.
magnus919/agent-skills
Design, review, simulate, and verify FPGA logic using explicit RTL contracts, clock and reset models, CDC analysis, timing constraints, and reproducible implementation evidence.
Works with
Categories
Clean, profile, validate, reshape, and document messy tabular, text, JSON, and relational data through an evidence-first, reproducible workflow. Data Cleaning is an agent skill from magnus919/agent-skills. Clean, profile, validate, reshape, and document messy tabular, text, JSON, and relational data through an evidence-first, reproducible workflow.
Data Cleaning fits situations like: preparing data for analysis; including AI-suggested repair; entity-match review; score calibration limits.
Run `npx skills add magnus919/agent-skills --skill data-cleaning -a claude-code`. Or copy the skill folder (data-cleaning in magnus919/agent-skills) into .claude/skills/data-cleaning in your project. Claude Code loads it when a task matches its description.
Run `npx skills add magnus919/agent-skills --skill data-cleaning -a codex`. Or copy the skill folder (data-cleaning in magnus919/agent-skills) into .agents/skills/data-cleaning in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add magnus919/agent-skills --skill data-cleaning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-cleaning, .gemini/skills/data-cleaning, .github/skills/data-cleaning and .opencode/skills/data-cleaning in your project.
Going by SKILL.md and its folder, Data Cleaning needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3. Compatibility (from SKILL.md): Works with any Agent Skills client. The bundled profiler requires Python 3.9+ and the standard library; ecosystem tools are optional..
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Data Cleaning is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.1k tokens (SKILL.md is roughly 8.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Data Cleaning: Credit Risk Data Cleaning (github/awesome-copilot, 40k stars), Authoritative Data Harvester (yushui2022/MathModel-Skill, 453 stars), Data Quality Frameworks (wshobson/agents, 40k stars) and Bio Batch Processing (GPTomics/bioSkills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
magnus919 (a GitHub user) maintains it in magnus919/agent-skills, which has 116 GitHub stars. The repository holds 130 skills in this directory. The repository was last updated on October 8, 2026.
Source: magnus919/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.