Credit Risk Data Cleaning
github/awesome-copilot
Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.
Sets up data quality checks with Great Expectations, dbt tests and data contracts, with checkpoints and pass-fail reports for pipelines.
$ npx skills add wshobson/agents --skill data-quality-frameworks -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install wshobson/agents data-quality-frameworks --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/data-engineering/skills/data-quality-frameworks .claude/skills/data-quality-frameworks && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "data-quality-frameworks" agent skill from https://github.com/wshobson/agents/tree/main/plugins/data-engineering/skills/data-quality-frameworks into .claude/skills/data-quality-frameworks/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-quality-frameworks", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/wshobson/agents/tree/main/plugins/data-engineering/skills/data-quality-frameworksType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add wshobson/agents --skill data-quality-frameworks -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install wshobson/agents data-quality-frameworks --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/data-engineering/skills/data-quality-frameworks .agents/skills/data-quality-frameworks && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "data-quality-frameworks" agent skill from https://github.com/wshobson/agents/tree/main/plugins/data-engineering/skills/data-quality-frameworks into .agents/skills/data-quality-frameworks/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-quality-frameworks", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill data-quality-frameworks -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install wshobson/agents data-quality-frameworks --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/data-engineering/skills/data-quality-frameworks .cursor/skills/data-quality-frameworks && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "data-quality-frameworks" agent skill from https://github.com/wshobson/agents/tree/main/plugins/data-engineering/skills/data-quality-frameworks into .cursor/skills/data-quality-frameworks/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-quality-frameworks", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/wshobson/agents.git --path plugins/data-engineering/skills/data-quality-frameworks--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add wshobson/agents --skill data-quality-frameworks -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install wshobson/agents data-quality-frameworks --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/data-engineering/skills/data-quality-frameworks .gemini/skills/data-quality-frameworks && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "data-quality-frameworks" agent skill from https://github.com/wshobson/agents/tree/main/plugins/data-engineering/skills/data-quality-frameworks into .gemini/skills/data-quality-frameworks/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-quality-frameworks", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install wshobson/agents data-quality-frameworksInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add wshobson/agents --skill data-quality-frameworks -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/data-engineering/skills/data-quality-frameworks .github/skills/data-quality-frameworks && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "data-quality-frameworks" agent skill from https://github.com/wshobson/agents/tree/main/plugins/data-engineering/skills/data-quality-frameworks into .github/skills/data-quality-frameworks/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-quality-frameworks", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill data-quality-frameworks -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install wshobson/agents data-quality-frameworks --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/data-engineering/skills/data-quality-frameworks .opencode/skills/data-quality-frameworks && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "data-quality-frameworks" agent skill from https://github.com/wshobson/agents/tree/main/plugins/data-engineering/skills/data-quality-frameworks into .opencode/skills/data-quality-frameworks/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-quality-frameworks", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
data-quality-frameworksSets up data quality checks with Great Expectations, dbt tests and data contracts, with checkpoints and pass-fail reports for pipelines.
This skill covers three ways to check data: Great Expectations validation, dbt tests and data contracts between teams. It frames quality through six dimensions, completeness, uniqueness, validity, accuracy, consistency and timeliness, and maps each to an example check such as a not-null or unique-values expectation.
It also describes a testing pyramid for data, with integration tests across tables at the top, and gives a quick start that installs Great Expectations and sets up a daily validation checkpoint in Python. A report-building example summarizes how many tables passed and lists the failed checks per table. Further patterns and worked examples sit in a details reference file.
2 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Data Quality Frameworks loads about 1.1k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 55 tokens; SKILL.md has 205 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 205 words, ~1,112 tokens.
.claude/skills/data-quality-frameworks/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Production patterns for implementing data quality with Great Expectations, dbt tests, and data contracts to ensure reliable data pipelines.
| Dimension | Description | Example Check |
|---|---|---|
| Completeness | No missing values | expect_column_values_to_not_be_null |
| Uniqueness | No duplicates | expect_column_values_to_be_unique |
| Validity | Values in expected range | expect_column_values_to_be_in_set |
| Accuracy | Data matches reality | Cross-reference validation |
| Consistency | No contradictions | expect_column_pair_values_A_to_be_greater_than_B |
| Timeliness | Data is recent | expect_column_max_to_be_between |
/\
/ \ Integration Tests (cross-table)
/────\
/ \ Unit Tests (single column)
/────────\
/ \ Schema Tests (structure)
/────────────\# Install
pip install great_expectations
# Initialize project
great_expectations init
# Create datasource
great_expectations datasource new# great_expectations/checkpoints/daily_validation.yml
import great_expectations as gx
# Create context
context = gx.get_context()
# Create expectation suite
suite = context.add_expectation_suite("orders_suite")
# Add expectations
suite.add_expectation(
gx.expectations.ExpectColumnValuesToNotBeNull(column="order_id")
)
suite.add_expectation(
gx.expectations.ExpectColumnValuesToBeUnique(column="order_id")
)
# Validate
results = context.run_checkpoint(checkpoint_name="daily_orders")Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.
report.append("")
for table, result in results.items():
status = "✅" if result.passed else "❌"
report.append(f"### {status} {table}")
report.append(f"- Expectations: {result.total_expectations}")
report.append(f"- Failed: {result.failed_expectations}")
if not result.passed:
report.append("- Failed checks:")
for detail in result.details:
if not detail["success"]:
report.append(f" - {detail['expectation']}: {detail['observed_value']}")
report.append("")
return "\n".join(report)context = gx.get_context() pipeline = DataQualityPipeline(context)
tables_to_validate = { "orders": "orders_suite", "customers": "customers_suite", "products": "products_suite", }
results = pipeline.run_all(tables_to_validate) report = pipeline.generate_report(results)
if not all(r.passed for r in results.values()): print(report) raise ValueError("Data quality checks failed!")
## Best Practices
### Do's
- **Test early** - Validate source data before transformations
- **Test incrementally** - Add tests as you find issues
- **Document expectations** - Clear descriptions for each test
- **Alert on failures** - Integrate with monitoring
- **Version contracts** - Track schema changes
### Don'ts
- **Don't test everything** - Focus on critical columns
- **Don't ignore warnings** - They often precede failures
- **Don't skip freshness** - Stale data is bad data
- **Don't hardcode thresholds** - Use dynamic baselines
- **Don't test in isolation** - Test relationships too© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in plugins/data-engineering/skills/data-quality-frameworks of wshobson/agents.
Open the folder on GitHubat commit 46891e7
We found 28 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 10 other GitHub owners. This page covers the copy in wshobson/agents, which our catalogue first saw on October 7, 2026.
Data Quality Frameworks next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Data Quality Frameworks this skillwshobson/agents | 40k | 10 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Credit Risk Data Cleaninggithub/awesome-copilot | 40k | 1 repos | ~1.5k | Automated safety check: Pass | MIT | |
| Dbt Parser Refreshyu-iskw/dbt-artifacts-parser | 118 | — | ~716 | Automated safety check: Pass | Apache-2.0 | |
| Suggesting Dbt Bouncer Checksgodatadriven/dbt-bouncer | 135 | — | ~1k | Automated safety check: Pass | MIT | |
| Authoritative Data Harvesteryushui2022/MathModel-Skill | 452 | 1 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Package Version Bumpyu-iskw/dbt-artifacts-parser | 118 | — | ~700 | Automated safety check: Pass | Apache-2.0 |
github/awesome-copilot
Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.
yu-iskw/dbt-artifacts-parser
Refreshes dbt artifact schemas from dbt-labs/dbt-core and regenerates Pydantic parser classes.
godatadriven/dbt-bouncer
Analyzes a dbt project and suggests dbt-bouncer checks that already pass (for existing projects) or a sensible starter config (for greenfield projects).
yushui2022/MathModel-Skill
Finds authoritative public data sources for modeling tasks, prefers official APIs and bulk downloads, and outputs a reproducible fetch and cleaning plan with citations.
yu-iskw/dbt-artifacts-parser
Bumps the dbt-artifacts-parser package semver, prepares a PyPI release, and aligns GitHub Releases with version.
benchflow-ai/skillsbench
World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.
wshobson/agents
Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.
wshobson/agents
Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.
wshobson/agents
Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.
wshobson/agents
Covers portfolio risk measurement with VaR, CVaR, Sharpe, Sortino and drawdown, plus guidance on limits, stress tests and tail risk.
wshobson/agents
Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.
wshobson/agents
Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.
Categories
Sets up data quality checks with Great Expectations, dbt tests and data contracts, with checkpoints and pass-fail reports for pipelines. This skill covers three ways to check data: Great Expectations validation, dbt tests and data contracts between teams. It frames quality through six dimensions, completeness, uniqueness, validity, accuracy, consistency and timeliness, and maps each to an example check such as a not-null or unique-values expectation.
Data Quality Frameworks fits situations like: adding data quality checks to a pipeline; setting up Great Expectations validation and checkpoints; building out a dbt test suite for models; agreeing data contracts between producing and consuming teams.
Run `npx skills add wshobson/agents --skill data-quality-frameworks -a claude-code`. Or copy the skill folder (plugins/data-engineering/skills/data-quality-frameworks in wshobson/agents) into .claude/skills/data-quality-frameworks in your project. Claude Code loads it when a task matches its description.
Run `npx skills add wshobson/agents --skill data-quality-frameworks -a codex`. Or copy the skill folder (plugins/data-engineering/skills/data-quality-frameworks in wshobson/agents) into .agents/skills/data-quality-frameworks in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill data-quality-frameworks -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-quality-frameworks, .gemini/skills/data-quality-frameworks, .github/skills/data-quality-frameworks and .opencode/skills/data-quality-frameworks in your project.
Going by SKILL.md and its folder, Data Quality Frameworks needs the command-line tools its instructions call (pip). Our summary lists: Python with great_expectations installed through pip.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Data Quality Frameworks is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.1k tokens (SKILL.md is roughly 4.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Data Quality Frameworks: Credit Risk Data Cleaning (github/awesome-copilot, 40k stars), Dbt Parser Refresh (yu-iskw/dbt-artifacts-parser, 118 stars), Suggesting Dbt Bouncer Checks (godatadriven/dbt-bouncer, 135 stars) and Authoritative Data Harvester (yushui2022/MathModel-Skill, 452 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,254 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.
Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.