Agent skill

Data Governance

by cbrock84 in cbrock84/headcount

Establishes ownership, definitions, quality, access, and lineage for the organization's data.

MITAuto-check passedData & Analytics

Install Data Governance

skills CLI
$ npx skills add cbrock84/headcount --skill data-governance -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cbrock84/headcount data-governance --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cbrock84/headcount.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/data-analytics/skills/data-governance .claude/skills/data-governance && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-governance
GitHub stars
2k
Token cost
~864 tokens
SKILL.md length
466 words
Files
2 (incl. references)
Skills in repo
178
Repo updated
First seen
Licence
MIT

At a glance

Establishes ownership, definitions, quality, access, and lineage for the organization's data.

  • Tasks that involve Data governance
  • SKILL.md covers Start with definitions, not…, Ownership, Quality, measured rather than… and Access, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Data cleaning

What it does

Data Governance is an agent skill from cbrock84/headcount. Establishes ownership, definitions, quality, access, and lineage for the organization's data. Use this when metrics disagree between teams, when nobody knows which dataset is authoritative, when setting up data ownership or access policy, when data quality is unreliable, or before opening a dataset to a wider audience.

Its SKILL.md is about 860 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/sources.md`).

It sits in Data & Analytics, covering Data governance and Data cleaning. The repository describes itself as: An agent organization structured as a company — 15+ departments, 125+ skills, each independently installable, citing the standards and regulators that settle the question. Runs… The licence is MIT.

When your agent uses it

  • Tasks that involve Data governance
  • Tasks that involve Data cleaning

Example prompts

  • “Use the data-governance skill to establish ownership, definitions, quality, access, and lineage for the organization's data”
  • “/data-governance”

What it can do on your machine

Read from SKILL.md and the folder at commit 98d1c17. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Governance loads about 864 tokens when it runs, and up to ~1.3k if it reads all its reference files. Until then it costs about 84 tokens; SKILL.md has 466 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~864
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from cbrock84/headcount at commit 98d1c17, republished under its MIT licence (© cbrock84). 466 words, ~864 tokens.

Download SKILL.mdSave it as .claude/skills/data-governance/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
data-governance
description
Establishes ownership, definitions, quality, access, and lineage for the organization's data. Use this when metrics disagree between teams, when nobody knows which dataset is authoritative, when setting up data ownership or access policy, when data quality is unreliable, or before opening a dataset to a wider audience.

Data governance

Governance has a reputation for bureaucracy because it is usually implemented as approval queues. Done properly it is the opposite: it makes data usable without asking anyone.

Start with definitions, not policy

The highest-value governance artifact is a metric dictionary. For each business metric:

  • The plain-language definition — what it counts, and what it deliberately excludes.
  • The computation, unambiguously: source table, filters, time grain, timezone.
  • The owner — a person who decides when it is disputed.
  • Known caveats — when it is misleading, and what changed historically.

Most metric disputes dissolve once both parties read the same definition and discover they were measuring different things. Almost none require a policy.

Watch the ones that look obvious. "Active customer," "revenue," and "signup" each have half a dozen defensible definitions, and the ambiguity surfaces at the worst moment.

Ownership

Every dataset has a named owner accountable for its quality and access — a person, not a team. Unowned datasets decay, and nobody notices until a decision is made on stale data.

The owner should sit with the business meaning, not with the pipeline. The team that generates the data understands what it means; the platform team understands how it moves.

Quality, measured rather than asserted

Test data like code, continuously, and alert on failures:

  • Freshness — did it arrive when expected?
  • Volume — is the row count within its normal range? A silent drop to zero is the classic failure.
  • Uniqueness and nullity on key fields.
  • Referential integrity across joins.
  • Distribution — has the shape shifted in a way nothing explains?

The point is finding breakage before a decision is made on it. A pipeline that fails loudly is better than one that silently produces yesterday's numbers.

Show full SKILL.md (186 more words)Show less

Access

Default to open for internal, non-personal data. Restrictive-by-default drives the shadow spreadsheet layer, which is genuinely less safe than a governed warehouse.

Personal, financial, and regulated data are the exception: least privilege, purpose stated, reviewed periodically, with Legal & Risk involved on anything with a lawful-basis question.

Lineage

Know where a number came from and what feeds it. Without lineage, you cannot answer the two questions that matter during an incident: what broke upstream, and what downstream is now wrong.

Sources

references/sources.md in this skill lists the outside authorities that settle the questions here — what each one is authoritative for, and what you may do with it. Check them before answering on anything they cover, and cite what you used. Most are free to read and not free to reproduce; the use note on each is binding.

Never

  • Let two systems each claim to be the source of truth for the same fact.
  • Fix a data-quality issue in a dashboard. Fix it upstream or it recurs in every other consumer.
  • Retire a dataset because it looks unused — you cannot see every consumer. Deprecate, announce, then remove.

© cbrock84, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in plugins/data-analytics/skills/data-governance of cbrock84/headcount.

  • SKILL.md
  • references/sources.md

Open the folder on GitHubat commit 98d1c17

Compare with similar skills

Data Governance next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Governance compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Governance this skillcbrock84/headcount2k—~864Automated safety check: PassMIT
Data Quality Frameworkswshobson/agents40k11 repos~1.1kAutomated safety check: PassMIT
Research Data Feasibility and Leakage ChecksLight0305/Light-skills640—~4.9kAutomated safety check: PassMIT
Datalineage Summarygoogle/skills21k—~1.7kAutomated safety check: PassApache-2.0
Monte Carlo Context Detectionsickn33/agentic-awesome-skills47k1 repos~2.6kAutomated safety check: WarnMIT
Querying AWS Sagemaker Catalogaws/agent-toolkit-for-aws2.8k—~2.6kAutomated safety check: PassApache-2.0

Similar skills

  • Sets up data quality checks with Great Expectations, dbt tests and data contracts, with checkpoints and pass-fail reports for pipelines.

    40k GitHub starsUsed in 11 repos~1.1k tokens
    Data & AnalyticsAuto-check passed
  • Finds usable public datasets, judges whether the data can support a research idea, and checks train and test splits for leakage before results are trusted.

    640 GitHub stars~4.9k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Datalineage Summary

    google/skills

    Official

    Summarizes Google Cloud Data Lineage graphs to help users debug data quality issues and understand data provenance for BQ/GCS.

    21k GitHub stars~1.7k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Monte Carlo Context Detection

    sickn33/agentic-awesome-skills

    Route data-related requests to the right Monte Carlo skill or workflow.

    47k GitHub starsUsed in 1 repo~2.6k tokens
    Data & AnalyticsAuto-check: warnings
  • Querying AWS Sagemaker Catalog

    aws/agent-toolkit-for-aws

    Official

    Runs SQL analytics on SageMaker Catalog asset metadata tables exported as Apache Iceberg in S3 Tables.

    2.8k GitHub stars~2.6k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Data Quality

    JoelLewis/finance_skills

    Design and operate data quality programs for financial data — validation rules, pricing validation, data lineage, exception management, profiling, and governance.

    205 GitHub stars~11k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed

More from cbrock84/headcount

All 178 skills in this repo
  • Agent Hierarchy

    cbrock84/headcount

    Designs orchestrator-and-subagent hierarchies for a repository — splitting agents by exclusive write surface, pairing every producer with an independent auditor, and enforcing the split with a…

    2k GitHub stars~1.2k tokensUpdated 21 days ago
    Auto-check passed
  • Access And Identity

    cbrock84/headcount

    Designs and audits who can reach what — authentication, authorization models, privileged access, service credentials, and joiner-mover-leaver process.

    2k GitHub stars~1.1k tokensUpdated 21 days ago
    Auto-check passed
  • Account Based Marketing

    cbrock84/headcount

    Concentrates marketing and sales effort on a named set of accounts rather than on volume — qualifying whether the model fits your economics at all, building the account list and the buying group…

    2k GitHub stars~1.2k tokensUpdated 21 days ago
    Auto-check passed
  • Activation

    cbrock84/headcount

    Gets new users from signup to first real value — signup flow, onboarding, time-to-value, and the early experience that determines whether someone becomes a user or a lapsed account.

    2k GitHub stars~865 tokensUpdated 21 days ago
    Auto-check passed
  • AI ML Governance

    cbrock84/headcount

    Governs models and AI systems in production — intended use, evaluation, monitoring, human oversight, documentation, and the decision to deploy or retire.

    2k GitHub stars~1k tokensUpdated 21 days ago
    Auto-check passed
  • AI Research Analyst

    cbrock84/headcount

    Produces executive-level research — market sizing, competitor mapping, trend analysis, and strategic intelligence — grounded in cited sources with the confidence in each claim made explicit.

    2k GitHub stars~916 tokensUpdated 21 days ago
    Auto-check passed

Questions about Data Governance

What does Data Governance do?

Establishes ownership, definitions, quality, access, and lineage for the organization's data. Data Governance is an agent skill from cbrock84/headcount. Establishes ownership, definitions, quality, access, and lineage for the organization's data.

When should I use Data Governance?

Data Governance fits situations like: tasks that involve Data governance; tasks that involve Data cleaning.

How do I install Data Governance in Claude Code?

Run `npx skills add cbrock84/headcount --skill data-governance -a claude-code`. Or copy the skill folder (plugins/data-analytics/skills/data-governance in cbrock84/headcount) into .claude/skills/data-governance in your project. Claude Code loads it when a task matches its description.

How do I install Data Governance in Codex?

Run `npx skills add cbrock84/headcount --skill data-governance -a codex`. Or copy the skill folder (plugins/data-analytics/skills/data-governance in cbrock84/headcount) into .agents/skills/data-governance in your project. Codex loads it when a task matches its description.

Can I use Data Governance in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cbrock84/headcount --skill data-governance -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-governance, .gemini/skills/data-governance, .github/skills/data-governance and .opencode/skills/data-governance in your project.

What does Data Governance need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Governance is instructions for the agent only.

Does Data Governance access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Governance safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Governance use?

Data Governance is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Governance use?

About 864 tokens (SKILL.md is roughly 3.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 426 tokens, read only when the agent opens those files.

What are the alternatives to Data Governance?

Skills that share tags, products or a category with Data Governance: Data Quality Frameworks (wshobson/agents, 40k stars), Research Data Feasibility and Leakage Checks (Light0305/Light-skills, 640 stars), Datalineage Summary (google/skills, 21k stars) and Monte Carlo Context Detection (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Governance?

cbrock84 (a GitHub user) maintains it in cbrock84/headcount, which has 2,016 GitHub stars. The repository holds 178 skills in this directory. The repository was last updated on September 17, 2026.

Source: cbrock84/headcount on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.