Agent skill

Chief Data Officer Advisor

by alirezarezvani in alirezarezvani/claude-skills

Chief Data Officer advisory for startups: AI training data rights and consent provenance, data product strategy (warehouse vs lakehouse vs mesh, build-vs-buy), B2B customer-data-as-asset valuation…

MITAuto-check passedBusiness, Finance & HR

Install Chief Data Officer Advisor

skills CLI
$ npx skills add alirezarezvani/claude-skills --skill chief-data-officer-advisor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install alirezarezvani/claude-skills chief-data-officer-advisor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/alirezarezvani/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/c-level-advisor/skills/chief-data-officer-advisor .claude/skills/chief-data-officer-advisor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
chief-data-officer-advisor
GitHub stars
28k
Token cost
~2.9k tokens
SKILL.md length
1,090 words
Files
8 (incl. scripts, references)
Skills in repo
342
Repo updated
First seen
Licence
MIT

At a glance

Chief Data Officer advisory for startups: AI training data rights and consent provenance, data product strategy (warehouse vs lakehouse vs mesh, build-vs-buy), B2B customer-data-as-asset valuation…

  • Works in 4 steps: AI Training Data Rights → Data Product Strategy → B2B Customer-Data-as-Asset → …
  • Deciding whether to train models on customer data
  • SKILL.md covers Keywords, Quick Start, Key Questions (ask these first) and Core Responsibilities, plus 4 more sections
  • Runs Python scripts from its folder; calls python

What it does

Chief Data Officer Advisor is an agent skill from alirezarezvani/claude-skills. Chief Data Officer advisory for startups: AI training data rights and consent provenance, data product strategy (warehouse vs lakehouse vs mesh, build-vs-buy), B2B customer-data-as-asset valuation and M&A readiness, data team org evolution. Use when deciding whether to train models on customer data, choosing data architecture, valuing data for fundraising or M&A, sequencing data hires, or when user mentions CDO, chief data officer, data strategy, data mesh, lakehouse, training data, data product, data…

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `references/ai_training_data_rights.md`, `references/customer_data_as_asset.md` and `references/data_product_strategy.md`).

It sits in Business, Finance & HR, covering Data warehousing, Data pipelines and ETL and Product strategy. The repository describes itself as: 380 Claude Code skills & agent skills & plugins (30+ Agents, 70+ custom commands, 380+ skills, customizable references, scripts)for Claude Code, Codex, Gemini CLI, Cursor, and 8… The licence is MIT.

When your agent uses it

  • Deciding whether to train models on customer data
  • Choosing data architecture
  • Valuing data for fundraising
  • Sequencing data hires

Example prompts

  • “/chief-data-officer-advisor”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. AI Training Data Rights
  2. Data Product Strategy
  3. B2B Customer-Data-as-Asset
  4. Data Team Org Evolution

What it can do on your machine

Read from SKILL.md and the folder at commit 19392f7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Chief Data Officer Advisor loads about 2.9k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 159 tokens; SKILL.md has 1,090 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~159
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~12k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from alirezarezvani/claude-skills at commit 19392f7, republished under its MIT licence (© alirezarezvani). 1,090 words, ~2,893 tokens.

Download SKILL.mdSave it as .claude/skills/chief-data-officer-advisor/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
chief-data-officer-advisor
description
Chief Data Officer advisory for startups: AI training data rights and consent provenance, data product strategy (warehouse vs lakehouse vs mesh, build-vs-buy), B2B customer-data-as-asset valuation and M&A readiness, data team org evolution. Use when deciding whether to train models on customer data, choosing data architecture, valuing data for fundraising or M&A, sequencing data hires, or when user mentions CDO, chief data officer, data strategy, data mesh, lakehouse, training data, data product, data monetization, or customer data asset. NOT a tactical data engineering skill — strategic decisions only.
license
MIT
metadata.version
1.0.0
metadata.author
Alireza Rezvani
metadata.category
c-level
metadata.domain
chief-data-officer-leadership
metadata.updated
2026-05-12
metadata.python-tools
ai_training_data_audit.py, data_product_strategy_picker.py, data_asset_valuator.py
metadata.frameworks
training-data-rights-matrix, data-product-strategy, customer-data-as-asset, data-team-org-evolution

Chief Data Officer Advisor

Strategic data leadership for startup CDOs and founders without one. Four decisions, no surveys:

  1. Can we train our model on this data? — origin × consent × use-case matrix
  2. Warehouse, lakehouse, or mesh — and what do we build vs buy? — stage-driven architecture
  3. What is our customer data worth? — strategic value + M&A multiplier + productization paths
  4. What data role do we hire next? — stage-to-role map, centralize-vs-embed trigger

This skill does not cover tactical data engineering. For schema design, observability, query optimization, RAG, or ML platform implementation, see engineering/database-designer/, engineering/observability-designer/, engineering/data-quality-auditor/, engineering/sql-database-assistant/, engineering/rag-architect/, engineering/llm-cost-optimizer/.

Keywords

CDO, chief data officer, AI training data, consent provenance, training rights, GDPR Article 6 lawful basis, GDPR Article 22, EU AI Act high-risk, ePrivacy, copyright fair use, hiQ v. LinkedIn, scraped data, synthetic data, data product, data mesh, lakehouse, medallion architecture, dbt, Snowflake, BigQuery, Databricks, Fivetran, Airbyte, reverse ETL, feature store, customer data as asset, data monetization, data productization, anonymization, k-anonymity, differential privacy, M&A data diligence, data org, analytics engineer, data engineer, data scientist, data product manager, centralize vs embed, hub and spoke

Quick Start

bash
# Audit data sources for AI training eligibility
python scripts/ai_training_data_audit.py                              # uses embedded sample
python scripts/ai_training_data_audit.py path/to/sources.json

# Pick data architecture + build-vs-buy + sequencing
python scripts/data_product_strategy_picker.py                        # uses embedded Series A SaaS
python scripts/data_product_strategy_picker.py path/to/profile.json

# Value the customer data corpus + productization viability
python scripts/data_asset_valuator.py                                 # uses embedded B2B sample
python scripts/data_asset_valuator.py path/to/corpus.json

Key Questions (ask these first)

  • What decision does this data drive? (If none, why are we collecting it?)
  • What's the consent provenance of every source we want to train on? (TOS-only is not the same as explicit opt-in.)
  • Who are the internal data consumers, and how many distinct domains do they span? (Drives centralize-vs-embed and warehouse-vs-mesh.)
  • In an M&A scenario, is our data a moat or a liability? (Customer carve-outs in MSAs can flip the answer.)
  • Are we hiring an analytics engineer or a data scientist next? (They solve different problems; founders confuse them.)
  • Have we run an anonymization audit before any external sharing? (k-anonymity ≥ 5 is the floor, not the ceiling.)

Core Responsibilities

1. AI Training Data Rights

The 2026 question every startup is facing: can we use customer data to train our model?

The answer is rarely binary. It depends on three independent dimensions:

DimensionValues
Origin1st-party-explicit-opt-in / 1st-party-TOS-only / partner-licensed / scraped / synthetic
Data classAnonymous aggregate / behavioral / PII / 3rd-party content / regulated (PHI, PCI, kids)
Use caseIn-product personalization / fine-tune our model / train foundation model / external sharing

Each combination produces GO / MITIGATE / NO-GO. Run ai_training_data_audit.py on a JSON inventory of sources.

See references/ai_training_data_rights.md for the full matrix + GDPR Art. 6 lawful basis decision tree + EU AI Act high-risk triggers.

2. Data Product Strategy

Architecture choice (warehouse vs lakehouse vs mesh) is stage-driven, not preference-driven:

  • Warehouse only (Snowflake / BigQuery / Postgres): ≤5 data consumers, <2TB, no ML use cases
  • Lakehouse (warehouse + object storage, often Databricks or Snowflake-with-Iceberg): 5–25 data consumers, 2TB–1PB, 1–3 ML use cases
  • Data mesh: 25+ data consumers across 4+ domains, federated ownership culture in place

Build vs buy is decided per layer:

LayerBuy unlessBuild only if
Storage / warehouseNever build(You’re a data infra company)
ELT / ingestNever buildSource isn’t supported by Fivetran/Airbyte
Modeling (dbt)Always buildThis is your IP
BI / dashboardsBuy at <100 consumersEmbedded analytics for customers
Feature storeDefer until 3+ prod modelsThen build OR buy Tecton/Hopsworks
ML platformDefer until 5+ prod modelsThen buy SageMaker/Vertex/Databricks

Run data_product_strategy_picker.py for a stage-specific recommendation. See references/data_product_strategy.md for kill criteria per architecture and the build-vs-buy decision tree.

3. B2B Customer-Data-as-Asset

The shift: at Series B+, customer data is no longer just operational — it’s an asset that can be:

  • A defensibility moat (replicating requires years of customer cohort)
  • An M&A multiplier (1.2x–2x ARR uplift for strategic buyers)
  • A direct revenue stream (anonymized industry benchmarks, embedding endpoints, licensing)

But it can also be a liability:

  • 47/380 customers with MSA carve-outs makes productization legally infeasible
  • Anonymization audits often reveal re-identification risk above tolerable thresholds
  • Regulatory exposure increases linearly with productization (GDPR Art. 28 processors vs Art. 26 joint controllers)

Run data_asset_valuator.py with corpus characteristics to get strategic value score + productization paths + risk-adjusted value.

See references/customer_data_as_asset.md for the valuation framework, M&A diligence prep checklist, and contractual constraint audit pattern.

Show full SKILL.md (434 more words)Show less
4. Data Team Org Evolution

The wrong question: "Should we hire a data scientist?" The right question: "What’s the next decision we can’t make because we lack data, and what role unblocks that?"

Stage-to-role map (B2B SaaS baseline):

StageFirst hireThenThen
Pre-seed / seedFounder-as-analyst (SQL + spreadsheets)——
Series A (Series A)AnalystAnalytics engineer (dbt)—
Series BData engineerSenior analyst (embedded in GTM)Data PM (if 3+ teams need data)
GrowthManager of analyticsML engineer (if model is core)Head of Data
Late-stageHead of Data → CDOSpecialized: BI, MLE, DPOFederated owners per domain (mesh)

Centralize-vs-embed trigger: when 3+ functional areas (sales, marketing, product, ops, CS) need bespoke data weekly, the central team becomes the bottleneck. Move to hub-and-spoke (central platform + embedded analysts) before that becomes a hiring crisis.

See references/data_team_org_evolution.md.

Workflows

Workflow 1: AI Training Decision (1 hour)

Goal: Decide whether a specific data source can train a specific use case.

bash
# 1. Build sources.json with one entry per data source
# 2. Run the audit
python scripts/ai_training_data_audit.py sources.json
# 3. For each MITIGATE: assign owner + remediation
# 4. For each NO-GO: document the kill reason for the legal log
# 5. Cross-check with cs-general-counsel-advisor on top-3 mitigation items
# 6. Log via /cs:decide
Workflow 2: Architecture Decision (1 day)

Goal: Pick warehouse / lakehouse / mesh and the build-vs-buy split for the next 12 months.

bash
python scripts/data_product_strategy_picker.py profile.json
# Cross-check with cs-cto-advisor on engineering capacity
# Cross-check with cs-cfo-advisor on 3-year TCO
# Log via /cs:decide; consider /cs:freeze 90 if signing a multi-year SaaS contract
Workflow 3: Data Asset Valuation for M&A Prep (3 days)

Goal: Value the data corpus and prepare for due diligence.

  1. Inventory the corpus: size, freshness, exclusivity, customer overlap, contractual restrictions
  2. Run data_asset_valuator.py
  3. Run the M&A diligence prep checklist in customer_data_as_asset.md
  4. Surface contractual carve-outs to cs-general-counsel-advisor for re-papering plan
  5. Decide productization path (benchmark report / embedding endpoint / direct license)
  6. Log via /cs:decide
Workflow 4: Data Team Roadmap (1 week)

Goal: Build the next 18 months of data hires aligned to business decisions.

  1. List the top 5 decisions the business can’t make today due to missing data or analysis
  2. Map each decision to the role that unblocks it
  3. Sequence hires (one role at a time, ramp before next)
  4. Cross-check with cs-chro-advisor on comp bands and leveling
  5. Identify the centralize-vs-embed trigger date

Output Standards (when invoked via cs-cdo-advisor)

**Bottom Line:** [one sentence — decision and rationale]
**The Decision:** [one of the 4 framings]
**The Evidence:** [numbers, not adjectives]
**How to Act:** [3 concrete next steps]
**Your Decision:** [the call only the founder can make]

Adjacent Skills

  • c-level-advisor/skills/cto-advisor/ — architecture capacity, scaling cliffs
  • c-level-advisor/skills/ciso-advisor/ — data security, threat modeling for productized data
  • c-level-advisor/skills/general-counsel-advisor/ — contractual constraints, DPA, training-data rights
  • c-level-advisor/skills/cfo-advisor/ — build-vs-buy TCO, M&A valuation math
  • c-level-advisor/skills/chro-advisor/ — data team hiring, leveling, comp
  • engineering/skills/database-designer/ — tactical schema design
  • engineering/skills/rag-architect/ — tactical AI/RAG implementation
  • engineering/llm-cost-optimizer/ — model cost management

References


Version: 1.0.0 Status: Production Ready Disclaimer: Decisions touching training data rights, data productization, or M&A data diligence should involve qualified counsel. This skill surfaces decisions and tradeoffs — it does not replace legal review.

© alirezarezvani, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references) in c-level-advisor/skills/chief-data-officer-advisor of alirezarezvani/claude-skills.

  • SKILL.md
  • references/ai_training_data_rights.md
  • references/customer_data_as_asset.md
  • references/data_product_strategy.md
  • references/data_team_org_evolution.md
  • scripts/ai_training_data_audit.py
  • scripts/data_asset_valuator.py
  • scripts/data_product_strategy_picker.py

Open the folder on GitHubat commit 19392f7

Compare with similar skills

Chief Data Officer Advisor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Chief Data Officer Advisor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Chief Data Officer Advisor this skillalirezarezvani/claude-skills28k—~2.9kAutomated safety check: PassMIT
Modeling Warehouse FoundationsPostHog/posthog40k—~2.1kAutomated safety check: PassCustom licence
C Level AdvisorLeoYeAI/openclaw-master-skills2.2k—~1.6kAutomated safety check: PassMIT
The Spiralegregore-labs/egregore291—~527Automated safety check: PassMIT
Sec 10k AnalysisOctagonAI/octagon-mcp-server147—~318Automated safety check: PassMIT
Data Engineeringrohitg00/awesome-claude-code-toolkit2.7k—~1.7kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Shared foundations for building reusable data models in PostHog, on either of two stacks: PostHog-native data-warehouse views / materialized views (HogQL, via the view- MCP tools), or an external…

    40k GitHub stars~2.1k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • C Level Advisor

    LeoYeAI/openclaw-master-skills

    Provides strategic business advice by channelling the perspectives of 10 executive roles — CEO, CTO, COO, CPO, CMO, CFO, CRO, CISO, CHRO, and Executive Mentor — across decisions, trade-offs, and org…

    2.2k GitHub stars~1.6k tokensUpdated 2 mo ago
    Business, Finance & HRAuto-check passed
  • The Spiral

    egregore-labs/egregore

    A generative epistemology engine that transforms intuitions into rigorous, communicable output through structured Socratic dialogue.

    291 GitHub stars~527 tokensUpdated 9 days ago
    Business, Finance & HRAuto-check passed
  • Sec 10k Analysis

    OctagonAI/octagon-mcp-server

    Analyze 10-K annual filings to extract business model, financial priorities, risk factors, segments, and notable changes.

    147 GitHub stars~318 tokensUpdated 3 mo ago
    Business, Finance & HRAuto-check passed
  • Data Engineering

    rohitg00/awesome-claude-code-toolkit

    Data engineering patterns for ETL pipelines, data warehousing, Apache Spark, and data quality validation

    2.7k GitHub stars~1.7k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check passed
  • Company Research

    claude-office-skills/skills

    Conduct comprehensive company research and due diligence. An agent skill from claude-office-skills/skills.

    505 GitHub stars~2.2k tokensUpdated 8 mo ago
    Sales & SupportAuto-check passed

More from alirezarezvani/claude-skills

All 342 skills in this repo
  • Agile Product Owner

    alirezarezvani/claude-skills

    Writes INVEST-checked user stories with acceptance criteria, splits epics, plans sprints from velocity and ranks the backlog with a weighted score.

    28k GitHub starsUsed in 3 repos~3.2k tokens
    Auto-check passed
  • Product Strategist

    alirezarezvani/claude-skills

    OKR cascade toolkit for product leaders: generates aligned company-to-team OKRs from five strategy types and scores how well they line up.

    28k GitHub starsUsed in 2 repos~1.8k tokens
    Auto-check passed
  • App Store Optimization

    alirezarezvani/claude-skills

    App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store.

    28k GitHub starsUsed in 1 repo~4.2k tokens
    Auto-check passed
  • AWS Solution Architect

    alirezarezvani/claude-skills

    Design AWS architectures for startups using serverless patterns and IaC templates.

    28k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Campaign Analytics

    alirezarezvani/claude-skills

    Calculates attribution, funnel and ROI figures for marketing campaigns with three Python scripts that need only the standard library.

    28k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Code to PRD

    alirezarezvani/claude-skills

    Reverse-engineers a frontend, backend or fullstack codebase into a product requirements document with per-page docs, an enum dictionary and an API inventory.

    28k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed

Questions about Chief Data Officer Advisor

What does Chief Data Officer Advisor do?

Chief Data Officer advisory for startups: AI training data rights and consent provenance, data product strategy (warehouse vs lakehouse vs mesh, build-vs-buy), B2B customer-data-as-asset valuation…. Chief Data Officer Advisor is an agent skill from alirezarezvani/claude-skills. Chief Data Officer advisory for startups: AI training data rights and consent provenance, data product strategy (warehouse vs lakehouse vs mesh, build-vs-buy), B2B customer-data-as-asset valuation and M&A readiness, data team org evolution.

When should I use Chief Data Officer Advisor?

Chief Data Officer Advisor fits situations like: deciding whether to train models on customer data; choosing data architecture; valuing data for fundraising; sequencing data hires.

How do I install Chief Data Officer Advisor in Claude Code?

Run `npx skills add alirezarezvani/claude-skills --skill chief-data-officer-advisor -a claude-code`. Or copy the skill folder (c-level-advisor/skills/chief-data-officer-advisor in alirezarezvani/claude-skills) into .claude/skills/chief-data-officer-advisor in your project. Claude Code loads it when a task matches its description.

How do I install Chief Data Officer Advisor in Codex?

Run `npx skills add alirezarezvani/claude-skills --skill chief-data-officer-advisor -a codex`. Or copy the skill folder (c-level-advisor/skills/chief-data-officer-advisor in alirezarezvani/claude-skills) into .agents/skills/chief-data-officer-advisor in your project. Codex loads it when a task matches its description.

Can I use Chief Data Officer Advisor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alirezarezvani/claude-skills --skill chief-data-officer-advisor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/chief-data-officer-advisor, .gemini/skills/chief-data-officer-advisor, .github/skills/chief-data-officer-advisor and .opencode/skills/chief-data-officer-advisor in your project.

What does Chief Data Officer Advisor need to run?

Going by SKILL.md and its folder, Chief Data Officer Advisor needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Chief Data Officer Advisor access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Chief Data Officer Advisor safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Chief Data Officer Advisor use?

Chief Data Officer Advisor is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Chief Data Officer Advisor use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.2k tokens, read only when the agent opens those files.

What are the alternatives to Chief Data Officer Advisor?

Skills that share tags, products or a category with Chief Data Officer Advisor: Modeling Warehouse Foundations (PostHog/posthog, 40k stars), C Level Advisor (LeoYeAI/openclaw-master-skills, 2.2k stars), The Spiral (egregore-labs/egregore, 291 stars) and Sec 10k Analysis (OctagonAI/octagon-mcp-server, 147 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Chief Data Officer Advisor?

alirezarezvani (a GitHub user) maintains it in alirezarezvani/claude-skills, which has 27,938 GitHub stars. The repository holds 342 skills in this directory. The repository was last updated on August 30, 2026.

Source: alirezarezvani/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.