Agent skill

Data Investigation

by Kilo-Org in Kilo-Org/kilo-marketplace

A workflow for conducting rigorous, reproducible data investigations and ad-hoc analyses.

Apache-2.0Auto-check passedDevelopment

Install Data Investigation

skills CLI
$ npx skills add Kilo-Org/kilo-marketplace --skill data-investigation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Kilo-Org/kilo-marketplace data-investigation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Kilo-Org/kilo-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/data-investigation .claude/skills/data-investigation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-investigation
GitHub stars
190
Token cost
~1.5k tokens
SKILL.md length
660 words
Files
1
Skills in repo
85
Repo updated
First seen
Licence
Apache-2.0

At a glance

A workflow for conducting rigorous, reproducible data investigations and ad-hoc analyses.

  • Works in 5 steps: Frame Before Querying → Build Queries In Escalating Specificity → SQL Hygiene Rules → …
  • Tasks that involve Root cause analysis
  • SKILL.md covers Purpose, Phase 1: Frame Before Querying, Phase 2: Build Queries In… and Phase 3: SQL Hygiene Rules, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Investigation is an agent skill from Kilo-Org/kilo-marketplace. A workflow for conducting rigorous, reproducible data investigations and ad-hoc analyses. This skill should be used when investigating a business question, explaining a metric anomaly, validating a hypothesis, performing root cause analysis, or preparing a one-off investigation dashboard or memo.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Root cause analysis. It works with SQL. The repository describes itself as: Kilo Marketplace - A curated collection of Skills, MCP Servers, and Modes for enhancing AI agent capabilities across the Kilo ecosystem—including Kilo Code (VS Code extension)… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Root cause analysis

Example prompts

  • “/data-investigation”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Frame Before Querying
  2. Build Queries In Escalating Specificity
  3. SQL Hygiene Rules
  4. Communication Rules
  5. Output Standard

What it can do on your machine

Read from SKILL.md and the folder at commit ff51758. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are sql).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Investigation loads about 1.5k tokens when it runs. Until then it costs about 79 tokens; SKILL.md has 660 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~79
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Kilo-Org/kilo-marketplace at commit ff51758, republished under its Apache-2.0 licence (© Kilo-Org). 660 words, ~1,480 tokens.

Download SKILL.mdSave it as .claude/skills/data-investigation/SKILL.md (or your agent's skills folder).
name
data-investigation
description
A workflow for conducting rigorous, reproducible data investigations and ad-hoc analyses. This skill should be used when investigating a business question, explaining a metric anomaly, validating a hypothesis, performing root cause analysis, or preparing a one-off investigation dashboard or memo.
metadata.category
data
metadata.author
Pedro / Kilo Code

Data Investigation

Use this skill to produce investigations that are fast, correct, reproducible, and communicate a clear conclusion rather than a pile of charts.

Purpose

Every investigation should be answerable in one sentence before the first SQL query is written.

Phase 1: Frame Before Querying

1. Write the one-sentence answer first

Before writing any SQL, write the sentence the conclusion is expected to be.

Example: The cohort size gap is a definition problem rather than a product behavior problem.

If the sentence cannot be written, the question is not yet understood.

2. Classify the investigation type
TypeTriggerApproach
Gap analysisWhy do A and B not match?Establish the gap, localize it, explain it
Root causeWhy did this metric change?Confirm real, isolate segment, align timing, validate mechanism
Hypothesis testIs X causing Y?Define what must be true, test sub-claims, confirm or reject
Feasibility checkIs this number trustworthy?Check grain, joins, nulls, definition overlap
3. State 2-3 competing hypotheses before querying

Never investigate with a single hypothesis. That creates confirmation bias.

Order hypotheses by plausibility and note which one is currently expected to be correct and why.

Phase 2: Build Queries In Escalating Specificity

1. Establish first, explain second

Step 1 always confirms the anomaly is real and measures its magnitude. Do not jump to cause until the effect is confirmed.

sql
select
    <time_bucket>,
    <source_a_count> as metric_a,
    <source_b_count> as metric_b,
    <source_a_count> - <source_b_count> as gap,
    round(100.0 * (<source_a_count> - <source_b_count>) / nullif(<source_a_count>, 0), 1) as pct_gap
from ...
order by 1
2. Localize by breaking one dimension at a time

After confirming the gap, break it down along one dimension per step:

  • By time: when did it appear?
  • By segment: who is affected?
  • By signal/source: which path is missing?
3. Align timing with upstream changes

Once the anomaly start date is known, check:

  • git commits to relevant models, ETL jobs, or application code
  • schema migrations or source changes
  • metric definition changes
  • external events such as launches, campaigns, or pricing changes

A credible root cause must explain why the metric changed when it did.

4. Validate the mechanism

After identifying a candidate cause, confirm it mechanically:

  • Does the affected population match the predicted population?
  • What does the counterfactual look like?
  • Does a second independent signal support the explanation?

Do not stop at correlation.

5. Quantify each hypothesis before concluding

For every candidate explanation, produce a number.

Examples:

  • H1 accounts for 1,240 users
  • H2 accounts for 1,050 users

If a hypothesis cannot be quantified, it is not yet validated.

Phase 3: SQL Hygiene Rules

Show full SKILL.md (264 more words)Show less
Use stable time bounds for investigations

Investigation SQL should be reproducible. Favor fixed bounds over rolling windows unless the analysis is intentionally operational and evergreen.

sql
-- Prefer fixed investigation scope
where created_at >= '2026-02-16'
  and created_at < '2026-04-04'
Add explicit grain comments in important CTEs
sql
-- grain: one row per user
with users as (
    ...
)
Prefer named comparison outputs over raw aggregates

Bad:

sql
select count(*) from ...

Better:

sql
select
    count(*) as users_in_scope,
    count(distinct org_id) as orgs_in_scope
from ...
Keep diagnostic queries small and discardable

Diagnostic SQL exists to answer one question. Do not build giant reusable queries too early.

Phase 4: Communication Rules

Lead with the answer

The first sentence of the written output should state the conclusion.

Bad:

I looked at several possible explanations for the discrepancy.

Good:

The discrepancy is caused by internal users being excluded from the dashboard query but included in the warehouse baseline.

Separate evidence from interpretation

Use a structure like:

  1. Conclusion
  2. Evidence
  3. Why alternative hypotheses were rejected
  4. Recommended next action
Include the limiting assumption

Every investigation should state the biggest assumption that could change the answer.

Example:

This conclusion assumes backend event timestamps are complete for the affected week; if ingestion was delayed, the gap may be overstated.

Phase 5: Output Standard

Every completed investigation should leave behind:

  1. One-sentence answer
  2. Final supporting SQL or notebook
  3. A note on rejected hypotheses
  4. Any follow-up question or unresolved ambiguity

Common Failure Modes

  • Treating a metric definition dispute as a product issue
  • Comparing sources with different grain or freshness
  • Accepting the first plausible explanation without quantification
  • Writing one giant query too early
  • Ending with charts instead of an answer

© Kilo-Org, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/data-investigation of Kilo-Org/kilo-marketplace.

Open the folder on GitHubat commit ff51758

Compare with similar skills

Data Investigation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Investigation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Investigation this skillKilo-Org/kilo-marketplace190—~1.5kAutomated safety check: PassApache-2.0
Review PRapache/shardingsphere21k—~6.4kAutomated safety check: PassApache-2.0
SQL Root Cause Analysiszj-unicom-ai/UniEmployee358—~433Automated safety check: PassMIT
Analyze Issueapache/shardingsphere21k—~6.5kAutomated safety check: PassApache-2.0
Logfire Querypydantic/skills140—~2.2kAutomated safety check: PassMIT
Perfetto Trace Analysisjameshnsears/QuoteUnquote1002 repos~1.7kAutomated safety check: PassApache-2.0

Similar skills

  • Review PR

    apache/shardingsphere

    Review Apache ShardingSphere or user-authorized downstream pull requests and PR discussions from public or authorized repository evidence.

    21k GitHub stars~6.4k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • SQL Root Cause Analysis

    zj-unicom-ai/UniEmployee

    SQL 版归因分析技能。当用户问为什么、指标异常、趋势下滑/增长、KPI 未达标、营收/订单/转化/成本波动时使用. An agent skill from zj-unicom-ai/UniEmployee.

    358 GitHub stars~433 tokensUpdated today
    DevelopmentAuto-check passed
  • Analyze Issue

    apache/shardingsphere

    Used to analyze Apache ShardingSphere community issues. An agent skill from apache/shardingsphere.

    21k GitHub stars~6.5k tokensUpdated yesterday
    DatabasesAuto-check passed
  • Logfire Query

    pydantic/skills

    Official

    Query and analyze Logfire telemetry data — traces, logs, spans, metrics, summaries, and SQL results.

    140 GitHub stars~2.2k tokensUpdated 7 days ago
    DatabasesAuto-check passed
  • Perfetto Trace Analysis

    jameshnsears/QuoteUnquote

    Analyzes Perfetto traces to find the root cause of latency, memory, or jank issues in Android apps.

    100 GitHub starsUsed in 2 repos~1.7k tokens
    MobileAuto-check passed
  • Diagnose Clickhouse Errors

    FrankChen021/datastoria

    Diagnose ClickHouse runtime query failures when the user wants database-level cause and fix guidance from an error or numeric error code, not source-code root cause analysis.

    327 GitHub stars~610 tokensUpdated 2 mo ago
    DatabasesAuto-check passed

More from Kilo-Org/kilo-marketplace

All 85 skills in this repo
  • AzureML Project Scaffolding

    Kilo-Org/kilo-marketplace

    Sets up and maintains AzureML-ready Python projects as uv workspaces with devcontainers, a Makefile and job YAML, so local runs match cloud jobs and experiments stay reproducible.

    190 GitHub stars~3.1k tokensUpdated 10 days ago
    Auto-check: notes
  • Jupyter Notebook Builder

    Kilo-Org/kilo-marketplace

    Creates, inspects, edits and runs Jupyter notebooks, scaffolding experiment or tutorial notebooks from templates and preferring a Jupyter MCP server over raw JSON edits.

    190 GitHub stars~1.3k tokensUpdated 10 days ago
    Auto-check passed
  • Tableau Dashboard Creator

    Kilo-Org/kilo-marketplace

    Takes a plain-language dashboard request through brand setup, data exploration, planning, an interactive HTML mock and a Tableau implementation spec.

    190 GitHub stars~3.8k tokensUpdated 10 days ago
    Auto-check: notes
  • Elasticsearch File Ingest

    Kilo-Org/kilo-marketplace

    Ingest and transform data files (CSV/JSON/Parquet/Arrow IPC) into Elasticsearch with stream processing and custom transforms.

    190 GitHub stars~2.8k tokensUpdated 10 days ago
    Auto-check passed
  • Nifi Flow Layout

    Kilo-Org/kilo-marketplace

    A skill your agent uses when arranging Apache NiFi processors, process groups, ports, comments, numbering, crossing connections, dense fan-in/fan-out, or reusable readable canvas layouts.

    190 GitHub stars~1.5k tokensUpdated 10 days ago
    Auto-check passed
  • Splunk Ingest Processor Setup

    Kilo-Org/kilo-marketplace

    Render Cisco Data Fabric ingest-time routing workflows and Splunk Cloud Platform Ingest Processor setup plans with SPL2 pipelines, source types, destinations, lifecycle handoffs, queue and…

    190 GitHub stars~1.2k tokensUpdated 10 days ago
    Auto-check passed

Works with

Categories

Questions about Data Investigation

What does Data Investigation do?

A workflow for conducting rigorous, reproducible data investigations and ad-hoc analyses. Data Investigation is an agent skill from Kilo-Org/kilo-marketplace. A workflow for conducting rigorous, reproducible data investigations and ad-hoc analyses.

When should I use Data Investigation?

Data Investigation fits situations like: tasks that involve Root cause analysis.

How do I install Data Investigation in Claude Code?

Run `npx skills add Kilo-Org/kilo-marketplace --skill data-investigation -a claude-code`. Or copy the skill folder (skills/data-investigation in Kilo-Org/kilo-marketplace) into .claude/skills/data-investigation in your project. Claude Code loads it when a task matches its description.

How do I install Data Investigation in Codex?

Run `npx skills add Kilo-Org/kilo-marketplace --skill data-investigation -a codex`. Or copy the skill folder (skills/data-investigation in Kilo-Org/kilo-marketplace) into .agents/skills/data-investigation in your project. Codex loads it when a task matches its description.

Can I use Data Investigation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Kilo-Org/kilo-marketplace --skill data-investigation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-investigation, .gemini/skills/data-investigation, .github/skills/data-investigation and .opencode/skills/data-investigation in your project.

What does Data Investigation need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Investigation is instructions for the agent only.

Does Data Investigation access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Data Investigation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Investigation use?

Data Investigation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Investigation use?

About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Investigation?

Skills that share tags, products or a category with Data Investigation: Review PR (apache/shardingsphere, 21k stars), SQL Root Cause Analysis (zj-unicom-ai/UniEmployee, 358 stars), Analyze Issue (apache/shardingsphere, 21k stars) and Logfire Query (pydantic/skills, 140 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Investigation?

Kilo-Org (a GitHub organization) maintains it in Kilo-Org/kilo-marketplace, which has 190 GitHub stars. The repository holds 85 skills in this directory. The repository was last updated on September 28, 2026.

Source: Kilo-Org/kilo-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.