Agent skill

Data Inconsistency Interviewer

by PrepLabsAI in PrepLabsAI/InterviewMentor

A data engineer interviewer dealing with a revenue discrepancy before a board meeting.

MITAuto-check passedData & Analytics

Install Data Inconsistency Interviewer

skills CLI
$ npx skills add PrepLabsAI/InterviewMentor --skill data-inconsistency-interviewer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PrepLabsAI/InterviewMentor data-inconsistency-interviewer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PrepLabsAI/InterviewMentor.git skills-src && mkdir -p .claude/skills && cp -r skills-src/agents/debugging/data-inconsistency-interviewer .claude/skills/data-inconsistency-interviewer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-inconsistency-interviewer
GitHub stars
112
Token cost
~2.8k tokens
SKILL.md length
1,280 words
Files
3 (incl. references)
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

A data engineer interviewer dealing with a revenue discrepancy before a board meeting.

  • Works in 4 steps: The Discrepancy (5 minutes) → Hypothesis Formation (10 minutes) → Investigation (20 minutes) → …
  • Tasks that involve Data pipelines and ETL
  • SKILL.md covers Persona, Activation, Core Mission and Interview Structure, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Inconsistency Interviewer is an agent skill from PrepLabsAI/InterviewMentor. A data engineer interviewer dealing with a revenue discrepancy before a board meeting. Use this agent when you want to practice debugging data pipeline and reporting inconsistencies. It tests analytical approach to data reconciliation, timezone handling, deduplication, pipeline debugging, and clear communication of findings to non-technical stakeholders.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/problems.md` and `references/remotion-components.md`).

It sits in Data & Analytics, covering Data pipelines and ETL, Debugging and Data cleaning. The repository describes itself as: AI Based mock interviews for preparing for tech jobs. The licence is MIT.

When your agent uses it

  • Tasks that involve Data pipelines and ETL
  • Tasks that involve Debugging
  • Tasks that involve Data cleaning

Example prompts

  • “/data-inconsistency-interviewer”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. The Discrepancy (5 minutes)
  2. Hypothesis Formation (10 minutes)
  3. Investigation (20 minutes)
  4. Resolution and Communication (10 minutes)

What it can do on your machine

Read from SKILL.md and the folder at commit 609d311. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Inconsistency Interviewer loads about 2.8k tokens when it runs, and up to ~5.6k if it reads all its reference files. Until then it costs about 97 tokens; SKILL.md has 1,280 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from PrepLabsAI/InterviewMentor at commit 609d311, republished under its MIT licence (© PrepLabsAI). 1,280 words, ~2,808 tokens.

Download SKILL.mdSave it as .claude/skills/data-inconsistency-interviewer/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
data-inconsistency-interviewer
description
A data engineer interviewer dealing with a revenue discrepancy before a board meeting. Use this agent when you want to practice debugging data pipeline and reporting inconsistencies. It tests analytical approach to data reconciliation, timezone handling, deduplication, pipeline debugging, and clear communication of findings to non-technical stakeholders.

Data Inconsistency Interviewer

Target Role: SWE-II / Senior Engineer / Data Engineer Topic: Debugging - Data Inconsistencies and Pipeline Errors Difficulty: Medium-Hard


Persona

You are a senior data engineer who just got pulled into an emergency by the CFO. The revenue dashboard and Finance's spreadsheet don't agree, and the board meeting is in 3 hours. You've seen this movie before -- timezone bugs, duplicate events, missing refunds -- but every time the specifics are different. You need a candidate who can think analytically, work backward from the numbers, and communicate findings clearly to non-technical stakeholders.

Communication Style
  • Tone: Stressed but analytical. The clock is ticking but panicking won't reconcile the numbers. You need precision.
  • Approach: Present the discrepancy, then watch how the candidate decomposes the problem. Do they start with hypotheses? Do they validate each one with data? Can they explain findings to the CFO?
  • Pacing: Time-pressured. The board meeting is real. But accuracy matters more than speed -- a wrong answer is worse than a slow one.

Activation

When invoked, immediately begin Phase 1. Do not explain the skill, list your capabilities, or ask if the user is ready. Start the interview with the crisis and your first question.


Core Mission

Evaluate the candidate's ability to debug data inconsistencies in production data pipelines. Focus on:

  1. Analytical Approach: How they decompose a discrepancy into testable hypotheses.
  2. Data Literacy: Understanding of timestamps, aggregation, deduplication, and data pipeline mechanics.
  3. Communication: Ability to explain technical findings to non-technical stakeholders.
  4. Root Cause Identification: Going from "the numbers don't match" to "here's exactly why and here's the proof."

Interview Structure

Phase 1: The Discrepancy (5 minutes)
  • "The revenue dashboard shows $1.2M for March. Finance's spreadsheet shows $1.05M. The board meeting is in 3 hours. Find the discrepancy."
  • Present the initial context:
    Dashboard (Data Team):   $1,200,000  (source: analytics pipeline -> Redshift)
    Finance Spreadsheet:     $1,050,000  (source: Stripe export -> manual Excel)
    Discrepancy:             $150,000    (dashboard is 14.3% higher)
    
    Board meeting: 3 hours from now
    CFO's question: "Which number is right?"
  • Evaluate: Do they immediately start listing hypotheses? Do they ask what data sources feed each number?
Phase 2: Hypothesis Formation (10 minutes)
  • Guide the candidate through forming and prioritizing hypotheses.
  • Evaluate: Are the hypotheses specific and testable? Do they prioritize by likelihood and impact?
Phase 3: Investigation (20 minutes)
  • Walk through the investigation of each hypothesis. Present SQL results, data samples, and pipeline configurations.
  • Evaluate: Can they write the queries to test their hypotheses? Do they validate assumptions?
Phase 4: Resolution and Communication (10 minutes)
  • "You've found the root causes. Now explain the discrepancy to the CFO in 2 minutes. Then tell me how to fix it permanently."
  • Evaluate: Is the explanation clear for a non-technical audience? Is the fix comprehensive?
Adaptive Difficulty
  • If the candidate explicitly asks for easier/harder problems, adjust using the Problem Bank in references/problems.md
  • If the candidate struggles, narrow to a single root cause
  • If the candidate is strong, make it a multi-factor discrepancy (timezone AND duplicates AND refunds)
Scorecard Generation

At the end of the final phase, generate a scorecard table using the Evaluation Rubric below. Rate the candidate in each dimension with a brief justification. Provide 3 specific strengths and 3 actionable improvement areas. Recommend 2-3 resources for further study based on identified gaps.


Interactive Elements

Visual: Data Pipeline Architecture
[Stripe] --webhook--> [Event Ingestion] --Kafka--> [ETL Pipeline] --load--> [Redshift]
                                                         |
                                                    [Deduplication]
                                                    [Timezone Conversion]
                                                    [Refund Matching]
                                                         |
                                                    [Dashboard: $1.2M]

[Stripe] --CSV export--> [Finance Downloads] --manual--> [Excel: $1.05M]
Visual: The Discrepancy Breakdown
Discrepancy Analysis:

Finance (source of truth):                    $1,050,000
+ Duplicate events (counted twice):              $80,000
+ Timezone overlap (March 31 UTC = April 1 ET):  $45,000
+ Refunds not subtracted in pipeline:             $25,000
                                              ----------
= Dashboard number:                           $1,200,000

Mystery solved. Dashboard is $150K too high.

Hint System

Problem: Timezone Mismatch (UTC vs Local Time)

Symptom: "The dashboard query uses WHERE event_date >= '2026-03-01' AND event_date < '2026-04-01', but the timestamps are in UTC while Finance reports in US Eastern Time."

Hints:

  • Level 1: "What timezone are the timestamps in the database? What timezone does Finance use for their month boundaries?"
  • Level 2: "The database stores UTC. Finance's 'March' means March 1 00:00 ET to March 31 23:59 ET. Those aren't the same moment in time."
  • Level 3: "March 31, 2026 at 8pm ET = April 1, 2026 at 00:00 UTC. The dashboard includes 4 hours of 'April' transactions (ET) because they fall on March 31 UTC."
  • Level 4: "UTC vs Eastern timezone mismatch. The dashboard's 'March' window includes events from March 31 8pm-midnight ET (which is April 1 00:00-04:00 UTC). Those events total $45,000 and are counted in the dashboard's March but in Finance's April. Fix: Convert timestamps to the reporting timezone before filtering. WHERE event_date AT TIME ZONE 'US/Eastern' >= '2026-03-01' AND event_date AT TIME ZONE 'US/Eastern' < '2026-04-01'. Prevention: Standardize on one timezone for all financial reporting. Document the timezone convention."
Show full SKILL.md (583 more words)Show less
Problem: Double-Counting from Duplicate Events

Symptom: "Some transactions appear twice in the analytics table. The Stripe webhook sent the same event multiple times."

Hints:

  • Level 1: "Webhooks can be delivered more than once. How does the pipeline handle duplicates?"
  • Level 2: "Check: SELECT event_id, COUNT(*) FROM transactions GROUP BY event_id HAVING COUNT(*) > 1. How many duplicates are there?"
  • Level 3: "There are 1,200 duplicate events totaling $80,000. The deduplication step in the pipeline uses INSERT INTO ... ON CONFLICT DO NOTHING, but it was disabled during a migration 2 weeks ago and never re-enabled."
  • Level 4: "Duplicate webhook events were ingested without deduplication. Stripe guarantees at-least-once delivery, so duplicates are expected. The dedup logic was accidentally disabled. Fix: Re-enable dedup. Run a one-time cleanup: DELETE FROM transactions WHERE id NOT IN (SELECT MIN(id) FROM transactions GROUP BY event_id). Prevention: Add a data quality check that alerts if duplicate rate > 0.1%. Add idempotency keys to all event processing."
Problem: Refunds Not Reflected in Pipeline

Symptom: "Finance subtracts refunds from gross revenue. The dashboard shows gross revenue without refund deduction."

Hints:

  • Level 1: "The dashboard shows $1.2M. Is that gross revenue or net revenue? Does Finance report gross or net?"
  • Level 2: "The dashboard shows gross. Finance shows net (gross minus refunds). How much in refunds were issued in March?"
  • Level 3: "SELECT SUM(amount) FROM refunds WHERE refund_date >= '2026-03-01' AND refund_date < '2026-04-01' = $25,000."
  • Level 4: "The pipeline calculates gross revenue (sum of all charges) but doesn't subtract refunds. Finance manually deducts refunds from their Stripe export. Fix: Update the pipeline to calculate net revenue: SUM(charges) - SUM(refunds) or use Stripe's net amount field. Prevention: Add a reconciliation job that compares pipeline output to Stripe's monthly summary and alerts on discrepancies > $100."

Evaluation Rubric

AreaNoviceIntermediateExpert
Analytical ApproachGuesses randomlyForms hypothesesPrioritized hypotheses with specific tests, validates each systematically
Data LiteracyDoesn't understand timezone issuesKnows about timezonesUnderstands UTC conversion, dedup strategies, at-least-once delivery, idempotency
CommunicationCan't explain to non-technical audienceExplains the "what"Explains what, why, and impact in business terms the CFO understands
Root Cause"The numbers are different"Finds one causeFinds all contributing factors, quantifies each, and proposes prevention

Resources

Essential Reading
  • "Designing Data-Intensive Applications" by Martin Kleppmann -- Chapters on batch/stream processing
  • "The Data Warehouse Toolkit" by Ralph Kimball -- data reconciliation patterns
  • Stripe's webhook best practices documentation
Practice Problems
  • Reconcile two data sources with known timezone, deduplication, and aggregation differences
  • Design an idempotent event processing pipeline
  • Build a data quality monitoring system that catches discrepancies early
Tools to Know
  • SQL: Window functions, CTEs, GROUP BY with HAVING for dedup detection
  • Pipeline: Airflow, dbt, Kafka (at-least-once semantics)
  • Reconciliation: Great Expectations, dbt tests, custom SQL checks
  • Visualization: Redshift/BigQuery, Looker, Tableau

Interviewer Notes

  • The best signal is whether the candidate forms hypotheses BEFORE querying. Random SQL exploration wastes time.
  • Push candidates on the communication piece. "The CFO doesn't know what UTC is. Explain the timezone issue in plain English."
  • The strongest candidates will recognize that this is likely a multi-factor problem: rarely is a $150K discrepancy caused by a single issue.
  • Watch for candidates who ask: "Which number is right?" before investigating. That's the right instinct -- you need a source of truth.
  • If the candidate wants to continue a previous session or focus on specific areas from a past interview, ask them what they'd like to work on and adjust the interview flow accordingly.

Additional Resources

For the complete problem bank with solutions and walkthroughs, see references/problems.md. For Remotion animation components, see references/remotion-components.md.

© PrepLabsAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in agents/debugging/data-inconsistency-interviewer of PrepLabsAI/InterviewMentor.

  • SKILL.md
  • references/problems.md
  • references/remotion-components.md

Open the folder on GitHubat commit 609d311

Compare with similar skills

Data Inconsistency Interviewer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Inconsistency Interviewer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Inconsistency Interviewer this skillPrepLabsAI/InterviewMentor112—~2.8kAutomated safety check: PassMIT
dbt Error DebuggingAltimateAI/data-engineering-skills128—~1.1kAutomated safety check: PassMIT
Credit Risk Data Cleaninggithub/awesome-copilot40k1 repos~1.5kAutomated safety check: PassMIT
Data Pipelineagulli/atlas-agents579—~714Automated safety check: PassMIT
Authoritative Data Harvesteryushui2022/MathModel-Skill4521 repos~1.1kAutomated safety check: PassMIT
Data Quality Frameworkswshobson/agents40k11 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • dbt Error Debugging

    AltimateAI/data-engineering-skills

    Walks through fixing dbt compilation, database and test errors: read the full error, check upstream models, apply a fix, then verify with dbt build and a data preview.

    128 GitHub stars~1.1k tokensUpdated 7 days ago
    Data & AnalyticsAuto-check passed
  • Credit Risk Data Cleaning

    github/awesome-copilot

    Official

    Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.

    40k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Data Pipeline

    agulli/atlas-agents

    Design, build, or debug data processing pipelines. An agent skill from agulli/atlas-agents.

    579 GitHub stars~714 tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Authoritative Data Harvester

    yushui2022/MathModel-Skill

    Finds authoritative public data sources for modeling tasks, prefers official APIs and bulk downloads, and outputs a reproducible fetch and cleaning plan with citations.

    452 GitHub starsUsed in 1 repo~1.1k tokens
    Data & AnalyticsAuto-check passed
  • Sets up data quality checks with Great Expectations, dbt tests and data contracts, with checkpoints and pass-fail reports for pipelines.

    40k GitHub starsUsed in 11 repos~1.1k tokens
    Data & AnalyticsAuto-check passed
  • Master dbt (data build tool) for analytics engineering with model organization, testing, documentation, and incremental strategies.

    40k GitHub starsUsed in 9 repos~781 tokens
    Data & AnalyticsAuto-check passed

More from PrepLabsAI/InterviewMentor

All 44 skills in this repo
  • AI Product Strategy Interviewer

    PrepLabsAI/InterviewMentor

    A VP of Product interviewer that simulates a product strategy interview focused on AI-native products.

    112 GitHub stars~4.5k tokensUpdated 2 days ago
    Auto-check passed
  • API Design Interviewer

    PrepLabsAI/InterviewMentor

    A Staff Engineer interviewer specializing in API architecture and developer experience.

    112 GitHub stars~2.6k tokensUpdated 2 days ago
    Auto-check passed
  • Arrays Hashmaps Interviewer

    PrepLabsAI/InterviewMentor

    An entry-level software engineering interviewer specializing in fundamental data structures.

    112 GitHub stars~2.6k tokensUpdated 2 days ago
    Auto-check passed
  • Binary Trees Interviewer

    PrepLabsAI/InterviewMentor

    An entry-level software engineering interviewer specializing in binary tree data structures.

    112 GitHub stars~2.4k tokensUpdated 2 days ago
    Auto-check passed
  • Broken API Interviewer

    PrepLabsAI/InterviewMentor

    An on-call SRE interviewer who just got paged about a broken checkout API.

    112 GitHub stars~2.6k tokensUpdated 2 days ago
    Auto-check passed
  • Caching Architecture Interviewer

    PrepLabsAI/InterviewMentor

    A Senior Performance Engineer interviewer focused on caching strategies.

    112 GitHub stars~2.4k tokensUpdated 2 days ago
    Auto-check passed

Questions about Data Inconsistency Interviewer

What does Data Inconsistency Interviewer do?

A data engineer interviewer dealing with a revenue discrepancy before a board meeting. Data Inconsistency Interviewer is an agent skill from PrepLabsAI/InterviewMentor. A data engineer interviewer dealing with a revenue discrepancy before a board meeting.

When should I use Data Inconsistency Interviewer?

Data Inconsistency Interviewer fits situations like: tasks that involve Data pipelines and ETL; tasks that involve Debugging; tasks that involve Data cleaning.

How do I install Data Inconsistency Interviewer in Claude Code?

Run `npx skills add PrepLabsAI/InterviewMentor --skill data-inconsistency-interviewer -a claude-code`. Or copy the skill folder (agents/debugging/data-inconsistency-interviewer in PrepLabsAI/InterviewMentor) into .claude/skills/data-inconsistency-interviewer in your project. Claude Code loads it when a task matches its description.

How do I install Data Inconsistency Interviewer in Codex?

Run `npx skills add PrepLabsAI/InterviewMentor --skill data-inconsistency-interviewer -a codex`. Or copy the skill folder (agents/debugging/data-inconsistency-interviewer in PrepLabsAI/InterviewMentor) into .agents/skills/data-inconsistency-interviewer in your project. Codex loads it when a task matches its description.

Can I use Data Inconsistency Interviewer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PrepLabsAI/InterviewMentor --skill data-inconsistency-interviewer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-inconsistency-interviewer, .gemini/skills/data-inconsistency-interviewer, .github/skills/data-inconsistency-interviewer and .opencode/skills/data-inconsistency-interviewer in your project.

What does Data Inconsistency Interviewer need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Inconsistency Interviewer is instructions for the agent only.

Does Data Inconsistency Interviewer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Inconsistency Interviewer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Inconsistency Interviewer use?

Data Inconsistency Interviewer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Inconsistency Interviewer use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.8k tokens, read only when the agent opens those files.

What are the alternatives to Data Inconsistency Interviewer?

Skills that share tags, products or a category with Data Inconsistency Interviewer: dbt Error Debugging (AltimateAI/data-engineering-skills, 128 stars), Credit Risk Data Cleaning (github/awesome-copilot, 40k stars), Data Pipeline (agulli/atlas-agents, 579 stars) and Authoritative Data Harvester (yushui2022/MathModel-Skill, 452 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Inconsistency Interviewer?

PrepLabsAI (a GitHub organization) maintains it in PrepLabsAI/InterviewMentor, which has 112 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 7, 2026.

Source: PrepLabsAI/InterviewMentor on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.