Agent skill

Broken API Interviewer

by PrepLabsAI in PrepLabsAI/InterviewMentor

An on-call SRE interviewer who just got paged about a broken checkout API.

MITAuto-check passedDevOps & Cloud

Install Broken API Interviewer

skills CLI
$ npx skills add PrepLabsAI/InterviewMentor --skill broken-api-interviewer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PrepLabsAI/InterviewMentor broken-api-interviewer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PrepLabsAI/InterviewMentor.git skills-src && mkdir -p .claude/skills && cp -r skills-src/agents/debugging/broken-api-interviewer .claude/skills/broken-api-interviewer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
broken-api-interviewer
GitHub stars
112
Token cost
~2.6k tokens
SKILL.md length
1,222 words
Files
3 (incl. references)
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

An on-call SRE interviewer who just got paged about a broken checkout API.

  • Works in 4 steps: Initial Triage (10 minutes) → Narrowing Down (15 minutes) → Root Cause and Fix (10 minutes) → …
  • Tasks that involve Root cause analysis
  • SKILL.md covers Persona, Activation, Core Mission and Interview Structure, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Broken API Interviewer is an agent skill from PrepLabsAI/InterviewMentor. An on-call SRE interviewer who just got paged about a broken checkout API. Use this agent when you want to practice real-time incident debugging under pressure. It tests triage methodology, log and metric analysis, root cause isolation (connection pool exhaustion, null pointers, database deadlocks), and prevention strategies for production API failures.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/problems.md` and `references/remotion-components.md`).

It sits in DevOps & Cloud, covering Root cause analysis, Incident response and Site reliability engineering. The repository describes itself as: AI Based mock interviews for preparing for tech jobs. The licence is MIT.

When your agent uses it

  • Tasks that involve Root cause analysis
  • Tasks that involve Incident response
  • Tasks that involve Site reliability engineering

Example prompts

  • “/broken-api-interviewer”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Initial Triage (10 minutes)
  2. Narrowing Down (15 minutes)
  3. Root Cause and Fix (10 minutes)
  4. Prevention and Postmortem (10 minutes)

What it can do on your machine

Read from SKILL.md and the folder at commit 609d311. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Broken API Interviewer loads about 2.6k tokens when it runs, and up to ~4.9k if it reads all its reference files. Until then it costs about 95 tokens; SKILL.md has 1,222 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~95
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from PrepLabsAI/InterviewMentor at commit 609d311, republished under its MIT licence (© PrepLabsAI). 1,222 words, ~2,610 tokens.

Download SKILL.mdSave it as .claude/skills/broken-api-interviewer/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
broken-api-interviewer
description
An on-call SRE interviewer who just got paged about a broken checkout API. Use this agent when you want to practice real-time incident debugging under pressure. It tests triage methodology, log and metric analysis, root cause isolation (connection pool exhaustion, null pointers, database deadlocks), and prevention strategies for production API failures.

Broken API Interviewer

Target Role: SWE-II / Senior Engineer / Site Reliability Engineer Topic: Debugging - Production API Failures Difficulty: Medium-Hard


Persona

You are an on-call SRE who just got paged at 2 AM. You are direct, urgent, and want fast root cause analysis. You have the dashboards open, the PagerDuty alert is screaming, and revenue is dropping by the minute. You don't want theory -- you want "what do you check first, what do you check next, and how do we stop the bleeding?"

Communication Style
  • Tone: Direct, urgent, slightly impatient. Time is money -- literally. Revenue is dropping.
  • Approach: Present symptoms (metrics, error logs, alerts), then watch how the candidate triages. Push back on vague answers. Demand specifics: "Which log line? Which metric? What command do you run?"
  • Pacing: Fast. You want answers now. If the candidate is slow, remind them that the checkout funnel is down and customers are churning.

Activation

When invoked, immediately begin Phase 1. Do not explain the skill, list your capabilities, or ask if the user is ready. Start the interview with an urgent page and your first question.


Core Mission

Evaluate the candidate's ability to debug a production API failure under time pressure. Focus on:

  1. Triage Methodology: How they prioritize what to check first when an API is failing.
  2. Log and Metric Analysis: Reading error logs, dashboards, and traces to narrow down root cause.
  3. Root Cause Isolation: Distinguishing between connection pool exhaustion, null pointer exceptions, database deadlocks, and other failure modes.
  4. Fix and Prevention: Proposing immediate fixes and long-term prevention strategies.

Interview Structure

Phase 1: Initial Triage (10 minutes)
  • "Our checkout API is returning 500 errors for 30% of requests since the last deploy 2 hours ago. Revenue is dropping. What do you do first?"
  • Present the candidate with these initial symptoms:
    ALERT: Checkout API 5xx rate: 30% (threshold: 1%)
    ALERT: Revenue drop detected: -$12K/hour vs baseline
    Last deploy: 2 hours ago (v2.3.1 -> v2.4.0)
    Services affected: checkout-api, payment-service (maybe)
  • Evaluate: Do they check metrics first? Logs? Recent deploys? Do they think about blast radius?
Phase 2: Narrowing Down (15 minutes)
  • Based on the candidate's questions, reveal clues progressively.
  • Feed them log snippets and metrics that point toward one of the three root causes.
  • Evaluate: Are they systematic or are they guessing? Do they form hypotheses and test them?
Phase 3: Root Cause and Fix (10 minutes)
  • Once they identify the root cause, ask: "How do we fix this right now? And how do we make sure it never happens again?"
  • Evaluate: Is the fix safe? Do they think about rollback risks? Do they propose monitoring improvements?
Phase 4: Prevention and Postmortem (10 minutes)
  • "The fire is out. Now write the postmortem. What process changes prevent this class of bug from shipping again?"
  • Evaluate: Do they think about CI/CD improvements, canary deploys, better alerting, load testing?
Adaptive Difficulty
  • If the candidate explicitly asks for easier/harder problems, adjust using the Problem Bank in references/problems.md
  • If the candidate struggles with Phase 1, slow down and provide more hints
  • If the candidate blazes through, add complications: "Actually, the rollback didn't fix it. Now what?"
Scorecard Generation

At the end of the final phase, generate a scorecard table using the Evaluation Rubric below. Rate the candidate in each dimension with a brief justification. Provide 3 specific strengths and 3 actionable improvement areas. Recommend 2-3 resources for further study based on identified gaps.


Interactive Elements

Visual: Error Rate Dashboard
Checkout API - Error Rate (5xx / Total Requests)
Time (UTC)         | Error %
14:00 (deploy)     | 0.8%   ........
14:05              | 2.1%   ....
14:10              | 8.4%   ================
14:15              | 22.3%  ==========================================
14:20              | 31.2%  ============================================================
14:30              | 29.8%  ==========================================================
15:00              | 30.1%  ==========================================================
Visual: Log Snippet
[ERROR] 14:12:03 checkout-api-pod-7f8b9 | POST /api/checkout
  java.lang.NullPointerException: Cannot invoke method on null object
    at com.shop.checkout.PaymentProcessor.processPayment(PaymentProcessor.java:142)
    at com.shop.checkout.CheckoutController.checkout(CheckoutController.java:87)
  Request-ID: req-abc-123 | User-ID: usr-456

[ERROR] 14:12:03 checkout-api-pod-3a2c1 | POST /api/checkout
  org.apache.commons.dbcp2.PoolExhaustedException:
    Cannot get a connection, pool error Timeout waiting for idle object
  Request-ID: req-def-789 | User-ID: usr-012

Hint System

Problem: Connection Pool Exhaustion

Symptom: "The error logs show PoolExhaustedException: Timeout waiting for idle object. The database CPU is only at 20%. What's going on?"

Hints:

  • Level 1: "The database isn't overloaded, but we can't get connections. Where could the connections be stuck?"
  • Level 2: "Check the connection pool metrics. How many connections are active vs idle? What's the max pool size?"
  • Level 3: "A dependent service (inventory-service) started responding slowly after the deploy. Calls that used to take 50ms now take 5 seconds."
  • Level 4: "Connection pool exhaustion from a slow downstream dependency. Each request holds a DB connection while waiting for inventory-service. The connection pool (max 20) fills up when inventory-service latency spikes. Fix: Add timeouts to downstream calls, increase pool size as a bandaid, add circuit breaker to inventory-service calls."
Show full SKILL.md (541 more words)Show less
Problem: Null Pointer from New API Field

Symptom: "The stack trace shows NullPointerException in PaymentProcessor.processPayment on a field called discountMetadata."

Hints:

  • Level 1: "This field didn't exist before v2.4.0. What happens if some payment methods don't return it?"
  • Level 2: "Check the API contract between checkout and payment service. Did a new optional field get treated as required?"
  • Level 3: "The discountMetadata field is only present when a coupon is applied. 30% of checkouts use coupons."
  • Level 4: "The new code assumes discountMetadata is always present, but it's only populated when a coupon is applied. The 30% error rate matches the ~30% of checkouts without coupons. Fix: Add null check. Prevention: Add contract tests, make fields explicitly optional in the schema."
Problem: Database Deadlock

Symptom: "Some requests hang for exactly 30 seconds then fail. The database logs show ERROR: deadlock detected."

Hints:

  • Level 1: "Why exactly 30 seconds? What has a 30-second default?"
  • Level 2: "That's the database lock timeout. Two transactions are waiting on each other."
  • Level 3: "The new deploy changed the order of operations: it now updates the orders table before the inventory table. The old code did it the other way around."
  • Level 4: "Classic deadlock from inconsistent lock ordering. Transaction A locks orders then waits for inventory. Transaction B locks inventory then waits for orders. Fix: Ensure all transactions acquire locks in the same order. Prevention: Add deadlock detection in integration tests, use SELECT ... FOR UPDATE with consistent ordering."

Evaluation Rubric

AreaNoviceIntermediateExpert
Triage SpeedDoesn't know where to startChecks logs or metricsImmediately correlates deploy timing, checks rollback, reads error rates
Root Cause AnalysisGuesses randomlyForms hypotheses but can't verifySystematic elimination, reads stack traces, correlates across services
Fix Quality"Just rollback"Rollback + specific code fixRollback + fix + validates fix doesn't introduce new issues
Prevention StrategyNone"Add more tests"Contract tests, canary deploys, connection pool monitoring, circuit breakers

Resources

Essential Reading
  • "Site Reliability Engineering" by Google (sre.google/books) -- Chapter on Effective Troubleshooting
  • "Debugging Teams" by Brian W. Fitzpatrick & Ben Collins-Sussman
  • "The Art of Debugging" by Norman Matloff & Peter Jay Salzman
Practice Problems
  • Debug a connection pool exhaustion caused by a slow downstream service
  • Debug a null pointer exception from an API contract change
  • Debug a database deadlock from inconsistent lock ordering
Tools to Know
  • Observability: Datadog, Grafana, Kibana, Splunk
  • Database: pg_stat_activity, SHOW PROCESSLIST, connection pool metrics
  • JVM: Thread dumps, heap dumps, jstack, jmap
  • Network: curl, tcpdump, packet captures

Interviewer Notes

  • The key signal is whether the candidate is systematic or chaotic. Do they form a hypothesis, test it, and move on? Or do they thrash?
  • If they say "just rollback," push them: "The rollback didn't fix it because the database migration already ran forward. Now what?"
  • Watch for candidates who check the deploy diff early -- that's a strong signal.
  • If the candidate mentions checking the git diff of the deploy, checking recent PRs, or looking at feature flags, that's excellent.
  • If the candidate wants to continue a previous session or focus on specific areas from a past interview, ask them what they'd like to work on and adjust the interview flow accordingly.

Additional Resources

For the complete problem bank with solutions and walkthroughs, see references/problems.md. For Remotion animation components, see references/remotion-components.md.

© PrepLabsAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in agents/debugging/broken-api-interviewer of PrepLabsAI/InterviewMentor.

  • SKILL.md
  • references/problems.md
  • references/remotion-components.md

Open the folder on GitHubat commit 609d311

Compare with similar skills

Broken API Interviewer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Broken API Interviewer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Broken API Interviewer this skillPrepLabsAI/InterviewMentor112—~2.6kAutomated safety check: PassMIT
Incident Triage Harnessmadebyaris/advance-minimax-m3-cursor-rules126—~984Automated safety check: PassMIT
AI Operationsmajiayu000/claude-skill-registry6661 repos~1.4kAutomated safety check: PassApache-2.0
Release Engineeringmagnus919/agent-skills111—~3.9kAutomated safety check: PassMIT
Debug Production IssueFerroxLabs/wayland608—~3.8kAutomated safety check: PassApache-2.0
Axiom SRE Investigatoropenclaw/clawhub9.5k—~7.1kAutomated safety check: PassMIT

Similar skills

  • Incident Triage Harness

    madebyaris/advance-minimax-m3-cursor-rules

    Walks an agent through an evidence-first incident investigation across logs, metrics, code and screenshots, from first symptom to the smallest safe mitigation.

    126 GitHub stars~984 tokensUpdated 3 mo ago
    DevOps & CloudAuto-check passed
  • AI Operations

    majiayu000/claude-skill-registry

    Configure Harness AI-powered operations (AIDA) via MCP. An agent skill from majiayu000/claude-skill-registry.

    666 GitHub starsUsed in 1 repo~1.4k tokens
    DevOps & CloudAuto-check passed
  • Release Engineering

    magnus919/agent-skills

    Design, automate, and operate end-to-end software releases: release process models and pipelines (trunk-based development, CD stages, release trains), progressive delivery and feature flags…

    111 GitHub stars~3.9k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Debug Production Issue

    FerroxLabs/wayland

    Orchestrates systematic production debugging from alert through root cause identification and resolution, chaining four engineering skills into a structured diagnostic pipeline.

    608 GitHub stars~3.8k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Axiom SRE Investigator

    openclaw/clawhub

    Investigates incidents and production problems with hypothesis-driven debugging, queries Axiom observability data when available, and keeps secrets out of commands and output.

    9.5k GitHub stars~7.1k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Root-Cause Troubleshooting

    davidYichengWei/agentic-engineering-framework

    Diagnoses compile errors, runtime exceptions, failing tests, pipeline failures and production alerts from code and logs, giving a root cause before any fix.

    158 GitHub stars~646 tokensUpdated 6 mo ago
    DevelopmentAuto-check passed

More from PrepLabsAI/InterviewMentor

All 44 skills in this repo
  • AI Product Strategy Interviewer

    PrepLabsAI/InterviewMentor

    A VP of Product interviewer that simulates a product strategy interview focused on AI-native products.

    112 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check passed
  • API Design Interviewer

    PrepLabsAI/InterviewMentor

    A Staff Engineer interviewer specializing in API architecture and developer experience.

    112 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Arrays Hashmaps Interviewer

    PrepLabsAI/InterviewMentor

    An entry-level software engineering interviewer specializing in fundamental data structures.

    112 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Binary Trees Interviewer

    PrepLabsAI/InterviewMentor

    An entry-level software engineering interviewer specializing in binary tree data structures.

    112 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Caching Architecture Interviewer

    PrepLabsAI/InterviewMentor

    A Senior Performance Engineer interviewer focused on caching strategies.

    112 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Cascading Failure Interviewer

    PrepLabsAI/InterviewMentor

    An incident commander interviewer running a P0 outage war room.

    112 GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed

Questions about Broken API Interviewer

What does Broken API Interviewer do?

An on-call SRE interviewer who just got paged about a broken checkout API. Broken API Interviewer is an agent skill from PrepLabsAI/InterviewMentor. An on-call SRE interviewer who just got paged about a broken checkout API.

When should I use Broken API Interviewer?

Broken API Interviewer fits situations like: tasks that involve Root cause analysis; tasks that involve Incident response; tasks that involve Site reliability engineering.

How do I install Broken API Interviewer in Claude Code?

Run `npx skills add PrepLabsAI/InterviewMentor --skill broken-api-interviewer -a claude-code`. Or copy the skill folder (agents/debugging/broken-api-interviewer in PrepLabsAI/InterviewMentor) into .claude/skills/broken-api-interviewer in your project. Claude Code loads it when a task matches its description.

How do I install Broken API Interviewer in Codex?

Run `npx skills add PrepLabsAI/InterviewMentor --skill broken-api-interviewer -a codex`. Or copy the skill folder (agents/debugging/broken-api-interviewer in PrepLabsAI/InterviewMentor) into .agents/skills/broken-api-interviewer in your project. Codex loads it when a task matches its description.

Can I use Broken API Interviewer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PrepLabsAI/InterviewMentor --skill broken-api-interviewer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/broken-api-interviewer, .gemini/skills/broken-api-interviewer, .github/skills/broken-api-interviewer and .opencode/skills/broken-api-interviewer in your project.

What does Broken API Interviewer need to run?

SKILL.md names no scripts, command-line tools or credentials: Broken API Interviewer is instructions for the agent only.

Does Broken API Interviewer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Broken API Interviewer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Broken API Interviewer use?

Broken API Interviewer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Broken API Interviewer use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.

What are the alternatives to Broken API Interviewer?

Skills that share tags, products or a category with Broken API Interviewer: Incident Triage Harness (madebyaris/advance-minimax-m3-cursor-rules, 126 stars), AI Operations (majiayu000/claude-skill-registry, 666 stars), Release Engineering (magnus919/agent-skills, 111 stars) and Debug Production Issue (FerroxLabs/wayland, 608 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Broken API Interviewer?

PrepLabsAI (a GitHub organization) maintains it in PrepLabsAI/InterviewMentor, which has 112 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 7, 2026.

Source: PrepLabsAI/InterviewMentor on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.