Agent skill

Failure Oriented Instrumentation

by ArabelaTso in ArabelaTso/Skills-4-SE

Selectively instruments code to capture runtime data for debugging failures and bugs.

Apache-2.0Auto-check passedDevelopment

Install Failure Oriented Instrumentation

skills CLI
$ npx skills add ArabelaTso/Skills-4-SE --skill failure-oriented-instrumentation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ArabelaTso/Skills-4-SE failure-oriented-instrumentation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ArabelaTso/Skills-4-SE.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/failure-oriented-instrumentation .claude/skills/failure-oriented-instrumentation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
failure-oriented-instrumentation
GitHub stars
253
Token cost
~2.1k tokens
SKILL.md length
672 words
Files
5 (incl. references)
Skills in repo
150
Repo updated
First seen
Licence
Apache-2.0

At a glance

Selectively instruments code to capture runtime data for debugging failures and bugs.

  • Works in 6 steps: Analyze the Failure → Determine Instrumentation Strategy → Select Instrumentation Patterns → …
  • Investigating crashes
  • SKILL.md covers Workflow, Quick Start Examples, Instrumentation Guidelines and Language-Specific References, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Failure Oriented Instrumentation is an agent skill from ArabelaTso/Skills-4-SE. Selectively instruments code to capture runtime data for debugging failures and bugs. Use when investigating crashes, exceptions, unexpected behavior, test failures, or performance issues. Analyzes stack traces and error messages to identify suspicious code regions, then adds targeted logging, tracing, and assertions to capture variable values, execution paths, timing, and conditional branches. Supports Python, JavaScript/TypeScript, Java, and C/C++.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/c-cpp.md`, `references/java.md` and `references/javascript.md`).

It sits in Development, covering Debugging and Failing and flaky tests. It works with Java, JavaScript, Python and C++. The repository describes itself as: A curated list of 180+ useful Claude Skills for Software Engineering and resources for customizing AI for SE workflows. The licence is Apache-2.0.

When your agent uses it

  • Investigating crashes
  • Unexpected behavior
  • Performance issues

Example prompts

  • “/failure-oriented-instrumentation”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Analyze the Failure
  2. Determine Instrumentation Strategy
  3. Select Instrumentation Patterns
  4. Insert Instrumentation
  5. Run and Collect Data
  6. Analyze Results

What it can do on your machine

Read from SKILL.md and the folder at commit 4f38503. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are java, python and javascript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Failure Oriented Instrumentation loads about 2.1k tokens when it runs, and up to ~6.3k if it reads all its reference files. Until then it costs about 122 tokens; SKILL.md has 672 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~122
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ArabelaTso/Skills-4-SE at commit 4f38503, republished under its Apache-2.0 licence (© ArabelaTso). 672 words, ~2,122 tokens.

Download SKILL.mdSave it as .claude/skills/failure-oriented-instrumentation/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
failure-oriented-instrumentation
description
Selectively instruments code to capture runtime data for debugging failures and bugs. Use when investigating crashes, exceptions, unexpected behavior, test failures, or performance issues. Analyzes stack traces and error messages to identify suspicious code regions, then adds targeted logging, tracing, and assertions to capture variable values, execution paths, timing, and conditional branches. Supports Python, JavaScript/TypeScript, Java, and C/C++.

Failure-Oriented Instrumentation

Strategically instrument code to capture high-signal runtime data for debugging failures, focusing only on suspicious regions rather than comprehensive instrumentation.

Workflow

1. Analyze the Failure

Gather and analyze failure information:

  • Error message and exception type
  • Stack trace showing call chain
  • Failure location (file, line, function)
  • Reproduction steps or test case
  • Expected vs actual behavior

Identify suspicious code regions:

  • Functions in the stack trace
  • Code paths leading to the failure
  • Variables involved in the error
  • Conditional branches that may affect the outcome
2. Determine Instrumentation Strategy

Based on the failure type, choose instrumentation targets:

For crashes/exceptions:

  • Function entry/exit in stack trace
  • Variable values before the crash
  • Conditional branches leading to error path
  • Exception handling blocks

For incorrect results:

  • Variable values at key computation points
  • Conditional branch decisions
  • Loop iterations and state changes
  • Function return values

For performance issues:

  • Timing information for slow operations
  • Loop iteration counts
  • Resource allocation/deallocation
  • Function call frequency

For intermittent failures:

  • State variables that may cause non-determinism
  • Thread/concurrency information
  • External dependencies (I/O, network, time)
  • Retry logic and error recovery paths
3. Select Instrumentation Patterns

Choose appropriate patterns based on language and context. See language-specific references:

Common patterns:

  • Function entry/exit logging
  • Variable value tracking
  • Conditional branch tracking
  • Loop iteration monitoring
  • Timing measurements
  • Assertions for invariants
4. Insert Instrumentation

Apply instrumentation to identified code regions:

Minimal approach (start here):

  • Instrument only the immediate failure location
  • Add 2-3 key variable logs
  • Track the critical conditional branch

Expanded approach (if minimal is insufficient):

  • Instrument entire call chain from stack trace
  • Add comprehensive variable tracking
  • Monitor all branches and loops in suspicious functions

Principles:

  • Start minimal, expand as needed
  • Focus on high-signal data (variables that affect control flow)
  • Avoid instrumenting stable, well-tested code
  • Minimize performance overhead
5. Run and Collect Data

Execute the instrumented code:

  • Run the failing test case or reproduction steps
  • Capture all instrumentation output (logs, traces)
  • Ensure instrumentation doesn't change behavior (except performance)
6. Analyze Results

Review captured data to identify root cause:

  • Compare variable values against expectations
  • Identify which branch was taken and why
  • Look for unexpected state transitions
  • Check timing for performance issues
  • Correlate multiple data points

Quick Start Examples

Example 1: NullPointerException in Java

Failure:

NullPointerException at UserService.java:45
  at UserService.processUser(UserService.java:45)
  at UserController.handleRequest(UserController.java:23)

Instrumentation:

java
public void processUser(String userId) {
    logger.debug("ENTER processUser: userId={}", userId);

    User user = userRepository.findById(userId);
    logger.debug("Retrieved user: {}", user);  // Check if null

    if (user == null) {
        logger.warn("User not found for userId={}", userId);
        return;
    }

    String email = user.getEmail();  // Line 45 - was failing here
    logger.debug("User email: {}", email);

    sendNotification(email);
}
Example 2: Incorrect Calculation in Python

Failure:

AssertionError: Expected 100, got 95
  at test_calculate_total (test_billing.py:12)
  at calculate_total (billing.py:34)

Instrumentation:

python
def calculate_total(items, discount_rate):
    logger.debug(f"ENTER calculate_total: items={items}, discount_rate={discount_rate}")

    subtotal = sum(item.price for item in items)
    logger.debug(f"Subtotal: {subtotal}")

    if discount_rate > 0:
        logger.debug(f"Applying discount: rate={discount_rate}")
        discount = subtotal * discount_rate
        logger.debug(f"Discount amount: {discount}")
    else:
        logger.debug("No discount applied")
        discount = 0

    total = subtotal - discount
    logger.debug(f"Final total: {total}")

    return total
Example 3: Intermittent Test Failure in JavaScript

Failure:

Test "should process async data" fails randomly
Expected: data processed
Actual: timeout

Instrumentation:

javascript
async function processAsyncData(dataId) {
    console.log(`ENTER processAsyncData: dataId=${dataId}, time=${Date.now()}`);

    const data = await fetchData(dataId);
    console.log(`Fetched data: ${JSON.stringify(data)}, time=${Date.now()}`);

    if (!data) {
        console.warn(`No data returned for dataId=${dataId}`);
        return null;
    }

    const processed = await processData(data);
    console.log(`Processed data: ${JSON.stringify(processed)}, time=${Date.now()}`);

    return processed;
}

Instrumentation Guidelines

Show full SKILL.md (284 more words)Show less
What to Instrument

High priority:

  • Functions in the stack trace
  • Variables mentioned in error messages
  • Conditional branches near the failure
  • Loop conditions and iteration variables
  • Function parameters and return values

Medium priority:

  • State variables that affect control flow
  • Resource allocations (memory, files, connections)
  • External dependencies (API calls, database queries)
  • Error handling and recovery logic

Low priority:

  • Stable utility functions
  • Simple getters/setters
  • Well-tested library code
  • Performance-critical hot paths (unless investigating performance)
What NOT to Instrument
  • Code unrelated to the failure
  • Third-party libraries (unless suspected)
  • Trivial operations (assignments, simple math)
  • Code that would generate excessive output
  • Security-sensitive operations (passwords, tokens)
Instrumentation Best Practices
  1. Use appropriate log levels: DEBUG for detailed tracing, INFO for key events, ERROR for exceptions
  2. Include context: Variable names, values, types, and relevant state
  3. Mark entry/exit: Clear boundaries for function execution
  4. Timestamp when relevant: For timing and ordering issues
  5. Avoid side effects: Instrumentation should not change program behavior
  6. Clean up after: Remove temporary instrumentation once bug is fixed
  7. Consider permanence: Some instrumentation may be valuable long-term for observability

Language-Specific References

Load these references for detailed instrumentation patterns:

Tips

  • Start small: Instrument one function, run, analyze, then expand if needed
  • Be specific: Log exact variable values, not just "processing data"
  • Use structured logging: Include variable names and context, not just values
  • Test instrumentation: Verify logs appear and contain expected information
  • Iterate quickly: Add instrumentation, run, analyze, repeat
  • Consider performance: For production code, use conditional compilation or feature flags
  • Document findings: Comment why instrumentation was added and what it revealed

© ArabelaTso, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in skills/failure-oriented-instrumentation of ArabelaTso/Skills-4-SE.

  • SKILL.md
  • references/c-cpp.md
  • references/java.md
  • references/javascript.md
  • references/python.md

Open the folder on GitHubat commit 4f38503

Compare with similar skills

Failure Oriented Instrumentation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Failure Oriented Instrumentation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Failure Oriented Instrumentation this skillArabelaTso/Skills-4-SE253—~2.1kAutomated safety check: PassApache-2.0
MCP Debuggerdebugmcp/mcp-debugger171—~3.8kAutomated safety check: PassMIT
Dbgtheodo-group/debug-that158—~2.2kAutomated safety check: PassMIT
CodeQL Security Scantrailofbits/skills7.4k—~4.6kAutomated safety check: NotesCC-BY-SA-4.0
Fory Performance Optimizationapache/fory4.6k—~2.2kAutomated safety check: PassApache-2.0
Cross-Language Coding Standardszereight/gitlab-mcp2k1 repos~1.4kAutomated safety check: PassMIT

Similar skills

  • MCP Debugger

    debugmcp/mcp-debugger

    A skill your agent uses when investigating a bug, failing test, or unexpected runtime behavior and the mcp-debugger MCP server is available — drives real step-through debuggers (breakpoints, stack…

    171 GitHub stars~3.8k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Dbg

    theodo-group/debug-that

    Debug applications using the dbg CLI debugger. An agent skill from theodo-group/debug-that.

    158 GitHub stars~2.2k tokensUpdated today
    DevelopmentAuto-check passed
  • CodeQL Security Scan

    trailofbits/skills

    Official

    Scans a codebase for vulnerabilities with CodeQL's data flow and taint tracking in run-all or important-only modes, including data extensions for project-specific sources and sinks.

    7.4k GitHub stars~4.6k tokensUpdated yesterday
    SecurityAuto-check: notes
  • Run profile-driven bottleneck optimization across Apache Fory implementations (Java, C++, Python/Cython, Go, Rust, Swift, C, JavaScript/TypeScript, Dart, Kotlin, Scala).

    4.6k GitHub stars~2.2k tokensUpdated yesterday
    MobileAuto-check passed
  • Shared reference for naming, function size, complexity and error handling rules that reviewer agents apply across TypeScript, Python, Go, Rust, Java, C# and Swift.

    2k GitHub starsUsed in 1 repo~1.4k tokens
    DevelopmentAuto-check passed
  • Supercov

    supercorp-ai/supercov

    Measures test coverage and code quality in a repository with the supercov CLI, and turns what it finds into small, focused tests or fixes.

    150 GitHub starsUsed in 1 repo~415 tokens
    Testing & QAAuto-check passed

More from ArabelaTso/Skills-4-SE

All 150 skills in this repo
  • Framework Migration Assistant

    ArabelaTso/Skills-4-SE

    Automatically migrate Python web applications between frameworks (Flask → FastAPI, Django → FastAPI).

    253 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Metamorphic Test Generator

    ArabelaTso/Skills-4-SE

    Generate test cases using metamorphic testing by applying transformations based on metamorphic properties.

    253 GitHub stars~798 tokensUpdated 1 mo ago
    Auto-check passed
  • Reproduction Trace Instrumenter

    ArabelaTso/Skills-4-SE

    Instruments programs to capture execution traces specifically for reproducing reported bugs, enabling consistent replay and diagnosis of failures.

    253 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Spring Mvc To Boot Migrator

    ArabelaTso/Skills-4-SE

    Automatically migrate Spring MVC applications to Spring Boot.

    253 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed
  • State Snapshot Instrumenter

    ArabelaTso/Skills-4-SE

    Instrument programs (Python, C/C++, Java) to capture snapshots of key program states at runtime, including variables, memory, and call stacks.

    253 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Failure Oriented Instrumentation

What does Failure Oriented Instrumentation do?

Selectively instruments code to capture runtime data for debugging failures and bugs. Failure Oriented Instrumentation is an agent skill from ArabelaTso/Skills-4-SE. Selectively instruments code to capture runtime data for debugging failures and bugs.

When should I use Failure Oriented Instrumentation?

Failure Oriented Instrumentation fits situations like: investigating crashes; unexpected behavior; performance issues.

How do I install Failure Oriented Instrumentation in Claude Code?

Run `npx skills add ArabelaTso/Skills-4-SE --skill failure-oriented-instrumentation -a claude-code`. Or copy the skill folder (skills/failure-oriented-instrumentation in ArabelaTso/Skills-4-SE) into .claude/skills/failure-oriented-instrumentation in your project. Claude Code loads it when a task matches its description.

How do I install Failure Oriented Instrumentation in Codex?

Run `npx skills add ArabelaTso/Skills-4-SE --skill failure-oriented-instrumentation -a codex`. Or copy the skill folder (skills/failure-oriented-instrumentation in ArabelaTso/Skills-4-SE) into .agents/skills/failure-oriented-instrumentation in your project. Codex loads it when a task matches its description.

Can I use Failure Oriented Instrumentation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ArabelaTso/Skills-4-SE --skill failure-oriented-instrumentation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/failure-oriented-instrumentation, .gemini/skills/failure-oriented-instrumentation, .github/skills/failure-oriented-instrumentation and .opencode/skills/failure-oriented-instrumentation in your project.

What does Failure Oriented Instrumentation need to run?

SKILL.md names no scripts, command-line tools or credentials: Failure Oriented Instrumentation is instructions for the agent only. Our summary lists: Python 3.

Does Failure Oriented Instrumentation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Failure Oriented Instrumentation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Failure Oriented Instrumentation use?

Failure Oriented Instrumentation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Failure Oriented Instrumentation use?

About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.2k tokens, read only when the agent opens those files.

What are the alternatives to Failure Oriented Instrumentation?

Skills that share tags, products or a category with Failure Oriented Instrumentation: MCP Debugger (debugmcp/mcp-debugger, 171 stars), Dbg (theodo-group/debug-that, 158 stars), CodeQL Security Scan (trailofbits/skills, 7.4k stars) and Fory Performance Optimization (apache/fory, 4.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Failure Oriented Instrumentation?

ArabelaTso (a GitHub user) maintains it in ArabelaTso/Skills-4-SE, which has 253 GitHub stars. The repository holds 150 skills in this directory. The repository was last updated on August 21, 2026.

Source: ArabelaTso/Skills-4-SE on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.