Agent skill

Counterexample Debugger

by ArabelaTso in ArabelaTso/Skills-4-SE

Debug proof failures using counterexamples from Nitpick (Isabelle) or QuickChick (Coq) to identify specification errors, missing preconditions, and proof strategy issues.

Apache-2.0Auto-check passedDevelopment

Install Counterexample Debugger

skills CLI
$ npx skills add ArabelaTso/Skills-4-SE --skill counterexample-debugger -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ArabelaTso/Skills-4-SE counterexample-debugger --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ArabelaTso/Skills-4-SE.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/counterexample-debugger .claude/skills/counterexample-debugger && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
counterexample-debugger
GitHub stars
253
Token cost
~1.9k tokens
SKILL.md length
763 words
Files
4 (incl. references)
Skills in repo
151
Repo updated
First seen
Licence
Apache-2.0

At a glance

Debug proof failures using counterexamples from Nitpick (Isabelle) or QuickChick (Coq) to identify specification errors, missing preconditions, and proof strategy issues.

  • Works in 6 steps: Receive Counterexample Information → Choose Target System → Analyze the Counterexample → …
  • A proof attempt fails and you need to understand why
  • SKILL.md covers Workflow, Counterexample Analysis Patterns, Tool-Specific Guidance and Common Root Causes, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Counterexample Debugger is an agent skill from ArabelaTso/Skills-4-SE. Debug proof failures using counterexamples from Nitpick (Isabelle) or QuickChick (Coq) to identify specification errors, missing preconditions, and proof strategy issues. Use when: (1) A proof attempt fails and you need to understand why, (2) Counterexamples are generated by Nitpick or QuickChick, (3) Specifications may be incorrect or incomplete, (4) Theorems need validation before proving, (5) Missing preconditions or lemmas need identification, or (6) Proof failures need explanation and correction suggestions…

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/examples.md`, `references/nitpick_guide.md` and `references/quickchick_guide.md`).

It sits in Development. The repository describes itself as: A curated list of 180+ useful Claude Skills for Software Engineering and resources for customizing AI for SE workflows. The licence is Apache-2.0.

When your agent uses it

  • A proof attempt fails and you need to understand why
  • Counterexamples are generated by Nitpick
  • Specifications may be incorrect
  • Theorems need validation before proving

Example prompts

  • “/counterexample-debugger”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Receive Counterexample Information
  2. Choose Target System
  3. Analyze the Counterexample
  4. Explain the Failure
  5. Suggest Corrections
  6. Recommend Next Steps

What it can do on your machine

Read from SKILL.md and the folder at commit 4f38503. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are isabelle and coq).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Counterexample Debugger loads about 1.9k tokens when it runs, and up to ~7.5k if it reads all its reference files. Until then it costs about 147 tokens; SKILL.md has 763 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~147
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ArabelaTso/Skills-4-SE at commit 4f38503, republished under its Apache-2.0 licence (© ArabelaTso). 763 words, ~1,889 tokens.

Download SKILL.mdSave it as .claude/skills/counterexample-debugger/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
counterexample-debugger
description
Debug proof failures using counterexamples from Nitpick (Isabelle) or QuickChick (Coq) to identify specification errors, missing preconditions, and proof strategy issues. Use when: (1) A proof attempt fails and you need to understand why, (2) Counterexamples are generated by Nitpick or QuickChick, (3) Specifications may be incorrect or incomplete, (4) Theorems need validation before proving, (5) Missing preconditions or lemmas need identification, or (6) Proof failures need explanation and correction suggestions. Supports both Isabelle/HOL and Coq equally.

Counterexample-Guided Proof Debugger

Analyze counterexamples from Nitpick or QuickChick to explain proof failures and suggest corrections to specifications or proofs.

Workflow

1. Receive Counterexample Information

Identify what information is provided:

  • Counterexample output: From Nitpick or QuickChick
  • Failed theorem: The statement that couldn't be proven
  • Proof attempt: Any partial proof or tactics tried
  • Context: Definitions and lemmas involved
2. Choose Target System

Determine which proof assistant is being used:

  • Isabelle/HOL with Nitpick: Finite model finder
  • Coq with QuickChick: Property-based testing
  • Both: Provide analysis for both systems
3. Analyze the Counterexample

Examine the counterexample systematically:

Verify the counterexample:

  • Manually compute the result for the given values
  • Confirm it actually violates the theorem
  • Check if it's a genuine counterexample or tool limitation

Identify the violation:

  • Which part of the theorem fails?
  • What values cause the failure?
  • Is it an edge case or fundamental issue?

Determine the root cause:

  • Missing precondition?
  • Incorrect specification?
  • Wrong quantifier order?
  • Implementation bug?
  • Off-by-one error?
  • Type constraint issue?
4. Explain the Failure

Provide clear explanation:

What went wrong:

  • Describe why the counterexample violates the theorem
  • Show the computation step-by-step
  • Highlight the specific point of failure

Why it happened:

  • Explain the underlying cause
  • Identify the conceptual error
  • Note any common patterns (empty list, boundary values, etc.)

Impact assessment:

  • Is the theorem fundamentally wrong?
  • Does it need preconditions?
  • Is the specification incomplete?
5. Suggest Corrections

Provide actionable fixes based on the root cause:

For missing preconditions:

isabelle
(* Before *)
lemma "hd xs ∈ set xs"

(* After *)
lemma "xs ≠ [] ⟹ hd xs ∈ set xs"

For incorrect specifications:

coq
(* Before: uses < instead of <= *)
x < y && is_sorted (y :: ys)

(* After *)
x <= y && is_sorted (y :: ys)

For quantifier order:

isabelle
(* Before *)
"∃y. ∀x. P x y"

(* After *)
"∀x. ∃y. P x y"

For incomplete specifications:

coq
(* Before: only checks sortedness *)
is_sorted (sort l)

(* After: also checks permutation *)
is_sorted (sort l) && permutation l (sort l)
6. Recommend Next Steps

Guide the user on what to do:

Retest with fix:

  • Run Nitpick/QuickChick again
  • Verify no counterexample found
  • Check if fix is sufficient

Identify additional issues:

  • Are there other edge cases?
  • Do other lemmas need fixing?
  • Is the specification now complete?

Proceed with proof:

  • If no counterexample, attempt proof
  • Suggest proof strategy
  • Identify needed helper lemmas

Counterexample Analysis Patterns

Pattern 1: Empty Structures

Symptom: Counterexample is [], {}, or None

Common causes:

  • Missing non-empty precondition
  • Undefined behavior on empty input
  • Base case not handled

Fix: Add precondition or handle empty case explicitly

Pattern 2: Boundary Values

Symptom: Counterexample is 0, 1, or type limits

Common causes:

  • Off-by-one errors
  • Boundary condition not considered
  • Edge case in arithmetic

Fix: Adjust bounds or add special case handling

Pattern 3: Duplicate Elements

Symptom: Counterexample has repeated values like [0, 0]

Common causes:

  • Using < instead of ≤
  • Assuming distinctness
  • Set vs. multiset confusion

Fix: Use appropriate comparison or add distinctness assumption

Pattern 4: Small Counterexamples

Symptom: Very small counterexample (1-2 elements)

Common causes:

  • Fundamental specification error
  • Wrong base case
  • Incorrect inductive step

Fix: Review base definitions and inductive structure

Pattern 5: Type-Specific Values

Symptom: Counterexample at type boundaries

Common causes:

  • Type constraints not considered
  • Overflow/underflow issues
  • Finite vs. infinite types

Fix: Add type constraints or adjust specification

Tool-Specific Guidance

Show full SKILL.md (305 more words)Show less
Nitpick (Isabelle/HOL)

For detailed Nitpick usage and interpretation:

Key points:

  • Searches for finite models
  • Configurable cardinality bounds
  • May miss counterexamples beyond bounds
  • "No counterexample" ≠ proof
QuickChick (Coq)

For detailed QuickChick usage and interpretation:

Key points:

  • Random testing with shrinking
  • Configurable test count
  • May miss rare counterexamples
  • "Success" ≠ proof

Common Root Causes

Specification Errors

Symptoms:

  • Counterexample shows spec doesn't match intent
  • Multiple unrelated counterexamples
  • Spec too weak or too strong

Fixes:

  • Strengthen postconditions
  • Add completeness requirements
  • Review specification against intent
Missing Preconditions

Symptoms:

  • Counterexample is edge case
  • Empty structures or boundary values
  • Undefined behavior

Fixes:

  • Add non-empty constraints
  • Add type bounds
  • Add well-formedness conditions
Quantifier Issues

Symptoms:

  • Counterexample shows wrong order
  • Existential/universal confusion
  • Skolem constant issues

Fixes:

  • Swap quantifier order
  • Review logical structure
  • Check variable dependencies
Implementation Bugs

Symptoms:

  • Counterexample shows function doesn't work
  • Output doesn't match specification
  • Logic error in definition

Fixes:

  • Fix the implementation
  • Review algorithm correctness
  • Test implementation separately

Debugging Checklist

When analyzing a counterexample:

  1. Verify manually: Compute the result for the counterexample values
  2. Identify violation: Which part of the theorem fails?
  3. Find root cause: Why does it fail? (precondition, spec, impl, quantifiers)
  4. Suggest fix: What needs to change?
  5. Check completeness: Are there other issues?
  6. Retest: Run tool again after fix
  7. Proceed: If no counterexample, attempt proof

Examples

For complete debugging examples including:

  • Incorrect sortedness definition
  • Missing preconditions
  • Wrong quantifier order
  • Off-by-one errors
  • Incomplete specifications

See examples.md

Tips

  • Trust the counterexample: If tool finds one, investigate thoroughly
  • Verify manually: Always check the counterexample by hand
  • Look for patterns: Empty, boundary, duplicates are common
  • Fix root cause: Don't just patch symptoms
  • Retest after fixing: Ensure fix is complete
  • Use tools early: Run Nitpick/QuickChick before proving
  • Iterate: May need multiple rounds of debugging
  • Document assumptions: Make preconditions explicit
  • Check completeness: Ensure specification is sufficient

© ArabelaTso, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/counterexample-debugger of ArabelaTso/Skills-4-SE.

  • SKILL.md
  • references/examples.md
  • references/nitpick_guide.md
  • references/quickchick_guide.md

Open the folder on GitHubat commit 4f38503

Compare with similar skills

Counterexample Debugger next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Counterexample Debugger compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Counterexample Debugger this skillArabelaTso/Skills-4-SE253—~1.9kAutomated safety check: PassApache-2.0
Finishing a Development Branchobra/superpowers296k5 repos~1.9kAutomated safety check: PassMIT
Typescript Advanced Typesrolling-scopes/rsschool-app10k24 repos~4.2kAutomated safety check: PassMPL-2.0
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Code Review ChecklistshareAI-lab/learn-claude-code78k5 repos~1.1kAutomated safety check: PassMIT
Greplooponyx-dot-app/onyx32k4 repos~3.3kAutomated safety check: PassMIT

Similar skills

  • Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.

    296k GitHub starsUsed in 5 repos~1.9k tokens
    DevelopmentAuto-check passed
  • Typescript Advanced Types

    rolling-scopes/rsschool-app

    Master TypeScript's advanced type system including generics, conditional types, mapped types, template literals, and utility types for building type-safe applications.

    10k GitHub starsUsed in 24 repos~4.2k tokens
    DevelopmentAuto-check passed
  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 5 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Greploop

    onyx-dot-app/onyx

    Iteratively improves a PR (GitHub), MR (GitLab), or shelved changelist (Perforce) until Greptile gives it a 5/5 confidence score with zero unresolved comments.

    32k GitHub starsUsed in 4 repos~3.3k tokens
    DevelopmentAuto-check passed
  • Guidelines

    akash-network/node

    Behavioral guidelines to reduce common LLM coding mistakes. An agent skill from akash-network/node.

    1.1k GitHub starsUsed in 22 repos~577 tokens
    DevelopmentAuto-check passed

More from ArabelaTso/Skills-4-SE

All 151 skills in this repo
  • Framework Migration Assistant

    ArabelaTso/Skills-4-SE

    Automatically migrate Python web applications between frameworks (Flask → FastAPI, Django → FastAPI).

    253 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Metamorphic Test Generator

    ArabelaTso/Skills-4-SE

    Generate test cases using metamorphic testing by applying transformations based on metamorphic properties.

    253 GitHub stars~798 tokensUpdated 1 mo ago
    Auto-check passed
  • Reproduction Trace Instrumenter

    ArabelaTso/Skills-4-SE

    Instruments programs to capture execution traces specifically for reproducing reported bugs, enabling consistent replay and diagnosis of failures.

    253 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Spring Mvc To Boot Migrator

    ArabelaTso/Skills-4-SE

    Automatically migrate Spring MVC applications to Spring Boot.

    253 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed
  • State Snapshot Instrumenter

    ArabelaTso/Skills-4-SE

    Instrument programs (Python, C/C++, Java) to capture snapshots of key program states at runtime, including variables, memory, and call stacks.

    253 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed

Categories

Questions about Counterexample Debugger

What does Counterexample Debugger do?

Debug proof failures using counterexamples from Nitpick (Isabelle) or QuickChick (Coq) to identify specification errors, missing preconditions, and proof strategy issues. Counterexample Debugger is an agent skill from ArabelaTso/Skills-4-SE. Debug proof failures using counterexamples from Nitpick (Isabelle) or QuickChick (Coq) to identify specification errors, missing preconditions, and proof strategy issues.

When should I use Counterexample Debugger?

Counterexample Debugger fits situations like: A proof attempt fails and you need to understand why; counterexamples are generated by Nitpick; specifications may be incorrect; theorems need validation before proving.

How do I install Counterexample Debugger in Claude Code?

Run `npx skills add ArabelaTso/Skills-4-SE --skill counterexample-debugger -a claude-code`. Or copy the skill folder (skills/counterexample-debugger in ArabelaTso/Skills-4-SE) into .claude/skills/counterexample-debugger in your project. Claude Code loads it when a task matches its description.

How do I install Counterexample Debugger in Codex?

Run `npx skills add ArabelaTso/Skills-4-SE --skill counterexample-debugger -a codex`. Or copy the skill folder (skills/counterexample-debugger in ArabelaTso/Skills-4-SE) into .agents/skills/counterexample-debugger in your project. Codex loads it when a task matches its description.

Can I use Counterexample Debugger in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ArabelaTso/Skills-4-SE --skill counterexample-debugger -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/counterexample-debugger, .gemini/skills/counterexample-debugger, .github/skills/counterexample-debugger and .opencode/skills/counterexample-debugger in your project.

What does Counterexample Debugger need to run?

SKILL.md names no scripts, command-line tools or credentials: Counterexample Debugger is instructions for the agent only.

Does Counterexample Debugger access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Counterexample Debugger safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Counterexample Debugger use?

Counterexample Debugger is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Counterexample Debugger use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.6k tokens, read only when the agent opens those files.

What are the alternatives to Counterexample Debugger?

Skills that share tags, products or a category with Counterexample Debugger: Finishing a Development Branch (obra/superpowers, 296k stars), Typescript Advanced Types (rolling-scopes/rsschool-app, 10k stars), PR Babysitter (openinterpreter/openinterpreter, 69k stars) and Code Review Checklist (shareAI-lab/learn-claude-code, 78k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Counterexample Debugger?

ArabelaTso (a GitHub user) maintains it in ArabelaTso/Skills-4-SE, which has 253 GitHub stars. The repository holds 151 skills in this directory. The repository was last updated on August 21, 2026.

Source: ArabelaTso/Skills-4-SE on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.