Agent skill

Debug Diff

by afnanenayet in afnanenayet/diffsitter

Diagnose unexpected diff output by tracing the pipeline from parsing through AST processing to hunk generation.

MITAuto-check: notesDevelopment

Install Debug Diff

skills CLI
$ npx skills add afnanenayet/diffsitter --skill debug-diff -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install afnanenayet/diffsitter debug-diff --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/afnanenayet/diffsitter.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/debug-diff .claude/skills/debug-diff && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
debug-diff
GitHub stars
2.4k
Token cost
~2.4k tokens
SKILL.md length
1,066 words
Files
1
Skills in repo
4
Repo updated
First seen
Licence
MIT

At a glance

Diagnose unexpected diff output by tracing the pipeline from parsing through AST processing to hunk generation.

  • Works in 8 steps: Language Detection → Tree-sitter Parse → AST Leaf Extraction → …
  • Diffsitter produces wrong
  • SKILL.md covers Pipeline Overview, Stage-by-Stage Debugging, Quick Diagnostic Commands and Common Failure Modes Summary
  • Calls jq

What it does

Debug Diff is an agent skill from afnanenayet/diffsitter. Diagnose unexpected diff output by tracing the pipeline from parsing through AST processing to hunk generation. Use when diffsitter produces wrong or surprising results.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development. The repository describes itself as: A tree-sitter based AST difftool to get meaningful semantic diffs. The licence is MIT.

When your agent uses it

  • Diffsitter produces wrong
  • Surprising results

Example prompts

  • “/debug-diff”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Bash

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Language Detection
  2. Tree-sitter Parse
  3. AST Leaf Extraction
  4. Node Filtering
  5. Grapheme Splitting and Whitespace Stripping
  6. Myers Diff
  7. Hunk Assembly
  8. Rendering

What it can do on your machine

Read from SKILL.md and the folder at commit a5a6681. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Debug Diff loads about 2.4k tokens when it runs. Until then it costs about 45 tokens; SKILL.md has 1,066 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~45
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Grep, Glob, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from afnanenayet/diffsitter at commit a5a6681, republished under its MIT licence (© afnanenayet). 1,066 words, ~2,360 tokens.

Download SKILL.mdSave it as .claude/skills/debug-diff/SKILL.md (or your agent's skills folder).
name
debug-diff
description
Diagnose unexpected diff output by tracing the pipeline from parsing through AST processing to hunk generation. Use when diffsitter produces wrong or surprising results.
allowed-tools
Read, Grep, Glob, Bash
user-invocable
true
argument-hint
[description] What unexpected behavior are you seeing?

Debugging Unexpected Diff Output in diffsitter

When diffsitter produces wrong or surprising results, trace the issue through the pipeline. The diff pipeline has 7 stages, and problems can originate at any of them.

Ask the user for the following information if not already provided:

  • The two files being compared (or representative snippets)
  • The language/file extension
  • Any custom config being used
  • What output they expected vs what they got

Pipeline Overview

The full pipeline (see src/bin/diffsitter.rs, function run_diff):

1. Language detection
2. Tree-sitter parse (file -> AST Tree)
3. AST leaf extraction (Tree -> Vec<VectorLeaf>)
4. Node filtering (exclude_kinds / include_kinds)
5. Grapheme splitting + whitespace stripping (VectorLeaf -> Vec<Entry>)
6. Myers diff (Vec<Entry> x 2 -> Vec<EditType>)
7. Hunk assembly (Vec<EditType> -> RichHunks)
8. Rendering (RichHunks -> terminal output)

Stage-by-Stage Debugging

Stage 1: Language Detection

File: src/parse.rs, function resolve_language_str

The language is resolved from the file extension via the FILE_EXTS phf_map, with optional user overrides from GrammarConfig.file_associations.

Common issues:

  • Extension not mapped (e.g., .jsx maps to "tsx", .h maps to "c" not "cpp")
  • User override in config shadowing the default

Debug: Check what language is resolved:

sh
# Enable debug logging to see language resolution
RUST_LOG=debug diffsitter old_file new_file 2>&1 | grep -i "deduced language"

Or read FILE_EXTS in src/parse.rs to verify the extension mapping.

Stage 2: Tree-sitter Parse

File: src/parse.rs, function parse_file

Creates a Parser, sets the language, reads the file to a string, and calls parser.parse(&text, None).

Common issues:

  • Grammar ABI version mismatch (AbiOutOfRange error) -- the grammar was compiled against an incompatible tree-sitter version.
  • Parse failure returns None from parser.parse(), resulting in a TSParseFailure error.
  • File encoding issues (tree-sitter expects UTF-8).

Debug: If parsing succeeds but the tree looks wrong, the grammar itself may have a bug for that language construct. Check tree-sitter's own playground or CLI to inspect the AST.

Stage 3: AST Leaf Extraction

File: src/input_processing.rs, function build

Recursively walks the tree via node.children(&mut cursor). Collects leaf nodes (nodes with child_count() == 0) OR nodes whose kind() matches a pseudo_leaf_types entry.

Key behaviors:

  • Empty byte ranges are skipped (node.byte_range().is_empty())
  • Nodes that are pure newlines (after removing \n, \r, \r\n) are skipped -- this is a workaround for the Go parser
  • Pseudo-leaf types (configured in input_processing.pseudo_leaf_types) treat certain non-leaf nodes as leaves. Default: {"markdown": {"inline"}}. This is critical for text-heavy documents.

Common issues:

  • Missing diffs in markdown/prose: Check if pseudo_leaf_types includes the right node types for that language. Without "inline" for markdown, large text blocks are treated as single atoms.
  • Unexpected nodes included: The grammar may expose more leaf nodes than expected (e.g., punctuation, delimiters).
Stage 4: Node Filtering

File: src/input_processing.rs, method TreeSitterProcessor::should_include_node

Filters nodes based on exclude_kinds and include_kinds from the config:

  • exclude_kinds takes precedence: if a node's kind() is in this set, it is excluded.
  • include_kinds: if set, only nodes whose kind() is in this set are included (unless also excluded).
  • If neither is set, all nodes pass through.

Common issues:

  • User config has exclude_kinds or include_kinds that filters out relevant nodes.
  • The node kind() string doesn't match what the user expects (tree-sitter kind names are grammar-specific).

Debug: Check what node kinds exist for a language using tree-sitter's node types. The kind() strings come from the grammar definition.

Stage 5: Grapheme Splitting and Whitespace Stripping

File: src/input_processing.rs, method VectorLeaf::split_on_graphemes

If split_graphemes is enabled (default: true), each leaf's text is split into individual Unicode graphemes, each becoming its own Entry with precise row/column positions. If strip_whitespace is enabled (default: true), whitespace-only graphemes are skipped.

Common issues:

  • Whitespace-only diffs not showing: strip_whitespace: true (the default) means pure whitespace/indentation changes are invisible. This is by design for AST-based diffing. If the user wants to see whitespace changes, they need "strip-whitespace": false in their config.
  • Performance with large files: split_graphemes: true generates many entries for large text nodes. Setting "split-graphemes": false trades granularity for speed.
  • Line position bugs: The grapheme splitter tracks row/column positions. If line_offset == 0, it offsets from the node's start_position().column. Otherwise it resets the column to idx (the byte offset within the line). Bugs here manifest as incorrect column highlighting in the rendered output.
Show full SKILL.md (461 more words)Show less
Stage 6: Myers Diff

File: src/diff.rs, struct Myers, method diff

Implements the classic Myers "An O(ND) Difference Algorithm" with divide-and-conquer via middle snake. Takes two &[Entry] slices, produces Vec<EditType<&Entry>> where EditType is either Addition or Deletion.

Key optimizations:

  • Common prefix/suffix are skipped before running the main algorithm (via common_prefix_len / common_suffix_len which use get_unchecked for speed).
  • Entry equality is based on kind_id AND text (see the PartialEq impl) -- two nodes are equal only if they have the same tree-sitter kind AND identical text content.

Common issues:

  • Too many diffs reported: If entries that should be equal are not, check if kind_id differs between them. Two nodes with the same text but different grammar kinds (e.g., identifier vs type_identifier in Rust) are considered different.
  • No diffs reported when expected: If entries that should differ are comparing as equal, verify the text content after whitespace stripping.
Stage 7: Hunk Assembly

File: src/diff.rs, struct RichHunksBuilder

Converts the flat edit script into grouped hunks. Each EditType::Addition becomes DocumentType::New, each EditType::Deletion becomes DocumentType::Old. Consecutive edits on adjacent lines are grouped into the same Hunk. Non-adjacent edits start a new hunk.

Common issues:

  • PriorLine or PriorColumn errors: These indicate the edit script produced entries in non-ascending order, which is a bug in the diff or input processing stage.
  • Hunks splitting unexpectedly: If edits on adjacent lines end up in separate hunks, there may be a gap in line numbers caused by filtered-out nodes.
Stage 8: Rendering

File: src/render/unified.rs (for unified renderer), src/render/json.rs (for JSON)

Common issues with unified renderer:

  • Line index out of bounds: The print_hunk method accesses lines[line_index]. If line_index >= lines.len(), it logs an error and skips the line (in release) or panics via debug_assert! (in debug).
  • Column range panics: print_line indexes into the line text using byte column ranges from entries. If the entry positions don't align with the actual text bytes, this can panic with a slice bounds error.
  • Missing terminal colors: If output is piped, term_info may be None. The unified renderer still works but won't have terminal width info for title formatting.

Quick Diagnostic Commands

sh
# Run with full debug logging
RUST_LOG=trace diffsitter old_file new_file 2>debug.log

# Output as JSON to inspect raw diff data
diffsitter --renderer json old_file new_file | jq .

# Force a specific language
diffsitter --file-type python old_file new_file

# Run without config to eliminate config issues
diffsitter --no-config old_file new_file

Common Failure Modes Summary

SymptomLikely StageCheck
"Unsupported extension" error1 (Language detection)Extension in FILE_EXTS?
Parse error / empty AST2 (Tree-sitter parse)Grammar ABI compatible? File is valid UTF-8?
No diff when content clearly changed5 (Whitespace stripping)strip_whitespace removing the changes?
Missing diffs in markdown/prose3 (Leaf extraction)pseudo_leaf_types configured for language?
Diffs include too many trivial nodes4 (Node filtering)Set exclude_kinds for noise (e.g., "comment", "string")
Wrong columns highlighted5 (Grapheme splitting)Grapheme position calculation bug
Panic on "index out of bounds"7/8 (Hunk assembly / Rendering)Line index mismatch between entry positions and actual text
Extremely slow on large files5/6 (Graphemes / Myers)Try "split-graphemes": false or check if diff is O(ND) worst-case

© afnanenayet, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/debug-diff of afnanenayet/diffsitter.

Open the folder on GitHubat commit a5a6681

Compare with similar skills

Debug Diff next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Debug Diff compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Debug Diff this skillafnanenayet/diffsitter2.4k—~2.4kAutomated safety check: NotesMIT
Vercel Composition Patternssupabase/supabase111k58 repos~726Automated safety check: PassMIT
Finishing a Development Branchobra/superpowers297k5 repos~1.9kAutomated safety check: PassMIT
Typescript Advanced Typesrolling-scopes/rsschool-app10k25 repos~4.2kAutomated safety check: PassMPL-2.0
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Code Review ChecklistshareAI-lab/learn-claude-code78k5 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • Official

    React composition patterns that scale. An agent skill from supabase/supabase.

    111k GitHub starsUsed in 58 repos~726 tokens
    DevelopmentAuto-check passed
  • Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.

    297k GitHub starsUsed in 5 repos~1.9k tokens
    DevelopmentAuto-check passed
  • Typescript Advanced Types

    rolling-scopes/rsschool-app

    Master TypeScript's advanced type system including generics, conditional types, mapped types, template literals, and utility types for building type-safe applications.

    10k GitHub starsUsed in 25 repos~4.2k tokens
    DevelopmentAuto-check passed
  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 5 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Greploop

    onyx-dot-app/onyx

    Iteratively improves a PR (GitHub), MR (GitLab), or shelved changelist (Perforce) until Greptile gives it a 5/5 confidence score with zero unresolved comments.

    32k GitHub starsUsed in 4 repos~3.3k tokens
    DevelopmentAuto-check passed

More from afnanenayet/diffsitter

  • Add Grammar

    afnanenayet/diffsitter

    Step-by-step guide for adding a new tree-sitter language grammar to diffsitter.

    2.4k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check: notes
  • Add Renderer

    afnanenayet/diffsitter

    Step-by-step guide for adding a new diff output renderer to diffsitter.

    2.4k GitHub stars~1.4k tokensUpdated yesterday
    Auto-check: notes
  • Rust Expert

    afnanenayet/diffsitter

    Expert Rust guidance for diffsitter: tree-sitter FFI patterns, lifetime management, unsafe optimization, edition 2024 idioms.

    2.4k GitHub stars~2.4k tokensUpdated yesterday
    Auto-check: notes

Categories

Questions about Debug Diff

What does Debug Diff do?

Diagnose unexpected diff output by tracing the pipeline from parsing through AST processing to hunk generation. Debug Diff is an agent skill from afnanenayet/diffsitter. Diagnose unexpected diff output by tracing the pipeline from parsing through AST processing to hunk generation.

When should I use Debug Diff?

Debug Diff fits situations like: diffsitter produces wrong; surprising results.

How do I install Debug Diff in Claude Code?

Run `npx skills add afnanenayet/diffsitter --skill debug-diff -a claude-code`. Or copy the skill folder (.claude/skills/debug-diff in afnanenayet/diffsitter) into .claude/skills/debug-diff in your project. Claude Code loads it when a task matches its description.

How do I install Debug Diff in Codex?

Run `npx skills add afnanenayet/diffsitter --skill debug-diff -a codex`. Or copy the skill folder (.claude/skills/debug-diff in afnanenayet/diffsitter) into .agents/skills/debug-diff in your project. Codex loads it when a task matches its description.

Can I use Debug Diff in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add afnanenayet/diffsitter --skill debug-diff -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debug-diff, .gemini/skills/debug-diff, .github/skills/debug-diff and .opencode/skills/debug-diff in your project.

What does Debug Diff need to run?

Going by SKILL.md and its folder, Debug Diff needs the command-line tools its instructions call (jq). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash.

Does Debug Diff access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Debug Diff safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Debug Diff use?

Debug Diff is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Debug Diff use?

About 2.4k tokens (SKILL.md is roughly 9.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Debug Diff?

Skills that share tags, products or a category with Debug Diff: Vercel Composition Patterns (supabase/supabase, 111k stars), Finishing a Development Branch (obra/superpowers, 297k stars), Typescript Advanced Types (rolling-scopes/rsschool-app, 10k stars) and PR Babysitter (openinterpreter/openinterpreter, 69k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Debug Diff?

afnanenayet (a GitHub user) maintains it in afnanenayet/diffsitter, which has 2,405 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 8, 2026.

Source: afnanenayet/diffsitter on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.