Agent skill

Lttng Tracing Root Cause Analysis

by isc-projects in isc-projects/bind9

Methodology for root-causing hard concurrency / memory-ordering bugs (intermittent races, use-after-free, RCU/lock-free publish-order defects, "impossible" stale reads) with LTTng flight-recorder…

MPL-2.0Auto-check passedDevelopment

Install Lttng Tracing Root Cause Analysis

skills CLI
$ npx skills add isc-projects/bind9 --skill lttng-tracing-root-cause-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install isc-projects/bind9 lttng-tracing-root-cause-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/isc-projects/bind9.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/lttng-tracing-root-cause-analysis .claude/skills/lttng-tracing-root-cause-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
lttng-tracing-root-cause-analysis
GitHub stars
780
Token cost
~2.5k tokens
SKILL.md length
1,322 words
Files
1
Skills in repo
9
Repo updated
First seen
Licence
MPL-2.0

At a glance

Methodology for root-causing hard concurrency / memory-ordering bugs (intermittent races, use-after-free, RCU/lock-free publish-order defects, "impossible" stale reads) with LTTng flight-recorder…

  • Works in 3 steps: Flight-recorder (snapshot) session,… → Emit the violation from the culprit,… → Read the window
  • Tasks that involve Root cause analysis
  • SKILL.md covers The setup: snapshot +…, Instrumentation discipline, Reading the trace — the… and After you've found it, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Lttng Tracing Root Cause Analysis is an agent skill from isc-projects/bind9. Methodology for root-causing hard concurrency / memory-ordering bugs (intermittent races, use-after-free, RCU/lock-free publish-order defects, "impossible" stale reads) with LTTng flight-recorder (snapshot) tracing — when static analysis, printf, and a debugger all fall short. Covers the snapshot+violation+abort setup, tracepoint instrumentation discipline (why tracepoints not printf), how to enrich a violation event so the trace is self-diagnosing, and the trace-reading patterns that crack these bugs (notably: a…

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Root cause analysis and Static analysis and SAST. The repository describes itself as: Archived mirror of https://gitlab.isc.org/isc-projects/bind9, please submit issues and PR/MRs in the GitLab. The licence is MPL-2.0.

When your agent uses it

  • Tasks that involve Root cause analysis
  • Tasks that involve Static analysis and SAST

Example prompts

  • “impossible”
  • “paradox”
  • “/lttng-tracing-root-cause-analysis”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Flight-recorder (snapshot) session, small per-CPU buffers. Snapshot mode
  2. Emit the violation from the culprit, then snapshot, then abort. At the
  3. Read the window

What it can do on your machine

Read from SKILL.md and the folder at commit 11d3a37. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are c).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Lttng Tracing Root Cause Analysis loads about 2.5k tokens when it runs. Until then it costs about 156 tokens; SKILL.md has 1,322 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~156
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from isc-projects/bind9 at commit 11d3a37, republished under its MPL-2.0 licence (© isc-projects). 1,322 words, ~2,527 tokens.

Download SKILL.mdSave it as .claude/skills/lttng-tracing-root-cause-analysis/SKILL.md (or your agent's skills folder).
name
lttng-tracing-root-cause-analysis
description
Methodology for root-causing hard concurrency / memory-ordering bugs (intermittent races, use-after-free, RCU/lock-free publish-order defects, "impossible" stale reads) with LTTng flight-recorder (snapshot) tracing — when static analysis, printf, and a debugger all fall short. Covers the snapshot+violation+abort setup, tracepoint instrumentation discipline (why tracepoints not printf), how to enrich a violation event so the trace is self-diagnosing, and the trace-reading patterns that crack these bugs (notably: a stale-after-write "paradox" is a happens-before gap, not a timing bug).

LTTng flight-recorder root-cause analysis

When a concurrency bug is intermittent and the assertion fires deep in a hot path, the three usual tools fail in three different ways:

  • Static reading can't tell you the interleaving that actually happened.
  • printf perturbs the timing — the µs-scale window you're hunting often vanishes when you add I/O — and floods you with output from the wrong threads.
  • A debugger stops the world; the race won't reproduce under a breakpoint, and you can't single-step a 192-thread interleaving.

LTTng flight-recorder (snapshot) mode is the tool that fits: near-zero overhead ring buffers per CPU, a global high-resolution clock so events from different CPUs are comparable, and an on-demand dump of exactly the window leading up to the failure. You instrument the culprit to fire a violation tracepoint, dump the snapshot, and abort; then you read the last events before the abort and the interleaving walks you to the cause. This is the tool of last resort for concurrency bugs — reach for it once you've ruled out the cheap explanations.

The setup: snapshot + violation + abort

  1. Flight-recorder (snapshot) session, small per-CPU buffers. Snapshot mode keeps a rolling overwrite buffer in memory and only writes to disk when you ask. Start small so the dump is a tight window around the failure:

    lttng create mysess --snapshot
    lttng enable-channel --userspace --subbuf-size=64K --num-subbuf=4 ch
    lttng enable-event   --userspace --channel ch 'myprovider:*'
    lttng start

    64 KiB/CPU is the CLAUDE.md default, and small is right for two reasons, not one: (a) the dump is a tight window around the failure, so it decodes fast and you read only the relevant events; (b) — the one that actually matters for reproducing the bug — a 64 KiB ring stays resident in L2, so the tracepoint stores don't evict the working set into L3/DRAM. A large (multi-MiB) ring pollutes the cache and perturbs the very µs-scale race window you're hunting; the bug can stop reproducing under heavy tracing for the same reason it stops under printf. Keep the ring small to keep the timing faithful. If your per-step tracepoints (below) are high-rate and the interesting window scrolls out before the violation, FIRST cut the event rate — disable the flooding per-iteration event and recover its data another way (e.g. from a core, or a single enriched violation event) — and only bump the subbuf size as a last resort (e.g. 256K × 4 = 1 MiB/CPU), as little as you need; a bigger buffer means both more events to read and more timing disturbance.

  2. Emit the violation from the culprit, then snapshot, then abort. At the exact check that detects the corruption, fire an enriched tracepoint, persist the in-memory ring, and crash so nothing overwrites the window:

    c
    if (corruption_detected) {
        FT_TP(violation, /* discriminating state — see below */);
        (void) system("lttng snapshot record 1>&2");
        abort();
    }

    Gate all of this behind a build flag (e.g. -DFT_ENABLE_TRACING) so it compiles out of production and your normal test matrix.

  3. Read the window:

    lttng stop
    babeltrace2 ~/lttng-traces/mysess-*/ > trace.txt

Instrumentation discipline

  • Tracepoints, never printf — timing matters. The bug lives in a sub-microsecond window; printf's I/O perturbs it out of existence and serializes threads. Tracepoint emission is a few hundred ns into a lock-free per-CPU buffer.

  • Make the violation event self-diagnosing. Don't just record "it failed" — record the state that discriminates between hypotheses. For a bad pointer, the high-value fields are usually:

    • the object's identity and a round-trip check (e.g. resolve the object by its own back-reference and compare): if it round-trips to itself the object is valid; if not, it's recycled / stale memory. This one field instantly separates "use-after-free of recycled memory" from "valid object, wrong links."
    • liveness counters (child count, refcount): zero/garbage ⇒ freed.
    • the relevant back-pointers (parent, prev) so you can see which links were and weren't wired.

    These let you classify the failure from the violation event alone, before you even read the surrounding window.

  • Add enter/step tracepoints to follow the algorithm. One tracepoint at the entry of the suspect routine and one per iteration of its core loop (carrying the loop variables) reconstruct the control flow that reached the violation — you see the path, not just the endpoint.

  • Mind the LTTNG_UST_TP_ARGS limit. lttng-ust caps a tracepoint at ~10 argument pairs. Exceed it and you get a cryptic macro error like unknown type name 'LTTNG_UST__TP_EXPROTOconst' (the arg-count machinery ran off the end). Keep a violation event ≤ ~8 fields; drop redundant ones (e.g. a field that's always NULL at the violation, or one a round-trip already implies). Pointer fields use lttng_ust_field_integer_hex(uintptr_t, name, (uintptr_t) val); counters use lttng_ust_field_integer(...).

Show full SKILL.md (611 more words)Show less

Reading the trace — the patterns that crack it

  • Read the full window, all CPUs, with ns timestamps and raw addresses. Then grep by address to pull every event touching the culprit object(s) across all threads, in time order. This reconstructs the cross-thread interleaving that no static reading could show. Note the cpu_id on each event to separate the writer thread from the reader thread.

  • Distinguish trace markers from the actual memory operation. A tracepoint at a function's entry fires before the store inside it. Don't read the tracepoint timestamp as the store's timestamp — find the event that corresponds to the real rcu_assign / publish (often a different, later marker). Mis-attributing the store's time sends you chasing ghosts.

  • THE key pattern — the stale-after-write "paradox" is a happens-before gap, not a timing bug. If the trace shows a field written at time T and read stale at T+Δ on the same object with no intervening write anywhere, that is not a contradiction and not "the store didn't land yet" (Δ can be microseconds). It means the reader reached that field through a pointer that was published before the field's store, so there is no release-consume edge carrying the store to the reader — the stale read is legal at any wall-clock delta. Treat the paradox as a signal: find which earlier publish anchored the reader's data-dependency (consume) chain, and you've found the mis-ordered publish. The fix is to publish the field before the pointer that lets readers reach it (see the rcu-mutation skill: wire back-pointers before the forward/back-channel publish; fresh edges before the live re-parent edge).

  • Walk backwards from the abort. The violation event is the last thing in the buffer. The few events just before it — on any CPU — are the proximate cause. Follow the addresses upward until the picture is consistent.

After you've found it

  • Strip the temporary enter/step/violation tracepoints and the system("lttng snapshot record") + abort() from the code before committing (they were scaffolding; the build flag kept them out of the matrix, but don't leave dead diagnostic noise in the source). Keep the durable, low-rate tracepoints if they have ongoing value.
  • A correct invariant you discovered while instrumenting may deserve to become a permanent assertion / verify-pass — but only commit it once the code actually satisfies it, or it turns the tree red on a pre-existing, non-destructive gap (scope it as separate work).

Worked example (userspace-rcu fractal trie — holder != NULL)

Symptom: an ordered-traversal reader intermittently hit assert(holder != NULL) in an up-walk (ft_skip_reanchor) under empty→rebuild churn — ~88% repro, but no static reading found it.

  1. Added a reanchor_violation tracepoint (the reached node, a round-trip metadata_to_item check, its child count, and the relevant back-pointers) + reanchor_enter/reanchor_step to follow the up-walk, all behind -DFT_ENABLE_TRACING; snapshot + abort on the violation.

  2. The violation event alone said: round-trip == self (so valid, not recycled), child-count == 1 (live), back-pointer set — yet parent == NULL. So: a valid, live node, reachable, with an unwired parent.

  3. The paradox: the writer set that node's parent at T, the reader read NULL at T+2.3µs, same metadata object, no intervening write. → happens-before gap.

  4. The full window (grepping the node's address across CPUs) showed the writer publishing a recompacted cluster into the live tree by setting a live re-parented child's back-pointer (a back-channel publish) before wiring a fresh sibling child's parent. The reader entered via the live child, so its consume chain anchored before the fresh-parent store → legal stale NULL.

  5. Fix: at publish, wire the fresh edge (and the cluster top's own back-pointer) first and the live re-parent edge last. ~88% failure → 0/96.

LTTng didn't just confirm a hypothesis — the enriched violation event and the all-CPU window generated the explanation that static analysis had missed.

© isc-projects, MPL-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/lttng-tracing-root-cause-analysis of isc-projects/bind9.

Open the folder on GitHubat commit 11d3a37

Compare with similar skills

Lttng Tracing Root Cause Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Lttng Tracing Root Cause Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Lttng Tracing Root Cause Analysis this skillisc-projects/bind9780—~2.5kAutomated safety check: PassMPL-2.0
Codexqa Rootcause Analyzeropenqa-cn/codexqa152—~2.6kAutomated safety check: PassApache-2.0
Counterexample ExplainerArabelaTso/Skills-4-SE253—~3.5kAutomated safety check: PassApache-2.0
Structural Code Search with ast-grepwarp-drive-data/warp-drive3.2k5 repos~2.4kAutomated safety check: PassMIT
Code Design Rationale Investigatorcursor/plugins10k9 repos~2.6kAutomated safety check: PassNone
ast-grep Structural Searchcode-yeongyu/oh-my-openagent70k—~3.3kAutomated safety check: PassMIT

Similar skills

  • Diagnoses exception root causes from stack traces, logs, call-chain dumps, and debug output using the CodexQA CLI for structured repo analysis.

    152 GitHub stars~2.6k tokensUpdated 7 days ago
    DevelopmentAuto-check passed
  • Counterexample Explainer

    ArabelaTso/Skills-4-SE

    Explain why counterexamples violate specifications by analyzing formal specifications (temporal logic, invariants, pre/postconditions, code contracts), informal requirements (user stories…

    253 GitHub stars~3.5k tokensUpdated 1 mo ago
    Product & Project ManagementAuto-check passed
  • Structural Code Search with ast-grep

    warp-drive-data/warp-drive

    Turns natural-language code queries into ast-grep rules for structural search, testing each rule against an example file before running it on a codebase.

    3.2k GitHub starsUsed in 5 repos~2.4k tokens
    DevelopmentAuto-check passed
  • Official

    Digs into why code is shaped the way it is by checking git history, pull requests and connected tools in parallel, then reporting a cited read on the tradeoffs.

    10k GitHub starsUsed in 9 repos~2.6k tokens
    DevelopmentAuto-check passed
  • ast-grep Structural Search

    code-yeongyu/oh-my-openagent

    Searches and rewrites code by syntax-tree shape across 25 languages with ast-grep, for codemods, structural queries and YAML lint rules, using a Python wrapper script.

    70k GitHub stars~3.3k tokensUpdated today
    DevelopmentAuto-check passed
  • Decides whether an OpenLogi device problem on macOS is a privacy-permission (TCC) problem, using agent log lines, and says which identity needs which grant.

    23k GitHub stars~2.5k tokensUpdated today
    DevelopmentAuto-check: notes

More from isc-projects/bind9

All 9 skills in this repo
  • Isc Mem Allocator

    isc-projects/bind9

    BIND 9's memory allocator wrapper (iscmem memory contexts and iscmempool fixed-size pools).

    780 GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Rcu Mutation

    isc-projects/bind9

    Correct discipline for mutating an RCU / lock-free pointer-based data structure (trie, tree, list, graph) that has concurrent readers — build a new node cluster invisibly, publish it, then reclaim…

    780 GitHub stars~3.8k tokensUpdated yesterday
    Auto-check passed
  • Struct Layout Analysis

    isc-projects/bind9

    Measure and fix C struct layout in BIND 9 — pahole on the build's DWARF for sizes, padding holes, and cacheline boundaries, plus the house cacheline-padding idiom.

    780 GitHub stars~461 tokensUpdated yesterday
    Auto-check passed
  • Tweak Release Notes

    isc-projects/bind9

    Review and refine ISC BIND 9 release notes for a new version — audit audience/action tags against the actual change substance, verify every covered issue is closed, and rewrite the auto-generated…

    780 GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Bind Mr Description

    isc-projects/bind9

    Drafting BIND 9 merge-request titles and descriptions — they feed the generated release notes, so the audience is system administrators.

    780 GitHub stars~674 tokensUpdated yesterday
    Auto-check passed
  • Isc Async Scheduling

    isc-projects/bind9

    How BIND 9 schedules callbacks — iscjobrun (same loop), iscasyncrun (any thread → any loop), iscworkenqueue (offload to a worker thread).

    780 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Lttng Tracing Root Cause Analysis

What does Lttng Tracing Root Cause Analysis do?

Methodology for root-causing hard concurrency / memory-ordering bugs (intermittent races, use-after-free, RCU/lock-free publish-order defects, "impossible" stale reads) with LTTng flight-recorder…. Lttng Tracing Root Cause Analysis is an agent skill from isc-projects/bind9. Methodology for root-causing hard concurrency / memory-ordering bugs (intermittent races, use-after-free, RCU/lock-free publish-order defects, "impossible" stale reads) with LTTng flight-recorder (snapshot) tracing — when static analysis, printf, and a debugger all fall short.

When should I use Lttng Tracing Root Cause Analysis?

Lttng Tracing Root Cause Analysis fits situations like: tasks that involve Root cause analysis; tasks that involve Static analysis and SAST.

How do I install Lttng Tracing Root Cause Analysis in Claude Code?

Run `npx skills add isc-projects/bind9 --skill lttng-tracing-root-cause-analysis -a claude-code`. Or copy the skill folder (.agents/skills/lttng-tracing-root-cause-analysis in isc-projects/bind9) into .claude/skills/lttng-tracing-root-cause-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Lttng Tracing Root Cause Analysis in Codex?

Run `npx skills add isc-projects/bind9 --skill lttng-tracing-root-cause-analysis -a codex`. Or copy the skill folder (.agents/skills/lttng-tracing-root-cause-analysis in isc-projects/bind9) into .agents/skills/lttng-tracing-root-cause-analysis in your project. Codex loads it when a task matches its description.

Can I use Lttng Tracing Root Cause Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add isc-projects/bind9 --skill lttng-tracing-root-cause-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/lttng-tracing-root-cause-analysis, .gemini/skills/lttng-tracing-root-cause-analysis, .github/skills/lttng-tracing-root-cause-analysis and .opencode/skills/lttng-tracing-root-cause-analysis in your project.

What does Lttng Tracing Root Cause Analysis need to run?

SKILL.md names no scripts, command-line tools or credentials: Lttng Tracing Root Cause Analysis is instructions for the agent only.

Does Lttng Tracing Root Cause Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Lttng Tracing Root Cause Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Lttng Tracing Root Cause Analysis use?

Lttng Tracing Root Cause Analysis is published under the MPL-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Lttng Tracing Root Cause Analysis use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Lttng Tracing Root Cause Analysis?

Skills that share tags, products or a category with Lttng Tracing Root Cause Analysis: Codexqa Rootcause Analyzer (openqa-cn/codexqa, 152 stars), Counterexample Explainer (ArabelaTso/Skills-4-SE, 253 stars), Structural Code Search with ast-grep (warp-drive-data/warp-drive, 3.2k stars) and Code Design Rationale Investigator (cursor/plugins, 10k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Lttng Tracing Root Cause Analysis?

isc-projects (a GitHub organization) maintains it in isc-projects/bind9, which has 780 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 8, 2026.

Source: isc-projects/bind9 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.