Agent skill

Brendangregg Use Tsa

by sickn33 in sickn33/agentic-awesome-skills

Methodical performance troubleshooting and root-cause analysis with Brendan Gregg's USE and TSA methods, plus evidence-backed RCA and postmortem reports.

MITAuto-check passedDevelopment

Install Brendangregg Use Tsa

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill brendangregg-use-tsa -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills brendangregg-use-tsa --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/brendangregg-use-tsa .claude/skills/brendangregg-use-tsa && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
brendangregg-use-tsa
GitHub stars
47k
Used in
1 other repo
Token cost
~2.8k tokens
SKILL.md length
1,122 words
Files
1
Skills in repo
1,497
Repo updated
First seen
Licence
MIT

At a glance

Methodical performance troubleshooting and root-cause analysis with Brendan Gregg's USE and TSA methods, plus evidence-backed RCA and postmortem reports.

  • Works in 8 steps: Problem Statement → 60-Second Triage (Linux) → USE Sweep (resource-oriented) → …
  • Tasks that involve Root cause analysis
  • SKILL.md covers Overview, When to Use This Skill, How It Works and Examples, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Brendangregg Use Tsa is an agent skill from sickn33/agentic-awesome-skills. Methodical performance troubleshooting and root-cause analysis with Brendan Gregg's USE and TSA methods, plus evidence-backed RCA and postmortem reports.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Root cause analysis and Runbooks and postmortems. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Tasks that involve Root cause analysis
  • Tasks that involve Runbooks and postmortems

Example prompts

  • “/brendangregg-use-tsa”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Problem Statement
  2. 60-Second Triage (Linux)
  3. USE Sweep (resource-oriented)
  4. TSA Sweep (thread-oriented)
  5. Drill Down
  6. Confirm Root Cause
  7. Fix and Verify
  8. Report

What it can do on your machine

Read from SKILL.md and the folder at commit b84d35a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • brendangregg.com
    • github.com
    • queue.acm.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Brendangregg Use Tsa loads about 2.8k tokens when it runs. Until then it costs about 44 tokens; SKILL.md has 1,122 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~44
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit b84d35a, republished under its MIT licence (© sickn33). 1,122 words, ~2,770 tokens.

Download SKILL.mdSave it as .claude/skills/brendangregg-use-tsa/SKILL.md (or your agent's skills folder).
name
brendangregg-use-tsa
description
Methodical performance troubleshooting and root-cause analysis with Brendan Gregg's USE and TSA methods, plus evidence-backed RCA and postmortem reports.
category
devops
risk
safe
source
community
source_repo
thecsdoctor/brendangregg-use-tsa-skill
source_type
community
date_added
2026-07-28
author
thecsdoctor
tags
performance, troubleshooting, root-cause-analysis, linux, observability, sre, postmortem
tools
claude, cursor, gemini, codex
license
MIT

Brendan Gregg USE+TSA Performance Analysis

Overview

A fixed, evidence-first procedure for system performance debugging, root-cause analysis (RCA), and incident reporting, distilled from Brendan Gregg's published methodologies. Instead of running whichever commands happen to be familiar, the agent poses questions first and then finds metrics to answer them: the USE Method (Utilization, Saturation, Errors) sweeps every resource, the TSA Method (Thread State Analysis) decomposes thread time, and off-CPU analysis plus flame graphs drill into what the sweeps find. Every investigation ends in a structured triage note, RCA report, or postmortem where each claim traces to a command and its output.

This skill adapts material from the community repository thecsdoctor/brendangregg-use-tsa-skill (full checklists, reference library, and report templates live there).

When to Use This Skill

  • Use when a server, VM, or container is "slow" and the cause is unknown
  • Use when latency or throughput regressed after a deploy, config change, or load shift
  • Use when CPU, memory, disk, or network metrics look abnormal and need interpretation
  • Use when an application hangs or threads pile up
  • Use when the user asks for debugging, triage, or root-cause analysis of a performance issue
  • Use when an incident needs an RCA report or a blameless postmortem with an evidence trail

How It Works

Step 0: Problem Statement

Define the problem before measuring. Ask: What makes you think there is a problem? Has it ever performed well? What changed recently (software, hardware, load)? Can it be expressed as latency or run time — quantify it. Who else is affected? What is the environment (OS, versions, config, container/VM limits)?

Step 1: 60-Second Triage (Linux)

Run the ten-command sweep, checking errors and saturation first (easiest to interpret), then utilization. Record every exonerated resource.

bash
uptime                 # load trend (includes uninterruptible I/O on Linux)
dmesg | tail           # kernel errors: oom-killer, SYN flooding, hardware
vmstat 1               # r > CPU count = CPU saturation; si/so = swapping; wa = disk
mpstat -P ALL 1        # per-CPU imbalance (single hot CPU = single-threaded app)
pidstat 1              # per-process CPU over time
iostat -xz 1           # await (app-suffered latency), avgqu-sz, %util
free -m                # memory; buffers/cache near zero hurts
sar -n DEV 1           # NIC throughput vs link limit
sar -n TCP,ETCP 1      # active/passive connections, retransmits
top                    # spot variable load
Step 2: USE Sweep (resource-oriented)

For every resource, check Utilization, Saturation, and Errors. Iterate CPUs, memory capacity, network interfaces, storage I/O and capacity, controllers, interconnects — plus software resources (mutex locks, thread pools, process/file-descriptor capacity) and imposed limits (cgroup quotas, hypervisor caps, ulimits). Check errors before utilization. Interpretations: 100% utilization is usually a bottleneck (confirm via saturation); any non-zero saturation can be a problem; non-zero, still-increasing error counters are worth investigating; and a clean sweep is a result — it narrows the search space.

Step 3: TSA Sweep (thread-oriented)

For each thread of interest, split time into: Executing / Runnable / Anonymous Paging / Sleeping / Lock / Idle. Investigate states from most to least frequent with state-appropriate tools. If more than ~10% of time is Runnable or Anonymous Paging, fix those first — latency states can be tuned to zero. Linux instruments: /proc/PID/schedstat run_delay and perf sched latency (Runnable), vmstat si/so and per-process min_flt (Paging), offcputime/cpudist from bcc (Sleeping), /proc/lock_stat and valgrind --tool=drd (Lock), pidstat/flame graphs (Executing).

Step 4: Drill Down

Follow the biggest contributor: Executing → CPU profile + flame graph; Sleeping/Lock → off-CPU stacks (offcputime -p PID, render with flamegraph.pl --color=io); latency complaints → time-division decomposition; microservices → RED method (Rate, Errors, Duration). Prefer eBPF in-kernel aggregation over per-event dumps; start with sub-second traces in production.

Step 5: Confirm Root Cause

State the causal chain (trigger → mechanism → symptom) with every link evidence-backed. Keep falsifiable hypotheses on record even when ruled out. Ask "why" up to five times. Would removing this cause prevent recurrence? Does it explain all primary evidence?

Step 6: Fix and Verify

Apply the cheapest effective fix (mantra order: don't do it → cache it → do it less → do it later → off-peak → concurrently → cheaper). Re-measure with the same instruments as the evidence and show before/after. "Deployed" is not "verified".

Step 7: Report

Produce the report the situation calls for — triage note, RCA report, or full postmortem (summary, impact, root cause, detection, investigation log, evidence table, resolution, prevention actions). Absolute dates everywhere; unknowns marked as known-unknowns.

Examples

Example 1: "This server feels slow"
User: prod-web-02 feels slow. Triage it and tell me what you ruled out.

Agent: runs the 60s sweep → dmesg shows oom-killer events at 09:41 UTC;
vmstat si/so non-zero; free -m shows 120MB free with page cache near zero.
Conclusion: memory capacity saturation (USE), host CPU/disk/network exonerated
with numbers. Report lists each exonerated resource next to its evidence.

Explanation: Errors-and-saturation-first finds the OOM events in step 1, and the exonerated resources stay on the record.

Example 2: Post-deploy latency regression
User: API p99 went 95ms → 1.9s after the 14:02 deploy. Root cause + RCA.

Agent: host sweep clean (CPU 48%, no iowait, 0 retransmits) → TSA on app
threads shows 61% Runnable on a half-idle host → checks resource controls:
/sys/fs/cgroup cpu.max = 1.5 CPUs, cpu.stat nr_throttled +54k/min → cgroup
CPU throttling after the replica increase. Fix: raise limit; verify:
nr_throttled 0/s for 72h, p99 110ms under 1.4x load. RCA report includes the
causal chain, the ruled-out hypotheses, and the command→output table.

Explanation: Runnable-dominant TSA on an under-utilized host is the signature of a resource-control limit, not a busy machine — the method routes around the wrong diagnosis.

Show full SKILL.md (460 more words)Show less

Best Practices

  • ✅ Do: Diagnose with read-only commands before changing anything
  • ✅ Do: Check errors and saturation before utilization — they interpret fastest
  • ✅ Do: Quantify everything ("p99 240ms → 2.1s", "run-queue 9 on 4 CPUs")
  • ✅ Do: Record what was ruled out, with the evidence — exoneration narrows the search
  • ✅ Do: Re-measure after the fix with the same instruments as the evidence
  • ❌ Don't: Change tunables at random until the symptom stops (drunk-man anti-method)
  • ❌ Don't: Trust low average utilization to rule out saturation — bursts hide in long intervals
  • ❌ Don't: Treat "package installed" or "dashboard green" as "working" — verify runtime state
  • ❌ Don't: Blame a component another team owns without data (blame-someone-else anti-method)

Limitations

  • This skill does not replace environment-specific validation, testing, or expert review.
  • Some metrics require privileges or tooling that may be absent (eBPF/bcc needs Linux ≥ 4.8 and usually root; perf needs perf_events access; sar needs sysstat). Missing instruments are reported as known-unknowns, not silently skipped.
  • The deepest checklists target Linux; other OSes follow the same resource × metric matrix with different instruments.
  • Stop and ask for clarification if required inputs, permissions, or safety boundaries are missing.

Security & Safety Notes

  • Diagnostics are read-only first. Any remediation (config edits, restarts, limit changes) requires explicit user confirmation before execution — the skill's own golden rules mandate this gate.
  • Production tracing has overhead: scheduler events can reach millions/sec. The skill instructs eBPF in-kernel aggregation over per-event dumping, starting with sub-second traces while watching system CPU.
  • All commands shown are standard local observability tools (vmstat, iostat, sar, perf, bcc tools, /proc reads); there are no network fetches, no credential handling, and no destructive examples. Intended usage is on systems the user is authorized to operate.

Common Pitfalls

  • Problem: Linux load averages look alarming but the CPUs are idle. Solution: Linux load includes uninterruptible (usually disk) tasks — check vmstat "r" for CPU saturation and iostat await for disk instead.
  • Problem: Host CPU looks fine but the application starves. Solution: Check resource controls, not just the host: cgroup cpu.max and cpu.stat nr_throttled (Runnable-dominant TSA is the tell).
  • Problem: "Time spent in MySQL" sends the investigation into the database. Solution: Component timers are request-oriented; run TSA on the threads — the time may be Runnable (a noisy neighbor), not execution.
  • Problem: Off-CPU stacks are polluted with nonsense frames on a busy box. Solution: Filter involuntary context switches: offcputime --state 2 (TASK_UNINTERRUPTIBLE) and fix frame pointers (-fomit-frame-pointer breaks user stacks).
  • @devops-troubleshooter - Broader DevOps incident response; use this skill for the performance-methodology core
  • @incident-responder - General incident command workflow; pairs with this skill's evidence discipline
  • @application-performance-performance-optimization - Application-level optimization after systemic bottlenecks are ruled out

Additional Resources

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/brendangregg-use-tsa of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit b84d35a

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Brendangregg Use Tsa next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Brendangregg Use Tsa compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Brendangregg Use Tsa this skillsickn33/agentic-awesome-skills47k1 repos~2.8kAutomated safety check: PassMIT
Root-Cause Analysis Writerassafkip/kipi-system112—~796Automated safety check: PassMIT
GitHub Issue Summaryascend-ai-coding/awesome-ascend-skills174—~1.4kAutomated safety check: PassNone
Post-Incident DebriefVeryGoodOpenSource/vgv-wingspan109—~1.9kAutomated safety check: PassMIT
Post Mortemhanamizuki/solopreneur152—~1.7kAutomated safety check: PassMIT
Post Mortemthananon/9arm-skills3.2k—~3.4kAutomated safety check: PassNone

Similar skills

  • Root-Cause Analysis Writer

    assafkip/kipi-system

    Writes a structured, blameless root-cause analysis for a defect that escaped a test or gate, separating surface from structural causes and linting the result.

    112 GitHub stars~796 tokensUpdated yesterday
    DevelopmentAuto-check passed
  • GitHub Issue Summary

    ascend-ai-coding/awesome-ascend-skills

    Analyze closed GitHub issues to create troubleshooting case studies with root cause analysis and lessons learned.

    174 GitHub stars~1.4k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Post-Incident Debrief

    VeryGoodOpenSource/vgv-wingspan

    Produces a blameless post-incident debrief with timeline, root cause and follow-up actions after an outage, failed release or significant bug, while details are fresh.

    109 GitHub stars~1.9k tokensUpdated 3 days ago
    DevOps & CloudAuto-check passed
  • Post Mortem

    hanamizuki/solopreneur

    Trace when a bug was introduced, find the root cause commit, understand why it happened, and produce a structured post-mortem report.

    152 GitHub stars~1.7k tokensUpdated 14 days ago
    DevOps & CloudAuto-check passed
  • Post Mortem

    thananon/9arm-skills

    Write the canonical engineering record of a fixed bug — root cause, mechanism, fix, validation, and how it slipped through.

    3.2k GitHub stars~3.4k tokensUpdated 3 mo ago
    DevOps & CloudAuto-check passed
  • After Action Report

    rampstackco/claude-skills

    Run a structured after-action review (postmortem, retrospective) on a launch, incident, or completed project to capture timeline, root cause analysis, contributing factors, and actionable lessons.

    945 GitHub stars~2.5k tokensUpdated 4 days ago
    Product & Project ManagementAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,497 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Questions about Brendangregg Use Tsa

What does Brendangregg Use Tsa do?

Methodical performance troubleshooting and root-cause analysis with Brendan Gregg's USE and TSA methods, plus evidence-backed RCA and postmortem reports. Brendangregg Use Tsa is an agent skill from sickn33/agentic-awesome-skills. Methodical performance troubleshooting and root-cause analysis with Brendan Gregg's USE and TSA methods, plus evidence-backed RCA and postmortem reports.

When should I use Brendangregg Use Tsa?

Brendangregg Use Tsa fits situations like: tasks that involve Root cause analysis; tasks that involve Runbooks and postmortems.

How do I install Brendangregg Use Tsa in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill brendangregg-use-tsa -a claude-code`. Or copy the skill folder (skills/brendangregg-use-tsa in sickn33/agentic-awesome-skills) into .claude/skills/brendangregg-use-tsa in your project. Claude Code loads it when a task matches its description.

How do I install Brendangregg Use Tsa in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill brendangregg-use-tsa -a codex`. Or copy the skill folder (skills/brendangregg-use-tsa in sickn33/agentic-awesome-skills) into .agents/skills/brendangregg-use-tsa in your project. Codex loads it when a task matches its description.

Can I use Brendangregg Use Tsa in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill brendangregg-use-tsa -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/brendangregg-use-tsa, .gemini/skills/brendangregg-use-tsa, .github/skills/brendangregg-use-tsa and .opencode/skills/brendangregg-use-tsa in your project.

What does Brendangregg Use Tsa need to run?

SKILL.md names no scripts, command-line tools or credentials: Brendangregg Use Tsa is instructions for the agent only.

Does Brendangregg Use Tsa access the network?

SKILL.md names 3 domains. As links in the text: brendangregg.com, github.com and queue.acm.org. This is read from the text; nothing was executed.

Is Brendangregg Use Tsa safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Brendangregg Use Tsa use?

Brendangregg Use Tsa is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Brendangregg Use Tsa use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Brendangregg Use Tsa?

Skills that share tags, products or a category with Brendangregg Use Tsa: Root-Cause Analysis Writer (assafkip/kipi-system, 112 stars), GitHub Issue Summary (ascend-ai-coding/awesome-ascend-skills, 174 stars), Post-Incident Debrief (VeryGoodOpenSource/vgv-wingspan, 109 stars) and Post Mortem (hanamizuki/solopreneur, 152 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Brendangregg Use Tsa?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,405 GitHub stars. The repository holds 1,497 skills in this directory. The repository was last updated on October 9, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.