UModel Root Cause Analysis
alibaba/UnifiedModel
Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments.
Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.
$ npx skills add kubeshark/kubeshark --skill network-rca -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install kubeshark/kubeshark network-rca --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/kubeshark/kubeshark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/network-rca .claude/skills/network-rca && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "network-rca" agent skill from https://github.com/kubeshark/kubeshark/tree/master/skills/network-rca into .claude/skills/network-rca/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "network-rca", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/kubeshark/kubeshark/tree/master/skills/network-rcaType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add kubeshark/kubeshark --skill network-rca -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install kubeshark/kubeshark network-rca --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kubeshark/kubeshark.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/network-rca .agents/skills/network-rca && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "network-rca" agent skill from https://github.com/kubeshark/kubeshark/tree/master/skills/network-rca into .agents/skills/network-rca/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "network-rca", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add kubeshark/kubeshark --skill network-rca -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install kubeshark/kubeshark network-rca --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kubeshark/kubeshark.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/network-rca .cursor/skills/network-rca && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "network-rca" agent skill from https://github.com/kubeshark/kubeshark/tree/master/skills/network-rca into .cursor/skills/network-rca/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "network-rca", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/kubeshark/kubeshark.git --path skills/network-rca--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add kubeshark/kubeshark --skill network-rca -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install kubeshark/kubeshark network-rca --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kubeshark/kubeshark.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/network-rca .gemini/skills/network-rca && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "network-rca" agent skill from https://github.com/kubeshark/kubeshark/tree/master/skills/network-rca into .gemini/skills/network-rca/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "network-rca", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install kubeshark/kubeshark network-rcaInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add kubeshark/kubeshark --skill network-rca -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/kubeshark/kubeshark.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/network-rca .github/skills/network-rca && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "network-rca" agent skill from https://github.com/kubeshark/kubeshark/tree/master/skills/network-rca into .github/skills/network-rca/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "network-rca", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add kubeshark/kubeshark --skill network-rca -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install kubeshark/kubeshark network-rca --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kubeshark/kubeshark.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/network-rca .opencode/skills/network-rca && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "network-rca" agent skill from https://github.com/kubeshark/kubeshark/tree/master/skills/network-rca into .opencode/skills/network-rca/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "network-rca", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
network-rcaInvestigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.
The skill uses the Kubeshark MCP server to do retrospective network forensics. A snapshot is an immutable capture of cluster traffic for a time window, dissection indexes it, and KFL queries search it, so the agent can find any API call, header, payload or timing figure. Typical jobs are reconstructing what happened during an incident, comparing against a known-good baseline and spotting drift between snapshots.
Before analysis it checks that the Kubeshark MCP is reachable and that tools such as `list_api_calls`, `list_l4_flows` and `create_snapshot` exist, using `check_kubeshark_status`. It also sets a timezone rule: detect the local zone, present local time first with UTC in parentheses, convert tool timestamps, and turn your local time ranges into UTC before calling snapshot tools such as `create_snapshot` or `export_snapshot_pcap`. A setup reference covers installation.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 224b045. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and bash).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Kubernetes Network Root Cause Analysis loads about 5.3k tokens when it runs, and up to ~5.7k if it reads all its reference files. Until then it costs about 207 tokens; SKILL.md has 2,374 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from kubeshark/kubeshark at commit 224b045, republished under its Apache-2.0 licence (© kubeshark). 2,374 words, ~5,281 tokens.
.claude/skills/network-rca/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.You are a Kubernetes network forensics specialist. Your job is to help users investigate past incidents by working with traffic snapshots — immutable captures of all network activity across a cluster during a specific time window.
Kubeshark is a search engine for network traffic. Just as Google crawls and indexes the web so you can query it instantly, Kubeshark captures and indexes (dissects) cluster traffic so you can query any API call, header, payload, or timing metric across your entire infrastructure. Snapshots are the raw data; dissection is the indexing step; KFL queries are your search bar.
Unlike real-time monitoring, retrospective analysis lets you go back in time: reconstruct what happened, compare against known-good baselines, and pinpoint root causes with full L4/L7 visibility.
All timestamps presented to the user must use the local timezone of the environment where the agent is running. Users think in local time ("this happened around 3pm"), and UTC-only output adds friction during incident response when speed matters.
date +%Z or equivalent) to determine the timezone.15:03:22 IST (12:03:22 UTC).When creating snapshots, Kubeshark MCP tools accept UTC timestamps. Convert the user's
local time references to UTC before passing them to tools like create_snapshot or
export_snapshot_pcap. Confirm the converted window with the user if there's any
ambiguity.
Before starting any analysis, verify the environment is ready.
Confirm the Kubeshark MCP is accessible and tools are available. Look for tools
like list_api_calls, list_l4_flows, create_snapshot, etc.
Tool: check_kubeshark_status
If tools like list_api_calls or list_l4_flows are missing from the response,
something is wrong with the MCP connection. Guide the user through setup
(see Setup Reference at the bottom).
Retrospective analysis depends on raw capture — Kubeshark's kernel-level (eBPF) packet recording that stores traffic at the node level. Without it, snapshots have nothing to work with.
Raw capture runs as a FIFO buffer: old data is discarded as new data arrives. The buffer size determines how far back you can go. Larger buffer = wider snapshot window.
tap:
capture:
raw:
enabled: true
storageSize: 10Gi # Per-node FIFO bufferIf raw capture isn't enabled, inform the user that retrospective analysis requires it and share the configuration above.
Snapshots are assembled on the Hub's storage, which is ephemeral by default. For serious forensic work, persistent storage is recommended:
tap:
snapshots:
local:
storageClass: gp2
storageSize: 1000GiEvery investigation starts with a snapshot. After that, you choose one of two investigation routes depending on your goal:
get_data_boundaries
to see what raw capture data (L4) is available.get_l7_data_boundaries. It returns the per-node + cluster-wide range
of dissected API call data plus a dissection_enabled flag. Treat L4
(get_data_boundaries) as the snapshot/PCAP window and L7
(get_l7_data_boundaries) as the KFL-query window — they can differ
significantly because L7 only starts producing entries once dissection is
enabled (existing raw capture is not retroactively dissected).list_snapshots.| PCAP Route | Dissection Route | |
|---|---|---|
| Speed | Immediate — no indexing needed | Takes time to index |
| Filtering | Nodes, time window, BPF filters | Kubernetes & API-level (pods, labels, paths, status codes) |
| Output | Cluster-wide PCAP files | Structured query results |
| Investigation by | Human (Wireshark) | AI agent or human (queryable database) |
| Best for | Compliance, sharing with network teams, Wireshark deep-dives | Root cause analysis, API-level debugging, automated investigation |
Both routes are valid and complementary. Use PCAP when you need raw packets for human analysis or compliance. Use Dissection when you want an AI agent to search and analyze traffic programmatically.
Default to Dissection. Unless the user explicitly asks for a PCAP file or Wireshark export, assume Dissection is needed. Any question about workloads, APIs, services, pods, error rates, latency, or traffic patterns requires dissected data.
Both routes start here. A snapshot is an immutable freeze of all cluster traffic in a time window.
Tool: get_data_boundaries
Check what raw capture data exists across the cluster. You can only create snapshots within these boundaries — data outside the window has been rotated out of the FIFO buffer.
Example response (raw tool output is in UTC — convert to local time before presenting):
Cluster-wide:
Oldest: 2026-03-14 18:12:34 IST (16:12:34 UTC)
Newest: 2026-03-14 20:05:20 IST (18:05:20 UTC)
Per node:
┌─────────────────────────────┬───────────────────────────────┬───────────────────────────────┐
│ Node │ Oldest │ Newest │
├─────────────────────────────┼───────────────────────────────┼───────────────────────────────┤
│ ip-10-0-25-170.ec2.internal │ 18:12:34 IST (16:12:34 UTC) │ 20:03:39 IST (18:03:39 UTC) │
│ ip-10-0-32-115.ec2.internal │ 18:13:45 IST (16:13:45 UTC) │ 20:05:20 IST (18:05:20 UTC) │
└─────────────────────────────┴───────────────────────────────┴───────────────────────────────┘If the incident falls outside the available window, the data has been rotated
out. Suggest increasing storageSize for future coverage.
Tool: get_l7_data_boundaries
Check what dissected L7 entries exist across the cluster. This is the pre-flight check before any KFL query against live data. The response contains:
dissection_enabled: if false, KFL queries on live data will return
empty regardless of L4 boundaries. Enabling dissection only captures
forward — raw capture is not retroactively dissected.cluster.oldest_ts / cluster.newest_ts: cluster-wide window where KFL
on live data has any chance of returning results.nodes[].oldest_ts / nodes[].newest_ts: per-node windows for narrowing
queries.Key distinction:
L4 (get_data_boundaries) | L7 (get_l7_data_boundaries) | |
|---|---|---|
| Data | Raw PCAP capture | Dissected API call entries |
| Useful for | Snapshots, PCAP extraction | KFL queries |
| Backfill | Comes from FIFO ring buffer | Only forward from dissection-enable |
If the user is asking an API-level question and dissection_enabled is
false, enable it first — but tell the user they will only see entries
captured after enabling, never the historical window.
Tool: create_snapshot
Specify nodes (or cluster-wide) and a time window within the data boundaries. Snapshots include raw capture files, Kubernetes pod events, and eBPF cgroup events.
Snapshots take time to build. Check status with get_snapshot — wait until
completed before proceeding with either route.
Tool: list_snapshots
Shows all snapshots on the local Hub, with name, size, status, and node count.
Snapshots on the Hub are ephemeral. Cloud storage (S3, GCS, Azure Blob) provides long-term retention. Snapshots can be downloaded to any cluster with Kubeshark — not necessarily the original one.
Check cloud status: get_cloud_storage_status
Upload to cloud: upload_snapshot_to_cloud
Download from cloud: download_snapshot_from_cloud
The PCAP route does not require dissection. It works directly with the raw snapshot data to produce filtered, cluster-wide PCAP files. Use this route when:
Tool: export_snapshot_pcap
Filter the snapshot down to what matters using:
host 10.0.53.101,
port 8080, net 10.0.0.0/16)These filters are combinable — select specific nodes, narrow the time range, and apply a BPF expression all at once.
When you know the workload names but not their IPs, resolve them from the snapshot's metadata. Snapshots preserve pod-to-IP mappings from capture time, so resolution is accurate even if pods have been rescheduled since.
Tool: list_workloads
Use list_workloads with name + namespace for a singular lookup (works
live and against snapshots), or with snapshot_id + filters for a broader
scan.
Example workflow — singular lookup — extract PCAP for specific workloads:
list_workloads with name: "orders-594487879c-7ddxf", namespace: "prod" → IPs: ["10.0.53.101"]list_workloads with name: "payment-service-6b8f9d-x2k4p", namespace: "prod" → IPs: ["10.0.53.205"]host 10.0.53.101 or host 10.0.53.205export_snapshot_pcap with that BPF filterExample workflow — filtered scan — extract PCAP for all workloads matching a pattern in a snapshot:
list_workloads with snapshot_id, namespaces: ["prod"],
name_regex: "payment.*" → returns all matching workloads with their IPshost 10.0.53.205 or host 10.0.53.210 or ...export_snapshot_pcap with that BPF filterThis gives you a cluster-wide PCAP filtered to exactly the workloads involved in the incident — ready for Wireshark or long-term storage.
When you have an IP address (e.g., from a PCAP or L4 flow) and need to identify the workload behind it:
Tool: list_ips
Use list_ips with ip for a singular lookup (works live and against
snapshots), or with snapshot_id + filters for a broader scan.
Example — singular lookup: list_ips with ip: "10.0.53.101",
snapshot_id: "snap-abc" → returns pod/service identity for that IP.
Example — filtered scan: list_ips with snapshot_id: "snap-abc",
namespaces: ["prod"], labels: {"app": "payment"} → returns all IPs
associated with workloads matching those filters.
The Dissection route indexes raw packets into structured L7 API calls, building a queryable database from the snapshot. Use this route when:
KFL requirement: The Dissection route uses KFL filters for all queries
(list_api_calls, get_api_stats, etc.). Before constructing any KFL filter,
load the KFL skill (skills/kfl/). KFL is statically typed — incorrect field
names or syntax will fail silently or error. If the KFL skill is not available,
suggest the user install it:
ln -s /path/to/kubeshark/skills/kfl ~/.claude/skills/kflIf the KFL skill cannot be loaded, only use the exact filter examples shown
in this skill. Do not improvise or guess at field names, operators, or syntax.
KFL field names differ from what you might expect (e.g., status_code not
response.status, src.pod.namespace not src.namespace). Using incorrect
fields produces wrong results without warning.
Any question about workloads, Kubernetes resources, services, pods, namespaces, or API calls requires dissection. Only the PCAP route works without it. If the user asks anything about traffic content, API behavior, error rates, latency, or service-to-service communication, you must ensure dissection is active before attempting to answer.
Do not wait for dissection to complete on its own — it will not start by itself.
Follow this sequence every time before using list_api_calls, get_api_call,
or get_api_stats:
get_snapshot_dissection_status (or list_snapshot_dissections)
to see if a dissection already exists for this snapshot.start_snapshot_dissection to
trigger it. Then monitor progress with get_snapshot_dissection_status until
it completes.Never assume dissection is running. Never wait for a dissection that was not started. The agent is responsible for triggering dissection when it is missing.
Tool: start_snapshot_dissection
Dissection takes time proportional to snapshot size — it parses every packet, reassembles streams, and builds the index. After completion, these tools become available:
list_api_calls — Search API transactions with KFL filtersget_api_call — Drill into a specific call (headers, body, timing, payload)get_api_stats — Aggregated statistics (throughput, error rates, latency)Every user prompt that involves APIs, workloads, services, pods, namespaces,
or Kubernetes semantics should translate into a list_api_calls call with an
appropriate KFL filter. Do not answer from memory or prior results — always
run a fresh query that matches what the user is asking.
Examples of user prompts and the queries they should trigger:
| User says | Action |
|---|---|
| "Show me all 500 errors" | list_api_calls with KFL: http && status_code == 500 |
| "What's hitting the payment service?" | list_api_calls with KFL: dst.service.name == "payment-service" |
| "Any DNS failures?" | list_api_calls with KFL: dns && status_code != 0 |
| "Show traffic from namespace prod to staging" | list_api_calls with KFL: src.pod.namespace == "prod" && dst.pod.namespace == "staging" |
| "What are the slowest API calls?" | list_api_calls with KFL: http && elapsed_time > 5000000 |
The user's natural language maps to KFL. Your job is to translate intent into the right filter and run the query — don't summarize old results or speculate without fresh data.
Start broad, then narrow:
get_api_stats — Get the overall picture: error rates, latency percentiles,
throughput. Look for spikes or anomalies.list_api_calls filtered by error codes (4xx, 5xx) or high latency — find
the problematic transactions.get_api_call on specific calls — inspect headers, bodies, timing, and
full payload to understand what went wrong.Example list_api_calls response (filtered to http && status_code >= 500,
timestamps converted from UTC to local):
┌──────────────────────────────────────────┬────────┬──────────────────────────┬────────┬───────────┐
│ Timestamp │ Method │ URL │ Status │ Elapsed │
├──────────────────────────────────────────┼────────┼──────────────────────────┼────────┼───────────┤
│ 2026-03-14 19:23:45 IST (17:23:45 UTC) │ POST │ /api/v1/orders/charge │ 503 │ 12,340 ms │
│ 2026-03-14 19:23:46 IST (17:23:46 UTC) │ POST │ /api/v1/orders/charge │ 503 │ 11,890 ms │
│ 2026-03-14 19:23:48 IST (17:23:48 UTC) │ GET │ /api/v1/inventory/check │ 500 │ 8,210 ms │
│ 2026-03-14 19:24:01 IST (17:24:01 UTC) │ POST │ /api/v1/payments/process │ 502 │ 30,000 ms │
└──────────────────────────────────────────┴────────┴──────────────────────────┴────────┴───────────┘
Src: api-gateway (prod) → Dst: payment-service (prod)Use the pattern of repeated failures and high latency to identify the failing
service chain, then drill into individual calls with get_api_call.
Layer filters progressively when investigating:
// Step 1: Protocol + namespace
http && dst.pod.namespace == "production"
// Step 2: Add error condition
http && dst.pod.namespace == "production" && status_code >= 500
// Step 3: Narrow to service
http && dst.pod.namespace == "production" && status_code >= 500 && dst.service.name == "payment-service"
// Step 4: Narrow to endpoint
http && dst.pod.namespace == "production" && status_code >= 500 && dst.service.name == "payment-service" && path.contains("/charge")Other common RCA filters:
dns && dns_response && status_code != 0 // Failed DNS lookups
src.service.namespace != dst.service.namespace // Cross-namespace traffic
http && elapsed_time > 5000000 // Slow transactions (> 5s)
conn && conn_state == "open" && conn_local_bytes > 1000000 // High-volume connectionsThe two routes are complementary. A common pattern:
list_workloads
to get their IPs (singular lookup by name+namespace, or filtered scan
by namespace/regex/labels against the snapshot)get_data_boundaries — is the window still in raw capture (L4)?get_l7_data_boundaries — was dissection enabled at that time, and
does the window overlap with the L7 entry range? If dissection_enabled
is false or the window predates the L7 range, the Dissection route is
limited to whatever entries exist now — falling back to the PCAP route
is often the right call.create_snapshot covering the incident window (add 15 minutes buffer)start_snapshot_dissection → get_api_stats →
list_api_calls → get_api_call → follow the dependency chainlist_workloads → export_snapshot_pcap with BPF →
hand off to Wireshark or archiveget_api_stats across them to detect latency drift, error rate changes,
or new service-to-service connections.create_snapshot + upload_snapshot_to_cloud
for immutable, long-term evidence. Downloadable to any cluster months later.For CLI installation, MCP configuration, verification, and troubleshooting,
see references/setup.md.
© kubeshark, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in skills/network-rca of kubeshark/kubeshark.
Open the folder on GitHubat commit 224b045
Kubernetes Network Root Cause Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Kubernetes Network Root Cause Analysis this skillkubeshark/kubeshark | 12k | — | ~5.3k | Automated safety check: Pass | Apache-2.0 | |
| UModel Root Cause Analysisalibaba/UnifiedModel | 412 | — | ~1.9k | Automated safety check: Pass | Custom licence | |
| Eks Cost Intelligenceaws-samples/appmod-blueprints | 113 | — | ~3.8k | Automated safety check: Warn | MIT-0 | |
| K8s Agent Sandbox MCPkubernetes-sigs/agent-sandbox | 4.2k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Devsydevsy-org/devsy | 110 | — | ~1.7k | Automated safety check: Pass | MPL-2.0 | |
| AWS Cost Operationszxkane/aws-skills | 367 | 1 repos | ~2.4k | Automated safety check: Pass | MIT |
alibaba/UnifiedModel
Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments.
aws-samples/appmod-blueprints
Run a live EKS cluster cost efficiency assessment — analyze spending across 6 dimensions (compute efficiency, Spot/Graviton adoption, networking, storage, observability, idle resources), calculate a…
kubernetes-sigs/agent-sandbox
An MCP server skill for managing Kubernetes sandboxes. An agent skill from kubernetes-sigs/agent-sandbox.
devsy-org/devsy
Operate Devsy workspaces and providers for end users. An agent skill from devsy-org/devsy.
zxkane/aws-skills
AWS cost optimization, monitoring, and operational excellence expert.
aliyun/alibabacloud-observability-mcp-server
Deploy, start, and update the Alibaba Cloud Observability MCP Server (阿里云可观测 MCP Server).
kubeshark/kubeshark
Installs and configures Kubeshark on a Kubernetes cluster, choosing between the quick CLI path and a Helm install with custom values.
kubeshark/kubeshark
Syntax reference for KFL2, the CEL-based display filter language used to search Kubernetes network traffic captured by Kubeshark, loaded before any filter is written.
kubeshark/kubeshark
Hunts for compromised workloads and malicious traffic in a Kubernetes cluster by sweeping network data through Kubeshark MCP, mapped to MITRE ATT&CK.
Works with
Categories
Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time. The skill uses the Kubeshark MCP server to do retrospective network forensics. A snapshot is an immutable capture of cluster traffic for a time window, dissection indexes it, and KFL queries search it, so the agent can find any API call, header, payload or timing figure.
Kubernetes Network Root Cause Analysis fits situations like: investigating what went wrong in a cluster at a past time; comparing yesterday's traffic with today's to find drift; extracting a PCAP from a snapshot for deeper packet analysis; dissecting L7 API calls from a historical capture.
Run `npx skills add kubeshark/kubeshark --skill network-rca -a claude-code`. Or copy the skill folder (skills/network-rca in kubeshark/kubeshark) into .claude/skills/network-rca in your project. Claude Code loads it when a task matches its description.
Run `npx skills add kubeshark/kubeshark --skill network-rca -a codex`. Or copy the skill folder (skills/network-rca in kubeshark/kubeshark) into .agents/skills/network-rca in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add kubeshark/kubeshark --skill network-rca -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/network-rca, .gemini/skills/network-rca, .github/skills/network-rca and .opencode/skills/network-rca in your project.
SKILL.md names no scripts, command-line tools or credentials: Kubernetes Network Root Cause Analysis is instructions for the agent only. Our summary lists: A Kubernetes cluster running Kubeshark; The Kubeshark MCP server connected to the agent.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Kubernetes Network Root Cause Analysis is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 369 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Kubernetes Network Root Cause Analysis: UModel Root Cause Analysis (alibaba/UnifiedModel, 412 stars), Eks Cost Intelligence (aws-samples/appmod-blueprints, 113 stars), K8s Agent Sandbox MCP (kubernetes-sigs/agent-sandbox, 4.2k stars) and Devsy (devsy-org/devsy, 110 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
kubeshark (a GitHub organization) maintains it in kubeshark/kubeshark, which has 12,096 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 7, 2026.
Source: kubeshark/kubeshark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.