Agent skill

Catc Troubleshoot

by automateyournetwork in automateyournetwork/netclaw

Catalyst Center troubleshooting workflows - device unreachable investigation, client connectivity issues, interface down analysis, site-wide outage triage, wireless roaming problems, integration…

Apache-2.0Auto-check passedDevOps & Cloud

Install Catc Troubleshoot

skills CLI
$ npx skills add automateyournetwork/netclaw --skill catc-troubleshoot -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install automateyournetwork/netclaw catc-troubleshoot --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/automateyournetwork/netclaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/workspace/skills/catc-troubleshoot .claude/skills/catc-troubleshoot && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
catc-troubleshoot
GitHub stars
676
Token cost
~7.1k tokens
SKILL.md length
1,749 words
Files
1
Skills in repo
120
Repo updated
First seen
Licence
Apache-2.0

At a glance

Catalyst Center troubleshooting workflows - device unreachable investigation, client connectivity issues, interface down analysis, site-wide outage triage, wireless roaming problems, integration…

  • Works in 12 steps: Identify All Unreachable Devices → Check If It Is Site-Localized → Inspect the Specific Device → …
  • The device is managed by Catalyst Center and you want its controller/assurance view: a device is unreachable
  • SKILL.md covers Catalyst Center MCP Server, How to Call Tools, When to Use and Troubleshooting Principles, plus 4 more sections
  • Calls python3

What it does

Catc Troubleshoot is an agent skill from automateyournetwork/netclaw. Catalyst Center troubleshooting workflows - device unreachable investigation, client connectivity issues, interface down analysis, site-wide outage triage, wireless roaming problems, integration with pyATS for CLI-level diagnostics. Use when the device is managed by Catalyst Center and you want its controller/assurance view: a device is unreachable, a user reports connectivity problems, an interface is down, a site has an outage, or wireless clients have roaming issues. For direct CLI-level troubleshooting on…

Its SKILL.md is about 7.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Incident response. The repository describes itself as: An AI agent that claws through your network. The licence is Apache-2.0.

When your agent uses it

  • The device is managed by Catalyst Center and you want its controller/assurance view: a device is unreachable
  • A user reports connectivity problems
  • An interface is down
  • A site has an outage

Example prompts

  • “/catc-troubleshoot”

Requirements

  • Python 3

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Identify All Unreachable Devices
  2. Check If It Is Site-Localized
  3. Inspect the Specific Device
  4. Check Interfaces on the Unreachable Device (If CatC Has Cached Data)
  5. Escalate to pyATS for Live Validation
  6. Get Time Range
  7. Look Up the Client
  8. Analyze Client Data
  9. Check the Access Device
  10. Escalate to pyATS for Switch-Port Level Troubleshooting
  11. Get All Interfaces on the Device
  12. Identify Problem Interfaces

What it can do on your machine

Read from SKILL.md and the folder at commit aa90e7d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Catc Troubleshoot loads about 7.1k tokens when it runs. Until then it costs about 152 tokens; SKILL.md has 1,749 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~152
When it runs · the whole SKILL.md, loaded when a task matches
~7.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from automateyournetwork/netclaw at commit aa90e7d, republished under its Apache-2.0 licence (© automateyournetwork). 1,749 words, ~7,122 tokens.

Download SKILL.mdSave it as .claude/skills/catc-troubleshoot/SKILL.md (or your agent's skills folder).
name
catc-troubleshoot
description
Catalyst Center troubleshooting workflows - device unreachable investigation, client connectivity issues, interface down analysis, site-wide outage triage, wireless roaming problems, integration with pyATS for CLI-level diagnostics. Use when the device is managed by Catalyst Center and you want its controller/assurance view: a device is unreachable, a user reports connectivity problems, an interface is down, a site has an outage, or wireless clients have roaming issues. For direct CLI-level troubleshooting on devices not managed by Catalyst Center, use `pyats-troubleshoot` instead.
license
Apache-2.0
user-invocable
true

Catalyst Center Troubleshooting Workflows

Catalyst Center MCP Server

All Catalyst Center tool calls use this invocation pattern:

CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 -u $CATC_MCP_SCRIPT

Variable shorthand used throughout this document:

CATC_CMD="CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 -u $CATC_MCP_SCRIPT"

How to Call Tools

Use the $MCP_CALL protocol handler to invoke MCP tools:

bash
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" TOOL_NAME 'ARGS_JSON'

When to Use

  • Help desk escalation: user reports "I can't connect"
  • NOC alert: device unreachable in Catalyst Center
  • Client connectivity complaints (wired or wireless)
  • Site-wide outage affecting multiple users
  • Wireless roaming issues or poor signal complaints
  • Interface down investigations
  • Post-change verification when something goes wrong
  • Correlation between Catalyst Center data and live device state via pyATS

Troubleshooting Principles

  1. Define the problem -- What exactly is reported? Single user, multiple users, entire site? Wired or wireless? When did it start?
  2. Scope the impact -- Use Catalyst Center client counts and device reachability to quantify: how many clients affected? How many devices unreachable?
  3. Gather data from the controller -- Catalyst Center provides a controller-level view that covers hundreds of devices simultaneously. Start here.
  4. Correlate with live device state -- When Catalyst Center data is stale or insufficient, escalate to pyATS for real-time CLI data.
  5. Isolate the fault domain -- Is it the client, the access layer, the distribution/core, or a WAN/upstream issue?
  6. Resolve and verify -- Fix the issue, confirm via both Catalyst Center and pyATS, document.

Issue 1: Device Unreachable

Trigger: Catalyst Center shows one or more devices with reachabilityStatus: Unreachable.

Decision Tree
Device Unreachable in Catalyst Center
|
+-- Step 1: Confirm the scope
|   How many devices are unreachable?
|   Are they at the same site?
|   |
|   +-- Single device unreachable
|   |   --> Go to "Single Device Investigation"
|   |
|   +-- Multiple devices at same site
|   |   --> Go to "Site-Wide Outage Triage" (Issue 4)
|   |
|   +-- Multiple devices across sites
|       --> Go to "Catalyst Center or Upstream Issue"
|
+-- Step 2: Single Device Investigation
|   |
|   +-- Check device details in CatC
|   +-- Check last update time
|   +-- Check collection status and error code
|   +-- Attempt pyATS connectivity test
|   |
|   +-- pyATS connects successfully?
|   |   |
|   |   +-- YES: CatC polling issue (SNMP/NETCONF credentials, ACL)
|   |   +-- NO: Device is truly unreachable
|   |       |
|   |       +-- Ping from adjacent device (pyATS)
|   |       +-- Check upstream interface state
|   |       +-- Check ARP/MAC table on upstream switch
|   |       +-- Physical layer: power, cable, SFP
Step 1: Identify All Unreachable Devices
bash
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" fetch_devices '{"reachabilityStatus":["Unreachable"]}'

Record for each device: hostname, managementIpAddress, platformId, locationName, lastUpdated, collectionStatus, errorCode.

Step 2: Check If It Is Site-Localized
bash
# Get the site hierarchy to understand the topology
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" fetch_sites '{}'

Group unreachable devices by locationName. If multiple devices at the same site are unreachable, skip to Issue 4 (Site-Wide Outage).

Step 3: Inspect the Specific Device
bash
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" fetch_devices '{"hostname":["UNREACHABLE-SW-01"]}'

Analyze:

  • lastUpdated -- How long has it been unreachable? Minutes = recent event. Hours/days = chronic issue.
  • collectionStatus -- Could Not Synchronize vs Partial Collection Failure vs Managed
  • errorCode -- Specific error from Catalyst Center's collection engine
Step 4: Check Interfaces on the Unreachable Device (If CatC Has Cached Data)
bash
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" fetch_interfaces '{"device_id":"<UUID-from-step-3>"}'

NOTE: This returns the last-known interface state. If the device became unreachable recently, this data may still reflect the pre-failure state. Look for interfaces that were already showing errors before the device went dark.

Step 5: Escalate to pyATS for Live Validation

When $PYATS_MCP_SCRIPT and $PYATS_TESTBED_PATH are available and the device is in the pyATS testbed:

bash
# Attempt to connect and get basic state
PYATS_TESTBED_PATH=$PYATS_TESTBED_PATH python3 $MCP_CALL "${PYATS_PYTHON:-python3} -u $PYATS_MCP_SCRIPT" pyats_run_show_command '{"device_name":"UNREACHABLE-SW-01","command":"show ip interface brief"}'

If pyATS connects but CatC cannot reach the device:

  • The device is alive but Catalyst Center's polling is blocked
  • Check SNMP community strings, NETCONF/RESTCONF configuration
  • Check if an ACL on the device is blocking the CatC management IP
  • Verify the management VRF routing to the CatC appliance

If pyATS also cannot connect:

  • The device is truly unreachable from the management network
  • Ping from an adjacent device in pyATS:
bash
PYATS_TESTBED_PATH=$PYATS_TESTBED_PATH python3 $MCP_CALL "${PYATS_PYTHON:-python3} -u $PYATS_MCP_SCRIPT" pyats_ping_from_network_device '{"device_name":"UPSTREAM-SW-01","command":"ping 10.1.10.1"}'
  • Check the upstream switch's interface to the unreachable device:
bash
PYATS_TESTBED_PATH=$PYATS_TESTBED_PATH python3 $MCP_CALL "${PYATS_PYTHON:-python3} -u $PYATS_MCP_SCRIPT" pyats_run_show_command '{"device_name":"UPSTREAM-SW-01","command":"show interfaces GigabitEthernet1/0/1"}'
  • Check ARP and MAC table on the upstream switch:
bash
PYATS_TESTBED_PATH=$PYATS_TESTBED_PATH python3 $MCP_CALL "${PYATS_PYTHON:-python3} -u $PYATS_MCP_SCRIPT" pyats_run_show_command '{"device_name":"UPSTREAM-SW-01","command":"show arp"}'
Common Root Causes: Device Unreachable
Root CauseCatC IndicatorspyATS IndicatorsResolution
Device powered offUnreachable, no updatesConnection refusedCheck power, PDU, UPS
Management interface downUnreachableCannot SSHVerify mgmt interface config via console
SNMP credential mismatchPartial Collection FailurepyATS connects OKFix SNMP community on device or CatC
ACL blocking CatCUnreachable from CatCpyATS connects OKUpdate ACL to permit CatC management IPs
Routing issue to mgmt subnetUnreachablePing fails from adjacent deviceCheck routing table, management VRF
Upstream link failureMultiple devices unreachableUpstream interface downPhysical layer: cable, SFP, port
Device crash/reloadRecently unreachableBoot messages in logsCheck show logging, show version uptime

Issue 2: Client Connectivity Problems

Trigger: User reports cannot connect to the network, slow connectivity, or intermittent drops.

Decision Tree
Client Connectivity Issue
|
+-- Step 1: Identify the client
|   Do we have the MAC address, IP address, or username?
|   |
|   +-- MAC known --> get_client_details_by_mac
|   +-- IP known --> get_clients_list with ipv4_address filter, then get details by MAC
|   +-- Only username/location --> get_clients_list filtered by site
|
+-- Step 2: Is the client visible in CatC?
|   |
|   +-- YES: Client is associating/authenticating
|   |   |
|   |   +-- Check health score
|   |   +-- Wired or wireless?
|   |   |   |
|   |   |   +-- WIRED --> Check connected switch port, VLAN, duplex
|   |   |   +-- WIRELESS --> Check RSSI, SNR, band, AP, SSID
|   |   |
|   |   +-- Check IP assignment (DHCP working?)
|   |   +-- Check connected network device status
|   |
|   +-- NO: Client not visible
|       |
|       +-- Is the access switch/AP reachable?
|       +-- Is the client physically connected (cable/Wi-Fi enabled)?
|       +-- Is 802.1X/MAB failing? (check ISE if available)
|       +-- Is the SSID broadcasting?
Step 1: Get Time Range
bash
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" get_api_compatible_time_range '{"time_window":"last 4 hours"}'
Step 2: Look Up the Client

By MAC address (preferred):

bash
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" get_client_details_by_mac '{"client_mac_address":"AA:BB:CC:DD:EE:FF","start_time":1705312800000,"end_time":1705399200000,"view":["Wireless","WirelessHealth"]}'

By IP address:

bash
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" get_clients_list '{"start_time":1705312800000,"end_time":1705399200000,"ipv4_address":["10.1.50.42"]}'
Step 3: Analyze Client Data

For wired clients, check:

  • connectedNetworkDeviceName -- Which switch is the client on?
  • connectedNetworkDeviceInterfaceName -- Which port?
  • vlanId -- Correct VLAN assignment?
  • ipv4Address -- Valid IP? Or APIPA (169.254.x.x) indicating DHCP failure?
  • healthScore -- Score 0-10. Below 7 = issue.

For wireless clients, check:

  • ssid -- Connected to the correct SSID?
  • band -- 2.4 GHz or 5 GHz?
  • rssi -- Signal strength (see RSSI table below)
  • snr -- Signal-to-noise ratio
  • dataRate -- Current data rate in Mbps
  • connectedNetworkDeviceName -- Which AP?
  • healthScore -- Client health score
Step 4: Check the Access Device

If the client's access switch or AP shows issues:

bash
# Is the access device reachable?
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" fetch_devices '{"hostname":["ACC-SW-01"]}'

# Get the switch's interfaces to check the client-facing port
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" fetch_interfaces '{"device_id":"<switch-UUID>"}'
Step 5: Escalate to pyATS for Switch-Port Level Troubleshooting

When CatC data is insufficient (e.g., you need to check port security, 802.1X session state, or MAC address table):

bash
# Check the specific switchport
PYATS_TESTBED_PATH=$PYATS_TESTBED_PATH python3 $MCP_CALL "${PYATS_PYTHON:-python3} -u $PYATS_MCP_SCRIPT" pyats_run_show_command '{"device_name":"ACC-SW-01","command":"show interfaces GigabitEthernet1/0/15"}'

# Check authentication sessions (802.1X/MAB)
PYATS_TESTBED_PATH=$PYATS_TESTBED_PATH python3 $MCP_CALL "${PYATS_PYTHON:-python3} -u $PYATS_MCP_SCRIPT" pyats_run_show_command '{"device_name":"ACC-SW-01","command":"show authentication sessions"}'

# Check MAC address table for the client
PYATS_TESTBED_PATH=$PYATS_TESTBED_PATH python3 $MCP_CALL "${PYATS_PYTHON:-python3} -u $PYATS_MCP_SCRIPT" pyats_run_show_command '{"device_name":"ACC-SW-01","command":"show mac address-table"}'

# Check DHCP snooping bindings
PYATS_TESTBED_PATH=$PYATS_TESTBED_PATH python3 $MCP_CALL "${PYATS_PYTHON:-python3} -u $PYATS_MCP_SCRIPT" pyats_run_show_command '{"device_name":"ACC-SW-01","command":"show ip dhcp snooping binding"}'

# Check port-security
PYATS_TESTBED_PATH=$PYATS_TESTBED_PATH python3 $MCP_CALL "${PYATS_PYTHON:-python3} -u $PYATS_MCP_SCRIPT" pyats_run_show_command '{"device_name":"ACC-SW-01","command":"show port-security"}'
Common Root Causes: Client Connectivity
Root CauseCatC IndicatorspyATS/CLI IndicatorsResolution
DHCP exhaustionClient has 169.254.x.x IPshow ip dhcp pool shows 0 freeExpand DHCP scope or reduce lease time
Port security violationClient not visibleerr-disabled state on portClear port-security, investigate
802.1X failureClient not visibleAuthentication failed in ISECheck credentials, certificate, RADIUS
Wrong VLANClient visible but no connectivityPort in wrong VLANCorrect VLAN assignment
Duplex mismatchPoor health score (wired)Half-duplex on one endSet both ends to auto or match manually
Wireless: low RSSILow health score, low RSSIN/A (wireless)Move closer to AP, add AP coverage
Wireless: co-channelIntermittent dropsN/A (wireless)Adjust channel plan, reduce AP power
AP overloadedMultiple clients with issuesN/A (wireless)Load balance, add AP capacity
Upstream link failureMultiple clients affectedInterface down on distributionPhysical layer or routing issue

Issue 3: Interface Down Analysis

Trigger: Catalyst Center shows interfaces in admin up / operationally down state, or CatC alerts on interface status changes.

Step 1: Get All Interfaces on the Device
bash
# Find the device first
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" fetch_devices '{"hostname":["DIST-SW-01"]}'

# Fetch interfaces
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" fetch_interfaces '{"device_id":"<UUID>"}'
Step 2: Identify Problem Interfaces

Filter the interface response for:

  • adminStatus: "UP" AND status: "down" -- Operationally failed interfaces (admin enabled but link is down)
  • High error counters in the interface data
  • Interfaces with no description (potential unused/misconfigured ports)
Step 3: Classify the Interface
Interface TypeImpactUrgency
Uplink to distribution/coreCRITICAL: Affects all downstream devices and clientsImmediate
Inter-switch link (trunk)HIGH: May break VLAN spanning, STP reconvergenceImmediate
Access port (user-facing)MEDIUM: Single user affectedStandard
Management interfaceHIGH: Loses management access to the deviceUrgent
LoopbackHIGH if used for routing (OSPF RID, BGP update-source)Urgent
Step 4: Escalate to pyATS for Detailed Interface Diagnostics
bash
# Detailed interface statistics
PYATS_TESTBED_PATH=$PYATS_TESTBED_PATH python3 $MCP_CALL "${PYATS_PYTHON:-python3} -u $PYATS_MCP_SCRIPT" pyats_run_show_command '{"device_name":"DIST-SW-01","command":"show interfaces TenGigabitEthernet1/0/1"}'

# Check for recent log messages about the interface
PYATS_TESTBED_PATH=$PYATS_TESTBED_PATH python3 $MCP_CALL "${PYATS_PYTHON:-python3} -u $PYATS_MCP_SCRIPT" pyats_show_logging '{"device_name":"DIST-SW-01"}'

# Check SFP/transceiver status
PYATS_TESTBED_PATH=$PYATS_TESTBED_PATH python3 $MCP_CALL "${PYATS_PYTHON:-python3} -u $PYATS_MCP_SCRIPT" pyats_run_show_command '{"device_name":"DIST-SW-01","command":"show interfaces transceiver"}'

# Check EtherChannel status if applicable
PYATS_TESTBED_PATH=$PYATS_TESTBED_PATH python3 $MCP_CALL "${PYATS_PYTHON:-python3} -u $PYATS_MCP_SCRIPT" pyats_run_show_command '{"device_name":"DIST-SW-01","command":"show etherchannel summary"}'

# Check spanning-tree for blocked ports
PYATS_TESTBED_PATH=$PYATS_TESTBED_PATH python3 $MCP_CALL "${PYATS_PYTHON:-python3} -u $PYATS_MCP_SCRIPT" pyats_run_show_command '{"device_name":"DIST-SW-01","command":"show spanning-tree"}'
Show full SKILL.md (729 more words)Show less
Interface Down Root Causes
SymptomCatC DataCLI InvestigationLikely Cause
admin UP, oper DOWNStatus mismatch in CatCNo link pulse on interfaceBad cable, SFP, remote end shut
CRC errors incrementingError counters in interface datashow interfaces CRC countFaulty cable, bad SFP, duplex mismatch
Err-disabledPort not visible or shows errorsshow interfaces status err-disabledPort-security, BPDU guard, storm-control
Flapping (up/down cycles)Multiple status changesshow logging UPDOWN messagesLoose cable, auto-negotiation failure
STP blockedPort up but no trafficshow spanning-tree BLK stateSTP topology issue, loop detected

Issue 4: Site-Wide Outage Triage

Trigger: Multiple users at a single site report connectivity loss, or Catalyst Center shows multiple devices unreachable at the same location.

Decision Tree
Site-Wide Outage
|
+-- Step 1: Quantify the impact
|   How many devices unreachable at this site?
|   How many clients affected?
|   |
|   +-- ALL devices unreachable at site
|   |   --> WAN link failure, upstream router, or site power outage
|   |
|   +-- SOME devices unreachable
|   |   --> Distribution layer or IDF/MDF issue
|   |
|   +-- Devices reachable but clients can't connect
|       --> DHCP, VLAN, wireless controller, or authentication issue
|
+-- Step 2: Identify the failure boundary
|   Which devices ARE still reachable?
|   What is the common upstream for the failed devices?
|
+-- Step 3: Investigate the upstream device
|   Check interfaces, routing, logs on the common upstream
|
+-- Step 4: Resolve and verify
    Fix the issue, confirm all devices come back, verify client counts recover
Step 1: Assess Site Device Status
bash
# Get all devices at the affected site
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" fetch_devices '{"locationName":["Global/USA/NYC/Floor3"]}'

Categorize the results:

  • Count reachable vs unreachable devices
  • Identify device roles (ACCESS, DISTRIBUTION, CORE)
  • If the distribution switch is unreachable, all downstream access switches will also be unreachable
Step 2: Check Client Impact
bash
# Get time range
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" get_api_compatible_time_range '{"time_window":"last 1 hours"}'

# Count clients at the affected site NOW
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" get_clients_count '{"start_time":1705312800000,"end_time":1705399200000,"site_hierarchy":["Global/USA/NYC/Floor3"]}'

# Compare with a baseline (e.g., same time window yesterday)
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" get_api_compatible_time_range '{"time_window":"yesterday"}'

Impact assessment:

  • Current client count vs expected baseline
  • If client count dropped to 0 or near-0: complete site outage
  • If client count dropped by 50%: partial outage (one IDF or floor affected)
Step 3: Find the Failure Boundary

Check the upstream/distribution devices that serve the affected site:

bash
# Check the distribution switch
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" fetch_devices '{"hostname":["DIST-SW-NYC"]}'

# Check its interfaces
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" fetch_interfaces '{"device_id":"<DIST-UUID>"}'
Step 4: Escalate to pyATS for the Reachable Upstream

If the distribution/core device is still reachable:

bash
# Check all interfaces on the distribution switch
PYATS_TESTBED_PATH=$PYATS_TESTBED_PATH python3 $MCP_CALL "${PYATS_PYTHON:-python3} -u $PYATS_MCP_SCRIPT" pyats_run_show_command '{"device_name":"DIST-SW-NYC","command":"show ip interface brief"}'

# Check routing table -- are routes to the affected site present?
PYATS_TESTBED_PATH=$PYATS_TESTBED_PATH python3 $MCP_CALL "${PYATS_PYTHON:-python3} -u $PYATS_MCP_SCRIPT" pyats_run_show_command '{"device_name":"DIST-SW-NYC","command":"show ip route"}'

# Check OSPF/BGP adjacencies
PYATS_TESTBED_PATH=$PYATS_TESTBED_PATH python3 $MCP_CALL "${PYATS_PYTHON:-python3} -u $PYATS_MCP_SCRIPT" pyats_run_show_command '{"device_name":"DIST-SW-NYC","command":"show ip ospf neighbor"}'

# Check for recent events in the logs
PYATS_TESTBED_PATH=$PYATS_TESTBED_PATH python3 $MCP_CALL "${PYATS_PYTHON:-python3} -u $PYATS_MCP_SCRIPT" pyats_show_logging '{"device_name":"DIST-SW-NYC"}'

# Ping the unreachable access switches from the distribution
PYATS_TESTBED_PATH=$PYATS_TESTBED_PATH python3 $MCP_CALL "${PYATS_PYTHON:-python3} -u $PYATS_MCP_SCRIPT" pyats_ping_from_network_device '{"device_name":"DIST-SW-NYC","command":"ping 10.1.30.1"}'
Site Outage Root Causes
ScopeLikely CauseKey EvidenceResolution
All devices + all clientsSite power outageAll devices unreachable, no response to pingsDispatch facilities team, check UPS/PDU
All devices + all clientsWAN link failureEdge router reachable but uplink downCheck ISP, activate backup WAN if available
One floor/IDFDistribution switch failureDist switch unreachable, access switches behind it downPower cycle, console access, hardware swap
One floor/IDFTrunk link failureAccess switches up but no VLAN connectivityCheck trunk port, SFP, cable between access and dist
Devices up but no clientsDHCP failureClients getting APIPA addressesCheck DHCP server, ip helper-address, DHCP relay
Devices up, wireless clients downWLC issueAPs reachable but no SSID broadcastCheck WLC status, AP join status
Devices up, some clients downVLAN issueSpecific VLAN clients affectedCheck VLAN trunking, SVI status

Issue 5: Wireless Client Roaming Issues

Trigger: Wireless users report drops when moving between floors or areas, or CatC shows poor wireless health scores.

Step 1: Get Client History
bash
# Extended time range to capture roaming events
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" get_api_compatible_time_range '{"time_window":"last 8 hours"}'

# Get detailed client info with wireless views
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" get_client_details_by_mac '{"client_mac_address":"AA:BB:CC:DD:EE:FF","start_time":1705312800000,"end_time":1705399200000,"view":["Wireless","WirelessHealth"]}'
Step 2: Analyze Roaming Indicators

Look for:

  • AP changes -- Frequent AP name changes indicate roaming events
  • RSSI drops -- RSSI dipping below -70 dBm before roaming = sticky client (not roaming early enough)
  • Band changes -- Client flipping between 2.4 GHz and 5 GHz during movement
  • SSID consistency -- Client should stay on the same SSID during roaming
Step 3: Check AP Coverage at the Problem Area
bash
# List all APs at the affected site
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" fetch_devices '{"family":["Unified AP"],"locationName":["Global/USA/NYC/Floor3"]}'

# Count wireless clients per AP to find overloaded APs
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" get_clients_list '{"start_time":1705312800000,"end_time":1705399200000,"client_type":"wireless","site_hierarchy":["Global/USA/NYC/Floor3"]}'
Roaming Issue Root Causes
SymptomLikely CauseResolution
Client holds onto far AP (sticky client)802.11k/v not enabled or client doesn't support itEnable Optimized Roaming, BSS Transition Management
Client drops during roam802.11r (FT) not enabled, or PMK caching disabledEnable Fast Transition (FT), enable CCKM/PMK caching
Client roams to 2.4 GHzBand steering not aggressive enoughIncrease band steering threshold, check client capability
Dead zone between APsRF coverage gapAdd AP, increase power (carefully), adjust antenna
Client bounces between 2 APsEqual signal from both APsReduce power on one AP, adjust cell boundaries

Issue 6: Catalyst Center Collection or API Issues

Trigger: CatC data seems stale, APIs return unexpected errors, or collectionStatus shows failures across many devices.

Check Collection Failures
bash
CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" fetch_devices '{"collectionStatus":["Partial Collection Failure"]}'

CCC_HOST=$CCC_HOST CCC_USER=$CCC_USER CCC_PWD=$CCC_PWD python3 $MCP_CALL "python3 -u $CATC_MCP_SCRIPT" fetch_devices '{"collectionStatus":["Could Not Synchronize"]}'
API Time Range Errors

If client API calls return error 14013 (start time > 30 days ago):

  • Always use get_api_compatible_time_range first to ensure the time range is valid
  • The tool automatically clamps start times to the 30-day boundary

If client API calls return error 14006 (data not ready for endTime):

  • The get_clients_list, get_client_details_by_mac, and get_clients_count tools automatically retry with the API-suggested adjusted endTime
  • This is normal behavior when querying very recent data (within the last few minutes)

Troubleshooting Report Format

Always produce a structured findings report after completing a troubleshooting session:

Troubleshooting Report
=======================
Catalyst Center: $CCC_HOST
Timestamp: YYYY-MM-DD HH:MM UTC
Ticket/Case: [reference number if applicable]

Problem Statement
-----------------
[What was reported, by whom, when it started]

Impact Assessment
-----------------
Devices affected: X unreachable out of Y total at site
Clients affected: ~Z clients (estimated from client count delta)
Sites affected: [list]
Severity: CRITICAL / HIGH / MEDIUM / LOW

Investigation Timeline
-----------------------
1. [HH:MM] Checked CatC device reachability -- found X devices unreachable
2. [HH:MM] Identified site-localized issue at Global/USA/NYC/Floor3
3. [HH:MM] Checked distribution switch DIST-SW-NYC -- reachable
4. [HH:MM] Found TenGig1/0/1 (uplink to Floor3 IDF) down/down
5. [HH:MM] pyATS logs show %LINK-3-UPDOWN at 14:23 UTC
6. [HH:MM] SFP transceiver showing rx power below threshold

Root Cause
----------
Failing SFP transceiver on DIST-SW-NYC TenGigabitEthernet1/0/1
(uplink to Floor3 access layer IDF)

Resolution
----------
[Steps taken to resolve, or escalation path if unresolved]

Verification
-----------
- All Floor3 access switches returned to Reachable in CatC
- Client count at Floor3 recovered to baseline (335 clients)
- No further interface flaps observed in 30-minute monitoring window

Preventive Measures
--------------------
- Schedule optical monitoring for all uplink SFPs
- Add redundant uplink to Floor3 IDF (single point of failure identified)

GAIT Audit Trail

After completing any troubleshooting session, record the findings and resolution in GAIT:

bash
python3 $MCP_CALL "python3 -u $GAIT_MCP_SCRIPT" gait_record_turn '{"user_text":"Example only: replace with the actual authorized request.","assistant_text":"Troubleshooting: Site-wide outage at Global/USA/NYC/Floor3. Root cause: Failing SFP on DIST-SW-NYC Te1/0/1. Impact: 335 clients, 5 access switches unreachable for 47 minutes. Resolution: SFP replaced, services restored. Preventive: Redundant uplink recommended.","artifacts":[]}'

Audit examples are illustrative. Replace request, outcomes, identifiers and counts with observed session evidence; do not record these example results as facts. Inspect MCP isError, returned ok, and the recorded turn with gait_show when validating a new client/schema. Follow gait-session-tracking for branch checkout.

© automateyournetwork, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in workspace/skills/catc-troubleshoot of automateyournetwork/netclaw.

Open the folder on GitHubat commit aa90e7d

Compare with similar skills

Catc Troubleshoot next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Catc Troubleshoot compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Catc Troubleshoot this skillautomateyournetwork/netclaw676—~7.1kAutomated safety check: PassApache-2.0
Kubernetes Network Root Cause Analysiskubeshark/kubeshark12k—~5.3kAutomated safety check: PassApache-2.0
UModel Root Cause Analysisalibaba/UnifiedModel415—~1.9kAutomated safety check: PassCustom licence
Learningskortix-ai/suna20k—~1.1kAutomated safety check: PassCustom licence
Nix Config Debugryan4yin/nix-config2.1k—~1.2kAutomated safety check: PassMIT
Oncallpigweed-project/pigweed548—~963Automated safety check: PassApache-2.0

Similar skills

  • Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.

    12k GitHub stars~5.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • UModel Root Cause Analysis

    alibaba/UnifiedModel

    Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments.

    415 GitHub stars~1.9k tokensUpdated 17 days ago
    DevOps & CloudAuto-check passed
  • Learnings

    kortix-ai/suna

    The project's episodic memory: a timestamped ledger of rules paid for with real outages and near-misses, one entry per incident.

    20k GitHub stars~1.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Nix Config Debug

    ryan4yin/nix-config

    A skill your agent uses when something here is broken or stops working: an eval or build error, a failed activation, a dead or restarting unit, a mihomo or DNS outage, an unreachable host or MicroVM…

    2.1k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Oncall

    pigweed-project/pigweed

    Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).

    548 GitHub stars~963 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Axiom SRE Investigator

    openclaw/clawhub

    Investigates incidents and production problems with hypothesis-driven debugging, queries Axiom observability data when available, and keeps secrets out of commands and output.

    9.5k GitHub stars~7.1k tokensUpdated yesterday
    DevOps & CloudAuto-check passed

More from automateyournetwork/netclaw

All 120 skills in this repo
  • EVE-NG Lab Topology Design

    automateyournetwork/netclaw

    Entry point for designing EVE-NG network labs: classifies the request, gathers missing requirements, proposes options and validates the resulting topology.

    677 GitHub stars~612 tokensUpdated today
    Auto-check passed
  • ACI Policy Change Deployment

    automateyournetwork/netclaw

    Deploys Cisco ACI policy changes only behind an approved ServiceNow Change Request, capturing pre and post-change fault baselines and rolling back automatically on a fault delta.

    677 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Cisco ACI Fabric Health Audit

    automateyournetwork/netclaw

    Runs a phased health audit of a Cisco ACI fabric through MCP tools: node status, links, tenant and policy review, faults and endpoint learning.

    677 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Anta Validation

    automateyournetwork/netclaw

    Validate Arista EOS network state against ANTA's pre-built 208-test catalogue, with structured pass/fail verdicts.

    677 GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Arista Cvp

    automateyournetwork/netclaw

    Arista CloudVision Portal (CVP) automation via REST API — device inventory, events, connectivity monitoring, tag management (4 tools).

    677 GitHub stars~2.2k tokensUpdated today
    Auto-check: notes
  • AWS Cloud Monitoring

    automateyournetwork/netclaw

    AWS CloudWatch monitoring — metrics, alarms, log queries, VPC flow log analysis, network performance.

    677 GitHub stars~1k tokensUpdated today
    Auto-check passed

Categories

Questions about Catc Troubleshoot

What does Catc Troubleshoot do?

Catalyst Center troubleshooting workflows - device unreachable investigation, client connectivity issues, interface down analysis, site-wide outage triage, wireless roaming problems, integration…. Catc Troubleshoot is an agent skill from automateyournetwork/netclaw. Catalyst Center troubleshooting workflows - device unreachable investigation, client connectivity issues, interface down analysis, site-wide outage triage, wireless roaming problems, integration with pyATS for CLI-level diagnostics.

When should I use Catc Troubleshoot?

Catc Troubleshoot fits situations like: the device is managed by Catalyst Center and you want its controller/assurance view: a device is unreachable; A user reports connectivity problems; an interface is down; A site has an outage.

How do I install Catc Troubleshoot in Claude Code?

Run `npx skills add automateyournetwork/netclaw --skill catc-troubleshoot -a claude-code`. Or copy the skill folder (workspace/skills/catc-troubleshoot in automateyournetwork/netclaw) into .claude/skills/catc-troubleshoot in your project. Claude Code loads it when a task matches its description.

How do I install Catc Troubleshoot in Codex?

Run `npx skills add automateyournetwork/netclaw --skill catc-troubleshoot -a codex`. Or copy the skill folder (workspace/skills/catc-troubleshoot in automateyournetwork/netclaw) into .agents/skills/catc-troubleshoot in your project. Codex loads it when a task matches its description.

Can I use Catc Troubleshoot in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add automateyournetwork/netclaw --skill catc-troubleshoot -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/catc-troubleshoot, .gemini/skills/catc-troubleshoot, .github/skills/catc-troubleshoot and .opencode/skills/catc-troubleshoot in your project.

What does Catc Troubleshoot need to run?

Going by SKILL.md and its folder, Catc Troubleshoot needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Catc Troubleshoot access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Catc Troubleshoot safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Catc Troubleshoot use?

Catc Troubleshoot is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Catc Troubleshoot use?

About 7.1k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Catc Troubleshoot?

Skills that share tags, products or a category with Catc Troubleshoot: Kubernetes Network Root Cause Analysis (kubeshark/kubeshark, 12k stars), UModel Root Cause Analysis (alibaba/UnifiedModel, 415 stars), Learnings (kortix-ai/suna, 20k stars) and Nix Config Debug (ryan4yin/nix-config, 2.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Catc Troubleshoot?

automateyournetwork (a GitHub user) maintains it in automateyournetwork/netclaw, which has 676 GitHub stars. The repository holds 120 skills in this directory. The repository was last updated on October 9, 2026.

Source: automateyournetwork/netclaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.