Official agent skill

Troubleshoot Platform

by aws-samples in aws-samples/appmod-blueprints

Systematic troubleshooting for the PEEKS workshop platform — EKS clusters, Terraform state, ingress, load balancers, MCP tool failures, YAML validation.

OfficialMIT-0Auto-check passedDevOps & Cloud

Install Troubleshoot Platform

skills CLI
$ npx skills add aws-samples/appmod-blueprints --skill troubleshoot-platform -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aws-samples/appmod-blueprints troubleshoot-platform --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aws-samples/appmod-blueprints.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.kiro/skills/troubleshoot-platform .claude/skills/troubleshoot-platform && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
troubleshoot-platform
GitHub stars
113
Token cost
~2k tokens
SKILL.md length
736 words
Files
2 (incl. references)
Skills in repo
9
Repo updated
First seen
Licence
MIT-0

At a glance

Systematic troubleshooting for the PEEKS workshop platform — EKS clusters, Terraform state, ingress, load balancers, MCP tool failures, YAML validation.

  • Works in 4 steps: Verify the Problem Exists → Investigate Systematically → Apply Fixes → …
  • Something is broken
  • SKILL.md covers Overview, Parameters and Workflow
  • Calls kubectl, terraform and jq

What it does

Troubleshoot Platform is an agent skill from aws-samples/appmod-blueprints, published by the product's own GitHub organization. Systematic troubleshooting for the PEEKS workshop platform — EKS clusters, Terraform state, ingress, load balancers, MCP tool failures, YAML validation. Use when something is broken, not deploying, or behaving unexpectedly. Also covers resource deletion safety. Do NOT use for KRO-specific issues — use troubleshoot-kro instead. Do NOT use for addon configuration — use manage-addons instead.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/deletion-safety.md`).

It sits in DevOps & Cloud, covering Cloud networking, Infrastructure as code and MCP servers. It works with Terraform and Argo CD. The licence is MIT-0.

When your agent uses it

  • Something is broken
  • Behaving unexpectedly
  • KRO-specific issues — use troubleshoot-kro instead
  • Addon configuration — use manage-addons instead

Example prompts

  • “/troubleshoot-platform”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Verify the Problem Exists
  2. Investigate Systematically
  3. Apply Fixes
  4. Resource Deletion Safety

What it can do on your machine

Read from SKILL.md and the folder at commit 42ea29c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kubectl
    • terraform
    • jq
    • yq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use kubectl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Troubleshoot Platform loads about 2k tokens when it runs, and up to ~2.3k if it reads all its reference files. Until then it costs about 104 tokens; SKILL.md has 736 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~104
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aws-samples/appmod-blueprints at commit 42ea29c, republished under its MIT-0 licence (© aws-samples). 736 words, ~1,982 tokens.

Download SKILL.mdSave it as .claude/skills/troubleshoot-platform/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
troubleshoot-platform
description
Systematic troubleshooting for the PEEKS workshop platform — EKS clusters, Terraform state, ingress, load balancers, MCP tool failures, YAML validation. Use when something is broken, not deploying, or behaving unexpectedly. Also covers resource deletion safety. Do NOT use for KRO-specific issues — use troubleshoot-kro instead. Do NOT use for addon configuration — use manage-addons instead.

Troubleshoot Platform

Overview

Systematic troubleshooting methodology for the workshop environment. Prioritizes investigation over immediate fixes and enforces safety gates on destructive operations.

Parameters

  • symptom (required): Description of what's broken or unexpected
  • component (optional): "eks", "terraform", "ingress", "argocd", "mcp", or "general"

Workflow

1. Verify the Problem Exists

Constraints:

  • You MUST verify the actual problem exists before starting troubleshooting because symptoms may be transient
  • You MUST validate YAML syntax after modifying any YAML file: yq eval '.' <file> > /dev/null for config files, kubectl apply --dry-run=client -f <file> for K8s manifests
  • You MUST search EKS troubleshooting guide and documentation with 3-4 different query variations using EKS MCP tools before attempting fixes because the answer is usually documented
2. Investigate Systematically

For ArgoCD issues (ALWAYS check first):

Before any other investigation, check for stuck ArgoCD state:

bash
# Check for apps stuck in Deleting state
kubectl get applications.argoproj.io -n argocd -o jsonpath='{range .items[*]}{.metadata.name}{" del="}{.metadata.deletionTimestamp}{"\n"}{end}' | grep "del=2"

# Check for stuck operations (Running > 5 min)
kubectl get applications.argoproj.io -n argocd -o json | jq -r --arg now "$(date -u +%Y-%m-%dT%H:%M:%SZ)" '.items[] | select(.status.operationState.phase == "Running" and ((.status.operationState.startedAt // .metadata.creationTimestamp) | strptime("%Y-%m-%dT%H:%M:%SZ") | mktime) < (($now | strptime("%Y-%m-%dT%H:%M:%SZ") | mktime) - 300)) | .metadata.name'

# Check for ComparisonError (git cache stale)
kubectl get applications.argoproj.io -n argocd -o json | jq -r '.items[] | select(.status.operationState.message // "" | contains("ComparisonError")) | .metadata.name + ": " + .status.operationState.message'

Recovery procedures:

  1. Stuck operations (Running > 5 min): Terminate and re-trigger

    bash
    kubectl patch applications.argoproj.io <app> -n argocd --type merge -p '{"status":{"operationState":null}}'
    kubectl annotate applications.argoproj.io <app> -n argocd argocd.argoproj.io/refresh=hard --overwrite
  2. Apps stuck Deleting: The EKS ArgoCD Capability controller re-adds finalizers. You cannot force-delete while the controller is running. Options:

    • Wait for the controller to complete cleanup (may take 5-10 min)
    • If the target cluster no longer exists, delete the cluster secret so the controller stops trying:
      bash
      kubectl delete secret <cluster-name> -n argocd
    • For apps targeting the hub itself: the controller will eventually succeed once it confirms resources are gone
  3. Git cache stale ("app path does not exist"): The EKS ArgoCD Capability caches git repos. Force refresh:

    bash
    kubectl annotate applications.argoproj.io <app> -n argocd argocd.argoproj.io/refresh=hard --overwrite

    If that doesn't work, delete and let the ApplicationSet recreate:

    bash
    kubectl patch applications.argoproj.io <app> -n argocd --type merge -p '{"metadata":{"finalizers":[]}}'
    kubectl delete applications.argoproj.io <app> -n argocd --wait=false
  4. All apps disappeared (ExternalSecret flicker): The hub cluster secret was momentarily reset, causing ArgoCD to delete all apps. Recovery:

    bash
    # Re-apply root-appset
    kubectl apply -f gitops/bootstrap/root-appset.yaml
    # Wait for bootstrap to regenerate child appsets
    # Then run argocd-sync to recover stuck apps
    ./scripts/argocd-sync.sh
  5. Use the recovery script:

    bash
    ./scripts/argocd-sync.sh          # refresh all + terminate stuck ops
    ./scripts/argocd-sync.sh <app>    # refresh specific app

Key insight: With EKS ArgoCD Capability, you CANNOT force-remove finalizers — the managed controller re-adds them. The only way to unstick a deleting app is to fix the underlying issue (remove stale cluster secrets, wait for resource cleanup) or wait for the controller timeout.

CRITICAL: You MUST ALWAYS ask the user before deleting any ArgoCD Application or ApplicationSet. Deleting an app with preserveResourcesOnDeletion: false (the default) will CASCADE-DELETE all Kubernetes resources it manages. This can destroy entire namespaces, deployments, secrets, and PVCs. Even deleting a single ApplicationSet can wipe out dozens of apps and their resources across multiple clusters. NEVER delete ArgoCD apps without explicit user confirmation.

For infrastructure issues:

  • Use terraform state list then terraform state show <resource> before checking AWS CLI because Terraform state is the source of truth
  • For networking issues, examine Terraform config files first
  • You MUST NOT use get_eks_vpc_config MCP tool because it returns incomplete data — use Terraform state instead
Show full SKILL.md (316 more words)Show less

For EKS issues:

  • You MUST check cluster config with describe-cluster to confirm AutoMode status before assuming missing controllers because EKS Auto Mode means Karpenter is NOT running as pods
  • After subnet tag changes, you MUST recreate LoadBalancer services to trigger new ALB creation

For ingress-nginx webhook issues:

  • If Ingress creation fails with x509: certificate signed by unknown authority, check ingress-nginx-admission ValidatingWebhookConfiguration for empty caBundle
  • Fix: CA_BUNDLE=$(kubectl get secret ingress-nginx-admission -n ingress-nginx -o jsonpath='{.data.ca}') && kubectl patch validatingwebhookconfiguration ingress-nginx-admission --type='json' -p="[{\"op\":\"replace\",\"path\":\"/webhooks/0/clientConfig/caBundle\",\"value\":\"${CA_BUNDLE}\"}]"

For MCP tool failures:

  • If EKS MCP tools fail, provide equivalent kubectl or AWS CLI commands as fallback
  • If Terraform MCP tools fail, use terraform state commands for inspection only
  • You MUST NOT run terraform apply or terraform destroy directly — drive the configured cluster provider via task install / task destroy because they handle provider selection, backend config, and state management

Constraints:

  • You MUST wait at least 150 seconds after a new AWS Load Balancer shows "active" before testing connectivity because DNS propagation takes up to 5 minutes
  • You MUST explicitly acknowledge tool usage and summarize key findings
  • If initial investigation is inconclusive, you MUST return to AWS docs with refined queries
3. Apply Fixes

Constraints:

  • You SHOULD update Terraform config and use deployment scripts rather than creating resources manually because manual resources drift
  • If Terraform plan shows unexpected destroy operations, you MUST STOP and investigate state drift before proceeding
  • If multiple issues are found, you MUST address them in dependency order: infrastructure → platform → application
4. Resource Deletion Safety

See references/deletion-safety.md for the full deletion approval process.

Constraints:

  • You MUST NOT execute any delete command (kubectl delete, terraform destroy, ArgoCD app delete) without explicit user confirmation because these may affect production resources
  • You MUST explain what will be deleted and the impact before asking for confirmation
  • You MUST consider non-destructive alternatives first (update, refresh, sync) before suggesting deletion
  • You MUST NOT assume silence or implicit approval means "yes"

© aws-samples, MIT-0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in .kiro/skills/troubleshoot-platform of aws-samples/appmod-blueprints.

  • SKILL.md
  • references/deletion-safety.md

Open the folder on GitHubat commit 42ea29c

Compare with similar skills

Troubleshoot Platform next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Troubleshoot Platform compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Troubleshoot Platform this skillaws-samples/appmod-blueprints113—~2kAutomated safety check: PassMIT-0
Ocioracle/skills873—~2.4kAutomated safety check: PassUPL-1.0
Deploying Applicationsancoleman/ai-design-components526—~3.1kAutomated safety check: PassMIT
Provisioning Infrastructuretelagod/code-abyss243—~250Automated safety check: PassMIT
Platform Engineeringmagnus919/agent-skills113—~2.4kAutomated safety check: PassMIT
Devops Automatorcuriositech/some_claude_skills243—~1.8kAutomated safety check: PassMIT

Similar skills

  • Oci

    oracle/skills

    Official

    Oracle Cloud Infrastructure guidance for designing, operating, and troubleshooting OCI services, including OCI Kubernetes Engine (OKE), OCI Internet of Things Platform, OCI Functions deployment and…

    873 GitHub stars~2.4k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Deploying Applications

    ancoleman/ai-design-components

    Deployment patterns from Kubernetes to serverless and edge functions.

    526 GitHub stars~3.1k tokensUpdated 10 mo ago
    DevOps & CloudAuto-check passed
  • Provisioning Infrastructure

    telagod/code-abyss

    Cloud-native infrastructure knowledge reference covering Kubernetes, Helm, Kustomize, Operators, CRDs, GitOps (ArgoCD, Flux), and IaC (Terraform, Pulumi, CDK).

    243 GitHub stars~250 tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • Platform Engineering

    magnus919/agent-skills

    A skill your agent uses when building or operating internal developer platforms: infrastructure as code, CI/CD, container orchestration, service networking, secrets, and observability, or when…

    113 GitHub stars~2.4k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Devops Automator

    curiositech/some_claude_skills

    Expert DevOps engineer for CI/CD, IaC, Kubernetes, and deployment automation.

    243 GitHub stars~1.8k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Official

    Analyze Terraform plan JSON output for AzureRM Provider to distinguish between false-positive diffs (order-only changes in Set-type attributes) and actual resource changes.

    40k GitHub starsUsed in 1 repo~547 tokens
    DevOps & CloudAuto-check passed

More from aws-samples/appmod-blueprints

All 9 skills in this repo
  • Eks Best Practices

    aws-samples/appmod-blueprints

    Official

    Advisory guidance for Amazon EKS architecture and configuration decisions — compute strategy, networking, security, reliability, cost, autoscaling, observability, multi-tenancy, and upgrade planning.

    113 GitHub stars~5k tokensUpdated today
    Auto-check passed
  • Eks Recon

    aws-samples/appmod-blueprints

    Official

    EKS cluster reconnaissance and environment discovery. An agent skill from aws-samples/appmod-blueprints.

    113 GitHub stars~4.7k tokensUpdated today
    Auto-check: warnings
  • Troubleshoot Kro

    aws-samples/appmod-blueprints

    Official

    Troubleshoot Kro ResourceGraphDefinition (RGD) issues — stuck instances, ACK resource failures, IAM trust policy problems, resource conflicts.

    113 GitHub stars~986 tokensUpdated today
    Auto-check passed
  • Eks Upgrade Check

    aws-samples/appmod-blueprints

    Official

    Assess EKS cluster upgrade readiness — run automated checks across 8 areas (version, breaking changes, deprecated APIs, add-on compatibility, node readiness, workload risks, AWS Insights, upgrade…

    113 GitHub stars~2.4k tokensUpdated today
    Auto-check: warnings
  • Eks Platform Engineering

    aws-samples/appmod-blueprints

    Official

    A skill your agent uses whenever someone is designing or building an Internal Developer Platform (IDP) or doing platform engineering on Amazon EKS — phrased as "build a developer platform"…

    113 GitHub stars~4.6k tokensUpdated today
    Auto-check passed
  • Eks Security

    aws-samples/appmod-blueprints

    Official

    A skill your agent uses whenever someone needs security or compliance guidance for Amazon EKS — phrased as "CIS Benchmark for EKS", "HIPAA / PCI-DSS / FedRAMP / SOC 2 / GDPR on EKS", "harden my EKS…

    113 GitHub stars~4.7k tokensUpdated today
    Auto-check passed

Categories

Questions about Troubleshoot Platform

What does Troubleshoot Platform do?

Systematic troubleshooting for the PEEKS workshop platform — EKS clusters, Terraform state, ingress, load balancers, MCP tool failures, YAML validation. Troubleshoot Platform is an agent skill from aws-samples/appmod-blueprints, published by the product's own GitHub organization. Systematic troubleshooting for the PEEKS workshop platform — EKS clusters, Terraform state, ingress, load balancers, MCP tool failures, YAML validation.

When should I use Troubleshoot Platform?

Troubleshoot Platform fits situations like: something is broken; behaving unexpectedly; KRO-specific issues — use troubleshoot-kro instead; addon configuration — use manage-addons instead.

How do I install Troubleshoot Platform in Claude Code?

Run `npx skills add aws-samples/appmod-blueprints --skill troubleshoot-platform -a claude-code`. Or copy the skill folder (.kiro/skills/troubleshoot-platform in aws-samples/appmod-blueprints) into .claude/skills/troubleshoot-platform in your project. Claude Code loads it when a task matches its description.

How do I install Troubleshoot Platform in Codex?

Run `npx skills add aws-samples/appmod-blueprints --skill troubleshoot-platform -a codex`. Or copy the skill folder (.kiro/skills/troubleshoot-platform in aws-samples/appmod-blueprints) into .agents/skills/troubleshoot-platform in your project. Codex loads it when a task matches its description.

Can I use Troubleshoot Platform in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aws-samples/appmod-blueprints --skill troubleshoot-platform -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/troubleshoot-platform, .gemini/skills/troubleshoot-platform, .github/skills/troubleshoot-platform and .opencode/skills/troubleshoot-platform in your project.

What does Troubleshoot Platform need to run?

Going by SKILL.md and its folder, Troubleshoot Platform needs the command-line tools its instructions call (kubectl, terraform, jq and yq).

Does Troubleshoot Platform access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Troubleshoot Platform safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Troubleshoot Platform use?

Troubleshoot Platform is published under the MIT-0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Troubleshoot Platform use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 318 tokens, read only when the agent opens those files.

What are the alternatives to Troubleshoot Platform?

Skills that share tags, products or a category with Troubleshoot Platform: Oci (oracle/skills, 873 stars), Deploying Applications (ancoleman/ai-design-components, 526 stars), Provisioning Infrastructure (telagod/code-abyss, 243 stars) and Platform Engineering (magnus919/agent-skills, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Troubleshoot Platform?

aws-samples (a GitHub organization, an official publisher) maintains it in aws-samples/appmod-blueprints, which has 113 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 7, 2026.

Source: aws-samples/appmod-blueprints on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.