Kcli Cluster Deployment
karmab/kcli
Guides deployment and management of Kubernetes clusters with kcli.
A skill your agent uses when a user asks to assess, plan, or validate an Amazon EKS cluster upgrade.
$ npx skills add aws/tools-for-devops-agent --skill eks-upgrade-readiness -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install aws/tools-for-devops-agent eks-upgrade-readiness --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eks-upgrade-readiness .claude/skills/eks-upgrade-readiness && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eks-upgrade-readiness" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/eks-upgrade-readiness into .claude/skills/eks-upgrade-readiness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eks-upgrade-readiness", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/aws/tools-for-devops-agent/tree/main/skills/eks-upgrade-readinessType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add aws/tools-for-devops-agent --skill eks-upgrade-readiness -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install aws/tools-for-devops-agent eks-upgrade-readiness --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/eks-upgrade-readiness .agents/skills/eks-upgrade-readiness && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eks-upgrade-readiness" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/eks-upgrade-readiness into .agents/skills/eks-upgrade-readiness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eks-upgrade-readiness", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aws/tools-for-devops-agent --skill eks-upgrade-readiness -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install aws/tools-for-devops-agent eks-upgrade-readiness --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/eks-upgrade-readiness .cursor/skills/eks-upgrade-readiness && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eks-upgrade-readiness" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/eks-upgrade-readiness into .cursor/skills/eks-upgrade-readiness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eks-upgrade-readiness", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/aws/tools-for-devops-agent.git --path skills/eks-upgrade-readiness--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add aws/tools-for-devops-agent --skill eks-upgrade-readiness -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install aws/tools-for-devops-agent eks-upgrade-readiness --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/eks-upgrade-readiness .gemini/skills/eks-upgrade-readiness && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eks-upgrade-readiness" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/eks-upgrade-readiness into .gemini/skills/eks-upgrade-readiness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eks-upgrade-readiness", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install aws/tools-for-devops-agent eks-upgrade-readinessInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add aws/tools-for-devops-agent --skill eks-upgrade-readiness -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/eks-upgrade-readiness .github/skills/eks-upgrade-readiness && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eks-upgrade-readiness" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/eks-upgrade-readiness into .github/skills/eks-upgrade-readiness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eks-upgrade-readiness", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aws/tools-for-devops-agent --skill eks-upgrade-readiness -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install aws/tools-for-devops-agent eks-upgrade-readiness --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/eks-upgrade-readiness .opencode/skills/eks-upgrade-readiness && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eks-upgrade-readiness" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/eks-upgrade-readiness into .opencode/skills/eks-upgrade-readiness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eks-upgrade-readiness", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eks-upgrade-readinessA skill your agent uses when a user asks to assess, plan, or validate an Amazon EKS cluster upgrade.
Eks Upgrade Readiness is an agent skill from aws/tools-for-devops-agent, published by the product's own GitHub organization. Use this skill when a user asks to assess, plan, or validate an Amazon EKS cluster upgrade. Activate on requests mentioning "EKS upgrade", "Kubernetes version upgrade", "upgrade readiness", "upgrade plan", "pre-upgrade check", "version skew", "deprecated API", "addon compatibility", "node group upgrade", "control plane upgrade", "EKS end of support", "EKS extended support", "Karpenter drift", "kubelet version skew", or "blue-green cluster migration". It runs a pre-upgrade assessment per the AWS EKS Best Practices…
Its SKILL.md is about 7.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 18 other files, including reference files (for example `.skilleval.yaml`, `CHANGELOG.md` and `README.md`).
It sits in DevOps & Cloud, covering Code migrations, Deployment and Site reliability engineering. It works with Amazon Web Services and Kubernetes. The repository describes itself as: Open-source tools for AWS DevOps Agent - extend DevOps Agent with ready-to-use skills, custom agents, and other tools, for incident response, root cause analysis, and operational…. The licence is Apache-2.0.
12 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit ddda70b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
kubectlawsjqhelmterraformpulumiFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.aws.amazon.comkubernetes.iogithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eks Upgrade Readiness loads about 7.7k tokens when it runs, and up to ~28k if it reads all its reference files. Until then it costs about 258 tokens; SKILL.md has 2,807 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from aws/tools-for-devops-agent at commit ddda70b, republished under its Apache-2.0 licence (© aws). 2,807 words, ~7,722 tokens.
.claude/skills/eks-upgrade-readiness/SKILL.md (or your agent's skills folder). This skill also uses 16 other files; get the full folder from GitHub.Assess and plan Amazon EKS cluster upgrades with comprehensive pre-upgrade validation aligned with the EKS Best Practices Guide.
Activate this skill when the user asks to:
Before doing anything, load references/safety-invariants.md. It defines
the knowledge hierarchy, hard rules, operation classification, and uncertainty
handling. Keep it in context for the entire assessment.
describe*, list*, get*.
The agent does NOT execute mutating APIs. Mutations are in Step 14 and
require explicit operator approval.Uses references/required-check-registry.yaml to track checks performed,
skipped, or blocked. EC = checks_performed / total_applicable × 100%.
EC < 50% produces a mandatory warning.
| Level | Meaning | When to Use |
|---|---|---|
| HIGH (90%+) | Confirmed from authoritative source | EKS Insights API, direct kubectl query, AWS API response |
| MEDIUM (60-89%) | Inferred from available data | Partial kubectl access, version matching heuristics |
| LOW (30-59%) | Limited data, possible gaps | No kubectl, no logging enabled, partial API access |
| UNKNOWN | Cannot determine | Tool unavailable, no data, access denied |
False-positive guards:
Verdict rules (evaluate applicable gates only; N/A gates are excluded):
Format: [PASS|FAIL|WARN|UNKNOWN|N/A] (confidence: HIGH) — <evidence>
AWS IAM — see README.md "Prerequisites → IAM Permissions" for the full
read-only action list (eks:Describe*, eks:List*, ec2:Describe*,
autoscaling:Describe*, iam:GetRole, servicequotas:GetServiceQuota).
Kubernetes RBAC (only if kubectl access is available — the assessment
still runs on AWS APIs alone without it, at lower confidence for CRD/Helm/PDB
checks). Read-only ClusterRole covering every kubectl get/describe used
in this skill:
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: eks-upgrade-readiness-readonly
rules:
- apiGroups: [""]
resources:
- nodes
- pods
- configmaps
- secrets
- events
- persistentvolumeclaims
- certificatesigningrequests
verbs: ["get", "list", "watch"]
- apiGroups: ["apps"]
resources: ["deployments", "statefulsets", "daemonsets", "replicasets"]
verbs: ["get", "list", "watch"]
- apiGroups: ["policy"]
resources: ["poddisruptionbudgets"]
verbs: ["get", "list", "watch"]
- apiGroups: ["apiextensions.k8s.io"]
resources: ["customresourcedefinitions"]
verbs: ["get", "list", "watch"]
- apiGroups: ["admissionregistration.k8s.io"]
resources:
- validatingwebhookconfigurations
- mutatingwebhookconfigurations
verbs: ["get", "list", "watch"]
- apiGroups: ["karpenter.sh"]
resources: ["nodepools", "nodeclaims"]
verbs: ["get", "list", "watch"]
- apiGroups: ["karpenter.k8s.aws"]
resources: ["ec2nodeclasses"]
verbs: ["get", "list", "watch"]
- apiGroups: ["crd.k8s.amazonaws.com"]
resources: ["eniconfigs"]
verbs: ["get", "list", "watch"]
- apiGroups: ["storage.k8s.io"]
resources: ["storageclasses", "csinodes"]
verbs: ["get", "list", "watch"]Bind with a ClusterRoleBinding to the identity the agent assumes (e.g. via
IRSA/Pod Identity or an EKS access entry). secrets read access is required
only for the Helm stored-manifest scan (Step 4) — omit that rule and accept
UNKNOWN on Helm checks if a customer's security policy disallows it.
aws eks describe-cluster --name <cluster-name> --region <region>Extract: cluster.version, platformVersion, status (must be ACTIVE),
kubernetesNetworkConfig, logging.clusterLogging (audit log must be enabled),
resourcesVpcConfig.subnetIds, tags (IaC ownership detection).
Determine target version: ask user or default to current + 1 minor. Confirm target is in standard support via the EKS release calendar.
Check these — failures are BLOCKERs:
aws ec2 describe-subnets with cluster subnet IDs.eks.amazonaws.com trust.aws service-quotas get-service-quota.VPC CNI mode and capacity-input detection:
kubectl get ds aws-node -n kube-system -o json | jq '
.spec.template.spec.containers[0].env[]
| select(.name | test("ENABLE_PREFIX_DELEGATION|AWS_VPC_K8S_CNI_CUSTOM_NETWORK_CFG|ENABLE_POD_ENI|WARM_IP_TARGET|MINIMUM_IP_TARGET|WARM_ENI_TARGET|WARM_PREFIX_TARGET"))
| {name, value}'Do not treat "mode detected" as capacity validated. First calculate the
managed-node-group surge using references/capacity-planning.md, distribute it
by the node group's AZ placement, then verify the relevant subnet/ENI resource
for the selected mode. Record the inputs and calculations as evidence; a missing
mode-specific input is UNKNOWN, not PASS.
| Mode | Required assessment before a node surge | Pass condition |
|---|---|---|
| Standard IPv4 | Inspect WARM_IP_TARGET, MINIMUM_IP_TARGET, and WARM_ENI_TARGET on aws-node; use node status.allocatable.pods and current pod count to calculate the additional secondary-IP demand for every surge node. | Every node subnet has enough free IPv4 addresses for its share of surge nodes, their primary ENIs, and configured warm/allocatable pod-IP demand. |
| Prefix delegation | Confirm ENABLE_PREFIX_DELEGATION=true; each IPv4 prefix consumes a /28 (16 addresses). Calculate required additional prefixes as ceil(additional_pod_ips / 16) per affected subnet/AZ. | floor(availableIpAddressCount / 16) covers the needed prefixes after allowing for node primary addresses and the configured warm-prefix target. |
| Custom networking | Confirm AWS_VPC_K8S_CNI_CUSTOM_NETWORK_CFG=true, enumerate ENIConfig resources, and map each node AZ to its spec.subnet and security groups. | The ENIConfig pod subnet, not only the cluster/node subnet, has capacity for the surge pod-IP demand in every used AZ. |
| Security Groups for Pods | Confirm ENABLE_POD_ENI=true, inspect trunk ENIs and the instance-type-specific branch-ENI limit for each node type. | Required branch ENIs/pod slots for surge workloads with pod SGs do not exceed the published limit for any instance type. Do not use generic ENI limits as a substitute. |
| IPv6 | Confirm IPv6 family and Nitro-compatible node/Fargate support. IPv6 pod addressing does not consume IPv4 pod IPs, but nodes still need valid ENI/subnet capacity. | Node ENI and subnet capacity cover surge nodes; custom networking is not assumed because it is unsupported with IPv6. |
# Custom networking: inspect every AZ-to-pod-subnet mapping
kubectl get eniconfig -o json | jq '.items[] | {name: .metadata.name, subnet: .spec.subnet, securityGroups: .spec.securityGroups}'
# Security Groups for Pods: inspect trunk/branch interfaces after mode detection
aws ec2 describe-network-interfaces \
--filters Name=interface-type,Values=trunk,branch \
--query 'NetworkInterfaces[].{type:InterfaceType,subnet:SubnetId,instance:Attachment.InstanceId,status:Status}'Operator/CI tooling that is too far behind the target version causes confusing failures during and after the upgrade. Check installed versions where available:
kubectl version --client -o json # client minor version
eksctl version # if eksctl-managed
helm version --short # if Helm-managed workloads| Tool | Skew Rule | Risk if Violated |
|---|---|---|
| kubectl | Must be within ±1 minor of the target kube-apiserver version (upstream Kubernetes version skew policy) | Unrecognized fields, API calls silently rejected or misinterpreted |
| eksctl | Must support the target EKS version (check release notes for the version that added support) | eksctl upgrade commands fail or use stale defaults |
| Helm | 3.8+ recommended for OCI registry support; otherwise not EKS-version-gated | Chart operations may fail independent of the cluster upgrade |
| Terraform AWS provider | Must be new enough to support any target-version-specific attributes in use (e.g. upgrade_policy, compute_config for Auto Mode) — check the provider changelog for the attribute | terraform apply fails validation or silently ignores the new attribute |
Do not hardcode exact version floors here — they shift every EKS release. Report the installed version, the rule, and a WARN if it cannot be confirmed current; never treat "tool not detected" as PASS.
Primary authoritative signal. Always query first.
aws eks list-insights --cluster-name <cluster> \
--filter '{"categories":["UPGRADE_READINESS"],"kubernetesVersions":["<target>"]}'
aws eks describe-insight --cluster-name <cluster> --id <insight-id>| Status | Gate | Action |
|---|---|---|
| ERROR | FAIL | Must fix before upgrade |
| WARNING | WARN | Recommended fix |
| PASSING | PASS | No action |
| UNKNOWN | UNKNOWN | EKS could not evaluate the check; investigate and refresh |
| None returned | UNKNOWN | Continue other checks |
Freshness gate: For every returned summary, capture lastRefreshTime and
lastTransitionTime, then call DescribeInsight to collect the status,
affected resources, and recommendation. Treat data as stale when
lastRefreshTime is more than 24 hours old at assessment time, or predates a
known relevant workload/addon change. A stale, missing, inaccessible, or
unpaginated insight set is UNKNOWN; never reuse it as PASS. The assessment
must not call StartInsightsRefresh because this skill is read-only. Instead,
ask the operator to refresh insights through an approved workflow and rerun the
assessment after the refresh completes.
Critical: Insights does NOT cover Helm stored manifests, CRD deprecations, StatefulSets, Karpenter, service quotas, PDBs, or capacity planning.
Check version-specific removal gates relevant to user's target:
Helm stored manifests — the #1 missed blocker. Scan latest deployed
release secrets for deprecated apiVersion lines:
kubectl get secrets -A -l owner=helm,status=deployed
# Decode: base64 -d | base64 -d | gunzip | jq -r '.manifest'Third-party CRD deprecations — check Istio, cert-manager, Karpenter, Flux, Argo, Prometheus Operator versions against known deprecation timelines.
Tools: kubent, pluto detect-all-in-cluster, helm mapkubeapis --dry-run.
See references/api-deprecations.md for full removal schedule.
Managed addons: Use live API to build compatibility matrix:
aws eks list-addons --cluster-name <cluster>
aws eks describe-addon --cluster-name <cluster> --addon-name <name>
aws eks describe-addon-versions --addon-name <name> --kubernetes-version <target>Self-managed addons: First compare ListAddons with the actual
kube-system workloads. Explicitly detect the three core addons: aws-node
(VPC CNI), coredns, and kube-proxy. If a core component is absent from the
managed-addon inventory but present in-cluster, mark it self-managed/custom and
inspect its image, args, and configuration before assessing target support:
kubectl -n kube-system get daemonset aws-node kube-proxy -o json | \
jq '.items[] | {name: .metadata.name, images: [.spec.template.spec.containers[].image], args: [.spec.template.spec.containers[].args]}'
kubectl -n kube-system get deployment coredns -o json | \
jq '.items[] | {name: .metadata.name, images: [.spec.template.spec.containers[].image], args: [.spec.template.spec.containers[].args]}'
kubectl -n kube-system get configmap coredns aws-node -o yamlFor a custom CoreDNS Corefile, run the Corefile migration check. For VPC CNI, validate custom environment/config-map values and the mode-specific capacity gate in Step 2. For kube-proxy, validate its deployed mode and version against its upstream support policy. Then scan other self-managed components (aws-load-balancer, external-dns, metrics-server, cluster-autoscaler, cert-manager, ingress-nginx, argocd, flux); extract image tags and validate against their K8s support matrices.
Pod Identity Agent (eks-pod-identity-agent): Check like any other managed
addon via DescribeAddon / DescribeAddonVersions — its version gates which
association features are available (e.g. multiple associations per pod,
target IAM role sessions). If not installed but kubectl get pods -A -o json
shows service accounts with eks.amazonaws.com/role-arn annotations instead,
the cluster is on IRSA, not Pod Identity — note this for blue-green planning
(see references/upgrade-troubleshooting.md → Identity Migration Considerations).
Upgrade order: Pre-CP: Karpenter, Cluster Autoscaler, incompatible webhooks. Post-CP: kube-proxy → vpc-cni → coredns → CSI drivers → others → self-managed.
See references/addon-version-matrix.md for static fallback reference.
Inventory ALL node populations. See references/data-plane-inventory.md for
complete detection commands.
aws eks list-nodegroups + describe-nodegroup for eachdescribe-auto-scaling-groupscluster.computeConfig.enabledaws eks list-fargate-profiles + describe eachkubectl get nodes — confirm within skew windowVersion skew requires two independent predicates:
kubelet_minor <= current_control_plane_minor). A node
already newer than the current API server is an invalid state and must be
corrected before planning the upgrade.Any node violating either predicate is a FAIL.
If AL2 detected and target ≥1.33: CRITICAL blocker (EKS releases AL2 AMIs only through 1.32). If AL2 is detected with a target <1.33: WARNING — upstream Amazon Linux 2 reaches end of life on June 30, 2026.
Assess: bootstrap method (bootstrap.sh vs nodeadm), custom AMIs, user data compatibility (yum→dnf, kubelet-extra-args→NodeConfig), cgroup v2 workload compatibility, IMDSv2 readiness.
See references/al2-al2023-migration.md for full detection commands and
migration strategy.
Pre-CP: Karpenter (if needed), Cluster Autoscaler (must match target),
admission webhooks with failurePolicy: Fail, custom controllers using
deprecated APIs.
Post-CP: Standard addon and node group upgrade order (Step 5).
Webhook check:
kubectl get validatingwebhookconfigurations -o json | jq '.items[] | select(.webhooks[].failurePolicy == "Fail")'
kubectl get mutatingwebhookconfigurations -o json | jq '.items[] | select(.webhooks[].failurePolicy == "Fail")'PDB blockers: maxUnavailable: 0, minAvailable == replicas, orphaned PDBs:
kubectl get pdb -A -o json | jq '.items[] | select(.status.disruptionsAllowed == 0)'Pre-drain safety (DRAIN-01 to DRAIN-06): Bare pods, emptyDir data loss,
custom finalizers, EBS AZ-pinning, webhook deadlock, CoreDNS SPOF.
See references/pre-drain-safety.md for full detection commands.
TopologySpreadConstraints: Flag multi-replica deployments without topology spread.
StatefulSet safety: Check terminationGracePeriodSeconds != 0, PVC retention
policy, single-replica without PDB, update strategy.
Scaled-to-zero workloads: Detect and flag for separate validation.
Confirm healthy steady state before upgrade. Failures compound on unhealthy clusters.
Record baselines for post-upgrade comparison.
Fargate pods upgrade when redeployed after CP upgrade. All Fargate pods must be restarted post-upgrade. Restart command is in Step 14 (mutation, requires approval).
Detect management method to route remediation correctly:
| Detection | Management Plane | Mutation Routing |
|---|---|---|
| ACK CRD + Cluster CR | ACK | Patch ACK Cluster CR |
ACK CR with kro.run/owned | KRO over ACK | Patch kro instance |
Tags: terraform:* | Terraform | Update .tf, terraform apply |
Tags: aws:cloudformation:* | CloudFormation | Update template, stack update |
Tags: aws:cdk:* | CDK | Update construct, cdk deploy |
Tags: eksctl.cluster.k8s.io/* | eksctl | Update config, eksctl upgrade |
Labels: argocd.argoproj.io/* | ArgoCD | Update Git source, sync |
Labels: kustomize.toolkit.fluxcd.io/* | Flux | Update Git source, reconcile |
Tags: pulumi:* | Pulumi | Update program, pulumi up |
| None found | unknown | Block mutations until confirmed |
Route ALL remediation through the owning tool — never suggest direct AWS CLI when IaC is detected (causes drift).
During upgrades, autoscalers can interfere with rolling replacement. Check current Karpenter consolidation config and Cluster Autoscaler scale-down state. Recommend pausing both before node rotation and re-enabling after completion.
Pause commands are in Step 14 (mutations, require operator approval).
⚠️ ALL commands in this section are MUTATIONS. The agent MUST NOT execute these — present as a playbook for operator review.
helm mapkubeapis + helm upgradeDescribeAddon output and
configurationValues; use --resolve-conflicts PRESERVE to retain reviewed
custom configuration, or OVERWRITE only after approving replacement with
EKS defaults and recording rollback steps. OVERWRITE can discard custom
configuration.kubectl rollout restart across namespaceskubectl patch pdb (revert after upgrade)kubectl annotate nodepools --all "karpenter.sh/do-not-disrupt=true"scale-down-enabled=falseupdate-nodegroup-version, Karpenter via
EC2NodeClass patch (drift), self-managed via launch template updatePresent as validation checklist for operator:
Execution Order:
Rollback Matrix:
| Component | Reversibility | Method |
|---|---|---|
| Control plane | CONDITIONAL (7-day window) | aws eks update-cluster-version --kubernetes-version <N-1> |
| Addons | FULL | Downgrade to previous version |
| MNG | PARTIAL | Can halt; completed nodes stay |
| Karpenter nodes | FULL | Revert EC2NodeClass |
| Self-managed | FULL | Revert launch template |
| Fargate | FULL | Redeploy previous config |
Rollback eligibility has two phases:
aws eks list-insights --cluster-name <cluster> --filter '{"categories":["ROLLBACK_READINESS"]}',
paginate, then describe-insight for each entry. ERROR blocks a normal
rollback; UNKNOWN means EKS could not evaluate readiness and also blocks a
normal rollback. Only PASSING insights support an eligible rollback.This assessment reports the result but never performs update-cluster-version
or a forced rollback.
## EKS Upgrade Readiness Report
**Cluster:** <name> (<region>)
**Current Version:** <current>
**Target Version:** <target>
**Assessment Date:** <date>
**Management Plane:** <detected>
**Evidence Completeness:** <X>% (<performed>/<applicable>)
**Overall Readiness:** READY / NOT READY / READY WITH WARNINGS / CANNOT DETERMINE
### Pre-Upgrade Health Baseline
- [PASS/FAIL] (confidence: HIGH) All nodes Ready
- [PASS/FAIL] (confidence: HIGH) No pending CSRs
- [PASS/FAIL] (confidence: HIGH) No crash-looping system pods
- [PASS/FAIL] (confidence: HIGH) DNS resolution working
- [PASS/FAIL] (confidence: HIGH) Metrics server responding
### Infrastructure Prerequisites
- [PASS/FAIL] (confidence: HIGH) Subnet IP availability (mode: <type>)
- [PASS/FAIL] (confidence: HIGH) EKS IAM role valid
- [PASS/FAIL/N/A] (confidence: HIGH) KMS key access
- [PASS/FAIL] (confidence: HIGH) EC2 vCPU quota headroom
- [PASS/FAIL] (confidence: HIGH) EBS volume quota headroom
### EKS Upgrade Insights
- [PASS/FAIL/UNKNOWN] (confidence: HIGH) <summary>
### Data Plane Inventory
- Managed Node Groups: <count> (versions: <list>)
- Self-Managed ASGs: <count> (versions: <list>)
- Karpenter NodePools: <count> (version: <ver>)
- Fargate Profiles: <count>
- Total Nodes: <count>
### Blockers (must fix)
1. [FAIL] (confidence: HIGH) <description> — <remediation>
### Warnings (recommended)
1. [WARN] (confidence: MEDIUM) <description> — <recommendation>
### Passing Checks
1. [PASS] (confidence: HIGH) <description>
### Unknown / Not Assessed
1. [UNKNOWN] <gate> — <reason>
### Upgrade Plan
<execution order from Step 16>
### Rollback Window
- Rollback eligibility: ELIGIBLE / NOT ELIGIBLE / CHECK AFTER UPGRADE
- Window: 7 days from CP upgrade completion
- Note: Add-ons and node groups must be rolled back BEFORE CP
### Pre-Drain Risks
- Bare pods (DRAIN-01): <count>
- EmptyDir data loss (DRAIN-02): <count>
- EBS AZ-pinning (DRAIN-04): <count>
- Webhook deadlock (DRAIN-05): <assessment>
- CoreDNS SPOF (DRAIN-06): <status>
### Estimated Timeline
- Control plane: ~30 min
- Addons: ~5 min each
- Node groups: ~<X> min per group
- Total: ~<Y> minWhen the operator asks for a structured result (CI/CD gating, scripted polling, dashboards), emit this JSON alongside — never instead of — the markdown report. Every gate in the markdown report must have a matching entry; the JSON is a serialization of the same evidence, not a summary.
{
"cluster": "<name>",
"region": "<region>",
"assessmentTimestamp": "<ISO-8601>",
"currentVersion": "<current>",
"targetVersion": "<target>",
"overallVerdict": "READY | READY_WITH_WARNINGS | NOT_READY | CANNOT_DETERMINE",
"evidenceCompletenessPct": 0,
"gates": [
{
"id": "<check-id from required-check-registry.yaml, e.g. NODE-04, ADDON-02, INFRA-01>",
"name": "<human-readable check name>",
"status": "PASS | FAIL | WARN | UNKNOWN | N_A",
"confidence": "HIGH | MEDIUM | LOW",
"evidence": "<short evidence string, same as markdown bullet>",
"remediation": "<remediation text, or null if PASS>",
"checkedAt": "<ISO-8601>"
}
],
"rollback": {
"eligible": true,
"windowExpiresAt": "<ISO-8601 or null>"
}
}gates[].id maps 1:1 to the IDs in references/required-check-registry.yaml
(prefixes: PF- pre-flight, INFRA- infrastructure, NODE- node assessment,
ADDON- addon assessment, WKLD- workload assessment, KARP- Karpenter,
DRAIN- pre-drain safety, ROLL- rollback), so a CI pipeline can gate on
specific check categories (e.g. fail only on NODE-* or ADDON-* FAILs,
warn-only on others) instead of just the overall verdict. overallVerdict
follows the same rules as the markdown report — it is never READY while
any gate is UNKNOWN.
See references/ directory for:
safety-invariants.md — Hard safety rules, knowledge hierarchy, operation classificationrequired-check-registry.yaml — All 60+ checks with IDs, categories, and severitypre-flight-checks.yaml — Blocking vs warning checks, timeouts, soak periods, rollback conditionsapi-deprecations.md — Full K8s API removal schedule by versionaddon-version-matrix.md — EKS addon compatibility per version (static fallback)capacity-planning.md — FDCR/ODCR and surge capacity guidanceupgrade-troubleshooting.md — Common failures, feature removals, and toolskarpenter-checks.md — Full 14-check Karpenter registry (KARP-01 to KARP-14)pre-drain-safety.md — DRAIN-01 to DRAIN-06 detection and remediational2-al2023-migration.md — AL2→AL2023 migration assessment detailsdata-plane-inventory.md — MNG, self-managed, Karpenter, Auto Mode, Fargate inventory commands© aws, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 16 other files (references) in skills/eks-upgrade-readiness of aws/tools-for-devops-agent.
Open the folder on GitHubat commit ddda70b
Eks Upgrade Readiness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eks Upgrade Readiness this skillaws/tools-for-devops-agent | 102 | — | ~7.7k | Automated safety check: Pass | Apache-2.0 | |
| Kcli Cluster Deploymentkarmab/kcli | 653 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Aspire DeploymentCommunityToolkit/Aspire | 629 | — | ~4.5k | Automated safety check: Notes | MIT | |
| Deployment Automationaiskillstore/marketplace | 430 | 1 repos | ~3k | Automated safety check: Notes | None | |
| Discover Infrarand/cc-polymath | 181 | — | ~783 | Automated safety check: Pass | MIT | |
| Senior DevOps Toolkitmaslennikov-ig/claude-code-orchestrator-kit | 260 | 6 repos | ~1.1k | Automated safety check: Notes | Custom licence |
karmab/kcli
Guides deployment and management of Kubernetes clusters with kcli.
CommunityToolkit/Aspire
WORKFLOW SKILL — Deploy Aspire apps from AppHost models to Docker Compose, Kubernetes, Azure, AWS, or preview Radius.
aiskillstore/marketplace
Automate application deployment to cloud platforms and servers.
rand/cc-polymath
Automatically discover cloud, infrastructure, deployment, and container skills when working with AWS, GCP, Azure, Docker, Kubernetes, Terraform, Netlify, Heroku, serverless, or IaC
maslennikov-ig/claude-code-orchestrator-kit
Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup…
langbot-app/LangBot
Deploys and configures a LangBot instance with Docker Compose or Kubernetes, covering config.yaml, the Box sandbox runtime, the plugin runtime and the global API key.
aws/tools-for-devops-agent
A skill your agent uses for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances.
aws/tools-for-devops-agent
ALWAYS use this skill in the beginning of any incident investigation, root cause analysis, or operational troubleshooting.
aws/tools-for-devops-agent
AWS Database Migration Service (DMS) operational review and troubleshooting skill.
aws/tools-for-devops-agent
Performs a comprehensive Amazon ECS operations review across the 6 review pillars (Resiliency & HA, Observability, Security, Operations, Performance, Additional Analysis) using read-only AWS APIs…
aws/tools-for-devops-agent
Comprehensive Amazon RDS and Aurora operational review aligned with the AWS Well-Architected Framework and RDS/Aurora best practices.
aws/tools-for-devops-agent
Amazon SageMaker AI Operational Review. An agent skill from aws/tools-for-devops-agent.
Works with
Categories
A skill your agent uses when a user asks to assess, plan, or validate an Amazon EKS cluster upgrade. Eks Upgrade Readiness is an agent skill from aws/tools-for-devops-agent, published by the product's own GitHub organization. Use this skill when a user asks to assess, plan, or validate an Amazon EKS cluster upgrade.
Eks Upgrade Readiness fits situations like: A user asks to assess; validate an Amazon EKS cluster upgrade; general EKS troubleshooting unrelated to version upgrades; EKS Anywhere/Outpost clusters.
Run `npx skills add aws/tools-for-devops-agent --skill eks-upgrade-readiness -a claude-code`. Or copy the skill folder (skills/eks-upgrade-readiness in aws/tools-for-devops-agent) into .claude/skills/eks-upgrade-readiness in your project. Claude Code loads it when a task matches its description.
Run `npx skills add aws/tools-for-devops-agent --skill eks-upgrade-readiness -a codex`. Or copy the skill folder (skills/eks-upgrade-readiness in aws/tools-for-devops-agent) into .agents/skills/eks-upgrade-readiness in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aws/tools-for-devops-agent --skill eks-upgrade-readiness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eks-upgrade-readiness, .gemini/skills/eks-upgrade-readiness, .github/skills/eks-upgrade-readiness and .opencode/skills/eks-upgrade-readiness in your project.
Going by SKILL.md and its folder, Eks Upgrade Readiness needs the command-line tools its instructions call (kubectl, aws, jq, helm, terraform and pulumi).
SKILL.md names 3 domains. As links in the text: docs.aws.amazon.com, kubernetes.io and github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Eks Upgrade Readiness is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.7k tokens (SKILL.md is roughly 31k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 20k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Eks Upgrade Readiness: Kcli Cluster Deployment (karmab/kcli, 653 stars), Aspire Deployment (CommunityToolkit/Aspire, 629 stars), Deployment Automation (aiskillstore/marketplace, 430 stars) and Discover Infra (rand/cc-polymath, 181 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
aws (a GitHub organization, an official publisher) maintains it in aws/tools-for-devops-agent, which has 102 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 8, 2026.
Source: aws/tools-for-devops-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.