Official agent skill

Msk Operations

by aws in aws/tools-for-devops-agent

Amazon MSK Provisioned operations, troubleshooting, and health assessment for Standard and Express brokers.

OfficialApache-2.0Auto-check passedBackend & APIs

Install Msk Operations

skills CLI
$ npx skills add aws/tools-for-devops-agent --skill msk-operations -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aws/tools-for-devops-agent msk-operations --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/msk-operations .claude/skills/msk-operations && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
msk-operations
GitHub stars
100
Token cost
~6.6k tokens
SKILL.md length
2,510 words
Files
15 (incl. references)
Skills in repo
31
Repo updated
First seen
Licence
Apache-2.0

At a glance

Amazon MSK Provisioned operations, troubleshooting, and health assessment for Standard and Express brokers.

  • Works in 8 steps: Identify Target Clusters → Determine Broker Type Per Cluster → Collect Cluster Configuration → …
  • The user mentions Amazon MSK
  • SKILL.md covers When to Use, Broker Type Determination, Critical Warnings and Quick Diagnostics, plus 6 more sections
  • Calls aws

What it does

Msk Operations is an agent skill from aws/tools-for-devops-agent, published by the product's own GitHub organization. Amazon MSK Provisioned operations, troubleshooting, and health assessment for Standard and Express brokers. Use whenever the user mentions Amazon MSK, MSK Provisioned, MSK Standard/Express brokers, Apache Kafka on AWS, kafka. / express. instance types, or the AWS/Kafka CloudWatch namespace. Covers MSK performance issues (high CPU, produce/fetch latency, TrafficShaping), consumer lag, storage/EBS issues, rolling restarts, Kafka version upgrades, SECURITYPATCHING, BROKERUPDATE, CloudWatch alarm design, Kafka client…

Its SKILL.md is about 6.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 16 other files, including reference files (for example `.skilleval.yaml`, `CHANGELOG.md` and `README.md`).

It sits in Backend & APIs, covering Event-driven systems, Infrastructure as code and Code migrations. It works with Apache Kafka, Amazon Web Services, Amazon DynamoDB and AWS CloudFormation. The repository describes itself as: Open-source tools for AWS DevOps Agent - extend DevOps Agent with ready-to-use skills, custom agents, and other tools, for incident response, root cause analysis, and operational…. The licence is Apache-2.0.

When your agent uses it

  • The user mentions Amazon MSK
  • MSK Provisioned
  • MSK Standard/Express brokers
  • Apache Kafka on AWS

Example prompts

  • “/msk-operations”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Identify Target Clusters
  2. Determine Broker Type Per Cluster
  3. Collect Cluster Configuration
  4. Detect Monitoring Level and Gaps
  5. Collect CloudWatch Metrics (7-Day Historical)
  6. Alarm Coverage
  7. Analyze Against Best Practices
  8. Generate the Report

What it can do on your machine

Read from SKILL.md and the folder at commit ddda70b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • aws

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.aws.amazon.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Msk Operations loads about 6.6k tokens when it runs, and up to ~22k if it reads all its reference files. Until then it costs about 238 tokens; SKILL.md has 2,510 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~238
When it runs · the whole SKILL.md, loaded when a task matches
~6.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~22k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aws/tools-for-devops-agent at commit ddda70b, republished under its Apache-2.0 licence (© aws). 2,510 words, ~6,615 tokens.

Download SKILL.mdSave it as .claude/skills/msk-operations/SKILL.md (or your agent's skills folder). This skill also uses 14 other files; get the full folder from GitHub.
name
msk-operations
description
Amazon MSK Provisioned operations, troubleshooting, and health assessment for Standard and Express brokers. Use whenever the user mentions Amazon MSK, MSK Provisioned, MSK Standard/Express brokers, Apache Kafka on AWS, `kafka.*` / `express.*` instance types, or the `AWS/Kafka` CloudWatch namespace. Covers MSK performance issues (high CPU, produce/fetch latency, TrafficShaping), consumer lag, storage/EBS issues, rolling restarts, Kafka version upgrades, SECURITY_PATCHING, BROKER_UPDATE, CloudWatch alarm design, Kafka client (producer/consumer) tuning, under-replicated partitions, unexpected broker reboots, and full MSK operational reviews / health checks / best-practices audits. Do NOT use for MSK Connect, MSK Serverless, or MSK Replicator. Do NOT use for authoring CloudFormation, CDK, or Terraform templates. Do NOT use for other AWS services (RDS, Aurora, S3, DynamoDB, Kinesis, Lambda, EC2) unless MSK is explicitly named.
metadata.author
kjjanaki
metadata.version
1.0.2
metadata.aws-devops-agent-skills.agent-t
Chat tasks, Evaluation
metadata.aws-devops-agent-skills.aws-ser
Amazon MSK
metadata.aws-devops-agent-skills.technic
Analytics

Amazon MSK Operations

Operate, troubleshoot, and assess Amazon MSK (Managed Streaming for Apache Kafka) Provisioned clusters — both Standard and Express broker types. This skill covers day-to-day operations (health assessments, monitoring setup) and ad-hoc incident response (performance degradation, consumer lag, storage full, unexpected broker reboots).

When to Use

Activate this skill when the user asks to:

  • Review, audit, or assess an MSK cluster for best practices, health, or operational readiness.
  • Troubleshoot an MSK cluster problem: high CPU, high produce/fetch latency, consumer lag, broker storage running out, TrafficShaping events, under-replicated partitions, or an unexpected broker restart.
  • Set up MSK monitoring: choose a monitoring level, create recommended CloudWatch alarms and dashboards, understand the metrics available in the AWS/Kafka namespace.
  • Plan an MSK maintenance event: rolling restart, Kafka version upgrade, security patching, broker instance type change.
  • Advise on Kafka client (producer / consumer) configuration when the client is connecting to an MSK cluster.

Do not activate this skill for MSK Connect, MSK Serverless, or MSK Replicator — those are separate services with their own operational surfaces.

Broker Type Determination

Determine the broker type first — many checks differ between Standard and Express.

aws kafka describe-cluster-v2 --cluster-arn <cluster-arn>

Check ClusterInfo.Provisioned.BrokerNodeGroupInfo.InstanceType:

  • Starts with kafka. (e.g. kafka.m5.large, kafka.m7g.xlarge) → Standard broker.
  • Starts with express. (e.g. express.m7g.large) → Express broker.
Key Standard vs Express differences

Standard brokers use customer-managed EBS volumes for storage. You choose instance types (kafka.m5.*, kafka.m7g.*), provision EBS, and manage storage scaling. Standard brokers have scheduled maintenance windows.

Express brokers provide fully managed, pay-as-you-go storage with no EBS provisioning. Instance types are prefixed with express.m7g.*. Express brokers offer up to 3× more throughput per broker than Standard, and have no maintenance windows. Express enforces a fixed replication factor of 3 and min.insync.replicas=2 — you cannot create topics with RF=1.

Critical Warnings

  • This skill is read-only. Every command in this file and in references/ that mutates cluster state — update-broker-storage, create-configuration, update-cluster-configuration, update-monitoring, put-metric-alarm, reboot-broker, and any partition reassignment — is a recommendation for the operator to run after review. Present these as proposed remediations with expected impact and preconditions; do NOT execute them, and do NOT imply that the agent will run them.
  • NEVER reboot brokers while UnderReplicatedPartitions > 0 (Standard only — Express brokers do not emit URP). This risks data loss and extended outages.
  • NEVER recommend partition reassignment without first checking replication status. Reassignment during URP compounds the problem.
  • linger.ms=0 is the #1 cause of "high CPU" on MSK. ALWAYS check client batch configuration before recommending broker scaling.
  • EBS throughput ceilings are invisible in Kafka metrics — ALWAYS check EBS volume metrics (VolumeReadBytes, VolumeWriteBytes, VolumeQueueLength) when diagnosing Standard broker latency.
  • Express brokers have NO customer-managed EBS — do NOT recommend EBS expansion or provisioned EBS throughput for Express clusters.
  • Express brokers enforce fixed RF=3 and min.insync.replicas=2 — do NOT attempt to create topics with RF=1 on Express. If RF=1 is needed, use Standard brokers.

Quick Diagnostics

These five checks cover the most common MSK issues. Use them before loading a reference file.

  1. CpuUser + CpuSystem > 60%: Check RequestHandlerAvgIdlePercent (PER_BROKER monitoring level). If < 30%, request threads are saturated. Check client batch.size and linger.ms before recommending scaling.

  2. KafkaDataLogsDiskUsed > 85% (Standard only): Recommend to the operator that EBS be expanded via aws kafka update-broker-storage (do not execute). Identify high-growth topics via per-topic BytesInPerSec to size the increase. Express clusters use StorageUsed metric instead and storage is fully managed.

  3. UnderReplicatedPartitions > 0 (Standard only): Check if a maintenance operation or broker restart is in progress. If URP is decreasing, wait for recovery. Do NOT restart brokers or reassign partitions during URP. Express brokers do not emit this metric — monitor ProduceThrottleTime, FetchThrottleTime, and consumer lag instead.

  4. Consumer OffsetLag / MaxOffsetLag increasing: Determine if broker-side (high ProduceTotalTimeMsMean, CPU saturation) or client-side (slow processing, insufficient consumers). Per-partition lag from PER_TOPIC_PER_PARTITION monitoring level helps isolate hot partitions.

  5. BytesInPerSec near throughput ceiling: For Standard, check EBS volume type and calculate: BytesInPerSec × ReplicationFactor vs volume throughput limit. For Express, check against the per-broker sustained performance limits in the MSK quotas.

Which Reference Do You Need?

Route to a reference file based on the customer intent. Read the reference in full before answering — do not paraphrase from memory.

Customer IntentReference
High CPU, high produce/fetch latency, slow cluster, TrafficShapingreferences/troubleshoot-performance.md
Consumer lag increasing, rebalance storms, stuck consumer groupsreferences/troubleshoot-consumer-lag.md
Disk filling up, retention planning, tiered storage, EBS scalingreferences/manage-storage.md
Setting up monitoring level, dashboards, recommended CloudWatch alarmsreferences/monitor-and-alarm.md
Rolling restart impact, patching, Kafka version upgrades, maintenance resiliencereferences/maintenance-operations.md
Producer / consumer configuration, IAM / SCRAM / TLS auth for clientsreferences/configure-clients.md

For sizing questions (broker count, instance type choice, monthly cost), refer the user to the Amazon MSK best practices — right-size your cluster documentation. Do not size from memory.

Operational Review Workflow

Use this workflow when the user asks for a review, audit, health check, or assessment of an MSK cluster. The routing table above handles ad-hoc troubleshooting; this section produces a consistent, comprehensive report.

Follow the steps in order for each target cluster. Do not skip steps. If a step cannot be completed (e.g. a metric requires a higher monitoring level than the cluster has enabled), record the gap in the report rather than silently omitting the check.

Step 1 — Identify Target Clusters

Ask the user which MSK clusters to review. Accept any of:

  • Specific cluster names or ARNs and regions
  • "all clusters" in specific regions
  • "all MSK clusters in all regions"

If no scope is given, default to all configured account regions. Enumerate clusters with aws kafka list-clusters-v2 per region.

Step 2 — Determine Broker Type Per Cluster

For each cluster:

aws kafka describe-cluster-v2 --cluster-arn <cluster-arn>

Read ClusterInfo.Provisioned.BrokerNodeGroupInfo.InstanceType. Standard brokers (kafka.*) and Express brokers (express.*) require different checks in the later steps — some metrics only exist on one type.

Step 3 — Collect Cluster Configuration

For each cluster, gather:

aws kafka describe-cluster-v2 --cluster-arn <arn>
aws kafka list-nodes --cluster-arn <arn>
aws kafka get-bootstrap-brokers --cluster-arn <arn>
aws kafka list-cluster-operations-v2 --cluster-arn <arn>   # last 30 days
aws kafka describe-configuration-revision \
     --arn <configuration-arn> --revision <revision>       # if a custom config is applied

Capture:

  • Cluster: state, Kafka version, number of broker nodes, AZ distribution (ZoneIds), storage mode (EBS / Tiered), current version.
  • Broker: instance type, EBS volume size (Standard), provisioned throughput (if any).
  • Encryption: EncryptionInTransit.ClientBroker (TLS / TLS_PLAINTEXT / PLAINTEXT), EncryptionInTransit.InCluster, EncryptionAtRest.DataVolumeKMSKeyId.
  • Auth: ClientAuthentication.Sasl.Iam.Enabled, ClientAuthentication.Sasl.Scram.Enabled, ClientAuthentication.Tls.Enabled, ClientAuthentication.Unauthenticated.Enabled.
  • Public access: BrokerNodeGroupInfo.ConnectivityInfo.PublicAccess.Type.
  • Monitoring level: EnhancedMonitoring (DEFAULT / PER_BROKER / PER_TOPIC_PER_BROKER / PER_TOPIC_PER_PARTITION).
  • Logging: LoggingInfo.BrokerLogs (CloudWatch / S3 / Firehose destinations and their Enabled flags).
  • Open monitoring: OpenMonitoring.Prometheus.JmxExporter.EnabledInBroker, NodeExporter.EnabledInBroker.
  • Recent operations: From list-cluster-operations-v2, note any SECURITY_PATCHING, BROKER_UPDATE, UPDATE_CLUSTER_CONFIGURATION, UPDATE_STORAGE, or UPDATE_MONITORING events in the review window.
Step 4 — Detect Monitoring Level and Gaps

The EnhancedMonitoring value from Step 3 determines which checks are available. At DEFAULT, most per-broker health metrics are still available (CPU, disk, network, partitions, connections, memory, TrafficShaping), but the following checks are not possible without upgrading:

  • ReplicationBytesInPerSec / ReplicationBytesOutPerSec (inter-broker replication load)
  • RequestHandlerAvgIdlePercent / NetworkProcessorAvgIdlePercent (thread pool saturation)
  • BwInAllowanceExceeded / BwOutAllowanceExceeded (detailed bandwidth breaches)
  • VolumeQueueLength (EBS I/O queue depth — Standard only)
  • VolumeReadBytes / VolumeWriteBytes (EBS throughput utilization — Standard only)
  • IAM connection metrics (IAMTooManyConnections)

Record the monitoring level and list any dimensions that will be scored partially or skipped. Recommend upgrading to PER_BROKER if any dimension is degraded by the current level.

Step 5 — Collect CloudWatch Metrics (7-Day Historical)

Namespace: AWS/Kafka. Dimensions: Cluster Name and Broker ID for per-broker metrics; Cluster Name and Consumer Group and Topic for consumer lag; Cluster Name only for cluster-wide metrics.

Use one cloudwatch.GetMetricData batch per cluster where possible. Period: 3600 (1 hour). StartTime: 7 days ago. EndTime: now.

5.1 Cluster-wide (both broker types)
MetricStatPurpose
ActiveControllerCountSumMust be exactly 1
OfflinePartitionsCountMaximumMust be 0
GlobalPartitionCountMaximumTotal leader partitions
GlobalTopicCountMaximumTotal topics
5.2 Per-broker — Standard (kafka.*)
MetricStatThreshold
CpuUser + CpuSystemAverage, Maximum< 60% avg
KafkaDataLogsDiskUsedAverage, Maximum< 70% avg, < 85% max
PartitionCountMaximum≤ recommended limit for broker size
LeaderCountMaximumCompare across brokers; skew < 10%
UnderReplicatedPartitionsMaximum0 in steady state
UnderMinIsrPartitionCountMaximum0
BytesInPerSec / BytesOutPerSecAverage, Maximumvs baseline bandwidth
ConnectionCountAverage, MaximumCompare across brokers
HeapMemoryAfterGCMaximum< 60%
TrafficShapingSumMust be 0
NetworkRxDropped / NetworkTxDropped / NetworkRxErrors / NetworkTxErrorsSumMust be 0
ReplicationBytesInPerSec / ReplicationBytesOutPerSecAverageRequires PER_BROKER
RequestHandlerAvgIdlePercent / NetworkProcessorAvgIdlePercentAverage> 30% (PER_BROKER)
VolumeQueueLengthAverage, MaximumAvg < 1 (PER_BROKER)
VolumeReadBytes + VolumeWriteBytesSumvs EBS baseline throughput (PER_BROKER)
5.3 Per-broker — Express (express.*)
MetricStatThreshold
CpuUser + CpuSystemAverage, Maximum< 60% avg
StorageUsedMaximumFully managed — flag if trending against per-broker quota
PartitionCountMaximum≤ recommended limit for broker size
LeaderCountMaximumCompare across brokers
BytesInPerSec / BytesOutPerSecAverage, Maximumvs Express per-broker ingress/egress quotas
ProduceThrottleTime / FetchThrottleTimeMaximumShould be 0
ClientConnectionCountAverage, Maximumvs listener quota (see references/monitor-and-alarm.md)

Express brokers do NOT emit: UnderReplicatedPartitions, UnderMinIsrPartitionCount, HeapMemoryAfterGC, TrafficShaping, Volume* metrics, KafkaDataLogsDiskUsed, ProduceMessageConversionsPerSec, FetchMessageConversionsPerSec.

5.4 Consumer lag (both broker types)

Per consumer group (identified from the customer or from the broker logs):

MetricStatNotes
SumOffsetLag per (Consumer Group, Topic)MaximumDEFAULT level
MaxOffsetLag per (Consumer Group, Topic)MaximumDEFAULT level
EstimatedMaxTimeLag per (Consumer Group, Topic)MaximumDEFAULT level
OffsetLag per (Consumer Group, Topic, Partition)MaximumRequires PER_TOPIC_PER_PARTITION
Step 6 — Alarm Coverage

Collect existing alarms:

aws cloudwatch describe-alarms --namespace AWS/Kafka

Compare against the 13-alarm recommended set (details in references/monitor-and-alarm.md) and produce a coverage table (present / missing / firing).

Step 7 — Analyze Against Best Practices

Assign a severity to every finding: CRITICAL / HIGH / MEDIUM / LOW / INFO (see Severity Definitions later in this file).

Evaluate across seven dimensions. Some checks are skipped for Express — noted inline.

Show full SKILL.md (1,021 more words)Show less
7.1 Cluster Configuration
  • Cluster state is ACTIVE. MAINTENANCE / UPDATING is transient; anything else is a finding.
  • Deployed across 3 AZs (Standard: broker count multiple of AZ count).
  • Kafka version within N-2 of the latest supported.
  • Enhanced monitoring at PER_BROKER or higher — DEFAULT → MEDIUM.
  • Storage mode: Tiered storage enabled for topics with long retention (Standard).
7.2 Security
  • Encryption in transit: TLS. TLS_PLAINTEXT → MEDIUM (prod: HIGH). PLAINTEXT → CRITICAL.
  • Encryption at rest: enabled with KMS. Customer-managed KMS preferred over AWS-managed for regulated workloads.
  • Authentication: at least one of IAM / SCRAM / mTLS enabled. Unauthenticated.Enabled=true in production → CRITICAL.
  • Public access: SERVICE_PROVIDED_EIPS in production → HIGH.
7.3 Logging & Monitoring
  • Broker logs enabled to at least one destination (CloudWatch / S3 / Firehose). All disabled → HIGH.
  • Open monitoring (Prometheus JMX + Node exporter) — informational.
  • Alarm coverage from Step 6: missing alarms → MEDIUM each; missing critical alarms (Active Controller, Offline Partitions, Disk, CPU) → HIGH.
  • Any alarm currently in ALARM state → HIGH (surface in report header).
7.4 Partition Health
  • PartitionCount per broker ≤ recommended limit for the broker instance type. Above recommended but below max → MEDIUM. Above max → HIGH (blocks update operations).
  • LeaderCount variance across brokers < 10%. 10-25% → MEDIUM. > 25% → HIGH.
  • UnderReplicatedPartitions > 0 sustained (Standard only) → HIGH. Transient during a SECURITY_PATCHING / BROKER_UPDATE operation from Step 3 → INFO — do NOT flag.
  • UnderMinIsrPartitionCount > 0 (Standard only) → CRITICAL.
  • Config-level ISR risk: topics with min.insync.replicas >= replication.factor are a configuration error the CloudWatch metric will not surface. If the user has provided topic configs, flag any such topic as HIGH.
7.5 Compute Health
  • CpuUser + CpuSystem avg < 60%. 60-70% → MEDIUM. > 70% → HIGH.
  • CPU variance across brokers < 20%. 20-50% → MEDIUM. > 50% → HIGH.
  • HeapMemoryAfterGC < 60% (Standard only). 60-80% → MEDIUM. > 80% → HIGH.
  • RequestHandlerAvgIdlePercent > 30% (PER_BROKER). 10-30% → MEDIUM. < 10% → HIGH.
  • NetworkProcessorAvgIdlePercent > 30% (PER_BROKER). 10-30% → MEDIUM. < 10% → HIGH.
7.6 Network Health
  • TrafficShaping = 0 (Standard only). Any non-zero → HIGH.
  • BwInAllowanceExceeded / BwOutAllowanceExceeded = 0 (PER_BROKER). Non-zero → HIGH.
  • Throughput variance across brokers < 20%. 20-50% → MEDIUM. > 50% → HIGH.
  • Total per-broker throughput < 60% of baseline bandwidth (see references/troubleshoot-performance.md for baseline table). 60-70% → MEDIUM. > 70% → HIGH.
  • Express: ProduceThrottleTime / FetchThrottleTime > 0 → HIGH.
7.7 Storage Health

Standard only (Express storage is managed — Express clusters only get the StorageUsed check):

  • KafkaDataLogsDiskUsed < 70%. 70-85% → HIGH. > 85% → CRITICAL.
  • VolumeQueueLength avg < 1 (PER_BROKER). 1-5 → MEDIUM. > 5 → HIGH.
  • EBS auto-scaling configured, or disk headroom > 30% → PASS. No auto-scaling AND disk > 50% → MEDIUM. No auto-scaling AND disk > 70% → HIGH.
  • EBS throughput (VolumeReadBytes + VolumeWriteBytes) < 60% of instance baseline (PER_BROKER). 60-70% → MEDIUM. > 70% → HIGH.

Express:

  • StorageUsed per broker vs Express per-broker storage quota (see MSK Express quotas). Approaching quota → MEDIUM.
Step 8 — Generate the Report

Generate a separate report artifact per cluster reviewed.

Artifact naming: msk-review-<cluster-name>-<YYYY-MM-DD>.md Example: msk-review-prod-orders-2026-04-29.md

Report structure:

Report Header
# MSK Operational Review — <cluster-name>
Account: <account-id> | Region: <region> | Date: <YYYY-MM-DD>
Broker Type: Standard/Express | Instance Type: <type> | Broker Count: <n> | AZs: <n>
Kafka Version: <version> | Monitoring Level: <level>
Executive Summary
  • Health: ✅ HEALTHY / ⚠️ WARNINGS / ❌ CRITICAL
  • Finding counts by severity
  • Top 3 CRITICAL/HIGH items
Configuration Snapshot

| Item | Value | | Cluster state / version | … | | Broker type / instance / count / AZs | … | | Storage | mode, size (Standard), provisioned throughput (if any) | | Encryption | in-transit, in-cluster, at-rest (KMS) | | Authentication | IAM / SCRAM / mTLS / Unauthenticated flags | | Public access | DISABLED / SERVICE_PROVIDED_EIPS | | Monitoring level | DEFAULT / PER_BROKER / PER_TOPIC_PER_BROKER / PER_TOPIC_PER_PARTITION | | Logging | destinations enabled |

Findings by Dimension

For each of the 7 dimensions (7.1-7.7):

| # | Finding | Severity | Current State | Recommendation |

If a dimension was skipped or partial due to monitoring level or broker type, say so explicitly in a note above the table for that dimension.

CloudWatch Metrics (7-Day)

| Metric | Stat | 7-Day Avg | 7-Day Max | Status | Finding |

Alarm Coverage

| # | Recommended Alarm | Metric | Threshold | Priority | Status |

Also list alarms currently in ALARM state with timestamps.

Recent Cluster Operations (Last 30 Days)

From list-cluster-operations-v2 in Step 3:

| Operation Type | Start Time | End Time | State |

Call out any operations that would explain transient metric anomalies.

Priority Matrix

| # | Finding | Severity | Dimension | Effort | Impact |

Sorted by severity.

Next Steps
  • Immediate (CRITICAL / HIGH — 7 days)
  • Short-term (MEDIUM — 30 days)
  • Long-term (LOW — 90 days)

Common CLI Recipes

Describe cluster:

aws kafka describe-cluster-v2 --cluster-arn <cluster-arn>

List brokers:

aws kafka list-nodes --cluster-arn <cluster-arn>

Get bootstrap brokers:

aws kafka get-bootstrap-brokers --cluster-arn <cluster-arn>

List recent cluster operations (patching, config updates, storage changes):

aws kafka list-cluster-operations-v2 --cluster-arn <cluster-arn>

Expand Standard broker storage — Operator-run (recommend, do not execute):

aws kafka update-broker-storage \
  --cluster-arn <cluster-arn> \
  --current-version <cluster-version> \
  --target-broker-ebs-volume-info '[{"KafkaBrokerNodeId": "All", "VolumeSizeGB": <target-size>}]'

Get a CloudWatch metric (example: CpuUser per broker):

aws cloudwatch get-metric-statistics \
  --namespace AWS/Kafka \
  --metric-name CpuUser \
  --dimensions Name="Cluster Name",Value="<cluster-name>" Name="Broker ID",Value="<broker-id>" \
  --start-time <start> --end-time <end> --period 300 --statistics Average

Create a cluster configuration (server.properties) — Operator-run (recommend, do not execute):

The --server-properties argument MUST be a real Kafka properties file with one key=value per line, separated by actual newline characters — NOT the literal two-character escape sequence \n. The MSK API accepts the bytes as-is; if you pass "k1=v1\nk2=v2" as a single string with escaped newlines, MSK stores ONE invalid property line and the cluster will fail to apply it.

Recommended pattern: write the properties to a local file with real newlines, then pass it via fileb:// so the CLI uploads the raw bytes verbatim. Verify by reading the revision back with describe-configuration-revision and base64-decoding ServerProperties — you should see one property per line.

cat > server.properties <<'EOF'
auto.create.topics.enable=false
default.replication.factor=3
min.insync.replicas=2
unclean.leader.election.enable=false
num.io.threads=32
num.network.threads=16
log.retention.hours=168
EOF

aws kafka create-configuration \
  --name <config-name> \
  --kafka-versions "3.6.0" \
  --server-properties fileb://server.properties

Common Error Reference

ErrorCauseFix
aws kafka update-broker-storage returns "storage is optimizing"Previous storage expansion still in cool-down (minimum 6 hours)Wait for optimization to complete. Check cluster state with describe-cluster-v2.
ClusterState is MAINTENANCEStandard brokers undergoing patching. Express brokers stay ACTIVE during maintenance.Wait for cluster to return to ACTIVE. Do not perform update operations during MAINTENANCE.
Consumer receives GROUP_COORDINATOR_NOT_AVAILABLECoordinator broker is temporarily unavailable during rolling restart or overloadedRetry with backoff. Check if maintenance is in progress via list-cluster-operations-v2.
NotEnoughReplicasException on produceFewer brokers in ISR than min.insync.replicas (default: 2)Check UnderReplicatedPartitions (Standard only). For Express, check ProduceThrottleTime and broker health instead — URP is not available. If a broker is down for maintenance, this is transient. Do NOT lower min.insync.replicas to work around this.

Severity Definitions (for review-style reports)

SeverityDefinitionSLA
CRITICALImmediate risk to availability, security, or data integrityFix within 24–48 hours
HIGHSignificant gap that could lead to incidentsFix within 1 week
MEDIUMNotable improvement opportunityPlan within 30 days
LOWMinor optimization or hardeningAddress when convenient
INFOObservation, no action requiredN/A

Additional Resources

© aws, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 14 other files (references) in skills/msk-operations of aws/tools-for-devops-agent.

  • SKILL.md
  • .skilleval.yaml
  • CHANGELOG.md
  • README.md
  • evals/benchmark.json
  • evals/eval_queries.json
  • evals/evals.json
  • evals/report.json
  • evals/trigger_report.json
  • references/configure-clients.md
  • references/maintenance-operations.md
  • references/manage-storage.md
  • references/monitor-and-alarm.md
  • references/troubleshoot-consumer-lag.md
  • references/troubleshoot-performance.md

Open the folder on GitHubat commit ddda70b

Compare with similar skills

Msk Operations next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Msk Operations compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Msk Operations this skillaws/tools-for-devops-agent100—~6.6kAutomated safety check: PassApache-2.0
AWS Advisordiegosouzapw/awesome-omni-skills159—~4.3kAutomated safety check: PassMIT
Amazon Elasticacheaws/agent-toolkit-for-aws2.8k—~4.5kAutomated safety check: PassApache-2.0
AWS Solution Architectborghei/Claude-Skills881—~1.8kAutomated safety check: PassMIT
AWS Solution Architectalirezarezvani/claude-skills28k1 repos~2.5kAutomated safety check: PassMIT
New Event Sourceaws/aws-lambda-dotnet1.7k—~3kAutomated safety check: PassApache-2.0

Similar skills

  • AWS Advisor

    diegosouzapw/awesome-omni-skills

    AWS Advisor workflow skill. An agent skill from diegosouzapw/awesome-omni-skills.

    159 GitHub stars~4.3k tokensUpdated 3 mo ago
    DevOps & CloudAuto-check passed
  • Amazon Elasticache

    aws/agent-toolkit-for-aws

    Official

    Activate when developers have latent caching needs: slow API responses, database read bottlenecks, DynamoDB throttling or cost, RDS/Aurora scaling pressure, Bedrock latency or cost, or adding a…

    2.8k GitHub stars~4.5k tokensUpdated today
    Backend & APIsAuto-check passed
  • AWS Solution Architect

    borghei/Claude-Skills

    Design AWS serverless architectures for startups with IaC. An agent skill from borghei/Claude-Skills.

    881 GitHub stars~1.8k tokensUpdated today
    DevOps & CloudAuto-check passed
  • AWS Solution Architect

    alirezarezvani/claude-skills

    Design AWS architectures for startups using serverless patterns and IaC templates.

    28k GitHub starsUsed in 1 repo~2.5k tokens
    Backend & APIsAuto-check passed
  • New Event Source

    aws/aws-lambda-dotnet

    Official

    Add a new AWS event source attribute (e.g., Kinesis, Kafka, MQ) to the Lambda .NET Annotations framework, including the attribute class, source generator integration, CloudFormation writer, unit…

    1.7k GitHub stars~3k tokensUpdated today
    Testing & QAAuto-check passed
  • AWS Cloud Patterns

    rohitg00/awesome-claude-code-toolkit

    AWS cloud patterns for Lambda, ECS, S3, DynamoDB, and Infrastructure as Code with CDK/Terraform

    2.7k GitHub stars~1.1k tokensUpdated 4 mo ago
    DevOps & CloudAuto-check passed

More from aws/tools-for-devops-agent

All 31 skills in this repo
  • Sagemaker AI Ops Review

    aws/tools-for-devops-agent

    Official

    Amazon SageMaker AI Operational Review. An agent skill from aws/tools-for-devops-agent.

    100 GitHub starsUsed in 1 repo~3.9k tokens
    Auto-check passed
  • Aiml GPU Training Cluster Investigation

    aws/tools-for-devops-agent

    Official

    A skill your agent uses for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances.

    100 GitHub stars~5.4k tokensUpdated today
    Auto-check passed
  • AWS Health Events

    aws/tools-for-devops-agent

    Official

    ALWAYS use this skill in the beginning of any incident investigation, root cause analysis, or operational troubleshooting.

    100 GitHub stars~4.6k tokensUpdated today
    Auto-check passed
  • Database Migration Service Expertise

    aws/tools-for-devops-agent

    Official

    AWS Database Migration Service (DMS) operational review and troubleshooting skill.

    100 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ecs Operation Review

    aws/tools-for-devops-agent

    Official

    Performs a comprehensive Amazon ECS operations review across the 6 review pillars (Resiliency & HA, Observability, Security, Operations, Performance, Additional Analysis) using read-only AWS APIs…

    100 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Rds Operation Review

    aws/tools-for-devops-agent

    Official

    Comprehensive Amazon RDS and Aurora operational review aligned with the AWS Well-Architected Framework and RDS/Aurora best practices.

    100 GitHub stars~4.8k tokensUpdated today
    Auto-check passed

Questions about Msk Operations

What does Msk Operations do?

Amazon MSK Provisioned operations, troubleshooting, and health assessment for Standard and Express brokers. Msk Operations is an agent skill from aws/tools-for-devops-agent, published by the product's own GitHub organization. Amazon MSK Provisioned operations, troubleshooting, and health assessment for Standard and Express brokers.

When should I use Msk Operations?

Msk Operations fits situations like: the user mentions Amazon MSK; MSK Provisioned; MSK Standard/Express brokers; apache Kafka on AWS.

How do I install Msk Operations in Claude Code?

Run `npx skills add aws/tools-for-devops-agent --skill msk-operations -a claude-code`. Or copy the skill folder (skills/msk-operations in aws/tools-for-devops-agent) into .claude/skills/msk-operations in your project. Claude Code loads it when a task matches its description.

How do I install Msk Operations in Codex?

Run `npx skills add aws/tools-for-devops-agent --skill msk-operations -a codex`. Or copy the skill folder (skills/msk-operations in aws/tools-for-devops-agent) into .agents/skills/msk-operations in your project. Codex loads it when a task matches its description.

Can I use Msk Operations in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aws/tools-for-devops-agent --skill msk-operations -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/msk-operations, .gemini/skills/msk-operations, .github/skills/msk-operations and .opencode/skills/msk-operations in your project.

What does Msk Operations need to run?

Going by SKILL.md and its folder, Msk Operations needs the command-line tools its instructions call (aws).

Does Msk Operations access the network?

SKILL.md names 1 domain. As links in the text: docs.aws.amazon.com. This is read from the text; nothing was executed.

Is Msk Operations safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Msk Operations use?

Msk Operations is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Msk Operations use?

About 6.6k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 15k tokens, read only when the agent opens those files.

What are the alternatives to Msk Operations?

Skills that share tags, products or a category with Msk Operations: AWS Advisor (diegosouzapw/awesome-omni-skills, 159 stars), Amazon Elasticache (aws/agent-toolkit-for-aws, 2.8k stars), AWS Solution Architect (borghei/Claude-Skills, 881 stars) and AWS Solution Architect (alirezarezvani/claude-skills, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Msk Operations?

aws (a GitHub organization, an official publisher) maintains it in aws/tools-for-devops-agent, which has 100 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 8, 2026.

Source: aws/tools-for-devops-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.