AWS Advisor
diegosouzapw/awesome-omni-skills
AWS Advisor workflow skill. An agent skill from diegosouzapw/awesome-omni-skills.
Amazon MSK Provisioned operations, troubleshooting, and health assessment for Standard and Express brokers.
$ npx skills add aws/tools-for-devops-agent --skill msk-operations -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install aws/tools-for-devops-agent msk-operations --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/msk-operations .claude/skills/msk-operations && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "msk-operations" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/msk-operations into .claude/skills/msk-operations/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "msk-operations", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/aws/tools-for-devops-agent/tree/main/skills/msk-operationsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add aws/tools-for-devops-agent --skill msk-operations -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install aws/tools-for-devops-agent msk-operations --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/msk-operations .agents/skills/msk-operations && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "msk-operations" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/msk-operations into .agents/skills/msk-operations/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "msk-operations", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aws/tools-for-devops-agent --skill msk-operations -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install aws/tools-for-devops-agent msk-operations --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/msk-operations .cursor/skills/msk-operations && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "msk-operations" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/msk-operations into .cursor/skills/msk-operations/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "msk-operations", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/aws/tools-for-devops-agent.git --path skills/msk-operations--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add aws/tools-for-devops-agent --skill msk-operations -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install aws/tools-for-devops-agent msk-operations --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/msk-operations .gemini/skills/msk-operations && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "msk-operations" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/msk-operations into .gemini/skills/msk-operations/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "msk-operations", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install aws/tools-for-devops-agent msk-operationsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add aws/tools-for-devops-agent --skill msk-operations -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/msk-operations .github/skills/msk-operations && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "msk-operations" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/msk-operations into .github/skills/msk-operations/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "msk-operations", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aws/tools-for-devops-agent --skill msk-operations -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install aws/tools-for-devops-agent msk-operations --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/msk-operations .opencode/skills/msk-operations && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "msk-operations" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/msk-operations into .opencode/skills/msk-operations/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "msk-operations", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
msk-operationsAmazon MSK Provisioned operations, troubleshooting, and health assessment for Standard and Express brokers.
Msk Operations is an agent skill from aws/tools-for-devops-agent, published by the product's own GitHub organization. Amazon MSK Provisioned operations, troubleshooting, and health assessment for Standard and Express brokers. Use whenever the user mentions Amazon MSK, MSK Provisioned, MSK Standard/Express brokers, Apache Kafka on AWS, kafka. / express. instance types, or the AWS/Kafka CloudWatch namespace. Covers MSK performance issues (high CPU, produce/fetch latency, TrafficShaping), consumer lag, storage/EBS issues, rolling restarts, Kafka version upgrades, SECURITYPATCHING, BROKERUPDATE, CloudWatch alarm design, Kafka client…
Its SKILL.md is about 6.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 16 other files, including reference files (for example `.skilleval.yaml`, `CHANGELOG.md` and `README.md`).
It sits in Backend & APIs, covering Event-driven systems, Infrastructure as code and Code migrations. It works with Apache Kafka, Amazon Web Services, Amazon DynamoDB and AWS CloudFormation. The repository describes itself as: Open-source tools for AWS DevOps Agent - extend DevOps Agent with ready-to-use skills, custom agents, and other tools, for incident response, root cause analysis, and operational…. The licence is Apache-2.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit ddda70b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
awsFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.aws.amazon.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Msk Operations loads about 6.6k tokens when it runs, and up to ~22k if it reads all its reference files. Until then it costs about 238 tokens; SKILL.md has 2,510 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from aws/tools-for-devops-agent at commit ddda70b, republished under its Apache-2.0 licence (© aws). 2,510 words, ~6,615 tokens.
.claude/skills/msk-operations/SKILL.md (or your agent's skills folder). This skill also uses 14 other files; get the full folder from GitHub.Operate, troubleshoot, and assess Amazon MSK (Managed Streaming for Apache Kafka) Provisioned clusters — both Standard and Express broker types. This skill covers day-to-day operations (health assessments, monitoring setup) and ad-hoc incident response (performance degradation, consumer lag, storage full, unexpected broker reboots).
Activate this skill when the user asks to:
AWS/Kafka
namespace.Do not activate this skill for MSK Connect, MSK Serverless, or MSK Replicator — those are separate services with their own operational surfaces.
Determine the broker type first — many checks differ between Standard and Express.
aws kafka describe-cluster-v2 --cluster-arn <cluster-arn>Check ClusterInfo.Provisioned.BrokerNodeGroupInfo.InstanceType:
kafka. (e.g. kafka.m5.large, kafka.m7g.xlarge) → Standard broker.express. (e.g. express.m7g.large) → Express broker.Standard brokers use customer-managed EBS volumes for storage. You choose
instance types (kafka.m5.*, kafka.m7g.*), provision EBS, and manage storage
scaling. Standard brokers have scheduled maintenance windows.
Express brokers provide fully managed, pay-as-you-go storage with no EBS
provisioning. Instance types are prefixed with express.m7g.*. Express brokers
offer up to 3× more throughput per broker than Standard, and have no maintenance
windows. Express enforces a fixed replication factor of 3 and
min.insync.replicas=2 — you cannot create topics with RF=1.
references/
that mutates cluster state — update-broker-storage, create-configuration,
update-cluster-configuration, update-monitoring, put-metric-alarm,
reboot-broker, and any partition reassignment — is a recommendation for
the operator to run after review. Present these as proposed remediations
with expected impact and preconditions; do NOT execute them, and do NOT
imply that the agent will run them.UnderReplicatedPartitions > 0 (Standard only —
Express brokers do not emit URP). This risks data loss and extended outages.linger.ms=0 is the #1 cause of "high CPU" on MSK. ALWAYS check client
batch configuration before recommending broker scaling.VolumeReadBytes, VolumeWriteBytes, VolumeQueueLength)
when diagnosing Standard broker latency.min.insync.replicas=2 — do NOT
attempt to create topics with RF=1 on Express. If RF=1 is needed, use Standard
brokers.These five checks cover the most common MSK issues. Use them before loading a reference file.
CpuUser + CpuSystem > 60%: Check RequestHandlerAvgIdlePercent
(PER_BROKER monitoring level). If < 30%, request threads are saturated. Check
client batch.size and linger.ms before recommending scaling.
KafkaDataLogsDiskUsed > 85% (Standard only): Recommend to the
operator that EBS be expanded via aws kafka update-broker-storage (do
not execute). Identify high-growth topics via per-topic BytesInPerSec to
size the increase. Express clusters use StorageUsed metric instead and
storage is fully managed.
UnderReplicatedPartitions > 0 (Standard only): Check if a maintenance
operation or broker restart is in progress. If URP is decreasing, wait for
recovery. Do NOT restart brokers or reassign partitions during URP. Express
brokers do not emit this metric — monitor ProduceThrottleTime,
FetchThrottleTime, and consumer lag instead.
Consumer OffsetLag / MaxOffsetLag increasing: Determine if broker-side
(high ProduceTotalTimeMsMean, CPU saturation) or client-side (slow
processing, insufficient consumers). Per-partition lag from
PER_TOPIC_PER_PARTITION monitoring level helps isolate hot partitions.
BytesInPerSec near throughput ceiling: For Standard, check EBS volume
type and calculate: BytesInPerSec × ReplicationFactor vs volume throughput
limit. For Express, check against the per-broker sustained performance limits
in the MSK quotas.
Route to a reference file based on the customer intent. Read the reference in full before answering — do not paraphrase from memory.
| Customer Intent | Reference |
|---|---|
| High CPU, high produce/fetch latency, slow cluster, TrafficShaping | references/troubleshoot-performance.md |
| Consumer lag increasing, rebalance storms, stuck consumer groups | references/troubleshoot-consumer-lag.md |
| Disk filling up, retention planning, tiered storage, EBS scaling | references/manage-storage.md |
| Setting up monitoring level, dashboards, recommended CloudWatch alarms | references/monitor-and-alarm.md |
| Rolling restart impact, patching, Kafka version upgrades, maintenance resilience | references/maintenance-operations.md |
| Producer / consumer configuration, IAM / SCRAM / TLS auth for clients | references/configure-clients.md |
For sizing questions (broker count, instance type choice, monthly cost), refer the user to the Amazon MSK best practices — right-size your cluster documentation. Do not size from memory.
Use this workflow when the user asks for a review, audit, health check, or assessment of an MSK cluster. The routing table above handles ad-hoc troubleshooting; this section produces a consistent, comprehensive report.
Follow the steps in order for each target cluster. Do not skip steps. If a step cannot be completed (e.g. a metric requires a higher monitoring level than the cluster has enabled), record the gap in the report rather than silently omitting the check.
Ask the user which MSK clusters to review. Accept any of:
If no scope is given, default to all configured account regions. Enumerate
clusters with aws kafka list-clusters-v2 per region.
For each cluster:
aws kafka describe-cluster-v2 --cluster-arn <cluster-arn>Read ClusterInfo.Provisioned.BrokerNodeGroupInfo.InstanceType. Standard
brokers (kafka.*) and Express brokers (express.*) require different checks
in the later steps — some metrics only exist on one type.
For each cluster, gather:
aws kafka describe-cluster-v2 --cluster-arn <arn>
aws kafka list-nodes --cluster-arn <arn>
aws kafka get-bootstrap-brokers --cluster-arn <arn>
aws kafka list-cluster-operations-v2 --cluster-arn <arn> # last 30 days
aws kafka describe-configuration-revision \
--arn <configuration-arn> --revision <revision> # if a custom config is appliedCapture:
ZoneIds), storage mode (EBS / Tiered), current version.EncryptionInTransit.ClientBroker (TLS / TLS_PLAINTEXT /
PLAINTEXT), EncryptionInTransit.InCluster, EncryptionAtRest.DataVolumeKMSKeyId.ClientAuthentication.Sasl.Iam.Enabled,
ClientAuthentication.Sasl.Scram.Enabled, ClientAuthentication.Tls.Enabled,
ClientAuthentication.Unauthenticated.Enabled.BrokerNodeGroupInfo.ConnectivityInfo.PublicAccess.Type.EnhancedMonitoring (DEFAULT / PER_BROKER /
PER_TOPIC_PER_BROKER / PER_TOPIC_PER_PARTITION).LoggingInfo.BrokerLogs (CloudWatch / S3 / Firehose destinations
and their Enabled flags).OpenMonitoring.Prometheus.JmxExporter.EnabledInBroker,
NodeExporter.EnabledInBroker.list-cluster-operations-v2, note any
SECURITY_PATCHING, BROKER_UPDATE, UPDATE_CLUSTER_CONFIGURATION,
UPDATE_STORAGE, or UPDATE_MONITORING events in the review window.The EnhancedMonitoring value from Step 3 determines which checks are
available. At DEFAULT, most per-broker health metrics are still available
(CPU, disk, network, partitions, connections, memory, TrafficShaping), but the
following checks are not possible without upgrading:
ReplicationBytesInPerSec / ReplicationBytesOutPerSec (inter-broker
replication load)RequestHandlerAvgIdlePercent / NetworkProcessorAvgIdlePercent (thread pool
saturation)BwInAllowanceExceeded / BwOutAllowanceExceeded (detailed bandwidth
breaches)VolumeQueueLength (EBS I/O queue depth — Standard only)VolumeReadBytes / VolumeWriteBytes (EBS throughput utilization — Standard
only)IAMTooManyConnections)Record the monitoring level and list any dimensions that will be scored
partially or skipped. Recommend upgrading to PER_BROKER if any dimension is
degraded by the current level.
Namespace: AWS/Kafka. Dimensions: Cluster Name and Broker ID for
per-broker metrics; Cluster Name and Consumer Group and Topic for
consumer lag; Cluster Name only for cluster-wide metrics.
Use one cloudwatch.GetMetricData batch per cluster where possible.
Period: 3600 (1 hour). StartTime: 7 days ago. EndTime: now.
| Metric | Stat | Purpose |
|---|---|---|
ActiveControllerCount | Sum | Must be exactly 1 |
OfflinePartitionsCount | Maximum | Must be 0 |
GlobalPartitionCount | Maximum | Total leader partitions |
GlobalTopicCount | Maximum | Total topics |
kafka.*)| Metric | Stat | Threshold |
|---|---|---|
CpuUser + CpuSystem | Average, Maximum | < 60% avg |
KafkaDataLogsDiskUsed | Average, Maximum | < 70% avg, < 85% max |
PartitionCount | Maximum | ≤ recommended limit for broker size |
LeaderCount | Maximum | Compare across brokers; skew < 10% |
UnderReplicatedPartitions | Maximum | 0 in steady state |
UnderMinIsrPartitionCount | Maximum | 0 |
BytesInPerSec / BytesOutPerSec | Average, Maximum | vs baseline bandwidth |
ConnectionCount | Average, Maximum | Compare across brokers |
HeapMemoryAfterGC | Maximum | < 60% |
TrafficShaping | Sum | Must be 0 |
NetworkRxDropped / NetworkTxDropped / NetworkRxErrors / NetworkTxErrors | Sum | Must be 0 |
ReplicationBytesInPerSec / ReplicationBytesOutPerSec | Average | Requires PER_BROKER |
RequestHandlerAvgIdlePercent / NetworkProcessorAvgIdlePercent | Average | > 30% (PER_BROKER) |
VolumeQueueLength | Average, Maximum | Avg < 1 (PER_BROKER) |
VolumeReadBytes + VolumeWriteBytes | Sum | vs EBS baseline throughput (PER_BROKER) |
express.*)| Metric | Stat | Threshold |
|---|---|---|
CpuUser + CpuSystem | Average, Maximum | < 60% avg |
StorageUsed | Maximum | Fully managed — flag if trending against per-broker quota |
PartitionCount | Maximum | ≤ recommended limit for broker size |
LeaderCount | Maximum | Compare across brokers |
BytesInPerSec / BytesOutPerSec | Average, Maximum | vs Express per-broker ingress/egress quotas |
ProduceThrottleTime / FetchThrottleTime | Maximum | Should be 0 |
ClientConnectionCount | Average, Maximum | vs listener quota (see references/monitor-and-alarm.md) |
Express brokers do NOT emit: UnderReplicatedPartitions,
UnderMinIsrPartitionCount, HeapMemoryAfterGC, TrafficShaping, Volume*
metrics, KafkaDataLogsDiskUsed, ProduceMessageConversionsPerSec,
FetchMessageConversionsPerSec.
Per consumer group (identified from the customer or from the broker logs):
| Metric | Stat | Notes |
|---|---|---|
SumOffsetLag per (Consumer Group, Topic) | Maximum | DEFAULT level |
MaxOffsetLag per (Consumer Group, Topic) | Maximum | DEFAULT level |
EstimatedMaxTimeLag per (Consumer Group, Topic) | Maximum | DEFAULT level |
OffsetLag per (Consumer Group, Topic, Partition) | Maximum | Requires PER_TOPIC_PER_PARTITION |
Collect existing alarms:
aws cloudwatch describe-alarms --namespace AWS/KafkaCompare against the 13-alarm recommended set (details in
references/monitor-and-alarm.md) and produce a coverage table (present / missing / firing).
Assign a severity to every finding: CRITICAL / HIGH / MEDIUM / LOW / INFO (see Severity Definitions later in this file).
Evaluate across seven dimensions. Some checks are skipped for Express — noted inline.
ACTIVE. MAINTENANCE / UPDATING is transient; anything
else is a finding.PER_BROKER or higher — DEFAULT → MEDIUM.TLS. TLS_PLAINTEXT → MEDIUM (prod: HIGH).
PLAINTEXT → CRITICAL.Unauthenticated.Enabled=true in production → CRITICAL.SERVICE_PROVIDED_EIPS in production → HIGH.ALARM state → HIGH (surface in report header).PartitionCount per broker ≤ recommended limit for the broker instance
type. Above recommended but below max → MEDIUM. Above max → HIGH (blocks
update operations).LeaderCount variance across brokers < 10%. 10-25% → MEDIUM. > 25% → HIGH.UnderReplicatedPartitions > 0 sustained (Standard only) → HIGH. Transient
during a SECURITY_PATCHING / BROKER_UPDATE operation from Step 3 →
INFO — do NOT flag.UnderMinIsrPartitionCount > 0 (Standard only) → CRITICAL.min.insync.replicas >= replication.factor
are a configuration error the CloudWatch metric will not surface. If the
user has provided topic configs, flag any such topic as HIGH.CpuUser + CpuSystem avg < 60%. 60-70% → MEDIUM. > 70% → HIGH.HeapMemoryAfterGC < 60% (Standard only). 60-80% → MEDIUM. > 80% → HIGH.RequestHandlerAvgIdlePercent > 30% (PER_BROKER). 10-30% → MEDIUM. < 10% → HIGH.NetworkProcessorAvgIdlePercent > 30% (PER_BROKER). 10-30% → MEDIUM. < 10% → HIGH.TrafficShaping = 0 (Standard only). Any non-zero → HIGH.BwInAllowanceExceeded / BwOutAllowanceExceeded = 0 (PER_BROKER). Non-zero → HIGH.references/troubleshoot-performance.md for baseline table). 60-70% →
MEDIUM. > 70% → HIGH.ProduceThrottleTime / FetchThrottleTime > 0 → HIGH.Standard only (Express storage is managed — Express clusters only get the
StorageUsed check):
KafkaDataLogsDiskUsed < 70%. 70-85% → HIGH. > 85% → CRITICAL.VolumeQueueLength avg < 1 (PER_BROKER). 1-5 → MEDIUM. > 5 → HIGH.VolumeReadBytes + VolumeWriteBytes) < 60% of instance
baseline (PER_BROKER). 60-70% → MEDIUM. > 70% → HIGH.Express:
StorageUsed per broker vs Express per-broker storage quota (see MSK
Express quotas). Approaching quota → MEDIUM.Generate a separate report artifact per cluster reviewed.
Artifact naming: msk-review-<cluster-name>-<YYYY-MM-DD>.md
Example: msk-review-prod-orders-2026-04-29.md
Report structure:
# MSK Operational Review — <cluster-name>
Account: <account-id> | Region: <region> | Date: <YYYY-MM-DD>
Broker Type: Standard/Express | Instance Type: <type> | Broker Count: <n> | AZs: <n>
Kafka Version: <version> | Monitoring Level: <level>| Item | Value | | Cluster state / version | … | | Broker type / instance / count / AZs | … | | Storage | mode, size (Standard), provisioned throughput (if any) | | Encryption | in-transit, in-cluster, at-rest (KMS) | | Authentication | IAM / SCRAM / mTLS / Unauthenticated flags | | Public access | DISABLED / SERVICE_PROVIDED_EIPS | | Monitoring level | DEFAULT / PER_BROKER / PER_TOPIC_PER_BROKER / PER_TOPIC_PER_PARTITION | | Logging | destinations enabled |
For each of the 7 dimensions (7.1-7.7):
| # | Finding | Severity | Current State | Recommendation |
If a dimension was skipped or partial due to monitoring level or broker type, say so explicitly in a note above the table for that dimension.
| Metric | Stat | 7-Day Avg | 7-Day Max | Status | Finding |
| # | Recommended Alarm | Metric | Threshold | Priority | Status |
Also list alarms currently in ALARM state with timestamps.
From list-cluster-operations-v2 in Step 3:
| Operation Type | Start Time | End Time | State |
Call out any operations that would explain transient metric anomalies.
| # | Finding | Severity | Dimension | Effort | Impact |
Sorted by severity.
Describe cluster:
aws kafka describe-cluster-v2 --cluster-arn <cluster-arn>List brokers:
aws kafka list-nodes --cluster-arn <cluster-arn>Get bootstrap brokers:
aws kafka get-bootstrap-brokers --cluster-arn <cluster-arn>List recent cluster operations (patching, config updates, storage changes):
aws kafka list-cluster-operations-v2 --cluster-arn <cluster-arn>Expand Standard broker storage — Operator-run (recommend, do not execute):
aws kafka update-broker-storage \
--cluster-arn <cluster-arn> \
--current-version <cluster-version> \
--target-broker-ebs-volume-info '[{"KafkaBrokerNodeId": "All", "VolumeSizeGB": <target-size>}]'Get a CloudWatch metric (example: CpuUser per broker):
aws cloudwatch get-metric-statistics \
--namespace AWS/Kafka \
--metric-name CpuUser \
--dimensions Name="Cluster Name",Value="<cluster-name>" Name="Broker ID",Value="<broker-id>" \
--start-time <start> --end-time <end> --period 300 --statistics AverageCreate a cluster configuration (server.properties) — Operator-run (recommend, do not execute):
The --server-properties argument MUST be a real Kafka properties file with
one key=value per line, separated by actual newline characters — NOT the
literal two-character escape sequence \n. The MSK API accepts the bytes as-is;
if you pass "k1=v1\nk2=v2" as a single string with escaped newlines, MSK
stores ONE invalid property line and the cluster will fail to apply it.
Recommended pattern: write the properties to a local file with real newlines,
then pass it via fileb:// so the CLI uploads the raw bytes verbatim. Verify by
reading the revision back with describe-configuration-revision and
base64-decoding ServerProperties — you should see one property per line.
cat > server.properties <<'EOF'
auto.create.topics.enable=false
default.replication.factor=3
min.insync.replicas=2
unclean.leader.election.enable=false
num.io.threads=32
num.network.threads=16
log.retention.hours=168
EOF
aws kafka create-configuration \
--name <config-name> \
--kafka-versions "3.6.0" \
--server-properties fileb://server.properties| Error | Cause | Fix |
|---|---|---|
aws kafka update-broker-storage returns "storage is optimizing" | Previous storage expansion still in cool-down (minimum 6 hours) | Wait for optimization to complete. Check cluster state with describe-cluster-v2. |
ClusterState is MAINTENANCE | Standard brokers undergoing patching. Express brokers stay ACTIVE during maintenance. | Wait for cluster to return to ACTIVE. Do not perform update operations during MAINTENANCE. |
Consumer receives GROUP_COORDINATOR_NOT_AVAILABLE | Coordinator broker is temporarily unavailable during rolling restart or overloaded | Retry with backoff. Check if maintenance is in progress via list-cluster-operations-v2. |
NotEnoughReplicasException on produce | Fewer brokers in ISR than min.insync.replicas (default: 2) | Check UnderReplicatedPartitions (Standard only). For Express, check ProduceThrottleTime and broker health instead — URP is not available. If a broker is down for maintenance, this is transient. Do NOT lower min.insync.replicas to work around this. |
| Severity | Definition | SLA |
|---|---|---|
| CRITICAL | Immediate risk to availability, security, or data integrity | Fix within 24–48 hours |
| HIGH | Significant gap that could lead to incidents | Fix within 1 week |
| MEDIUM | Notable improvement opportunity | Plan within 30 days |
| LOW | Minor optimization or hardening | Address when convenient |
| INFO | Observation, no action required | N/A |
© aws, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 14 other files (references) in skills/msk-operations of aws/tools-for-devops-agent.
Open the folder on GitHubat commit ddda70b
Msk Operations next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Msk Operations this skillaws/tools-for-devops-agent | 100 | — | ~6.6k | Automated safety check: Pass | Apache-2.0 | |
| AWS Advisordiegosouzapw/awesome-omni-skills | 159 | — | ~4.3k | Automated safety check: Pass | MIT | |
| Amazon Elasticacheaws/agent-toolkit-for-aws | 2.8k | — | ~4.5k | Automated safety check: Pass | Apache-2.0 | |
| AWS Solution Architectborghei/Claude-Skills | 881 | — | ~1.8k | Automated safety check: Pass | MIT | |
| AWS Solution Architectalirezarezvani/claude-skills | 28k | 1 repos | ~2.5k | Automated safety check: Pass | MIT | |
| New Event Sourceaws/aws-lambda-dotnet | 1.7k | — | ~3k | Automated safety check: Pass | Apache-2.0 |
diegosouzapw/awesome-omni-skills
AWS Advisor workflow skill. An agent skill from diegosouzapw/awesome-omni-skills.
aws/agent-toolkit-for-aws
Activate when developers have latent caching needs: slow API responses, database read bottlenecks, DynamoDB throttling or cost, RDS/Aurora scaling pressure, Bedrock latency or cost, or adding a…
borghei/Claude-Skills
Design AWS serverless architectures for startups with IaC. An agent skill from borghei/Claude-Skills.
alirezarezvani/claude-skills
Design AWS architectures for startups using serverless patterns and IaC templates.
aws/aws-lambda-dotnet
Add a new AWS event source attribute (e.g., Kinesis, Kafka, MQ) to the Lambda .NET Annotations framework, including the attribute class, source generator integration, CloudFormation writer, unit…
rohitg00/awesome-claude-code-toolkit
AWS cloud patterns for Lambda, ECS, S3, DynamoDB, and Infrastructure as Code with CDK/Terraform
aws/tools-for-devops-agent
Amazon SageMaker AI Operational Review. An agent skill from aws/tools-for-devops-agent.
aws/tools-for-devops-agent
A skill your agent uses for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances.
aws/tools-for-devops-agent
ALWAYS use this skill in the beginning of any incident investigation, root cause analysis, or operational troubleshooting.
aws/tools-for-devops-agent
AWS Database Migration Service (DMS) operational review and troubleshooting skill.
aws/tools-for-devops-agent
Performs a comprehensive Amazon ECS operations review across the 6 review pillars (Resiliency & HA, Observability, Security, Operations, Performance, Additional Analysis) using read-only AWS APIs…
aws/tools-for-devops-agent
Comprehensive Amazon RDS and Aurora operational review aligned with the AWS Well-Architected Framework and RDS/Aurora best practices.
Categories
Amazon MSK Provisioned operations, troubleshooting, and health assessment for Standard and Express brokers. Msk Operations is an agent skill from aws/tools-for-devops-agent, published by the product's own GitHub organization. Amazon MSK Provisioned operations, troubleshooting, and health assessment for Standard and Express brokers.
Msk Operations fits situations like: the user mentions Amazon MSK; MSK Provisioned; MSK Standard/Express brokers; apache Kafka on AWS.
Run `npx skills add aws/tools-for-devops-agent --skill msk-operations -a claude-code`. Or copy the skill folder (skills/msk-operations in aws/tools-for-devops-agent) into .claude/skills/msk-operations in your project. Claude Code loads it when a task matches its description.
Run `npx skills add aws/tools-for-devops-agent --skill msk-operations -a codex`. Or copy the skill folder (skills/msk-operations in aws/tools-for-devops-agent) into .agents/skills/msk-operations in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aws/tools-for-devops-agent --skill msk-operations -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/msk-operations, .gemini/skills/msk-operations, .github/skills/msk-operations and .opencode/skills/msk-operations in your project.
Going by SKILL.md and its folder, Msk Operations needs the command-line tools its instructions call (aws).
SKILL.md names 1 domain. As links in the text: docs.aws.amazon.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Msk Operations is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.6k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 15k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Msk Operations: AWS Advisor (diegosouzapw/awesome-omni-skills, 159 stars), Amazon Elasticache (aws/agent-toolkit-for-aws, 2.8k stars), AWS Solution Architect (borghei/Claude-Skills, 881 stars) and AWS Solution Architect (alirezarezvani/claude-skills, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
aws (a GitHub organization, an official publisher) maintains it in aws/tools-for-devops-agent, which has 100 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 8, 2026.
Source: aws/tools-for-devops-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.