Diagnose and fix common ABCA issues: deployment failures, preflight errors, authentication problems, agent failures, and build issues.

OfficialMIT-0Auto-check passedDevOps & Cloud

Install Troubleshoot

skills CLI
$ npx skills add aws-samples/sample-autonomous-cloud-coding-agents --skill troubleshoot -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aws-samples/sample-autonomous-cloud-coding-agents troubleshoot --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aws-samples/sample-autonomous-cloud-coding-agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/docs/abca-plugin/skills/troubleshoot .claude/skills/troubleshoot && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
troubleshoot
GitHub stars
158
Token cost
~2.7k tokens
SKILL.md length
1,107 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
MIT-0

At a glance

Diagnose and fix common ABCA issues: deployment failures, preflight errors, authentication problems, agent failures, and build issues.

  • Works in 6 steps: Build/Compilation — TypeScript errors,… → Deployment — CDK deploy/synth failures,… → Authentication — Cognito errors, token… → …
  • The user says troubleshoot
  • SKILL.md covers Step 1: Identify the Problem…, Build/Compilation Issues, Deployment Issues and Authentication Issues, plus 4 more sections
  • Calls aws, node and mise; needs GITHUB_TOKEN

What it does

Troubleshoot is an agent skill from aws-samples/sample-autonomous-cloud-coding-agents, published by the product's own GitHub organization. Diagnose and fix common ABCA issues: deployment failures, preflight errors, authentication problems, agent failures, and build issues. Use when the user says "troubleshoot", "debug", "not working", "error", "failed", "help me fix", "preflightfailed", "task failed", "deploy failed", "auth error", "401", "422", "503", or describes something not working as expected.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Deployment and Authentication. It works with Amazon Web Services. The repository describes itself as: Autonomous background coding agents on AWS. Turn tasks into pull requests via isolated runtimes, with built-in orchestration, observability, and governance. The licence is MIT-0.

When your agent uses it

  • The user says troubleshoot
  • Preflightfailed
  • Describes something not working as expected

Example prompts

  • “troubleshoot”
  • “not working”
  • “failed”
  • “/troubleshoot”

Requirements

  • Docker
  • A credential in GITHUB_TOKEN

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Build/Compilation — TypeScript errors, test failures, lint issues
  2. Deployment — CDK deploy/synth failures, CloudFormation errors
  3. Authentication — Cognito errors, token issues, 401 responses
  4. Task Submission — 422 errors, validation failures, guardrail blocks
  5. Task Execution — Preflight failures, agent failures, timeouts
  6. Local Agent Testing — Docker issues, run.sh problems

What it can do on your machine

Read from SKILL.md and the folder at commit dfcde8d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • aws
    • node
    • mise
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.aws.amazon.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GITHUB_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Troubleshoot loads about 2.7k tokens when it runs. Until then it costs about 95 tokens; SKILL.md has 1,107 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~95
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aws-samples/sample-autonomous-cloud-coding-agents at commit dfcde8d, republished under its MIT-0 licence (© aws-samples). 1,107 words, ~2,734 tokens.

Download SKILL.mdSave it as .claude/skills/troubleshoot/SKILL.md (or your agent's skills folder).
name
troubleshoot
description
Diagnose and fix common ABCA issues: deployment failures, preflight errors, authentication problems, agent failures, and build issues. Use when the user says "troubleshoot", "debug", "not working", "error", "failed", "help me fix", "preflight_failed", "task failed", "deploy failed", "auth error", "401", "422", "503", or describes something not working as expected.

ABCA Troubleshooting

You are diagnosing an issue with the ABCA platform. Follow a systematic approach: gather symptoms, check the most common causes, and apply targeted fixes.

Running the CLI: commands below call node cli/lib/bin/bgagent.js …. In a non-interactive or mise-managed shell node may not be on PATH — prefix with mise exec --. Ironically, node: command not found is itself a common symptom (the shell hasn't activated mise); that's a missing prefix, not a broken install.

Step 1: Identify the Problem Category

Determine which area the issue falls into:

  1. Build/Compilation — TypeScript errors, test failures, lint issues
  2. Deployment — CDK deploy/synth failures, CloudFormation errors
  3. Authentication — Cognito errors, token issues, 401 responses
  4. Task Submission — 422 errors, validation failures, guardrail blocks
  5. Task Execution — Preflight failures, agent failures, timeouts
  6. Local Agent Testing — Docker issues, run.sh problems

Build/Compilation Issues

bash
export MISE_EXPERIMENTAL=1
mise //cdk:compile 2>&1 | tail -50  # TypeScript errors
mise //cdk:test 2>&1 | tail -50     # Test failures

Common causes:

  • Missing mise run install after pulling changes
  • yarn: command not found — Run corepack enable && corepack prepare yarn@1.22.22 --activate
  • Type mismatches after editing cdk/src/handlers/shared/types.ts without updating cli/src/types.ts

Deployment Issues

bash
# Check CloudFormation events for the failed stack
aws cloudformation describe-stack-events --stack-name backgroundagent-dev \
  --query 'StackEvents[?ResourceStatus==`CREATE_FAILED` || ResourceStatus==`UPDATE_FAILED`].[LogicalResourceId,ResourceStatusReason]' \
  --output table

Common causes:

  • Docker not running — Required for CDK asset bundling
  • Missing CDK bootstrap — Run mise //cdk:bootstrap
  • IAM permission issues — Check aws sts get-caller-identity
  • Region mismatch — Ensure consistent region across all commands

Authentication Issues

bash
# Verify credentials
aws sts get-caller-identity

# Check Cognito user exists
aws cognito-idp admin-get-user \
  --user-pool-id $USER_POOL_ID \
  --username user@example.com

Common causes:

  • "App client does not exist" — Region mismatch between CLI config and stack deployment
  • Token expired — Re-authenticate with bgagent login
  • 401 on API calls — Token not included or malformed in Authorization header
  • User not created — Self-signup is disabled; admin must create users

Task Submission Issues (422 / 400)

"Repository not onboarded" / REPO_NOT_ONBOARDED (422):

  • The repo isn't registered. Fastest fix: bgagent repo onboard <owner/repo> (operator path — writes the RepoTable record at runtime, no redeploy). A CDK Blueprint is only needed for declarative config. Use the onboard-repo skill for details.
  • Also confirm the owner/repo matches exactly what you pass to bgagent submit --repo.

"GUARDRAIL_BLOCKED" (400):

  • Task description triggered Bedrock Guardrails content screening
  • Review and rephrase the task description to remove potentially flagged content

Validation errors:

  • Check required fields: repo is required, plus at least one of issue_number, task_description, pr_number
  • max_turns range: 1-500
  • max_budget_usd range: $0.01-$100

Task Execution Issues

bash
# Check task events for details
node cli/lib/bin/bgagent.js events <TASK_ID> --output json

preflight_failed:

  • GitHub PAT lacks permissions for the repo
  • Repository doesn't exist or is private without proper token scope
  • Check event reason and detail fields for specifics
  • Verify PAT: fine-grained token must include the target repository with Contents (read/write), Pull Requests (read/write), Issues (read)

task_failed / task completes with 0 tokens and no PR:

  • Agent encountered an error during execution
  • Check CloudWatch logs for the session:
    bash
    aws logs filter-log-events \
      --log-group-name "/aws/vendedlogs/bedrock-agentcore/runtime/APPLICATION_LOGS/jean_cloude" \
      --filter-pattern "<TASK_ID>" \
      --region us-west-2 --query 'events[*].message' --output text
  • Common: repo build/test commands not documented in CLAUDE.md

403 "not authorized to perform bedrock:InvokeModelWithResponseStream":

  • The repo's model_id is a model the runtime IAM role wasn't granted. The runtime only has grantInvoke for the models in the stack's configured set — read it from the BedrockModelIds stack output rather than a list here (Sonnet 4.6, Opus 4.8, Opus 5, Haiku 4.5 by default).
  • Quick fix: point the repo at an already-granted model — bgagent repo onboard <owner/repo> --model global.anthropic.claude-opus-5 (no redeploy).
  • To add a new model to the runtime: grant it in the stack and redeploy. The model set is the shared list in cdk/src/constructs/bedrock-models.ts — add the model via the bedrockModels CDK context (cdk.json) so both the AgentCore and ECS backends grant it (#433). Adding a model also requires account-level Bedrock access for it (separate from IAM — see the next row).

Model not enabled / "not available on your Bedrock deployment" (often immediate failure, few turns, zero or near-zero tokens):

  • IAM is necessary but not sufficient. The AgentCore role may already have bedrock:InvokeModel*, but the account must also satisfy Amazon Bedrock model access: Marketplace subscription flow on first serverless use (with aws-marketplace:Subscribe / ViewSubscriptions where needed), Anthropic first-time use details (PutUseCaseForModelAccess or the console model catalog), and a valid payment method for Marketplace-backed models.
  • Use an inference profile ID in the Blueprint / DynamoDB model_id when Bedrock requires it for on-demand invocation (for example global.anthropic.claude-opus-5 for global Opus 5). See Use an inference profile in model invocation. Raw anthropic.* IDs often hit "on-demand not supported" or wrong routing — see the 400 row below.
  • Cross-Region profiles route across Regions in a geography; ensure IAM and any SCPs allow Bedrock in all destination Regions for that profile. See Supported Regions and models for inference profiles.
  • Task status: When the Claude CLI reports a terminal error via ResultMessage.is_error, the agent marks the task FAILED (not COMPLETED) and persists error_message in DynamoDB.
Show full SKILL.md (386 more words)Show less

400 "Invocation with on-demand throughput isn't supported":

  • The Blueprint modelId uses a raw foundation model ID (e.g. anthropic.claude-opus-4-8)
  • Fix: change to the inference profile ID, prefixed with the geography the stack grants — its BedrockGeoRegion output (e.g. global.anthropic.claude-opus-4-8) — then update DynamoDB via redeploy. A prefix from a different geography raises AccessDenied rather than this 400, since the IAM grant is scoped per geography.

503 "Too many connections" / task completes with 0 tokens after long duration:

  • Bedrock is throttling model invocations. The agent retries for minutes then gives up.
  • Symptoms: task runs for 10-15 minutes, may end with COMPLETED if the SDK does not flag ResultMessage.is_error (unlike hard Bedrock entitlement errors, which surface as FAILED once the CLI sets is_error on the result)
  • Diagnosis:
    1. Check application logs for "text": "API Error: 503 Too many connections"
    2. Check what model_id is actually being passed — the DynamoDB record may have a stale model override:
      bash
      aws dynamodb get-item \
        --table-name <RepoTableName> \
        --key '{"repo": {"S": "owner/repo"}}' \
        --query 'Item.model_id' --output text
  • Causes:
    • Stale model_id in DynamoDB (most common) — the Blueprint onUpdate only sets fields present in props; removing a modelId prop does NOT remove the field from DynamoDB. The task keeps using the old model.
    • Bedrock service-level throttling for the specific model (Opus-class models have tighter limits than Sonnet or Haiku)
    • Account quota limits reached
  • Fix:
    1. Check and fix the DynamoDB record first — remove stale model_id if present:
      bash
      aws dynamodb update-item \
        --table-name <RepoTableName> \
        --key '{"repo": {"S": "owner/repo"}}' \
        --update-expression "REMOVE model_id"
    2. If model_id is correct, wait and retry — throttling is often transient
    3. Switch to a model with higher availability (Haiku 4.5 > Sonnet 4.6 > Opus)
    4. Request a Bedrock quota increase for InvokeModel RPM on your model

task_timed_out:

  • 9-hour maximum exceeded
  • Consider reducing scope or increasing max_turns for complex tasks
  • Check if the agent is stuck in a loop (review logs)

Concurrency limit:

  • Default: 3 concurrent tasks per user
  • Wait for running tasks to complete or cancel them

Local Agent Testing Issues

bash
# Verify Docker is running
docker info

# Test locally with dry run
DRY_RUN=1 ./agent/run.sh "owner/repo" "Test task"

Common causes:

  • Missing environment variables: GITHUB_TOKEN, AWS_REGION
  • Docker not running or insufficient resources (needs 2 vCPU, 8 GB RAM)
  • Missing AWS credentials for Bedrock access

Diagnostic Commands Quick Reference

bash
# Stack status
aws cloudformation describe-stacks --stack-name backgroundagent-dev --query 'Stacks[0].StackStatus'

# Stack outputs
aws cloudformation describe-stacks --stack-name backgroundagent-dev --query 'Stacks[0].Outputs' --output table

# Task status (use --verbose for HTTP-level debug output)
node cli/lib/bin/bgagent.js --verbose status <TASK_ID>
node cli/lib/bin/bgagent.js events <TASK_ID> --output json

# Watch task progress in real time
node cli/lib/bin/bgagent.js watch <TASK_ID>

# Download full execution trace (task must have been submitted with --trace)
node cli/lib/bin/bgagent.js trace download <TASK_ID>

# List running tasks
node cli/lib/bin/bgagent.js list --status RUNNING

# Build health
mise run build

Tip: Add --verbose to any bgagent command to see the full HTTP request/response cycle on stderr. This is the fastest way to diagnose auth, network, or API contract issues.

© aws-samples, MIT-0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in docs/abca-plugin/skills/troubleshoot of aws-samples/sample-autonomous-cloud-coding-agents.

Open the folder on GitHubat commit dfcde8d

Compare with similar skills

Troubleshoot next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Troubleshoot compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Troubleshoot this skillaws-samples/sample-autonomous-cloud-coding-agents158—~2.7kAutomated safety check: PassMIT-0
AWS Cdk Developmentzxkane/aws-skills3672 repos~2.5kAutomated safety check: PassMIT
Senior DevOps Toolkitmaslennikov-ig/claude-code-orchestrator-kit2606 repos~1.1kAutomated safety check: NotesCustom licence
Ecspressokayac/ecspresso1.1k—~1.4kAutomated safety check: PassMIT
Spa Create Configsplunk/splunk-platform-automator138—~3.5kAutomated safety check: PassProprietary
Kcli Cluster Deploymentkarmab/kcli653—~1.5kAutomated safety check: PassApache-2.0

Similar skills

  • AWS Cdk Development

    zxkane/aws-skills

    AWS Cloud Development Kit (CDK) expert for building cloud infrastructure with TypeScript/Python.

    367 GitHub starsUsed in 2 repos~2.5k tokens
    DevOps & CloudAuto-check passed
  • Senior DevOps Toolkit

    maslennikov-ig/claude-code-orchestrator-kit

    Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup…

    260 GitHub starsUsed in 6 repos~1.1k tokens
    DevOps & CloudAuto-check: notes
  • Ecspresso

    kayac/ecspresso

    ECS deployment tool - deploy, manage, and troubleshoot ECS services

    1.1k GitHub stars~1.4k tokensUpdated 4 days ago
    DevOps & CloudAuto-check passed
  • Spa Create Config

    splunk/splunk-platform-automator

    A skill your agent uses when creating or updating splunkconfig.yml, designing Splunk Enterprise lab topology, multisite IDXC, SHC layout, architecture plan before config, or AWS Terraform block for…

    138 GitHub stars~3.5k tokensUpdated 3 days ago
    DevOps & CloudAuto-check passed
  • Guides deployment and management of Kubernetes clusters with kcli.

    653 GitHub stars~1.5k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Spa Add Test Scenario

    splunk/splunk-platform-automator

    A skill your agent uses when adding app scope/routing test coverage (deployer, CM, DS, direct).

    138 GitHub stars~2.2k tokensUpdated 3 days ago
    DevOps & CloudAuto-check passed

More from aws-samples/sample-autonomous-cloud-coding-agents

  • Deploy

    aws-samples/sample-autonomous-cloud-coding-agents

    Official

    Deploy, diff, or destroy the ABCA CDK stack. An agent skill from aws-samples/sample-autonomous-cloud-coding-agents.

    158 GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Onboard Repo

    aws-samples/sample-autonomous-cloud-coding-agents

    Official

    Onboard a new GitHub repository to the ABCA platform so the agent can target it.

    158 GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Setup

    aws-samples/sample-autonomous-cloud-coding-agents

    Official

    Guided installation and first-time setup for ABCA. An agent skill from aws-samples/sample-autonomous-cloud-coding-agents.

    158 GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Submit Task

    aws-samples/sample-autonomous-cloud-coding-agents

    Official

    Submit a coding task to the ABCA platform via CLI or REST API.

    158 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Status

    aws-samples/sample-autonomous-cloud-coding-agents

    Official

    Check ABCA platform status — stack health, running tasks, and recent task history.

    158 GitHub stars~438 tokensUpdated today
    Auto-check: notes

Categories

Questions about Troubleshoot

What does Troubleshoot do?

Diagnose and fix common ABCA issues: deployment failures, preflight errors, authentication problems, agent failures, and build issues. Troubleshoot is an agent skill from aws-samples/sample-autonomous-cloud-coding-agents, published by the product's own GitHub organization. Diagnose and fix common ABCA issues: deployment failures, preflight errors, authentication problems, agent failures, and build issues.

When should I use Troubleshoot?

Troubleshoot fits situations like: the user says troubleshoot; preflightfailed; describes something not working as expected.

How do I install Troubleshoot in Claude Code?

Run `npx skills add aws-samples/sample-autonomous-cloud-coding-agents --skill troubleshoot -a claude-code`. Or copy the skill folder (docs/abca-plugin/skills/troubleshoot in aws-samples/sample-autonomous-cloud-coding-agents) into .claude/skills/troubleshoot in your project. Claude Code loads it when a task matches its description.

How do I install Troubleshoot in Codex?

Run `npx skills add aws-samples/sample-autonomous-cloud-coding-agents --skill troubleshoot -a codex`. Or copy the skill folder (docs/abca-plugin/skills/troubleshoot in aws-samples/sample-autonomous-cloud-coding-agents) into .agents/skills/troubleshoot in your project. Codex loads it when a task matches its description.

Can I use Troubleshoot in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aws-samples/sample-autonomous-cloud-coding-agents --skill troubleshoot -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/troubleshoot, .gemini/skills/troubleshoot, .github/skills/troubleshoot and .opencode/skills/troubleshoot in your project.

What does Troubleshoot need to run?

Going by SKILL.md and its folder, Troubleshoot needs the command-line tools its instructions call (aws, node, mise and docker) and credentials named GITHUB_TOKEN. Our summary lists: Docker; A credential in GITHUB_TOKEN.

Does Troubleshoot access the network?

SKILL.md names 1 domain. As links in the text: docs.aws.amazon.com. This is read from the text; nothing was executed.

Is Troubleshoot safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Troubleshoot use?

Troubleshoot is published under the MIT-0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Troubleshoot use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Troubleshoot?

Skills that share tags, products or a category with Troubleshoot: AWS Cdk Development (zxkane/aws-skills, 367 stars), Senior DevOps Toolkit (maslennikov-ig/claude-code-orchestrator-kit, 260 stars), Ecspresso (kayac/ecspresso, 1.1k stars) and Spa Create Config (splunk/splunk-platform-automator, 138 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Troubleshoot?

aws-samples (a GitHub organization, an official publisher) maintains it in aws-samples/sample-autonomous-cloud-coding-agents, which has 158 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 8, 2026.

Source: aws-samples/sample-autonomous-cloud-coding-agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.