Foresight
Hmbown/Wizards-of-the-Ghosts
Foresight produces bounded forecasts with explicit uncertainty to guide decisions.
A skill your agent uses when your agent or environment is broken — wrong answers, errors, timeouts, tool failures, or CLI issues.
$ npx skills add aws/agent-toolkit-for-aws --skill agents-debug -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install aws/agent-toolkit-for-aws agents-debug --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/aws-agents/skills/agents-debug .claude/skills/agents-debug && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agents-debug" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/plugins/aws-agents/skills/agents-debug into .claude/skills/agents-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agents-debug", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/aws/agent-toolkit-for-aws/tree/main/plugins/aws-agents/skills/agents-debugType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add aws/agent-toolkit-for-aws --skill agents-debug -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install aws/agent-toolkit-for-aws agents-debug --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/aws-agents/skills/agents-debug .agents/skills/agents-debug && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agents-debug" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/plugins/aws-agents/skills/agents-debug into .agents/skills/agents-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agents-debug", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aws/agent-toolkit-for-aws --skill agents-debug -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install aws/agent-toolkit-for-aws agents-debug --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/aws-agents/skills/agents-debug .cursor/skills/agents-debug && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agents-debug" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/plugins/aws-agents/skills/agents-debug into .cursor/skills/agents-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agents-debug", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/aws/agent-toolkit-for-aws.git --path plugins/aws-agents/skills/agents-debug--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add aws/agent-toolkit-for-aws --skill agents-debug -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install aws/agent-toolkit-for-aws agents-debug --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/aws-agents/skills/agents-debug .gemini/skills/agents-debug && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agents-debug" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/plugins/aws-agents/skills/agents-debug into .gemini/skills/agents-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agents-debug", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install aws/agent-toolkit-for-aws agents-debugInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add aws/agent-toolkit-for-aws --skill agents-debug -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/aws-agents/skills/agents-debug .github/skills/agents-debug && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agents-debug" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/plugins/aws-agents/skills/agents-debug into .github/skills/agents-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agents-debug", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aws/agent-toolkit-for-aws --skill agents-debug -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install aws/agent-toolkit-for-aws agents-debug --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/aws-agents/skills/agents-debug .opencode/skills/agents-debug && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agents-debug" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/plugins/aws-agents/skills/agents-debug into .opencode/skills/agents-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agents-debug", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agents-debugA skill your agent uses when your agent or environment is broken — wrong answers, errors, timeouts, tool failures, or CLI issues.
Agents Debug is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Use when your agent or environment is broken — wrong answers, errors, timeouts, tool failures, or CLI issues. Reads traces and logs to diagnose root causes. Also checks prerequisites when the CLI itself isn't working. Triggers on: "agent not working", "wrong answer", "agent error", "tool call failing", "debug agent", "check logs", "read traces", "broken", "500 error", "424 error", "model access denied", "command not found", "stuck in DELETING", "maxVms exceeded", "cold start diagnosis", "cold start slow"…
Its SKILL.md is about 7.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/doctor.md`).
It sits in DevOps & Cloud, covering Root cause analysis and Observability. The repository describes itself as: Official, AWS-supported MCP servers, skills, and plugins to help AI agents build on AWS. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 188af2f. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadGrepGlobBashFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
jqawspipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
console.aws.amazon.comgithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
GEMINI_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agents Debug loads about 7.7k tokens when it runs, and up to ~9.2k if it reads all its reference files. Until then it costs about 209 tokens; SKILL.md has 3,279 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Read, Grep, Glob, BashAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from aws/agent-toolkit-for-aws at commit 188af2f, republished under its Apache-2.0 licence (© aws). 3,279 words, ~7,732 tokens.
.claude/skills/agents-debug/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Diagnose why your AgentCore agent or environment isn't working correctly.
agentcore command not found or prerequisites are missingDo NOT use for:
agents-deployagents-get-startedagents-optimize$ARGUMENTS is optional:
/agents-debug # interactive — describe what's wrong
/agents-debug traces # read and explain recent traces
/agents-debug logs # search recent logs for errors
/agents-debug memory # diagnose memory recall issues specifically
/agents-debug doctor # check environment prerequisitesIf the developer's issue is about the CLI itself (command not found, prerequisites, environment setup), load references/doctor.md and follow its diagnostic checklist.
If the issue is about agent behavior (wrong answers, errors, timeouts, tool failures), continue with Step 1 below.
Run agentcore --version. This skill requires v0.9.0 or later. If the version is older, tell the developer to run agentcore update before proceeding.
Ask (or infer from context):
"What's happening?
- The agent returns an error message
- The agent returns a wrong or unhelpful answer
- A specific tool call is failing
- Memory isn't working (agent doesn't remember things)
- The agent is slow or timing out
- I want to understand what the agent did in a specific session"
Don't ask the developer to paste logs — read them directly.
# List recent traces
agentcore traces list --runtime <AgentName> --since 1h
# Get the most recent trace ID
agentcore traces list --runtime <AgentName> --since 1h --limit 1
# Download and read the trace
agentcore traces get <traceId> --runtime <AgentName>
# Search logs for errors
agentcore logs --runtime <AgentName> --since 1h --level error
# Search logs for a specific pattern
agentcore logs --runtime <AgentName> --since 2h --query "timeout"
agentcore logs --runtime <AgentName> --since 2h --query "model access"Important: CloudWatch put-to-get latency is ~10 seconds end-to-end — that's the delay from when a span is emitted to when it's readable by agentcore traces get or agentcore run eval. There is no separate "trace ingested but eval not ready yet" window; the same ingestion step unlocks both paths. Older skills and docs said 30–60s for traces and 2–5 minutes for evals — both are stale. If you just invoked the agent, wait ~15 seconds and both trace reads and evals will work.
Read agentcore/agentcore.json to get the agent name if not provided.
Most common cause: The model isn't enabled in the Bedrock console for your region.
Fix:
Second cause: The execution role is missing bedrock:InvokeModel.
Check:
aws iam simulate-principal-policy \
--policy-source-arn $(agentcore status --json | jq -r '.runtimes[0].executionRoleArn') \
--action-names bedrock:InvokeModel \
--resource-arns "arn:aws:bedrock:*::foundation-model/*"Third cause: Cross-region inference profile requires model access in all regions.
Model IDs starting with a geographic prefix are cross-region inference profiles that route requests within that geography:
| Prefix | Geography | Example destination regions |
|---|---|---|
us. | United States | us-east-1, us-east-2, us-west-2 |
eu. | Europe | eu-central-1, eu-west-1, eu-west-2, eu-west-3 |
apac. | Asia Pacific | ap-northeast-1, ap-southeast-1, ap-southeast-2, ap-south-1 |
global. | All commercial regions worldwide | All supported regions |
The AgentCore CLI scaffolds global. by default (e.g., global.anthropic.claude-sonnet-4-5-20250929-v1:0). All prefixes require model access enabled in every destination region the profile covers. For us. profiles, enable in all US regions; for eu., all EU regions; for global., all supported regions. Not all models support all prefixes — global. is currently available for select models only. Use global. for maximum throughput when available, or a geographic prefix when data residency requirements constrain where inference can run. Check the Bedrock inference profiles docs for current model × prefix availability.
Step 1: Find the failing tool call in the trace:
agentcore traces get <traceId> --runtime <AgentName>Look for tool call entries with error status.
Step 2: Check the gateway status:
agentcore status --type gateway
agentcore fetch access --name <AgentName> --type agentStep 3: Common tool call failures:
Gateway URL not set (local dev):
The AGENTCORE_GATEWAY_*_URL env var is only set after deploy. In agentcore dev, gateway tools aren't available. This is expected — the agent should handle this gracefully.
Auth failure on tool call:
agentcore logs --runtime <AgentName> --since 1h --query "auth"Check that the credential is configured correctly: agentcore status --type credential
Lambda function error: The Lambda itself is failing. Check Lambda logs directly:
aws logs tail /aws/lambda/<function-name> --since 1hPolicy denial: If a policy engine is attached, check policy decision logs:
agentcore logs --runtime <AgentName> --since 1h --query "policy"
agentcore status --type policy-engineStep 1: Read the trace to see the agent's reasoning:
agentcore traces get <traceId> --runtime <AgentName>The trace shows the model's reasoning steps, tool calls made, and the final response. Look for:
Step 2: Check if memory is involved:
If the agent should be using memory context but isn't, see the "Symptom: Memory not persisting" section later in this skill, or load references/doctor.md if this is an environment issue.
Step 3: Common causes:
Memory not persisting across sessions (LTM):
agentcore status --type memory --json | jq '.memories[].strategies'Wait 5–30 seconds after a session ends — LTM extraction is async. The agent must finish its session before facts are extracted.
Use UUIDs (v4) for session IDs — the platform requires a minimum of 33 characters. Short IDs like "session-1" cause LTM to fail silently. agentcore invoke generates compliant IDs by default.
Verify the memory resource is ACTIVE:
agentcore status --type memoryMemory not loading at session start:
MEMORY_*_ID env var is set:agentcore status --type memory --json | jq '.memories[].id'Verify the actor_id is consistent across sessions — memory is scoped per actor.
Check the namespace paths in your retrieval config match the namespaces used when writing.
Step 1: Check the trace for where time is being spent:
agentcore traces get <traceId> --runtime <AgentName>Look for long-running steps — model calls, tool calls, memory operations.
Step 2: Common timeout causes:
Slow agent initialization: If the first invocation after an idle period is slow but subsequent requests are fast, the agent is spending too much time initializing. Check for heavy imports at module level, database connections in global scope, or MCP client initialization during startup. Move expensive setup into the request handler or use lazy initialization. See the agents-harden skill for optimization guidance.
Model call timeout: The model is taking too long. Consider using a faster model for time-sensitive operations (e.g., Haiku instead of Sonnet for simple tasks).
Tool call timeout: The Lambda or external API is slow. Check the tool's own logs.
Memory retrieval timeout: Semantic search can be slow for large memory stores. Consider reducing top_k in your retrieval config.
VPC connectivity issue: If the agent is in a VPC, check security group rules and route tables. See agents-build (loads references/vpc.md) for VPC-specific debugging.
ServiceQuotaExceededException: maxVms limit exceeded (despite low observed concurrency)Your CloudWatch "concurrent sessions" metric shows modest numbers (maybe 30–50) but InvokeAgentRuntime calls return ServiceQuotaExceededException: maxVms limit exceeded.
What's actually happening: CloudWatch's concurrent-sessions metric is not the same as live microVM count. The maxVms quota counts all environments your account has active — including ones that finished their invocation but haven't been reclaimed yet. Idle-but-not-yet-reclaimed environments count against the quota until idleRuntimeSessionTimeout expires (default 900 seconds / 15 minutes) or you explicitly stop them.
If your code uses a new session ID per request and doesn't call StopRuntimeSession, every request leaves an environment sitting idle for 15 minutes counting against the quota.
Fix order (try in this order before requesting a quota increase):
Call StopRuntimeSession after each logical request completes. If you're not going to send more requests on this session, stop it explicitly.
client.stop_runtime_session(
agentRuntimeArn=runtime_arn,
runtimeSessionId=session_id,
)Reuse session IDs across related requests. If a user interaction produces multiple backend calls, route them to the same session instead of generating a new session ID per call.
Lower idleRuntimeSessionTimeout. If your sessions are short-lived and you can't add StopRuntimeSession everywhere, lower the timeout by editing the runtime's lifecycleConfiguration in agentcore/agentcore.json and running agentcore deploy.
Only after the above, request a quota increase. See agents-harden (loads references/limits.md) — request it through the Service Quotas console (Amazon Bedrock AgentCore), not by filing a support ticket directly.
See agents-harden Session lifecycle management section for the full pattern.
This usually means the agent container failed to start or crashed during initialization.
Step 1: Check the agent logs for startup errors:
agentcore logs --runtime <AgentName> --since 30m --level errorStep 2: Common causes:
Missing Python dependency: The agent code imports a package not in pyproject.toml. The container starts but crashes on first request. Fix: add the dependency and redeploy.
Entrypoint crash: The main.py throws an exception during import or app.run(). Check logs for the traceback.
Container image pull failure: If using Container build, the ECR image may not exist or the execution role lacks ecr:BatchGetImage. Check:
agentcore status --runtime <AgentName> --jsonMemory resource not ACTIVE: If the agent code assumes memory is available but the memory resource is still in CREATING state, the entrypoint may fail. Check:
agentcore status --type memoryInitialization timeout: The agent takes too long to be ready for its first request — heavy imports at module level, synchronous database connections, or MCP client initialization during startup can exceed the service's health-check window. The symptom looks like a 424 on the first invoke but healthy on subsequent ones. Fix: move expensive setup out of module level, use lazy initialization, or warm the agent before production traffic. See agents-harden Initialization time section for patterns.
Usually not an agent bug — the dev server is on a different port than you expect.
Default ports agentcore dev binds:
| Protocol | Default |
|---|---|
| HTTP | 8080 |
| MCP | 8000 |
| A2A | 9000 |
When the default is occupied (second dev session, a lingering process from a previous run, another service on 8080), the CLI auto-increments silently: 8080 → 8081 → 8082. A test harness or curl script hardcoded to 8080 will get Connection refused (curl exit code 7) while the agent is running fine on 8082.
Diagnose in this order:
Read the CLI banner that agentcore dev prints — it shows the actual bound port and URL. This is always the source of truth.
If the banner is gone (terminal cleared, running in background), check the log file:
tail -20 agentcore/.cli/logs/dev/*.logOr find the process directly:
# macOS / Linux
ps aux | grep -E 'agentcore dev|uvicorn' | grep -v grep
lsof -iTCP -sTCP:LISTEN -n -P | grep -E '8080|8081|8082|8000|9000'Fix options:
agentcore dev --port 8080lsof -tiTCP:8080 -sTCP:LISTEN | xargs killThis is also a common source of "works locally one day, fails the next" reports — the port shifted between runs.
Step 1: Verify the auth type matches the target type. This is the most common gateway error — using the wrong outbound auth for the target:
| Target type | Valid outbound auth |
|---|---|
mcp-server | none, oauth, or IAM (SigV4 via API) |
lambda-function-arn | IAM only (automatic) |
open-api-schema | oauth or api-key (required) |
api-gateway | none, api-key, or IAM |
smithy-model | IAM or oauth |
Step 2: Check for expired OAuth tokens. If the gateway target uses OAuth, the access token may have expired. Look for auth-related errors:
agentcore logs --runtime <AgentName> --since 1h --query "auth"
agentcore logs --runtime <AgentName> --since 1h --query "401"
agentcore logs --runtime <AgentName> --since 1h --query "403"If tokens are expiring, verify the OAuth credential provider's token endpoint is reachable and the client credentials are still valid. For MCP server targets with OAuth, the gateway handles token refresh automatically — if it's failing, the credential provider config may be wrong.
Step 3: Check the credential is configured:
agentcore status --type credential
agentcore status --type gateway --jsonWait ~15 seconds — there's a short delay (typically ~10s) between invocation and trace availability.
If still no traces after ~30 seconds:
agentcore logs --runtime <AgentName> --since 1hThis is the most common observability issue, especially for Container/Docker builds.
AgentCore doesn't capture raw stdout. It uses OpenTelemetry to ship logs to CloudWatch. Three things must be true:
1. Your entrypoint must be wrapped with opentelemetry-instrument.
CodeZip builds do this automatically. Docker/Container builds need it added manually — this is the #1 thing people miss.
In your Dockerfile CMD:
# ✅ Correct — wrapped with opentelemetry-instrument
CMD ["opentelemetry-instrument", "python", "main.py"]
# ❌ Wrong — no OTEL wrapper, logs won't appear
CMD ["python", "main.py"]2. Your runtime IAM role needs CloudWatch and X-Ray permissions:
logs:CreateLogGroup
logs:CreateLogStream
logs:PutLogEvents → scoped to /aws/bedrock-agentcore/runtimes/*
xray:PutTelemetryRecords
xray:PutTraceSegments → scoped to *If using the AgentCore CLI with CodeZip, the CDK scaffold adds these automatically. If using a custom role or Container build, verify they're present.
3. Use Python's logging module, not print().
OTEL hooks into logging automatically — no custom handlers needed. print() statements won't appear in CloudWatch.
import logging
logger = logging.getLogger(__name__)
logger.setLevel(logging.INFO)
# ✅ This appears in CloudWatch
logger.info("Processing request")
# ❌ This does NOT appear in CloudWatch
print("Processing request")Also verify: CloudWatch Transaction Search is enabled in your account. Without it, traces and spans won't appear in the GenAI Observability dashboard.
A common pattern: a runtime deployed via Terraform, CDK, or a custom IAM role works correctly (returns responses) but no CloudWatch log streams appear — while the same agent code deployed via the AgentCore Console logs fine.
This is almost always an IAM scoping issue. The execution role for a runtime deployed via the Console gets broad CloudWatch permissions by default. IaC templates often scope those permissions narrowly to /aws/bedrock-agentcore/runtimes/*, which breaks log stream creation.
The fix: logs:DescribeLogGroups must have Resource: "*", not a scoped resource. The other logs actions can be scoped to the runtime's log group.
{
"Effect": "Allow",
"Action": [
"logs:DescribeLogGroups"
],
"Resource": "*"
},
{
"Effect": "Allow",
"Action": [
"logs:CreateLogGroup",
"logs:CreateLogStream",
"logs:PutLogEvents"
],
"Resource": "arn:aws:logs:<REGION>:<ACCOUNT_ID>:log-group:/aws/bedrock-agentcore/runtimes/*:*"
}After updating the execution role's IAM policy, redeploy the runtime with agentcore deploy to pick up the new permissions.
Your agent uses SSE or long-polling responses and the connection drops mid-stream. Symptoms in client code:
RemoteProtocolError: peer closed connection without sending complete message bodyIncompleteRead exception while iterating the stream[DONE] event, response just stopsRoot cause: Infrastructure-layer idle timeout on streaming connections. If no data flows on the response stream for several minutes (a silent period while a tool executes, for example), a load balancer in front of the runtime terminates the TCP connection.
The timeout is on data flowing through the stream, not on the request total duration. As long as you emit bytes periodically, the connection stays open.
Fix: emit keepalive events during long-running tool executions.
Python pattern for a streaming entrypoint:
import asyncio
import json
from bedrock_agentcore.runtime import BedrockAgentCoreApp
app = BedrockAgentCoreApp()
async def emit_keepalive(tool_task):
"""Yield heartbeat events every 30s while tool_task is running."""
while not tool_task.done():
yield f"data: {json.dumps({'type': 'heartbeat'})}\n\n"
try:
await asyncio.wait_for(asyncio.shield(tool_task), timeout=30)
except asyncio.TimeoutError:
continue # tool still running, emit another heartbeat
@app.entrypoint
async def invoke(payload, context):
async def stream():
tool_task = asyncio.create_task(run_long_tool(payload))
# Emit heartbeats while the tool runs
async for event in emit_keepalive(tool_task):
yield event
# Tool completed — emit the real result
result = await tool_task
yield f"data: {json.dumps({'type': 'result', 'content': result})}\n\n"
yield "data: [DONE]\n\n"
return stream()Pick a heartbeat interval of ~30 seconds. Too long risks hitting the idle timeout; too short wastes bandwidth.
On the client side, filter heartbeat events before surfacing bytes to the user:
for chunk in response.iter_lines():
if not chunk:
continue
data = json.loads(chunk.removeprefix(b"data: "))
if data.get("type") == "heartbeat":
continue # ignore keepalives
# process real eventsAlternative: use the SDK's async task API for fire-and-forget patterns. If the client doesn't need to wait for the result, register the work via add_async_task / complete_async_task and return the invocation immediately. See agents-harden Long-running background tasks section.
You run multiple agent invocations in parallel with unique runtimeSessionId values, but the AI Observability dashboard groups them as one session — making it impossible to isolate a single run. Data plane logs show the session IDs are correctly unique 1:1 with request IDs, but the trace view still merges them.
Most common cause: the caller isn't enabling Active Tracing, so upstream spans arrive with Sampled=0. AgentCore respects upstream trace-sampling decisions by default. If the parent context says "don't sample," spans drop and concurrent invocations can appear merged in the dashboard.
Fix by caller type:
Lambda caller: Enable Active Tracing on the Lambda function.
aws lambda update-function-configuration \
--function-name my-caller-function \
--tracing-config Mode=ActiveOr in the Lambda console: Configuration → Monitoring and operations tools → AWS X-Ray → Active tracing.
ECS / EC2 / container caller: Initialize the AWS X-Ray SDK and ensure outbound calls to AgentCore are instrumented. For Python, use aws-xray-sdk and patch the SDK:
from aws_xray_sdk.core import xray_recorder, patch_all
patch_all() # patches boto3, requests, etc.Direct SDK caller without X-Ray: If you can't enable upstream tracing, force the runtime to sample by setting an environment variable on the agent:
OTEL_TRACES_SAMPLER=always_onThis makes the runtime sample every trace regardless of the parent context's sampling decision. Trade-off: higher tracing costs, but the traces are correct.
If traces show only a single top-level AgentCore.Runtime.Invoke span with no child spans, check the ARN your caller is using. The invoke target should be the agent runtime ARN:
arn:aws:bedrock-agentcore:<region>:<account>:runtime/<runtime-name>Not the endpoint ARN:
arn:aws:bedrock-agentcore:<region>:<account>:runtime/<runtime-name>/runtime-endpoint/DEFAULTInvoking with the endpoint ARN can bypass the full trace instrumentation path. This is a subtle trap — both ARNs produce successful responses, but only the agent ARN produces complete traces.
You called DeleteAgentRuntime, got a successful response with status: DELETING, and the runtime has been stuck in that state for more than 30 minutes. Attempting to delete the default endpoint separately returns ConflictException: Default endpoints are removed when you delete the agent.
What's happening: The deletion workflow is stuck on the service side. Retrying DeleteAgentRuntime won't help — the call succeeds immediately (returning DELETING) but the back-end workflow is the thing that's stuck. Customer-side tooling can't force-complete it.
What to do:
agentRuntimeId)requestId and timestamp of the original DeleteAgentRuntime call (from CloudTrail)Orphaned resources from a stuck deletion (ENIs, workload identities) may need manual cleanup from the service team as part of the same case.
LangGraph — model format:
Older versions of langchain-aws required the model ID without the cross-region prefix. Recent versions may support cross-region inference profiles — check your installed version:
pip show langchain-aws | grep VersionIf you hit model errors with LangGraph, try the non-prefixed ID:
# If cross-region prefix errors in your langchain-aws version:
llm = init_chat_model("anthropic.claude-sonnet-4-5-20250929-v1:0", model_provider="bedrock_converse")
# If your version supports cross-region profiles (us. = US, eu. = Europe, apac. = Asia Pacific, global. = worldwide):
llm = init_chat_model("global.anthropic.claude-sonnet-4-5-20250929-v1:0", ...)Verify against the current langchain-aws release notes: https://github.com/langchain-ai/langchain-aws/releases — cross-region inference profile support has been evolving.
Google ADK — Gemini only:
ADK only works with Gemini models. If you're seeing model errors with ADK, check that GEMINI_API_KEY is set and you're using a gemini-* model ID.
A2A agents — wrong port: A2A servers must run on port 9000. If your A2A agent isn't responding, check it's not accidentally running on 8080.
A trace shows the full execution path of one agent invocation. Key sections:
# Download trace to a file for detailed inspection
agentcore traces get <traceId> --runtime <AgentName> --output trace.json
cat trace.json | jq '.trace.orchestrationTrace.modelInvocationOutput'Once you've identified the root cause, hand off to the skill that owns the fix:
| Root cause | Hand off to | Detail |
|---|---|---|
| Memory misconfigured (wrong strategy, namespace, wiring) | agents-build | Load references/memory.md |
| Agent invocation from app not working (auth, URL, streaming) | agents-build | Load references/integrate.md |
| VPC connectivity (can't reach RDS, no internet, AZ error) | agents-build | Load references/vpc.md |
| Multi-agent delegation not working | agents-build | Load references/multi-agent.md |
| Custom request headers not reaching agent code | agents-build | Load references/request-headers.md |
| Cross-account invocation from an app in another account | agents-build | Load references/integrate.md (cross-account section) |
| Gateway auth misconfigured (401, wrong auth type) | agents-connect | Gateway auth matrix |
| Gateway target type question (Lambda vs OpenAPI vs MCP vs API Gateway) | agents-connect | "What Gateway is and isn't" section |
| Policy denying unexpectedly (Cedar, access denied on tool) | agents-connect | Load references/policy.md |
| Observability not set up (no logs, no traces appearing) | agents-optimize | Load references/observability.md |
| Cold start / initialization too slow | agents-harden | Initialization time section |
Session lifecycle / maxVms / StopRuntimeSession | agents-harden | Session lifecycle management section |
| Long-running background tasks being reclaimed | agents-harden | Long-running background tasks section |
JWT inbound auth failing (403, allowedClients/allowedAudience, issuer mismatch) | agents-harden | Inbound auth section |
| Throttling / quota error / limit increase request | agents-harden | Load references/limits.md |
| Deploy artifact stale or wrong version | agents-deploy | Redeploy workflow |
| Environment broken (CLI, credentials, Node, uv) | Load references/doctor.md | Self-contained in this skill |
State the diagnosis clearly, then tell the developer which skill to use next. If the agent can load the referenced skill in the same session, do so.
© aws, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in plugins/aws-agents/skills/agents-debug of aws/agent-toolkit-for-aws.
Open the folder on GitHubat commit 188af2f
Agents Debug next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agents Debug this skillaws/agent-toolkit-for-aws | 2.8k | — | ~7.7k | Automated safety check: Notes | Apache-2.0 | |
| ForesightHmbown/Wizards-of-the-Ghosts | 109 | — | ~779 | Automated safety check: Pass | CC0-1.0 | |
| Kubernetes Network Root Cause Analysiskubeshark/kubeshark | 12k | — | ~5.3k | Automated safety check: Pass | Apache-2.0 | |
| UModel Root Cause Analysisalibaba/UnifiedModel | 412 | — | ~1.9k | Automated safety check: Pass | Custom licence | |
| Kubernetes Troubleshooting with Inspektor Gadgetinspektor-gadget/inspektor-gadget | 2.9k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Axiom SRE Investigatoropenclaw/clawhub | 9.5k | — | ~7.1k | Automated safety check: Pass | MIT |
Hmbown/Wizards-of-the-Ghosts
Foresight produces bounded forecasts with explicit uncertainty to guide decisions.
kubeshark/kubeshark
Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.
alibaba/UnifiedModel
Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments.
inspektor-gadget/inspektor-gadget
Traces what the kernel is doing for a misbehaving pod using Inspektor Gadget's eBPF tools, tagged with namespace, pod, container and node, without changing workloads.
openclaw/clawhub
Investigates incidents and production problems with hypothesis-driven debugging, queries Axiom observability data when available, and keeps secrets out of commands and output.
different-ai/openwork
Traces an opaque production error in an OpenWork build to its cause using local server logs and Sentry, names the regressing PR and files a report.
aws/agent-toolkit-for-aws
Entry point for AI-agent work on AWS: pick a runtime, plan a migration for existing workloads, and build an executable POC — one phased flow.
aws/agent-toolkit-for-aws
A skill your agent uses to extend an existing agent project with memory, app integration, VPC, multi-agent, migration, model, browser, code interpreter, payments, or resource removal.
aws/agent-toolkit-for-aws
Migrates vibe-coded web applications to AWS. An agent skill from aws/agent-toolkit-for-aws.
aws/agent-toolkit-for-aws
Deploy an event-driven workflow that routes S3 uploads to either Lambda or Fargate via Step Functions based on file size.
aws/agent-toolkit-for-aws
Deploys, queries, and debugs AWS Marketplace usage-based (PAYG) metering — the pipeline (ResolveCustomer, BatchMeterUsage, EventBridge via SAM) and querying/debugging metering records, statuses…
aws/agent-toolkit-for-aws
A skill your agent uses when THIS agent needs to pay for x402-protected content at runtime: hitting a paywall mid-task, settling it via AgentCore Payments, and applying operator-defined spend limits.
Categories
A skill your agent uses when your agent or environment is broken — wrong answers, errors, timeouts, tool failures, or CLI issues. Agents Debug is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Use when your agent or environment is broken — wrong answers, errors, timeouts, tool failures, or CLI issues.
Agents Debug fits situations like: environment is broken — wrong answers; : agent not working; tool call failing; model access denied.
Run `npx skills add aws/agent-toolkit-for-aws --skill agents-debug -a claude-code`. Or copy the skill folder (plugins/aws-agents/skills/agents-debug in aws/agent-toolkit-for-aws) into .claude/skills/agents-debug in your project. Claude Code loads it when a task matches its description.
Run `npx skills add aws/agent-toolkit-for-aws --skill agents-debug -a codex`. Or copy the skill folder (plugins/aws-agents/skills/agents-debug in aws/agent-toolkit-for-aws) into .agents/skills/agents-debug in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aws/agent-toolkit-for-aws --skill agents-debug -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agents-debug, .gemini/skills/agents-debug, .github/skills/agents-debug and .opencode/skills/agents-debug in your project.
Going by SKILL.md and its folder, Agents Debug needs the command-line tools its instructions call (jq, aws and pip) and credentials named GEMINI_API_KEY. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash.
SKILL.md names 2 domains. As links in the text: console.aws.amazon.com and github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Agents Debug is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.7k tokens (SKILL.md is roughly 31k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.5k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Agents Debug: Foresight (Hmbown/Wizards-of-the-Ghosts, 109 stars), Kubernetes Network Root Cause Analysis (kubeshark/kubeshark, 12k stars), UModel Root Cause Analysis (alibaba/UnifiedModel, 412 stars) and Kubernetes Troubleshooting with Inspektor Gadget (inspektor-gadget/inspektor-gadget, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
aws (a GitHub organization, an official publisher) maintains it in aws/agent-toolkit-for-aws, which has 2,825 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on October 7, 2026.
Source: aws/agent-toolkit-for-aws on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.