Weavebench Cua Reproduce
AMAP-ML/LongHorizon-Harness
Reproduce CUA-Harness experiments on WeaveBench from a GitHub checkout.
Autonomous coding agent. An agent skill from aws-samples/sample-strands-agent-with-agentcore.
$ npx skills add aws-samples/sample-strands-agent-with-agentcore --skill code-agent -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install aws-samples/sample-strands-agent-with-agentcore code-agent --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/aws-samples/sample-strands-agent-with-agentcore.git skills-src && mkdir -p .claude/skills && cp -r skills-src/chatbot-app/agentcore/skills/code-agent .claude/skills/code-agent && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "code-agent" agent skill from https://github.com/aws-samples/sample-strands-agent-with-agentcore/tree/main/chatbot-app/agentcore/skills/code-agent into .claude/skills/code-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "code-agent", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/aws-samples/sample-strands-agent-with-agentcore/tree/main/chatbot-app/agentcore/skills/code-agentType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add aws-samples/sample-strands-agent-with-agentcore --skill code-agent -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install aws-samples/sample-strands-agent-with-agentcore code-agent --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws-samples/sample-strands-agent-with-agentcore.git skills-src && mkdir -p .agents/skills && cp -r skills-src/chatbot-app/agentcore/skills/code-agent .agents/skills/code-agent && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "code-agent" agent skill from https://github.com/aws-samples/sample-strands-agent-with-agentcore/tree/main/chatbot-app/agentcore/skills/code-agent into .agents/skills/code-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "code-agent", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aws-samples/sample-strands-agent-with-agentcore --skill code-agent -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install aws-samples/sample-strands-agent-with-agentcore code-agent --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws-samples/sample-strands-agent-with-agentcore.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/chatbot-app/agentcore/skills/code-agent .cursor/skills/code-agent && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "code-agent" agent skill from https://github.com/aws-samples/sample-strands-agent-with-agentcore/tree/main/chatbot-app/agentcore/skills/code-agent into .cursor/skills/code-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "code-agent", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/aws-samples/sample-strands-agent-with-agentcore.git --path chatbot-app/agentcore/skills/code-agent--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add aws-samples/sample-strands-agent-with-agentcore --skill code-agent -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install aws-samples/sample-strands-agent-with-agentcore code-agent --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws-samples/sample-strands-agent-with-agentcore.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/chatbot-app/agentcore/skills/code-agent .gemini/skills/code-agent && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "code-agent" agent skill from https://github.com/aws-samples/sample-strands-agent-with-agentcore/tree/main/chatbot-app/agentcore/skills/code-agent into .gemini/skills/code-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "code-agent", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install aws-samples/sample-strands-agent-with-agentcore code-agentInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add aws-samples/sample-strands-agent-with-agentcore --skill code-agent -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/aws-samples/sample-strands-agent-with-agentcore.git skills-src && mkdir -p .github/skills && cp -r skills-src/chatbot-app/agentcore/skills/code-agent .github/skills/code-agent && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "code-agent" agent skill from https://github.com/aws-samples/sample-strands-agent-with-agentcore/tree/main/chatbot-app/agentcore/skills/code-agent into .github/skills/code-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "code-agent", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aws-samples/sample-strands-agent-with-agentcore --skill code-agent -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install aws-samples/sample-strands-agent-with-agentcore code-agent --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws-samples/sample-strands-agent-with-agentcore.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/chatbot-app/agentcore/skills/code-agent .opencode/skills/code-agent && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "code-agent" agent skill from https://github.com/aws-samples/sample-strands-agent-with-agentcore/tree/main/chatbot-app/agentcore/skills/code-agent into .opencode/skills/code-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "code-agent", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
code-agentAutonomous coding agent. An agent skill from aws-samples/sample-strands-agent-with-agentcore.
Code Agent is an agent skill from aws-samples/sample-strands-agent-with-agentcore, published by the product's own GitHub organization. Autonomous coding agent. Delegate any task that involves understanding, writing, or running code — from a GitHub issue, a bug report, or a user request. It explores, implements, and verifies on its own.
Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `DESIGN.md`, `IMPLEMENT.md` and `REVIEW.md`).
It sits in Testing & QA, covering QA and bug reports. It works with GitHub. The repository describes itself as: Reference architecture for agentic AI chatbots with Strands Agents and Amazon Bedrock AgentCore. The licence is MIT.
Read from SKILL.md and the folder at commit 6dd9d13. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are xml).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Code Agent loads about 3.1k tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 1,618 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from aws-samples/sample-strands-agent-with-agentcore at commit 6dd9d13, republished under its MIT licence (© aws-samples). 1,618 words, ~3,122 tokens.
.claude/skills/code-agent/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.An autonomous coding agent. It doesn't just write code on demand — it thinks through problems, forms its own plan, reads the existing codebase to understand context, implements solutions iteratively, and verifies they work before finishing.
Given a goal, it will:
Brief it like you'd brief a capable engineer: describe what you want to achieve, not how to do it.
| Code Agent | Code Interpreter | |
|---|---|---|
| Nature | Autonomous agent (Claude Code) | Sandboxed execution environment |
| Best for | Multi-file projects, refactoring, test suites | Quick scripts, data analysis, prototyping |
| File persistence | All files auto-synced to S3, accessible via workspace tools | Files in /mnt/workspace persist across interpreter restarts |
| Session state | Files + conversation persist across sessions | Variables persist within one interpreter session; workspace files persist for the chat session |
| Autonomy | Plans, writes, runs, and iterates independently | You write the code it executes |
| Use when | You need an engineer to solve a problem end-to-end | You need to run a specific piece of code |
The code agent runs in an isolated container dedicated solely to this session. Its filesystem, running processes, and local ports are completely separate from your own environment — do not attempt to access its paths or local servers via browser or other tools.
User-uploaded session files are synchronized from the canonical Workspace
inputs/ prefix before every Code Agent task and are available under
inputs/<filename>. Treat inputs/ as read-only; generated files belong
elsewhere in the Code Agent workspace.
Every code_agent call must include workspace_paths. Pass the exact
inputs/<filename> paths required by the task, or [] when the task does not
use session attachments. The task fails before execution if a required file is
not synchronized; do not ask the Code Agent to reconstruct a missing source.
Trust the code agent's reasoning and autonomy — delegate not just implementation but also testing, verification, and iteration. Only step in when there's a genuine constraint the agent cannot resolve on its own; in that case, surface it to the user and decide together.
You give direction and verify results. The agent explores, implements, and checks in when it hits a genuine decision point.
Trust the agent to deliver. Don't over-specify the how — focus on the what. For complex tasks, break work into phases and steer between turns. Surface critical design decisions to the user early, then execute autonomously.
The code agent can read the entire workspace. What it can't do is reach outside it. That's where you add value.
Your job is to bring in what the agent can't get on its own:
What you should NOT be doing:
Reading a file to spot-check the agent's output is fine. Spending time reading 10 files to diagnose a problem yourself — then handing the agent a pre-solved task — is not. That's the agent's job.
| You (orchestrator) provide | Code agent discovers on its own |
|---|---|
| What the user wants — goals, constraints, preferences | How to implement — codebase structure, existing patterns, design decisions |
| External context the agent can't reach — API docs, user requirements, npm/registry info | Internal context from the workspace — file layout, dependencies, coding conventions |
| Resolved decisions — framework choice, scope boundaries | Implementation decisions — variable naming, module structure, error strategies |
When the code agent encounters a requirements-level question it can't resolve from the codebase alone (e.g., "should this be public or internal?", "which auth provider?"), it will surface it. That's the right behavior — resolve it and pass the answer back. Don't try to pre-answer every possible question; let the agent ask when it genuinely needs direction.
The goal is to deliver the best possible result with minimal friction. The key is how you (orchestrator) and the code agent collaborate — not just fire-and-forget.
One at a time. The code agent runs as a single process against one workspace. Always wait for the current call to complete before making the next one. Never issue parallel
code_agentcalls — they will conflict and produce broken results.
Timeout awareness. Each code agent call has a ~30-minute practical limit. For large tasks, break them into focused phases (explore → implement → test) rather than sending a single massive request. If a task might exceed this, split it proactively — don't wait for a timeout error.
code_agent(task="Fix the typo in src/config.ts line 42: 'recieve' → 'receive'",
workspace_paths=[],
task_complexity="low")code_agent(task="Add input validation to the /api/users endpoint.
Validate email format and required fields. Add tests.",
workspace_paths=[],
task_complexity="medium")# Turn 1: Explore & plan
code_agent(task="Explore how auth works and propose a plan for adding JWT.
Do NOT modify files yet.", workspace_paths=[], task_complexity="high")
# Review the plan the agent returns — does the approach make sense?
# Turn 2: Implement
code_agent(task="Implement JWT middleware with httpOnly cookies.",
workspace_paths=[])
# Turn 3: Integrate
code_agent(task="Apply middleware to routes. Exclude /api/public.",
workspace_paths=[])
# Turn 4: Verify
code_agent(task="Run full test suite and fix any failures.",
workspace_paths=[])
# → Report to userUse your judgment. The complexity of the delegation should match the complexity of the task. Don't over-orchestrate simple work, but don't fire-and-forget complex multi-file changes either.
For complex tasks, the orchestrator and code agent naturally go back and forth. This happens autonomously — the user doesn't need to be involved in each turn:
Turn 1: Explore → Agent returns findings + proposed plan
Turn 2: Implement core → Agent returns results
Turn 3: Fix issue found in Turn 2 → Agent iterates
Turn 4: Run tests → All pass
→ Report to user: "JWT auth added. 4 files changed, 12 tests pass."The user sees real-time terminal progress throughout. They only get pulled in if a genuine design decision emerges that the code agent can't resolve from the codebase alone.
Before diving into implementation, scan for genuine ambiguities that only the user can resolve:
If you spot these, ask the user before delegating implementation. But most tasks don't need this — if the codebase and user request are clear enough, just proceed.
Important: Ask only what the user must decide. Don't ask about implementation details the agent can figure out. Don't ask "should I proceed?" — just proceed after resolving any genuine decision point.
When the code agent finishes, summarize concisely. Do NOT pass through raw code, full file contents, or verbose agent output. The user sees the code agent's terminal activity in real-time — they don't need it repeated.
Include:
Do NOT include:
Example format:
Files changed:
- src/middleware/rateLimiter.ts — added rate limiting logic (new file)
- tests/rateLimiter.test.ts — added 4 tests; all pass
Verified: ran full test suite (42 tests, 0 failures)
Note: rate limit is currently per-IP. If per-user-ID is needed later,
the key function can be swapped without touching routes.→ DESIGN.md — requirements capture, scope decisions, trade-off escalation → IMPLEMENT.md — stepwise delegation, steering, correctness verification → REVIEW.md — iterative review, complexity-based depth, known issue checklist
compact_session=True — before a new task in a long session. Summarizes history, saves tokens, preserves context.reset_session=True — only when switching to a completely unrelated project. Clears history, keeps workspace files.A long conversation that handles multiple unrelated tasks is a liability — earlier context bleeds into later tasks and causes subtle wrong assumptions. When switching to a significantly different task (e.g., bug fix → new feature, frontend → backend), use compact_session=True to summarize and reset context. This is especially important when the nature of the work changes, not just the file being edited.
| Delegate to code_agent | Handle directly |
|---|---|
| Implement from a GitHub issue or feature request | Explain how an algorithm works |
| Investigate code to figure out an implementation approach | Write a short standalone snippet |
| Fix a failing test or bug | Answer a syntax or API question |
| Refactor a module | Simple code review without changes |
| Analyze uploaded source files | Generate a one-off script with no files |
| Run tests and fix failures | Summarize what code does |
| Scaffold following project conventions |
Files uploaded by the user are automatically available in the workspace:
task = "Unzip the uploaded my-project.zip and summarize the architecture."Only use this when requirements are already fully resolved and you need explicit acceptance criteria. For most tasks, a plain description works better.
<task>
<objective>Verifiable "done" state.</objective>
<scope>What area of the system to work within. What to leave alone.</scope>
<context>API signatures, versions, prior research findings.</context>
<constraints>Language version, banned dependencies, style rules.</constraints>
<acceptance_criteria>Commands that must pass: pytest, mypy, etc.</acceptance_criteria>
</task>Code Agent:
© aws-samples, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files in chatbot-app/agentcore/skills/code-agent of aws-samples/sample-strands-agent-with-agentcore.
Open the folder on GitHubat commit 6dd9d13
Code Agent next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Code Agent this skillaws-samples/sample-strands-agent-with-agentcore | 194 | — | ~3.1k | Automated safety check: Pass | MIT | |
| Weavebench Cua ReproduceAMAP-ML/LongHorizon-Harness | 1.7k | — | ~1.6k | Automated safety check: Pass | MIT | |
| Evidence-Driven Testingmichaelshimeles/skills | 1.3k | 1 repos | ~3.9k | Automated safety check: Pass | None | |
| Create GitHub IssueNVIDIA/OpenShell | 15k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Triage IssuesClickHouse/clickhouse-java | 1.6k | — | ~904 | Automated safety check: Pass | Apache-2.0 | |
| Gentle AI Issue CreationGentleman-Programming/gentle-shell | 1.2k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 |
AMAP-ML/LongHorizon-Harness
Reproduce CUA-Harness experiments on WeaveBench from a GitHub checkout.
michaelshimeles/skills
Records an annotated screen recording of the agent testing an app hands-on, then posts the video and a results summary to the PR and tracker issue.
NVIDIA/OpenShell
Create GitHub issues using the gh CLI. An agent skill from NVIDIA/OpenShell.
ClickHouse/clickhouse-java
Analyzes a single GitHub issue at a time. An agent skill from ClickHouse/clickhouse-java.
Gentleman-Programming/gentle-shell
Create and triage GitHub issues from repository evidence. An agent skill from Gentleman-Programming/gentle-shell.
termio-sh/termio
Diagnose a termio hang, beachball, crash, or 'it froze again' from the evidence macOS and termio actually leave behind — live process samples, crash and CPU-burn reports, the unified log, the…
aws-samples/sample-strands-agent-with-agentcore
Guide users through a structured workflow for co-authoring documentation.
aws-samples/sample-strands-agent-with-agentcore
Test and prototype code in a sandboxed environment. An agent skill from aws-samples/sample-strands-agent-with-agentcore.
aws-samples/sample-strands-agent-with-agentcore
Deep research with structured reports and charts. An agent skill from aws-samples/sample-strands-agent-with-agentcore.
aws-samples/sample-strands-agent-with-agentcore
Read and write files in the shared session workspace. An agent skill from aws-samples/sample-strands-agent-with-agentcore.
aws-samples/sample-strands-agent-with-agentcore
Delegate one independent, context-heavy task to an isolated asynchronous analyst or reviewer.
aws-samples/sample-strands-agent-with-agentcore
Create hand-drawn style diagrams and flowcharts using Excalidraw
Works with
Categories
Autonomous coding agent. An agent skill from aws-samples/sample-strands-agent-with-agentcore. Code Agent is an agent skill from aws-samples/sample-strands-agent-with-agentcore, published by the product's own GitHub organization. Autonomous coding agent.
Code Agent fits situations like: tasks that involve QA and bug reports.
Run `npx skills add aws-samples/sample-strands-agent-with-agentcore --skill code-agent -a claude-code`. Or copy the skill folder (chatbot-app/agentcore/skills/code-agent in aws-samples/sample-strands-agent-with-agentcore) into .claude/skills/code-agent in your project. Claude Code loads it when a task matches its description.
Run `npx skills add aws-samples/sample-strands-agent-with-agentcore --skill code-agent -a codex`. Or copy the skill folder (chatbot-app/agentcore/skills/code-agent in aws-samples/sample-strands-agent-with-agentcore) into .agents/skills/code-agent in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aws-samples/sample-strands-agent-with-agentcore --skill code-agent -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/code-agent, .gemini/skills/code-agent, .github/skills/code-agent and .opencode/skills/code-agent in your project.
SKILL.md names no scripts, command-line tools or credentials: Code Agent is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Code Agent is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Code Agent: Weavebench Cua Reproduce (AMAP-ML/LongHorizon-Harness, 1.7k stars), Evidence-Driven Testing (michaelshimeles/skills, 1.3k stars), Create GitHub Issue (NVIDIA/OpenShell, 15k stars) and Triage Issues (ClickHouse/clickhouse-java, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
aws-samples (a GitHub organization, an official publisher) maintains it in aws-samples/sample-strands-agent-with-agentcore, which has 194 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 6, 2026.
Source: aws-samples/sample-strands-agent-with-agentcore on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.