Crabbox
openclaw/openclaw
Crabbox and Blacksmith Testbox remote testing: isolation, cross-platform E2E, diagnostics, cleanup.
Run one already-written new-e2e test locally and triage the setup failures that stop it — "run the containers e2e tests", "my e2e run fails before any test starts".
$ npx skills add DataDog/datadog-agent --skill run-e2e -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install DataDog/datadog-agent run-e2e --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/DataDog/datadog-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/run-e2e .claude/skills/run-e2e && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "run-e2e" agent skill from https://github.com/DataDog/datadog-agent/tree/main/.agents/skills/run-e2e into .claude/skills/run-e2e/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-e2e", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/DataDog/datadog-agent/tree/main/.agents/skills/run-e2eType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add DataDog/datadog-agent --skill run-e2e -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install DataDog/datadog-agent run-e2e --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/DataDog/datadog-agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/run-e2e .agents/skills/run-e2e && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "run-e2e" agent skill from https://github.com/DataDog/datadog-agent/tree/main/.agents/skills/run-e2e into .agents/skills/run-e2e/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-e2e", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add DataDog/datadog-agent --skill run-e2e -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install DataDog/datadog-agent run-e2e --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/DataDog/datadog-agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/run-e2e .cursor/skills/run-e2e && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "run-e2e" agent skill from https://github.com/DataDog/datadog-agent/tree/main/.agents/skills/run-e2e into .cursor/skills/run-e2e/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-e2e", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/DataDog/datadog-agent.git --path .agents/skills/run-e2e--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add DataDog/datadog-agent --skill run-e2e -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install DataDog/datadog-agent run-e2e --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/DataDog/datadog-agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/run-e2e .gemini/skills/run-e2e && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "run-e2e" agent skill from https://github.com/DataDog/datadog-agent/tree/main/.agents/skills/run-e2e into .gemini/skills/run-e2e/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-e2e", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install DataDog/datadog-agent run-e2eInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add DataDog/datadog-agent --skill run-e2e -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/DataDog/datadog-agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/run-e2e .github/skills/run-e2e && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "run-e2e" agent skill from https://github.com/DataDog/datadog-agent/tree/main/.agents/skills/run-e2e into .github/skills/run-e2e/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-e2e", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add DataDog/datadog-agent --skill run-e2e -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install DataDog/datadog-agent run-e2e --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/DataDog/datadog-agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/run-e2e .opencode/skills/run-e2e && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "run-e2e" agent skill from https://github.com/DataDog/datadog-agent/tree/main/.agents/skills/run-e2e into .opencode/skills/run-e2e/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-e2e", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
run-e2eRun one already-written new-e2e test locally and triage the setup failures that stop it — "run the containers e2e tests", "my e2e run fails before any test starts".
Run E2E is an agent skill from DataDog/datadog-agent, published by the product's own GitHub organization. Run one already-written new-e2e test locally and triage the setup failures that stop it — "run the containers e2e tests", "my e2e run fails before any test starts".
Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `references/devenv.md`, `references/flags.md` and `references/setup.md`).
It sits in Testing & QA, covering End-to-end testing. It works with Amazon Web Services. The repository describes itself as: Main repository for Datadog Agent. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 20eff25. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashReadAskUserQuestionFrom allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonpulumigitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Run E2E loads about 2.2k tokens when it runs, and up to ~7.1k if it reads all its reference files. Until then it costs about 43 tokens; SKILL.md has 1,108 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Bash, Read, AskUserQuestionAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from DataDog/datadog-agent at commit 20eff25, republished under its Apache-2.0 licence (© DataDog). 1,108 words, ~2,247 tokens.
.claude/skills/run-e2e/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Run a single new-e2e target with dda inv -- new-e2e-tests.run. Most targets provision real
infrastructure in a cloud account — usually AWS, sometimes GCP or Azure — though the framework also has
local provisioners that cost nothing but time. Run duration is a property of the target: minutes for one
VM, considerably longer for a Kubernetes cluster. Either way, aim for one correct run rather than a fast
iteration loop.
| Reference | Load when |
|---|---|
references/devenv.md | devenv_e2e.py exits non-zero, or the container behaves differently from the host |
references/setup.md | Before offering to run dda inv -- e2e.setup |
references/troubleshooting.md | Any failure before the first --- PASS/--- FAIL line |
references/flags.md | The request needs more than a target and a test name |
A target is a package path relative to /test/new-e2e/, like ./tests/agent-subcommands/flare; the
task resolves against that module, so repeating the prefix is wrong.
Requests usually name a test, not a package — "run the flare e2e test" is the normal shape. Use a target if given, otherwise ask, offering candidates you already know. Never search the tree: a guessed target provisions the wrong thing and you find out after paying for it.
Anchor a supplied test name (--run '^TestFlareSuite$'), or TestFlare also selects TestFlareOpts.
test -n "$WORKSPACE_NAME" && echo IN_WORKSPACE
test -f /.started && echo IN_DEVENV || echo ON_HOSTThe dev env entrypoint creates /.started. IN_WORKSPACE → 3C — a workspace already provides
the environment the dev env would build, so run directly on it without one. ON_HOST → step 3A,
IN_DEVENV → step 3B, --host → 3C. devenv_e2e.py up also refuses to run when
WORKSPACE_NAME is set (exit 7), so this cannot slip through to 3A.
python .agents/skills/run-e2e/scripts/devenv_e2e.py up --jsonStarts a dev env at id e2e-run if needed, gives it the host's E2E config and keypair, establishes
Pulumi's backend, checks AWS access, and prints the run_prefix for step 5. Idempotent, so a reused env
pays the setup cost once. Add --no-aws-check for a locally-provisioned target to skip the SSO
acceptance; the host still needs AWS config, because the run task requires it whatever the target is.
The AWS check needs the user present — authorizing a new container means completing an SSO flow whose
browser tab opens on their desktop. Warn them, and if it gives up, relay the aws-vault login it prints
and rerun up. That is normal on a new env.
Every failure prints an actionable message; relay it. The table picks the reference file and says which machine the remedy belongs on.
| Exit | Meaning | What to do |
|---|---|---|
| 0 | Ready | Step 4, using the printed run_prefix |
| 2 | Host has no usable ~/.test_infra_config.yaml | Read references/setup.md, run dda inv -- e2e.setup --team=<github-team> on the host (ask the user for their team first; never the interactive form), retry |
| 3 | The container would not hold this working tree | Follow the printed remedy; references/devenv.md per case. Offer --host if the checkout cannot be used |
| 4 | The container cannot authenticate to AWS | Run the printed aws-vault login inside the env, then retry |
| 5 | Already inside a dev env | Step 2 misread the marker; go to 3B |
| 6 | Env is in error, so its stacks cannot be checked | Do not remove it for them; relay the message, which says when recreating is safe |
| 7 | WORKSPACE_NAME is set: this is a workspace, not a host that needs a dev env | Step 3C — run directly on the workspace |
| other | No dedicated remedy | Relay the message; references/troubleshooting.md |
Azure and GCP targets are not handled — only AWS credentials reach the container. Use --host.
test -f ~/.test_infra_config.yaml && echo CONFIG_OK || echo CONFIG_MISSING
pulumi whoami >/dev/null 2>&1 && echo BACKEND_OK || echo BACKEND_MISSINGBACKEND_MISSING → PULUMI_SKIP_UPDATE_CHECK=true dda inv -- e2e.setup --no-interactive. Do not install
Pulumi; the image ships it, only its plugins and backend are missing. CONFIG_MISSING → stop; this env
has no E2E identity, so have them recreate it with devenv_e2e.py up from the host. Never the
interactive dda inv -- e2e.setup here — see references/setup.md.
This container needs its own AWS authorization, which nothing on the host provides. Have them run
aws-vault login sso-agent-sandbox-account-admin-8h first rather than discovering it ten minutes in.
Then step 4 with a suffix identifying them, because stack names take the container user name dd and
would otherwise collide. git config user.email is set from the host.
E2E_STACK_NAME_SUFFIX=<you> dda inv -- new-e2e-tests.run --targets=<target> [--run <regex>] [flags]--host escape hatchCheck ~/.test_infra_config.yaml, pulumi whoami, and a live AWS session — here the host's own counts —
then run dda inv -- new-e2e-tests.run directly. Faster on a configured Linux or macOS machine, and
Pulumi state survives there. Not the default because an unconfigured or Windows host fails in ways the
dev env does not.
Get an explicit yes for: the exact command, target and --run, where it runs, what it provisions and in
which account, roughly how long, and whether the stack is destroyed afterwards. If you cannot tell what
it provisions, say so — that is worth confirming before paying for it.
Use the run_prefix from step 3A verbatim and append the test command. Do not add container paths of
your own: on Windows, Git Bash rewrites them before they reach the container, and everything the run
needs is already in the config the bootstrap installed.
dda env dev run -t linux-container --id e2e-run -- env <env_args...> \
dda inv -- new-e2e-tests.run --targets=<target> [--run <regex>] [flags]Start it with run_in_background: true; these outlast a foreground Bash call. The first run in a fresh
env is much the slowest — the test binary compiles from a cold cache before any infrastructure is
touched, so several minutes of silence is normal.
Report once a minute while it runs. Poll the background output about every minute and post a
one-line status update: which phase it is in (compiling · provisioning · running tests · tearing down)
and the last meaningful output line. Long silent stretches are inherent to the run — the user cannot
see your terminal, so from their side an agent that goes quiet for 10 minutes looks exactly like a
hung one. Silence from the tool is normal; silence from you is not. If the phase has not changed,
say so and note the elapsed time (still provisioning, ~6m in, no new output) rather than skipping
the update.
### E2E run — <target> [--run <regex>]
- Where: dev env `e2e-run` | host
- Command: <exact command as executed>
- Result: PASS | FAIL | SETUP FAILURE (failed before any test ran)
- Duration: <mm:ss>
- Stack: <name> — destroyed | kept (--keep-stack)
- Failures:
- <TestSuite/TestName> — <one-line reason>
- Diagnostics: <path> (+ the `docker cp` to retrieve it, if it ran in a dev env)
- Next step: <the single most useful action>For a SETUP FAILURE, take the symptom to references/troubleshooting.md rather than reporting raw
stderr.
The env is reusable and costs nothing idle, so leave it unless asked. When asked:
python .agents/skills/run-e2e/scripts/devenv_e2e.py downIt exits 6 rather than removing an env whose Pulumi stacks are live or uncheckable — that state exists
nowhere else, so an orphaned cluster is the cost of getting this wrong. It prints the destroy command.
--force overrides it; only reach for that once you have confirmed nothing is running. --keep-stack
implies keeping the env too.
"run TestVMSuite in ./examples" — the default path, from a host that may not be configured
python .agents/skills/run-e2e/scripts/devenv_e2e.py up --json
dda env dev run -t linux-container --id e2e-run -- env E2E_STACK_NAME_SUFFIX=alice \
dda inv -- new-e2e-tests.run --targets=./examples --run='^TestVMSuite$'"just run it here, my machine is already set up" —
--host, skipping the container
dda inv -- new-e2e-tests.run --targets=./examples --run='^TestVMSuite$'© DataDog, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (scripts, references) in .agents/skills/run-e2e of DataDog/datadog-agent.
Open the folder on GitHubat commit 20eff25
Run E2E next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Run E2E this skillDataDog/datadog-agent | 3.8k | — | ~2.2k | Automated safety check: Notes | Apache-2.0 | |
| Crabboxopenclaw/openclaw | 392k | — | ~3.7k | Automated safety check: Pass | MIT | |
| New Integgo-to-k/cdkd | 143 | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | |
| Web Application Testinganthropics/skills | 180k | 51 repos | ~966 | Automated safety check: Pass | Apache-2.0 | |
| OpenHarness End-to-End EvalsHKUDS/OpenHarness | 16k | 1 repos | ~2.1k | Automated safety check: Notes | MIT | |
| playwright-cli Browser Automationgithub/gh-aw | 5.4k | 24 repos | ~2.8k | Automated safety check: Pass | MIT |
openclaw/openclaw
Crabbox and Blacksmith Testbox remote testing: isolation, cross-platform E2E, diagnostics, cleanup.
go-to-k/cdkd
Scaffold a new integration test for cdkd. An agent skill from go-to-k/cdkd.
anthropics/skills
Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.
HKUDS/OpenHarness
Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.
github/gh-aw
Drives a real browser from the command line with playwright-cli to open pages, interact, mock requests, save state and work with Playwright tests.
appsmithorg/appsmith
Writes a Playwright end-to-end test from a prompt, runs it against a live Appsmith deployment and retries with fixes up to three times until it passes.
DataDog/datadog-agent
Classify a failed CI as either caused by an active incident, flakiness, or a true code regression.
DataDog/datadog-agent
Run a structured discovery session to build an Allium specification through conversation.
DataDog/datadog-agent
Monitor the current PR's GitLab pipeline to completion, then report success, auto-fix, or investigate a failure.
DataDog/datadog-agent
A skill your agent uses when an engineer or manager asks to recap, summarize, or post an update on a Jira Epic — a progress update for an in-progress Epic (how far along it is, what's shipped so…
DataDog/datadog-agent
Explains a lading.yaml config file from the regression test suite, using the lading Rust source as ground truth for field meanings and defaults.
DataDog/datadog-agent
Extract an Allium specification from an existing codebase. An agent skill from DataDog/datadog-agent.
Works with
Categories
Run one already-written new-e2e test locally and triage the setup failures that stop it — "run the containers e2e tests", "my e2e run fails before any test starts". Run E2E is an agent skill from DataDog/datadog-agent, published by the product's own GitHub organization. Run one already-written new-e2e test locally and triage the setup failures that stop it — "run the containers e2e tests", "my e2e run fails before any test starts".
Run E2E fits situations like: tasks that involve End-to-end testing.
Run `npx skills add DataDog/datadog-agent --skill run-e2e -a claude-code`. Or copy the skill folder (.agents/skills/run-e2e in DataDog/datadog-agent) into .claude/skills/run-e2e in your project. Claude Code loads it when a task matches its description.
Run `npx skills add DataDog/datadog-agent --skill run-e2e -a codex`. Or copy the skill folder (.agents/skills/run-e2e in DataDog/datadog-agent) into .agents/skills/run-e2e in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add DataDog/datadog-agent --skill run-e2e -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run-e2e, .gemini/skills/run-e2e, .github/skills/run-e2e and .opencode/skills/run-e2e in your project.
Going by SKILL.md and its folder, Run E2E needs Python for the scripts in its folder and the command-line tools its instructions call (python, pulumi and git). Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Bash, Read, AskUserQuestion.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Run E2E is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.2k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Run E2E: Crabbox (openclaw/openclaw, 392k stars), New Integ (go-to-k/cdkd, 143 stars), Web Application Testing (anthropics/skills, 180k stars) and OpenHarness End-to-End Evals (HKUDS/OpenHarness, 16k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
DataDog (a GitHub organization, an official publisher) maintains it in DataDog/datadog-agent, which has 3,757 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 8, 2026.
Source: DataDog/datadog-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.