Official agent skill

E2E Audit

by DataDog in DataDog/datadog-agent

Judge whether Agent behavior belongs in a new-e2e, integration, or unit test

OfficialApache-2.0Auto-check: notesTesting & QA

Install E2E Audit

skills CLI
$ npx skills add DataDog/datadog-agent --skill e2e-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install DataDog/datadog-agent e2e-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/DataDog/datadog-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/e2e-audit .claude/skills/e2e-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
e2e-audit
GitHub stars
3.8k
Token cost
~1.9k tokens
SKILL.md length
1,036 words
Files
1
Skills in repo
35
Repo updated
First seen
Licence
Apache-2.0

At a glance

Judge whether Agent behavior belongs in a new-e2e, integration, or unit test

  • Works in 5 steps: List the observable claims that the… → For each assertion, state the failure it… → Identify the smallest environment that… → …
  • Tasks that involve End-to-end testing
  • SKILL.md covers Principle, Procedure and Output
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

E2E Audit is an agent skill from DataDog/datadog-agent, published by the product's own GitHub organization. Judge whether Agent behavior belongs in a new-e2e, integration, or unit test

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering End-to-end testing and Unit testing. The repository describes itself as: Main repository for Datadog Agent. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve End-to-end testing
  • Tasks that involve Unit testing

Example prompts

  • “/e2e-audit”

Requirements

  • Pre-approved tools (allowed-tools): Read, Glob, Grep, Bash

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. List the observable claims that the proposed test would assert.
  2. For each assertion, state the failure it is intended to catch.
  3. Identify the smallest environment that preserves each failure mode.
  4. Classify each assertion. If E2E is justified, name the real boundary that
  5. Give an overall verdict based on the assertion requiring the broadest real

What it can do on your machine

Read from SKILL.md and the folder at commit 20eff25. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Glob
    • Grep
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

E2E Audit loads about 1.9k tokens when it runs. Until then it costs about 22 tokens; SKILL.md has 1,036 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~22
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Glob, Grep, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from DataDog/datadog-agent at commit 20eff25, republished under its Apache-2.0 licence (© DataDog). 1,036 words, ~1,867 tokens.

Download SKILL.mdSave it as .claude/skills/e2e-audit/SKILL.md (or your agent's skills folder).
name
e2e-audit
description
Judge whether Agent behavior belongs in a new-e2e, integration, or unit test
allowed-tools
Read, Glob, Grep, Bash
argument-hint
<behavior description> | <path-to-test-file-or-dir> [more paths...]
model
sonnet

Decide whether assertions about proposed or existing Agent behavior belong in a full new-e2e test, an integration test, or a unit test. Produce only a verdict and rationale. Never edit, move, or delete tests.

Principle

Analyze observable claims and failure modes, not entire test files or individual assert or require calls. Group checks that validate the same contract as one assertion, including implicit claims such as successful installation, startup, or command execution.

Choose the cheapest test that preserves the boundary and failure mode that matter. Do not ask only whether an assertion can be made with mocks; ask what the assertion would stop validating if its dependencies were replaced.

E2E justified

Use new-e2e when the deployed Agent or its real environment is part of the behavior being validated. This includes boundaries such as:

  • installation, packaging, permissions, or merged deployed configuration;
  • process or service lifecycle, cross-process communication, or CLI behavior that depends on the running Agent;
  • kernel, operating-system, cloud, container-orchestrator, or network behavior that cannot be represented faithfully by a local dependency;
  • user-visible data reaching fakeintake when the Agent's real assembly, configuration, encoding, or forwarding path is material to the test.

A behavior is not E2E-worthy merely because the current test reaches it through SSH, a CLI, or remote infrastructure. If those layers add no relevant coverage, use a lower-level test.

Should be an integration test

Use an integration test when the important boundary can be preserved locally, for example by wiring components with fx.Test, using a local fakeintake, or using a real local daemon, driver, or hardware dependency. A real local dependency does not by itself require new-e2e.

Should be a unit test

Use a unit test when the behavior is isolated logic and does not require a real component graph, process, or external dependency.

Layered coverage and residual E2E value

Duplicating an E2E assertion in a unit or integration test is useful when the lower-level test preserves its failure mode: it provides faster PR feedback, more deterministic failures, and easier debugging. Do not treat this useful duplication as waste by itself.

After identifying lower-level coverage, reevaluate what the E2E test uniquely validates. Repeating the same assertion through the deployed Agent can still be valuable when it catches assembly, configuration, packaging, lifecycle, or forwarding failures that the lower-level test cannot. If nearly all material assertions have equivalent lower-level coverage and the E2E test preserves no meaningful additional boundary, its feedback no longer justifies its provisioning, runtime, and maintenance cost; recommend removing it. Base this decision on residual failure coverage, not only the number of duplicated assertions.

Procedure

Proposed behavior
  1. List the observable claims that the proposed test would assert.
  2. For each assertion, state the failure it is intended to catch.
  3. Identify the smallest environment that preserves each failure mode.
  4. Classify each assertion. If E2E is justified, name the real boundary that would be lost in a lower-level test.
  5. Give an overall verdict based on the assertion requiring the broadest real boundary.

If the description does not establish the relevant assertions or boundaries, ask for the missing information or return an explicitly uncertain verdict.

Show full SKILL.md (518 more words)Show less
Existing tests
  1. Resolve every concrete suite. For directories, find all *_test.go files. Follow shared suites, setup code, helpers, and provisioners rather than judging files in isolation.
  2. Read each test and subtest, including setup, gating, environment updates, and cleanup. Inspect the production code it exercises before deciding that an assertion can be tested at a lower level.
  3. Inventory the observable claims in each test. Include implicit claims from setup and lifecycle operations, and group checks that cover the same contract.
  4. For each assertion, state the failure it catches, identify the smallest environment that preserves that failure, note equivalent lower-level coverage, and classify it using the principle above.
  5. Give the suite-level verdict:
    • If any assertion needs the deployed Agent boundary, the suite remains E2E; name the assertion and boundary that determine this verdict.
    • If none does, recommend an integration or unit test as appropriate.
  6. Recommend duplicating suitable assertions at a lower level when that would provide faster PR feedback, more deterministic failures, or easier debugging, even if the suite initially remains E2E.
  7. After accounting for that lower-level coverage, identify the failure modes and real boundaries that only the E2E test preserves. If nearly all material assertions are duplicated and no meaningful E2E-only boundary remains, recommend removing the E2E test. Weigh its residual value against runtime, provisioning, flakiness, and maintenance cost; do not assume all cost is paid only once at suite setup.
  8. When reviewing several suites, briefly note redundant provisioners or equivalent coverage that could be consolidated.

For large reviews, inspect files in parallel if possible, then verify and synthesize the results.

Output

Use one of these verdicts:

  • E2E justified
  • Should be an integration test
  • Should be a unit test

For proposed behavior, classify each expected assertion and return an overall verdict with a short reason. For existing tests, classify each material assertion and return one verdict per concrete suite, naming the assertion and boundary that determine it. State what the E2E test uniquely validates after accounting for lower-level coverage, or say that no material E2E-only boundary remains. When several assertions differ, use a concise table with these columns: assertion, failure caught, smallest environment, and classification. Add only material uncertainty, lower-level candidates, or consolidation opportunities.

Examples

Input: Verify that installing the Agent package creates a running service with the expected permissions and that data reaches fakeintake after a reboot.

Output: E2E justified — The package installation, service lifecycle, permissions, reboot, and forwarding path are the behavior under test; a lower-level test would not preserve those deployed-system boundaries.

Input: Verify that the assembled Agent components transform a payload and send it to a local fakeintake.

Output: Should be an integration test — A locally assembled component graph and fakeintake preserve the component wiring and payload boundary without provisioning remote infrastructure.

Input: Verify that the configuration parser rejects a negative timeout and applies the default when the field is absent.

Output: Should be a unit test — This is isolated parsing and validation logic that does not require a real component graph, process, or external dependency.

Do not propose implementation changes unless the user asks for them.

© DataDog, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/e2e-audit of DataDog/datadog-agent.

Open the folder on GitHubat commit 20eff25

Compare with similar skills

E2E Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

E2E Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
E2E Audit this skillDataDog/datadog-agent3.8k—~1.9kAutomated safety check: NotesApache-2.0
TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph11211 repos~2.4kAutomated safety check: PassNone
Contractssamchon/nestia2.2k—~1.3kAutomated safety check: PassMIT
Atdd Mutateswingerman/engineer154—~2.7kAutomated safety check: PassMIT
E2E Scenario Testingobra/dotfiles116—~1.8kAutomated safety check: PassNone
TDD Workflowaffaan-m/ECC275k2 repos~2kAutomated safety check: PassMIT

Similar skills

  • TDD Workflow

    hellangleZ/burn-in-cceverywhere-ralph

    A skill your agent uses when writing new features, fixing bugs, or refactoring code.

    112 GitHub starsUsed in 11 repos~2.4k tokens
    Testing & QAAuto-check passed
  • Contracts

    samchon/nestia

    Defines self-acknowledgments for production declarations and tests.

    2.2k GitHub stars~1.3k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Atdd Mutate

    swingerman/engineer

    A skill your agent uses to add a third validation layer to the ATDD workflow — after acceptance tests verify WHAT and unit tests verify HOW, mutation testing verifies the tests actually catch bugs.

    154 GitHub stars~2.7k tokensUpdated 15 days ago
    Testing & QAAuto-check passed
  • A skill your agent uses when verifying a running application end-to-end through its real interface — a web UI, a CLI, or a TUI — by writing and executing agent-run "scenario cards" against a freshly…

    116 GitHub stars~1.8k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • TDD Workflow

    affaan-m/ECC

    新機能の作成、バグ修正、コードのリファクタリング時にこのスキルを使用します。ユニット、統合、E2Eテストを含む80%以上のカバレッジでテスト駆動開発を強制します。

    275k GitHub starsUsed in 2 repos~2k tokens
    Testing & QAAuto-check passed
  • TDD Workflow

    affaan-m/ECC

    새 기능 작성, 버그 수정 또는 코드 리팩터링 시 이 스킬을 사용하세요. An agent skill from affaan-m/ECC.

    275k GitHub starsUsed in 2 repos~2.1k tokens
    Testing & QAAuto-check passed

More from DataDog/datadog-agent

All 35 skills in this repo
  • Triage CI Failure

    DataDog/datadog-agent

    Official

    Classify a failed CI as either caused by an active incident, flakiness, or a true code regression.

    3.8k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Elicit

    DataDog/datadog-agent

    Official

    Run a structured discovery session to build an Allium specification through conversation.

    3.8k GitHub starsUsed in 1 repo~3.7k tokens
    Auto-check passed
  • Follow PR

    DataDog/datadog-agent

    Official

    Monitor the current PR's GitLab pipeline to completion, then report success, auto-fix, or investigate a failure.

    3.8k GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Create Epic Recap

    DataDog/datadog-agent

    Official

    A skill your agent uses when an engineer or manager asks to recap, summarize, or post an update on a Jira Epic — a progress update for an in-progress Epic (how far along it is, what's shipped so…

    3.8k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Explain Lading Config

    DataDog/datadog-agent

    Official

    Explains a lading.yaml config file from the regression test suite, using the lading Rust source as ground truth for field meanings and defaults.

    3.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Distill

    DataDog/datadog-agent

    Official

    Extract an Allium specification from an existing codebase. An agent skill from DataDog/datadog-agent.

    3.8k GitHub starsUsed in 1 repo~7k tokens
    Auto-check passed

Categories

Questions about E2E Audit

What does E2E Audit do?

Judge whether Agent behavior belongs in a new-e2e, integration, or unit test. E2E Audit is an agent skill from DataDog/datadog-agent, published by the product's own GitHub organization.

When should I use E2E Audit?

E2E Audit fits situations like: tasks that involve End-to-end testing; tasks that involve Unit testing.

How do I install E2E Audit in Claude Code?

Run `npx skills add DataDog/datadog-agent --skill e2e-audit -a claude-code`. Or copy the skill folder (.agents/skills/e2e-audit in DataDog/datadog-agent) into .claude/skills/e2e-audit in your project. Claude Code loads it when a task matches its description.

How do I install E2E Audit in Codex?

Run `npx skills add DataDog/datadog-agent --skill e2e-audit -a codex`. Or copy the skill folder (.agents/skills/e2e-audit in DataDog/datadog-agent) into .agents/skills/e2e-audit in your project. Codex loads it when a task matches its description.

Can I use E2E Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add DataDog/datadog-agent --skill e2e-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/e2e-audit, .gemini/skills/e2e-audit, .github/skills/e2e-audit and .opencode/skills/e2e-audit in your project.

What does E2E Audit need to run?

SKILL.md names no scripts, command-line tools or credentials: E2E Audit is instructions for the agent only. Its frontmatter pre-approves these tools: Read, Glob, Grep, Bash.

Does E2E Audit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is E2E Audit safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does E2E Audit use?

E2E Audit is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does E2E Audit use?

About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to E2E Audit?

Skills that share tags, products or a category with E2E Audit: TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars), Contracts (samchon/nestia, 2.2k stars), Atdd Mutate (swingerman/engineer, 154 stars) and E2E Scenario Testing (obra/dotfiles, 116 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains E2E Audit?

DataDog (a GitHub organization, an official publisher) maintains it in DataDog/datadog-agent, which has 3,757 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 8, 2026.

Source: DataDog/datadog-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.