Official agent skill

E2E Test

by grafana in grafana/agento11y

Optional credential-free Hermes integration checks using an explicit loopback model provider and local telemetry receivers.

OfficialApache-2.0Auto-check: notesTesting & QA

Install E2E Test

skills CLI
$ npx skills add grafana/agento11y --skill e2e-test -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install grafana/agento11y e2e-test --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/grafana/agento11y.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/hermes/.agents/skills/e2e-test .claude/skills/e2e-test && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
e2e-test
GitHub stars
127
Token cost
~1.7k tokens
SKILL.md length
763 words
Files
13 (incl. scripts)
Skills in repo
4
Repo updated
First seen
Licence
Apache-2.0

At a glance

Optional credential-free Hermes integration checks using an explicit loopback model provider and local telemetry receivers.

  • Tasks that involve End-to-end testing
  • SKILL.md covers Safety boundary, Wrapper limitations, Isolated setup and Start loopback receivers and…, plus 2 more sections
  • Runs Python and Shell scripts from its folder; needs OPENAI_API_KEY

What it does

E2E Test is an agent skill from grafana/agento11y, published by the product's own GitHub organization. Optional credential-free Hermes integration checks using an explicit loopback model provider and local telemetry receivers. Inspect the scripts before running them; their defaults are not safe evidence of a local-only run.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including scripts (for example `scripts/check-generations.py`, `scripts/check-install.py` and `scripts/mock-provider.py`).

It sits in Testing & QA, covering End-to-end testing. It works with OpenTelemetry. The repository describes itself as: Actually Useful Agent Observability. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve End-to-end testing

Example prompts

  • “/e2e-test”

Requirements

  • Python 3
  • A Bash shell
  • A credential in OPENAI_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 9ab60bc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 12 files in scripts/ (Python and Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

E2E Test loads about 1.7k tokens when it runs. Until then it costs about 58 tokens; SKILL.md has 763 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~58
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:28
    Hermes's `.env` overrides process exports: never reuse an old test home.
  • NoteMentions a .env fileSKILL.md:61
    ered hooks. Keep the fresh home free of `.env`

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from grafana/agento11y at commit 9ab60bc, republished under its Apache-2.0 licence (© grafana). 763 words, ~1,735 tokens.

Download SKILL.mdSave it as .claude/skills/e2e-test/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.
name
e2e-test
description
Optional credential-free Hermes integration checks using an explicit loopback model provider and local telemetry receivers. Inspect the scripts before running them; their defaults are not safe evidence of a local-only run.

Optional Hermes end-to-end checks

Unit tests stub the SDK and invent hook payloads. These optional checks inspect real Hermes hooks and exported OTel data. They are not part of normal CI.

Run from plugins/hermes. Read each script before execution. Do not run Cloud verification, install packages, start servers, or invoke Hermes without approval.

Safety boundary

A local telemetry sink does not make the model provider local. A mock provider does not make telemetry local either. Explicitly configure all three destinations:

DestinationCredential-free choice
Model providerLOOPBACK: openai-api, mock-model, http://127.0.0.1:8799/v1, dummy key
Generation exportDisabled: AGENTO11Y_PROTOCOL=none, AGENTO11Y_AUTH_MODE=none, no token or tenant
Traces and metricsLOOPBACK: http://127.0.0.1:8801, no auth headers

Use a new temporary HOME and HERMES_HOME, not the user's real config. Clear the inherited environment so provider keys, telemetry aliases, per-signal endpoints, proxy settings, and AGENTO11Y_ENV_FILE cannot redirect the test. Hermes's .env overrides process exports: never reuse an old test home. Use synthetic content only. Hook-probe logs contain payloads before redaction.

Wrapper limitations

  • setup.sh installs packages and defaults its config to Anthropic. Its warning requesting an Anthropic key is not applicable to the loopback recipe.
  • run-hermes.sh sink redirects OTel only. Its provider defaults to Anthropic, so sink alone is neither credential-free nor a local-only test.
  • run-hermes.sh full leaves capture mode unset. It tests metadata_only, despite its name. The wrapper does not pass through capture mode, redaction, or automatic-tag environment switches.
  • run-mock.sh chooses a loopback provider but reads generation and OTLP destinations from the environment or AGENTO11Y_ENV_FILE. It forces basic auth and uses broad pkill matching. Do not use it for this recipe.
  • verify-backend.sh uses Cloud queries. It is outside this credential-free flow.
  • otlp-sink.py decodes spans and metrics, but logs only attribute names. It does not validate generation ingestion or secret-redaction values.

Isolated setup

Prerequisites: uv, Python 3.11+, and permission to download dependencies. The setup step can access package registries; the model and telemetry steps below use loopback. Run the steps in the same shell.

sh
S="$PWD/.agents/skills/e2e-test/scripts"
E2E_DIR=$(mktemp -d)
mkdir -p "$E2E_DIR/user"
env -i PATH="$PATH" HOME="$E2E_DIR/user" E2E_DIR="$E2E_DIR" PY_VERSION=3.13 "$S/setup.sh" 0.19.0

setup.sh installs plugins/hermes. Confirm it reports the agento11y entry point and registered hooks. Keep the fresh home free of .env files. No provider key is needed for the next step.

Start loopback receivers and run Hermes directly

Bypass the wrappers so privacy switches reach the Hermes process. The mock and OTLP server both bind to 127.0.0.1. Choose unused ports; if either server exits on startup, stop rather than connecting to an unknown listener.

sh
env -i PATH="$PATH" HOME="$E2E_DIR/user" E2E_DIR="$E2E_DIR" MOCK_SCRIPT=tool,ok "$E2E_DIR/.venv/bin/python" "$S/mock-provider.py" 8799 >"$E2E_DIR/mock.out" 2>&1 &
mock_pid=$!
env -i PATH="$PATH" HOME="$E2E_DIR/user" E2E_DIR="$E2E_DIR" "$E2E_DIR/.venv/bin/python" "$S/otlp-sink.py" 8801 >"$E2E_DIR/sink.out" 2>&1 &
sink_pid=$!
trap 'kill "$mock_pid" "$sink_pid" 2>/dev/null || true' EXIT
sleep 1
kill -0 "$mock_pid" "$sink_pid" || exit 1

env -i PATH="$PATH" HOME="$E2E_DIR/user" TERM=dumb E2E_DIR="$E2E_DIR" HERMES_HOME="$E2E_DIR/home" OPENAI_API_KEY=mock-key OPENAI_BASE_URL=http://127.0.0.1:8799/v1 AGENTO11Y_PROTOCOL=none AGENTO11Y_AUTH_MODE=none AGENTO11Y_CONTENT_CAPTURE_MODE=metadata_only AGENTO11Y_AUTO_CODING_AGENT_TAGS=false OTEL_EXPORTER_OTLP_ENDPOINT=http://127.0.0.1:8801 OTEL_EXPORTER_OTLP_INSECURE=true "$E2E_DIR/.venv/bin/hermes" -m mock-model --provider openai-api -z 'List the available skills, then reply OK.'

This checks only the OTel channel. The disabled generation channel is deliberate, not evidence that generation export works. To test neither channel, omit the OTLP variables and set AGENTO11Y_HERMES_OTEL_AUTO=false.

To test generation export, first provide a loopback receiver that implements /api/v1/generations:export and validates the SDK request. Explicitly set its loopback endpoint and HTTP protocol. The plugin's channel activation also needs a supported non-none auth mode; use dummy local credentials only. Keep OTel pointed to the local sink or disable it. The OTLP sink does not implement this ingest protocol; do not treat its generic HTTP 200 response as validation.

Show full SKILL.md (268 more words)Show less

What to inspect

Read mock.log, hooks.jsonl, and otlp-sink.log under the temporary directory. The mock should show a tool request followed by completion. The sink should show generation/tool spans and metrics. Missing spans can indicate a flush or hook problem; no Cloud sampling is involved here.

Check these cases with explicit settings on the direct Hermes invocation:

  • Capture mode unset, default, and invalid: all must resolve to metadata_only.
  • Explicit full: exercise content and shared redaction with synthetic secrets. Use a receiver that inspects values; this sink's attribute-name log is insufficient.
  • AGENTO11Y_REDACT_INPUT_MESSAGES=false: only prompt redaction turns off.
  • Automatic tags off: no automatic user/repo/branch and no cwd. Enable user,repo, then branch, and check metric labels and explicit-tag precedence.
  • Sampling rate zero: no generation or tool telemetry.
  • Retry scripts 429,429,ok, empty, 401, and scratchpad: restart the mock with the chosen MOCK_SCRIPT and inspect actual hooks, not assumed attempt counts.
  • Tool spans parent to the requesting generation, even though that parent ended.
  • Request clipping and cache reuse: vary HERMES_PLUGIN_PAYLOAD_MAX_CHARS. Reused sampling parameters must come from the same model.

One-shot mode disables logging and bypasses atexit, so missing plugin logs are expected. Some early-return paths can omit finalization and lose open records. Use interactive Hermes when diagnosing logging or exit hooks.

Repeat on the supported floor and proposed Hermes upgrades. The PyPI release's hook call sites are the contract; upstream HEAD can contain unreleased kwargs.

Cleanup

Stop only the PIDs started above. Review the synthetic logs, then remove only the temporary directory printed by printf '%s\n' "$E2E_DIR". Do not use broad pkill or delete a reused path. No Cloud data should have been written.

© grafana, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 12 other files (scripts) in plugins/hermes/.agents/skills/e2e-test of grafana/agento11y.

  • SKILL.md
  • scripts/check-generations.py
  • scripts/check-install.py
  • scripts/mock-provider.py
  • scripts/otlp-sink.py
  • scripts/probe-plugin/__init__.py
  • scripts/probe-plugin/plugin.yaml
  • scripts/run-hermes.sh
  • scripts/run-mock.sh
  • scripts/setup.sh
  • scripts/show-hooks.py
  • scripts/show-spans.py
  • scripts/verify-backend.sh

Open the folder on GitHubat commit 9ab60bc

Compare with similar skills

E2E Test next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

E2E Test compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
E2E Test this skillgrafana/agento11y127—~1.7kAutomated safety check: NotesApache-2.0
Error Handling And E2Ehome-operations/kopiur113—~2.7kAutomated safety check: PassAGPL-3.0
Quality Codevvedantb/eva1011 repos~797Automated safety check: PassMIT
Web Application Testinganthropics/skills180k51 repos~966Automated safety check: PassApache-2.0
OpenHarness End-to-End EvalsHKUDS/OpenHarness16k1 repos~2.1kAutomated safety check: NotesMIT
playwright-cli Browser Automationgithub/gh-aw5.4k24 repos~2.8kAutomated safety check: PassMIT

Similar skills

  • Error Handling And E2E

    home-operations/kopiur

    How Kopiur does strongly-typed, actionable error handling and end-to-end testing.

    113 GitHub stars~2.7k tokensUpdated today
    Testing & QAAuto-check passed
  • Quality Code

    vvedantb/eva

    A skill your agent uses when writing or reviewing TypeScript/full-stack code.

    101 GitHub starsUsed in 1 repo~797 tokens
    DevOps & CloudAuto-check passed
  • Web Application Testing

    anthropics/skills

    Official

    Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

    180k GitHub starsUsed in 51 repos~966 tokens
    Testing & QAAuto-check passed
  • Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.

    16k GitHub starsUsed in 1 repo~2.1k tokens
    Testing & QAAuto-check: notes
  • Official

    Drives a real browser from the command line with playwright-cli to open pages, interact, mock requests, save state and work with Playwright tests.

    5.4k GitHub starsUsed in 24 repos~2.8k tokens
    Testing & QAAuto-check passed
  • Writes a Playwright end-to-end test from a prompt, runs it against a live Appsmith deployment and retries with fixes up to three times until it passes.

    41k GitHub stars~2.9k tokensUpdated today
    Testing & QAAuto-check: notes

More from grafana/agento11y

  • Setup Local Guards

    grafana/agento11y

    Official

    Help choose, configure, and test local agento11y guard packs for coding-agent tool calls.

    127 GitHub stars~2.6k tokensUpdated today
    Auto-check: notes
  • Agento11y Experiments

    grafana/agento11y

    Official

    Run any Python LLM agent as an Agent Observability experiment using the public agento11y.experiments package: define a test suite, run an existing agent through typed trials, bind or record…

    127 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Agento11y Eval Starter

    grafana/agento11y

    Official

    Use early in an AI-agent project — before ship, before real traffic — to decide which evaluations to set up and to scaffold a starter experiment.

    127 GitHub stars~6.4k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about E2E Test

What does E2E Test do?

Optional credential-free Hermes integration checks using an explicit loopback model provider and local telemetry receivers. E2E Test is an agent skill from grafana/agento11y, published by the product's own GitHub organization. Optional credential-free Hermes integration checks using an explicit loopback model provider and local telemetry receivers.

When should I use E2E Test?

E2E Test fits situations like: tasks that involve End-to-end testing.

How do I install E2E Test in Claude Code?

Run `npx skills add grafana/agento11y --skill e2e-test -a claude-code`. Or copy the skill folder (plugins/hermes/.agents/skills/e2e-test in grafana/agento11y) into .claude/skills/e2e-test in your project. Claude Code loads it when a task matches its description.

How do I install E2E Test in Codex?

Run `npx skills add grafana/agento11y --skill e2e-test -a codex`. Or copy the skill folder (plugins/hermes/.agents/skills/e2e-test in grafana/agento11y) into .agents/skills/e2e-test in your project. Codex loads it when a task matches its description.

Can I use E2E Test in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add grafana/agento11y --skill e2e-test -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/e2e-test, .gemini/skills/e2e-test, .github/skills/e2e-test and .opencode/skills/e2e-test in your project.

What does E2E Test need to run?

Going by SKILL.md and its folder, E2E Test needs Python and a shell for the scripts in its folder and credentials named OPENAI_API_KEY. Our summary lists: Python 3; A Bash shell; A credential in OPENAI_API_KEY.

Does E2E Test access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is E2E Test safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does E2E Test use?

E2E Test is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does E2E Test use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to E2E Test?

Skills that share tags, products or a category with E2E Test: Error Handling And E2E (home-operations/kopiur, 113 stars), Quality Code (vvedantb/eva, 101 stars), Web Application Testing (anthropics/skills, 180k stars) and OpenHarness End-to-End Evals (HKUDS/OpenHarness, 16k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains E2E Test?

grafana (a GitHub organization, an official publisher) maintains it in grafana/agento11y, which has 127 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 8, 2026.

Source: grafana/agento11y on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.