Error Handling And E2E
home-operations/kopiur
How Kopiur does strongly-typed, actionable error handling and end-to-end testing.
Optional credential-free Hermes integration checks using an explicit loopback model provider and local telemetry receivers.
$ npx skills add grafana/agento11y --skill e2e-test -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install grafana/agento11y e2e-test --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/grafana/agento11y.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/hermes/.agents/skills/e2e-test .claude/skills/e2e-test && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "e2e-test" agent skill from https://github.com/grafana/agento11y/tree/main/plugins/hermes/.agents/skills/e2e-test into .claude/skills/e2e-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "e2e-test", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/grafana/agento11y/tree/main/plugins/hermes/.agents/skills/e2e-testType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add grafana/agento11y --skill e2e-test -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install grafana/agento11y e2e-test --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/grafana/agento11y.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/hermes/.agents/skills/e2e-test .agents/skills/e2e-test && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "e2e-test" agent skill from https://github.com/grafana/agento11y/tree/main/plugins/hermes/.agents/skills/e2e-test into .agents/skills/e2e-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "e2e-test", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add grafana/agento11y --skill e2e-test -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install grafana/agento11y e2e-test --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/grafana/agento11y.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/hermes/.agents/skills/e2e-test .cursor/skills/e2e-test && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "e2e-test" agent skill from https://github.com/grafana/agento11y/tree/main/plugins/hermes/.agents/skills/e2e-test into .cursor/skills/e2e-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "e2e-test", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/grafana/agento11y.git --path plugins/hermes/.agents/skills/e2e-test--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add grafana/agento11y --skill e2e-test -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install grafana/agento11y e2e-test --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/grafana/agento11y.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/hermes/.agents/skills/e2e-test .gemini/skills/e2e-test && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "e2e-test" agent skill from https://github.com/grafana/agento11y/tree/main/plugins/hermes/.agents/skills/e2e-test into .gemini/skills/e2e-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "e2e-test", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install grafana/agento11y e2e-testInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add grafana/agento11y --skill e2e-test -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/grafana/agento11y.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/hermes/.agents/skills/e2e-test .github/skills/e2e-test && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "e2e-test" agent skill from https://github.com/grafana/agento11y/tree/main/plugins/hermes/.agents/skills/e2e-test into .github/skills/e2e-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "e2e-test", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add grafana/agento11y --skill e2e-test -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install grafana/agento11y e2e-test --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/grafana/agento11y.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/hermes/.agents/skills/e2e-test .opencode/skills/e2e-test && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "e2e-test" agent skill from https://github.com/grafana/agento11y/tree/main/plugins/hermes/.agents/skills/e2e-test into .opencode/skills/e2e-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "e2e-test", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
e2e-testOptional credential-free Hermes integration checks using an explicit loopback model provider and local telemetry receivers.
E2E Test is an agent skill from grafana/agento11y, published by the product's own GitHub organization. Optional credential-free Hermes integration checks using an explicit loopback model provider and local telemetry receivers. Inspect the scripts before running them; their defaults are not safe evidence of a local-only run.
Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including scripts (for example `scripts/check-generations.py`, `scripts/check-install.py` and `scripts/mock-provider.py`).
It sits in Testing & QA, covering End-to-end testing. It works with OpenTelemetry. The repository describes itself as: Actually Useful Agent Observability. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 9ab60bc. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 12 files in scripts/ (Python and Shell), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPENAI_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
E2E Test loads about 1.7k tokens when it runs. Until then it costs about 58 tokens; SKILL.md has 763 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
Hermes's `.env` overrides process exports: never reuse an old test home.ered hooks. Keep the fresh home free of `.env`Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from grafana/agento11y at commit 9ab60bc, republished under its Apache-2.0 licence (© grafana). 763 words, ~1,735 tokens.
.claude/skills/e2e-test/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.Unit tests stub the SDK and invent hook payloads. These optional checks inspect real Hermes hooks and exported OTel data. They are not part of normal CI.
Run from plugins/hermes. Read each script before execution. Do not run Cloud
verification, install packages, start servers, or invoke Hermes without approval.
A local telemetry sink does not make the model provider local. A mock provider does not make telemetry local either. Explicitly configure all three destinations:
| Destination | Credential-free choice |
|---|---|
| Model provider | LOOPBACK: openai-api, mock-model, http://127.0.0.1:8799/v1, dummy key |
| Generation export | Disabled: AGENTO11Y_PROTOCOL=none, AGENTO11Y_AUTH_MODE=none, no token or tenant |
| Traces and metrics | LOOPBACK: http://127.0.0.1:8801, no auth headers |
Use a new temporary HOME and HERMES_HOME, not the user's real config. Clear the
inherited environment so provider keys, telemetry aliases, per-signal endpoints,
proxy settings, and AGENTO11Y_ENV_FILE cannot redirect the test.
Hermes's .env overrides process exports: never reuse an old test home.
Use synthetic content only. Hook-probe logs contain payloads before redaction.
setup.sh installs packages and defaults its config to Anthropic. Its warning
requesting an Anthropic key is not applicable to the loopback recipe.run-hermes.sh sink redirects OTel only. Its provider defaults to Anthropic,
so sink alone is neither credential-free nor a local-only test.run-hermes.sh full leaves capture mode unset. It tests
metadata_only, despite its name. The wrapper does not pass through capture
mode, redaction, or automatic-tag environment switches.run-mock.sh chooses a loopback provider but reads generation and OTLP
destinations from the environment or AGENTO11Y_ENV_FILE. It forces basic
auth and uses broad pkill matching. Do not use it for this recipe.verify-backend.sh uses Cloud queries. It is outside this credential-free flow.otlp-sink.py decodes spans and metrics, but logs only attribute names.
It does not validate generation ingestion or secret-redaction values.Prerequisites: uv, Python 3.11+, and permission to download dependencies.
The setup step can access package registries; the model and telemetry steps
below use loopback. Run the steps in the same shell.
S="$PWD/.agents/skills/e2e-test/scripts"
E2E_DIR=$(mktemp -d)
mkdir -p "$E2E_DIR/user"
env -i PATH="$PATH" HOME="$E2E_DIR/user" E2E_DIR="$E2E_DIR" PY_VERSION=3.13 "$S/setup.sh" 0.19.0setup.sh installs plugins/hermes. Confirm it reports the
agento11y entry point and registered hooks. Keep the fresh home free of .env
files. No provider key is needed for the next step.
Bypass the wrappers so privacy switches reach the Hermes process.
The mock and OTLP server both bind to 127.0.0.1. Choose unused ports; if either
server exits on startup, stop rather than connecting to an unknown listener.
env -i PATH="$PATH" HOME="$E2E_DIR/user" E2E_DIR="$E2E_DIR" MOCK_SCRIPT=tool,ok "$E2E_DIR/.venv/bin/python" "$S/mock-provider.py" 8799 >"$E2E_DIR/mock.out" 2>&1 &
mock_pid=$!
env -i PATH="$PATH" HOME="$E2E_DIR/user" E2E_DIR="$E2E_DIR" "$E2E_DIR/.venv/bin/python" "$S/otlp-sink.py" 8801 >"$E2E_DIR/sink.out" 2>&1 &
sink_pid=$!
trap 'kill "$mock_pid" "$sink_pid" 2>/dev/null || true' EXIT
sleep 1
kill -0 "$mock_pid" "$sink_pid" || exit 1
env -i PATH="$PATH" HOME="$E2E_DIR/user" TERM=dumb E2E_DIR="$E2E_DIR" HERMES_HOME="$E2E_DIR/home" OPENAI_API_KEY=mock-key OPENAI_BASE_URL=http://127.0.0.1:8799/v1 AGENTO11Y_PROTOCOL=none AGENTO11Y_AUTH_MODE=none AGENTO11Y_CONTENT_CAPTURE_MODE=metadata_only AGENTO11Y_AUTO_CODING_AGENT_TAGS=false OTEL_EXPORTER_OTLP_ENDPOINT=http://127.0.0.1:8801 OTEL_EXPORTER_OTLP_INSECURE=true "$E2E_DIR/.venv/bin/hermes" -m mock-model --provider openai-api -z 'List the available skills, then reply OK.'This checks only the OTel channel. The disabled generation channel is deliberate,
not evidence that generation export works. To test neither channel, omit the
OTLP variables and set AGENTO11Y_HERMES_OTEL_AUTO=false.
To test generation export, first provide a loopback receiver that implements
/api/v1/generations:export and validates the SDK request. Explicitly set its
loopback endpoint and HTTP protocol. The plugin's channel activation also needs
a supported non-none auth mode; use dummy local credentials only. Keep OTel
pointed to the local sink or disable it. The OTLP sink does not implement
this ingest protocol; do not treat its generic HTTP 200 response as validation.
Read mock.log, hooks.jsonl, and otlp-sink.log under the temporary directory.
The mock should show a tool request followed by completion. The sink should show
generation/tool spans and metrics. Missing spans can indicate a flush or hook
problem; no Cloud sampling is involved here.
Check these cases with explicit settings on the direct Hermes invocation:
default, and invalid: all must resolve to metadata_only.full: exercise content and shared redaction with synthetic secrets.
Use a receiver that inspects values; this sink's attribute-name log is insufficient.AGENTO11Y_REDACT_INPUT_MESSAGES=false: only prompt redaction turns off.cwd.
Enable user,repo, then branch, and check metric labels and explicit-tag precedence.429,429,ok, empty, 401, and scratchpad: restart the mock
with the chosen MOCK_SCRIPT and inspect actual hooks, not assumed attempt counts.HERMES_PLUGIN_PAYLOAD_MAX_CHARS.
Reused sampling parameters must come from the same model.One-shot mode disables logging and bypasses atexit, so missing plugin logs are expected. Some early-return paths can omit finalization and lose open records. Use interactive Hermes when diagnosing logging or exit hooks.
Repeat on the supported floor and proposed Hermes upgrades. The PyPI release's hook call sites are the contract; upstream HEAD can contain unreleased kwargs.
Stop only the PIDs started above. Review the synthetic logs, then remove only the
temporary directory printed by printf '%s\n' "$E2E_DIR". Do not use broad pkill
or delete a reused path. No Cloud data should have been written.
© grafana, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 12 other files (scripts) in plugins/hermes/.agents/skills/e2e-test of grafana/agento11y.
Open the folder on GitHubat commit 9ab60bc
E2E Test next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| E2E Test this skillgrafana/agento11y | 127 | — | ~1.7k | Automated safety check: Notes | Apache-2.0 | |
| Error Handling And E2Ehome-operations/kopiur | 113 | — | ~2.7k | Automated safety check: Pass | AGPL-3.0 | |
| Quality Codevvedantb/eva | 101 | 1 repos | ~797 | Automated safety check: Pass | MIT | |
| Web Application Testinganthropics/skills | 180k | 51 repos | ~966 | Automated safety check: Pass | Apache-2.0 | |
| OpenHarness End-to-End EvalsHKUDS/OpenHarness | 16k | 1 repos | ~2.1k | Automated safety check: Notes | MIT | |
| playwright-cli Browser Automationgithub/gh-aw | 5.4k | 24 repos | ~2.8k | Automated safety check: Pass | MIT |
home-operations/kopiur
How Kopiur does strongly-typed, actionable error handling and end-to-end testing.
vvedantb/eva
A skill your agent uses when writing or reviewing TypeScript/full-stack code.
anthropics/skills
Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.
HKUDS/OpenHarness
Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.
github/gh-aw
Drives a real browser from the command line with playwright-cli to open pages, interact, mock requests, save state and work with Playwright tests.
appsmithorg/appsmith
Writes a Playwright end-to-end test from a prompt, runs it against a live Appsmith deployment and retries with fixes up to three times until it passes.
grafana/agento11y
Help choose, configure, and test local agento11y guard packs for coding-agent tool calls.
grafana/agento11y
Run any Python LLM agent as an Agent Observability experiment using the public agento11y.experiments package: define a test suite, run an existing agent through typed trials, bind or record…
grafana/agento11y
Use early in an AI-agent project — before ship, before real traffic — to decide which evaluations to set up and to scaffold a starter experiment.
Works with
Categories
Optional credential-free Hermes integration checks using an explicit loopback model provider and local telemetry receivers. E2E Test is an agent skill from grafana/agento11y, published by the product's own GitHub organization. Optional credential-free Hermes integration checks using an explicit loopback model provider and local telemetry receivers.
E2E Test fits situations like: tasks that involve End-to-end testing.
Run `npx skills add grafana/agento11y --skill e2e-test -a claude-code`. Or copy the skill folder (plugins/hermes/.agents/skills/e2e-test in grafana/agento11y) into .claude/skills/e2e-test in your project. Claude Code loads it when a task matches its description.
Run `npx skills add grafana/agento11y --skill e2e-test -a codex`. Or copy the skill folder (plugins/hermes/.agents/skills/e2e-test in grafana/agento11y) into .agents/skills/e2e-test in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add grafana/agento11y --skill e2e-test -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/e2e-test, .gemini/skills/e2e-test, .github/skills/e2e-test and .opencode/skills/e2e-test in your project.
Going by SKILL.md and its folder, E2E Test needs Python and a shell for the scripts in its folder and credentials named OPENAI_API_KEY. Our summary lists: Python 3; A Bash shell; A credential in OPENAI_API_KEY.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
E2E Test is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with E2E Test: Error Handling And E2E (home-operations/kopiur, 113 stars), Quality Code (vvedantb/eva, 101 stars), Web Application Testing (anthropics/skills, 180k stars) and OpenHarness End-to-End Evals (HKUDS/OpenHarness, 16k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
grafana (a GitHub organization, an official publisher) maintains it in grafana/agento11y, which has 127 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 8, 2026.
Source: grafana/agento11y on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.