Opentelemetry
grafana/skills
Instrument any app with OpenTelemetry and ship metrics / logs / traces to Grafana Cloud or self-hosted Mimir / Loki / Tempo / Pyroscope.
Run any Python LLM agent as an Agent Observability experiment using the public agento11y.experiments package: define a test suite, run an existing agent through typed trials, bind or record…
$ npx skills add grafana/agento11y --skill agento11y-experiments -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install grafana/agento11y agento11y-experiments --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/grafana/agento11y.git skills-src && mkdir -p .claude/skills && cp -r skills-src/python/skills/agento11y-experiments .claude/skills/agento11y-experiments && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agento11y-experiments" agent skill from https://github.com/grafana/agento11y/tree/main/python/skills/agento11y-experiments into .claude/skills/agento11y-experiments/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agento11y-experiments", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/grafana/agento11y/tree/main/python/skills/agento11y-experimentsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add grafana/agento11y --skill agento11y-experiments -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install grafana/agento11y agento11y-experiments --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/grafana/agento11y.git skills-src && mkdir -p .agents/skills && cp -r skills-src/python/skills/agento11y-experiments .agents/skills/agento11y-experiments && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agento11y-experiments" agent skill from https://github.com/grafana/agento11y/tree/main/python/skills/agento11y-experiments into .agents/skills/agento11y-experiments/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agento11y-experiments", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add grafana/agento11y --skill agento11y-experiments -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install grafana/agento11y agento11y-experiments --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/grafana/agento11y.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/python/skills/agento11y-experiments .cursor/skills/agento11y-experiments && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agento11y-experiments" agent skill from https://github.com/grafana/agento11y/tree/main/python/skills/agento11y-experiments into .cursor/skills/agento11y-experiments/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agento11y-experiments", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/grafana/agento11y.git --path python/skills/agento11y-experiments--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add grafana/agento11y --skill agento11y-experiments -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install grafana/agento11y agento11y-experiments --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/grafana/agento11y.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/python/skills/agento11y-experiments .gemini/skills/agento11y-experiments && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agento11y-experiments" agent skill from https://github.com/grafana/agento11y/tree/main/python/skills/agento11y-experiments into .gemini/skills/agento11y-experiments/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agento11y-experiments", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install grafana/agento11y agento11y-experimentsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add grafana/agento11y --skill agento11y-experiments -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/grafana/agento11y.git skills-src && mkdir -p .github/skills && cp -r skills-src/python/skills/agento11y-experiments .github/skills/agento11y-experiments && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agento11y-experiments" agent skill from https://github.com/grafana/agento11y/tree/main/python/skills/agento11y-experiments into .github/skills/agento11y-experiments/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agento11y-experiments", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add grafana/agento11y --skill agento11y-experiments -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install grafana/agento11y agento11y-experiments --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/grafana/agento11y.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/python/skills/agento11y-experiments .opencode/skills/agento11y-experiments && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agento11y-experiments" agent skill from https://github.com/grafana/agento11y/tree/main/python/skills/agento11y-experiments into .opencode/skills/agento11y-experiments/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agento11y-experiments", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agento11y-experimentsRun any Python LLM agent as an Agent Observability experiment using the public agento11y.experiments package: define a test suite, run an existing agent through typed trials, bind or record…
Agento11y Experiments is an agent skill from grafana/agento11y, published by the product's own GitHub organization. Run any Python LLM agent as an Agent Observability experiment using the public agento11y.experiments package: define a test suite, run an existing agent through typed trials, bind or record generation I/O, grade outputs, and publish scores, including from stored Grafana test suites.
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud, covering Test generation, Observability and Monitoring and alerting. It works with Grafana and Python. The repository describes itself as: Actually Useful Agent Observability. The licence is Apache-2.0.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 447d692. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
AGENTO11Y_SERVICE_ACCOUNT_TOKENAGENTO11Y_AUTH_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agento11y Experiments loads about 2k tokens when it runs. Until then it costs about 76 tokens; SKILL.md has 494 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from grafana/agento11y at commit 447d692, republished under its Apache-2.0 licence (© grafana). 494 words, ~2,027 tokens.
.claude/skills/agento11y-experiments/SKILL.md (or your agent's skills folder).Use this skill when adding framework-free offline evaluation to a Python project.
The public SDK surface is agento11y.experiments; do not use removed v0 runner
APIs.
This is the reference for the run-side API. If you don't yet know which evaluators
you need or have no test cases, start with the agento11y-eval-starter skill — it reads
your agent, recommends evaluators, writes a starter suite, and generates a minimal
runner; come here for the deeper patterns (binding existing generations, auditable
LLM judges, cross-process verifiers, pass@k/pass^k).
The normal setup cost for an already instrumented agent should be small:
experiments from agento11y.TestSuite with TestCases.with exp.trial(case) as trial:.trial.record_io(...) when the harness owns the call.final score remains the legacy fallback.pip install "agento11y>=0.11.0"Required environment:
export AGENTO11Y_ENDPOINT=https://agento11y-prod-<region>.grafana.net
export AGENTO11Y_AUTH_TOKEN=<grafana-cloud-ingestion-api-key>
# Optional when the endpoint requires tenant-scoped basic auth.
export AGENTO11Y_AUTH_TENANT_ID=<stack-id>
# Optional UI host for deep links when it differs from AGENTO11Y_ENDPOINT.
export AGENTO11Y_GRAFANA_URL=https://<your-stack>.grafana.netLocal-suite experiment ingest uses only the Cloud ingestion API key. Stored
suite push/pull additionally uses AGENTO11Y_CONTROL_ENDPOINT and a Grafana
service-account token in AGENTO11Y_SERVICE_ACCOUNT_TOKEN.
Experimental OTel eval spans/events are disabled by default. Opt in only when asked:
with experiments.experiment("nightly", use_experimental_otel=True) as exp:
...from agento11y import experiments
suite = experiments.TestSuite(
suite_id="smoke",
name="Smoke",
version="2026-06-29",
test_cases=[
experiments.TestCase(test_case_id="capital-fr", input="Capital of France?", expected="Paris"),
],
)
verifier = experiments.Evaluator(evaluator_id="exact_match", version="2026-06-29", kind="deterministic")
with experiments.experiment(
"PR experiment",
experiment_id=f"pr-{git_sha}",
suite=suite,
planned_trial_count=len(suite.test_cases),
candidate={"git_sha": git_sha, "model_name": "gpt-4o-mini"},
tags=["ci"],
) as exp:
for case in suite.test_cases:
with exp.trial(case) as trial:
answer = call_your_agent(case.input)
# If normal instrumentation already created a conversation/generation,
# bind those ids instead of recording duplicate I/O.
# trial.bind_conversation(conversation_id)
# trial.bind_generation(generation_id, conversation_id=conversation_id)
trial.record_io(
input=case.input,
output=answer,
model_provider="openai",
model_name="gpt-4o-mini",
)
passed = str(case.expected).lower() in answer.lower()
trial.final_score(
1.0 if passed else 0.0,
passed=passed,
explanation=f"expected {case.expected!r}, got {answer!r}",
evaluator=verifier,
)
print(exp.url)The context manager upserts the run on enter, creates a typed trial per case,
exports buffered scores when each trial exits, and finalizes the run as
completed or failed.
Set planned_trial_count to the runner's exact post-filter, post-attempt trial
count. Do not derive it from the stored suite when the runner filters cases or
runs multiple attempts. Normal context-manager finalization omits
score_count, allowing Agent Observability to use its authoritative stored score count.
Each (test_case_id, attempt) pair must be unique within a run; increment
attempt for retries.
Stored suites use a Grafana service-account token in addition to the ingestion credential:
export AGENTO11Y_CONTROL_ENDPOINT=https://<stack>.grafana.net/a/grafana-agento11y-app
export AGENTO11Y_SERVICE_ACCOUNT_TOKEN=<grafana-service-account-token>Pull and run the latest published version while preserving exact suite provenance:
from agento11y import experiments
with experiments.experiment_from_suite(
"dashboard-regression",
version="latest_published",
experiment_id=f"pr-{git_sha}",
) as exp:
for case in exp.suite.cases:
with exp.trial(case) as trial:
answer = call_your_agent(case.input)
trial.final_score(answer == case.expected)Use TestSuite.from_yaml(...) and TestSuitesClient.push_suite(...) to manage
portable source-controlled suites. Pushes are additive by default; pass
prune=True to delete remote-only draft cases and publish=True to publish the
resulting version.
Use ReportRole.PRIMARY_VERDICT on the one score intended to drive the headline
and pass rate. Use ReportRole.DIAGNOSTIC on supporting scores that should remain
available for analysis without determining the verdict. An unannotated
trial.final_score(...) remains the backward-compatible fallback.
trial.score(
"answer_relevancy",
0.91,
passed=True,
report_role=experiments.ReportRole.PRIMARY_VERDICT,
)
trial.check_score(
"json_valid",
passed=is_valid_json(answer),
report_role=experiments.ReportRole.DIAGNOSTIC,
)Locally configured judges do not require a platform evaluator. LLMJudge
executes an injected model callable, and trial.evaluate_output publishes the grader
transcript and links it to the score automatically.
judge = experiments.LLMJudge(
evaluator_id="judge.correctness",
invoke=judge_model.invoke,
model_provider="anthropic",
model_name="claude-sonnet-4-5",
prompt_template="Input: {input}\nExpected: {expected}\nOutput: {output}\nReturn a JSON score.",
)
trial.evaluate_output(judge, input=case.input, expected=case.expected, output=answer)
regex = experiments.RegexJudge(evaluator_id="regex.answer", pattern=r"Paris", score_key="contains_answer")
trial.evaluate_output(regex, input=case.input, output=answer)Use TrialRef when a verifier runs in a separate process or container.
ref = trial.ref
env = ref.to_env()In the verifier:
from agento11y import experiments
client = experiments.Client(
endpoint=os.environ["AGENTO11Y_ENDPOINT"],
tenant_id=os.environ.get("AGENTO11Y_AUTH_TENANT_ID", ""),
ingest_token=os.environ["AGENTO11Y_AUTH_TOKEN"],
)
ref = experiments.TrialRef.from_env()
if ref is None:
raise RuntimeError("missing experiment trial environment")
trial = experiments.Trial.from_ref(client, ref)
trial.final_score(0.9, passed=True)
trial.close()For all v1 environment and worker-lifecycle changes, see the Experiments v2 migration guide.
experiment_id for CI retries.record_io(...) when the experiment harness is the only
instrumentation around the agent call./a/grafana-agento11y-app/experiments/runs/{experiment_id}.with exp.trial(...) block when one
malformed judge response should fail only that trial and the run should continue.© grafana, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in python/skills/agento11y-experiments of grafana/agento11y.
Open the folder on GitHubat commit 447d692
Agento11y Experiments next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agento11y Experiments this skillgrafana/agento11y | 128 | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| Opentelemetrygrafana/skills | 282 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Axiom Dashboard Builderopenclaw/clawhub | 9.5k | — | ~4.9k | Automated safety check: Pass | MIT | |
| Happy Infra Metrics and Grafanaslopus/happy | 24k | — | ~2k | Automated safety check: Notes | MIT | |
| OpenTelemetry Pipeline Metrics Speccomet-ml/opik | 22k | — | ~3.2k | Automated safety check: Pass | Apache-2.0 | |
| Logfire Instrumentationbasicmachines-co/basic-memory | 4.1k | — | ~2.3k | Automated safety check: Pass | AGPL-3.0 |
grafana/skills
Instrument any app with OpenTelemetry and ship metrics / logs / traces to Grafana Cloud or self-hosted Mimir / Loki / Tempo / Pyroscope.
openclaw/clawhub
Designs and deploys Axiom dashboards through the API, choosing chart types and writing APL or metrics queries, with templates and migration notes for Splunk and Grafana.
slopus/happy
Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API.
comet-ml/opik
Specifies how to instrument an opik-backend pipeline with per-stage OpenTelemetry metrics for throughput, latency, errors and queue delay by workspace.
basicmachines-co/basic-memory
Adds Pydantic Logfire tracing, logging and metrics to Python, JavaScript or TypeScript and Rust projects, with the correct setup order and library extras.
letsrevel/revel-backend
A skill your agent uses when investigating production behaviour from logs — a 500/error in prod, a failing or stuck Celery task, tracing one request/traceid/user across services, or confirming a…
grafana/agento11y
Optional credential-free Hermes integration checks using an explicit loopback model provider and local telemetry receivers.
grafana/agento11y
Help choose, configure, and test local agento11y guard packs for coding-agent tool calls.
grafana/agento11y
Use early in an AI-agent project — before ship, before real traffic — to decide which evaluations to set up and to scaffold a starter experiment.
Categories
Run any Python LLM agent as an Agent Observability experiment using the public agento11y.experiments package: define a test suite, run an existing agent through typed trials, bind or record…. Agento11y Experiments is an agent skill from grafana/agento11y, published by the product's own GitHub organization.experiments package: define a test suite, run an existing agent through typed trials, bind or record generation I/O, grade outputs, and publish scores, including from stored Grafana test suites.
Agento11y Experiments fits situations like: tasks that involve Test generation; tasks that involve Observability; tasks that involve Monitoring and alerting.
Run `npx skills add grafana/agento11y --skill agento11y-experiments -a claude-code`. Or copy the skill folder (python/skills/agento11y-experiments in grafana/agento11y) into .claude/skills/agento11y-experiments in your project. Claude Code loads it when a task matches its description.
Run `npx skills add grafana/agento11y --skill agento11y-experiments -a codex`. Or copy the skill folder (python/skills/agento11y-experiments in grafana/agento11y) into .agents/skills/agento11y-experiments in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add grafana/agento11y --skill agento11y-experiments -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agento11y-experiments, .gemini/skills/agento11y-experiments, .github/skills/agento11y-experiments and .opencode/skills/agento11y-experiments in your project.
Going by SKILL.md and its folder, Agento11y Experiments needs the command-line tools its instructions call (pip) and credentials named AGENTO11Y_SERVICE_ACCOUNT_TOKEN and AGENTO11Y_AUTH_TOKEN. Our summary lists: Python 3; A credential in AGENTO11Y_AUTH_TOKEN; A credential in AGENTO11Y_SERVICE_ACCOUNT_TOKEN.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agento11y Experiments is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Agento11y Experiments: Opentelemetry (grafana/skills, 282 stars), Axiom Dashboard Builder (openclaw/clawhub, 9.5k stars), Happy Infra Metrics and Grafana (slopus/happy, 24k stars) and OpenTelemetry Pipeline Metrics Spec (comet-ml/opik, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
grafana (a GitHub organization, an official publisher) maintains it in grafana/agento11y, which has 128 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 9, 2026.
Source: grafana/agento11y on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.