Official agent skill

Agento11y Experiments

by grafana in grafana/agento11y

Run any Python LLM agent as an Agent Observability experiment using the public agento11y.experiments package: define a test suite, run an existing agent through typed trials, bind or record…

OfficialApache-2.0Auto-check passedDevOps & Cloud

Install Agento11y Experiments

skills CLI
$ npx skills add grafana/agento11y --skill agento11y-experiments -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install grafana/agento11y agento11y-experiments --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/grafana/agento11y.git skills-src && mkdir -p .claude/skills && cp -r skills-src/python/skills/agento11y-experiments .claude/skills/agento11y-experiments && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agento11y-experiments
GitHub stars
128
Token cost
~2k tokens
SKILL.md length
494 words
Files
1
Skills in repo
4
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run any Python LLM agent as an Agent Observability experiment using the public agento11y.experiments package: define a test suite, run an existing agent through typed trials, bind or record…

  • Works in 5 steps: Import experiments from agento11y. → Define a TestSuite with TestCases. → Wrap the existing agent call in with… → …
  • Tasks that involve Test generation
  • SKILL.md covers Setup, Recommended Pattern, Stored Suites and Scoring, plus 2 more sections
  • Calls pip; needs AGENTO11Y_SERVICE_ACCOUNT_TOKEN and AGENTO11Y_AUTH_TOKEN

What it does

Agento11y Experiments is an agent skill from grafana/agento11y, published by the product's own GitHub organization. Run any Python LLM agent as an Agent Observability experiment using the public agento11y.experiments package: define a test suite, run an existing agent through typed trials, bind or record generation I/O, grade outputs, and publish scores, including from stored Grafana test suites.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Test generation, Observability and Monitoring and alerting. It works with Grafana and Python. The repository describes itself as: Actually Useful Agent Observability. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Test generation
  • Tasks that involve Observability
  • Tasks that involve Monitoring and alerting

Example prompts

  • “/agento11y-experiments”

Requirements

  • Python 3
  • A credential in AGENTO11Y_AUTH_TOKEN
  • A credential in AGENTO11Y_SERVICE_ACCOUNT_TOKEN

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Import experiments from agento11y.
  2. Define a TestSuite with TestCases.
  3. Wrap the existing agent call in with exp.trial(case) as trial:.
  4. Bind the generation/conversation ids your normal instrumentation already
  5. Emit one primary-verdict score and any supporting diagnostic scores; an

What it can do on your machine

Read from SKILL.md and the folder at commit 447d692. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • AGENTO11Y_SERVICE_ACCOUNT_TOKEN
    • AGENTO11Y_AUTH_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agento11y Experiments loads about 2k tokens when it runs. Until then it costs about 76 tokens; SKILL.md has 494 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~76
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from grafana/agento11y at commit 447d692, republished under its Apache-2.0 licence (© grafana). 494 words, ~2,027 tokens.

Download SKILL.mdSave it as .claude/skills/agento11y-experiments/SKILL.md (or your agent's skills folder).
name
agento11y-experiments
description
Run any Python LLM agent as an Agent Observability experiment using the public agento11y.experiments package: define a test suite, run an existing agent through typed trials, bind or record generation I/O, grade outputs, and publish scores, including from stored Grafana test suites.

Agent Observability experiments

Use this skill when adding framework-free offline evaluation to a Python project. The public SDK surface is agento11y.experiments; do not use removed v0 runner APIs.

This is the reference for the run-side API. If you don't yet know which evaluators you need or have no test cases, start with the agento11y-eval-starter skill — it reads your agent, recommends evaluators, writes a starter suite, and generates a minimal runner; come here for the deeper patterns (binding existing generations, auditable LLM judges, cross-process verifiers, pass@k/pass^k).

The normal setup cost for an already instrumented agent should be small:

  1. Import experiments from agento11y.
  2. Define a TestSuite with TestCases.
  3. Wrap the existing agent call in with exp.trial(case) as trial:.
  4. Bind the generation/conversation ids your normal instrumentation already produced, or call trial.record_io(...) when the harness owns the call.
  5. Emit one primary-verdict score and any supporting diagnostic scores; an unannotated final score remains the legacy fallback.

Setup

bash
pip install "agento11y>=0.11.0"

Required environment:

bash
export AGENTO11Y_ENDPOINT=https://agento11y-prod-<region>.grafana.net
export AGENTO11Y_AUTH_TOKEN=<grafana-cloud-ingestion-api-key>

# Optional when the endpoint requires tenant-scoped basic auth.
export AGENTO11Y_AUTH_TENANT_ID=<stack-id>

# Optional UI host for deep links when it differs from AGENTO11Y_ENDPOINT.
export AGENTO11Y_GRAFANA_URL=https://<your-stack>.grafana.net

Local-suite experiment ingest uses only the Cloud ingestion API key. Stored suite push/pull additionally uses AGENTO11Y_CONTROL_ENDPOINT and a Grafana service-account token in AGENTO11Y_SERVICE_ACCOUNT_TOKEN.

Experimental OTel eval spans/events are disabled by default. Opt in only when asked:

python
with experiments.experiment("nightly", use_experimental_otel=True) as exp:
    ...
python
from agento11y import experiments

suite = experiments.TestSuite(
    suite_id="smoke",
    name="Smoke",
    version="2026-06-29",
    test_cases=[
        experiments.TestCase(test_case_id="capital-fr", input="Capital of France?", expected="Paris"),
    ],
)
verifier = experiments.Evaluator(evaluator_id="exact_match", version="2026-06-29", kind="deterministic")

with experiments.experiment(
    "PR experiment",
    experiment_id=f"pr-{git_sha}",
    suite=suite,
    planned_trial_count=len(suite.test_cases),
    candidate={"git_sha": git_sha, "model_name": "gpt-4o-mini"},
    tags=["ci"],
) as exp:
    for case in suite.test_cases:
        with exp.trial(case) as trial:
            answer = call_your_agent(case.input)

            # If normal instrumentation already created a conversation/generation,
            # bind those ids instead of recording duplicate I/O.
            # trial.bind_conversation(conversation_id)
            # trial.bind_generation(generation_id, conversation_id=conversation_id)
            trial.record_io(
                input=case.input,
                output=answer,
                model_provider="openai",
                model_name="gpt-4o-mini",
            )

            passed = str(case.expected).lower() in answer.lower()
            trial.final_score(
                1.0 if passed else 0.0,
                passed=passed,
                explanation=f"expected {case.expected!r}, got {answer!r}",
                evaluator=verifier,
            )

print(exp.url)

The context manager upserts the run on enter, creates a typed trial per case, exports buffered scores when each trial exits, and finalizes the run as completed or failed.

Set planned_trial_count to the runner's exact post-filter, post-attempt trial count. Do not derive it from the stored suite when the runner filters cases or runs multiple attempts. Normal context-manager finalization omits score_count, allowing Agent Observability to use its authoritative stored score count. Each (test_case_id, attempt) pair must be unique within a run; increment attempt for retries.

Show full SKILL.md (212 more words)Show less

Stored Suites

Stored suites use a Grafana service-account token in addition to the ingestion credential:

bash
export AGENTO11Y_CONTROL_ENDPOINT=https://<stack>.grafana.net/a/grafana-agento11y-app
export AGENTO11Y_SERVICE_ACCOUNT_TOKEN=<grafana-service-account-token>

Pull and run the latest published version while preserving exact suite provenance:

python
from agento11y import experiments

with experiments.experiment_from_suite(
    "dashboard-regression",
    version="latest_published",
    experiment_id=f"pr-{git_sha}",
) as exp:
    for case in exp.suite.cases:
        with exp.trial(case) as trial:
            answer = call_your_agent(case.input)
            trial.final_score(answer == case.expected)

Use TestSuite.from_yaml(...) and TestSuitesClient.push_suite(...) to manage portable source-controlled suites. Pushes are additive by default; pass prune=True to delete remote-only draft cases and publish=True to publish the resulting version.

Scoring

Use ReportRole.PRIMARY_VERDICT on the one score intended to drive the headline and pass rate. Use ReportRole.DIAGNOSTIC on supporting scores that should remain available for analysis without determining the verdict. An unannotated trial.final_score(...) remains the backward-compatible fallback.

python
trial.score(
    "answer_relevancy",
    0.91,
    passed=True,
    report_role=experiments.ReportRole.PRIMARY_VERDICT,
)
trial.check_score(
    "json_valid",
    passed=is_valid_json(answer),
    report_role=experiments.ReportRole.DIAGNOSTIC,
)

Locally configured judges do not require a platform evaluator. LLMJudge executes an injected model callable, and trial.evaluate_output publishes the grader transcript and links it to the score automatically.

python
judge = experiments.LLMJudge(
    evaluator_id="judge.correctness",
    invoke=judge_model.invoke,
    model_provider="anthropic",
    model_name="claude-sonnet-4-5",
    prompt_template="Input: {input}\nExpected: {expected}\nOutput: {output}\nReturn a JSON score.",
)
trial.evaluate_output(judge, input=case.input, expected=case.expected, output=answer)

regex = experiments.RegexJudge(evaluator_id="regex.answer", pattern=r"Paris", score_key="contains_answer")
trial.evaluate_output(regex, input=case.input, output=answer)

Cross-Process Evaluation

Use TrialRef when a verifier runs in a separate process or container.

python
ref = trial.ref
env = ref.to_env()

In the verifier:

python
from agento11y import experiments

client = experiments.Client(
    endpoint=os.environ["AGENTO11Y_ENDPOINT"],
    tenant_id=os.environ.get("AGENTO11Y_AUTH_TENANT_ID", ""),
    ingest_token=os.environ["AGENTO11Y_AUTH_TOKEN"],
)
ref = experiments.TrialRef.from_env()
if ref is None:
    raise RuntimeError("missing experiment trial environment")
trial = experiments.Trial.from_ref(client, ref)
trial.final_score(0.9, passed=True)
trial.close()

For all v1 environment and worker-lifecycle changes, see the Experiments v2 migration guide.

Gotchas

  • Use a stable experiment_id for CI retries.
  • Prefer binding existing conversation/generation ids when the agent is already instrumented; use record_io(...) when the experiment harness is the only instrumentation around the agent call.
  • The Grafana UI route is /a/grafana-agento11y-app/experiments/runs/{experiment_id}.
  • Catch evaluator failures outside each with exp.trial(...) block when one malformed judge response should fail only that trial and the run should continue.

© grafana, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in python/skills/agento11y-experiments of grafana/agento11y.

Open the folder on GitHubat commit 447d692

Compare with similar skills

Agento11y Experiments next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agento11y Experiments compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agento11y Experiments this skillgrafana/agento11y128—~2kAutomated safety check: PassApache-2.0
Opentelemetrygrafana/skills282—~1.7kAutomated safety check: PassApache-2.0
Axiom Dashboard Builderopenclaw/clawhub9.5k—~4.9kAutomated safety check: PassMIT
Happy Infra Metrics and Grafanaslopus/happy24k—~2kAutomated safety check: NotesMIT
OpenTelemetry Pipeline Metrics Speccomet-ml/opik22k—~3.2kAutomated safety check: PassApache-2.0
Logfire Instrumentationbasicmachines-co/basic-memory4.1k—~2.3kAutomated safety check: PassAGPL-3.0

Similar skills

  • Opentelemetry

    grafana/skills

    Official

    Instrument any app with OpenTelemetry and ship metrics / logs / traces to Grafana Cloud or self-hosted Mimir / Loki / Tempo / Pyroscope.

    282 GitHub stars~1.7k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Axiom Dashboard Builder

    openclaw/clawhub

    Designs and deploys Axiom dashboards through the API, choosing chart types and writing APL or metrics queries, with templates and migration notes for Splunk and Grafana.

    9.5k GitHub stars~4.9k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API.

    24k GitHub stars~2k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Specifies how to instrument an opik-backend pipeline with per-stage OpenTelemetry metrics for throughput, latency, errors and queue delay by workspace.

    22k GitHub stars~3.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Logfire Instrumentation

    basicmachines-co/basic-memory

    Adds Pydantic Logfire tracing, logging and metrics to Python, JavaScript or TypeScript and Rust projects, with the correct setup order and library extras.

    4.1k GitHub stars~2.3k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Loki Logs

    letsrevel/revel-backend

    A skill your agent uses when investigating production behaviour from logs — a 500/error in prod, a failing or stuck Celery task, tracing one request/traceid/user across services, or confirming a…

    110 GitHub stars~918 tokensUpdated 3 days ago
    DevOps & CloudAuto-check: notes

More from grafana/agento11y

  • E2E Test

    grafana/agento11y

    Official

    Optional credential-free Hermes integration checks using an explicit loopback model provider and local telemetry receivers.

    128 GitHub stars~1.7k tokensUpdated today
    Auto-check: notes
  • Setup Local Guards

    grafana/agento11y

    Official

    Help choose, configure, and test local agento11y guard packs for coding-agent tool calls.

    128 GitHub stars~2.6k tokensUpdated today
    Auto-check: notes
  • Agento11y Eval Starter

    grafana/agento11y

    Official

    Use early in an AI-agent project — before ship, before real traffic — to decide which evaluations to set up and to scaffold a starter experiment.

    128 GitHub stars~6.4k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Agento11y Experiments

What does Agento11y Experiments do?

Run any Python LLM agent as an Agent Observability experiment using the public agento11y.experiments package: define a test suite, run an existing agent through typed trials, bind or record…. Agento11y Experiments is an agent skill from grafana/agento11y, published by the product's own GitHub organization.experiments package: define a test suite, run an existing agent through typed trials, bind or record generation I/O, grade outputs, and publish scores, including from stored Grafana test suites.

When should I use Agento11y Experiments?

Agento11y Experiments fits situations like: tasks that involve Test generation; tasks that involve Observability; tasks that involve Monitoring and alerting.

How do I install Agento11y Experiments in Claude Code?

Run `npx skills add grafana/agento11y --skill agento11y-experiments -a claude-code`. Or copy the skill folder (python/skills/agento11y-experiments in grafana/agento11y) into .claude/skills/agento11y-experiments in your project. Claude Code loads it when a task matches its description.

How do I install Agento11y Experiments in Codex?

Run `npx skills add grafana/agento11y --skill agento11y-experiments -a codex`. Or copy the skill folder (python/skills/agento11y-experiments in grafana/agento11y) into .agents/skills/agento11y-experiments in your project. Codex loads it when a task matches its description.

Can I use Agento11y Experiments in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add grafana/agento11y --skill agento11y-experiments -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agento11y-experiments, .gemini/skills/agento11y-experiments, .github/skills/agento11y-experiments and .opencode/skills/agento11y-experiments in your project.

What does Agento11y Experiments need to run?

Going by SKILL.md and its folder, Agento11y Experiments needs the command-line tools its instructions call (pip) and credentials named AGENTO11Y_SERVICE_ACCOUNT_TOKEN and AGENTO11Y_AUTH_TOKEN. Our summary lists: Python 3; A credential in AGENTO11Y_AUTH_TOKEN; A credential in AGENTO11Y_SERVICE_ACCOUNT_TOKEN.

Does Agento11y Experiments access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Agento11y Experiments safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agento11y Experiments use?

Agento11y Experiments is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agento11y Experiments use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agento11y Experiments?

Skills that share tags, products or a category with Agento11y Experiments: Opentelemetry (grafana/skills, 282 stars), Axiom Dashboard Builder (openclaw/clawhub, 9.5k stars), Happy Infra Metrics and Grafana (slopus/happy, 24k stars) and OpenTelemetry Pipeline Metrics Spec (comet-ml/opik, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agento11y Experiments?

grafana (a GitHub organization, an official publisher) maintains it in grafana/agento11y, which has 128 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 9, 2026.

Source: grafana/agento11y on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.