Sls Dashboard Builder
alibaba/loongsuite-pilot
当任务需要创建、修改、扩展或重组阿里云 SLS 的 dashboard JSON 或可导入的大盘配置时使用;尤其适用于线上大盘、强对比的分析看板、已校验的查询包,或需要专业中文标签与指标定义的运维向大盘。
Use early in an AI-agent project — before ship, before real traffic — to decide which evaluations to set up and to scaffold a starter experiment.
$ npx skills add grafana/agento11y --skill agento11y-eval-starter -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install grafana/agento11y agento11y-eval-starter --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/grafana/agento11y.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agento11y-eval-starter .claude/skills/agento11y-eval-starter && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agento11y-eval-starter" agent skill from https://github.com/grafana/agento11y/tree/main/skills/agento11y-eval-starter into .claude/skills/agento11y-eval-starter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agento11y-eval-starter", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/grafana/agento11y/tree/main/skills/agento11y-eval-starterType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add grafana/agento11y --skill agento11y-eval-starter -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install grafana/agento11y agento11y-eval-starter --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/grafana/agento11y.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/agento11y-eval-starter .agents/skills/agento11y-eval-starter && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agento11y-eval-starter" agent skill from https://github.com/grafana/agento11y/tree/main/skills/agento11y-eval-starter into .agents/skills/agento11y-eval-starter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agento11y-eval-starter", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add grafana/agento11y --skill agento11y-eval-starter -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install grafana/agento11y agento11y-eval-starter --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/grafana/agento11y.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/agento11y-eval-starter .cursor/skills/agento11y-eval-starter && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agento11y-eval-starter" agent skill from https://github.com/grafana/agento11y/tree/main/skills/agento11y-eval-starter into .cursor/skills/agento11y-eval-starter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agento11y-eval-starter", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/grafana/agento11y.git --path skills/agento11y-eval-starter--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add grafana/agento11y --skill agento11y-eval-starter -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install grafana/agento11y agento11y-eval-starter --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/grafana/agento11y.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/agento11y-eval-starter .gemini/skills/agento11y-eval-starter && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agento11y-eval-starter" agent skill from https://github.com/grafana/agento11y/tree/main/skills/agento11y-eval-starter into .gemini/skills/agento11y-eval-starter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agento11y-eval-starter", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install grafana/agento11y agento11y-eval-starterInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add grafana/agento11y --skill agento11y-eval-starter -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/grafana/agento11y.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/agento11y-eval-starter .github/skills/agento11y-eval-starter && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agento11y-eval-starter" agent skill from https://github.com/grafana/agento11y/tree/main/skills/agento11y-eval-starter into .github/skills/agento11y-eval-starter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agento11y-eval-starter", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add grafana/agento11y --skill agento11y-eval-starter -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install grafana/agento11y agento11y-eval-starter --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/grafana/agento11y.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/agento11y-eval-starter .opencode/skills/agento11y-eval-starter && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agento11y-eval-starter" agent skill from https://github.com/grafana/agento11y/tree/main/skills/agento11y-eval-starter into .opencode/skills/agento11y-eval-starter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agento11y-eval-starter", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agento11y-eval-starterUse early in an AI-agent project — before ship, before real traffic — to decide which evaluations to set up and to scaffold a starter experiment.
Agento11y Eval Starter is an agent skill from grafana/agento11y, published by the product's own GitHub organization. Use early in an AI-agent project — before ship, before real traffic — to decide which evaluations to set up and to scaffold a starter experiment. Reads the agent's own code (system prompt, tools, task), recommends specific evaluators with reasons that cite real lines, and writes a labeled draft test suite as an Agent Observability suite YAML. It also assesses how runnable the agent is: for an easily-invoked agent it generates a runner stub (runexperiment.py, or run-experiment.ts for a TypeScript agent) with one…
Its SKILL.md is about 6.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud, covering Observability, Test generation and Prompt engineering. It works with TypeScript and Grafana. The repository describes itself as: Actually Useful Agent Observability. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 447d692. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
javaFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
AGENTO11Y_AUTH_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agento11y Eval Starter loads about 6.4k tokens when it runs. Until then it costs about 214 tokens; SKILL.md has 2,458 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from grafana/agento11y at commit 447d692, republished under its Apache-2.0 licence (© grafana). 2,458 words, ~6,352 tokens.
.claude/skills/agento11y-eval-starter/SKILL.md (or your agent's skills folder).Help a developer who has an AI agent but no evaluation set up yet. The hard part isn't running an experiment — it's knowing what to evaluate and having cases to test against before there's any traffic. Answer both, grounded in the agent's actual code.
Always produce:
Then, depending on how runnable the agent is (Step 1):
run_experiment.py, or
run-experiment.ts for a TypeScript agent) that wires the suite to the SDK
with one hole to fill — and optionally run it (Step 6), only with permission.
For an agent that needs a harness or full runtime, point to the existing eval infra instead
of a runner that can't actually call it.This skill is language-agnostic — the reading, recommending, and YAML it produces do not
depend on the agent's language. What differs is how runnable the agent is: recommendations +
YAML always apply, but the runner (Step 4) and the optional run (Step 6) adapt to whether the
agent has a clean function seam or needs a harness / full stack. For deeper run-side patterns
(binding existing generations, cross-process verifiers) point to the per-language run skill
(Python: agento11y-experiments).
AGENTO11Y_ENDPOINT + credentials (Grafana Cloud); if
they are not set, ask for them, do not invent an endpoint.localhost);
the developer supplies their Cloud endpoint and token.Find and read these in the target repo, and record the file path and line range of each:
Every recommendation must cite one of these locations.
Also assess runnability — how hard it is to invoke this agent for one input and get its output — because it decides what Step 4 produces. Classify it as one of:
run_agent.Signals that push toward in-process/full-stack: the agent isn't Python; tools hit real
APIs/datasources; there is already a dedicated eval harness, Docker stack, or
scenario/ground-truth files in the repo. If a repo already has eval scenarios with expected
outputs, note them — they are better test cases than anything generated, and later steps
should point to or reuse them.
Separately, note the agent's language, because the agento11y experiments SDK exists in Python, Go, and JavaScript/TypeScript (not Java or .NET yet). This is independent of runnability — a Java agent can be trivially runnable yet have no experiments SDK in its language. If the agent is Python, Go, or TypeScript, the runner is native. If it is Java or .NET, say so plainly: the runner must be Python, Go, or TypeScript (calling the agent across a process boundary), or the developer waits for experiments support in their language. Do not imply a native Java or .NET experiments API exists.
The SDK models an evaluator as an id plus a kind. There are two kinds:
llm_judge — a model grades the output against a rubric (relevance, helpfulness,
groundedness, tone, task completion, format, safety, and similar judgment calls).deterministic — code decides pass/fail (exact/substring match, JSON validity, schema
or regex shape, length bounds, "not empty", a required field is present).Pick 3–6 (not more). Map each to what you read in Step 1, and give a one-line why per pick
that cites a file:line. Common mappings:
| If the agent… | Evaluator | Kind |
|---|---|---|
| answers open-ended user requests | relevance, helpfulness | llm_judge |
| retrieves / cites sources / does RAG | groundedness | llm_judge |
| must stay on-topic / in-scope | task_adherence | llm_judge |
| must emit JSON or a fixed shape | json_valid / schema_match | deterministic |
| must follow a specific output format | format_adherence | llm_judge |
| calls tools | tool_call_correct | llm_judge (or deterministic if the correct call is checkable in code) |
| handles user data that could be echoed | pii_leak | llm_judge |
| produces public-facing text | toxicity | llm_judge |
| must always return something | response_not_empty | deterministic |
Evaluator ids are yours to choose — pick clear, stable ids. Be concrete about what "defining"
each one means, because it is not an Agent Observability control-plane action here: in the offline SDK flow
an evaluator is code the developer writes in the runner (Step 4). An llm_judge is a
function that calls a model and returns a score; a deterministic one is a plain code check.
experiments.Evaluator(evaluator_id=..., kind=...) is only the label attached to the score, not the
logic. (Forking a predefined template is a separate online-eval path, not needed here.)
One exception, for tenants that already have evaluators: if the developer points at an evaluator
that exists in Agent Observability already, the Python runner can bind the trial to the agent's
conversation and let that evaluator grade it, instead of writing the check in the runner:
trial.bind_conversation(conversation_id) then trial.evaluate("<existing-evaluator-id>"). The
trial then closes as completed with no local final_score. Do not create the evaluator to make
this work, and do not suggest it when the tenant has none; this skill only uses what is already
there.
This skill targets the offline phase: run these evaluators as offline experiments against the suite from Step 3, before there is traffic. Do not recommend live/online evaluation to an agent with no traffic.
Once the agent ships and has traffic, the same evaluation criteria can also be applied online — as Agent Observability rules over ingested conversations, or as SDK guard hooks on the request path — but those are separate Agent Observability surfaces, configured elsewhere, and out of scope for this skill. Mention this only as a one-line "next, once you have traffic" note; do not instruct on it here.
Write a suite file in the target repo (suggest evals/<agent>-starter.yaml). It must load
with the SDK's TestSuite.from_yaml(...), so match this schema exactly:
suite_id (required), plus optional name, version, description, tags,
changelog, and cases (a list). version defaults to 1.0.0.id (required), plus optional name, description, tags, category,
input, expected, weight, metadata. input and expected are free-form (a string
or a mapping).Derive cases from the agent's real task. Produce at least 6, weighted toward edge and
adversarial (generated cases skew easy). Use category values happy, edge,
adversarial. Keep the header comment verbatim.
Every case must actually reach the agent — test the agent, not its harness. From Step 1, note where the entrypoint parses/validates input before the model runs (e.g. a JSON parse, a CLI arg check). Do not generate cases that fail in that pre-agent layer (malformed JSON, wrong CLI flags) as if they tested the agent; if such a boundary is worth covering, label it clearly as a harness case in its notes, or leave it out.
# STARTER DRAFT — review before use. Generated from your agent code (<file:line refs>).
# NOT validated. Add your own real cases; the edge/adversarial cases need your judgment on
# expected behavior. Loads with agento11y TestSuite.from_yaml(...).
suite_id: <agent>-starter
name: <Agent> starter suite
version: 1.0.0
cases:
- id: happy-basic-request
category: happy
tags: [smoke]
input:
prompt: "<a real request this agent is built to answer>"
expected: "<what a good answer looks like, or a rubric note>"
- id: edge-underspecified
category: edge
input:
prompt: "<a vague / multi-part / boundary request>"
expected: "<how the agent should handle ambiguity>"
- id: adversarial-prompt-injection
category: adversarial
input:
prompt: "<an injection / out-of-scope / data-extraction attempt>"
expected: "<agent should refuse / stay in scope / not leak>"The suite YAML alone does not run anything. The agento11y SDK stores and aggregates scores, but it does not run the agent or compute the evaluators — the developer writes both. Left at just the YAML, a developer new to offline eval is still blocked ("now what?"). So generate a minimal bootstrap runner — just enough to get one experiment running.
This is deliberately the simplest path, not the full run-side API. The agento11y-experiments
skill is the reference for everything beyond bootstrap — binding an already-instrumented
agent's real generations/conversations, auditable LLM-judge grading, cross-process verifiers
(TrialRef), pass@k/pass^k. Don't reproduce those patterns here; generate the minimal runner
and point to agento11y-experiments for depth.
What you generate depends on the runnability you assessed in Step 1:
evals/run_experiment.py for a Python or Go
agent, evals/run-experiment.ts for a TypeScript agent. One hole to fill (run_agent).
A Go agent gets the Python runner, because this skill ships no Go template.run_agent that can't actually call the agent (e.g.
a Go agent). Generate the same experiment wiring, but make run_agent shell out to a small
harness in the agent's language (or write that harness), and say plainly the seam is the
injectable LLM client. Point to any existing test that already invokes the agent in-process
as the template.Also branch on language (the experiments SDK covers Python, Go, and JavaScript/TypeScript):
run_experiment.py, the Go
agento11y package, or the @grafana/agento11y/experiments subpath).java -jar your-agent.jar and reads its output), or the
developer waits for experiments support in their language. Offer the subprocess bridge only as a
labeled option with its cost (serializing input to the CLI, parsing output), not as a clean
default.For an easy Python or Go agent, write evals/run_experiment.py. It must:
TestSuite.from_yaml(...).experiments.experiment(...)) and one trial per case.run_agent(case) — this is the
one hole the developer fills; wire it to the real entrypoint you found in Step 1.llm_judge — a real model
call that returns a JSON {score, passed, explanation}), so they see the shape and can copy
it for the others. Reference the rest by name in a comment; do not stub all of them.trial.record_io(...)) and emit trial.final_score(...) with the evaluator.Keep the header verbatim, and be honest in it about what still needs doing:
#!/usr/bin/env python3
"""STARTER RUNNER — generated by agento11y-eval-starter, review before use.
Runs <agent> over evals/<agent>-starter.yaml as an Agent Observability experiment and publishes scores.
You still need to: (1) fill run_agent(case) to call YOUR agent; (2) tune the sketched
judge; (3) set real credentials — AGENTO11Y_ENDPOINT + AGENTO11Y_AUTH_TOKEN for your Grafana Cloud
stack. The SDK stores scores; it does not run the agent or the judge.
Set AGENTO11Y_INGEST_ACTOR to a stable value: the run and its trials must share one actor, or
trial creation fails with "401: experiment is owned by another actor".
AGENTO11Y_ENDPOINT=... AGENTO11Y_AUTH_TOKEN=... AGENTO11Y_INGEST_ACTOR=ingest:sdk/python \
python evals/run_experiment.py
"""
import json, os, time
from pathlib import Path
from dotenv import load_dotenv
from agento11y import experiments
load_dotenv()
SUITE = Path(__file__).parent / "<agent>-starter.yaml"
def run_agent(case: experiments.TestCase) -> str:
"""THE ONE HOLE YOU FILL — call your agent for this case, return its output text."""
raise NotImplementedError("wire this to your agent entrypoint (see Step 1 refs)")
def judge_<evaluator>(case_input, output) -> tuple[float, bool, str]:
"""Sketched llm_judge — a model call returning (score 0-1, passed, explanation). Tune it."""
import litellm
prompt = f"Grade <what this evaluator checks>. Return JSON {{\"score\":0-1,\"passed\":bool,\"explanation\":\"...\"}}.\n\nInput:\n{case_input}\n\nOutput:\n{output}"
model = os.getenv("GRADER_MODEL") or os.getenv("MODEL_NAME") # a LIVE model id; no default (dead ids 404)
text = litellm.completion(model=model, messages=[{"role": "user", "content": prompt}],
temperature=0, max_tokens=300).choices[0].message.content or "{}"
s, e = text.find("{"), text.rfind("}")
d = json.loads(text[s:e + 1]) if s >= 0 else {}
score = max(0.0, min(1.0, float(d.get("score", 0.0))))
return score, bool(d.get("passed", score >= 0.6)), str(d.get("explanation", ""))
def main() -> None:
suite = experiments.TestSuite.from_yaml(str(SUITE))
verifier = experiments.Evaluator(evaluator_id="<evaluator>", version="draft-0", kind="llm_judge")
candidate = {
"agent_name": "<agent>",
# Always send a declared agent_version. Without it Agent Observability auto-derives a version from
# the system-prompt hash, so you can't reliably attribute scores to a version or
# compare versions. Replace "v1" with your real version (git tag, prompt version,
# semver...) — the "v1" fallback is a placeholder, not a version worth comparing.
"agent_version": os.getenv("AGENT_VERSION", "v1"),
"git_sha": os.getenv("GIT_SHA", ""),
"model_name": os.getenv("MODEL_NAME", ""),
}
with experiments.experiment(name="<agent> starter", experiment_id=f"<agent>-starter-{int(time.time())}",
suite=suite, candidate=candidate, tags=["starter"],
actor=os.getenv("AGENTO11Y_INGEST_ACTOR", "ingest:sdk/python")) as exp:
for case in suite.test_cases:
with exp.trial(case) as trial:
out = run_agent(case)
trial.record_io(input=json.dumps(case.input), output=out,
model_provider="<provider>", model_name=os.getenv("MODEL_NAME", ""))
score, passed, why = judge_<evaluator>(case.input, out)
trial.final_score(score, passed=passed, explanation=why, evaluator=verifier)
print(f" {case.test_case_id}: score={score:.2f} passed={passed}")
print(f"\nExperiment: {exp.url}")
if __name__ == "__main__":
main()For an easy TypeScript agent, write evals/run-experiment.ts against the
@grafana/agento11y/experiments subpath. Same shape and same one hole, in JS spellings:
parseSuiteYAML(...) loads the suite, withExperiment(client, {...}, ...) and
experiment.withTrial(case, ...) own the lifecycle, trial.recordIO({input, output}) records
I/O, and trial.finalScore(score, {passed, explanation, evaluator}) publishes the verdict:
/**
* STARTER RUNNER - generated by agento11y-eval-starter, review before use.
*
* Runs <agent> over evals/<agent>-starter.yaml as an Agent Observability experiment and
* publishes scores.
*
* You still need to: (1) fill runAgent(case) to call YOUR agent; (2) tune the sketched judge;
* (3) set real credentials - AGENTO11Y_ENDPOINT + AGENTO11Y_AUTH_TOKEN for your Grafana Cloud
* stack. The SDK stores scores; it does not run the agent or the judge.
*
* Set AGENTO11Y_INGEST_ACTOR to a stable value: the run and its trials must share one actor,
* or trial creation fails with "401: experiment is owned by another actor".
*
* AGENTO11Y_ENDPOINT=... AGENTO11Y_AUTH_TOKEN=... AGENTO11Y_INGEST_ACTOR=ingest:sdk/js \
* npx tsx evals/run-experiment.ts
*/
import { readFileSync } from 'node:fs';
import { ExperimentsClient, parseSuiteYAML, type TestCase, withExperiment } from '@grafana/agento11y/experiments';
const suite = parseSuiteYAML(readFileSync(new URL('./<agent>-starter.yaml', import.meta.url), 'utf8'));
const verifier = { evaluatorId: '<evaluator>', version: 'draft-0', kind: 'llm_judge' } as const;
/** THE ONE HOLE YOU FILL - call your agent for this case and return its output text. */
async function runAgent(testCase: TestCase): Promise<string> {
throw new Error('wire this to your agent entrypoint (see Step 1 refs)');
}
/** Sketched llm_judge - one model call returning a 0-1 score, a verdict, and an explanation. */
async function judge(input: unknown, output: string): Promise<{ score: number; passed: boolean; why: string }> {
// Replace with your model call, reading the model id from GRADER_MODEL or MODEL_NAME.
// Use a live model id; there is no default, and a dead id 404s.
throw new Error('tune this judge');
}
// endpoint, tenantId, and ingestToken fall back to AGENTO11Y_ENDPOINT,
// AGENTO11Y_AUTH_TENANT_ID, and AGENTO11Y_AUTH_TOKEN.
const client = new ExperimentsClient({ actor: process.env.AGENTO11Y_INGEST_ACTOR ?? 'ingest:sdk/js' });
await withExperiment(
client,
{
// A fresh id per run: reusing one created by another actor fails with 401.
experimentId: `<agent>-starter-${Math.floor(Date.now() / 1000)}`,
name: '<agent> starter',
suite,
tags: ['starter'],
candidate: {
agentName: '<agent>',
// Always declare a version. Without it Agent Observability derives one from the
// system-prompt hash, and scores cannot be attributed to a version.
agentVersion: process.env.AGENT_VERSION ?? 'v1',
modelName: process.env.MODEL_NAME ?? '',
},
},
async (experiment) => {
for (const testCase of suite.testCases) {
await experiment.withTrial(testCase, async (trial) => {
const output = await runAgent(testCase);
trial.recordIO({ input: JSON.stringify(testCase.input), output });
const { score, passed, why } = await judge(testCase.input, output);
trial.finalScore(score, { passed, explanation: why, evaluator: verifier });
console.log(` ${testCase.testCaseId}: score=${score} passed=${passed}`);
});
}
console.log(`Experiment: ${experiment.url}`);
},
);Use a fresh experiment_id per run (a timestamp works) — reusing an id created by a
different auth actor fails with 401: experiment is owned by another actor.
Output, in this order:
why (with file:line).
Add a one-line "once you have traffic, these criteria can also run online (Agent Observability rules or
guard hooks) — separate surfaces, not set up here."evals/<agent>-starter.yaml and the runner:
evals/run_experiment.py, or evals/run-experiment.ts for TypeScript), and a one-line
reminder to review the edge/adversarial cases and add real ones.run_agent(case) (runAgent in TypeScript),
tune the sketched judge
(and add the other recommended evaluators the same way), and set credentials
(AGENTO11Y_ENDPOINT + AGENTO11Y_AUTH_TOKEN). State the boundary explicitly: this skill only
bootstraps the first run; for anything past that — binding an already-instrumented agent's
real generations, auditable LLM-judge grading, cross-process verifiers, repeated-sampling
metrics — the agento11y-experiments skill is the reference.Only offer this for an easy agent (clean function seam) in Python, Go, or TypeScript
(the languages with an experiments SDK). For in-process/full-stack agents, or agents in a
language with no experiments SDK (Java/.NET), don't offer to run — point to the existing
harness/infra (or the subprocess-bridge option) and stop; a real run there is out of scope for
this skill.
After the summary, offer to run the starter experiment for them — do not run automatically. Ask: "Want me to try running this now?" Only proceed if they say yes.
If they accept:
run_agent(case) — wire it to the real entrypoint from Step 1, so the runner
actually calls their agent.AGENTO11Y_ENDPOINT + AGENTO11Y_AUTH_TOKEN. If either is unset, ask the developer for it
proactively — AGENTO11Y_AUTH_TOKEN is their Grafana Cloud ingestion API key (Cloud portal →
stack → API keys), AGENTO11Y_ENDPOINT their stack endpoint. Never mint, generate, or
fabricate a token yourself, and never invent an endpoint — the developer owns the
credential and supplies it; you only read it from the environment or ask for it.MODEL_NAME is a live model (dead model ids fail with a 404 not_found).AGENTO11Y_INGEST_ACTOR so run and trials share one actor (else 401: owned by another actor).AGENT_VERSION (in the candidate). Without it Agent Observability auto-derives a version
from the system-prompt hash, and the developer can't attribute scores to a version or
compare versions — which is the whole point of the agent's Quality view. Confirm a real
value (git tag / prompt version / semver), don't leave it defaulted.exp.url, and note this published one experiment's
scores (no tenant evaluators/rules/guards were created).If they decline, stop after the summary.
© grafana, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/agento11y-eval-starter of grafana/agento11y.
Open the folder on GitHubat commit 447d692
Agento11y Eval Starter next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agento11y Eval Starter this skillgrafana/agento11y | 128 | — | ~6.4k | Automated safety check: Pass | Apache-2.0 | |
| Sls Dashboard Builderalibaba/loongsuite-pilot | 201 | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| Loongsuite Pilot Insightalibaba/loongsuite-pilot | 201 | — | ~944 | Automated safety check: Pass | Apache-2.0 | |
| Langchain Otel Observabilityjeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~3.6k | Automated safety check: Pass | MIT | |
| OpenTelemetry Pipeline Metrics Speccomet-ml/opik | 22k | — | ~3.2k | Automated safety check: Pass | Apache-2.0 | |
| Liveblog Devliveblog/liveblog | 119 | — | ~1.9k | Automated safety check: Pass | AGPL-3.0 |
alibaba/loongsuite-pilot
当任务需要创建、修改、扩展或重组阿里云 SLS 的 dashboard JSON 或可导入的大盘配置时使用;尤其适用于线上大盘、强对比的分析看板、已校验的查询包,或需要专业中文标签与指标定义的运维向大盘。
alibaba/loongsuite-pilot
基于 LoongSuite Pilot / AI Coding Agent 日志生成事件洞察、组织洞察、数据质量、研发效能和 AI Native 使用类 SLS 报表时使用;包含 AI Coding 事件表语义,以及团队报表可选的部门维表、deptuser 组织关系、指标口径和公共 CTE,通常与 sls-dashboard-builder 一起使用。
jeremylongshore/tons-of-skills-marketplace
Wire LangChain 1.0 / LangGraph 1.0 traces into an OpenTelemetry-native backend (Jaeger, Honeycomb, Grafana Tempo, Datadog) with LLM-specific SLOs, safe prompt-content policy, and subgraph-aware span…
comet-ml/opik
Specifies how to instrument an opik-backend pipeline with per-stage OpenTelemetry metrics for throughput, latency, errors and queue delay by workspace.
liveblog/liveblog
Run a local Liveblog development environment. An agent skill from liveblog/liveblog.
openclaw/clawhub
Investigates incidents and production problems with hypothesis-driven debugging, queries Axiom observability data when available, and keeps secrets out of commands and output.
grafana/agento11y
Optional credential-free Hermes integration checks using an explicit loopback model provider and local telemetry receivers.
grafana/agento11y
Help choose, configure, and test local agento11y guard packs for coding-agent tool calls.
grafana/agento11y
Run any Python LLM agent as an Agent Observability experiment using the public agento11y.experiments package: define a test suite, run an existing agent through typed trials, bind or record…
Works with
Categories
Use early in an AI-agent project — before ship, before real traffic — to decide which evaluations to set up and to scaffold a starter experiment. Agento11y Eval Starter is an agent skill from grafana/agento11y, published by the product's own GitHub organization. Use early in an AI-agent project — before ship, before real traffic — to decide which evaluations to set up and to scaffold a starter experiment.
Agento11y Eval Starter fits situations like: tasks that involve Observability; tasks that involve Test generation; tasks that involve Prompt engineering.
Run `npx skills add grafana/agento11y --skill agento11y-eval-starter -a claude-code`. Or copy the skill folder (skills/agento11y-eval-starter in grafana/agento11y) into .claude/skills/agento11y-eval-starter in your project. Claude Code loads it when a task matches its description.
Run `npx skills add grafana/agento11y --skill agento11y-eval-starter -a codex`. Or copy the skill folder (skills/agento11y-eval-starter in grafana/agento11y) into .agents/skills/agento11y-eval-starter in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add grafana/agento11y --skill agento11y-eval-starter -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agento11y-eval-starter, .gemini/skills/agento11y-eval-starter, .github/skills/agento11y-eval-starter and .opencode/skills/agento11y-eval-starter in your project.
Going by SKILL.md and its folder, Agento11y Eval Starter needs the command-line tools its instructions call (java) and credentials named AGENTO11Y_AUTH_TOKEN. Our summary lists: Python 3; Docker.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agento11y Eval Starter is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.4k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Agento11y Eval Starter: Sls Dashboard Builder (alibaba/loongsuite-pilot, 201 stars), Loongsuite Pilot Insight (alibaba/loongsuite-pilot, 201 stars), Langchain Otel Observability (jeremylongshore/tons-of-skills-marketplace, 2.8k stars) and OpenTelemetry Pipeline Metrics Spec (comet-ml/opik, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
grafana (a GitHub organization, an official publisher) maintains it in grafana/agento11y, which has 128 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 9, 2026.
Source: grafana/agento11y on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.