Eval
agentevals-dev/agentevals
Evaluate and score agent behavior against a golden reference.
Analyze & debug GenAI/LLM apps: token cost & caching by prompt, model & provider; latency/errors; agent & tool loops/failures; conversations; guardrails; evaluations; OpenTelemetry/dt-evals setup.
$ npx skills add Dynatrace/dynatrace-for-ai --skill dt-obs-genai -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Dynatrace/dynatrace-for-ai dt-obs-genai --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Dynatrace/dynatrace-for-ai.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/dt-obs-genai .claude/skills/dt-obs-genai && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "dt-obs-genai" agent skill from https://github.com/Dynatrace/dynatrace-for-ai/tree/main/skills/dt-obs-genai into .claude/skills/dt-obs-genai/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dt-obs-genai", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Dynatrace/dynatrace-for-ai/tree/main/skills/dt-obs-genaiType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Dynatrace/dynatrace-for-ai --skill dt-obs-genai -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Dynatrace/dynatrace-for-ai dt-obs-genai --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Dynatrace/dynatrace-for-ai.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/dt-obs-genai .agents/skills/dt-obs-genai && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "dt-obs-genai" agent skill from https://github.com/Dynatrace/dynatrace-for-ai/tree/main/skills/dt-obs-genai into .agents/skills/dt-obs-genai/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dt-obs-genai", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Dynatrace/dynatrace-for-ai --skill dt-obs-genai -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Dynatrace/dynatrace-for-ai dt-obs-genai --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Dynatrace/dynatrace-for-ai.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/dt-obs-genai .cursor/skills/dt-obs-genai && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "dt-obs-genai" agent skill from https://github.com/Dynatrace/dynatrace-for-ai/tree/main/skills/dt-obs-genai into .cursor/skills/dt-obs-genai/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dt-obs-genai", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Dynatrace/dynatrace-for-ai.git --path skills/dt-obs-genai--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Dynatrace/dynatrace-for-ai --skill dt-obs-genai -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Dynatrace/dynatrace-for-ai dt-obs-genai --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Dynatrace/dynatrace-for-ai.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/dt-obs-genai .gemini/skills/dt-obs-genai && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "dt-obs-genai" agent skill from https://github.com/Dynatrace/dynatrace-for-ai/tree/main/skills/dt-obs-genai into .gemini/skills/dt-obs-genai/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dt-obs-genai", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Dynatrace/dynatrace-for-ai dt-obs-genaiInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Dynatrace/dynatrace-for-ai --skill dt-obs-genai -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Dynatrace/dynatrace-for-ai.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/dt-obs-genai .github/skills/dt-obs-genai && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "dt-obs-genai" agent skill from https://github.com/Dynatrace/dynatrace-for-ai/tree/main/skills/dt-obs-genai into .github/skills/dt-obs-genai/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dt-obs-genai", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Dynatrace/dynatrace-for-ai --skill dt-obs-genai -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Dynatrace/dynatrace-for-ai dt-obs-genai --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Dynatrace/dynatrace-for-ai.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/dt-obs-genai .opencode/skills/dt-obs-genai && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "dt-obs-genai" agent skill from https://github.com/Dynatrace/dynatrace-for-ai/tree/main/skills/dt-obs-genai into .opencode/skills/dt-obs-genai/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dt-obs-genai", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
dt-obs-genaiAnalyze & debug GenAI/LLM apps: token cost & caching by prompt, model & provider; latency/errors; agent & tool loops/failures; conversations; guardrails; evaluations; OpenTelemetry/dt-evals setup.
Dt Obs Genai is an agent skill from Dynatrace/dynatrace-for-ai. Analyze & debug GenAI/LLM apps: token cost & caching by prompt, model & provider; latency/errors; agent & tool loops/failures; conversations; guardrails; evaluations; OpenTelemetry/dt-evals setup.
Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including reference files (for example `references/agent-signals.md`, `references/conversation-analytics.md` and `references/cost-and-tokens.md`).
It sits in AI & LLM Engineering, covering Observability, LLM cost and token optimization and LLM evaluation. It works with OpenTelemetry. The repository describes itself as: Skills, prompts, and instructions for building AI agents on top of Dynatrace production context. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 4f9aa71. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are dql).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Dt Obs Genai loads about 4.5k tokens when it runs, and up to ~21k if it reads all its reference files. Until then it costs about 52 tokens; SKILL.md has 1,438 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Dynatrace/dynatrace-for-ai at commit 4f9aa71, republished under its Apache-2.0 licence (© Dynatrace). 1,438 words, ~4,544 tokens.
.claude/skills/dt-obs-genai/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.Analyze AI Observability signals from customer GenAI applications using DQL — golden signals, LLM signals, token and cost analytics (with usage attribution and prompt-caching economics), agent signals (including loop/runaway detection and Smartscape topology), conversation/session-level analytics, guardrails, and evaluation quality.
Use this skill for observability questions about customer GenAI applications — anything
reading OpenTelemetry GenAI spans (gen_ai.*) or LLM evaluation bizevents. Example triggers:
Davis CoPilot/MCP telemetry (dt-platform), generic service metrics (dt-obs-services), logs (dt-obs-logs), or non-GenAI distributed tracing (dt-obs-tracing).
When suggesting follow-up questions (e.g., "give me one example question per topic"), use these canonical, ready-to-ask phrasings — one per capability. Keep suggestions single-clause and avoid the literal phrase "content filter" (use "blocked or safety-filtered" instead); overly long, multi-clause questions can be rejected.
The four classic observability signals — traffic, errors, latency, and saturation — apply
directly to GenAI applications. Traffic is request throughput over time; errors are spans
where span.status_code == "error"; latency is the duration field (a Grail duration
value — divide by the 1ms literal, duration / 1ms, for a numeric millisecond value);
saturation is proxied by total token throughput per minute (input + output tokens combined).
fetch spans, from: now()-24h
| filter isNotNull(gen_ai.request.model)
| summarize total = count(), errors = countIf(span.status_code == "error"), by: {gen_ai.request.model}
| fieldsAdd error_rate_pct = if(total > 0, errors * 100.0 / total, else: 0.0)
| sort error_rate_pct desc→ Full traffic, latency, and saturation queries: See references/golden-signals.md
LLM signals describe which model and provider served each request, what operation type
was invoked (chat, execute_tool, invoke_agent, create_agent), and how tokens were
consumed. Use these to benchmark provider latency, compare model performance, and
understand the token distribution across model-provider combinations.
fetch spans, from: now()-24h
| filter isNotNull(gen_ai.request.model)
| summarize p95_ms = percentile(duration, 95) / 1ms, requests = count(), by: {gen_ai.provider.name}
| sort p95_ms desc→ Slowest models, token usage by model: See references/llm-signals.md
Token consumption is the primary cost driver. Dynatrace stores gen_ai.usage.input_tokens
and gen_ai.usage.output_tokens on every span — there is no stored cost field; estimated
cost must be derived by multiplying token sums by the per-model price you supply. Use these
queries to identify the highest-spend model-provider combinations and detect token-burn spikes.
fetch spans, from: now()-24h
| filter isNotNull(gen_ai.usage.input_tokens) or isNotNull(gen_ai.usage.output_tokens)
| summarize input_tokens = sum(gen_ai.usage.input_tokens), output_tokens = sum(gen_ai.usage.output_tokens), total_tokens = sum(gen_ai.usage.input_tokens) + sum(gen_ai.usage.output_tokens), by: {gen_ai.provider.name, gen_ai.request.model}
| sort total_tokens desc→ Token spikes, cost estimation, most expensive prompts, usage attribution, prompt-caching economics: See references/cost-and-tokens.md
GenAI agents emit spans for each tool invocation (execute_tool), agent step
(invoke_agent), and agent creation (create_agent). Use agent signals to identify
which tools are called most often and which agents are failing. For structural
questions — which agents, models, and providers exist and how they connect — query the
GenAI Smartscape entities (GENAI_AGENT, GENAI_MODEL, GENAI_PROVIDER, GENAI_SERVICE)
instead of scanning spans; this is a feature-flag-gated preview.
fetch spans, from: now()-24h
| filter isNotNull(gen_ai.agent.name)
| summarize total = count(), errors = countIf(span.status_code == "error"), by: {gen_ai.agent.name}
| fieldsAdd error_rate_pct = if(total > 0, errors * 100.0 / total, else: 0.0)
| sort errors desc→ Tool usage, failing agents, agent step latency, loop/runaway detection, Smartscape topology: See references/agent-signals.md
Per-span and per-trace signals measure one request or one turn. When the application
propagates gen_ai.conversation.id, you can roll spans up to the session level — cost
per conversation, how deep conversations run, and which sessions are runaway-expensive or
error-prone. This is the unit that matters for chargeback and user-perceived reliability.
fetch spans, from: now()-24h
| filter isNotNull(gen_ai.conversation.id)
| summarize turns = countDistinct(trace.id), total_tokens = sum(gen_ai.usage.input_tokens) + sum(gen_ai.usage.output_tokens), errors = countIf(span.status_code == "error"), by: {gen_ai.conversation.id}
| sort total_tokens desc→ Cost/depth per conversation, session error rate: See references/conversation-analytics.md
Guardrails surface as gen_ai.response.finish_reasons on the span — content_filter means a
safety filter blocked or redacted output, length means the response was truncated at the
token limit — and as the proactive safety evaluators (prompt-injection, pii-leakage,
toxicity, bias) in the evaluation bizevents. Use these to quantify blocked and truncated
responses and tie them back to the LLM-judge safety verdicts.
fetch spans, from: now()-24h
| filter isNotNull(gen_ai.response.finish_reasons)
| fieldsAdd finish_reason = gen_ai.response.finish_reasons
| expand finish_reason
| summarize calls = count(), by: {finish_reason, gen_ai.request.model}
| sort calls desc→ Blocked (safety-filtered), truncated (length), finish-reason breakdown: See references/guardrails.md
Evaluation results are captured as bizevents (not spans) with
event.type == "gen_ai.evaluation.result". Each evaluator emits one bizevent per
response, carrying the score, pass/fail label, explanation, and the exact Q&A pair.
Use evaluation queries to monitor quality dimensions and surface failed responses
with the LLM judge's reasoning. Each bizevent also carries the trace.id of the run that
produced the evaluated response, so you can pivot from a quality failure to the spans that
caused it.
fetch bizevents, from: now()-24h
| filter event.type == "gen_ai.evaluation.result"
| filter gen_ai.evaluation.score.label == "fail"
| fields timestamp, gen_ai.evaluation.name, gen_ai.evaluation.score.value, gen_ai.evaluation.explanation, gen_ai.evaluation.input.question, gen_ai.evaluation.input.answer
| sort timestamp desc→ Quality scores, failed evaluations, fail rates: See references/evaluations.md
When any signal query returns no rows, do not report "no data found" — first confirm whether the application sends GenAI telemetry at all. These two presence checks show which signal families are present:
fetch spans, from: now()-24h
| summarize
has_genai = countIf(isNotNull(gen_ai.request.model)),
has_tokens = countIf(isNotNull(gen_ai.usage.input_tokens) or isNotNull(gen_ai.usage.output_tokens)),
has_agents = countIf(isNotNull(gen_ai.agent.name)),
has_tools = countIf(gen_ai.operation.name == "execute_tool"),
has_conversation = countIf(isNotNull(gen_ai.conversation.id)),
has_finish_reason = countIf(isNotNull(gen_ai.response.finish_reasons)),
has_cached_tokens = countIf(isNotNull(gen_ai.usage.cache_read.input_tokens) or isNotNull(gen_ai.usage.cache_creation.input_tokens)),
total = count()
fetch bizevents, from: now()-24h
| filter event.type == "gen_ai.evaluation.result"
| summarize evals = count()If has_genai is zero, report that the application appears not to be instrumented for AI
Observability yet — not "no data found". If has_genai is non-zero but a specific family
(has_tokens, has_agents, has_tools, has_conversation, has_finish_reason,
has_cached_tokens, evals) is zero, only that signal type is missing — for example
has_conversation == 0 means session-level analytics are unavailable because the app does
not propagate a conversation id, and has_cached_tokens == 0 means prompt-caching telemetry
is not being reported. These optional families may use different attribute names depending on
the provider/SDK; verify before reporting them absent.
When a user asks for analysis, proceed immediately with sensible defaults. Do not ask for parameter values you can reasonably assume.
Default values when not specified:
| Parameter | Default | Rationale |
|---|---|---|
| Timeframe | Last 24 h (from: now()-24h) | Covers a full operational day without being too narrow |
| Model scope | All models (no model filter) | Shows the full picture; user can narrow after seeing results |
| Provider scope | All providers | Same rationale as model scope |
| Token threshold | None | Show all — let the data reveal the outliers |
Exception — cost prices. Per-model prices are the one input you cannot default (there is no cost field in the data). Ask the user for them before estimating USD; never use prices from memory. See cost-and-tokens.md.
When any signal query returns no rows, run the two presence checks in the
Empty-State Check capability above before responding — never reply "no data found".
If has_genai is zero, report that the application appears not to be instrumented for AI
Observability yet; if only a specific family is zero, say which signal type is missing.
This skill covers AI Observability signals for customer GenAI applications only.
Product documentation and configuration how-to questions (e.g., "How do I configure
the Dynatrace OTLP endpoint?") go to ask-dynatrace-docs — this skill does not
contain product configuration how-tos.
Map user requests and prompt-starter phrasings to capabilities:
| User Request / Prompt Starter | Capability | Reference File |
|---|---|---|
| "Understand AI Observability signals" | All signal categories overview | This SKILL.md |
| "Analyze LLM latency and errors", "LLM errors", "error rate by model" | Golden Signals | golden-signals.md |
| "Which models are slowest right now?", "compare latency across providers" | LLM Signals | llm-signals.md |
| "Show token usage by model", "token usage spikes" | Cost and Tokens | cost-and-tokens.md |
| "Break down cost by model and provider", "which prompts are most expensive?" | Cost and Tokens | cost-and-tokens.md |
| "Trace a failing agent run", "show failed tool calls" | Agent Signals | agent-signals.md |
| "Break down agent steps by latency" | Agent Signals | agent-signals.md |
| "Map agent topology", "which models does this agent use?", "list GenAI agents/models/providers" | Agent Signals (Smartscape) | agent-signals.md |
| "Is an agent stuck in a loop?", "find runaway agents", "what caused the token spike?" | Agent Signals (loops) | agent-signals.md |
| "Cost per conversation", "most expensive sessions", "how deep do conversations run?" | Conversation Analytics | conversation-analytics.md |
| "Stitch together an agent trajectory", "filter by session id", "connect traces across a session" | Conversation Analytics | conversation-analytics.md |
| "How often are responses blocked/filtered?", "are responses being truncated?", "finish reasons" | Guardrails | guardrails.md |
| "Cost by application/user/tenant", "who is driving token spend?" | Cost and Tokens (attribution) | cost-and-tokens.md |
| "Do I have prompt caching?", "cache hit rate", "caching savings" | Cost and Tokens (caching) | cost-and-tokens.md |
| "Summarize evaluation quality scores", "show low-scoring responses", "show failed evaluations" | Evaluation Quality | evaluations.md |
| "What signals am I missing?", "why is there no data?" | Empty-State Check | This SKILL.md |
1. Run token usage by model and provider (cost-and-tokens.md → "Token usage by model and provider")
2. Identify the top model-provider combinations by total_tokens
3. For the top offenders, run token usage spikes to check for abnormal time windows
4. Use the cost-estimation template in "Most expensive prompts and models" to estimate USD spend — ask the user for per-model prices first (see "Exception — cost prices" under Agent Instructions)
5. Check for prompt-size outliers: high input_tokens / output_tokens ratio indicates large context windows
6. Attribute spend to a consumer (cost-and-tokens.md → "Usage attribution") and check whether prompt caching is enabled and effective (cost-and-tokens.md → "Prompt caching economics")1. Run token usage spikes (cost-and-tokens.md → "Token usage spikes") to find the abnormal time window
2. Within that window, run repeated-tool-calls and runaway-turn queries (agent-signals.md → "Agent loops and runaway detection")
3. For a flagged trace.id, open the trace to see what the agent looped on
4. If conversation ids are present, re-run the loop query grouped by gen_ai.conversation.id to catch cross-turn loops (conversation-analytics.md)1. Run failing agent activity query (agent-signals.md → "Failing agent activity")
2. Sort by errors desc to find the most error-prone agent
3. Take the trace.id from a failing span and open in Dynatrace distributed-tracing view
4. Check agent steps by latency (agent-signals.md) to see which operation type is slowest1. Run the guardrails presence check (guardrails.md) to confirm finish reasons are recorded
2. Run the finish-reason breakdown, then the blocked (content_filter) and truncated (length) queries
3. Correlate safety-filter spikes with the prompt-injection evaluator and truncation with answer-completeness failures (evaluations.md)
4. For a specific block or failure, take the trace.id and pivot to the originating spans (evaluations.md → "Correlating evaluations to traces")1. Run evaluation quality scores (evaluations.md → "Evaluation quality scores") to rank evaluators by avg_score asc
2. Focus on the lowest-scoring evaluator
3. Run failed evaluations (evaluations.md → "Failed evaluations") to surface the exact Q&A pairs and LLM judge explanations
4. Use the "Fail rate by evaluator" query to see how many responses fail each evaluator and the share of total evaluations
5. To root-cause a specific failure, take its trace.id and pivot to the originating spans (evaluations.md → "Correlating evaluations to traces")gen_ai.conversation.id)© Dynatrace, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 7 other files (references) in skills/dt-obs-genai of Dynatrace/dynatrace-for-ai.
Open the folder on GitHubat commit 4f9aa71
Dt Obs Genai next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Dt Obs Genai this skillDynatrace/dynatrace-for-ai | 163 | — | ~4.5k | Automated safety check: Pass | Apache-2.0 | |
| Evalagentevals-dev/agentevals | 163 | — | ~904 | Automated safety check: Pass | Apache-2.0 | |
| Agent Kill Switchvivekchand/clawmetry | 426 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Clawmetry Selfcheckvivekchand/clawmetry | 426 | — | ~515 | Automated safety check: Pass | MIT | |
| Deploying Scalable Agentsmicrosoft/ai-agents-for-beginners | 77k | — | ~1.5k | Automated safety check: Pass | MIT | |
| Deploying Scalable Agentsmicrosoft/ai-agents-for-beginners | 77k | — | ~1.6k | Automated safety check: Pass | MIT |
agentevals-dev/agentevals
Evaluate and score agent behavior against a golden reference.
vivekchand/clawmetry
Give the human an off switch and a cost meter for the coding agents on this machine, using ClawMetry.
vivekchand/clawmetry
Read your own agent telemetry from ClawMetry (waste, progress, cost) and act on it before finishing a task.
microsoft/ai-agents-for-beginners
Take a working agent prototype to a scalable, observable production deployment on Microsoft Foundry.
microsoft/ai-agents-for-beginners
Take one working agent prototype go scalable, observable production deployment for Microsoft Foundry.
DenisSergeevitch/agents-best-practices
A skill your agent uses when designing, generating an MVP blueprint for, auditing, troubleshooting, refactoring, or explaining an agentic harness for any domain.
Dynatrace/dynatrace-for-ai
Analyze dashboards and notebooks using Davis analyzers — anomaly detection, novelty scoring, and correlation.
Dynatrace/dynatrace-for-ai
Set up the Dynatrace iOS SDK (OneAgent) in an iOS project using Swift Package Manager.
Dynatrace/dynatrace-for-ai
End-to-end Dynatrace alerting lifecycle — anomaly detector setup and model selection (static threshold, adaptive baseline, seasonal baseline), alert event storage in Grail, problem grouping and…
Dynatrace/dynatrace-for-ai
AWS cloud resource monitoring including EC2, RDS, Lambda, ECS/EKS, VPC networking, load balancers, S3, DynamoDB, SQS/SNS, and cost optimization.
Dynatrace/dynatrace-for-ai
3rd-party test and monitor result ingestion into Dynatrace Grail via the platform events ingest API (platform/ingest/custom/events/).
Dynatrace/dynatrace-for-ai
DAVIS problem analysis including root cause identification, impact assessment, and correlation with other telemetry.
Works with
Categories
Analyze & debug GenAI/LLM apps: token cost & caching by prompt, model & provider; latency/errors; agent & tool loops/failures; conversations; guardrails; evaluations; OpenTelemetry/dt-evals setup. Dt Obs Genai is an agent skill from Dynatrace/dynatrace-for-ai. Analyze & debug GenAI/LLM apps: token cost & caching by prompt, model & provider; latency/errors; agent & tool loops/failures; conversations; guardrails; evaluations; OpenTelemetry/dt-evals setup.
Dt Obs Genai fits situations like: tasks that involve Observability; tasks that involve LLM cost and token optimization; tasks that involve LLM evaluation.
Run `npx skills add Dynatrace/dynatrace-for-ai --skill dt-obs-genai -a claude-code`. Or copy the skill folder (skills/dt-obs-genai in Dynatrace/dynatrace-for-ai) into .claude/skills/dt-obs-genai in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Dynatrace/dynatrace-for-ai --skill dt-obs-genai -a codex`. Or copy the skill folder (skills/dt-obs-genai in Dynatrace/dynatrace-for-ai) into .agents/skills/dt-obs-genai in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Dynatrace/dynatrace-for-ai --skill dt-obs-genai -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dt-obs-genai, .gemini/skills/dt-obs-genai, .github/skills/dt-obs-genai and .opencode/skills/dt-obs-genai in your project.
SKILL.md names no scripts, command-line tools or credentials: Dt Obs Genai is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Dt Obs Genai is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 16k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Dt Obs Genai: Eval (agentevals-dev/agentevals, 163 stars), Agent Kill Switch (vivekchand/clawmetry, 426 stars), Clawmetry Selfcheck (vivekchand/clawmetry, 426 stars) and Deploying Scalable Agents (microsoft/ai-agents-for-beginners, 77k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Dynatrace (a GitHub organization) maintains it in Dynatrace/dynatrace-for-ai, which has 163 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 1, 2026.
Source: Dynatrace/dynatrace-for-ai on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.