Official agent skill

GPU Live Metric Validation

by DataDog in DataDog/datadog-agent

Validate live GPU metrics on clusters running the Agent version under test and investigate missing metrics or tag failures.

OfficialApache-2.0Auto-check passedDevOps & Cloud

Install GPU Live Metric Validation

skills CLI
$ npx skills add DataDog/datadog-agent --skill gpu-live-metric-validation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install DataDog/datadog-agent gpu-live-metric-validation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/DataDog/datadog-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/gpu-live-metric-validation .claude/skills/gpu-live-metric-validation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gpu-live-metric-validation
GitHub stars
3.8k
Token cost
~1.1k tokens
SKILL.md length
468 words
Files
1
Skills in repo
35
Repo updated
First seen
Licence
Apache-2.0

At a glance

Validate live GPU metrics on clusters running the Agent version under test and investigate missing metrics or tag failures.

  • Works in 3 steps: Ask the user which supported orgs to… → Run the validator for each selected org → Record the selected org, Agent version…
  • Asked to validate GPU metrics on live clusters
  • SKILL.md covers Purpose, Required input, Validation workflow and Investigating findings, plus 2 more sections
  • Needs DD_API_KEY and DD_APP_KEY

What it does

GPU Live Metric Validation is an agent skill from DataDog/datadog-agent, published by the product's own GitHub organization. Validate live GPU metrics on clusters running the Agent version under test and investigate missing metrics or tag failures. Use when asked to validate GPU metrics on live clusters, check a GPU Agent release, or run dda inv gpu.validate-metrics.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud. It works with Datadog. The repository describes itself as: Main repository for Datadog Agent. The licence is Apache-2.0.

When your agent uses it

  • Asked to validate GPU metrics on live clusters
  • Check a GPU Agent release
  • Run dda inv gpu.validate-metrics

Example prompts

  • “/gpu-live-metric-validation”

Requirements

  • A credential in DD_API_KEY
  • A credential in DD_APP_KEY

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Ask the user which supported orgs to validate.
  2. Run the validator for each selected org
  3. Record the selected org, Agent version wildcard, lookback window, and

What it can do on your machine

Read from SKILL.md and the folder at commit a706f1a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DD_API_KEY
    • DD_APP_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

GPU Live Metric Validation loads about 1.1k tokens when it runs. Until then it costs about 68 tokens; SKILL.md has 468 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~68
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from DataDog/datadog-agent at commit a706f1a, republished under its Apache-2.0 licence (© DataDog). 468 words, ~1,077 tokens.

Download SKILL.mdSave it as .claude/skills/gpu-live-metric-validation/SKILL.md (or your agent's skills folder).
name
gpu-live-metric-validation
description
Validate live GPU metrics on clusters running the Agent version under test and investigate missing metrics or tag failures. Use when asked to validate GPU metrics on live clusters, check a GPU Agent release, or run dda inv gpu.validate-metrics.
<!-- @format -->

GPU Live Metric Validation

Purpose

Run GPU metric validation against Kubernetes clusters that are exclusively running the Agent version under test during the selected validation window, then investigate any findings.

Required input

Ask the user which Datadog orgs to validate before running any queries. The supported values match tasks/gpu.py:

  • prod (app.datadoghq.com)
  • staging (ddstaging.datadoghq.com)

Validation workflow

  1. Ask the user which supported orgs to validate.

  2. Run the validator for each selected org:

    bash
    dda inv gpu.validate-metrics \
      --org <prod-or-staging> \
      --lookback-seconds <window-seconds>

    By default, the task derives the Agent image-tag wildcard from the current release branch and its latest release-candidate tag. It passes that wildcard to the validator, which queries datadog.agent.running grouped by kube_cluster_name,image_tag, selects clusters whose nonempty image tags all match the wildcard, and ANDs that cluster selection with the GPU configuration filters.

    Use --agent-version <wildcard> to override the derived version; it is required outside release branches (N.N.x), where derivation fails. Use --metric-filter <filter> only for an additional scope; it is ANDed with the version-derived cluster filter.

  3. Record the selected org, Agent version wildcard, lookback window, and validation output before interpreting any findings.

Investigating findings

For follow-up Datadog queries, use dd-auth to select the target. pup consumes the injected DD_API_KEY, DD_APP_KEY, and DD_SITE; do not combine this workflow with pup --org.

  1. Start from a failing GPU metric and group it by the smallest useful set of dimensions, normally gpu_uuid, host, and kube_cluster_name. Add the failed tag as a group-by dimension when investigating a tag failure.

    bash
    dd-auth --domain <target-domain> -- \
      pup --no-agent metrics query \
      --query='count:gpu.<metric>{<validation-filter>} by {gpu_uuid,host,kube_cluster_name,<failed-tag>}' \
      --from=<validation-window> --to=now
  2. Identify patterns before drawing conclusions: whether failures are limited to a GPU architecture, device mode, GPU model, host, cluster, workload, or Agent image tag. Confirm the affected cluster's Agent image tag with:

    bash
    dd-auth --domain <target-domain> -- \
      pup --no-agent metrics query \
      --query='sum:datadog.agent.running{<cluster-filter>} by {kube_cluster_name,image_tag}' \
      --from=<validation-window> --to=now
  3. Preserve the exact queries and affected GPU/host/cluster identifiers in the summary.

Show full SKILL.md (185 more words)Show less

Candidate investigations by failure type

  • Missing metric: compare the affected GPU configuration with the metric's expected support and query gpu.device.total for the same scope to confirm that devices were present.
  • Missing required tag: group the failing metric by the required tag and affected GPU/host/cluster. Compare a related GPU metric to determine whether the absence is metric-specific.
  • Invalid tag value: group by the invalid tag and affected GPU/host/cluster. Check whether a series has multiple values for the same tag key before treating a comma-separated value as a single emitted tag value.
  • Unknown or extra tag: query a workload/container metric such as container.cpu.usage for the same pod and container. Compare its tags with the GPU metric to determine whether the tag originates from the workload.
  • If a workload tag points to a Kubernetes resource, inspect the corresponding Kubernetes object and its labels and annotations before assigning a source.

Constraints

  • Use an explicit, bounded validation window.
  • Use dd-auth to select the target for pup queries.
  • Do not validate clusters that reported a different nonempty Agent image tag in the validation window.
  • Do not use pup --org with dd-auth.

© DataDog, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/gpu-live-metric-validation of DataDog/datadog-agent.

Open the folder on GitHubat commit a706f1a

Compare with similar skills

GPU Live Metric Validation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

GPU Live Metric Validation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
GPU Live Metric Validation this skillDataDog/datadog-agent3.8k—~1.1kAutomated safety check: PassApache-2.0
Apm IntegrationsDataDog/dd-trace-js837—~3kAutomated safety check: PassCustom licence
Dd IdpDataDog/pup1k—~2kAutomated safety check: PassApache-2.0
Datadog Data Source GeneratorDataDog/terraform-provider-datadog468—~2.7kAutomated safety check: PassMPL-2.0
Apm IntegrationsDataDog/dd-trace-java736—~3.7kAutomated safety check: NotesApache-2.0
Azure FunctionsDataDog/dd-trace-dotnet573—~4.7kAutomated safety check: PassApache-2.0

Similar skills

  • Apm Integrations

    DataDog/dd-trace-js

    Official

    A skill your agent uses when adding, debugging, fixing, or modifying instrumentation and plugins for third-party libraries in dd-trace-js.

    837 GitHub stars~3k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Dd Idp

    DataDog/pup

    Official

    Find, filter, count, and connect software, teams, engineering work and delivery, infrastructure, and operational or security records through Pup's read-only Datadog entity graph.

    1k GitHub stars~2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Datadog Data Source Generator

    DataDog/terraform-provider-datadog

    Official

    Generates a Datadog Terraform provider data source from an OpenAPI operation with tfgen and opens a review-ready GitHub PR with a risk scan and testing guide.

    468 GitHub stars~2.7k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Apm Integrations

    DataDog/dd-trace-java

    Official

    Write a new library instrumentation end-to-end. An agent skill from DataDog/dd-trace-java.

    736 GitHub stars~3.7k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Azure Functions

    DataDog/dd-trace-dotnet

    Official

    Dev/test workflow for tracer engineers working on the Datadog .NET tracer — build a local Datadog.AzureFunctions NuGet package, deploy it to a test Azure Function App, trigger it, and analyze…

    573 GitHub stars~4.7k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Tool Connector

    ZhixiangLuo/10xProductivity

    Connect any tool you use at work to your agent — including internal company tools, custom-built systems, deployment portals, incident trackers, internal knowledge bases, HR systems, and commercial…

    478 GitHub stars~925 tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed

More from DataDog/datadog-agent

All 35 skills in this repo
  • Triage CI Failure

    DataDog/datadog-agent

    Official

    Classify a failed CI as either caused by an active incident, flakiness, or a true code regression.

    3.8k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Elicit

    DataDog/datadog-agent

    Official

    Run a structured discovery session to build an Allium specification through conversation.

    3.8k GitHub starsUsed in 1 repo~3.7k tokens
    Auto-check passed
  • Follow PR

    DataDog/datadog-agent

    Official

    Monitor the current PR's GitLab pipeline to completion, then report success, auto-fix, or investigate a failure.

    3.8k GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Create Epic Recap

    DataDog/datadog-agent

    Official

    A skill your agent uses when an engineer or manager asks to recap, summarize, or post an update on a Jira Epic — a progress update for an in-progress Epic (how far along it is, what's shipped so…

    3.8k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Explain Lading Config

    DataDog/datadog-agent

    Official

    Explains a lading.yaml config file from the regression test suite, using the lading Rust source as ground truth for field meanings and defaults.

    3.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Distill

    DataDog/datadog-agent

    Official

    Extract an Allium specification from an existing codebase. An agent skill from DataDog/datadog-agent.

    3.8k GitHub starsUsed in 1 repo~7k tokens
    Auto-check passed

Works with

Categories

Questions about GPU Live Metric Validation

What does GPU Live Metric Validation do?

Validate live GPU metrics on clusters running the Agent version under test and investigate missing metrics or tag failures. GPU Live Metric Validation is an agent skill from DataDog/datadog-agent, published by the product's own GitHub organization. Validate live GPU metrics on clusters running the Agent version under test and investigate missing metrics or tag failures.

When should I use GPU Live Metric Validation?

GPU Live Metric Validation fits situations like: asked to validate GPU metrics on live clusters; check a GPU Agent release; run dda inv gpu.validate-metrics.

How do I install GPU Live Metric Validation in Claude Code?

Run `npx skills add DataDog/datadog-agent --skill gpu-live-metric-validation -a claude-code`. Or copy the skill folder (.agents/skills/gpu-live-metric-validation in DataDog/datadog-agent) into .claude/skills/gpu-live-metric-validation in your project. Claude Code loads it when a task matches its description.

How do I install GPU Live Metric Validation in Codex?

Run `npx skills add DataDog/datadog-agent --skill gpu-live-metric-validation -a codex`. Or copy the skill folder (.agents/skills/gpu-live-metric-validation in DataDog/datadog-agent) into .agents/skills/gpu-live-metric-validation in your project. Codex loads it when a task matches its description.

Can I use GPU Live Metric Validation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add DataDog/datadog-agent --skill gpu-live-metric-validation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gpu-live-metric-validation, .gemini/skills/gpu-live-metric-validation, .github/skills/gpu-live-metric-validation and .opencode/skills/gpu-live-metric-validation in your project.

What does GPU Live Metric Validation need to run?

Going by SKILL.md and its folder, GPU Live Metric Validation needs credentials named DD_API_KEY and DD_APP_KEY. Our summary lists: A credential in DD_API_KEY; A credential in DD_APP_KEY.

Does GPU Live Metric Validation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is GPU Live Metric Validation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does GPU Live Metric Validation use?

GPU Live Metric Validation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does GPU Live Metric Validation use?

About 1.1k tokens (SKILL.md is roughly 4.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to GPU Live Metric Validation?

Skills that share tags, products or a category with GPU Live Metric Validation: Apm Integrations (DataDog/dd-trace-js, 837 stars), Dd Idp (DataDog/pup, 1k stars), Datadog Data Source Generator (DataDog/terraform-provider-datadog, 468 stars) and Apm Integrations (DataDog/dd-trace-java, 736 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains GPU Live Metric Validation?

DataDog (a GitHub organization, an official publisher) maintains it in DataDog/datadog-agent, which has 3,759 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 9, 2026.

Source: DataDog/datadog-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.