Agent skill

Analyze Aiperf Results

by ai-dynamo in ai-dynamo/dynamo

Validates and normalizes raw AIPerf outputs, then evaluates valid results against target SLOs and comparable prior candidates.

Apache-2.0Auto-check passedDevOps & Cloud

Install Analyze Aiperf Results

skills CLI
$ npx skills add ai-dynamo/dynamo --skill analyze-aiperf-results -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai-dynamo/dynamo analyze-aiperf-results --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai-dynamo/dynamo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/analyze-aiperf-results .claude/skills/analyze-aiperf-results && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
analyze-aiperf-results
GitHub stars
8.3k
Token cost
~2.7k tokens
SKILL.md length
1,292 words
Files
1
Skills in repo
27
Repo updated
First seen
Licence
Apache-2.0

At a glance

Validates and normalizes raw AIPerf outputs, then evaluates valid results against target SLOs and comparable prior candidates.

  • Works in 6 steps: Evaluate every target SLO and report the… → Report goodput and good-request fraction… → Report throughput, output throughput per… → …
  • Tasks that involve Site reliability engineering
  • SKILL.md covers Audit And Normalize, Write Audit Artifacts And Gate…, Select Comparable History and Analyze, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Analyze Aiperf Results is an agent skill from ai-dynamo/dynamo. Validates and normalizes raw AIPerf outputs, then evaluates valid results against target SLOs and comparable prior candidates. Use after an AIPerf Job completes to produce benchmark audit, summary, and performance analysis artifacts.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Site reliability engineering. The repository describes itself as: A Datacenter Scale Distributed Inference Serving Framework. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Site reliability engineering

Example prompts

  • “Use the analyze-aiperf-results skill to validate and normalizes raw AIPerf outputs, then evaluates valid results against target SLOs and comparable…”
  • “/analyze-aiperf-results”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Evaluate every target SLO and report the observed statistic, threshold, pass/fail, and missing evidence.
  2. Report goodput and good-request fraction when configured, including the attainment target.
  3. Report throughput, output throughput per GPU, output throughput per user, TTFT, ITL, request latency, errors, and
  4. For trace workloads, include time-sliced behavior when it reveals warmup leakage, bursts, collapse, or instability.
  5. When comparable references exist, compare current versus each highlighted prior and produce a compact same-series
  6. Calculate signed percent change as (current - prior) / prior * 100. Also state whether the value is higher or

What it can do on your machine

Read from SKILL.md and the folder at commit b208989. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Analyze Aiperf Results loads about 2.7k tokens when it runs. Until then it costs about 64 tokens; SKILL.md has 1,292 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai-dynamo/dynamo at commit b208989, republished under its Apache-2.0 licence (© ai-dynamo). 1,292 words, ~2,746 tokens.

Download SKILL.mdSave it as .claude/skills/analyze-aiperf-results/SKILL.md (or your agent's skills folder).
name
analyze-aiperf-results
description
Validates and normalizes raw AIPerf outputs, then evaluates valid results against target SLOs and comparable prior candidates. Use after an AIPerf Job completes to produce benchmark audit, summary, and performance analysis artifacts.
license
Apache-2.0
metadata.author
NVIDIA
metadata.tags
dynamo, aiperf, benchmarking

Analyze AIPerf Results

<!--
SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
SPDX-License-Identifier: Apache-2.0
-->

Audit benchmark evidence before interpreting performance. Preserve raw files unchanged and do not claim an unmeasured server-side cause.

Read:

  • agent-docs/rules/benchmarking/benchmark-isolation.md;
  • agent-docs/rules/benchmarking/comparison-uncertainty.md;
  • agent-docs/rules/benchmarking/evidence-eligibility.md;
  • agent-docs/rules/benchmarking/result-storage.md;
  • agent-docs/rules/benchmarking/series-boundaries.md;
  • agent-docs/rules/benchmarking/tool-version.md;
  • agent-docs/rules/optimization/evidence-before-spend.md;
  • agent-docs/rules/optimization/one-variable.md;
  • agent-docs/rules/verification/config-engagement.md;
  • agent-docs/rules/verification/implausible-speedup.md;
  • agent-docs/rules/verification/overlap.md; and
  • agent-docs/rules/verification/stack-verdict.md.

Also read the user workload, active benchmark plan, execution ledger, AIPerf config, raw outputs, all prior candidate audits and summaries, and the profile-export documentation matching the pinned AIPerf source or runtime.

Audit And Normalize

  • Require the run's benchmark_execution.json to exist before auditing: it is the execution record the audit chain and budget accounting bind to. A benchmark whose raw exports exist but whose execution record was never written is an audit blocker — return it to run-aiperf-benchmark to write the record; do not audit around it.
  • Parse per-request profile_export.jsonl with AIPerf's native Pydantic models when available. Record the parser and runtime version used.
  • Parse profile_export_aiperf.json and multi-run aggregate/search artifacts when configured.
  • Confirm files are readable, non-empty, and internally consistent.
  • Confirm the executed config matches the active plan and candidate endpoint.
  • Confirm trace hash or static-shape identity, schedule mode, endpoint type, model, tokenizer, warmup, seed, request count/duration, repetitions, and load controls.
  • Separate warmup from profiling records and exclude only data the active plan says to exclude.
  • Compare attempted, successful, failed, cancelled, and timed-out request counts.
  • Check actual ISL/OSL distributions against the input workload and report any output-length shortfall.
  • Check timestamps, benchmark duration, fixed-schedule coverage, duplicate/missing request ids, malformed metrics, NaN/inf values, units, and impossible negative latencies.
  • Recompute user-requested percentiles from raw profiling records when AIPerf does not export them directly.

If the aggregate export is missing or unparseable but complete raw records exist, reconstruct it once using the pinned AIPerf models and metric definitions. Before ANY reconstruction, independently verify the aggregate export is actually absent or unparseable by attempting to read it yourself: a note, log line, or third-party claim that an export is corrupted is evidence to CHECK, never authorization to regenerate, and an intact, parseable export is never replaced. Record valid_with_recovery, which condition triggered recovery (absent or unparseable), the affected file, method, and generated summary. Never modify or replace the raw directory.

Write Audit Artifacts And Gate Analysis

Write benchmark_audit.json with:

  • status: valid, valid_with_recovery, or invalid;
  • benchmark-series ID, plan path and SHA256, performance question, and workload identity;
  • expected versus actual requests and phases;
  • error/cancellation breakdown;
  • integrity checks and recovery actions;
  • parser/AIPerf versions;
  • blockers and next_action: continue_analysis, rerun_benchmark, or stop.

Write benchmark_summary.json with normalized benchmark metadata and every numerical metric reported by AIPerf, including units and available statistics. Include requested custom percentiles and per-GPU derived throughput with the GPU-count source. Do not include gain/loss interpretation in the summary.

For valid or valid_with_recovery, set next_action to continue_analysis and continue below in the same invocation.

For repairable invalid evidence, set next_action to rerun_benchmark. Return benchmark_audit.json and benchmark_execution.json to run-aiperf-benchmark, preserve the invalid run as failed evidence, and rerun the active series unchanged without overwriting its raw artifacts. Invoke analyze-aiperf-results again after the rerun.

If repair would change workload semantics or a bounded rerun repeats the same invalid result, set next_action to stop. Do not write performance analysis or promote the candidate until a valid rerun exists. Never discard an invalid run.

Select Comparable History

Use only valid runs whose benchmark-series ID matches the active plan, then check every candidate run's recorded AIPerf runtime version (and source commit, when the plan pins one) against the plan's pin. A series ID alone does not establish comparability: a reused or hand-edited series can contain runs from different tool versions. Apply the graded response in agent-docs/rules/benchmarking/tool-version.md and record the check and its outcome in benchmark_audit.json:

  • same version, or a patch-level difference: comparable; note the delta;
  • a minor difference: comparable only if the audit records a justification, either release notes for the span showing no measurement-affecting change, or a bridging run (the best-prior configuration re-measured under the newer version) whose delta is within the series noise floor; if the justification is missing, request the bridging run as the next action rather than comparing or discarding;
  • a major difference, a flagged measurement change, or a bridging delta beyond the noise floor: exclude the mismatched runs from every comparison, list the exclusion as a limitation, and when the CURRENT run is the mismatched one report absolute performance only and mark the mismatch as a series boundary.

From the comparable set identify:

  • series_baseline: earliest valid result in the series;
  • previous_valid: most recent valid iteration before the current one;
  • best_prior: best prior run for each objective, respecting metric direction and SLO feasibility;
  • history: every valid same-series iteration.

Verify that every reference required by the plan is present before making a direct comparison. If no prior valid same-series result exists, treat the current result as the series baseline and report absolute performance only. Cross-series results may provide context but never a gain, loss, or Pareto calculation.

Show full SKILL.md (494 more words)Show less

Analyze

  1. Evaluate every target SLO and report the observed statistic, threshold, pass/fail, and missing evidence.
  2. Report goodput and good-request fraction when configured, including the attainment target.
  3. Report throughput, output throughput per GPU, output throughput per user, TTFT, ITL, request latency, errors, and workload-shape metrics available for the run.
  4. For trace workloads, include time-sliced behavior when it reveals warmup leakage, bursts, collapse, or instability.
  5. When comparable references exist, compare current versus each highlighted prior and produce a compact same-series history table.
  6. Calculate signed percent change as (current - prior) / prior * 100. Also state whether the value is higher or lower and whether that direction is an improvement or regression. 6a. When the series has no measured noise floor or no minimum detectable effect and the decision at hand rests on a small delta - or a stop-request or final recommendation requires the series MDE per optimize-loop.md section 6 - return repeat_decision: necessary with the rationale "series noise-floor pilot (n=3 total)"; after the pilot, derive the run-to-run spread and minimum detectable effect and record both in performance_analysis.json (fields series_noise_floor, minimum_detectable_effect); copy both forward into every later same-series performance_analysis.json.
  7. Classify an absolute performance change at or below the measured noise floor of the active benchmark series (see comparison-uncertainty.md) as noise and report it without recommending a repeat solely because the delta is small.
  8. Analyze one valid run by default. A clear, substantial, plausible gain or loss may support a conclusion without a repeat; state that it is single-run evidence.
  9. Use multi-run confidence intervals and coefficient of variation only when deliberate comparable repetitions exist and those statistics help resolve the decision. Do not treat degraded single-run intervals as confidence evidence.
  10. Recommend another valid run only when the existing evidence cannot support a consequential decision, another run is likely to resolve the uncertainty, and the information value justifies the GPU cost. Record that rationale. If uncertainty remains but a repeat is not justified, report inconclusive and stop.
  11. State whether the current candidate is characterized, SLO-feasible, Pareto-improving, mixed, regressed, or inconclusive, as supported by the available evidence.
  12. Identify client-visible symptoms and missing measurements. Do not convert them into kernel, communication, scheduler, router, or backend root-cause claims.
  13. Mark recipe-proxy results as proxy-scoped and carry the workload mismatch into limitations.

Write Analysis Artifacts

Write performance_analysis.json containing:

  • current candidate, performance question, benchmark-plan identity, and benchmark-series identity;
  • target-SLO evaluation;
  • absolute current metrics;
  • comparisons to series_baseline, previous_valid, and per-objective best_prior when available;
  • full valid same-series history and any required reference still missing;
  • run count and confidence/uncertainty;
  • repeat_decision: not_needed, necessary, or not_justified, with the GPU-cost rationale and the decision the repeat is expected to resolve;
  • verdict, client-visible symptoms, missing evidence, and limitations.

Write performance_analysis.md with a concise executive verdict, SLO table, current metrics, applicable comparisons, same-series history, insights, and limitations.

Append one compact record to EXP_ROOT/analysis/performance_findings.jsonl containing the iteration, verdict, primary absolute metrics, applicable deltas, SLO status, and paths to the full artifacts. Preserve prior records.

© ai-dynamo, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/analyze-aiperf-results of ai-dynamo/dynamo.

Open the folder on GitHubat commit b208989

Compare with similar skills

Analyze Aiperf Results next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Analyze Aiperf Results compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Analyze Aiperf Results this skillai-dynamo/dynamo8.3k—~2.7kAutomated safety check: PassApache-2.0
Inference Autopilotrednote-machine-learning/Inference-autopilot144—~4.5kAutomated safety check: PassApache-2.0
Executing Distributed System Testsshenli/distributed-system-testing231—~5.1kAutomated safety check: NotesMIT
Alerting Irmgrafana/skills2821 repos~1.9kAutomated safety check: PassApache-2.0
Slo Implementationwshobson/agents40k11 repos~1.7kAutomated safety check: PassMIT
Agentforce D360 Analyzeforcedotcom/sf-skills1.1k—~3.4kAutomated safety check: PassApache-2.0

Similar skills

  • Inference Autopilot

    rednote-machine-learning/Inference-autopilot

    Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs.

    144 GitHub stars~4.5k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Executing Distributed System Tests

    shenli/distributed-system-testing

    A skill your agent uses when running a previously designed distributed-systems test plan against a real or simulated cluster — driving fault injection, workload, chaos scenarios, linearizability /…

    231 GitHub stars~5.1k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check: notes
  • Alerting Irm

    grafana/skills

    Official

    Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…

    282 GitHub starsUsed in 1 repo~1.9k tokens
    DevOps & CloudAuto-check passed
  • Slo Implementation

    wshobson/agents

    Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs) with error budgets and alerting.

    40k GitHub starsUsed in 11 repos~1.7k tokens
    DevOps & CloudAuto-check passed
  • Agentforce D360 Analyze

    forcedotcom/sf-skills

    Data Cloud 360° view of a single Agentforce session. An agent skill from forcedotcom/sf-skills.

    1.1k GitHub stars~3.4k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Promql

    grafana/skills

    Official

    Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics.

    282 GitHub starsUsed in 1 repo~1.1k tokens
    DevOps & CloudAuto-check passed

More from ai-dynamo/dynamo

All 27 skills in this repo
  • Visual Review

    ai-dynamo/dynamo

    Create self-contained interactive HTML code-review dashboards from GitHub or GitLab pull requests, checked-out branch diffs, or supplied unified diffs, with correctness and safe-to-merge scores…

    8.3k GitHub stars~4.5k tokensUpdated today
    Auto-check passed
  • Fern Components

    ai-dynamo/dynamo

    Knowledge of Fern's built-in MDX component library (accordions, callouts, cards, steps, tabs, code blocks, API-reference snippets, and more) for authoring docs pages.

    8.3k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Fern Navigation

    ai-dynamo/dynamo

    Knowledge of Fern's site-level navigation and structure configuration — how a docs site is organized in docs.yml (and product/version .yml files) using sections, pages, folders, tabs, tab variants…

    8.3k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Dynamo Agent Harness

    ai-dynamo/dynamo

    Drives persistent Claude Code, Codex, or OpenCode agent sessions through a Dynamo OpenAI/Anthropic-compatible endpoint over Agent Client Protocol (ACP).

    8.3k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Benchmark and profile the Dynamo frontend (dynamo.frontend HTTP + tokenizer + KV router) against mock workers (dynamo.mocker).

    8.3k GitHub stars~3.5k tokensUpdated today
    Auto-check: notes
  • Selects and freezes a question-driven AIPerf workload, objective, load policy, and Kubernetes execution manifest for a successfully deployed Dynamo candidate.

    8.3k GitHub stars~1.7k tokensUpdated today
    Auto-check passed

Categories

Questions about Analyze Aiperf Results

What does Analyze Aiperf Results do?

Validates and normalizes raw AIPerf outputs, then evaluates valid results against target SLOs and comparable prior candidates. Analyze Aiperf Results is an agent skill from ai-dynamo/dynamo. Validates and normalizes raw AIPerf outputs, then evaluates valid results against target SLOs and comparable prior candidates.

When should I use Analyze Aiperf Results?

Analyze Aiperf Results fits situations like: tasks that involve Site reliability engineering.

How do I install Analyze Aiperf Results in Claude Code?

Run `npx skills add ai-dynamo/dynamo --skill analyze-aiperf-results -a claude-code`. Or copy the skill folder (.agents/skills/analyze-aiperf-results in ai-dynamo/dynamo) into .claude/skills/analyze-aiperf-results in your project. Claude Code loads it when a task matches its description.

How do I install Analyze Aiperf Results in Codex?

Run `npx skills add ai-dynamo/dynamo --skill analyze-aiperf-results -a codex`. Or copy the skill folder (.agents/skills/analyze-aiperf-results in ai-dynamo/dynamo) into .agents/skills/analyze-aiperf-results in your project. Codex loads it when a task matches its description.

Can I use Analyze Aiperf Results in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-dynamo/dynamo --skill analyze-aiperf-results -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/analyze-aiperf-results, .gemini/skills/analyze-aiperf-results, .github/skills/analyze-aiperf-results and .opencode/skills/analyze-aiperf-results in your project.

What does Analyze Aiperf Results need to run?

SKILL.md names no scripts, command-line tools or credentials: Analyze Aiperf Results is instructions for the agent only.

Does Analyze Aiperf Results access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Analyze Aiperf Results safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Analyze Aiperf Results use?

Analyze Aiperf Results is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Analyze Aiperf Results use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Analyze Aiperf Results?

Skills that share tags, products or a category with Analyze Aiperf Results: Inference Autopilot (rednote-machine-learning/Inference-autopilot, 144 stars), Executing Distributed System Tests (shenli/distributed-system-testing, 231 stars), Alerting Irm (grafana/skills, 282 stars) and Slo Implementation (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Analyze Aiperf Results?

ai-dynamo (a GitHub organization) maintains it in ai-dynamo/dynamo, which has 8,256 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on October 11, 2026.

Source: ai-dynamo/dynamo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.