Official agent skill

K6 Trend Analysis

by grafana in grafana/skills

Analyze Grafana Cloud k6 test run trends over time. An agent skill from grafana/skills.

OfficialApache-2.0Auto-check passedDevOps & Cloud

Install K6 Trend Analysis

skills CLI
$ npx skills add grafana/skills --skill k6-trend-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install grafana/skills k6-trend-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/grafana/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/grafana-k6/k6-trend-analysis .claude/skills/k6-trend-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
k6-trend-analysis
GitHub stars
282
Token cost
~4.8k tokens
SKILL.md length
2,128 words
Files
1
Skills in repo
51
Repo updated
First seen
Licence
Apache-2.0

At a glance

Analyze Grafana Cloud k6 test run trends over time. An agent skill from grafana/skills.

  • Works in 10 steps: Identify the test → Determine the analysis window → Fetch runs in the window → …
  • The user asks about test performance trends
  • SKILL.md covers What this skill does NOT do, Dependencies, Workflow and Report template, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

K6 Trend Analysis is an agent skill from grafana/skills, published by the product's own GitHub organization. Analyze Grafana Cloud k6 test run trends over time. Detects slow metric drift (e.g., P95 latency creeping up while still passing thresholds), computes headroom to thresholds, flags anomalies, and recommends threshold tightening. Use when the user asks about test performance trends, wants to know if metrics are degrading, asks whether thresholds should be tightened, or wants a health check across recent runs for a specific test. Trigger on phrases like "how is my test trending", "is P95 getting worse", "check for…

Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Load testing, Monitoring and alerting and Forecasting and time series. It works with Grafana. The licence is Apache-2.0.

When your agent uses it

  • The user asks about test performance trends
  • Wants to know if metrics are degrading
  • Asks whether thresholds should be tightened
  • Wants a health check across recent runs for a specific test

Example prompts

  • “how is my test trending”
  • “is P95 getting worse”
  • “check for performance regression”
  • “/k6-trend-analysis”

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. Identify the test
  2. Determine the analysis window
  3. Fetch runs in the window
  4. Fetch all metrics and their types
  5. Fetch per-run aggregate values
  6. Extract thresholds
  7. Compute trends
  8. Headroom analysis
  9. Service-side correlation (when warranted)
  10. Present the report

What it can do on your machine

Read from SKILL.md and the folder at commit 1ccacf2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

K6 Trend Analysis loads about 4.8k tokens when it runs. Until then it costs about 204 tokens; SKILL.md has 2,128 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~204
When it runs · the whole SKILL.md, loaded when a task matches
~4.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from grafana/skills at commit 1ccacf2, republished under its Apache-2.0 licence (© grafana). 2,128 words, ~4,774 tokens.

Download SKILL.mdSave it as .claude/skills/k6-trend-analysis/SKILL.md (or your agent's skills folder).
name
k6-trend-analysis
description
Analyze Grafana Cloud k6 test run trends over time. Detects slow metric drift (e.g., P95 latency creeping up while still passing thresholds), computes headroom to thresholds, flags anomalies, and recommends threshold tightening. Use when the user asks about test performance trends, wants to know if metrics are degrading, asks whether thresholds should be tightened, or wants a health check across recent runs for a specific test. Trigger on phrases like "how is my test trending", "is P95 getting worse", "check for performance regression", "should I tighten thresholds", "are my tests degrading", "show me trends for test X", "analyze my k6 test runs", or "is my test getting slower". Also trigger when a user asks to check all tests in a project -- run this skill once per test and synthesize.

k6 Trend Analysis

Analyze metric trends across multiple runs of a Grafana Cloud k6 test to catch degradation early -- before thresholds breach and alerts fire. A P95 at 380ms against a 500ms threshold is "green" today, but if it was 250ms a month ago, something is quietly degrading; this skill surfaces that drift and recommends action.

What this skill does NOT do

  • Deep-dive into a single run's failure: load k6-cloud-investigate-test
  • Edit scripts or apply threshold changes: load k6-test-maintenance
  • Create new test scripts: use the appropriate test creation workflow
  • Query service-side metrics directly: this skill hands off to debug-with-grafana when observability correlation is needed

Dependencies

This skill delegates all GCk6 API mechanics to k6-manage. Read it before executing any API call -- it covers auth, path construction (the doubled cloud/cloud/ prefix), pagination with @nextLink, the spill envelope, and metric query syntax. Do not duplicate that knowledge here.

Tools used: gcx (via k6-manage patterns).


Workflow

Follow these steps in order. Present findings at the end -- do not apply changes.

Step 1: Identify the test

The user provides one of:

  • A GCk6 URL (extract the load test ID from the path)
  • A test ID directly
  • A test name (search via gcx k6 tests list or the v6 API)

Confirm the test exists by fetching its metadata. Record the id, name, project_id, and created timestamp. You will need the load test ID (not a run ID) for the multi-run metric endpoints.

Step 2: Determine the analysis window

Default to last 30 days. Adjust if:

  • The user requests a specific window
  • The test has very few runs (<5 in 30 days) -- widen the window and note this
  • The test runs very frequently (hundreds of runs in 30 days) -- consider sampling or narrowing. Ask the user if the volume is extreme (>200 runs)

State the window explicitly: "Analyzing runs from {start_date} to {end_date}."

Step 3: Fetch runs in the window

List all runs for the test using the v6 API with $orderby=created desc pagination (see k6-manage Section 3 for the @nextLink loop pattern). Filter to runs within the analysis window.

For each run, record:

  • id, created, ended
  • result (passed/failed/timed_out)
  • status (created/queued/initializing/running/finished/aborted)
  • note (if present -- users sometimes annotate runs)

Discard runs that did not reach finished status (aborted, timed_out, etc.) unless the user specifically asks about them -- incomplete runs produce unreliable metric aggregates.

Count the runs. If fewer than 3 usable runs exist, inform the user that trend analysis is not meaningful with this sample size and suggest widening the window or waiting for more runs.

Step 4: Fetch all metrics and their types

Before querying values, discover what metrics the test emits. Use the multi-run metric listing endpoint (k6-manage references/metrics.md Section 2):

GET /cloud/v5/load_tests/{loadTestId}/metrics(test_run_ids=[{id1},{id2},...])

Pass a representative subset of run IDs (the first and last few) to catch metrics that may have been added or removed over the window. Record each metric's name and type (counter, gauge, trend, rate).

Group metrics by type -- the query method must match the metric type (see metrics.md "Query methods"). Using the wrong method returns empty results.

Step 5: Fetch per-run aggregate values

For each metric, query its aggregate value across all runs in the window using the multi-run aggregate endpoint (metrics.md Section 8):

GET /cloud/v5/load_tests/{loadTestId}/query_aggregate_k6(
  query='<method>',
  metric='<metric_name>',
  test_run_ids=[{id1},{id2},...]
)

Concrete gcx form (proxy prefix per k6-manage §2; -o json avoids the spill envelope):

bash
LT=<load_test_id>; IDS="123,124,125"
gcx --context <ctx> api "/api/plugins/k6-app/resources/cloud/cloud/v5/load_tests/$LT/query_aggregate_k6(query='histogram_quantile(0.95)',metric='http_req_duration',test_run_ids=[$IDS])" -o json

Choose the aggregate method based on metric type:

Metric typePrimary methodWhat it captures
trendhistogram_quantile(0.95)P95 latency -- the most common SLO target
trendhistogram_quantile(0.5)Median -- shows typical behavior
trendhistogram_avgMean -- sensitive to outliers
counterincreaseTotal count per run
rateratioSuccess/failure ratio
gaugemaxPeak value per run

For trend-type metrics (latencies), query multiple quantiles (P50, P90, P95, P99) to see if degradation is uniform or concentrated in the tail.

Always break down by request grouping for multi-target tests. A single aggregate p95 across a whole test obscures regressions confined to one endpoint or one page -- a 3x slowdown on one URL can be invisible at the test-level p95 if the test hits many URLs. Default to grouped queries:

Metric patternDefault groupingWhy
browser_web_vital_* (LCP, FCP, CLS, TTFB, INP, FID)by (url)Each navigation emits its own web vital; per-URL trends pin regressions to a specific page.
http_req_duration, http_req_failed, http_reqsby (name, status) if requests are tagged with name; otherwise by (url, status)Different endpoints have different baselines; mixing them hides per-endpoint drift.
iteration_duration (multi-scenario tests)by (scenario)Browser and protocol scenarios have very different durations; mixing them is meaningless.
Custom trends with tagsby (<the tag>)Whatever the user tagged on is presumably what they care about.

Discover available labels first via the labels endpoint (metrics.md §4) and label-values endpoint (metrics.md §5) before constructing the grouped query. Don't assume labels from reading the script -- a tag rename or refactor can silently shift the label space. Example:

# List labels for a specific metric on a representative run
GET /cloud/v5/test_runs/{id}/labels?match[]=browser_web_vital_lcp

# Then enumerate values for a useful label
GET /cloud/v5/test_runs/{id}/label/url/values

Only fall back to the bare aggregate (no by) for metrics where grouping adds no information -- e.g., vus, load_generator_cpu_percent, single-target tests where every request hits the same URL.

The response includes test_run_id as a label -- use this to map each value back to its run timestamp from Step 3. Grouped queries return one series per (group, run) combination; flatten into a tidy "rid x group" table for the trend computation in Step 7.

If there are too many run IDs to fit in a single URL (hundreds), batch the queries into groups of 50 run IDs and merge the results.

Step 6: Extract thresholds

Thresholds define what "passing" means. Fetch the test's current script (via k6-manage Section 5) and parse the export const options = { thresholds: {...} } block. Record each threshold's:

  • Metric name and selector (e.g., http_req_duration{name:homepage})
  • Condition (e.g., p(95)<500)
  • Whether abortOnFail is set

Also check the most recent run's threshold results from the run data to see which thresholds are currently passing vs. failing.

Not all metrics will have thresholds -- that's fine. Metrics without thresholds still get trend analysis; they just won't have headroom calculations.

Detect inflection points and rule out script changes deterministically. Before drawing conclusions about a regression, look for discontinuities in the metric values -- sudden jumps or drops that align across multiple metrics on the same date. When you spot one, do not guess whether the script changed. The run-bundled script endpoint gives a deterministic answer in seconds:

  1. sha256-diff the bundled scripts at the boundary. Fetch the snapshot from the last "before" run and the first "after" run via GET /cloud/v6/test_runs/{id}/script (k6-manage §5, "Two distinct script endpoints" -- use the run-scoped endpoint, not the load-test one, since the latter only shows the current version). Compare with shasum -a 256.

    bash
    gcx --context <ctx> api /api/plugins/k6-app/resources/cloud/cloud/v6/test_runs/<before_id>/script > /tmp/before.bin
    gcx --context <ctx> api /api/plugins/k6-app/resources/cloud/cloud/v6/test_runs/<after_id>/script > /tmp/after.bin
    shasum -a 256 /tmp/before.bin /tmp/after.bin

    For thoroughness on multi-run inflections, hash every run across the transition -- if N consecutive runs share one hash and N more share another, you have a clean before/after boundary. If hashes change mid-stream, the test was edited multiple times.

  2. If the sha256 differs, the test script changed -- the inflection may be a test-side artifact, not a service regression. Split the analysis into distinct eras at the boundary and compute trends within each era separately. Comparing metrics across script changes produces misleading trends -- a P95 drop from 3,000ms to 150ms is not an "improvement" if the script simply stopped hitting a slow endpoint. State the eras explicitly in the report and focus recommendations on the most recent era.

  3. If the sha256 matches, the script is byte-identical and the regression is external to the test (service-side, infrastructure, or load-zone). This is high-confidence information -- carry it into Step 9 (service-side correlation) instead of leaving "did the test change?" as an open question.

This sha256 diff is a 5-second deterministic check that rules out a huge class of causes. Run it at every detected inflection, not just when the user asks.

For each metric, build a time-ordered series of (run_timestamp, value) pairs. Then compute:

Basic statistics:

  • Mean, standard deviation, min, max across all runs
  • Current value (most recent run)
  • Baseline value (oldest run in window, or mean of first 3 runs for stability)

Trend direction: Split the runs into two halves (first half and second half of the time window). Compare the mean of each half:

  • Degrading: second-half mean is worse by >10% (higher for latency/errors, lower for success rates)
  • Improving: second-half mean is better by >10%
  • Stable: change is within 10%
  • Volatile: standard deviation exceeds 25% of the mean, regardless of direction

The 10% and 25% thresholds are starting points. If the user's test has very tight tolerances or very noisy metrics, adjust and explain the reasoning.

Rate of change: Express as percentage change per week: ((recent_mean - baseline_mean) / baseline_mean) * 100 / weeks_in_window. This normalizes across different window sizes.

Anomaly detection: Flag any run where the metric value is more than 2 standard deviations from the overall mean. These are potential inflection points worth investigating individually.

Show full SKILL.md (674 more words)Show less
Step 8: Headroom analysis

For metrics that have thresholds defined (from Step 6), compute headroom:

headroom_pct = ((threshold_value - current_value) / threshold_value) * 100

Classify headroom:

  • Comfortable (>50%): well within limits
  • Adequate (20-50%): healthy but worth monitoring
  • Thin (<20%): at risk of breaching if trends continue
  • Breached (<0%): already failing

For degrading metrics with thin headroom, estimate when the threshold will be breached if the current rate of change continues:

weeks_until_breach = headroom_absolute / rate_of_change_per_week

This is a rough projection, not a prediction -- present it as "at the current rate of degradation, this metric could breach in approximately N weeks."

Step 9: Service-side correlation (when warranted)

If Step 7 reveals degradation, offer (don't auto-run) to correlate with service-side data to separate service degradation (fix the service), test-environment changes (load-zone latency, LG exhaustion), and script changes (a slower edit). Present findings and ask first.

Hand off to debug-with-grafana (via gcx) to query the service's Prometheus metrics and Loki logs for the same window; look for service error-rate changes, upstream latency, resource pressure (CPU/memory/pools), and deploys coinciding with inflection points.

When the service has no observability (third-party or another team's service -- no Prometheus job, Loki stream, or probe), fall back to client-side signal, which still localises the regression:

  • Browser tests: compare Tempo iteration traces (k6-manage §7, /api/v1/tempo/api/search + /traces/{id}) for a before/after run. The span-name rollup, slowest spans, and per-URL navigation durations pin the regression to a URL, locator action, or asset; web vitals are web_vital.* span attributes -- read them directly.
  • Protocol tests: compare http_req_duration by (name, status) and http_req_failed by (name, status); a regression on one named request points at that endpoint. Split http_req_waiting (server time) vs http_req_receiving (transfer) to separate slow processing from slow download.
  • Payload corroboration (both): query data_received / browser_data_received. Latency up + payload up -> server content changes; latency up + payload stable -> server processing changes; intermittent payload -> flaky cache or A/B test.

Prefer server-side correlation when available -- it answers "why" directly; the fallbacks answer "where" and "what kind of change", and make a good ticket for the owning team.

Step 10: Present the report

Always present findings as a structured report. Never apply changes directly.


Report template

Use this structure. Omit sections that don't apply (e.g., skip "Anomalies" if none were detected).

markdown
# Trend Analysis: {test_name}

**Test ID**: {test_id}
**Analysis window**: {start_date} to {end_date} ({N} runs analyzed)
**Overall health**: {Healthy | Watch | Degrading | Critical}

## Run Summary

| Period | Runs | Passed | Failed | Pass Rate |
|--------|------|--------|--------|-----------|
| First half | N | N | N | N% |
| Second half | N | N | N | N% |

## Metric Trends

| Metric | Type | Current | Baseline | Change | Trend | Threshold | Headroom |
|--------|------|---------|----------|--------|-------|-----------|----------|
| http_req_duration (P95) | trend | 380ms | 250ms | +52% | Degrading | 500ms | 24% |
| http_req_failed | rate | 0.8% | 0.3% | +167% | Degrading | 1% | 20% |
| http_reqs | counter | 15,230 | 15,100 | +0.9% | Stable | - | - |

## Flagged Issues

### 1. {metric_name}: {classification}
- **Current**: {value} | **Baseline**: {value} | **Change**: {pct}%
- **Threshold**: {threshold} | **Headroom**: {pct}%
- **Rate of change**: {pct}% per week
- **Projected breach**: ~{N} weeks at current rate
- **Anomalous runs**: {run_ids with dates, if any}

## Threshold Recommendations

| Metric | Current Threshold | Recommended | Rationale |
|--------|-------------------|-------------|-----------|
| http_req_duration | p(95)<500 | p(95)<420 | Current P95 is 380ms; tightening to 420ms gives 10% headroom from current performance while surfacing further degradation early |

## Suggested Next Steps

- [ ] **Investigate service side**: P95 latency has increased 52% -- consider
      loading `debug-with-grafana` to check service health
- [ ] **Deep-dive run {run_id}**: anomalous P95 spike on {date} -- consider
      loading `k6-cloud-investigate-test` for this run
- [ ] **Tighten thresholds**: 2 metrics have >30% headroom that could be
      tightened -- consider loading `k6-test-maintenance` to apply changes

Overall health classification

Derive the overall health from the worst-case metric:

  • Healthy: all metrics stable or improving, headroom comfortable or adequate
  • Watch: at least one metric degrading but headroom still adequate
  • Degrading: at least one metric degrading with thin headroom
  • Critical: at least one metric has breached its threshold, or multiple metrics are degrading with thin headroom

Threshold recommendations

When recommending threshold changes, follow these principles:

  • Only tighten, never loosen unless the user asks. Loosening thresholds masks problems.
  • Target 10-20% headroom above the recent P95 of the metric. Enough room for normal variance but tight enough to catch real degradation.
  • Use the second-half mean as the baseline, not the single most recent run (which could be an outlier).
  • Respect the user's intent: if thresholds are currently very loose (>100% headroom), they may be intentionally permissive. Mention the opportunity to tighten but don't push hard -- the user knows their context.
  • Consider grouped thresholds: if the test uses tag-based thresholds (e.g., http_req_duration{name:homepage}), recommend per-endpoint thresholds where the trends differ between endpoints.

Gotchas

IssueDetail
Zero-observation thresholdsA threshold with zero observations passes by default in k6. If a metric appears to pass but has no data, flag it -- the threshold is not actually being evaluated.
Metric type changes across runsIf a metric's type changed between runs (e.g., script refactor), the multi-run aggregate endpoint uses the latest type. Earlier runs queried with the wrong method return empty. Flag this if detected.
Incomplete runs skew trendsAborted or timed-out runs typically have shorter durations and fewer iterations, producing unrepresentative metric values. Exclude them by default.
LG resource metricsload_generator_cpu_percent and load_generator_file_handles trending up may indicate the test is outgrowing its load generator allocation, not that the service is degrading. Call this out separately.
Rate metric directionFor ratio-type rate metrics (like check pass rates), "degrading" means the value is decreasing (fewer passes), which is the opposite direction from latency metrics.

© grafana, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/grafana-k6/k6-trend-analysis of grafana/skills.

Open the folder on GitHubat commit 1ccacf2

Compare with similar skills

K6 Trend Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

K6 Trend Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
K6 Trend Analysis this skillgrafana/skills282—~4.8kAutomated safety check: PassApache-2.0
DashboardingKilo-Org/kilo-marketplace190—~2.4kAutomated safety check: PassApache-2.0
Monitoring ExpertJeffallan/claude-skills12k—~1.6kAutomated safety check: PassMIT
Promql CLIsamber/cc-skills227—~1.9kAutomated safety check: PassMIT
Magic MouthHmbown/Wizards-of-the-Ghosts110—~836Automated safety check: PassCC0-1.0
Axiom Dashboard Builderopenclaw/clawhub9.5k—~4.9kAutomated safety check: PassMIT

Similar skills

  • Dashboarding

    Kilo-Org/kilo-marketplace

    Create, modify, and organise Grafana dashboards including panels, variables, transformations, and alerting.

    190 GitHub stars~2.4k tokensUpdated 12 days ago
    DevOps & CloudAuto-check passed
  • Monitoring Expert

    Jeffallan/claude-skills

    Sets up application monitoring: structured logs, Prometheus metrics, OpenTelemetry tracing, Grafana dashboards, alert rules and load tests with k6 or Artillery.

    12k GitHub stars~1.6k tokensUpdated 7 days ago
    DevOps & CloudAuto-check passed
  • Promql CLI

    samber/cc-skills

    CLI for querying Prometheus and PromQL-compatible engines (Thanos, Cortex, VictoriaMetrics, Grafana Mimir, Grafana Tempo...) — instant queries, range queries, metric discovery (metrics/labels/meta…

    227 GitHub stars~1.9k tokensUpdated 10 days ago
    DevOps & CloudAuto-check passed
  • Magic Mouth

    Hmbown/Wizards-of-the-Ghosts

    Magic Mouth is trigger → message. An agent skill from Hmbown/Wizards-of-the-Ghosts.

    110 GitHub stars~836 tokensUpdated 6 mo ago
    Backend & APIsAuto-check passed
  • Axiom Dashboard Builder

    openclaw/clawhub

    Designs and deploys Axiom dashboards through the API, choosing chart types and writing APL or metrics queries, with templates and migration notes for Splunk and Grafana.

    9.5k GitHub stars~4.9k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • QA Dashboard

    petrkindlmann/qa-skills

    Build and visualize QA dashboards and reports with Allure Report, Grafana, and ReportPortal.

    170 GitHub stars~4.7k tokensUpdated 4 mo ago
    DevOps & CloudAuto-check passed

More from grafana/skills

All 51 skills in this repo
  • K6 Docs

    grafana/skills

    Official

    Write or review k6 documentation across the three k6 repositories - k6-DefinitelyTyped (TypeScript types), k6-docs (user documentation), and k6 (release notes / changelog).

    282 GitHub stars~678 tokensUpdated 2 days ago
    Auto-check passed
  • Alerting Irm

    grafana/skills

    Official

    Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…

    282 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Dashboarding

    grafana/skills

    Official

    Build, modify, and ship Grafana dashboards as JSON via the HTTP API — panel types (timeseries / stat / gauge / table / heatmap / logs / traces / node-graph), gridPos 24-column layout, units…

    282 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • K6 Perf Test Website

    grafana/skills

    Official

    A skill your agent uses when the user wants to performance-test, load-test, or stress-test a public website end-to-end with k6.

    282 GitHub stars~3.3k tokensUpdated 2 days ago
    Auto-check passed
  • Promql

    grafana/skills

    Official

    Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics.

    282 GitHub starsUsed in 1 repo~1.1k tokens
    Auto-check passed
  • Adaptive Metrics

    grafana/skills

    Official

    Cut Grafana Cloud Metrics cost by shrinking active-series count with Adaptive Metrics aggregation rules — auto-recommendations from query history, custom exact/regex rules, label-drop config…

    282 GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed

Works with

Questions about K6 Trend Analysis

What does K6 Trend Analysis do?

Analyze Grafana Cloud k6 test run trends over time. An agent skill from grafana/skills. K6 Trend Analysis is an agent skill from grafana/skills, published by the product's own GitHub organization. Analyze Grafana Cloud k6 test run trends over time.

When should I use K6 Trend Analysis?

K6 Trend Analysis fits situations like: the user asks about test performance trends; wants to know if metrics are degrading; asks whether thresholds should be tightened; wants a health check across recent runs for a specific test.

How do I install K6 Trend Analysis in Claude Code?

Run `npx skills add grafana/skills --skill k6-trend-analysis -a claude-code`. Or copy the skill folder (skills/grafana-k6/k6-trend-analysis in grafana/skills) into .claude/skills/k6-trend-analysis in your project. Claude Code loads it when a task matches its description.

How do I install K6 Trend Analysis in Codex?

Run `npx skills add grafana/skills --skill k6-trend-analysis -a codex`. Or copy the skill folder (skills/grafana-k6/k6-trend-analysis in grafana/skills) into .agents/skills/k6-trend-analysis in your project. Codex loads it when a task matches its description.

Can I use K6 Trend Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add grafana/skills --skill k6-trend-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/k6-trend-analysis, .gemini/skills/k6-trend-analysis, .github/skills/k6-trend-analysis and .opencode/skills/k6-trend-analysis in your project.

What does K6 Trend Analysis need to run?

SKILL.md names no scripts, command-line tools or credentials: K6 Trend Analysis is instructions for the agent only.

Does K6 Trend Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is K6 Trend Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does K6 Trend Analysis use?

K6 Trend Analysis is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does K6 Trend Analysis use?

About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to K6 Trend Analysis?

Skills that share tags, products or a category with K6 Trend Analysis: Dashboarding (Kilo-Org/kilo-marketplace, 190 stars), Monitoring Expert (Jeffallan/claude-skills, 12k stars), Promql CLI (samber/cc-skills, 227 stars) and Magic Mouth (Hmbown/Wizards-of-the-Ghosts, 110 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains K6 Trend Analysis?

grafana (a GitHub organization, an official publisher) maintains it in grafana/skills, which has 282 GitHub stars. The repository holds 51 skills in this directory. The repository was last updated on October 8, 2026.

Source: grafana/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.