Agent skill

Testing Performance And Load

by jaktestowac in jaktestowac/awesome-copilot-for-testers

Designs and runs performance and load tests: workload modelling from real traffic, thresholds tied to SLOs, warmup and ramp shapes, percentile-based analysis, and lightweight CI perf checks with k6…

MITAuto-check passedDevOps & Cloud

Install Testing Performance And Load

skills CLI
$ npx skills add jaktestowac/awesome-copilot-for-testers --skill testing-performance-and-load -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaktestowac/awesome-copilot-for-testers testing-performance-and-load --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaktestowac/awesome-copilot-for-testers.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/testing-performance-and-load .claude/skills/testing-performance-and-load && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
testing-performance-and-load
GitHub stars
116
Token cost
~2.8k tokens
SKILL.md length
1,556 words
Files
5
Skills in repo
13
Repo updated
First seen
Licence
MIT

At a glance

Designs and runs performance and load tests: workload modelling from real traffic, thresholds tied to SLOs, warmup and ramp shapes, percentile-based analysis, and lightweight CI perf checks with k6…

  • Works in 7 steps: Name the question → Model the workload → Set thresholds from the SLO → …
  • A feature has latency
  • SKILL.md covers When to Use, Operating Principles, Workflow and Common Failure Modes, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Testing Performance And Load is an agent skill from jaktestowac/awesome-copilot-for-testers. Designs and runs performance and load tests: workload modelling from real traffic, thresholds tied to SLOs, warmup and ramp shapes, percentile-based analysis, and lightweight CI perf checks with k6 or Artillery. Use when a feature has latency or throughput requirements, when "it feels slow" needs to become a number, when a launch needs a capacity check, or when a performance result needs interpreting rather than just collecting.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `resources/k6-recipes.md`, `resources/perf-report-template.md` and `resources/thresholds-and-slos.md`).

It sits in DevOps & Cloud, covering Load testing and Site reliability engineering. The repository describes itself as: 👨💻 Instructions, prompts, and chat modes to help You with test automation for GitHub Copilot 🤖. The licence is MIT.

When your agent uses it

  • A feature has latency
  • Throughput requirements
  • It feels slow needs to become a number
  • A launch needs a capacity check

Example prompts

  • “it feels slow”
  • “Use the testing-performance-and-load skill to design and runs performance and load tests: workload modelling from real traffic, thresholds tied to…”
  • “/testing-performance-and-load”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Name the question
  2. Model the workload
  3. Set thresholds from the SLO
  4. Prepare the environment
  5. Warm up, then measure
  6. Analyze
  7. Report and act

What it can do on your machine

Read from SKILL.md and the folder at commit 8910672. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are javascript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Testing Performance And Load loads about 2.8k tokens when it runs. Until then it costs about 115 tokens; SKILL.md has 1,556 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~115
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaktestowac/awesome-copilot-for-testers at commit 8910672, republished under its MIT licence (© jaktestowac). 1,556 words, ~2,814 tokens.

Download SKILL.mdSave it as .claude/skills/testing-performance-and-load/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
testing-performance-and-load
description
Designs and runs performance and load tests: workload modelling from real traffic, thresholds tied to SLOs, warmup and ramp shapes, percentile-based analysis, and lightweight CI perf checks with k6 or Artillery. Use when a feature has latency or throughput requirements, when "it feels slow" needs to become a number, when a launch needs a capacity check, or when a performance result needs interpreting rather than just collecting.
argument-hint
System under test, the concern (latency, throughput, capacity, endurance), real traffic data if available, and the SLO or target if one exists
user-invocable
true

Testing Performance and Load

Use this skill when a system's speed or capacity is in question and the answer needs to be a number someone can act on.

Most performance testing produces numbers nobody uses. The two causes are always the same: a workload that does not resemble reality, and a result reported as an average. A test that hammers one endpoint with a flat 500 virtual users tells you how the system responds to something that will never happen, and a mean response time of 200ms is compatible with one user in twenty waiting four seconds.

When to Use

  • a feature has a stated latency or throughput requirement
  • "the app feels slow" needs to become a measurement
  • a launch, campaign, or migration needs a capacity check
  • a performance regression is suspected between two releases
  • a performance result exists and needs interpreting
  • a CI check should catch obvious regressions before they ship

Operating Principles

  • Model the workload before writing the script. Test shape comes from real traffic: which endpoints, in what ratio, with what think time, at what concurrency.
  • Percentiles, never averages. Report p50, p95, p99. The mean describes nobody's experience and hides the tail entirely.
  • A threshold without an SLO is a guess. Derive the number from what users need or what the business promised, and say which.
  • One variable per run. Comparing two runs that differ in load, data volume, and code changes tells you nothing about any of them.
  • Environment differences invalidate the number, not just weaken it. A result from a quarter-size environment is a shape, not a capacity.
  • Find the bottleneck, do not just report the symptom. "p95 is 3 seconds" is an observation; "p95 is 3 seconds because the product query has no index on category_id" is a finding.

Workflow

Phase 0: Name the question

Different questions need different tests. Pick one per run.

QuestionTest typeShape
Is it fast enough under normal load?Load testSteady state at expected concurrency
Where does it break?Stress testRamp until failure
Does it survive a sudden surge?Spike testStep change up, then down
Does it degrade over hours?Soak testSteady load for 4 to 24 hours
Did this release make it slower?Regression comparisonIdentical shape, two builds
How much capacity do we have?Capacity testRamp to the SLO breach point

A test that tries to answer all six answers none. If "it is slow" is all you have, start with a load test at expected concurrency and let the result narrow the question.

Phase 1: Model the workload

The phase that decides whether the result means anything. Work through ./resources/workload-model.md.

From real traffic where possible:

  • Endpoint mix: the actual ratio, from access logs or analytics. Reads usually outnumber writes by an order of magnitude, and synthetic tests usually get this wrong.
  • Concurrency: derived from throughput and think time, not guessed. concurrent users ≈ arrival rate × session duration.
  • Think time: real users pause. A test with zero think time is a different system entirely.
  • Data distribution: most requests hit a small hot set. A test with uniformly random ids has a cache hit rate nothing like production.
  • Session shape: users log in, browse, and act. A test that only calls the checkout endpoint skips the state that makes checkout slow.
  • Peak profile: daily peak, weekly peak, and the campaign or event shape if there is one.

When real traffic is unavailable, state the assumptions explicitly and mark the result as assumption-dependent. An assumed workload is workable; an unstated one is not.

Phase 2: Set thresholds from the SLO

Each threshold traces to something: a user need, a documented SLO, a competitor benchmark, or a current baseline you intend to hold.

./resources/thresholds-and-slos.md covers deriving them. The shape:

js
thresholds: {
  'http_req_duration{name:checkout}': ['p(95)<800', 'p(99)<2000'],
  'http_req_failed': ['rate<0.001'],
  'checkout_completed': ['count>0'],
}

Rules:

  • per-endpoint, not global. A global p95 mixes a health check with a report export and describes neither.
  • include an error rate threshold. A fast 500 is not a pass, and a load test with no error threshold will report one as success.
  • include at least one business-outcome counter. "Requests succeeded" is not the same as "orders were created", and a test that never completes a checkout can look perfect.
Phase 3: Prepare the environment

Record every difference from production before running anything. The differences bound what the result can claim.

  • Sizing: instance count, CPU, memory, relative to production
  • Data volume: a database with 2,000 rows and one with 4 million behave differently, and the difference is not linear
  • Data shape: long-lived accounts, large carts, deep histories
  • Caches: cold or warm, and warmed how
  • Dependencies: real, stubbed, or rate-limited. A stubbed dependency that responds in 1ms removes the queueing that production has.
  • Network: the load generator's location relative to the system
  • Autoscaling: enabled or fixed, and with what policy

Then decide where the load generator runs. A generator on the same machine as the system under test competes with it for CPU, and the result measures the contention.

Phase 4: Warm up, then measure
  • Warmup: run at low load until JIT compilation, connection pools, and caches settle. Exclude the warmup from the results, and say how long it was.
  • Ramp: reach target load over minutes, not instantly, unless the spike is the subject.
  • Steady state: long enough for the metric to stabilize. Ten minutes is usually the minimum for anything meaningful.
  • Ramp-down: watch for errors during the descent; connection pool problems often surface here.

Monitor the system, not only the client. Client-side response times tell you something is slow; server CPU, memory, database connections, queue depth, and GC pauses tell you why. A run with no server-side observation produces a symptom and no cause.

Show full SKILL.md (618 more words)Show less
Phase 5: Analyze

Report distributions.

  • p50 is the typical experience
  • p95 is the experience of a noticeable minority
  • p99 is where the timeouts and abandonments live
  • max is one data point, and it is usually a GC pause or a cold start

Look for the shapes in ./resources/perf-report-template.md:

  • the knee: throughput plateaus while response time climbs; that is the capacity limit
  • the sawtooth: periodic latency spikes, usually GC or a cache expiry
  • the drift: response time climbing over a soak; usually a leak or unbounded growth
  • the cliff: sudden failure past a threshold, usually a pool or a connection limit

Correlate every latency feature with a server-side metric before naming a cause. A hypothesis without a correlating metric is a guess, and performance work built on a guessed bottleneck is expensive.

Phase 6: Report and act

Use ./resources/perf-report-template.md. It requires:

  • the workload model, so the reader can judge whether it resembles reality
  • environment differences from production, stated up front
  • results as percentiles per endpoint, with the error rate
  • the bottleneck, with the evidence for it
  • what the result does not tell you
  • a recommendation

Add a lightweight CI check when there is something worth protecting: a short run at modest load, thresholds set generously enough not to be flaky, and a comparison against the last release. Recipes in ./resources/k6-recipes.md. Its job is catching an order-of-magnitude regression, not measuring capacity.

Common Failure Modes

  • reporting the average, which describes nobody and hides the tail
  • a flat load with no think time and no ramp, which no real traffic resembles
  • uniformly random test data, giving a cache behaviour production never has
  • testing one endpoint in isolation when the slowness comes from contention between several
  • no error rate threshold, so a run that returned 500s at speed is reported as a pass
  • a test that never completes a business transaction while every request succeeds
  • results from an environment a quarter the size, presented as capacity
  • no warmup, so the first minute's cold-start latency poisons the percentiles
  • naming a bottleneck from client-side numbers alone
  • a CI perf gate so tight it flakes, then gets disabled and never re-enabled

Resource Map

  • ./resources/workload-model.md - deriving the endpoint mix, concurrency, think time, and data distribution from real traffic, with a worked example
  • ./resources/thresholds-and-slos.md - deriving thresholds from SLOs, percentile choice, error budgets, and per-endpoint targets
  • ./resources/k6-recipes.md - k6 scripts for load, stress, spike, and soak shapes, plus the CI regression check and an Artillery equivalent
  • ./resources/perf-report-template.md - report structure, the four curve shapes and what each means, and a worked example
  • assessing-release-readiness - when a performance result feeds a go/no-go decision
  • analyzing-quality-metrics - when performance should be trended across releases rather than measured once
  • testing-api-contracts - when the endpoints under load also need shape verification
  • designing-test-data - when the load test needs a realistic hot set and data distribution
  • handling-sensitive-test-data - when load test data is derived from production traffic
  • automating-ci-test-pipelines (planned) - for running the CI regression check and storing its history
  • tech-debt-analysis - when the bottleneck is structural rather than a single query

Definition of Done

This skill is complete when:

  • the question is one of load, stress, spike, soak, regression, or capacity, and only one
  • the workload model states endpoint mix, concurrency, think time, and data distribution, with its source or its assumptions
  • thresholds are per-endpoint and each traces to an SLO, a user need, or a stated baseline
  • an error rate threshold and at least one business-outcome counter are included
  • environment differences from production are recorded before the run
  • warmup is excluded from the results and its duration stated
  • results are reported as percentiles, never as averages
  • every latency finding is correlated with a server-side metric before a cause is named
  • the report says what the result does not tell you

© jaktestowac, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in skills/testing-performance-and-load of jaktestowac/awesome-copilot-for-testers.

  • SKILL.md
  • resources/k6-recipes.md
  • resources/perf-report-template.md
  • resources/thresholds-and-slos.md
  • resources/workload-model.md

Open the folder on GitHubat commit 8910672

Compare with similar skills

Testing Performance And Load next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Testing Performance And Load compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Testing Performance And Load this skilljaktestowac/awesome-copilot-for-testers116—~2.8kAutomated safety check: PassMIT
Synthetic Monitoringpetrkindlmann/qa-skills163—~5.8kAutomated safety check: PassMIT
Performance Testingkid-sid/claude-spellbook189—~2.5kAutomated safety check: PassMIT
Capacity And Cost Engineeringmagnus919/agent-skills111—~4.7kAutomated safety check: PassMIT
Capacity PlannerFerroxLabs/wayland608—~4.4kAutomated safety check: PassApache-2.0
Performance Engineeringancoleman/ai-design-components526—~2.9kAutomated safety check: NotesMIT

Similar skills

  • Synthetic Monitoring

    petrkindlmann/qa-skills

    Scheduled probes that run CONTINUOUSLY after release. An agent skill from petrkindlmann/qa-skills.

    163 GitHub stars~5.8k tokensUpdated 3 mo ago
    DevOps & CloudAuto-check passed
  • Performance Testing

    kid-sid/claude-spellbook

    A skill your agent uses when load testing a service before launch or after a significant traffic change — writing k6 or Locust scripts, setting SLO-based pass/fail thresholds, diagnosing bottlenecks…

    189 GitHub stars~2.5k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • Capacity And Cost Engineering

    magnus919/agent-skills

    Model technical capacity, unit cost, and budget constraints connected to demand, performance, and reliability decisions.

    111 GitHub stars~4.7k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Capacity Planner

    FerroxLabs/wayland

    Capacity planning expertise covering load testing methodologies, autoscaling policies, resource forecasting, performance budgets, cost-capacity curves, bottleneck identification, queue theory…

    608 GitHub stars~4.4k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Performance Engineering

    ancoleman/ai-design-components

    When validating system performance under load, identifying bottlenecks through profiling, or optimizing application responsiveness.

    526 GitHub stars~2.9k tokensUpdated 10 mo ago
    Testing & QAAuto-check: notes
  • Afrexai Performance Engineering

    LeoYeAI/openclaw-master-skills

    Complete performance engineering system — profiling, optimization, load testing, capacity planning, and performance culture.

    2.2k GitHub stars~7.1k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed

More from jaktestowac/awesome-copilot-for-testers

All 13 skills in this repo
  • API Playwright Test Developer

    jaktestowac/awesome-copilot-for-testers

    Writes and reviews API automation tests with Playwright Test, covering setup/teardown, assertions, data management, and hybrid API+UI flows.

    116 GitHub stars~2k tokensUpdated 1 mo ago
    Auto-check passed
  • Assessing Comprehension Debt

    jaktestowac/awesome-copilot-for-testers

    Measures the risk that code shipped without anyone understanding it: a teach-back attestation on high-risk changes, a risk band from changed-code complexity, diff size and whether a human…

    116 GitHub stars~2.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Creating Orchestration Packs

    jaktestowac/awesome-copilot-for-testers

    Creates agent orchestration packs: cooperating .agent.md files with an orchestrator, subagents, matched handoffs, minimal tool grants, and a shared handoff packet contract.

    116 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Creating Plugins

    jaktestowac/awesome-copilot-for-testers

    Packages repository skills as installable Copilot plugins: marketplace registration, plugin.json manifests, generated skill copies, and the sync check CI enforces.

    116 GitHub stars~3.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Governing Quality Waivers

    jaktestowac/awesome-copilot-for-testers

    Turns "we will skip this check for now" into a dated, attributed, expiring waiver with a stated reason and owner, inventories the silent skips already hiding in a repo - skipped tests, disabled lint…

    116 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Recording Change Intent

    jaktestowac/awesome-copilot-for-testers

    Requires an externalised rationale for high-risk changes - new public exports, new endpoints, auth edits, migrations, removed guards - recorded as an Intent commit trailer, an ADR reference, or a…

    116 GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Testing Performance And Load

What does Testing Performance And Load do?

Designs and runs performance and load tests: workload modelling from real traffic, thresholds tied to SLOs, warmup and ramp shapes, percentile-based analysis, and lightweight CI perf checks with k6…. Testing Performance And Load is an agent skill from jaktestowac/awesome-copilot-for-testers. Designs and runs performance and load tests: workload modelling from real traffic, thresholds tied to SLOs, warmup and ramp shapes, percentile-based analysis, and lightweight CI perf checks with k6 or Artillery.

When should I use Testing Performance And Load?

Testing Performance And Load fits situations like: A feature has latency; throughput requirements; it feels slow needs to become a number; A launch needs a capacity check.

How do I install Testing Performance And Load in Claude Code?

Run `npx skills add jaktestowac/awesome-copilot-for-testers --skill testing-performance-and-load -a claude-code`. Or copy the skill folder (skills/testing-performance-and-load in jaktestowac/awesome-copilot-for-testers) into .claude/skills/testing-performance-and-load in your project. Claude Code loads it when a task matches its description.

How do I install Testing Performance And Load in Codex?

Run `npx skills add jaktestowac/awesome-copilot-for-testers --skill testing-performance-and-load -a codex`. Or copy the skill folder (skills/testing-performance-and-load in jaktestowac/awesome-copilot-for-testers) into .agents/skills/testing-performance-and-load in your project. Codex loads it when a task matches its description.

Can I use Testing Performance And Load in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaktestowac/awesome-copilot-for-testers --skill testing-performance-and-load -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/testing-performance-and-load, .gemini/skills/testing-performance-and-load, .github/skills/testing-performance-and-load and .opencode/skills/testing-performance-and-load in your project.

What does Testing Performance And Load need to run?

SKILL.md names no scripts, command-line tools or credentials: Testing Performance And Load is instructions for the agent only.

Does Testing Performance And Load access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Testing Performance And Load safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Testing Performance And Load use?

Testing Performance And Load is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Testing Performance And Load use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Testing Performance And Load?

Skills that share tags, products or a category with Testing Performance And Load: Synthetic Monitoring (petrkindlmann/qa-skills, 163 stars), Performance Testing (kid-sid/claude-spellbook, 189 stars), Capacity And Cost Engineering (magnus919/agent-skills, 111 stars) and Capacity Planner (FerroxLabs/wayland, 608 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Testing Performance And Load?

jaktestowac (a GitHub user) maintains it in jaktestowac/awesome-copilot-for-testers, which has 116 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on August 26, 2026.

Source: jaktestowac/awesome-copilot-for-testers on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.