Agent skill

Performance Testing

by kid-sid in kid-sid/claude-spellbook

A skill your agent uses when load testing a service before launch or after a significant traffic change — writing k6 or Locust scripts, setting SLO-based pass/fail thresholds, diagnosing bottlenecks…

MITAuto-check passedTesting & QA

Install Performance Testing

skills CLI
$ npx skills add kid-sid/claude-spellbook --skill performance-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install kid-sid/claude-spellbook performance-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/kid-sid/claude-spellbook.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/performance-testing .claude/skills/performance-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
performance-testing
GitHub stars
189
Token cost
~2.5k tokens
SKILL.md length
636 words
Files
1
Skills in repo
54
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when load testing a service before launch or after a significant traffic change — writing k6 or Locust scripts, setting SLO-based pass/fail thresholds, diagnosing bottlenecks…

  • Works in 4 steps: Run load test against staging with… → Record p50 / p95 / p99 and error rate → Set regression threshold: fail if p99… → …
  • Load testing a service before launch
  • SKILL.md covers When to Activate, Test Type Decision Table, k6 and Locust (Python), plus 5 more sections
  • Calls go; needs API_TOKEN and K6_CLOUD_TOKEN

What it does

Performance Testing is an agent skill from kid-sid/claude-spellbook. Use when load testing a service before launch or after a significant traffic change — writing k6 or Locust scripts, setting SLO-based pass/fail thresholds, diagnosing bottlenecks under load, or integrating performance tests into CI.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Load testing and Site reliability engineering. The repository describes itself as: A curated collection of skills, prompts, and workflows that extend Claude's capabilities — your personal grimoire for AI-powered development. The licence is MIT.

When your agent uses it

  • Load testing a service before launch
  • After a significant traffic change — writing k6
  • Setting SLO-based pass/fail thresholds
  • Diagnosing bottlenecks under load

Example prompts

  • “/performance-testing”

Requirements

  • Python 3
  • A credential in API_TOKEN
  • A credential in LOAD_TEST_TOKEN

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Run load test against staging with production-like traffic shape
  2. Record p50 / p95 / p99 and error rate
  3. Set regression threshold: fail if p99 degrades > 20% from baseline
  4. Set SLO threshold: fail if p99 exceeds SLO target

What it can do on your machine

Read from SKILL.md and the folder at commit a7c2ac9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • go

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • API_TOKEN
    • K6_CLOUD_TOKEN
    • LOAD_TEST_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Performance Testing loads about 2.5k tokens when it runs. Until then it costs about 63 tokens; SKILL.md has 636 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from kid-sid/claude-spellbook at commit a7c2ac9, republished under its MIT licence (© kid-sid). 636 words, ~2,547 tokens.

Download SKILL.mdSave it as .claude/skills/performance-testing/SKILL.md (or your agent's skills folder).
name
performance-testing
description
Use when load testing a service before launch or after a significant traffic change — writing k6 or Locust scripts, setting SLO-based pass/fail thresholds, diagnosing bottlenecks under load, or integrating performance tests into CI.

Performance Testing

Load and performance testing validates that your system meets latency and throughput requirements under realistic and extreme traffic conditions.

When to Activate

  • Load testing an API before a product launch
  • Setting up k6 or Locust for a project
  • Writing Go benchmark functions for critical code paths
  • Defining SLO-based pass/fail thresholds for load tests
  • Identifying bottlenecks under load (pool exhaustion, N+1, GC pressure)
  • Adding performance regression detection to a CI/CD pipeline

Test Type Decision Table

TypeDescriptionLoad shapeGoalWhen to run
LoadSimulate expected trafficRamp to normal, holdVerify baseline meets SLOPre-launch, nightly
StressPush beyond capacityRamp past normalFind breaking pointBefore scaling decisions
SoakSustained load over timeConstant for 1–4 hoursDetect memory leaks, pool exhaustionWeekly
SpikeSudden burst0 → peak instantlyTest autoscaling, queue bufferingBefore planned events
VolumeLarge datasets, normal loadNormal rps, huge dataFind data-size bottlenecksWhen data volume increases

k6

Script Structure
javascript
import http from 'k6/http';
import { check, sleep } from 'k6';
import { Rate, Trend } from 'k6/metrics';

const errorRate = new Rate('errors');
const paymentDuration = new Trend('payment_duration');

export const options = {
  stages: [
    { duration: '2m', target: 50 },   // ramp up
    { duration: '5m', target: 50 },   // hold
    { duration: '2m', target: 100 },  // ramp up further
    { duration: '5m', target: 100 },  // hold
    { duration: '2m', target: 0 },    // ramp down
  ],
  thresholds: {
    // SLO-based pass/fail: test fails if these are breached
    'http_req_duration': ['p(95)<500', 'p(99)<1000'],
    'http_req_failed':   ['rate<0.01'],
    'errors':            ['rate<0.05'],
  },
};

export default function () {
  const res = http.post(
    'https://api.example.com/payments',
    JSON.stringify({ amount: 100, currency: 'USD' }),
    {
      headers: {
        'Content-Type': 'application/json',
        Authorization: `Bearer ${__ENV.API_TOKEN}`,
      },
    }
  );

  const ok = check(res, {
    'status is 201':          (r) => r.status === 201,
    'response time < 500ms':  (r) => r.timings.duration < 500,
  });

  errorRate.add(!ok);
  paymentDuration.add(res.timings.duration);
  sleep(1);  // think time between requests
}
Scenarios (Mixed Workloads)
javascript
export const options = {
  scenarios: {
    browse: {
      executor: 'constant-vus',
      vus: 100,
      duration: '10m',
      exec: 'browseProducts',
    },
    checkout: {
      executor: 'ramping-arrival-rate',
      startRate: 10,
      timeUnit: '1s',
      stages: [{ duration: '5m', target: 50 }],
      preAllocatedVUs: 60,
      exec: 'checkout',
    },
  },
};

export function browseProducts() { /* ... */ }
export function checkout() { /* ... */ }
Running k6
bash
k6 run script.js
k6 run --vus 100 --duration 10m script.js

# Export to InfluxDB + Grafana for dashboards
k6 run --out influxdb=http://localhost:8086/k6 script.js

# Cloud execution
k6 cloud script.js

Locust (Python)

python
from locust import HttpUser, task, between

class PaymentUser(HttpUser):
    wait_time = between(1, 3)

    def on_start(self):
        """Called once per VU — authenticate"""
        res = self.client.post('/auth/token', json={
            'email': 'test@example.com',
            'password': 'password',
        })
        self.token = res.json()['access_token']

    @task(3)  # weight 3: 3× more frequent than weight-1 tasks
    def browse_products(self):
        with self.client.get(
            '/products',
            headers=self._auth(),
            name='/products',          # group dynamic URLs
            catch_response=True,
        ) as res:
            if res.status_code != 200:
                res.failure(f"Got {res.status_code}")

    @task(1)
    def create_payment(self):
        self.client.post(
            '/payments',
            json={'amount': 100},
            headers=self._auth(),
        )

    def _auth(self):
        return {'Authorization': f'Bearer {self.token}'}
bash
# Headless CI mode
locust -f locustfile.py \
  --headless -u 100 -r 10 --run-time 5m \
  --host https://api.example.com \
  --csv results          # outputs results_stats.csv, results_failures.csv

Go Benchmarks

go
package payment_test

import (
    "fmt"
    "testing"
)

func BenchmarkProcessPayment(b *testing.B) {
    svc := NewPaymentService(testDB)
    b.ResetTimer()   // don't count setup time
    b.ReportAllocs() // show allocations/op in output

    for i := 0; i < b.N; i++ {
        _, err := svc.ProcessPayment(ctx, Payment{Amount: 100})
        if err != nil {
            b.Fatal(err)
        }
    }
}

// Sub-benchmarks for different scenarios
func BenchmarkProcessPayment_Sizes(b *testing.B) {
    for _, amount := range []float64{1, 100, 10_000} {
        b.Run(fmt.Sprintf("amount=%.0f", amount), func(b *testing.B) {
            for i := 0; i < b.N; i++ {
                svc.ProcessPayment(ctx, Payment{Amount: amount})
            }
        })
    }
}
bash
# Run benchmarks
go test -bench=. -benchmem -benchtime=10s ./...
# Output: BenchmarkProcessPayment-8  50000  23456 ns/op  1024 B/op  12 allocs/op

# Compare before/after a change
go test -bench=. -count=10 -benchmem ./... > before.txt
# ... make the change ...
go test -bench=. -count=10 -benchmem ./... > after.txt
benchstat before.txt after.txt

SLO-Based Pass/Fail Criteria

Defining Thresholds from SLOs

Base thresholds on your production SLOs — not arbitrary numbers.

javascript
// If SLO: p99 < 500ms, error rate < 0.1%
thresholds: {
  'http_req_duration': ['p(50)<100', 'p(95)<300', 'p(99)<500'],
  'http_req_failed':   ['rate<0.001'],
}
Establishing a Baseline
  1. Run load test against staging with production-like traffic shape
  2. Record p50 / p95 / p99 and error rate
  3. Set regression threshold: fail if p99 degrades > 20% from baseline
  4. Set SLO threshold: fail if p99 exceeds SLO target
Bottleneck Identification Under Load
SymptomLikely causeHow to confirmFix
Latency climbs with VU countConnection pool exhaustedCheck pool wait metricIncrease pool / add PgBouncer
Error spikes at N rpsThread / goroutine limitCheck active connectionsTune concurrency config
Memory grows during soakMemory leak / large cacheHeap profile during testFix leak, tune GC
High latency, low CPUN+1 queriesCount DB queries per requestAdd eager loading
CPU > 90%Compute bottleneckCPU flame graphOptimize hot path, add cache
Latency spikes periodicallyGC pause (JVM/Go)GC log analysisTune GC, reduce allocations

CI Integration

When to Run
TypeFrequencyTriggerFailure action
Smoke perf (5 VUs, 1 min)Every PRPR CIFail PR if p99 > 2× baseline
Full load testNightlyCronAlert on Slack
Stress testWeeklyCronReport only
GitHub Actions Example
yaml
jobs:
  load-test:
    runs-on: ubuntu-latest
    if: github.event_name == 'schedule'
    steps:
      - uses: actions/checkout@v4

      - name: Run k6 load test
        uses: grafana/k6-action@v0.3.0
        with:
          filename: tests/load/payment.js
        env:
          API_TOKEN: ${{ secrets.LOAD_TEST_TOKEN }}
          K6_CLOUD_TOKEN: ${{ secrets.K6_CLOUD_TOKEN }}

      - name: Upload results
        if: always()
        uses: actions/upload-artifact@v4
        with:
          name: k6-results-${{ github.run_id }}
          path: results/

See also: performance, observability, ci-cd

Show full SKILL.md (267 more words)Show less

Red Flags

  • Symmetric ramp-up/ramp-down without a sustained plateau — spike-then-ramp-down misses memory leaks and GC pressure; hold at target RPS for ≥10 min in steady state
  • Asserting only on HTTP 200 — a cached error page or open circuit breaker returns 200; use check() to assert on specific response body fields, not just the status code
  • Single load generator machine for high VU counts — one machine saturates its NIC before the target; use distributed execution (k6 cloud, multiple Locust workers) above ~500 VUs
  • No baseline before the test — without a pre-change baseline you can't tell whether 300ms p99 is a regression or always was that way
  • Load test traffic escaping into production — test traffic that bypasses rate limits can trigger real customer alerts; isolate by dedicated API key, IP allowlist, or a separate environment
  • Zero think time between requests — real users pause between actions; 0ms think time inflates effective concurrency 5–10×, producing false bottlenecks that don't exist in production
  • Setting SLO thresholds from the first test run — first-run numbers are noisy; run 3+ tests under stable conditions before codifying a regression threshold

Checklist

  • Test type chosen (load/stress/soak/spike) matches the specific question being answered
  • k6 / Locust thresholds tied to SLO values — not made-up numbers
  • Baseline measured before setting regression thresholds
  • Test users and data isolated from production
  • Think time (sleep) included in VU scripts for realistic simulation
  • k6 check() used for per-request assertions (not just global thresholds)
  • Go benchmarks include b.ReportAllocs() and b.ResetTimer()
  • benchstat used to compare before/after for Go performance changes
  • Bottleneck identification checklist followed when tests fail
  • Load test results stored as CI artifacts for trending over time

© kid-sid, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/performance-testing of kid-sid/claude-spellbook.

Open the folder on GitHubat commit a7c2ac9

Compare with similar skills

Performance Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Performance Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Performance Testing this skillkid-sid/claude-spellbook189—~2.5kAutomated safety check: PassMIT
Capacity PlannerFerroxLabs/wayland608—~4.4kAutomated safety check: PassApache-2.0
Performance Engineeringancoleman/ai-design-components526—~2.9kAutomated safety check: NotesMIT
Afrexai Performance EngineeringLeoYeAI/openclaw-master-skills2.2k—~7.1kAutomated safety check: PassMIT
Load Testing Patternsvibeeval/vibecosystem531—~1.6kAutomated safety check: PassMIT
Performance EngineerFerroxLabs/wayland608—~5.1kAutomated safety check: PassApache-2.0

Similar skills

  • Capacity Planner

    FerroxLabs/wayland

    Capacity planning expertise covering load testing methodologies, autoscaling policies, resource forecasting, performance budgets, cost-capacity curves, bottleneck identification, queue theory…

    608 GitHub stars~4.4k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Performance Engineering

    ancoleman/ai-design-components

    When validating system performance under load, identifying bottlenecks through profiling, or optimizing application responsiveness.

    526 GitHub stars~2.9k tokensUpdated 10 mo ago
    Testing & QAAuto-check: notes
  • Afrexai Performance Engineering

    LeoYeAI/openclaw-master-skills

    Complete performance engineering system — profiling, optimization, load testing, capacity planning, and performance culture.

    2.2k GitHub stars~7.1k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • Load Testing Patterns

    vibeeval/vibecosystem

    k6 script templates, load profiles, response time thresholds, SLO validation, and performance testing strategies.

    531 GitHub stars~1.6k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • Performance Engineer

    FerroxLabs/wayland

    Becomes a senior performance engineer who identifies bottlenecks, designs optimization strategies, and conducts load testing using systematic profiling and benchmarking methodology.

    608 GitHub stars~5.1k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • HTTP Load Profiler

    zebbern/claude-code-guide

    Run stepped HTTP load tests with ab/wrk, ramping concurrency levels to collect p50/p90/p99 latency, detect performance inflection points, and recommend optimal concurrency.

    4.6k GitHub stars~1.6k tokensUpdated today
    Testing & QAAuto-check: notes

More from kid-sid/claude-spellbook

All 54 skills in this repo
  • Accessibility

    kid-sid/claude-spellbook

    A skill your agent uses when building or reviewing UI components for keyboard and screen reader compatibility, adding ARIA to custom widgets, auditing a page for WCAG AA conformance, or preparing…

    189 GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Agentex

    kid-sid/claude-spellbook

    A skill your agent uses when building, wiring, or debugging an Agentex agent — choosing agent type, configuring acp.py and manifest.yaml, using adk.messages or adk.state, or resolving…

    189 GitHub stars~2.2k tokensUpdated 2 mo ago
    Auto-check: notes
  • AI Engineer

    kid-sid/claude-spellbook

    A skill your agent uses when building production LLM applications — designing RAG pipelines, choosing vector databases, implementing agent orchestration, optimizing cost, or adding AI safety…

    189 GitHub stars~3.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Angular

    kid-sid/claude-spellbook

    A skill your agent uses when building or refactoring Angular applications — choosing between signals, RxJS, and NgRx for state, configuring routing with guards and lazy loading, optimizing change…

    189 GitHub stars~5k tokensUpdated 2 mo ago
    Auto-check passed
  • API Design

    kid-sid/claude-spellbook

    A skill your agent uses when designing new REST endpoints, reviewing an existing API contract, adding pagination or filtering, planning a versioning strategy, or building a public or partner-facing…

    189 GitHub stars~3.6k tokensUpdated 2 mo ago
    Auto-check passed
  • Auth

    kid-sid/claude-spellbook

    A skill your agent uses when implementing login flows, issuing or validating JWTs, setting up OAuth2/OIDC with a provider, designing role-based or attribute-based access control, securing API…

    189 GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Performance Testing

What does Performance Testing do?

A skill your agent uses when load testing a service before launch or after a significant traffic change — writing k6 or Locust scripts, setting SLO-based pass/fail thresholds, diagnosing bottlenecks…. Performance Testing is an agent skill from kid-sid/claude-spellbook. Use when load testing a service before launch or after a significant traffic change — writing k6 or Locust scripts, setting SLO-based pass/fail thresholds, diagnosing bottlenecks under load, or integrating performance tests into CI.

When should I use Performance Testing?

Performance Testing fits situations like: load testing a service before launch; after a significant traffic change — writing k6; setting SLO-based pass/fail thresholds; diagnosing bottlenecks under load.

How do I install Performance Testing in Claude Code?

Run `npx skills add kid-sid/claude-spellbook --skill performance-testing -a claude-code`. Or copy the skill folder (skills/performance-testing in kid-sid/claude-spellbook) into .claude/skills/performance-testing in your project. Claude Code loads it when a task matches its description.

How do I install Performance Testing in Codex?

Run `npx skills add kid-sid/claude-spellbook --skill performance-testing -a codex`. Or copy the skill folder (skills/performance-testing in kid-sid/claude-spellbook) into .agents/skills/performance-testing in your project. Codex loads it when a task matches its description.

Can I use Performance Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add kid-sid/claude-spellbook --skill performance-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/performance-testing, .gemini/skills/performance-testing, .github/skills/performance-testing and .opencode/skills/performance-testing in your project.

What does Performance Testing need to run?

Going by SKILL.md and its folder, Performance Testing needs the command-line tools its instructions call (go) and credentials named API_TOKEN, K6_CLOUD_TOKEN and LOAD_TEST_TOKEN. Our summary lists: Python 3; A credential in API_TOKEN; A credential in LOAD_TEST_TOKEN.

Does Performance Testing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Performance Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Performance Testing use?

Performance Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Performance Testing use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Performance Testing?

Skills that share tags, products or a category with Performance Testing: Capacity Planner (FerroxLabs/wayland, 608 stars), Performance Engineering (ancoleman/ai-design-components, 526 stars), Afrexai Performance Engineering (LeoYeAI/openclaw-master-skills, 2.2k stars) and Load Testing Patterns (vibeeval/vibecosystem, 531 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Performance Testing?

kid-sid (a GitHub user) maintains it in kid-sid/claude-spellbook, which has 189 GitHub stars. The repository holds 54 skills in this directory. The repository was last updated on August 5, 2026.

Source: kid-sid/claude-spellbook on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.