Agent skill

Performance Testing

by proffesor-for-testing in proffesor-for-testing/agentic-qe

Profiles application performance under load using k6, Artillery, or JMeter to measure latency, throughput, and error rates.

MITAuto-check passedTesting & QA

Install Performance Testing

skills CLI
$ npx skills add proffesor-for-testing/agentic-qe --skill performance-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install proffesor-for-testing/agentic-qe performance-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/proffesor-for-testing/agentic-qe.git skills-src && mkdir -p .claude/skills && cp -r skills-src/assets/skills/performance-testing .claude/skills/performance-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
performance-testing
GitHub stars
494
Token cost
~2.4k tokens
SKILL.md length
609 words
Files
7 (incl. scripts, references)
Skills in repo
111
Repo updated
First seen
Licence
MIT

At a glance

Profiles application performance under load using k6, Artillery, or JMeter to measure latency, throughput, and error rates.

  • Works in 5 steps: DEFINE SLOs: p95 response time,… → IDENTIFY critical paths: revenue flows,… → CREATE realistic scenarios: user… → …
  • Planning load tests
  • SKILL.md covers Quick Reference Card, Defining SLOs, Realistic Scenarios and Common Bottlenecks, plus 10 more sections
  • Calls node

What it does

Performance Testing is an agent skill from proffesor-for-testing/agentic-qe. Profiles application performance under load using k6, Artillery, or JMeter to measure latency, throughput, and error rates. Use when planning load tests, stress tests, soak tests, benchmarking APIs, or identifying performance bottlenecks.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `config.json`, `evals/performance-testing.yaml` and `references/k6-patterns.md`).

It sits in Testing & QA, covering Load testing. The repository describes itself as: Agentic QE Fleet is an open-source AI-powered QA/QE platform designed for use with Coding Agents (works best with Claude Code) featuring specialized agents and skills to support… The licence is MIT.

When your agent uses it

  • Planning load tests
  • Benchmarking APIs
  • Identifying performance bottlenecks

Example prompts

  • “Use the performance-testing skill to profile application performance under load using k6, Artillery, or JMeter to measure latency, throughput, and…”
  • “/performance-testing”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. DEFINE SLOs: p95 response time, throughput, error rate targets
  2. IDENTIFY critical paths: revenue flows, high-traffic pages, key APIs
  3. CREATE realistic scenarios: user journeys, think time, varied data
  4. EXECUTE with monitoring: CPU, memory, DB queries, network
  5. ANALYZE bottlenecks and fix before production

What it can do on your machine

Read from SKILL.md and the folder at commit 829d030. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Performance Testing loads about 2.4k tokens when it runs, and up to ~2.8k if it reads all its reference files. Until then it costs about 65 tokens; SKILL.md has 609 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~65
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from proffesor-for-testing/agentic-qe at commit 829d030, republished under its MIT licence (© proffesor-for-testing). 609 words, ~2,381 tokens.

Download SKILL.mdSave it as .claude/skills/performance-testing/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
performance-testing
description
Profiles application performance under load using k6, Artillery, or JMeter to measure latency, throughput, and error rates. Use when planning load tests, stress tests, soak tests, benchmarking APIs, or identifying performance bottlenecks.
category
specialized-testing
priority
high
tokenEstimate
1100
agents
qe-performance-tester, qe-quality-analyzer, qe-production-intelligence
implementation_status
optimized
optimization_version
1
last_optimized
2025-12-02
quick_reference_card
true
tags
performance, load-testing, stress-testing, scalability, k6, bottlenecks
trust_tier
3

Performance Testing

<default_to_action> When testing performance or planning load tests:

  1. DEFINE SLOs: p95 response time, throughput, error rate targets
  2. IDENTIFY critical paths: revenue flows, high-traffic pages, key APIs
  3. CREATE realistic scenarios: user journeys, think time, varied data
  4. EXECUTE with monitoring: CPU, memory, DB queries, network
  5. ANALYZE bottlenecks and fix before production

Quick Test Type Selection:

  • Expected load validation → Load testing
  • Find breaking point → Stress testing
  • Sudden traffic spike → Spike testing
  • Memory leaks, resource exhaustion → Endurance/soak testing
  • Horizontal/vertical scaling → Scalability testing

Critical Success Factors:

  • Performance is a feature, not an afterthought
  • Test early and often, not just before release
  • Focus on user-impacting bottlenecks </default_to_action>

Quick Reference Card

When to Use
  • Before major releases
  • After infrastructure changes
  • Before scaling events (Black Friday)
  • When setting SLAs/SLOs
Test Types
TypePurposeWhen
LoadExpected trafficEvery release
StressBeyond capacityQuarterly
SpikeSudden surgeBefore events
EnduranceMemory leaksAfter code changes
ScalabilityScaling validationInfrastructure changes
Key Metrics
MetricTargetWhy
p95 response< 200msUser experience
Throughput10k req/minCapacity
Error rate< 0.1%Reliability
CPU< 70%Headroom
Memory< 80%Stability
Tools
  • k6: Modern, JS-based, CI/CD friendly
  • JMeter: Enterprise, feature-rich
  • Artillery: Simple YAML configs
  • Gatling: Scala, great reporting
Agent Coordination
  • qe-performance-tester: Load test orchestration
  • qe-quality-analyzer: Results analysis
  • qe-production-intelligence: Production comparison

Defining SLOs

Bad: "The system should be fast" Good: "p95 response time < 200ms under 1,000 concurrent users"

javascript
export const options = {
  thresholds: {
    http_req_duration: ['p(95)<200'],  // 95% < 200ms
    http_req_failed: ['rate<0.01'],     // < 1% failures
  },
};

Realistic Scenarios

Bad: Every user hits homepage repeatedly Good: Model actual user behavior

javascript
// Realistic distribution
// 40% browse, 30% search, 20% details, 10% checkout
export default function () {
  const action = Math.random();
  if (action < 0.4) browse();
  else if (action < 0.7) search();
  else if (action < 0.9) viewProduct();
  else checkout();

  sleep(randomInt(1, 5)); // Think time
}

Common Bottlenecks

Database

Symptoms: Slow queries under load, connection pool exhaustion Fixes: Add indexes, optimize N+1 queries, increase pool size, read replicas

N+1 Queries
javascript
// BAD: 100 orders = 101 queries
const orders = await Order.findAll();
for (const order of orders) {
  const customer = await Customer.findById(order.customerId);
}

// GOOD: 1 query
const orders = await Order.findAll({ include: [Customer] });
Synchronous Processing

Problem: Blocking operations in request path (sending email during checkout) Fix: Use message queues, process async, return immediately

Memory Leaks

Detection: Endurance testing, memory profiling Common causes: Event listeners not cleaned, caches without eviction

External Dependencies

Solutions: Aggressive timeouts, circuit breakers, caching, graceful degradation


k6 CI/CD Example

javascript
// performance-test.js
import http from 'k6/http';
import { check, sleep } from 'k6';

export const options = {
  stages: [
    { duration: '1m', target: 50 },   // Ramp up
    { duration: '3m', target: 50 },   // Steady
    { duration: '1m', target: 0 },    // Ramp down
  ],
  thresholds: {
    http_req_duration: ['p(95)<200'],
    http_req_failed: ['rate<0.01'],
  },
};

export default function () {
  const res = http.get('https://api.example.com/products');
  check(res, {
    'status is 200': (r) => r.status === 200,
    'response time < 200ms': (r) => r.timings.duration < 200,
  });
  sleep(1);
}
yaml
# GitHub Actions
- name: Run k6 test
  uses: grafana/k6-action@v0.3.0
  with:
    filename: performance-test.js

Analyzing Results

Good Results
Load: 1,000 users | p95: 180ms | Throughput: 5,000 req/s
Error rate: 0.05% | CPU: 65% | Memory: 70%
Problems
Load: 1,000 users | p95: 3,500ms ❌ | Throughput: 500 req/s ❌
Error rate: 5% ❌ | CPU: 95% ❌ | Memory: 90% ❌
Root Cause Analysis
  1. Correlate metrics: When response time spikes, what changes?
  2. Check logs: Errors, warnings, slow queries
  3. Profile code: Where is time spent?
  4. Monitor resources: CPU, memory, disk
  5. Trace requests: End-to-end flow

Show full SKILL.md (253 more words)Show less

Anti-Patterns

❌ Anti-Pattern✅ Better
Testing too lateTest early and often
Unrealistic scenariosModel real user behavior
0 to 1000 users instantlyRamp up gradually
No monitoring during testsMonitor everything
No baselineEstablish and track trends
One-time testingContinuous performance testing

Agent-Assisted Performance Testing

typescript
// Comprehensive load test
await Task("Load Test", {
  target: 'https://api.example.com',
  scenarios: {
    checkout: { vus: 100, duration: '5m' },
    search: { vus: 200, duration: '5m' },
    browse: { vus: 500, duration: '5m' }
  },
  thresholds: {
    'http_req_duration': ['p(95)<200'],
    'http_req_failed': ['rate<0.01']
  }
}, "qe-performance-tester");

// Bottleneck analysis
await Task("Analyze Bottlenecks", {
  testResults: perfTest,
  metrics: ['cpu', 'memory', 'db_queries', 'network']
}, "qe-performance-tester");

// CI integration
await Task("CI Performance Gate", {
  mode: 'smoke',
  duration: '1m',
  vus: 10,
  failOn: { 'p95_response_time': 300, 'error_rate': 0.01 }
}, "qe-performance-tester");

Agent Coordination Hints

Memory Namespace
aqe/performance/
├── results/*       - Test execution results
├── baselines/*     - Performance baselines
├── bottlenecks/*   - Identified bottlenecks
└── trends/*        - Historical trends
Fleet Coordination
typescript
const perfFleet = await FleetManager.coordinate({
  strategy: 'performance-testing',
  agents: [
    'qe-performance-tester',
    'qe-quality-analyzer',
    'qe-production-intelligence',
    'qe-deployment-readiness'
  ],
  topology: 'sequential'
});

Pre-Production Checklist

  • Load test passed (expected traffic)
  • Stress test passed (2-3x expected)
  • Spike test passed (sudden surge)
  • Endurance test passed (24+ hours)
  • Database indexes in place
  • Caching configured
  • Monitoring and alerting set up
  • Performance baseline established


Remember

Performance is a feature: Test it like functionality Test continuously: Not just before launch Monitor production: Synthetic + real user monitoring Fix what matters: Focus on user-impacting bottlenecks Trend over time: Catch degradation early

With Agents: Agents automate load testing, analyze bottlenecks, and compare with production. Use agents to maintain performance at scale.

Run History

After each performance test run, append results to run-history.json in this skill directory:

bash
node -e "
const fs = require('fs');
const h = JSON.parse(fs.readFileSync('.claude/skills/performance-testing/run-history.json'));
h.runs.push({date: new Date().toISOString().split('T')[0], scenario: 'load', p95_ms: P95, throughput_rps: RPS, error_rate_pct: ERR});
fs.writeFileSync('.claude/skills/performance-testing/run-history.json', JSON.stringify(h, null, 2));
"

Read run-history.json before each run — compare with baselines. Alert if p95 increases >20% from baseline.

Gotchas

  • k6 scripts generated by agent often hardcode base URLs — use environment variables for portability
  • Load tests in containers hit resource limits before app limits — ensure container has 2x the resources of target
  • Agent forgets to include think time between requests — without it, load is unrealistically bursty
  • P95 vs P99 matters — agent defaults to averages which hide tail latency problems
  • Baseline comparison requires consistent environment — CI runner variance can cause 20%+ noise

© proffesor-for-testing, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references) in assets/skills/performance-testing of proffesor-for-testing/agentic-qe.

  • SKILL.md
  • config.json
  • evals/performance-testing.yaml
  • references/k6-patterns.md
  • run-history.json
  • schemas/output.json
  • scripts/validate-config.json

Open the folder on GitHubat commit 829d030

Compare with similar skills

Performance Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Performance Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Performance Testing this skillproffesor-for-testing/agentic-qe494—~2.4kAutomated safety check: PassMIT
Writing Livekit Scenarioslivekit-examples/agent-starter-python2641 repos~2.5kAutomated safety check: PassMIT
Go Testingcxuu/golang-skills1701 repos~1.3kAutomated safety check: PassApache-2.0
Goalcraftgrp06/goalcraft102—~3.8kAutomated safety check: PassMIT
Thinking Partnermattnowdev/thinking-partner205—~4.4kAutomated safety check: PassMIT
Challengeblueberrycongee/termcanvas406—~1.5kAutomated safety check: PassMIT

Similar skills

  • Writing Livekit Scenarios

    livekit-examples/agent-starter-python

    Creates and maintains the scenarios a LiveKit agent simulation runs, and wires the agent to consume them.

    264 GitHub starsUsed in 1 repo~2.5k tokens
    Testing & QAAuto-check passed
  • Go Testing

    cxuu/golang-skills

    A skill your agent uses when writing, reviewing, or improving Go test code — including table-driven tests, subtests, parallel tests, test helpers, test doubles, and assertions with cmp.Diff.

    170 GitHub starsUsed in 1 repo~1.3k tokens
    Testing & QAAuto-check passed
  • Goalcraft

    grp06/goalcraft

    Turn a rough draft, vague ambition, or messy task brief into a powerful Codex /goal objective for persistent, evidence-checked work.

    102 GitHub stars~3.8k tokensUpdated 4 mo ago
    Testing & QAAuto-check passed
  • Thinking Partner

    mattnowdev/thinking-partner

    A deterministic thinking partner that challenges assumptions and applies mental models to sharpen decisions, solve problems, and think more clearly.

    205 GitHub stars~4.4k tokensUpdated 6 mo ago
    Testing & QAAuto-check passed
  • Challenge

    blueberrycongee/termcanvas

    Adversarial review skill. An agent skill from blueberrycongee/termcanvas.

    406 GitHub stars~1.5k tokensUpdated 4 mo ago
    Testing & QAAuto-check passed
  • Vision

    kunchenguid/vision

    Draft and stress-test a VISION.md for a repository, then iterate with the author on an interactive review board until approved.

    329 GitHub stars~2.9k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed

More from proffesor-for-testing/agentic-qe

All 111 skills in this repo
  • Contract Testing

    proffesor-for-testing/agentic-qe

    Consumer-driven contract testing for microservices using Pact, schema validation, API versioning, and backward compatibility testing.

    494 GitHub stars~1.8k tokensUpdated 3 days ago
    Auto-check passed
  • Mutation Testing

    proffesor-for-testing/agentic-qe

    Test quality validation through mutation testing, assessing test suite effectiveness by introducing code mutations and measuring kill rate.

    494 GitHub stars~1.7k tokensUpdated 3 days ago
    Auto-check passed
  • Code Review Quality

    proffesor-for-testing/agentic-qe

    Conduct context-driven code reviews focusing on quality, testability, and maintainability.

    494 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Security Testing

    proffesor-for-testing/agentic-qe

    Scans for security vulnerabilities including XSS, SQL injection, CSRF, and auth flaws using OWASP Top 10 methodology.

    494 GitHub stars~2.7k tokensUpdated 3 days ago
    Auto-check: notes
  • Database Testing

    proffesor-for-testing/agentic-qe

    Database schema validation, data integrity testing, migration testing, transaction isolation, and query performance.

    494 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Accessibility Testing

    proffesor-for-testing/agentic-qe

    WCAG 2.2 compliance testing, screen reader validation, and inclusive design verification.

    494 GitHub stars~2k tokensUpdated 3 days ago
    Auto-check passed

Categories

Questions about Performance Testing

What does Performance Testing do?

Profiles application performance under load using k6, Artillery, or JMeter to measure latency, throughput, and error rates. Performance Testing is an agent skill from proffesor-for-testing/agentic-qe. Profiles application performance under load using k6, Artillery, or JMeter to measure latency, throughput, and error rates.

When should I use Performance Testing?

Performance Testing fits situations like: planning load tests; benchmarking APIs; identifying performance bottlenecks.

How do I install Performance Testing in Claude Code?

Run `npx skills add proffesor-for-testing/agentic-qe --skill performance-testing -a claude-code`. Or copy the skill folder (assets/skills/performance-testing in proffesor-for-testing/agentic-qe) into .claude/skills/performance-testing in your project. Claude Code loads it when a task matches its description.

How do I install Performance Testing in Codex?

Run `npx skills add proffesor-for-testing/agentic-qe --skill performance-testing -a codex`. Or copy the skill folder (assets/skills/performance-testing in proffesor-for-testing/agentic-qe) into .agents/skills/performance-testing in your project. Codex loads it when a task matches its description.

Can I use Performance Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add proffesor-for-testing/agentic-qe --skill performance-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/performance-testing, .gemini/skills/performance-testing, .github/skills/performance-testing and .opencode/skills/performance-testing in your project.

What does Performance Testing need to run?

Going by SKILL.md and its folder, Performance Testing needs the command-line tools its instructions call (node).

Does Performance Testing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Performance Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Performance Testing use?

Performance Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Performance Testing use?

About 2.4k tokens (SKILL.md is roughly 9.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 423 tokens, read only when the agent opens those files.

What are the alternatives to Performance Testing?

Skills that share tags, products or a category with Performance Testing: Writing Livekit Scenarios (livekit-examples/agent-starter-python, 264 stars), Go Testing (cxuu/golang-skills, 170 stars), Goalcraft (grp06/goalcraft, 102 stars) and Thinking Partner (mattnowdev/thinking-partner, 205 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Performance Testing?

proffesor-for-testing (a GitHub user) maintains it in proffesor-for-testing/agentic-qe, which has 494 GitHub stars. The repository holds 111 skills in this directory. The repository was last updated on October 4, 2026.

Source: proffesor-for-testing/agentic-qe on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.