Onboarding Validation
open-edge-platform/edge-ai-suites
Validate the get-started experience of Open Edge Platform (OEP) software components from the perspective of a first-time user.
Safe-release techniques DURING rollout: feature flags, progressive rollouts, canary analysis, guardrail metrics, production smoke tests, and synthetic users.
$ npx skills add petrkindlmann/qa-skills --skill testing-in-production -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install petrkindlmann/qa-skills testing-in-production --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/testing-in-production .claude/skills/testing-in-production && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "testing-in-production" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/testing-in-production into .claude/skills/testing-in-production/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "testing-in-production", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/petrkindlmann/qa-skills/tree/main/skills/testing-in-productionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add petrkindlmann/qa-skills --skill testing-in-production -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install petrkindlmann/qa-skills testing-in-production --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/testing-in-production .agents/skills/testing-in-production && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "testing-in-production" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/testing-in-production into .agents/skills/testing-in-production/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "testing-in-production", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add petrkindlmann/qa-skills --skill testing-in-production -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install petrkindlmann/qa-skills testing-in-production --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/testing-in-production .cursor/skills/testing-in-production && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "testing-in-production" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/testing-in-production into .cursor/skills/testing-in-production/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "testing-in-production", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/petrkindlmann/qa-skills.git --path skills/testing-in-production--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add petrkindlmann/qa-skills --skill testing-in-production -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install petrkindlmann/qa-skills testing-in-production --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/testing-in-production .gemini/skills/testing-in-production && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "testing-in-production" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/testing-in-production into .gemini/skills/testing-in-production/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "testing-in-production", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install petrkindlmann/qa-skills testing-in-productionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add petrkindlmann/qa-skills --skill testing-in-production -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/testing-in-production .github/skills/testing-in-production && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "testing-in-production" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/testing-in-production into .github/skills/testing-in-production/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "testing-in-production", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add petrkindlmann/qa-skills --skill testing-in-production -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install petrkindlmann/qa-skills testing-in-production --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/testing-in-production .opencode/skills/testing-in-production && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "testing-in-production" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/testing-in-production into .opencode/skills/testing-in-production/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "testing-in-production", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
testing-in-productionSafe-release techniques DURING rollout: feature flags, progressive rollouts, canary analysis, guardrail metrics, production smoke tests, and synthetic users.
Testing In Production is an agent skill from petrkindlmann/qa-skills. Safe-release techniques DURING rollout: feature flags, progressive rollouts, canary analysis, guardrail metrics, production smoke tests, and synthetic users. Bridges QA and SRE practices. Use when: "feature flag testing," "canary deploy," "progressive rollout," "guardrail metrics," "dark launch," "safe rollout." Not for: scheduled probes that run continuously after release — use synthetic-monitoring. Not for: designing tests from prod telemetry — use observability-driven-testing. Related: release-readiness…
Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/patterns.md` and `references/rollout-policy.md`).
It sits in DevOps & Cloud, covering Deployment, QA and bug reports and Feature launches and release readiness. The repository describes itself as: 50 QA and test-automation skills for Claude Code, Codex, Cursor, and any Agent Skills Standard runtime. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b3bb61b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
npxFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Testing In Production loads about 5.3k tokens when it runs, and up to ~8k if it reads all its reference files. Until then it costs about 151 tokens; SKILL.md has 2,511 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from petrkindlmann/qa-skills at commit b3bb61b, republished under its MIT licence (© petrkindlmann). 2,511 words, ~5,269 tokens.
.claude/skills/testing-in-production/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.<objective>
Production is the only environment that is production. Every other environment is an approximation — staging never replicates real data volume, traffic, or third-party quirks. This skill covers how to validate quality in production safely: controlled blast radius, automated rollback, guardrail metrics, and smoke tests that catch problems before users do. The recurring failure it prevents: shipping to 100% of users with no flag to flip, no baseline to compare against, and no tested way back.
</objective>
| Situation | Go to |
|---|---|
| Shipping a feature behind a flag | Feature Flag Testing |
| Ramping traffic 1% → 100% with gates | Progressive Rollout + references/rollout-policy.md |
| Need post-deploy checks on every release | Production Smoke Tests + references/patterns.md |
| Deciding what numbers gate the rollout | Guardrail Metrics |
| New code path with no user-visible change yet | Dark Launches |
| Proving the rollback actually works | Verification |
Check .agents/qa-project-context.md first. If it exists, use it as context and skip questions already answered there.
Feature flag system:
Rollout capability:
Monitoring maturity:
Production access and safety:
Staging approximates production. It does not replicate production's data volume, traffic patterns, third-party integrations, infrastructure quirks, or user behavior. Testing in production is not reckless — it is realistic. The question is not whether to test in production, but how to do it safely.
Every production test must answer: "If this goes wrong, how many users are affected?" The answer must be as small as possible. Feature flags, canary deploys, and traffic splitting exist to shrink the blast radius from 100% to 1% or less.
Before any production test begins, the rollback mechanism must be identified, tested, and fast. "Disable the flag" is a good rollback plan. "Redeploy the previous version" is acceptable. "We'll figure it out" is not a plan. A rollback you have never fired is a hypothesis, not a plan — see Verification for how to prove it works.
You cannot test in production without monitoring. If you cannot measure error rates, latency, and business metrics in real time, you cannot detect problems. Fix monitoring gaps before adding production tests.
Production tests must never corrupt real user data, send real notifications to real users, charge real payment methods, or create side effects that require manual cleanup. Synthetic accounts, test flags, and isolated resources are mandatory.
Feature flags are the safest mechanism for production testing. They decouple deployment from release and provide instant rollback.
Every flagged feature needs tests in both states. The flag-off path is the rollback path and must work flawlessly — use setFeatureFlag(name, true|false, { userId: TEST_USER_ID }) to drive both. See references/patterns.md for the full ON/OFF test pair.
Flags are not just on or off. They transition through states, and each transition must be validated.
Flag lifecycle:
Created → Targeting internal users → Canary (1%) → Partial (10-50%) → Full (100%) → Cleanup (removed)
Test at each stage:
- Internal: Feature works for internal accounts, hidden from external
- Canary: Metrics are comparable between flag-on and flag-off cohorts
- Partial: No performance degradation at scale
- Full: All user segments work correctly
- Cleanup: Code with flag removed behaves identically to flag-onFlags left in code become technical debt. Run a weekly CI job that queries the flag provider for flags that are 100% rolled out and older than 14 days. These are candidates for code cleanup — remove the flag branching logic and retain only the enabled path. Removing the flag without removing the dead code is half a cleanup.
When multiple flags interact, test the combinations that matter. Do not test all 2^N combinations — focus on flags that affect the same user flow (e.g. new-checkout, express-pay, discount-engine-v2 all touch checkout). Pick the critical combinations: all-new, a representative mixed state, and all-legacy. See references/patterns.md for the combination-test loop.
Vendor-native canary analysis. Before hand-rolling the rollout-policy YAML below, check whether your platform already does it: LaunchDarkly Guarded Rollouts (auto-monitored progressive rollouts with metric-based auto-rollback; uses a frequentist sequential-testing analysis model since early 2026), Statsig Auto-tune, Argo Rollouts AnalysisRun, Flagger, Harness Continuous Verification. If you have one, prefer it — the integration with your metrics and rollback mechanics is cheaper than maintaining a custom analysis loop.
AI feature rollout is its own pattern: model variant + prompt as a flag value, with cost guardrails and a kill switch. LaunchDarkly AI Configs / AgentControl (AI Configs rebranded under the AgentControl umbrella, announced May 2026) is the documented path for shipping LLM features behind progressive rollout. See
release-readinessfor the full pattern.
A structured rollout with explicit promotion criteria at each stage.
| Stage | Traffic | Hold Time | Key Checks |
|---|---|---|---|
| Canary | 1% | 15-30 min | Error rate, crash rate, exceptions |
| Early adopters | 10% | 1-2 hours | Latency P95, conversion rate |
| Partial | 50% | 2-4 hours | All guardrails, business metrics |
| Full | 100% | 24 hours monitoring | Long-tail issues, batch job compatibility |
Define machine-checkable conditions for advancing between stages — hold_duration plus metric conditions (error_rate_5xx < 0.5%, latency_p95 < 500ms, crash_rate == 0). Automatic rollback fires when guardrails are breached, with no human approval needed: error_rate_5xx > 2x_baseline for 5m, latency_p99 > 3x_baseline for 5m, crash_rate > 0.1% for 2m, each notifying on-call. Gate on error-budget burn rate, not only raw multipliers, so slow burns that still blow the SLO are caught. See references/rollout-policy.md for the full promotion YAML, rollback triggers, and SLO-gate config.
Auto-rollback firing is not the same as the incident being resolved. After a rollback triggers, do not declare "recovered" until you have confirmed the system is actually back:
GET /api/health returns healthy and the previous version).Only when all four hold is the rollback verified. See Verification for the staging dry-run that proves this chain before first production use.
Run immediately after every deployment as a pipeline stage (not only in pre-deploy CI). These verify core functionality works with production configuration, data, and infrastructure: a /api/health check, the authentication flow with synthetic credentials, core data loading, and search. Configure retries: 1 and a timeout so a flaky post-deploy run doesn't block the pipeline on the first blip. See references/patterns.md for the full production-smoke.spec.ts.
Production test accounts must be clearly distinguishable from real users: a reserved email pattern (smoke-test+{env}@yourcompany.com), an is_synthetic = true flag, and exclusion from analytics, billing, and email campaigns. Create them via admin API, and prefer short-lived OIDC / workload-identity tokens over long-lived passwords. See references/patterns.md for the full conventions.
Production smoke tests must read, not write. When writes are unavoidable, clean up in fixture teardown so cleanup runs whether the test passes or fails.
Do not call test.afterEach() inside a test() body — Playwright registers hooks at describe/file scope, so a hook registered mid-test never schedules teardown and throws test.afterEach() can only be called in a describe block. The data leaks. Use an auto-cleanup fixture that records created resource IDs and deletes them on teardown (or try/finally for a one-off script). See references/patterns.md for the fixture-based create-verify-cleanup pattern.
| Category | Metric | Comparison Method | Alert Threshold |
|---|---|---|---|
| Errors | HTTP 5xx rate | vs. pre-deploy baseline | >2x baseline for 5 min |
| Errors | Unhandled exception count | vs. pre-deploy baseline | Any new exception type |
| Latency | P50 response time | vs. pre-deploy baseline | >1.5x baseline |
| Latency | P95 response time | vs. pre-deploy baseline | >2x baseline |
| Latency | P99 response time | vs. pre-deploy baseline | >3x baseline |
| Business | Conversion rate | vs. 7-day average | Drop >5% |
| Business | Revenue per session | vs. 7-day average | Drop >10% |
| Client | Crash rate (mobile) | vs. previous release | >0.1% increase |
| Client | JavaScript error rate | vs. pre-deploy baseline | >2x baseline |
| Infra | CPU utilization | absolute | >80% sustained |
| Infra | Memory utilization | absolute | >85% sustained |
Compare canary metrics against a control group running the previous version, not against historical data alone.
Comparison approaches (best to worst):
1. Canary vs. control: split traffic, compare groups in real time (best)
2. Before/after: compare post-deploy metrics to pre-deploy window (good)
3. Historical: compare to same time last week (acceptable for trends)
4. Absolute thresholds: fixed thresholds regardless of baseline (fragile)For business metrics (conversion, revenue), small sample sizes produce noisy results. Wait for statistical significance before drawing conclusions.
Minimum sample sizes for rollout decisions:
- Error rate: 1,000 requests (errors are rare events, need volume)
- Latency: 500 requests (more stable, converges faster)
- Conversion rate: 5,000 sessions (business metrics have high variance)
- Crash rate: 10,000 app launches (crashes are rare events)
Rule of thumb: if you don't have enough traffic at 1% to reach
significance in 30 minutes, increase to 5% or extend the hold window.Dark launches deploy new functionality to production but hide it from users. Real production traffic exercises the new code path without user-visible impact.
Duplicate incoming requests to the new service. Compare responses without returning the new response to the user.
Request flow:
User → Load Balancer → Production Service (returns response to user)
↘ Shadow Service (processes request, logs result, discards)
What to compare:
- Response status codes: shadow should match production
- Response body: diff for semantic equivalence (ignore timestamps, IDs)
- Latency: shadow should not be significantly slower
- Error rate: shadow should not produce more errorsFor migrations (new database, new algorithm, new service), run both the old and new path in production. The old path returns the result to the user; the new path runs asynchronously, logs differences, and discards its result. Track the match rate over time — target 99%+ match before cutting over.
Shadow launch timeline:
Week 1: Deploy shadow, start comparing, expect <50% match
Week 2: Fix mismatches, match rate should climb to 90%+
Week 3: Match rate stable at 99%+, handle remaining edge cases
Week 4: Cut over: shadow becomes primary, old becomes shadow
Week 5: Remove old path after 1 week of stabilityRunning production tests without dashboards and alerts is flying blind. You will not know if your tests caused an issue until a user reports it. Fix: Monitoring is a prerequisite. Before adding any production test, verify you can see error rates, latency, and key business metrics in real time. Set up alerts before the first test runs.
"We'll deploy a fix if something goes wrong" is not a rollback plan. Under pressure, fixes take longer, introduce new bugs, and extend the outage. Fix: Every production test or rollout must have a documented rollback mechanism that takes less than 5 minutes to execute — feature flag disable, previous deployment, or traffic reroute — and a verification step that confirms it actually recovered the system (see Verification).
Production tests that create real orders, send real emails, or modify real user data are not tests — they are incidents waiting to happen. Fix: Use synthetic accounts flagged as test data. Use sandbox modes for payment and email. Clean up created data in fixture teardown so it always runs. If a test cannot be made non-destructive, it does not belong in production.
A cleanup step placed after an assertion never runs when the assertion fails, leaking exactly the test data it was meant to remove. Calling test.afterEach() inside a test() body is the same trap — it throws or is ignored.
Fix: Put teardown in an auto-cleanup fixture or a finally block so it runs on pass and fail alike. See references/patterns.md.
Production testing supplements pre-production testing. It does not replace it. If staging is broken and you are "testing in production" because it is the only working environment, fix staging first. Fix: Maintain a working pre-production environment. Use production testing for what only production can validate: real traffic, real data volumes, real third-party integrations.
Deploying to 1% of traffic but not comparing canary metrics against a control group misses the entire point. You are just deploying slowly, not detecting problems. Fix: Always compare canary metrics against a baseline. Use side-by-side dashboards or automated canary analysis tools (Kayenta, Argo Rollouts analysis).
Flags that are fully rolled out but never removed accumulate. After a year, you have 200 flags with unknown interactions, and every code path has branching logic that nobody understands. Fix: Every flag gets an expiration date at creation time. After full rollout + 2 weeks of stability, remove the flag and its dead branch. Track flag age and alert when flags exceed their expiration.
Prove the rollback path actually fires before the first production use — a rollback you have never triggered is a hypothesis. Smallest check first:
error_rate_5xx above 2x_baseline, or fail the health check 3 times) on a staged rollout wired to the same automatic_rollback policy. Confirm the rollback action fires within its for: window.status === 'healthy' and the previous version. Query the flag platform/deploy that the previous build is serving — don't assume.npx playwright test production-smoke.spec.ts exits 0.If steps 1–4 cannot be demonstrated in staging, the rollback is unverified and the rollout is not ready.
Agent shortcut: vendor MCP servers exist for LaunchDarkly, GrowthBook, Unleash, Flagsmith, Statsig, and Harness FME, letting an AI agent flip flags and read rollout metrics directly during these checks rather than driving the dashboard by hand.
references/)production-smoke.spec.ts suite, synthetic-account conventions, and the fixture-based non-destructive create-verify-cleanup pattern.© petrkindlmann, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in skills/testing-in-production of petrkindlmann/qa-skills.
Open the folder on GitHubat commit b3bb61b
Testing In Production next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Testing In Production this skillpetrkindlmann/qa-skills | 170 | — | ~5.3k | Automated safety check: Pass | MIT | |
| Onboarding Validationopen-edge-platform/edge-ai-suites | 140 | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | |
| Frontmcp Production Readinessagentfront/frontmcp | 146 | — | ~6.5k | Automated safety check: Pass | Apache-2.0 | |
| Deploying Scalable Agentsmicrosoft/ai-agents-for-beginners | 77k | — | ~1.5k | Automated safety check: Pass | MIT | |
| Deploying Scalable Agentsmicrosoft/ai-agents-for-beginners | 77k | — | ~1.6k | Automated safety check: Pass | MIT | |
| Release Itwondelai/skills | 2.4k | — | ~4k | Automated safety check: Pass | MIT |
open-edge-platform/edge-ai-suites
Validate the get-started experience of Open Edge Platform (OEP) software components from the perspective of a first-time user.
agentfront/frontmcp
Pre-production audit, hardening, and go-live checklists for FrontMCP servers.
microsoft/ai-agents-for-beginners
Take a working agent prototype to a scalable, observable production deployment on Microsoft Foundry.
microsoft/ai-agents-for-beginners
Take one working agent prototype go scalable, observable production deployment for Microsoft Foundry.
wondelai/skills
Build production-ready systems with stability patterns: circuit breakers, bulkheads, timeouts, and retry logic.
rsmdt/the-startup
Unified platform operations guidance for CI/CD pipeline design, deployment strategies, observability, SLI/SLOs, and incident-ready rollouts.
petrkindlmann/qa-skills
Test for WCAG 2.2 AA compliance with axe-core + Playwright, keyboard navigation audits, screen reader testing, ARIA pattern validation, and legal compliance mapping (ADA, EAA, Section 508).
petrkindlmann/qa-skills
Goal-driven E2E testing where a browser agent (Playwright MCP / computer-use) reads a natural-language goal and explores the app via the accessibility tree to assert outcomes — no pre-written script.
petrkindlmann/qa-skills
Use AI to write NEW test code from specs, PRDs, user stories, code diffs, bug reports, or OpenAPI specs.
petrkindlmann/qa-skills
Test REST and GraphQL APIs with Playwright APIRequestContext, Supertest, or standalone HTTP clients.
petrkindlmann/qa-skills
Design CI/CD pipelines that run test suites. An agent skill from petrkindlmann/qa-skills.
petrkindlmann/qa-skills
Test for regulatory compliance: GDPR/CMP consent verification, Google Consent Mode v2, Global Privacy Control (GPC), CCPA/US state opt-out, EU AI Act Article 50 transparency, Better Ads Standards…
Categories
Safe-release techniques DURING rollout: feature flags, progressive rollouts, canary analysis, guardrail metrics, production smoke tests, and synthetic users. Testing In Production is an agent skill from petrkindlmann/qa-skills. Safe-release techniques DURING rollout: feature flags, progressive rollouts, canary analysis, guardrail metrics, production smoke tests, and synthetic users.
Testing In Production fits situations like: : feature flag testing; progressive rollout; guardrail metrics; safe rollout. Not for: scheduled probes that run continuously after release — use synthetic-monitoring.
Run `npx skills add petrkindlmann/qa-skills --skill testing-in-production -a claude-code`. Or copy the skill folder (skills/testing-in-production in petrkindlmann/qa-skills) into .claude/skills/testing-in-production in your project. Claude Code loads it when a task matches its description.
Run `npx skills add petrkindlmann/qa-skills --skill testing-in-production -a codex`. Or copy the skill folder (skills/testing-in-production in petrkindlmann/qa-skills) into .agents/skills/testing-in-production in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add petrkindlmann/qa-skills --skill testing-in-production -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/testing-in-production, .gemini/skills/testing-in-production, .github/skills/testing-in-production and .opencode/skills/testing-in-production in your project.
Going by SKILL.md and its folder, Testing In Production needs the command-line tools its instructions call (npx).
SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Testing In Production is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.7k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Testing In Production: Onboarding Validation (open-edge-platform/edge-ai-suites, 140 stars), Frontmcp Production Readiness (agentfront/frontmcp, 146 stars), Deploying Scalable Agents (microsoft/ai-agents-for-beginners, 77k stars) and Deploying Scalable Agents (microsoft/ai-agents-for-beginners, 77k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
petrkindlmann (a GitHub user) maintains it in petrkindlmann/qa-skills, which has 170 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on June 10, 2026.
Source: petrkindlmann/qa-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.