Web Application Testing
anthropics/skills
Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.
Empirical performance comparison playbook for poku covering hypothesis lenses, a standalone harness with a byte-identical correctness gate, variant generation, sequential hyperfine measurement…
$ npx skills add wellwelwel/poku --skill benchmarking -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install wellwelwel/poku benchmarking --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/wellwelwel/poku.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/benchmarking .claude/skills/benchmarking && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "benchmarking" agent skill from https://github.com/wellwelwel/poku/tree/main/.claude/skills/benchmarking into .claude/skills/benchmarking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmarking", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/wellwelwel/poku/tree/main/.claude/skills/benchmarkingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add wellwelwel/poku --skill benchmarking -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install wellwelwel/poku benchmarking --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wellwelwel/poku.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/benchmarking .agents/skills/benchmarking && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "benchmarking" agent skill from https://github.com/wellwelwel/poku/tree/main/.claude/skills/benchmarking into .agents/skills/benchmarking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmarking", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wellwelwel/poku --skill benchmarking -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install wellwelwel/poku benchmarking --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wellwelwel/poku.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/benchmarking .cursor/skills/benchmarking && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "benchmarking" agent skill from https://github.com/wellwelwel/poku/tree/main/.claude/skills/benchmarking into .cursor/skills/benchmarking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmarking", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/wellwelwel/poku.git --path .claude/skills/benchmarking--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add wellwelwel/poku --skill benchmarking -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install wellwelwel/poku benchmarking --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wellwelwel/poku.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/benchmarking .gemini/skills/benchmarking && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "benchmarking" agent skill from https://github.com/wellwelwel/poku/tree/main/.claude/skills/benchmarking into .gemini/skills/benchmarking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmarking", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install wellwelwel/poku benchmarkingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add wellwelwel/poku --skill benchmarking -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/wellwelwel/poku.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/benchmarking .github/skills/benchmarking && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "benchmarking" agent skill from https://github.com/wellwelwel/poku/tree/main/.claude/skills/benchmarking into .github/skills/benchmarking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmarking", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wellwelwel/poku --skill benchmarking -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install wellwelwel/poku benchmarking --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wellwelwel/poku.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/benchmarking .opencode/skills/benchmarking && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "benchmarking" agent skill from https://github.com/wellwelwel/poku/tree/main/.claude/skills/benchmarking into .opencode/skills/benchmarking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmarking", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
benchmarkingEmpirical performance comparison playbook for poku covering hypothesis lenses, a standalone harness with a byte-identical correctness gate, variant generation, sequential hyperfine measurement…
Benchmarking is an agent skill from wellwelwel/poku. Empirical performance comparison playbook for poku covering hypothesis lenses, a standalone harness with a byte-identical correctness gate, variant generation, sequential hyperfine measurement, composition re-testing, and anomaly forensics. Use when measuring, comparing, or proving the performance of any project logic, when the user asks for benchmarks or optimization validation, or when a micro-optimization claim needs evidence before touching src/.
Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `templates/fixtures.ts`, `templates/gen_variants.py` and `templates/run.ts`).
It sits in Testing & QA. The repository describes itself as: 🐷 Poku makes testing easy for Node.js, Bun, Deno, and you at the same time. The licence is MIT.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit e02df0b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (TypeScript and Python), which the agent can run.
Shell commands in SKILL.md call:
npmbunnodeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Benchmarking loads about 3k tokens when it runs. Until then it costs about 117 tokens; SKILL.md has 1,471 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from wellwelwel/poku at commit e02df0b, republished under its MIT licence (© wellwelwel). 1,471 words, ~3,029 tokens.
.claude/skills/benchmarking/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Target-agnostic methodology to measure and prove micro-optimizations in any hot path of this project. Theory does not count: every claim must survive a byte-identical correctness gate and a statistically clean benchmark before touching src/.
grep -rn "fnName" src/), and the call frequency of each. Frequency decides the verdict: a 5x win on a function called once per run has zero real impact, and a lookup table built at module load never amortizes for it. The final report must weigh gains against real call counts.Three independent lenses, each producing structured hypotheses:
instanceof vs Array.isArray vs getPrototypeOf, typeof chain ordering, inlining budgets (function size), builtin fast paths, deopt triggers.Each hypothesis states: claim, concrete replacement code, the workload where the effect appears, and a behaviorChange boolean. With subagents, run the lenses in parallel (independent and read-only).
One directory per investigation under the gitignored temp/:
mkdir -p ./temp/bench/<target>/impl ./temp/bench/<target>/results
cp .claude/skills/benchmarking/templates/* ./temp/bench/<target>/./temp/bench/<target>/
run.ts static: generic driver, runs unmodified
gen_variants.py static skeleton: fill only the `variants` dict
fixtures.ts per target: fill every TODO(target) marker
impl/baseline.ts per target: faithful port of the current code, zero project imports
impl/<variant>.ts per target: generated, one file per hypothesis, minimal diff over baseline
results/*.json per comparison: hyperfine exportsRunner: bun --bun run.ts ... (~11ms startup). Fallback only when Bun is not available: node --import=tsx run.ts ... (~125ms). Never mix runners between the two sides of a comparison. The runner picks the engine: Bun measures JavaScriptCore, Node + tsx measures V8.
run.ts is generic because fixtures.ts exports a fixed contract: TargetModule (surface of the module under test), verifyCases (string per case), benchModes (iterations plus a checksum-returning run), and resetSideEffects/sideEffectsSnapshot for shared counters.
Requirements:
time bun --bun run.ts bench-x impl/baseline.ts.gen_variants.py generates each variant from the baseline via exact-match replacements with occurrence-count assertions. Fill only the variants dict, copying old blocks verbatim from impl/baseline.ts (indentation included). A not-found or miscounted block aborts generation: that is the protection against silently benchmarking a baseline-identical variant. Never hand-edit near-identical variant files.
One comparison at a time. hyperfine already runs its warmups and runs sequentially, so isolation only breaks at the orchestration level: never start a comparison while another is running.
hyperfine --warmup 3 --runs 10 \
-n baseline 'bun --bun run.ts bench-x impl/baseline.ts' \
-n variant 'bun --bun run.ts bench-x impl/variant.ts' \
--export-json results/variant.json--warmup 3 --runs 10 is the project standard. Same runner on both sides, always.for ... await), never parallel().combined variant and re-measure against baseline: gains do not add up.When a combination behaves unexpectedly, isolate before discarding:
node --trace-deopt ... | grep -c deoptimizing against baseline.--max-opt=0. A tie interpreted plus divergence with JIT on means an optimization-quality cliff, not algorithmic cost. Document the toxic shape.src/**/*.ts symbols under the TargetModule contract, then:bun --bun run.ts verify impl/real-src.tsimpl/baseline.ts. The canonical numbers remain the Phase 4-6 ones.npm run typecheck && npm run lint:fix && npm run build && npm test.Transferable heuristics, measured on Apple Silicon (engine noted when it matters). Re-validate before relying on them in another engine.
Confirmed wins:
getPrototypeOf(x) === Object.prototype check before instanceof chains lets a dominant plain-object path skip them (-3.7%, V8).includes uses a vectorized path and discards most lines in one pass (-19.5%, V8).acc += part beats parts.push(part) + join (-15%, V8).Confirmed non-wins (do not "fix" these):
localeCompare() with no arguments has a V8 fast path. A cached Intl.Collator().compare is 2.8x slower. Default sort() was neutral, so dropping locale order buys nothing.split('\n') loses to a manual indexOf/substring loop (+18%, V8)./\S/.test(line) loses to line.trim().length > 0 as a blank-line test (+6%, V8).Date.prototype.toTimeString().slice() builds a far larger string than three padded getters./ 1eN division with * 1e-N multiplication is a correctness bug, not an optimization: results diverge by 1 ULP on a large share of the domain (engine-independent IEEE behavior).toFixed replacements change rounding on binary-float edges (0.015 formats as 0.01, round-half-up gives 0.02).Set allocation, manual key quoting: all within noise (V8).Documented JIT cliff (V8): memoized indexOf(needle, offset) positions combined with a rope accumulator in the same loop measured 2.1x slower, while each ingredient alone was faster. Identical when interpreted (--max-opt=0), so the combined function shape pessimizes under TurboFan. Re-measure every composition.
./temp/bench/<target>/, fill baseline, fixtures, and calibrationvariants dict, never by hand--warmup 3 --runs 10, sequential, same runner both sides, export JSON--trace-deopt, --max-opt=0)© wellwelwel, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files in .claude/skills/benchmarking of wellwelwel/poku.
Open the folder on GitHubat commit e02df0b
Benchmarking next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Benchmarking this skillwellwelwel/poku | 1.2k | — | ~3k | Automated safety check: Pass | MIT | |
| Web Application Testinganthropics/skills | 180k | 51 repos | ~966 | Automated safety check: Pass | Apache-2.0 | |
| Diagnosing Bugsfossasia/eventyay-interpretation | 1.6k | 32 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| TDDpietheinstrengholt/rssmonster | 564 | 30 repos | ~906 | Automated safety check: Pass | MIT | |
| TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph | 112 | 11 repos | ~2.4k | Automated safety check: Pass | None | |
| TDDsanity-io/sanity | 6.4k | 20 repos | ~1k | Automated safety check: Pass | MIT |
anthropics/skills
Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.
fossasia/eventyay-interpretation
Diagnosis loop for hard bugs and performance regressions. An agent skill from fossasia/eventyay-interpretation.
pietheinstrengholt/rssmonster
Test-driven development. An agent skill from pietheinstrengholt/rssmonster.
hellangleZ/burn-in-cceverywhere-ralph
A skill your agent uses when writing new features, fixing bugs, or refactoring code.
sanity-io/sanity
Test-driven development with red-green-refactor loop. An agent skill from sanity-io/sanity.
Ibrahim-3d/orchestrator-supaconductor
A skill your agent uses when working with Conductor's context-driven development methodology, managing project context artifacts, or understanding the relationship between product.md, tech-stack.md…
wellwelwel/poku
Architecture deep-dive for poku covering project structure, execution flow, config discovery, runtime detection, and competitive context.
wellwelwel/poku
Documentation deep-dive for the poku website covering Docusaurus setup, language policy, feature-to-docs mapping, conventions, JSDoc, and the changelog.
wellwelwel/poku
Engineering deep-dive for poku covering performance, code, and security patterns, TypeScript config, build pipeline, developer experience, and CI/CD.
wellwelwel/poku
Run the mandatory pre-commit checks for the poku repository before staging a commit.
wellwelwel/poku
Testing deep-dive for poku covering test structure, commands, patterns, fixtures, utils, Docker compatibility, and coverage.
Categories
Empirical performance comparison playbook for poku covering hypothesis lenses, a standalone harness with a byte-identical correctness gate, variant generation, sequential hyperfine measurement…. Benchmarking is an agent skill from wellwelwel/poku. Empirical performance comparison playbook for poku covering hypothesis lenses, a standalone harness with a byte-identical correctness gate, variant generation, sequential hyperfine measurement, composition re-testing, and anomaly forensics.
Benchmarking fits situations like: proving the performance of any project logic; the user asks for benchmarks; optimization validation; A micro-optimization claim needs evidence before touching src/.
Run `npx skills add wellwelwel/poku --skill benchmarking -a claude-code`. Or copy the skill folder (.claude/skills/benchmarking in wellwelwel/poku) into .claude/skills/benchmarking in your project. Claude Code loads it when a task matches its description.
Run `npx skills add wellwelwel/poku --skill benchmarking -a codex`. Or copy the skill folder (.claude/skills/benchmarking in wellwelwel/poku) into .agents/skills/benchmarking in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wellwelwel/poku --skill benchmarking -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmarking, .gemini/skills/benchmarking, .github/skills/benchmarking and .opencode/skills/benchmarking in your project.
Going by SKILL.md and its folder, Benchmarking needs TypeScript and Python for the scripts in its folder and the command-line tools its instructions call (npm, bun and node). Our summary lists: Python 3; Node.js.
SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Benchmarking is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Benchmarking: Web Application Testing (anthropics/skills, 180k stars), Diagnosing Bugs (fossasia/eventyay-interpretation, 1.6k stars), TDD (pietheinstrengholt/rssmonster, 564 stars) and TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
wellwelwel (a GitHub user) maintains it in wellwelwel/poku, which has 1,183 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on June 18, 2026.
Source: wellwelwel/poku on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.