Agent skill

QA Session

by getlago in getlago/lago-front

A skill your agent uses when the operator wants to verify in the browser that a change works — asks "how do I test this locally?", "give me the steps", "QA this", references a worktree/PR/ticket…

AGPL-3.0Auto-check passedTesting & QA

Install QA Session

skills CLI
$ npx skills add getlago/lago-front --skill qa-session -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install getlago/lago-front qa-session --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/getlago/lago-front.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/qa-session .claude/skills/qa-session && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
qa-session
GitHub stars
163
Token cost
~3.8k tokens
SKILL.md length
2,086 words
Files
1
Skills in repo
17
Repo updated
First seen
Licence
AGPL-3.0

At a glance

A skill your agent uses when the operator wants to verify in the browser that a change works — asks "how do I test this locally?", "give me the steps", "QA this", references a worktree/PR/ticket…

  • Works in 8 steps: Args (both required) → Resolve the app under test → Preconditions → …
  • The operator wants to verify in the browser that a change works — asks how do I test this locally?
  • SKILL.md covers Step 0 — Args (both required), Step 1 — Resolve the app under…, Step 2 — Preconditions and Step 3 — Test plan (UI…, plus 5 more sections
  • Calls git, docker and gh

What it does

QA Session is an agent skill from getlago/lago-front. Use when the operator wants to verify in the browser that a change works — asks "how do I test this locally?", "give me the steps", "QA this", references a worktree/PR/ticket built in this session, or reports something not working during manual testing (button does nothing, section missing, block blank).

Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering QA and bug reports and Git worktrees. It works with Playwright. The repository describes itself as: Open Source Metering and Usage Based Billing. The licence is AGPL-3.0.

When your agent uses it

  • The operator wants to verify in the browser that a change works — asks how do I test this locally?
  • Give me the steps
  • References a worktree/PR/ticket built in this session
  • Reports something not working during manual testing (button does nothing

Example prompts

  • “how do I test this locally?”
  • “give me the steps”
  • “QA this”
  • “/qa-session”

Requirements

  • Docker

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Args (both required)
  2. Resolve the app under test
  3. Preconditions
  4. Test plan (UI language only)
  5. Control run on main (mandatory, do it FIRST)
  6. Verify the fix
  7. Triage anomalies live
  8. Record, then the verdict LAST

What it can do on your machine

Read from SKILL.md and the folder at commit 79b5b3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • docker
    • gh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, docker and gh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

QA Session loads about 3.8k tokens when it runs. Until then it costs about 79 tokens; SKILL.md has 2,086 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~79
When it runs · the whole SKILL.md, loaded when a task matches
~3.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from getlago/lago-front at commit 79b5b3d, republished under its AGPL-3.0 licence (© getlago). 2,086 words, ~3,844 tokens.

Download SKILL.mdSave it as .claude/skills/qa-session/SKILL.md (or your agent's skills folder).
name
qa-session
description
Use when the operator wants to verify in the browser that a change works — asks "how do I test this locally?", "give me the steps", "QA this", references a worktree/PR/ticket built in this session, or reports something not working during manual testing (button does nothing, section missing, block blank).

QA Session — browser verification of a change

Turn a diff into UI steps, reproduce the old bug on main, then prove the fix removes it. Every step is derived from the code (routes, labels, gates) — never guessed.

What QA has to disprove is the reported symptom, not the diff. A fix can be correct, its tests green, its own control run red-then-green, and still leave the ticket's bug in place — the diff answers the analysis, and the analysis can have found a real but different defect. Everything below is built so that failure mode cannot end in a PASS.

Step 0 — Args (both required)

  1. Mode: manual (operator clicks, you give steps + triage) or auto (you drive the browser).
  2. Target: ISSUE-ID, PR number, branch/worktree, or free-text description of the surface to test.

Either missing → AskUserQuestion, stop until answered.

auto mode drives two browsers and they are not interchangeable. Never the built-in Browser pane, and never type credentials in either.

  • Claude in Chrome (mcp__claude-in-chrome__*) carries the operator's session, so it reaches the app with no login. Use it for navigation, DOM and computed-style probes, and plain clicks. It has NO middle-click action at all, and its input channel can die mid-run, delivering zero events to the page while every call still returns success. Prefer browser_batch; re-read refs after every navigation (stale refs click the backdrop and close drawers).
  • Playwright (mcp__plugin_playwright_playwright__*) has real CDP input: browser_click takes button: "middle" and modifiers: ["ControlOrMeta"], browser_press_key sends genuine keypresses, browser_tabs lists what a gesture opened. Its browser is a separate profile with NO app session, so ask the operator to log in there once — after that the whole real-gesture round runs unattended.

Any check whose verdict depends on a browser default — middle-click, modifier-click, a keypress — MUST run on Playwright. A dispatched MouseEvent or KeyboardEvent exercises the app's handlers and nothing else, so a round driven that way has verified the handler, not the gesture: report it as such.

Prove the instrument before trusting a negative. Install a capture listener, fire one harmless click, and confirm an event actually reached the page. A dead input channel reads exactly like an app ignoring the click.

Step 1 — Resolve the app under test

Resolve the target down to a worktree name first — that name is what the slot registry is keyed on and what the container names derive from:

  • ISSUE-ID → $LOOP_STATE_DIR/<ISSUE-ID>/state.md (default ~/.claude/loop-state) for worktree, branch, port.
  • PR number or URL → its head branch, then the state dir whose branch: matches, same as loop-revise:
    bash
    BRANCH=$(gh pr view <PR> --json headRefName --jq .headRefName)
    grep -l "^branch: ${BRANCH}$" "${LOOP_STATE_DIR:-$HOME/.claude/loop-state}"/*/state.md
  • Branch, worktree name or free-text → the slot registry at the monorepo root, i.e. ../.worktree-slots from the front/ checkout (name:front_port:front_base:api_port:api_base), else docker ps.

lago-worktree keys a slot by the branch with / replaced by -, so normalize before looking one up — WT=${BRANCH//\//-} — otherwise every slashed branch misses its own slot and reads as "not running".

  • Front URL http://localhost:<front_port>; containers lago_front_wt_${SAN} and lago_api_wt_${SAN}, with SAN=$(echo "$WT" | tr '[:upper:]' '[:lower:]' | sed 's/[^a-z0-9]/_/g'). An empty api_port in the slot means the worktree has no API of its own and proxies the shared stack, whose API is lago_api_dev.
  • Stale code / 504 Outdated Optimize Dep / blank page → clear vite cache THEN restart (first load is slow):
    bash
    docker exec <container> sh -c 'rm -rf /app/node_modules/.vite' || true
    docker restart <container>

Step 2 — Preconditions

Read the ticket's own repro steps first, then the diff + spec — in that order, so the report frames the diff and not the other way round. Enumerate the ticket's attachments (get_issue → attachments, plus list_comments): a video, screenshot or Slack thread often carries the only repro there is. You cannot watch a video — say so explicitly, and ask the operator for the gesture and the surface rather than substituting a plausible one.

No ticket (PR, branch or free-text target): the report is whatever the operator handed over — PR body, their own message. Ask for the exact gesture, payload and surface when it is not there; never substitute a plausible one. Everywhere below, "ticket" and "spec" mean that source instead, and with no acceptance criteria the AC control (Step 4a) collapses into the report control — say so rather than deriving ACs from the diff.

Then check what gates the surface:

  • Feature flag — featureFlag: FeatureFlagEnum.X on the route in src/core/router/*; enum value = DB string. Flip it on the organization whose slug is in the URL under test, in the API container resolved in Step 1 (lago_api_wt_${SAN}, or lago_api_dev on the shared stack). Organization.first is an arbitrary row: it can enable the flag on an org nobody is testing and leave the tested route gated.
    bash
    docker exec -it <api_container> bin/rails runner 'o = Organization.find_by!(slug: "<slug>"); o.update!(feature_flags: (o.feature_flags | ["<flag>"])); puts o.feature_flags.inspect'
  • Permissions / premium — permissions: on the route, premiumIntegrations gates.
  • Data — what must exist (customer, plan, wallet) and the org slug (URLs are /<slug>/...).

Step 3 — Test plan (UI language only)

Numbered steps, each with: exact URL · exact visible label (button, menu item, drawer title, read from the component and translations/base.json) · expected result · old-bug behavior. No internals in the steps.

  • Check 1 is always the reporter's own path — their steps, their payload, their surface, described in their words. Diff-derived checks come after it, and are secondary.
  • If the spec's acceptance criteria don't cover the reporter's path, that is already a finding: say so in the plan, and treat the report as the thing to satisfy. An AC set that only describes the diff means the analysis may have scoped a different defect.
  • Target the exact code path the diff touched. Same feature ≠ same surface (in-editor preview toggle vs saved-version preview rebuild). Name which surface proves the fix.
  • Test on the product surface, not a dev harness. /design-system/* pages differ structurally from the real one — they grow instead of scrolling, have no fixed-height/overflow-auto wrapper, no aside, no real data. Those differences are exactly what hides layout, scroll and clipping defects. Using a harness for speed is fine; list the structural deltas, then re-run check 1 on the real surface.
  • Input fixtures: when the bug depends on content (paste, import, file), give the exact content, crafted against the diff — a generic fixture can silently fail to reproduce.
  • Declared deltas: pull intentional visual changes and known follow-ups from the spec/PR in as "expected — not a bug" lines, so the operator doesn't file them.

Step 4 — Control run on main (mandatory, do it FIRST)

Reproduce the bug on unfixed code before verifying the fix — otherwise a PASS proves nothing. Two separate controls, both on main:

  • (a) AC control — the spec's failing check. Proves the fix does what it claims.
  • (b) Report control — the reporter's own steps, payload and surface. Proves the fix addresses what was actually filed.

They are not interchangeable, and (a) passing says nothing about (b). If (b) does not reproduce on main — the reported gesture works fine on unfixed code — then the analysis found a real defect that is not the reported one: stop, say it plainly, and hunt the reported symptom before signing anything off. Verdict is capped at PARTIAL until (b) reproduces.

Same trick both times:

  1. Preferred (same worktree, git status --short must be empty) — revert the whole branch diff, which git checkout main -- <paths> cannot do for paths the branch adds, deletes or renames:

    bash
    git diff --binary main...HEAD > /tmp/qa-control.patch   # --binary: without it an image or font in the diff won't re-apply
    git apply -R /tmp/qa-control.patch   # worktree now runs main's code, vite HMR reloads
    # run the failing check, confirm the old behavior
    git apply /tmp/qa-control.patch      # back to the branch code — mandatory, see below

    Re-applying the same patch forward is what restores adds, deletes and renames symmetrically.

    The forward re-apply is not optional and never skipped. Every way out of the control run goes through it: the check reproducing, not reproducing, erroring, the browser leg dying mid-run in auto mode, and the stop above when (b) doesn't reproduce. Skip it and the branch is left running main's code under a dirty tree that reads as the operator's own uncommitted edits. Restore, confirm git status --short is empty, and only then report — including when reporting a stop.

  2. Alternative: another running worktree whose branch doesn't touch those files (git diff main...HEAD --stat) — they share the DB, so the same fixture URL works on both ports.

Record which route was used. If neither is possible, say so plainly and do not claim causality.

Show full SKILL.md (785 more words)Show less

Step 5 — Verify the fix

Same steps on the fixed app. Screenshot every expected-result checkpoint. Include a reload check when the value also arrives from a second path (hydration, refetch), so the fix isn't masking a regression there.

  • Presence is not visibility. !!document.querySelector('[data-test=x]') passes on an element rendered off-screen, clipped or collapsed. For anything floating (menu, toolbar, popper, tooltip, drawer) assert the geometry: rect inside the visible box of its scroll container, non-zero size, not covered.
  • When two code paths end at the same URL, the URL cannot tell them apart. Intercept history.pushState and history.replaceState and count the calls: that is what separates one navigation from two, and a suppressed navigation from a replace onto an identical target.
  • Scroll is a test dimension, not a detail: run the check at scrollTop 0 and with the container scrolled. An absolutely-positioned overlay inside a scrolled position: relative container is a standing trap — its offset must include scrollTop/scrollLeft, and at scroll 0 a broken one looks perfect.

Step 6 — Triage anomalies live

Read the handler BEFORE calling anything a bug.

SymptomCheck first
Button no-ops silentlyearly-return validation gate in the save handler
Menu entry / page missingfeature flag or permission on the route
Blank page, 504 Outdated Optimize Depvite cache → Step 1
Stale behavior after a rebuildcontainer running old code → restart
A real gesture does nothing at allprove the input channel is alive before blaming the app (Step 0)
Middle-click appears to move the current tab toore-run it once the page has settled: a click landing mid-hydration can do both
Click does nothing on part of a blockhitbox is the inner content, not the row
Menu / popper / tooltip "never opens"it may be in the DOM but positioned out of view — compare its rect with the scroll container's, check offsetParent, scrollTop in the offset math, clipping and z-index
Looks off (spacing, alignment)measure from the CSS source and fix the computed delta, never by eye
A control states an absence ("no default", "none available")re-read once the data settles — it must render nothing while its query is in flight, never assert the negative early
A raw enum or a repeated label in the UI (stripe, label == sublabel)a display mapping was skipped; find the value's label map instead of printing the field

src/ edits (CSS included) hot-reload — the operator just reloads. Restart only for dependency/config changes.

Step 7 — Record, then the verdict LAST

Append results to $LOOP_STATE_DIR/<ISSUE-ID>/qa.md, or $LOOP_STATE_DIR/qa/<branch-or-PR>.md when there is no ISSUE-ID: one row per check, mode, PASS/FAIL, the control outcome, what was deliberately not covered, anomalies with root cause. Genuine side-findings → propose as a separate ticket; never fix unasked.

Parity with main is not a finding. Before reporting anything as a defect, a risk, or a decision for the operator, run the same check on main. A gap that behaves identically there is out of scope: say so once and close it, never escalate it as a choice to be made. The same yardstick applies to a fix of your own — audit its blast radius, and when a defect can be corrected either at the call site or in a shared component, choose the call site.

Then close the reply with this block as the very last thing — nothing after it:

## Verdict: PASS | PARTIAL | FAIL
<one line why>
Control on main (AC): bug reproduced | not reproduced | not possible (<reason>)
Control on main (reported repro, reporter's own steps): reproduced | not reproduced | not possible (<reason>)
Surface: <product surface tested> (<dev harness only, if that is all that was covered>)
Next: <nothing to do | what needs another round | what is still broken>
  • PASS — every check passed, both controls reproduced on main, and check 1 ran on the product surface.
  • PARTIAL — fix works but a check was blocked, either control couldn't run, the reported repro didn't reproduce on main, only a dev harness was covered, or something new surfaced → another round needed.
  • FAIL — the target behavior is still broken on the fixed app.

Common mistakes

  • Skipping the control run, so "it works" doesn't prove the fix did it.
  • Building the control fixture from the spec's ACs instead of the reporter's steps — it then proves the fix matches the analysis, which is the one thing never in doubt.
  • Signing off a fix for a defect nobody reported while the filed symptom is still there.
  • Asserting an element exists instead of asserting it is visible where the user looks.
  • Verifying only on /design-system/*, whose layout can't reproduce the product surface's scroll or clipping.
  • auto mode on the built-in pane, or typing credentials in either browser.
  • Driving a real-gesture check with dispatched events, then reporting the gesture as verified.
  • Calling a check failed when it was the harness's input channel that was dead.
  • Reporting behaviour main already had as a finding, or handing it to the operator as a decision.
  • Guessing URLs and labels instead of reading routes, components, and translations.
  • Proving the fix on the wrong surface.
  • Calling a silent validation gate a bug.
  • Burying the verdict in the middle of the reply.

© getlago, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/qa-session of getlago/lago-front.

Open the folder on GitHubat commit 79b5b3d

Compare with similar skills

QA Session next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

QA Session compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
QA Session this skillgetlago/lago-front163—~3.8kAutomated safety check: PassAGPL-3.0
Diagnose Playwright Failure as Product Bugappsmithorg/appsmith41k—~1.5kAutomated safety check: PassApache-2.0
Evidence-Driven Testingmichaelshimeles/skills1.3k1 repos~3.9kAutomated safety check: PassNone
Verify Omnigent End-to-Endomnigent-ai/omnigent11k—~1.6kAutomated safety check: PassApache-2.0
Verifymorapelker/hive470—~709Automated safety check: PassMIT
Hands On Testktnyt/cclsp675—~1.7kAutomated safety check: PassMIT

Similar skills

  • Investigates a stubbornly failing Playwright test as a possible product bug, using error output, screenshots, traces and server code, and writes a structured bug report.

    41k GitHub stars~1.5k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Evidence-Driven Testing

    michaelshimeles/skills

    Records an annotated screen recording of the agent testing an app hands-on, then posts the video and a results summary to the PR and tracker issue.

    1.3k GitHub starsUsed in 1 repo~3.9k tokens
    Testing & QAAuto-check passed
  • Verify Omnigent End-to-End

    omnigent-ai/omnigent

    Spins up an isolated Omnigent server, runner and mock model to prove a user-facing behavior or bug fix with recorded evidence instead of reasoning from code.

    11k GitHub stars~1.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Verify

    morapelker/hive

    Build, launch, and drive this worktree's Hive app over CDP to verify a change end-to-end with playwright-cli

    470 GitHub stars~709 tokensUpdated 8 days ago
    Testing & QAAuto-check passed
  • Hands On Test

    ktnyt/cclsp

    Performs manual hands-on testing of a web application using playwright-cli.

    675 GitHub stars~1.7k tokensUpdated 7 mo ago
    Testing & QAAuto-check passed
  • Glance Test

    DebugBase/glance

    Run E2E browser tests on any web application using Glance MCP.

    156 GitHub stars~827 tokensUpdated 5 mo ago
    Testing & QAAuto-check passed

More from getlago/lago-front

All 17 skills in this repo
  • Babysit

    getlago/lago-front

    A skill your agent uses when asked to babysit, monitor, shepherd, or keep working on a GitHub pull request until it is green, review-ready, approved, mergeable, or ready to merge.

    163 GitHub stars~5.2k tokensUpdated today
    Auto-check passed
  • Cve Doctor

    getlago/lago-front

    Triage a CVE / Dependabot alert in a JS/TS project and recommend the least-invasive fix.

    163 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Extract Section To Drawer

    getlago/lago-front

    Extract a Formik form section into a TanStack Form drawer with Zod validation, following the plan form migration pattern.

    163 GitHub stars~4k tokensUpdated today
    Auto-check: notes
  • Loop Build

    getlago/lago-front

    Phase 2 of the loop pipeline for lago-front. An agent skill from getlago/lago-front.

    163 GitHub stars~3.5k tokensUpdated today
    Auto-check: notes
  • Loop Clean

    getlago/lago-front

    Cleanup phase of the loop pipeline for lago-front, for the worktree layout only.

    163 GitHub stars~816 tokensUpdated today
    Auto-check passed
  • Loop Flywheel

    getlago/lago-front

    Harvest phase of the loop pipeline for lago-front. An agent skill from getlago/lago-front.

    163 GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about QA Session

What does QA Session do?

A skill your agent uses when the operator wants to verify in the browser that a change works — asks "how do I test this locally?", "give me the steps", "QA this", references a worktree/PR/ticket…. QA Session is an agent skill from getlago/lago-front.", "give me the steps", "QA this", references a worktree/PR/ticket built in this session, or reports something not working during manual testing (button does nothing, section missing, block blank).

When should I use QA Session?

QA Session fits situations like: the operator wants to verify in the browser that a change works — asks how do I test this locally?; give me the steps; references a worktree/PR/ticket built in this session; reports something not working during manual testing (button does nothing.

How do I install QA Session in Claude Code?

Run `npx skills add getlago/lago-front --skill qa-session -a claude-code`. Or copy the skill folder (.agents/skills/qa-session in getlago/lago-front) into .claude/skills/qa-session in your project. Claude Code loads it when a task matches its description.

How do I install QA Session in Codex?

Run `npx skills add getlago/lago-front --skill qa-session -a codex`. Or copy the skill folder (.agents/skills/qa-session in getlago/lago-front) into .agents/skills/qa-session in your project. Codex loads it when a task matches its description.

Can I use QA Session in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add getlago/lago-front --skill qa-session -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/qa-session, .gemini/skills/qa-session, .github/skills/qa-session and .opencode/skills/qa-session in your project.

What does QA Session need to run?

Going by SKILL.md and its folder, QA Session needs the command-line tools its instructions call (git, docker and gh). Our summary lists: Docker.

Does QA Session access the network?

SKILL.md contains no URLs. Its commands use git, docker and gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is QA Session safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does QA Session use?

QA Session is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does QA Session use?

About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to QA Session?

Skills that share tags, products or a category with QA Session: Diagnose Playwright Failure as Product Bug (appsmithorg/appsmith, 41k stars), Evidence-Driven Testing (michaelshimeles/skills, 1.3k stars), Verify Omnigent End-to-End (omnigent-ai/omnigent, 11k stars) and Verify (morapelker/hive, 470 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains QA Session?

getlago (a GitHub organization) maintains it in getlago/lago-front, which has 163 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on October 7, 2026.

Source: getlago/lago-front on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.