Perform quality assurance on code changes after the research-phase - plan-phase - execute-phase workflow.

MITAuto-check: notesTesting & QA

Install QA From Execute

skills CLI
$ npx skills add alchemiststudiosDOTai/harness-engineering --skill qa-from-execute -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install alchemiststudiosDOTai/harness-engineering qa-from-execute --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/alchemiststudiosDOTai/harness-engineering.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/qa-from-execute .claude/skills/qa-from-execute && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
qa-from-execute
GitHub stars
105
Token cost
~2.5k tokens
SKILL.md length
841 words
Files
1
Skills in repo
8
Repo updated
First seen
Licence
MIT

At a glance

Perform quality assurance on code changes after the research-phase - plan-phase - execute-phase workflow.

  • Works in 6 steps: Load Execute Context → Identify Changed Areas → Apply QA Checklist Per Changed Area → …
  • Tasks that involve QA and bug reports
  • SKILL.md covers CRITICAL BOUNDARIES, When to Use, Workflow and Finding Severity Levels, plus 3 more sections
  • Calls mypy, npm and jq

What it does

QA From Execute is an agent skill from alchemiststudiosDOTai/harness-engineering. Perform quality assurance on code changes after the research-phase - plan-phase - execute-phase workflow. STRICTLY QA only—no coding, no fixes, no source-code changes. Focus on changed areas only, emphasizing control/data flow correctness.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering QA and bug reports. The repository describes itself as: harness-engineering discussion of shortcuts, automation, hacks and overall productivity with code agents like claude code, codex, and other harness. The licence is MIT.

When your agent uses it

  • Tasks that involve QA and bug reports

Example prompts

  • “/qa-from-execute”

Requirements

  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Load Execute Context
  2. Identify Changed Areas
  3. Apply QA Checklist Per Changed Area
  4. Test & Contracts Analysis
  5. Secondary Scans (Optional)
  6. Write QA Report

What it can do on your machine

Read from SKILL.md and the folder at commit 7a9fa15. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • mypy
    • npm
    • jq
    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

QA From Execute loads about 2.5k tokens when it runs. Until then it costs about 64 tokens; SKILL.md has 841 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Write, Edit, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from alchemiststudiosDOTai/harness-engineering at commit 7a9fa15, republished under its MIT licence (© alchemiststudiosDOTai). 841 words, ~2,542 tokens.

Download SKILL.mdSave it as .claude/skills/qa-from-execute/SKILL.md (or your agent's skills folder).
name
qa-from-execute
description
Perform quality assurance on code changes after the research-phase -> plan-phase -> execute-phase workflow. STRICTLY QA only—no coding, no fixes, no source-code changes. Focus on changed areas only, emphasizing control/data flow correctness.
allowed-tools
Read, Write, Edit, Bash
writes-to
memory-bank/qa/
hard-guards
QA only - no coding or source-code changes, Write QA output to memory-bank/qa/, Review only the scope captured in memory-bank/execute/

QA From Execute

Evaluate code changes for correctness, risks, and quality. This skill performs read-only analysis of implemented work, producing a QA report without modifying code.

CRITICAL BOUNDARIES

ActivityStatus
QA Analysis✅ This skill
Code Changes❌ NO — Read only
Bug Fixes❌ NO — Report only
Execute❌ NO — Analysis only

This skill is STRICTLY for QA evaluation. Do not write code, do not fix issues, and do not perform the Execute phase. Analyze, evaluate, and report.

When to Use

Use this skill when:

  • The Execute phase is complete
  • Code has been written and needs quality evaluation
  • The task is to assess correctness of changes, not modify them
  • Pre-merge or post-implementation review is needed

Workflow

Step 1: Load Execute Context

Locate and read the execution log:

  • If a path is provided: Read from memory-bank/execute/<path>
  • If a topic is provided: Find the latest matching file in memory-bank/execute/

Extract:

  • Which files were modified
  • Which functions/endpoints were added or changed
  • What the acceptance criteria were
  • Any issues encountered during the Execute phase
Step 2: Identify Changed Areas

From the execution log, build a list of:

  • Files modified: Paths to all changed files
  • Functions changed: Public functions that were added or modified
  • Interfaces changed: API endpoints, CLI commands, public methods
  • State changes: Database schema, configuration, shared resources

Focus analysis ONLY on these changed areas. Do not review unchanged code.

Step 3: Apply QA Checklist Per Changed Area

For each changed file/function/endpoint, evaluate:

3.1 Inputs & Preconditions
CheckQuestion
ValidationAre all inputs validated before use?
Type safetyAre type assumptions explicit and checked?
Null/emptyAre null, undefined, and empty cases handled?
BoundariesAre min/max values, sizes, and limits enforced?
3.2 Control Flow
CheckQuestion
Branch coverageAre all branches reachable? Any dead code?
Fall-throughAre switch/case fall-throughs intentional?
Early returnsAre guard clauses used appropriately?
Loop terminationDo all loops have guaranteed termination?
3.3 Data Flow
CheckQuestion
InvariantsAre invariants preserved through transformations?
Mutation scopeIs mutation limited to appropriate scope?
Shared stateIs shared state access properly synchronized?
AliasingAre aliasing risks (multiple refs to same data) handled?
3.4 State & Transactions
CheckQuestion
IdempotencyIs the operation safe to retry?
AtomicityAre multi-step operations atomic?
RollbackIs there a path to undo partial changes?
ConcurrencyAre race conditions handled?
3.5 Error Handling
CheckQuestion
SpecificityAre exceptions specific (not broad catches)?
Retry logicIs transient failure handled with backoff?
Dead letterAre unprocessable items routed to DLQ/log?
Error contextDo errors include sufficient debugging info?
3.6 Contracts
CheckQuestion
Pre-conditionsAre pre-conditions documented and enforced?
Post-conditionsAre post-conditions guaranteed on success?
Schema driftDo request/response schemas match implementation?
VersioningAre breaking changes properly versioned?
3.7 Time & Locale
CheckQuestion
TimezonesAre datetime operations timezone-aware?
Monotonic timeIs elapsed time measured with monotonic clocks?
DSTAre daylight saving time transitions handled?
Format stabilityAre date/time formats consistent and unambiguous?
3.8 Resource Hygiene
CheckQuestion
File lifecycleAre files opened/closed properly (with statements)?
Connection poolingAre connections returned to pools?
TimeoutsDo all blocking operations have timeouts?
CancellationIs cancellation propagated through async chains?
Show full SKILL.md (335 more words)Show less
3.9 Edge Cases
CheckQuestion
Empty inputsIs empty/null input handled gracefully?
Max sizesAre large inputs bounded (pagination, limits)?
Partial failureIs partial failure detectable and recoverable?
Resource exhaustionAre OOM, disk full, quota exceeded handled?
3.10 Public Surface
CheckQuestion
Backward compatAre breaking changes intentional and documented?
OpenAPI alignmentDo implementations match OpenAPI/JSON schemas?
Type exportsAre public types exported and documented?
DeprecationAre deprecated items marked and alternatives provided?
Step 4: Test & Contracts Analysis

For each changed public function/endpoint:

  1. Map to test coverage

    • Run: pytest -q or equivalent
    • Run: coverage run -m pytest && coverage report --format=markdown
    • Identify which changed functions have tests
  2. Identify missing test cases

    • Error branches: Are failure paths tested?
    • Boundary conditions: Are min/max values tested?
    • Property invariants: Are data guarantees verified?
    • Mutation tests: Would incorrect code fail tests?
  3. Contract/API verification

    • Compare OpenAPI/JSON schema to implementation
    • Verify request/response DTOs match spec
    • Check for breaking field/enum changes
Step 5: Secondary Scans (Optional)

Run static analysis tools (read-only, report results):

bash
# Type checking
mypy . --ignore-missing-imports 2>/dev/null || echo "mypy not available"

# Security scan
bandit -r . -q 2>/dev/null || echo "bandit not available"

# Dependency audit
pip-audit 2>/dev/null || npm audit --json 2>/dev/null | jq '.metadata' || echo "audit not available"

Note findings without attempting fixes.

Step 6: Write QA Report

Create memory-bank/qa/YYYY-MM-DD_HH-MM-SS_<topic>_qa.md:

yaml
---
title: "<topic> – QA Report"
phase: QA
date: "YYYY-MM-DD HH:MM:SS"
owner: "<agent_or_user>"
parent_execute: "memory-bank/execute/<file>.md"
git_commit_at_qa: "<sha>"
tags: [qa, <topic>]
---

## Summary

| Metric | Count |
|--------|-------|
| Files reviewed | N |
| Functions reviewed | N |
| CRITICAL findings | N |
| WARNING findings | N |
| INFO findings | N |
| PASS (no issues) | N |

## Changed Areas Reviewed

### File: `path/to/file.py`

| Function/Class | Lines | Status |
|----------------|-------|--------|
| `function_name()` | L45-89 | ⚠️ WARNING |
| `ClassName` | L120-200 | ✅ PASS |

#### Findings for `function_name()`

| Severity | Category | Finding | Recommendation |
|----------|----------|---------|----------------|
| WARNING | Error Handling | Broad `except Exception` catch | Catch specific exceptions |
| INFO | Data Flow | Mutation of input parameter | Document or avoid |

### File: `path/to/another.js`

...

## Test Coverage Analysis

| Function | Has Tests | Coverage % | Missing Cases |
|----------|-----------|------------|---------------|
| `function_name()` | ✅ | 85% | Error branch, empty input |
| `another_function()` | ❌ | 0% | All cases |

## Contract/API Verification

| Endpoint | Schema Match | Breaking Changes |
|----------|--------------|------------------|
| `POST /api/items` | ✅ | None |
| `GET /api/items/:id` | ⚠️ | New required field |

## Static Analysis Summary

| Tool | Result |
|------|--------|
| mypy | N errors, M warnings |
| bandit | N low, M medium issues |
| pip-audit | N vulnerabilities |

## Risk Assessment

| Risk | Likelihood | Impact | Mitigation Status |
|------|------------|--------|-------------------|
| Race condition in shared state | Medium | High | Not mitigated |
| Missing error branch coverage | High | Medium | Not tested |

## Recommendations Summary

### Must Fix (CRITICAL)
1. [Description of critical issue]

### Should Fix (WARNING)
1. [Description of warning]

### Observations (INFO)
1. [Description of observation]

Finding Severity Levels

LevelDefinitionAction Required
CRITICALSecurity risk, data loss, or system instabilityMust fix before merge
WARNINGPotential bugs, maintainability issues, missing coverageShould fix, can defer
INFOStyle observations, suggestions, notesOptional
PASSNo issues foundNone

Constraints

ConstraintRule
NO CODE CHANGESNever write, modify, or delete code
NO FIXESReport issues, do not implement solutions
FOCUS ON CHANGESOnly review files listed in the execution log
READ-ONLY TOOLSUse tools that don't modify state
DOCUMENT FINDINGSEvery issue must be in the QA report

Subagent Usage

If additional analysis is needed:

With subagents available: Deploy maximum 3:

SubagentWhen to Deploy
antipattern-snifferReview changed code for anti-patterns and code smells
codebase-analyzerDeep analysis of specific function implementations
context-synthesisIdentify hidden dependencies affected by changes

Without subagents: Perform manual analysis following the checklist.

Handoff

After writing the QA report to memory-bank/qa/, hand off to the user for disposition.

Suggested next action:

text
Review memory-bank/qa/<file>.md and decide whether to accept the work or create follow-up planning.

© alchemiststudiosDOTai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/qa-from-execute of alchemiststudiosDOTai/harness-engineering.

Open the folder on GitHubat commit 7a9fa15

Compare with similar skills

QA From Execute next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

QA From Execute compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
QA From Execute this skillalchemiststudiosDOTai/harness-engineering105—~2.5kAutomated safety check: NotesMIT
Reproduce Chat Statesdifferent-ai/openwork24k—~673Automated safety check: PassCustom licence
Dynamo Jira TicketDynamoDS/Dynamo2k—~1.1kAutomated safety check: PassApache-2.0
Minimal Run And Auditlllllllama/RigorPilot-Skills4972 repos~691Automated safety check: PassMIT
Moav E2EMotherofallVPNs/MoaV448—~1.9kAutomated safety check: NotesMIT
Anchor Reprolynxlangya/techne1051 repos~1.2kAutomated safety check: PassMIT

Similar skills

  • Reproduce Chat States

    different-ai/openwork

    Fires known chat states in the running OpenWork desktop app, such as provider errors, retries and tool steps, so you can check how each renders.

    24k GitHub stars~673 tokensUpdated today
    Testing & QAAuto-check passed
  • Dynamo Jira Ticket

    DynamoDS/Dynamo

    Create structured Jira tickets for Dynamo from bug reports, failing tests, or feature requests.

    2k GitHub stars~1.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Minimal Run And Audit

    lllllllama/RigorPilot-Skills

    Rigor Run skill for README-first deep learning repo reproduction.

    497 GitHub starsUsed in 2 repos~691 tokens
    Testing & QAAuto-check passed
  • Moav E2E

    MotherofallVPNs/MoaV

    Run and debug MoaV's end-to-end tests — real protocol connectivity (client-test.sh) and the moav CLI smoke test — against a LIVE server, via the self-hosted e2e workflow or a local test VPS.

    448 GitHub stars~1.9k tokensUpdated yesterday
    Testing & QAAuto-check: notes
  • Anchor Repro

    lynxlangya/techne

    Reproduce a behavioral bug before fixing it, record the failing probe, and verify the fix with the same probe.

    105 GitHub starsUsed in 1 repo~1.2k tokens
    Testing & QAAuto-check passed
  • Creating A Coral Task

    Human-Agent-Society/CORAL

    Author a new CORAL task — the three pieces that must line up (task.yaml, seed/, a packaged grader/), the coral init → coral validate → smoke-test loop, and how to pick a grader pattern (stdout…

    1k GitHub stars~2.2k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed

More from alchemiststudiosDOTai/harness-engineering

All 8 skills in this repo
  • Ast Grep Setup

    alchemiststudiosDOTai/harness-engineering

    Set up ast-grep for a codebase with common TypeScript rules for detecting anti-patterns, enforcing best practices, and preventing bugs.

    105 GitHub stars~4.7k tokensUpdated 6 mo ago
    Auto-check passed
  • Research Phase

    alchemiststudiosDOTai/harness-engineering

    This skill should be used when mapping or researching a codebase to understand its structure, patterns, and architecture.

    105 GitHub stars~1.4k tokensUpdated 6 mo ago
    Auto-check: notes
  • Execute Phase

    alchemiststudiosDOTai/harness-engineering

    Execute implementation plans from .artifacts/plan/. An agent skill from alchemiststudiosDOTai/harness-engineering.

    105 GitHub stars~1.8k tokensUpdated 6 mo ago
    Auto-check passed
  • Harness Map

    alchemiststudiosDOTai/harness-engineering

    Map a repository's mechanical harness layers: canonical check command, local and CI gates, architecture boundaries, structural rules, behavioral verification, docs ratchets, evidence workflows, and…

    105 GitHub stars~1.8k tokensUpdated 6 mo ago
    Auto-check passed
  • Agents Md Mapper

    alchemiststudiosDOTai/harness-engineering

    This skill should be used when creating, refreshing, or validating a repository AGENTS.md so it stays concise, current, and grounded in repository evidence.

    105 GitHub stars~1.8k tokensUpdated 6 mo ago
    Auto-check passed
  • Differential Session Runner

    alchemiststudiosDOTai/harness-engineering

    Run or continue a differential debugging session between two implementations, traces, captures, or outputs.

    105 GitHub stars~1.7k tokensUpdated 6 mo ago
    Auto-check: notes

Categories

Questions about QA From Execute

What does QA From Execute do?

Perform quality assurance on code changes after the research-phase - plan-phase - execute-phase workflow. QA From Execute is an agent skill from alchemiststudiosDOTai/harness-engineering. Perform quality assurance on code changes after the research-phase - plan-phase - execute-phase workflow.

When should I use QA From Execute?

QA From Execute fits situations like: tasks that involve QA and bug reports.

How do I install QA From Execute in Claude Code?

Run `npx skills add alchemiststudiosDOTai/harness-engineering --skill qa-from-execute -a claude-code`. Or copy the skill folder (skills/qa-from-execute in alchemiststudiosDOTai/harness-engineering) into .claude/skills/qa-from-execute in your project. Claude Code loads it when a task matches its description.

How do I install QA From Execute in Codex?

Run `npx skills add alchemiststudiosDOTai/harness-engineering --skill qa-from-execute -a codex`. Or copy the skill folder (skills/qa-from-execute in alchemiststudiosDOTai/harness-engineering) into .agents/skills/qa-from-execute in your project. Codex loads it when a task matches its description.

Can I use QA From Execute in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alchemiststudiosDOTai/harness-engineering --skill qa-from-execute -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/qa-from-execute, .gemini/skills/qa-from-execute, .github/skills/qa-from-execute and .opencode/skills/qa-from-execute in your project.

What does QA From Execute need to run?

Going by SKILL.md and its folder, QA From Execute needs the command-line tools its instructions call (mypy, npm, jq and pytest). Its frontmatter pre-approves these tools: Read, Write, Edit, Bash.

Does QA From Execute access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is QA From Execute safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does QA From Execute use?

QA From Execute is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does QA From Execute use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to QA From Execute?

Skills that share tags, products or a category with QA From Execute: Reproduce Chat States (different-ai/openwork, 24k stars), Dynamo Jira Ticket (DynamoDS/Dynamo, 2k stars), Minimal Run And Audit (lllllllama/RigorPilot-Skills, 497 stars) and Moav E2E (MotherofallVPNs/MoaV, 448 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains QA From Execute?

alchemiststudiosDOTai (a GitHub organization) maintains it in alchemiststudiosDOTai/harness-engineering, which has 105 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on March 17, 2026.

Source: alchemiststudiosDOTai/harness-engineering on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.