Agent skill

Run Smoke Tests

by tamdogood in tamdogood/builder-essential-skills

Inspect an unfamiliar repository, interpret a broad Markdown user journey at runtime, operate the real product through its supported web, API, CLI, desktop, or mobile surface, and produce an…

MITAuto-check passedTesting & QA

Install Run Smoke Tests

skills CLI
$ npx skills add tamdogood/builder-essential-skills --skill run-smoke-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tamdogood/builder-essential-skills run-smoke-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tamdogood/builder-essential-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/run-smoke-tests .claude/skills/run-smoke-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
run-smoke-tests
GitHub stars
221
Token cost
~1.9k tokens
SKILL.md length
883 words
Files
7 (incl. references)
Skills in repo
18
Repo updated
First seen
Licence
MIT

At a glance

Inspect an unfamiliar repository, interpret a broad Markdown user journey at runtime, operate the real product through its supported web, API, CLI, desktop, or mobile surface, and produce an…

  • Works in 6 steps: Pass the repository-understanding gate → Validate the journey contract → Preflight without changing product… → …
  • Asked to run smoke tests
  • SKILL.md covers Operating contract, Workflow, Demonstration examples and Failure handling
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Run Smoke Tests is an agent skill from tamdogood/builder-essential-skills. Inspect an unfamiliar repository, interpret a broad Markdown user journey at runtime, operate the real product through its supported web, API, CLI, desktop, or mobile surface, and produce an auditable pass, fail, or blocked judgment with screenshots, logs, recordings, and a step timeline when available. Use when asked to run smoke tests, validate an end-to-end user journey, dogfood a product, test a release candidate, or execute a non-deterministic workflow that may include authentication or human checkpoints. Do…

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including reference files (for example `README.md`, `agents/openai.yaml` and `examples/developer-cli-first-run.smoke.md`).

It sits in Testing & QA, covering QA and bug reports and Customer journey mapping. The repository describes itself as: A repository for skills that are essential to my daily work. The licence is MIT.

When your agent uses it

  • Asked to run smoke tests
  • Validate an end-to-end user journey
  • Dogfood a product
  • Test a release candidate

Example prompts

  • “/run-smoke-tests”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Pass the repository-understanding gate
  2. Validate the journey contract
  3. Preflight without changing product behavior
  4. Interpret and execute the Markdown
  5. Judge once, retry carefully
  6. Clean up and hand off

What it can do on your machine

Read from SKILL.md and the folder at commit 1be9984. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Run Smoke Tests loads about 1.9k tokens when it runs, and up to ~2.5k if it reads all its reference files. Until then it costs about 154 tokens; SKILL.md has 883 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~154
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from tamdogood/builder-essential-skills at commit 1be9984, republished under its MIT licence (© tamdogood). 883 words, ~1,907 tokens.

Download SKILL.mdSave it as .claude/skills/run-smoke-tests/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
run-smoke-tests
description
Inspect an unfamiliar repository, interpret a broad Markdown user journey at runtime, operate the real product through its supported web, API, CLI, desktop, or mobile surface, and produce an auditable pass, fail, or blocked judgment with screenshots, logs, recordings, and a step timeline when available. Use when asked to run smoke tests, validate an end-to-end user journey, dogfood a product, test a release candidate, or execute a non-deterministic workflow that may include authentication or human checkpoints. Do not use as a substitute for deterministic unit, integration, or scenario tests.

Run Smoke Tests

Execute a broad user journey against a real build and leave enough evidence for a person to audit the judgment. Understand the host project before starting the product or touching its state.

Operating contract

  • Treat the Markdown as runtime instructions, not source to compile into code.
  • Test through a supported user-facing surface: web UI, API, CLI, desktop, mobile, or a documented combination.
  • Separate execution from repair. Do not edit the product, tests, or journey during a run. Finish the report before starting any requested fix.
  • Default to local, test, preview, or staging environments. Never mutate production, spend money, send real invitations, or change real user data without explicit authorization.
  • Make every pass or fail claim traceable to captured evidence.
  • Keep the first failed run. Do not erase evidence by retrying until green.

Workflow

1. Pass the repository-understanding gate

Read the nearest repository instructions, product README, runbook, contribution guide, manifests, launch scripts, existing smoke or end-to-end tests, and the smallest relevant product documentation. Search for auth setup, seed data, test accounts, feature flags, observability, artifact conventions, and cleanup commands. Trace the journey across the relevant UI, service, state, and external-system boundaries so failures can be attributed to the right part of the product.

Before launching anything, be able to state:

text
product and user surface:
authoritative journey source:
relevant architecture and system boundaries:
build and start commands:
target environment and data boundary:
interaction driver or available tools:
authentication and human checkpoints:
evidence capture capabilities:
safe cleanup path:
known abort conditions:

Do not guess missing commands, URLs, credentials, or permissions. If a missing item blocks safe execution, report it as a preflight blocker.

2. Validate the journey contract

Read references/smoke-format.md. Normalize supplied prose without silently adding product behavior.

Confirm that the journey:

  • represents one coherent user goal;
  • names its environment and allowed mutations;
  • defines critical and noncritical oracles;
  • identifies authentication, payment, security, or destructive checkpoints;
  • has a cleanup or quarantine plan;
  • names conditions that require an immediate abort.

Move small deterministic behaviors into scenario tests. Keep smoke coverage for the larger path and the seams between systems.

3. Preflight without changing product behavior

Build or launch the product with documented commands. Confirm health, test-data availability, driver connectivity, recording support, and enough storage for artifacts.

Prefer the repository's existing artifact directory. Otherwise use .context/smoke-runs/ when .context already exists; if it does not, create a temporary directory and report its absolute path. Do not add generated evidence to version control unless the user asks.

Create a run manifest containing:

text
journey ID and source path
commit or build identifier
environment and base URL or executable
start time and executor
human checkpoints
recording and log sources

Abort before step one when the build is unhealthy, the environment is wrong, the account or data boundary is unsafe, or required evidence capture is unavailable for a critical oracle.

4. Interpret and execute the Markdown

Follow the journey in order through the real product surface. For each step, record:

text
timestamp
instruction
action taken
observable result
evidence path or log reference
provisional judgment
deviation, latency, or uncertainty

Use normal user affordances. Do not call internal APIs to bypass UI steps unless the journey explicitly tests an API or uses setup APIs only for isolated fixtures.

Pause at a declared human checkpoint. State the exact action needed and resume from the same state after the human completes it. Never request that a user paste secrets into chat.

If the driver supports video, record the complete run and generate subtitles or a timestamped transcript from the step timeline. Otherwise capture screenshots, terminal output, structured logs, and state snapshots. Do not claim to have a video when the environment cannot produce one.

Show full SKILL.md (353 more words)Show less
5. Judge once, retry carefully

Assign one primary result:

  • PASS: every critical oracle was observed and no abort condition occurred;
  • FAIL: a critical oracle was violated or an unexpected product error prevented the goal;
  • BLOCKED: an external prerequisite, permission, environment, or human checkpoint prevented a valid judgment.

Add FLAKY as a flag when the same build and inputs produce inconsistent results.

Do not convert missing evidence into a pass. A noncritical issue may preserve a pass only when the report calls it out separately.

Retry at most once, and only when evidence points to an environmental or driver failure rather than a product failure. Preserve both runs and explain why the retry was allowed.

6. Clean up and hand off

Run only the cleanup authorized by the journey. If cleanup fails, quarantine the test data, record identifiers, and report the residue. Do not hide a cleanup failure behind a passing product judgment.

Write a report with:

  1. result, confidence, build, environment, and duration;
  2. a timestamped step table;
  3. critical oracle results with evidence links;
  4. human checkpoints and deviations;
  5. console, network, log, screenshot, and recording locations;
  6. cleanup result and residual data;
  7. candidate deterministic tests exposed by the run.

The final user message must distinguish product failures, test-environment failures, and uncertain judgments.

Demonstration examples

Use the examples to demonstrate runtime interpretation. Replace every command, URL, account, and expected result with evidence from the host repository.

Failure handling

  • Product cannot start: capture build and launch output, mark BLOCKED, and do not improvise a different environment.
  • Authentication unavailable: pause at the checkpoint or mark BLOCKED; never bypass it with unapproved credentials.
  • Product behavior diverges from the Markdown: preserve evidence and mark FAIL unless the behavior source is genuinely contradictory.
  • Evidence capture fails mid-run: abort when it affects a critical oracle; otherwise continue with the limitation visible.
  • A defect is found: finish the report first. Diagnose or fix only under a separate user request.

© tamdogood, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (references) in skills/run-smoke-tests of tamdogood/builder-essential-skills.

  • SKILL.md
  • README.md
  • agents/openai.yaml
  • examples/developer-cli-first-run.smoke.md
  • examples/saas-team-onboarding.sample-report.md
  • examples/saas-team-onboarding.smoke.md
  • references/smoke-format.md

Open the folder on GitHubat commit 1be9984

Compare with similar skills

Run Smoke Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Run Smoke Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Run Smoke Tests this skilltamdogood/builder-essential-skills221—~1.9kAutomated safety check: PassMIT
Verify Bbget-bb/bb4.2k—~3.2kAutomated safety check: PassMIT
Plugin TestNanmiCoder/dsh-auto-mode1641 repos~2.9kAutomated safety check: PassMIT
Solution Testingkid-sid/claude-spellbook190—~3.9kAutomated safety check: PassMIT
Reproduce Chat Statesdifferent-ai/openwork24k—~673Automated safety check: PassCustom licence
Dynamo Jira TicketDynamoDS/Dynamo2k—~1.1kAutomated safety check: PassApache-2.0

Similar skills

  • Verify Bb

    get-bb/bb

    Verify BB user journeys in an isolated source dev app using dev-browser@next and the matching source CLI.

    4.2k GitHub stars~3.2k tokensUpdated today
    Testing & QAAuto-check passed
  • Plugin Test

    NanmiCoder/dsh-auto-mode

    A skill your agent uses when writing or reviewing tests for DeepSeek Harness plugins, external DSH plugin packages, or package changes in the deepseek-harness repository.

    164 GitHub starsUsed in 1 repo~2.9k tokens
    Testing & QAAuto-check passed
  • Solution Testing

    kid-sid/claude-spellbook

    A skill your agent uses when writing Playwright E2E tests for critical user journeys, setting up post-deployment smoke tests, debugging flaky browser automation, or implementing BDD feature files…

    190 GitHub stars~3.9k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • Reproduce Chat States

    different-ai/openwork

    Fires known chat states in the running OpenWork desktop app, such as provider errors, retries and tool steps, so you can check how each renders.

    24k GitHub stars~673 tokensUpdated today
    Testing & QAAuto-check passed
  • Dynamo Jira Ticket

    DynamoDS/Dynamo

    Create structured Jira tickets for Dynamo from bug reports, failing tests, or feature requests.

    2k GitHub stars~1.1k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Moav E2E

    MotherofallVPNs/MoaV

    Run and debug MoaV's end-to-end tests — real protocol connectivity (client-test.sh) and the moav CLI smoke test — against a LIVE server, via the self-hosted e2e workflow or a local test VPS.

    449 GitHub stars~1.9k tokensUpdated 3 days ago
    Testing & QAAuto-check: notes

More from tamdogood/builder-essential-skills

All 18 skills in this repo
  • Create Marketing Kit

    tamdogood/builder-essential-skills

    Create truthful, human-centered marketing campaigns for an app or product, including positioning, channel copy, original artwork, editable layouts, README banners, and selective website integration.

    221 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Name Your Business

    tamdogood/builder-essential-skills

    Generate, refine, compare, and when needed validate distinctive names for startups, AI products, developer tools, protocols, open-source projects, apps, product families, local businesses, services…

    221 GitHub stars~4.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Session Profiler

    tamdogood/builder-essential-skills

    Profile and debug Hermes sessions from their JSONL transcripts.

    221 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Build Scenario Tests

    tamdogood/builder-essential-skills

    Inspect an unfamiliar repository, turn a focused Markdown behavior scenario into a deterministic test in the repository's native test stack, run it, and preserve traceability between intent and code.

    221 GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Create Skill

    tamdogood/builder-essential-skills

    Create or update a complete repository skill from a user's idea, including the workflow instructions, references, scripts or assets, agent metadata, skill-card artwork, cinematic banner artwork…

    221 GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Paper Opportunity Radar

    tamdogood/builder-essential-skills

    Run a cumulative daily or retrospective sweep of research papers on a chosen topic, audit their claims, methods, integrity signals, and independent support, then identify overlooked but feasible…

    221 GitHub stars~3.1k tokensUpdated 1 mo ago
    Auto-check passed

Categories

Questions about Run Smoke Tests

What does Run Smoke Tests do?

Inspect an unfamiliar repository, interpret a broad Markdown user journey at runtime, operate the real product through its supported web, API, CLI, desktop, or mobile surface, and produce an…. Run Smoke Tests is an agent skill from tamdogood/builder-essential-skills. Inspect an unfamiliar repository, interpret a broad Markdown user journey at runtime, operate the real product through its supported web, API, CLI, desktop, or mobile surface, and produce an auditable pass, fail, or blocked judgment with screenshots, logs, recordings, and a step timeline when available.

When should I use Run Smoke Tests?

Run Smoke Tests fits situations like: asked to run smoke tests; validate an end-to-end user journey; dogfood a product; test a release candidate.

How do I install Run Smoke Tests in Claude Code?

Run `npx skills add tamdogood/builder-essential-skills --skill run-smoke-tests -a claude-code`. Or copy the skill folder (skills/run-smoke-tests in tamdogood/builder-essential-skills) into .claude/skills/run-smoke-tests in your project. Claude Code loads it when a task matches its description.

How do I install Run Smoke Tests in Codex?

Run `npx skills add tamdogood/builder-essential-skills --skill run-smoke-tests -a codex`. Or copy the skill folder (skills/run-smoke-tests in tamdogood/builder-essential-skills) into .agents/skills/run-smoke-tests in your project. Codex loads it when a task matches its description.

Can I use Run Smoke Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tamdogood/builder-essential-skills --skill run-smoke-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run-smoke-tests, .gemini/skills/run-smoke-tests, .github/skills/run-smoke-tests and .opencode/skills/run-smoke-tests in your project.

What does Run Smoke Tests need to run?

SKILL.md names no scripts, command-line tools or credentials: Run Smoke Tests is instructions for the agent only.

Does Run Smoke Tests access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Run Smoke Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Run Smoke Tests use?

Run Smoke Tests is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Run Smoke Tests use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 561 tokens, read only when the agent opens those files.

What are the alternatives to Run Smoke Tests?

Skills that share tags, products or a category with Run Smoke Tests: Verify Bb (get-bb/bb, 4.2k stars), Plugin Test (NanmiCoder/dsh-auto-mode, 164 stars), Solution Testing (kid-sid/claude-spellbook, 190 stars) and Reproduce Chat States (different-ai/openwork, 24k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Run Smoke Tests?

tamdogood (a GitHub user) maintains it in tamdogood/builder-essential-skills, which has 221 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on August 16, 2026.

Source: tamdogood/builder-essential-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.