Agent skill

Integration Test

by tracewayapp in tracewayapp/traceway

Run a live-instance verification of traceway-cli that goes beyond the Go smoke suite — exercises real-data detail endpoints, TTY-default rendering, adaptive metric-name discovery, and emits a…

MITAuto-check passedTesting & QA

Install Integration Test

skills CLI
$ npx skills add tracewayapp/traceway --skill integration-test -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tracewayapp/traceway integration-test --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tracewayapp/traceway.git skills-src && mkdir -p .claude/skills && cp -r skills-src/cli/.claude/skills/integration-test .claude/skills/integration-test && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
integration-test
GitHub stars
1.6k
Token cost
~2.4k tokens
SKILL.md length
1,065 words
Files
1
Skills in repo
2
Repo updated
First seen
Licence
MIT

At a glance

Run a live-instance verification of traceway-cli that goes beyond the Go smoke suite — exercises real-data detail endpoints, TTY-default rendering, adaptive metric-name discovery, and emits a…

  • Works in 5 steps: Real-data detail endpoints — exceptions… → Adaptive metric-name discovery — walk a… → TTY-vs-pipe default — table rendering to… → …
  • Explicitly asks (e.g
  • SKILL.md covers Trigger, Relationship to the Go smoke…, Hard constraints and Pre-flight, plus 6 more sections
  • Calls jq, just and nix

What it does

Integration Test is an agent skill from tracewayapp/traceway. Run a live-instance verification of traceway-cli that goes beyond the Go smoke suite — exercises real-data detail endpoints, TTY-default rendering, adaptive metric-name discovery, and emits a human-readable coverage report. Invoke ONLY when the user explicitly asks (e.g. "run integration tests", "verify the CLI against stormwind"). Never invoke automatically after edits or commits. Assumes the user is already authenticated and a default project is configured.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Integration testing and Test coverage. It works with OpenTelemetry. The repository describes itself as: The only tool you need to know what is happening and how to fix it. The licence is MIT.

When your agent uses it

  • Explicitly asks (e.g
  • Tasks that involve Integration testing
  • Tasks that involve Test coverage

Example prompts

  • “run integration tests”
  • “verify the CLI against stormwind”
  • “/integration-test”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Real-data detail endpoints — exceptions show , populated metrics query --name with every aggregation + group-by.
  2. Adaptive metric-name discovery — walk a candidate list until one populates.
  3. TTY-vs-pipe default — table rendering to a real terminal.
  4. Coverage matrix report — a human-readable artifact, on demand.
  5. Safety doctrine — the forbidden-verb blocklist and confirmMutation env hygiene, applied to every probe.

What it can do on your machine

Read from SKILL.md and the folder at commit 3e5ca7d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jq
    • just
    • nix

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Integration Test loads about 2.4k tokens when it runs. Until then it costs about 120 tokens; SKILL.md has 1,065 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~120
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from tracewayapp/traceway at commit 3e5ca7d, republished under its MIT licence (© tracewayapp). 1,065 words, ~2,381 tokens.

Download SKILL.mdSave it as .claude/skills/integration-test/SKILL.md (or your agent's skills folder).
name
integration-test
description
Run a live-instance verification of traceway-cli that goes beyond the Go smoke suite — exercises real-data detail endpoints, TTY-default rendering, adaptive metric-name discovery, and emits a human-readable coverage report. Invoke ONLY when the user explicitly asks (e.g. "run integration tests", "verify the CLI against stormwind"). Never invoke automatically after edits or commits. Assumes the user is already authenticated and a default project is configured.

integration-test — traceway-cli

A repeatable protocol for verifying the CLI end-to-end against a live Traceway server, focused on the things the Go smoke suite can't easily cover.

Trigger

Invoke ONLY when the user explicitly asks. Never run as a side effect of edits, commits, or builds.

Relationship to the Go smoke suite

The just smoke-test target (test/smoke/*_test.go, build tag smoke) is the primary regression check. Against a live instance it already covers:

  • JSON shape of every list endpoint and profiles list.
  • Output-format coverage (json/table/yaml) for projects, profiles, exceptions, endpoints, logs.
  • Client-side enum validation (--search-type, --order-by, --sort-direction, --aggregation, --tag).
  • Time-range parsing edges (--since 7D, missing --to, --since + absolute mix, far-future windows).
  • metrics query missing --name, bogus name → empty series, malformed tag.
  • exceptions show <zeros> → exit 5 not_found.
  • --profile no-such-profile → exit 4; --project 0…0 → non-zero, no panic, no connection_failed.

Do not re-implement these here. If they regress, that's a Go-test bug, not a skill failure.

What this skill adds beyond just smoke-test:

  1. Real-data detail endpoints — exceptions show <captured-hash>, populated metrics query --name <real> with every aggregation + group-by.
  2. Adaptive metric-name discovery — walk a candidate list until one populates.
  3. TTY-vs-pipe default — table rendering to a real terminal.
  4. Coverage matrix report — a human-readable artifact, on demand.
  5. Safety doctrine — the forbidden-verb blocklist and confirmMutation env hygiene, applied to every probe.

Hard constraints

Read-only. No exceptions. Even if a subcommand looks safe by name, check --help for mutating flags before running.

Forbidden verbs and flags

Skip any subcommand whose name or --help mentions:

  • archive, unarchive, resolve, unresolve, mute, ack, acknowledge
  • create, delete, update, set, put, post, add, remove, rm
  • assign, claim, close, reopen
  • login, logout, token, rotate, regenerate
  • --archive, --resolve, --delete, --write, --mutate, --apply, --commit

If a new subcommand is ambiguous (sync, refresh, replay, export), do not run it — list it under "skipped — manual review". If --dry-run exists, still skip write-shaped subcommands.

Mutation safeguards

The CLI gates mutations via confirmMutation (cmd/traceway/querycommon.go). The harness MUST:

  • Never pass --yes.
  • unset TRACEWAY_ASSUME_YES at the top of the script.
  • Run with stdin from /dev/null.

So that if a forbidden verb slips through, the gate refuses with exit 2 usage_error instead of hanging on a prompt.

Pre-flight

Run in order. Stop if any fails.

  1. Build: nix develop --command go build -o ./bin/traceway ./cmd/traceway.
  2. Config exists (don't print — JWT inside): test -f "${XDG_CONFIG_HOME:-$HOME/.config}/traceway/config.json".
  3. Reachability + capture TW_PROJECT_ID:
    bash
    ./bin/traceway projects list --output json | jq -e 'type=="array" and length>=1' >/dev/null
    TW_PROJECT_ID=$(./bin/traceway projects list --output json | jq -r '.[0].id')

projects list --output json returns a bare array, not a {data, pagination} envelope.

Detail-endpoint probes

exceptions show <captured-hash>
  1. Capture a real hash:
    bash
    HASH=$(./bin/traceway exceptions list --since 720h --page-size 1 --output json | jq -r '.data[0].exceptionHash // empty')
    If empty, retry against other projects via --project <id>. If still empty, skip with reason no exception found across all projects.
  2. With a real hash: three output formats + --help. JSON shape: {group: {...}, occurrences: [...], pagination: {...}} — assert .group and .occurrences.
  3. Capture .occurrences[0].traceId if present for the logs probe below.
metrics query --name <real-metric> (adaptive)

Probe these names in order until one returns a populated series:

system.cpu.utilization
system.network.io
system.network.errors
system.network.dropped
http.server.duration
traceway.requests

If none populates, skip the live block with reason no live metric name found.

For the first metric that populates:

  • Three output formats + --help.
  • All aggregations: avg, sum, count, min, max, p50, p95, p99.
  • --interval-minutes 15.
  • --group-by direction for network metrics (splits __all__ into receive/transmit).

JSON shape: {results: [{name, unit, series: {<tag-key>: [{timestamp, value}, ...]}}]} — series is a map keyed by group tag, default key __all__.

logs query --trace-id <captured>

If a real trace id was captured above, run logs query --trace-id $TRACE --since 720h. Assert exit 0 and {data, pagination} shape.

Show full SKILL.md (524 more words)Show less
By-id detail commands (captured id + recordedAt)

These all require a timestamp flag; capture the id and its recordedAt together from exceptions show, then exercise them. All read-only.

  1. Capture one occurrence's id + recordedAt (+ optional trace/session ids):
    bash
    OCC=$(./bin/traceway exceptions show "$HASH" --output json | jq -c '.occurrences[0]')
    OID=$(jq -r '.id'                 <<<"$OCC")
    OTS=$(jq -r '.recordedAt'         <<<"$OCC")
    DT=$(jq -r '.traceId // empty'    <<<"$OCC")
    SID=$(jq -r '.sessionId // empty' <<<"$OCC")
  2. exceptions occurrence $OID --recorded-at $OTS — assert exit 0 and .exception.id == $OID.
  3. Required-flag enforcement (no live data needed): exceptions occurrence $OID with no --recorded-at → exit 2 usage_error; --recorded-at notadate → exit 2 invalid_timestamp; endpoints show not-a-uuid --recorded-at $OTS → exit 2 usage_error.
  4. If $DT is non-empty: traces show $DT --recorded-at $OTS → exit 0, .nodes is an array. If $SID is non-empty: sessions show $SID --started-at $OTS → exit 0, .session present.
  5. endpoints show / tasks show / ai-traces show need an id of their own type — capture one from a traces show node when available (.nodes[].endpoint.id + .endpoint.recordedAt, etc.); otherwise skip with reason no <type> id captured.

TTY-vs-pipe default

If script is available:

bash
script -q /dev/null ./bin/traceway projects list | head -20

Expect a table. Piping without script should yield JSON. Mark as "not verified" if script is absent.

Subcommand skip lists

Mutating (skip with reason forbidden verb): exceptions archive, exceptions unarchive, login, logout, profiles use (local mutation), projects use (local mutation of state.json).

Not in the CLI (skip with reason subcommand not in CLI) — kept so the report shows the gap if they ship: tasks list, sessions list, ai-traces list, traces list, metrics discover.

Observation and reporting

Classify each invocation:

ResultClassification
exit 0, valid JSON, expected keys presentpass
exit 0 but stdout empty when data expected, or JSON invalid / missing keysfail — schema
exit non-zero, clean message, expected errorpass — error case
exit non-zero with panic or stack tracefail — crash
exit non-zero on a happy pathfail — unexpected error
stderr non-empty on a passing commandwarn — noisy
> 10s on a list callwarn — slow

Drive the run from one ephemeral /tmp/*.sh script (not committed). Top of script:

bash
#!/usr/bin/env bash
set -u                            # never set -e
unset TRACEWAY_ASSUME_YES
exec </dev/null                   # no TTY for the suite
LOG=$(mktemp /tmp/traceway-it.XXXXXX.log)

Capture exact invocation, exit code, and first 20 lines of stdout/stderr per probe to $LOG.

End-of-run markdown report:

  1. Summary: N pass, M fail, W warn, S skipped + wall-clock.
  2. Failures table: command, classification, one-line excerpt.
  3. Warnings table: same shape.
  4. Skipped table: command + reason.
  5. Coverage matrix: dimensions exercised per probed command (— for not exercised).
  6. Smoke-suite pointer: note whether just smoke-test was run in this session and its result, so the report stands alone.
  7. Log path: location of $LOG.

Do not inline the full log.

What this skill does not do

  • It does not duplicate just smoke-test. Assume that suite's coverage is green; if you suspect regressions there, run smoke separately.
  • It does not write or update the CLI's Go test files.
  • It does not perform any login, token rotation, profile creation, or credential mutation.
  • It does not gate commits or CI.

When to expand

If a new read subcommand ships, add a probe block. If a new mutating subcommand ships, add it to the forbidden list. If a regression pattern shows up that is deterministic and stateless, push it down to test/smoke/*_test.go — not into this skill.

© tracewayapp, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in cli/.claude/skills/integration-test of tracewayapp/traceway.

Open the folder on GitHubat commit 3e5ca7d

Compare with similar skills

Integration Test next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Integration Test compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Integration Test this skilltracewayapp/traceway1.6k—~2.4kAutomated safety check: PassMIT
Test Writing WorkflowiOfficeAI/AionUi33k1 repos~1.2kAutomated safety check: PassApache-2.0
OpenROAD Module Test AdderThe-OpenROAD-Project/OpenROAD3.2k—~1.8kAutomated safety check: PassBSD-3-Clause
Write Testsgrafana/synthetic-monitoring-app171—~1.2kAutomated safety check: PassAGPL-3.0
Benchflow Experiment Reviewbenchflow-ai/benchflow353—~4kAutomated safety check: PassApache-2.0
NIC Testing Patternsnginx/kubernetes-ingress5.1k—~2.8kAutomated safety check: PassApache-2.0

Similar skills

  • Test Writing Workflow

    iOfficeAI/AionUi

    Sets the test-writing workflow for the repository: risk-first scenario lists, behavior-focused Vitest tests, a full run before each commit and a coverage target.

    33k GitHub starsUsed in 1 repo~1.2k tokens
    Testing & QAAuto-check passed
  • OpenROAD Module Test Adder

    The-OpenROAD-Project/OpenROAD

    Adds integration or unit tests to an OpenROAD module: writes the Tcl test, generates golden files and registers it in both CMake and Bazel.

    3.2k GitHub stars~1.8k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Write Tests

    grafana/synthetic-monitoring-app

    Official

    Write Jest integration and unit tests for the Grafana Synthetic Monitoring app using React Testing Library, MSW, and src/test helpers.

    171 GitHub stars~1.2k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Benchflow Experiment Review

    benchflow-ai/benchflow

    Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes.

    353 GitHub stars~4k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • NIC Testing Patterns

    nginx/kubernetes-ingress

    Testing conventions for the NGINX Ingress Controller repo: Go table-driven tests, mandatory snapshot regeneration, Helm tests and Python pytest integration tests.

    5.1k GitHub stars~2.8k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Backend Testing

    biersoeckli/QuickStack

    Create and review QuickStack backend unit and integration tests using the project's Vitest, Prisma, SQLite, and k3s conventions.

    361 GitHub stars~670 tokensUpdated yesterday
    Testing & QAAuto-check passed

More from tracewayapp/traceway

  • Traceway Setup

    tracewayapp/traceway

    Analyze and instrument repositories for Traceway observability.

    1.6k GitHub stars~17k tokensUpdated 2 days ago
    Auto-check: notes

Works with

Categories

Questions about Integration Test

What does Integration Test do?

Run a live-instance verification of traceway-cli that goes beyond the Go smoke suite — exercises real-data detail endpoints, TTY-default rendering, adaptive metric-name discovery, and emits a…. Integration Test is an agent skill from tracewayapp/traceway. Run a live-instance verification of traceway-cli that goes beyond the Go smoke suite — exercises real-data detail endpoints, TTY-default rendering, adaptive metric-name discovery, and emits a human-readable coverage report.

When should I use Integration Test?

Integration Test fits situations like: explicitly asks (e.g; tasks that involve Integration testing; tasks that involve Test coverage.

How do I install Integration Test in Claude Code?

Run `npx skills add tracewayapp/traceway --skill integration-test -a claude-code`. Or copy the skill folder (cli/.claude/skills/integration-test in tracewayapp/traceway) into .claude/skills/integration-test in your project. Claude Code loads it when a task matches its description.

How do I install Integration Test in Codex?

Run `npx skills add tracewayapp/traceway --skill integration-test -a codex`. Or copy the skill folder (cli/.claude/skills/integration-test in tracewayapp/traceway) into .agents/skills/integration-test in your project. Codex loads it when a task matches its description.

Can I use Integration Test in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tracewayapp/traceway --skill integration-test -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/integration-test, .gemini/skills/integration-test, .github/skills/integration-test and .opencode/skills/integration-test in your project.

What does Integration Test need to run?

Going by SKILL.md and its folder, Integration Test needs the command-line tools its instructions call (jq, just and nix).

Does Integration Test access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Integration Test safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Integration Test use?

Integration Test is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Integration Test use?

About 2.4k tokens (SKILL.md is roughly 9.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Integration Test?

Skills that share tags, products or a category with Integration Test: Test Writing Workflow (iOfficeAI/AionUi, 33k stars), OpenROAD Module Test Adder (The-OpenROAD-Project/OpenROAD, 3.2k stars), Write Tests (grafana/synthetic-monitoring-app, 171 stars) and Benchflow Experiment Review (benchflow-ai/benchflow, 353 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Integration Test?

tracewayapp (a GitHub organization) maintains it in tracewayapp/traceway, which has 1,600 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 5, 2026.

Source: tracewayapp/traceway on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.