Agent skill

Field Test

by cloudposse in cloudposse/atmos

Hands-on manual DX test pass of a feature or CLI command: read the real implementation and tests, hypothesize plausible user misunderstandings and misuse automated tests don't cover, build durable…

Apache-2.0Auto-check: warningsDevOps & Cloud

Install Field Test

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add cloudposse/atmos --skill field-test -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cloudposse/atmos field-test --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cloudposse/atmos.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/field-test .claude/skills/field-test && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
field-test
GitHub stars
1.4k
Token cost
~4.6k tokens
SKILL.md length
2,654 words
Files
1
Skills in repo
70
Repo updated
First seen
Licence
Apache-2.0

At a glance

Hands-on manual DX test pass of a feature or CLI command: read the real implementation and tests, hypothesize plausible user misunderstandings and misuse automated tests don't cover, build durable…

  • Works in 5 steps: Research before touching anything → Generate hypotheses, don't just wander → Build real, durable fixtures → …
  • DevOps & Cloud work in your project
  • SKILL.md covers Plan Mode, Phase 1 — Research before…, Phase 2 — Generate hypotheses,… and Phase 3 — Build real, durable…, plus 3 more sections
  • Calls git, gh and terraform

What it does

Field Test is an agent skill from cloudposse/atmos. Hands-on manual DX test pass of a feature or CLI command: read the real implementation and tests, hypothesize plausible user misunderstandings and misuse automated tests don't cover, build durable fixtures, execute for real against real state, and report ranked findings. Investigation only — never fixes anything found. Defaults to testing whatever the current branch changed vs its base branch when no explicit target is given. Invoke on explicit requests like 'field test X' / 'do a DX test pass on X' / 'find…

Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud. It works with Git. The repository describes itself as: Atmos is the open-source runtime for infrastructure — it builds, authenticates, and ships Terraform, OpenTofu, Packer, Ansible, Kubernetes, Helm, and containers the same way on… The licence is Apache-2.0.

When your agent uses it

  • DevOps & Cloud work in your project

Example prompts

  • “field test X”
  • “do a DX test pass on X”
  • “find vibe-coded slop in X”
  • “/field-test”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Research before touching anything
  2. Generate hypotheses, don't just wander
  3. Build real, durable fixtures
  4. Execute for real, with discipline
  5. Report

What it can do on your machine

Read from SKILL.md and the folder at commit bbe58a6. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • gh
    • terraform

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and gh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Field Test loads about 4.6k tokens when it runs. Until then it costs about 143 tokens; SKILL.md has 2,654 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~143
When it runs · the whole SKILL.md, loaded when a task matches
~4.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningTells the agent its actions are pre-authorized / not to stop for confirmationSKILL.md:72
    Mode` to request approval of that plan. Don't ask for approval any other way.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from cloudposse/atmos at commit bbe58a6, republished under its Apache-2.0 licence (© cloudposse). 2,654 words, ~4,648 tokens.

Download SKILL.mdSave it as .claude/skills/field-test/SKILL.md (or your agent's skills folder).
name
field-test
description
Hands-on manual DX test pass of a feature or CLI command: read the real implementation and tests, hypothesize plausible user misunderstandings and misuse automated tests don't cover, build durable fixtures, execute for real against real state, and report ranked findings. Investigation only — never fixes anything found. Defaults to testing whatever the current branch changed vs its base branch when no explicit target is given. Invoke on explicit requests like 'field test X' / 'do a DX test pass on X' / 'find vibe-coded slop in X' / 'field test this branch'.
argument-hint
Feature or command to test, e.g. 'atmos vendor pull' (omit to default to this branch's change)
metadata.copyright
Copyright Cloud Posse, LLC 2026
metadata.version
1.0.0

Field Test

Hands-on, adversarial test pass of $ARGUMENTS (the feature/command named when this skill was invoked, e.g. atmos vendor pull).

If no target was given, default to the change introduced on the current branch rather than asking. Resolve the actual pull-request base branch when one is available (gh pr view --json baseRefName -q .baseRefName for the current branch), falling back to the repository's default branch (gh repo view --json defaultBranchRef -q .defaultBranchRef.name) when no PR exists yet. Do NOT use the upstream tracking branch (@{u}) as the base — for a normal feature branch that tracks origin/<same-branch-name>, diffing against its own upstream produces an empty or near-empty diff, not the PR's actual changes, once the branch has been pushed.

A branch NAME from gh (e.g. develop) is not guaranteed to be a usable git ref in THIS checkout — shallow clones, detached HEADs, and worktrees with a narrow fetch refspec can have the name without the commits. Before running any diff, verify the base actually resolves here (git rev-parse --verify --quiet <candidate>^{commit}), trying in order: origin/<resolved-name>, <resolved-name> — take the first that verifies. Do not fall back to origin/main/main when resolved-name came from an actual PR base other than the default branch (e.g. a PR targeting develop) — diffing against main instead of the PR's real base compares against the wrong history and can silently miss real changes or include unrelated ones. origin/main/main are legitimate fallback candidates only in the no-PR case (gh pr view returned nothing, so resolved-name already IS the repository's default branch) or when gh itself isn't available at all. If a known, non-default PR base doesn't resolve as either origin/<resolved-name> or <resolved-name>, stop and ask the user which base to diff against — do not run `git diff

<base>...` against an unverified ref and let it fail with a confusing git error, and do not
silently substitute a different branch for a known PR base. Once a real base is confirmed, inspect
the FULL set of
changes relative to that base — content, not just a file-list summary: `git diff <base>...HEAD`
(the full patch, not `--stat`, since `--stat` only shows file names and line counts, not what
those lines actually do) for committed history; `git diff HEAD` for any staged/unstaged changes
to tracked files not yet committed; and `git ls-files --others --exclude-standard` to enumerate
untracked files — `git status --porcelain` lists their paths too but never their content, and
neither `git diff HEAD` nor a `--stat` summary includes untracked files at all, so a brand-new
implementation file can otherwise go completely unread. Read the actual content of every
untracked file this turns up, the same as any diff hunk. Derive the test target from all of
that — the CLI command(s), flag(s), config option(s), or subsystem the changed files implement —
and state explicitly what you inferred and why before proceeding to Phase 1. Only fall back to
asking the user if no candidate base ref resolves at all, there's truly nothing changed (clean
worktree, base equals HEAD, no untracked files), or the changes span multiple unrelated features
with no coherent single target (ask which one to focus on, don't silently pick one).

Goal: catch "vibe-coded slop" — behavior that looks fine in code review but breaks or misleads a real user — not to re-run what automated tests already cover. Anticipate plausible user misunderstandings, not just obvious bugs.

This pass is investigation only. Do not fix anything you find — see Phase 5. Report and stop; the user decides what to fix. Any follow-up fix work — including a plan to fix the findings, not just the implementation — must close with the fix-log skill, not this one.

Plan Mode

If plan mode is active when this skill is invoked (a <system-reminder> says so), Phase 3 (Build fixtures) and Phase 4 (Execute) can't run inline — writing fixture files and running real, sometimes state-mutating commands are exactly what plan mode exists to gate.

  • Run Phase 1 (Research) and Phase 2 (Generate hypotheses) as normal — both are already read-only and fit plan mode's constraints without modification.
  • Instead of proceeding into Phase 3/4, write the plan file: the target under test, the prioritized hypothesis list from Phase 2, the fixtures Phase 3 would build (note which extend/copy an existing fixture, per that phase's guidance), and which planned Phase 4 commands are read-only vs. state-mutating — the mutating ones are specifically what need sign-off.
  • Call ExitPlanMode to request approval of that plan. Don't ask for approval any other way.
  • Once approved, resume at Phase 3 using the approved plan as the fixture/execution blueprint.

Phase 1 — Research before touching anything

The goal is a map of "documented or plausible usage" minus "already tested" = what needs manual verification. This phase is broad, read-only research — delegate it to Agent subagent_type: "Explore" (1-3 agents in parallel, one per bullet below) rather than doing it all serially inline.

When defaulting to the current branch (no explicit target given), scope every bullet below to the target inferred from the branch diff — don't research the whole surrounding subsystem when the branch only touched one corner of it. If the branch's changed files span more than one command or package, treat each as a separate target to cover in Phase 2-4, prioritized by how much of the diff each accounts for.

  • Implementation — the actual code, not just its docs or the skill describing it. Inspect every changed production package identified by the diff — new business logic belongs in narrow pkg/ packages per this repo's conventions, but internal/exec/ is still where a large amount of existing logic lives during its ongoing migration, so a branch touching files there must still be read, not skipped. Check both cmd/<command>/ (thin call site) and whichever pkg//internal/exec/ package(s) it delegates to for the real logic and error paths. When defaulting from a branch diff, read the full diff content itself first (not just the post-change files, and not just a --stat summary) — the diff shows what changed from, which is where a regression or half-finished edge case would show up. Include untracked files (git ls-files --others --exclude-standard) in this reading pass too — they never appear in any diff at all.
  • Docs and skills — every relevant page under website/docs/cli/commands/, the matching .claude/skills/atmos-* skill(s) for the subsystem, and any README describing the feature. Note anything phrased with confidence you haven't independently confirmed against the code — docs and skills describe intended behavior, not necessarily current behavior. A capability can be documented on one page and missing from another that also covers it (a reference page, the skill's own summary, a related subsystem doc that mentions it in passing) — check every doc surface for the feature, not just the primary one. Also re-read each doc for internal self-contradiction: a rule stated in one paragraph can be silently overridden or contradicted by an example or a later paragraph in the same file.
  • Existing automated tests — unit tests colocated with the code, tests/test-cases/ fixtures, tests/testdata/ golden snapshots. For each, note exactly what it does and doesn't exercise (mocked vs. real execution, which flags/paths/backends are hit).
  • Every flag, config option, and documented action/mode — grep for them and list them. You will need to touch every one in Phase 4. Cross-reference mechanically: grep the command's flag registration calls in its Go source and diff that list against the flags documented in the corresponding .mdx's <dl> — every registered flag needs a matching <dt>, and every documented flag needs to actually be registered.

Phase 2 — Generate hypotheses, don't just wander

Before running anything, write down concrete things to try, prioritized by what a real user would plausibly do:

  • Every flag combination that seems natural but might not be validated (two flags that should be mutually exclusive; two config fields whose combination is never cross-checked).
  • Every flag's parsed value actually reaching the code path that would use it — not just that the flag exists, parses, and the command exits 0. A flag can be fully registered and documented yet silently dropped before the logic that should consume it (e.g. the command builds a fresh/empty config struct instead of the one built from parsed flags). Trace the value from flag definition to point of use for every flag, don't just confirm the command accepts it.
  • Every place the docs/skill claim something you haven't verified against actual code.
  • Any "safe-looking" command (plan/preview/--dry-run/list/describe) that might secretly mutate state or trigger side effects, if built the same way as a mutating command — Atmos has many of these pairs (e.g. terraform plan vs apply, vendor diff vs pull), so this is a high-yield category here specifically.
  • Any action/mode described in docs but not exercised by ANY test or example in the repo — those are the highest-yield targets; if nothing has ever run it for real, assume it's broken until you prove otherwise.
  • Copy-paste/misconfiguration scenarios — what happens if a user copies a working stack/component block and changes one field but forgets a related one?
  • Error messages — accurate, do they name the actual flags/values involved, do they suggest a fix (per this repo's error-builder/hint conventions)?
  • Idempotency/rerun-safety — run the same operation twice; does the second run behave correctly?
  • Determinism — run the same read-only command several times with no state change between runs — is the output identical every time?
  • Any command that does real work across multiple items (batch installs/updates, concurrent workers, multi-resource loops) — does it show live progress, or does it silently buffer everything and dump it all at once when the whole batch finishes? A command that takes 10+ seconds with zero output is indistinguishable from a hang to a real user. Also check whether its status lines actually use ui.Success/ui.Error/ui.Warning/ui.Info (icon + theme color) per CLAUDE.md's I/O and UI Usage section, rather than a hand-rolled glyph ("✓ %s", "✗ %s") printed through the plain ui.Writef/ui.Write — the two are easy to conflate since both compile and both "print a checkmark," but only the semantic function is themed/colored. This class of bug is easy to miss when every command in this pass has been run through the Bash tool — captured output can look fine in the transcript even when the real behavior (silent hang, unstyled text) would be obvious to a human watching a real terminal. This needs two separate checks, not one piped command doing double duty — piping through cat -v makes the command non-TTY, which exercises the non-live fallback renderer instead of the real live-progress path, and --force-color does not restore TTY behavior:
    • Live progress: run the command in a real pseudo-TTY, unpiped (e.g. script -q /dev/null build/atmos toolchain update --force-tty --force-color), and watch it — does a spinner/progress bar actually redraw in place, or does it silently buffer and dump everything at once?
    • ANSI styling: separately, pipe a --force-color run through cat -v (or grep for the raw \x1b[ / ^[[ escape sequence) to confirm color codes are actually present around each status line, not just plain text with a Unicode glyph. Piping is fine here since this check only cares about styling, not live-rendering behavior.
  • Any "N -> M" / diff-style report line — construct a case where N and M are the same value in different string forms (e.g. v1.2.3 vs 1.2.3, the one equivalence normalizeVersion actually handles by stripping a leading v — don't test casing or trailing-metadata variants unless the specific normalizer under test explicitly documents supporting them). A raw string-equality comparison will misreport a no-op as a change, which also tends to corrupt whatever summary/count line tallies outcomes — check the tally against the individual lines above it, don't just trust it.
Show full SKILL.md (761 more words)Show less

Phase 3 — Build real, durable fixtures

  • Prefer extending or copying an existing fixture (tests/test-cases/, examples/, demo/) over inventing one from scratch.
  • Build fixtures that exercise every documented capability, especially ones nothing in the repo currently exercises. Make them realistic, not minimal-to-the-point-of-artificial.
  • If real infrastructure/emulators are available for what you're testing, use them for at least one pass — this repo ships local AWS/GCP/Azure/Kubernetes/Vault/registry emulators for exactly this purpose (see the atmos-emulator skill). Don't rely solely on mocked/dry-run paths, since that's exactly what's already covered by automated tests.
  • Never manually edit golden snapshot files under tests/test-cases/, tests/testdata/, or tests/snapshots/ — regenerate them via -regenerate-snapshots per CLAUDE.md's Golden Snapshots section.
  • Keep fixtures that have lasting value (they close real coverage gaps); don't create scratch-and-delete throwaways unless truly one-off.

Phase 4 — Execute for real, with discipline

  • Run actual commands against a disposable fixture or emulator by default. Before any command that can mutate state, obtain explicit user confirmation. Run against shared or production state only with a documented backup and rollback plan. Don't reason abstractly about what "should" happen — observe what does happen. Build a fresh binary first (atmos build) if the change under test isn't already reflected in ./build/atmos.
  • Never pipe redirection into a command under test — per CLAUDE.md, piping breaks TTY detection, which can mask exactly the DX issues (interactive prompts, color, spinners) you're testing for.
  • Before every test, verify you're actually starting from a clean/expected state — don't assume. Stale state from a previous run (yours or a prior session's) will silently corrupt your results. Reset explicitly and confirm the reset worked (check a resource count/id changed, not just that a command exited 0).
  • When something surprises you, reduce it to the smallest reproducible case and verify the repro twice.
  • For any command with a summary/tally line (Updated N, up to date M, failed K), manually recount the individual lines above it and compare — don't trust the tally at face value. A miscounted summary is a strong signal the classification logic feeding it is wrong somewhere, not just a cosmetic issue.
  • If Phase 1 research made a claim, verify it live before trusting it — code-reading can miss control flow (e.g. assuming a flag is silently ignored when it actually errors, or vice versa). The same applies to any externally suggested fix or finding (a review comment, a prior report) — don't propagate its claimed root cause (e.g. a named sentinel error) without running the code to confirm it's actually correct. Correct the record explicitly when research — yours or someone else's — turns out wrong.
  • A test that only asserts require.Error (or similarly loose) without pinning down which error — ErrorIs/ErrorAs against the actual sentinel — is itself a field-test finding worth reporting, not just something to note in passing.
  • Test the happy path too, not just edge cases — confirm what's supposed to work actually does, so the report distinguishes real regressions from things that were never broken.

Phase 5 — Report

For every finding: exact repro command(s), expected vs. actual output, and severity (silent data loss/mutation > crash on reasonable input > confusing error message > cosmetic). Rank the report by severity, most dangerous first. Explicitly call out anything verified as working correctly too — a report that's only bad news is as misleading as one that's only good news.

Keep a running scratch log of findings as you go (in your own working notes/task list) rather than reconstructing everything at the end from memory — but don't commit that log. Per CLAUDE.md's Git section, scratch/research files never get committed; only the fixtures built in Phase 3 (if kept for lasting value) and this final report are durable output.

Do not fix anything found — this pass is investigation only. Stop and report; the user decides what to fix. If a plan gets made to fix any finding — whether right away or as a later follow-up, even in a different session — that plan and its implementation must close with the fix-log skill so the fix leaves a durable record under docs/fixes/. End by invoking the say skill — a completed test pass reaching a stopping point a human should review is exactly its trigger.

  • Plan Mode (above) — when invoked under plan mode, Phases 3-4 wait for approval via ExitPlanMode before building fixtures or executing.
  • Explore agent — Phase 1's broad read-only research.
  • atmos-emulator skill — real local infra for Phase 3/4 when the target touches AWS/GCP/Azure/Kubernetes.
  • docs skill — conventions for the CLI docs being cross-checked in Phase 1.
  • fix-log skill — required for any follow-up fix work, including a plan to fix findings, once the user decides what to fix; out of scope here.
  • say skill — end-of-pass notification.

© cloudposse, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/field-test of cloudposse/atmos.

Open the folder on GitHubat commit bbe58a6

Compare with similar skills

Field Test next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Field Test compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Field Test this skillcloudposse/atmos1.4k—~4.6kAutomated safety check: WarnApache-2.0
Repo Mirror Sourcesnetdata/netdata81k—~1.2kAutomated safety check: NotesGPL-3.0
CI Failure Triage and RepairChachamaru127/claude-code-harness3.2k1 repos~1.1kAutomated safety check: NotesMIT
Bfe Rd Workflowbfenetworks/bfe6.3k—~1.3kAutomated safety check: PassApache-2.0
Ssh Skillbadseal/ssh-skill535—~2.4kAutomated safety check: NotesNone
GreptimeDB Release RunbookGreptimeTeam/greptimedb6.7k—~1.4kAutomated safety check: PassApache-2.0

Similar skills

  • Repo Mirror Sources

    netdata/netdata

    Inspect Netdata-org source checkouts under NETDATAREPOSDIR, or set up and synchronize that mirror when requested.

    81k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • CI Failure Triage and Repair

    Chachamaru127/claude-code-harness

    Diagnoses failing CI pipelines and tests, deciding first whether the test or the implementation is at fault, and hands hard cases to a dedicated fixer subagent.

    3.2k GitHub starsUsed in 1 repo~1.1k tokens
    DevOps & CloudAuto-check: notes
  • Bfe Rd Workflow

    bfenetworks/bfe

    引导用户在 bfe 代码库中完成一次完整的功能研发流程,包括需求对齐、文档修改、代码实现、集成测试与回归验证. An agent skill from bfenetworks/bfe.

    6.3k GitHub stars~1.3k tokensUpdated 4 days ago
    DevOps & CloudAuto-check passed
  • Ssh Skill

    badseal/ssh-skill

    A skill your agent uses when a task requires SSH or SCP/SFTP behavior, a remote server, server alias/IP/hostname/user@host, bastion or jump-host access, remote command execution, upload/download…

    535 GitHub stars~2.4k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check: notes
  • GreptimeDB Release Runbook

    GreptimeTeam/greptimedb

    Runbook for publishing a GreptimeDB version: pick the release branch, verify the Cargo version, then tag, create the GitHub release and open the docs note PR.

    6.7k GitHub stars~1.4k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Crabbox Quickstart

    openclaw/crabbox

    Gets you running your repository's tests in a disposable Docker or Podman container on your own machine with Crabbox, with no account and no cloud spend.

    1.5k GitHub stars~1.6k tokensUpdated today
    DevOps & CloudAuto-check: notes

More from cloudposse/atmos

All 70 skills in this repo
  • Fix Log

    cloudposse/atmos

    A skill your agent uses when implementing, finishing, documenting, or reviewing a fix, repair, remediation, bug fix, debug-and-fix task, workflow fix, infrastructure fix, or any change that should…

    1.4k GitHub stars~685 tokensUpdated today
    Auto-check passed
  • Atmos Lint

    cloudposse/atmos

    Atmos Terraform linting with TFLint: standalone atmos terraform lint, component-aware config discovery and toolchain versions, TFLint rule configuration, and lifecycle hooks/CI findings.

    1.4k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Changelog

    cloudposse/atmos

    Blog post authoring for Atmos: MDX template, frontmatter, website/blog/tags.yml and authors.yml rules, problem-first framing, backtick-opening ban, optional cast embeds, and no-Go-internals leakage.

    1.4k GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Editions

    cloudposse/atmos

    Decide whether a PR's new or changed default needs edition-journal handling (pkg/edition, docs/prd/editions.md), and do the mechanical work if so: journal entries, the four-layer default check…

    1.4k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Atmos Migration

    cloudposse/atmos

    Migrate to Atmos from native Terraform, Terraform Workspaces, Terramate, Terragrunt, Make, Just, or Task; migrate tool versions from mise or Aqua CLI; migrate AWS/GCP/Azure CLI configs, Leapp…

    1.4k GitHub stars~5.1k tokensUpdated today
    Auto-check: warnings
  • PR Maintenance Loop

    cloudposse/atmos

    Start an hourly background loop that keeps the current branch's PR rebased, its addressed CodeRabbit threads resolved, its CI checks passing, its lint clean, its tests passing with adequate patch…

    1.4k GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Field Test

What does Field Test do?

Hands-on manual DX test pass of a feature or CLI command: read the real implementation and tests, hypothesize plausible user misunderstandings and misuse automated tests don't cover, build durable…. Field Test is an agent skill from cloudposse/atmos. Hands-on manual DX test pass of a feature or CLI command: read the real implementation and tests, hypothesize plausible user misunderstandings and misuse automated tests don't cover, build durable fixtures, execute for real against real state, and report ranked findings.

When should I use Field Test?

Field Test fits situations like: devOps & Cloud work in your project.

How do I install Field Test in Claude Code?

Run `npx skills add cloudposse/atmos --skill field-test -a claude-code`. Or copy the skill folder (.claude/skills/field-test in cloudposse/atmos) into .claude/skills/field-test in your project. Claude Code loads it when a task matches its description.

How do I install Field Test in Codex?

Run `npx skills add cloudposse/atmos --skill field-test -a codex`. Or copy the skill folder (.claude/skills/field-test in cloudposse/atmos) into .agents/skills/field-test in your project. Codex loads it when a task matches its description.

Can I use Field Test in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cloudposse/atmos --skill field-test -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/field-test, .gemini/skills/field-test, .github/skills/field-test and .opencode/skills/field-test in your project.

What does Field Test need to run?

Going by SKILL.md and its folder, Field Test needs the command-line tools its instructions call (git, gh and terraform).

Does Field Test access the network?

SKILL.md contains no URLs. Its commands use git and gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Field Test safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): tells the agent its actions are pre-authorized / not to stop for confirmation. Read the flagged lines before installing; the check is not a guarantee either way.

What licence does Field Test use?

Field Test is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Field Test use?

About 4.6k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Field Test?

Skills that share tags, products or a category with Field Test: Repo Mirror Sources (netdata/netdata, 81k stars), CI Failure Triage and Repair (Chachamaru127/claude-code-harness, 3.2k stars), Bfe Rd Workflow (bfenetworks/bfe, 6.3k stars) and Ssh Skill (badseal/ssh-skill, 535 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Field Test?

cloudposse (a GitHub organization) maintains it in cloudposse/atmos, which has 1,395 GitHub stars. The repository holds 70 skills in this directory. The repository was last updated on October 7, 2026.

Source: cloudposse/atmos on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.