Agent skill

CLI Testing

by vellum-ai in vellum-ai/vellum-assistant

Manually test a running Vellum assistant end-to-end purely from the CLI — no desktop app or web UI.

MITAuto-check passedDevOps & Cloud

Install CLI Testing

skills CLI
$ npx skills add vellum-ai/vellum-assistant --skill cli-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vellum-ai/vellum-assistant cli-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/cli-testing .claude/skills/cli-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cli-testing
GitHub stars
1.4k
Token cost
~1.6k tokens
SKILL.md length
658 words
Files
1
Skills in repo
108
Repo updated
First seen
Licence
MIT

At a glance

Manually test a running Vellum assistant end-to-end purely from the CLI — no desktop app or web UI.

  • Works in 6 steps: Prerequisites → Provide an LLM provider key (from the… → Hatch — default to a Docker hatch built… → …
  • Verifying assistant behavior
  • SKILL.md covers 0. Prerequisites, 1. Provide an LLM provider key…, 2. Hatch — default to a Docker… and 3. Verify functionality, plus 2 more sections
  • Needs ANTHROPIC_API_KEY and OPENAI_API_KEY

What it does

CLI Testing is an agent skill from vellum-ai/vellum-assistant. Manually test a running Vellum assistant end-to-end purely from the CLI — no desktop app or web UI. Hatch an instance, send messages, watch the reply, and tear it down. Use when verifying assistant behavior, reproducing a bug, or smoke-testing a change without the macOS/web clients.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Frontend development, Containers and End-to-end testing. It works with Docker and macOS. The repository describes itself as: An AI Assistant that’s easy to setup, does your work 24/7, knows your preferences and gets better over time. The licence is MIT.

When your agent uses it

  • Verifying assistant behavior
  • Reproducing a bug
  • Smoke-testing a change without the macOS/web clients

Example prompts

  • “/cli-testing”

Requirements

  • Docker
  • A credential in ANTHROPIC_API_KEY
  • A credential in OPENAI_API_KEY

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Prerequisites
  2. Provide an LLM provider key (from the environment)
  3. Hatch — default to a Docker hatch built from source
  4. Verify functionality
  5. Tear down
  6. Fallback: local mode (no Docker)

What it can do on your machine

Read from SKILL.md and the folder at commit c92ead1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ANTHROPIC_API_KEY
    • OPENAI_API_KEY
    • GEMINI_API_KEY
    • FIREWORKS_API_KEY
    • OPENROUTER_API_KEY
    • MINIMAX_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

CLI Testing loads about 1.6k tokens when it runs. Until then it costs about 74 tokens; SKILL.md has 658 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vellum-ai/vellum-assistant at commit c92ead1, republished under its MIT licence (© vellum-ai). 658 words, ~1,602 tokens.

Download SKILL.mdSave it as .claude/skills/cli-testing/SKILL.md (or your agent's skills folder).
name
cli-testing
description
Manually test a running Vellum assistant end-to-end purely from the CLI — no desktop app or web UI. Hatch an instance, send messages, watch the reply, and tear it down. Use when verifying assistant behavior, reproducing a bug, or smoke-testing a change without the macOS/web clients.

CLI Testing — Exercise the Assistant End-to-End

Drive a real assistant from the terminal only. The vellum CLI (cli/, package @vellumai/cli) manages instance lifecycle; vellum message / vellum events exercise a running instance. See cli/AGENTS.md and the root README.md § CLI for command reference.

0. Prerequisites

bash
export PATH="$HOME/.bun/bin:$PATH"   # bun + the linked `vellum` binary
vellum ps                            # sanity check the CLI resolves

If vellum is missing, run ./setup.sh from the repo root once (installs deps, links the vellum command). Docker must be running for the default flow below.

1. Provide an LLM provider key (from the environment)

Local-mode and Docker-mode instances need one LLM provider key. The CLI reads it straight from the host environment — just export it before hatching/setup:

bash
export ANTHROPIC_API_KEY=sk-ant-...   # or OPENAI_API_KEY / GEMINI_API_KEY /
                                      # FIREWORKS_API_KEY / OPENROUTER_API_KEY /
                                      # MINIMAX_API_KEY

In Devin sessions ANTHROPIC_API_KEY is typically already present in the environment — check with echo "${ANTHROPIC_API_KEY:0:7}" before asking for one. The CLI maps providers to env vars in cli/src/shared/provider-env-vars.ts.

2. Hatch — default to a Docker hatch built from source

Always default to --remote docker. It runs the assistant, gateway, and credential-executor in isolated containers that mirror production and keep the test off your host process table. Reserve --remote local (§5) for the rare case where Docker is unavailable.

Build from source — that's the point of testing. A bare vellum hatch --remote docker pulls the published platform images even when the CLI itself runs from your checkout, so it would test released code, not your changes. Source-build is opt-in via a flag (resolveDockerHatchMode in cli/src/lib/docker.ts):

  • --source <path> — build images once from the source tree at <path>, no watcher. Default for testing: picks up your current changes and is robust for a scripted one-shot run.
  • --watch — build from source and start a file-watcher that rebuilds the affected image on change (watches each service's src/, package.json, and Dockerfile). Use while iterating. The watcher is a long-lived foreground process, so prefer --source for unattended/scripted runs.
bash
vellum hatch --remote docker --source . --name clitest   # build from cwd
# → "Mode: build-from-source" then "Images (local build): vellum-assistant:local-clitest …"

If --source/--watch is passed but no full source tree is found (e.g. the CLI is running from a packaged app bundle), the CLI falls back to pulling the published images and says so — watch for that line if you expect a build. Building all three images takes ~1–2 min the first time.

Hatch attached — do not pass -d. An attached hatch leases the guardian token and configures the provider credential from your environment inline, then returns once the containers are healthy — no follow-up vellum setup needed. Detached mode (-d) defers the guardian-token lease, so a later vellum setup cannot authenticate against the gateway and fails with an invalid_signature 401. Confirm readiness with vellum ps (🟢 healthy) before messaging.

Show full SKILL.md (250 more words)Show less

3. Verify functionality

vellum message is async (returns a message id, not the reply — --json only adds {accepted, messageId}). vellum events streams the reply but is long-running, so background it, send, wait, then read.

Assert on a token the assistant must generate, never one you put in the prompt. vellum events echoes your prompt as **You:** <text> (cli/src/commands/events.ts), so grepping for a word that appears in the prompt passes even when the assistant never replied. Ask a question whose answer is absent from the prompt:

bash
( vellum events > /tmp/vel_events.log 2>&1 & )   # stream in background
sleep 2
vellum message "What is 6 multiplied by 7? Reply with only the number."
sleep 25                                          # let the assistant respond
pkill -f "vellum events"
grep -w 42 /tmp/vel_events.log                    # "42" is NOT in the prompt,
                                                  # so a match proves a real reply

The assistant's streamed reply is written as plain text (no **You:** prefix), so a match on a generated answer confirms the round-trip worked. If you must use a fixed sentinel string, strip the echoed prompt first (grep -v '^\*\*You:\*\*' /tmp/vel_events.log | grep <sentinel>).

Common verification commands
CommandPurpose
vellum psList instances + health (🟢 healthy), id, runtime URL, cloud
vellum message "<text>"Send a message (async; prints message id)
vellum eventsStream live events/replies (long-running — background it)
vellum logs -n 100Last 100 log lines; add -f to follow, -s assistant/-s gateway to filter
vellum clientInteractive terminal chat session (manual exploration)
vellum message --json "<text>"Send-ack as JSON ({accepted, messageId}) — the reply still arrives via vellum events, not here

4. Tear down

bash
vellum retire clitest --yes          # stops containers and removes the instance

retire is destructive (removes per-instance Docker volumes); always clean up test instances when done.

5. Fallback: local mode (no Docker)

Only when Docker is unavailable. Runs the daemon + gateway as plain host processes; configures the provider key automatically from the env at hatch time:

bash
vellum hatch --name clitest          # defaults to --remote local
# verify via the `vellum events` + generated-answer pattern in §3, then:
vellum retire clitest --yes

© vellum-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/cli-testing of vellum-ai/vellum-assistant.

Open the folder on GitHubat commit c92ead1

Compare with similar skills

CLI Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

CLI Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
CLI Testing this skillvellum-ai/vellum-assistant1.4k—~1.6kAutomated safety check: PassMIT
Borg Live Debugkaranhudia/borg-ui1.7k—~1.4kAutomated safety check: NotesAGPL-3.0
.NET Crash Dump Collectiondotnet/skills5.6k2 repos~1.1kAutomated safety check: PassMIT
Test Minecraft Exporterdirien/minecraft-prometheus-exporter142—~1.6kAutomated safety check: PassApache-2.0
Build Imageskubernetes-sigs/cloud-provider-azure294—~1.7kAutomated safety check: PassApache-2.0
Agentdock User Guideuvwt/agentdock1.2k—~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • Borg Live Debug

    karanhudia/borg-ui

    Live Borg debugging by exec-ing into the borg-web-ui Docker container.

    1.7k GitHub stars~1.4k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Official

    Configures automatic crash dumps or captures dumps from running processes for modern .NET apps on Linux, macOS and Windows, including Docker and Kubernetes.

    5.6k GitHub starsUsed in 2 repos~1.1k tokens
    DevOps & CloudAuto-check passed
  • Test Minecraft Exporter

    dirien/minecraft-prometheus-exporter

    End-to-end docker-compose test harness for the minecraft-prometheus-exporter.

    142 GitHub stars~1.6k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Build Images

    kubernetes-sigs/cloud-provider-azure

    Official

    Build cloud-provider-azure container images through the repo Makefile with explicit IMAGETAG and IMAGEREGISTRY inputs, optional make flag overrides, and opt-in bounded Docker or Podman retries.

    294 GitHub stars~1.7k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Agentdock User Guide

    uvwt/agentdock

    当用户询问 AgentDock 是什么、如何使用、配置在哪里、不同平台或安装方式怎样修改配置并生效、如何重启或验证配置、如何发现并配置 Codex/Claude/Grok 等 Coding Agent 的 ACP,以及常见运行问题时使用;覆盖 macOS Desktop、Windows Desktop、Linux 服务、Docker 和直接运行二进制,不用于源码开发与贡献流程。

    1.2k GitHub stars~1.6k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Plugin Test

    apache/skywalking-python

    Build Docker test images and run SkyWalking Python plugin/unit/e2e tests locally, mirroring the CI pipeline

    219 GitHub stars~2.3k tokensUpdated today
    DevOps & CloudAuto-check passed

More from vellum-ai/vellum-assistant

All 108 skills in this repo
  • Vellum GitHub App Setup

    vellum-ai/vellum-assistant

    Create and configure a GitHub App so the assistant can push commits, open PRs, and comment under its own bot identity.

    1.4k GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Discord App Setup

    vellum-ai/vellum-assistant

    Connect a Discord bot to the assistant via the Discord Gateway with guided application creation and intent configuration

    1.4k GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Sentry App Setup

    vellum-ai/vellum-assistant

    Create and configure a Sentry internal integration so the assistant can manage issues, alerts, and releases under its own identity

    1.4k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Memory Corpus Ingest

    vellum-ai/vellum-assistant

    Ingest a large dataset into memory as a skimmed map. An agent skill from vellum-ai/vellum-assistant.

    1.4k GitHub stars~3k tokensUpdated today
    Auto-check: notes
  • Plugin Builder

    vellum-ai/vellum-assistant

    A skill your agent uses when the user wants to build, scaffold, ship, or edit a Vellum plugin that bundles multiple surfaces (hooks, tools, skills, and more) into one installable package.

    1.4k GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Slack App Setup

    vellum-ai/vellum-assistant

    Connect a Slack app to the Vellum Assistant via Socket Mode.

    1.4k GitHub stars~2.5k tokensUpdated today
    Auto-check: warnings

Works with

Questions about CLI Testing

What does CLI Testing do?

Manually test a running Vellum assistant end-to-end purely from the CLI — no desktop app or web UI. CLI Testing is an agent skill from vellum-ai/vellum-assistant. Manually test a running Vellum assistant end-to-end purely from the CLI — no desktop app or web UI.

When should I use CLI Testing?

CLI Testing fits situations like: verifying assistant behavior; reproducing a bug; smoke-testing a change without the macOS/web clients.

How do I install CLI Testing in Claude Code?

Run `npx skills add vellum-ai/vellum-assistant --skill cli-testing -a claude-code`. Or copy the skill folder (.claude/skills/cli-testing in vellum-ai/vellum-assistant) into .claude/skills/cli-testing in your project. Claude Code loads it when a task matches its description.

How do I install CLI Testing in Codex?

Run `npx skills add vellum-ai/vellum-assistant --skill cli-testing -a codex`. Or copy the skill folder (.claude/skills/cli-testing in vellum-ai/vellum-assistant) into .agents/skills/cli-testing in your project. Codex loads it when a task matches its description.

Can I use CLI Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vellum-ai/vellum-assistant --skill cli-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cli-testing, .gemini/skills/cli-testing, .github/skills/cli-testing and .opencode/skills/cli-testing in your project.

What does CLI Testing need to run?

Going by SKILL.md and its folder, CLI Testing needs credentials named ANTHROPIC_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY and FIREWORKS_API_KEY. Our summary lists: Docker; A credential in ANTHROPIC_API_KEY; A credential in OPENAI_API_KEY.

Does CLI Testing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is CLI Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does CLI Testing use?

CLI Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does CLI Testing use?

About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to CLI Testing?

Skills that share tags, products or a category with CLI Testing: Borg Live Debug (karanhudia/borg-ui, 1.7k stars), .NET Crash Dump Collection (dotnet/skills, 5.6k stars), Test Minecraft Exporter (dirien/minecraft-prometheus-exporter, 142 stars) and Build Images (kubernetes-sigs/cloud-provider-azure, 294 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains CLI Testing?

vellum-ai (a GitHub organization) maintains it in vellum-ai/vellum-assistant, which has 1,397 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 7, 2026.

Source: vellum-ai/vellum-assistant on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.