Agent skill

Verify Omnigent End-to-End

by omnigent-ai in omnigent-ai/omnigent

Spins up an isolated Omnigent server, runner and mock model to prove a user-facing behavior or bug fix with recorded evidence instead of reasoning from code.

Apache-2.0Auto-check passedTesting & QA

Install Verify Omnigent End-to-End

skills CLI
$ npx skills add omnigent-ai/omnigent --skill verify-omnigent -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install omnigent-ai/omnigent verify-omnigent --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/omnigent-ai/omnigent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/feature-map/skills/verify-omnigent .claude/skills/verify-omnigent && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
verify-omnigent
GitHub stars
11k
Token cost
~1.6k tokens
SKILL.md length
799 words
Files
2 (incl. scripts)
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

Spins up an isolated Omnigent server, runner and mock model to prove a user-facing behavior or bug fix with recorded evidence instead of reasoning from code.

  • Works in 2 steps: Install dependencies, Chromium, and… → Start an instance and load its paths
  • Reproducing a user-facing bug before attempting a fix
  • SKILL.md covers Launch, Doctor, Drive and Evidence, plus 3 more sections
  • Calls uv, pnpm and python

What it does

This skill has two parts: an isolated instance wrapping a repro-environment module that starts a server, runner and mock model server on private ports with their own config, data, Claude and Codex directories, never touching the real home install, a running host daemon, or another developer's server; and a feature map listing every user-facing feature's entry points, the tests that drive them, and known traps, where a fix only counts as verified once every listed entry point for its feature has proof.

Starting an instance installs dependencies, Chromium, and builds the web UI, then launches the instance and waits roughly ten seconds until the server, runner and mock model all answer; the instance writes its root path so later shell sessions can find it, and it stops itself automatically after a lease that defaults to 90 minutes. A doctor command checks readiness, supervisor health, and whether the runner and model server still answer, and should run before the first drive and after any failed one.

Readiness does not include checking that the browser actually launches, and native-terminal journeys need their own CLI and terminal prerequisites such as tmux beyond just a browser install; in CI, a repro-environment exec command reuses the environment the workflow already started instead of spinning up a second one.

When your agent uses it

  • Reproducing a user-facing bug before attempting a fix
  • Proving a fix works across every entry point for its feature
  • Checking whether every native harness terminal journey is covered

Example prompts

  • “Start a verify-omnigent instance and reproduce this reported bug.”
  • “Prove the composer fix works in both the session and new-session surfaces.”
  • “Run the doctor check before I start driving this verification.”

Requirements

  • uv
  • pnpm
  • Playwright with Chromium
  • tmux (for native-terminal journeys)

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Install dependencies, Chromium, and build the web UI for the checkout
  2. Start an instance and load its paths

What it can do on your machine

Read from SKILL.md and the folder at commit 2e1cd15. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • uv
    • pnpm
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv and pnpm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Verify Omnigent End-to-End loads about 1.6k tokens when it runs. Until then it costs about 119 tokens; SKILL.md has 799 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~119
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from omnigent-ai/omnigent at commit 2e1cd15, republished under its Apache-2.0 licence (© omnigent-ai). 799 words, ~1,631 tokens.

Download SKILL.mdSave it as .claude/skills/verify-omnigent/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
verify-omnigent
description
Drive Omnigent the way a user does and prove a behavior with recorded evidence, using an isolated server, runner, and mock model plus a feature map of every user entry point. Load before reproducing a user-facing bug, before claiming a fix works, or when reviewing whether a change covered every surface (session vs. new-session composer, terminal strip vs. agent terminal, all twelve native harnesses). Covers the web UI, native harness terminals, and the CLI.

Verify Omnigent

Use this skill to see a behavior happen in the real app and to prove a change, not to reason about it from code. It has two parts:

  • An isolated instance. scripts/verify-env wraps python -m dev.repro_env: a server, runner, and mock model server on private ports, with their own config, data, Claude, and Codex directories. It never touches ~/.omnigent, a running host daemon, or another developer server.
  • A feature map. feature map lists each user-facing feature's entry points, the tests that drive them, and the traps. A fix is verified only when every entry point listed for its feature has proof.

Run all commands from the repository root. Put scripts/ on your path or call feature-map/skills/verify-omnigent/scripts/verify-env directly.

Launch

  1. Install dependencies, Chromium, and build the web UI for the checkout:

    sh
    uv sync --frozen --extra all --group test
    uv run --no-sync playwright install --with-deps chromium
    pnpm install --frozen-lockfile --filter web && pnpm --filter web run build
  2. Start an instance and load its paths:

    sh
    verify-env start          # waits until the runner is online, about 10 seconds
    eval "$(verify-env paths)"

    start writes VERIFY_ROOT to .omnigent/verify/current, so later shells find the same instance. The instance stops itself after its lease (default 90 minutes; --lease SECONDS, 60 to 21600).

Ready means the server answers, the runner reports online, and the mock model server answers. start fails with the environment's error and log path otherwise.

Ready does not check browser launch. Use the Chromium version installed by this checkout's Playwright package; if the environment provides it through PLAYWRIGHT_BROWSERS_PATH, keep that path available to the test process. --ui-skip-build reuses the existing bundle, so rebuild after frontend edits. Native-terminal journeys also need the relevant CLI and terminal prerequisites (such as tmux); a browser installation alone does not provide them.

In CI, the repro workflow already runs this environment. Use python -m dev.repro_env exec -- ... there instead of starting another.

Doctor

Run verify-env doctor before the first drive, after any failed drive, and whenever something looks off. It checks that the instance is ready, that its supervisor is alive, that the runner and model server answer, and it warns when the checkout has moved since launch. A warning about a moved checkout means the instance runs old code: stop it and start a new one.

Drive

  1. Open the matching file in feature map and list every entry point for the behavior in question.

  2. For each entry point, run the named test through the instance, recording on:

    sh
    verify-env run -- python -m pytest <tests/...py::test_name> \
      --ui-skip-build --video=on --screenshot=on \
      --output="$VERIFY_EVIDENCE/<feature>"

    Tests under tests/browser_ui/ need no instance: uv run pytest <test> --browser-ui-skip-build --video=on --output=.... Tests the feature file marks "own environment" also run with plain uv run pytest. Server/transport/component tests use plain pytest without browser recording or build flags.

  3. For an entry point with no test, hand-drive it with Playwright against OMNIGENT_REPRO_SERVER_URL inside verify-env run -- python <script>, and script model replies with the mock helpers named in the feature map README.

  4. To reproduce a bug, run the same journey on the unfixed code first and keep that evidence; then run it again on the fix.

A test that fails before reaching the user action has a setup failure, not a reproduction. Check its browser error and server/runner logs. When a test starts its own server or runner from inside an agent, check whether it inherited the parent's OMNIGENT_RUNNER_*, RUNNER_SERVER_URL, or OMNIGENT_PROCESS_LOG_FILE. Use an isolated child environment for that test; do not change the controlling agent's environment. Record setup failures, skipped variants, and unavailable credentials separately from behavior results.

Show full SKILL.md (259 more words)Show less

Evidence

Everything under $VERIFY_EVIDENCE is the proof, one directory per feature. Instance logs, the database, and the mock model's recorded requests stay in $VERIFY_ENV. For each claim, record the checkout commit, installed harness versions, feature file, entry point ID, command, and resulting artifact. Keep source review, component tests, real process checks, browser drives, and live provider checks distinct. A map contract pass verifies references and structure; it is not evidence that the journeys ran. The proof standards are in the feature map README. Mock runs prove Omnigent's integration with Claude and Codex, not a live vendor model.

Cleanup

verify-env stop stops only this instance's supervisor, which stops its own server, runner, and model processes and saves the model's request log. It never deletes $VERIFY_ROOT, so the evidence survives; remove old roots under .omnigent/verify/ or /tmp/verify-omnigent.* yourself once the evidence is attached. Never kill Omnigent processes by name: a developer's own server and host daemon may be running on the same machine.

Helpers

scripts/verify-env is the only helper. Its subcommands are start [--lease SECONDS], doctor, run -- COMMAND..., paths, and stop; run it with no arguments for usage.

On macOS the helper places the instance under /tmp/verify-omnigent.* when the checkout path is long. macOS limits Unix socket paths to 104 bytes, and the environment's socket relay only works around long paths on Linux.

Keeping the map honest

When a reviewer says a change missed a surface, add that surface to the feature file in the same change. The checks and the weekly upkeep job are described in Keeping the map current.

© omnigent-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in feature-map/skills/verify-omnigent of omnigent-ai/omnigent.

  • SKILL.md
  • scripts/verify-env

Open the folder on GitHubat commit 2e1cd15

Compare with similar skills

Verify Omnigent End-to-End next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Verify Omnigent End-to-End compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Verify Omnigent End-to-End this skillomnigent-ai/omnigent11k—~1.6kAutomated safety check: PassApache-2.0
CodexBar Live QAsteipete/CodexBar22k—~1.2kAutomated safety check: PassMIT
Senpi Agent QA Harnesscode-yeongyu/senpi472—~2.7kAutomated safety check: NotesMIT
tmux Real User TestingQwenLM/qwen-code28k—~2.3kAutomated safety check: PassApache-2.0
Glance TestDebugBase/glance156—~827Automated safety check: PassMIT
Create a Verification Skillcursor/plugins10k8 repos~1.5kAutomated safety check: PassNone

Similar skills

  • CodexBar Live QA

    steipete/CodexBar

    Runs live QA for the CodexBar app: provider usage matrix checks through its packaged CLI, config validation and menu checks, with 1Password-backed credentials handled safely.

    22k GitHub stars~1.2k tokensUpdated today
    Testing & QAAuto-check passed
  • Senpi Agent QA Harness

    code-yeongyu/senpi

    Checks changes to the senpi coding agent by driving the real CLI from source in an isolated sandbox, over RPC, terminal UI, mock model and CLI smoke channels.

    472 GitHub stars~2.7k tokensUpdated today
    Testing & QAAuto-check: notes
  • tmux Real User Testing

    QwenLM/qwen-code

    Drives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review.

    28k GitHub stars~2.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Glance Test

    DebugBase/glance

    Run E2E browser tests on any web application using Glance MCP.

    156 GitHub stars~827 tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Official

    Generates a project-local skill that launches your app, exercises a feature the way a user would and captures evidence, for web, CLI, API or desktop projects.

    10k GitHub starsUsed in 8 repos~1.5k tokens
    Testing & QAAuto-check passed
  • Twenty QA Scout

    twentyhq/twenty

    Browser QA for a pull request against a running Twenty app: scenarios drawn from the diff, run in a real browser, checked in the database and logs, and closed with a verdict and report.

    58k GitHub stars~2k tokensUpdated today
    Testing & QAAuto-check passed

More from omnigent-ai/omnigent

All 19 skills in this repo
  • Omnigent Docker Compose Deploy

    omnigent-ai/omnigent

    Brings up the Omnigent server and Postgres as a Docker compose stack on any Docker host, and covers the Dockerfile's runtime and host build targets for extending it to a new platform.

    11k GitHub stars~1.3k tokensUpdated today
    Auto-check: notes
  • Omnigent Framework Detection

    omnigent-ai/omnigent

    Scans Python agent code for framework imports and recommends the matching Omnigent executor type, or says when the framework is not natively supported yet.

    11k GitHub stars~610 tokensUpdated today
    Auto-check passed
  • Omnigent Load Test Runner

    omnigent-ai/omnigent

    Runs the Omnigent load test with real hosts and multi-turn sessions against a mocked LLM, then explains the latency results from summary.md.

    11k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Spins up a local Omnigent server and exercises the Antigravity (Gemini) SDK harness end to end: building agents, running real turns, smoke tests and bug-bashing.

    11k GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Omnigent Agent Builder

    omnigent-ai/omnigent

    Gives patterns for generating a minimal, valid Omnigent agent directory: the config.yaml fields, the right executor type, and the files each agent needs.

    11k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Copilot SDK E2E Dev

    omnigent-ai/omnigent

    Spin up a live local Omnigent server and exercise the GitHub Copilot SDK harness end-to-end — build copilot agents, run real turns, smoke-test, and bug-bash.

    11k GitHub stars~2.7k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Verify Omnigent End-to-End

What does Verify Omnigent End-to-End do?

Spins up an isolated Omnigent server, runner and mock model to prove a user-facing behavior or bug fix with recorded evidence instead of reasoning from code. This skill has two parts: an isolated instance wrapping a repro-environment module that starts a server, runner and mock model server on private ports with their own config, data, Claude and Codex directories, never touching the real home install, a running host daemon, or another developer's server; and a feature map listing every user-facing feature's entry points, the tests that drive them, and known traps, where a fix only counts as verified once every listed entry point for its feature has proof.

When should I use Verify Omnigent End-to-End?

Verify Omnigent End-to-End fits situations like: reproducing a user-facing bug before attempting a fix; proving a fix works across every entry point for its feature; checking whether every native harness terminal journey is covered.

How do I install Verify Omnigent End-to-End in Claude Code?

Run `npx skills add omnigent-ai/omnigent --skill verify-omnigent -a claude-code`. Or copy the skill folder (feature-map/skills/verify-omnigent in omnigent-ai/omnigent) into .claude/skills/verify-omnigent in your project. Claude Code loads it when a task matches its description.

How do I install Verify Omnigent End-to-End in Codex?

Run `npx skills add omnigent-ai/omnigent --skill verify-omnigent -a codex`. Or copy the skill folder (feature-map/skills/verify-omnigent in omnigent-ai/omnigent) into .agents/skills/verify-omnigent in your project. Codex loads it when a task matches its description.

Can I use Verify Omnigent End-to-End in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add omnigent-ai/omnigent --skill verify-omnigent -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/verify-omnigent, .gemini/skills/verify-omnigent, .github/skills/verify-omnigent and .opencode/skills/verify-omnigent in your project.

What does Verify Omnigent End-to-End need to run?

Going by SKILL.md and its folder, Verify Omnigent End-to-End needs the command-line tools its instructions call (uv, pnpm and python). Our summary lists: uv; pnpm; Playwright with Chromium; tmux (for native-terminal journeys).

Does Verify Omnigent End-to-End access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Verify Omnigent End-to-End safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Verify Omnigent End-to-End use?

Verify Omnigent End-to-End is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Verify Omnigent End-to-End use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Verify Omnigent End-to-End?

Skills that share tags, products or a category with Verify Omnigent End-to-End: CodexBar Live QA (steipete/CodexBar, 22k stars), Senpi Agent QA Harness (code-yeongyu/senpi, 472 stars), tmux Real User Testing (QwenLM/qwen-code, 28k stars) and Glance Test (DebugBase/glance, 156 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Verify Omnigent End-to-End?

omnigent-ai (a GitHub organization) maintains it in omnigent-ai/omnigent, which has 10,691 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 9, 2026.

Source: omnigent-ai/omnigent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.