Verify or reproduce visible product behavior by delegating to Oz's dedicated computer-use capability, requiring a native Oz video artifact for meaningful UI flows and durable Oz run/artifact links.

MITAuto-check passedProductivity & Automation

Install Verify Behavior

skills CLI
$ npx skills add warpdotdev-demos/cloud-factory-demo --skill verify-behavior -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install warpdotdev-demos/cloud-factory-demo verify-behavior --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/warpdotdev-demos/cloud-factory-demo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/verify-behavior .claude/skills/verify-behavior && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
verify-behavior
GitHub stars
320
Token cost
~2.2k tokens
SKILL.md length
1,008 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Verify or reproduce visible product behavior by delegating to Oz's dedicated computer-use capability, requiring a native Oz video artifact for meaningful UI flows and durable Oz run/artifact links.

  • Works in 4 steps: Read it (parent path, specs//PRODUCT.md,… → Build a checklist of exercisable… → reproduce: stories tied to the failure +… → …
  • Triage needs visual reproduction
  • SKILL.md covers Modes, Platform contract (do not…, When to use / skip and PRODUCT.md coverage, plus 4 more sections
  • Calls gh; reaches oz.warp.dev and oz.staging.warp.dev

What it does

Verify Behavior is an agent skill from warpdotdev-demos/cloud-factory-demo. Verify or reproduce visible product behavior by delegating to Oz's dedicated computer-use capability, requiring a native Oz video artifact for meaningful UI flows and durable Oz run/artifact links. Use when triage needs visual reproduction, implementation needs behavioral proof, review needs interactive confirmation, or any factory stage asks to verify UI/app behavior.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Productivity & Automation, covering Desktop control. The licence is MIT.

When your agent uses it

  • Triage needs visual reproduction
  • Implementation needs behavioral proof
  • Review needs interactive confirmation
  • Any factory stage asks to verify UI/app behavior

Example prompts

  • “/verify-behavior”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Read it (parent path, specs//PRODUCT.md, or issue/PR links).
  2. Build a checklist of exercisable user-facing stories. Newer issue comments win on conflict.
  3. reproduce: stories tied to the failure + path to reach it. verify: all in-scope stories (or one assigned story if fanning out).
  4. No PRODUCT.md → derive stories from the issue.

What it can do on your machine

Read from SKILL.md and the folder at commit ab21d0c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • oz.warp.dev
    • oz.staging.warp.dev

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Verify Behavior loads about 2.2k tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 1,008 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from warpdotdev-demos/cloud-factory-demo at commit ab21d0c, republished under its MIT licence (© warpdotdev-demos). 1,008 words, ~2,224 tokens.

Download SKILL.mdSave it as .claude/skills/verify-behavior/SKILL.md (or your agent's skills folder).
name
verify-behavior
description
Verify or reproduce visible product behavior by delegating to Oz's dedicated computer-use capability, requiring a native Oz video artifact for meaningful UI flows and durable Oz run/artifact links. Use when triage needs visual reproduction, implementation needs behavioral proof, review needs interactive confirmation, or any factory stage asks to verify UI/app behavior.

Verify behavior

Prove or disprove visible product behavior for bugs and greenfield features. Parents (triage, implementation, review) invoke this instead of driving the UI themselves.

Modes

  • reproduce — does the reported bug still happen on baseline? (usually default branch; triage)
  • verify — does the implemented change match expected behavior? (features and fixes; implementation/review)

Infer if unnamed: issue-only → reproduce; implementation/PR branch → verify.

Platform contract (do not invent alternatives)

Oz computer use is a dedicated capability, not a generic remote child and not a set of parent-callable capture APIs.

LayerWhat happens
Outer Oz cloud runMust already have computer use enabled (computer_use: true / --computer-use / computer_use_enabled: true). Factory implement workflows set this on the parent run.
Parent / orchestratorDelegates GUI work through the platform computer_use capability with a concrete task and expected evidence.
Permissionrequest_computer_use is emitted automatically when the computer-use subagent starts. Parents must not call it, simulate it, or document fake tool calls for it.
Computer-use subagentNative prompt already covers GUI actions, proactive start_recording / stop_recording, report_screenshot, and forbids ffmpeg / screen-capture CLIs. Do not re-teach capture plumbing or shell out to recorders.

Never:

  • Call or invent direct request_computer_use, start_recording, or stop_recording from the parent LLM
  • Launch generic run_agents remote children as a substitute for computer_use
  • Build custom recorder pipelines (ffmpeg, x11grab, avfoundation, Playwright video, etc.)
  • Claim gh ... --attach can upload Oz videos to GitHub (it cannot reliably attach native Oz recordings)

When to use / skip

Use for UI, browser, desktop, mobile, or other interactive behavior where visual proof helps.

Skip pure backend/CI/text-only work, or when credentials/state are unavailable (report the blocker). If this skill file is missing, parents continue without visual verification.

PRODUCT.md coverage

When PRODUCT.md exists it is the primary source of stories and acceptance criteria:

  1. Read it (parent path, specs/<issue-slug>/PRODUCT.md, or issue/PR links).
  2. Build a checklist of exercisable user-facing stories. Newer issue comments win on conflict.
  3. reproduce: stories tied to the failure + path to reach it. verify: all in-scope stories (or one assigned story if fanning out).
  4. No PRODUCT.md → derive stories from the issue.

Parent workflow

  1. Collect mode, issue/PR, branch/ref, setup commands/URLs, surface hints, PRODUCT.md path/excerpts, and expected behavior.
  2. Skip if not visually observable; report why.
  3. Confirm the current Oz run has computer use enabled. If it does not and you cannot enable it, stop with an explicit blocker — do not fake verification.
  4. For each verification unit (full checklist, or one story when fanning out), delegate once to the dedicated computer_use capability with a task that includes:
    • Mode (reproduce | verify)
    • Branch/ref and app setup
    • Exact story/repro path and acceptance checks
    • Instruction to exercise the critical path end-to-end
    • Expected evidence: native session video of the meaningful UI flow (subagent records proactively) plus reported screenshots of key states
    • Instruction to return observations and steps, not a vague pass/fail monologue
  5. Multi-story verify: fan out bounded parallel computer_use delegations (about 4–6), one key independent story each; stay single-delegation for reproduce, one story, sequential stories, or scarce credentials.
  6. Aggregate statuses: confirmed / partially confirmed / not reproduced / verified / partially verified / not verified / blocked.
  7. Fold evidence into triage comments, PR bodies, or review.json. Do not claim behavioral verification without evidence or an explicit blocker.
Delegation task shape
text
Mode: verify | reproduce
Branch/ref: <implementation head or baseline>
Setup: <install/start commands or URL>
Story or repro: <one checklist item or full path>
Checks: <what must be visibly true>
Evidence required:
- Native Oz computer-use recording of the critical interactive path
- Reported screenshots for baseline, key action, and outcome states
Return: steps taken, observations per check, blockers, artifact/run references

Prefer browser-oriented interaction when the product is a pure web app and the environment provides browser automation inside computer use; prefer full desktop computer use for native surfaces, OS dialogs, or multi-window flows. Record Channel: browser | desktop | hybrid and a one-line why in the parent summary.

Show full SKILL.md (432 more words)Show less

Evidence requirements

For every meaningful UI flow:

  1. Native Oz video artifact — produced inside the computer-use session via platform recording. Required unless the subagent reports that recording failed; quote that failure.
  2. Durable Oz links — include https://oz.warp.dev/runs/<run-id> (or https://oz.staging.warp.dev/runs/<run-id> on staging). Never app.warp.dev or /run/ (singular).
  3. Reported screenshots — supplementary keyframes via the subagent's report_screenshot path; useful, not a substitute for video on multi-step flows.
  4. PR / issue write-up — status, story checklist outcomes, Oz run link(s), and pointers to the native video/screenshot artifacts on the Oz run. Prefer platform PR artifact helpers (e.g. get_artifacts_for_pull_request_description) when available so managed screenshot/video markdown is embedded. Do not invent gh --attach video uploads.

Screenshots-only is allowed only when recording was attempted and failed, with the error quoted, or when the check is a single static state with no meaningful motion.

A verification that never delegated to computer_use on a computer-use-enabled run is incomplete.

Screenshot captions (required on GitHub)

Whenever a screenshot is posted back to a GitHub issue or PR — including platform-managed evidence blocks, PR body sections, issue comments, and any other evidence summary — put a concise caption immediately below each screenshot that states:

  1. The UI state or test case being verified (e.g. idle preview, in-flight submit, settled error)
  2. What the screenshot demonstrates relative to the acceptance check (e.g. overlay present with spinner; overlay cleared)

Apply captions for managed markdown (e.g. ![…](…) from artifact helpers) and for any hand-written evidence summary. Prefer short one-liners under each image; do not dump long prose. If the platform helper only supplies alt text, still add an explicit caption line under the image in the issue/PR write-up when you control that text.

Do not invent unsupported upload commands to satisfy captions. Captions travel with whatever evidence path the platform already supports.

Report shape

text
Behavior verification:
- Mode / change type / channel(s) / fan-out
- Status / issue-PR / branch-ref
- PRODUCT.md stories: n passed / failed / blocked / not run
- Oz run: https://oz.warp.dev/runs/<id>
- Evidence: native video (required for flows) + captioned screenshots (supplementary)
- Findings / next step

reproduce statuses: confirmed | partially confirmed | not reproduced | blocked
verify statuses: verified | partially verified | not verified | blocked

Guardrails

  • Delegate GUI work through computer_use; parents do not click through unless the user explicitly asks
  • Outer run must have computer use enabled; do not substitute generic remote children
  • Do not call request_computer_use or parent-level recording APIs
  • No ffmpeg or external recorders
  • Native Oz video required for meaningful UI flows; screenshots supplement
  • Every GitHub-posted screenshot needs a concise caption (state/case + what it shows)
  • Durable Oz run/artifact links on issue and PR write-ups; no fake GitHub binary-attach of Oz recordings
  • Features are first-class; PRODUCT.md drives coverage when present
  • Bounded fan-out; no secrets in prompts, screenshots, or reports
  • No claim of verification without evidence or an explicit blocker

Optional verify-behavior-local may specialize setup/surface/fan-out only — not weaken evidence, privacy, captions, or this contract.

© warpdotdev-demos, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/verify-behavior of warpdotdev-demos/cloud-factory-demo.

Open the folder on GitHubat commit ab21d0c

Compare with similar skills

Verify Behavior next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Verify Behavior compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Verify Behavior this skillwarpdotdev-demos/cloud-factory-demo320—~2.2kAutomated safety check: PassMIT
Vision SkillsAnionex/agent-vision-toolkit1.2k—~3.9kAutomated safety check: PassMIT
Mac Computer UseTo3akaRin/mac-computer-use1.1k—~495Automated safety check: PassMIT
Agent Managementautonomous-ai/Physical-AI-Operating-System381—~1.7kAutomated safety check: PassApache-2.0
Computer Usebam-bam-2/solo-skills3671 repos~915Automated safety check: PassMIT
Verify UI Change In CloudHeartcoolman/warp-cn119—~1.1kAutomated safety check: PassAGPL-3.0

Similar skills

  • Vision Skills

    Anionex/agent-vision-toolkit

    Local vision CLIs: glance (describe/ask/OCR an image), ground (locate a target, pixel box), detect (element inventory), trace (image to SVG geometry), crop (cut a pixel box to a file), and…

    1.2k GitHub stars~3.9k tokensUpdated 5 days ago
    Productivity & AutomationAuto-check passed
  • Mac Computer Use

    To3akaRin/mac-computer-use

    操作 macOS 桌面应用,探测窗口和自动化接口、截图、读取或修改辅助功能元素、执行鼠标键盘动作,以及通过 CDP 操作内嵌 Chromium 页面。适用于桌面应用自动化与界面验收;普通网页任务优先使用已有浏览器工具。

    1.1k GitHub stars~495 tokensUpdated 13 days ago
    Productivity & AutomationAuto-check passed
  • Agent Management

    autonomous-ai/Physical-AI-Operating-System

    Legacy Autonomous Buddy control for explicitly requested Buddy coding sessions.

    381 GitHub stars~1.7k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Computer Use

    bam-bam-2/solo-skills

    Use Orca's computer-use CLI to inspect and operate local desktop app windows through accessibility trees, screenshots, and safe UI actions.

    367 GitHub starsUsed in 1 repo~915 tokens
    Productivity & AutomationAuto-check passed
  • Verify UI Change In Cloud

    Heartcoolman/warp-cn

    Invoke this automatically after completing any user-facing client change, ONLY in non-sandboxed environments and local environments.

    119 GitHub stars~1.1k tokensUpdated 2 days ago
    Productivity & AutomationAuto-check passed
  • Verify Reel

    rselbach/reel

    Verify Reel, the macOS menu-bar screen recorder, by launching a disposable app bundle and driving its real UI with Computer Use.

    118 GitHub stars~2k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check passed

More from warpdotdev-demos/cloud-factory-demo

  • Improve Review PR

    warpdotdev-demos/cloud-factory-demo

    Daily outer loop that reviews human reactions to automated review-pr comments, synthesizes durable organizational knowledge, and opens a PR to update the review-pr skill when the feedback is worth…

    320 GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Review PR

    warpdotdev-demos/cloud-factory-demo

    Review a PR from local annotated-diff artifacts and write validated review.json for the workflow to publish.

    320 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Triage

    warpdotdev-demos/cloud-factory-demo

    Triage an incoming GitHub, Jira, Linear, or other issue-tracker issue against the current codebase and related open issues, then return a structured decision with exactly one…

    320 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Oz Cloud Factory Demo

    warpdotdev-demos/cloud-factory-demo

    Sets up a beginner-friendly Oz cloud software factory that automatically triages new GitHub issues, specs issues labeled ready-to-spec, implements issues labeled ready-to-implement, reviews PRs, and…

    320 GitHub stars~6.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Implementation

    warpdotdev-demos/cloud-factory-demo

    Implement a fix or feature from a GitHub, Jira, Linear, or other issue-tracker issue by fetching issue context, inspecting the current codebase, making code changes, validating them, verifying…

    320 GitHub stars~3.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Spec

    warpdotdev-demos/cloud-factory-demo

    Coordinate spec-driven development for a GitHub, Jira, Linear, or other issue-tracker issue marked ready-to-spec by using write-product-spec and write-tech-spec, creating PRODUCT.md and TECH.md…

    320 GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Verify Behavior

What does Verify Behavior do?

Verify or reproduce visible product behavior by delegating to Oz's dedicated computer-use capability, requiring a native Oz video artifact for meaningful UI flows and durable Oz run/artifact links. Verify Behavior is an agent skill from warpdotdev-demos/cloud-factory-demo. Verify or reproduce visible product behavior by delegating to Oz's dedicated computer-use capability, requiring a native Oz video artifact for meaningful UI flows and durable Oz run/artifact links.

When should I use Verify Behavior?

Verify Behavior fits situations like: triage needs visual reproduction; implementation needs behavioral proof; review needs interactive confirmation; any factory stage asks to verify UI/app behavior.

How do I install Verify Behavior in Claude Code?

Run `npx skills add warpdotdev-demos/cloud-factory-demo --skill verify-behavior -a claude-code`. Or copy the skill folder (.agents/skills/verify-behavior in warpdotdev-demos/cloud-factory-demo) into .claude/skills/verify-behavior in your project. Claude Code loads it when a task matches its description.

How do I install Verify Behavior in Codex?

Run `npx skills add warpdotdev-demos/cloud-factory-demo --skill verify-behavior -a codex`. Or copy the skill folder (.agents/skills/verify-behavior in warpdotdev-demos/cloud-factory-demo) into .agents/skills/verify-behavior in your project. Codex loads it when a task matches its description.

Can I use Verify Behavior in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add warpdotdev-demos/cloud-factory-demo --skill verify-behavior -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/verify-behavior, .gemini/skills/verify-behavior, .github/skills/verify-behavior and .opencode/skills/verify-behavior in your project.

What does Verify Behavior need to run?

Going by SKILL.md and its folder, Verify Behavior needs the command-line tools its instructions call (gh).

Does Verify Behavior access the network?

SKILL.md names 2 domains. In commands or code: oz.warp.dev and oz.staging.warp.dev; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Verify Behavior safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Verify Behavior use?

Verify Behavior is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Verify Behavior use?

About 2.2k tokens (SKILL.md is roughly 8.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Verify Behavior?

Skills that share tags, products or a category with Verify Behavior: Vision Skills (Anionex/agent-vision-toolkit, 1.2k stars), Mac Computer Use (To3akaRin/mac-computer-use, 1.1k stars), Agent Management (autonomous-ai/Physical-AI-Operating-System, 381 stars) and Computer Use (bam-bam-2/solo-skills, 367 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Verify Behavior?

warpdotdev-demos (a GitHub organization) maintains it in warpdotdev-demos/cloud-factory-demo, which has 320 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on August 12, 2026.

Source: warpdotdev-demos/cloud-factory-demo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.