Agent skill

Acceptance Evidence for Deliveries

by lobehub in lobehub/lobehub

Verifies a delivery end to end by driving the real product on a CLI, web, desktop or iOS Simulator surface, capturing evidence and publishing a round with the lh CLI.

Apache-2.0Auto-check passedTesting & QA

Install Acceptance Evidence for Deliveries

skills CLI
$ npx skills add lobehub/lobehub --skill acceptance -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install lobehub/lobehub acceptance --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/lobehub/lobehub.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/acceptance .claude/skills/acceptance && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
acceptance
GitHub stars
83k
Token cost
~9.7k tokens
SKILL.md length
4,243 words
Files
34 (incl. scripts, references)
Skills in repo
50
Repo updated
First seen
Licence
Apache-2.0

At a glance

Verifies a delivery end to end by driving the real product on a CLI, web, desktop or iOS Simulator surface, capturing evidence and publishing a round with the lh CLI.

  • Works in 6 steps: Use the named acceptance, or create one… → lh acceptance flow publish --file… → lh acceptance flow plan --flow creates a… → …
  • Proving that a finished change works before handing it over
  • SKILL.md covers Decide whether to execute…, Independent acceptance review…, Read the project layer first and Living logs — inject each by…, plus 10 more sections
  • Calls rg, git and bun; reaches lobehub.com

What it does

The agent acts as the builder of a delivery whose claims are judged by a separate review step against a plan, either one handed to the run or checks it writes itself. Any check that declares required evidence cannot pass on the agent's words alone: a missing artifact marks it uncertain and holds the delivery. The flow is to author or discover the plan, pick the surface, capture evidence, publish the round and check coverage.

Before starting, the agent decides whether to run at all. Creating or updating a PR, marking it ready or being asked to upload a report does not by itself trigger another verification run; it first looks at existing reports, evidence and published acceptance links, including a local .acceptances folder, and skips product acceptance for documentation-only or pure refactor changes while saying why.

Reference notes cover agent-browser use, web authentication, accessibility checks with axe, computer use, evidence rules, plan format, mock patterns, project adapters, recording through CDP, the iOS Simulator or native macOS, reports and resource guards. The folder is large and works in any repository, with or without a preconfigured verify plan.

When your agent uses it

  • Proving that a finished change works before handing it over
  • Capturing screenshots or recordings as evidence for a verify plan
  • Testing a desktop or Electron build end to end
  • Publishing a test report round with the lh CLI

Example prompts

  • “Verify the task end to end and collect evidence that the new export button works.”
  • “Test the desktop app and upload the recording to the acceptance report.”
  • “Write a verify plan for this change and publish the first round.”
  • “Check the existing acceptance reports before deciding whether to run again.”

Requirements

  • The lh CLI for publishing rounds
  • A runnable product on the chosen surface: CLI, web, desktop or iOS Simulator

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Use the named acceptance, or create one before publishing the flow. If none
  2. lh acceptance flow publish --file flow.json saves the
  3. lh acceptance flow plan --flow creates a draft round
  4. Share the acceptance link so the user can inspect the proposed nodes, branches
  5. Implement the work and exercise the real product, then use
  6. After all required checks are recorded and passed, run

What it can do on your machine

Read from SKILL.md and the folder at commit 1863542. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • rg
    • git
    • bun

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • lobehub.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Acceptance Evidence for Deliveries loads about 9.7k tokens when it runs, and up to ~45k if it reads all its reference files. Until then it costs about 190 tokens; SKILL.md has 4,243 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~190
When it runs · the whole SKILL.md, loaded when a task matches
~9.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~45k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from lobehub/lobehub at commit 1863542, republished under its Apache-2.0 licence (© lobehub). 4,243 words, ~9,729 tokens.

Download SKILL.mdSave it as .claude/skills/acceptance/SKILL.md (or your agent's skills folder). This skill also uses 33 other files; get the full folder from GitHub.
name
acceptance
description
End-to-end verification and self-evidence for a delivery in any repository, with or without a preconfigured verify plan. Discover an existing plan when one was handed to this run; otherwise author checks and publish a standalone acceptance. Pick the proving surface (CLI / web / desktop / iOS Simulator), drive the real product, capture visually confirmed evidence, and publish a round with the lh CLI. Triggers on 'verify the task', 'collect evidence', 'prove it works', 'upload evidence', 'verify plan', 'requiredEvidence', 'local test', 'manual test', 'test report', 'test with cli', 'test in electron', 'test desktop', or any local end-to-end verification task. Needs no ambient ids, and never depends on running inside a LobeHub conversation.
license
Apache-2.0
metadata.version
0.7.0

Acceptance (Builder Self-Evidence)

You are the builder for a delivery. A separate review step judges it against a plan — checks you author, or a verify plan handed to this run. A check that declares requiredEvidence cannot pass on your text alone: a missing artifact marks it uncertain and holds the delivery.

author (or discover) the plan  →  pick the surface  →  capture evidence  →  publish the round  →  self-check coverage

Decide whether to execute before starting a round

Creating or updating a PR, marking it ready, or being asked to upload a report must not by itself start another verification run. First inspect the requested scope and the task's existing reports, evidence, and published acceptance links (from the conversation, PR, or local .acceptances/ directory).

Delivery stateAction
Documentation/instruction-only change, or pure refactor/tooling change with no product behavior changeSkip product acceptance and briefly state why. Keep any applicable quality checks.
Gitlink-only syncDo not launch a fresh acceptance. Link the upstream change and its existing acceptance when available; disclose missing upstream evidence without claiming it passed. Cloud changes accompanying the sync are assessed separately.
Completed acceptance already published and still covers the deliveryReuse its URL and coverage. Do not create a round or rerun cases just for the PR.
Completed acceptance report and evidence exist locally and still cover the deliveryInspect coverage and artifacts, then upload that report using report.md. Preserve the original execution provenance; no product rerun, new plan, or repeated completed checker review is needed merely for upload.
The delivery was already exercised on the real product earlier in this session (observations and raw artifacts exist, no report yet)Do not rerun, re-plan, or open a checker stage. Write plan[] and cases[] from the observations already made, attach the original artifacts (logs, command output, captures) with their original provenance, disclose any required medium that was never captured instead of recapturing it, and ingest.
Product behavior lacks valid evidence, or relevant behavior changed after verificationExecute only the missing or affected outcomes, retain unaffected evidence with its original provenance, and publish according to the round rules below.

Evidence is reusable when its criteria cover the requested behavior, its artifacts are available and support the observations, and subsequent code, dependency, configuration, or environment changes do not invalidate those observations. Compare the relevant changes; a different commit SHA, rebase, PR event, or report publication status alone is not a reason to rerun. Failed/blocked checks and missing required evidence are not passes: repair or supplement those specific gaps. An explicit user request for fresh verification still takes precedence.

The execution, environment setup, plan/checker, and capture sections below apply when executing acceptance. For reuse or upload only, inspect the existing report and evidence and complete the necessary publication/coverage steps; do not boot services or replay completed cases. Uploading does not change when or against which implementation the evidence was captured.

Independent acceptance review (first round only)

The primary checks the environment, writes the plan, executes cases, inspects evidence, repairs failures, and publishes. Use one acceptance-checker agent at two points in the first acceptance round: give at most two feedback responses on the plan and cases before execution, then perform exactly one quick report/evidence check against the agreed criteria before publishing. A second plan check is optional, only to check the primary's revisions; there is no third plan-feedback response. Count the two stages separately. After either stage's limit, the primary owns remaining corrections and verification. The acceptance-checker is the plan gate — never ask the user to approve a plan; ask the user only for a user-owned prerequisite or a product decision that changes the plan. In both stages, the primary supplies an explicit file list and the relevant diff text or prepared diff artifact paths. The acceptance-checker limits code reading to these materials; it must not run git diff or discover its own scope. This does not restrict inspection of the plan, report, or evidence. During evidence review, use it only to identify the updates and the agreed cases whose evidence needs checking; the core task is checking the report against the plan and artifacts. Do not reopen requirements, expand into code review, or investigate implementation details. Return contradictions to the primary for explanation or repair. Follow-up rounds have no acceptance-checker: the primary re-runs, inspects, and publishes itself. Do not delegate execution or require per-case approval.

Read acceptance-checker.md for the input/output contract, review boundaries, and follow-up rules. Acceptance review supplements the primary's own checks and any configured verifier; it does not replace either. If delegation or required media inspection is unavailable, disclose the missing review and unverified claims rather than claiming independent acceptance.

Read the project layer first

Before touching an environment, check for .agents/acceptance/:

FileWhat it owns
PROJECT.mdStart/stop commands, ports, services, auth, surfaces, probes
PROCESS.mdThe run process: plan gate, execution rules, teardown
common-mistakes.mdProject living log — what earlier rounds got wrong here
probe-mock-patterns.mdProject living log — how to force state on this product

The project layer owns how this repository is run; this skill owns what a valid round is (plan, evidence, report, immutable round, the hard rule). On running, the project layer wins; on what may be published, this skill wins. Never invent a start command, port, or auth flow PROJECT.md answers; fix a divergence in the adapter during the run instead of working around it. No .agents/acceptance/ → bootstrap one first: project-adapter.md.

Living logs — inject each by its own shape

Both layers (this skill's generic copies and the project's own) are loaded once the target is known, silently:

  • common-mistakes.md — read its Checklist in full, now and again before marking any case pass. Pull an entry by id only when a checklist line applies to a case.
  • probe-mock-patterns.md — read the heading index, then pull only the entries this round needs. Pick by meaning, not keyword; rg over the body is the fallback.
bash
rg -n '^#{2,4} ' <file>          # the index, with line numbers
sed -n '<start>,<end>p' <file>   # one entry, in full

Record new project-specific learnings in the project layer only.

Two paths — no id is required

Every evidence command targets a round. The authored path is the default; you have an operation id only when the invocation names one. Never hunt the environment for one, and never report this skill inapplicable — a round without an operation is simply recorded as standalone.

You havePath
No plan — you author the checksWrite result.json + assets/, publish with lh acceptance run ingest — report.md
An operation id you were givenlh verify plan state, then result submit --operation per criterion — plan-format.md

Pass --subject (task:<id> / topic:<id> / document:<id>) only when the caller named one; otherwise ingest attaches the round itself when it can and creates a standalone acceptance when it cannot. On the first ingest, always supply --requirement "<one-sentence business goal>" — the durable goal of the whole acceptance, not this round's scope; it is immutable once recorded.

Prerequisites: lh is authed (lh acceptance run list --json returns [] or data; an auth error means stop and surface it), and only the UI driver the selected surface needs is installed — probe before adding dependencies, and never substitute a private agent plugin.

Optional user-journey flows

Before authoring checks, identify the independently reviewable user tasks in the requirement. Use those tasks as business groups, not the PR title or test surface. For example, reassignment, scheduled continuation, and failure recovery can be separate groups when the delivery covers all three; do not impose these groups on unrelated work. Each check should have an outcome the user can accept or reject independently. Keep shared entry/accessibility checks separate and avoid repeating their expectations across business checks.

When acceptance depends on a sequence of user states, publish its graph during planning, before implementation or verification begins. Keep the checklist paths above for independent checks; a graph is optional and does not replace evidence or human review.

For flow-based plans, each flow's title is its checks' default checklist category. Publish independent user journeys as separate flows in the same acceptance/run; use subflows for actual composed journeys. An umbrella flow containing checks for several independent tasks collapses them into one checklist group. Edges must describe real user transitions, not artificial links added to make unrelated checks reachable. Start at the user entry and follow the journey through outcomes and recovery; UUIDs identify nodes and must not encode business order. Read back the published plan and inspect its groups and reading order before execution.

For an existing acceptance that only needs different checklist groups, use lh acceptance regroup <acceptanceId> --file groups.json. Read the acceptance bundle first; write { expectedVersion, groups: [{ title, checkItemIds }] }, using the exact union checks[].id values and acceptance.metadata.checkGrouping.version (0 when absent). The groups replace the current presentation grouping; an empty list restores plan categories. Unassigned checks keep their plan category. This preserves check IDs, numbering, evidence and review history without creating a round. It does not change flow transitions or verification conditions. Do not move execution nodes or start a new round just to reorganize the checklist; those operations have different execution semantics.

  1. Use the named acceptance, or create one before publishing the flow. If none was named, first run lh acceptance create --help and confirm it shows Usage: lh acceptance create [options] and --requirement. Parent-command help or a zero exit code alone does not prove support. If unavailable, upgrade @lobehub/cli to a release that supports this command and check again; updating the skill alone does not upgrade the CLI. If still unavailable, report flow-first creation as blocked. Do not invent a subject ID, upload an empty report, or substitute lh acceptance run create (which creates a round).

    bash
    lh acceptance create --title "Checkout recovery" \
      --requirement "Customers can recover from a declined payment and complete checkout" --json

    --requirement is a required, nonblank durable business goal; --title is optional. Omit --subject for a fresh standalone subject, even when an ambient topic exists. Pass --subject task:<id>, topic:<id>, document:<id>, or standalone:<id> only for an explicitly supplied subject. Reusing a subject preserves its recorded requirement, title, and state; it does not reopen it. Creation does not create a verification round, report, results, or passing verdict.

    The JSON contains acceptanceId, acceptanceUrl, requirement, status, and subject: { subjectType, subjectId }. Use acceptanceId in all flow commands below, not subject.subjectId or a verification run ID. Share acceptanceUrl verbatim; it already uses the CLI's configured server.

    Write a JSON file with definition: { title, entryNodeId, nodes, edges }. Give nodes and edges stable UUIDs. Each node has id and exactly one of criterionId (existing check asset), check: { id, title, definition } (a check asset with steps, fixtures, preconditions and expected outcome), or subFlowId (another flow in this acceptance). Edges have id, sourceNodeId, targetNodeId, trigger, required, and optional condition. Every node must be reachable from the entry. Publish child flows before referencing them.

  2. lh acceptance flow publish <acceptanceId> --file flow.json saves the definition and returns flowId. To edit it, include that flowId and the current expectedHash in the file. lh acceptance flow view <acceptanceId> reads definitions, snapshots and results. Publishing does not execute checks. Revise a graph in place rather than publishing a second one; a superseded graph left behind still renders as its own journey with its own unexecuted checks. lh acceptance flow delete <acceptanceId> --flow <flowId> removes one that never should have existed, and only while it has no verified history: it is refused once a settled round has run it, or while another flow invokes it as a subflow.

  3. lh acceptance flow plan <acceptanceId> --flow <flowId> creates a draft round with the graph and its plan. While the round is only planned it follows the live graph: publishing an edit refreshes its snapshot and plan in place, and running flow plan again refreshes the same draft instead of opening another round. Add --run <verifyRunId> to attach another flow to the same draft. Read lh acceptance run get <verifyRunId> --json for the actual plan IDs: each branch and subflow invocation has its own checkItemId; never substitute the reusable asset ID.

  4. Share the acceptance link so the user can inspect the proposed nodes, branches and expected outcomes before implementation. Read and address any actionable feedback. Preparing a plan neither executes checks nor approves delivery; there is no separate flow-confirmation action. Continue within the user's authorized scope, or pause if the user explicitly asked to review before work. For requested changes, publish the revised definition with its flowId and expectedHash; the draft round follows automatically. Never open another round or another flow just to revise a plan that has not executed.

  5. Implement the work and exercise the real product, then use lh acceptance flow record <acceptanceId> --file result.json, containing verifyRunId, checkItemId, verdict (passed, failed, uncertain, or blocked) and observation. Record only what was observed. Use the returned result ID to attach required artifacts through lh acceptance run evidence (inspect its --help), following the same evidence rules as checklist checks.

  6. After all required checks are recorded and passed, run lh acceptance flow complete <acceptanceId> --run <verifyRunId>. Completion settles verification; it does not accept the delivery on the user's behalf. Read back the round and verify evidence coverage before handing it over.

To rerun the exact old graph, prepare a plan with --from-run <sourceVerifyRunId> and omit --run for a fresh round. This preserves the old definition and starts without results. Each replay starts as an unexecuted draft. A round is frozen by its first recorded result; only then does it keep its number. An lh acceptance run ingest that reaches an acceptance whose latest round is still a draft folds into that draft rather than opening a new round. Accepted or closed acceptances must be explicitly reopened before starting. Edges describe business transitions; they do not automatically schedule execution. Continue to read lh acceptance feedback <acceptanceId> --actionable before repairs and publish new rounds into the same acceptance.

HARD RULE — programmatic gates are NEVER acceptance checks

Every check MUST be an outcome a person decides about the delivery: what the user sees, hears, reads, or receives. These MUST NOT appear as a check, under any phrasing: unit / integration / regression / snapshot tests, coverage, type-check / tsc, lint / eslint, format, "compiles", "build passes", "CI is green". Run them, then report them as one line of narrative.

Enforced at ingest: every matching item (matched on title, category, AND method — "run bun run test" under a product-sounding title still matches) is dropped with a warning and summary recounted; a round of only such checks fails to publish. The line is the subject of the check, not who judged it: a CLI behavior asserted by a command is a fine check (verifier: "program"); "the suite is green" is not. Before writing any plan, ask of each draft check: would the user click accept/reject on this?

Rounds are immutable — repair means a NEW round

A published round is a permanent record. Never re-submit into a round after changing the code — publish the re-verification as the next round and let the acceptance page show the progression.

Before a repair round, read the aggregate with lh acceptance view <acceptanceId | type:id> --json. Omit checks whose latest userReview.action is accept; address non-stale rejects under their exact stable ids; when a check semantically replaces another, declare supersedes: ['old-id'] and repeat the full lineage in every later round that reuses the successor id. Pass --acceptance <acceptanceId> so the round joins the same history.

Show full SKILL.md (1,755 more words)Show less

Rules you will be tempted to skip

Not judgment calls — the moves an agent under pressure makes and must not. Each excuse below was made in a real round.

ExcuseReality
"Injection is hard; happy-path plus unit tests covers it"The error state was the goal. Walk the probe ladder (probe-mock-patterns.md A) before calling it blocked. (M2)
"The branch name says what to verify" / "Loading the living logs first…"The task lives in the user's words. Recover it, or confirm a labeled guess with one structured question — silently; never narrate setup. (M3, M21)
"The black frame is probably display sleep / a permission"Measure first: pixel brightness, the permission bit, an A/B with one variable toggled. Publish "confirmed by X" or "suspected", never a guess. (M4)
"Let me ask how they want it run" / "I'll click Sign in and you authorize" / "too small to screenshot"Environment mechanics are yours: full isolated run, auth by direct injection (never an interactive login — it hijacks the user's browser), a screenshot for every user-facing change. Ask only about the product decision. (M8)
"One more config edit and the env will boot" / "I'll mock it" / "I'll drive the rest myself"Timebox. Inventory running instances, probe for the real capability before mocking (a mock that records nothing is not in the path), re-delegate a dead subagent's remaining steps, revert experiments and ask. (M17)
"The fix is in and tests pass — verified"Reproduce the failure's precondition first, then verify with it held. A run that cannot fail proves nothing; "reproduces sometimes" means an unnamed precondition. When the mocked seam is the suspect, drop the mock. (M31)

Pick the surface by the user-visible outcome

Match the requirement to the cheapest surface that can prove the complete outcome, not merely the layer containing the code change. A backend fix for missing cards, stale lists, navigation, or another visible behavior still requires the consuming UI, its actual data response, and inspected screenshots. Database assertions and passing tests support that evidence; they do not replace it.

What your task changedSurfaceGuide
Backend / CLI / library / data logic with no UI outcomeCLI — stdout as text, zero UI flakinesssurfaces/cli.md
Web app frontend / styles / interactionsWeb (agent-browser → running web app)surfaces/web.md
New/changed API plus the UI consuming itWeb, full-stack (agent-browser + network capture)surfaces/web.md
Desktop-only behavior (native windows, IPC, packaged shell)Electron (agent-browser --cdp)surfaces/electron.md
Native macOS app / OS chrome agent-browser can't reachNative (osascript + screencapture, local macOS)surfaces/native.md
Native iOS behavior, gestures, device-size layoutiOS Simulator (sim-use/AXe + simctl)surfaces/ios-simulator.md
  • Use CLI alone only when the required outcome has no UI surface. If a visible outcome cannot be exercised, report that acceptance as incomplete instead of narrowing it to data checks. Use Electron only when the criterion depends on desktop-only code; iOS is driven by a Simulator HID/AX CLI, never host mouse — mark the case blocked if the CLI cannot express the gesture.
  • Structured data uses native visualizations (cases[].datasets + cases[].visualizations; raw CSV/JSON stays as evidence), not a PNG — report.md. A deliverable the user hears needs audio — evidence.md.
  • Auth is a gate scoped to the surface: authenticate that surface first or every capture lands on the sign-in page. Web: auth-web.md.
  • A UI round may price its interaction cost by recording KLM operator counts into interaction-trace.jsonl; optional, never hand-written — interaction-cost.md.

Every file submission MUST include a non-empty, reviewer-facing description (--desc for CLI submissions; description for tools and ingest entries). Identify what the file contains and what it demonstrates for this criterion. A filename, path, artifact id, or generic label such as "evidence" is not a sufficient description. This also applies when a text file is stored inline.

Shared rules for every artifact — media types, provenance, file vs inline, safety — are in evidence.md.

Keep checklist explanations brief

Write each check's observation and inline explanation in the user's language, usually 1–3 short sentences: what was done, what happened, and any limitation needed to judge that outcome. Do not paste the execution report into the check. Omit repeated titles, verdict labels, SHA/port/ID headers, environment boilerplate, and round-history explanations. Put shared setup and revision details once in the round report; keep commands, traces, raw output, and detailed reasoning in separate evidence attachments. Briefly disclose a limitation in the check when it changes the verdict; concision must not hide missing verification.

Example: “转派后,新 Agent 收到原对话上下文并创建了独立话题。刷新后消息仍保留。” For a failure, name the unmet outcome directly, without recounting the debugging process. Keep required evidence complete; shorten its presentation, not the work.

Final handoff (mandatory)

Cloud browser links use https://lobehub.com. For all acceptance, round, cleanup, and upgrade URLs in this skill (including instructions below that say "verbatim"), normalize LobeHub Cloud origins to https://lobehub.com, preserving the path, query, and fragment. Cloud hosts are lobehub.com and its subdomains. Keep self-hosted and development origins unchanged. This changes display links, not the CLI's configured API server. Keep the Skill installation resource at https://app.lobehub.com/acceptance/skill.md.

Close every browser session this run opened (agent-browser --session <name> close, web teardown) before handing off; a session left open keeps a full browser running indefinitely. Stop this run's resource guard (resource-guard.sh stop --state-dir <run state dir>) as well, and state in the round report whether it reached yellow or red and what that stopped; a run that hit red must say which checks it left blocked instead of passing.

Before declaring the task done, prove coverage: for each check with requiredEvidence, every declared type is present at least once. Report it explicitly; a missing type holds the delivery at uncertain no matter how good the work is.

Storage limits require a user-facing recovery handoff. For report ingest, atomic evidence upload, or result submission with a file, recognize recovery.reason: "storage_quota", failedEvidence[].reason: "storage_quota", or a storage_block: error. Do not stop at "upload failed" or "noted in the PR":

  • In the final response, state that storage limits blocked publication, distinguish locally observed results from uploaded evidence, and report the actual coverage. Include the saved acceptance/round links when available; do not invent them for an atomic submission that failed before saving a result.
  • Give both recovery options, in the user's language, using available recovery.cleanupUrl and recovery.upgradeUrl verbatim and following recovery.message, applying the Cloud browser-link rule above. Never delete user data automatically. Deletion is permanent.
    • Personal scope: clean up unneeded acceptances or upgrade the personal plan. Acceptance cleanup requires selecting "permanently delete all rounds, reports, and evidence files"; deleting only a record or evidence association does not free storage.
    • Workspace scope (recovery.scope: "workspace"): clean up that workspace's files or upgrade that workspace's plan. The cleanup link opens its resource library, not an acceptance list; do not invent an acceptance-purge checkbox there. Ask its owner/admin for cleanup or billing access. Personal cleanup or a personal upgrade does not resolve a workspace limit.
    • If the CLI reports unresolved workspace scope and omits recovery URLs, report that limitation and its scope-check instructions. Do not invent links or substitute personal pages.
  • For an older CLI without recovery metadata, resolve server and scope using lh doctor --offline --json and lh workspace current --json. Personal scope uses /acceptance and /settings/plans. For workspace scope, resolve its slug with lh workspace view --json, verify the returned ID matches the active workspace, and use /:workspaceSlug/resource and /:workspaceSlug/settings/plans; there is no /:workspaceSlug/acceptance route. If lookup fails, give scope-specific guidance without guessed links. Strip URL username/password when constructing display links. For LobeHub Cloud, personal cleanup uses https://lobehub.com/acceptance; personal plan upgrades use https://lobehub.com/settings/plans. Workspace resource and plan paths use https://lobehub.com. Keep self-hosted users on their configured server.
  • Preserve local reports, artifacts, and the returned retry instructions. Stop blind retries until the user has addressed storage. For a partially ingested report, retry only failed artifacts using failedEvidence[].retryArgs or retryCommand, not the whole ingest. For an atomic upload/submission that saved nothing, retry that command. Supplementing evidence does not change recorded verdicts; read back coverage and do not claim the delivery is complete while required evidence is missing.

The final response for a completed handoff MUST include the published acceptance URL together with the coverage result — never only a check-result id or a prose claim. Obtain the links from the path you actually executed:

  • Authored round: copy acceptanceUrl returned by lh acceptance run ingest --json verbatim.
  • Operation-plan round: follow the read-only plan handoff lookup. It resolves the supplied operation ID to its existing run, acceptance, and round using the CLI's actual server configuration. Copy its acceptanceUrl output. Do not run authored ingest, create another acceptance, or resubmit evidence merely to obtain a link.

Never guess a host, acceptance ID, or round index. The documented plan lookup is the only reconstruction needed for CLIs whose submission output contains only an internal run URL. If the run has no acceptance association or the lookup fails, report the handoff as blocked and preserve the submitted evidence; do not declare delivery complete or fabricate a link. Put no images, local paths, local file links, or internal run-page paths in the chat reply.

Write the link as a plain-text line, never inside a fenced or inline code block — the chat client only linkifies plain text, and a code block makes it unclickable. Hand off only the acceptance URL: the acceptance page opens on its latest round, so a separate per-round link adds nothing for the reader. Replace the placeholder below with the URL from the selected path:

Acceptance: <acceptanceUrl, verbatim> Coverage: 2/2 criteria, all required evidence uploaded

Portability rules

  • Engine-level capture over OS capture. agent-browser screenshot / dom / eval run headless; screencapture / osascript are macOS-only. iOS: xcrun simctl io over host-window capture. Rounds land under .acceptances/, which the CLI keeps out of git.
  • Upload as you go. Evidence keyed to its check mid-run survives a crash near the end.
  • Don't invent evidence. Capture only the types a check declares.
  • Chapter every recording. Log a mark for each step and each claim you verify on a frame, and disclose every anomaly you noticed as a flag; the reviewer seeks to them instead of watching the whole clip — video-chapters.md.

Reference map

For both acceptance-checker handoffs and review output, read acceptance-checker.md.

NeedReference
Bounding a run's memory useresource-guard.md
The project layer, bootstrapping an adapterproject-adapter.md
Mistakes checklist (read every round)common-mistakes.md
Forcing state, error injection, runtime probesprobe-mock-patterns.md
Authored rounds, result.json, ingestreport.md
Plan-driven rounds: schema, submit, coverageplan-format.md
Evidence media, provenance, submission, safetyevidence.md
Interaction cost overlayinteraction-cost.md
Web/Electron Chromium CLI commandsagent-browser.md
iOS Simulator driver CLI commandssim-use.md (preferred), axe.md (fallback)
Bundled CDP screenshot and macOS capture preflightscreenshot-helpers.md
Authenticated Web sessionauth-web.md
Native macOS / OS-owned stepcomputer-use.md
Video chapters: steps, checks, flags on a clipvideo-chapters.md
Temporal evidence: Web/Electron, iOS, nativerecording-cdp.md, recording-ios-simulator.md, recording-native-macos.md

© lobehub, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 33 other files (scripts, references) in .agents/skills/acceptance of lobehub/lobehub.

  • SKILL.md
  • LICENSE
  • agents/openai.yaml
  • references/acceptance-checker.md
  • references/agent-browser.md
  • references/auth-web.md
  • references/axe.md
  • references/common-mistakes.md
  • references/computer-use.md
  • references/evidence.md
  • references/interaction-cost.md
  • references/plan-format.md
  • references/probe-mock-patterns.md
  • references/project-adapter.md
  • references/recording-cdp.md
  • references/recording-ios-simulator.md
  • references/recording-native-macos.md
  • references/report.md
  • references/resource-guard.md
  • … and 15 more

Open the folder on GitHubat commit 1863542

Compare with similar skills

Acceptance Evidence for Deliveries next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Acceptance Evidence for Deliveries compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Acceptance Evidence for Deliveries this skilllobehub/lobehub83k—~9.7kAutomated safety check: PassApache-2.0
E2E Smoke TestAppsFlyerSDK/appsflyer-unity-plugin178—~325Automated safety check: PassMIT
Senpi Agent QA Harnesscode-yeongyu/senpi470—~2.7kAutomated safety check: NotesMIT
tmux Real User TestingQwenLM/qwen-code28k—~2.3kAutomated safety check: PassApache-2.0
Create a Verification Skillcursor/plugins10k8 repos~1.5kAutomated safety check: PassNone
Heavy Verify Loop for Peri TUIKonghaYao/peri223—~2.1kAutomated safety check: NotesApache-2.0

Similar skills

  • E2E Smoke Test

    AppsFlyerSDK/appsflyer-unity-plugin

    Run or review a basic end-to-end smoke test for the AppsFlyer Unity plugin on Android emulator or iOS simulator, covering startup, initialization, and basic event flow.

    178 GitHub stars~325 tokensUpdated 3 days ago
    Testing & QAAuto-check passed
  • Senpi Agent QA Harness

    code-yeongyu/senpi

    Checks changes to the senpi coding agent by driving the real CLI from source in an isolated sandbox, over RPC, terminal UI, mock model and CLI smoke channels.

    470 GitHub stars~2.7k tokensUpdated today
    Testing & QAAuto-check: notes
  • tmux Real User Testing

    QwenLM/qwen-code

    Drives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review.

    28k GitHub stars~2.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Official

    Generates a project-local skill that launches your app, exercises a feature the way a user would and captures evidence, for web, CLI, API or desktop projects.

    10k GitHub starsUsed in 8 repos~1.5k tokens
    Testing & QAAuto-check passed
  • Verifies and repairs a feature by using the real Peri terminal UI as a user would, looping verify, decide, fix and review until a fresh round shows no blockers.

    223 GitHub stars~2.1k tokensUpdated today
    Testing & QAAuto-check: notes
  • Installs Reticle's dev-only SDK in a running web app and verifies user-facing changes by driving a real flow, returning a verdict with the file and line to fix.

    1.2k GitHub stars~2.7k tokensUpdated today
    Testing & QAAuto-check: notes

More from lobehub/lobehub

All 50 skills in this repo
  • Builds single-file interactive HTML prototypes rendered with the real LobeHub UI components and written as production-style React, so they can later be split into files.

    83k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Git Worktree Cleanup

    lobehub/lobehub

    Audits stale Git worktrees and branches with a bundled script, classifies each one, and deletes only after you approve the exact candidates.

    83k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Maintains LobeHub's model-backed alint rule set: writing rules, removing false positives against real code, deciding warn versus error and tracking token cost.

    83k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Guides building LobeHub builtin agent tools, from the manifest and execution runtime to executors, chat UI renders and registry wiring.

    83k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Explains how LobeHub client code fetches data through services, SWR store hooks and cache keys, and when to avoid useEffect fetching or duplicated state.

    83k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Guides work on LobeHub's own product search: the shared search repository, provider choice, Elasticsearch mappings, change syncing and reindexing.

    83k GitHub stars~4.1k tokensUpdated today
    Auto-check passed

Works with

Questions about Acceptance Evidence for Deliveries

What does Acceptance Evidence for Deliveries do?

Verifies a delivery end to end by driving the real product on a CLI, web, desktop or iOS Simulator surface, capturing evidence and publishing a round with the lh CLI. The agent acts as the builder of a delivery whose claims are judged by a separate review step against a plan, either one handed to the run or checks it writes itself. Any check that declares required evidence cannot pass on the agent's words alone: a missing artifact marks it uncertain and holds the delivery.

When should I use Acceptance Evidence for Deliveries?

Acceptance Evidence for Deliveries fits situations like: proving that a finished change works before handing it over; capturing screenshots or recordings as evidence for a verify plan; testing a desktop or Electron build end to end; publishing a test report round with the lh CLI.

How do I install Acceptance Evidence for Deliveries in Claude Code?

Run `npx skills add lobehub/lobehub --skill acceptance -a claude-code`. Or copy the skill folder (.agents/skills/acceptance in lobehub/lobehub) into .claude/skills/acceptance in your project. Claude Code loads it when a task matches its description.

How do I install Acceptance Evidence for Deliveries in Codex?

Run `npx skills add lobehub/lobehub --skill acceptance -a codex`. Or copy the skill folder (.agents/skills/acceptance in lobehub/lobehub) into .agents/skills/acceptance in your project. Codex loads it when a task matches its description.

Can I use Acceptance Evidence for Deliveries in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add lobehub/lobehub --skill acceptance -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/acceptance, .gemini/skills/acceptance, .github/skills/acceptance and .opencode/skills/acceptance in your project.

What does Acceptance Evidence for Deliveries need to run?

Going by SKILL.md and its folder, Acceptance Evidence for Deliveries needs the command-line tools its instructions call (rg, git and bun). Our summary lists: The lh CLI for publishing rounds; A runnable product on the chosen surface: CLI, web, desktop or iOS Simulator.

Does Acceptance Evidence for Deliveries access the network?

SKILL.md names 1 domain. In commands or code: lobehub.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Acceptance Evidence for Deliveries safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Acceptance Evidence for Deliveries use?

Acceptance Evidence for Deliveries is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Acceptance Evidence for Deliveries use?

About 9.7k tokens (SKILL.md is roughly 39k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 36k tokens, read only when the agent opens those files.

What are the alternatives to Acceptance Evidence for Deliveries?

Skills that share tags, products or a category with Acceptance Evidence for Deliveries: E2E Smoke Test (AppsFlyerSDK/appsflyer-unity-plugin, 178 stars), Senpi Agent QA Harness (code-yeongyu/senpi, 470 stars), tmux Real User Testing (QwenLM/qwen-code, 28k stars) and Create a Verification Skill (cursor/plugins, 10k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Acceptance Evidence for Deliveries?

lobehub (a GitHub organization) maintains it in lobehub/lobehub, which has 83,023 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on October 7, 2026.

Source: lobehub/lobehub on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.