Agent skill

Task Verification Checkpoints

by HKUDS in HKUDS/OpenSpace

Verify output format, file coverage, and task alignment before completing

MITAuto-check passedAgent Workflows

Install Task Verification Checkpoints

skills CLI
$ npx skills add HKUDS/OpenSpace --skill task-verification-checkpoints -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HKUDS/OpenSpace task-verification-checkpoints --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/benchmarks/gdpval/skills/task-verification-checkpoints .claude/skills/task-verification-checkpoints && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
task-verification-checkpoints
GitHub stars
7.7k
Token cost
~1.3k tokens
SKILL.md length
297 words
Files
2
Skills in repo
199
Repo updated
First seen
Licence
MIT

At a glance

Verify output format, file coverage, and task alignment before completing

  • Works in 4 steps: Re-read the task prompt for explicit… → Check file extensions of all output files → Verify the actual file type matches the… → …
  • Tasks that involve Verification before completion
  • SKILL.md covers Overview, Three Critical Checkpoints, Pre-Completion Checklist and When to Apply, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Task Verification Checkpoints is an agent skill from HKUDS/OpenSpace. Verify output format, file coverage, and task alignment before completing

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.

It sits in Agent Workflows, covering Verification before completion. The repository describes itself as: "OpenSpace: The Skill Management Layer for AI Agents" -- https://open-space.cloud/. The licence is MIT.

When your agent uses it

  • Tasks that involve Verification before completion

Example prompts

  • “/task-verification-checkpoints”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Re-read the task prompt for explicit format requirements
  2. Check file extensions of all output files
  3. Verify the actual file type matches the extension
  4. Confirm against any format specifications (e.g., "revised PDF report")

What it can do on your machine

Read from SKILL.md and the folder at commit 3827781. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Task Verification Checkpoints loads about 1.3k tokens when it runs. Until then it costs about 26 tokens; SKILL.md has 297 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~26
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HKUDS/OpenSpace at commit 3827781, republished under its MIT licence (© HKUDS). 297 words, ~1,256 tokens.

Download SKILL.mdSave it as .claude/skills/task-verification-checkpoints/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
task-verification-checkpoints
description
Verify output format, file coverage, and task alignment before completing

Task Verification Checkpoints

Before declaring a task as <COMPLETE>, perform these verification checks to ensure the deliverable matches requirements and prevents wrong-task completion.

Overview

This skill prevents common failures where agents:

  • Create outputs in wrong formats (PDF vs PPTX vs DOCX)
  • Skip processing some reference files
  • Drift from the actual task requirements

Three Critical Checkpoints

Checkpoint 1: Output Format Verification

Question: Does the output file format match task requirements?

Actions:

  1. Re-read the task prompt for explicit format requirements
  2. Check file extensions of all output files
  3. Verify the actual file type matches the extension
  4. Confirm against any format specifications (e.g., "revised PDF report")

Example validation:

python
def verify_output_format(task_prompt, output_files):
    """Check if output format matches task requirements."""
    required_format = extract_format_requirement(task_prompt)  # e.g., 'pdf', 'pptx'
    
    for f in output_files:
        actual_ext = f.split('.')[-1].lower()
        if actual_ext != required_format:
            return False, f"Expected .{required_format}, got .{actual_ext}"
    
    return True, "Format verified"
Checkpoint 2: Reference File Coverage

Question: Were ALL reference/input files processed?

Actions:

  1. List all files mentioned in the task as inputs
  2. Confirm each was read/analyzed/used (check file access logs)
  3. Document which content from each file contributed to output
  4. Flag any unprocessed reference files

Example validation:

python
def verify_reference_coverage(task_files, accessed_files):
    """Ensure all reference files were actually processed."""
    missing = set(task_files) - set(accessed_files)
    if missing:
        return False, f"Unprocessed files: {missing}"
    return True, "All references processed"
Checkpoint 3: Task Alignment Verification

Question: Does the deliverable address the ACTUAL task prompt?

Actions:

  1. Re-read the original task description verbatim
  2. Compare output purpose against stated task goal
  3. Check for scope creep or task drift
  4. Verify the output solves the stated problem (not a different one)

Example validation:

python
def verify_task_alignment(task_prompt, output_description):
    """Confirm output addresses the actual task."""
    task_verbs = extract_action_verbs(task_prompt)  # e.g., 'review', 'revise', 'create'
    output_verbs = extract_action_verbs(output_description)
    
    if not set(task_verbs).issubset(set(output_verbs)):
        return False, "Output actions don't match task requirements"
    
    return True, "Task alignment verified"

Pre-Completion Checklist

Execute this checklist before marking ANY task complete:

python
def pre_completion_verification(task_prompt, output_files, reference_files):
    """
    Complete verification before declaring task done.
    Returns (success, message) tuple.
    """
    checks = []
    
    # Check 1: Format
    format_ok, format_msg = verify_output_format(task_prompt, output_files)
    checks.append(('Format', format_ok, format_msg))
    
    # Check 2: Coverage
    coverage_ok, coverage_msg = verify_reference_coverage(reference_files, get_accessed_files())
    checks.append(('Coverage', coverage_ok, coverage_msg))
    
    # Check 3: Alignment
    alignment_ok, alignment_msg = verify_task_alignment(task_prompt, get_output_summary())
    checks.append(('Alignment', alignment_ok, alignment_msg))
    
    # Report results
    all_passed = all(ok for _, ok, _ in checks)
    
    if not all_passed:
        print("❌ VERIFICATION FAILED - Do not mark complete")
        for name, ok, msg in checks:
            status = "✓" if ok else "✗"
            print(f"  {status} {name}: {msg}")
        return False, "Verification failed"
    
    print("✓ All verification checkpoints passed")
    return True, "Ready for completion"

When to Apply

Use this skill whenever:

  • Creating deliverables from input/reference files
  • Task specifies output format requirements
  • Multiple files must be processed/analyzed
  • Task involves transformation, revision, or synthesis
  • Always before any <COMPLETE> declaration

Common Failure Modes to Catch

Failure TypeExamplePrevention
Format mismatchCreated PPTX when PDF requestedCheckpoint 1
Incomplete coverageSkipped Photographs.zipCheckpoint 2
Task driftCreated new report instead of revising existingCheckpoint 3
Assumption errorsAssumed format without verificationAll checkpoints

Quick Reference

Before <COMPLETE>:
  □ Output format matches requirements?
  □ All reference files processed?
  □ Deliverable addresses actual task?
  
If ANY box unchecked → Do NOT complete, fix first.

© HKUDS, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in benchmarks/gdpval/skills/task-verification-checkpoints of HKUDS/OpenSpace.

  • SKILL.md
  • .skill_id

Open the folder on GitHubat commit 3827781

Compare with similar skills

Task Verification Checkpoints next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Task Verification Checkpoints compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Task Verification Checkpoints this skillHKUDS/OpenSpace7.7k—~1.3kAutomated safety check: PassMIT
Show Me Your Work Decision Logcursor/plugins10k9 repos~1.6kAutomated safety check: PassNone
Verification Before Completionfarm-fe/farm5.6k45 repos~1kAutomated safety check: PassMIT
Scope Creep Guardlennney/stop-that-shit2.5k1 repos~2kAutomated safety check: PassMIT
PUA High-Agency Governancetanweai/pua20k—~502Automated safety check: PassMIT
Incremental Implementationaddyosmani/agent-skills102k1 repos~2.3kAutomated safety check: PassMIT

Similar skills

  • Official

    Keeps a TSV decision log for long or unattended agent runs, one row per decision with what, why, evidence and result, so a reviewer can check the work later.

    10k GitHub starsUsed in 9 repos~1.6k tokens
    Agent WorkflowsAuto-check passed
  • A skill your agent uses when about to claim work is complete, fixed, or passing, before committing or creating PRs - requires running verification commands and confirming output before making any…

    5.6k GitHub starsUsed in 45 repos~1k tokens
    Agent WorkflowsAuto-check passed
  • Scope Creep Guard

    lennney/stop-that-shit

    Keeps an agent focused on the requested work by applying a five-step ladder that checks for direct solutions, real gaps and speculative defenses before adding anything.

    2.5k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed
  • Pushes an agent to keep verifying and changing approach after repeated failures, using a diagnosis line, evidence-based completion and confirmation before risky edits.

    20k GitHub stars~502 tokensUpdated 28 days ago
    Agent WorkflowsAuto-check passed
  • Incremental Implementation

    addyosmani/agent-skills

    Delivers a change in thin vertical slices, each implemented, tested, verified and committed before the next, using vertical, contract-first or risk-first slicing.

    102k GitHub starsUsed in 1 repo~2.3k tokens
    Agent WorkflowsAuto-check passed
  • PUA Loop

    tanweai/pua

    Runs an unattended iterate-until-verified loop in which a user-set verify command, not the agent's own claim, decides when the task is finished.

    20k GitHub stars~1.1k tokensUpdated 28 days ago
    Agent WorkflowsAuto-check passed

More from HKUDS/OpenSpace

All 199 skills in this repo
  • Walks through producing a master audio track plus stems in Python, from checking a reference file and timing sections by BPM to effects, a zip archive and final verification.

    7.7k GitHub stars~2.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Handle cascading data retrieval tool failures by falling back to embedded knowledge generation

    7.7k GitHub stars~765 tokensUpdated 1 mo ago
    Auto-check passed
  • Gives an agent a workaround when its code-execution sandbox keeps failing: save the Python script to a file and run it through the shell instead.

    7.7k GitHub stars~588 tokensUpdated 1 mo ago
    Auto-check passed
  • A recovery routine for agents whose sandboxed code runner keeps failing: save the Python script to disk, then run it through the shell and read the output.

    7.7k GitHub stars~652 tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback ladder for failed sandboxed code runs, plus the habit of fixing the working directory first so generated files land in the right place.

    7.7k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback workflow for executing Python code when executecodesandbox fails repeatedly

    7.7k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Task Verification Checkpoints

What does Task Verification Checkpoints do?

Verify output format, file coverage, and task alignment before completing. Task Verification Checkpoints is an agent skill from HKUDS/OpenSpace.

When should I use Task Verification Checkpoints?

Task Verification Checkpoints fits situations like: tasks that involve Verification before completion.

How do I install Task Verification Checkpoints in Claude Code?

Run `npx skills add HKUDS/OpenSpace --skill task-verification-checkpoints -a claude-code`. Or copy the skill folder (benchmarks/gdpval/skills/task-verification-checkpoints in HKUDS/OpenSpace) into .claude/skills/task-verification-checkpoints in your project. Claude Code loads it when a task matches its description.

How do I install Task Verification Checkpoints in Codex?

Run `npx skills add HKUDS/OpenSpace --skill task-verification-checkpoints -a codex`. Or copy the skill folder (benchmarks/gdpval/skills/task-verification-checkpoints in HKUDS/OpenSpace) into .agents/skills/task-verification-checkpoints in your project. Codex loads it when a task matches its description.

Can I use Task Verification Checkpoints in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUDS/OpenSpace --skill task-verification-checkpoints -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/task-verification-checkpoints, .gemini/skills/task-verification-checkpoints, .github/skills/task-verification-checkpoints and .opencode/skills/task-verification-checkpoints in your project.

What does Task Verification Checkpoints need to run?

SKILL.md names no scripts, command-line tools or credentials: Task Verification Checkpoints is instructions for the agent only. Our summary lists: Python 3.

Does Task Verification Checkpoints access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Task Verification Checkpoints safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Task Verification Checkpoints use?

Task Verification Checkpoints is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Task Verification Checkpoints use?

About 1.3k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Task Verification Checkpoints?

Skills that share tags, products or a category with Task Verification Checkpoints: Show Me Your Work Decision Log (cursor/plugins, 10k stars), Verification Before Completion (farm-fe/farm, 5.6k stars), Scope Creep Guard (lennney/stop-that-shit, 2.5k stars) and PUA High-Agency Governance (tanweai/pua, 20k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Task Verification Checkpoints?

HKUDS (a GitHub organization) maintains it in HKUDS/OpenSpace, which has 7,743 GitHub stars. The repository holds 199 skills in this directory. The repository was last updated on August 12, 2026.

Source: HKUDS/OpenSpace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.