Agent skill

Challenge Review

by werf in werf/werf

Independent challenge pass for a non-trivial or high-risk change.

Apache-2.0Auto-check passedDevOps & Cloud

Install Challenge Review

skills CLI
$ npx skills add werf/werf --skill challenge-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install werf/werf challenge-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/werf/werf.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/challenge-review .claude/skills/challenge-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
challenge-review
GitHub stars
4.7k
Token cost
~1.1k tokens
SKILL.md length
615 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Independent challenge pass for a non-trivial or high-risk change.

  • A fresh second opinion
  • SKILL.md covers Recover the contract first, Test the tests, Check-gaming and Inspection depth, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • A change touches tests

What it does

Challenge Review is an agent skill from werf/werf. Independent challenge pass for a non-trivial or high-risk change. Use for a fresh second opinion, when a change touches tests or verification infrastructure, or when invoked as /challenge-review. Attempts to disprove implementation claims through contract recovery, test falsification, and check-gaming detection; complements ordinary review.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering CI/CD. It works with Docker. The repository describes itself as: A solution for implementing efficient and consistent software delivery to Kubernetes facilitating best practices. The licence is Apache-2.0.

When your agent uses it

  • A fresh second opinion
  • A change touches tests
  • Verification infrastructure
  • Invoked as /challenge-review

Example prompts

  • “/challenge-review”

What it can do on your machine

Read from SKILL.md and the folder at commit 186a839. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Challenge Review loads about 1.1k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 615 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from werf/werf at commit 186a839, republished under its Apache-2.0 licence (© werf). 615 words, ~1,116 tokens.

Download SKILL.mdSave it as .claude/skills/challenge-review/SKILL.md (or your agent's skills folder).
name
challenge-review
description
Independent challenge pass for a non-trivial or high-risk change. Use for a fresh second opinion, when a change touches tests or verification infrastructure, or when invoked as /challenge-review. Attempts to disprove implementation claims through contract recovery, test falsification, and check-gaming detection; complements ordinary review.

Challenge Review

Treat the change as an untrusted implementation candidate. Readable code, passing tests, high coverage, and the author's confidence are not evidence of correctness. Coverage percentage is not evidence by itself.

Do not ask only:

Does this code look correct?

Also ask:

What evidence would fail if this implementation were wrong?

Recover the contract first

Identify intended behavior, acceptance criteria, and behavior that must remain unchanged. If the intended behavior is unclear, that is the first finding.

Do not invent a stronger contract than the task provides. A finding against a requirement the task never stated is noise.

Test the tests

Invoke the test-the-tests skill for this step: verify each test that carries weight by actually mutating the implementation (invert a condition, remove validation, suppress an error, skip a side effect, revert to the prior behavior) and confirming the test fails. A suite that cannot detect a plausible fault is weak even when it passes — passing and high coverage are not evidence by themselves.

Check-gaming

Flag changes that:

  • weaken or delete assertions;
  • update golden files without explaining the behavioral change;
  • skip, quarantine, or filter tests;
  • lower quality thresholds;
  • add test-only branches or detect CI/test environments;
  • hardcode fixture-specific answers;
  • rely entirely on mocks for critical behavior;
  • modify verification scripts together with the implementation;
  • present logs or reports without proving they belong to the reviewed commit.

Unexplained changes to tests or verification infrastructure are high risk by default.

Inspection depth

Direct inspection of the implementation is mandatory regardless of green checks when the change touches auth, secrets, crypto, billing, data deletion, migrations, concurrency, release or supply-chain logic, public APIs, persistent formats, or anything hard to roll back.

For low-risk mechanical changes with strong evidence, targeted inspection is enough. No line-by-line narration in either case.

Show full SKILL.md (322 more words)Show less

Independence matters

If you wrote the diff you are now reviewing, this pass is necessary but not sufficient. Self-review — even done adversarially, even by mutating your own tests — inherits your own design assumptions; it reliably catches localized bugs and weak tests, but is a poor substitute for a second opinion on whether the overall approach or architecture is sound. An independently-invoked reviewer with no memory of your rationale (a fresh subagent, or an external tool/model such as Codex) will more reliably surface issues you can't see because you already believe your own premises.

For anything more than a mechanical or low-risk change, escalate to an independently invoked reviewer before merging — don't treat your own adversarial pass as the final word. Say so explicitly in your findings when you are the diff's author and no second reviewer has looked yet, so the person deciding whether to merge knows that gap exists.

When that review comes back and you disagree, separate the finding from the fix it proposes: a correct finding often ships with a remedy that does not work, and verifying the mechanism tells you which half to reject. Say which one you are rejecting and show the evidence.

Two failure modes on your side. "It does not reduce the headline risk" is not sufficient grounds to decline a cheap fix that closes a silent-failure path — measure the cost, and if it is small, take it. And if a reviewer raises the same finding a second time, re-examine whether the fix is feasible instead of restating your position; the second request is evidence your explanation did not land, not that it needs repeating.

Output

Report only actionable findings. Each one: where the problem is visible, what fails and under which conditions, and the smallest concrete correction or missing proof. Style preferences are not defects.

Passing checks are claims. The review's job is to establish whether those checks are capable of disproving the implementation.

© werf, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/challenge-review of werf/werf.

Open the folder on GitHubat commit 186a839

Compare with similar skills

Challenge Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Challenge Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Challenge Review this skillwerf/werf4.7k—~1.1kAutomated safety check: PassApache-2.0
Senior DevOps Toolkitmaslennikov-ig/claude-code-orchestrator-kit2606 repos~1.1kAutomated safety check: NotesCustom licence
Megatron-LM Base Image BumpNVIDIA/Megatron-LM18k—~2.8kAutomated safety check: PassApache-2.0
GitHub Actions CreatorFNOSP/FlyNarwhal4961 repos~2.4kAutomated safety check: PassAGPL-3.0
Swig CI Reproswig/swig6.3k—~1.2kAutomated safety check: PassCustom licence
Megalinter Checknvuillam/npm-groovy-lint2481 repos~3.9kAutomated safety check: NotesMIT

Similar skills

  • Senior DevOps Toolkit

    maslennikov-ig/claude-code-orchestrator-kit

    Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup…

    260 GitHub starsUsed in 6 repos~1.1k tokens
    DevOps & CloudAuto-check: notes
  • Megatron-LM Base Image Bump

    NVIDIA/Megatron-LM

    Official

    Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up.

    18k GitHub stars~2.8k tokensUpdated today
    DevOps & CloudAuto-check passed
  • GitHub Actions Creator

    FNOSP/FlyNarwhal

    A skill your agent uses when the user wants to create, generate, or set up a GitHub Actions workflow.

    496 GitHub starsUsed in 1 repo~2.4k tokens
    DevOps & CloudAuto-check passed
  • Swig CI Repro

    swig/swig

    Reproduce a GitHub Actions Linux CI failure locally when it does not happen on your machine: a podman/docker image that mirrors the ubuntu-22.04 runner by reusing the real Tools/CI-linux-.sh install…

    6.3k GitHub stars~1.2k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Megalinter Check

    nvuillam/npm-groovy-lint

    Collect MegaLinter lint errors for the current repository. An agent skill from nvuillam/npm-groovy-lint.

    248 GitHub starsUsed in 1 repo~3.9k tokens
    DevOps & CloudAuto-check: notes
  • Maintains the DDNS project's GitHub Actions, Docker and Nuitka builds, packaging and release preparation without touching publishing credentials.

    4.7k GitHub stars~444 tokensUpdated yesterday
    DevOps & CloudAuto-check passed

More from werf/werf

  • werf conventions for branch names and commit messages. An agent skill from werf/werf.

    4.7k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Pull Request

    werf/werf

    Generates Pull Request titles and descriptions according to werf conventions.

    4.7k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Review

    werf/werf

    Code review of a pull request, branch, or diff. An agent skill from werf/werf.

    4.7k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • How to treat conclusions inherited from an earlier session — handover notes, prepared comments, verdict files, plans.

    4.7k GitHub stars~669 tokensUpdated today
    Auto-check passed
  • Session Retro

    werf/werf

    Analyze the current session for harness-worthy lessons — repeated corrections, discovered conventions, skill bugs — and turn them into concrete repo changes: docs, skills, task targets, linter…

    4.7k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Verify a test actually falsifies the behavior it claims to cover, via real mutation.

    4.7k GitHub stars~1k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Challenge Review

What does Challenge Review do?

Independent challenge pass for a non-trivial or high-risk change. Challenge Review is an agent skill from werf/werf. Independent challenge pass for a non-trivial or high-risk change.

When should I use Challenge Review?

Challenge Review fits situations like: A fresh second opinion; A change touches tests; verification infrastructure; invoked as /challenge-review.

How do I install Challenge Review in Claude Code?

Run `npx skills add werf/werf --skill challenge-review -a claude-code`. Or copy the skill folder (.agents/skills/challenge-review in werf/werf) into .claude/skills/challenge-review in your project. Claude Code loads it when a task matches its description.

How do I install Challenge Review in Codex?

Run `npx skills add werf/werf --skill challenge-review -a codex`. Or copy the skill folder (.agents/skills/challenge-review in werf/werf) into .agents/skills/challenge-review in your project. Codex loads it when a task matches its description.

Can I use Challenge Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add werf/werf --skill challenge-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/challenge-review, .gemini/skills/challenge-review, .github/skills/challenge-review and .opencode/skills/challenge-review in your project.

What does Challenge Review need to run?

SKILL.md names no scripts, command-line tools or credentials: Challenge Review is instructions for the agent only.

Does Challenge Review access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Challenge Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Challenge Review use?

Challenge Review is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Challenge Review use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Challenge Review?

Skills that share tags, products or a category with Challenge Review: Senior DevOps Toolkit (maslennikov-ig/claude-code-orchestrator-kit, 260 stars), Megatron-LM Base Image Bump (NVIDIA/Megatron-LM, 18k stars), GitHub Actions Creator (FNOSP/FlyNarwhal, 496 stars) and Swig CI Repro (swig/swig, 6.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Challenge Review?

werf (a GitHub organization) maintains it in werf/werf, which has 4,734 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 8, 2026.

Source: werf/werf on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.