Agent skill

Claw Score

by openclaw in openclaw/openclaw

Audit or refresh OpenClaw maturity scorecard docs from root taxonomy, maturity scores, and QA evidence artifacts without using maintainer discrawl data or committed inventory reports.

MITAuto-check passed

Install Claw Score

skills CLI
$ npx skills add openclaw/openclaw --skill claw-score -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install openclaw/openclaw claw-score --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/openclaw/openclaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/claw-score .claude/skills/claw-score && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
claw-score
GitHub stars
392k
Token cost
~2.5k tokens
SKILL.md length
1,213 words
Files
52 (incl. references)
Skills in repo
93
Repo updated
First seen
Licence
MIT

At a glance

Audit or refresh OpenClaw maturity scorecard docs from root taxonomy, maturity scores, and QA evidence artifacts without using maintainer discrawl data or committed inventory reports.

  • Works in 7 steps: Read the surface in taxonomy.yaml. → Read the surface completeness rubric under → Gather public repo evidence from docs,… → …
  • SKILL.md covers Authority, Source Model, Commands and Full Generation Runs, plus 5 more sections
  • Calls pnpm, node and gh

What it does

Claw Score is an agent skill from openclaw/openclaw. Audit or refresh OpenClaw maturity scorecard docs from root taxonomy, maturity scores, and QA evidence artifacts without using maintainer discrawl data or committed inventory reports.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 53 other files, including reference files (for example `references/completeness/agent-runtime-and-provider-execution.md`, `references/completeness/android-app.md` and `references/completeness/anthropic-provider-path.md`).

The repository describes itself as: The AI that really does things. Any OS. Any Platform. The lobster way. 🦞. The licence is MIT.

Example prompts

  • “/claw-score”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Read the surface in taxonomy.yaml.
  2. Read the surface completeness rubric under
  3. Gather public repo evidence from docs, source, tests, and QA scenario
  4. Prefer existing release profile qa-evidence.json artifacts for executed
  5. Update qa/maturity-scores.yaml only for Quality, Completeness, and LTS
  6. Run the schema validation command from this skill.
  7. Run pnpm check:docs if docs prose changed, and focused QA coverage checks

What it can do on your machine

Read from SKILL.md and the folder at commit 1eb5970. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pnpm
    • node
    • gh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pnpm and gh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Claw Score loads about 2.5k tokens when it runs, and up to ~23k if it reads all its reference files. Until then it costs about 49 tokens; SKILL.md has 1,213 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~23k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from openclaw/openclaw at commit 1eb5970, republished under its MIT licence (© openclaw). 1,213 words, ~2,548 tokens.

Download SKILL.mdSave it as .claude/skills/claw-score/SKILL.md (or your agent's skills folder). This skill also uses 51 other files; get the full folder from GitHub.
name
claw-score
description
Audit or refresh OpenClaw maturity scorecard docs from root taxonomy, maturity scores, and QA evidence artifacts without using maintainer discrawl data or committed inventory reports.

claw-score

Use this skill when working on the OpenClaw maturity scorecard in this repo. This is the openclaw-local version of the maintainer claw-score workflow: it keeps the taxonomy and scorecard concepts, but excludes discrawl and the old committed inventory/ report tree.

Authority

This skill owns the operational workflow for:

  • taxonomy.yaml
  • qa/maturity-scores.yaml
  • docs/concepts/qa-e2e-automation.md
  • qa/scenarios/index.yaml

Keep person-specific, maintainer-private, Discord archive, and discrawl facts out of this repo. If a score needs private evidence, use the redacted qa-evidence.json artifact shape generated by OpenClaw QA workflows.

Source Model

  • taxonomy.yaml is the hand-edited source of truth for surfaces, levels, QA profiles, categories, feature coverage IDs, docs refs, LTS overrides, and completeness-instruction paths.
  • Each feature has exactly one coverageIds entry. Keep that evidence ID unique to the feature; broader many-to-many evidence mapping is not part of the current taxonomy schema.
  • Coverage IDs use dotted namespace.behavior form, with lowercase alphanumeric/dash segments. Profile, surface, and category IDs may remain dashed or dotted.
  • Keep categories and feature names unique, product-shaped, and broader than raw coverage IDs. Do not promote generic IDs into standalone feature names.
  • Avoid duplicate coverage-ID bundles under different feature names in one category.
  • qa/maturity-scores.yaml is the committed aggregate source for Quality, Completeness, and LTS review state.
  • extensions/qa-lab/src/scorecard-taxonomy.ts exports readValidatedQaMaturityScoreSources; use it to validate score output.
  • Generated public docs are docs/maturity/scorecard.md and docs/maturity/taxonomy.md; both come from pnpm maturity:render. Do not hand-edit generated Markdown to change score results.
  • qa-evidence.json artifacts provide per-run QA scorecard evidence. Release profile artifacts are the source of truth for Coverage. They can enrich generated artifact docs, but they are not committed as inventory.

Commands

Run from the openclaw repo root.

Validate taxonomy YAML structure and the maturity score schema after source edits:

bash
node --import tsx --input-type=module <<'NODE'
import fs from "node:fs";
import YAML from "yaml";
import { readValidatedQaMaturityScoreSources } from "./extensions/qa-lab/src/scorecard-taxonomy.ts";

for (const file of ["taxonomy.yaml", "qa/scenarios/index.yaml"]) {
  YAML.parse(fs.readFileSync(file, "utf8"));
}
readValidatedQaMaturityScoreSources();
NODE

Check docs when touching docs prose:

bash
pnpm check:docs

Run focused QA/profile checks when changing coverage IDs or profile membership:

bash
pnpm openclaw qa coverage --json

Full Generation Runs

For a direct full scorecard run that publishes the generated-doc pull request, use floating main resolution by default:

bash
gh workflow run maturity-scorecard.yml \
  --repo openclaw/openclaw \
  --ref main \
  -f ref=main \
  -f expected_sha='' \
  -f publish_pull_request=true \
  -f allow_failures=true

Do not resolve main locally and pass that commit as both ref and expected_sha for an ordinary manual generation run. OpenClaw's main moves quickly, so the caller-selected commit can become stale before validation. The workflow then correctly rejects publication when the pull request base contains newer maturity inputs, and QA never starts.

With ref=main and a blank expected_sha, the workflow's floating_default_branch path fetches and freezes the current remote default branch inside validation before handing an immutable revision to downstream jobs. Use an explicit SHA only when the requested evidence must remain bound to that exact revision, such as a release-candidate workflow call or an artifact-only historical reproduction. If that exact-revision run also requests publication and main has changed relevant inputs, expect validation to fail and dispatch again from floating main instead.

Scoring Workflow

When asked to score or refresh a surface:

  1. Read the surface in taxonomy.yaml.
  2. Read the surface completeness rubric under .agents/skills/claw-score/references/completeness/.
  3. Gather public repo evidence from docs, source, tests, and QA scenario metadata.
  4. Prefer existing release profile qa-evidence.json artifacts for executed proof.
  5. Update qa/maturity-scores.yaml only for Quality, Completeness, and LTS review state backed by public or redacted artifact evidence.
  6. Run the schema validation command from this skill.
  7. Run pnpm check:docs if docs prose changed, and focused QA coverage checks if coverage IDs or profile membership changed.

For subjective score changes, make the smallest defensible edit and leave the evidence path in the PR or task summary. Keep manual prose in current docs and keep score data in qa/maturity-scores.yaml.

Show full SKILL.md (639 more words)Show less

Default Completeness Process

Completeness is scored against the intended operator-visible workflow for each category, not against test breadth or implementation quality. The completeness reference files under references/completeness/ define the category scope and any surface-specific variation from this default process.

By default, Completeness measures how fully OpenClaw exposes the intended surface capability set to the user, operator, author, or maintainer persona for that surface. Score whether each category delivers the full expected workflow, including setup, normal use, status or inspection, recovery, and important platform, provider, channel, security, or lifecycle variants where they apply.

Treat Surface-Specific Scoring Questions and Surface-Specific Guidance as higher-priority instructions for that surface. The surface instructions may flesh out, narrow, or intentionally conflict with the default ideas here; when they do, follow the surface instructions and make the score rationale reflect that surface-specific instruction. If a reference file does not include surface-specific questions or guidance, apply this default process to the surface's Category Scope.

For each category, ask:

  • Can the intended user or operator complete the category workflow end to end?
  • Are the taxonomy features present as supported capabilities rather than isolated implementation fragments?
  • Are the important lifecycle stages represented: setup, normal operation, status/inspection, recovery, and upgrade or removal where relevant?
  • Are the important environment, provider, platform, channel, or security branches present for this surface?
  • Do the known gaps leave major user-visible capability branches missing?

Default guidance:

  • Favor higher Completeness when the category supports the full operator-visible workflow described by taxonomy and category evidence.
  • Lower Completeness when only the happy path exists, when important variants are undocumented or unimplemented, or when recovery/status paths are missing.
  • Do not lower Completeness because tests are thin; that is Coverage.
  • Do not lower Completeness because implementation quality is fragile; that is Quality.

Default Completeness bands:

  • Clawesome (95-100): complete across expected workflows, variants, and recovery branches, with only minor polish gaps.
  • Stable (80-95): the expected workflow set is broadly present, with only bounded missing branches.
  • Beta (70-80): the main workflow exists, but meaningful branches or recovery paths are still absent.
  • Alpha (50-70): only a partial capability set is present; users can complete some core tasks but not the full expected workflow.
  • Experimental (0-50): the category exposes only fragments of the intended capability.

Decision Context

Record an optional decision beside score and label for surface and category Quality/Completeness, or beside supported for category LTS. In taxonomy.yaml, use optional level_decision beside the canonical surface level.

Each record contains value, rationale, reviewer, evidence_refs, and revalidate_when. Use an integer from 0–100 for Quality/Completeness, a boolean for LTS, and a declared taxonomy level ID for level_decision. Supply nonempty text fields and at least one evidence reference. Name the actual reviewer and the condition that should trigger another review.

Leave unavailable history absent: it is unknown, not an invitation to invent reviewers, rationale, or evidence. A record does not overwrite the current score, support flag, or canonical level. If its value differs, retain both; generated docs show a non-gating mismatch, including under strict input validation.

Do not attach decisions to Coverage, computed rollups, surface LTS summaries, or the copied level in score aggregates. Decision context does not change coverage identity, score calculations, support commitments, or release gates.

Score Semantics

  • Coverage: deterministic release validation coverage derived from the release profile qa-evidence.json.scorecard feature fulfillment data.
  • Quality: reliability, maintainability, operator safety, and regression confidence for the category.
  • Completeness: how much of the intended operator-visible workflow exists for the category. Use the default completeness process plus any surface-specific variation before changing this score.
  • LTS: derived from Quality, release-evidence Coverage, and human_lts_override; do not hand-edit generated Markdown to change LTS status.

Bands:

  • Clawesome: 95-100
  • Stable: 80-95
  • Beta: 70-80
  • Alpha: 50-70
  • Experimental: 0-50

Artifacts

Do not add the maintainer repo's docs/kevinslin/maturity-scorecard/inventory/ tree to openclaw. Evidence-enriched scorecard outputs belong in short-lived artifacts, not committed generated docs, unless this repo adds an explicit renderer/check workflow first.

© openclaw, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 51 other files (references) in .agents/skills/claw-score of openclaw/openclaw.

  • SKILL.md
  • references/completeness/agent-runtime-and-provider-execution.md
  • references/completeness/android-app.md
  • references/completeness/anthropic-provider-path.md
  • references/completeness/automation-cron-hooks-tasks-polling.md
  • references/completeness/browser-automation-and-exec-sandbox-tools.md
  • references/completeness/browser-control-ui-and-webchat.md
  • references/completeness/channel-framework.md
  • references/completeness/clawhub-and-external-plugin-distribution.md
  • references/completeness/cli-install-update-onboard-doctor.md
  • references/completeness/discord.md
  • references/completeness/docker-podman-hosting.md
  • references/completeness/feishu-qq-bot-wechat-yuanbao-zalo-zalo-personal-regional-channels.md
  • references/completeness/gateway-runtime.md
  • references/completeness/google-chat.md
  • references/completeness/google-provider-path.md
  • references/completeness/image-video-music-generation-tools.md
  • references/completeness/imessage-bluebubbles.md
  • references/completeness/ios-app.md
  • … and 33 more

Open the folder on GitHubat commit 1eb5970

Compare with similar skills

Claw Score next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Claw Score compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Claw Score this skillopenclaw/openclaw392k—~2.5kAutomated safety check: PassMIT
Harness Scoreruvnet/ruflo74k—~605Automated safety check: NotesMIT
Index Refreshpaperclipai/paperclip99k—~994Automated safety check: PassMIT
Meta Refreshthedaviddias/Front-End-Checklist74k—~434Automated safety check: PassMIT
Score Evalsickn33/agentic-awesome-skills47k1 repos~304Automated safety check: PassMIT
UI Scoresickn33/agentic-awesome-skills47k1 repos~1.8kAutomated safety check: PassMIT

Similar skills

  • Harness Score

    ruvnet/ruflo

    5-dimension harness readiness scorecard from metaharness score <path.

    74k GitHub stars~605 tokensUpdated today
    DevelopmentAuto-check: notes
  • Index Refresh

    paperclipai/paperclip

    A skill your agent uses when an LLM Wiki operation issue requests an index refresh.

    99k GitHub stars~994 tokensUpdated today
    Knowledge ManagementAuto-check passed
  • Meta Refresh

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing rendered HTML, interactive components, or design-system patterns related to Avoid meta refresh redirects.

    74k GitHub stars~434 tokensUpdated 2 days ago
    Frontend & DesignAuto-check passed
  • Score Eval

    sickn33/agentic-awesome-skills

    Imported skill score-eval from upstream source. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 1 repo~304 tokens
    EducationAuto-check passed
  • UI Score

    sickn33/agentic-awesome-skills

    Score a UI file's design quality 0-100 against StyleSeed's design language — per-category breakdown, the worst offenders, and a prioritized fix list.

    47k GitHub starsUsed in 1 repo~1.8k tokens
    Testing & QAAuto-check passed
  • Code Maturity Assessor

    trailofbits/skills

    Official

    Scores a smart contract or blockchain codebase across 9 maturity categories with evidence, then delivers a scorecard and a priority-ordered improvement roadmap.

    7.4k GitHub stars~1.8k tokensUpdated today
    SecurityAuto-check passed

More from openclaw/openclaw

All 93 skills in this repo
  • Openclaw Live Updater

    openclaw/openclaw

    Maintain the canonical live OpenClaw main checkout, macOS LaunchAgent-managed Gateway, local macOS app, exact-head main CI, and recurring full release validation.

    392k GitHub stars~3.7k tokensUpdated today
    Auto-check passed
  • Tmux

    openclaw/openclaw

    Control tmux sessions/panes for interactive CLIs: list, capture output, send keys, paste text, monitor prompts.

    392k GitHub starsUsed in 2 repos~640 tokens
    Auto-check passed
  • Feishu Doc

    openclaw/openclaw

    Feishu document read/write workflows. An agent skill from openclaw/openclaw.

    392k GitHub stars~516 tokensUpdated today
    Auto-check passed
  • Openclaw PR Maintainer

    openclaw/openclaw

    Review, triage, repair, or land OpenClaw issues and pull requests with current-source evidence and the native maintainer workflow.

    392k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Browser Automation

    openclaw/openclaw

    A skill your agent uses when controlling web pages with the OpenClaw browser tool, especially multi-step flows, login checks, tab management, or recovery from stale refs/timeouts.

    392k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Clawsweeper

    openclaw/openclaw

    A skill your agent uses for all ClawSweeper work: OpenClaw issue/PR sweep reports, repair jobs, cloud fix PRs, @clawsweeper maintainer mention commands, trusted ClawSweeper-reviewed…

    392k GitHub stars~3k tokensUpdated today
    Auto-check passed

Questions about Claw Score

What does Claw Score do?

Audit or refresh OpenClaw maturity scorecard docs from root taxonomy, maturity scores, and QA evidence artifacts without using maintainer discrawl data or committed inventory reports. Claw Score is an agent skill from openclaw/openclaw. Audit or refresh OpenClaw maturity scorecard docs from root taxonomy, maturity scores, and QA evidence artifacts without using maintainer discrawl data or committed inventory reports.

How do I install Claw Score in Claude Code?

Run `npx skills add openclaw/openclaw --skill claw-score -a claude-code`. Or copy the skill folder (.agents/skills/claw-score in openclaw/openclaw) into .claude/skills/claw-score in your project. Claude Code loads it when a task matches its description.

How do I install Claw Score in Codex?

Run `npx skills add openclaw/openclaw --skill claw-score -a codex`. Or copy the skill folder (.agents/skills/claw-score in openclaw/openclaw) into .agents/skills/claw-score in your project. Codex loads it when a task matches its description.

Can I use Claw Score in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add openclaw/openclaw --skill claw-score -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/claw-score, .gemini/skills/claw-score, .github/skills/claw-score and .opencode/skills/claw-score in your project.

What does Claw Score need to run?

Going by SKILL.md and its folder, Claw Score needs the command-line tools its instructions call (pnpm, node and gh).

Does Claw Score access the network?

SKILL.md contains no URLs. Its commands use gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Claw Score safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Claw Score use?

Claw Score is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Claw Score use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 20k tokens, read only when the agent opens those files.

What are the alternatives to Claw Score?

Skills that share tags, products or a category with Claw Score: Harness Score (ruvnet/ruflo, 74k stars), Index Refresh (paperclipai/paperclip, 99k stars), Meta Refresh (thedaviddias/Front-End-Checklist, 74k stars) and Score Eval (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Claw Score?

openclaw (a GitHub organization) maintains it in openclaw/openclaw, which has 391,610 GitHub stars. The repository holds 93 skills in this directory. The repository was last updated on October 8, 2026.

Source: openclaw/openclaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.