Knowledge Engineering Quality And Delivery
echoVic/blade-code
覆盖 Blade Code 跨测试、构建、资格验证、发布与双语文档的工程质量闭环. An agent skill from echoVic/blade-code.
AI Agent Skill unit testing framework. An agent skill from LeoYeAI/openclaw-master-skills.
$ npx skills add LeoYeAI/openclaw-master-skills --skill eval-skills -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills eval-skills --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-skills .claude/skills/eval-skills && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-skills" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/eval-skills into .claude/skills/eval-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-skills", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/eval-skillsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add LeoYeAI/openclaw-master-skills --skill eval-skills -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills eval-skills --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/eval-skills .agents/skills/eval-skills && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-skills" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/eval-skills into .agents/skills/eval-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-skills", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LeoYeAI/openclaw-master-skills --skill eval-skills -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills eval-skills --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/eval-skills .cursor/skills/eval-skills && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-skills" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/eval-skills into .cursor/skills/eval-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-skills", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/LeoYeAI/openclaw-master-skills.git --path skills/eval-skills--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add LeoYeAI/openclaw-master-skills --skill eval-skills -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills eval-skills --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/eval-skills .gemini/skills/eval-skills && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-skills" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/eval-skills into .gemini/skills/eval-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-skills", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install LeoYeAI/openclaw-master-skills eval-skillsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add LeoYeAI/openclaw-master-skills --skill eval-skills -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/eval-skills .github/skills/eval-skills && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-skills" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/eval-skills into .github/skills/eval-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-skills", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LeoYeAI/openclaw-master-skills --skill eval-skills -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills eval-skills --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/eval-skills .opencode/skills/eval-skills && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-skills" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/eval-skills into .opencode/skills/eval-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-skills", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-skillsAI Agent Skill unit testing framework. An agent skill from LeoYeAI/openclaw-master-skills.
Eval Skills is an agent skill from LeoYeAI/openclaw-master-skills. AI Agent Skill unit testing framework. A framework-agnostic toolkit for discovering, scaffolding, selecting, evaluating, and reporting on AI skills. Use this skill to assess skill quality before production, compare candidate skills on the same benchmark, enforce quality gates in CI/CD, and generate human-readable evaluation reports.
Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 209 other files, including scripts (for example `CHANGELOG.md`, `README.md` and `_meta.json`).
It sits in Testing & QA, covering Unit testing, Project scaffolding and Quality gates. The repository describes itself as: 🧠 Curated collection of 1209+ best OpenClaw skills — weekly updated by MyClaw.ai. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit e5199b5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval Skills loads about 4.3k tokens when it runs. Until then it costs about 87 tokens; SKILL.md has 1,302 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from LeoYeAI/openclaw-master-skills at commit e5199b5, republished under its MIT licence (© LeoYeAI). 1,302 words, ~4,260 tokens.
.claude/skills/eval-skills/SKILL.md (or your agent's skills folder). This skill also uses 201 other files; get the full folder from GitHub.AI Agent Skill unit testing framework — a framework-agnostic toolkit for discovering, scaffolding, selecting, evaluating, and reporting on AI skills.
This skill fills the L1 (Skill Unit Test) gap that LangSmith / DeepEval leave open: while those platforms focus on agent-level and trajectory-level evaluation (L2-L3), eval-skills targets the individual skill level, ensuring each building block meets quality standards before it ever enters an agent pipeline.
eval to verify it meets your quality gate.select to rank them on the same benchmark.report diff to detect regressions.--exit-on-fail to block merges that degrade skill quality.create to generate a ready-to-fill skeleton.Search for existing skills by keyword, tag, or adapter type.
eval-skills find \
--query "web search" \
--tag retrieval api \
--adapter http \
--min-completion 0.8 \
--skills-dir ./skills \
--limit 10| Option | Description | Default |
|---|---|---|
-q, --query <string> | Keyword search (matches name, description, tags) | — |
-t, --tag <tags...> | Filter by tags (intersection: skill must have ALL specified tags) | — |
-a, --adapter <type> | Filter by adapter type (http, subprocess, mcp) | — |
--min-completion <rate> | Minimum historical completion rate (0.0 ~ 1.0) | — |
--skills-dir <dir> | Directory to scan for skill.json files | ./skills |
--limit <n> | Maximum number of results | 20 |
Results are ranked by search relevance (when --query is provided) or by historical completion rate (descending).
Generate a skill skeleton from a template to bootstrap development.
eval-skills create \
--name my_api_skill \
--from-template http_request \
--output-dir ./skills \
--description "Fetches weather data from OpenWeather API"| Option | Description | Default |
|---|---|---|
--name <name> | Required. Skill name | — |
--from-template <tpl> | Template type: http_request, python_script, mcp_tool | http_request |
--output-dir <dir> | Output directory | ./skills |
--description <text> | Human-readable description embedded in skill.json | — |
Generated file structure:
skills/my_api_skill/
skill.json # Skill metadata (id, schemas, adapter config)
adapter.config.json # Adapter-specific configuration
tests/
basic.eval.json # A starter benchmark with one sample task
skill.py # (python_script template only) JSON-RPC entrypointRun benchmark evaluations against one or more skills. This is the core command.
eval-skills eval \
--skills ./skills/calculator/skill.json ./skills/search/ \
--benchmark coding-easy \
--concurrency 4 \
--timeout 30000 \
--retries 2 \
--runs 3 \
--evaluator exact \
--format json markdown html \
--output-dir ./reports \
--exit-on-fail --min-completion 0.8 \
--store ./eval-skills.db| Option | Description | Default |
|---|---|---|
--skills <paths...> | Required. Skill file(s) or directory(ies) | — |
--benchmark <id|path> | Built-in benchmark ID or path to benchmark.json | coding-easy |
--tasks <file> | Custom tasks JSON file (replaces benchmark) | — |
--concurrency <n> | Number of parallel task executions | 4 |
--timeout <ms> | Per-task timeout in milliseconds | 30000 |
--retries <n> | Retry count on task failure (with incremental backoff) | 0 |
--runs <n> | Repeat evaluation N times for consistency scoring | 1 |
--evaluator <type> | Default scorer type (see Scorer Types below) | exact |
--format <formats...> | Output formats: json, markdown, html | json markdown |
--output-dir <dir> | Report output directory | ./reports |
--exit-on-fail | Exit with code 1 if any skill falls below threshold | disabled |
--min-completion <rate> | Threshold for --exit-on-fail | 0.7 |
--dry-run | Validate configuration only; do not execute tasks | disabled |
--benchmarks-dir <dir> | Directory containing built-in benchmarks | ./benchmarks |
--store <path> | SQLite database path for persistent result storage | ./eval-skills.db |
-c, --config <path> | Path to eval-skills.config.yaml | auto-detected |
Evaluation flow:
--skills paths (supports both single skill.json and directories)--benchmark or --tasksskills x tasks x runs--concurrency, with timeout and retry)SkillCompletionReport per skill--output-dirFilter and rank skills based on evaluation reports using a multi-dimensional strategy.
eval-skills select \
--from ./skills \
--reports ./reports/eval-result.json \
--strategy ./strategy.yaml \
--min-completion 0.8 \
--top-k 5 \
--output ./selected.json| Option | Description | Default |
|---|---|---|
--from <path> | Required. Candidate skills directory or JSON file | — |
--reports <file> | Evaluation reports JSON file | — |
--strategy <file> | SelectStrategy YAML/JSON file | built-in default |
--min-completion <rate> | Override minimum completion rate filter | — |
--top-k <n> | Return only the top K results | all |
--output <file> | Write selected skills to file | stdout |
Selection pipeline: Filter (by completion rate, error rate, latency, adapter type, required tags) -> Score -> Rank (by compositeScore, completionRate, latency, or tokenCost) -> TopK
Example strategy.yaml:
filters:
minCompletionRate: 0.8
maxErrorRate: 0.1
maxLatencyP95Ms: 5000
adapterTypes: [http, subprocess]
requiredTags: [production-ready]
sortBy: compositeScore
order: desc
topK: 5Execute the full end-to-end pipeline: Find -> Eval -> Select -> Report in a single command.
eval-skills run \
--query "math" \
--benchmark coding-easy \
--skills-dir ./skills \
--top-k 3 \
--min-completion 0.7 \
--format json markdown \
--output-dir ./reportsThis command automates the entire process:
--skills-dir and optionally filters by --query--benchmark--min-completion, --top-k, and optional --strategy--formatseval-skills report convert \
--input ./reports/eval-result.json \
--format html \
--output ./reports/eval-result.htmlSupported output formats: markdown, html.
eval-skills report diff \
./reports/v1.json ./reports/v2.json \
--label-a "v1.0" --label-b "v2.0" \
--output ./reports/diff.mdGenerates a side-by-side delta table per skill showing changes in completion rate, error rate, P95 latency, and composite score with directional arrows.
eval-skills init --dir .Creates the project scaffold:
eval-skills.config.yaml — global configurationskills/ — directory for skill definitionsbenchmarks/ — directory for benchmark filesreports/ — directory for evaluation output# List all current configuration values
eval-skills config list
# Get a specific value (supports dot notation)
eval-skills config get llm.model
# Set a value (persisted to ~/.eval-skills/config.yaml)
eval-skills config set concurrency 8
eval-skills config set llm.model gpt-4o
eval-skills config set llm.temperature 0Configuration is resolved in priority order:
eval-skills.config.yaml in current directory~/.eval-skills/config.yamlconcurrency: 4, timeoutMs: 30000, outputDir: ./reports)Each task in a benchmark specifies an evaluator type. The scorer compares the skill's actual output against the expected output.
| Type | Aliases | Description | Score Range |
|---|---|---|---|
exact_match | exact | Strict equality comparison. Supports caseSensitive option. | 0 or 1 |
contains | — | Checks for the presence of all specified keywords in the output. Partial credit: matched_keywords / total_keywords. | 0.0 ~ 1.0 |
json_schema | schema | Validates output against a JSON Schema (using Ajv). | 0 or 1 |
llm_judge | — | Sends the output + expected rubric to an LLM (configurable model) for quality rating. | 0.0 ~ 1.0 |
custom | — | Loads a custom scorer from expectedOutput.customScorerPath. | 0.0 ~ 1.0 |
Every evaluation produces a SkillCompletionReport with these metrics:
| Metric | Description | Formula |
|---|---|---|
| Completion Rate | Fraction of tasks that passed | pass_count / total_count |
| Partial Score | Mean score across all tasks | mean(task_scores) |
| Error Rate | Fraction of tasks that errored or timed out | (error_count + timeout_count) / total_count |
| Consistency Score | Stability across multiple runs (requires --runs >= 2) | 1 - stddev(per_run_completion_rates) |
| P50 / P95 / P99 Latency | Response time percentiles | Sorted percentile of latencyMs |
| Composite Score | Weighted overall quality score | 0.5 * CR + 0.2 * (1 - latP95_norm) + 0.3 * (1 - ER) |
| ID | Domain | Tasks | Scoring | Description |
|---|---|---|---|---|
coding-easy | coding | 20 | mean / exact_match | Math expressions, string reversal, palindrome detection |
skill-quality | tool-use | 5 | mean / contains | Metadata completeness, description quality, structure checks |
web-search-basic | web | 8 | mean / contains + schema | Factual queries, keyword verification, structured output validation |
gaia-v1 | general | — | mean | Placeholder for GAIA benchmark Level 1 tasks |
toolbench-lite | tool-use | — | mean | Placeholder for ToolBench single-tool scenarios |
Create a benchmark.json file:
{
"id": "my-benchmark",
"name": "My Custom Benchmark",
"version": "1.0.0",
"domain": "general",
"scoringMethod": "mean",
"maxLatencyMs": 30000,
"metadata": { "source": "internal", "lastUpdated": "2026-02-28" },
"tasks": [
{
"id": "task_001",
"description": "Test basic addition",
"inputData": { "expression": "2+3" },
"expectedOutput": { "type": "exact", "value": "5" },
"evaluator": { "type": "exact" },
"timeoutMs": 10000,
"tags": ["math"]
},
{
"id": "task_002",
"description": "Test keyword presence",
"inputData": { "query": "TypeScript" },
"expectedOutput": { "type": "contains", "keywords": ["JavaScript", "Microsoft"] },
"evaluator": { "type": "contains", "caseSensitive": false },
"timeoutMs": 15000,
"tags": ["search"]
}
]
}eval-skills eval --skills ./my-skill/ --benchmark ./my-benchmark.jsonSkills communicate through adapters. The adapter type is specified in skill.json via adapterType.
| Adapter | Protocol | How it works | Key config |
|---|---|---|---|
http | REST POST | Sends POST { skillId, version, input } to skill.entrypoint. Supports Bearer / API-Key auth via env vars. | baseUrl, authType, authTokenEnvKey |
subprocess | JSON-RPC 2.0 over stdin/stdout | Spawns skill.entrypoint (e.g. python3 skill.py), writes JSON-RPC request to stdin, reads response from stdout. | command, args |
mcp | MCP Protocol | (Phase 2) Native Model Context Protocol integration via @modelcontextprotocol/sdk. | — |
# 1. Create a skill skeleton
eval-skills create --name my_calc --from-template python_script
# 2. Implement your logic in skills/my_calc/skill.py
# 3. Run evaluation against the coding-easy benchmark
eval-skills eval \
--skills ./skills/my_calc/skill.json \
--benchmark coding-easy \
--runs 3 \
--format json markdown
# 4. Review the report
cat ./reports/eval-result-*.md# 1. Discover candidates
eval-skills find --query "weather" --skills-dir ./skills
# 2. Evaluate all candidates on the same benchmark
eval-skills eval \
--skills ./skills/weather_v1 ./skills/weather_v2 ./skills/weather_v3 \
--benchmark web-search-basic \
--runs 3
# 3. Select the best
eval-skills select \
--from ./skills \
--reports ./reports/eval-result-*.json \
--min-completion 0.8 \
--top-k 2
# 4. Compare two versions
eval-skills report diff \
./reports/v1.json ./reports/v2.json \
--label-a "weather_v1" --label-b "weather_v2"eval-skills run \
--skills-dir ./skills \
--benchmark coding-easy \
--top-k 3 \
--min-completion 0.7 \
--format json markdown html \
--output-dir ./reports# In your CI pipeline — fail the build if completion rate drops below 80%
eval-skills eval \
--skills ./skills/production_skill \
--benchmark coding-easy \
--exit-on-fail \
--min-completion 0.8 \
--format json# Compare today's evaluation against the baseline
eval-skills report diff \
./reports/baseline.json ./reports/latest.json \
--label-a "baseline" --label-b "latest" \
--output ./reports/regression-check.mdAlways use --runs 3 or more when evaluating for production decisions. Single-run results can be noisy; the consistency score captures stability across runs.
Use --exit-on-fail in CI/CD pipelines to enforce quality gates. Set --min-completion to your acceptable threshold (recommended: 0.8 for production skills).
Create domain-specific custom benchmarks rather than relying solely on built-in ones. Your custom benchmark should reflect real-world inputs your skill will encounter.
Use report diff after every skill upgrade to catch regressions early. Compare the new evaluation against a saved baseline report.
Use --dry-run before long evaluations to validate your configuration (skill paths, benchmark resolution, task count) without actually executing tasks.
Persist results with --store to track skill quality over time. The SQLite store enables historical trend queries.
Start with --concurrency 1 when debugging a failing skill, then increase for production benchmarking.
Tag your benchmark tasks to enable per-category analysis (e.g., filter by math, string, edge-case).
Every skill must provide a skill.json that conforms to this structure:
{
"id": "my_skill_v1",
"name": "My Skill",
"version": "1.0.0",
"description": "Does something useful",
"tags": ["utility", "math"],
"inputSchema": {
"type": "object",
"properties": { "query": { "type": "string" } },
"required": ["query"]
},
"outputSchema": {
"type": "object",
"properties": { "result": { "type": "string" } }
},
"adapterType": "subprocess",
"entrypoint": "python3 skill.py",
"metadata": {
"author": "Your Name",
"license": "MIT",
"homepage": "https://github.com/you/my-skill"
}
}Validation rules:
id: lowercase alphanumeric with _ or -, non-emptyversion: semver format (X.Y.Z)adapterType: one of http, subprocess, mcp, langchain, customentrypoint: non-empty string (URL for http, command for subprocess)These options are available on all commands:
| Option | Description |
|---|---|
-c, --config <path> | Path to configuration file |
--json | JSON output format (CI-friendly) |
--no-color | Disable colored output |
-v, --verbose | Verbose logging |
--version | Show version |
-h, --help | Show help |
© LeoYeAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 201 other files (scripts) in skills/eval-skills of LeoYeAI/openclaw-master-skills.
Open the folder on GitHubat commit e5199b5
Eval Skills next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval Skills this skillLeoYeAI/openclaw-master-skills | 2.2k | — | ~4.3k | Automated safety check: Pass | MIT | |
| Knowledge Engineering Quality And DeliveryechoVic/blade-code | 181 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Shift Left Testingpetrkindlmann/qa-skills | 170 | — | ~6.7k | Automated safety check: Pass | MIT | |
| Cicd Pipeline Qe Orchestratorproffesor-for-testing/agentic-qe | 495 | — | ~2.7k | Automated safety check: Pass | MIT | |
| Palantir CI Integrationjeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~1.3k | Automated safety check: Pass | MIT | |
| Qcsd Cicd Swarmproffesor-for-testing/agentic-qe | 495 | — | ~2.3k | Automated safety check: Pass | MIT |
echoVic/blade-code
覆盖 Blade Code 跨测试、构建、资格验证、发布与双语文档的工程质量闭环. An agent skill from echoVic/blade-code.
petrkindlmann/qa-skills
Move quality earlier in the development lifecycle. An agent skill from petrkindlmann/qa-skills.
proffesor-for-testing/agentic-qe
Orchestrate quality engineering across CI/CD pipeline phases.
jeremylongshore/tons-of-skills-marketplace
Design and verify Foundry-native continuous-integration and review gates for Code Repositories and transforms.
proffesor-for-testing/agentic-qe
A skill your agent uses when enforcing CI/CD quality gates before release, running regression analysis, detecting flaky tests, or assessing deployment readiness in the QCSD Verification phase.
kid-sid/claude-spellbook
A skill your agent uses when verifying a fresh deployment, gating a CI pipeline before full test runs, checking that core user flows are reachable after a release, or building a minimal health-check…
LeoYeAI/openclaw-master-skills
Manages pipelines on a DevOps quality and efficiency platform through its OpenAPI: list workspaces and templates, create, update, run and cancel pipelines, and read run records.
LeoYeAI/openclaw-master-skills
Patches OpenClaw's Feishu extension so an edited document triggers an isolated agent session that reads the doc and replies inline, turning it into a live chat space.
LeoYeAI/openclaw-master-skills
Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.
LeoYeAI/openclaw-master-skills
Runs a brand's AI-search visibility work end to end: diagnosing how AI platforms represent it, repositioning it, producing AI-optimized content and monitoring ongoing mentions.
LeoYeAI/openclaw-master-skills
Installs and authenticates the gws CLI, then automates Gmail, Drive, Sheets, Calendar, Docs, Chat and Tasks with ready-made recipes, persona bundles and security audits.
LeoYeAI/openclaw-master-skills
Runs four advisor roles, a fitness coach, nutritionist, data analyst and TCM practitioner, to build a health profile and track workouts, diet and wellness over time.
Categories
AI Agent Skill unit testing framework. An agent skill from LeoYeAI/openclaw-master-skills. Eval Skills is an agent skill from LeoYeAI/openclaw-master-skills. AI Agent Skill unit testing framework.
Eval Skills fits situations like: assess skill quality before production; compare candidate skills on the same benchmark; enforce quality gates in CI/CD; generate human-readable evaluation reports.
Run `npx skills add LeoYeAI/openclaw-master-skills --skill eval-skills -a claude-code`. Or copy the skill folder (skills/eval-skills in LeoYeAI/openclaw-master-skills) into .claude/skills/eval-skills in your project. Claude Code loads it when a task matches its description.
Run `npx skills add LeoYeAI/openclaw-master-skills --skill eval-skills -a codex`. Or copy the skill folder (skills/eval-skills in LeoYeAI/openclaw-master-skills) into .agents/skills/eval-skills in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LeoYeAI/openclaw-master-skills --skill eval-skills -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-skills, .gemini/skills/eval-skills, .github/skills/eval-skills and .opencode/skills/eval-skills in your project.
Going by SKILL.md and its folder, Eval Skills needs the command-line tools its instructions call (python3).
SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Eval Skills is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Eval Skills: Knowledge Engineering Quality And Delivery (echoVic/blade-code, 181 stars), Shift Left Testing (petrkindlmann/qa-skills, 170 stars), Cicd Pipeline Qe Orchestrator (proffesor-for-testing/agentic-qe, 495 stars) and Palantir CI Integration (jeremylongshore/tons-of-skills-marketplace, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
LeoYeAI (a GitHub user) maintains it in LeoYeAI/openclaw-master-skills, which has 2,161 GitHub stars. The repository holds 1,235 skills in this directory. The repository was last updated on July 20, 2026.
Source: LeoYeAI/openclaw-master-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.