Agent skill

Eval Skills

by LeoYeAI in LeoYeAI/openclaw-master-skills

AI Agent Skill unit testing framework. An agent skill from LeoYeAI/openclaw-master-skills.

MITAuto-check passedTesting & QA

Install Eval Skills

skills CLI
$ npx skills add LeoYeAI/openclaw-master-skills --skill eval-skills -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install LeoYeAI/openclaw-master-skills eval-skills --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-skills .claude/skills/eval-skills && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-skills
GitHub stars
2.2k
Token cost
~4.3k tokens
SKILL.md length
1,302 words
Files
202 (incl. scripts)
Skills in repo
1,235
Repo updated
First seen
Licence
MIT

At a glance

AI Agent Skill unit testing framework. An agent skill from LeoYeAI/openclaw-master-skills.

  • Works in 8 steps: Find Skills → Create Skills → Evaluate Skills → …
  • Assess skill quality before production
  • SKILL.md covers When to Use This Skill, Capabilities, Scorer Types and Evaluation Metrics, plus 6 more sections
  • Calls python3; reaches github.com

What it does

Eval Skills is an agent skill from LeoYeAI/openclaw-master-skills. AI Agent Skill unit testing framework. A framework-agnostic toolkit for discovering, scaffolding, selecting, evaluating, and reporting on AI skills. Use this skill to assess skill quality before production, compare candidate skills on the same benchmark, enforce quality gates in CI/CD, and generate human-readable evaluation reports.

Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 209 other files, including scripts (for example `CHANGELOG.md`, `README.md` and `_meta.json`).

It sits in Testing & QA, covering Unit testing, Project scaffolding and Quality gates. The repository describes itself as: 🧠 Curated collection of 1209+ best OpenClaw skills — weekly updated by MyClaw.ai. The licence is MIT.

When your agent uses it

  • Assess skill quality before production
  • Compare candidate skills on the same benchmark
  • Enforce quality gates in CI/CD
  • Generate human-readable evaluation reports

Example prompts

  • “/eval-skills”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Find Skills
  2. Create Skills
  3. Evaluate Skills
  4. Select Skills
  5. Run Pipeline
  6. Generate & Compare Reports
  7. Initialize Project
  8. Manage Configuration

What it can do on your machine

Read from SKILL.md and the folder at commit e5199b5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Skills loads about 4.3k tokens when it runs. Until then it costs about 87 tokens; SKILL.md has 1,302 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~87
When it runs · the whole SKILL.md, loaded when a task matches
~4.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from LeoYeAI/openclaw-master-skills at commit e5199b5, republished under its MIT licence (© LeoYeAI). 1,302 words, ~4,260 tokens.

Download SKILL.mdSave it as .claude/skills/eval-skills/SKILL.md (or your agent's skills folder). This skill also uses 201 other files; get the full folder from GitHub.
name
eval-skills
description
AI Agent Skill unit testing framework. A framework-agnostic toolkit for discovering, scaffolding, selecting, evaluating, and reporting on AI skills. Use this skill to assess skill quality before production, compare candidate skills on the same benchmark, enforce quality gates in CI/CD, and generate human-readable evaluation reports.
version
0.1.0

eval-skills

AI Agent Skill unit testing framework — a framework-agnostic toolkit for discovering, scaffolding, selecting, evaluating, and reporting on AI skills.

This skill fills the L1 (Skill Unit Test) gap that LangSmith / DeepEval leave open: while those platforms focus on agent-level and trajectory-level evaluation (L2-L3), eval-skills targets the individual skill level, ensuring each building block meets quality standards before it ever enters an agent pipeline.

When to Use This Skill

  • Before deploying a new skill to production — run eval to verify it meets your quality gate.
  • When choosing between multiple candidate skills — run select to rank them on the same benchmark.
  • When a skill is upgraded — run report diff to detect regressions.
  • In CI/CD — use --exit-on-fail to block merges that degrade skill quality.
  • When bootstrapping a new skill — run create to generate a ready-to-fill skeleton.

Capabilities

1. Find Skills

Search for existing skills by keyword, tag, or adapter type.

bash
eval-skills find \
  --query "web search" \
  --tag retrieval api \
  --adapter http \
  --min-completion 0.8 \
  --skills-dir ./skills \
  --limit 10
OptionDescriptionDefault
-q, --query <string>Keyword search (matches name, description, tags)—
-t, --tag <tags...>Filter by tags (intersection: skill must have ALL specified tags)—
-a, --adapter <type>Filter by adapter type (http, subprocess, mcp)—
--min-completion <rate>Minimum historical completion rate (0.0 ~ 1.0)—
--skills-dir <dir>Directory to scan for skill.json files./skills
--limit <n>Maximum number of results20

Results are ranked by search relevance (when --query is provided) or by historical completion rate (descending).

2. Create Skills

Generate a skill skeleton from a template to bootstrap development.

bash
eval-skills create \
  --name my_api_skill \
  --from-template http_request \
  --output-dir ./skills \
  --description "Fetches weather data from OpenWeather API"
OptionDescriptionDefault
--name <name>Required. Skill name—
--from-template <tpl>Template type: http_request, python_script, mcp_toolhttp_request
--output-dir <dir>Output directory./skills
--description <text>Human-readable description embedded in skill.json—

Generated file structure:

skills/my_api_skill/
  skill.json            # Skill metadata (id, schemas, adapter config)
  adapter.config.json   # Adapter-specific configuration
  tests/
    basic.eval.json     # A starter benchmark with one sample task
  skill.py              # (python_script template only) JSON-RPC entrypoint
3. Evaluate Skills

Run benchmark evaluations against one or more skills. This is the core command.

bash
eval-skills eval \
  --skills ./skills/calculator/skill.json ./skills/search/ \
  --benchmark coding-easy \
  --concurrency 4 \
  --timeout 30000 \
  --retries 2 \
  --runs 3 \
  --evaluator exact \
  --format json markdown html \
  --output-dir ./reports \
  --exit-on-fail --min-completion 0.8 \
  --store ./eval-skills.db
OptionDescriptionDefault
--skills <paths...>Required. Skill file(s) or directory(ies)—
--benchmark <id|path>Built-in benchmark ID or path to benchmark.jsoncoding-easy
--tasks <file>Custom tasks JSON file (replaces benchmark)—
--concurrency <n>Number of parallel task executions4
--timeout <ms>Per-task timeout in milliseconds30000
--retries <n>Retry count on task failure (with incremental backoff)0
--runs <n>Repeat evaluation N times for consistency scoring1
--evaluator <type>Default scorer type (see Scorer Types below)exact
--format <formats...>Output formats: json, markdown, htmljson markdown
--output-dir <dir>Report output directory./reports
--exit-on-failExit with code 1 if any skill falls below thresholddisabled
--min-completion <rate>Threshold for --exit-on-fail0.7
--dry-runValidate configuration only; do not execute tasksdisabled
--benchmarks-dir <dir>Directory containing built-in benchmarks./benchmarks
--store <path>SQLite database path for persistent result storage./eval-skills.db
-c, --config <path>Path to eval-skills.config.yamlauto-detected

Evaluation flow:

  1. Load skills from --skills paths (supports both single skill.json and directories)
  2. Load benchmark tasks from --benchmark or --tasks
  3. Build the cartesian product: skills x tasks x runs
  4. Execute all task items concurrently (controlled by --concurrency, with timeout and retry)
  5. Score each result using the appropriate scorer
  6. Aggregate into SkillCompletionReport per skill
  7. Write reports to --output-dir
4. Select Skills

Filter and rank skills based on evaluation reports using a multi-dimensional strategy.

bash
eval-skills select \
  --from ./skills \
  --reports ./reports/eval-result.json \
  --strategy ./strategy.yaml \
  --min-completion 0.8 \
  --top-k 5 \
  --output ./selected.json
OptionDescriptionDefault
--from <path>Required. Candidate skills directory or JSON file—
--reports <file>Evaluation reports JSON file—
--strategy <file>SelectStrategy YAML/JSON filebuilt-in default
--min-completion <rate>Override minimum completion rate filter—
--top-k <n>Return only the top K resultsall
--output <file>Write selected skills to filestdout

Selection pipeline: Filter (by completion rate, error rate, latency, adapter type, required tags) -> Score -> Rank (by compositeScore, completionRate, latency, or tokenCost) -> TopK

Example strategy.yaml:

yaml
filters:
  minCompletionRate: 0.8
  maxErrorRate: 0.1
  maxLatencyP95Ms: 5000
  adapterTypes: [http, subprocess]
  requiredTags: [production-ready]
sortBy: compositeScore
order: desc
topK: 5
5. Run Pipeline

Execute the full end-to-end pipeline: Find -> Eval -> Select -> Report in a single command.

bash
eval-skills run \
  --query "math" \
  --benchmark coding-easy \
  --skills-dir ./skills \
  --top-k 3 \
  --min-completion 0.7 \
  --format json markdown \
  --output-dir ./reports

This command automates the entire process:

  1. Find — scans --skills-dir and optionally filters by --query
  2. Eval — evaluates all candidate skills against --benchmark
  3. Select — filters and ranks results using --min-completion, --top-k, and optional --strategy
  4. Report — generates output files in all requested --formats
6. Generate & Compare Reports
Convert report format
bash
eval-skills report convert \
  --input ./reports/eval-result.json \
  --format html \
  --output ./reports/eval-result.html

Supported output formats: markdown, html.

Diff two reports (regression detection)
bash
eval-skills report diff \
  ./reports/v1.json ./reports/v2.json \
  --label-a "v1.0" --label-b "v2.0" \
  --output ./reports/diff.md

Generates a side-by-side delta table per skill showing changes in completion rate, error rate, P95 latency, and composite score with directional arrows.

7. Initialize Project
bash
eval-skills init --dir .

Creates the project scaffold:

  • eval-skills.config.yaml — global configuration
  • skills/ — directory for skill definitions
  • benchmarks/ — directory for benchmark files
  • reports/ — directory for evaluation output
8. Manage Configuration
bash
# List all current configuration values
eval-skills config list

# Get a specific value (supports dot notation)
eval-skills config get llm.model

# Set a value (persisted to ~/.eval-skills/config.yaml)
eval-skills config set concurrency 8
eval-skills config set llm.model gpt-4o
eval-skills config set llm.temperature 0

Configuration is resolved in priority order:

  1. CLI flags (highest priority)
  2. eval-skills.config.yaml in current directory
  3. ~/.eval-skills/config.yaml
  4. Built-in defaults (concurrency: 4, timeoutMs: 30000, outputDir: ./reports)

Scorer Types

Each task in a benchmark specifies an evaluator type. The scorer compares the skill's actual output against the expected output.

TypeAliasesDescriptionScore Range
exact_matchexactStrict equality comparison. Supports caseSensitive option.0 or 1
contains—Checks for the presence of all specified keywords in the output. Partial credit: matched_keywords / total_keywords.0.0 ~ 1.0
json_schemaschemaValidates output against a JSON Schema (using Ajv).0 or 1
llm_judge—Sends the output + expected rubric to an LLM (configurable model) for quality rating.0.0 ~ 1.0
custom—Loads a custom scorer from expectedOutput.customScorerPath.0.0 ~ 1.0
Show full SKILL.md (479 more words)Show less

Evaluation Metrics

Every evaluation produces a SkillCompletionReport with these metrics:

MetricDescriptionFormula
Completion RateFraction of tasks that passedpass_count / total_count
Partial ScoreMean score across all tasksmean(task_scores)
Error RateFraction of tasks that errored or timed out(error_count + timeout_count) / total_count
Consistency ScoreStability across multiple runs (requires --runs >= 2)1 - stddev(per_run_completion_rates)
P50 / P95 / P99 LatencyResponse time percentilesSorted percentile of latencyMs
Composite ScoreWeighted overall quality score0.5 * CR + 0.2 * (1 - latP95_norm) + 0.3 * (1 - ER)

Built-in Benchmarks

IDDomainTasksScoringDescription
coding-easycoding20mean / exact_matchMath expressions, string reversal, palindrome detection
skill-qualitytool-use5mean / containsMetadata completeness, description quality, structure checks
web-search-basicweb8mean / contains + schemaFactual queries, keyword verification, structured output validation
gaia-v1general—meanPlaceholder for GAIA benchmark Level 1 tasks
toolbench-litetool-use—meanPlaceholder for ToolBench single-tool scenarios
Custom Benchmark

Create a benchmark.json file:

json
{
  "id": "my-benchmark",
  "name": "My Custom Benchmark",
  "version": "1.0.0",
  "domain": "general",
  "scoringMethod": "mean",
  "maxLatencyMs": 30000,
  "metadata": { "source": "internal", "lastUpdated": "2026-02-28" },
  "tasks": [
    {
      "id": "task_001",
      "description": "Test basic addition",
      "inputData": { "expression": "2+3" },
      "expectedOutput": { "type": "exact", "value": "5" },
      "evaluator": { "type": "exact" },
      "timeoutMs": 10000,
      "tags": ["math"]
    },
    {
      "id": "task_002",
      "description": "Test keyword presence",
      "inputData": { "query": "TypeScript" },
      "expectedOutput": { "type": "contains", "keywords": ["JavaScript", "Microsoft"] },
      "evaluator": { "type": "contains", "caseSensitive": false },
      "timeoutMs": 15000,
      "tags": ["search"]
    }
  ]
}
bash
eval-skills eval --skills ./my-skill/ --benchmark ./my-benchmark.json

Adapter Types

Skills communicate through adapters. The adapter type is specified in skill.json via adapterType.

AdapterProtocolHow it worksKey config
httpREST POSTSends POST { skillId, version, input } to skill.entrypoint. Supports Bearer / API-Key auth via env vars.baseUrl, authType, authTokenEnvKey
subprocessJSON-RPC 2.0 over stdin/stdoutSpawns skill.entrypoint (e.g. python3 skill.py), writes JSON-RPC request to stdin, reads response from stdout.command, args
mcpMCP Protocol(Phase 2) Native Model Context Protocol integration via @modelcontextprotocol/sdk.—

Workflow Examples

Evaluating a Single Skill
bash
# 1. Create a skill skeleton
eval-skills create --name my_calc --from-template python_script

# 2. Implement your logic in skills/my_calc/skill.py

# 3. Run evaluation against the coding-easy benchmark
eval-skills eval \
  --skills ./skills/my_calc/skill.json \
  --benchmark coding-easy \
  --runs 3 \
  --format json markdown

# 4. Review the report
cat ./reports/eval-result-*.md
Comparing Multiple Candidate Skills
bash
# 1. Discover candidates
eval-skills find --query "weather" --skills-dir ./skills

# 2. Evaluate all candidates on the same benchmark
eval-skills eval \
  --skills ./skills/weather_v1 ./skills/weather_v2 ./skills/weather_v3 \
  --benchmark web-search-basic \
  --runs 3

# 3. Select the best
eval-skills select \
  --from ./skills \
  --reports ./reports/eval-result-*.json \
  --min-completion 0.8 \
  --top-k 2

# 4. Compare two versions
eval-skills report diff \
  ./reports/v1.json ./reports/v2.json \
  --label-a "weather_v1" --label-b "weather_v2"
Full Pipeline (One Command)
bash
eval-skills run \
  --skills-dir ./skills \
  --benchmark coding-easy \
  --top-k 3 \
  --min-completion 0.7 \
  --format json markdown html \
  --output-dir ./reports
CI/CD Quality Gate
bash
# In your CI pipeline — fail the build if completion rate drops below 80%
eval-skills eval \
  --skills ./skills/production_skill \
  --benchmark coding-easy \
  --exit-on-fail \
  --min-completion 0.8 \
  --format json
Regression Detection
bash
# Compare today's evaluation against the baseline
eval-skills report diff \
  ./reports/baseline.json ./reports/latest.json \
  --label-a "baseline" --label-b "latest" \
  --output ./reports/regression-check.md

Best Practices

  1. Always use --runs 3 or more when evaluating for production decisions. Single-run results can be noisy; the consistency score captures stability across runs.

  2. Use --exit-on-fail in CI/CD pipelines to enforce quality gates. Set --min-completion to your acceptable threshold (recommended: 0.8 for production skills).

  3. Create domain-specific custom benchmarks rather than relying solely on built-in ones. Your custom benchmark should reflect real-world inputs your skill will encounter.

  4. Use report diff after every skill upgrade to catch regressions early. Compare the new evaluation against a saved baseline report.

  5. Use --dry-run before long evaluations to validate your configuration (skill paths, benchmark resolution, task count) without actually executing tasks.

  6. Persist results with --store to track skill quality over time. The SQLite store enables historical trend queries.

  7. Start with --concurrency 1 when debugging a failing skill, then increase for production benchmarking.

  8. Tag your benchmark tasks to enable per-category analysis (e.g., filter by math, string, edge-case).

Skill JSON Schema

Every skill must provide a skill.json that conforms to this structure:

json
{
  "id": "my_skill_v1",
  "name": "My Skill",
  "version": "1.0.0",
  "description": "Does something useful",
  "tags": ["utility", "math"],
  "inputSchema": {
    "type": "object",
    "properties": { "query": { "type": "string" } },
    "required": ["query"]
  },
  "outputSchema": {
    "type": "object",
    "properties": { "result": { "type": "string" } }
  },
  "adapterType": "subprocess",
  "entrypoint": "python3 skill.py",
  "metadata": {
    "author": "Your Name",
    "license": "MIT",
    "homepage": "https://github.com/you/my-skill"
  }
}

Validation rules:

  • id: lowercase alphanumeric with _ or -, non-empty
  • version: semver format (X.Y.Z)
  • adapterType: one of http, subprocess, mcp, langchain, custom
  • entrypoint: non-empty string (URL for http, command for subprocess)

Global Options

These options are available on all commands:

OptionDescription
-c, --config <path>Path to configuration file
--jsonJSON output format (CI-friendly)
--no-colorDisable colored output
-v, --verboseVerbose logging
--versionShow version
-h, --helpShow help

© LeoYeAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 201 other files (scripts) in skills/eval-skills of LeoYeAI/openclaw-master-skills.

  • SKILL.md
  • CHANGELOG.md
  • README.md
  • _meta.json
  • benchmarks/coding-easy/benchmark.json
  • benchmarks/gaia-v1/benchmark.json
  • benchmarks/skill-quality/benchmark.json
  • benchmarks/toolbench-lite/benchmark.json
  • benchmarks/web-search-basic/benchmark.json
  • docs/guides/ci-cd-integration.md
  • docs/guides/create-skill.md
  • docs/guides/custom-benchmark.md
  • docs/guides/quickstart.md
  • … and 189 more

Open the folder on GitHubat commit e5199b5

Compare with similar skills

Eval Skills next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Skills compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Skills this skillLeoYeAI/openclaw-master-skills2.2k—~4.3kAutomated safety check: PassMIT
Knowledge Engineering Quality And DeliveryechoVic/blade-code181—~1.4kAutomated safety check: PassMIT
Shift Left Testingpetrkindlmann/qa-skills170—~6.7kAutomated safety check: PassMIT
Cicd Pipeline Qe Orchestratorproffesor-for-testing/agentic-qe495—~2.7kAutomated safety check: PassMIT
Palantir CI Integrationjeremylongshore/tons-of-skills-marketplace2.8k—~1.3kAutomated safety check: PassMIT
Qcsd Cicd Swarmproffesor-for-testing/agentic-qe495—~2.3kAutomated safety check: PassMIT

Similar skills

  • 覆盖 Blade Code 跨测试、构建、资格验证、发布与双语文档的工程质量闭环. An agent skill from echoVic/blade-code.

    181 GitHub stars~1.4k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Shift Left Testing

    petrkindlmann/qa-skills

    Move quality earlier in the development lifecycle. An agent skill from petrkindlmann/qa-skills.

    170 GitHub stars~6.7k tokensUpdated 4 mo ago
    Testing & QAAuto-check passed
  • Cicd Pipeline Qe Orchestrator

    proffesor-for-testing/agentic-qe

    Orchestrate quality engineering across CI/CD pipeline phases.

    495 GitHub stars~2.7k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Palantir CI Integration

    jeremylongshore/tons-of-skills-marketplace

    Design and verify Foundry-native continuous-integration and review gates for Code Repositories and transforms.

    2.8k GitHub stars~1.3k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Qcsd Cicd Swarm

    proffesor-for-testing/agentic-qe

    A skill your agent uses when enforcing CI/CD quality gates before release, running regression analysis, detecting flaky tests, or assessing deployment readiness in the QCSD Verification phase.

    495 GitHub stars~2.3k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Smoke Testing

    kid-sid/claude-spellbook

    A skill your agent uses when verifying a fresh deployment, gating a CI pipeline before full test runs, checking that core user flows are reachable after a release, or building a minimal health-check…

    190 GitHub stars~1.7k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed

More from LeoYeAI/openclaw-master-skills

All 1,200 skills in this repo
  • DevOps Pipeline Management

    LeoYeAI/openclaw-master-skills

    Manages pipelines on a DevOps quality and efficiency platform through its OpenAPI: list workspaces and templates, create, update, run and cancel pipelines, and read run records.

    2.2k GitHub stars~4.2k tokensUpdated 2 mo ago
    Auto-check: notes
  • Feishu Document Collaboration

    LeoYeAI/openclaw-master-skills

    Patches OpenClaw's Feishu extension so an edited document triggers an isolated agent session that reads the doc and replies inline, turning it into a live chat space.

    2.2k GitHub stars~2k tokensUpdated 2 mo ago
    Auto-check passed
  • Files Memory System

    LeoYeAI/openclaw-master-skills

    Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.

    2.2k GitHub stars~3.8k tokensUpdated 2 mo ago
    Auto-check passed
  • GEO-Claw AI Visibility Agent

    LeoYeAI/openclaw-master-skills

    Runs a brand's AI-search visibility work end to end: diagnosing how AI platforms represent it, repositioning it, producing AI-optimized content and monitoring ongoing mentions.

    2.2k GitHub stars~4.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Google Workspace CLI

    LeoYeAI/openclaw-master-skills

    Installs and authenticates the gws CLI, then automates Gmail, Drive, Sheets, Calendar, Docs, Chat and Tasks with ready-made recipes, persona bundles and security audits.

    2.2k GitHub stars~2.6k tokensUpdated 2 mo ago
    Auto-check: notes
  • HealthFit Health Advisors

    LeoYeAI/openclaw-master-skills

    Runs four advisor roles, a fitness coach, nutritionist, data analyst and TCM practitioner, to build a health profile and track workouts, diet and wellness over time.

    2.2k GitHub stars~4.4k tokensUpdated 2 mo ago
    Auto-check passed

Categories

Questions about Eval Skills

What does Eval Skills do?

AI Agent Skill unit testing framework. An agent skill from LeoYeAI/openclaw-master-skills. Eval Skills is an agent skill from LeoYeAI/openclaw-master-skills. AI Agent Skill unit testing framework.

When should I use Eval Skills?

Eval Skills fits situations like: assess skill quality before production; compare candidate skills on the same benchmark; enforce quality gates in CI/CD; generate human-readable evaluation reports.

How do I install Eval Skills in Claude Code?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill eval-skills -a claude-code`. Or copy the skill folder (skills/eval-skills in LeoYeAI/openclaw-master-skills) into .claude/skills/eval-skills in your project. Claude Code loads it when a task matches its description.

How do I install Eval Skills in Codex?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill eval-skills -a codex`. Or copy the skill folder (skills/eval-skills in LeoYeAI/openclaw-master-skills) into .agents/skills/eval-skills in your project. Codex loads it when a task matches its description.

Can I use Eval Skills in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LeoYeAI/openclaw-master-skills --skill eval-skills -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-skills, .gemini/skills/eval-skills, .github/skills/eval-skills and .opencode/skills/eval-skills in your project.

What does Eval Skills need to run?

Going by SKILL.md and its folder, Eval Skills needs the command-line tools its instructions call (python3).

Does Eval Skills access the network?

SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Eval Skills safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Eval Skills use?

Eval Skills is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Skills use?

About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Skills?

Skills that share tags, products or a category with Eval Skills: Knowledge Engineering Quality And Delivery (echoVic/blade-code, 181 stars), Shift Left Testing (petrkindlmann/qa-skills, 170 stars), Cicd Pipeline Qe Orchestrator (proffesor-for-testing/agentic-qe, 495 stars) and Palantir CI Integration (jeremylongshore/tons-of-skills-marketplace, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Skills?

LeoYeAI (a GitHub user) maintains it in LeoYeAI/openclaw-master-skills, which has 2,161 GitHub stars. The repository holds 1,235 skills in this directory. The repository was last updated on July 20, 2026.

Source: LeoYeAI/openclaw-master-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.