Agent skill

Evaluation

by aiskillstore in aiskillstore/marketplace

Build evaluation frameworks for agent systems. An agent skill from aiskillstore/marketplace.

No licenceAuto-check passedAgent Workflows

Install Evaluation

skills CLI
$ npx skills add aiskillstore/marketplace --skill evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aiskillstore/marketplace evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aiskillstore/marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/sickn33/evaluation .claude/skills/evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluation
GitHub stars
430
Used in
6 other repos
Token cost
~2.8k tokens
SKILL.md length
1,171 words
Files
2
Skills in repo
1,085
Repo updated
First seen
Licence
None found

At a glance

Build evaluation frameworks for agent systems. An agent skill from aiskillstore/marketplace.

  • Works in 8 steps: Define quality dimensions relevant to… → Create rubrics with clear, actionable… → Build test sets from real usage patterns… → …
  • Testing agent performance systematically
  • SKILL.md covers When to Use This Skill, When to Use, Core Concepts and Detailed Topics, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Evaluation is an agent skill from aiskillstore/marketplace. Build evaluation frameworks for agent systems. Use when testing agent performance systematically, validating context engineering choices, or measuring improvements over time.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `skill-report.json`).

It sits in Agent Workflows, covering Context engineering. The repository describes itself as: Security-audited skills for Claude, Codex & Claude Code. One-click install, quality verified.

When your agent uses it

  • Testing agent performance systematically
  • Validating context engineering choices
  • Measuring improvements over time

Example prompts

  • “/evaluation”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Define quality dimensions relevant to your use case
  2. Create rubrics with clear, actionable level descriptions
  3. Build test sets from real usage patterns and edge cases
  4. Implement automated evaluation pipelines
  5. Establish baseline metrics before making changes
  6. Run evaluations on all significant changes
  7. Track metrics over time for trend analysis
  8. Supplement automated evaluation with human review

What it can do on your machine

Read from SKILL.md and the folder at commit 4ac52da. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluation loads about 2.8k tokens when it runs. Until then it costs about 46 tokens; SKILL.md has 1,171 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 1,171 words (~2,760 tokens).

“Build evaluation frameworks for agent systems”

— opening of SKILL.md by aiskillstore
name
evaluation
risk
safe
source
https://github.com/muratcankoylan/Agent-Skills-for-Context-Engineering/tree/main/skills/evaluation
date_added
2026-02-27

Read the full SKILL.md on GitHub

Files

SKILL.md and 1 other file in skills/sickn33/evaluation of aiskillstore/marketplace.

  • SKILL.md
  • skill-report.json

Open the folder on GitHubat commit 4ac52da

Used in 6 other repositories

We found 17 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 6 other GitHub owners. This page covers the copy in aiskillstore/marketplace, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluation this skillaiskillstore/marketplace4306 repos~2.8kAutomated safety check: PassNone
Context Mode Output Sandboxmksglu/context-mode26k—~4.1kAutomated safety check: PassCustom licence
Memori Long-Term MemoryMemoriLabs/Memori17k—~2kAutomated safety check: NotesCustom licence
Picoclaw Skill Creatorsipeed/picoclaw30k—~4.4kAutomated safety check: PassMIT
ccc Semantic Code Searchcocoindex-io/cocoindex-code2.7k—~938Automated safety check: PassApache-2.0
Context Mode for Antigravity CLImksglu/context-mode26k—~850Automated safety check: PassCustom licence

Similar skills

  • Context Mode Output Sandbox

    mksglu/context-mode

    Routes large command, file, API and browser output through context-mode tools so only the needed result enters the agent's context, instead of dumping it via Bash.

    26k GitHub stars~4.1k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Memori Long-Term Memory

    MemoriLabs/Memori

    Connects Claude Code to Memori Cloud for long-term memory, recalling stored context before substantive replies and saving new context afterward.

    17k GitHub stars~2k tokensUpdated 4 days ago
    Agent WorkflowsAuto-check: notes
  • Picoclaw Skill Creator

    sipeed/picoclaw

    Guidance for creating, updating and reviewing Picoclaw skills, from the SKILL.md structure to organizing bundled scripts, references and assets.

    30k GitHub stars~4.4k tokensUpdated 12 days ago
    Agent WorkflowsAuto-check passed
  • ccc Semantic Code Search

    cocoindex-io/cocoindex-code

    Semantic code search and index management with the ccc CLI: the agent initializes, indexes and queries the project by concept, filtering by language or path.

    2.7k GitHub stars~938 tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Routing rules for using context-mode MCP tools in Antigravity CLI: sandboxed code runs, file analysis, indexed search and web fetches that keep large output out of the conversation.

    26k GitHub stars~850 tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Token Optimizer

    alexgreensh/token-optimizer

    Audit a Claude Code or Codex setup for context-window waste, then fix it and measure the savings.

    2.5k GitHub stars~3.6k tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed

More from aiskillstore/marketplace

All 1,085 skills in this repo
  • Code Stats

    aiskillstore/marketplace

    Analyze codebase with tokei (fast line counts by language) and difft (semantic AST-aware diffs).

    430 GitHub starsUsed in 2 repos~697 tokens
    Auto-check: notes
  • Data Processing

    aiskillstore/marketplace

    Process JSON with jq and YAML/TOML with yq. An agent skill from aiskillstore/marketplace.

    430 GitHub starsUsed in 1 repo~720 tokens
    Auto-check: notes
  • Doc Scanner

    aiskillstore/marketplace

    Scans for project documentation files (AGENTS.md, CLAUDE.md, GEMINI.md, COPILOT.md, CURSOR.md, WARP.md, and 15+ other formats) and synthesizes guidance.

    430 GitHub starsUsed in 1 repo~644 tokens
    Auto-check: notes
  • File Search

    aiskillstore/marketplace

    Modern file and content search using fd, ripgrep (rg), and fzf.

    430 GitHub starsUsed in 1 repo~598 tokens
    Auto-check: notes
  • Find Replace

    aiskillstore/marketplace

    Modern find-and-replace using sd (simpler than sed) and batch replacement patterns.

    430 GitHub starsUsed in 1 repo~527 tokens
    Auto-check: notes
  • Investigating Codebases

    aiskillstore/marketplace

    Automatically activated when user asks how something works, wants to understand unfamiliar code, needs to explore a new codebase, or asks questions like "where is X implemented?", "how does Y…

    430 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check: notes

Categories

Questions about Evaluation

What does Evaluation do?

Build evaluation frameworks for agent systems. An agent skill from aiskillstore/marketplace. Evaluation is an agent skill from aiskillstore/marketplace. Build evaluation frameworks for agent systems.

When should I use Evaluation?

Evaluation fits situations like: testing agent performance systematically; validating context engineering choices; measuring improvements over time.

How do I install Evaluation in Claude Code?

Run `npx skills add aiskillstore/marketplace --skill evaluation -a claude-code`. Or copy the skill folder (skills/sickn33/evaluation in aiskillstore/marketplace) into .claude/skills/evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Evaluation in Codex?

Run `npx skills add aiskillstore/marketplace --skill evaluation -a codex`. Or copy the skill folder (skills/sickn33/evaluation in aiskillstore/marketplace) into .agents/skills/evaluation in your project. Codex loads it when a task matches its description.

Can I use Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aiskillstore/marketplace --skill evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluation, .gemini/skills/evaluation, .github/skills/evaluation and .opencode/skills/evaluation in your project.

What does Evaluation need to run?

SKILL.md names no scripts, command-line tools or credentials: Evaluation is instructions for the agent only. Our summary lists: Python 3.

Does Evaluation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evaluation use?

No licence was found for Evaluation or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Evaluation use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evaluation?

Skills that share tags, products or a category with Evaluation: Context Mode Output Sandbox (mksglu/context-mode, 26k stars), Memori Long-Term Memory (MemoriLabs/Memori, 17k stars), Picoclaw Skill Creator (sipeed/picoclaw, 30k stars) and ccc Semantic Code Search (cocoindex-io/cocoindex-code, 2.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluation?

aiskillstore (a GitHub organization) maintains it in aiskillstore/marketplace, which has 430 GitHub stars. The repository holds 1,085 skills in this directory. The repository was last updated on October 7, 2026.

Source: aiskillstore/marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.