Agent skill

Create Evaluations

by GAIK-project in GAIK-project/gaik-toolkit

Creates evaluation documentation for a GAIK component in both locations: evaluationlayer/evalmethods/{component}eval/ (README + optional script stubs) and…

MITAuto-check passedDocuments & Office

Install Create Evaluations

skills CLI
$ npx skills add GAIK-project/gaik-toolkit --skill create-evaluations -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GAIK-project/gaik-toolkit create-evaluations --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GAIK-project/gaik-toolkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/create-evaluations .claude/skills/create-evaluations && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
create-evaluations
GitHub stars
100
Token cost
~2.3k tokens
SKILL.md length
840 words
Files
3 (incl. references)
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Creates evaluation documentation for a GAIK component in both locations: evaluationlayer/evalmethods/{component}eval/ (README + optional script stubs) and…

  • Works in 5 steps: Context Parsing → Outline → Plan Review (never skip) → …
  • Tasks that involve Markdown
  • SKILL.md covers What this skill creates, Workflow, Hard Rules and References
  • Calls pnpm and git

What it does

Create Evaluations is an agent skill from GAIK-project/gaik-toolkit. Creates evaluation documentation for a GAIK component in both locations: evaluationlayer/evalmethods/{component}eval/ (README + optional script stubs) and guidancelayer/website/content/docs/evaluation-layer/{component}-eval.mdx (full user-facing page). Also updates index.mdx and meta.json to register the new page.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/mdx-template.md` and `references/readme-template.md`).

It sits in Documents & Office, covering Markdown and Technical documentation. The repository describes itself as: Python toolkit providing reusable AI/ML utilities: schema extraction, structured outputs, and production-ready components. The licence is MIT.

When your agent uses it

  • Tasks that involve Markdown
  • Tasks that involve Technical documentation

Example prompts

  • “Use the create-evaluations skill to create evaluation documentation for a GAIK component in both locations…”
  • “/create-evaluations”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Context Parsing
  2. Outline
  3. Plan Review (never skip)
  4. Generate Content
  5. Verification Summary

What it can do on your machine

Read from SKILL.md and the folder at commit e516ece. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pnpm
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pnpm and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Create Evaluations loads about 2.3k tokens when it runs, and up to ~6.1k if it reads all its reference files. Until then it costs about 85 tokens; SKILL.md has 840 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~85
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GAIK-project/gaik-toolkit at commit e516ece, republished under its MIT licence (© GAIK-project). 840 words, ~2,315 tokens.

Download SKILL.mdSave it as .claude/skills/create-evaluations/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
create-evaluations
description
Creates evaluation documentation for a GAIK component in both locations: evaluation_layer/eval_methods/{component}_eval/ (README + optional script stubs) and guidance_layer/website/content/docs/evaluation-layer/{component}-eval.mdx (full user-facing page). Also updates index.mdx and meta.json to register the new page.
argument-hint
[component-name] [evaluation context: metrics, results, methodology, error analysis, use cases]

Create Evaluations

Generates evaluation documentation for a GAIK component in two mirrored locations following the established pattern of transcription_eval and extraction_eval.

Requires user-provided context — the user must supply the evaluation details (metrics, results, methodology, errors, use cases) as part of their message or as attached content. This skill does not run evaluations itself; it documents them.


What this skill creates

FileAction
evaluation_layer/eval_methods/{component}_eval/README.mdCreate (or promote stub to full)
evaluation_layer/eval_methods/{component}_eval/requirements.txtCreate if scripts are requested
evaluation_layer/eval_methods/{component}_eval/*.pyCreate script stubs if requested
guidance_layer/website/content/docs/evaluation-layer/{component}-eval.mdxCreate (or promote stub to full)
guidance_layer/website/content/docs/evaluation-layer/index.mdxUpdate: add new entry, remove Coming Soon stub
guidance_layer/website/content/docs/evaluation-layer/meta.jsonUpdate: insert page before remaining stubs

Workflow

Phase 1 — Context Parsing
  1. Read the component source at implementation_layer/src/gaik/software_components/{component}/ to understand what the component does, its inputs, outputs, and configuration options.
  2. If a stub README exists at evaluation_layer/eval_methods/{component}_eval/README.md, read it.
  3. If a stub MDX exists at evaluation-layer/{component}-eval.mdx, read it.
  4. Map the user-provided context against the required sections for both output files (see references/readme-template.md and references/mdx-template.md for the full section lists).
  5. For each required section where no context was provided, ask the user:

    "No content was provided for [Section Name] (e.g. benchmarking results / error taxonomy / CLI usage). Do you want to supply it now, or continue and mark it N/A?"

    • If the user supplies content → incorporate it before generating files.
    • If the user skips → that section gets _N/A — to be completed._ as its body.
  6. Ask once: "Should I also generate Python evaluation script stubs, or README + website content only?"
Phase 2 — Outline

Present a two-column outline to the user:

README sections:          MDX sections:
1. Evaluation Metrics     1. The Problem
2. Tools / Code           2. How We Evaluate
3. Results                3. Benchmarking Results
4. Error Classification   4. Error Classification
5. Improvement Strategies 5. Real-World Applications
6. Reproduction / Usage   6. Quality Considerations
7. Integration snippet    7. Getting Started
8. Installation & Setup
9. Related Resources

Populated: [list sections that have content]
N/A:       [list sections that will be marked N/A]
Scripts:   [yes / no]
Phase 3 — Plan Review (never skip)

Present the outline and wait for explicit approval before writing any files. Adjust based on feedback. Do not start Phase 4 until the user confirms.

Phase 4 — Generate Content

Run all sub-steps in order. Do not commit.

4a — Implementation Layer README

Follow the structure in references/readme-template.md. Key rules:

  • Use ## for numbered top-level sections (## 1. Evaluation Metrics, ## 2. Evaluation Tools / Code, etc.)
  • Use ### for subsections (### 1.1 List of Metrics, ### 3.2 Performance Comparison Table, etc.)
  • Each metric gets: Definition, Formula (code block), Components, Business interpretation, Reference values table
  • Results section has a Markdown table with model names as rows and metric columns
  • Improvement strategies section has a mapping table: | Performance issue | Improvement strategy |
  • Integration section has a code snippet using the actual component's Python API
  • Sections with no user-provided content → _N/A — to be completed._

4b — Python Script Stubs (only if user said yes in Phase 1)

Create one stub per logical evaluation step following the pattern below. Add a requirements.txt.

python
"""
{component}_eval — {purpose of this script}.

Usage:
    python {script_name}.py <arg1> <arg2>
"""

from __future__ import annotations

import argparse
from pathlib import Path


def main(arg1: str, arg2: str) -> None:
    # TODO: implement evaluation logic
    pass


if __name__ == "__main__":
    parser = argparse.ArgumentParser(description=__doc__)
    parser.add_argument("arg1", help="...")
    parser.add_argument("arg2", help="...")
    args = parser.parse_args()
    main(args.arg1, args.arg2)

4c — Website MDX Page

Follow the structure in references/mdx-template.md. Key rules:

  • Frontmatter: title (display name) and description (one sentence, matches index entry)
  • Section order: The Problem → How We Evaluate → Benchmarking Results → Error Classification → Real-World Applications → Quality Considerations → Getting Started
  • Tone: business-first (why it matters), then technical detail
  • Challenges in "The Problem" → bullet list with bold lead term
  • Results → Markdown table + Key Findings bullet list below it
  • Error categories use bold Problem / Impact sub-headings (see extraction-eval pattern)
  • Quality Considerations: bold question/statement + one-sentence explanation per item
  • Getting Started: numbered list of 4–6 steps, ending with GitHub repo link
  • Do not use <Callout type="warn"> — that is only for Coming Soon stubs
  • N/A sections: use _N/A — to be completed._
Show full SKILL.md (301 more words)Show less

4d — Update index.mdx

In the Output Evaluation Methods section, under the Available Methods list, append the new entry after the last existing ### … Evaluation entry. If any Coming Soon <Callout type="warn"> stubs are present, insert the entry immediately before the first one instead (before its preceding ---):

markdown
### {Component Display Name} Evaluation

{One-sentence description matching the MDX `description` frontmatter.}

[View {Component Display Name} Evaluation →](/evaluation-layer/{component}-eval)

---

If a Coming Soon stub for this component already exists (### {Name} + Callout), replace that entire stub block with the new entry.

4e — Update meta.json

Insert "{component}-eval" in the pages array within the Output Evaluation Methods group — after the "---Output Evaluation Methods---" separator and after the last existing method slug (before any remaining stub slug, if present). The page key must match the .mdx filename exactly (minus .mdx).

Phase 5 — Verification Summary

After writing all files, print:

Files created / modified:
  ✓ evaluation_layer/eval_methods/{component}_eval/README.md  (N lines)
  ✓ guidance_layer/website/content/docs/evaluation-layer/{component}-eval.mdx  (N lines)
  ✓ guidance_layer/website/content/docs/evaluation-layer/index.mdx  (updated)
  ✓ guidance_layer/website/content/docs/evaluation-layer/meta.json  (updated)
  [✓ script stubs if generated]

Sections marked N/A: [list or "none"]

To verify:
  cd guidance_layer/website && pnpm dev
  → /evaluation-layer/{component}-eval  (confirm page renders)
  → /evaluation-layer               (confirm index entry and sidebar position)

Hard Rules

  • Never skip Phase 3. Always wait for explicit approval before writing files.
  • Never overwrite a full README or MDX (i.e. one that already has real content, not just a stub) without the user explicitly confirming they want to replace it.
  • meta.json page key must exactly match the .mdx filename (minus .mdx). A mismatch silently breaks sidebar navigation.
  • Full pages must not use <Callout type="warn"> — reserved for Coming Soon stubs only.
  • Do not commit. Leave all changes staged-but-uncommitted so the user can review with git diff.
  • Script stubs are scaffolds, not implementations. Never fabricate evaluation logic or invent metric results.
  • Do not invent results. If no benchmarking data was provided and the user said N/A, leave the results table as a placeholder, not made-up numbers.

References

  • references/readme-template.md — canonical implementation README structure with all section headings, subsection numbering, and formatting conventions
  • references/mdx-template.md — canonical website MDX structure with frontmatter, section order, MDX component rules, and link conventions

Pattern references (read these when in doubt):

  • evaluation_layer/eval_methods/transcription_eval/README.md — most complete README example
  • guidance_layer/website/content/docs/evaluation-layer/transcription-eval.mdx — most complete MDX example
  • guidance_layer/website/content/docs/evaluation-layer/extraction-eval.mdx — second MDX example (simpler results section)

© GAIK-project, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in .claude/skills/create-evaluations of GAIK-project/gaik-toolkit.

  • SKILL.md
  • references/mdx-template.md
  • references/readme-template.md

Open the folder on GitHubat commit e516ece

Compare with similar skills

Create Evaluations next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Create Evaluations compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Create Evaluations this skillGAIK-project/gaik-toolkit100—~2.3kAutomated safety check: PassMIT
Review Docsvideojs/video.js40k—~413Automated safety check: PassCustom licence
Rstack Docs Writerweb-infra-dev/rstest505—~361Automated safety check: PassMIT
Markdown Document StructurerArabelaTso/Skills-4-SE253—~2.2kAutomated safety check: PassApache-2.0
WooCommerce Markdown Guidelineswoocommerce/woocommerce11k1 repos~1.7kAutomated safety check: PassCustom licence
Docs Conventionsflet-dev/flet17k—~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • Review Docs

    videojs/video.js

    Review Video.js documentation without editing it. An agent skill from videojs/video.js.

    40k GitHub stars~413 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Rstack Docs Writer

    web-infra-dev/rstest

    Write or revise Markdown and MDX documentation, including READMEs, guides, and Rspress-based docs.

    505 GitHub stars~361 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Markdown Document Structurer

    ArabelaTso/Skills-4-SE

    Reorganizes markdown documents into well-structured, consistent format while preserving content and improving readability.

    253 GitHub stars~2.2k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • WooCommerce Markdown Guidelines

    woocommerce/woocommerce

    Rules for writing and editing markdown in the WooCommerce repository, with the project's markdownlint settings for headings, lists and code blocks.

    11k GitHub starsUsed in 1 repo~1.7k tokens
    DevelopmentAuto-check passed
  • Docs Conventions

    flet-dev/flet

    A skill your agent uses when writing or reviewing Flet documentation, including Python docstrings (Google style, reST roles, admonitions), Markdown docs (cross-references, images, code examples)…

    17k GitHub stars~1.6k tokensUpdated today
    DevelopmentAuto-check passed
  • Article Exporter

    actionbook/actionbook

    Export any web article to a local Obsidian-ready Markdown directory.

    1.6k GitHub stars~3.1k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed

More from GAIK-project/gaik-toolkit

All 15 skills in this repo
  • Brief To Slides

    GAIK-project/gaik-toolkit

    Builds a visual, editable PowerPoint (.pptx) deck with speaker-ready notes, exact timing, citations and a layout-checked design from a topic, an audience and a length, using only the user's own…

    100 GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed
  • Gaik Toolkit

    GAIK-project/gaik-toolkit

    GAIK toolkit overview and reference. An agent skill from GAIK-project/gaik-toolkit.

    100 GitHub stars~5.7k tokensUpdated yesterday
    Auto-check passed
  • Extracting Structured Data

    GAIK-project/gaik-toolkit

    Extracts structured data — fields, tables, line items — out of documents into a validated schema using the gaik toolkit, and designs schemas that stay inside provider limits and produce checkable…

    100 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Parsing Documents

    GAIK-project/gaik-toolkit

    Converts PDFs, scans, and Word documents into text or markdown with the gaik toolkit's parsers, choosing the parser that will not silently destroy the structure the downstream task depends on.

    100 GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed
  • Searching Documents

    GAIK-project/gaik-toolkit

    Builds and debugs retrieval with the gaik toolkit — PgVectorStore, Ranker, FinnishTextProcessor, RelevanceGate — as hybrid search: pgvector similarity plus Postgres full-text, fused by rank, and the…

    100 GitHub stars~4.2k tokensUpdated yesterday
    Auto-check passed
  • Construction Diary Creation

    GAIK-project/gaik-toolkit

    Extracts structured data from Finnish construction site daily diary audio recordings (Työmaapäiväkirja) and creates a formatted Word document with extracted fields.

    100 GitHub stars~3.6k tokensUpdated yesterday
    Auto-check passed

Questions about Create Evaluations

What does Create Evaluations do?

Creates evaluation documentation for a GAIK component in both locations: evaluationlayer/evalmethods/{component}eval/ (README + optional script stubs) and…. Create Evaluations is an agent skill from GAIK-project/gaik-toolkit.mdx (full user-facing page).

When should I use Create Evaluations?

Create Evaluations fits situations like: tasks that involve Markdown; tasks that involve Technical documentation.

How do I install Create Evaluations in Claude Code?

Run `npx skills add GAIK-project/gaik-toolkit --skill create-evaluations -a claude-code`. Or copy the skill folder (.claude/skills/create-evaluations in GAIK-project/gaik-toolkit) into .claude/skills/create-evaluations in your project. Claude Code loads it when a task matches its description.

How do I install Create Evaluations in Codex?

Run `npx skills add GAIK-project/gaik-toolkit --skill create-evaluations -a codex`. Or copy the skill folder (.claude/skills/create-evaluations in GAIK-project/gaik-toolkit) into .agents/skills/create-evaluations in your project. Codex loads it when a task matches its description.

Can I use Create Evaluations in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GAIK-project/gaik-toolkit --skill create-evaluations -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/create-evaluations, .gemini/skills/create-evaluations, .github/skills/create-evaluations and .opencode/skills/create-evaluations in your project.

What does Create Evaluations need to run?

Going by SKILL.md and its folder, Create Evaluations needs the command-line tools its instructions call (pnpm and git). Our summary lists: Python 3.

Does Create Evaluations access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Create Evaluations safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Create Evaluations use?

Create Evaluations is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Create Evaluations use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.8k tokens, read only when the agent opens those files.

What are the alternatives to Create Evaluations?

Skills that share tags, products or a category with Create Evaluations: Review Docs (videojs/video.js, 40k stars), Rstack Docs Writer (web-infra-dev/rstest, 505 stars), Markdown Document Structurer (ArabelaTso/Skills-4-SE, 253 stars) and WooCommerce Markdown Guidelines (woocommerce/woocommerce, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Create Evaluations?

GAIK-project (a GitHub organization) maintains it in GAIK-project/gaik-toolkit, which has 100 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 7, 2026.

Source: GAIK-project/gaik-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.