Agent skill

Evaluation Compiler Agent

by HakunaSama in HakunaSama/SkillMiner

Skill 编译 Agent(双产物:一个 skill + 该 skill 的评测任务). An agent skill from HakunaSama/SkillMiner.

No licenceAuto-check passed

Install Evaluation Compiler Agent

skills CLI
$ npx skills add HakunaSama/SkillMiner --skill evaluation-compiler-agent -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HakunaSama/SkillMiner evaluation-compiler-agent --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HakunaSama/SkillMiner.git skills-src && mkdir -p .claude/skills && cp -r skills-src/evaluation-compiler-agent-skill .claude/skills/evaluation-compiler-agent && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluation-compiler-agent
GitHub stars
100
Token cost
~1.8k tokens
SKILL.md length
452 words
Files
3 (incl. assets)
Skills in repo
3
Repo updated
First seen
Licence
None found

At a glance

Skill 编译 Agent(双产物:一个 skill + 该 skill 的评测任务). An agent skill from HakunaSama/SkillMiner.

  • Works in 12 steps: 针对同一主题簇的全部语义分析报告 → 可选辅助输入 → 双产物收敛 → …
  • SKILL.md covers 一、输入, 二、输出(两个产物), 三、你的核心目标 and 四、核心原则, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Evaluation Compiler Agent is an agent skill from HakunaSama/SkillMiner. Skill 编译 Agent(双产物:一个 skill + 该 skill 的评测任务)。 该 skill 的目标不是重新做语义发现,而是把"语义发现 Agent"针对同一个主题簇 产出的全部语义分析报告(可能来自该簇的多个批次),编译为两个配套产物: 1. 一个可执行的 skill 定义(单个 SKILL.md,多个能力维度作为其中的章节) 2. 这个 skill 对应的评测任务(用来考核"运用该 skill 做出的判断/处理"对不对, 同时也用来回归检验 skill 本身是否覆盖完整) 这两个产物是"一体两面、同步产出"的: - skill 讲"该怎么做" - 评测任务讲"怎么判断做得对不对" - 二者按"能力维度"一一对应,最后必须互相对照检查一致性 本版本强调: - 领域无关:适用于任意主题文档凝练出的语义报告 - 输入是第二个 agent 针对同一主题簇产出的全部语义分析报告(含各批次报告 + 结构化缺口清单) - 输出是一个 skill + 一个评测任务(不是多个独立 skill,不是旧的五类评分协议族) - 多个维度以"章节"形式并列在同一个 SKILL.md 内 -…

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including assets (for example `assets/evaluation_template.md` and `assets/skill_template.md`).

Example prompts

  • “语义发现 Agent”
  • “运用该 skill 做出的判断/处理”
  • “一体两面、同步产出”
  • “/evaluation-compiler-agent”

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. 针对同一主题簇的全部语义分析报告
  2. 可选辅助输入
  3. 双产物收敛
  4. 跨批次合并去重(关键)
  5. 维度内聚自检
  6. 可追溯(禁止臆造)
  7. 置信分级(关键,服务于"高质量")
  8. skill 与评测任务一一对应
  9. 覆盖完整
  10. 保留边界与不确定性
  11. :汇总全部语义分析报告
  12. :跨批次合并去重

What it can do on your machine

Read from SKILL.md and the folder at commit 56d4610. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluation Compiler Agent loads about 1.8k tokens when it runs. Until then it costs about 165 tokens; SKILL.md has 452 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~165
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 452 words (~1,819 tokens).

name
evaluation-compiler-agent

Read the full SKILL.md on GitHub

Files

SKILL.md and 2 other files (assets) in evaluation-compiler-agent-skill of HakunaSama/SkillMiner.

  • SKILL.md
  • assets/evaluation_template.md
  • assets/skill_template.md

Open the folder on GitHubat commit 56d4610

Compare with similar skills

Evaluation Compiler Agent next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluation Compiler Agent compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluation Compiler Agent this skillHakunaSama/SkillMiner100—~1.8kAutomated safety check: PassNone
Arize Evaluatorgithub/awesome-copilot40k2 repos~8.1kAutomated safety check: NotesMIT
LLM Evaluationdavila7/claude-code-templates32k13 repos~3.5kAutomated safety check: PassMIT
Agent Evaluationsickn33/agentic-awesome-skills47k1 repos~2kAutomated safety check: PassMIT
EvaluatorsArize-ai/phoenix12k—~1.7kAutomated safety check: PassCustom licence
Agent Evaluation Reportingsickn33/agentic-awesome-skills47k1 repos~2.1kAutomated safety check: PassMIT

Similar skills

  • Arize Evaluator

    github/awesome-copilot

    Official

    Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

    40k GitHub starsUsed in 2 repos~8.1k tokens
    AI & LLM EngineeringAuto-check: notes
  • LLM Evaluation

    davila7/claude-code-templates

    Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

    32k GitHub starsUsed in 13 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Evaluation

    sickn33/agentic-awesome-skills

    Evaluate agent behavior with versioned cases and explicit verifiers.

    47k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed
  • Evaluators

    Arize-ai/phoenix

    Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output.

    12k GitHub stars~1.7k tokensUpdated today
    EducationAuto-check passed
  • Agent Evaluation Reporting

    sickn33/agentic-awesome-skills

    A skill your agent uses when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.

    47k GitHub starsUsed in 1 repo~2.1k tokens
    Agent WorkflowsAuto-check passed
  • Official

    Author continuously-running online evaluations in PostHog AI observability, grounded in real failure modes you've identified.

    40k GitHub stars~6.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from HakunaSama/SkillMiner

Questions about Evaluation Compiler Agent

What does Evaluation Compiler Agent do?

Skill 编译 Agent(双产物:一个 skill + 该 skill 的评测任务). An agent skill from HakunaSama/SkillMiner. Evaluation Compiler Agent is an agent skill from HakunaSama/SkillMiner. Skill 编译 Agent(双产物:一个 skill + 该 skill 的评测任务)。 该 skill 的目标不是重新做语义发现,而是把"语义发现 Agent"针对同一个主题簇 产出的全部语义分析报告(可能来自该簇的多个批次),编译为两个配套产物: 1.

How do I install Evaluation Compiler Agent in Claude Code?

Run `npx skills add HakunaSama/SkillMiner --skill evaluation-compiler-agent -a claude-code`. Or copy the skill folder (evaluation-compiler-agent-skill in HakunaSama/SkillMiner) into .claude/skills/evaluation-compiler-agent in your project. Claude Code loads it when a task matches its description.

How do I install Evaluation Compiler Agent in Codex?

Run `npx skills add HakunaSama/SkillMiner --skill evaluation-compiler-agent -a codex`. Or copy the skill folder (evaluation-compiler-agent-skill in HakunaSama/SkillMiner) into .agents/skills/evaluation-compiler-agent in your project. Codex loads it when a task matches its description.

Can I use Evaluation Compiler Agent in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HakunaSama/SkillMiner --skill evaluation-compiler-agent -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluation-compiler-agent, .gemini/skills/evaluation-compiler-agent, .github/skills/evaluation-compiler-agent and .opencode/skills/evaluation-compiler-agent in your project.

What does Evaluation Compiler Agent need to run?

SKILL.md names no scripts, command-line tools or credentials: Evaluation Compiler Agent is instructions for the agent only.

Does Evaluation Compiler Agent access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evaluation Compiler Agent safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evaluation Compiler Agent use?

No licence was found for Evaluation Compiler Agent or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Evaluation Compiler Agent use?

About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evaluation Compiler Agent?

Skills that share tags, products or a category with Evaluation Compiler Agent: Arize Evaluator (github/awesome-copilot, 40k stars), LLM Evaluation (davila7/claude-code-templates, 32k stars), Agent Evaluation (sickn33/agentic-awesome-skills, 47k stars) and Evaluators (Arize-ai/phoenix, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluation Compiler Agent?

HakunaSama (a GitHub user) maintains it in HakunaSama/SkillMiner, which has 100 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on July 26, 2026.

Source: HakunaSama/SkillMiner on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.