Agent skill

Eval

by jh941213 in jh941213/my-cc-harness

코드 산출물을 4축(기능/품질/독창성/보안)으로 평가하고 점수 산출. An agent skill from jh941213/my-cc-harness.

No licenceAuto-check: notes

Install Eval

skills CLI
$ npx skills add jh941213/my-cc-harness --skill eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jh941213/my-cc-harness eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jh941213/my-cc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval .claude/skills/eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval
GitHub stars
126
Token cost
~304 tokens
SKILL.md length
97 words
Files
1
Skills in repo
66
Repo updated
First seen
Licence
None found

At a glance

코드 산출물을 4축(기능/품질/독창성/보안)으로 평가하고 점수 산출. An agent skill from jh941213/my-cc-harness.

  • Works in 3 steps: Evaluator 에이전트 스폰 → 결과 확인 → CONDITIONAL/FAIL 시
  • SKILL.md covers 실행 프로세스 and pass@k 멱등성 테스트 (선택)
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Eval is an agent skill from jh941213/my-cc-harness. 코드 산출물을 4축(기능/품질/독창성/보안)으로 평가하고 점수 산출. Evaluator 에이전트를 스폰하여 독립 평가 실행. Triggers on: eval, 평가, 품질 점수, 코드 평가, quality score. NOT for: 코드 작성, 구현, 리뷰.

Its SKILL.md is about 300 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

Example prompts

  • “/eval”

Requirements

  • Pre-approved tools (allowed-tools): Read, Bash, Grep, Glob, Agent

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Evaluator 에이전트 스폰
  2. 결과 확인
  3. CONDITIONAL/FAIL 시

What it can do on your machine

Read from SKILL.md and the folder at commit e9210e2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash
    • Grep
    • Glob
    • Agent

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval loads about 304 tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 97 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~38
When it runs · the whole SKILL.md, loaded when a task matches
~304

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash, Grep, Glob, Agent

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 97 words (~304 tokens).

“Generator(구현자)와 분리된 Evaluator 에이전트를 스폰하여 산출물을 독립 평가합니다.”

— opening of SKILL.md by jh941213
name
eval
allowed-tools
Read, Bash, Grep, Glob, Agent
user-invocable
true
disable-model-invocation
false

Read the full SKILL.md on GitHub

Files

Just SKILL.md in skills/eval of jh941213/my-cc-harness.

Open the folder on GitHubat commit e9210e2

Compare with similar skills

Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval this skilljh941213/my-cc-harness126—~304Automated safety check: NotesNone
Evalalirezarezvani/claude-skills28k1 repos~618Automated safety check: PassMIT
Evalhashgraph-online/awesome-codex-plugins1.2k—~2kAutomated safety check: PassApache-2.0
Evalmikeyobrien/rho371—~9.7kAutomated safety check: PassMIT
Evalai-analyst-lab/ai-analyst304—~2.4kAutomated safety check: PassMIT
Evalagentevals-dev/agentevals162—~904Automated safety check: PassApache-2.0

Similar skills

  • Eval

    alirezarezvani/claude-skills

    Evaluate and rank agent results by metric or LLM judge for an AgentHub session.

    28k GitHub starsUsed in 1 repo~618 tokens
    AI & LLM EngineeringAuto-check passed
  • Eval

    hashgraph-online/awesome-codex-plugins

    Quality and performance evaluation with baseline comparison.

    1.2k GitHub stars~2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Eval

    mikeyobrien/rho

    Plan and run conversational AI agent evaluations with test generation and analysis.

    371 GitHub stars~9.7k tokensUpdated 6 days ago
    Testing & QAAuto-check passed
  • Eval

    ai-analyst-lab/ai-analyst

    Evaluate a named AI Analyst configuration across a frozen suite.

    304 GitHub stars~2.4k tokensUpdated 6 days ago
    Auto-check passed
  • Eval

    agentevals-dev/agentevals

    Evaluate and score agent behavior against a golden reference.

    162 GitHub stars~904 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    affaan-m/ECC

    Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k…

    274k GitHub stars~2.2k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from jh941213/my-cc-harness

All 66 skills in this repo
  • API Design Principles

    jh941213/my-cc-harness

    REST 및 GraphQL API 설계 원칙 가이드. An agent skill from jh941213/my-cc-harness.

    126 GitHub starsUsed in 19 repos~3.4k tokens
    Auto-check passed
  • Tailwind Design System

    jh941213/my-cc-harness

    Build scalable design systems with Tailwind CSS, design tokens, component libraries, and responsive patterns.

    126 GitHub starsUsed in 9 repos~4.7k tokens
    Auto-check passed
  • Docs Architecture

    jh941213/my-cc-harness

    Generate/update architecture docs — ARCHITECTURE.md (codemap), architecture diagrams (C4 mermaid), ADRs (MADR), data model ERD.

    126 GitHub stars~923 tokensUpdated 2 mo ago
    Auto-check: notes
  • Auto Memory

    jh941213/my-cc-harness

    A skill your agent uses when starting substantial work in a repo (implementation, fixes, deploys, debugging), when first exploring a new repo, or when finishing work that produced reusable knowledge.

    126 GitHub stars~890 tokensUpdated 2 mo ago
    Auto-check passed
  • Autodev

    jh941213/my-cc-harness

    Ralph Loop based autonomous development loop. An agent skill from jh941213/my-cc-harness.

    126 GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check: notes
  • Async Python Patterns

    jh941213/my-cc-harness

    Python asyncio 및 async/await 패턴 가이드. An agent skill from jh941213/my-cc-harness.

    126 GitHub starsUsed in 13 repos~4.7k tokens
    Auto-check passed

Questions about Eval

What does Eval do?

코드 산출물을 4축(기능/품질/독창성/보안)으로 평가하고 점수 산출. An agent skill from jh941213/my-cc-harness. Eval is an agent skill from jh941213/my-cc-harness. 코드 산출물을 4축(기능/품질/독창성/보안)으로 평가하고 점수 산출.

How do I install Eval in Claude Code?

Run `npx skills add jh941213/my-cc-harness --skill eval -a claude-code`. Or copy the skill folder (skills/eval in jh941213/my-cc-harness) into .claude/skills/eval in your project. Claude Code loads it when a task matches its description.

How do I install Eval in Codex?

Run `npx skills add jh941213/my-cc-harness --skill eval -a codex`. Or copy the skill folder (skills/eval in jh941213/my-cc-harness) into .agents/skills/eval in your project. Codex loads it when a task matches its description.

Can I use Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jh941213/my-cc-harness --skill eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval, .gemini/skills/eval, .github/skills/eval and .opencode/skills/eval in your project.

What does Eval need to run?

SKILL.md names no scripts, command-line tools or credentials: Eval is instructions for the agent only. Its frontmatter pre-approves these tools: Read, Bash, Grep, Glob, Agent.

Does Eval access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Eval use?

No licence was found for Eval or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Eval use?

About 304 tokens (SKILL.md is roughly 1.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval?

Skills that share tags, products or a category with Eval: Eval (alirezarezvani/claude-skills, 28k stars), Eval (hashgraph-online/awesome-codex-plugins, 1.2k stars), Eval (mikeyobrien/rho, 371 stars) and Eval (ai-analyst-lab/ai-analyst, 304 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval?

jh941213 (a GitHub user) maintains it in jh941213/my-cc-harness, which has 126 GitHub stars. The repository holds 66 skills in this directory. The repository was last updated on August 3, 2026.

Source: jh941213/my-cc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.