Agent skill

Evaluation Anchor Checker

by WILLOSCAR in WILLOSCAR/research-units-pipeline-skills

Audit and rewrite evaluation/numeric claims to ensure they carry minimal protocol context (task + metric + constraint) and avoid underspecified model naming.

No licenceAuto-check passed

Install Evaluation Anchor Checker

skills CLI
$ npx skills add WILLOSCAR/research-units-pipeline-skills --skill evaluation-anchor-checker -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install WILLOSCAR/research-units-pipeline-skills evaluation-anchor-checker --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/WILLOSCAR/research-units-pipeline-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.codex/skills/evaluation-anchor-checker .claude/skills/evaluation-anchor-checker && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluation-anchor-checker
GitHub stars
513
Token cost
~1.4k tokens
SKILL.md length
528 words
Files
4 (incl. scripts, references, assets)
Skills in repo
107
Repo updated
First seen
Licence
None found

At a glance

Audit and rewrite evaluation/numeric claims to ensure they carry minimal protocol context (task + metric + constraint) and avoid underspecified model naming.

  • SKILL.md covers Triggers & routing, Inputs, Outputs and Recommended slot in the survey…, plus 7 more sections
  • Runs Python scripts from its folder; calls uv

What it does

Evaluation Anchor Checker is an agent skill from WILLOSCAR/research-units-pipeline-skills. Audit and rewrite evaluation/numeric claims to ensure they carry minimal protocol context (task + metric + constraint) and avoid underspecified model naming.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts, reference files and assets (for example `assets/numeric_hygiene.json`, `references/numeric_hygiene.md` and `scripts/run.py`).

The repository describes itself as: Research pipelines as semantic execution units: each skill declares inputs/outputs, acceptance criteria, and guardrails. Evidence-first methodology prevents hollow writing…

Example prompts

  • “/evaluation-anchor-checker”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit c92912a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluation Anchor Checker loads about 1.4k tokens when it runs, and up to ~1.6k if it reads all its reference files. Until then it costs about 46 tokens; SKILL.md has 528 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 528 words (~1,438 tokens).

“Purpose: fix a reviewer-magnet failure mode in agent surveys:”

— opening of SKILL.md by WILLOSCAR
name
evaluation-anchor-checker

Read the full SKILL.md on GitHub

Files

SKILL.md and 3 other files (scripts, references, assets) in .codex/skills/evaluation-anchor-checker of WILLOSCAR/research-units-pipeline-skills.

  • SKILL.md
  • assets/numeric_hygiene.json
  • references/numeric_hygiene.md
  • scripts/run.py

Open the folder on GitHubat commit c92912a

Compare with similar skills

Evaluation Anchor Checker next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluation Anchor Checker compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluation Anchor Checker this skillWILLOSCAR/research-units-pipeline-skills513—~1.4kAutomated safety check: PassNone
Claimsruvnet/ruflo74k2 repos~1.1kAutomated safety check: PassMIT
AI Claim Checkeriflytek/skillhub5.2k—~1.2kAutomated safety check: PassCC-BY-SA-4.0
Arize Evaluatorgithub/awesome-copilot40k1 repos~8.1kAutomated safety check: NotesMIT
LLM Evaluationdavila7/claude-code-templates32k12 repos~3.5kAutomated safety check: PassMIT
Agent Evaluationsickn33/agentic-awesome-skills47k1 repos~2kAutomated safety check: PassMIT

Similar skills

  • Claims

    ruvnet/ruflo

    Claims-based authorization for agents and operations. An agent skill from ruvnet/ruflo.

    74k GitHub starsUsed in 2 repos~1.1k tokens
    Backend & APIsAuto-check passed
  • AI Claim Checker

    iflytek/skillhub

    Breaks AI-generated text into checkable claims, verifies them against independent sources and labels each one, with an optional exercise for learners.

    5.2k GitHub stars~1.2k tokensUpdated today
    Research & ScienceAuto-check passed
  • Arize Evaluator

    github/awesome-copilot

    Official

    Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

    40k GitHub starsUsed in 1 repo~8.1k tokens
    AI & LLM EngineeringAuto-check: notes
  • LLM Evaluation

    davila7/claude-code-templates

    Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

    32k GitHub starsUsed in 12 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Evaluation

    sickn33/agentic-awesome-skills

    Evaluate agent behavior with versioned cases and explicit verifiers.

    47k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed
  • Evaluators

    Arize-ai/phoenix

    Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output.

    12k GitHub stars~1.7k tokensUpdated today
    EducationAuto-check passed

More from WILLOSCAR/research-units-pipeline-skills

All 107 skills in this repo
  • Appendix Table Writer

    WILLOSCAR/research-units-pipeline-skills

    Curate reader-facing survey tables for the Appendix (clean layout + high information density), using only in-scope evidence and existing citation keys.

    513 GitHub stars~1.8k tokensUpdated 4 days ago
    Auto-check passed
  • Artifact Contract Auditor

    WILLOSCAR/research-units-pipeline-skills

    Audit one research Workspace for declared Unit outputs and Pipeline target Artifacts, writing output/CONTRACTREPORT.md; use for mid-Run coverage snapshots or final delivery completeness, not deep…

    513 GitHub stars~918 tokensUpdated 4 days ago
    Auto-check passed
  • Arxiv Search

    WILLOSCAR/research-units-pipeline-skills

    Retrieve arXiv paper metadata with keyword queries or import an offline arXiv export, and save results as JSONL (papers/papersraw.jsonl).

    513 GitHub stars~1.5k tokensUpdated 4 days ago
    Auto-check passed
  • Chapter Lead Writer

    WILLOSCAR/research-units-pipeline-skills

    Write H2 chapter lead blocks (sections/S<secidlead.md) that preview the chapter's comparison lens and connect its H3 subsections, without adding new facts.

    513 GitHub stars~1.5k tokensUpdated 4 days ago
    Auto-check passed
  • Chapter Skeleton

    WILLOSCAR/research-units-pipeline-skills

    Build a retrieval-informed chapter skeleton (outline/chapterskeleton.yml) from taxonomy/core scope before stable H3 decomposition.

    513 GitHub stars~475 tokensUpdated 4 days ago
    Auto-check passed
  • Dedupe Rank

    WILLOSCAR/research-units-pipeline-skills

    A skill your agent uses when a broad paper candidate pool needs deterministic deduplication and a stable core set.

    513 GitHub stars~438 tokensUpdated 4 days ago
    Auto-check passed

Questions about Evaluation Anchor Checker

What does Evaluation Anchor Checker do?

Audit and rewrite evaluation/numeric claims to ensure they carry minimal protocol context (task + metric + constraint) and avoid underspecified model naming. Evaluation Anchor Checker is an agent skill from WILLOSCAR/research-units-pipeline-skills. Audit and rewrite evaluation/numeric claims to ensure they carry minimal protocol context (task + metric + constraint) and avoid underspecified model naming.

How do I install Evaluation Anchor Checker in Claude Code?

Run `npx skills add WILLOSCAR/research-units-pipeline-skills --skill evaluation-anchor-checker -a claude-code`. Or copy the skill folder (.codex/skills/evaluation-anchor-checker in WILLOSCAR/research-units-pipeline-skills) into .claude/skills/evaluation-anchor-checker in your project. Claude Code loads it when a task matches its description.

How do I install Evaluation Anchor Checker in Codex?

Run `npx skills add WILLOSCAR/research-units-pipeline-skills --skill evaluation-anchor-checker -a codex`. Or copy the skill folder (.codex/skills/evaluation-anchor-checker in WILLOSCAR/research-units-pipeline-skills) into .agents/skills/evaluation-anchor-checker in your project. Codex loads it when a task matches its description.

Can I use Evaluation Anchor Checker in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add WILLOSCAR/research-units-pipeline-skills --skill evaluation-anchor-checker -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluation-anchor-checker, .gemini/skills/evaluation-anchor-checker, .github/skills/evaluation-anchor-checker and .opencode/skills/evaluation-anchor-checker in your project.

What does Evaluation Anchor Checker need to run?

Going by SKILL.md and its folder, Evaluation Anchor Checker needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Evaluation Anchor Checker access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Evaluation Anchor Checker safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Evaluation Anchor Checker use?

No licence was found for Evaluation Anchor Checker or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Evaluation Anchor Checker use?

About 1.4k tokens (SKILL.md is roughly 5.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 207 tokens, read only when the agent opens those files.

What are the alternatives to Evaluation Anchor Checker?

Skills that share tags, products or a category with Evaluation Anchor Checker: Claims (ruvnet/ruflo, 74k stars), AI Claim Checker (iflytek/skillhub, 5.2k stars), Arize Evaluator (github/awesome-copilot, 40k stars) and LLM Evaluation (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluation Anchor Checker?

WILLOSCAR (a GitHub user) maintains it in WILLOSCAR/research-units-pipeline-skills, which has 513 GitHub stars. The repository holds 107 skills in this directory. The repository was last updated on October 5, 2026.

Source: WILLOSCAR/research-units-pipeline-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.