Jd Gap Analysis
starkyru/learn-ai
Analyze a job description (pasted text OR a URL) and find the AI/ML/GenAI topics it requires that this learn-ai course does NOT yet cover.
Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines.
$ npx skills add wshobson/agents --skill eval-harness-first -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install wshobson/agents eval-harness-first --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/llm-finetuning/skills/eval-harness-first .claude/skills/eval-harness-first && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-harness-first" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/eval-harness-first into .claude/skills/eval-harness-first/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-harness-first", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/eval-harness-firstType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add wshobson/agents --skill eval-harness-first -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install wshobson/agents eval-harness-first --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/llm-finetuning/skills/eval-harness-first .agents/skills/eval-harness-first && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-harness-first" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/eval-harness-first into .agents/skills/eval-harness-first/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-harness-first", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill eval-harness-first -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install wshobson/agents eval-harness-first --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/llm-finetuning/skills/eval-harness-first .cursor/skills/eval-harness-first && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-harness-first" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/eval-harness-first into .cursor/skills/eval-harness-first/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-harness-first", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/wshobson/agents.git --path plugins/llm-finetuning/skills/eval-harness-first--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add wshobson/agents --skill eval-harness-first -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install wshobson/agents eval-harness-first --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/llm-finetuning/skills/eval-harness-first .gemini/skills/eval-harness-first && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-harness-first" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/eval-harness-first into .gemini/skills/eval-harness-first/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-harness-first", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install wshobson/agents eval-harness-firstInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add wshobson/agents --skill eval-harness-first -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/llm-finetuning/skills/eval-harness-first .github/skills/eval-harness-first && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-harness-first" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/eval-harness-first into .github/skills/eval-harness-first/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-harness-first", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill eval-harness-first -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install wshobson/agents eval-harness-first --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/llm-finetuning/skills/eval-harness-first .opencode/skills/eval-harness-first && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-harness-first" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/eval-harness-first into .opencode/skills/eval-harness-first/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-harness-first", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-harness-firstBuild the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines.
Eval Harness First is an agent skill from wshobson/agents. Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines. Use when starting a fine-tuning effort, when converting traces into an eval set, or when calibrating a judge against human labels.
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/grader-templates.md` and `references/judge-calibration.md`).
It sits in AI & LLM Engineering, covering LLM evaluation, Fine-tuning and Performance reviews. The repository describes itself as: Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi. The licence is MIT.
8 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval Harness First loads about 2k tokens when it runs, and up to ~6.2k if it reads all its reference files. Until then it costs about 74 tokens; SKILL.md has 938 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 938 words, ~1,958 tokens.
.claude/skills/eval-harness-first/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.The Phase 0 gate for the whole plugin:
finetuning-method-selection and every downstream
skill assume this harness exists before a training
config gets written. The harness is not a run-end
side artifact — it is the data-curation engine. The
same labeled traces that build the goldens feed
training data, minus an explicit holdout.
Input: production/agent traces if they exist, or
a task spec if they don't, plus labelers willing to
grade ≥100 examples.
Output format: the eval/ directory below —
goldens, graders, drift suite, and the base-model
baseline that later phases gate on.
No eval harness, no fine-tune. Skip to a training config and there is nothing to measure against, nothing to catch regressions, and no labeled data to train on. The flywheel:
eval/goldens.jsonl
ID stays excluded from training data by ID.Steps 2–4 build the harness; steps 5–8 are why it must exist first — it is both the training data source and the checkpoint's exit gate.
eval/goldens.jsonl, diff it in review, tag it per
release. It doubles as the CI regression suite.One grader per failure bucket from error analysis — not one for the whole eval set. A single blended score hides which bucket regressed.
references/grader-templates.md.Any bucket routed to an LLM-judge needs calibration before its verdicts count for anything beyond exploration — a hard prerequisite, not a nice-to-have. N/A when no bucket routes to a judge — an all-deterministic harness has nothing to calibrate; state that rather than leaving this section unaddressed.
references/judge-calibration.md.Before Phase 1 (method selection) starts, run the full harness — goldens plus the capability-drift suite — against the unmodified base model. This is the number every later checkpoint gets compared against.
eval/baseline-<model>.json is the gate token. No
baseline file, no comparison basis for
checkpoint-promotion — a checkpoint that "looks
better" against nothing measured isn't a finding.
eval/
├── goldens.jsonl # labeled traces + synthetic goldens, versioned
├── graders/ # one module per failure bucket
│ ├── schema_compliance.py
│ ├── exact_match.py
│ └── rubric_judge.py
├── drift-suite.yaml # frozen benchmarks + 200-500 domain-adjacent items
└── baseline-<model>.json # gate token: harness + drift suite vs the base model
runs/
└── <run-id>/
└── results.json # per-run harness output, one per checkpointeval/ persists across runs and lives outside
runs/ — the fixed measuring stick, not a run
artifact. runs/ is disposable; eval/ is not.
Never let a run script write into eval/. Canonical
location: every per-trace results.json — the
Phase 0 baseline included — lives at
runs/<run-id>/results.json, never under
eval/runs/...; an instruction requesting the
latter is wrong, not this contract.
Before finetuning-method-selection, confirm:
eval/goldens.jsonl committed and versioned.eval/drift-suite.yaml frozen.eval/baseline-<model>.json written.Missing any of the six (or its stated N/A)? Not
Phase 0 complete — /finetune checks the baseline
file before a run.
General-purpose evaluation guidance (dashboards, A/B
testing, non-fine-tuning harnesses) lives in the
llm-application-dev plugin's llm-evaluation
skill — this skill covers only the fine-tuning
coupling: goldens that double as training data, and
the baseline that gates a checkpoint.
finetuning-method-selection — routes here first.dataset-curation — formats these traces into
training rows.trace-to-training-data — turns graded traces into
training examples.checkpoint-promotion — consumes
baseline-<model>.json, re-runs this harness on
each candidate checkpoint.references/grader-templates.md — runnable grader
examples per shape, plus a drift-suite.yaml
example and MMLU logprob-scoring note.references/judge-calibration.md — the
calibration protocol, including the all-
deterministic N/A path.© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in plugins/llm-finetuning/skills/eval-harness-first of wshobson/agents.
Open the folder on GitHubat commit 46891e7
Eval Harness First next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval Harness First this skillwshobson/agents | 40k | — | ~2k | Automated safety check: Pass | MIT | |
| Jd Gap Analysisstarkyru/learn-ai | 107 | — | ~1.9k | Automated safety check: Pass | MIT | |
| ML Training Run VerifierLeeroo-AI/superml | 195 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | |
| Advanced Evaluationguanyang/open-agent-hub | 977 | 2 repos | ~4.2k | Automated safety check: Pass | MIT | |
| Audit Sft Data Qualitytokenbender/agent-guides | 367 | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | |
| Fine-Tuning ExpertJeffallan/claude-skills | 12k | — | ~1.7k | Automated safety check: Pass | MIT |
starkyru/learn-ai
Analyze a job description (pasted text OR a URL) and find the AI/ML/GenAI topics it requires that this learn-ai course does NOT yet cover.
Leeroo-AI/superml
Checks training code, configs and math against documented framework behavior before an expensive run, citing a knowledge base or official docs for every claim.
guanyang/open-agent-hub
This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated…
tokenbender/agent-guides
Audit supervised fine-tuning datasets against the behavior and task they are meant to teach.
Jeffallan/claude-skills
Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.
maziyarpanahi/openmed
Evaluate an OpenMed de-identification or clinical NER model against the leakage-first release gates G1a through G8, which gate releases on residual PHI leakage rather than on F1.
wshobson/agents
Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.
wshobson/agents
Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.
wshobson/agents
Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.
wshobson/agents
Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.
wshobson/agents
Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks.
wshobson/agents
Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.
Categories
Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines. Eval Harness First is an agent skill from wshobson/agents. Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines.
Eval Harness First fits situations like: starting a fine-tuning effort; converting traces into an eval set; calibrating a judge against human labels.
Run `npx skills add wshobson/agents --skill eval-harness-first -a claude-code`. Or copy the skill folder (plugins/llm-finetuning/skills/eval-harness-first in wshobson/agents) into .claude/skills/eval-harness-first in your project. Claude Code loads it when a task matches its description.
Run `npx skills add wshobson/agents --skill eval-harness-first -a codex`. Or copy the skill folder (plugins/llm-finetuning/skills/eval-harness-first in wshobson/agents) into .agents/skills/eval-harness-first in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill eval-harness-first -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-harness-first, .gemini/skills/eval-harness-first, .github/skills/eval-harness-first and .opencode/skills/eval-harness-first in your project.
SKILL.md names no scripts, command-line tools or credentials: Eval Harness First is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Eval Harness First is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Eval Harness First: Jd Gap Analysis (starkyru/learn-ai, 107 stars), ML Training Run Verifier (Leeroo-AI/superml, 195 stars), Advanced Evaluation (guanyang/open-agent-hub, 977 stars) and Audit Sft Data Quality (tokenbender/agent-guides, 367 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,305 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.
Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.