Eval Creator CI
pskoett/pskoett-ai-skills
[Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows).
Set up CIAgent regression testing for the AI agent in this repo — write a runner, record golden baselines, generate a test spec, and verify it.
$ npx skills add davepoon/buildwithclaude --skill onboard -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install davepoon/buildwithclaude onboard --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/davepoon/buildwithclaude.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ciagent/skills/onboard .claude/skills/onboard && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "onboard" agent skill from https://github.com/davepoon/buildwithclaude/tree/main/plugins/ciagent/skills/onboard into .claude/skills/onboard/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "onboard", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/davepoon/buildwithclaude/tree/main/plugins/ciagent/skills/onboardType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add davepoon/buildwithclaude --skill onboard -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install davepoon/buildwithclaude onboard --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/davepoon/buildwithclaude.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/ciagent/skills/onboard .agents/skills/onboard && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "onboard" agent skill from https://github.com/davepoon/buildwithclaude/tree/main/plugins/ciagent/skills/onboard into .agents/skills/onboard/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "onboard", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add davepoon/buildwithclaude --skill onboard -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install davepoon/buildwithclaude onboard --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/davepoon/buildwithclaude.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/ciagent/skills/onboard .cursor/skills/onboard && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "onboard" agent skill from https://github.com/davepoon/buildwithclaude/tree/main/plugins/ciagent/skills/onboard into .cursor/skills/onboard/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "onboard", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/davepoon/buildwithclaude.git --path plugins/ciagent/skills/onboard--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add davepoon/buildwithclaude --skill onboard -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install davepoon/buildwithclaude onboard --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/davepoon/buildwithclaude.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/ciagent/skills/onboard .gemini/skills/onboard && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "onboard" agent skill from https://github.com/davepoon/buildwithclaude/tree/main/plugins/ciagent/skills/onboard into .gemini/skills/onboard/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "onboard", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install davepoon/buildwithclaude onboardInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add davepoon/buildwithclaude --skill onboard -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/davepoon/buildwithclaude.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/ciagent/skills/onboard .github/skills/onboard && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "onboard" agent skill from https://github.com/davepoon/buildwithclaude/tree/main/plugins/ciagent/skills/onboard into .github/skills/onboard/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "onboard", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add davepoon/buildwithclaude --skill onboard -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install davepoon/buildwithclaude onboard --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/davepoon/buildwithclaude.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/ciagent/skills/onboard .opencode/skills/onboard && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "onboard" agent skill from https://github.com/davepoon/buildwithclaude/tree/main/plugins/ciagent/skills/onboard into .opencode/skills/onboard/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "onboard", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
onboardSet up CIAgent regression testing for the AI agent in this repo — write a runner, record golden baselines, generate a test spec, and verify it.
Onboard is an agent skill from davepoon/buildwithclaude. Set up CIAgent regression testing for the AI agent in this repo — write a runner, record golden baselines, generate a test spec, and verify it. Use when the user asks to add tests, evals, or regression testing for their AI agent, or to set up CIAgent.
Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering QA and bug reports and LLM evaluation. The repository describes itself as: A single hub to find Claude Skills, Agents, Commands, Hooks, Plugins, and Marketplace collections to extend Claude Code, Claude Desktop, Agent SDK and OpenClaw. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 616deb5. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
Bash(ciagent *)Bash(pip install *)Bash(python -c *)ReadGrepGlobWriteEditFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pippythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Onboard loads about 1.3k tokens when it runs. Until then it costs about 65 tokens; SKILL.md has 638 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from davepoon/buildwithclaude at commit 616deb5, republished under its MIT licence (© davepoon). 638 words, ~1,334 tokens.
.claude/skills/onboard/SKILL.md (or your agent's skills folder).You are setting up CIAgent (pip install ciagent) so this repo's AI agent has
recorded golden baselines and a runnable regression suite. The end state: the
user can run ciagent test --runs 3 and see a stability report for their agent.
Work through the steps in order. Do not skip the cost gate in step 4.
openai, anthropic, langgraph,
langchain) and for the function or endpoint that takes a user message and
returns the agent's answer.pip install "ciagent[openai]", [anthropic], [langgraph], or [all].ciagent --version then ciagent doctor (it reports what is
missing; a missing spec is expected at this point).Create agentci_runner.py at the repo root (or inside the package if the repo
has one clear package):
def run_for_agentci(query: str) -> str:
"""CIAgent entry point: one query in, final answer text out."""
# import the user's agent and invoke it ONCE, no chat history
...
return final_answer_textRules:
python -c "from agentci_runner import run_for_agentci; print(run_for_agentci('hello'))".Write agentci_queries.txt, one query per line — 8 to 15 queries:
Recording baselines runs the real agent once per query, on the user's API keys.
State the query count and a cost ballpark, and ask the user to confirm
before step 5. If there are no API keys or the user declines: write
agentci_spec.yaml by hand instead (same queries, runner: set), validate with
ciagent test --mock, and tell the user which step to resume later.
ciagent bootstrap --runner agentci_runner:run_for_agentci \
--queries agentci_queries.txt --agent <agent-name> --yesThis runs every query, saves each trace as a golden baseline under
./baselines/<agent-name>/, and writes agentci_spec.yaml with path and cost
budgets derived from the recorded traces. Read the printed answers as they
stream by — if an answer is visibly wrong, that query should not be golden:
fix the agent or the query, delete that baseline file, and rerun.
The generated spec has path and cost budgets but no correctness checks. Add a
correctness: block per query, derived from the recorded baseline answers and
the repo's docs/KB — never from what you wish the agent said:
correctness:
expected_in_answer: ["30 days"] # hard facts, AND
any_expected_in_answer: ["$9.95", "9.95"] # phrasing variants, OR
not_in_answer: ["I don't know"] # forbidden contentCheck facts, not phrasing. If the repo has a knowledge-base directory, run
ciagent generate-checks --kb <dir> --dry-run and review its candidates —
every surviving candidate was already validated against the recorded goldens.
ciagent test --mock # structure check, zero API calls
ciagent test --yes --format json # live run (covered by step 4 approval)
ciagent test --runs 3 --yes # stability reportExit codes: 0 = pass (flaky-but-passing is 0), 1 = correctness failure in every
run, 2 = infra/config error. In the stability report, flips labeled
agent-variance mean the agent's answer changed (an agent problem); flips
labeled judge-flake mean the eval itself is unstable (a check/judge problem).
If a check fails, fix the agent or fix a factually wrong check. Do not loosen a correct check to make the run green — report the failure to the user instead.
ciagent init scaffolds a GitHub Actions workflow (add --hook for a
pre-push hook if the user wants it).agentci_queries.txt, agentci_spec.yaml, baselines/,
and the workflow.ciagent test --runs 3 is the
command to watch after future agent changes.© davepoon, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in plugins/ciagent/skills/onboard of davepoon/buildwithclaude.
Open the folder on GitHubat commit 616deb5
Onboard next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Onboard this skilldavepoon/buildwithclaude | 3.6k | — | ~1.3k | Automated safety check: Pass | MIT | |
| Eval Creator CIpskoett/pskoett-ai-skills | 314 | — | ~2.2k | Automated safety check: Pass | None | |
| RAG Observability Evalssickn33/agentic-awesome-skills | 47k | 2 repos | ~3.1k | Automated safety check: Pass | MIT | |
| EvaluationPrimeIntellect-ai/prime-envs | 130 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Eval Driven Devgithub/awesome-copilot | 40k | 1 repos | ~4.4k | Automated safety check: Warn | MIT | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT |
pskoett/pskoett-ai-skills
[Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows).
sickn33/agentic-awesome-skills
Monitor and evaluate RAG systems with retrieval quality metrics, groundedness checks, hallucination detection, and continuous regression testing.
PrimeIntellect-ai/prime-envs
Install and run a verifiers environment — smoke testing during development and full benchmark evals.
github/awesome-copilot
Improve AI application with evaluation-driven development. An agent skill from github/awesome-copilot.
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
davepoon/buildwithclaude
Build, update, and apply iOS design specifications using Apple Human Interface Guidelines (HIG) source data.
davepoon/buildwithclaude
Download YouTube videos with customizable quality and format options.
davepoon/buildwithclaude
A skill your agent uses when the user asks to "analyze video", "watch this video", "what happens in this video", "describe this clip", "review this footage", "classify these videos", "compare…
davepoon/buildwithclaude
Discover Atlas Cloud image and video models, inspect their live schemas, and submit one confirmed media generation request with bounded GET polling.
davepoon/buildwithclaude
面向没有编程经验的用户,把想法做成可试用的浏览器插件,并完成检查、商店材料、审核提交和上线验证;也用于继续已有插件、排错和发布新版。用户说“帮我做个插件”“把插件上架”“继续我的插件”时使用。普通网站开发、仅查询插件知识不触发。
davepoon/buildwithclaude
Toolkit for creating animated GIFs optimized for Slack, with validators for size constraints and composable animation primitives.
Categories
Set up CIAgent regression testing for the AI agent in this repo — write a runner, record golden baselines, generate a test spec, and verify it. Onboard is an agent skill from davepoon/buildwithclaude. Set up CIAgent regression testing for the AI agent in this repo — write a runner, record golden baselines, generate a test spec, and verify it.
Onboard fits situations like: the user asks to add tests; regression testing for their AI agent.
Run `npx skills add davepoon/buildwithclaude --skill onboard -a claude-code`. Or copy the skill folder (plugins/ciagent/skills/onboard in davepoon/buildwithclaude) into .claude/skills/onboard in your project. Claude Code loads it when a task matches its description.
Run `npx skills add davepoon/buildwithclaude --skill onboard -a codex`. Or copy the skill folder (plugins/ciagent/skills/onboard in davepoon/buildwithclaude) into .agents/skills/onboard in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davepoon/buildwithclaude --skill onboard -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/onboard, .gemini/skills/onboard, .github/skills/onboard and .opencode/skills/onboard in your project.
Going by SKILL.md and its folder, Onboard needs the command-line tools its instructions call (pip and python). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash(ciagent *), Bash(pip install *), Bash(python -c *), Read, Grep, Glob, Write, Edit.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Onboard is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Onboard: Eval Creator CI (pskoett/pskoett-ai-skills, 314 stars), RAG Observability Evals (sickn33/agentic-awesome-skills, 47k stars), Evaluation (PrimeIntellect-ai/prime-envs, 130 stars) and Eval Driven Dev (github/awesome-copilot, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
davepoon (a GitHub user) maintains it in davepoon/buildwithclaude, which has 3,605 GitHub stars. The repository holds 246 skills in this directory. The repository was last updated on October 9, 2026.
Source: davepoon/buildwithclaude on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.