Agent skill

Onboard

by davepoon in davepoon/buildwithclaude

Set up CIAgent regression testing for the AI agent in this repo — write a runner, record golden baselines, generate a test spec, and verify it.

MITAuto-check passedAI & LLM Engineering

Install Onboard

skills CLI
$ npx skills add davepoon/buildwithclaude --skill onboard -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install davepoon/buildwithclaude onboard --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/davepoon/buildwithclaude.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ciagent/skills/onboard .claude/skills/onboard && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
onboard
GitHub stars
3.6k
Token cost
~1.3k tokens
SKILL.md length
638 words
Files
1
Skills in repo
246
Repo updated
First seen
Licence
MIT

At a glance

Set up CIAgent regression testing for the AI agent in this repo — write a runner, record golden baselines, generate a test spec, and verify it.

  • Works in 8 steps: Find the agent and install CIAgent → Write the runner → Choose queries → …
  • The user asks to add tests
  • SKILL.md covers 1. Find the agent and install…, 2. Write the runner, 3. Choose queries and 4. Cost gate — ask before…, plus 4 more sections
  • Calls pip and python

What it does

Onboard is an agent skill from davepoon/buildwithclaude. Set up CIAgent regression testing for the AI agent in this repo — write a runner, record golden baselines, generate a test spec, and verify it. Use when the user asks to add tests, evals, or regression testing for their AI agent, or to set up CIAgent.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering QA and bug reports and LLM evaluation. The repository describes itself as: A single hub to find Claude Skills, Agents, Commands, Hooks, Plugins, and Marketplace collections to extend Claude Code, Claude Desktop, Agent SDK and OpenClaw. The licence is MIT.

When your agent uses it

  • The user asks to add tests
  • Regression testing for their AI agent

Example prompts

  • “/onboard”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Bash(ciagent *), Bash(pip install *), Bash(python -c *), Read, Grep, Glob, Write, Edit

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Find the agent and install CIAgent
  2. Write the runner
  3. Choose queries
  4. Cost gate — ask before running live
  5. Record golden baselines
  6. Add correctness checks
  7. Verify
  8. Wire CI and hand off

What it can do on your machine

Read from SKILL.md and the folder at commit 616deb5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(ciagent *)
    • Bash(pip install *)
    • Bash(python -c *)
    • Read
    • Grep
    • Glob
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Onboard loads about 1.3k tokens when it runs. Until then it costs about 65 tokens; SKILL.md has 638 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~65
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from davepoon/buildwithclaude at commit 616deb5, republished under its MIT licence (© davepoon). 638 words, ~1,334 tokens.

Download SKILL.mdSave it as .claude/skills/onboard/SKILL.md (or your agent's skills folder).
name
onboard
description
Set up CIAgent regression testing for the AI agent in this repo — write a runner, record golden baselines, generate a test spec, and verify it. Use when the user asks to add tests, evals, or regression testing for their AI agent, or to set up CIAgent.
allowed-tools
Bash(ciagent *), Bash(pip install *), Bash(python -c *), Read, Grep, Glob, Write, Edit

Onboard CIAgent into this repo

You are setting up CIAgent (pip install ciagent) so this repo's AI agent has recorded golden baselines and a runnable regression suite. The end state: the user can run ciagent test --runs 3 and see a stability report for their agent.

Work through the steps in order. Do not skip the cost gate in step 4.

1. Find the agent and install CIAgent

  • Locate the agent: search for LLM SDK usage (openai, anthropic, langgraph, langchain) and for the function or endpoint that takes a user message and returns the agent's answer.
  • Install with the matching extra so trace capture hooks the SDK: pip install "ciagent[openai]", [anthropic], [langgraph], or [all].
  • Sanity check: ciagent --version then ciagent doctor (it reports what is missing; a missing spec is expected at this point).

2. Write the runner

Create agentci_runner.py at the repo root (or inside the package if the repo has one clear package):

python
def run_for_agentci(query: str) -> str:
    """CIAgent entry point: one query in, final answer text out."""
    # import the user's agent and invoke it ONCE, no chat history
    ...
    return final_answer_text

Rules:

  • Return the final answer string. CIAgent wraps the call in its own trace capture, so LLM calls and tool calls are recorded automatically — do not build Trace objects unless the repo already produces them.
  • Fresh context per call: no shared history between queries.
  • Reuse the repo's own config/env loading so the runner works from the repo root.
  • Verify it imports and answers before going further: python -c "from agentci_runner import run_for_agentci; print(run_for_agentci('hello'))".

3. Choose queries

Write agentci_queries.txt, one query per line — 8 to 15 queries:

  • Cover the agent's main jobs (mine the README, docs, knowledge base, prompts, and existing tests for what it is supposed to handle).
  • Include at least 2 out-of-scope queries the agent should refuse or deflect.
  • Prefer queries whose correct answers contain hard facts (prices, dates, limits, names) — those become deterministic checks in step 6.

4. Cost gate — ask before running live

Recording baselines runs the real agent once per query, on the user's API keys. State the query count and a cost ballpark, and ask the user to confirm before step 5. If there are no API keys or the user declines: write agentci_spec.yaml by hand instead (same queries, runner: set), validate with ciagent test --mock, and tell the user which step to resume later.

Show full SKILL.md (273 more words)Show less

5. Record golden baselines

bash
ciagent bootstrap --runner agentci_runner:run_for_agentci \
  --queries agentci_queries.txt --agent <agent-name> --yes

This runs every query, saves each trace as a golden baseline under ./baselines/<agent-name>/, and writes agentci_spec.yaml with path and cost budgets derived from the recorded traces. Read the printed answers as they stream by — if an answer is visibly wrong, that query should not be golden: fix the agent or the query, delete that baseline file, and rerun.

6. Add correctness checks

The generated spec has path and cost budgets but no correctness checks. Add a correctness: block per query, derived from the recorded baseline answers and the repo's docs/KB — never from what you wish the agent said:

yaml
correctness:
  expected_in_answer: ["30 days"]          # hard facts, AND
  any_expected_in_answer: ["$9.95", "9.95"] # phrasing variants, OR
  not_in_answer: ["I don't know"]           # forbidden content

Check facts, not phrasing. If the repo has a knowledge-base directory, run ciagent generate-checks --kb <dir> --dry-run and review its candidates — every surviving candidate was already validated against the recorded goldens.

7. Verify

bash
ciagent test --mock                      # structure check, zero API calls
ciagent test --yes --format json         # live run (covered by step 4 approval)
ciagent test --runs 3 --yes              # stability report

Exit codes: 0 = pass (flaky-but-passing is 0), 1 = correctness failure in every run, 2 = infra/config error. In the stability report, flips labeled agent-variance mean the agent's answer changed (an agent problem); flips labeled judge-flake mean the eval itself is unstable (a check/judge problem).

If a check fails, fix the agent or fix a factually wrong check. Do not loosen a correct check to make the run green — report the failure to the user instead.

8. Wire CI and hand off

  • ciagent init scaffolds a GitHub Actions workflow (add --hook for a pre-push hook if the user wants it).
  • Commit: runner, agentci_queries.txt, agentci_spec.yaml, baselines/, and the workflow.
  • Tell the user: how many goldens were recorded, the suite score, anything flaky (with its flip source), and that ciagent test --runs 3 is the command to watch after future agent changes.

© davepoon, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/ciagent/skills/onboard of davepoon/buildwithclaude.

Open the folder on GitHubat commit 616deb5

Compare with similar skills

Onboard next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Onboard compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Onboard this skilldavepoon/buildwithclaude3.6k—~1.3kAutomated safety check: PassMIT
Eval Creator CIpskoett/pskoett-ai-skills314—~2.2kAutomated safety check: PassNone
RAG Observability Evalssickn33/agentic-awesome-skills47k2 repos~3.1kAutomated safety check: PassMIT
EvaluationPrimeIntellect-ai/prime-envs130—~4.6kAutomated safety check: PassApache-2.0
Eval Driven Devgithub/awesome-copilot40k1 repos~4.4kAutomated safety check: WarnMIT
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT

Similar skills

  • Eval Creator CI

    pskoett/pskoett-ai-skills

    [Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows).

    314 GitHub stars~2.2k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • RAG Observability Evals

    sickn33/agentic-awesome-skills

    Monitor and evaluate RAG systems with retrieval quality metrics, groundedness checks, hallucination detection, and continuous regression testing.

    47k GitHub starsUsed in 2 repos~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Evaluation

    PrimeIntellect-ai/prime-envs

    Install and run a verifiers environment — smoke testing during development and full benchmark evals.

    130 GitHub stars~4.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Eval Driven Dev

    github/awesome-copilot

    Official

    Improve AI application with evaluation-driven development. An agent skill from github/awesome-copilot.

    40k GitHub starsUsed in 1 repo~4.4k tokens
    Testing & QAAuto-check: warnings
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed

More from davepoon/buildwithclaude

All 246 skills in this repo
  • iOS Hig Design Guide

    davepoon/buildwithclaude

    Build, update, and apply iOS design specifications using Apple Human Interface Guidelines (HIG) source data.

    3.6k GitHub stars~735 tokensUpdated today
    Auto-check passed
  • Video Downloader

    davepoon/buildwithclaude

    Download YouTube videos with customizable quality and format options.

    3.6k GitHub starsUsed in 1 repo~871 tokens
    Auto-check passed
  • Qwen Vision

    davepoon/buildwithclaude

    A skill your agent uses when the user asks to "analyze video", "watch this video", "what happens in this video", "describe this clip", "review this footage", "classify these videos", "compare…

    3.6k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Atlas Cloud Media

    davepoon/buildwithclaude

    Discover Atlas Cloud image and video models, inspect their live schemas, and submit one confirmed media generation request with bounded GET polling.

    3.6k GitHub stars~852 tokensUpdated today
    Auto-check passed
  • Browser Extension Launch

    davepoon/buildwithclaude

    面向没有编程经验的用户,把想法做成可试用的浏览器插件,并完成检查、商店材料、审核提交和上线验证;也用于继续已有插件、排错和发布新版。用户说“帮我做个插件”“把插件上架”“继续我的插件”时使用。普通网站开发、仅查询插件知识不触发。

    3.6k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Slack Gif Creator

    davepoon/buildwithclaude

    Toolkit for creating animated GIFs optimized for Slack, with validators for size constraints and composable animation primitives.

    3.6k GitHub starsUsed in 12 repos~4.3k tokens
    Auto-check passed

Questions about Onboard

What does Onboard do?

Set up CIAgent regression testing for the AI agent in this repo — write a runner, record golden baselines, generate a test spec, and verify it. Onboard is an agent skill from davepoon/buildwithclaude. Set up CIAgent regression testing for the AI agent in this repo — write a runner, record golden baselines, generate a test spec, and verify it.

When should I use Onboard?

Onboard fits situations like: the user asks to add tests; regression testing for their AI agent.

How do I install Onboard in Claude Code?

Run `npx skills add davepoon/buildwithclaude --skill onboard -a claude-code`. Or copy the skill folder (plugins/ciagent/skills/onboard in davepoon/buildwithclaude) into .claude/skills/onboard in your project. Claude Code loads it when a task matches its description.

How do I install Onboard in Codex?

Run `npx skills add davepoon/buildwithclaude --skill onboard -a codex`. Or copy the skill folder (plugins/ciagent/skills/onboard in davepoon/buildwithclaude) into .agents/skills/onboard in your project. Codex loads it when a task matches its description.

Can I use Onboard in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davepoon/buildwithclaude --skill onboard -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/onboard, .gemini/skills/onboard, .github/skills/onboard and .opencode/skills/onboard in your project.

What does Onboard need to run?

Going by SKILL.md and its folder, Onboard needs the command-line tools its instructions call (pip and python). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash(ciagent *), Bash(pip install *), Bash(python -c *), Read, Grep, Glob, Write, Edit.

Does Onboard access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Onboard safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Onboard use?

Onboard is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Onboard use?

About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Onboard?

Skills that share tags, products or a category with Onboard: Eval Creator CI (pskoett/pskoett-ai-skills, 314 stars), RAG Observability Evals (sickn33/agentic-awesome-skills, 47k stars), Evaluation (PrimeIntellect-ai/prime-envs, 130 stars) and Eval Driven Dev (github/awesome-copilot, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Onboard?

davepoon (a GitHub user) maintains it in davepoon/buildwithclaude, which has 3,605 GitHub stars. The repository holds 246 skills in this directory. The repository was last updated on October 9, 2026.

Source: davepoon/buildwithclaude on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.