Agent Eval Engineering
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
Runs and debugs evaluations of how Chatbox models answer questions about large attached files, using synthetic and real long-document fixtures.
$ npx skills add chatboxai/chatbox --skill session-rag-eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install chatboxai/chatbox session-rag-eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/chatboxai/chatbox.git skills-src && mkdir -p .claude/skills && cp -r skills-src/scripts/session-rag-eval .claude/skills/session-rag-eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "session-rag-eval" agent skill from https://github.com/chatboxai/chatbox/tree/main/scripts/session-rag-eval into .claude/skills/session-rag-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "session-rag-eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/chatboxai/chatbox/tree/main/scripts/session-rag-evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add chatboxai/chatbox --skill session-rag-eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install chatboxai/chatbox session-rag-eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/chatboxai/chatbox.git skills-src && mkdir -p .agents/skills && cp -r skills-src/scripts/session-rag-eval .agents/skills/session-rag-eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "session-rag-eval" agent skill from https://github.com/chatboxai/chatbox/tree/main/scripts/session-rag-eval into .agents/skills/session-rag-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "session-rag-eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add chatboxai/chatbox --skill session-rag-eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install chatboxai/chatbox session-rag-eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/chatboxai/chatbox.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/scripts/session-rag-eval .cursor/skills/session-rag-eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "session-rag-eval" agent skill from https://github.com/chatboxai/chatbox/tree/main/scripts/session-rag-eval into .cursor/skills/session-rag-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "session-rag-eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/chatboxai/chatbox.git --path scripts/session-rag-eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add chatboxai/chatbox --skill session-rag-eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install chatboxai/chatbox session-rag-eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/chatboxai/chatbox.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/scripts/session-rag-eval .gemini/skills/session-rag-eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "session-rag-eval" agent skill from https://github.com/chatboxai/chatbox/tree/main/scripts/session-rag-eval into .gemini/skills/session-rag-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "session-rag-eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install chatboxai/chatbox session-rag-evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add chatboxai/chatbox --skill session-rag-eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/chatboxai/chatbox.git skills-src && mkdir -p .github/skills && cp -r skills-src/scripts/session-rag-eval .github/skills/session-rag-eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "session-rag-eval" agent skill from https://github.com/chatboxai/chatbox/tree/main/scripts/session-rag-eval into .github/skills/session-rag-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "session-rag-eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add chatboxai/chatbox --skill session-rag-eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install chatboxai/chatbox session-rag-eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/chatboxai/chatbox.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/scripts/session-rag-eval .opencode/skills/session-rag-eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "session-rag-eval" agent skill from https://github.com/chatboxai/chatbox/tree/main/scripts/session-rag-eval into .opencode/skills/session-rag-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "session-rag-eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
session-rag-evalRuns and debugs evaluations of how Chatbox models answer questions about large attached files, using synthetic and real long-document fixtures.
The evaluation checks two sides of model behavior: the model should call `query_session_attachment` when the answer depends on a large uploaded file, and should skip retrieval for clearly unrelated requests. The preferred route is the Chatbox conversation-flow harness, because it exercises the real renderer, config, license, local API, file upload, indexing, tool registration and persisted messages. Fixtures live in a separate repository and include synthetic long documents with hidden needle facts, Wikipedia-derived real documents, direct and implicit positive cases, and unrelated negative cases.
You start the local API, build with `USE_LOCAL_API=true` so the renderer origin is compiled into the bundle, and run `pnpm eval:session-rag:chatbox` with a named case. The harness copies the real `config.json` into an isolated temporary user-data directory, injects a default chat model if one is missing and keeps the session RAG sqlite database in a separate path. A list of known failure modes covers an attachment reported as too large, requests that reach production, Electron hanging at startup and files shown as still indexing.
Read from SKILL.md and the folder at commit 0ac6385. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (JavaScript), which the agent can run.
Shell commands in SKILL.md call:
pnpmnodeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pnpm, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Chatbox Session RAG Eval loads about 758 tokens when it runs. Until then it costs about 31 tokens; SKILL.md has 276 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from chatboxai/chatbox at commit 0ac6385, republished under its GPL-3.0 licence (© chatboxai). 276 words, ~758 tokens.
.claude/skills/session-rag-eval/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Use this skill when validating file question-answering behavior for large chat attachments.
Verify both sides of model behavior:
query_session_attachment when the answer depends on an uploaded large file.Prefer the Chatbox conversation-flow harness for product validation because it exercises the real renderer, config, license, local API, file upload, indexing, tool registration, and persisted messages.
Default fixture repo:
../../chatbox-session-rag-eval-fixturesFixture types:
Regenerate fixtures in the fixture repo:
node scripts/generate-fixtures.mjs
node scripts/fetch-real-fixtures.mjsStart the local API before running the harness. Then build with USE_LOCAL_API=true; the renderer API origin is compiled
into the bundle.
USE_LOCAL_API=true node ./node_modules/electron-vite/bin/electron-vite.js build --mode development
pnpm eval:session-rag:chatbox -- --case long-citrine-threshold --keep-user-dataThe harness copies the real config.json into an isolated temporary userDataDir, injects a temporary default chat model
if missing, and stores the session RAG sqlite DB at a separate temporary path.
session_attachment_embedding capability is false or unavailable for the
active local API/license path.USE_LOCAL_API=true.sessionAttachmentIndexStatus: "ready" before context building.pnpm eval:session-rag:chatbox -- --case long-citrine-threshold
pnpm eval:session-rag:chatbox -- --case implicit-citrine-current-policy
pnpm eval:session-rag:chatbox -- --case real-wiki-apollo-implicit-landing-site
pnpm eval:session-rag:chatbox -- --case multi-turn-real-wiki-apollo-followup
pnpm eval:session-rag:chatbox -- --case unrelated-simple-math
pnpm eval:session-rag:chatbox -- --case real-wiki-unrelated-capitalFor fast model-behavior iteration without Electron:
pnpm eval:session-rag -- --dry-run
pnpm eval:session-rag -- --case implicit-citrine-current-policy© chatboxai, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files in scripts/session-rag-eval of chatboxai/chatbox.
Open the folder on GitHubat commit 0ac6385
Chatbox Session RAG Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Chatbox Session RAG Eval this skillchatboxai/chatbox | 42k | — | ~758 | Automated safety check: Pass | GPL-3.0 | |
| Agent Eval Engineeringlangchain-ai/langchain-skills | 1.3k | — | ~4k | Automated safety check: Pass | MIT | |
| Evaluate RAGai-evals-course/evals-skills | 1.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Synthetic Eval Data Generatorai-evals-course/evals-skills | 1.5k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| Paperclip Evalspaperclipai/paperclip | 98k | — | ~839 | Automated safety check: Pass | MIT | |
| Run EvalsDevin-AXIS/iPolloWork | 6.7k | — | ~851 | Automated safety check: Pass | Custom licence |
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
ai-evals-course/evals-skills
Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.
ai-evals-course/evals-skills
Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.
paperclipai/paperclip
Choose, inspect, validate, and report Paperclip Runner or Product E2E evaluations while preserving evidence, provenance, cost, and failure classification.
Devin-AXIS/iPolloWork
do e2e tests, run e2e, validate feature, prove it works, PR proof, frame proof, pnpm evals.
Jeffallan/claude-skills
Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.
chatboxai/chatbox
Translates new or changed i18n keys from a Chatbox Pro diff, staged changes or a commit range, writing the locale JSON files directly with a built-in glossary.
chatboxai/chatbox
Cherry-picks commits from the chatbox-pro repo into the open chatbox repo, skipping mobile-only files and keeping the open package name.
Works with
Categories
Runs and debugs evaluations of how Chatbox models answer questions about large attached files, using synthetic and real long-document fixtures. The evaluation checks two sides of model behavior: the model should call `query_session_attachment` when the answer depends on a large uploaded file, and should skip retrieval for clearly unrelated requests. The preferred route is the Chatbox conversation-flow harness, because it exercises the real renderer, config, license, local API, file upload, indexing, tool registration and persisted messages.
Chatbox Session RAG Eval fits situations like: checking that a model queries a large uploaded file instead of guessing; confirming the model skips retrieval on unrelated questions; debugging an evaluation run that hangs or reports an attachment as still indexing.
Run `npx skills add chatboxai/chatbox --skill session-rag-eval -a claude-code`. Or copy the skill folder (scripts/session-rag-eval in chatboxai/chatbox) into .claude/skills/session-rag-eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add chatboxai/chatbox --skill session-rag-eval -a codex`. Or copy the skill folder (scripts/session-rag-eval in chatboxai/chatbox) into .agents/skills/session-rag-eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add chatboxai/chatbox --skill session-rag-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/session-rag-eval, .gemini/skills/session-rag-eval, .github/skills/session-rag-eval and .opencode/skills/session-rag-eval in your project.
Going by SKILL.md and its folder, Chatbox Session RAG Eval needs JavaScript for the scripts in its folder and the command-line tools its instructions call (pnpm and node). Our summary lists: A Chatbox source checkout built with `USE_LOCAL_API=true`; pnpm and Node; The separate session RAG fixtures repository.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Chatbox Session RAG Eval is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 758 tokens (SKILL.md is roughly 3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Chatbox Session RAG Eval: Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars), Evaluate RAG (ai-evals-course/evals-skills, 1.5k stars), Synthetic Eval Data Generator (ai-evals-course/evals-skills, 1.5k stars) and Paperclip Evals (paperclipai/paperclip, 98k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
chatboxai (a GitHub organization) maintains it in chatboxai/chatbox, which has 41,956 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 24, 2026.
Source: chatboxai/chatbox on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.