Agent skill

Chatbox Session RAG Eval

by chatboxai in chatboxai/chatbox

Runs and debugs evaluations of how Chatbox models answer questions about large attached files, using synthetic and real long-document fixtures.

GPL-3.0Auto-check passedAI & LLM Engineering

Install Chatbox Session RAG Eval

skills CLI
$ npx skills add chatboxai/chatbox --skill session-rag-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install chatboxai/chatbox session-rag-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/chatboxai/chatbox.git skills-src && mkdir -p .claude/skills && cp -r skills-src/scripts/session-rag-eval .claude/skills/session-rag-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
session-rag-eval
GitHub stars
42k
Token cost
~758 tokens
SKILL.md length
276 words
Files
4
Skills in repo
3
Repo updated
First seen
Licence
GPL-3.0

At a glance

Runs and debugs evaluations of how Chatbox models answer questions about large attached files, using synthetic and real long-document fixtures.

  • Checking that a model queries a large uploaded file instead of guessing
  • SKILL.md covers Goal, Fixtures, Chatbox Flow and Known Failure Modes, plus 1 more section
  • Runs JavaScript scripts from its folder; calls pnpm and node
  • Confirming the model skips retrieval on unrelated questions

What it does

The evaluation checks two sides of model behavior: the model should call `query_session_attachment` when the answer depends on a large uploaded file, and should skip retrieval for clearly unrelated requests. The preferred route is the Chatbox conversation-flow harness, because it exercises the real renderer, config, license, local API, file upload, indexing, tool registration and persisted messages. Fixtures live in a separate repository and include synthetic long documents with hidden needle facts, Wikipedia-derived real documents, direct and implicit positive cases, and unrelated negative cases.

You start the local API, build with `USE_LOCAL_API=true` so the renderer origin is compiled into the bundle, and run `pnpm eval:session-rag:chatbox` with a named case. The harness copies the real `config.json` into an isolated temporary user-data directory, injects a default chat model if one is missing and keeps the session RAG sqlite database in a separate path. A list of known failure modes covers an attachment reported as too large, requests that reach production, Electron hanging at startup and files shown as still indexing.

When your agent uses it

  • Checking that a model queries a large uploaded file instead of guessing
  • Confirming the model skips retrieval on unrelated questions
  • Debugging an evaluation run that hangs or reports an attachment as still indexing

Example prompts

  • “Run the long-citrine-threshold case through the Chatbox flow harness and keep the user data.”
  • “The attachment says it is too large. Check whether session_attachment_embedding is enabled for my local API.”
  • “Regenerate the synthetic and real long-document fixtures, then rerun the implicit-citrine-current-policy case.”

Requirements

  • A Chatbox source checkout built with `USE_LOCAL_API=true`
  • pnpm and Node
  • The separate session RAG fixtures repository

What it can do on your machine

Read from SKILL.md and the folder at commit 0ac6385. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • pnpm
    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pnpm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Chatbox Session RAG Eval loads about 758 tokens when it runs. Until then it costs about 31 tokens; SKILL.md has 276 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~31
When it runs · the whole SKILL.md, loaded when a task matches
~758

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from chatboxai/chatbox at commit 0ac6385, republished under its GPL-3.0 licence (© chatboxai). 276 words, ~758 tokens.

Download SKILL.mdSave it as .claude/skills/session-rag-eval/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
session-rag-eval
description
Run and debug Chatbox session attachment RAG model evaluation with synthetic and real long-file fixtures.

Session Attachment RAG Eval

Use this skill when validating file question-answering behavior for large chat attachments.

Goal

Verify both sides of model behavior:

  • The model calls query_session_attachment when the answer depends on an uploaded large file.
  • The model avoids retrieval for clearly unrelated user requests.

Prefer the Chatbox conversation-flow harness for product validation because it exercises the real renderer, config, license, local API, file upload, indexing, tool registration, and persisted messages.

Fixtures

Default fixture repo:

bash
../../chatbox-session-rag-eval-fixtures

Fixture types:

  • synthetic long documents with hidden needle facts;
  • Wikipedia-derived real long documents with source attribution;
  • direct positive cases;
  • implicit positive cases where the user does not explicitly ask to search the file;
  • unrelated negative cases.

Regenerate fixtures in the fixture repo:

bash
node scripts/generate-fixtures.mjs
node scripts/fetch-real-fixtures.mjs

Chatbox Flow

Start the local API before running the harness. Then build with USE_LOCAL_API=true; the renderer API origin is compiled into the bundle.

bash
USE_LOCAL_API=true node ./node_modules/electron-vite/bin/electron-vite.js build --mode development
pnpm eval:session-rag:chatbox -- --case long-citrine-threshold --keep-user-data

The harness copies the real config.json into an isolated temporary userDataDir, injects a temporary default chat model if missing, and stores the session RAG sqlite DB at a separate temporary path.

Known Failure Modes

  • "This attachment is too large..." means session_attachment_embedding capability is false or unavailable for the active local API/license path.
  • If requests go to production, rebuild with USE_LOCAL_API=true.
  • If Electron hangs at startup, check stale CDP/Electron processes. The harness now allocates a free CDP port.
  • If the assistant says the file is still indexing after DB says ready, inspect the persisted user message. The message file must have sessionAttachmentIndexStatus: "ready" before context building.
  • Do not trust a log line that says the message was submitted unless the textarea cleared and the session contains the user message.

Useful Cases

bash
pnpm eval:session-rag:chatbox -- --case long-citrine-threshold
pnpm eval:session-rag:chatbox -- --case implicit-citrine-current-policy
pnpm eval:session-rag:chatbox -- --case real-wiki-apollo-implicit-landing-site
pnpm eval:session-rag:chatbox -- --case multi-turn-real-wiki-apollo-followup
pnpm eval:session-rag:chatbox -- --case unrelated-simple-math
pnpm eval:session-rag:chatbox -- --case real-wiki-unrelated-capital

For fast model-behavior iteration without Electron:

bash
pnpm eval:session-rag -- --dry-run
pnpm eval:session-rag -- --case implicit-citrine-current-policy

© chatboxai, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in scripts/session-rag-eval of chatboxai/chatbox.

  • SKILL.md
  • README.md
  • run-chatbox-flow-eval.mjs
  • run-session-rag-eval.mjs

Open the folder on GitHubat commit 0ac6385

Compare with similar skills

Chatbox Session RAG Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Chatbox Session RAG Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Chatbox Session RAG Eval this skillchatboxai/chatbox42k—~758Automated safety check: PassGPL-3.0
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
Evaluate RAGai-evals-course/evals-skills1.5k—~1.9kAutomated safety check: PassApache-2.0
Synthetic Eval Data Generatorai-evals-course/evals-skills1.5k—~1.4kAutomated safety check: PassApache-2.0
Paperclip Evalspaperclipai/paperclip98k—~839Automated safety check: PassMIT
Run EvalsDevin-AXIS/iPolloWork6.7k—~851Automated safety check: PassCustom licence

Similar skills

  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Evaluate RAG

    ai-evals-course/evals-skills

    Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.

    1.5k GitHub stars~1.9k tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check passed
  • Synthetic Eval Data Generator

    ai-evals-course/evals-skills

    Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.

    1.5k GitHub stars~1.4k tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check passed
  • Paperclip Evals

    paperclipai/paperclip

    Choose, inspect, validate, and report Paperclip Runner or Product E2E evaluations while preserving evidence, provenance, cost, and failure classification.

    98k GitHub stars~839 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Run Evals

    Devin-AXIS/iPolloWork

    do e2e tests, run e2e, validate feature, prove it works, PR proof, frame proof, pnpm evals.

    6.7k GitHub stars~851 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • RAG Architect

    Jeffallan/claude-skills

    Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

    12k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed

More from chatboxai/chatbox

  • Chatbox i18n Translator

    chatboxai/chatbox

    Translates new or changed i18n keys from a Chatbox Pro diff, staged changes or a commit range, writing the locale JSON files directly with a built-in glossary.

    42k GitHub stars~508 tokensUpdated 13 days ago
    Auto-check passed
  • Cherry-picks commits from the chatbox-pro repo into the open chatbox repo, skipping mobile-only files and keeping the open package name.

    42k GitHub stars~1.2k tokensUpdated 13 days ago
    Auto-check passed

Works with

Questions about Chatbox Session RAG Eval

What does Chatbox Session RAG Eval do?

Runs and debugs evaluations of how Chatbox models answer questions about large attached files, using synthetic and real long-document fixtures. The evaluation checks two sides of model behavior: the model should call `query_session_attachment` when the answer depends on a large uploaded file, and should skip retrieval for clearly unrelated requests. The preferred route is the Chatbox conversation-flow harness, because it exercises the real renderer, config, license, local API, file upload, indexing, tool registration and persisted messages.

When should I use Chatbox Session RAG Eval?

Chatbox Session RAG Eval fits situations like: checking that a model queries a large uploaded file instead of guessing; confirming the model skips retrieval on unrelated questions; debugging an evaluation run that hangs or reports an attachment as still indexing.

How do I install Chatbox Session RAG Eval in Claude Code?

Run `npx skills add chatboxai/chatbox --skill session-rag-eval -a claude-code`. Or copy the skill folder (scripts/session-rag-eval in chatboxai/chatbox) into .claude/skills/session-rag-eval in your project. Claude Code loads it when a task matches its description.

How do I install Chatbox Session RAG Eval in Codex?

Run `npx skills add chatboxai/chatbox --skill session-rag-eval -a codex`. Or copy the skill folder (scripts/session-rag-eval in chatboxai/chatbox) into .agents/skills/session-rag-eval in your project. Codex loads it when a task matches its description.

Can I use Chatbox Session RAG Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add chatboxai/chatbox --skill session-rag-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/session-rag-eval, .gemini/skills/session-rag-eval, .github/skills/session-rag-eval and .opencode/skills/session-rag-eval in your project.

What does Chatbox Session RAG Eval need to run?

Going by SKILL.md and its folder, Chatbox Session RAG Eval needs JavaScript for the scripts in its folder and the command-line tools its instructions call (pnpm and node). Our summary lists: A Chatbox source checkout built with `USE_LOCAL_API=true`; pnpm and Node; The separate session RAG fixtures repository.

Does Chatbox Session RAG Eval access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Chatbox Session RAG Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Chatbox Session RAG Eval use?

Chatbox Session RAG Eval is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Chatbox Session RAG Eval use?

About 758 tokens (SKILL.md is roughly 3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Chatbox Session RAG Eval?

Skills that share tags, products or a category with Chatbox Session RAG Eval: Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars), Evaluate RAG (ai-evals-course/evals-skills, 1.5k stars), Synthetic Eval Data Generator (ai-evals-course/evals-skills, 1.5k stars) and Paperclip Evals (paperclipai/paperclip, 98k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Chatbox Session RAG Eval?

chatboxai (a GitHub organization) maintains it in chatboxai/chatbox, which has 41,956 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 24, 2026.

Source: chatboxai/chatbox on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.