Agent skill

Ragu Build

by RaguTeam in RaguTeam/RAGU

Interview the user about their RAGU use case, select an appropriate RAGU pipeline, and generate a validated ragubuild.yaml plus a runnable build<name.py script.

MITAuto-check: notesKnowledge Management

Install Ragu Build

skills CLI
$ npx skills add RaguTeam/RAGU --skill ragu-build -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install RaguTeam/RAGU ragu-build --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/RaguTeam/RAGU.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/ragu-build .claude/skills/ragu-build && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ragu-build
GitHub stars
136
Token cost
~3.8k tokens
SKILL.md length
1,932 words
Files
6 (incl. scripts, references, assets)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Interview the user about their RAGU use case, select an appropriate RAGU pipeline, and generate a validated ragubuild.yaml plus a runnable build<name.py script.

  • Works in 5 steps: Inventory → Interview → Decision summary → …
  • Knowledge Management work in your project
  • SKILL.md covers Skill resources, 0.1 Read the module map, 0.2 Inspect the project and 0.3 Credentials boundary, plus 13 more sections
  • Runs Python scripts from its folder; calls docker

What it does

Ragu Build is an agent skill from RaguTeam/RAGU. Interview the user about their RAGU use case, select an appropriate RAGU pipeline, and generate a validated ragubuild.yaml plus a runnable build<name.py script.

Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts, reference files and assets (for example `.claude-plugin/plugin.json`, `assets/build_template.py` and `references/decision-matrix.md`).

It sits in Knowledge Management. The repository describes itself as: Modular GraphRAG framework. The licence is MIT.

When your agent uses it

  • Knowledge Management work in your project

Example prompts

  • “/ragu-build”

Requirements

  • Python 3
  • Docker

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Inventory
  2. Interview
  3. Decision summary
  4. Generation
  5. Final report

What it can do on your machine

Read from SKILL.md and the folder at commit abd3f29. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ragu Build loads about 3.8k tokens when it runs, and up to ~9.2k if it reads all its reference files. Until then it costs about 44 tokens; SKILL.md has 1,932 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~44
When it runs · the whole SKILL.md, loaded when a task matches
~3.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:116
    * `.env` contents;

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from RaguTeam/RAGU at commit abd3f29, republished under its MIT licence (© RaguTeam). 1,932 words, ~3,838 tokens.

Download SKILL.mdSave it as .claude/skills/ragu-build/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
ragu-build
description
Interview the user about their RAGU use case, select an appropriate RAGU pipeline, and generate a validated ragu_build.yaml plus a runnable build_<name>.py script.

Assembling a RAGU build

Goal: determine what the user actually needs through a short interview, then produce two artifacts:

  • ragu_build.yaml — the recorded decisions and rationale;
  • build_<name>.py — a working RAGU build script.

The RAGU library is already documented: every module ships a README.md describing its role in the pipeline and providing examples.

This skill is a decision process and navigation map, not a retelling of that documentation.

Do not read module sources unless explicitly allowed below. Do not summarize whole READMEs. Read narrowly, guided by references/module-map.md.

Conduct the interview in whatever language the user writes in.

Skill resources

All paths under:

  • references/
  • assets/
  • scripts/

are relative to this skill's own directory, not to the user's project root.

Resolve the skill directory using whatever mechanism the current agent environment provides. Never assume that the skill lives inside the user's repository or that its resources are reachable through paths relative to the current working directory.


Phase 0. Inventory

Before asking the user anything, infer everything that can reasonably be discovered from the project.

This removes unnecessary interview questions.

0.1 Read the module map

Read:

references/module-map.md

Use it only as a navigation map for RAGU components and documentation.

0.2 Inspect the project

Inspect relevant project state in one batch where practical.

Corpus

If the user named a data directory, inspect:

  • file extensions;
  • file count where useful;
  • total size.

For example:

bash
find <dir> -type f | sed 's/.*\.//' | sort | uniq -c
du -sh <dir>

This establishes:

  • approximate corpus size;
  • which file types are present;
  • how much of the corpus RAGU can actually ingest.

Do not recursively inspect document contents unless necessary.

Existing infrastructure

Look for already-running or already-configured storage infrastructure.

Useful signals include:

bash
docker ps

and project-local evidence such as:

  • qdrant_storage/;
  • Qdrant configuration;
  • Neo4j configuration;
  • existing storage-related compose services.

Use only obvious project configuration. Do not search through unrelated files.

Local model capability

If available, run:

bash
nvidia-smi

Use the result only to determine whether local GPU-backed models are realistic.

Failure or absence of nvidia-smi is not itself an error.

0.3 Credentials boundary

Do not search for:

  • API keys;
  • credentials;
  • .env contents;
  • secrets.

Which model provider or runtime the user intends to use is an interview question, not something to discover from private credentials.

0.4 State findings

Before starting the interview, tell the user in one short statement what was discovered and what will be treated as given.

For example:

Your data directory contains about 900 text files totaling ~40 MB. No running vector or graph storage is visible, and a local NVIDIA GPU is available. I'll plan around those facts unless you want to override them.

Anything already established in Phase 0 must not be asked again.

Present inferred facts as correctable facts rather than questions.


Phase 1. Interview

The purpose of the interview is to eliminate incompatible pipeline branches with as few questions as possible.

Interview mechanism

Use the current environment's interactive question mechanism when one is available and appropriate.

Otherwise ask directly in chat.

Do not depend on any specific platform tool name such as AskUserQuestion.

Ask one or two questions at a time.

Never dump the whole interview into a single form or message.

Interview rules

  • Phrase questions using example queries and example data, not RAGU implementation terminology.

  • Avoid terms such as:

    • "extractive";
    • "abstractive";
    • "hybrid retrieval";
    • "community detection";
    • internal RAGU class names.

Technical terminology may appear in the explanation after the user answers, but not unnecessarily in the question itself.

Q1 is the deliberate exception: whether to build a graph is asked explicitly because this decision dominates both cost and architecture, and many users already know whether they want one.

If they do not know, use Q1b.

After each answer, briefly explain which pipeline branches were eliminated.

Order questions by branching factor.

Use at most six primary questions.

If a question's answer follows from:

  • Phase 0;
  • an earlier interview answer;
  • something the user already stated;

skip it.

Missing answers

Never let an unanswered question stall the deliverable.

If the user:

  • skips a question;
  • gives an unusable answer;
  • selects an equivalent of "I'll specify later" but does not specify it;

ask once more.

If the answer is still unavailable:

  1. choose the most defensible default;
  2. clearly state that it is an assumption, not a finding;
  3. continue.

Every such assumption must appear in all relevant outputs.

In the Phase 2 decision table, write:

ASSUMPTION

in the based on what column.

In ragu_build.yaml, mark it with:

yaml
# ASSUMPTION

In build_<name>.py, keep the assumed value as a named constant near the top of the script.

Do not bury assumed values deep inside the implementation.

In the final report, list unresolved assumptions as things the user should verify before the first real run.

An unanswered question should cost one line in the final report.

It must never cost the whole deliverable.


Interview decision order

Exact wording and options live in:

references/decision-matrix.md

section:

Questions

Use that wording when available.

The decision sequence is:

#AboutEliminates
1whether to build a graph, including its costgraph / flat index
1bhow answers are distributed across documents — only if Q1 is "not sure"graph / flat index
2examples of typical queriessearch engine
3exact terms, codes, IDs or part numbers in querieslexical / sparse retrieval
4where models runLLM / embedder client
5corpus size and update patternstorage backends
6extraction quality versus cost — graph builds onlyartifact extractor

Q1 branching

If Q1 selects a vector-only / flat-index build:

  • do not ask Q6;
  • skip graph-specific decisions that no longer matter.

Q1b runs only when Q1 is effectively:

not sure

Never ask both Q1 and Q1b as independent decisions.


Phase 2. Decision summary

Before generating any files, show the user a concise decision table with:

choicewhybased on what

The based on what column must identify the source of each decision:

  • a specific user answer;
  • a Phase 0 finding;
  • ASSUMPTION.

Also explicitly state important components that are not included and why.

For example:

No Global engine — none of the example queries asked for corpus-wide themes or summaries. It can be added later without changing the ingestion strategy.

Do not list every conceivable unused RAGU component. Mention exclusions only when they represent meaningful architectural branches.

Wait for user confirmation before Phase 3.

If the user changes one decision:

  • revise that decision;
  • revise decisions that depend on it;
  • do not restart the whole interview unless the change genuinely invalidates all previous answers.

Phase 3. Generation

After the user confirms the decision summary, generate the build.

3.1 Read the decision matrix

Read:

references/decision-matrix.md

in full.

It is the source of truth for:

  • exact imports;
  • RAGU class names;
  • constructor signatures;
  • stock build structure;
  • component substitutions.

Take every class name and constructor parameter from the matrix rather than from memory.

3.2 Resolve missing component details

If the matrix does not contain enough detail for a selected component:

  1. use references/module-map.md to locate the corresponding RAGU module;
  2. read only the relevant section of that module's README.md.

Do not read the whole README unless the required information cannot otherwise be located.

Do not inspect source code merely for convenience.

Show full SKILL.md (775 more words)Show less
Last-resort signature lookup

If a required constructor detail exists in neither:

  • the decision matrix;
  • the relevant README;

then inspect only the real constructor signature.

For example:

bash
grep -n "def __init__" -A 25 <file>

Use this only as a last resort.

Do not explore implementation internals.

Never invent constructor parameters.


3.3 Generate ragu_build.yaml

Write:

ragu_build.yaml

It must record the selected build decisions.

For every meaningful decision include:

  • selected option;
  • one-line rationale;
  • enough information to reconstruct why the pipeline was chosen.

Keep rejected architectural alternatives out of the Python script; they belong here when worth recording.

Mark unresolved defaults explicitly:

yaml
# ASSUMPTION

3.4 Generate build_<name>.py

Start from:

assets/build_template.py

The template represents stock build B:

graph + local search.

Adapt it to the decisions selected during the interview.

Use:

  • Part 2 of references/decision-matrix.md for exact signatures;
  • Part 3 for the shape of other stock builds.

A comparable hand-written example is:

examples/extract_with_llm_and_local_search.py

when it exists in the user's RAGU checkout.

Script style

Keep the generated script intentionally flat.

Preferred structure:

  1. imports;
  2. named configuration constants;
  3. async def main(...);
  4. top-to-bottom construction and execution;
  5. if __name__ == "__main__" guard.

Do not create helper functions that are called only once unless they materially improve correctness.

Do not include:

  • commented-out alternatives;
  • unused imports;
  • dead code;
  • speculative components;
  • rejected pipeline choices.

Keep choices that were considered but not selected in ragu_build.yaml.

Assumptions

Any unresolved assumption must remain visible as a named constant near the top.

For example:

python
# ASSUMPTION: user did not specify the collection name.
COLLECTION_NAME = "ragu"

Indexing safety

The validator runs main() for real.

Therefore:

  • construct components inside main();
  • keep expensive corpus indexing behind an explicit flag;
  • do not index the corpus unconditionally at import time;
  • do not initiate expensive work merely by importing the generated module.

The generated script is for the user to run intentionally.


3.5 Validate the build

The validator belongs to this skill at:

scripts/validate_build.py

Resolve this path relative to the skill's own directory.

Never assume it is relative to the user's project.

Choose Python

Run validation using a Python interpreter capable of importing the user's ragu installation.

Prefer, in order:

  1. an active $VIRTUAL_ENV;
  2. project .venv/bin/python;
  3. project venv/bin/python;
  4. python3.

Confirm that the selected interpreter can import RAGU before relying on the validator.

A typical check is:

bash
if [ -n "$VIRTUAL_ENV" ] && [ -x "$VIRTUAL_ENV/bin/python" ]; then
    PY="$VIRTUAL_ENV/bin/python"
elif [ -x ".venv/bin/python" ]; then
    PY=".venv/bin/python"
elif [ -x "venv/bin/python" ]; then
    PY="venv/bin/python"
else
    PY="python3"
fi

"$PY" -c "import ragu"

Then run:

bash
"$PY" <skill-dir>/scripts/validate_build.py build_<name>.py

Use the actual resolved skill directory in place of <skill-dir>.

Validation failure

If no available interpreter can import ragu:

  • say so clearly;
  • do not pretend validation succeeded;
  • do not claim the generated build is checked.

Stop trying to execute the validator.

The files may still be generated, but the final report must identify validation as blocked.

Validator behavior

The validator:

  • checks RAGU calls against real signatures;
  • runs the generated script's main();
  • replaces external/network/model calls with local stand-ins.

Validation must not intentionally send requests to external model providers or execute real model inference.

Fix every validator error caused by the generated script before reporting success.

Do not suppress or ignore validator failures merely to finish the task.


Phase 4. Final report

Report the result concisely.

Open with anything that blocks the first real run, including:

  • unresolved assumptions;
  • placeholder values;
  • missing Python environment;
  • services that need to be running;
  • model/provider configuration still required.

Then state:

  • where ragu_build.yaml was written;
  • where build_<name>.py was written;
  • whether validation passed;
  • how to run the generated script;
  • approximate first-build cost;
  • approximate first-build duration;
  • the first things to adjust if answer quality is poor.

Cost and duration estimates must be presented as estimates, with the assumptions behind them.

Do not imply precision that the available corpus size, model provider, hardware, or extraction strategy does not support.


Boundaries

Input formats

RAGU ingests plain text and nothing else.

RAGU does not itself provide:

  • PDF parsing;
  • office-document parsing;
  • OCR;
  • ASR;
  • image understanding;
  • video transcription.

If the user's corpus contains:

  • PDFs;
  • Word documents;
  • presentations;
  • images;
  • audio;
  • video;
  • other non-text formats;

say plainly that RAGU cannot ingest those files directly.

Converting them to text is a separate preprocessing step outside the build produced by this skill.

Never generate a build that pretends unsupported files can be read directly.

Library modifications

This skill produces:

  • configuration;
  • a runnable script.

It does not modify the RAGU library itself.

Do not patch RAGU source code as part of this workflow.

Expensive execution

Do not:

  • build the real graph;
  • index the full corpus;
  • call paid models;
  • burn tokens;
  • start large extraction jobs;

unless the user explicitly asks for execution.

Generating and locally validating the script is allowed.

The resulting script is the user's build to run.

Source inspection

Prefer information sources in this order:

  1. references/decision-matrix.md;
  2. references/module-map.md;
  3. relevant narrow sections of module README.md;
  4. exact constructor signature inspection as a last resort.

Do not browse RAGU implementation source for architecture understanding.

Never invent APIs, class names, arguments, or constructor parameters.

© RaguTeam, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references, assets) in .agents/skills/ragu-build of RaguTeam/RAGU.

  • SKILL.md
  • .claude-plugin/plugin.json
  • assets/build_template.py
  • references/decision-matrix.md
  • references/module-map.md
  • scripts/validate_build.py

Open the folder on GitHubat commit abd3f29

Compare with similar skills

Ragu Build next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ragu Build compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ragu Build this skillRaguTeam/RAGU136—~3.8kAutomated safety check: NotesMIT
Qmdalsk1992/CloddsBot3k3 repos~1.2kAutomated safety check: PassMIT
Gitnexus CLIaws-samples/sample-kolya-br-proxy10611 repos~822Automated safety check: PassMIT-0
Open Notebookagent-skills-hub/agent-skills-hub1123 repos~2.6kAutomated safety check: PassMIT
Gate Checkundefined-ui/second-brain-os1k—~854Automated safety check: PassMIT
Goal Testundefined-ui/second-brain-os1k—~754Automated safety check: PassMIT

Similar skills

  • Qmd

    alsk1992/CloddsBot

    Local hybrid search for markdown notes and docs. An agent skill from alsk1992/CloddsBot.

    3k GitHub starsUsed in 3 repos~1.2k tokens
    Knowledge ManagementAuto-check passed
  • Gitnexus CLI

    aws-samples/sample-kolya-br-proxy

    Official

    A skill your agent uses when the user needs to run GitNexus CLI commands like analyze/index a repo, check status, clean the index, generate a wiki, or list indexed repos.

    106 GitHub starsUsed in 11 repos~822 tokens
    Knowledge ManagementAuto-check passed
  • Open Notebook

    agent-skills-hub/agent-skills-hub

    Drives a self-hosted Open Notebook instance to organize sources into notebooks, chat with documents, generate notes and multi-speaker podcasts, and search across material.

    112 GitHub starsUsed in 3 repos~2.6k tokens
    Knowledge ManagementAuto-check passed
  • Gate Check

    undefined-ui/second-brain-os

    Find the decisions in a pipeline that do not need the expensive model and propose the gate for each: a rule, a classic classifier, or a small model, with fail-closed routing.

    1k GitHub stars~854 tokensUpdated yesterday
    Knowledge ManagementAuto-check passed
  • Goal Test

    undefined-ui/second-brain-os

    Turn a vague task into a testable definition of done and generate an executable goal-test script for it, optionally with a bounded retry loop around a headless agent.

    1k GitHub stars~754 tokensUpdated yesterday
    Knowledge ManagementAuto-check passed
  • Using Mineecho Skills

    Health-Yang/MineEcho

    技能系统导航。当用户询问"你能做什么"、"有什么功能"、"有什么技能"或不确定如何完成某个任务时,使用此技能列出所有可用技能并建议最合适的技能。此技能是每个对话开始时默认加载的,用于技能发现。

    220 GitHub stars~124 tokensUpdated 4 mo ago
    Knowledge ManagementAuto-check passed

Questions about Ragu Build

What does Ragu Build do?

Interview the user about their RAGU use case, select an appropriate RAGU pipeline, and generate a validated ragubuild.yaml plus a runnable build<name.py script. Ragu Build is an agent skill from RaguTeam/RAGU.py script.

When should I use Ragu Build?

Ragu Build fits situations like: knowledge Management work in your project.

How do I install Ragu Build in Claude Code?

Run `npx skills add RaguTeam/RAGU --skill ragu-build -a claude-code`. Or copy the skill folder (.agents/skills/ragu-build in RaguTeam/RAGU) into .claude/skills/ragu-build in your project. Claude Code loads it when a task matches its description.

How do I install Ragu Build in Codex?

Run `npx skills add RaguTeam/RAGU --skill ragu-build -a codex`. Or copy the skill folder (.agents/skills/ragu-build in RaguTeam/RAGU) into .agents/skills/ragu-build in your project. Codex loads it when a task matches its description.

Can I use Ragu Build in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add RaguTeam/RAGU --skill ragu-build -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ragu-build, .gemini/skills/ragu-build, .github/skills/ragu-build and .opencode/skills/ragu-build in your project.

What does Ragu Build need to run?

Going by SKILL.md and its folder, Ragu Build needs Python for the scripts in its folder and the command-line tools its instructions call (docker). Our summary lists: Python 3; Docker.

Does Ragu Build access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Ragu Build safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Ragu Build use?

Ragu Build is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ragu Build use?

About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.4k tokens, read only when the agent opens those files.

What are the alternatives to Ragu Build?

Skills that share tags, products or a category with Ragu Build: Qmd (alsk1992/CloddsBot, 3k stars), Gitnexus CLI (aws-samples/sample-kolya-br-proxy, 106 stars), Open Notebook (agent-skills-hub/agent-skills-hub, 112 stars) and Gate Check (undefined-ui/second-brain-os, 1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ragu Build?

RaguTeam (a GitHub organization) maintains it in RaguTeam/RAGU, which has 136 GitHub stars. The repository was last updated on October 4, 2026.

Source: RaguTeam/RAGU on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.