Agent skill

Creating Kb

by oaustegard in oaustegard/claude-skills

Builds a portable, embedding-free knowledgebase from a set of files and delivers it as a self-contained .skill bundle (BM25 index + bundled searcher + query protocol).

MITAuto-check passedAI & LLM Engineering

Install Creating Kb

skills CLI
$ npx skills add oaustegard/claude-skills --skill creating-kb -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install oaustegard/claude-skills creating-kb --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/creating-kb .claude/skills/creating-kb && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
creating-kb
GitHub stars
150
Token cost
~1.5k tokens
SKILL.md length
671 words
Files
8 (incl. scripts)
Skills in repo
66
Repo updated
First seen
Licence
MIT

At a glance

Builds a portable, embedding-free knowledgebase from a set of files and delivers it as a self-contained .skill bundle (BM25 index + bundled searcher + query protocol).

  • Works in 3 steps: Gather the sources → Build the bundle → Deliver
  • A user wants to turn uploaded files
  • SKILL.md covers Workflow, Choosing chunk size, Verifying the bundle and What ships in the bundle, plus 1 more section
  • Runs JavaScript and Python scripts from its folder; calls node and npm

What it does

Creating Kb is an agent skill from oaustegard/claude-skills. Builds a portable, embedding-free knowledgebase from a set of files and delivers it as a self-contained .skill bundle (BM25 index + bundled searcher + query protocol). Use when a user wants to turn uploaded files, a folder, or a corpus into a searchable knowledgebase they can hand to any agent — phrased as "make a knowledgebase", "build a KB skill", "package these docs for retrieval", "create a searchable bundle", or references to a .skill KB. The output runs anywhere with Node or Python — no model, no install…

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts (for example `CHANGELOG.md`, `scripts/build_lexkb.js` and `scripts/bundle_SKILL.md`).

It sits in AI & LLM Engineering, covering Embeddings. It works with GitHub and Python. The repository describes itself as: My collection of Claude skills. The licence is MIT.

When your agent uses it

  • A user wants to turn uploaded files
  • A corpus into a searchable knowledgebase they can hand to any agent — phrased as make a knowledgebase
  • Build a KB skill
  • Package these docs for retrieval

Example prompts

  • “make a knowledgebase”
  • “build a KB skill”
  • “package these docs for retrieval”
  • “/creating-kb”

Requirements

  • Python 3
  • Node.js

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Gather the sources
  2. Build the bundle
  3. Deliver

What it can do on your machine

Read from SKILL.md and the folder at commit cf49d47. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 5 files in scripts/ (JavaScript and Python), which the agent can run.

    Shell commands in SKILL.md call:

    • node
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Creating Kb loads about 1.5k tokens when it runs. Until then it costs about 165 tokens; SKILL.md has 671 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~165
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from oaustegard/claude-skills at commit cf49d47, republished under its MIT licence (© oaustegard). 671 words, ~1,494 tokens.

Download SKILL.mdSave it as .claude/skills/creating-kb/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
creating-kb
description
Builds a portable, embedding-free knowledgebase from a set of files and delivers it as a self-contained `.skill` bundle (BM25 index + bundled searcher + query protocol). Use when a user wants to turn uploaded files, a folder, or a corpus into a searchable knowledgebase they can hand to any agent — phrased as "make a knowledgebase", "build a KB skill", "package these docs for retrieval", "create a searchable bundle", or references to a `.skill` KB. The output runs anywhere with Node or Python — no model, no install, no network. Distinct from `bm25` (ephemeral in-session search) and `building-github-index` (markdown project-knowledge index).
metadata.version
0.2.0

creating-kb

Turn a pile of files into a portable, deployable knowledgebase. The output is a .skill bundle — an ordinary zip — containing a BM25 inverted index, the chunk text, a pure-Node searcher, and a query protocol. It has no embedding model and no semantic search: retrieval is lexical, and the consuming agent supplies the semantic layer by expanding the query at search time. That is what makes the bundle portable — any agent that can run node can query it with no npm install, no model download, and no network.

The whole toolchain is JavaScript so one implementation serves both this builder and the in-browser packer. Build with the bundled script; do not hand-roll the index.

SCRIPTS=/mnt/skills/user/creating-kb/scripts
node $SCRIPTS/build_lexkb.js CORPUS_DIR --out /tmp/kb --name my-kb --zip

Workflow

1. Gather the sources

Collect the files into one directory. In a Claude.ai chat, uploads land in /mnt/user-data/uploads/ — point the builder there. Otherwise use any path the user names. Supported extensions default to txt,md,html,htm; pass --ext to change them.

This MVP interface is bounded by how many files a chat can accept. For a large corpus, stage the files in a directory first, or use the browser packer (built from the same scripts) that runs entirely client-side.

2. Build the bundle
bash
SCRIPTS=/mnt/skills/user/creating-kb/scripts
node $SCRIPTS/build_lexkb.js /mnt/user-data/uploads \
  --out /tmp/kb --name my-kb --zip \
  --source "human description of the corpus"

The script chunks each file, builds the BM25 index, writes the bundle dir (SKILL.md + search.js + index.json + chunks.jsonl), and — with --zip — emits my-kb.skill next to --out.

3. Deliver

Move the .skill to the outputs directory and give the user a download link:

bash
cp /tmp/my-kb.skill /mnt/user-data/outputs/
markdown
[Download my-kb.skill](computer:///mnt/user-data/outputs/my-kb.skill)

Tell the user how to deploy it: unzip into an agent's skill directory (or upload it as a skill). The bundle's own SKILL.md then drives querying — the consuming agent reads it, expands each question into search terms, and runs the bundled search.js. No further setup.

Choosing chunk size

The retrieval unit and the reasoning unit are decoupled, which makes chunk size a low-stakes choice. search.py/search.js rank on the whole chunk (best recall) but return only the query-densest passage of it by default (--snippet, ~1200 chars), so a big chunk does not flood the consuming agent's context with surrounding noise. Index for recall; the searcher handles signal.

--target-chars controls chunk size (whole paragraphs are packed up to the target; --target-chars 0 makes each file one chunk). Lexical BM25 tolerates — and on a real-corpus sweep slightly preferred — larger chunks than embedding-based retrieval, because there is no vector to dilute: BM25 scores individual term presence with length normalization, so a big chunk still ranks on the exact terms it contains.

  • Default: --target-chars 0 (whole document). Best recall, fewest chunks; the snippet return keeps reasoning context focused.
  • Long, multi-topic files where you want tighter citation units: 1500–4000.
  • 500 only if you need very fine-grained chunk ids and accept more chunks.
Show full SKILL.md (228 more words)Show less

Verifying the bundle

Test before delivering. Run a query against the freshly built bundle and confirm it returns sensible hits:

bash
node /tmp/kb/search.js --query "a representative question" \
  --core "key term" --expand "synonym" --k 3

Each hit's text is the query-focused passage by default; add --snippet 0 to inspect a full chunk.

search.js prints JSON {"hits": [...]}. Confirm the right chunks surface.

What ships in the bundle

FileRole
SKILL.mdthe query protocol the consuming agent follows (expand → search → cite)
search.js / search.pyequivalent BM25 + RM3 + metadata-filter searchers; return query-focused passages (matched sentences kept in neighbour context, merged); the agent runs whichever runtime it has
index.jsonprecomputed inverted index (postings, df, doc lengths, BM25 params)
chunks.jsonlchunk text + structured metadata

Both searchers are thin readers of the same neutral JSON index, so the bundle runs in a Node-only or a Python-only consumer. Metadata stays structured (not folded into the indexed text), which lets the consuming agent filter on it (--filter section=blog, --filter date>=2025).

Scripts

  • scripts/build_lexkb.js — chunker + BM25 index builder + .skill writer.
  • scripts/search.js — the JS runtime searcher, copied verbatim into every bundle. It owns the tokenizer; the builder imports it so index and queries tokenize identically.
  • scripts/search.py — the Python runtime searcher, copied verbatim into every bundle; a thin reader of the same neutral JSON index, parity-pinned to search.js (identical results on a shared index).
  • scripts/zipstore.js — pure-JS ZIP-STORED writer (used by the builder; shared with the in-browser packer).
  • scripts/bundle_SKILL.md — the query-side SKILL.md template written into each bundle.

© oaustegard, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts) in creating-kb of oaustegard/claude-skills.

  • SKILL.md
  • CHANGELOG.md
  • scripts/build_lexkb.js
  • scripts/bundle_SKILL.md
  • scripts/search.js
  • scripts/search.py
  • scripts/zipstore.js
  • test_parity.py

Open the folder on GitHubat commit cf49d47

Compare with similar skills

Creating Kb next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Creating Kb compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Creating Kb this skilloaustegard/claude-skills150—~1.5kAutomated safety check: PassMIT
Copilot SDKgithub/awesome-copilot40k4 repos~6.3kAutomated safety check: PassMIT
Multimodal Embedding Serving Useropen-edge-platform/edge-ai-libraries169—~1.7kAutomated safety check: PassApache-2.0
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k7 repos~2.3kAutomated safety check: PassMIT
Embeddings via 9Routerdecolua/9router30k—~604Automated safety check: PassMIT
Esmfold2JimLiu/science-skills2274 repos~2.5kAutomated safety check: PassApache-2.0

Similar skills

  • Copilot SDK

    github/awesome-copilot

    Official

    Build agentic applications with GitHub Copilot SDK. An agent skill from github/awesome-copilot.

    40k GitHub starsUsed in 4 repos~6.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Multimodal Embedding Serving User

    open-edge-platform/edge-ai-libraries

    Deploy and consume the Multimodal Embedding Serving microservice — bring it up with setup.sh + docker compose (from a repo clone, or by fetching those same files from GitHub when no clone exists)…

    169 GitHub stars~1.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Embeddings via 9Router

    decolua/9router

    Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.

    30k GitHub stars~604 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    227 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Precheck PR

    vllm-project/vllm-omni

    Self-check your branch before creating a PR — catch dead code, prevent new model-specific Python examples, verify accuracy/perf claims, validate PR title format, and confirm merge readiness.

    7.1k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from oaustegard/claude-skills

All 66 skills in this repo
  • Vega-Lite Interactive Charts

    oaustegard/claude-skills

    Builds interactive Vega-Lite charts from uploaded data: analyzes the fields, picks five to ten fitting chart types, and produces a React artifact with the data embedded inline.

    150 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Single-File HTML Composer

    oaustegard/claude-skills

    Builds self-contained single-file HTML pages such as reports, decks, postmortems, flowcharts and prototypes from a small spec using a bundled Python composer and templates.

    150 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Declauding

    oaustegard/claude-skills

    Rewrites model-sounding prose into plain technical writing and checks that every claim survives, for PR text, docs, commit messages and similar drafts.

    150 GitHub stars~5.2k tokensUpdated yesterday
    Auto-check passed
  • Preact Developer

    oaustegard/claude-skills

    Guides building standards-based Preact apps with native-first choices, HTM syntax, import maps and vendored ESM, from single-file demos to larger builds.

    150 GitHub stars~4.6k tokensUpdated yesterday
    Auto-check passed
  • Bluesky Zeitgeist Sampler

    oaustegard/claude-skills

    Deprecated sampler that captures short windows of the Bluesky firehose, clusters trending terms and builds an HTML report; replaced by the browsing-bluesky skill.

    150 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Adversarial Review Before Shipping

    oaustegard/claude-skills

    Has a fresh-context adversary attack a blog post, recommendation, analysis brief or piece of code before you ship it, using a profile suited to that kind of artifact.

    150 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Creating Kb

What does Creating Kb do?

Builds a portable, embedding-free knowledgebase from a set of files and delivers it as a self-contained .skill bundle (BM25 index + bundled searcher + query protocol). Creating Kb is an agent skill from oaustegard/claude-skills.skill bundle (BM25 index + bundled searcher + query protocol).

When should I use Creating Kb?

Creating Kb fits situations like: A user wants to turn uploaded files; A corpus into a searchable knowledgebase they can hand to any agent — phrased as make a knowledgebase; build a KB skill; package these docs for retrieval.

How do I install Creating Kb in Claude Code?

Run `npx skills add oaustegard/claude-skills --skill creating-kb -a claude-code`. Or copy the skill folder (creating-kb in oaustegard/claude-skills) into .claude/skills/creating-kb in your project. Claude Code loads it when a task matches its description.

How do I install Creating Kb in Codex?

Run `npx skills add oaustegard/claude-skills --skill creating-kb -a codex`. Or copy the skill folder (creating-kb in oaustegard/claude-skills) into .agents/skills/creating-kb in your project. Codex loads it when a task matches its description.

Can I use Creating Kb in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oaustegard/claude-skills --skill creating-kb -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/creating-kb, .gemini/skills/creating-kb, .github/skills/creating-kb and .opencode/skills/creating-kb in your project.

What does Creating Kb need to run?

Going by SKILL.md and its folder, Creating Kb needs JavaScript and Python for the scripts in its folder and the command-line tools its instructions call (node and npm). Our summary lists: Python 3; Node.js.

Does Creating Kb access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Creating Kb safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Creating Kb use?

Creating Kb is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Creating Kb use?

About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Creating Kb?

Skills that share tags, products or a category with Creating Kb: Copilot SDK (github/awesome-copilot, 40k stars), Multimodal Embedding Serving User (open-edge-platform/edge-ai-libraries, 169 stars), Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Embeddings via 9Router (decolua/9router, 30k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Creating Kb?

oaustegard (a GitHub user) maintains it in oaustegard/claude-skills, which has 150 GitHub stars. The repository holds 66 skills in this directory. The repository was last updated on October 8, 2026.

Source: oaustegard/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.