Wiki Ingest
paperclipai/paperclip
A skill your agent uses when an operation issue asks to ingest a captured raw/ source into the LLM Wiki, or the user says "ingest <slug".
Ingest a large dataset into memory as a skimmed map. An agent skill from vellum-ai/vellum-assistant.
$ npx skills add vellum-ai/vellum-assistant --skill memory-corpus-ingest -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install vellum-ai/vellum-assistant memory-corpus-ingest --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/memory-corpus-ingest .claude/skills/memory-corpus-ingest && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "memory-corpus-ingest" agent skill from https://github.com/vellum-ai/vellum-assistant/tree/main/skills/memory-corpus-ingest into .claude/skills/memory-corpus-ingest/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "memory-corpus-ingest", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/vellum-ai/vellum-assistant/tree/main/skills/memory-corpus-ingestType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add vellum-ai/vellum-assistant --skill memory-corpus-ingest -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install vellum-ai/vellum-assistant memory-corpus-ingest --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/memory-corpus-ingest .agents/skills/memory-corpus-ingest && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "memory-corpus-ingest" agent skill from https://github.com/vellum-ai/vellum-assistant/tree/main/skills/memory-corpus-ingest into .agents/skills/memory-corpus-ingest/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "memory-corpus-ingest", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vellum-ai/vellum-assistant --skill memory-corpus-ingest -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install vellum-ai/vellum-assistant memory-corpus-ingest --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/memory-corpus-ingest .cursor/skills/memory-corpus-ingest && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "memory-corpus-ingest" agent skill from https://github.com/vellum-ai/vellum-assistant/tree/main/skills/memory-corpus-ingest into .cursor/skills/memory-corpus-ingest/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "memory-corpus-ingest", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/vellum-ai/vellum-assistant.git --path skills/memory-corpus-ingest--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add vellum-ai/vellum-assistant --skill memory-corpus-ingest -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install vellum-ai/vellum-assistant memory-corpus-ingest --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/memory-corpus-ingest .gemini/skills/memory-corpus-ingest && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "memory-corpus-ingest" agent skill from https://github.com/vellum-ai/vellum-assistant/tree/main/skills/memory-corpus-ingest into .gemini/skills/memory-corpus-ingest/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "memory-corpus-ingest", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install vellum-ai/vellum-assistant memory-corpus-ingestInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add vellum-ai/vellum-assistant --skill memory-corpus-ingest -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/memory-corpus-ingest .github/skills/memory-corpus-ingest && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "memory-corpus-ingest" agent skill from https://github.com/vellum-ai/vellum-assistant/tree/main/skills/memory-corpus-ingest into .github/skills/memory-corpus-ingest/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "memory-corpus-ingest", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vellum-ai/vellum-assistant --skill memory-corpus-ingest -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install vellum-ai/vellum-assistant memory-corpus-ingest --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/memory-corpus-ingest .opencode/skills/memory-corpus-ingest && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "memory-corpus-ingest" agent skill from https://github.com/vellum-ai/vellum-assistant/tree/main/skills/memory-corpus-ingest into .opencode/skills/memory-corpus-ingest/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "memory-corpus-ingest", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
memory-corpus-ingestIngest a large dataset into memory as a skimmed map. An agent skill from vellum-ai/vellum-assistant.
Memory Corpus Ingest is an agent skill from vellum-ai/vellum-assistant. Ingest a large dataset into memory as a skimmed map. Cold-store the raw files under a workspace imports directory, census them into a slice plan, skim each slice into compact map pages that point back at the raw files, ingest the map with the memory ingest CLI, and author a drill-in retrieval skill so the corpus stays searchable on demand. For recording archives, transcript collections, document dumps, and any corpus too large to hold in memory directly.
Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/fathom.md`, `references/map-page-template.md` and `scripts/inventory.test.ts`). Compatibility notes: Designed for Vellum personal assistants
The repository describes itself as: An AI Assistant that’s easy to setup, does your work 24/7, knows your preferences and gets better over time. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 33cc983. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (TypeScript), which the agent can run.
Shell commands in SKILL.md call:
rgrsyncbunFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use rsync, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Designed for Vellum personal assistants
From compatibility in the SKILL.md frontmatter.
Memory Corpus Ingest loads about 3k tokens when it runs, and up to ~6.5k if it reads all its reference files. Until then it costs about 120 tokens; SKILL.md has 1,343 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
find /path/to/raw-corpus \( -name '.env*' -o -name '*.key' -o -name '*.pem' \rsync -a --exclude='.env*' --exclude='*.key' --exclude='*.pem' \Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from vellum-ai/vellum-assistant at commit 33cc983, republished under its MIT licence (© vellum-ai). 1,343 words, ~3,015 tokens.
.claude/skills/memory-corpus-ingest/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.Bring a large dataset into the assistant's working knowledge without stuffing it into memory. The model is a library: the workspace holds the stacks (the raw files, cold and complete), memory holds the card catalog (a small set of map pages that say what exists, when it is from, and where to look), and a purpose-built retrieval skill is the librarian that walks to the right shelf on demand.
Two invariants drive everything below:
memory/concepts/ except the map pages, and nothing is ever appended to memory/buffer.md (bulk buffer appends trip the consolidation burst guard and per-run caps; the map bypasses the buffer entirely via assistant memory ingest).Identify the source and its size before committing:
du -sh /path/to/raw-corpus
find /path/to/raw-corpus -type f | wc -lTell the user what will happen: the raw files move into the workspace, a bounded number of summarization passes read them once to build the map, the map is ingested into memory, and a lookup skill is authored for drill-in. Skimming a large corpus is real LLM work that costs time and money; confirm before starting. For Fathom recording exports, read references/fathom.md first for format discovery and slicing guidance.
Land the raw files under an imports directory in the workspace, one directory per source:
Screen for credentials BEFORE copying: an arbitrary corpus can carry secret
material, and anything landed under imports/ becomes reachable by workspace
tools, backups, and retrieval flows.
cd "$VELLUM_WORKSPACE_DIR"
# 1a. Screen for secret-bearing FILE NAMES; review every hit with the user.
find /path/to/raw-corpus \( -name '.env*' -o -name '*.key' -o -name '*.pem' \
-o -name '*credential*' -o -name '*secret*' -o -name 'cookies*' \
-o -path '*tokens*' -o -path '*oauth*' \) -print
# 1b. Screen file CONTENTS for credential shapes. --hidden and --no-ignore
# matter: rg skips dotfiles and gitignored paths by default, which is
# exactly where credentials live. Capture the FULL list (no truncation):
# every file named here must be excluded below or cleaned with the user
# before it lands.
rg -l -i --hidden --no-ignore \
"api[_-]?key|access[_-]?token|refresh[_-]?token|client[_-]?secret|password\s*[=:]|passwd|bearer |AKIA[0-9A-Z]{16}|BEGIN [A-Z ]*PRIVATE KEY" \
/path/to/raw-corpus > /tmp/corpus-secret-hits.txt
cat /tmp/corpus-secret-hits.txt
# 2. Build rsync exclusions from the content hits (paths relative to the
# corpus root), then copy with ALL flagged paths excluded.
sed 's|^/path/to/raw-corpus/||' /tmp/corpus-secret-hits.txt > /tmp/corpus-secret-exclusions.txt
mkdir -p imports/<source>
rsync -a --exclude='.env*' --exclude='*.key' --exclude='*.pem' \
--exclude='tokens/' --exclude='oauth/' --exclude='cookies*' \
--exclude-from=/tmp/corpus-secret-exclusions.txt \
/path/to/raw-corpus/ imports/<source>/Rules:
memory/. The cold store is imports/<source>/; the map is the only thing that enters memory.Census the corpus and produce a machine-readable slice plan:
mkdir -p "$VELLUM_WORKSPACE_DIR/imports/<source>/.staging"
bun run {baseDir}/scripts/inventory.ts "$VELLUM_WORKSPACE_DIR/imports/<source>" \
> "$VELLUM_WORKSPACE_DIR/imports/<source>/.staging/plan.json"The script prints a human census to stderr (file count, total size, extension mix, date range) and a JSON plan to stdout: { files, totalBytes, byExtension, dateRange, suggestedSlices }. Each suggested slice is a date-windowed group of files sized for one skim pass. Review the plan before skimming:
For each slice in the plan, run one summarization pass that reads the slice's files and writes one staged map page:
<slug>.md (for example imports/fathom/.staging/fathom-recordings-2025-q1.md). The slug is the filename minus .md.references/map-page-template.md exactly: lead that stands alone as the retrieval card, a Raw data: pointer line in the lead, ## sections per topic or time slice, and source: / origin_date: / ref_files: / links: frontmatter.kind: index) whose links: enumerate all slice pages.Bound the fan-out. The number of skim passes equals the number of slices in the plan, full stop. Run them a few at a time (parallelism of 3 or 4 is plenty). Never spawn one pass per file, never let a pass recursively spawn more passes, and never re-skim slices that already have a staged page unless their content changed. An unbounded fan-out over a large corpus is the expensive failure mode of this skill.
A skim pass extracts what a future search needs to route: decisions, recurring topics, named people and projects, date spans, open threads. It does not transcribe. Verbatim content stays cold; the map records that it exists and where.
Dry-run first, review, then ingest for real:
cd "$VELLUM_WORKSPACE_DIR"
assistant memory ingest --dir imports/<source>/.staging --dry-runThe dry run validates every page and reports per-page results without writing. Fix anything reported invalid (bad frontmatter, bad slug), and resolve any warning about a links: or [[wikilink]] target that is neither on disk nor in the staged set (stage the missing page, or make the reference plain prose; retrieval drops a link whose target page does not exist), then run without --dry-run. Notes:
memory.v3.live or memory.v2.enabled).--overwrite is passed; use --overwrite when re-running after edits.After ingest, the maintain job picks up the new pages and embeds their sections; the map becomes retrievable shortly after, with each page's card dated by its origin_date rather than the import day.
The map tells the model a slice exists; the drill-in skill is how it follows the pointer. Author a workspace skill for this corpus under $VELLUM_WORKSPACE_DIR/skills/<source>-lookup/:
SKILL.md with frontmatter per the Agent Skills spec (name matching the directory, keyword-rich description) and activation hints drawn from the intents you actually observed while scoping (for example "when the user asks what was said in a meeting", "when a question needs the <source> archive").scripts/ with actually runnable search commands over the cold store, not prose descriptions of searching. A minimal helper is a ripgrep wrapper scoped to the imports directory with date filtering:#!/usr/bin/env bash
# search.sh <pattern> [YYYY-MM] scoped search over the <source> cold store
DIR="$VELLUM_WORKSPACE_DIR/imports/<source>"
if [ -n "$2" ]; then
# Two globs: a slash-free pattern matches the date in any basename, and
# the "**/" prefixed pattern matches the date in a directory segment at
# any depth (a leading "*" cannot cross path separators).
rg -i -C 3 "$1" "$DIR" --glob "*$2*" --glob "**/*$2*/**"
else
rg -i -C 3 "$1" "$DIR"
fiConfirm the new skill appears in assistant skills list (workspace skills are picked up from the workspace skills/ directory).
Run 3 to 5 representative queries a real user would ask of this corpus (mix a routing question, a specific-fact question, and a date-scoped question). For each, check:
Raw data: pointer or invokes the drill-in skill to reach the actual files.If a query routes to nothing, the relevant lead is not doing its card job; rewrite it and re-ingest that page with --overwrite. For large ingests, assistant memory v3 eval can additionally gate the change, but it is a two-corpus comparison requiring --snapshot, --staging, and --out; use it only when you captured a pre-ingest snapshot of memory/concepts/ to compare against. Otherwise the query checks above are the verification.
memory/concepts/ or memory/buffer.md. Map pages only, via assistant memory ingest.imports/<source>/ and is read-only after landing.Raw data: line) live in the page lead, because the card renderer injects the lead plus a section list; a pointer buried in a section is invisible at routing time.origin_date: is the content's own date, never the import date.references/map-page-template.md: the exact article shape for map pages, with a full example. Read before writing any staged page.references/fathom.md: Fathom recording archives specifically; export discovery, speaker and date extraction, slicing, and what belongs in the map versus the cold store.© vellum-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (scripts, references) in skills/memory-corpus-ingest of vellum-ai/vellum-assistant.
Open the folder on GitHubat commit 33cc983
Memory Corpus Ingest next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Memory Corpus Ingest this skillvellum-ai/vellum-assistant | 1.4k | — | ~3k | Automated safety check: Notes | MIT | |
| Wiki Ingestpaperclipai/paperclip | 99k | — | ~933 | Automated safety check: Pass | MIT | |
| Token Mapnexu-io/open-design | 100k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| DatasetsArize-ai/phoenix | 12k | — | ~1.6k | Automated safety check: Pass | Custom licence | |
| Maps Geographyasgeirtj/system_prompts_leaks | 69k | — | ~717 | Automated safety check: Pass | CC0-1.0 | |
| Market Ingestruvnet/ruflo | 74k | — | ~529 | Automated safety check: Notes | MIT |
paperclipai/paperclip
A skill your agent uses when an operation issue asks to ingest a captured raw/ source into the LLM Wiki, or the user says "ingest <slug".
nexu-io/open-design
Map an extracted Figma / source-code token bag onto the active OD design system, producing a deterministic mapping the generate stage can consume.
Arize-ai/phoenix
Understand what a Phoenix dataset is and reason well about its examples, outputs, splits, and how it feeds evaluators and experiments.
asgeirtj/system_prompts_leaks
Accurate maps from real geo data — use for any map, or whenever geography would make a good graphic for a deliverable
ruvnet/ruflo
Ingest and normalize market data into OHLCV vectors with HNSW indexing
onyx-dot-app/onyx
Use the Onyx feature map (.agents/feature-map/) to learn what a product surface does, the code behind it, and what a change can break.
vellum-ai/vellum-assistant
Create and configure a GitHub App so the assistant can push commits, open PRs, and comment under its own bot identity.
vellum-ai/vellum-assistant
Connect a Discord bot to the assistant via the Discord Gateway with guided application creation and intent configuration
vellum-ai/vellum-assistant
Create and configure a Sentry internal integration so the assistant can manage issues, alerts, and releases under its own identity
vellum-ai/vellum-assistant
A skill your agent uses when the user wants to build, scaffold, ship, or edit a Vellum plugin that bundles multiple surfaces (hooks, tools, skills, and more) into one installable package.
vellum-ai/vellum-assistant
Connect a Slack app to the Vellum Assistant via Socket Mode.
vellum-ai/vellum-assistant
Shop on Amazon and Amazon Fresh through your browser. An agent skill from vellum-ai/vellum-assistant.
Ingest a large dataset into memory as a skimmed map. An agent skill from vellum-ai/vellum-assistant. Memory Corpus Ingest is an agent skill from vellum-ai/vellum-assistant. Ingest a large dataset into memory as a skimmed map.
Run `npx skills add vellum-ai/vellum-assistant --skill memory-corpus-ingest -a claude-code`. Or copy the skill folder (skills/memory-corpus-ingest in vellum-ai/vellum-assistant) into .claude/skills/memory-corpus-ingest in your project. Claude Code loads it when a task matches its description.
Run `npx skills add vellum-ai/vellum-assistant --skill memory-corpus-ingest -a codex`. Or copy the skill folder (skills/memory-corpus-ingest in vellum-ai/vellum-assistant) into .agents/skills/memory-corpus-ingest in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vellum-ai/vellum-assistant --skill memory-corpus-ingest -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/memory-corpus-ingest, .gemini/skills/memory-corpus-ingest, .github/skills/memory-corpus-ingest and .opencode/skills/memory-corpus-ingest in your project.
Going by SKILL.md and its folder, Memory Corpus Ingest needs TypeScript for the scripts in its folder and the command-line tools its instructions call (rg, rsync and bun). Our summary lists: Node.js. Compatibility (from SKILL.md): Designed for Vellum personal assistants.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Memory Corpus Ingest is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.4k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Memory Corpus Ingest: Wiki Ingest (paperclipai/paperclip, 99k stars), Token Map (nexu-io/open-design, 100k stars), Datasets (Arize-ai/phoenix, 12k stars) and Maps Geography (asgeirtj/system_prompts_leaks, 69k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
vellum-ai (a GitHub organization) maintains it in vellum-ai/vellum-assistant, which has 1,408 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 9, 2026.
Source: vellum-ai/vellum-assistant on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.