Agent skill

Memory Corpus Ingest

by vellum-ai in vellum-ai/vellum-assistant

Ingest a large dataset into memory as a skimmed map. An agent skill from vellum-ai/vellum-assistant.

MITAuto-check: notes

Install Memory Corpus Ingest

skills CLI
$ npx skills add vellum-ai/vellum-assistant --skill memory-corpus-ingest -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vellum-ai/vellum-assistant memory-corpus-ingest --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/memory-corpus-ingest .claude/skills/memory-corpus-ingest && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
memory-corpus-ingest
GitHub stars
1.4k
Token cost
~3k tokens
SKILL.md length
1,343 words
Files
5 (incl. scripts, references)
Skills in repo
108
Repo updated
First seen
Licence
MIT

At a glance

Ingest a large dataset into memory as a skimmed map. An agent skill from vellum-ai/vellum-assistant.

  • Works in 7 steps: Scope and confirm → Cold-store the raw corpus → Inventory and slice plan → …
  • SKILL.md covers Procedure, Hard rules and References
  • Runs TypeScript scripts from its folder; calls rg, rsync and bun

What it does

Memory Corpus Ingest is an agent skill from vellum-ai/vellum-assistant. Ingest a large dataset into memory as a skimmed map. Cold-store the raw files under a workspace imports directory, census them into a slice plan, skim each slice into compact map pages that point back at the raw files, ingest the map with the memory ingest CLI, and author a drill-in retrieval skill so the corpus stays searchable on demand. For recording archives, transcript collections, document dumps, and any corpus too large to hold in memory directly.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/fathom.md`, `references/map-page-template.md` and `scripts/inventory.test.ts`). Compatibility notes: Designed for Vellum personal assistants

The repository describes itself as: An AI Assistant that’s easy to setup, does your work 24/7, knows your preferences and gets better over time. The licence is MIT.

Example prompts

  • “/memory-corpus-ingest”

Requirements

  • Node.js
  • Compatibility (from SKILL.md): Designed for Vellum personal assistants

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Scope and confirm
  2. Cold-store the raw corpus
  3. Inventory and slice plan
  4. Skim each slice into a staged map page
  5. Ingest the map
  6. Author the drill-in retrieval skill
  7. Verify

What it can do on your machine

Read from SKILL.md and the folder at commit 33cc983. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (TypeScript), which the agent can run.

    Shell commands in SKILL.md call:

    • rg
    • rsync
    • bun

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use rsync, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Vellum personal assistants

    From compatibility in the SKILL.md frontmatter.

Context cost

Memory Corpus Ingest loads about 3k tokens when it runs, and up to ~6.5k if it reads all its reference files. Until then it costs about 120 tokens; SKILL.md has 1,343 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~120
When it runs · the whole SKILL.md, loaded when a task matches
~3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:54
    find /path/to/raw-corpus \( -name '.env*' -o -name '*.key' -o -name '*.pem' \
  • NoteMentions a .env fileSKILL.md:72
    rsync -a --exclude='.env*' --exclude='*.key' --exclude='*.pem' \

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from vellum-ai/vellum-assistant at commit 33cc983, republished under its MIT licence (© vellum-ai). 1,343 words, ~3,015 tokens.

Download SKILL.mdSave it as .claude/skills/memory-corpus-ingest/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
memory-corpus-ingest
description
Ingest a large dataset into memory as a skimmed map. Cold-store the raw files under a workspace imports directory, census them into a slice plan, skim each slice into compact map pages that point back at the raw files, ingest the map with the memory ingest CLI, and author a drill-in retrieval skill so the corpus stays searchable on demand. For recording archives, transcript collections, document dumps, and any corpus too large to hold in memory directly.
compatibility
Designed for Vellum personal assistants
metadata.emoji
🗺️

Corpus Ingest

Bring a large dataset into the assistant's working knowledge without stuffing it into memory. The model is a library: the workspace holds the stacks (the raw files, cold and complete), memory holds the card catalog (a small set of map pages that say what exists, when it is from, and where to look), and a purpose-built retrieval skill is the librarian that walks to the right shelf on demand.

Two invariants drive everything below:

  1. Raw data never enters the memory corpus. Nothing from the dataset is written into memory/concepts/ except the map pages, and nothing is ever appended to memory/buffer.md (bulk buffer appends trip the consolidation burst guard and per-run caps; the map bypasses the buffer entirely via assistant memory ingest).
  2. The map stays small. Roughly 10 to 50 pages regardless of corpus size. If the corpus doubles, the pages get denser or the slices get coarser; the page count does not double.

Procedure

Step 1: Scope and confirm

Identify the source and its size before committing:

bash
du -sh /path/to/raw-corpus
find /path/to/raw-corpus -type f | wc -l

Tell the user what will happen: the raw files move into the workspace, a bounded number of summarization passes read them once to build the map, the map is ingested into memory, and a lookup skill is authored for drill-in. Skimming a large corpus is real LLM work that costs time and money; confirm before starting. For Fathom recording exports, read references/fathom.md first for format discovery and slicing guidance.

Step 2: Cold-store the raw corpus

Land the raw files under an imports directory in the workspace, one directory per source:

Screen for credentials BEFORE copying: an arbitrary corpus can carry secret material, and anything landed under imports/ becomes reachable by workspace tools, backups, and retrieval flows.

bash
cd "$VELLUM_WORKSPACE_DIR"
# 1a. Screen for secret-bearing FILE NAMES; review every hit with the user.
find /path/to/raw-corpus \( -name '.env*' -o -name '*.key' -o -name '*.pem' \
  -o -name '*credential*' -o -name '*secret*' -o -name 'cookies*' \
  -o -path '*tokens*' -o -path '*oauth*' \) -print

# 1b. Screen file CONTENTS for credential shapes. --hidden and --no-ignore
#     matter: rg skips dotfiles and gitignored paths by default, which is
#     exactly where credentials live. Capture the FULL list (no truncation):
#     every file named here must be excluded below or cleaned with the user
#     before it lands.
rg -l -i --hidden --no-ignore \
  "api[_-]?key|access[_-]?token|refresh[_-]?token|client[_-]?secret|password\s*[=:]|passwd|bearer |AKIA[0-9A-Z]{16}|BEGIN [A-Z ]*PRIVATE KEY" \
  /path/to/raw-corpus > /tmp/corpus-secret-hits.txt
cat /tmp/corpus-secret-hits.txt

# 2. Build rsync exclusions from the content hits (paths relative to the
#    corpus root), then copy with ALL flagged paths excluded.
sed 's|^/path/to/raw-corpus/||' /tmp/corpus-secret-hits.txt > /tmp/corpus-secret-exclusions.txt
mkdir -p imports/<source>
rsync -a --exclude='.env*' --exclude='*.key' --exclude='*.pem' \
  --exclude='tokens/' --exclude='oauth/' --exclude='cookies*' \
  --exclude-from=/tmp/corpus-secret-exclusions.txt \
  /path/to/raw-corpus/ imports/<source>/

Rules:

  • Never place raw corpus files under memory/. The cold store is imports/<source>/; the map is the only thing that enters memory.
  • Never land credentials in the cold store. Exclude secret-bearing files during the copy; if the content screen finds embedded live tokens inside otherwise-wanted files, pause and resolve them with the user before landing those files.
  • Treat the cold store as read-only once landed. The map pages and the drill-in skill both point at these paths; moving files later breaks every pointer.
  • If the corpus is huge, check free disk space first and copy in batches. Prefer copy over move until the user confirms the original can be released.
Step 3: Inventory and slice plan

Census the corpus and produce a machine-readable slice plan:

bash
mkdir -p "$VELLUM_WORKSPACE_DIR/imports/<source>/.staging"
bun run {baseDir}/scripts/inventory.ts "$VELLUM_WORKSPACE_DIR/imports/<source>" \
  > "$VELLUM_WORKSPACE_DIR/imports/<source>/.staging/plan.json"

The script prints a human census to stderr (file count, total size, extension mix, date range) and a JSON plan to stdout: { files, totalBytes, byExtension, dateRange, suggestedSlices }. Each suggested slice is a date-windowed group of files sized for one skim pass. Review the plan before skimming:

  • If the slice count is outside roughly 10 to 50, adjust: merge sparse adjacent slices or split dense ones. The plan is a suggestion, not a contract.
  • Files with no recognizable date cluster on file mtime; if mtimes are all import-day (a fresh copy), pick slices by directory or topic instead and say so in the map.
Step 4: Skim each slice into a staged map page

For each slice in the plan, run one summarization pass that reads the slice's files and writes one staged map page:

  • Output goes to the staging directory as <slug>.md (for example imports/fathom/.staging/fathom-recordings-2025-q1.md). The slug is the filename minus .md.
  • Every page follows references/map-page-template.md exactly: lead that stands alone as the retrieval card, a Raw data: pointer line in the lead, ## sections per topic or time slice, and source: / origin_date: / ref_files: / links: frontmatter.
  • Also write one corpus index page (kind: index) whose links: enumerate all slice pages.

Bound the fan-out. The number of skim passes equals the number of slices in the plan, full stop. Run them a few at a time (parallelism of 3 or 4 is plenty). Never spawn one pass per file, never let a pass recursively spawn more passes, and never re-skim slices that already have a staged page unless their content changed. An unbounded fan-out over a large corpus is the expensive failure mode of this skill.

A skim pass extracts what a future search needs to route: decisions, recurring topics, named people and projects, date spans, open threads. It does not transcribe. Verbatim content stays cold; the map records that it exists and where.

Show full SKILL.md (619 more words)Show less
Step 5: Ingest the map

Dry-run first, review, then ingest for real:

bash
cd "$VELLUM_WORKSPACE_DIR"
assistant memory ingest --dir imports/<source>/.staging --dry-run

The dry run validates every page and reports per-page results without writing. Fix anything reported invalid (bad frontmatter, bad slug), and resolve any warning about a links: or [[wikilink]] target that is neither on disk nor in the staged set (stage the missing page, or make the reference plain prose; retrieval drops a link whose target page does not exist), then run without --dry-run. Notes:

  • Requires concept-page memory (memory.v3.live or memory.v2.enabled).
  • Existing slugs are skipped unless --overwrite is passed; use --overwrite when re-running after edits.
  • Pages go up in batches of 200 (a map should never get near that).
  • If the command fails because the consolidation lock is held, it names the holder; wait for that run to finish and retry.

After ingest, the maintain job picks up the new pages and embeds their sections; the map becomes retrievable shortly after, with each page's card dated by its origin_date rather than the import day.

Step 6: Author the drill-in retrieval skill

The map tells the model a slice exists; the drill-in skill is how it follows the pointer. Author a workspace skill for this corpus under $VELLUM_WORKSPACE_DIR/skills/<source>-lookup/:

  • SKILL.md with frontmatter per the Agent Skills spec (name matching the directory, keyword-rich description) and activation hints drawn from the intents you actually observed while scoping (for example "when the user asks what was said in a meeting", "when a question needs the <source> archive").
  • scripts/ with actually runnable search commands over the cold store, not prose descriptions of searching. A minimal helper is a ripgrep wrapper scoped to the imports directory with date filtering:
bash
#!/usr/bin/env bash
# search.sh <pattern> [YYYY-MM]  scoped search over the <source> cold store
DIR="$VELLUM_WORKSPACE_DIR/imports/<source>"
if [ -n "$2" ]; then
  # Two globs: a slash-free pattern matches the date in any basename, and
  # the "**/" prefixed pattern matches the date in a directory segment at
  # any depth (a leading "*" cannot cross path separators).
  rg -i -C 3 "$1" "$DIR" --glob "*$2*" --glob "**/*$2*/**"
else
  rg -i -C 3 "$1" "$DIR"
fi
  • The skill body should name the cold-store root, the slice layout, and the map pages' slugs so the model can go from a card to a file in one step.

Confirm the new skill appears in assistant skills list (workspace skills are picked up from the workspace skills/ directory).

Step 7: Verify

Run 3 to 5 representative queries a real user would ask of this corpus (mix a routing question, a specific-fact question, and a date-scoped question). For each, check:

  1. A map card is selected for the turn (the slice page or the corpus index surfaces).
  2. The assistant follows the card's Raw data: pointer or invokes the drill-in skill to reach the actual files.
  3. The answer cites content that exists in the cold store.

If a query routes to nothing, the relevant lead is not doing its card job; rewrite it and re-ingest that page with --overwrite. For large ingests, assistant memory v3 eval can additionally gate the change, but it is a two-corpus comparison requiring --snapshot, --staging, and --out; use it only when you captured a pre-ingest snapshot of memory/concepts/ to compare against. Otherwise the query checks above are the verification.

Hard rules

  • Raw corpus content never enters memory/concepts/ or memory/buffer.md. Map pages only, via assistant memory ingest.
  • The cold store lives under imports/<source>/ and is read-only after landing.
  • Map size is 10 to 50 pages regardless of corpus size.
  • Drill-in pointers (the Raw data: line) live in the page lead, because the card renderer injects the lead plus a section list; a pointer buried in a section is invisible at routing time.
  • origin_date: is the content's own date, never the import date.
  • Skim fan-out is bounded by the slice plan; one pass per slice, a few in flight at a time.
  • Always dry-run the ingest before the real run.

References

  • references/map-page-template.md: the exact article shape for map pages, with a full example. Read before writing any staged page.
  • references/fathom.md: Fathom recording archives specifically; export discovery, speaker and date extraction, slicing, and what belongs in the map versus the cold store.

© vellum-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in skills/memory-corpus-ingest of vellum-ai/vellum-assistant.

  • SKILL.md
  • references/fathom.md
  • references/map-page-template.md
  • scripts/inventory.test.ts
  • scripts/inventory.ts

Open the folder on GitHubat commit 33cc983

Compare with similar skills

Memory Corpus Ingest next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Memory Corpus Ingest compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Memory Corpus Ingest this skillvellum-ai/vellum-assistant1.4k—~3kAutomated safety check: NotesMIT
Wiki Ingestpaperclipai/paperclip99k—~933Automated safety check: PassMIT
Token Mapnexu-io/open-design100k—~1.4kAutomated safety check: PassApache-2.0
DatasetsArize-ai/phoenix12k—~1.6kAutomated safety check: PassCustom licence
Maps Geographyasgeirtj/system_prompts_leaks69k—~717Automated safety check: PassCC0-1.0
Market Ingestruvnet/ruflo74k—~529Automated safety check: NotesMIT

Similar skills

  • Wiki Ingest

    paperclipai/paperclip

    A skill your agent uses when an operation issue asks to ingest a captured raw/ source into the LLM Wiki, or the user says "ingest <slug".

    99k GitHub stars~933 tokensUpdated today
    Knowledge ManagementAuto-check passed
  • Token Map

    nexu-io/open-design

    Map an extracted Figma / source-code token bag onto the active OD design system, producing a deterministic mapping the generate stage can consume.

    100k GitHub stars~1.4k tokensUpdated today
    Frontend & DesignAuto-check passed
  • Datasets

    Arize-ai/phoenix

    Understand what a Phoenix dataset is and reason well about its examples, outputs, splits, and how it feeds evaluators and experiments.

    12k GitHub stars~1.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Maps Geography

    asgeirtj/system_prompts_leaks

    Accurate maps from real geo data — use for any map, or whenever geography would make a good graphic for a deliverable

    69k GitHub stars~717 tokensUpdated yesterday
    Auto-check passed
  • Market Ingest

    ruvnet/ruflo

    Ingest and normalize market data into OHLCV vectors with HNSW indexing

    74k GitHub stars~529 tokensUpdated today
    DatabasesAuto-check: notes
  • Feature Map

    onyx-dot-app/onyx

    Use the Onyx feature map (.agents/feature-map/) to learn what a product surface does, the code behind it, and what a change can break.

    32k GitHub stars~459 tokensUpdated yesterday
    Auto-check passed

More from vellum-ai/vellum-assistant

All 108 skills in this repo
  • Vellum GitHub App Setup

    vellum-ai/vellum-assistant

    Create and configure a GitHub App so the assistant can push commits, open PRs, and comment under its own bot identity.

    1.4k GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Discord App Setup

    vellum-ai/vellum-assistant

    Connect a Discord bot to the assistant via the Discord Gateway with guided application creation and intent configuration

    1.4k GitHub stars~4.2k tokensUpdated yesterday
    Auto-check passed
  • Sentry App Setup

    vellum-ai/vellum-assistant

    Create and configure a Sentry internal integration so the assistant can manage issues, alerts, and releases under its own identity

    1.4k GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Plugin Builder

    vellum-ai/vellum-assistant

    A skill your agent uses when the user wants to build, scaffold, ship, or edit a Vellum plugin that bundles multiple surfaces (hooks, tools, skills, and more) into one installable package.

    1.4k GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Slack App Setup

    vellum-ai/vellum-assistant

    Connect a Slack app to the Vellum Assistant via Socket Mode.

    1.4k GitHub stars~2.5k tokensUpdated yesterday
    Auto-check: warnings
  • Amazon

    vellum-ai/vellum-assistant

    Shop on Amazon and Amazon Fresh through your browser. An agent skill from vellum-ai/vellum-assistant.

    1.4k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed

Questions about Memory Corpus Ingest

What does Memory Corpus Ingest do?

Ingest a large dataset into memory as a skimmed map. An agent skill from vellum-ai/vellum-assistant. Memory Corpus Ingest is an agent skill from vellum-ai/vellum-assistant. Ingest a large dataset into memory as a skimmed map.

How do I install Memory Corpus Ingest in Claude Code?

Run `npx skills add vellum-ai/vellum-assistant --skill memory-corpus-ingest -a claude-code`. Or copy the skill folder (skills/memory-corpus-ingest in vellum-ai/vellum-assistant) into .claude/skills/memory-corpus-ingest in your project. Claude Code loads it when a task matches its description.

How do I install Memory Corpus Ingest in Codex?

Run `npx skills add vellum-ai/vellum-assistant --skill memory-corpus-ingest -a codex`. Or copy the skill folder (skills/memory-corpus-ingest in vellum-ai/vellum-assistant) into .agents/skills/memory-corpus-ingest in your project. Codex loads it when a task matches its description.

Can I use Memory Corpus Ingest in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vellum-ai/vellum-assistant --skill memory-corpus-ingest -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/memory-corpus-ingest, .gemini/skills/memory-corpus-ingest, .github/skills/memory-corpus-ingest and .opencode/skills/memory-corpus-ingest in your project.

What does Memory Corpus Ingest need to run?

Going by SKILL.md and its folder, Memory Corpus Ingest needs TypeScript for the scripts in its folder and the command-line tools its instructions call (rg, rsync and bun). Our summary lists: Node.js. Compatibility (from SKILL.md): Designed for Vellum personal assistants.

Does Memory Corpus Ingest access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Memory Corpus Ingest safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Memory Corpus Ingest use?

Memory Corpus Ingest is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Memory Corpus Ingest use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.4k tokens, read only when the agent opens those files.

What are the alternatives to Memory Corpus Ingest?

Skills that share tags, products or a category with Memory Corpus Ingest: Wiki Ingest (paperclipai/paperclip, 99k stars), Token Map (nexu-io/open-design, 100k stars), Datasets (Arize-ai/phoenix, 12k stars) and Maps Geography (asgeirtj/system_prompts_leaks, 69k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Memory Corpus Ingest?

vellum-ai (a GitHub organization) maintains it in vellum-ai/vellum-assistant, which has 1,408 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 9, 2026.

Source: vellum-ai/vellum-assistant on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.