Agent skill

Corpus Building

by scholay in scholay/skills

Extract, structure, and quality-check evidence from registered source objects.

MITAuto-check passedWriting & Content

Install Corpus Building

skills CLI
$ npx skills add scholay/skills --skill corpus-building -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install scholay/skills corpus-building --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/scholay/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/corpus-building .claude/skills/corpus-building && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
corpus-building
GitHub stars
141
Token cost
~1.5k tokens
SKILL.md length
811 words
Files
4 (incl. references)
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Extract, structure, and quality-check evidence from registered source objects.

  • Works in 6 steps: Create the source card first → Run OCR → Slice into evidence cards → …
  • Converting processed documents into searchable evidence cards
  • SKILL.md covers The core problem, Key concepts, Workflow and Rules, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Corpus Building is an agent skill from scholay/skills. Extract, structure, and quality-check evidence from registered source objects. Use when converting processed documents into searchable evidence cards, managing OCR and translation layers, running quality scoring, handling noise and recovery, or maintaining a corpus index.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/ocr-noise.md`, `references/quality-scoring.md` and `references/slice-schema.md`).

It sits in Writing & Content, covering Translation. The repository describes itself as: Open academic AI skills maintained by Scholay. The licence is MIT.

When your agent uses it

  • Converting processed documents into searchable evidence cards
  • Managing OCR and translation layers
  • Running quality scoring
  • Handling noise and recovery

Example prompts

  • “/corpus-building”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Create the source card first
  2. Run OCR
  3. Slice into evidence cards
  4. Run quality scoring
  5. Handle noise and blocked pages
  6. Validate and rebuild

What it can do on your machine

Read from SKILL.md and the folder at commit cd61bd9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Corpus Building loads about 1.5k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 72 tokens; SKILL.md has 811 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~72
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from scholay/skills at commit cd61bd9, republished under its MIT licence (© scholay). 811 words, ~1,525 tokens.

Download SKILL.mdSave it as .claude/skills/corpus-building/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
corpus-building
description
Extract, structure, and quality-check evidence from registered source objects. Use when converting processed documents into searchable evidence cards, managing OCR and translation layers, running quality scoring, handling noise and recovery, or maintaining a corpus index.

Corpus Building

The core problem

A folder of PDFs is not a corpus. A corpus is a collection of addressable evidence units — each one traceable to a specific page of a specific source, with enough context to be used in writing without reopening the original document.

This skill covers the full path from source object to evidence card, including OCR, translation, quality scoring, and the noise problems that appear at scale.

Key concepts

Source card — one record per registered source. Holds metadata, processing status, and links to all derived evidence.

Text slice — one evidence unit extracted from a specific page range of a source. Stores the original text, cleaned OCR text, and working translation as separate layers. Never collapses them.

Image slice — one image evidence unit (a page scan, figure, plan, photograph, table, or diagram). Stores file path, description, and status.

The separation rule: a slice stores what the source says. AI annotations store what AI thinks about the slice. Do not mix them.

Workflow

1. Create the source card first

Before creating any slices, create a source card for the registered source object. Minimum fields: source ID, title, language, file path, topics, status fields for OCR and translation.

2. Run OCR

For each source, generate per-page text output.

Default path: use a local OCR engine (Tesseract or equivalent). For multilingual documents, set the language pack explicitly — running an incorrect language pack produces script-mismatch noise, not just bad quality.

Cloud OCR path (optional): use a cloud OCR API for higher fidelity on difficult scans. Reserve this for pages that local OCR consistently fails on, not as a default for all sources.

Output per page: raw OCR text file. Keep raw output separate from any cleaned version.

3. Slice into evidence cards

For each page (or logical text unit), create a text slice card with:

  • Slice ID (SLC-####-T###)
  • Source ID and page reference
  • Original text (from source)
  • OCR-cleaned text (if different)
  • Working translation (in your project's working language)
  • Status fields: OCR status, translation status, verification status

For each page image, create a baseline image slice card with:

  • Slice ID (SLC-####-I###)
  • Source ID and page number
  • Image file path
  • Image type (page scan, figure, plan, photo, table, diagram)
  • Description status

On image descriptions: use a multimodal model to inspect the actual image pixels. Do not use OCR output as a substitute for visual inspection. Label OCR-derived descriptions as ocr_assisted_draft — they are lower trust than a genuine visual description.

4. Run quality scoring

After slicing, score each page on:

  • OCR quality: are there repeated characters, symbol noise, script artifacts?
  • Translation quality: is the working language layer present and readable?

Assign a triage priority:

  • P1 — likely broken; inspect immediately before using
  • P2 — needs review; often image-heavy or missing translation
  • P3 — usable but worth spot-checking
  • OK — acceptable for browsing

Calibrate before scaling: after your first new source type, inspect all P1 pages plus a sample of P2 and OK. If the scoring looks wrong, fix the scorer before running 50 more sources.

Show full SKILL.md (314 more words)Show less
5. Handle noise and blocked pages

At scale, some pages will have OCR artifacts severe enough to quarantine:

  • Repeated-character runs
  • Script mismatch (wrong language pack)
  • Symbol floods

Before bulk quarantine, sample the flagged pages visually. Common false positives:

  • Table-of-contents pages with dotted leaders (…………) — readable, just unusual
  • Pages in a different script than you expected — fix the language setting, not the page

For genuinely noisy pages, mark them excluded and log them in a blocked pages list with a reason code. Revisit after fixing the underlying cause (wrong language pack, poor scan quality, etc.) rather than leaving them permanently blocked.

6. Validate and rebuild

After every batch, validate that:

  • Every slice card in your folder has a row in your index
  • No duplicate slice IDs exist
  • No index rows reference missing files

Rebuild any downstream views (app data, search indexes) from the validated corpus, not from cached state.

Rules

  • Markdown/text cards are the source of truth. CSV/database indexes are rebuildable projections.
  • Never let a translation overwrite the original text or the OCR-cleaned text. Keep all three layers.
  • Never use OCR output as a substitute for visual inspection of images.
  • Do not create slices from unregistered source objects.
  • Do not run quality scoring and then ignore the results — calibrate the scorer and act on P1 pages.
  • Over-quarantine is as harmful as under-quarantine. Sample before bulk-blocking.
  • Auto-generated page slices are browsing inventory, not citation-ready evidence. Mark them accordingly.
  • Validate after every batch. Do not continue slicing when validation fails.

Status vocabulary

OCR status: not_started / in_progress / raw_ocr / cleaned / verified / blocked / not_needed

Translation status: not_started / machine_draft / agent_draft / pending_translation / reviewed / quote_ready / blocked / not_needed

Verification status: draft / checked / reviewed

Image description status: not_started / ocr_assisted_draft / visual_draft / reviewed / caption_ready

References

  • Read references/slice-schema.md for required fields, status vocabulary, and granularity rules.
  • Read references/quality-scoring.md for triage priority semantics, risk tags, and calibration guidance.
  • Read references/ocr-noise.md for common noise patterns, false-positive quarantine risks, and recovery steps.

© scholay, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in corpus-building of scholay/skills.

  • SKILL.md
  • references/ocr-noise.md
  • references/quality-scoring.md
  • references/slice-schema.md

Open the folder on GitHubat commit cd61bd9

Compare with similar skills

Corpus Building next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Corpus Building compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Corpus Building this skillscholay/skills141—~1.5kAutomated safety check: PassMIT
Cuimao TranslatorCuimao777/cuimao-translator486—~3.2kAutomated safety check: PassNone
Nature-Style Academic PolishingYuan1z0825/nature-skills47k—~1.5kAutomated safety check: PassApache-2.0
Visa Doc Translateaffaan-m/ECC276k4 repos~1kAutomated safety check: PassMIT
Fantasia Locale Translations Inputvishiri/fantasia-archive409—~1.1kAutomated safety check: PassGPL-3.0
Translation Diff ExportDevolutions/UniGetUI26k—~1.1kAutomated safety check: PassMIT

Similar skills

  • Cuimao Translator

    Cuimao777/cuimao-translator

    Translate English PDF books into natural Chinese with three quality modes.

    486 GitHub stars~3.2k tokensUpdated 5 mo ago
    Writing & ContentAuto-check passed
  • Nature-Style Academic Polishing

    Yuan1z0825/nature-skills

    Polishes, translates or tightens existing academic prose and fixes manuscript LaTeX layout while keeping facts, terminology and evidence boundaries intact.

    47k GitHub stars~1.5k tokensUpdated today
    Writing & ContentAuto-check passed
  • Visa Doc Translate

    affaan-m/ECC

    Translate visa document images (bank deposit, employment, income, and retirement certificates; HEIC, PNG, or JPG) into English via OCR and produce a bilingual PDF pairing the original image with a…

    276k GitHub starsUsed in 4 repos~1k tokens
    Writing & ContentAuto-check passed
  • Fantasia Locale Translations Input

    vishiri/fantasia-archive

    FaLocaleTranslationsInput element and per-locale string maps for Project Settings world names, document template titles, world appendix, layout group names, and placement nicknames.

    409 GitHub stars~1.1k tokensUpdated yesterday
    Writing & ContentAuto-check passed
  • Translation Diff Export

    Devolutions/UniGetUI

    Compares UniGetUI JSON locale files against English, identifies untranslated or source-changed keys, and generates patch, reference, and handoff files for a target language.

    26k GitHub stars~1.1k tokensUpdated today
    Writing & ContentAuto-check passed
  • Sync Translations

    symfony/symfony

    Synchronize translation catalogs across maintained Symfony branches: find messages that newer branches added to the English catalogs but that are still missing from the oldest maintained branch…

    31k GitHub stars~1.9k tokensUpdated today
    Writing & ContentAuto-check passed

More from scholay/skills

  • Evidence Review

    scholay/skills

    Assess and promote corpus slices toward citation readiness. An agent skill from scholay/skills.

    141 GitHub stars~1.1k tokensUpdated 3 mo ago
    Auto-check passed
  • Literature Discovery

    scholay/skills

    Systematically discover and shortlist academic literature for a research project.

    141 GitHub stars~864 tokensUpdated 3 mo ago
    Auto-check passed
  • Research Writing

    scholay/skills

    Draft and maintain academic papers from verified corpus evidence, manage bibliography, and produce output in any target format.

    141 GitHub stars~1.3k tokensUpdated 3 mo ago
    Auto-check passed
  • Source Intake

    scholay/skills

    Organize, classify, and register raw research materials into a stable, trackable source layer.

    141 GitHub stars~1.1k tokensUpdated 3 mo ago
    Auto-check passed

Questions about Corpus Building

What does Corpus Building do?

Extract, structure, and quality-check evidence from registered source objects. Corpus Building is an agent skill from scholay/skills. Extract, structure, and quality-check evidence from registered source objects.

When should I use Corpus Building?

Corpus Building fits situations like: converting processed documents into searchable evidence cards; managing OCR and translation layers; running quality scoring; handling noise and recovery.

How do I install Corpus Building in Claude Code?

Run `npx skills add scholay/skills --skill corpus-building -a claude-code`. Or copy the skill folder (corpus-building in scholay/skills) into .claude/skills/corpus-building in your project. Claude Code loads it when a task matches its description.

How do I install Corpus Building in Codex?

Run `npx skills add scholay/skills --skill corpus-building -a codex`. Or copy the skill folder (corpus-building in scholay/skills) into .agents/skills/corpus-building in your project. Codex loads it when a task matches its description.

Can I use Corpus Building in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scholay/skills --skill corpus-building -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/corpus-building, .gemini/skills/corpus-building, .github/skills/corpus-building and .opencode/skills/corpus-building in your project.

What does Corpus Building need to run?

SKILL.md names no scripts, command-line tools or credentials: Corpus Building is instructions for the agent only.

Does Corpus Building access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Corpus Building safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Corpus Building use?

Corpus Building is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Corpus Building use?

About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.5k tokens, read only when the agent opens those files.

What are the alternatives to Corpus Building?

Skills that share tags, products or a category with Corpus Building: Cuimao Translator (Cuimao777/cuimao-translator, 486 stars), Nature-Style Academic Polishing (Yuan1z0825/nature-skills, 47k stars), Visa Doc Translate (affaan-m/ECC, 276k stars) and Fantasia Locale Translations Input (vishiri/fantasia-archive, 409 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Corpus Building?

scholay (a GitHub user) maintains it in scholay/skills, which has 141 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on June 17, 2026.

Source: scholay/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.