Agent skill

Source Intake

by scholay in scholay/skills

Organize, classify, and register raw research materials into a stable, trackable source layer.

MITAuto-check passedDocuments & Office

Install Source Intake

skills CLI
$ npx skills add scholay/skills --skill source-intake -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install scholay/skills source-intake --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/scholay/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/source-intake .claude/skills/source-intake && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
source-intake
GitHub stars
141
Token cost
~1.1k tokens
SKILL.md length
563 words
Files
3 (incl. references)
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Organize, classify, and register raw research materials into a stable, trackable source layer.

  • Works in 6 steps: Inventory before touching → Expand all archives → Check for duplicates → …
  • Processing newly collected PDFs
  • SKILL.md covers The core problem, Object types, Workflow and Rules, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Source Intake is an agent skill from scholay/skills. Organize, classify, and register raw research materials into a stable, trackable source layer. Use when processing newly collected PDFs, archives, image folders, or mixed downloads; assigning stable IDs; maintaining intake manifests; or deciding what type of object each material is before any processing begins.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/archive-first.md` and `references/naming-conventions.md`).

It sits in Documents & Office. The repository describes itself as: Open academic AI skills maintained by Scholay. The licence is MIT.

When your agent uses it

  • Processing newly collected PDFs
  • Mixed downloads
  • Assigning stable IDs
  • Maintaining intake manifests

Example prompts

  • “/source-intake”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Inventory before touching
  2. Expand all archives
  3. Check for duplicates
  4. Assign stable IDs
  5. Normalize
  6. Record derivatives separately

What it can do on your machine

Read from SKILL.md and the folder at commit cd61bd9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Source Intake loads about 1.1k tokens when it runs, and up to ~2.8k if it reads all its reference files. Until then it costs about 82 tokens; SKILL.md has 563 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from scholay/skills at commit cd61bd9, republished under its MIT licence (© scholay). 563 words, ~1,069 tokens.

Download SKILL.mdSave it as .claude/skills/source-intake/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
source-intake
description
Organize, classify, and register raw research materials into a stable, trackable source layer. Use when processing newly collected PDFs, archives, image folders, or mixed downloads; assigning stable IDs; maintaining intake manifests; or deciding what type of object each material is before any processing begins.

Source Intake

The core problem

Research materials accumulate in chaotic ways: some arrive as standalone PDFs, some inside zip archives, some as image folders, some as web pages saved to disk. Processing them without first understanding what you have creates a false baseline — you think you've ingested everything when the most important items are still locked inside a zip file.

Rule zero: expand all archives before claiming anything is ingested.

Object types

Before assigning an ID to anything, classify it:

TypeDescription
source_pdfA PDF that can be processed directly
derived_pdfA PDF you generate from images, scans, or web pages
image_groupA folder of images that form a logical unit (scanned pages, figures)
datasetStructured data files (CSV, XLSX, JSON, GIS)
reference_fileRights, permissions, licenses, or metadata files
archiveZIP, RAR, or other compressed container — expand first, then classify contents
ambiguousNeeds a decision before it can be registered

Workflow

1. Inventory before touching

Before doing anything else, list what you have and where it came from. Keep a raw inventory file. Do not rename or move anything in the raw layer.

2. Expand all archives

Expand every archive into a staging area. Note the parent archive ID for each extracted file so you can trace it back later.

Update these manifests after expansion:

  • expanded_inventory — what came out of which archive
  • pending_pdf_candidates — PDFs waiting for registration
  • pending_image_groups — image folders waiting for a handling decision
3. Check for duplicates

Before assigning a new ID, check whether this item is already registered. Match on: DOI or URL first → title + year → manual review.

4. Assign stable IDs

Once an item is confirmed unique and classified, assign a stable ID. Recommended format:

  • Raw objects: RAW-#### (source as received, untouched)
  • Processed source objects: SRC-#### (stable, normalized copy)

Keep both IDs in your manifest so the chain from raw to processed is always traceable.

Show full SKILL.md (254 more words)Show less
5. Normalize

Create one clean, stable copy of each source object:

  • PDF: copy with a stable filename into your normalized folder
  • Image group: document the group strategy before doing anything else (will this become a derived PDF, or stay as individual images?)
  • Dataset: register and note the format; do not force into PDF
  • Web page: preserve the URL and access date; create a PDF capture if needed
6. Record derivatives separately

If you generate page images, OCR text, or other derivatives from a source object, record them in a separate manifest — not in the raw inventory. Keep the chain: raw → normalized → derivative.

Rules

  • Never rename or overwrite files in the raw layer. All transformations happen in a separate staging or working layer.
  • Do not register a source object before checking for duplicates.
  • Do not force image groups into PDF form by default — decide the handling strategy first.
  • Do not force datasets, license files, or ambiguous objects into PDF form.
  • Manifests must make parent–child relationships explicit (which archive did this PDF come from?).
  • Record the source before generating any derivatives.
  • If a source is ambiguous, mark it status: pending_decision rather than guessing.

Manifest minimum fields

For each registered source object, record at minimum:

  • stable ID (SRC-####)
  • object type (from the type table above)
  • original raw path
  • normalized/working path
  • parent raw ID (if from an archive)
  • language
  • status
  • notes

References

  • Read references/naming-conventions.md for ID formats, filename patterns, and status vocabulary.
  • Read references/archive-first.md for why archive expansion must come before registration, and common mistakes when skipping it.

© scholay, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in source-intake of scholay/skills.

  • SKILL.md
  • references/archive-first.md
  • references/naming-conventions.md

Open the folder on GitHubat commit cd61bd9

Compare with similar skills

Source Intake next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Source Intake compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Source Intake this skillscholay/skills141—~1.1kAutomated safety check: PassMIT
Markdown Article FormatterJimLiu/baoyu-skills26k6 repos~3.5kAutomated safety check: PassMIT
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Obsidian MarkdownAtmosphere/atmosphere3.8k20 repos~1.3kAutomated safety check: PassApache-2.0
DOCXrvdbreemen/OTGW-firmware20733 repos~4.3kAutomated safety check: PassProprietary
Gzh Designisjiamu/gzh-design-skill3.9k1 repos~2.2kAutomated safety check: PassAGPL-3.0

Similar skills

  • Markdown Article Formatter

    JimLiu/baoyu-skills

    Reformats plain text or Markdown articles with frontmatter, a title, a summary, headings, bold, lists and code blocks, and saves a separate formatted copy.

    26k GitHub starsUsed in 6 repos~3.5k tokens
    Documents & OfficeAuto-check passed
  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Obsidian Markdown

    Atmosphere/atmosphere

    Create and edit Obsidian Flavored Markdown with wikilinks, embeds, callouts, properties, and other Obsidian-specific syntax.

    3.8k GitHub starsUsed in 20 repos~1.3k tokens
    Documents & OfficeAuto-check passed
  • DOCX

    rvdbreemen/OTGW-firmware

    A skill your agent uses whenever the user wants to create, read, edit, or manipulate Word documents (.docx files).

    207 GitHub starsUsed in 33 repos~4.3k tokens
    Documents & OfficeAuto-check passed
  • Gzh Design

    isjiamu/gzh-design-skill

    微信公众号文章排版引擎,将 Markdown 转换为可直接粘贴到公众号编辑器的 HTML。主题风格从 references/theme-index.md 注册的自定义主题库中选取,自动章节编号、关键词下划线标记、引言卡片、目录导航、代码块、图片/GIF、作者签名。支持 Markdown / Word(.docx) / PDF / 纯文本输入(非 Markdown…

    3.9k GitHub starsUsed in 1 repo~2.2k tokens
    Documents & OfficeAuto-check passed
  • Reads, creates and edits Word .docx files with python-docx, and drops to raw OOXML for tracked changes, comments and byte-exact edits.

    41k GitHub stars~2.5k tokensUpdated 3 days ago
    Documents & OfficeAuto-check passed

More from scholay/skills

  • Corpus Building

    scholay/skills

    Extract, structure, and quality-check evidence from registered source objects.

    141 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Evidence Review

    scholay/skills

    Assess and promote corpus slices toward citation readiness. An agent skill from scholay/skills.

    141 GitHub stars~1.1k tokensUpdated 3 mo ago
    Auto-check passed
  • Literature Discovery

    scholay/skills

    Systematically discover and shortlist academic literature for a research project.

    141 GitHub stars~864 tokensUpdated 3 mo ago
    Auto-check passed
  • Research Writing

    scholay/skills

    Draft and maintain academic papers from verified corpus evidence, manage bibliography, and produce output in any target format.

    141 GitHub stars~1.3k tokensUpdated 3 mo ago
    Auto-check passed

Questions about Source Intake

What does Source Intake do?

Organize, classify, and register raw research materials into a stable, trackable source layer. Source Intake is an agent skill from scholay/skills. Organize, classify, and register raw research materials into a stable, trackable source layer.

When should I use Source Intake?

Source Intake fits situations like: processing newly collected PDFs; mixed downloads; assigning stable IDs; maintaining intake manifests.

How do I install Source Intake in Claude Code?

Run `npx skills add scholay/skills --skill source-intake -a claude-code`. Or copy the skill folder (source-intake in scholay/skills) into .claude/skills/source-intake in your project. Claude Code loads it when a task matches its description.

How do I install Source Intake in Codex?

Run `npx skills add scholay/skills --skill source-intake -a codex`. Or copy the skill folder (source-intake in scholay/skills) into .agents/skills/source-intake in your project. Codex loads it when a task matches its description.

Can I use Source Intake in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scholay/skills --skill source-intake -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/source-intake, .gemini/skills/source-intake, .github/skills/source-intake and .opencode/skills/source-intake in your project.

What does Source Intake need to run?

SKILL.md names no scripts, command-line tools or credentials: Source Intake is instructions for the agent only.

Does Source Intake access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Source Intake safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Source Intake use?

Source Intake is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Source Intake use?

About 1.1k tokens (SKILL.md is roughly 4.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.8k tokens, read only when the agent opens those files.

What are the alternatives to Source Intake?

Skills that share tags, products or a category with Source Intake: Markdown Article Formatter (JimLiu/baoyu-skills, 26k stars), Markitdown (ImCa0/just-laws, 781 stars), Obsidian Markdown (Atmosphere/atmosphere, 3.8k stars) and DOCX (rvdbreemen/OTGW-firmware, 207 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Source Intake?

scholay (a GitHub user) maintains it in scholay/skills, which has 141 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on June 17, 2026.

Source: scholay/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.