Agent skill

Archive Crawler

by inbrainfun in inbrainfun/inbrain

Universal archivist for personal file archives (Dropbox/B2/Gmail-takeout/local-mount/hard-drive-dump).

Custom licenceAuto-check passedData & Analytics

Install Archive Crawler

skills CLI
$ npx skills add inbrainfun/inbrain --skill archive-crawler -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install inbrainfun/inbrain archive-crawler --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/inbrainfun/inbrain.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/archive-crawler .claude/skills/archive-crawler && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
archive-crawler
GitHub stars
142
Used in
1 other repo
Token cost
~2.6k tokens
SKILL.md length
832 words
Files
2
Skills in repo
52
Repo updated
First seen
Licence
Custom licence

At a glance

Universal archivist for personal file archives (Dropbox/B2/Gmail-takeout/local-mount/hard-drive-dump).

  • Works in 3 steps: Inventory → Crawl → Ingest
  • Tasks that involve Web scraping
  • SKILL.md covers Safety gate (REQUIRED, no…, What this is, Concepts and Protocol, plus 6 more sections
  • Calls python3; reaches schemas.openxmlformats.org

What it does

Archive Crawler is an agent skill from inbrainfun/inbrain. Universal archivist for personal file archives (Dropbox/B2/Gmail-takeout/local-mount/hard-drive-dump). Filters for high-value content (the user's own writing, ideas, relationships) and surfaces it interactively. REFUSES TO RUN without an explicit inbrain.yml archive-crawler.scanpaths: allow-list.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.

It sits in Data & Analytics, covering Web scraping and Email management. It works with Dropbox and Gmail.

When your agent uses it

  • Tasks that involve Web scraping
  • Tasks that involve Email management

Example prompts

  • “/archive-crawler”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Inventory
  2. Crawl
  3. Ingest

What it can do on your machine

Read from SKILL.md and the folder at commit 5231990. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • schemas.openxmlformats.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Archive Crawler loads about 2.6k tokens when it runs. Until then it costs about 79 tokens; SKILL.md has 832 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~79
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 832 words (~2,640 tokens).

“archive-crawler refuses to run unless archive-crawler.scan_paths: is explicitly set in inbrain.yml. This is a deliberate safety fence against the agent over-scoping a scan and ingesting sensitive content (tax PDFs, medical records, credentials).”

— opening of SKILL.md by inbrainfun, Custom licence
name
archive-crawler
version
0.1.0
triggers
crawl my archive, find gold in my archive, archive crawler, scan my dropbox for, mine my old files for
mutating
true
writes_pages
true
writes_to
originals/, personal/, ideas/

Read the full SKILL.md on GitHub

Files

SKILL.md and 1 other file in skills/archive-crawler of inbrainfun/inbrain.

  • SKILL.md
  • routing-eval.jsonl

Open the folder on GitHubat commit 5231990

Used in 1 other repository

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in inbrainfun/inbrain, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Archive Crawler next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Archive Crawler compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Archive Crawler this skillinbrainfun/inbrain1421 repos~2.6kAutomated safety check: PassCustom licence
SupercompressSupercompress/Supercompress106—~519Automated safety check: PassMIT
SupercompressSupercompress/Supercompress106—~393Automated safety check: PassMIT
Abm Outboundsundial-org/awesome-openclaw-skills663—~1.8kAutomated safety check: PassNone
Tmuxtrpc-group/trpc-agent-go1.8k24 repos~868Automated safety check: PassApache-2.0
Google WorkspaceNousResearch/hermes-agent252k3 repos~3.5kAutomated safety check: PassMIT

Similar skills

  • Supercompress

    Supercompress/Supercompress

    Always-on context compression for Grok Build. An agent skill from Supercompress/Supercompress.

    106 GitHub stars~519 tokensUpdated yesterday
    Productivity & AutomationAuto-check passed
  • Supercompress

    Supercompress/Supercompress

    Compress bulky coding-agent context (tool dumps, logs, diffs, files, scrapes) before it burns tokens.

    106 GitHub stars~393 tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Abm Outbound

    sundial-org/awesome-openclaw-skills

    Multi-channel ABM automation that turns LinkedIn URLs into coordinated outbound campaigns.

    663 GitHub stars~1.8k tokensUpdated 7 mo ago
    Marketing & SEOAuto-check passed
  • Tmux

    trpc-group/trpc-agent-go

    Remote-control tmux sessions for interactive CLIs by sending keystrokes and scraping pane output.

    1.8k GitHub starsUsed in 24 repos~868 tokens
    Data & AnalyticsAuto-check passed
  • Google Workspace

    NousResearch/hermes-agent

    Gmail, Calendar, Drive, Docs, Sheets via gws CLI or Python. An agent skill from NousResearch/hermes-agent.

    252k GitHub starsUsed in 3 repos~3.5k tokens
    Documents & OfficeAuto-check passed
  • Gmail Inbox Watcher

    googleworkspace/cli

    Streams new Gmail messages as NDJSON from the gws command line tool using Google Pub/Sub, with label filters, batch settings and optional per-message files.

    31k GitHub starsUsed in 1 repo~476 tokens
    Productivity & AutomationAuto-check passed

More from inbrainfun/inbrain

All 52 skills in this repo
  • Academic Verify

    inbrainfun/inbrain

    Verify a research claim or academic citation by tracing it through publication → methodology → raw data → independent replication.

    142 GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Cold Start

    inbrainfun/inbrain

    Day-one data bootstrapping for a new brain. An agent skill from inbrainfun/inbrain.

    142 GitHub starsUsed in 1 repo~4.7k tokens
    Auto-check passed
  • Enrich

    inbrainfun/inbrain

    Enrich brain pages with tiered enrichment protocol. An agent skill from inbrainfun/inbrain.

    142 GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed
  • Idea Lineage

    inbrainfun/inbrain

    Trace one idea's evolution through the brain: first mention, best articulation, related concepts, reversals, contradictions, abandoned branches, and the current live version.

    142 GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed
  • Maintain

    inbrainfun/inbrain

    Brain health checks: back-link enforcement, citation audit, filing validation, stale info detection, orphan pages, and benchmarks.

    142 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Schema Author

    inbrainfun/inbrain

    Evolve your brain's schema pack. An agent skill from inbrainfun/inbrain.

    142 GitHub starsUsed in 1 repo~3.3k tokens
    Auto-check passed

Works with

Questions about Archive Crawler

What does Archive Crawler do?

Universal archivist for personal file archives (Dropbox/B2/Gmail-takeout/local-mount/hard-drive-dump). Archive Crawler is an agent skill from inbrainfun/inbrain. Universal archivist for personal file archives (Dropbox/B2/Gmail-takeout/local-mount/hard-drive-dump).

When should I use Archive Crawler?

Archive Crawler fits situations like: tasks that involve Web scraping; tasks that involve Email management.

How do I install Archive Crawler in Claude Code?

Run `npx skills add inbrainfun/inbrain --skill archive-crawler -a claude-code`. Or copy the skill folder (skills/archive-crawler in inbrainfun/inbrain) into .claude/skills/archive-crawler in your project. Claude Code loads it when a task matches its description.

How do I install Archive Crawler in Codex?

Run `npx skills add inbrainfun/inbrain --skill archive-crawler -a codex`. Or copy the skill folder (skills/archive-crawler in inbrainfun/inbrain) into .agents/skills/archive-crawler in your project. Codex loads it when a task matches its description.

Can I use Archive Crawler in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add inbrainfun/inbrain --skill archive-crawler -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/archive-crawler, .gemini/skills/archive-crawler, .github/skills/archive-crawler and .opencode/skills/archive-crawler in your project.

What does Archive Crawler need to run?

Going by SKILL.md and its folder, Archive Crawler needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Archive Crawler access the network?

SKILL.md names 1 domain. In commands or code: schemas.openxmlformats.org; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Archive Crawler safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Archive Crawler use?

Archive Crawler has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Archive Crawler use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Archive Crawler?

Skills that share tags, products or a category with Archive Crawler: Supercompress (Supercompress/Supercompress, 106 stars), Supercompress (Supercompress/Supercompress, 106 stars), Abm Outbound (sundial-org/awesome-openclaw-skills, 663 stars) and Tmux (trpc-group/trpc-agent-go, 1.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Archive Crawler?

inbrainfun (a GitHub user) maintains it in inbrainfun/inbrain, which has 142 GitHub stars. The repository holds 52 skills in this directory. The repository was last updated on July 17, 2026.

Source: inbrainfun/inbrain on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.