Agent skill

Exploring Codebases

by oaustegard in oaustegard/claude-skills

First-encounter orientation on a repository nobody here has worked in yet.

MITAuto-check passed

Install Exploring Codebases

skills CLI
$ npx skills add oaustegard/claude-skills --skill exploring-codebases -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install oaustegard/claude-skills exploring-codebases --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/exploring-codebases .claude/skills/exploring-codebases && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
exploring-codebases
GitHub stars
150
Token cost
~2.3k tokens
SKILL.md length
973 words
Files
4 (incl. references)
Skills in repo
93
Repo updated
First seen
Licence
MIT

At a glance

First-encounter orientation on a repository nobody here has worked in yet.

  • Works in 5 steps: Setup (once per session) → Get the repo (tarball, not per-file) → Structural scan → …
  • I just cloned this
  • SKILL.md covers Workflow, When to Use This vs Other Skills, Delegating to subagents and Notes
  • Calls uv, curl and python3; reaches api.github.com; needs GH_TOKEN

What it does

Exploring Codebases is an agent skill from oaustegard/claude-skills. First-encounter orientation on a repository nobody here has worked in yet. Runs a fixed five-step workflow — venv setup, tarball fetch, tree-sitting structural scan, featuring synthesis, then reasoning over the two — and yields an account of what the repo contains and how it is arranged, optionally written out as FEATURES.md. Use for "I just cloned this", "what is this repo", "what does this do", "explore this repo", "give me an orientation", "what are the main features", "review what's new in this repo", or…

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `CHANGELOG.md`, `README.md` and `references/subagent-delegation.md`).

It works with GitHub and Python. The repository describes itself as: My collection of Claude skills. The licence is MIT.

When your agent uses it

  • I just cloned this
  • What is this repo
  • What does this do
  • Explore this repo

Example prompts

  • “I just cloned this”
  • “what is this repo”
  • “what does this do”
  • “/exploring-codebases”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Setup (once per session)
  2. Get the repo (tarball, not per-file)
  3. Structural scan
  4. Feature synthesis
  5. Reason about the combined output

What it can do on your machine

Read from SKILL.md and the folder at commit 559a6cd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • curl
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GH_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Exploring Codebases loads about 2.3k tokens when it runs, and up to ~2.9k if it reads all its reference files. Until then it costs about 233 tokens; SKILL.md has 973 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~233
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from oaustegard/claude-skills at commit 559a6cd, republished under its MIT licence (© oaustegard). 973 words, ~2,322 tokens.

Download SKILL.mdSave it as .claude/skills/exploring-codebases/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
exploring-codebases
description
First-encounter orientation on a repository nobody here has worked in yet. Runs a fixed five-step workflow — venv setup, tarball fetch, tree-sitting structural scan, featuring synthesis, then reasoning over the two — and yields an account of what the repo contains and how it is arranged, optionally written out as _FEATURES.md. Use for "I just cloned this", "what is this repo", "what does this do", "explore this repo", "give me an orientation", "what are the main features", "review what's new in this repo", or before starting work in a codebase you have not seen. This is the divergent what's-here skill. Route elsewhere for: a named symbol, a file's structure or a line range (tree-sitting); all callers of a Python symbol (searching-codebases); teaching a human the codebase through exercises (orienting-codebases); fetching or cloning a repo without analysing it (accessing-github-repos, cloning-project).
metadata.version
2.5.2

Exploring Codebases

Exploratory code analysis for unfamiliar repositories. Orchestrates tree-sitting (structural) and featuring (semantic) over a local copy.

Workflow

Five numbered steps, in order. Do not skip step 0.

0. Setup (once per session)
bash
uv venv /home/claude/.venv 2>/dev/null
uv pip install tree-sitter --python /home/claude/.venv/bin/python
export PYTHON=/home/claude/.venv/bin/python
export TREESIT=/mnt/skills/user/tree-sitting/scripts/treesit.py
export GATHER=/mnt/skills/user/featuring/scripts/gather.py

If step 2's --stats reports Symbols: 0 on a repo you know contains code, the tree-sitter core package isn't installed — come back here and install it (the engine bundles its own grammars and does NOT use tree-sitter-language-pack). Treesit exits 0 and prints no error in that case, so zero symbols is the only signal you get. There is no Errors: line: that one appears for parse failures, and an absent parser never reaches parsing. The full signal is in the tree-sitting skill's Setup section.

1. Get the repo (tarball, not per-file)
bash
OWNER=...
REPO=...
REF=main                    # branch name, tag, or SHA. For a PR: pull/N/head
curl -sL -H "Authorization: Bearer $GH_TOKEN" \
  "https://api.github.com/repos/$OWNER/$REPO/tarball/$REF" -o /tmp/$REPO.tar.gz
mkdir -p /tmp/$REPO && tar -xzf /tmp/$REPO.tar.gz -C /tmp/$REPO --strip-components=1
ls /tmp/$REPO | head        # sanity check — did extraction land?

One HTTP call gets the whole repo. Do NOT curl README, cat files, or fetch via contents/PATH first — they're in the tarball. The Authorization header is only needed for private repos; public repos work without it.

Ref selection matters. If exploring a feature branch, PR, or tag, set REF accordingly. The default main will silently give you stale code if the question is about an unmerged branch.

2. Structural scan
bash
$PYTHON $TREESIT /tmp/$REPO --stats

Read the output. It gives file counts, symbol counts, languages, and per-directory symbol density. This IS the orienting artifact — treat it as the product of this step, not warm-up.

Drill only if you have a specific question. For pure "what is this repo" exploration, skip drilling and go to step 3 — featuring surfaces the interesting paths for you. Drill when a user asked about a specific subsystem, or when step 3's output raises a question that needs source.

When you do drill, batch queries in one invocation. Every treesit call pays the full scan cost. Multiple queries added to the same command share that scan and each additional query adds ~0ms. If you're about to make a second treesit call on the same path, fold it into the first.

bash
# GOOD — one scan, three answers
$PYTHON $TREESIT /tmp/$REPO --path=SUBDIR --detail=full \
  'find:*Handler*:function' 'source:main' 'refs:Config'

# BAD — three scans, three answers (3× the cost for the same information)
$PYTHON $TREESIT /tmp/$REPO --path=SUBDIR --detail=full
$PYTHON $TREESIT /tmp/$REPO 'find:*Handler*:function'
$PYTHON $TREESIT /tmp/$REPO 'refs:Config'
3. Feature synthesis

Pick the mode from your DELIVERABLE, before you run it.

Your deliverableCommandSize
Your own understanding — a review, an orientation read, answering a question--orient~115 lines
A written _FEATURES.md that must cite every symbolfull outputthousands of lines
bash
# Default. Complexity assessment, decomposition ranking, directory tree, entry points.
$PYTHON $GATHER /tmp/$REPO --skip tests,.github,node_modules --orient

# Only when you are about to WRITE the inventory into a file:
$PYTHON $GATHER /tmp/$REPO --skip tests,.github,node_modules --source-budget 8000

Output includes a "Candidate areas for sub-files (by symbol density)" list near the top — that's your drill-target picker, ranked.

Never pipe the full output through head. If you are about to truncate it, --orient was the correct mode and you have paid for thousands of lines you will not read. One review's full gather ran to 5,697 lines and was cut at line 120; every finding in it came from treesit drilling and targeted reads instead. --orient returns the ~115 lines that get used. The full mode's symbol inventory exists to be CITED, not read.

4. Reason about the combined output

Synthesize 2+3: capabilities, feature groups, architecture, entry points, anomalies. Produce _FEATURES.md when warranted. This is the LLM step; everything before was mechanical.

When to Use This vs Other Skills

SituationUse
"I just cloned this, what is it?"exploring-codebases (this skill)
"Where is the retry logic?"searching-codebases
"Find all files matching class.*Error"searching-codebases
"Show me the symbols in auth.py"tree-sitting directly
"Which files are most about CSRF / sessions / queryset filtering?"bm25
"Rank these docs by relevance to a multi-word concept"bm25
"Document what this codebase does"featuring directly
"Teach me this codebase" (a human is learning)orienting-codebases
"Get me this repo" — fetch, no analysisaccessing-github-repos, cloning-project

Exploring is the divergent skill — you don't know what you're looking for yet. Searching is the convergent skill — you know what you want.

orienting-codebases runs the same tree-sitting + featuring pipeline and is the nearest thing in the catalogue to this skill. The split is the audience: this one builds Claude's understanding so work can proceed; that one builds the user's understanding through guided exercises and HTML artifacts. If nobody is being taught, this is the right skill.

Show full SKILL.md (321 more words)Show less
Pairing bm25 with this workflow

Once steps 2–3 have surfaced the rough shape of the repo, bm25 is the natural complement when you want ranked content search beyond grep and beyond exact-symbol lookup. It ranks files by lexical relevance to a multi-word query, which is useful for "what's this codebase actually about when I search for X?" — particularly when you don't yet know the symbol name to feed to tree-sitting.

bash
BM25=/mnt/skills/user/bm25/scripts/bm25.py

# Pass multiple queries — index builds once, all queries reuse it
python3 $BM25 /tmp/$REPO 'auth flow' 'session backend' 'middleware pipeline' \
  --exclude 'tests/*' --exclude '*/tests/*' --top-k 5

Two patterns that pair especially well:

  1. bm25 → tree-sitting. Use bm25 to find the top-ranked files for a concept; then tree-sitting source:Symbol:path/to/file.py to read the actual implementation.
  2. bm25 with --exclude 'tests/*'. Test directories tend to dominate keyword queries because test names redundantly mention domain terms. Excluding them up front lands you on implementation files.

bm25 is corpus-agnostic — it'll also work on project knowledge stores or uploads/ if your exploration spans docs, transcripts, or PDFs.

Delegating to subagents

Only when the repo is large (>1000 files or several distinct subsystems) and this environment exposes a subagent tool (Agent/Task in Claude Code and CCotw). Claude.ai chat and bare-skill runs have none: run steps 2-4 inline and skip this entirely. Never simulate fan-out by other means when the tool is absent.

Steps 2-3 stay inline either way. Only step 4's judgment work fans out, one agent per subsystem, and a subagent inherits nothing -- not the conversation, not this file, not the knowledge that scan artifacts are already on disk. Read references/subagent-delegation.md before writing the first agent prompt; it carries the four things every prompt must include and what happens when they are missing.

Notes

  • Large repos (>100 files): use --skip tests,vendored,docs,... in step 2 to focus the scan.
  • Monorepos: treat each package/service as a separate exploration. Generate per-subsystem _FEATURES.md files linked from a root index.
  • Drill heuristics (if step 2 drilling is warranted): directories with high symbol-to-file ratio (dense logic), entry-point names (main, cli, app, server, routes), files with many imports (integration points).

© oaustegard, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in exploring-codebases of oaustegard/claude-skills.

  • SKILL.md
  • CHANGELOG.md
  • README.md
  • references/subagent-delegation.md

Open the folder on GitHubat commit 559a6cd

Compare with similar skills

Exploring Codebases next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Exploring Codebases compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Exploring Codebases this skilloaustegard/claude-skills150—~2.3kAutomated safety check: PassMIT
GitHub Deep Researchbytedance/deer-flow83k5 repos~1.3kAutomated safety check: PassMIT
Update V8 Versionopeninterpreter/openinterpreter69k2 repos~845Automated safety check: PassApache-2.0
Merge Dependabot PRsonyx-dot-app/onyx32k1 repos~2.2kAutomated safety check: PassMIT
Create Cuda Python Pull RequestNVIDIA/cuda-python3.4k—~1.1kAutomated safety check: PassApache-2.0
Final Release Reviewopenai/openai-agents-python30k—~5.4kAutomated safety check: PassMIT

Similar skills

  • GitHub Deep Research

    bytedance/deer-flow

    Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.

    83k GitHub starsUsed in 5 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Update V8 Version

    openinterpreter/openinterpreter

    Bumps the pinned v8 and rusty_v8 versions in Codex, validates the release-candidate path with the v8-canary check, and traces failures to upstream build changes.

    69k GitHub starsUsed in 2 repos~845 tokens
    DevOps & CloudAuto-check passed
  • Merge Dependabot PRs

    onyx-dot-app/onyx

    Triages and lands a batch of open Dependabot PRs in the Onyx repo, where main is gated exclusively by GitHub's merge queue: approves and enqueues green PRs, closes superseded duplicates, fixes…

    32k GitHub starsUsed in 1 repo~2.2k tokens
    DevelopmentAuto-check passed
  • Official

    Create a CUDA Python pull request from an approved personal or organization-owned fork, including the GitHub CLI GraphQL fallback for renamed organization-owned forks.

    3.4k GitHub stars~1.1k tokensUpdated today
    DevelopmentAuto-check passed
  • Final Release Review

    openai/openai-agents-python

    Official

    Assess a Python SDK release candidate or release plan against the previous release and recommend ship or block.

    30k GitHub stars~5.4k tokensUpdated today
    Product & Project ManagementAuto-check passed
  • Setup PR Branch

    Nuitka/Nuitka

    Set up a local branch that tracks a fork PR and can be pushed back to it, without fetching the whole fork.

    15k GitHub stars~1.1k tokensUpdated today
    DevelopmentAuto-check passed

More from oaustegard/claude-skills

All 93 skills in this repo
  • Bluesky Zeitgeist Sampler

    oaustegard/claude-skills

    Deprecated sampler that captures short windows of the Bluesky firehose, clusters trending terms and builds an HTML report; replaced by the browsing-bluesky skill.

    150 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Vega-Lite Interactive Charts

    oaustegard/claude-skills

    Builds interactive Vega-Lite charts from uploaded data: analyzes the fields, picks five to ten fitting chart types, and produces a React artifact with the data embedded inline.

    150 GitHub stars~2.1k tokensUpdated 5 days ago
    Auto-check passed
  • Single-File HTML Composer

    oaustegard/claude-skills

    Builds self-contained single-file HTML pages such as reports, decks, postmortems, flowcharts and prototypes from a small spec using a bundled Python composer and templates.

    150 GitHub stars~3.2k tokensUpdated 5 days ago
    Auto-check passed
  • Declauding

    oaustegard/claude-skills

    Rewrites model-sounding prose into plain technical writing and checks that every claim survives, for PR text, docs, commit messages and similar drafts.

    150 GitHub stars~5.1k tokensUpdated 5 days ago
    Auto-check passed
  • Forecasting Reverso

    oaustegard/claude-skills

    Zero-shot univariate time series forecasting using the Reverso foundation model (NumPy/Numba CPU-only inference).

    150 GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Preact Developer

    oaustegard/claude-skills

    Guides building standards-based Preact apps with native-first choices, HTM syntax, import maps and vendored ESM, from single-file demos to larger builds.

    150 GitHub stars~4.6k tokensUpdated 5 days ago
    Auto-check passed

Works with

Questions about Exploring Codebases

What does Exploring Codebases do?

First-encounter orientation on a repository nobody here has worked in yet. Exploring Codebases is an agent skill from oaustegard/claude-skills. First-encounter orientation on a repository nobody here has worked in yet.

When should I use Exploring Codebases?

Exploring Codebases fits situations like: I just cloned this; what is this repo; what does this do; explore this repo.

How do I install Exploring Codebases in Claude Code?

Run `npx skills add oaustegard/claude-skills --skill exploring-codebases -a claude-code`. Or copy the skill folder (exploring-codebases in oaustegard/claude-skills) into .claude/skills/exploring-codebases in your project. Claude Code loads it when a task matches its description.

How do I install Exploring Codebases in Codex?

Run `npx skills add oaustegard/claude-skills --skill exploring-codebases -a codex`. Or copy the skill folder (exploring-codebases in oaustegard/claude-skills) into .agents/skills/exploring-codebases in your project. Codex loads it when a task matches its description.

Can I use Exploring Codebases in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oaustegard/claude-skills --skill exploring-codebases -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/exploring-codebases, .gemini/skills/exploring-codebases, .github/skills/exploring-codebases and .opencode/skills/exploring-codebases in your project.

What does Exploring Codebases need to run?

Going by SKILL.md and its folder, Exploring Codebases needs the command-line tools its instructions call (uv, curl and python3) and credentials named GH_TOKEN. Our summary lists: Python 3.

Does Exploring Codebases access the network?

SKILL.md names 1 domain. In commands or code: api.github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Exploring Codebases safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Exploring Codebases use?

Exploring Codebases is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Exploring Codebases use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 564 tokens, read only when the agent opens those files.

What are the alternatives to Exploring Codebases?

Skills that share tags, products or a category with Exploring Codebases: GitHub Deep Research (bytedance/deer-flow, 83k stars), Update V8 Version (openinterpreter/openinterpreter, 69k stars), Merge Dependabot PRs (onyx-dot-app/onyx, 32k stars) and Create Cuda Python Pull Request (NVIDIA/cuda-python, 3.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Exploring Codebases?

oaustegard (a GitHub user) maintains it in oaustegard/claude-skills, which has 150 GitHub stars. The repository holds 93 skills in this directory. The repository was last updated on October 2, 2026.

Source: oaustegard/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.