Agent skill

Searching Codebases

by oaustegard in oaustegard/claude-skills

Binding-resolved Python symbol queries via pyright — every true caller (--refs), the real definition (--def), or an inferred signature (--hover) of a .py symbol, excluding the same-named false…

MITAuto-check passedFrontend & Design

Install Searching Codebases

skills CLI
$ npx skills add oaustegard/claude-skills --skill searching-codebases -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install oaustegard/claude-skills searching-codebases --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/searching-codebases .claude/skills/searching-codebases && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
searching-codebases
GitHub stars
150
Token cost
~2.1k tokens
SKILL.md length
859 words
Files
15 (incl. scripts)
Skills in repo
93
Repo updated
First seen
Licence
MIT

At a glance

Binding-resolved Python symbol queries via pyright — every true caller (--refs), the real definition (--def), or an inferred signature (--hover) of a .py symbol, excluding the same-named false…

  • A task needs ALL callers
  • SKILL.md covers When NOT to use this skill, Prerequisites, Primary Command and Search Modes, plus 7 more sections
  • Runs Python scripts from its folder; calls python3, uv and rg
  • Users of a named Python symbol and grep would over-match

What it does

Searching Codebases is an agent skill from oaustegard/claude-skills. Binding-resolved Python symbol queries via pyright — every true caller (--refs), the real definition (--def), or an inferred signature (--hover) of a .py symbol, excluding the same-named false positives text search cannot tell apart. Use when a task needs ALL callers or users of a named Python symbol and grep would over-match. Python only, and only for the caller/definition question: measured 2026-07-04 on real issue-localization tasks, plain ripgrep tied or beat the semantic and indexed-regex tiers at 4-60x less…

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 18 other files, including scripts (for example `CHANGELOG.md`, `README.md` and `scripts/code_rag.py`).

It sits in Frontend & Design, covering Internationalization. It works with Python. The repository describes itself as: My collection of Claude skills. The licence is MIT.

When your agent uses it

  • A task needs ALL callers
  • Users of a named Python symbol and grep would over-match

Example prompts

  • “/searching-codebases”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 559a6cd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 7 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • uv
    • rg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Searching Codebases loads about 2.1k tokens when it runs. Until then it costs about 151 tokens; SKILL.md has 859 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~151
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from oaustegard/claude-skills at commit 559a6cd, republished under its MIT licence (© oaustegard). 859 words, ~2,072 tokens.

Download SKILL.mdSave it as .claude/skills/searching-codebases/SKILL.md (or your agent's skills folder). This skill also uses 14 other files; get the full folder from GitHub.
name
searching-codebases
description
Binding-resolved Python symbol queries via pyright — every true caller (--refs), the real definition (--def), or an inferred signature (--hover) of a .py symbol, excluding the same-named false positives text search cannot tell apart. Use when a task needs ALL callers or users of a named Python symbol and grep would over-match. Python only, and only for the caller/definition question: measured 2026-07-04 on real issue-localization tasks, plain ripgrep tied or beat the semantic and indexed-regex tiers at 4-60x less wall-clock, so those tiers survive below without being a default.
metadata.version
2.5.0

Searching Codebases

Find code in any codebase by pattern or concept. One entry point, two search strategies, automatic routing.

When NOT to use this skill

Python callers and definitions only. Every other code question is cheaper elsewhere, and the measurement below is why this skill is edge-case-only.

SituationUse
What symbols does this file contain?tree-sitting
Where is X defined, so I can read it?tree-sitting
Any structure question, any languagetree-sitting
A literal string or regexplain ripgrep
Which files are most about a concept?bm25
First look at an unfamiliar repoexploring-codebases

Non-Python code has no pyright binding resolution to offer, so there is no version of this skill that applies to it.

Prerequisites

bash
uv tool install ripgrep

tree-sitting installs automatically when needed — for --expand context expansion and for the binding-resolved --refs/--def/--hover tier, which uses it to resolve symbol positions. Only the bare tree-sitter package is fetched; the language grammars ship bundled.

Primary Command

bash
SKILL_DIR=/mnt/skills/user/searching-codebases

python3 $SKILL_DIR/scripts/search.py SOURCE "query1" ["query2" ...] [OPTIONS]

SOURCE is any of:

  • Local directory path
  • GitHub URL (downloads tarball automatically)
  • uploads (uses /mnt/user-data/uploads/)
  • project (uses /mnt/project/)
  • Path to a .zip or .tar.gz archive

Search Modes

Regex mode (patterns, identifiers, literal text):

bash
python3 $SKILL_DIR/scripts/search.py ./repo "def handle_error"
python3 $SKILL_DIR/scripts/search.py ./repo "class.*Exception" --regex
python3 $SKILL_DIR/scripts/search.py ./repo "TODO|FIXME|HACK"

Semantic mode (concepts, natural language):

bash
python3 $SKILL_DIR/scripts/search.py ./repo "retry logic with backoff" --semantic
python3 $SKILL_DIR/scripts/search.py ./repo "authentication flow"
python3 $SKILL_DIR/scripts/search.py ./repo "error handling strategy"

Auto-detection: short queries and code-like tokens → regex. Multi-word natural language → semantic. Override with --regex or --semantic.

Binding-resolved mode (Python only — pyright via the python-lsp skill):

bash
python3 $SKILL_DIR/scripts/search.py ./repo --refs SYMBOL    # find all real uses
python3 $SKILL_DIR/scripts/search.py ./repo --def SYMBOL     # go-to-definition
python3 $SKILL_DIR/scripts/search.py ./repo --hover SYMBOL   # inferred type/signature

Regex mode matches text, so a cross-reference for a function false-positives on shadowed and same-named-but-unrelated symbols. --refs is binding-resolved: pyright excludes the unrelated same-named symbol and follows imports. Use it when you need a true "find all callers/users" for a .py symbol, not a text grep.

The tier is engaged lazily — pyright's index cost is paid only when you ask for --refs/--def/--hover, never on ordinary searches. It is Python-only; for non-.py sources, or when pyright/node is unavailable, it prints a one-line degradation note and falls back to the regex text path. Each takes a single bare symbol name and is mutually exclusive with the other two and with text queries.

Options

  • --regex / --semantic: Force search mode
  • --refs SYMBOL / --def SYMBOL / --hover SYMBOL: Binding-resolved Python queries via pyright (see Binding-resolved mode above)
  • --expand: Return full function bodies via tree-sitting AST context
  • --benchmark: Compare indexed regex vs brute-force ripgrep
  • --branch NAME: Git branch for GitHub URLs (default: main)
  • --skip DIRS: Comma-separated directories to skip
  • --json: Machine-readable output
  • -v: Show index stats and query routing decisions

How It Works

Regex search builds a sparse n-gram inverted index over all files. Queries are decomposed into literal fragments, looked up in the index to identify candidate files (typically 90-99% reduction), then verified with ripgrep. Frequency-weighted n-grams make rare character sequences more selective.

Semantic search builds a TF-IDF index over code chunks (functions, classes, structural entries). Queries are ranked by cosine similarity.

Context expansion (--expand) uses tree-sitting's AST cache to identify function/class boundaries, returning complete structural units rather than line fragments. On first use, tree-sitting scans the repo (~700ms for 250 files); subsequent expansions are sub-millisecond.

Small codebases (< 20 files) skip indexing entirely — direct ripgrep is faster when there's nothing to narrow.

Mixed Queries

Multiple queries can use different modes in a single invocation. Each query is auto-routed independently, and indexes are built once per mode:

bash
python3 $SKILL_DIR/scripts/search.py ./repo \
  "class.*Error" \
  "error recovery strategy" \
  "def retry"
Show full SKILL.md (334 more words)Show less

Dependencies

  • tree-sitting: Provides AST context expansion for --expand and the symbol→position resolution that seeds the binding-resolved tier (--refs/--def/--hover). Auto-installs the bare tree-sitter package when either is used (grammars are bundled). Regex and semantic search work without it.
  • ripgrep: Required for regex verification. Install via uv tool install ripgrep.
  • scikit-learn: Required for semantic mode. Installs automatically.
  • python-lsp: Provides the binding-resolved tier (--refs/--def/--hover). Self-bootstraps pyright on first use and requires system node (v18+). Not required — without it those flags degrade to the regex text path.

When to Use — narrow, by design

The ONE recommended use: binding-resolved Python symbol queries.

  • "find all callers of X" / "where is X really defined" for a .py symbol, when same-named-but-unrelated symbols would pollute a text grep. Empirical basis: rg get on psf/requests returned 232 hits, 224 of them false; --refs get excluded all 224 (2026-06-15).

When NOT to Use — which is most of the time

Everything else. Measured head-to-head on real issue-localization tasks (7 scikit-learn issues with merged fix-PRs, gold = PR diff files, 2026-07-04, replicating the file-discovery metric of arXiv:2602.11988):

  • Literal tokens / identifiers: naive rg -l tied or beat the indexed tier on recall@10 in every instance, at 0.4s vs 25s.
  • Concept / natural-language search: the TF-IDF semantic tier never beat identifier grep — not even on identifier-poor issues, which are themselves rare (~0.3% of merged-PR traffic in the sample).
  • First encounter / "what is this repo": use exploring-codebases.
  • Repos under ~20 files: read them.

The self-test before invoking: would plain rg return the same answer? If yes, use rg. The indexed-regex and semantic tiers are retained for completeness and for corpora where they may yet earn their cost (very large repos, non-code document collections), but they carry the burden of proof.

Files

  • scripts/search.py — Entry point, query routing, output formatting
  • scripts/resolve.py — Input source resolution (GitHub, uploads, archives)
  • scripts/context.py — tree-sitting-based AST context expansion
  • scripts/ngram_index.py — Sparse n-gram inverted index, regex decomposition
  • scripts/sparse_ngrams.py — Core n-gram algorithms, frequency weights
  • scripts/code_rag.py — TF-IDF semantic search over code chunks
  • scripts/lsp_refs.py — Binding-resolved Python tier: symbol→position resolution (tree-sitting), pyright queries (python-lsp), soft fallback

© oaustegard, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 14 other files (scripts) in searching-codebases of oaustegard/claude-skills.

  • SKILL.md
  • CHANGELOG.md
  • README.md
  • scripts/code_rag.py
  • scripts/context.py
  • scripts/lsp_refs.py
  • scripts/ngram_index.py
  • scripts/resolve.py
  • scripts/search.py
  • scripts/sparse_ngrams.py
  • tests/fixture/pkg/__init__.py
  • tests/fixture/pkg/models.py
  • tests/fixture/pkg/other.py
  • tests/fixture/pkg/service.py
  • tests/test_lsp_refs.py

Open the folder on GitHubat commit 559a6cd

Compare with similar skills

Searching Codebases next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Searching Codebases compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Searching Codebases this skilloaustegard/claude-skills150—~2.1kAutomated safety check: PassMIT
Claude Desktop Chinese Localizationjavaht/claude-desktop-zh-cn7.5k—~1.6kAutomated safety check: PassMIT
Ok Script I18nAliceJump/ok-end-field554—~828Automated safety check: PassAGPL-3.0
I18n Development4thfever/cultivation-world-simulator2.1k—~451Automated safety check: PassCustom licence
Ok Script I18nbaoxin1100/ok-kes101—~903Automated safety check: PassNone
I18nbrickbots/PiFinder250—~2.8kAutomated safety check: PassGPL-3.0

Similar skills

  • Claude Desktop Chinese Localization

    javaht/claude-desktop-zh-cn

    Adds missing Simplified and Traditional Chinese translations to the Claude Desktop Chinese patch across three layers, then checks how many mappings actually hit.

    7.5k GitHub stars~1.6k tokensUpdated 2 days ago
    Frontend & DesignAuto-check passed
  • Ok Script I18n

    AliceJump/ok-end-field

    Maintain gettext translations for ok-script task UI and runtime messages.

    554 GitHub stars~828 tokensUpdated today
    Frontend & DesignAuto-check passed
  • I18n Development

    4thfever/cultivation-world-simulator

    国际化 (i18n) 开发指南。在添加新文本、创建物品/事件、修改翻译或管理 PO/MO 文件时使用. An agent skill from 4thfever/cultivation-world-simulator.

    2.1k GitHub stars~451 tokensUpdated 1 mo ago
    Frontend & DesignAuto-check passed
  • Ok Script I18n

    baoxin1100/ok-kes

    Add, sync, repair, and compile gettext translations for ok-script Python task classes and task metadata.

    101 GitHub stars~903 tokensUpdated 6 days ago
    Frontend & DesignAuto-check passed
  • I18n

    brickbots/PiFinder

    PiFinder's internationalization (i18n) workflow — marking strings for translation, running the Babel extract/update/compile pipeline, adding or updating language translations, and filling in missing…

    250 GitHub stars~2.8k tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • Fluent Migration

    BrowserWorks/waterfox-android

    A skill your agent uses when a patch or local changes rename, restructure, move, or replace Fluent (.ftl) strings - or migrate legacy .properties strings to Fluent - and you need a migration recipe…

    376 GitHub stars~4.2k tokensUpdated 15 days ago
    Frontend & DesignAuto-check passed

More from oaustegard/claude-skills

All 93 skills in this repo
  • Bluesky Zeitgeist Sampler

    oaustegard/claude-skills

    Deprecated sampler that captures short windows of the Bluesky firehose, clusters trending terms and builds an HTML report; replaced by the browsing-bluesky skill.

    150 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Vega-Lite Interactive Charts

    oaustegard/claude-skills

    Builds interactive Vega-Lite charts from uploaded data: analyzes the fields, picks five to ten fitting chart types, and produces a React artifact with the data embedded inline.

    150 GitHub stars~2.1k tokensUpdated 6 days ago
    Auto-check passed
  • Single-File HTML Composer

    oaustegard/claude-skills

    Builds self-contained single-file HTML pages such as reports, decks, postmortems, flowcharts and prototypes from a small spec using a bundled Python composer and templates.

    150 GitHub stars~3.2k tokensUpdated 6 days ago
    Auto-check passed
  • Declauding

    oaustegard/claude-skills

    Rewrites model-sounding prose into plain technical writing and checks that every claim survives, for PR text, docs, commit messages and similar drafts.

    150 GitHub stars~5.1k tokensUpdated 6 days ago
    Auto-check passed
  • Forecasting Reverso

    oaustegard/claude-skills

    Zero-shot univariate time series forecasting using the Reverso foundation model (NumPy/Numba CPU-only inference).

    150 GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Preact Developer

    oaustegard/claude-skills

    Guides building standards-based Preact apps with native-first choices, HTM syntax, import maps and vendored ESM, from single-file demos to larger builds.

    150 GitHub stars~4.6k tokensUpdated 6 days ago
    Auto-check passed

Works with

Questions about Searching Codebases

What does Searching Codebases do?

Binding-resolved Python symbol queries via pyright — every true caller (--refs), the real definition (--def), or an inferred signature (--hover) of a .py symbol, excluding the same-named false…. Searching Codebases is an agent skill from oaustegard/claude-skills.py symbol, excluding the same-named false positives text search cannot tell apart.

When should I use Searching Codebases?

Searching Codebases fits situations like: A task needs ALL callers; users of a named Python symbol and grep would over-match.

How do I install Searching Codebases in Claude Code?

Run `npx skills add oaustegard/claude-skills --skill searching-codebases -a claude-code`. Or copy the skill folder (searching-codebases in oaustegard/claude-skills) into .claude/skills/searching-codebases in your project. Claude Code loads it when a task matches its description.

How do I install Searching Codebases in Codex?

Run `npx skills add oaustegard/claude-skills --skill searching-codebases -a codex`. Or copy the skill folder (searching-codebases in oaustegard/claude-skills) into .agents/skills/searching-codebases in your project. Codex loads it when a task matches its description.

Can I use Searching Codebases in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oaustegard/claude-skills --skill searching-codebases -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/searching-codebases, .gemini/skills/searching-codebases, .github/skills/searching-codebases and .opencode/skills/searching-codebases in your project.

What does Searching Codebases need to run?

Going by SKILL.md and its folder, Searching Codebases needs Python for the scripts in its folder and the command-line tools its instructions call (python3, uv and rg). Our summary lists: Python 3.

Does Searching Codebases access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Searching Codebases safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Searching Codebases use?

Searching Codebases is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Searching Codebases use?

About 2.1k tokens (SKILL.md is roughly 8.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Searching Codebases?

Skills that share tags, products or a category with Searching Codebases: Claude Desktop Chinese Localization (javaht/claude-desktop-zh-cn, 7.5k stars), Ok Script I18n (AliceJump/ok-end-field, 554 stars), I18n Development (4thfever/cultivation-world-simulator, 2.1k stars) and Ok Script I18n (baoxin1100/ok-kes, 101 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Searching Codebases?

oaustegard (a GitHub user) maintains it in oaustegard/claude-skills, which has 150 GitHub stars. The repository holds 93 skills in this directory. The repository was last updated on October 2, 2026.

Source: oaustegard/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.