Repomix Codebase Explorer
yamadashy/repomix
Packs a local or remote repository into a single AI-friendly file with the Repomix CLI, then reads and searches that output to explain structure, find patterns or report metrics.
Adds tree-sitter support for a new language to the codegraph project, then tests it and benchmarks extraction quality and retrieval value on real repositories.
$ npx skills add colbymchenry/codegraph --skill add-lang -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install colbymchenry/codegraph add-lang --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/colbymchenry/codegraph.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/add-lang .claude/skills/add-lang && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "add-lang" agent skill from https://github.com/colbymchenry/codegraph/tree/main/.claude/skills/add-lang into .claude/skills/add-lang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-lang", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/colbymchenry/codegraph/tree/main/.claude/skills/add-langType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add colbymchenry/codegraph --skill add-lang -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install colbymchenry/codegraph add-lang --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/colbymchenry/codegraph.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/add-lang .agents/skills/add-lang && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "add-lang" agent skill from https://github.com/colbymchenry/codegraph/tree/main/.claude/skills/add-lang into .agents/skills/add-lang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-lang", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add colbymchenry/codegraph --skill add-lang -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install colbymchenry/codegraph add-lang --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/colbymchenry/codegraph.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/add-lang .cursor/skills/add-lang && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "add-lang" agent skill from https://github.com/colbymchenry/codegraph/tree/main/.claude/skills/add-lang into .cursor/skills/add-lang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-lang", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/colbymchenry/codegraph.git --path .claude/skills/add-lang--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add colbymchenry/codegraph --skill add-lang -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install colbymchenry/codegraph add-lang --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/colbymchenry/codegraph.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/add-lang .gemini/skills/add-lang && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "add-lang" agent skill from https://github.com/colbymchenry/codegraph/tree/main/.claude/skills/add-lang into .gemini/skills/add-lang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-lang", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install colbymchenry/codegraph add-langInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add colbymchenry/codegraph --skill add-lang -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/colbymchenry/codegraph.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/add-lang .github/skills/add-lang && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "add-lang" agent skill from https://github.com/colbymchenry/codegraph/tree/main/.claude/skills/add-lang into .github/skills/add-lang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-lang", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add colbymchenry/codegraph --skill add-lang -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install colbymchenry/codegraph add-lang --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/colbymchenry/codegraph.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/add-lang .opencode/skills/add-lang && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "add-lang" agent skill from https://github.com/colbymchenry/codegraph/tree/main/.claude/skills/add-lang into .opencode/skills/add-lang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-lang", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
add-langAdds tree-sitter support for a new language to the codegraph project, then tests it and benchmarks extraction quality and retrieval value on real repositories.
This is a maintainer workflow for the codegraph project. Given a lowercase language token such as lua, elixir or zig, the agent wires a tree-sitter grammar and an extractor into codegraph's extraction pipeline, writes tests, and benchmarks extraction quality and retrieval value on 3 popular real-world repositories. If the language is already supported, it skips the wiring steps and goes straight to benchmarking, noting in its report that no code changed.
The work follows a checklist in order. The agent first looks for a grammar in the tree-sitter-wasms package or vendors a .wasm file, then health-checks the grammar with a check-grammar script before writing any extractor, because a grammar that is present can still be unusable, for example an old ABI that corrupts the shared WASM heap. It then dumps the grammar's AST node types, builds and links the local development build for the benchmark, which spawns real claude runs, and updates the docs.
It runs fully autonomously from the codegraph repo root, using node, git, gh and a logged-in claude CLI, and it never commits, pushes, publishes or tags, leaving every change for you to review.
10 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b635dd4. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
npmnodeclaudemakenpxghgitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npm, npx, gh and git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Add a Language to CodeGraph loads about 2.7k tokens when it runs. Until then it costs about 81 tokens; SKILL.md has 1,123 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from colbymchenry/codegraph at commit b635dd4, republished under its MIT licence (© colbymchenry). 1,123 words, ~2,685 tokens.
.claude/skills/add-lang/SKILL.md (or your agent's skills folder).Wire a new tree-sitter language into codegraph's extraction pipeline, prove it extracts real symbols on popular repos, and prove it beats no-codegraph for an agent. Runs fully autonomously — pick repos, benchmark, update docs, then report. Never commit, push, publish, or tag (house rule); leave all changes for the user to review.
The argument is the language token used throughout the Language union, e.g.
lua, elixir, zig. If none was given, ask which language. Use the lowercase
single-token form everywhere (csharp, not c#).
node, git, gh, and a logged-in
claude CLI (the benchmark spawns real claude -p runs).Copy this checklist and work through it in order:
- [ ] 1. Resolve language; bail early if already supported (just benchmark)
- [ ] 2. Find a grammar + health-check it (ABI / heap corruption)
- [ ] 3. Discover the grammar's AST node types (dump-ast.mjs)
- [ ] 4. Wire the language (4 files; sometimes a 5th core touch)
- [ ] 5. Build + verify-extraction loop until PASS
- [ ] 6. Add extraction tests; make them green
- [ ] 7. Auto-pick 3 popular repos by size tier; add to corpus.json
- [ ] 8. Benchmark all 3: extraction + with/without A/B
- [ ] 9. Update README + CHANGELOG
- [ ] 10. Report; do NOT commitCheck whether the language is already wired: look for the token in the
LANGUAGES const (src/types.ts) and the EXTRACTORS map
(src/extraction/languages/index.ts). If it is already supported (e.g.
typescript, rust), skip Steps 2–6 and go straight to benchmarking
(Steps 7–8) to validate/measure it — note in the report that no code changed.
ls node_modules/tree-sitter-wasms/out/ | grep -i <lang> # csharp -> c_sharpgrammars.ts resolves it from
tree-sitter-wasms automatically. (Many languages: elixir, zig, ocaml,
solidity, toml, yaml, …).wasm into src/extraction/wasm/ (like pascal /
scala / lua) and add the token to the vendored branch in Step 4.Always health-check before writing an extractor — a present grammar can still be unusable:
node scripts/add-lang/check-grammar.mjs <lang> path/to/valid-sample.<ext>It prints the grammar's ABI version and parses a valid sample many times in a multi-grammar runtime. If it FAILs (ERROR trees on valid code — an old ABI corrupting the shared WASM heap, which silently drops nested calls/imports on every file after the first; e.g. the tree-sitter-wasms Lua grammar is ABI 13 and fails), do NOT use that wasm. Vendor a newer (ABI 14/15) build instead:
npm pack @tree-sitter-grammars/tree-sitter-<lang> # often ships a prebuilt *.wasm
# or build one: npx tree-sitter build --wasm (needs Docker/emscripten)
cp <the>.wasm src/extraction/wasm/tree-sitter-<lang>.wasmthen add the token to the vendored branch in Step 4 and re-run check-grammar on the vendored path until it PASSes. If you cannot obtain a healthy wasm, STOP and tell the user.
Get a representative source file (write a small sample covering functions,
classes/structs, imports, enums; or curl a raw file from a known repo), then:
node scripts/add-lang/dump-ast.mjs <lang> path/to/sample.<ext>
# vendored grammar: pass the wasm path instead of the token
node scripts/add-lang/dump-ast.mjs src/extraction/wasm/tree-sitter-<lang>.wasm sample.<ext>The frequency table + field names (name:, parameters:, body:,
return_type:) tell you what to map. Open the existing extractor closest to the
language's paradigm as a model: rust.ts/scala.ts (functional, traits),
java.ts/csharp.ts (OO), python.ts/ruby.ts (scripting), go.ts
(top-level methods + receivers).
These are exact, fragile wiring — match the existing style precisely:
src/types.ts — TWO edits:'<lang>', to the LANGUAGES const (before 'unknown');'**/*.<ext>', to DEFAULT_CONFIG.include. Don't skip this — it's
the file-scan allowlist; without the glob, codegraph init finds 0
files even though detection/extraction are wired.src/extraction/grammars.ts — three maps:WASM_GRAMMAR_FILES: <lang>: 'tree-sitter-<lang>.wasm',EXTENSION_MAP: each file extension → '<lang>' (e.g. '.lua': 'lua',)getLanguageDisplayName: <lang>: '<Display Name>',<lang> to the
(lang === 'pascal' || lang === 'scala' || …) wasm-path branch.src/extraction/languages/<lang>.ts — new file exporting
export const <lang>Extractor: LanguageExtractor = { … }. Map the node types
from Step 3. Required fields: functionTypes, classTypes, methodTypes,
interfaceTypes, structTypes, enumTypes, typeAliasTypes,
importTypes, callTypes, variableTypes, nameField, bodyField,
paramsField. Add hooks as the grammar needs them (getSignature,
getVisibility, isExported, extractImport, visitNode, getReceiverType,
interfaceKind, enumMemberTypes, etc. — see
src/extraction/tree-sitter-types.ts).src/extraction/languages/index.ts — import { <lang>Extractor } from './<lang>'; and add <lang>: <lang>Extractor, to EXTRACTORS.Sometimes a 5th, core touch in src/extraction/tree-sitter.ts — variable
extraction has per-language branches in extractVariable (the generic fallback
only finds direct identifier/variable_declarator children). If the grammar
nests declared names (e.g. Lua's variable_declaration → variable_list), add a
} else if (this.language === '<lang>') branch there, mirroring the existing
ts/python/go ones. Import forms that aren't a distinct node (Lua/Ruby require
is a call) are handled in the extractor's visitNode hook instead.
npm run build # tsc + copy-assets (copies any vendored *.wasm into dist/)Index a small sample repo and check extraction:
( cd <sample-repo> && codegraph init -i )
node scripts/add-lang/verify-extraction.mjs <sample-repo> <lang>verify-extraction.mjs fails (exit 1) if the language isn't detected or only
file/import nodes were produced — the classic symptom of wrong node-type
names. On FAIL or a thin WARN: re-run dump-ast.mjs on a richer file, fix the
mappings in <lang>.ts, npm run build, re-index, re-verify. Repeat until
PASS.
Add to __tests__/extraction.test.ts, modeled on the Rust Extraction block:
detectLanguage assertion in describe('Language Detection')describe('<Lang> Extraction') block asserting functions/classes/imports
are extracted from an inline source string.npx vitest run __tests__/extraction.test.tsGreen before continuing.
Pick without asking. Find candidates, then curate 3 that are genuinely
<lang>-dominant, one per size tier:
gh search repos --language=<lang> --sort=stars --limit 40 \
--json fullName,stargazerCount,descriptionTiers (match corpus.json): Small <~150 files · **Medium** ~150–1500 ·
**Large** >~1500. Skip repos that are tagged <lang> but mostly another
language. Write one cross-file architecture question per repo (the kind that
needs tracing across files). Add a "<Language>" block to
.claude/skills/agent-eval/corpus.json (fields: name, repo, size,
files, question) so /agent-eval can reuse them.
Make the dev build the codegraph on PATH once, then loop:
npm run build && ./scripts/local-install.sh
scripts/add-lang/bench.sh <lang> <name> <url> "<question>" headless # ×3bench.sh clones (shared /tmp/codegraph-corpus), wipes + indexes, runs
verify-extraction.mjs, then the with/without retrieval A/B via
scripts/agent-eval/run-all.sh (skips the paid A/B if extraction is broken).
Read each parse-run.mjs summary printed by run-all.sh: tool calls, file
Reads, Grep/Bash, codegraph-tool calls, duration, and cost — for both the
with and without arms. After the loop, restore the dev link if needed:
./scripts/local-install.sh.
<Lang> to the "19+ Languages" feature bullet, and add a
row to the Supported Languages table:
| <Lang> | \.ext` | Full support (classes, methods, …) |`.## [Unreleased] section at the top (above the
latest version) with ### Added → a user-perspective bullet, e.g.
"CodeGraph now indexes <Lang> (.ext) — functions, classes, imports, and
call edges." If ## [Unreleased] already exists, append under it. (It's
folded into the next versioned block at release time.)Summarize for review:
.wasm).verify-extraction result.with vs without (tool calls, file Reads, cost) and a
one-line verdict — did codegraph reduce effort, and did both arms reach a
correct answer?Hand the changes to the user. Do not run git commit/push or publish —
releases go through the GitHub Actions Release workflow.
claude -p runs (opus, --max-budget-usd),
2 arms × 3 repos. The corpus dir /tmp/codegraph-corpus is shared with
/agent-eval, so clones are reused across runs.*.wasm must live in src/extraction/wasm/ — copy-assets (run by
npm run build) ships it; otherwise it won't be in dist/.© colbymchenry, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/add-lang of colbymchenry/codegraph.
Open the folder on GitHubat commit b635dd4
Add a Language to CodeGraph next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Add a Language to CodeGraph this skillcolbymchenry/codegraph | 74k | — | ~2.7k | Automated safety check: Pass | MIT | |
| Repomix Codebase Exploreryamadashy/repomix | 29k | — | ~2.7k | Automated safety check: Pass | MIT | |
| Birdview Architecture MapQiuner/birdview | 727 | — | ~2.9k | Automated safety check: Pass | MIT | |
| Acquire Codebase Knowledgegithub/awesome-copilot | 40k | 1 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Codebase Contexthomarr-labs/homarr | 5k | — | ~679 | Automated safety check: Pass | Apache-2.0 | |
| tilth Code Reading CLIjahala/tilth | 352 | — | ~1.1k | Automated safety check: Pass | MIT |
yamadashy/repomix
Packs a local or remote repository into a single AI-friendly file with the Repomix CLI, then reads and searches that output to explain structure, find patterns or report metrics.
Qiuner/birdview
Shows an evidence-linked map of a codebase's architecture and constraints, highlighting the modules an AI plans to change before it edits anything.
github/awesome-copilot
Maps an unfamiliar codebase into seven evidence-backed documents in docs/codebase/, using a scan script and templates, for onboarding or architecture write-ups.
homarr-labs/homarr
Navigate Homarr's monorepo architecture and reuse shared packages.
jahala/tilth
Replaces grep, cat, find and ls with the tilth CLI, which returns AST-aware outlines, definitions, usages and callers across many languages in one call.
Q00/ouroboros
Scans a directory for existing git repositories and worktrees, then registers and manages which ones serve as default context during interviews.
colbymchenry/codegraph
Benchmarks how much CodeGraph helps a coding agent on a real repository, comparing runs with and without it for a chosen local or published version.
Categories
Adds tree-sitter support for a new language to the codegraph project, then tests it and benchmarks extraction quality and retrieval value on real repositories. This is a maintainer workflow for the codegraph project. Given a lowercase language token such as lua, elixir or zig, the agent wires a tree-sitter grammar and an extractor into codegraph's extraction pipeline, writes tests, and benchmarks extraction quality and retrieval value on 3 popular real-world repositories.
Add a Language to CodeGraph fits situations like: adding support for a new programming language to codegraph; benchmarking the extraction quality of an already supported language; checking that a tree-sitter grammar is healthy before writing an extractor.
Run `npx skills add colbymchenry/codegraph --skill add-lang -a claude-code`. Or copy the skill folder (.claude/skills/add-lang in colbymchenry/codegraph) into .claude/skills/add-lang in your project. Claude Code loads it when a task matches its description.
Run `npx skills add colbymchenry/codegraph --skill add-lang -a codex`. Or copy the skill folder (.claude/skills/add-lang in colbymchenry/codegraph) into .agents/skills/add-lang in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add colbymchenry/codegraph --skill add-lang -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-lang, .gemini/skills/add-lang, .github/skills/add-lang and .opencode/skills/add-lang in your project.
Going by SKILL.md and its folder, Add a Language to CodeGraph needs the command-line tools its instructions call (npm, node, claude, make, npx and gh). Our summary lists: The codegraph repository root as the working directory; node, git and gh on the PATH; A logged-in claude CLI for the benchmark runs.
SKILL.md contains no URLs. Its commands use npm, npx, gh and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Add a Language to CodeGraph is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Add a Language to CodeGraph: Repomix Codebase Explorer (yamadashy/repomix, 29k stars), Birdview Architecture Map (Qiuner/birdview, 727 stars), Acquire Codebase Knowledge (github/awesome-copilot, 40k stars) and Codebase Context (homarr-labs/homarr, 5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
colbymchenry (a GitHub user) maintains it in colbymchenry/codegraph, which has 73,544 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 7, 2026.
Source: colbymchenry/codegraph on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.