Agent skill

Add a Language to CodeGraph

by colbymchenry in colbymchenry/codegraph

Adds tree-sitter support for a new language to the codegraph project, then tests it and benchmarks extraction quality and retrieval value on real repositories.

MITAuto-check passedDevelopment

Install Add a Language to CodeGraph

skills CLI
$ npx skills add colbymchenry/codegraph --skill add-lang -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install colbymchenry/codegraph add-lang --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/colbymchenry/codegraph.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/add-lang .claude/skills/add-lang && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-lang
GitHub stars
74k
Token cost
~2.7k tokens
SKILL.md length
1,123 words
Files
1
Skills in repo
2
Repo updated
First seen
Licence
MIT

At a glance

Adds tree-sitter support for a new language to the codegraph project, then tests it and benchmarks extraction quality and retrieval value on real repositories.

  • Works in 10 steps: Resolve + short-circuit → Find a grammar, then health-check it → Discover AST node types → …
  • Adding support for a new programming language to codegraph
  • SKILL.md covers Prerequisites, Workflow and Notes
  • Calls npm, node and claude

What it does

This is a maintainer workflow for the codegraph project. Given a lowercase language token such as lua, elixir or zig, the agent wires a tree-sitter grammar and an extractor into codegraph's extraction pipeline, writes tests, and benchmarks extraction quality and retrieval value on 3 popular real-world repositories. If the language is already supported, it skips the wiring steps and goes straight to benchmarking, noting in its report that no code changed.

The work follows a checklist in order. The agent first looks for a grammar in the tree-sitter-wasms package or vendors a .wasm file, then health-checks the grammar with a check-grammar script before writing any extractor, because a grammar that is present can still be unusable, for example an old ABI that corrupts the shared WASM heap. It then dumps the grammar's AST node types, builds and links the local development build for the benchmark, which spawns real claude runs, and updates the docs.

It runs fully autonomously from the codegraph repo root, using node, git, gh and a logged-in claude CLI, and it never commits, pushes, publishes or tags, leaving every change for you to review.

When your agent uses it

  • Adding support for a new programming language to codegraph
  • Benchmarking the extraction quality of an already supported language
  • Checking that a tree-sitter grammar is healthy before writing an extractor

Example prompts

  • “Run /add-lang zig and report how it does on real repositories.”
  • “Add OCaml support to codegraph and benchmark it on three popular repos.”
  • “Check whether Elixir is already wired into codegraph and just run the benchmark.”

Requirements

  • The codegraph repository root as the working directory
  • node, git and gh on the PATH
  • A logged-in claude CLI for the benchmark runs

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. Resolve + short-circuit
  2. Find a grammar, then health-check it
  3. Discover AST node types
  4. Wire the language (4 files)
  5. Build + verify loop
  6. Tests
  7. Auto-pick 3 repos + corpus
  8. Benchmark all 3 (extraction + A/B)
  9. Docs + CHANGELOG
  10. Report (do NOT commit)

What it can do on your machine

Read from SKILL.md and the folder at commit b635dd4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm
    • node
    • claude
    • make
    • npx
    • gh
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, npx, gh and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add a Language to CodeGraph loads about 2.7k tokens when it runs. Until then it costs about 81 tokens; SKILL.md has 1,123 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~81
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from colbymchenry/codegraph at commit b635dd4, republished under its MIT licence (© colbymchenry). 1,123 words, ~2,685 tokens.

Download SKILL.mdSave it as .claude/skills/add-lang/SKILL.md (or your agent's skills folder).
name
add-lang
description
Add tree-sitter language support to codegraph end-to-end — wire the grammar + extractor, write tests, then benchmark extraction quality and retrieval value on 3 popular real-world repos. Use when the user runs /add-lang <language> or asks to add/support a new language (e.g. Lua, Elixir, Zig, OCaml) in codegraph.

Add a language to CodeGraph

Wire a new tree-sitter language into codegraph's extraction pipeline, prove it extracts real symbols on popular repos, and prove it beats no-codegraph for an agent. Runs fully autonomously — pick repos, benchmark, update docs, then report. Never commit, push, publish, or tag (house rule); leave all changes for the user to review.

The argument is the language token used throughout the Language union, e.g. lua, elixir, zig. If none was given, ask which language. Use the lowercase single-token form everywhere (csharp, not c#).

Prerequisites

  • Run from the codegraph repo root. node, git, gh, and a logged-in claude CLI (the benchmark spawns real claude -p runs).
  • The benchmark uses the local dev build — Step 8 builds + links it on PATH.

Workflow

Copy this checklist and work through it in order:

- [ ] 1. Resolve language; bail early if already supported (just benchmark)
- [ ] 2. Find a grammar + health-check it (ABI / heap corruption)
- [ ] 3. Discover the grammar's AST node types (dump-ast.mjs)
- [ ] 4. Wire the language (4 files; sometimes a 5th core touch)
- [ ] 5. Build + verify-extraction loop until PASS
- [ ] 6. Add extraction tests; make them green
- [ ] 7. Auto-pick 3 popular repos by size tier; add to corpus.json
- [ ] 8. Benchmark all 3: extraction + with/without A/B
- [ ] 9. Update README + CHANGELOG
- [ ] 10. Report; do NOT commit
Step 1 — Resolve + short-circuit

Check whether the language is already wired: look for the token in the LANGUAGES const (src/types.ts) and the EXTRACTORS map (src/extraction/languages/index.ts). If it is already supported (e.g. typescript, rust), skip Steps 2–6 and go straight to benchmarking (Steps 7–8) to validate/measure it — note in the report that no code changed.

Step 2 — Find a grammar, then health-check it
bash
ls node_modules/tree-sitter-wasms/out/ | grep -i <lang>   # csharp -> c_sharp
  • Present → likely off-the-shelf; grammars.ts resolves it from tree-sitter-wasms automatically. (Many languages: elixir, zig, ocaml, solidity, toml, yaml, …)
  • Absent → vendor a .wasm into src/extraction/wasm/ (like pascal / scala / lua) and add the token to the vendored branch in Step 4.

Always health-check before writing an extractor — a present grammar can still be unusable:

bash
node scripts/add-lang/check-grammar.mjs <lang> path/to/valid-sample.<ext>

It prints the grammar's ABI version and parses a valid sample many times in a multi-grammar runtime. If it FAILs (ERROR trees on valid code — an old ABI corrupting the shared WASM heap, which silently drops nested calls/imports on every file after the first; e.g. the tree-sitter-wasms Lua grammar is ABI 13 and fails), do NOT use that wasm. Vendor a newer (ABI 14/15) build instead:

bash
npm pack @tree-sitter-grammars/tree-sitter-<lang>   # often ships a prebuilt *.wasm
# or build one: npx tree-sitter build --wasm   (needs Docker/emscripten)
cp <the>.wasm src/extraction/wasm/tree-sitter-<lang>.wasm

then add the token to the vendored branch in Step 4 and re-run check-grammar on the vendored path until it PASSes. If you cannot obtain a healthy wasm, STOP and tell the user.

Step 3 — Discover AST node types

Get a representative source file (write a small sample covering functions, classes/structs, imports, enums; or curl a raw file from a known repo), then:

bash
node scripts/add-lang/dump-ast.mjs <lang> path/to/sample.<ext>
# vendored grammar: pass the wasm path instead of the token
node scripts/add-lang/dump-ast.mjs src/extraction/wasm/tree-sitter-<lang>.wasm sample.<ext>

The frequency table + field names (name:, parameters:, body:, return_type:) tell you what to map. Open the existing extractor closest to the language's paradigm as a model: rust.ts/scala.ts (functional, traits), java.ts/csharp.ts (OO), python.ts/ruby.ts (scripting), go.ts (top-level methods + receivers).

Step 4 — Wire the language (4 files)

These are exact, fragile wiring — match the existing style precisely:

  1. src/types.ts — TWO edits:
    • add '<lang>', to the LANGUAGES const (before 'unknown');
    • add '**/*.<ext>', to DEFAULT_CONFIG.include. Don't skip this — it's the file-scan allowlist; without the glob, codegraph init finds 0 files even though detection/extraction are wired.
  2. src/extraction/grammars.ts — three maps:
    • WASM_GRAMMAR_FILES: <lang>: 'tree-sitter-<lang>.wasm',
    • EXTENSION_MAP: each file extension → '<lang>' (e.g. '.lua': 'lua',)
    • getLanguageDisplayName: <lang>: '<Display Name>',
    • vendored only: add <lang> to the (lang === 'pascal' || lang === 'scala' || …) wasm-path branch.
  3. src/extraction/languages/<lang>.ts — new file exporting export const <lang>Extractor: LanguageExtractor = { … }. Map the node types from Step 3. Required fields: functionTypes, classTypes, methodTypes, interfaceTypes, structTypes, enumTypes, typeAliasTypes, importTypes, callTypes, variableTypes, nameField, bodyField, paramsField. Add hooks as the grammar needs them (getSignature, getVisibility, isExported, extractImport, visitNode, getReceiverType, interfaceKind, enumMemberTypes, etc. — see src/extraction/tree-sitter-types.ts).
  4. src/extraction/languages/index.ts — import { <lang>Extractor } from './<lang>'; and add <lang>: <lang>Extractor, to EXTRACTORS.

Sometimes a 5th, core touch in src/extraction/tree-sitter.ts — variable extraction has per-language branches in extractVariable (the generic fallback only finds direct identifier/variable_declarator children). If the grammar nests declared names (e.g. Lua's variable_declaration → variable_list), add a } else if (this.language === '<lang>') branch there, mirroring the existing ts/python/go ones. Import forms that aren't a distinct node (Lua/Ruby require is a call) are handled in the extractor's visitNode hook instead.

Step 5 — Build + verify loop
bash
npm run build            # tsc + copy-assets (copies any vendored *.wasm into dist/)

Index a small sample repo and check extraction:

bash
( cd <sample-repo> && codegraph init -i )
node scripts/add-lang/verify-extraction.mjs <sample-repo> <lang>

verify-extraction.mjs fails (exit 1) if the language isn't detected or only file/import nodes were produced — the classic symptom of wrong node-type names. On FAIL or a thin WARN: re-run dump-ast.mjs on a richer file, fix the mappings in <lang>.ts, npm run build, re-index, re-verify. Repeat until PASS.

Show full SKILL.md (441 more words)Show less
Step 6 — Tests

Add to __tests__/extraction.test.ts, modeled on the Rust Extraction block:

  • a detectLanguage assertion in describe('Language Detection')
  • a describe('<Lang> Extraction') block asserting functions/classes/imports are extracted from an inline source string.
bash
npx vitest run __tests__/extraction.test.ts

Green before continuing.

Step 7 — Auto-pick 3 repos + corpus

Pick without asking. Find candidates, then curate 3 that are genuinely <lang>-dominant, one per size tier:

bash
gh search repos --language=<lang> --sort=stars --limit 40 \
  --json fullName,stargazerCount,description

Tiers (match corpus.json): Small <~150 files · **Medium** ~150–1500 · **Large** >~1500. Skip repos that are tagged <lang> but mostly another language. Write one cross-file architecture question per repo (the kind that needs tracing across files). Add a "<Language>" block to .claude/skills/agent-eval/corpus.json (fields: name, repo, size, files, question) so /agent-eval can reuse them.

Step 8 — Benchmark all 3 (extraction + A/B)

Make the dev build the codegraph on PATH once, then loop:

bash
npm run build && ./scripts/local-install.sh
scripts/add-lang/bench.sh <lang> <name> <url> "<question>" headless   # ×3

bench.sh clones (shared /tmp/codegraph-corpus), wipes + indexes, runs verify-extraction.mjs, then the with/without retrieval A/B via scripts/agent-eval/run-all.sh (skips the paid A/B if extraction is broken). Read each parse-run.mjs summary printed by run-all.sh: tool calls, file Reads, Grep/Bash, codegraph-tool calls, duration, and cost — for both the with and without arms. After the loop, restore the dev link if needed: ./scripts/local-install.sh.

Step 9 — Docs + CHANGELOG
  • README.md: add <Lang> to the "19+ Languages" feature bullet, and add a row to the Supported Languages table: | <Lang> | \.ext` | Full support (classes, methods, …) |`.
  • CHANGELOG.md: add an ## [Unreleased] section at the top (above the latest version) with ### Added → a user-perspective bullet, e.g. "CodeGraph now indexes <Lang> (.ext) — functions, classes, imports, and call edges." If ## [Unreleased] already exists, append under it. (It's folded into the next versioned block at release time.)
Step 10 — Report (do NOT commit)

Summarize for review:

  • Files changed: the 4 wiring edits + new extractor + tests + README + CHANGELOG + corpus.json (+ any vendored .wasm).
  • Extraction per repo: files / nodes / edges / verify-extraction result.
  • A/B per repo: with vs without (tool calls, file Reads, cost) and a one-line verdict — did codegraph reduce effort, and did both arms reach a correct answer?
  • Gaps / follow-ups (node types not yet mapped, resolution edges missing, framework routes, etc.).

Hand the changes to the user. Do not run git commit/push or publish — releases go through the GitHub Actions Release workflow.

Notes

  • The A/B spawns real paid claude -p runs (opus, --max-budget-usd), 2 arms × 3 repos. The corpus dir /tmp/codegraph-corpus is shared with /agent-eval, so clones are reused across runs.
  • Any new *.wasm must live in src/extraction/wasm/ — copy-assets (run by npm run build) ships it; otherwise it won't be in dist/.
  • An index must be served by the same binary that built it. Step 8 builds + links the dev build first, so this holds.
  • If a grammar can't be obtained, or extraction can't reach PASS, STOP and report — don't ship a half-wired language.

© colbymchenry, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/add-lang of colbymchenry/codegraph.

Open the folder on GitHubat commit b635dd4

Compare with similar skills

Add a Language to CodeGraph next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add a Language to CodeGraph compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add a Language to CodeGraph this skillcolbymchenry/codegraph74k—~2.7kAutomated safety check: PassMIT
Repomix Codebase Exploreryamadashy/repomix29k—~2.7kAutomated safety check: PassMIT
Birdview Architecture MapQiuner/birdview727—~2.9kAutomated safety check: PassMIT
Acquire Codebase Knowledgegithub/awesome-copilot40k1 repos~2.3kAutomated safety check: PassMIT
Codebase Contexthomarr-labs/homarr5k—~679Automated safety check: PassApache-2.0
tilth Code Reading CLIjahala/tilth352—~1.1kAutomated safety check: PassMIT

Similar skills

  • Repomix Codebase Explorer

    yamadashy/repomix

    Packs a local or remote repository into a single AI-friendly file with the Repomix CLI, then reads and searches that output to explain structure, find patterns or report metrics.

    29k GitHub stars~2.7k tokensUpdated 6 days ago
    DevelopmentAuto-check passed
  • Shows an evidence-linked map of a codebase's architecture and constraints, highlighting the modules an AI plans to change before it edits anything.

    727 GitHub stars~2.9k tokensUpdated today
    DevelopmentAuto-check passed
  • Acquire Codebase Knowledge

    github/awesome-copilot

    Official

    Maps an unfamiliar codebase into seven evidence-backed documents in docs/codebase/, using a scan script and templates, for onboarding or architecture write-ups.

    40k GitHub starsUsed in 1 repo~2.3k tokens
    DevelopmentAuto-check passed
  • Codebase Context

    homarr-labs/homarr

    Navigate Homarr's monorepo architecture and reuse shared packages.

    5k GitHub stars~679 tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Replaces grep, cat, find and ls with the tilth CLI, which returns AST-aware outlines, definitions, usages and callers across many languages in one call.

    352 GitHub stars~1.1k tokensUpdated 12 days ago
    DevelopmentAuto-check passed
  • Scans a directory for existing git repositories and worktrees, then registers and manages which ones serve as default context during interviews.

    6.2k GitHub stars~2.2k tokensUpdated 2 days ago
    DevelopmentAuto-check passed

More from colbymchenry/codegraph

  • CodeGraph Agent Eval

    colbymchenry/codegraph

    Benchmarks how much CodeGraph helps a coding agent on a real repository, comparing runs with and without it for a chosen local or published version.

    74k GitHub stars~950 tokensUpdated 2 days ago
    Auto-check passed

Categories

Questions about Add a Language to CodeGraph

What does Add a Language to CodeGraph do?

Adds tree-sitter support for a new language to the codegraph project, then tests it and benchmarks extraction quality and retrieval value on real repositories. This is a maintainer workflow for the codegraph project. Given a lowercase language token such as lua, elixir or zig, the agent wires a tree-sitter grammar and an extractor into codegraph's extraction pipeline, writes tests, and benchmarks extraction quality and retrieval value on 3 popular real-world repositories.

When should I use Add a Language to CodeGraph?

Add a Language to CodeGraph fits situations like: adding support for a new programming language to codegraph; benchmarking the extraction quality of an already supported language; checking that a tree-sitter grammar is healthy before writing an extractor.

How do I install Add a Language to CodeGraph in Claude Code?

Run `npx skills add colbymchenry/codegraph --skill add-lang -a claude-code`. Or copy the skill folder (.claude/skills/add-lang in colbymchenry/codegraph) into .claude/skills/add-lang in your project. Claude Code loads it when a task matches its description.

How do I install Add a Language to CodeGraph in Codex?

Run `npx skills add colbymchenry/codegraph --skill add-lang -a codex`. Or copy the skill folder (.claude/skills/add-lang in colbymchenry/codegraph) into .agents/skills/add-lang in your project. Codex loads it when a task matches its description.

Can I use Add a Language to CodeGraph in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add colbymchenry/codegraph --skill add-lang -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-lang, .gemini/skills/add-lang, .github/skills/add-lang and .opencode/skills/add-lang in your project.

What does Add a Language to CodeGraph need to run?

Going by SKILL.md and its folder, Add a Language to CodeGraph needs the command-line tools its instructions call (npm, node, claude, make, npx and gh). Our summary lists: The codegraph repository root as the working directory; node, git and gh on the PATH; A logged-in claude CLI for the benchmark runs.

Does Add a Language to CodeGraph access the network?

SKILL.md contains no URLs. Its commands use npm, npx, gh and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Add a Language to CodeGraph safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Add a Language to CodeGraph use?

Add a Language to CodeGraph is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add a Language to CodeGraph use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Add a Language to CodeGraph?

Skills that share tags, products or a category with Add a Language to CodeGraph: Repomix Codebase Explorer (yamadashy/repomix, 29k stars), Birdview Architecture Map (Qiuner/birdview, 727 stars), Acquire Codebase Knowledge (github/awesome-copilot, 40k stars) and Codebase Context (homarr-labs/homarr, 5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add a Language to CodeGraph?

colbymchenry (a GitHub user) maintains it in colbymchenry/codegraph, which has 73,544 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 7, 2026.

Source: colbymchenry/codegraph on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.