Codegraph QA
QwenLM/qwen-code-examples
Use CodeScope to analyze any indexed codebase via its graph database (neug) and vector index (zvec).
Answers questions about code structure, history, bugs and PR risk using a CodeScope knowledge graph and semantic index built from the repository.
$ npx skills add QwenLM/qwen-code --skill codegraph -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install QwenLM/qwen-code codegraph --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.qwen/skills/codegraph .claude/skills/codegraph && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "codegraph" agent skill from https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/codegraph into .claude/skills/codegraph/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "codegraph", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/codegraphType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add QwenLM/qwen-code --skill codegraph -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install QwenLM/qwen-code codegraph --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.qwen/skills/codegraph .agents/skills/codegraph && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "codegraph" agent skill from https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/codegraph into .agents/skills/codegraph/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "codegraph", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add QwenLM/qwen-code --skill codegraph -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install QwenLM/qwen-code codegraph --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.qwen/skills/codegraph .cursor/skills/codegraph && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "codegraph" agent skill from https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/codegraph into .cursor/skills/codegraph/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "codegraph", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/QwenLM/qwen-code.git --path .qwen/skills/codegraph--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add QwenLM/qwen-code --skill codegraph -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install QwenLM/qwen-code codegraph --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.qwen/skills/codegraph .gemini/skills/codegraph && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "codegraph" agent skill from https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/codegraph into .gemini/skills/codegraph/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "codegraph", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install QwenLM/qwen-code codegraphInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add QwenLM/qwen-code --skill codegraph -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .github/skills && cp -r skills-src/.qwen/skills/codegraph .github/skills/codegraph && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "codegraph" agent skill from https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/codegraph into .github/skills/codegraph/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "codegraph", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add QwenLM/qwen-code --skill codegraph -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install QwenLM/qwen-code codegraph --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.qwen/skills/codegraph .opencode/skills/codegraph && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "codegraph" agent skill from https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/codegraph into .opencode/skills/codegraph/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "codegraph", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
codegraphAnswers questions about code structure, history, bugs and PR risk using a CodeScope knowledge graph and semantic index built from the repository.
CodeScope indexes source code into a two-layer knowledge graph, structure (functions, calls, imports, classes, modules) and evolution (commits, file changes, function modifications), plus semantic embeddings for every function. It supports Python, JavaScript and TypeScript, C and Java. The skill uses it for call graphs, dependencies, dead code, hotspots, module coupling, architecture reports, semantic search, impact analysis and UML class diagrams.
It can also fetch GitHub issues and trace bugs to likely code locations, and review open PRs by scoring risk, detecting cross-PR conflicts, finding auto-merge candidates and applying labels. Setup is pip install codegraph-ai, then codegraph status to check the index and codegraph init to create one. Companion files cover bug analysis, patterns, PR analysis and the graph schema. The skill applies when a .codegraph index exists in the workspace or when you want to create one.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 4970bfa. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
ghpippythongitFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
hf-mirror.commodelscope.cnFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
GITHUB_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
CodeScope Codebase Graph Analysis loads about 9.3k tokens when it runs. Until then it costs about 182 tokens; SKILL.md has 2,363 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from QwenLM/qwen-code at commit 4970bfa, republished under its Apache-2.0 licence (© QwenLM). 2,363 words, ~9,266 tokens.
.claude/skills/codegraph/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.CodeScope indexes source code into a two-layer knowledge graph — structure (functions, calls, imports, classes, modules) and evolution (commits, file changes, function modifications) — plus semantic embeddings for every function. Supports Python, JavaScript/TypeScript, C, and Java (including Hadoop-scale repositories with 8K+ files). This combination enables analyses that grep, LSP, or pure vector search cannot do alone. It can also fetch GitHub issues and trace bugs to code, and review open PRs — scoring per-PR risk, detecting cross-PR conflicts, identifying auto-merge candidates, and applying GitHub labels.
.codegraph directory (or similar index) exists in the workspacepip install codegraph-ai# Create Python virtural environment
python -m venv .venv
source .venv/bin/activate
# Point to a pre-built database (skip indexing)
export CODESCOPE_DB_DIR="/path/to/.linux_db"
# Offline mode for HuggingFace models
export HF_HUB_OFFLINE="1"
# Fallback when HuggingFace is unreachable (e.g., network issues in China)
# Use HF mirror or ModelScope for sentence-transformers models:
export HF_ENDPOINT="https://hf-mirror.com"
# https://www.modelscope.cn/models/sentence-transformers/all-MiniLM-L6-v2codegraph status --db $CODESCOPE_DB_DIRIf no index exists, create one:
codegraph init --repo . --lang auto --commits 500Supported languages: python, c, javascript, typescript, java, or auto (auto-detects from file extensions).
The --commits flag ingests git history (for evolution queries). Without it, only structural analysis is available. Add --backfill-limit 200 to also compute function-level MODIFIES edges (slower but enables change_attribution and co_change).
To add git history to an existing index (without re-indexing structure):
codegraph ingest --repo . --db $CODESCOPE_DB_DIR --commits 500
codegraph ingest --repo . --db $CODESCOPE_DB_DIR --backfill-limit 200 # add MODIFIES edges onlyUse the CLI for status and reports:
codegraph status --db $CODESCOPE_DB_DIR
codegraph analyze --db $CODESCOPE_DB_DIR --output report.mdUse the Python API for queries and custom analyses:
import os
os.environ['HF_HUB_OFFLINE'] = '1' # required
from codegraph.core import CodeScope
cs = CodeScope(os.environ['CODESCOPE_DB_DIR'])
# Cypher query
rows = list(cs.conn.execute('''
MATCH (caller:Function)-[:CALLS]->(f:Function {name: "free_irq"})
RETURN caller.name, caller.file_path LIMIT 10
'''))
for r in rows:
print(r)
cs.close() # always close when doneThe Python API is more powerful — it gives you raw Cypher access and lets you chain queries.
These are the building blocks for any custom analysis:
| Method | What it does |
|---|---|
cs.conn.execute(cypher) | Run any Cypher query against the graph — returns list of tuples |
cs.vector_only_search(query, topk=10) | Semantic search over all function embeddings — returns [{id, score}] |
cs.summary() | Print a human-readable overview of the indexed codebase |
| Method | What it does |
|---|---|
cs.impact(func_name, change_desc, max_hops=3) | Find callers up to N hops, ranked by semantic relevance to the change |
cs.hotspots(topk=10) | Rank functions by structural risk (fan-in × fan-out) |
cs.dead_code() | Find functions with zero callers (excluding entry points) |
cs.circular_deps() | Detect circular import chains at file level |
cs.module_coupling(topk=10) | Find cross-module coupling pairs with call counts |
cs.bridge_functions(topk=30) | Find functions called from the most distinct modules |
cs.layer_discovery(topk=30) | Auto-discover infrastructure / mid / consumer layers |
cs.stability_analysis(topk=50) | Correlate fan-in with modification frequency |
cs.class_hierarchy(class_name=None) | Return inheritance tree for a class (or all classes) |
CodeScope extracts three UML relationship types from class fields and type annotations during indexing:
| Relationship | UML symbol | Meaning | How detected |
|---|---|---|---|
COMPOSES | *-- filled diamond | Strong ownership — field always holds an instance | Non-optional field assigned a constructed object |
AGGREGATES | o-- open diamond | Optional/weak reference — may be None | Optional[X], X | None, or assigned None |
INHERITS | <|-- hollow arrow | Subclass extends parent | class A(B) |
# Get all composition relationships (A strongly owns B)
list(cs.conn.execute('MATCH (c1:Class)-[:COMPOSES]->(c2:Class) RETURN c1.name, c2.name'))
# Get all aggregation relationships (A optionally holds B)
list(cs.conn.execute('MATCH (c1:Class)-[:AGGREGATES]->(c2:Class) RETURN c1.name, c2.name'))
# How many objects does a class directly own?
list(cs.conn.execute(
'MATCH (c:Class {name: "Llama"})-[:COMPOSES]->(t:Class) RETURN t.name'
))
# Full dependency graph for a class (composition + aggregation + inheritance)
list(cs.conn.execute(
'MATCH (c:Class {name: "GPUModelRunner"})-[r:COMPOSES|AGGREGATES]->(t:Class) '
'RETURN type(r), t.name'
))Generating a Mermaid class diagram:
inherits = list(cs.conn.execute('MATCH (c1:Class)-[:INHERITS]->(c2:Class) RETURN c1.name, c2.name'))
composes = list(cs.conn.execute('MATCH (c1:Class)-[:COMPOSES]->(c2:Class) RETURN c1.name, c2.name'))
aggregates = list(cs.conn.execute('MATCH (c1:Class)-[:AGGREGATES]->(c2:Class) RETURN c1.name, c2.name'))
print('classDiagram')
for src, tgt in inherits: print(f' {tgt} <|-- {src}') # parent <|-- child
for src, tgt in composes: print(f' {src} *-- {tgt}') # owner *-- owned
for src, tgt in aggregates: print(f' {src} o-- {tgt}') # holder o-- optionalScale reference:
| Project | Classes | INHERITS | COMPOSES | AGGREGATES | Index time |
|---|---|---|---|---|---|
| llama-cpp-python | 128 | 18 | 8 | 4 | ~2s |
| vllm | 4,002 | 2,185 | 3,217 | 149 | ~50s |
| Method | What it does |
|---|---|
cs.similar(function, scope, topk=10) | Find functions similar to a given function within a module scope |
cs.cross_locate(query, topk=10) | Find semantically related functions, then reveal call-chain connections |
cs.semantic_cross_pollination(query, topk=15) | Find similar functions across distant subsystems |
--commits during init)| Method | What it does |
|---|---|
cs.change_attribution(func_name, file_path=None, limit=20) | Which commits modified a function? (requires backfill) |
cs.co_change(func_name, file_path=None, min_commits=2, topk=10) | Functions that are always modified together |
cs.intent_search(query, topk=10) | Find commits matching a natural-language intent |
cs.commit_modularity(topk=20) | Score commits by how many modules they touch |
cs.hot_cold_map(topk=30) | Module modification density |
from codegraph.analyzer import generate_report
report = generate_report(cs) # full architecture analysis as markdownOr via CLI:
codegraph analyze --output reports/analysis.mdThe report covers: overview stats, subsystem distribution, top modules, architectural layers (with Mermaid diagrams), bridge functions, fan-in/fan-out hotspots, cross-module coupling, evolution hotspots, and dead code density.
CodeScope includes a full Java adapter that handles enterprise-scale repositories like Apache Hadoop (~8K files, ~97K functions indexed in ~3.5 minutes).
| Element | Graph Node/Edge | Notes |
|---|---|---|
| Classes | Class node | Includes generics, annotations |
| Interfaces | Class node | extends → INHERITS edge |
| Enums | Class node | Enum methods extracted |
| Methods | Function node | Full generic signatures, JavaDoc |
| Constructors | Function node (name=<init>) | Including super() calls |
| Method calls | CALLS edge | Receiver context preserved (obj.method()) |
new expressions | CALLS edge to ClassName.<init> | Constructor invocations |
| Imports | IMPORTS edge (file→file) | Single, wildcard, static |
| Inner classes | Class node (name=Outer.Inner) | Prefixed with outer class |
| Inheritance | INHERITS edge | extends + implements |
codegraph init --repo /path/to/java-project --lang java --commits 500Or with auto-detection (auto-detects .java files):
codegraph init --repo /path/to/java-project --lang autoBy default, these directories are excluded when indexing Java projects: target/, build/, .gradle/, .idea/, .settings/, bin/, out/, test/, tests/, src/test/.
# Find all classes that extend a specific class
list(cs.conn.execute("""
MATCH (c:Class)-[:INHERITS]->(p:Class {name: 'FileSystem'})
RETURN c.name, c.file_path
"""))
# Find all methods in a specific class
list(cs.conn.execute("""
MATCH (c:Class {name: 'DefaultParser'})-[:HAS_METHOD]->(f:Function)
RETURN f.name, f.signature
"""))
# Find constructor call chains
list(cs.conn.execute("""
MATCH (f:Function)-[:CALLS]->(init:Function {name: '<init>'})
WHERE init.class_name = 'Configuration'
RETURN f.name, f.file_path LIMIT 10
"""))CodeScope can fetch GitHub issues and map them to code using the graph + vector infrastructure. This is the core workflow for answering questions like "why does this project have so many bugs?" or "where in the code does this bug come from?"
gh CLI must be installed and authenticated (gh auth login)# Analyze a specific GitHub issue against the indexed code graph
result = cs.analyze_issue("owner", "repo", 1234, topk=10)
print(result.format_report())This:
cross_locate) to find related codeimpact()# Analyze top-k bug issues and get aggregated hotspot data
results = cs.analyze_top_bugs("owner", "repo", k=10, label="bug")
for r in results:
print(f"#{r.issue.number}: {r.issue.title}")
for c in r.candidates[:3]:
print(f" {c.function_name} ({c.file_path}) score={c.score:.2f}")# Fetch and parse a single issue (no graph needed)
codegraph fetch-issue owner repo 1234
# Fetch top-k bugs from a repo
codegraph fetch-bugs owner repo --top 10 --label bug
# Analyze a single bug against the code graph
codegraph analyze-bug owner repo 1234 --db .codegraph --topk 10
# Batch analyze top bugs
codegraph analyze-bugs owner repo --db .codegraph --top 10 --label bugFor custom analysis pipelines, the components can be used individually:
from codegraph.issue_fetcher import fetch_and_parse_issue
from codegraph.bug_locator import (
resolve_paths_to_files,
find_semantic_matches,
trace_callers,
rank_root_causes,
analyze_bug,
)
# Fetch and parse (with caching)
issue = fetch_and_parse_issue("owner", "repo", 1234)
print(issue.extracted_paths) # file paths found in body
print(issue.extracted_funcs) # function names from stack traces
print(issue.linked_commits) # merge commit SHAs from linked PRs
# Match paths to graph nodes
path_matches = resolve_paths_to_files(cs, issue.extracted_paths)
# Semantic search using issue description
semantic_matches = find_semantic_matches(cs, f"{issue.title}\n{issue.body}")
# Trace callers of mentioned functions
caller_traces = trace_callers(cs, issue.extracted_funcs, max_hops=2)
# Combine into ranked candidates
candidates = rank_root_causes(path_matches, semantic_matches, caller_traces, issue.extracted_funcs)Root cause candidates are scored by combining multiple signals:
| Signal | Score | Description |
|---|---|---|
| Direct mention | +1.0 | Function name appears in issue body/stack trace |
| File path match | +0.8 | Function is in a file mentioned in the issue |
| Semantic match | +score | Raw cosine similarity (0.0-1.0) from cross_locate |
| Caller relationship | +0.5/hops | Function calls a mentioned function (decays with distance) |
Parsed issues are cached at ~/.codegraph/issue_cache/{owner}_{repo}_{number}.json. Cache hits skip the GitHub API call entirely (sub-millisecond). To force a refresh, pass use_cache=False or use --no-cache on CLI.
from codegraph.issue_cache import clear_cache
clear_cache(owner="openclaw", repo="openclaw") # clear specific repo
clear_cache() # clear allThe parser automatically extracts file paths and function names from stack traces in Python, C/C++, JavaScript/Node.js, Go, and Rust formats. It also extracts func_name() references in backticks and inline code.
CodeScope can analyze open PRs against the indexed code graph to compute structural risk scores, detect cross-PR conflicts, and generate prioritized review reports.
gh CLI must be installed and authenticated (gh auth login)GITHUB_TOKEN environment variable recommended to avoid rate limitingTwo subcommands: prepare (analyze + write to DB) and label (apply GitHub labels + comments).
# Phase 1: Analyze PRs, detect conflicts, write to graph DB (full rebuild)
# Pipeline: cross-PR analysis → single-PR risk scoring → report + labels
codegraph pr-review prepare --db .codegraph
# Filter by author during prepare:
codegraph pr-review prepare --db .codegraph --author someone
# Override auto-detected GitHub repo (owner/repo):
codegraph pr-review prepare --db .codegraph --repo owner/repo
# Skip per-PR risk scoring (conflict-only, faster):
codegraph pr-review prepare --db .codegraph --skip-single-pr
# Phase 2: Apply labels and post conflict comments from graph DB
codegraph pr-review label --db .codegraph
# Label with dry-run (preview without API calls):
codegraph pr-review label --db .codegraph --dry-runRequired arg: --db. Local repo path derived from --db parent. GitHub repo auto-detected from git remote get-url origin (or specified via --repo). Optional: --author, --output, --skip-single-pr (prepare); --dry-run (label).
For programmatic use within the same Python process, use PRReview — a
high-level wrapper that manages CodeScope lifecycle automatically.
from codegraph.pr_api import PRReview
# Full pipeline in 2 lines
with PRReview(db=".codegraph") as pr:
pr.prepare() # fetch PRs → graph DB → scoring → report
pr.label(dry_run=True) # preview labels without API calls
# Query after prepare (works across sessions once DB has data)
with PRReview(db=".codegraph") as pr:
# Conflicts
pr.conflict_prs_of("100") # → ["101", "102"]
# Risk
pr.risk("100") # → {"number": "100", "risk_level": "HIGH", ...}
# Classification
pr.auto_merge_candidates() # → [{"number": "200", ...}, ...]
pr.conflicting_groups() # → [["100", "101"], ["103"]]
# All PRs in DB
pr.all_prs() # → [{"number": "100", ...}, ...]
# Functions changed by a specific PR (added / modified / deleted)
import json
cs = pr._open_cs()
rows = list(cs.conn.execute(
f"MATCH (pr:PR {{id: {json.dumps('439')}}})-[c:CHANGES]->(f:Function) "
f"RETURN c.info AS change_type, f.name, f.file_path "
f"ORDER BY c.info, f.name"
))
for change_type, name, path in rows:
print(f" [{change_type}] {name} ({path})")
# change_type: 'hunk' (modified), 'new' (added), 'deleted', 'related' (newly calls)All query methods return structured Python objects — no text parsing
required. The CLI and Python API share the same underlying implementation
(run_prepare / run_label / graph DB), so you can prepare via CLI
and query via Python, or vice versa.
For lower-level components (PRScorer, CrossPRAnalyzer, etc.), see:
from codegraph.pr_analysis import GitHubClient, GraphAnalyzer, PRScorer, CrossPRAnalyzer
gh = GitHubClient(repo='owner/repo')
scorer = PRScorer(GraphAnalyzer(cs, repo_dir), repo_dir, gh)
result = scorer.analyze(gh.pr_to_entry(pr), output_dir='/tmp') # risk_score, risk_level, peak_blast...
cross = CrossPRAnalyzer(cs, repo_dir, gh)
cross.prepare(pr_ids) # index PR nodes into graph
cross.connected_components() # {root: [pr_ids]} — detects conflicts
cross.update_pr_labels(assignments) # persist labels to graph DB
# Load PR results from graph DB (no GitHub API needed)
all_results, components = cross.load_from_graph()
# Build and apply labels from analysis results
from codegraph.pr_labeler import build_label_assignments, apply_labels
assignments = build_label_assignments(all_results, components)
apply_labels(assignments, repo='owner/repo', create_labels=True)For detailed workflows, Cypher patterns, and CrossPRAnalyzer query dimensions, see pr-analysis.md.
Risk levels: CRITICAL (≥12), HIGH (≥7), MEDIUM (≥3), LOW (<3), UNKNOWN (when --skip-single-pr). Key signals: blast_radius (3.0×), no_test_coverage (2.0×), interface_change (2.5×), dead_code (1.5×).
After running codegraph pr-review prepare, run codegraph pr-review label to apply category labels to GitHub PRs and post conflict comments:
# Apply labels and post conflict comments:
codegraph pr-review label --db .codegraph
# Preview without making API calls:
codegraph pr-review label --db .codegraph --dry-runThe label subcommand reads PR labels from the graph DB (pr.label column) — no re-analysis needed. For conflicting PRs (labelled conflicting-group-N), it also posts a comment on the GitHub PR listing shared functions and other conflicting PRs.
Labels are computed during prepare from the analysis results (connected components + risk scores) and persisted to PR nodes in the graph DB (pr.label column, semicolon-delimited).
Label scheme:
| Category | Label | Color |
|---|---|---|
| Auto-merge Candidates (Part 1) | auto-merge-candidate | Green |
| Independent Review (Part 2) | independent-review | Yellow |
| Conflicting Group N (Part 3) | conflicting-group-N | Red/Orange/Blue |
| Any conflicting PR (Part 3) | conflicting-pr | Red |
PR-specific follow-up questions are automatically included in codegraph explore when PR nodes exist in the graph DB (i.e., after codegraph pr-review prepare). PR exploration is a question template set integrated into explore. To query a specific PR's details (conflicts, changed functions), use the PRReview Python API.
# After pr-review prepare, explore includes PR questions automatically:
codegraph explore --db .codegraph --top 15
# Interactive exploration (including PR follow-up questions):
codegraph explore --db .codegraph
# Focus on PR-specific questions (use reviewer role):
codegraph explore --db .codegraph --role reviewer
# Filter to only architecture questions (exclude PR patterns):
codegraph explore --db .codegraph --type architecture
# Filter to only risk questions:
codegraph explore --db .codegraph --type risk
# Filter to only PR review questions:
codegraph explore --db .codegraph --type pr-review --role reviewerThe --type filter controls which question categories appear:
all (default): all categories mixed togetherarchitecture: structural design questions (fan-in, coupling, cycles)risk: risk-focused questions (structural risk + PR risk)evolution: git history questions (change attribution, modification patterns)hotspot: frequently modified code questionspr-review: PR-specific questions (impact, conflicts, test coverage)When --type pr-review is specified, only PR-related questions are shown.
The key decision is: does the user want an exact structural answer, a fuzzy semantic one, or a bug-to-code mapping?
| User asks... | Best approach |
|---|---|
"Who calls free_irq?" | Cypher: MATCH (c:Function)-[:CALLS]->(f:Function {name: 'free_irq'}) RETURN c.name, c.file_path |
| "Find functions related to memory allocation" | cs.vector_only_search("memory allocation") or cs.cross_locate("memory allocation") |
| "What's the most complex function?" | cs.hotspots(topk=1) |
| "Is there dead code in the networking stack?" | cs.dead_code() then filter by file path |
"How has schedule() changed recently?" | cs.change_attribution("schedule", "kernel/sched/core.c") |
| "Which modules are tightly coupled?" | cs.module_coupling(topk=20) |
| "Generate a full architecture report" | codegraph analyze or generate_report(cs) |
"What's the architectural role of mm/?" | cs.layer_discovery() then find mm entries |
| "Which functions act as API boundaries?" | cs.bridge_functions(topk=30) |
| "Find commits about fixing race conditions" | cs.intent_search("fix race condition") |
"What functions are always changed together with kmalloc?" | cs.co_change("kmalloc") |
| "Why does this project have so many bugs?" | cs.analyze_top_bugs("owner", "repo", k=10) then aggregate hotspots |
| "Analyze issue #1234 from GitHub" | cs.analyze_issue("owner", "repo", 1234) |
| "What code is related to this bug?" | cs.analyze_issue(...) or manual cross_locate(bug_description) |
| "Find the root cause of the crash in issue #42" | cs.analyze_issue("owner", "repo", 42) |
| "Which modules have the most bugs?" | cs.analyze_top_bugs(...) then aggregate by file/module |
| "Index this Java project" | codegraph init --repo . --lang java |
| "What classes extend FileSystem in Hadoop?" | Cypher: MATCH (c:Class)-[:INHERITS]->(p:Class {name: 'FileSystem'}) RETURN c.name, c.file_path |
| "Find all constructors called in this module" | Cypher: MATCH (f:Function)-[:CALLS]->(init:Function {name: '<init>'}) WHERE f.file_path CONTAINS 'module' RETURN ... |
| "Draw a class diagram / show class UML" | Query COMPOSES, AGGREGATES, INHERITS edges and render as Mermaid classDiagram |
"What does Llama own / compose?" | Cypher: MATCH (c:Class {name:'Llama'})-[:COMPOSES]->(t:Class) RETURN t.name |
"Which class holds a reference to KVCacheManager?" | Cypher: MATCH (c:Class)-[:COMPOSES|AGGREGATES]->(t:Class {name:'KVCacheManager'}) RETURN c.name |
"Show all optional dependencies of GPUModelRunner" | Cypher: MATCH (c:Class {name:'GPUModelRunner'})-[:AGGREGATES]->(t:Class) RETURN t.name |
| "Review all open PRs and generate report" | codegraph pr-review prepare --db ... |
| "Which PRs can be auto-merged?" | Run pr-review prepare, check Part 1 of report |
| "Are there conflicting PRs?" | Run pr-review prepare, check Part 3 (connected components) |
| "What's the risk of PR #42?" | PRScorer.analyze(entry) for per-PR scoring |
| "What's the blast radius of this PR?" | PRScorer.analyze(entry) → result['peak_blast'] and call graph viz |
| "Which PRs modify the same function?" | CrossPRAnalyzer.connected_components() → same-function edge type |
| "Label PRs with their review category" | codegraph pr-review label --db ... |
| "Post conflict comments on PRs" | codegraph pr-review label --db ... (automatic for conflicting PRs) |
| "Preview labels/comments without applying" | codegraph pr-review label --db ... --dry-run |
| "Explore PR follow-up questions interactively" | codegraph explore --db .codegraph (auto-includes PR patterns if prepare was run) |
| "Query a specific PR's conflicts" | PRReview.conflict_prs_of("42") — returns list of conflicting PR numbers |
| "Query a specific PR's changed functions" | Cypher: MATCH (pr:PR {id: '42'})-[c:CHANGES]->(f:Function) RETURN c.info, f.name, f.file_path |
| "Compare two PRs for overlap" | Cypher: MATCH (pr1:PR {id: '42'})-[c1:CHANGES]->(f:Function)<-[c2:CHANGES]-(pr2:PR {id: '43'}) RETURN f.name, f.file_path |
| "Show only architecture questions" | codegraph explore --db .codegraph --type architecture |
| "Show only PR review questions" | codegraph explore --db .codegraph --type pr-review --role reviewer |
| "Show top PR risk questions" | codegraph explore --db .codegraph --top 15 --role reviewer |
| "Full PR review pipeline: analyze, label, explore" | 1) codegraph pr-review prepare 2) codegraph pr-review label 3) codegraph explore --db .codegraph |
For novel investigations not covered by pre-built methods, compose raw Cypher queries. See patterns.md for templates. For bug analysis patterns, see bug-analysis.md.
When writing Cypher queries, these filters prevent misleading results:
f.is_historical = 0 — exclude deleted/renamed functions that are still in the graph as historical recordsf.is_external = 0 (on File nodes) — exclude system headers/library filesc.version_tag = 'bf' — only backfilled commits have MODIFIES edges; non-backfilled commits only have TOUCHES (file-level) edgesLIMIT — large codebases can return hundreds of thousands of rowsBefore running evolution queries, check what's available:
# How many commits are indexed?
list(cs.conn.execute("MATCH (c:Commit) RETURN count(c)"))
# How many have MODIFIES edges (backfilled)?
list(cs.conn.execute("MATCH (c:Commit) WHERE c.version_tag = 'bf' RETURN count(c)"))If no commits exist, evolution methods will return empty results — guide the user to run codegraph ingest first. If commits exist but aren't backfilled, TOUCHES (file-level) queries still work but MODIFIES (function-level) queries won't.
| Error | Cause | Fix |
|---|---|---|
Database locked | Crashed process left neug lock | rm <db>/graph.db/neugdb.lock |
Can't open lock file | zvec LOCK file deleted | touch <db>/vectors/LOCK |
Can't lock read-write collection | Another process holds lock | Kill the other process |
recovery idmap failed | Stale WAL files | Remove empty .log files from <db>/vectors/idmap.0/ |
| HuggingFace model download fails | Network/firewall blocks huggingface.co | Use HF_ENDPOINT="https://hf-mirror.com" or ModelScope (see Getting Started tip) |
The CLI auto-cleans lock issues on startup when possible.
© QwenLM, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files in .qwen/skills/codegraph of QwenLM/qwen-code.
Open the folder on GitHubat commit 4970bfa
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in QwenLM/qwen-code, which our catalogue first saw on October 7, 2026.
CodeScope Codebase Graph Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| CodeScope Codebase Graph Analysis this skillQwenLM/qwen-code | 28k | 1 repos | ~9.3k | Automated safety check: Pass | Apache-2.0 | |
| Codegraph QAQwenLM/qwen-code-examples | 143 | — | ~4.3k | Automated safety check: Pass | None | |
| Code To Chartrongxinzy/RongxinAI | 154 | — | ~596 | Automated safety check: Pass | MIT | |
| Smart PR ReviewLeoYeAI/openclaw-master-skills | 2.2k | — | ~2.7k | Automated safety check: Notes | MIT | |
| Test Revieweraxelixlabs/axelix | 147 | — | ~3.2k | Automated safety check: Pass | LGPL-3.0 | |
| Code Reviewerjewbetcha/opentrace | 116 | 2 repos | ~1.1k | Automated safety check: Notes | MIT |
QwenLM/qwen-code-examples
Use CodeScope to analyze any indexed codebase via its graph database (neug) and vector index (zvec).
rongxinzy/RongxinAI
解析代码仓库的 import/依赖关系,自动生成架构图、流程图和组织架构图,输出 Mermaid 文本或 SVG 图片。支持 Python、JavaScript、TypeScript、Go 和 Java 项目。当用户需要可视化代码结构、分析模块依赖、生成架构文档,或提及“代码架构图”、“依赖关系图”、“流程图”、“组织架构图”、“Mermaid 图”等关键词时触发。
LeoYeAI/openclaw-master-skills
Opinionated AI code reviewer — not a yes-machine. An agent skill from LeoYeAI/openclaw-master-skills.
axelixlabs/axelix
Reviews test code in GitHub pull requests for isolation, public-API contract coverage, AAA structure, and correct exception assertions.
jewbetcha/opentrace
Comprehensive code review skill for TypeScript, JavaScript, Python, Swift, Kotlin, Go.
zereight/gitlab-mcp
Shared reference for naming, function size, complexity and error handling rules that reviewer agents apply across TypeScript, Python, Go, Rust, Java, C# and Swift.
QwenLM/qwen-code
Reproduces a feature from Codex or Claude Code in Qwen Code by running the reference agent under capture, reading the traces, then implementing matching behavior.
QwenLM/qwen-code
Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic.
QwenLM/qwen-code
Scheduled CI skill that scans a repository for small, certain docs, test and code hygiene issues and fixes them on one branch with a commit per finding.
QwenLM/qwen-code
Builds a rebranded Qwen Code desktop package from the Tauri shell using only a brand id and a logo, with sensible derived defaults.
QwenLM/qwen-code
Walks through capturing and comparing V8 heap snapshots to find memory leaks in the Qwen Code Node.js CLI, using tmux and the chrome-devtools CLI.
QwenLM/qwen-code
Drives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review.
Works with
Categories
Answers questions about code structure, history, bugs and PR risk using a CodeScope knowledge graph and semantic index built from the repository. CodeScope indexes source code into a two-layer knowledge graph, structure (functions, calls, imports, classes, modules) and evolution (commits, file changes, function modifications), plus semantic embeddings for every function. It supports Python, JavaScript and TypeScript, C and Java.
CodeScope Codebase Graph Analysis fits situations like: finding callers, callees and dependency chains in a large codebase; locating dead code, hotspots or tightly coupled modules; tracing a bug report to the most relevant code; scoring open PRs for risk and cross-PR conflicts.
Run `npx skills add QwenLM/qwen-code --skill codegraph -a claude-code`. Or copy the skill folder (.qwen/skills/codegraph in QwenLM/qwen-code) into .claude/skills/codegraph in your project. Claude Code loads it when a task matches its description.
Run `npx skills add QwenLM/qwen-code --skill codegraph -a codex`. Or copy the skill folder (.qwen/skills/codegraph in QwenLM/qwen-code) into .agents/skills/codegraph in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add QwenLM/qwen-code --skill codegraph -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/codegraph, .gemini/skills/codegraph, .github/skills/codegraph and .opencode/skills/codegraph in your project.
Going by SKILL.md and its folder, CodeScope Codebase Graph Analysis needs the command-line tools its instructions call (gh, pip, python and git) and credentials named GITHUB_TOKEN. Our summary lists: Python with the codegraph-ai package; A .codegraph index, created with codegraph init; GitHub access for issue and PR analysis.
SKILL.md names 2 domains. In commands or code: hf-mirror.com and modelscope.cn; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
CodeScope Codebase Graph Analysis is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 9.3k tokens (SKILL.md is roughly 37k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with CodeScope Codebase Graph Analysis: Codegraph QA (QwenLM/qwen-code-examples, 143 stars), Code To Chart (rongxinzy/RongxinAI, 154 stars), Smart PR Review (LeoYeAI/openclaw-master-skills, 2.2k stars) and Test Reviewer (axelixlabs/axelix, 147 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
QwenLM (a GitHub organization) maintains it in QwenLM/qwen-code, which has 28,337 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 7, 2026.
Source: QwenLM/qwen-code on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.