Agent skill

Build Knowledge Graph

by the-palindrome in the-palindrome/ml-knowledge-graph

Recursively builds a knowledge graph of mathematical concepts and ML algorithms.

MITAuto-check passedKnowledge Management

Install Build Knowledge Graph

skills CLI
$ npx skills add the-palindrome/ml-knowledge-graph --skill build-knowledge-graph -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install the-palindrome/ml-knowledge-graph build-knowledge-graph --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/the-palindrome/ml-knowledge-graph.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/build-knowledge-graph .claude/skills/build-knowledge-graph && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
build-knowledge-graph
GitHub stars
151
Token cost
~3.3k tokens
SKILL.md length
1,166 words
Files
2 (incl. scripts)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Recursively builds a knowledge graph of mathematical concepts and ML algorithms.

  • Works in 3 steps: Initialize → Iterate: decompose pending concepts → Finalize
  • The user asks to build a knowledge graph
  • SKILL.md covers Node format, Tool, Workflow and Concept decomposition rules, plus 4 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Build Knowledge Graph is an agent skill from the-palindrome/ml-knowledge-graph. Recursively builds a knowledge graph of mathematical concepts and ML algorithms. Use when the user asks to "build a knowledge graph", "decompose concepts", "map dependencies between algorithms", or provides seed concepts for graph expansion. Each node is a precisely definable concept (algorithm, theorem, mathematical object) with prerequisite edges.

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/graph.py`).

It sits in Knowledge Management, covering Knowledge graphs. The repository describes itself as: Knowledge graph explorer for machine learning. The licence is MIT.

When your agent uses it

  • The user asks to build a knowledge graph
  • Decompose concepts
  • Map dependencies between algorithms
  • Provides seed concepts for graph expansion

Example prompts

  • “build a knowledge graph”
  • “decompose concepts”
  • “map dependencies between algorithms”
  • “/build-knowledge-graph”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Initialize
  2. Iterate: decompose pending concepts
  3. Finalize

What it can do on your machine

Read from SKILL.md and the folder at commit c7e20bf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Build Knowledge Graph loads about 3.3k tokens when it runs. Until then it costs about 93 tokens; SKILL.md has 1,166 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~93
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from the-palindrome/ml-knowledge-graph at commit c7e20bf, republished under its MIT licence (© the-palindrome). 1,166 words, ~3,312 tokens.

Download SKILL.mdSave it as .claude/skills/build-knowledge-graph/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
build-knowledge-graph
description
Recursively builds a knowledge graph of mathematical concepts and ML algorithms. Use when the user asks to "build a knowledge graph", "decompose concepts", "map dependencies between algorithms", or provides seed concepts for graph expansion. Each node is a precisely definable concept (algorithm, theorem, mathematical object) with prerequisite edges.

Build Knowledge Graph

Recursively decompose ML algorithms and mathematical concepts into a directed knowledge graph. Each node is a precisely definable concept -- an algorithm, theorem, mathematical object, or operation (e.g. "eigenvalue", "transformer layer", "decision tree", "directed acyclic graph"). Nodes are never fields, disciplines, or vague terms.

The final graph must be a DAG (no directed cycles).

Node format

json
{
  "id": "a0237bba",
  "label": "singular value decomposition",
  "to": ["9f9d6c30", "d2155e0e"],
  "from": ["bf594ee2", "1c77c3c7"],
  "category": "Linear and multilinear algebra; matrix theory",
  "definition": "A factorization M = U S V^T where U, V are orthogonal and S is diagonal with non-negative entries...",
  "long_description": "The singular value decomposition (SVD) expresses any m × n real (or complex) matrix A as A = UΣV^T ...",
  "_depth": 3,
  "_pagerank": 0.00024,
  "_degree_centrality": 0.00120,
  "_betweenness_centrality": 0.0,
  "_descendant_ratio": 0.99759,
  "_prerequisite_ratio": 0.0,
  "_reachability_ratio": 0.99519
}
FieldDescription
idDeterministic 8-char hex SHA-256 of the lowercase label
labelCanonical lowercase name
toIDs of nodes that depend on this concept (this node is a prerequisite OF those)
fromIDs of nodes this concept depends on (prerequisites OF this node)
categoryMSC 2020 category (math) or CS task taxonomy (ML/CS)
definitionPrecise mathematical definition (1-3 sentences)
long_descriptionExtended mathematical description (200-300 words) covering formal definition, key properties, theorems, relationships, and intuition
_depthInternal structural depth (0 for nodes with no prerequisites; otherwise 1 + max(prereq depth))
_pagerankPageRank score on directed prerequisite graph
_degree_centralityNormalized directed degree centrality
_betweenness_centralityDirected betweenness centrality
_descendant_ratioDescendant count divided by number of nodes at strictly higher depth
_prerequisite_ratioPrerequisite count divided by number of nodes at strictly lower depth
_reachability_ratio(descendant count + prerequisite count) / total nodes

Tool

python3 .claude/skills/build-knowledge-graph/scripts/graph.py <subcommand> [args]
SubcommandPurpose
init <path>Create empty graph file
add-seeds <path> <s1> <s2> ...Add seed concepts at depth 0
pending <path> [--limit N] [--max-depth D]Show next N unvisited concepts
ingest <path> [--max-depth D]Read JSON decompositions from stdin, update graph
finalize <path>Post-process and enforce DAG: eliminate cycles, compute metrics on complete DAG, transitive-reduce, remove orphans, deduplicate, sort
stats <path>Print graph statistics

Workflow

1. Initialize
bash
GRAPH="knowledge_graph.json"
GR="python3 .claude/skills/build-knowledge-graph/scripts/graph.py"

$GR init "$GRAPH"
$GR add-seeds "$GRAPH" "concept 1" "concept 2" "concept 3"

If a graph file already exists and the user wants to resume/extend, skip init and just run pending to see what remains.

2. Iterate: decompose pending concepts

Repeat until pending prints NO_PENDING:

Step A -- Get the next batch:

bash
$GR pending "$GRAPH" --limit 10

Step B -- For every concept in the batch, produce a decomposition. Analyze them all at once. For each concept determine:

  1. Definition: precise mathematical/algorithmic definition (1-3 sentences).
  2. Long description: an extended mathematical description (200-300 words) covering the formal definition, key properties and theorems, relationship to other concepts, computational considerations, and intuition. Write at the level of a graduate mathematics or ML textbook. Use LaTeX-style notation for formulas.
  3. Prerequisites: specific concepts directly used in the definition. Each must be a single, precisely definable concept. Use canonical lowercase names. Decompose deeply into mathematical foundations -- if a concept relies on "vector", "function", "real numbers", "limit", "supremum", etc., list them.
  4. Category: one MSC 2020 top-level name (math) or one CS/ML task discipline (CS).

Format as a JSON array:

json
[
  {
    "label": "concept name",
    "definition": "...",
    "long_description": "...",
    "prerequisites": ["prereq a", "prereq b"],
    "category": "Category Name"
  }
]

Step C -- Pipe the JSON into ingest:

bash
cat <<'BATCH' | $GR ingest "$GRAPH"
[ ... JSON array ... ]
BATCH

Step D -- Read the ingest output to see how many nodes are pending. Go back to Step A.

3. Finalize
bash
$GR finalize "$GRAPH"
$GR stats "$GRAPH"

Always run finalize before delivering results. finalize enforces that the resulting graph is a DAG (cycles removed + transitive reduction applied). finalize also computes and saves node metrics on the complete DAG (after cycle elimination and before transitive reduction), then writes the reduced DAG. The final saved graph must include _pagerank, _degree_centrality, _betweenness_centrality, _descendant_ratio, _prerequisite_ratio, and _reachability_ratio on every node.

Report the final node count, edge count, top categories, and confirm that node metrics were computed and saved from the complete graph.

Concept decomposition rules

These rules are critical for graph quality:

  1. Nodes must be precise concepts, not fields.

    • Yes: "matrix multiplication", "softmax function", "cross-entropy loss", "convolution"
    • No: "linear algebra", "calculus", "deep learning", "statistics"
  2. Prerequisites must appear in the definition. Only list concepts that someone must understand to parse the definition. Do not list tangentially related concepts.

  3. Use canonical lowercase names: "rectified linear unit", "batch normalization", "bayes theorem", "singular value decomposition".

  4. Terminal concepts (do not decompose further -- the bare logical and set-theoretic bedrock): set, element, natural numbers, logical conjunction, logical disjunction, logical negation, logical implication, universal quantifier, existential quantifier, equality.

    Everything above this level must be decomposed. For example:

    • "logarithm" -> "inverse function", "exponential function"
    • "dot product" -> "vector", "multiplication", "summation"
    • "vector" -> "vector space"
    • "vector space" -> "field", "abelian group", "scalar multiplication"
    • "function" -> "relation", "domain", "codomain"
    • "relation" -> "ordered pair", "subset", "cartesian product"
    • "ordered pair" -> "set"
    • "real numbers" -> "complete ordered field", "dedekind cut" or "cauchy sequence"
    • "limit" -> "epsilon-delta definition", "real numbers", "absolute value"
    • "absolute value" -> "real numbers", "function"
    • "addition" -> "binary operation", "natural numbers" (then extends to integers, rationals, reals)
    • "multiplication" -> "binary operation", "natural numbers"
    • "summation" -> "addition", "index set", "sequence"
    • "sequence" -> "function", "natural numbers"

    The goal is a graph that bottoms out at foundational mathematics (set theory, logic, basic algebraic structures) rather than stopping at calculus-level concepts.

  5. No self-references: a concept cannot list itself as a prerequisite.

  6. Prefer specificity: "convolutional layer" decomposes into "convolution", "activation function", "bias vector" -- not into "neural network" or "deep learning".

  7. Result must be acyclic: the final delivered graph must be a DAG. Always run finalize to enforce cycle elimination and transitive reduction.

Show full SKILL.md (366 more words)Show less

Formatting guidelines

Both definition and long_description are Markdown strings with LaTeX math. Follow these rules strictly:

LaTeX math
  • Use $...$ for inline math and $$...$$ for display math.
  • All variable names, operators, and formulas must be in LaTeX, never plain ASCII math.
    • Yes: $f(x) = \sum_{i=1}^{n} w_i x_i$
    • No: f(x) = sum_i w_i x_i or f(x) = Σ wᵢxᵢ
  • Use proper LaTeX commands for symbols:
    • Greek letters: $\alpha$, $\beta$, $\Sigma$, $\epsilon$, $\theta$, $\lambda$
    • Operators: $\sum$, $\prod$, $\int$, $\nabla$, $\partial$, $\max$, $\min$, $\arg\max$, $\arg\min$
    • Relations: $\leq$, $\geq$, $\neq$, $\in$, $\subset$, $\subseteq$, $\forall$, $\exists$, $\implies$, $\iff$
    • Decorations: $\hat{y}$, $\bar{x}$, $\tilde{w}$, $\mathbf{x}$ (bold vectors), $\mathbb{R}$ (number sets), $\mathcal{L}$ (loss/Lagrangian)
    • Delimiters: $\left( ... \right)$, $\| \mathbf{x} \|$ for norms
    • Text in math: $\text{softmax}$, $\operatorname{ReLU}$
  • Use \text{} or \operatorname{} for multi-letter function names inside math mode — never bare words.
  • Prefer \lVert \cdot \rVert or \| \cdot \| for norms, not || ||.
Markdown structure (for long_description)
  • Use bold for key terms on first introduction.
  • Use bullet lists for enumerating properties or variants.
  • Separate logical sections (definition, properties, intuition) with line breaks.
  • Keep paragraphs short (2-4 sentences each).
  • Do not use headings (#, ##) inside descriptions — the description is a single node's content.
Example

definition (short):

"The **sigmoid function** is defined as $\\sigma(x) = \\frac{1}{1 + e^{-x}}$, mapping $\\mathbb{R} \\to (0, 1)$."

long_description (extended):

"The **sigmoid function** (also called the logistic function) is the smooth, monotonically increasing map $\\sigma : \\mathbb{R} \\to (0, 1)$ defined by\n\n$$\\sigma(x) = \\frac{1}{1 + e^{-x}}.$$\n\nIt arises naturally as the canonical link function for Bernoulli-distributed responses in generalized linear models, converting log-odds to probabilities.\n\n**Key properties:**\n\n- **Symmetry:** $\\sigma(-x) = 1 - \\sigma(x)$.\n- **Derivative:** $\\sigma'(x) = \\sigma(x)(1 - \\sigma(x))$, which is maximal at $x = 0$ (value $\\frac{1}{4}$) and vanishes as $|x| \\to \\infty$.\n- **Inverse:** The logit function $\\sigma^{-1}(p) = \\ln\\frac{p}{1-p}$.\n- **Limits:** $\\lim_{x \\to -\\infty} \\sigma(x) = 0$ and $\\lim_{x \\to +\\infty} \\sigma(x) = 1$.\n\nIn neural networks, the sigmoid was historically the default activation function but has been largely replaced by ReLU and its variants in hidden layers due to the **vanishing gradient problem**: for large $|x|$, $\\sigma'(x) \\approx 0$, causing gradients to shrink exponentially through deep layers during backpropagation.\n\nThe sigmoid remains standard in **output layers for binary classification**, where the output $\\sigma(\\mathbf{w}^\\top \\mathbf{x} + b)$ is interpreted as $P(y = 1 \\mid \\mathbf{x})$, and in **gating mechanisms** (LSTM, GRU, mixture of experts) where a value in $(0, 1)$ controls information flow.\n\nComputationally, care must be taken to evaluate $\\sigma$ in a numerically stable way, using $\\sigma(x) = e^x / (1 + e^x)$ for $x < 0$ and the standard form for $x \\geq 0$ to avoid overflow."

Category taxonomy

Mathematics -- use MSC 2020 top-level category names:

  • Combinatorics
  • Number theory
  • Linear and multilinear algebra; matrix theory
  • Real functions
  • Measure and integration
  • Probability theory and stochastic processes
  • Numerical analysis
  • Operations research, mathematical programming
  • Statistics
  • Calculus of variations and optimal control
  • Ordinary differential equations
  • Partial differential equations
  • Functional analysis
  • Approximations and expansions
  • Information and communication theory
  • (and other MSC 2020 categories as appropriate)

CS/ML -- use the narrowest applicable discipline:

  • Machine learning
  • Deep learning
  • Computer vision
  • Natural language processing
  • Reinforcement learning
  • Information theory
  • Neural and evolutionary computing
  • Pattern recognition
  • Optimization
  • Signal processing
  • Large language models
  • Computation and language
  • Artificial intelligence
  • Robotics
  • (and other CS task areas as appropriate)

Parameters

ParameterDefaultDescription
--max-depth100Max structural depth used when selecting pending nodes and ingest expansion
--limit10-20Batch size per iteration (adjust for speed vs. thoroughness)
Output pathknowledge_graph.jsonDefault graph file

Resumability

The graph file is the checkpoint. If a build is interrupted, simply run pending on the existing file to pick up where you left off. New seeds can be added to an existing graph with add-seeds.

© the-palindrome, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in .claude/skills/build-knowledge-graph of the-palindrome/ml-knowledge-graph.

  • SKILL.md
  • scripts/graph.py

Open the folder on GitHubat commit c7e20bf

Compare with similar skills

Build Knowledge Graph next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Build Knowledge Graph compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Build Knowledge Graph this skillthe-palindrome/ml-knowledge-graph151—~3.3kAutomated safety check: PassMIT
Obsidian Canvas BoardsAgriciDaniel/claude-obsidian15k—~1.4kAutomated safety check: PassMIT
Ontology1mancompany/OneManCompany4412 repos~1.5kAutomated safety check: PassApache-2.0
Knowledge Graphgnomeria/usbtree691—~1.5kAutomated safety check: PassMIT
Graphagenticnotetaking/arscontexta3.5k—~4.9kAutomated safety check: NotesMIT
LLM Wiki Knowledge GraphEgonex-AI/Understand-Anything86k—~1.5kAutomated safety check: PassMIT

Similar skills

  • Obsidian Canvas Boards

    AgriciDaniel/claude-obsidian

    Creates, inspects and updates Obsidian JSON Canvas boards in a vault, with text, file, link, group and edge nodes, using safe recoverable edits.

    15k GitHub stars~1.4k tokensUpdated 28 days ago
    Knowledge ManagementAuto-check passed
  • Ontology

    1mancompany/OneManCompany

    Typed knowledge graph for structured agent memory and composable skills.

    441 GitHub starsUsed in 2 repos~1.5k tokens
    Knowledge ManagementAuto-check passed
  • Knowledge Graph

    gnomeria/usbtree

    Set up and maintain a lightweight, file-based knowledge graph of the repo — entities, typed relations, decisions, gotchas — so agents load context fast instead of re-exploring the codebase every…

    691 GitHub stars~1.5k tokensUpdated 1 mo ago
    Knowledge ManagementAuto-check passed
  • Graph

    agenticnotetaking/arscontexta

    Interactive knowledge graph analysis. An agent skill from agenticnotetaking/arscontexta.

    3.5k GitHub stars~4.9k tokensUpdated 7 mo ago
    Knowledge ManagementAuto-check: notes
  • LLM Wiki Knowledge Graph

    Egonex-AI/Understand-Anything

    Detects a Karpathy-pattern LLM wiki and builds an interactive knowledge graph with entities, implicit relationships and topic clusters.

    86k GitHub stars~1.5k tokensUpdated today
    Knowledge ManagementAuto-check passed
  • Gitnexus Guide

    aws-samples/sample-kolya-br-proxy

    Official

    A skill your agent uses when the user asks about GitNexus itself — available tools, how to query the knowledge graph, MCP resources, graph schema, or workflow reference.

    106 GitHub starsUsed in 11 repos~867 tokens
    Knowledge ManagementAuto-check passed

Questions about Build Knowledge Graph

What does Build Knowledge Graph do?

Recursively builds a knowledge graph of mathematical concepts and ML algorithms. Build Knowledge Graph is an agent skill from the-palindrome/ml-knowledge-graph. Recursively builds a knowledge graph of mathematical concepts and ML algorithms.

When should I use Build Knowledge Graph?

Build Knowledge Graph fits situations like: the user asks to build a knowledge graph; decompose concepts; map dependencies between algorithms; provides seed concepts for graph expansion.

How do I install Build Knowledge Graph in Claude Code?

Run `npx skills add the-palindrome/ml-knowledge-graph --skill build-knowledge-graph -a claude-code`. Or copy the skill folder (.claude/skills/build-knowledge-graph in the-palindrome/ml-knowledge-graph) into .claude/skills/build-knowledge-graph in your project. Claude Code loads it when a task matches its description.

How do I install Build Knowledge Graph in Codex?

Run `npx skills add the-palindrome/ml-knowledge-graph --skill build-knowledge-graph -a codex`. Or copy the skill folder (.claude/skills/build-knowledge-graph in the-palindrome/ml-knowledge-graph) into .agents/skills/build-knowledge-graph in your project. Codex loads it when a task matches its description.

Can I use Build Knowledge Graph in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add the-palindrome/ml-knowledge-graph --skill build-knowledge-graph -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/build-knowledge-graph, .gemini/skills/build-knowledge-graph, .github/skills/build-knowledge-graph and .opencode/skills/build-knowledge-graph in your project.

What does Build Knowledge Graph need to run?

Going by SKILL.md and its folder, Build Knowledge Graph needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Build Knowledge Graph access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Build Knowledge Graph safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Build Knowledge Graph use?

Build Knowledge Graph is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Build Knowledge Graph use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Build Knowledge Graph?

Skills that share tags, products or a category with Build Knowledge Graph: Obsidian Canvas Boards (AgriciDaniel/claude-obsidian, 15k stars), Ontology (1mancompany/OneManCompany, 441 stars), Knowledge Graph (gnomeria/usbtree, 691 stars) and Graph (agenticnotetaking/arscontexta, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Build Knowledge Graph?

the-palindrome (a GitHub organization) maintains it in the-palindrome/ml-knowledge-graph, which has 151 GitHub stars. The repository was last updated on April 15, 2026.

Source: the-palindrome/ml-knowledge-graph on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.