Agent skill

Kg Builder

by Mathews-Tom in Mathews-Tom/armory

Designs and builds knowledge graphs from documents — ontology modeling with domain/range constraints, entity/relation/event extraction, entity resolution, provenance and supersession, and GraphRAG…

MITAuto-check passedKnowledge Management

Install Kg Builder

skills CLI
$ npx skills add Mathews-Tom/armory --skill kg-builder -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Mathews-Tom/armory kg-builder --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/kg-builder .claude/skills/kg-builder && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
kg-builder
GitHub stars
328
Token cost
~2.9k tokens
SKILL.md length
1,288 words
Files
9 (incl. scripts, references)
Skills in repo
80
Repo updated
First seen
Licence
MIT

At a glance

Designs and builds knowledge graphs from documents — ontology modeling with domain/range constraints, entity/relation/event extraction, entity resolution, provenance and supersession, and GraphRAG…

  • Works in 4 steps: Design (do not skip) → Extract → Consolidate → …
  • Asked to build a knowledge graph
  • SKILL.md covers Reference Files, The deterministic / LLM boundary, Workflow and Output, plus 2 more sections
  • Runs Python scripts from its folder; calls uv

What it does

Kg Builder is an agent skill from Mathews-Tom/armory. Designs and builds knowledge graphs from documents — ontology modeling with domain/range constraints, entity/relation/event extraction, entity resolution, provenance and supersession, and GraphRAG serving. Use when asked to "build a knowledge graph", "design an ontology", "extract entities and relations", "deduplicate entities", "entity resolution", "add GraphRAG", or "graph memory for an agent". NOT for multi-agent task graphs or agent orchestration, use task-decomposer.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including scripts and reference files (for example `evals/cases.yaml`, `references/extraction.md` and `references/fusion.md`).

It sits in Knowledge Management, covering Knowledge graphs. The repository describes itself as: Curated, production-grade skills for AI coding agents. Battle-tested workflows for developers who use AI seriously. The licence is MIT.

When your agent uses it

  • Asked to build a knowledge graph
  • Design an ontology
  • Extract entities and relations
  • Deduplicate entities

Example prompts

  • “build a knowledge graph”
  • “design an ontology”
  • “extract entities and relations”
  • “/kg-builder”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Design (do not skip)
  2. Extract
  3. Consolidate
  4. Serve and maintain

What it can do on your machine

Read from SKILL.md and the folder at commit 4594fb7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Kg Builder loads about 2.9k tokens when it runs, and up to ~14k if it reads all its reference files. Until then it costs about 122 tokens; SKILL.md has 1,288 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~122
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~14k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Mathews-Tom/armory at commit 4594fb7, republished under its MIT licence (© Mathews-Tom). 1,288 words, ~2,897 tokens.

Download SKILL.mdSave it as .claude/skills/kg-builder/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
kg-builder
description
Designs and builds knowledge graphs from documents — ontology modeling with domain/range constraints, entity/relation/event extraction, entity resolution, provenance and supersession, and GraphRAG serving. Use when asked to "build a knowledge graph", "design an ontology", "extract entities and relations", "deduplicate entities", "entity resolution", "add GraphRAG", or "graph memory for an agent". NOT for multi-agent task graphs or agent orchestration, use task-decomposer.
metadata.version
1.0.0
metadata.category
data
metadata.tags
knowledge-graph, ontology, entity-resolution, graphrag, provenance
metadata.difficulty
advanced
metadata.phase
build

KG Builder

A knowledge graph is a product with a schema, not a pile of triples. Quality comes from pipeline order: model the domain before extracting, validate during extraction, fuse before storing, and attach provenance to every fact from the first write.

This skill covers the full build — value test, ontology, extraction, quality gate, entity resolution, serving, and maintenance — plus the boundary question that decides whether the result is trustworthy: which stages are deterministic code and which are LLM judgment.

Scope note. This is about knowledge graphs — what an agent remembers. It is not about task graphs, agent orchestration, or multi-agent topology.

Reference Files

FileContentsLoad When
references/ontology-design.mdCompetency questions, entity/relation types, domain/range, storage choicePhase 1
references/extraction.mdSource routing, NER/RE/EE prompt patterns, validation, failure modesPhase 2
references/fusion.mdBlocking, matching layers, merge policy, threshold bandsPhase 3
references/serving.mdGraphRAG retrieval, path queries, community summaries, query layerPhase 4
references/provenance-and-supersession.mdClaim model, append-only updates, contradiction handling, audit trailPhase 1 and Phase 4

The deterministic / LLM boundary

Decide this before writing code. Code owns control flow, identity, validation, and merges. The model gets contained judgments behind a typed interface, each with a measured baseline.

StageDeterministic (code)LLM judgment (measure it)
Source routingformat detection, structured mapping—
Entity extractionspan capture, type validation, dictionary matching"what entities are in this text"
Relation extractiondomain/range enforcement, endpoint checks"which relation does this sentence assert"
Quality gatesampling, scoring, thresholds—
Blockingkey generation, candidate pairing—
Matchingstring/attribute/structure scoringambiguous middle band only
Mergecanonical selection, edge union, lineage— (never let a model own a merge)
Servingtraversal, subgraph selection, serializationthe agent's own reasoning

Measure every LLM surface against a prompt-only baseline before trusting it. This is not theoretical caution. In a pre-registered real-model evaluation of an LLM-adjudicated dedup and contradiction loop, the loop trailed a plain prompt-only baseline by 0.28–0.33 on detection and safety across every provider cell tested. Adjudication that is not measured is decoration.

Workflow

Phase 1 — Design (do not skip)
  1. Value test. A graph pays off when queries are multi-hop ("who worked with X on projects using Y"), when entities recur across documents, or when the relationships are the data. If every query is a single-hop lookup or an aggregation, use a table and stop here. Write the kill criterion down before continuing.
  2. Competency questions. Write the 10–20 questions the graph must answer. These are the ontology's spec and its test suite. Anything you cannot path through the finished schema is a missing type or relation.
  3. Ontology. 5–15 entity types, 10–30 relation types, each relation with explicit domain and range. Precise verb names (ACQUIRED, DEPENDS_ON) — never RELATED_TO. Keep it in ontology.yaml as the single source of truth; every extraction prompt embeds it verbatim.
  4. Storage and identity. Choose property graph (default), RDF/OWL (interop, description-logic reasoning), or typed edges in SQLite (<50K nodes). Decide now how time and provenance attach to every fact — retrofitting provenance after fusion is effectively impossible.

Validate the schema before extracting anything:

bash
uv run scripts/validate_ontology.py ontology.yaml
Phase 2 — Extract
  1. Route by source type. Structured sources (databases, CSVs, APIs) map column → type in deterministic code with no model involved. Semi-structured sources (HTML tables, infoboxes) get per-layout parsers. Only unstructured text enters the LLM pipeline. Running NLP over already-structured data is the classic waste.
  2. Entities. Dictionaries and exact rules first for closed vocabularies — free, deterministic, perfect precision. LLM extraction for open text, with the ontology in the prompt. Always capture surface form, canonical guess, type, source pointer, and confidence.
  3. Relations. Extract only between entities that already passed step 6; never let relation extraction invent endpoints. Constrain output to the ontology's relation list and validate domain/range in code. Require a verbatim evidence quote that asserts the relation — co-occurrence is not assertion ("Musk discussed Twitter" is not OWNS).
  4. Events. For dynamic domains, extract events as first-class nodes (trigger + typed arguments + time), never flattened into pairwise edges — flattening loses which acquisition happened at which price.
Phase 3 — Consolidate
  1. Quality gate. Sample 50 items and score entity precision and relation precision before fusing anything. Target ≥90% precision. Fix the prompt or the rules, then re-run — never hand-patch the output. Recall improves with more passes; bad precision poisons the graph permanently.
  2. Fusion. Blocking → matching → merge. Blocking avoids O(n²) comparison; matching scores string, attribute, and structural evidence (two J. Smith nodes sharing three coauthors and an affiliation are one person; identical names with disjoint neighborhoods are not); merge policy is deterministic code that keeps the canonical name, unions aliases and edges, preserves conflicting values with provenance, and records merged_from for undo. An erroneous merge is far more damaging than a missed one — it silently fuses two entities' entire edge sets. Auto-merge only above the high band; queue the middle for review.

Check the blocking strategy against labeled pairs before running it at scale:

bash
uv run scripts/blocking_report.py candidates.jsonl --labels matches.jsonl
Show full SKILL.md (482 more words)Show less
Phase 4 — Serve and maintain
  1. Serving. Entity-link the query, expand 1–2 hops, serialize the subgraph as compact (head)-[REL {time, source}]->(tail) lines grouped by head. For multi-hop questions retrieve paths between the query's entities, not neighborhoods around each — the path is the answer skeleton. Cluster and pre-summarize for "what are the themes" questions.
  2. Maintenance. New facts supersede rather than overwrite: keep the prior claim with status, supersedes, and validity interval. When a new fact contradicts a stored one, keep both with time and provenance and prefer the newer at retrieval. Re-run fusion periodically — unmaintained memory graphs rot exactly like unfused extractions.

Output

Deliver these artifacts, in this order:

ArtifactContents
competency.mdThe 10–20 questions, each marked answerable or blocked
ontology.yamlEntity types, relation types with domain/range, event argument schemas
extraction/Per-source-type prompts and deterministic mappings
quality-report.mdSampled entity and relation precision, with sample size and method
fusion-report.mdBlocking reduction ratio, pair recall, merge counts per band
The graphNodes and edges, every one carrying source, extracted_at, confidence

Report precision as a sampled estimate with its sample size. A precision number without a stated sampling method is a vibe.

Working rules

  • Schema first, always. Extraction without an ontology produces a word cloud with arrows. If the user resists schema design, induce a minimal 5-type ontology from three sample documents and show it for approval — never skip to extraction.
  • Provenance on every fact. source, extracted_at, confidence. Non-negotiable; fusion and trust both depend on it.
  • Pilot before scale. Run 10 documents through all four phases first. The pilot exposes ontology gaps at 1% of the cost.
  • Never auto-accept an induced schema. LLM-proposed ontologies overfit their sample documents. Prune to the minimal set that answers the competency questions.
  • The LLM is stage machinery, not the pipeline. It slots into extraction and the ambiguous matching band. The surrounding schema, validation, and merge policy are what make the output a knowledge graph rather than a transcript.

Errors and troubleshooting

SymptomCauseFix
Graph full of Concept/Thing nodesExtracted without an ontologyPhase 1 first, then re-extract
Same person appears as four nodesNo canonical-form rule; fusion skippedDefine the rule in ontology.yaml; run Phase 3
Confident but wrong relationsCo-occurrence treated as assertionRequire evidence quotes; enforce domain/range in code
Events flattened into edge soupNo event argument schemaPromote events to first-class nodes with typed arguments
Precision collapses as sources growOne prompt drifting across document typesPer-source-type prompts; run the quality gate per source
Fusion merges two real entitiesThreshold too low; no structural layerRaise the auto-merge band; add neighborhood comparison; undo via merged_from
Multi-hop answers are wrong but fluentUnfused duplicates break pathsRe-run fusion; paths cannot cross duplicate boundaries
GraphRAG returns noiseHop expansion too wideCap at 2 hops, or re-rank; retrieve paths, not neighborhoods
Retrieval is stale after updatesFacts overwritten instead of supersededAdopt the claim model in references/provenance-and-supersession.md

© Mathews-Tom, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references) in skills/kg-builder of Mathews-Tom/armory.

  • SKILL.md
  • evals/cases.yaml
  • references/extraction.md
  • references/fusion.md
  • references/ontology-design.md
  • references/provenance-and-supersession.md
  • references/serving.md
  • scripts/blocking_report.py
  • scripts/validate_ontology.py

Open the folder on GitHubat commit 4594fb7

Compare with similar skills

Kg Builder next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Kg Builder compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Kg Builder this skillMathews-Tom/armory328—~2.9kAutomated safety check: PassMIT
Obsidian Canvas BoardsAgriciDaniel/claude-obsidian15k—~1.4kAutomated safety check: PassMIT
Ontology1mancompany/OneManCompany4412 repos~1.5kAutomated safety check: PassApache-2.0
Knowledge Graphgnomeria/usbtree691—~1.5kAutomated safety check: PassMIT
Graphagenticnotetaking/arscontexta3.5k—~4.9kAutomated safety check: NotesMIT
LLM Wiki Knowledge GraphEgonex-AI/Understand-Anything86k—~1.5kAutomated safety check: PassMIT

Similar skills

  • Obsidian Canvas Boards

    AgriciDaniel/claude-obsidian

    Creates, inspects and updates Obsidian JSON Canvas boards in a vault, with text, file, link, group and edge nodes, using safe recoverable edits.

    15k GitHub stars~1.4k tokensUpdated 28 days ago
    Knowledge ManagementAuto-check passed
  • Ontology

    1mancompany/OneManCompany

    Typed knowledge graph for structured agent memory and composable skills.

    441 GitHub starsUsed in 2 repos~1.5k tokens
    Knowledge ManagementAuto-check passed
  • Knowledge Graph

    gnomeria/usbtree

    Set up and maintain a lightweight, file-based knowledge graph of the repo — entities, typed relations, decisions, gotchas — so agents load context fast instead of re-exploring the codebase every…

    691 GitHub stars~1.5k tokensUpdated 1 mo ago
    Knowledge ManagementAuto-check passed
  • Graph

    agenticnotetaking/arscontexta

    Interactive knowledge graph analysis. An agent skill from agenticnotetaking/arscontexta.

    3.5k GitHub stars~4.9k tokensUpdated 7 mo ago
    Knowledge ManagementAuto-check: notes
  • LLM Wiki Knowledge Graph

    Egonex-AI/Understand-Anything

    Detects a Karpathy-pattern LLM wiki and builds an interactive knowledge graph with entities, implicit relationships and topic clusters.

    86k GitHub stars~1.5k tokensUpdated today
    Knowledge ManagementAuto-check passed
  • Gitnexus Guide

    aws-samples/sample-kolya-br-proxy

    Official

    A skill your agent uses when the user asks about GitNexus itself — available tools, how to query the knowledge graph, MCP resources, graph schema, or workflow reference.

    106 GitHub starsUsed in 11 repos~867 tokens
    Knowledge ManagementAuto-check passed

More from Mathews-Tom/armory

All 80 skills in this repo
  • Architecture Reviewer

    Mathews-Tom/armory

    Architecture reviews across 7 dimensions (structural, scalability, enterprise readiness, performance, security, ops, data) with scored reports.

    328 GitHub stars~4.6k tokensUpdated 3 days ago
    Auto-check passed
  • Concept To Image

    Mathews-Tom/armory

    Turn concepts into static HTML visuals exported as PNG or SVG files via HTML/CSS/SVG.

    328 GitHub stars~2.6k tokensUpdated 3 days ago
    Auto-check passed
  • Watch

    Mathews-Tom/armory

    A skill your agent uses when analyzing an existing video URL or local recording: "watch this video", "analyze youtube video", "summarize this video", "youtube transcript", "find this moment", "what…

    328 GitHub stars~2.8k tokensUpdated 3 days ago
    Auto-check passed
  • Code Refiner

    Mathews-Tom/armory

    Deep code simplification and refactoring preserving behavior across Python, Go, TypeScript, Rust.

    328 GitHub stars~3.1k tokensUpdated 3 days ago
    Auto-check passed
  • Concept To Video

    Mathews-Tom/armory

    Turn concepts into animated explainer videos using Manim (Python) with MP4/GIF output, audio overlay, multi-scene composition.

    328 GitHub stars~4.9k tokensUpdated 3 days ago
    Auto-check passed
  • Decision Map

    Mathews-Tom/armory

    Maps the unresolved architecture, policy, and scope decisions that must be answered before planning can start: one durable decision ticket per question on the issue tracker, typed and blocker-linked…

    328 GitHub stars~2.7k tokensUpdated 3 days ago
    Auto-check passed

Questions about Kg Builder

What does Kg Builder do?

Designs and builds knowledge graphs from documents — ontology modeling with domain/range constraints, entity/relation/event extraction, entity resolution, provenance and supersession, and GraphRAG…. Kg Builder is an agent skill from Mathews-Tom/armory. Designs and builds knowledge graphs from documents — ontology modeling with domain/range constraints, entity/relation/event extraction, entity resolution, provenance and supersession, and GraphRAG serving.

When should I use Kg Builder?

Kg Builder fits situations like: asked to build a knowledge graph; design an ontology; extract entities and relations; deduplicate entities.

How do I install Kg Builder in Claude Code?

Run `npx skills add Mathews-Tom/armory --skill kg-builder -a claude-code`. Or copy the skill folder (skills/kg-builder in Mathews-Tom/armory) into .claude/skills/kg-builder in your project. Claude Code loads it when a task matches its description.

How do I install Kg Builder in Codex?

Run `npx skills add Mathews-Tom/armory --skill kg-builder -a codex`. Or copy the skill folder (skills/kg-builder in Mathews-Tom/armory) into .agents/skills/kg-builder in your project. Codex loads it when a task matches its description.

Can I use Kg Builder in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Mathews-Tom/armory --skill kg-builder -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kg-builder, .gemini/skills/kg-builder, .github/skills/kg-builder and .opencode/skills/kg-builder in your project.

What does Kg Builder need to run?

Going by SKILL.md and its folder, Kg Builder needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Kg Builder access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Kg Builder safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Kg Builder use?

Kg Builder is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Kg Builder use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 11k tokens, read only when the agent opens those files.

What are the alternatives to Kg Builder?

Skills that share tags, products or a category with Kg Builder: Obsidian Canvas Boards (AgriciDaniel/claude-obsidian, 15k stars), Ontology (1mancompany/OneManCompany, 441 stars), Knowledge Graph (gnomeria/usbtree, 691 stars) and Graph (agenticnotetaking/arscontexta, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Kg Builder?

Mathews-Tom (a GitHub user) maintains it in Mathews-Tom/armory, which has 328 GitHub stars. The repository holds 80 skills in this directory. The repository was last updated on October 6, 2026.

Source: Mathews-Tom/armory on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.