Agent skill

Knowledge Graph Construction

by wentorai in wentorai/research-plugins

Build research knowledge graphs for literature synthesis and RAG systems

MITAuto-check passedKnowledge Management

Install Knowledge Graph Construction

skills CLI
$ npx skills add wentorai/research-plugins --skill knowledge-graph-construction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins knowledge-graph-construction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tools/knowledge-graph/knowledge-graph-construction .claude/skills/knowledge-graph-construction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
knowledge-graph-construction
GitHub stars
298
Used in
1 other repo
Token cost
~2.6k tokens
SKILL.md length
358 words
Files
1
Skills in repo
405
Repo updated
First seen
Licence
MIT

At a glance

Build research knowledge graphs for literature synthesis and RAG systems

  • Tasks that involve Knowledge graphs
  • SKILL.md covers Overview, Knowledge Graph Fundamentals, Entity and Relation Extraction and Graph Storage and Querying, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Retrieval-augmented generation

What it does

Knowledge Graph Construction is an agent skill from wentorai/research-plugins. Build research knowledge graphs for literature synthesis and RAG systems

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Knowledge Management, covering Knowledge graphs and Retrieval-augmented generation. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve Knowledge graphs
  • Tasks that involve Retrieval-augmented generation

Example prompts

  • “/knowledge-graph-construction”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and yaml).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • neo4j.com
    • networkx.org
    • docs.openalex.org
    • docs.llamaindex.ai
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Knowledge Graph Construction loads about 2.6k tokens when it runs. Until then it costs about 25 tokens; SKILL.md has 358 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~25
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 358 words, ~2,626 tokens.

Download SKILL.mdSave it as .claude/skills/knowledge-graph-construction/SKILL.md (or your agent's skills folder).
name
knowledge-graph-construction
description
Build research knowledge graphs for literature synthesis and RAG systems

Knowledge Graph Construction Guide

Overview

Knowledge graphs (KGs) organize information as networks of entities and relationships, making them powerful tools for research synthesis, literature exploration, and AI-augmented retrieval. In academic contexts, knowledge graphs can represent relationships between papers, authors, methods, datasets, findings, and concepts -- enabling queries like "Which methods have been applied to dataset X?" or "What are the common limitations reported across studies of Y?"

This guide covers building knowledge graphs for research applications: defining schemas (ontologies), extracting entities and relations from text, storing and querying graph data, and integrating knowledge graphs with Retrieval Augmented Generation (RAG) systems for AI-powered research assistants.

Whether you are building a personal research knowledge base, constructing a domain-specific literature graph, or developing a RAG system for an academic chatbot, these patterns provide a solid foundation.

Knowledge Graph Fundamentals

Core Components
ComponentDefinitionResearch Example
Entity (Node)A distinct concept or objectPaper, Author, Method, Dataset
Relation (Edge)A typed connection between entities"cites", "uses_method", "evaluates_on"
PropertyAn attribute of an entity or relationPaper.year, Author.affiliation
Ontology/SchemaFormal definition of entity and relation typesResearch ontology defining valid types
Designing a Research Ontology
yaml
# research_ontology.yaml
entities:
  Paper:
    properties: [title, year, doi, abstract, venue]
  Author:
    properties: [name, affiliation, orcid]
  Method:
    properties: [name, description, category]
  Dataset:
    properties: [name, domain, size, url]
  Finding:
    properties: [description, metric, value, significance]
  Concept:
    properties: [name, definition, domain]

relations:
  CITES:
    from: Paper
    to: Paper
  AUTHORED_BY:
    from: Paper
    to: Author
  USES_METHOD:
    from: Paper
    to: Method
  EVALUATES_ON:
    from: Paper
    to: Dataset
  REPORTS_FINDING:
    from: Paper
    to: Finding
  RELATED_TO:
    from: Concept
    to: Concept
  INTRODUCES:
    from: Paper
    to: Method

Entity and Relation Extraction

LLM-Based Extraction

Using a large language model to extract structured knowledge from paper abstracts:

python
import json
from openai import OpenAI

client = OpenAI()

EXTRACTION_PROMPT = """Extract entities and relationships from this research paper abstract.

Return JSON with:
- entities: list of {type, name, properties}
- relations: list of {source, relation, target}

Entity types: Paper, Method, Dataset, Finding, Concept
Relation types: USES_METHOD, EVALUATES_ON, REPORTS_FINDING, RELATED_TO, INTRODUCES

Abstract: {abstract}

Respond ONLY with valid JSON."""

def extract_from_abstract(abstract, paper_title):
    response = client.chat.completions.create(
        model="gpt-4o",
        messages=[
            {"role": "system", "content": "You are a research knowledge extraction system."},
            {"role": "user", "content": EXTRACTION_PROMPT.format(abstract=abstract)}
        ],
        response_format={"type": "json_object"},
        temperature=0
    )

    result = json.loads(response.choices[0].message.content)

    # Add the paper itself as an entity
    result['entities'].insert(0, {
        'type': 'Paper',
        'name': paper_title,
        'properties': {'abstract': abstract[:200]}
    })

    return result
SpaCy + Custom NER for Domain-Specific Extraction
python
import spacy
from spacy.tokens import Span

nlp = spacy.load("en_core_web_trf")

# Register custom entity types
@spacy.Language.component("research_entities")
def research_entity_component(doc):
    # Pattern-based recognition for methods
    method_patterns = [
        "random forest", "gradient boosting", "neural network",
        "transformer", "attention mechanism", "BERT", "GPT",
        "convolutional", "recurrent", "GAN"
    ]

    new_ents = list(doc.ents)
    for token in doc:
        for pattern in method_patterns:
            if pattern.lower() in doc[token.i:token.i+3].text.lower():
                span = doc.char_span(token.idx, token.idx + len(pattern),
                                     label="METHOD")
                if span and span not in new_ents:
                    new_ents.append(span)
    doc.ents = spacy.util.filter_spans(new_ents)
    return doc

nlp.add_pipe("research_entities", after="ner")
Show full SKILL.md (145 more words)Show less

Graph Storage and Querying

Neo4j (Production)
python
from neo4j import GraphDatabase

class ResearchGraph:
    def __init__(self, uri, user, password):
        self.driver = GraphDatabase.driver(uri, auth=(user, password))

    def add_paper(self, paper):
        with self.driver.session() as session:
            session.run("""
                MERGE (p:Paper {doi: $doi})
                SET p.title = $title, p.year = $year, p.abstract = $abstract
            """, **paper)

    def add_citation(self, citing_doi, cited_doi):
        with self.driver.session() as session:
            session.run("""
                MATCH (a:Paper {doi: $citing})
                MATCH (b:Paper {doi: $cited})
                MERGE (a)-[:CITES]->(b)
            """, citing=citing_doi, cited=cited_doi)

    def add_method_usage(self, paper_doi, method_name):
        with self.driver.session() as session:
            session.run("""
                MATCH (p:Paper {doi: $doi})
                MERGE (m:Method {name: $method})
                MERGE (p)-[:USES_METHOD]->(m)
            """, doi=paper_doi, method=method_name)

    def find_papers_using_method(self, method_name):
        with self.driver.session() as session:
            result = session.run("""
                MATCH (p:Paper)-[:USES_METHOD]->(m:Method {name: $method})
                RETURN p.title AS title, p.year AS year, p.doi AS doi
                ORDER BY p.year DESC
            """, method=method_name)
            return [dict(record) for record in result]

    def find_common_methods(self, doi1, doi2):
        with self.driver.session() as session:
            result = session.run("""
                MATCH (p1:Paper {doi: $doi1})-[:USES_METHOD]->(m:Method)
                      <-[:USES_METHOD]-(p2:Paper {doi: $doi2})
                RETURN m.name AS method
            """, doi1=doi1, doi2=doi2)
            return [record['method'] for record in result]
NetworkX (Lightweight / Prototyping)
python
import networkx as nx
import json

def build_research_graph(extracted_data_list):
    """Build a NetworkX graph from extracted paper data."""
    G = nx.MultiDiGraph()

    for data in extracted_data_list:
        for entity in data['entities']:
            G.add_node(
                entity['name'],
                type=entity['type'],
                **entity.get('properties', {})
            )

        for rel in data['relations']:
            G.add_edge(
                rel['source'],
                rel['target'],
                relation=rel['relation']
            )

    return G

# Query the graph
def get_method_landscape(G):
    """Find which methods are most used across papers."""
    methods = [n for n, d in G.nodes(data=True) if d.get('type') == 'Method']
    method_usage = {}
    for method in methods:
        papers = [n for n in G.predecessors(method)
                  if G.nodes[n].get('type') == 'Paper']
        method_usage[method] = len(papers)
    return sorted(method_usage.items(), key=lambda x: x[1], reverse=True)

Knowledge Graph + RAG Integration

Combining knowledge graphs with retrieval augmented generation creates powerful research assistants:

python
def kg_rag_query(question, graph, embedding_model, llm):
    """Answer a research question using KG-enhanced RAG."""

    # Step 1: Extract entities from the question
    question_entities = extract_entities(question)

    # Step 2: Retrieve relevant subgraph
    subgraph_nodes = set()
    for entity in question_entities:
        if entity in graph:
            # Get 2-hop neighborhood
            neighbors = nx.ego_graph(graph, entity, radius=2)
            subgraph_nodes.update(neighbors.nodes())

    # Step 3: Format context from subgraph
    context_parts = []
    for node in subgraph_nodes:
        node_data = graph.nodes[node]
        edges = list(graph.edges(node, data=True))
        context_parts.append(
            f"{node} ({node_data.get('type', 'Unknown')}): "
            f"{', '.join(f'{e[2].get(\"relation\", \"related_to\")} {e[1]}' for e in edges[:5])}"
        )
    context = '\n'.join(context_parts[:20])

    # Step 4: Generate answer with LLM
    prompt = f"""Based on the following knowledge graph context, answer the question.

Context:
{context}

Question: {question}

Provide a detailed answer citing specific papers, methods, and findings from the context."""

    return llm.generate(prompt)

Best Practices

  • Start with a clear schema. Define your entity types and relations before extracting data. A schema change later requires re-processing.
  • Use persistent identifiers. DOIs for papers, ORCIDs for authors, and canonical names for methods prevent duplicate nodes.
  • Validate extracted triples. LLM extraction is imperfect. Sample and manually verify 5-10% of extractions.
  • Enrich with external data. Link your KG to OpenAlex, CrossRef, or Wikidata for additional metadata.
  • Version your graph. Export snapshots regularly and track changes over time.
  • Design queries before building. Know what questions you want to answer before deciding on the schema.

References

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/tools/knowledge-graph/knowledge-graph-construction of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Knowledge Graph Construction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Knowledge Graph Construction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Knowledge Graph Construction this skillwentorai/research-plugins2981 repos~2.6kAutomated safety check: PassMIT
GraphmemorybradAGI/GraphMemory160—~3.1kAutomated safety check: PassMIT
Neo4j Document Import Skillneo4j-contrib/neo4j-skills114—~5.4kAutomated safety check: NotesMIT
Cortexdb Memory Hermesliliang-cn/cortexdb274—~1.7kAutomated safety check: PassMIT
Cortexdb Memory Openclawliliang-cn/cortexdb274—~1.6kAutomated safety check: PassMIT
Docsmint Document ManagerHiAi-gg/docsmint118—~584Automated safety check: PassApache-2.0

Similar skills

  • Graphmemory

    bradAGI/GraphMemory

    Build and query embedded GraphRAG knowledge graphs with DuckDB-backed vector, full-text, and hybrid search.

    160 GitHub stars~3.1k tokensUpdated 5 mo ago
    Knowledge ManagementAuto-check passed
  • Neo4j Document Import Skill

    neo4j-contrib/neo4j-skills

    Ingests unstructured and semi-structured documents into Neo4j as a knowledge graph.

    114 GitHub stars~5.4k tokensUpdated yesterday
    Knowledge ManagementAuto-check: notes
  • Cortexdb Memory Hermes

    liliang-cn/cortexdb

    Give a Python agent (such as Hermes Agent by Nous Research) durable, local-first memory plus a queryable SPARQL knowledge graph, backed by CortexDB through its gRPC sidecar and the cortexdb-client…

    274 GitHub stars~1.7k tokensUpdated yesterday
    Knowledge ManagementAuto-check passed
  • Cortexdb Memory Openclaw

    liliang-cn/cortexdb

    Give a Node.js agent (such as OpenClaw) durable, local-first memory plus a queryable SPARQL knowledge graph, backed by CortexDB through its gRPC sidecar and the cortexdb-client npm package.

    274 GitHub stars~1.6k tokensUpdated yesterday
    Knowledge ManagementAuto-check passed
  • Manage and research DocsMint documents through its scoped MCP tools, including categories, folders, hybrid search, GraphRAG, rerank, and index refresh.

    118 GitHub stars~584 tokensUpdated 9 days ago
    Knowledge ManagementAuto-check passed
  • Scholar RAG

    joshzyj/open-scholar-skill

    Build and query a local vector database + GraphRAG over your entire reference library (Zotero or a PDF folder) for literature review.

    168 GitHub stars~7.4k tokensUpdated 19 days ago
    Research & ScienceAuto-check: notes

More from wentorai/research-plugins

All 405 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Questions about Knowledge Graph Construction

What does Knowledge Graph Construction do?

Build research knowledge graphs for literature synthesis and RAG systems. Knowledge Graph Construction is an agent skill from wentorai/research-plugins.

When should I use Knowledge Graph Construction?

Knowledge Graph Construction fits situations like: tasks that involve Knowledge graphs; tasks that involve Retrieval-augmented generation.

How do I install Knowledge Graph Construction in Claude Code?

Run `npx skills add wentorai/research-plugins --skill knowledge-graph-construction -a claude-code`. Or copy the skill folder (skills/tools/knowledge-graph/knowledge-graph-construction in wentorai/research-plugins) into .claude/skills/knowledge-graph-construction in your project. Claude Code loads it when a task matches its description.

How do I install Knowledge Graph Construction in Codex?

Run `npx skills add wentorai/research-plugins --skill knowledge-graph-construction -a codex`. Or copy the skill folder (skills/tools/knowledge-graph/knowledge-graph-construction in wentorai/research-plugins) into .agents/skills/knowledge-graph-construction in your project. Codex loads it when a task matches its description.

Can I use Knowledge Graph Construction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill knowledge-graph-construction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/knowledge-graph-construction, .gemini/skills/knowledge-graph-construction, .github/skills/knowledge-graph-construction and .opencode/skills/knowledge-graph-construction in your project.

What does Knowledge Graph Construction need to run?

SKILL.md names no scripts, command-line tools or credentials: Knowledge Graph Construction is instructions for the agent only. Our summary lists: Python 3.

Does Knowledge Graph Construction access the network?

SKILL.md names 5 domains. As links in the text: neo4j.com, networkx.org, docs.openalex.org, docs.llamaindex.ai and github.com. This is read from the text; nothing was executed.

Is Knowledge Graph Construction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Knowledge Graph Construction use?

Knowledge Graph Construction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Knowledge Graph Construction use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Knowledge Graph Construction?

Skills that share tags, products or a category with Knowledge Graph Construction: Graphmemory (bradAGI/GraphMemory, 160 stars), Neo4j Document Import Skill (neo4j-contrib/neo4j-skills, 114 stars), Cortexdb Memory Hermes (liliang-cn/cortexdb, 274 stars) and Cortexdb Memory Openclaw (liliang-cn/cortexdb, 274 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Knowledge Graph Construction?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.