Graphify
EBISPOT/ols4
A skill your agent uses for any question about a codebase, its architecture, file relationships, or project content — especially when graphify-out/ exists, where the question should be treated as a…
Bulk-import a large RDF graph (code graph, corpus, GitHub history, etc.) into a DKG node's working memory.
$ npx skills add OriginTrail/dkg --skill dkg-importer -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install OriginTrail/dkg dkg-importer --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/OriginTrail/dkg.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/cli/skills/dkg-importer .claude/skills/dkg-importer && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "dkg-importer" agent skill from https://github.com/OriginTrail/dkg/tree/main/packages/cli/skills/dkg-importer into .claude/skills/dkg-importer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dkg-importer", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/OriginTrail/dkg/tree/main/packages/cli/skills/dkg-importerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add OriginTrail/dkg --skill dkg-importer -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install OriginTrail/dkg dkg-importer --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OriginTrail/dkg.git skills-src && mkdir -p .agents/skills && cp -r skills-src/packages/cli/skills/dkg-importer .agents/skills/dkg-importer && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "dkg-importer" agent skill from https://github.com/OriginTrail/dkg/tree/main/packages/cli/skills/dkg-importer into .agents/skills/dkg-importer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dkg-importer", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add OriginTrail/dkg --skill dkg-importer -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install OriginTrail/dkg dkg-importer --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OriginTrail/dkg.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/packages/cli/skills/dkg-importer .cursor/skills/dkg-importer && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "dkg-importer" agent skill from https://github.com/OriginTrail/dkg/tree/main/packages/cli/skills/dkg-importer into .cursor/skills/dkg-importer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dkg-importer", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/OriginTrail/dkg.git --path packages/cli/skills/dkg-importer--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add OriginTrail/dkg --skill dkg-importer -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install OriginTrail/dkg dkg-importer --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OriginTrail/dkg.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/packages/cli/skills/dkg-importer .gemini/skills/dkg-importer && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "dkg-importer" agent skill from https://github.com/OriginTrail/dkg/tree/main/packages/cli/skills/dkg-importer into .gemini/skills/dkg-importer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dkg-importer", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install OriginTrail/dkg dkg-importerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add OriginTrail/dkg --skill dkg-importer -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/OriginTrail/dkg.git skills-src && mkdir -p .github/skills && cp -r skills-src/packages/cli/skills/dkg-importer .github/skills/dkg-importer && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "dkg-importer" agent skill from https://github.com/OriginTrail/dkg/tree/main/packages/cli/skills/dkg-importer into .github/skills/dkg-importer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dkg-importer", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add OriginTrail/dkg --skill dkg-importer -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install OriginTrail/dkg dkg-importer --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OriginTrail/dkg.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/packages/cli/skills/dkg-importer .opencode/skills/dkg-importer && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "dkg-importer" agent skill from https://github.com/OriginTrail/dkg/tree/main/packages/cli/skills/dkg-importer into .opencode/skills/dkg-importer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dkg-importer", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
dkg-importerBulk-import a large RDF graph (code graph, corpus, GitHub history, etc.) into a DKG node's working memory.
Dkg Importer is an agent skill from OriginTrail/dkg. Bulk-import a large RDF graph (code graph, corpus, GitHub history, etc.) into a DKG node's working memory. Use this skill when you need to push more than a few thousand triples in a single import — it codifies the chunking budgets, the assertion-loop shape, the resumability manifest, and the canonical URI rules so your importer converges with every other importer in the workspace.
Its SKILL.md is about 8.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Knowledge Management, covering Knowledge graphs. It works with GitHub and Ethereum. The repository describes itself as: OriginTrail Decentralized Knowledge Graph (DKG) is a decentralized knowledge infrastructure for multi-agent AI memory — enabling agents to publish, verify, and query shared… The licence is Apache-2.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 493e6a1. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
curlFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
wikidata.orgFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
DKG_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Dkg Importer loads about 8.4k tokens when it runs. Until then it costs about 99 tokens; SKILL.md has 2,768 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from OriginTrail/dkg at commit 493e6a1, republished under its Apache-2.0 licence (© OriginTrail). 2,768 words, ~8,360 tokens.
.claude/skills/dkg-importer/SKILL.md (or your agent's skills folder).This skill is the agent-readable manual for bulk imports against a DKG V10 node. If you are about to write more than a few thousand triples in one logical operation — a code graph, a Markdown corpus, a GitHub issue archive, a domain-specific dataset — read this first. It documents the contract every existing in-tree importer follows, so the graphs you produce join naturally with graphs other agents and the scanners produce.
For the general node API surface (auth, contextGraphs, SWM/VM publish, SPARQL)
see packages/cli/skills/dkg-node/SKILL.md. This skill
sits one layer above: it assumes you already know how to call dkg_knowledge_asset_*
and focuses on how to call them at scale, repeatedly, without losing data on
restart and without fragmenting the graph against parallel producers.
The daemon's /api/knowledge-assets create + /api/knowledge-assets/<name>/{wm/write,swm/share} loop is the
chunked-write API. There is no /api/import/bulk and there will not be one
(see ADR 0002 for
the rejected-alternative analysis). To push a large graph you call the loop
many times, with each call staying under fixed budgets.
| Constant | Value | Where it lands |
|---|---|---|
CHUNK | 5,000 quads | Per POST /api/knowledge-assets/<name>/wm/write call |
ROOT_CHUNK | 400 URIs | Initial per-call batch for POST /api/knowledge-assets/<name>/swm/share; halve on the 4 MiB encoded-payload cap |
| Max concurrent writes within one assertion | 1 (sequential) | The daemon does not parallelise intra-assertion writes; the manifest in §3 tracks per-assertion state anyway |
| Max concurrent assertions | 4 | Safe across assertions; keeps memory bounded for laptop-class nodes |
These constants are conservative: a 5,000-quad N-Quads payload serialises at
roughly 1.0-1.5 MB, well under the daemon's 10 MB MAX_BODY_BYTES cap. Going
larger gives no throughput win and risks a 413 on URIs that serialise on the
heavy end.
These are the three hard caps you will hit if you push past the constants
above, with the verbatim error text the daemon emits. The HTTP body caps live
in packages/cli/src/daemon/http-utils.ts;
the gossip cap lives in
packages/core/src/constants.ts.
| Endpoint | Cap | Constant | Trigger | Error response |
|---|---|---|---|---|
POST /api/knowledge-assets/<name>/wm/write | 10 MB request body | MAX_BODY_BYTES | N-Quads payload too large | HTTP 413 "Request body too large (>10485760 bytes)" |
POST /api/knowledge-assets/<name>/swm/share | 256 KB request body | SMALL_BODY_BYTES | entities array too long (~4,000+ URIs at 60-char average) | HTTP 413 "Request body too large (>262144 bytes)" |
POST /api/knowledge-assets/<name>/swm/share | 4 MiB gossip message | DKG_GOSSIP_MAX_MESSAGE_BYTES | Promoted assertion's encoded payload exceeds 4 MiB | HTTP 500 "Promoted assertion too large for gossip (XXXX KB, limit 4 MB). Promote fewer entities per call." |
The two /swm/share caps are independent: the 256 KB body cap is on the
request you send (URI count × URI length); the 4 MiB gossip cap is on the
assertion that ends up in SWM (triples × N-Quads length). It is possible
to hit the gossip cap with a single-URI entities array, if the assertion
under that root is large enough. entities: "all" triggers it most often
because it asks the daemon to gossip every root in one message; for any
assertion above ~12k typical-sized triples, expect to split.
In practice this means a robust importer needs two independent halve-and-
retry paths on /swm/share: one for 413 (shrink the entities array) and
one for 500 (shrink the per-root scope or switch from "all" to explicit
batches of N ≤ 400 root URIs). See §5 Error handling
for the recipes.
Self-tune from /api/status. Future versions of the daemon advertise their
current per-call limits at /api/status under an importLimits block. If
present, use those values — they reflect any operator-side tuning. If absent
(older daemon), use the constants above verbatim.
For each logical slice of triples (one slice ≈ one source artefact: one file, one PR, one document, one record group):
POST /api/knowledge-assets { name, subGraphName, contextGraphId }
POST /api/knowledge-assets/<name>/wm/write { quads: [...] } ── one or more times
POST /api/knowledge-assets/<name>/swm/share { entities: [...] }Reference implementation — see scripts/lib/dkg-daemon.mjs
for DkgClient. writeAssertion auto-chunks at a conservative 500-triple
default (override via the second-argument batchSize); promote does not
chunk — split the entities array yourself before calling it for big imports.
import { DkgClient } from './scripts/lib/dkg-daemon.mjs';
const client = new DkgClient({ token: process.env.DKG_TOKEN });
await client.ensureProject({ id: 'my-corpus', name: 'My Corpus' });
await client.ensureSubGraph(client.cgId, 'code');
async function ensureAssertion(client, body) {
try {
await client.request('POST', '/api/knowledge-assets', body);
} catch (err) {
if (err.status === 400 && /already exists/i.test(JSON.stringify(err.body ?? err.message))) {
return;
}
throw err;
}
}
for (const partition of partitions) { // one source artefact
const triples = generateTriples(partition); // ≤ tens of thousands typical
const assertionName = `import-${partition.slug}`;
await ensureAssertion(client, {
contextGraphId: client.cgId,
name: assertionName,
subGraphName: 'code',
});
await client.writeAssertion({ // auto-chunks at 500 quads
contextGraphId: client.cgId,
assertionName,
subGraphName: 'code',
triples,
}, { batchSize: 5000 }); // bump if your triples are small
const entities = rootUrisFor(partition);
for (let i = 0; i < entities.length; i += 400) { // chunk promote ourselves
await client.promote({
contextGraphId: client.cgId,
assertionName,
subGraphName: 'code',
entities: entities.slice(i, i + 400),
});
}
}import os
import requests
PORT = int(os.environ.get('DKG_PORT', '9200'))
TOKEN_PATH = os.path.expanduser('~/.dkg/auth.token')
with open(TOKEN_PATH) as f:
token = f.read().strip()
H = {'Authorization': f'Bearer {token}', 'Content-Type': 'application/json'}
BASE = f'http://localhost:{PORT}/api'
CHUNK = 5000
ROOT_CHUNK = 400
def ensure_assertion(cg, name, sg):
res = requests.post(f'{BASE}/knowledge-assets',
headers=H, json={'contextGraphId': cg, 'name': name, 'subGraphName': sg})
if res.status_code == 400 and 'already exists' in res.text.lower():
return
res.raise_for_status()
def write_assertion(cg, name, sg, triples, entities):
ensure_assertion(cg, name, sg)
for i in range(0, len(triples), CHUNK):
requests.post(f'{BASE}/knowledge-assets/{name}/wm/write',
headers=H, json={'contextGraphId': cg, 'subGraphName': sg,
'quads': triples[i:i+CHUNK]}).raise_for_status()
for i in range(0, len(entities), ROOT_CHUNK):
requests.post(f'{BASE}/knowledge-assets/{name}/swm/share',
headers=H, json={'contextGraphId': cg, 'subGraphName': sg,
'entities': entities[i:i+ROOT_CHUNK]}).raise_for_status()A 10,000-partition import that fails on partition 7,453 must not start over
from partition 1. The pattern is a small RDF manifest assertion the importer
maintains as it works. The reference implementation lives in
scripts/lib/manifest.mjs.
import { createImportManifest, markPartitionStatus, loadImportManifest, pendingPartitions }
from './scripts/lib/manifest.mjs';
const importId = 'my-corpus-2026-01-15';
const partitions = enumerateSourceArtefacts().map((p) => p.key); // strings, one per slice
await createImportManifest({
client, importId, partitions, subGraphName: 'meta',
});createImportManifest writes a single urn:dkg:import:<id> assertion to the
meta sub-graph listing every partition with initialStatus = "pending".
Manifests follow the chunking contract automatically: the import root and
every partition URI are promoted to SWM in chunks of ROOT_CHUNK ≤ 400 so a
peer node (or this node after a restart) can read the manifest back from SWM
to resume — promoting only the root would leave partition triples in WM only.
for (const part of partitions) {
await markPartitionStatus({
client, importId, partitionKey: part, status: 'in_progress', subGraphName: 'meta',
});
try {
await importOne(part);
await markPartitionStatus({ client, importId, partitionKey: part, status: 'done', subGraphName: 'meta' });
} catch (err) {
await markPartitionStatus({ client, importId, partitionKey: part, status: 'failed', subGraphName: 'meta' });
throw err; // or continue, depending on policy
}
}Status events are append-only — each markPartitionStatus call writes a fresh
StatusEvent triple with a timestamp and promotes BOTH the partition root
and the new event root to SWM so peers (or a resume on a different node)
see the progress, not just the local WM. (The promote root list is [partIri, evIri] because the new partIri imp:statusEvent evIri edge has subject
partIri; promoting only evIri would leave that edge in WM and a peer-side
loadImportManifest would never observe the new status.)
loadImportManifest resolves the "current" status as the latest event per
partition using a standard "max row" SPARQL pattern (FILTER NOT EXISTS
against any later event), which avoids the classic SAMPLE+MAX
decorrelation foot-gun. This append-only pattern also avoids needing SPARQL
DELETE/INSERT and gives you a complete audit trail for free.
const { partitions: state } = await loadImportManifest({ client, importId, subGraphName: 'meta' });
const pending = pendingPartitions(state);
for (const part of pending) {
await importOne(part.key); // pick up where we left off
}If your import is producing nodes that other producers also produce — files, packages, GitHub PRs, etc. — reuse their URIs, don't fork a new namespace.
Canonical patterns (ADR 0003):
urn:dkg:code:package:<pkgName> Package (workspace name)
urn:dkg:code:file:<pkgName>/<relPath> Source file (relPath ≡ path inside the package)
urn:dkg:github:repo:<owner>/<name> GitHub repo node
urn:dkg:github:pr:<owner>/<name>/<num> GitHub PR
urn:dkg:github:issue:<owner>/<name>/<num> GitHub issue
urn:dkg:import:<id> Your own import manifestEncoding rule: every path segment is encodeURIComponent'd. A file with
spaces, @, +, parens, etc. would otherwise produce an IRI Oxigraph
rejects with Invalid IRI code point.
Pre-mint check:
dkg_memory_search with the unnormalised label.If you're producing the canonical code-graph triples for the workspace's own
packages, use the helpers in scripts/lib/ontology.mjs
rather than redeclaring class/property IRIs.
/wm/write (MAX_BODY_BYTES = 10 MB)You exceeded the request-body cap with too many N-Quads. Halve and retry:
try {
await client.writeOne(slice);
} catch (err) {
if (err.status !== 413) throw err;
// Halve the chunk size for the next attempt; exponential backoff is fine.
await client.writeOne(slice.slice(0, slice.length / 2));
await client.writeOne(slice.slice(slice.length / 2));
}If you hit 413 frequently on /wm/write, check /api/status for the daemon's
current importLimits and tune your CHUNK constant down. Don't paper over
it by bumping retries.
/swm/share (SMALL_BODY_BYTES = 256 KB)This is a different 413 from the /wm/write one: the share route uses a
smaller body limit because its requests should be small (just root URIs).
You hit it by sending too many URIs in entities — roughly 4,000+ URIs at
typical lengths. Recovery is to shrink the entities array, not the
underlying assertion:
async function promoteRoots(assertion, roots, batchSize = 400) {
for (let i = 0; i < roots.length; i += batchSize) {
const batch = roots.slice(i, i + batchSize);
try {
await client.request('POST', `/api/knowledge-assets/${assertion}/swm/share`, {
contextGraphId: client.cgId, subGraphName, entities: batch,
});
} catch (err) {
if (err.status !== 413) throw err;
// Smaller batch and retry the same range.
i -= batchSize;
batchSize = Math.max(50, Math.floor(batchSize / 2));
}
}
}/swm/share with "too large for gossip"The assertion you're promoting serialises to an encoded payload above 4 MiB.
This is independent of how many root URIs you pass — even
entities: ["<one-uri>"] can trip this if that one root's transitive
triple set is bigger than 4 MiB. The error message is verbatim:
HTTP 500 "Promoted assertion too large for gossip
(XXXX KB, limit 4 MB). Promote fewer entities per call."Recovery is to shrink the per-promote scope, which means: if you were
promoting with entities: "all", switch to an explicit URI batch sized so
its transitive triples land under 4 MiB. There is no formula because triple
fan-out varies; start around 200-400 roots per call for code graphs and
50-100 for prose corpora with long string literals, then halve on the cap.
async function promoteAllInBatches(assertion, allRoots) {
let batch = 400;
for (let i = 0; i < allRoots.length; i += batch) {
try {
await promoteRoots(assertion, allRoots.slice(i, i + batch));
} catch (err) {
if (!/too large for gossip/.test(err.message)) throw err;
i -= batch;
batch = Math.max(50, Math.floor(batch / 2));
}
}
}If both 413 and 500-gossip fire on the same import, you need both recovery
loops — they're orthogonal. As of PR #4 the daemon ships an
in-process async promote queue (POST /api/knowledge-assets/<name>/swm/share-async)
that removes both failure modes by promoting one classifying-and-retrying
job at a time in the background. See §6 Async promote queue
for the new contract — for a fresh importer, prefer the async path and skip
both halve-and-retry loops above.
Token problem, not a chunking problem. See
dkg-node/SKILL.md §4 "If you get 401 or 403 on a
protected route, diagnose in this order" — call GET /api/agent/identity
to confirm who the daemon thinks you are.
Standard retry with exponential backoff. The daemon does not implement
idempotency tokens. wm/write is safe to retry with the same payload
(duplicate triples are deduped server-side), and retrying swm/share
is safe too. Raw POST /api/knowledge-assets (create) returns HTTP 400 when the
assertion already exists; higher-level helpers can normalize that into
idempotent success by treating an already exists response as reuse.
WM survives restarts under the Oxigraph persistence contract
(the archived incident report characterises the bug fixed in OriginTrail/dkg#636-639). On resume,
loadImportManifest gives you the "where was I?" answer; if a particular
assertion's WM state is partial, you can either:
POST /api/knowledge-assets (create) "already exists" as reuse
(or call a helper that does), then re-run wm/write to re-assert the
same triples without duplication.POST /api/knowledge-assets/<name>/wm/discard
and start over from your last done partition.Rule 4: rootEntity ... already existsThis is the #1 trap for "real-world" graph importers — Wikidata, schema.org,
Graphify-style code graphs, EPCIS event streams, anything where the same
subject URI legitimately appears across many logical artefacts. It fires from
the daemon's autoPartition step (during finalize: true on create, or as
part of /api/knowledge-assets/{name}/vm/publish) and looks like:
HTTP 400 "Rule 4 violation: rootEntity <http://www.wikidata.org/entity/Q2831>
already exists as the root of knowledge collection 17 in context graph 4.
Use POST /api/update to extend the existing knowledge collection."The rule: every Knowledge Asset (KA) within a context graph has exactly
one root entity, and a given subject URI can be the root of at most one KA
per CG. Multiple KAs sharing a root would make on-chain ownership /
attribution ambiguous, so the contract enforces uniqueness. The error
message's /api/update hint is correct if you want to extend the existing
KA — but for a bulk import producing many KAs that mention the same
entities ("Michael Jackson appears in 500 of my 5,000 album KAs"), updating
isn't what you want. You want each KA to have its own unrelated root.
The fix — partition-scoped blank-node rewrite. Before submitting quads
for partition N, rewrite every Wikidata / external URI to a partition-scoped
blank node and anchor them under a single, unique-per-partition root:
function buildPartitionQuads(partitionIdx, rawQuads, anchorUri) {
// 1. Mint one partition-scoped anchor — this becomes the KA's sole root.
// URI is unique per partition; blank nodes underneath it inherit
// partition scope so Q2831 in partition 17 != Q2831 in partition 18
// from the contract's perspective.
const anchor = `<${anchorUri}>`; // e.g. urn:dkg:miles-stress:partition:17
const blankFor = new Map(); // subject-URI -> deterministic _:bN
let bnCounter = 0;
const blankNodeFor = (uri) => {
if (!blankFor.has(uri)) {
// Deterministic skolem-ish label keeps the rewrite repeatable across
// resume runs without coordinating state.
blankFor.set(uri, `_:p${partitionIdx}_b${bnCounter++}`);
}
return blankFor.get(uri);
};
// 2. Rewrite every non-anchor URI in the subject (and object, when an IRI)
// position to its partition-scoped blank node.
//
// Detect IRIs generically via the RFC 3986 scheme grammar rather than
// hard-coding a scheme list. Earlier drafts checked only `http` / `urn:`,
// which silently misses valid RDF IRIs that use other schemes — `did:`,
// `ipfs:`, `tag:`, `file:`, plain-IRI imports etc. — and lets colliding
// root entities leak through to keep hitting Rule 4 on subsequent
// partitions. If your parser exposes `term.termType === 'NamedNode'`,
// prefer that over the regex.
const ABS_IRI = /^[A-Za-z][A-Za-z0-9+\-.]*:/;
const isIri = (s) => ABS_IRI.test(s);
const out = [];
for (const { s, p, o } of rawQuads) {
const subj = isIri(s) ? blankNodeFor(s) : s;
const obj = (o.kind === 'iri' && o.value !== anchorUri)
? blankNodeFor(o.value)
: serializeObject(o);
out.push(`${subj} <${p}> ${obj} .`);
}
// 3. Link the anchor to every rewritten root with `<anchor> stress:contains <_:bN>`
// so the KA's transitive triple set is reachable from the single root.
for (const blank of new Set(blankFor.values())) {
out.push(`${anchor} <urn:dkg:stress:contains> ${blank} .`);
}
out.push(`${anchor} a <urn:dkg:stress:Partition> .`);
return out;
}The result: each KA has one root (the anchor), every Wikidata URI inside
appears only as a blank-node label, and partitions sharing entities don't
collide. Battle-tested in scripts/testnet-publish-stress/publish-loop.mjs
(Base Sepolia, miles-publish-stress-26may, 5000-partition stress run); see
that file for a full reference implementation including pace-control,
checkpointing and retry-with-unique-name.
If your data has a natural "real" root that's already unique per artefact (e.g. an EPCIS event ID, a GitHub PR URL, a build ID), use that as the anchor instead of minting a synthetic one — the blank-node rewrite still applies for everything under it.
Both the synchronous per-KA /api/knowledge-assets/{name}/vm/publish and the async promote queue
run through autoPartition, so this trap exists on both paths. Fix it
at the importer level, before any quads reach the daemon.
As of PR #4 in the async-promote-queue series the daemon ships an in-process
queue that converts the synchronous POST /api/knowledge-assets/<name>/swm/share
round-trip into a fire-and-forget enqueue. For bulk imports — where the
synchronous promote round-trip is the bottleneck — this is the recommended
path. See docs/specs/SPEC_ASYNC_PROMOTE_QUEUE.md
for the design and packages/cli/skills/dkg-node/SKILL.md §8 for the
in-daemon worker configuration.
The synchronous /swm/share route blocks on SWM insert + gossip publish. For a
multi-thousand-partition import that's the single biggest source of wall-clock
time: tens of minutes spent in the 413 / 500-gossip halve-and-retry recipes in
§5. The async route returns HTTP 202 { jobId, state: "queued" } immediately; an in-daemon worker dequeues, runs the same promote
logic, and writes the result back into a Control Graph for inspection.
You still need to chunk writes at CHUNK=5,000 per §1 — that cap hasn't
changed. The async queue specifically targets the promote half of the
loop.
| Method | Route | Purpose |
|---|---|---|
POST | /api/knowledge-assets/<name>/swm/share-async | Enqueue. Body: { contextGraphId, entities?: [...] | "all", subGraphName? }. Returns 202 { jobId, state: "queued", enqueuedAt }. Returns 409 { existingJobId } if there is already an active job for the same (contextGraphId, subGraphName, name). |
GET | /api/knowledge-assets/swm/share-jobs | List jobs. Query: state=queued,running,failed_retrying,succeeded,failed (comma-separated), contextGraphId=..., limit=N. Returns { jobs: [...] }. |
GET | /api/knowledge-assets/swm/share-jobs/<jobId> | Read one job: state, attempt.count, commitMarker, result, attempt.lastError with classification: transient|cap_exceeded|fatal. |
DELETE | /api/knowledge-assets/swm/share-jobs/<jobId> | Cancel a queued / failed_retrying job. 409 if the job is running (let the lease expire). |
POST | /api/knowledge-assets/swm/share-jobs/<jobId>/recover | Re-queue a failed job after fixing whatever was wrong (subdivide an over-large entity set, restart an upstream, etc.). |
Identical to §2 right up to the promote step, then swap the synchronous call for the enqueue + poll-or-fire-and-forget pattern:
for (const part of partitions) {
await markPartitionStatus({ client, importId, partitionKey: part.key, status: 'in_progress', subGraphName: 'meta' });
// CREATE + WRITE — unchanged from §2.
await client.request('POST', '/api/knowledge-assets', { name: part.assertion, subGraphName: part.subGraphName, contextGraphId: client.cgId });
for (const slice of chunks(part.quads, 5000)) {
await client.writeAssertion({ contextGraphId: client.cgId, assertionName: part.assertion, subGraphName: part.subGraphName, triples: slice });
}
// PROMOTE — async path. Returns immediately; the worker takes over.
const { jobId } = await client.request(
'POST',
`/api/knowledge-assets/${encodeURIComponent(part.assertion)}/swm/share-async`,
{ contextGraphId: client.cgId, subGraphName: part.subGraphName, entities: part.roots },
);
// For a typical importer: track the jobId in your own state and move on.
// Mark partition `done` only after the job reaches state="succeeded".
await trackAsyncPromote({ client, jobId, partitionKey: part.key });
}trackAsyncPromote can either:
GET /api/knowledge-assets/swm/share-jobs/<jobId> on a backoff until
state === "succeeded" (call markPartitionStatus(..., 'done') then) or
state === "failed" (recover or escalate per §6 below). Use a 250-1000ms
interval — the worker polls the queue at ~100ms by default.jobId in the manifest as a partitionStatus
metadata field and let a separate reconciliation pass at the end of the
import walk all in-flight jobs to a terminal state.attempt.lastError.classification)The worker classifies every failure into one of three buckets. Read it from
GET /api/knowledge-assets/swm/share-jobs/<jobId>:
| Classification | Retry? | Typical cause | Importer action |
|---|---|---|---|
transient | yes (until maxRetries=5 reached) | fetch failed / ECONNRESET / timeout | Wait — the worker auto-retries with backoff. No-op for the importer until the job leaves failed_retrying. |
cap_exceeded | no | Promoted assertion too large for gossip (4 MiB) or Request body too large (256 KB) | Re-enqueue with a smaller entities slice (the queue can't subdivide on its own — that's a future enhancement). Same halve-and-retry shape as the synchronous recipe, just applied at the queue layer. |
fatal | no | Bad request, missing assertion, etc. | Inspect the error message, fix the cause, then POST /api/knowledge-assets/swm/share-jobs/<jobId>/recover to re-queue. |
Note that cap_exceeded jobs reach state: "failed", not failed_retrying,
because retrying the same payload would just hit the cap again. The importer
is on the hook for subdivision — see §5 for the same
halve-and-retry recipe, except you re-enqueue rather than retry inline.
/swm/shareThe synchronous route is not deprecated. Use it when:
Otherwise, prefer /swm/share-async. The contract is identical (same body
shape, same chunking budgets, same per-root semantics) — the only difference
is when the SWM insert lands.
A running import can be inspected without interrupting it:
# Everything still queued for this context graph
curl -H "Authorization: Bearer $DKG_TOKEN" \
"http://localhost:9200/api/knowledge-assets/swm/share-jobs?contextGraphId=$CG_ID&state=queued,running"
# Anything that failed and is waiting on operator action
curl -H "Authorization: Bearer $DKG_TOKEN" \
"http://localhost:9200/api/knowledge-assets/swm/share-jobs?state=failed&contextGraphId=$CG_ID"This is the queue-level view; per-partition state still lives in the
manifest in §3. The two are complementary: the manifest tracks logical
progress (pending → in_progress → done); the queue tracks the
mechanical promote step. A partition stays in_progress while its
promote job is queued / running.
/wm/write call. It will hit 413
and you'll learn the chunk size the slow way.owl:sameAs (ADR 0003 §Reconciliation)
is the recovery path, not the steady state.await Promise.all(partitions.map(importOne)) with N > 4. The
daemon serialises intra-assertion writes anyway; >4 concurrent assertions
just inflates memory pressure without throughput gain./api/knowledge-assets/{name}/vm/publish mid-import. That's the SWM → VM
on-chain transition (costs TRAC, human-gated). It is not the
/swm/share (WM → SWM share) step. Confusing the two is the most common
"where did my money go?" mistake.Rule 4"
before any quads reach the daemon. The error message will tell you to use
/api/update, which is correct for "extend an existing KA" but wrong for
"produce many independent KAs that happen to mention the same entities".1. Decide your import id and partition keys (one per source artefact).
2. createImportManifest({ client, importId, partitions, subGraphName: 'meta' })
3. For each partition (≤ 4 concurrent):
a. markPartitionStatus(..., 'in_progress')
b. POST /api/knowledge-assets { name, subGraphName, contextGraphId }
c. POST /api/knowledge-assets/<name>/wm/write { quads } // chunks of ≤ 5000
d. POST /api/knowledge-assets/<name>/swm/share { entities } // chunks of ≤ 400 URIs
e. markPartitionStatus(..., 'done')
4. On 413: halve chunk + retry.
5. On crash: loadImportManifest → pendingPartitions → resume from step 3.
6. (Optional, human-gated) per-KA /api/knowledge-assets/{name}/vm/publish mints a sealed KA SWM → VM.1. Decide your import id and partition keys.
2. createImportManifest({ client, importId, partitions, subGraphName: 'meta' })
3. For each partition (≤ 4 concurrent):
a. markPartitionStatus(..., 'in_progress')
b. POST /api/knowledge-assets { name, subGraphName, contextGraphId }
c. POST /api/knowledge-assets/<name>/wm/write { quads } // chunks of ≤ 5000
d. POST /api/knowledge-assets/<name>/swm/share-async { entities } // returns 202 { jobId }
e. Persist jobId; do NOT mark partition done yet.
4. Reconciliation pass: for each in-flight jobId,
GET /api/knowledge-assets/swm/share-jobs/<jobId> until state ∈ {succeeded, failed}.
- succeeded → markPartitionStatus(..., 'done')
- failed + classification=cap_exceeded → subdivide entities + re-enqueue
- failed + classification=fatal → fix root cause + POST .../recover
- failed_retrying → wait; worker will auto-retry transient errors
5. On crash: loadImportManifest → pendingPartitions → resume from step 3.
6. (Optional, human-gated) per-KA /api/knowledge-assets/{name}/vm/publish mints a sealed KA SWM → VM.scripts/lib/manifest.mjs — reference manifest implementationscripts/lib/dkg-daemon.mjs — DkgClient with built-in chunkingscripts/lib/ontology.mjs — canonical code:* ontology constantspackages/cli/skills/dkg-node/SKILL.md — node API surface (auth, CGs, SWM/VM, SPARQL), incl. §8 async promote queue worker config© OriginTrail, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in packages/cli/skills/dkg-importer of OriginTrail/dkg.
Open the folder on GitHubat commit 493e6a1
Dkg Importer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Dkg Importer this skillOriginTrail/dkg | 144 | — | ~8.4k | Automated safety check: Pass | Apache-2.0 | |
| GraphifyEBISPOT/ols4 | 105 | 6 repos | ~9.5k | Automated safety check: Pass | Apache-2.0 | |
| Orchestratorwp-media/wp-rocket | 767 | — | ~10k | Automated safety check: Pass | GPL-2.0 | |
| Release Roundethereumjs/ethereumjs-monorepo | 2.8k | — | ~2k | Automated safety check: Pass | None | |
| Squash BugbotLFDT-Lineth/lineth-monorepo | 126 | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | |
| CodeScope Codebase Graph AnalysisQwenLM/qwen-code | 28k | 1 repos | ~9.3k | Automated safety check: Pass | Apache-2.0 |
EBISPOT/ols4
A skill your agent uses for any question about a codebase, its architecture, file relationships, or project content — especially when graphify-out/ exists, where the question should be treated as a…
wp-media/wp-rocket
User-facing entry point for the wp-rocket issue workflow. An agent skill from wp-media/wp-rocket.
ethereumjs/ethereumjs-monorepo
Runs a coordinated EthereumJS npm release round in six human-gated phases — intent and readiness, CHANGELOG, version bump, publish (human executes), post-publish verification, and announcements.
LFDT-Lineth/lineth-monorepo
Triage unresolved bot review comments on a GitHub PR. An agent skill from LFDT-Lineth/lineth-monorepo.
QwenLM/qwen-code
Answers questions about code structure, history, bugs and PR risk using a CodeScope knowledge graph and semantic index built from the repository.
QwenLM/qwen-code-examples
Use CodeScope to analyze any indexed codebase via its graph database (neug) and vector index (zvec).
Categories
Bulk-import a large RDF graph (code graph, corpus, GitHub history, etc.) into a DKG node's working memory. Dkg Importer is an agent skill from OriginTrail/dkg.) into a DKG node's working memory.
Dkg Importer fits situations like: you need to push more than a few thousand triples in a single import — it codifies the chunking budgets; the assertion-loop shape; the resumability manifest; the canonical URI rules so your importer converges with every other importer in the workspace.
Run `npx skills add OriginTrail/dkg --skill dkg-importer -a claude-code`. Or copy the skill folder (packages/cli/skills/dkg-importer in OriginTrail/dkg) into .claude/skills/dkg-importer in your project. Claude Code loads it when a task matches its description.
Run `npx skills add OriginTrail/dkg --skill dkg-importer -a codex`. Or copy the skill folder (packages/cli/skills/dkg-importer in OriginTrail/dkg) into .agents/skills/dkg-importer in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OriginTrail/dkg --skill dkg-importer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dkg-importer, .gemini/skills/dkg-importer, .github/skills/dkg-importer and .opencode/skills/dkg-importer in your project.
Going by SKILL.md and its folder, Dkg Importer needs the command-line tools its instructions call (curl) and credentials named DKG_TOKEN. Our summary lists: Python 3; A credential in DKG_TOKEN.
SKILL.md names 1 domain. In commands or code: wikidata.org; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Dkg Importer is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 8.4k tokens (SKILL.md is roughly 33k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Dkg Importer: Graphify (EBISPOT/ols4, 105 stars), Orchestrator (wp-media/wp-rocket, 767 stars), Release Round (ethereumjs/ethereumjs-monorepo, 2.8k stars) and Squash Bugbot (LFDT-Lineth/lineth-monorepo, 126 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
OriginTrail (a GitHub organization) maintains it in OriginTrail/dkg, which has 144 GitHub stars. The repository was last updated on October 9, 2026.
Source: OriginTrail/dkg on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.