Neo4j Document Import Skill
neo4j-contrib/neo4j-skills
Ingests unstructured and semi-structured documents into Neo4j as a knowledge graph.
Re-ingest-correctness SOP for production RAG. An agent skill from agentsope/SkillAlchemy.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-idempotent-ingestion -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-idempotent-ingestion --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentsop-idempotent-ingestion .claude/skills/agentsop-idempotent-ingestion && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agentsop-idempotent-ingestion" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-idempotent-ingestion into .claude/skills/agentsop-idempotent-ingestion/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-idempotent-ingestion", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-idempotent-ingestionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-idempotent-ingestion -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-idempotent-ingestion --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/agentsop-idempotent-ingestion .agents/skills/agentsop-idempotent-ingestion && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agentsop-idempotent-ingestion" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-idempotent-ingestion into .agents/skills/agentsop-idempotent-ingestion/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-idempotent-ingestion", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-idempotent-ingestion -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-idempotent-ingestion --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/agentsop-idempotent-ingestion .cursor/skills/agentsop-idempotent-ingestion && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agentsop-idempotent-ingestion" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-idempotent-ingestion into .cursor/skills/agentsop-idempotent-ingestion/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-idempotent-ingestion", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/agentsope/SkillAlchemy.git --path skills/agentsop-idempotent-ingestion--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-idempotent-ingestion -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-idempotent-ingestion --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/agentsop-idempotent-ingestion .gemini/skills/agentsop-idempotent-ingestion && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agentsop-idempotent-ingestion" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-idempotent-ingestion into .gemini/skills/agentsop-idempotent-ingestion/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-idempotent-ingestion", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install agentsope/SkillAlchemy agentsop-idempotent-ingestionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add agentsope/SkillAlchemy --skill agentsop-idempotent-ingestion -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/agentsop-idempotent-ingestion .github/skills/agentsop-idempotent-ingestion && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-idempotent-ingestion" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-idempotent-ingestion into .github/skills/agentsop-idempotent-ingestion/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-idempotent-ingestion", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-idempotent-ingestion -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-idempotent-ingestion --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/agentsop-idempotent-ingestion .opencode/skills/agentsop-idempotent-ingestion && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-idempotent-ingestion" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-idempotent-ingestion into .opencode/skills/agentsop-idempotent-ingestion/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-idempotent-ingestion", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agentsop-idempotent-ingestionRe-ingest-correctness SOP for production RAG. An agent skill from agentsope/SkillAlchemy.
Agentsop Idempotent Ingestion is an agent skill from agentsope/SkillAlchemy. Re-ingest-correctness SOP for production RAG. Activate when a calling agent builds, reviews, or debugs an ingestion pipeline that runs more than once over a changing corpus — scheduled re-index, incremental updates, CI re-ingest, or a "retrieval has duplicates / shows deleted docs" bug. Encodes the rule — ingestion must be idempotent: a document's content hash decides insert/update/skip, so re-running over unchanged docs is a no-op — plus the docstore + doc-hash upsert machinery (LlamaIndex IngestionPipeline +…
Its SKILL.md is about 6.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `README.md`, `intermediate/operation_candidates.json` and `references/R1-source-evidence.md`).
It sits in AI & LLM Engineering, covering Operations and SOPs, Building AI agents and Retrieval-augmented generation. It works with LlamaIndex and LangChain. The repository describes itself as: From thought to skill. From signal to structure. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit d0f0355. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
developers.llamaindex.aipython.langchain.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agentsop Idempotent Ingestion loads about 6.8k tokens when it runs, and up to ~10k if it reads all its reference files. Until then it costs about 209 tokens; SKILL.md has 3,065 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from agentsope/SkillAlchemy at commit d0f0355, republished under its MIT licence (© agentsope). 3,065 words, ~6,850 tokens.
.claude/skills/agentsop-idempotent-ingestion/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.Third-person operating model for a coder agent that owns ingestion correctness across repeated runs. The audience is the LLM agent writing or reviewing the pipeline code — not the end user.
One sentence: The interesting run is the second one. A correct pipeline hashes each document, looks the hash up in a docstore, and inserts / updates / skips accordingly — so re-running over unchanged docs changes nothing.
This skill is an ENHANCE overlay over [[llamaindex]]. The base skill names
IngestionPipeline with docstore + UPSERTS_AND_DELETE as a hardening
must-have (OP-08, anti-pattern A10) but stops at "use it". This overlay
surfaces the correctness contract — what idempotency actually means, how the
hash decides the action, and how it generalises beyond LlamaIndex.
Activate this skill whenever any of the following holds:
VectorStoreIndex.from_documents,
pipeline.run(documents=...), vectorstore.add_documents, a custom embedRecordManager, "keep the index in sync with the source".Do not activate when:
Three principles. Violating any one means a second run can duplicate or corrupt the index — regardless of how good the retrieval stack downstream is.
Re-running the pipeline over an unchanged document is a no-op. Run it once, run it a thousand times — same index. The mechanism: each document is reduced to a stable content hash (and a stable doc id); that hash is looked up in a docstore (a key→hash record store separate from the vector store); the lookup result decides the action.
doc_id seen? hash changed?
│ │
┌────────────┴───────────┐ ┌──────┴──────┐
no yes no yes
│ │ │ │
INSERT (embed + upsert) │ SKIP UPDATE
└──────────────►(delete old chunks,
re-embed, upsert)The hash is the whole game. Without it, the pipeline cannot tell "I have seen
this exact content" from "this is new" — so it re-embeds and re-adds
everything, and the vector store accumulates duplicates. (LlamaIndex base skill
failure #3 / A10; developers.llamaindex.ai/.../loading/ingestion_pipeline/.)
The vector store holds embeddings keyed by node id. The docstore holds, per
document, {doc_id → content_hash} (and the node ids derived from it). They are
two stores with two jobs:
| Store | Holds | Answers |
|---|---|---|
| Docstore | doc_id → hash, doc→node mapping | "have I seen this content before?" |
| Vector store | node_id → embedding + metadata | "what is semantically near this query?" |
The dedup decision happens against the docstore, before any embedding
call — so an unchanged corpus costs zero embedding tokens on re-run. If the
docstore is ephemeral (in-memory, lost on restart), every cold start looks like
a first run and re-embeds the world. The docstore must be persisted to the
same durability tier as the vector store (SimpleDocumentStore.persist(),
Redis, MongoDB, Postgres). A docstore that doesn't survive the process is not a
docstore.
Insert and update are driven by documents that are present. A document that
disappeared from the source emits no event — so a naive upsert pipeline
never learns it should remove the orphaned chunks. Stale content keeps getting
retrieved long after the source file is gone. Propagating deletes requires the
pipeline to compare the full set of doc_ids seen this run against the set in
the docstore, and purge the difference. In LlamaIndex this is the _AND_DELETE
half of DocstoreStrategy.UPSERTS_AND_DELETE; in LangChain it is
cleanup="full" / cleanup="incremental". Choosing the upsert-only strategy
silently accepts stale ghosts — sometimes correct (append-only corpus), usually
not.
Five stages. Each gates the next. Stop and remediate at the first failure. The gate that matters most is Stage 4 — prove the second run is a no-op.
Before code, answer:
doc_id? It must survive across runs —
a file path, a CMS id, a URL. Not a random uuid generated at load time
(that makes every run look new). LlamaIndex derives a hash from content +
doc_id; if doc_id is unstable, idempotency breaks even with a docstore.from llama_index.core.ingestion import IngestionPipeline, DocstoreStrategy
from llama_index.core.storage.docstore import SimpleDocumentStore
from llama_index.core.node_parser import SentenceSplitter
from llama_index.embeddings.openai import OpenAIEmbedding
pipeline = IngestionPipeline(
transformations=[
SentenceSplitter(chunk_size=1024, chunk_overlap=20),
OpenAIEmbedding(model="text-embedding-3-small"),
],
docstore=SimpleDocumentStore(), # the dedup ledger
vector_store=vector_store, # the embedding store
docstore_strategy=DocstoreStrategy.UPSERTS_AND_DELETE,
)For production, swap SimpleDocumentStore for a server-backed one
(RedisDocumentStore, MongoDocumentStore, PostgresDocumentStore) so the
ledger is shared across workers and survives restarts. A SimpleDocumentStore
that is never .persist()-ed is the #1 cause of "it re-embeds everything every
night".
docs = SimpleDirectoryReader("./data", filename_as_id=True).load_data()
# filename_as_id=True → doc.id_ is the file path → stable across runs.
# LlamaIndex then computes the content hash internally per node.
pipeline.run(documents=docs)Hard rules:
doc_id comes from a stable external identity (path / CMS id / URL),
never a fresh uuid per load. Set filename_as_id=True or assign doc.id_
explicitly.datetime.now() into metadata) — that makes
the hash change every run and defeats the skip path. Keep volatile metadata
out of the hashed payload.| Strategy | Re-run behaviour | Use when |
|---|---|---|
UPSERTS | new+changed docs upserted; deletes not propagated | append-mostly corpus; deletes handled out-of-band |
UPSERTS_AND_DELETE | new+changed upserted; vanished docs purged | the corpus is a mirror of a source that can shrink |
DUPLICATES_ONLY | dedup identical docs within one run; no cross-run update | one-shot batch that may contain dupes; no live sync |
Default for any system that mirrors a mutable source: UPSERTS_AND_DELETE.
DUPLICATES_ONLY is not a re-ingest strategy — it only de-dupes within a
single .run(); a second .run() with the same docs will not skip them.
Before merging, the PR must include — and CI must run — a test that runs the pipeline twice and asserts the second run did nothing observable:
def test_reingest_is_idempotent(pipeline, docs, vector_store):
n1 = len(pipeline.run(documents=docs)) # cold run
count_after_first = vector_store.count()
n2 = len(pipeline.run(documents=docs)) # identical re-run
count_after_second = vector_store.count()
assert n2 == 0, "re-run re-processed unchanged docs"
assert count_after_second == count_after_first, "re-run duplicated vectors"pipeline.run() returns the list of nodes it actually processed; on a
clean idempotent re-run over unchanged input that list is empty. If it
isn't, the docstore is missing, ephemeral, or the doc ids / hashes are
unstable — go back to Stage 1-2. Then add the edit case (change one doc → exactly
that doc's chunks replaced) and, for UPSERTS_AND_DELETE, the delete case
(remove one doc → its chunks gone, others untouched). See OP-06 / OP-07.
storage_context.persist() /
server-backed store auto-persists) — otherwise the next cold start re-embeds
everything.{docs_in, inserted, updated, skipped, deleted, embedding_calls}. On a steady-state corpus, inserted+updated+deleted should
trend to ~0 and skipped ≈ docs_in. A run that re-embeds everything every
night is the loud signal that idempotency is broken.Format: Trigger / Action / Output / Evidence. Machine-readable in
intermediate/operation_candidates.json.
IngestionPipeline / ingestion path that will run more than
once with no docstore wired.docstore= (persisted: Redis/Mongo/Postgres in prod,
SimpleDocumentStore only for tests) plus docstore_strategy=. The docstore
is the dedup ledger; without it dedup is structurally impossible.developers.llamaindex.ai/python/framework/module_guides/loading/ingestion_pipeline/; [[llamaindex]] OP-08 / A10.filename_as_id=True or assign doc.id_ from a stable
external identity (path / CMS id / URL). Never a fresh uuid per load.Document.doc_id / SimpleDirectoryReader(filename_as_id=True) docs.UPSERTS (append-mostly), UPSERTS_AND_DELETE (mirror of a
shrinkable source — the production default), DUPLICATES_ONLY (in-run dedup
only, not cross-run). Decide from Stage 0 change shape, not by default.DocstoreStrategy enum; LlamaIndex ingestion docs.UPSERTS_AND_DELETE; the pipeline diffs the doc_ids seen this
run against the docstore and purges orphans. (LangChain: cleanup="full" for
a full snapshot run, cleanup="incremental" for streamed batches.)DocstoreStrategy.UPSERTS_AND_DELETE; LangChain index(..., cleanup=...) docs.run() returns 0 processed nodes and the vector count is unchanged (Stage 4).pipeline.run() returns processed nodes; direct application of
Principle 1.UPSERTS_AND_DELETE) → assert its chunks gone, neighbours intact.UPSERTS_AND_DELETE.SimpleDocumentStore used, or docstore not persisted after run.storage_context.persist(persist_dir=...) after each run, or use
a server-backed docstore that auto-persists and is shared across workers.{doc_id → sha256(content)} table next to the vector
store; on each doc compute the hash, compare, and INSERT / UPDATE
(delete-then-add chunks) / SKIP; track seen ids for delete propagation.IngestionPipeline / RecordManager automate; see
R2 cross-framework table.困境: A 200-page handbook had one paragraph edited. The naive idempotent pipeline hashes at the document level → the whole doc's hash changes → all its chunks are deleted and re-embedded. Correct for safety, but on a large doc a one-line edit triggers a full re-embed of that document. Should the pipeline diff at the chunk level instead?
约束:
SentenceSplitter re-segments the whole doc when text shifts, so chunk
boundaries (and ids) move even for unedited text downstream of the edit.决策步骤:
doc_id.结果: Move the granularity by splitting documents, not by hand-rolling chunk-diffing. Document-level hashing stays the default; the fix is upstream in how the source is segmented into docs.
可提取的操作: OP-02, OP-03, OP-04.
困境: 50 files were removed from the source folder. The upsert-only pipeline
leaves their chunks in the vector store forever — they keep surfacing in
retrieval and the LLM cites documents that no longer exist. Switch to
UPSERTS_AND_DELETE? But what if a run sees a partial corpus (a flaky reader
returned only half the files) — would it then delete everything legitimately
missing-this-run?
约束:
UPSERTS_AND_DELETE purges any doc_id in the docstore that was not seen
in the current run. If the run's input is incomplete, that purge is
catastrophic — it deletes live content.cleanup="full"); for streamed
/ partial batches you need cleanup="incremental" (per-batch, scoped by
source).决策步骤:
UPSERTS_AND_DELETE (LlamaIndex) / cleanup="full"
(LangChain). Streamed batches scoped by source → cleanup="incremental".deleted=true metadata +
query-time filter) over hard delete, so erasure is auditable and reversible.结果: Purge — but only behind a "this is a complete snapshot" guard.
Idempotency for deletes is UPSERTS_AND_DELETE plus a partial-input circuit
breaker, never a blind diff-and-delete.
可提取的操作: OP-05, OP-07, OP-08.
| # | Anti-pattern | Why it's wrong | Correct move |
|---|---|---|---|
| A1 | VectorStoreIndex.from_documents(docs) re-run every cron tick | No docstore → re-embeds + re-adds everything; vector store grows without bound; duplicates pollute top-k | IngestionPipeline(docstore=..., docstore_strategy=UPSERTS_AND_DELETE) (OP-01) |
| A2 | IngestionPipeline(...) with no docstore= | The dedup decision has nowhere to look up the hash — same as no idempotency | Pass a persisted docstore ([[llamaindex]] A10) |
| A3 | SimpleDocumentStore() never .persist()-ed | Ledger lost on restart; every cold start looks like a first run | Persist after each run, or use Redis/Mongo/Postgres docstore (OP-08) |
| A4 | Random / per-load doc_id (fresh uuid each run) | Same content hashes to a new key each run → nothing ever skips | Stable external id: filename_as_id=True / explicit doc.id_ (OP-02) |
| A5 | Volatile metadata (timestamps, run id) in the hashed payload | Hash changes every run even for unchanged content → constant re-embed | Keep volatile metadata out of the hashed body |
| A6 | DUPLICATES_ONLY treated as a re-ingest strategy | Only dedupes within one run; a second run does not skip | Use UPSERTS / UPSERTS_AND_DELETE for cross-run (OP-04) |
| A7 | Upsert-only on a corpus whose source files get deleted | Stale chunks linger; LLM cites deleted docs | UPSERTS_AND_DELETE behind a snapshot guard (OP-05, Dilemma 2) |
| A8 | Blind UPSERTS_AND_DELETE on a possibly-partial run input | A flaky reader returning half the files purges live content | Gate delete on "complete snapshot"; circuit-break on big drops |
| A9 | No twice-run test in CI | First duplicate / ghost found by a user in production | OP-06 mandatory on every PR |
| A10 | Swapping embedding model and expecting incremental upsert to "fix" old vectors | Old vectors were embedded by the old model; mixed space breaks retrieval | Full re-embed; tag index with model name+version ([[llamaindex]] A2) |
UPSERTS (or even
plain add) is sufficient; delete propagation is dead weight.from_documents(...) inside anything that runs on a schedule.IngestionPipeline(...) with no docstore= keyword.SimpleDocumentStore() and no persist( anywhere in the file.doc.id_ = str(uuid4()) / loader without filename_as_id.datetime.now() / time.time() written into document metadata before
ingestion.docstore_strategy=DocstoreStrategy.DUPLICATES_ONLY on a live-sync pipeline.The same "hash → docstore → insert/update/skip" contract across the common
stacks. Compact below; full runnable code in references/R2-cross-framework.md.
Verified against current docs (May 2026).
IngestionPipeline(docstore=..., vector_store=..., docstore_strategy=DocstoreStrategy.UPSERTS_AND_DELETE). Hash is on the
document (content + doc_id); pipeline.run(documents=docs) returns only the
nodes it processed — empty list = clean no-op. Persist the docstore
(Redis/Mongo/Postgres in prod) to keep idempotency across restarts.index(docs, SQLRecordManager(...), vectorstore, cleanup="full", source_id_key="source"). The RecordManager is the
docstore-equivalent (a hash per source_id in a durable store). Returns
{num_added, num_updated, num_skipped, num_deleted}; on an unchanged re-run
num_skipped == len(docs) and the rest are 0. cleanup="full" propagates
deletes for a complete snapshot, "incremental" for source-scoped batches.{doc_id: sha256(text)} table; per doc:
prev == h → SKIP, prev set → delete-then-add (UPDATE), else INSERT; after
the loop purge stored_keys − seen (DELETE); persist(). This is exactly
what the two frameworks automate.| Stack | Mechanism | Unchanged re-run | Deletes propagate? |
|---|---|---|---|
LlamaIndex IngestionPipeline + docstore | document hash in docstore | run() returns [] | only with UPSERTS_AND_DELETE |
LangChain index() + RecordManager | hash per source_id in SQL/record store | num_skipped == len(docs) | only with cleanup="full"/"incremental" |
| Manual hash table | {doc_id: sha256} table | inner loop all continue | only if you diff seen vs stored keys |
from_documents / raw add (no ledger) | none | re-embeds + re-adds → duplicates | never |
The bottom row is the anti-pattern (A1). Every correct row has the same shape: a persisted hash ledger separate from the vector store, consulted before embedding, with an explicit delete-propagation switch.
DocstoreStrategy):
https://developers.llamaindex.ai/python/framework/module_guides/loading/ingestion_pipeline/cleanup):
https://python.langchain.com/docs/how_to/indexing/references/R1-source-evidence.md — claims traced to base [[llamaindex]] skill + primary docs.references/R2-cross-framework.md — full LlamaIndex / LangChain / manual code with delete-propagation variants.intermediate/operation_candidates.json — OP-01..09 in machine-readable Trigger / Action / Output / Evidence form.[[llamaindex]] — base RAG SOP. This overlay enhances OP-08
(IngestionWithDocstore) and anti-pattern A10 with the full
re-ingest-correctness contract.© agentsope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in skills/agentsop-idempotent-ingestion of agentsope/SkillAlchemy.
Open the folder on GitHubat commit d0f0355
Agentsop Idempotent Ingestion next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agentsop Idempotent Ingestion this skillagentsope/SkillAlchemy | 436 | — | ~6.8k | Automated safety check: Pass | MIT | |
| Neo4j Document Import Skillneo4j-contrib/neo4j-skills | 114 | — | ~5.4k | Automated safety check: Notes | MIT | |
| Neo4j Graphrag Skillneo4j-contrib/neo4j-skills | 114 | — | ~4.2k | Automated safety check: Notes | MIT | |
| Mem0 Platform SDKmem0ai/mem0 | 67k | 1 repos | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| SynalinksSynaLinks/synalinks-skills | 907 | — | ~4.8k | Automated safety check: Pass | Apache-2.0 | |
| Dive Into LangGraphluochang212/dive-into-langgraph | 457 | — | ~837 | Automated safety check: Notes | Custom licence |
neo4j-contrib/neo4j-skills
Ingests unstructured and semi-structured documents into Neo4j as a knowledge graph.
neo4j-contrib/neo4j-skills
Build GraphRAG retrieval pipelines on Neo4j using the neo4j-graphrag Python package (v1.22.0+).
mem0ai/mem0
Adds persistent memory to AI apps with the Mem0 Python and TypeScript SDKs: store, search, update and delete user memories, with framework integrations.
SynaLinks/synalinks-skills
A skill your agent uses for anything involving the Synalinks neuro-symbolic LM framework (Keras-inspired): DataModel/Field/Input, JSON operators (+ & | ^ ~), synalinks.ops…
luochang212/dive-into-langgraph
A Chinese-language guide and reference for building agents with LangGraph 1.0, from a first ReAct agent through middleware, memory, MCP, RAG and web search.
kanchengw/cnllm
Guide for upgrading Stripe API versions and SDKs. An agent skill from kanchengw/cnllm.
agentsope/SkillAlchemy
SOP for terminal-based, git-native AI pair programming with Aider (git work-tree + tree-sitter repo-map + edit-format + human-in-loop REPL).
agentsope/SkillAlchemy
Coder-agent working-file budget discipline: keep the editable working set (files you /add into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only…
agentsope/SkillAlchemy
Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…
agentsope/SkillAlchemy
SOP for building multi-agent systems with CrewAI — role-based collaboration, sequential/hierarchical processes, Flows, memory, delegation.
agentsope/SkillAlchemy
SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable.
agentsope/SkillAlchemy
Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.
Works with
Re-ingest-correctness SOP for production RAG. An agent skill from agentsope/SkillAlchemy. Agentsop Idempotent Ingestion is an agent skill from agentsope/SkillAlchemy. Re-ingest-correctness SOP for production RAG.
Agentsop Idempotent Ingestion fits situations like: tasks that involve Operations and SOPs; tasks that involve Building AI agents; tasks that involve Retrieval-augmented generation.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-idempotent-ingestion -a claude-code`. Or copy the skill folder (skills/agentsop-idempotent-ingestion in agentsope/SkillAlchemy) into .claude/skills/agentsop-idempotent-ingestion in your project. Claude Code loads it when a task matches its description.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-idempotent-ingestion -a codex`. Or copy the skill folder (skills/agentsop-idempotent-ingestion in agentsope/SkillAlchemy) into .agents/skills/agentsop-idempotent-ingestion in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentsope/SkillAlchemy --skill agentsop-idempotent-ingestion -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentsop-idempotent-ingestion, .gemini/skills/agentsop-idempotent-ingestion, .github/skills/agentsop-idempotent-ingestion and .opencode/skills/agentsop-idempotent-ingestion in your project.
SKILL.md names no scripts, command-line tools or credentials: Agentsop Idempotent Ingestion is instructions for the agent only. Our summary lists: Python 3.
SKILL.md names 2 domains. As links in the text: developers.llamaindex.ai and python.langchain.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agentsop Idempotent Ingestion is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.8k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.1k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Agentsop Idempotent Ingestion: Neo4j Document Import Skill (neo4j-contrib/neo4j-skills, 114 stars), Neo4j Graphrag Skill (neo4j-contrib/neo4j-skills, 114 stars), Mem0 Platform SDK (mem0ai/mem0, 67k stars) and Synalinks (SynaLinks/synalinks-skills, 907 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
agentsope (a GitHub user) maintains it in agentsope/SkillAlchemy, which has 436 GitHub stars. The repository holds 46 skills in this directory. The repository was last updated on October 9, 2026.
Source: agentsope/SkillAlchemy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.