Agent skill

Agentsop Idempotent Ingestion

by agentsope in agentsope/SkillAlchemy

Re-ingest-correctness SOP for production RAG. An agent skill from agentsope/SkillAlchemy.

MITAuto-check passedAI & LLM Engineering

Install Agentsop Idempotent Ingestion

skills CLI
$ npx skills add agentsope/SkillAlchemy --skill agentsop-idempotent-ingestion -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentsope/SkillAlchemy agentsop-idempotent-ingestion --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentsop-idempotent-ingestion .claude/skills/agentsop-idempotent-ingestion && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agentsop-idempotent-ingestion
GitHub stars
436
Token cost
~6.8k tokens
SKILL.md length
3,065 words
Files
5 (incl. references)
Skills in repo
46
Repo updated
First seen
Licence
MIT

At a glance

Re-ingest-correctness SOP for production RAG. An agent skill from agentsope/SkillAlchemy.

  • Works in 7 steps: 何时激活 (Activation Rules) → 核心心智模型 (Core Mental Model) → SOP 工作流 (Agentic Protocol) → …
  • Tasks that involve Operations and SOPs
  • SKILL.md covers 1. 何时激活 (Activation Rules), 2. 核心心智模型 (Core Mental Model), 3. SOP 工作流 (Agentic Protocol) and 4. 操作模型 (Operation Models), plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Agentsop Idempotent Ingestion is an agent skill from agentsope/SkillAlchemy. Re-ingest-correctness SOP for production RAG. Activate when a calling agent builds, reviews, or debugs an ingestion pipeline that runs more than once over a changing corpus — scheduled re-index, incremental updates, CI re-ingest, or a "retrieval has duplicates / shows deleted docs" bug. Encodes the rule — ingestion must be idempotent: a document's content hash decides insert/update/skip, so re-running over unchanged docs is a no-op — plus the docstore + doc-hash upsert machinery (LlamaIndex IngestionPipeline +…

Its SKILL.md is about 6.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `README.md`, `intermediate/operation_candidates.json` and `references/R1-source-evidence.md`).

It sits in AI & LLM Engineering, covering Operations and SOPs, Building AI agents and Retrieval-augmented generation. It works with LlamaIndex and LangChain. The repository describes itself as: From thought to skill. From signal to structure. The licence is MIT.

When your agent uses it

  • Tasks that involve Operations and SOPs
  • Tasks that involve Building AI agents
  • Tasks that involve Retrieval-augmented generation

Example prompts

  • “retrieval has duplicates / shows deleted docs”
  • “/agentsop-idempotent-ingestion”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. 何时激活 (Activation Rules)
  2. 核心心智模型 (Core Mental Model)
  3. SOP 工作流 (Agentic Protocol)
  4. 操作模型 (Operation Models)
  5. 困境决策案例 (Dilemma Cases)
  6. 反模式与边界 (Anti-patterns & Boundaries)
  7. 跨框架对照 (Cross-Framework Reference Table)

What it can do on your machine

Read from SKILL.md and the folder at commit d0f0355. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • developers.llamaindex.ai
    • python.langchain.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agentsop Idempotent Ingestion loads about 6.8k tokens when it runs, and up to ~10k if it reads all its reference files. Until then it costs about 209 tokens; SKILL.md has 3,065 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~209
When it runs · the whole SKILL.md, loaded when a task matches
~6.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~10k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agentsope/SkillAlchemy at commit d0f0355, republished under its MIT licence (© agentsope). 3,065 words, ~6,850 tokens.

Download SKILL.mdSave it as .claude/skills/agentsop-idempotent-ingestion/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
agentsop-idempotent-ingestion
description
Re-ingest-correctness SOP for production RAG. Activate when a calling agent builds, reviews, or debugs an ingestion pipeline that runs more than once over a changing corpus — scheduled re-index, incremental updates, CI re-ingest, or a "retrieval has duplicates / shows deleted docs" bug. Encodes the rule — **ingestion must be idempotent: a document's content hash decides insert/update/skip, so re-running over unchanged docs is a no-op** — plus the docstore + doc-hash upsert machinery (LlamaIndex `IngestionPipeline` + `DocstoreStrategy`), the delete-propagation problem, and cross-framework equivalents (LangChain `index()` + `RecordManager`, manual hash table). ENHANCE overlay over [[llamaindex]]: the IngestionPipeline exists in the base skill but the re-ingest-correctness contract is not surfaced.
version
0.1.0
trigger_keywords
idempotent ingestion, re-ingest, re-index, docstore, doc hash, upsert, duplicate chunks, IngestionPipeline, RecordManager, incremental update, scheduled reindex
when_to_use
any ingestion pipeline that will run more than once (cron, CI, webhook, manual re-run), a corpus where source documents are added, edited, or deleted over…
when_not_to_use
a one-shot index built once and never refreshed (static corpus, throwaway notebook), the corpus is small enough to fully rebuild from scratch in seconds and…

Idempotent Ingestion · Re-Ingest-Correctness SOP

Third-person operating model for a coder agent that owns ingestion correctness across repeated runs. The audience is the LLM agent writing or reviewing the pipeline code — not the end user.

One sentence: The interesting run is the second one. A correct pipeline hashes each document, looks the hash up in a docstore, and inserts / updates / skips accordingly — so re-running over unchanged docs changes nothing.

This skill is an ENHANCE overlay over [[llamaindex]]. The base skill names IngestionPipeline with docstore + UPSERTS_AND_DELETE as a hardening must-have (OP-08, anti-pattern A10) but stops at "use it". This overlay surfaces the correctness contract — what idempotency actually means, how the hash decides the action, and how it generalises beyond LlamaIndex.


1. 何时激活 (Activation Rules)

Activate this skill whenever any of the following holds:

  1. The codebase contains an ingestion call (VectorStoreIndex.from_documents, pipeline.run(documents=...), vectorstore.add_documents, a custom embed
    • upsert loop) and that call will execute more than once over a corpus that can change between runs.
  2. The user mentions any of: re-ingest, re-index, scheduled / nightly index refresh, incremental document updates, docstore, doc hash, upsert, RecordManager, "keep the index in sync with the source".
  3. A bug report says "I see the same answer chunk twice", "retrieval returns duplicates", "deleted a file but it still shows up in answers", "the index keeps growing every night even though nothing changed".
  4. PR review: new ingestion code that builds the index with no docstore / no hash-keyed dedup and the corpus is live (LlamaIndex base skill A10).
  5. A production RAG is about to ship and the ingestion side has only ever been tested as a cold first run.

Do not activate when:

  • The index is built once from a static corpus and never refreshed — there is no second run to make idempotent.
  • The corpus is tiny and rebuilding from scratch each run is measurably cheaper than maintaining a docstore (rare; verify, don't assume).
  • The task is purely query-time (retrieval, rerank, synthesis) with no ingestion path in scope.

2. 核心心智模型 (Core Mental Model)

Three principles. Violating any one means a second run can duplicate or corrupt the index — regardless of how good the retrieval stack downstream is.

Principle 1 — Ingestion must be idempotent; the hash is the decision

Re-running the pipeline over an unchanged document is a no-op. Run it once, run it a thousand times — same index. The mechanism: each document is reduced to a stable content hash (and a stable doc id); that hash is looked up in a docstore (a key→hash record store separate from the vector store); the lookup result decides the action.

            doc_id seen?           hash changed?
                │                       │
   ┌────────────┴───────────┐    ┌──────┴──────┐
   no                       yes  no            yes
   │                         │   │              │
INSERT (embed + upsert)      │  SKIP        UPDATE
                             └──────────────►(delete old chunks,
                                              re-embed, upsert)

The hash is the whole game. Without it, the pipeline cannot tell "I have seen this exact content" from "this is new" — so it re-embeds and re-adds everything, and the vector store accumulates duplicates. (LlamaIndex base skill failure #3 / A10; developers.llamaindex.ai/.../loading/ingestion_pipeline/.)

Principle 2 — The docstore is a separate ledger, not the vector store

The vector store holds embeddings keyed by node id. The docstore holds, per document, {doc_id → content_hash} (and the node ids derived from it). They are two stores with two jobs:

StoreHoldsAnswers
Docstoredoc_id → hash, doc→node mapping"have I seen this content before?"
Vector storenode_id → embedding + metadata"what is semantically near this query?"

The dedup decision happens against the docstore, before any embedding call — so an unchanged corpus costs zero embedding tokens on re-run. If the docstore is ephemeral (in-memory, lost on restart), every cold start looks like a first run and re-embeds the world. The docstore must be persisted to the same durability tier as the vector store (SimpleDocumentStore.persist(), Redis, MongoDB, Postgres). A docstore that doesn't survive the process is not a docstore.

Principle 3 — Deletes don't propagate for free

Insert and update are driven by documents that are present. A document that disappeared from the source emits no event — so a naive upsert pipeline never learns it should remove the orphaned chunks. Stale content keeps getting retrieved long after the source file is gone. Propagating deletes requires the pipeline to compare the full set of doc_ids seen this run against the set in the docstore, and purge the difference. In LlamaIndex this is the _AND_DELETE half of DocstoreStrategy.UPSERTS_AND_DELETE; in LangChain it is cleanup="full" / cleanup="incremental". Choosing the upsert-only strategy silently accepts stale ghosts — sometimes correct (append-only corpus), usually not.


3. SOP 工作流 (Agentic Protocol)

Five stages. Each gates the next. Stop and remediate at the first failure. The gate that matters most is Stage 4 — prove the second run is a no-op.

Stage 0 — Frame the re-run

Before code, answer:

  1. Run cadence: how does the pipeline get re-triggered? (cron / webhook / CI / manual). If the honest answer is "never, it's one-shot" → this skill is over-engineering; stop.
  2. Change shape: do source docs get added only, or also edited and deleted? This decides upsert-only vs upsert-and-delete (Principle 3).
  3. Doc identity: what is the stable doc_id? It must survive across runs — a file path, a CMS id, a URL. Not a random uuid generated at load time (that makes every run look new). LlamaIndex derives a hash from content + doc_id; if doc_id is unstable, idempotency breaks even with a docstore.
  4. Durability tier: where does the docstore live so it survives restarts?
Stage 1 — Attach a persisted docstore
python
from llama_index.core.ingestion import IngestionPipeline, DocstoreStrategy
from llama_index.core.storage.docstore import SimpleDocumentStore
from llama_index.core.node_parser import SentenceSplitter
from llama_index.embeddings.openai import OpenAIEmbedding

pipeline = IngestionPipeline(
    transformations=[
        SentenceSplitter(chunk_size=1024, chunk_overlap=20),
        OpenAIEmbedding(model="text-embedding-3-small"),
    ],
    docstore=SimpleDocumentStore(),          # the dedup ledger
    vector_store=vector_store,               # the embedding store
    docstore_strategy=DocstoreStrategy.UPSERTS_AND_DELETE,
)

For production, swap SimpleDocumentStore for a server-backed one (RedisDocumentStore, MongoDocumentStore, PostgresDocumentStore) so the ledger is shared across workers and survives restarts. A SimpleDocumentStore that is never .persist()-ed is the #1 cause of "it re-embeds everything every night".

Stage 2 — Pin stable doc ids and let the pipeline hash
python
docs = SimpleDirectoryReader("./data", filename_as_id=True).load_data()
# filename_as_id=True → doc.id_ is the file path → stable across runs.
# LlamaIndex then computes the content hash internally per node.
pipeline.run(documents=docs)

Hard rules:

  • doc_id comes from a stable external identity (path / CMS id / URL), never a fresh uuid per load. Set filename_as_id=True or assign doc.id_ explicitly.
  • Do not mutate document text in a non-deterministic transformation before the hash is taken (e.g. injecting datetime.now() into metadata) — that makes the hash change every run and defeats the skip path. Keep volatile metadata out of the hashed payload.
  • The embedding model is part of the index identity: changing it requires a full re-embed (LlamaIndex base skill A2), not an incremental upsert.
Stage 3 — Choose the upsert strategy deliberately
StrategyRe-run behaviourUse when
UPSERTSnew+changed docs upserted; deletes not propagatedappend-mostly corpus; deletes handled out-of-band
UPSERTS_AND_DELETEnew+changed upserted; vanished docs purgedthe corpus is a mirror of a source that can shrink
DUPLICATES_ONLYdedup identical docs within one run; no cross-run updateone-shot batch that may contain dupes; no live sync

Default for any system that mirrors a mutable source: UPSERTS_AND_DELETE. DUPLICATES_ONLY is not a re-ingest strategy — it only de-dupes within a single .run(); a second .run() with the same docs will not skip them.

Stage 4 — Prove the second run is a no-op (the gate)

Before merging, the PR must include — and CI must run — a test that runs the pipeline twice and asserts the second run did nothing observable:

python
def test_reingest_is_idempotent(pipeline, docs, vector_store):
    n1 = len(pipeline.run(documents=docs))      # cold run
    count_after_first = vector_store.count()

    n2 = len(pipeline.run(documents=docs))       # identical re-run
    count_after_second = vector_store.count()

    assert n2 == 0,                       "re-run re-processed unchanged docs"
    assert count_after_second == count_after_first, "re-run duplicated vectors"

pipeline.run() returns the list of nodes it actually processed; on a clean idempotent re-run over unchanged input that list is empty. If it isn't, the docstore is missing, ephemeral, or the doc ids / hashes are unstable — go back to Stage 1-2. Then add the edit case (change one doc → exactly that doc's chunks replaced) and, for UPSERTS_AND_DELETE, the delete case (remove one doc → its chunks gone, others untouched). See OP-06 / OP-07.

Stage 5 — Schedule, persist, observe
  • Persist the docstore after every run (storage_context.persist() / server-backed store auto-persists) — otherwise the next cold start re-embeds everything.
  • Run on a schedule, not by hand (LlamaIndex base skill OP-08): cron / Airflow / serverless cron. The whole point of idempotency is that a missed-or-doubled trigger is harmless.
  • Log per run: {docs_in, inserted, updated, skipped, deleted, embedding_calls}. On a steady-state corpus, inserted+updated+deleted should trend to ~0 and skipped ≈ docs_in. A run that re-embeds everything every night is the loud signal that idempotency is broken.

4. 操作模型 (Operation Models)

Format: Trigger / Action / Output / Evidence. Machine-readable in intermediate/operation_candidates.json.

OP-01 AttachDocstore
  • Trigger: any IngestionPipeline / ingestion path that will run more than once with no docstore wired.
  • Action: pass docstore= (persisted: Redis/Mongo/Postgres in prod, SimpleDocumentStore only for tests) plus docstore_strategy=. The docstore is the dedup ledger; without it dedup is structurally impossible.
  • Output: pipeline that can answer "have I seen this content?" before embedding.
  • Evidence: developers.llamaindex.ai/python/framework/module_guides/loading/ingestion_pipeline/; [[llamaindex]] OP-08 / A10.
OP-02 PinStableDocId
  • Trigger: docs loaded with random / per-run ids; "re-run re-adds everything" symptom.
  • Action: set filename_as_id=True or assign doc.id_ from a stable external identity (path / CMS id / URL). Never a fresh uuid per load.
  • Output: the same source document hashes to the same key across runs → the skip path can fire.
  • Evidence: LlamaIndex Document.doc_id / SimpleDirectoryReader(filename_as_id=True) docs.
OP-03 HashBeforeEmbed
  • Trigger: designing or reviewing the dedup decision point.
  • Action: ensure the content hash is computed and looked up in the docstore before the embedding transformation runs, so unchanged docs cost zero embedding tokens. In LlamaIndex this is automatic given a docstore; in a manual pipeline, hash first, embed only on miss/change.
  • Output: re-run cost on a steady corpus ≈ 0 embedding calls.
  • Evidence: IngestionPipeline dedup section; LangChain indexing API ("avoids re-writing unchanged content, avoids re-computing embeddings").
OP-04 PickUpsertStrategy
  • Trigger: choosing pipeline behaviour for the corpus's change shape.
  • Action: UPSERTS (append-mostly), UPSERTS_AND_DELETE (mirror of a shrinkable source — the production default), DUPLICATES_ONLY (in-run dedup only, not cross-run). Decide from Stage 0 change shape, not by default.
  • Output: documented strategy + one-line rationale.
  • Evidence: DocstoreStrategy enum; LlamaIndex ingestion docs.
OP-05 PropagateDeletes
  • Trigger: source docs can be deleted and stale chunks must not be retrieved.
  • Action: use UPSERTS_AND_DELETE; the pipeline diffs the doc_ids seen this run against the docstore and purges orphans. (LangChain: cleanup="full" for a full snapshot run, cleanup="incremental" for streamed batches.)
  • Output: deleted source docs → their vectors removed; retrieval stays in sync.
  • Evidence: DocstoreStrategy.UPSERTS_AND_DELETE; LangChain index(..., cleanup=...) docs.
OP-06 ReingestNoOpTest
  • Trigger: any ingestion code change merges.
  • Action: run the pipeline twice over identical input; assert the second run() returns 0 processed nodes and the vector count is unchanged (Stage 4).
  • Output: regression gate proving idempotency is wired through.
  • Evidence: pipeline.run() returns processed nodes; direct application of Principle 1.
OP-07 IncrementalUpdateTest
  • Trigger: corpus supports edits and/or deletes.
  • Action: edit one doc → assert exactly that doc's old chunks are gone and new ones present, others untouched. Delete one doc (under UPSERTS_AND_DELETE) → assert its chunks gone, neighbours intact.
  • Output: proof that update/delete are surgical, not full-rebuild.
  • Evidence: hash-change → update path; delete path of UPSERTS_AND_DELETE.
OP-08 PersistDocstore
  • Trigger: SimpleDocumentStore used, or docstore not persisted after run.
  • Action: storage_context.persist(persist_dir=...) after each run, or use a server-backed docstore that auto-persists and is shared across workers.
  • Output: cold start reads the ledger → does not re-embed the world.
  • Evidence: LlamaIndex storage / persistence docs; Redis/Mongo/Postgres docstore integrations.
OP-09 ManualHashUpsert
  • Trigger: no framework available (raw vector-store SDK, custom loader).
  • Action: maintain a {doc_id → sha256(content)} table next to the vector store; on each doc compute the hash, compare, and INSERT / UPDATE (delete-then-add chunks) / SKIP; track seen ids for delete propagation.
  • Output: framework-free idempotent ingestion with the same contract.
  • Evidence: this is what IngestionPipeline / RecordManager automate; see R2 cross-framework table.

5. 困境决策案例 (Dilemma Cases)

Show full SKILL.md (1,244 more words)Show less
Dilemma 1 — A document changed slightly: re-embed all, or diff?

困境: A 200-page handbook had one paragraph edited. The naive idempotent pipeline hashes at the document level → the whole doc's hash changes → all its chunks are deleted and re-embedded. Correct for safety, but on a large doc a one-line edit triggers a full re-embed of that document. Should the pipeline diff at the chunk level instead?

约束:

  • LlamaIndex's docstore dedup keys on the document hash by default — any change re-processes the whole document's nodes.
  • Chunk-level diffing requires stable chunk identity across edits, but SentenceSplitter re-segments the whole doc when text shifts, so chunk boundaries (and ids) move even for unedited text downstream of the edit.
  • Re-embedding one document is usually cheap; re-embedding the whole corpus every run (the failure this skill prevents) is the expensive thing.

决策步骤:

  1. Default to document-level hashing — it is correct and simple. A per-document re-embed on edit is acceptable for most corpora.
  2. If single documents are huge and edited frequently, split the source into smaller logical documents (per section / per page) so the hash granularity matches the edit granularity — each becomes its own doc_id.
  3. Only build true chunk-level content-addressed dedup (hash each chunk, stable chunk id) if profiling proves per-document re-embed is the real bottleneck. It is rarely worth the complexity.

结果: Move the granularity by splitting documents, not by hand-rolling chunk-diffing. Document-level hashing stays the default; the fix is upstream in how the source is segmented into docs.

可提取的操作: OP-02, OP-03, OP-04.

Dilemma 2 — Source documents were deleted: purge from the index, or keep?

困境: 50 files were removed from the source folder. The upsert-only pipeline leaves their chunks in the vector store forever — they keep surfacing in retrieval and the LLM cites documents that no longer exist. Switch to UPSERTS_AND_DELETE? But what if a run sees a partial corpus (a flaky reader returned only half the files) — would it then delete everything legitimately missing-this-run?

约束:

  • UPSERTS_AND_DELETE purges any doc_id in the docstore that was not seen in the current run. If the run's input is incomplete, that purge is catastrophic — it deletes live content.
  • Therefore delete-propagation is only safe when the run input is a complete snapshot of the source (LangChain calls this cleanup="full"); for streamed / partial batches you need cleanup="incremental" (per-batch, scoped by source).
  • Stale chunks are a correctness bug (wrong citations); over-aggressive delete is an availability bug (lost content). Both are bad.

决策步骤:

  1. Confirm the run input is a full snapshot. If the reader can return a partial set on failure, gate the delete: abort the run (don't purge) if the seen-doc count drops more than a sane threshold vs the previous run.
  2. Snapshot + complete → UPSERTS_AND_DELETE (LlamaIndex) / cleanup="full" (LangChain). Streamed batches scoped by source → cleanup="incremental".
  3. For regulated corpora, prefer soft delete (deleted=true metadata + query-time filter) over hard delete, so erasure is auditable and reversible.

结果: Purge — but only behind a "this is a complete snapshot" guard. Idempotency for deletes is UPSERTS_AND_DELETE plus a partial-input circuit breaker, never a blind diff-and-delete.

可提取的操作: OP-05, OP-07, OP-08.


6. 反模式与边界 (Anti-patterns & Boundaries)

Top anti-patterns (instant red flags in code review)
#Anti-patternWhy it's wrongCorrect move
A1VectorStoreIndex.from_documents(docs) re-run every cron tickNo docstore → re-embeds + re-adds everything; vector store grows without bound; duplicates pollute top-kIngestionPipeline(docstore=..., docstore_strategy=UPSERTS_AND_DELETE) (OP-01)
A2IngestionPipeline(...) with no docstore=The dedup decision has nowhere to look up the hash — same as no idempotencyPass a persisted docstore ([[llamaindex]] A10)
A3SimpleDocumentStore() never .persist()-edLedger lost on restart; every cold start looks like a first runPersist after each run, or use Redis/Mongo/Postgres docstore (OP-08)
A4Random / per-load doc_id (fresh uuid each run)Same content hashes to a new key each run → nothing ever skipsStable external id: filename_as_id=True / explicit doc.id_ (OP-02)
A5Volatile metadata (timestamps, run id) in the hashed payloadHash changes every run even for unchanged content → constant re-embedKeep volatile metadata out of the hashed body
A6DUPLICATES_ONLY treated as a re-ingest strategyOnly dedupes within one run; a second run does not skipUse UPSERTS / UPSERTS_AND_DELETE for cross-run (OP-04)
A7Upsert-only on a corpus whose source files get deletedStale chunks linger; LLM cites deleted docsUPSERTS_AND_DELETE behind a snapshot guard (OP-05, Dilemma 2)
A8Blind UPSERTS_AND_DELETE on a possibly-partial run inputA flaky reader returning half the files purges live contentGate delete on "complete snapshot"; circuit-break on big drops
A9No twice-run test in CIFirst duplicate / ghost found by a user in productionOP-06 mandatory on every PR
A10Swapping embedding model and expecting incremental upsert to "fix" old vectorsOld vectors were embedded by the old model; mixed space breaks retrievalFull re-embed; tag index with model name+version ([[llamaindex]] A2)
Boundaries — when this skill is not the right move
  • B1 One-shot index, never refreshed → no second run; skip.
  • B2 Corpus tiny enough that full rebuild each run is measurably cheaper than docstore maintenance → rebuild is fine; but profile before deciding.
  • B3 Append-only event log that never edits or deletes → UPSERTS (or even plain add) is sufficient; delete propagation is dead weight.
  • B4 The "duplicate results" bug is actually a chunk-overlap or retriever top_k issue, not a re-ingest issue → diagnose first; don't add a docstore to a problem it doesn't solve.
PR review smells
  • from_documents(...) inside anything that runs on a schedule.
  • IngestionPipeline(...) with no docstore= keyword.
  • SimpleDocumentStore() and no persist( anywhere in the file.
  • doc.id_ = str(uuid4()) / loader without filename_as_id.
  • datetime.now() / time.time() written into document metadata before ingestion.
  • docstore_strategy=DocstoreStrategy.DUPLICATES_ONLY on a live-sync pipeline.
  • A re-ingest function with no test that calls it twice.

7. 跨框架对照 (Cross-Framework Reference Table)

The same "hash → docstore → insert/update/skip" contract across the common stacks. Compact below; full runnable code in references/R2-cross-framework.md. Verified against current docs (May 2026).

  • LlamaIndex — IngestionPipeline(docstore=..., vector_store=..., docstore_strategy=DocstoreStrategy.UPSERTS_AND_DELETE). Hash is on the document (content + doc_id); pipeline.run(documents=docs) returns only the nodes it processed — empty list = clean no-op. Persist the docstore (Redis/Mongo/Postgres in prod) to keep idempotency across restarts.
  • LangChain — index(docs, SQLRecordManager(...), vectorstore, cleanup="full", source_id_key="source"). The RecordManager is the docstore-equivalent (a hash per source_id in a durable store). Returns {num_added, num_updated, num_skipped, num_deleted}; on an unchanged re-run num_skipped == len(docs) and the rest are 0. cleanup="full" propagates deletes for a complete snapshot, "incremental" for source-scoped batches.
  • Manual — keep a persisted {doc_id: sha256(text)} table; per doc: prev == h → SKIP, prev set → delete-then-add (UPDATE), else INSERT; after the loop purge stored_keys − seen (DELETE); persist(). This is exactly what the two frameworks automate.
7.x Side-by-side: "what happens on the second run?"
StackMechanismUnchanged re-runDeletes propagate?
LlamaIndex IngestionPipeline + docstoredocument hash in docstorerun() returns []only with UPSERTS_AND_DELETE
LangChain index() + RecordManagerhash per source_id in SQL/record storenum_skipped == len(docs)only with cleanup="full"/"incremental"
Manual hash table{doc_id: sha256} tableinner loop all continueonly if you diff seen vs stored keys
from_documents / raw add (no ledger)nonere-embeds + re-adds → duplicatesnever

The bottom row is the anti-pattern (A1). Every correct row has the same shape: a persisted hash ledger separate from the vector store, consulted before embedding, with an explicit delete-propagation switch.


References

Primary docs (cited inline above)
Companion files
  • references/R1-source-evidence.md — claims traced to base [[llamaindex]] skill + primary docs.
  • references/R2-cross-framework.md — full LlamaIndex / LangChain / manual code with delete-propagation variants.
  • intermediate/operation_candidates.json — OP-01..09 in machine-readable Trigger / Action / Output / Evidence form.
Base skill (overlay relationship)
  • [[llamaindex]] — base RAG SOP. This overlay enhances OP-08 (IngestionWithDocstore) and anti-pattern A10 with the full re-ingest-correctness contract.

© agentsope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in skills/agentsop-idempotent-ingestion of agentsope/SkillAlchemy.

  • SKILL.md
  • README.md
  • intermediate/operation_candidates.json
  • references/R1-source-evidence.md
  • references/R2-cross-framework.md

Open the folder on GitHubat commit d0f0355

Compare with similar skills

Agentsop Idempotent Ingestion next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agentsop Idempotent Ingestion compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agentsop Idempotent Ingestion this skillagentsope/SkillAlchemy436—~6.8kAutomated safety check: PassMIT
Neo4j Document Import Skillneo4j-contrib/neo4j-skills114—~5.4kAutomated safety check: NotesMIT
Neo4j Graphrag Skillneo4j-contrib/neo4j-skills114—~4.2kAutomated safety check: NotesMIT
Mem0 Platform SDKmem0ai/mem067k1 repos~2.2kAutomated safety check: PassApache-2.0
SynalinksSynaLinks/synalinks-skills907—~4.8kAutomated safety check: PassApache-2.0
Dive Into LangGraphluochang212/dive-into-langgraph457—~837Automated safety check: NotesCustom licence

Similar skills

  • Neo4j Document Import Skill

    neo4j-contrib/neo4j-skills

    Ingests unstructured and semi-structured documents into Neo4j as a knowledge graph.

    114 GitHub stars~5.4k tokensUpdated yesterday
    Knowledge ManagementAuto-check: notes
  • Neo4j Graphrag Skill

    neo4j-contrib/neo4j-skills

    Build GraphRAG retrieval pipelines on Neo4j using the neo4j-graphrag Python package (v1.22.0+).

    114 GitHub stars~4.2k tokensUpdated yesterday
    Knowledge ManagementAuto-check: notes
  • Adds persistent memory to AI apps with the Mem0 Python and TypeScript SDKs: store, search, update and delete user memories, with framework integrations.

    67k GitHub starsUsed in 1 repo~2.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Synalinks

    SynaLinks/synalinks-skills

    A skill your agent uses for anything involving the Synalinks neuro-symbolic LM framework (Keras-inspired): DataModel/Field/Input, JSON operators (+ & | ^ ~), synalinks.ops…

    907 GitHub stars~4.8k tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check passed
  • Dive Into LangGraph

    luochang212/dive-into-langgraph

    A Chinese-language guide and reference for building agents with LangGraph 1.0, from a first ReAct agent through middleware, memory, MCP, RAG and web search.

    457 GitHub stars~837 tokensUpdated 29 days ago
    AI & LLM EngineeringAuto-check: notes
  • Upgrade Stripe

    kanchengw/cnllm

    Guide for upgrading Stripe API versions and SDKs. An agent skill from kanchengw/cnllm.

    173 GitHub starsUsed in 3 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed

More from agentsope/SkillAlchemy

All 46 skills in this repo
  • Agentsop Aider

    agentsope/SkillAlchemy

    SOP for terminal-based, git-native AI pair programming with Aider (git work-tree + tree-sitter repo-map + edit-format + human-in-loop REPL).

    436 GitHub stars~3.5k tokensUpdated 2 days ago
    Auto-check passed
  • Agentsop Context Scope Discipline

    agentsope/SkillAlchemy

    Coder-agent working-file budget discipline: keep the editable working set (files you /add into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only…

    436 GitHub stars~3k tokensUpdated 2 days ago
    Auto-check passed
  • Agentsop Cost Tiered Models

    agentsope/SkillAlchemy

    Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…

    436 GitHub stars~3k tokensUpdated 2 days ago
    Auto-check passed
  • Agentsop Crewai

    agentsope/SkillAlchemy

    SOP for building multi-agent systems with CrewAI — role-based collaboration, sequential/hierarchical processes, Flows, memory, delegation.

    436 GitHub stars~4.8k tokensUpdated 2 days ago
    Auto-check passed
  • Agentsop Dify

    agentsope/SkillAlchemy

    SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable.

    436 GitHub stars~5.4k tokensUpdated 2 days ago
    Auto-check: notes
  • Agentsop Multiscale Chunking

    agentsope/SkillAlchemy

    Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.

    436 GitHub stars~4.9k tokensUpdated 2 days ago
    Auto-check passed

Questions about Agentsop Idempotent Ingestion

What does Agentsop Idempotent Ingestion do?

Re-ingest-correctness SOP for production RAG. An agent skill from agentsope/SkillAlchemy. Agentsop Idempotent Ingestion is an agent skill from agentsope/SkillAlchemy. Re-ingest-correctness SOP for production RAG.

When should I use Agentsop Idempotent Ingestion?

Agentsop Idempotent Ingestion fits situations like: tasks that involve Operations and SOPs; tasks that involve Building AI agents; tasks that involve Retrieval-augmented generation.

How do I install Agentsop Idempotent Ingestion in Claude Code?

Run `npx skills add agentsope/SkillAlchemy --skill agentsop-idempotent-ingestion -a claude-code`. Or copy the skill folder (skills/agentsop-idempotent-ingestion in agentsope/SkillAlchemy) into .claude/skills/agentsop-idempotent-ingestion in your project. Claude Code loads it when a task matches its description.

How do I install Agentsop Idempotent Ingestion in Codex?

Run `npx skills add agentsope/SkillAlchemy --skill agentsop-idempotent-ingestion -a codex`. Or copy the skill folder (skills/agentsop-idempotent-ingestion in agentsope/SkillAlchemy) into .agents/skills/agentsop-idempotent-ingestion in your project. Codex loads it when a task matches its description.

Can I use Agentsop Idempotent Ingestion in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentsope/SkillAlchemy --skill agentsop-idempotent-ingestion -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentsop-idempotent-ingestion, .gemini/skills/agentsop-idempotent-ingestion, .github/skills/agentsop-idempotent-ingestion and .opencode/skills/agentsop-idempotent-ingestion in your project.

What does Agentsop Idempotent Ingestion need to run?

SKILL.md names no scripts, command-line tools or credentials: Agentsop Idempotent Ingestion is instructions for the agent only. Our summary lists: Python 3.

Does Agentsop Idempotent Ingestion access the network?

SKILL.md names 2 domains. As links in the text: developers.llamaindex.ai and python.langchain.com. This is read from the text; nothing was executed.

Is Agentsop Idempotent Ingestion safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agentsop Idempotent Ingestion use?

Agentsop Idempotent Ingestion is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agentsop Idempotent Ingestion use?

About 6.8k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.1k tokens, read only when the agent opens those files.

What are the alternatives to Agentsop Idempotent Ingestion?

Skills that share tags, products or a category with Agentsop Idempotent Ingestion: Neo4j Document Import Skill (neo4j-contrib/neo4j-skills, 114 stars), Neo4j Graphrag Skill (neo4j-contrib/neo4j-skills, 114 stars), Mem0 Platform SDK (mem0ai/mem0, 67k stars) and Synalinks (SynaLinks/synalinks-skills, 907 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agentsop Idempotent Ingestion?

agentsope (a GitHub user) maintains it in agentsope/SkillAlchemy, which has 436 GitHub stars. The repository holds 46 skills in this directory. The repository was last updated on October 9, 2026.

Source: agentsope/SkillAlchemy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.