RAG Architect
Jeffallan/claude-skills
Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.
Security-first SOP for multi-tenant RAG systems. An agent skill from agentsope/SkillAlchemy.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-multi-tenant-rag -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-multi-tenant-rag --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentsop-multi-tenant-rag .claude/skills/agentsop-multi-tenant-rag && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agentsop-multi-tenant-rag" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-multi-tenant-rag into .claude/skills/agentsop-multi-tenant-rag/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-multi-tenant-rag", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-multi-tenant-ragType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-multi-tenant-rag -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-multi-tenant-rag --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/agentsop-multi-tenant-rag .agents/skills/agentsop-multi-tenant-rag && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agentsop-multi-tenant-rag" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-multi-tenant-rag into .agents/skills/agentsop-multi-tenant-rag/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-multi-tenant-rag", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-multi-tenant-rag -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-multi-tenant-rag --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/agentsop-multi-tenant-rag .cursor/skills/agentsop-multi-tenant-rag && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agentsop-multi-tenant-rag" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-multi-tenant-rag into .cursor/skills/agentsop-multi-tenant-rag/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-multi-tenant-rag", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/agentsope/SkillAlchemy.git --path skills/agentsop-multi-tenant-rag--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-multi-tenant-rag -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-multi-tenant-rag --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/agentsop-multi-tenant-rag .gemini/skills/agentsop-multi-tenant-rag && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agentsop-multi-tenant-rag" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-multi-tenant-rag into .gemini/skills/agentsop-multi-tenant-rag/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-multi-tenant-rag", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install agentsope/SkillAlchemy agentsop-multi-tenant-ragInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add agentsope/SkillAlchemy --skill agentsop-multi-tenant-rag -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/agentsop-multi-tenant-rag .github/skills/agentsop-multi-tenant-rag && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-multi-tenant-rag" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-multi-tenant-rag into .github/skills/agentsop-multi-tenant-rag/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-multi-tenant-rag", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-multi-tenant-rag -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-multi-tenant-rag --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/agentsop-multi-tenant-rag .opencode/skills/agentsop-multi-tenant-rag && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-multi-tenant-rag" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-multi-tenant-rag into .opencode/skills/agentsop-multi-tenant-rag/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-multi-tenant-rag", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agentsop-multi-tenant-ragSecurity-first SOP for multi-tenant RAG systems. An agent skill from agentsope/SkillAlchemy.
Agentsop Multi Tenant RAG is an agent skill from agentsope/SkillAlchemy. Security-first SOP for multi-tenant RAG systems. Activate when a calling agent is building, reviewing, or debugging any retrieval pipeline whose vector store is shared across more than one user, organisation, workspace, customer, or permission scope. Encodes the single non-negotiable rule — filter at the vector store query, never after retrieval / never after rerank — together with the per-vendor query-time filter APIs (Pinecone namespaces + $eq/$in, Weaviate multiTenancyConfig + tenant handle, Qdrant istenant…
Its SKILL.md is about 9.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `intermediate/research_notes.md`, `references/R1-vendor-filter-cheatsheet.md` and `references/R2-cross-tenant-test-recipes.md`).
It sits in AI & LLM Engineering, covering Vector databases, Retrieval-augmented generation and Multi-tenancy. It works with pgvector, Pinecone, LangChain and LlamaIndex. The repository describes itself as: From thought to skill. From signal to structure. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit d0f0355. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
kubectlFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
qdrant.techdocs.pinecone.iodevelopers.llamaindex.aipinecone.iodocs.weaviate.ioweaviate.iodocs.trychroma.comcookbook.chromadb.devpython.langchain.comwe45.comcsoonline.comchristian-schneider.netkiteworks.combeyondscale.techFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agentsop Multi Tenant RAG loads about 9.8k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 210 tokens; SKILL.md has 3,918 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from agentsope/SkillAlchemy at commit d0f0355, republished under its MIT licence (© agentsope). 3,918 words, ~9,783 tokens.
.claude/skills/agentsop-multi-tenant-rag/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Third-person operating model for a coder agent that owns retrieval correctness across tenant boundaries. The audience is the LLM agent writing or reviewing the code — not the end user.
One sentence: Isolation lives at the vector store query boundary, not at the model. Anything that reaches the LLM's context window has already leaked.
Activate this skill whenever any of the following holds:
vector_store.query,
query_points, similarity_search, as_retriever().retrieve(...), raw
pgvector ORDER BY embedding <-> $1) and the corpus serves more than
one tenant, customer, organisation, workspace, user, or permission scope.tenant_id, org_id, user_id, "cross-customer", "shared index",
"knowledge base per team".nodes / documents
list after retrieval.Do not activate when:
Three principles. If a design violates any of them the system is exploitable, regardless of how good the LLM prompt is.
The boundary is the vector store query call. Whatever crosses that boundary is in the trust set. An LLM, a reranker, or a postprocessor that "filters out" foreign-tenant chunks is operating inside the breach: those tokens have already been embedded, retrieved, scored, and exposed to attacker-controlled prompts. CVE-2024-41892 (Pinecone, 2024) is the canonical demonstration — RBAC checks executed after retrieval allowed sentinel data to cross namespace boundaries before the access check fired. (See we45, CSO Online references.)
Operational corollary: any code shaped
results = vs.query(...); results = [r for r in results if r.metadata["tenant"] == ctx.tenant]is a security defect, not a style nit. The leak already happened — you only hid it from the user.
A query-time filter is only enforceable if every vector carries the key. The two failure shapes:
tenant_id becomes a
shared global record returned to all tenants — it matches no filter
predicate that requires equality.tenant_id derived from the document body / LLM extraction
rather than from the request's authenticated session — attacker can craft
document content that re-labels itself.The tenant key must come from the authenticated session at ingestion time and be written into the immutable payload / metadata column. Treat it like a foreign key to your tenants table.
Most production-grade stacks layer two mechanisms:
| Layer | Mechanism | What it stops |
|---|---|---|
| Storage | Per-tenant namespace / shard / collection / RLS policy | Operator bugs, mis-routed queries, ops mistakes |
| Query | filter={tenant_id: $session.tenant} on every call | App bugs, namespace selection mistakes, cross-shard joins |
Pinecone's official guidance is "namespace per tenant in serverless"; Weaviate
mandates a tenant handle on every CRUD when multiTenancyConfig.enabled=true;
Qdrant marks the field as is_tenant=true and requires Filter.must on
read; pgvector pairs schema partitioning with RLS. The principle is the same:
one layer is one bug away from failing open.
Eight stages. Each gates the next. Stop and remediate at the first failure.
Before any code change, write down:
tenant_id,
org_id, workspace_id, user_id, acl_label). Pick one — do not mix.10k → payload partition + index is mandatory). Drives Stage 2 choice.
Artifact: a 5-line THREAT_MODEL.md snippet co-located with the retriever.
| Vendor | Recommended primitive | When |
|---|---|---|
| Pinecone | One namespace per tenant in a serverless index | Default; cheapest query (1 RU per 1 GB / namespace); offboarding = delete_namespace |
| Weaviate | multiTenancyConfig.enabled=True on the collection | Native, per-shard isolation; mandatory tenant handle on CRUD |
| Qdrant | Single collection, payload key with is_tenant=true, keyword index | Scales to millions of tenants; physical co-location, logical isolation |
| Chroma | where={"tenant_id": "..."} on every query/get/update/delete | Smallest stacks; cookbook calls this the "naive multi-tenancy strategy" — accept its limits |
| pgvector | Postgres schema or row-level security (RLS) on tenant_id column | When the rest of the app already lives in Postgres |
| Milvus | partition_key on tenant field | Same shape as Qdrant is_tenant |
Decision rule: prefer the strongest physical primitive your tenant cardinality can afford, then always add the filter at query time anyway (Principle 3).
# Canonical shape — vendor-agnostic
node.metadata = {
"tenant_id": session.tenant_id, # MUST come from authenticated session
"doc_id": doc.id,
"source": doc.uri,
"ingested_at": now_utc(),
# ... domain fields
}
assert node.metadata["tenant_id"], "refuse to embed without tenant_id"Hard rules:
tenant_id is missing or empty — fail loud, do
not default.tenant_id is never taken from document content; never mutable
post-ingestion.$eq silently.create_payload_index(field="tenant_id", schema="keyword", is_tenant=True)).collection.tenants.create([Tenant(name=t)]))
before any write.Every retrieval call must accept a tenant key from the session and pass it into the vendor's filter argument. Templates in OP-04 through OP-08.
# Generic LlamaIndex shape — works across Pinecone, Qdrant, Weaviate, Chroma
from llama_index.core.vector_stores import (
MetadataFilter, MetadataFilters, FilterOperator,
)
filters = MetadataFilters(filters=[
MetadataFilter(key="tenant_id", value=session.tenant_id,
operator=FilterOperator.EQ),
])
retriever = index.as_retriever(similarity_top_k=8, filters=filters)If the framework or vendor SDK does not expose a query-time filter, switch the framework or vendor — do not patch with a post-filter.
Before merging, the PR must include — and CI must run — a test that:
If the unfiltered top-1 is not a foreign-tenant doc, the test is dishonest — rewrite the corpus to make foreign-tenant semantically closer, otherwise the test always passes vacuously.
See OP-09 for a runnable scaffold.
Layer at least one of:
USING (tenant_id = current_setting('app.current_tenant')::uuid)).A single layer is one config bug away from open.
Every retrieval call logs:
{ts, request_id, session.tenant_id, query_hash,
vs.namespace_or_collection, filter_clause,
returned_count, returned_tenant_ids_set}Then add a synchronous assertion in the request path:
assert returned_tenant_ids_set.issubset({session.tenant_id, GLOBAL_TENANT})This converts a silent leak into a loud 500 — the right failure mode.
Walk every component after retrieval and confirm none of them:
cache_key = hash(query) is a leak; correct is cache_key = hash((tenant_id, query))).Reranker / NodePostprocessor / synthesizer can only safely narrow the set; they never restore foreign tenants and never invent context — confirm by reading the code path.
Format: Trigger / Action / Output / Evidence. Vendor-specific filter syntax canonicalised against current docs (May 2026).
tenant_id into metadata.tenant_id (from authenticated session) to every node's
metadata at chunk creation. Validate non-empty before add() /
upsert(). Make the field part of the ingestion schema.is_tenant); >10k → tiered (hot collection +
cold archive). Pair with query filter (Principle 3).tenant_id = request.json["tenant_id"]
or similar untrusted source.tenant_id to the authenticated principal (JWT claim,
session row, signed header). Pass through one chokepoint (Context /
RequestState) — never read user input again downstream.get_tenant_id(ctx) accessor used everywhere; no
string-typed tenant ids floating through function args.filter={"tenant_id": {"$eq": tenant_id}}. Avoid $in lists >10,000 (hard cap).index.query(
namespace=tenant_id, # primary isolation
vector=embedding,
top_k=8,
filter={"tenant_id": {"$eq": tenant_id}}, # defence-in-depth
include_metadata=True,
)delete_namespace(tenant_id) for
offboarding.docs.pinecone.io/guides/index-data/implement-multitenancy,
docs.pinecone.io/troubleshooting/namespaces-vs-metadata-filtering.multiTenancyConfig(enabled=True) on the collection;
create one tenant per customer; pass tenant= on every read/write.coll = client.collections.get("Docs").with_tenant(session.tenant_id)
coll.query.near_vector(vector=emb, limit=8)auto_tenant_activation=True if tenant set is sparse.docs.weaviate.io/weaviate/manage-collections/multi-tenancy;
Rethinking Vector Search at Scale (Weaviate blog).is_tenant=True;
filter on every search.client.create_payload_index(
collection_name="docs",
field_name="tenant_id",
field_schema=models.KeywordIndexParams(
type="keyword", is_tenant=True))
client.query_points(
collection_name="docs",
query=emb, limit=8,
query_filter=models.Filter(must=[
models.FieldCondition(
key="tenant_id",
match=models.MatchValue(value=session.tenant_id))]),
)qdrant.tech/documentation/manage-data/multitenancy/;
qdrant.tech/articles/multitenancy/;
qdrant.tech/documentation/examples/llama-index-multitenancy/.collection.query(
query_embeddings=[emb], n_results=8,
where={"$and": [
{"tenant_id": {"$eq": session.tenant_id}},
{"deleted": {"$eq": False}},
]},
)where= on get, update, delete.docs.trychroma.com/docs/querying-collections/metadata-filtering;
cookbook.chromadb.dev/strategies/multi-tenancy/naive-multi-tenancy/.tenant_id column; enable RLS; create a policy.ALTER TABLE chunks ENABLE ROW LEVEL SECURITY;
CREATE POLICY tenant_isolation ON chunks
USING (tenant_id = current_setting('app.current_tenant')::uuid);
-- per-request, before the ORDER BY embedding <-> $1:
SET LOCAL app.current_tenant = '...';SELECT *.def test_cross_tenant_isolation(retriever_factory):
ingest("alpha", ["alpha-secret about widgets X1, X2"])
ingest("beta", ["beta confidential roadmap for Q3"])
# Query as alpha for a topic beta owns
r_alpha = retriever_factory(tenant="alpha").retrieve("Q3 roadmap")
assert all(n.metadata["tenant_id"] == "alpha" for n in r_alpha)
assert len(r_alpha) >= 0 # zero is fine; foreign is not
# And the reverse
r_beta = retriever_factory(tenant="beta").retrieve("widget X1")
assert all(n.metadata["tenant_id"] == "beta" for n in r_beta)bad = [n for n in nodes
if n.metadata.get("tenant_id") not in {ctx.tenant, GLOBAL_TENANT}]
if bad:
log.critical("cross_tenant_leak", tenant=ctx.tenant, bad=bad)
raise SecurityError("retrieval boundary violated")VectorIndexAutoRetriever with a VectorStoreInfo describing
filterable fields. Pin tenant_id server-side — the auto-retriever
decides other filters, but tenant_id is always injected from the
session.MetadataFilters + AutoRetriever docs;
Principle 2 (tenant from session, not from content).delete_namespace is O(1)).docs.pinecone.io/troubleshooting/namespaces-vs-metadata-filtering.困境: Re-embedding the same public chunk (e.g. shared regulation text) once per tenant doubles cost and complicates updates. But sharing a vector across tenants means it lives in a "global" pool that must be readable by all — opening a path for poisoned global content to reach every tenant.
约束:
tenant_id="GLOBAL" is by definition in scope for
every tenant's filter — it is not isolated, it is intentionally shared.决策步骤:
global namespace / collection with
write-restricted ingestion (only platform admins, signed source) and is
served via a second retrieval call, not by widening the tenant filter.filter={"tenant_id": {"$in": [tenant, "GLOBAL"]}} —
that pattern hides which pool a chunk came from in downstream logs.结果: Two pools, two retrieval calls, one application-layer merge. Cost deduplicated for public content; private content stays in private isolation; provenance is auditable.
可提取的操作: OP-01, OP-04..08 (per-pool), OP-10.
困境: Tenant A has 1M chunks, tenant B has 30 chunks. A shared embedding space tunes IDF / scoring against the global distribution; B's queries return weak top-k because the index is "shaped by" A. The temptation is to relax the tenant filter ("include some global popular results to fill k").
约束:
决策步骤:
结果: A tiered architecture: dedicated collections for the top-N
tenants, shared collection with is_tenant=true for the long tail. Filter
at query is preserved in both tiers.
可提取的操作: OP-02, OP-06, OP-12.
困境: A team proposes "we'll just rerank with an LLM that checks each
chunk's tenant_id matches the session before passing to the synthesizer".
Cheap, generic, frames the filter as redundant.
约束:
决策步骤:
结果: Filter at query (deterministic, gated by Stage 4 test) + runtime assert (deterministic) + reranker (relevance only). LLM-as-judge is not in the security path.
可提取的操作: OP-10; reject any post-filter scheme.
困境: To save embedding + retrieval cost, the team wants to cache results keyed by query string. If two users of the same tenant ask the same question they share. Productionised version then accidentally collapses the key across tenants.
约束:
决策步骤:
hash((tenant_id, query_normalised, filter_clause, index_version)).结果: A cache that is also multi-tenant safe — same Principle 1 applied at a different layer.
可提取的操作: Stage 7; OP-10 still required after cache hit.
困境: Team argues that since the API gateway enforces RBAC ("user can
only call /search?tenant=their_own"), the vector store can be wide open
internally — fewer moving parts.
约束:
kubectl exec.决策步骤:
结果: Gateway RBAC is additional, not load-bearing. The vector store itself enforces tenancy.
可提取的操作: OP-10, OP-08 (RLS where applicable), Stage 5.
| # | Anti-pattern | Why it's wrong | Correct move |
|---|---|---|---|
| A1 | results = vs.query(...); results = [r for r in results if r.metadata["tenant"] == ctx.tenant] | Post-filter; the foreign content was already retrieved, scored, and (next step) injected into LLM context | Filter at query: pass tenant into vendor's filter= / where= / query_filter= |
| A2 | Trusting the LLM with system_prompt += "only use chunks where tenant=X" | LLMs don't enforce; chunks already in context | Never inside the context window |
| A3 | tenant_id derived from request JSON body or query param | Attacker-controlled; trivial IDOR | Bind to authenticated session/JWT claim once at the chokepoint |
| A4 | tenant_id extracted from document content at ingestion | Document body is attacker-controlled (for any user-uploaded doc) | Source from upload-time session, write immutable |
| A5 | Same vector index, no tenant field at all, "we filter by ACL later" | Without the key in metadata, the filter is impossible at query | OP-01: embed at ingest, refuse without key |
| A6 | Rerank LLM as the security boundary | Foreign content already in rerank model's context; prompt-injectable | Rerank narrows only; filter is the boundary (Dilemma 3) |
| A7 | Cache keyed on query alone, not (tenant, query) | Cross-tenant cache hit returns foreign tenant's answer | OP cache-key contract (Dilemma 4) |
| A8 | One shared "global" namespace queried with {"tenant_id": {"$in": [t, "global"]}} and no provenance tracking | Loses audit trail; lets a poisoned global chunk reach all tenants invisibly | Two pools, two queries, explicit provenance (Dilemma 1) |
| A9 | "We have one index per customer so we don't need a filter" | One config typo / one shared client reused across customers and isolation collapses | Defence in depth: namespace + filter both (Principle 3) |
| A10 | No cross-tenant property test in CI | First leak detected by a customer | OP-09 mandatory; runs on every PR |
| A11 | Schema drift — tenant_id is sometimes int, sometimes UUID, sometimes string | $eq silently fails to match across types | Schema-validate metadata at ingest (Pydantic) and assert on read |
| A12 | Verbose logging of full chunk text in shared observability backend | Foreign tenant content readable by ops staff of other tenants | Log hashes / IDs; redact body in shared sinks |
similarity_search(query) without filter= argument and the corpus is multi-tenant.vs.query(... namespace=request.json["ns"]) — namespace from user input.if node.metadata["tenant_id"] != ctx.tenant: continue anywhere after a retrieval call.filter= — tests pass, prod leaks.tenant_id typed as str in one file and int in another — type drift = silent filter miss.OPEN_TO_PUBLIC / __all__ sentinel in metadata used as a fallback when the filter returns empty.How the same "filter at query" surface looks across the common stacks. All verified against current docs (May 2026); URLs in References.
from llama_index.core.vector_stores import (
MetadataFilter, MetadataFilters, FilterOperator,
)
filters = MetadataFilters(
filters=[MetadataFilter(key="tenant_id",
value=ctx.tenant_id,
operator=FilterOperator.EQ)],
# condition=FilterCondition.AND for multi-clause
)
retriever = index.as_retriever(similarity_top_k=8, filters=filters)Supported operators: ==, !=, >, <, >=, <=, in, nin,
text_match. Caveat: the in-memory default vector store does not honour
filters — use Qdrant/Chroma/Pinecone/Weaviate/pgvector for any real
multi-tenant deployment.
docs = vectorstore.similarity_search(
query="Q3 roadmap", k=8,
filter={"tenant_id": ctx.tenant_id}, # simple equality
)
# Chroma-style operators:
docs = vectorstore.similarity_search(
query="Q3 roadmap", k=8,
filter={"$and": [
{"tenant_id": {"$eq": ctx.tenant_id}},
{"deleted": {"$eq": False}},
]},
)Filter syntax is vendor-passed-through — Chroma uses $and/$or/$eq,
Pinecone uses its own dialect, etc. The filter is forwarded to the vendor;
LangChain does not enforce.
index.query(
namespace=ctx.tenant_id, # primary isolation
vector=emb, top_k=8,
filter={"tenant_id": {"$eq": ctx.tenant_id}, # defence-in-depth
"doc_type": {"$in": ["policy", "wiki"]}},
include_metadata=True,
)Operators: $eq, $ne, $gt, $gte, $lt, $lte, $in (≤10000 values),
$nin, $and, $or. Disk-based metadata filtering (2025) lets high-
cardinality filters scale without memory blow-up.
collection = client.collections.get("Docs").with_tenant(ctx.tenant_id)
res = collection.query.near_vector(
near_vector=emb, limit=8,
filters=Filter.by_property("doc_type").equal("policy"),
)multiTenancyConfig.enabled=True makes the tenant= handle mandatory — the
client cannot issue a tenant-less query.
client.query_points(
collection_name="docs",
query=emb, limit=8,
query_filter=models.Filter(must=[
models.FieldCondition(key="tenant_id",
match=models.MatchValue(value=ctx.tenant_id)),
models.FieldCondition(key="doc_type",
match=models.MatchValue(value="policy")),
]),
)With is_tenant=True on the payload index, Qdrant co-locates per-tenant
vectors on shared shards and accelerates the filter.
collection.query(
query_embeddings=[emb], n_results=8,
where={"$and": [
{"tenant_id": {"$eq": ctx.tenant_id}},
{"deleted": {"$eq": False}},
]},
)Same where= on get/update/delete. Operators: $eq, $ne, $gt, $gte, $lt, $lte, $in, $nin, $contains, $not_contains, $and, $or.
-- Per-request:
SET LOCAL app.current_tenant = :tenant_id;
SELECT chunk_id, body, 1 - (embedding <=> :query_vec) AS score
FROM chunks
ORDER BY embedding <=> :query_vec
LIMIT 8;
-- RLS policy filters tenant_id; no WHERE clause needed in app code.The RLS policy (OP-08) makes the filter unforgeable from the application code.
| Stack | Forgot the filter → what happens |
|---|---|
| Pinecone (namespace only) | Wrong namespace = wrong tenant; namespace argument missing = default "" namespace (cross-tenant if you wrote there) |
| Weaviate (MT enabled) | Client raises — no tenant = no operation |
Qdrant (is_tenant only) | Returns all tenants' vectors — silent leak |
| Chroma | Returns all docs — silent leak |
| pgvector + RLS | Returns nothing for unauthenticated tenant context — safe by default |
| LlamaIndex / LangChain wrappers | Forwards whatever you (don't) pass — same as underlying vendor |
This table is the argument for Principle 3: pick a stack where the failure mode is "loud" (Weaviate, pgvector+RLS) or add OP-10 runtime assertion on stacks where the failure mode is "silent" (Qdrant, Chroma, raw Pinecone-without-namespace).
references/R1-vendor-filter-cheatsheet.md — exhaustive vendor filter
syntax with code blocks; safe to paste into PR descriptions.references/R2-cross-tenant-test-recipes.md — five copy-pasteable
property tests (LlamaIndex/LangChain × Pinecone/Qdrant/Chroma).intermediate/research_notes.md — raw research dump (provenance).© agentsope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in skills/agentsop-multi-tenant-rag of agentsope/SkillAlchemy.
Open the folder on GitHubat commit d0f0355
Agentsop Multi Tenant RAG next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agentsop Multi Tenant RAG this skillagentsope/SkillAlchemy | 466 | — | ~9.8k | Automated safety check: Pass | MIT | |
| RAG ArchitectJeffallan/claude-skills | 12k | — | ~2k | Automated safety check: Pass | MIT | |
| RAG Implementationwshobson/agents | 40k | 9 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Hunt RAG Vectorelementalsouls/Claude-BugHunter | 4.8k | — | ~2.6k | Automated safety check: Pass | MIT | |
| Pinecone Vector DatabaseOrchestra-Research/AI-Research-SKILLs | 13k | 5 repos | ~2k | Automated safety check: Pass | MIT | |
| RAG Patternssoftspark/ai-toolkit | 179 | — | ~1.8k | Automated safety check: Pass | Apache-2.0 |
Jeffallan/claude-skills
Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.
wshobson/agents
Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline.
elementalsouls/Claude-BugHunter
Hunt vector-store / embedding-layer weaknesses in RAG pipelines (OWASP LLM08 Vector and Embedding Weaknesses) — persistent corpus poisoning that survives across sessions and users (distinct from…
Orchestra-Research/AI-Research-SKILLs
Shows how to use Pinecone, a managed vector database, for production RAG, semantic search and recommendations: indexes, upserts, queries, filters and namespaces.
softspark/ai-toolkit
RAG: embeddings, chunking, hybrid search (BM25+vector), reranking, CRAG, multi-hop.
neo4j-contrib/neo4j-skills
Build GraphRAG retrieval pipelines on Neo4j using the neo4j-graphrag Python package (v1.22.0+).
agentsope/SkillAlchemy
SOP for terminal-based, git-native AI pair programming with Aider (git work-tree + tree-sitter repo-map + edit-format + human-in-loop REPL).
agentsope/SkillAlchemy
Coder-agent working-file budget discipline: keep the editable working set (files you /add into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only…
agentsope/SkillAlchemy
Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…
agentsope/SkillAlchemy
SOP for building multi-agent systems with CrewAI — role-based collaboration, sequential/hierarchical processes, Flows, memory, delegation.
agentsope/SkillAlchemy
SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable.
agentsope/SkillAlchemy
Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.
Security-first SOP for multi-tenant RAG systems. An agent skill from agentsope/SkillAlchemy. Agentsop Multi Tenant RAG is an agent skill from agentsope/SkillAlchemy. Security-first SOP for multi-tenant RAG systems.
Agentsop Multi Tenant RAG fits situations like: tasks that involve Vector databases; tasks that involve Retrieval-augmented generation; tasks that involve Multi-tenancy.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-multi-tenant-rag -a claude-code`. Or copy the skill folder (skills/agentsop-multi-tenant-rag in agentsope/SkillAlchemy) into .claude/skills/agentsop-multi-tenant-rag in your project. Claude Code loads it when a task matches its description.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-multi-tenant-rag -a codex`. Or copy the skill folder (skills/agentsop-multi-tenant-rag in agentsope/SkillAlchemy) into .agents/skills/agentsop-multi-tenant-rag in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentsope/SkillAlchemy --skill agentsop-multi-tenant-rag -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentsop-multi-tenant-rag, .gemini/skills/agentsop-multi-tenant-rag, .github/skills/agentsop-multi-tenant-rag and .opencode/skills/agentsop-multi-tenant-rag in your project.
Going by SKILL.md and its folder, Agentsop Multi Tenant RAG needs the command-line tools its instructions call (kubectl). Our summary lists: Python 3.
SKILL.md names 14 domains. As links in the text: qdrant.tech, docs.pinecone.io, developers.llamaindex.ai, pinecone.io, docs.weaviate.io, weaviate.io, docs.trychroma.com, cookbook.chromadb.dev, python.langchain.com, we45.com, csoonline.com, christian-schneider.net, kiteworks.com and beyondscale.tech. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agentsop Multi Tenant RAG is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 9.8k tokens (SKILL.md is roughly 39k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Agentsop Multi Tenant RAG: RAG Architect (Jeffallan/claude-skills, 12k stars), RAG Implementation (wshobson/agents, 40k stars), Hunt RAG Vector (elementalsouls/Claude-BugHunter, 4.8k stars) and Pinecone Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
agentsope (a GitHub user) maintains it in agentsope/SkillAlchemy, which has 466 GitHub stars. The repository holds 46 skills in this directory. The repository was last updated on October 9, 2026.
Source: agentsope/SkillAlchemy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.