Official agent skill

Pinecone RAG

by github in github/awesome-copilot

Build production RAG pipelines and persistent agent memory using Pinecone as the vector database backend.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Pinecone RAG

skills CLI
$ npx skills add github/awesome-copilot --skill pinecone-rag -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install github/awesome-copilot pinecone-rag --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/pinecone-rag .claude/skills/pinecone-rag && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pinecone-rag
GitHub stars
40k
Token cost
~2.4k tokens
SKILL.md length
544 words
Files
1
Skills in repo
417
Repo updated
First seen
Licence
Apache-2.0

At a glance

Build production RAG pipelines and persistent agent memory using Pinecone as the vector database backend.

  • Works in 4 steps: Choose your index configuration → Embed and upsert documents → Choose retrieval strategy → …
  • The user mentions Pinecone
  • SKILL.md covers Before you start — ask one…, Step 1 — Choose your index…, Step 2 — Embed and upsert… and Step 3 — Choose retrieval…, plus 5 more sections
  • Needs PINECONE_API_KEY

What it does

Pinecone RAG is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Build production RAG pipelines and persistent agent memory using Pinecone as the vector database backend. ALWAYS USE THIS SKILL when the user mentions Pinecone, wants to index documents for semantic search, build a retrieval-augmented generation system, store agent memory across sessions, implement hybrid search, or connect an LLM to a searchable knowledge base — even if they don't say "Pinecone" explicitly. Also use when the user asks about vector databases for RAG, namespace isolation for multi-tenant agents…

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: pinecone=6.0.0, Python 3.10+

It sits in AI & LLM Engineering, covering Vector databases and Retrieval-augmented generation. It works with Pinecone and pgvector. The repository describes itself as: Community-contributed instructions, agents, skills, and configurations to help you make the most of GitHub Copilot. The licence is Apache-2.0.

When your agent uses it

  • The user mentions Pinecone
  • Wants to index documents for semantic search
  • Build a retrieval-augmented generation system
  • Store agent memory across sessions

Example prompts

  • “t say”
  • “/pinecone-rag”

Requirements

  • Python 3
  • A credential in PINECONE_API_KEY
  • Compatibility (from SKILL.md): pinecone>=6.0.0, Python 3.10+

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Choose your index configuration
  2. Embed and upsert documents
  3. Choose retrieval strategy
  4. Wire it together and test end to end

What it can do on your machine

Read from SKILL.md and the folder at commit 727ff2e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • PINECONE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    pinecone>=6.0.0, Python 3.10+

    From compatibility in the SKILL.md frontmatter.

Context cost

Pinecone RAG loads about 2.4k tokens when it runs. Until then it costs about 183 tokens; SKILL.md has 544 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~183
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from github/awesome-copilot at commit 727ff2e, republished under its Apache-2.0 licence (© github). 544 words, ~2,405 tokens.

Download SKILL.mdSave it as .claude/skills/pinecone-rag/SKILL.md (or your agent's skills folder).
name
pinecone-rag
description
Build production RAG pipelines and persistent agent memory using Pinecone as the vector database backend. ALWAYS USE THIS SKILL when the user mentions Pinecone, wants to index documents for semantic search, build a retrieval-augmented generation system, store agent memory across sessions, implement hybrid search, or connect an LLM to a searchable knowledge base — even if they don't say "Pinecone" explicitly. Also use when the user asks about vector databases for RAG, namespace isolation for multi-tenant agents, embedding pipelines, or scaling a knowledge base beyond what local storage can handle. DO NOT use for local-only vector stores (Chroma, FAISS, pgvector) or pure keyword search with no semantic component.
compatibility
pinecone>=6.0.0, Python 3.10+
license
Apache-2.0

Pinecone RAG Skill

This skill guides you through building a production RAG pipeline or persistent agent memory system using Pinecone. Follow the workflow from start to finish — don't skip steps or jump to code before understanding what the user actually needs.

Before you start — ask one question

Before writing any code, identify which of these two use cases applies:

A — RAG over documents: User wants to index a corpus (PDFs, docs, code, web pages) and retrieve relevant chunks to ground LLM responses.

B — Agent memory: User wants an agent to remember facts, decisions, or context across sessions or across multiple agents sharing a knowledge base.

The setup is similar but the namespace strategy and retrieval patterns differ. If the user hasn't said, ask: "Is this for document retrieval, agent memory, or both?" Then follow the relevant workflow below.


Step 1 — Choose your index configuration

Pick the index type before writing any code. Getting this wrong means re-creating the index later.

Serverless (recommended for most cases)

python
from pinecone import Pinecone, ServerlessSpec

pc = Pinecone(api_key="PINECONE_API_KEY")

if "my-index" not in pc.list_indexes().names():
    pc.create_index(
        name="my-index",
        dimension=1536,        # must match your embedding model exactly
        metric="cosine",
        spec=ServerlessSpec(cloud="aws", region="us-east-1")
    )
index = pc.Index("my-index")

Pod-based (for consistent high-throughput production)

python
from pinecone import PodSpec

pc.create_index(
    name="my-index-prod",
    dimension=1536,
    metric="cosine",
    spec=PodSpec(environment="us-east1-gcp", pod_type="p1.x1")
)

Dimension quick reference — match this exactly to your embedding model:

ModelDimension
text-embedding-3-small1536
text-embedding-3-large3072
voyage-3 / voyage-multimodal-31024
BAAI/bge-large-en-v1.51024
intfloat/multilingual-e5-large (Arabic, Malay, Chinese)1024

Checkpoint: Index exists, dimension matches embedding model, index.describe_index_stats() returns without error.


Step 2 — Embed and upsert documents

Always batch upserts — never upsert one vector at a time.

python
from openai import OpenAI

client = OpenAI()

def embed(texts: list[str]) -> list[list[float]]:
    res = client.embeddings.create(model="text-embedding-3-small", input=texts)
    return [r.embedding for r in res.data]

def upsert_docs(index, docs: list[dict], namespace: str = "default"):
    """docs = [{"id": "...", "text": "...", "metadata": {...}}]"""
    BATCH = 100
    for i in range(0, len(docs), BATCH):
        batch = docs[i:i + BATCH]
        vecs = [
            {
                "id": d["id"],
                "values": emb,
                "metadata": {**d.get("metadata", {}), "text": d["text"]}
            }
            for d, emb in zip(batch, embed([d["text"] for d in batch]))
        ]
        index.upsert(vectors=vecs, namespace=namespace)

Always store the original text in metadata — this avoids a second lookup at retrieval time.

Checkpoint: index.describe_index_stats() shows vector count > 0 in the target namespace.


Step 3 — Choose retrieval strategy

Dense (semantic) search — use for most cases
python
def search(index, query: str, top_k: int = 5, namespace: str = "default",
           filter: dict = None) -> list[dict]:
    [q_emb] = embed([query])
    results = index.query(
        vector=q_emb, top_k=top_k, namespace=namespace,
        include_metadata=True, filter=filter
    )
    return [{"text": m.metadata["text"], "score": m.score, "id": m.id}
            for m in results.matches]
Hybrid search (semantic + BM25 keyword) — use when corpus has exact terminology

Use hybrid when the domain has precise terms that semantic search misses: legal citations, medical codes, product SKUs, API method names.

python
from pinecone_text.sparse import BM25Encoder

bm25 = BM25Encoder().default()
bm25.fit([d["text"] for d in docs])  # fit once on your corpus

def hybrid_search(index, query: str, top_k: int = 5, alpha: float = 0.7):
    """alpha=1.0 is pure dense; alpha=0.0 is pure sparse."""
    dense = [v * alpha for v in embed([query])[0]]
    sparse_raw = bm25.encode_queries(query)
    sparse = {
        "indices": sparse_raw["indices"],
        "values": [v * (1 - alpha) for v in sparse_raw["values"]]
    }
    return index.query(vector=dense, sparse_vector=sparse,
                       top_k=top_k, include_metadata=True).matches
Metadata filtering — use to scope results before semantic ranking
python
# Exact match
results = index.query(vector=emb, filter={"source": {"$eq": "confluence"}})

# Combined filter
results = index.query(vector=emb, filter={
    "$and": [
        {"category": {"$eq": "engineering"}},
        {"language": {"$in": ["en", "ar"]}}
    ]
})

Checkpoint: A test query returns relevant results with scores > 0.7 for clearly matching content.


Step 4A — Full RAG pipeline (document use case)

python
def rag_answer(index, question: str, namespace: str = "default",
               model: str = "gpt-4o-mini") -> str:
    hits = search(index, question, top_k=5, namespace=namespace)
    context = "\n\n".join(h["text"] for h in hits)

    return client.chat.completions.create(
        model=model,
        messages=[
            {
                "role": "system",
                "content": (
                    "Answer using only the provided context. "
                    "If the answer isn't in the context, say so.\n\n"
                    f"Context:\n{context}"
                )
            },
            {"role": "user", "content": question}
        ]
    ).choices[0].message.content

Show full SKILL.md (221 more words)Show less

Step 4B — Agent memory (memory use case)

Use namespaces to isolate each agent's or user's memories completely. Namespace per agent prevents memory bleed across users or sessions.

python
import time, hashlib

def remember(index, agent_id: str, content: str,
             memory_type: str = "fact"):
    """Store a memory for an agent."""
    mem_id = hashlib.md5(
        f"{agent_id}{content}{time.time()}".encode()
    ).hexdigest()
    [emb] = embed([content])
    index.upsert(
        vectors=[{
            "id": mem_id,
            "values": emb,
            "metadata": {
                "text": content,
                "type": memory_type,
                "timestamp": time.time(),
                "agent_id": agent_id
            }
        }],
        namespace=f"agent_{agent_id}"
    )

def recall(index, agent_id: str, query: str,
           top_k: int = 5) -> list[str]:
    """Recall relevant memories for an agent."""
    return [h["text"] for h in
            search(index, query, top_k=top_k,
                   namespace=f"agent_{agent_id}")]

def forget(index, agent_id: str):
    """Wipe all memories for an agent (e.g., on user request)."""
    index.delete(delete_all=True, namespace=f"agent_{agent_id}")

Step 5 — Wire it together and test end to end

Run a quick smoke test before integrating into the larger system:

python
# Smoke test
upsert_docs(index, [
    {"id": "t1", "text": "Pinecone is a vector database for semantic search."},
    {"id": "t2", "text": "RAG combines retrieval with language model generation."},
])

hits = search(index, "What is Pinecone?")
assert hits[0]["score"] > 0.7, f"Expected high similarity, got {hits[0]['score']}"
print("Smoke test passed:", hits[0]["text"])

Checkpoint: Smoke test passes. End-to-end: index → upsert → query → LLM response works without errors.


Common pitfalls — fix these before they become bugs

  • Dimension mismatch: always verify len(embed(["test"])[0]) matches the index dimension before your first upsert.
  • Missing text in metadata: if you don't store "text" in metadata, you'll need a second lookup to get the actual content at query time.
  • Single-vector upserts in a loop: always batch in chunks of 100.
  • No namespace strategy: decide upfront — one namespace per user/agent prevents cross-tenant data leaks that are hard to fix later.
  • Fitting BM25 on a small corpus: BM25 needs a representative corpus to build good term frequencies. Fit on at least a few hundred documents.

When NOT to use this skill

Use a different approach when:

  • The dataset fits in memory and latency doesn't matter → use FAISS or Chroma
  • You're already on PostgreSQL and want to avoid a new service → use pgvector
  • You need sub-5ms p99 latency with no external API calls → local vector store
  • The user explicitly wants a different vector DB (Weaviate, Qdrant, etc.)

© github, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/pinecone-rag of github/awesome-copilot.

Open the folder on GitHubat commit 727ff2e

Compare with similar skills

Pinecone RAG next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pinecone RAG compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pinecone RAG this skillgithub/awesome-copilot40k—~2.4kAutomated safety check: PassApache-2.0
RAG Implementationwshobson/agents40k9 repos~1.1kAutomated safety check: PassMIT
Hunt RAG Vectorelementalsouls/Claude-BugHunter4.8k—~2.6kAutomated safety check: PassMIT
RAG ArchitectFerroxLabs/wayland608—~4.2kAutomated safety check: PassApache-2.0
RAG ArchitectJeffallan/claude-skills12k1 repos~2kAutomated safety check: PassMIT
Using Vector Databasesancoleman/ai-design-components5261 repos~3.5kAutomated safety check: PassMIT

Similar skills

  • RAG Implementation

    wshobson/agents

    Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline.

    40k GitHub starsUsed in 9 repos~1.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Hunt RAG Vector

    elementalsouls/Claude-BugHunter

    Hunt vector-store / embedding-layer weaknesses in RAG pipelines (OWASP LLM08 Vector and Embedding Weaknesses) — persistent corpus poisoning that survives across sessions and users (distinct from…

    4.8k GitHub stars~2.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • RAG Architect

    FerroxLabs/wayland

    RAG system design covering document chunking strategies, embedding model selection, vector database selection (Pinecone, Weaviate, Chroma, pgvector), retrieval strategies (hybrid search…

    608 GitHub stars~4.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • RAG Architect

    Jeffallan/claude-skills

    Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

    12k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • Using Vector Databases

    ancoleman/ai-design-components

    Vector database implementation for AI/ML applications, semantic search, and RAG systems.

    526 GitHub starsUsed in 1 repo~3.5k tokens
    DatabasesAuto-check passed
  • Agentsop Multi Tenant RAG

    agentsope/SkillAlchemy

    Security-first SOP for multi-tenant RAG systems. An agent skill from agentsope/SkillAlchemy.

    457 GitHub stars~9.8k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from github/awesome-copilot

All 417 skills in this repo
  • Acquire Codebase Knowledge

    github/awesome-copilot

    Official

    Maps an unfamiliar codebase into seven evidence-backed documents in docs/codebase/, using a scan script and templates, for onboarding or architecture write-ups.

    40k GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Azure Architecture Autopilot

    github/awesome-copilot

    Official

    Designs Azure infrastructure from a natural-language description, or diagrams an existing resource group, then refines the design through conversation and deploys it with Bicep.

    40k GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Draw.io Diagram Generator

    github/awesome-copilot

    Official

    Generates, edits and validates draw.io files with correct mxGraph XML, covering flowcharts, architecture, sequence, ER and UML class diagrams.

    40k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Credit Risk Data Cleaning

    github/awesome-copilot

    Official

    Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.

    40k GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Daily Focus Board

    github/awesome-copilot

    Official

    Builds a warm, browser-based daily focus board the user updates by talking to their agent, with Eisenhower priorities, a brain-dump box and kind not-today carryover.

    40k GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Python Pypi Package Builder

    github/awesome-copilot

    Official

    End-to-end skill for building, testing, linting, versioning, and publishing a production-grade Python library to PyPI.

    40k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Questions about Pinecone RAG

What does Pinecone RAG do?

Build production RAG pipelines and persistent agent memory using Pinecone as the vector database backend. Pinecone RAG is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Build production RAG pipelines and persistent agent memory using Pinecone as the vector database backend.

When should I use Pinecone RAG?

Pinecone RAG fits situations like: the user mentions Pinecone; wants to index documents for semantic search; build a retrieval-augmented generation system; store agent memory across sessions.

How do I install Pinecone RAG in Claude Code?

Run `npx skills add github/awesome-copilot --skill pinecone-rag -a claude-code`. Or copy the skill folder (skills/pinecone-rag in github/awesome-copilot) into .claude/skills/pinecone-rag in your project. Claude Code loads it when a task matches its description.

How do I install Pinecone RAG in Codex?

Run `npx skills add github/awesome-copilot --skill pinecone-rag -a codex`. Or copy the skill folder (skills/pinecone-rag in github/awesome-copilot) into .agents/skills/pinecone-rag in your project. Codex loads it when a task matches its description.

Can I use Pinecone RAG in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add github/awesome-copilot --skill pinecone-rag -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pinecone-rag, .gemini/skills/pinecone-rag, .github/skills/pinecone-rag and .opencode/skills/pinecone-rag in your project.

What does Pinecone RAG need to run?

Going by SKILL.md and its folder, Pinecone RAG needs credentials named PINECONE_API_KEY. Our summary lists: Python 3; A credential in PINECONE_API_KEY. Compatibility (from SKILL.md): pinecone>=6.0.0, Python 3.10+.

Does Pinecone RAG access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Pinecone RAG safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Pinecone RAG use?

Pinecone RAG is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pinecone RAG use?

About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Pinecone RAG?

Skills that share tags, products or a category with Pinecone RAG: RAG Implementation (wshobson/agents, 40k stars), Hunt RAG Vector (elementalsouls/Claude-BugHunter, 4.8k stars), RAG Architect (FerroxLabs/wayland, 608 stars) and RAG Architect (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pinecone RAG?

github (a GitHub organization, an official publisher) maintains it in github/awesome-copilot, which has 39,748 GitHub stars. The repository holds 417 skills in this directory. The repository was last updated on October 7, 2026.

Source: github/awesome-copilot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.