Pgvector Semantic Search
timescale/pg-aiguide
A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.
Comprehensive toolkit for developing with the CocoIndex library.
$ npx skills add davila7/claude-code-templates --skill cocoindex -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install davila7/claude-code-templates cocoindex --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .claude/skills && cp -r skills-src/cli-tool/components/skills/development/cocoindex .claude/skills/cocoindex && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "cocoindex" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/development/cocoindex into .claude/skills/cocoindex/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cocoindex", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/development/cocoindexType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add davila7/claude-code-templates --skill cocoindex -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install davila7/claude-code-templates cocoindex --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .agents/skills && cp -r skills-src/cli-tool/components/skills/development/cocoindex .agents/skills/cocoindex && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "cocoindex" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/development/cocoindex into .agents/skills/cocoindex/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cocoindex", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add davila7/claude-code-templates --skill cocoindex -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install davila7/claude-code-templates cocoindex --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/cli-tool/components/skills/development/cocoindex .cursor/skills/cocoindex && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "cocoindex" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/development/cocoindex into .cursor/skills/cocoindex/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cocoindex", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/davila7/claude-code-templates.git --path cli-tool/components/skills/development/cocoindex--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add davila7/claude-code-templates --skill cocoindex -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install davila7/claude-code-templates cocoindex --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/cli-tool/components/skills/development/cocoindex .gemini/skills/cocoindex && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "cocoindex" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/development/cocoindex into .gemini/skills/cocoindex/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cocoindex", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install davila7/claude-code-templates cocoindexInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add davila7/claude-code-templates --skill cocoindex -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .github/skills && cp -r skills-src/cli-tool/components/skills/development/cocoindex .github/skills/cocoindex && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "cocoindex" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/development/cocoindex into .github/skills/cocoindex/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cocoindex", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add davila7/claude-code-templates --skill cocoindex -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install davila7/claude-code-templates cocoindex --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/cli-tool/components/skills/development/cocoindex .opencode/skills/cocoindex && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "cocoindex" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/development/cocoindex into .opencode/skills/cocoindex/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cocoindex", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
cocoindexComprehensive toolkit for developing with the CocoIndex library.
Cocoindex is an agent skill from davila7/claude-code-templates. Comprehensive toolkit for developing with the CocoIndex library. Use when users need to create data transformation pipelines (flows), write custom functions, or operate flows via CLI or API. Covers building ETL workflows for AI data processing, including embedding documents into vector databases, building knowledge graphs, creating search indexes, or processing data streams with incremental updates.
Its SKILL.md is about 6.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/api_operations.md`, `references/cli_operations.md` and `references/custom_functions.md`).
It sits in AI & LLM Engineering, covering Vector databases, Embeddings and Data cleaning. It works with PostgreSQL. The repository describes itself as: CLI tool for configuring and monitoring Claude Code. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 14680ec. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
psqlFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
cocoindex.ioAlso links to:
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPENAI_API_KEYANTHROPIC_API_KEYGOOGLE_API_KEYVOYAGE_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Cocoindex loads about 6.4k tokens when it runs, and up to ~19k if it reads all its reference files. Until then it costs about 103 tokens; SKILL.md has 1,636 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
**Guide user to create `.env` file:**- Check `.env` has `COCOINDEX_DATABASE_URL`- Set global limits in `.env`: `COCOINDEX_SOURCE_MAX_INFLIGHT_ROWS`Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from davila7/claude-code-templates at commit 14680ec, republished under its MIT licence (© davila7). 1,636 words, ~6,410 tokens.
.claude/skills/cocoindex/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.CocoIndex is an ultra-performant real-time data transformation framework for AI with incremental processing. This skill enables building indexing flows that extract data from sources, apply transformations (chunking, embedding, LLM extraction), and export to targets (vector databases, graph databases, relational databases).
Core capabilities:
Key features:
For detailed documentation: https://cocoindex.io/docs/ Search documentation: https://cocoindex.io/docs/search?q=url%20encoded%20keyword
Use when users request:
Ask clarifying questions to understand:
Data source:
Transformations:
Target:
Guide user to add CocoIndex with appropriate extras to their project based on their needs:
Required dependency:
cocoindex - Core functionality, CLI, and most built-in functionsOptional extras (add as needed):
cocoindex[embeddings] - For SentenceTransformer embeddings (when using SentenceTransformerEmbed)cocoindex[colpali] - For ColPali image/document embeddings (when using ColPaliEmbedImage or ColPaliEmbedQuery)cocoindex[lancedb] - For LanceDB target (when exporting to LanceDB)cocoindex[embeddings,lancedb] - Multiple extras can be combinedWhat's included:
embeddings extra: SentenceTransformers library for local embedding modelscolpali extra: ColPali engine for multimodal document/image embeddingslancedb extra: LanceDB client library for LanceDB vector database supportUsers can install using their preferred package manager (pip, uv, poetry, etc.) or add to pyproject.toml.
For installation details: https://cocoindex.io/docs/getting_started/installation
Check existing environment first:
Check if COCOINDEX_DATABASE_URL exists in environment variables
postgres://cocoindex:cocoindex@localhost/cocoindexFor flows requiring LLM APIs (embeddings, extraction):
Guide user to create .env file:
# Database connection (required - internal storage)
COCOINDEX_DATABASE_URL=postgres://cocoindex:cocoindex@localhost/cocoindex
# LLM API keys (add the ones you need)
OPENAI_API_KEY=sk-... # For OpenAI (generation + embeddings)
ANTHROPIC_API_KEY=sk-ant-... # For Anthropic (generation only)
GOOGLE_API_KEY=... # For Gemini (generation + embeddings)
VOYAGE_API_KEY=pa-... # For Voyage (embeddings only)
# Ollama requires no API key (local)For more LLM options: https://cocoindex.io/docs/ai/llm
Create basic project structure:
# main.py
from dotenv import load_dotenv
import cocoindex
@cocoindex.flow_def(name="FlowName")
def my_flow(flow_builder: cocoindex.FlowBuilder, data_scope: cocoindex.DataScope):
# Flow definition here
pass
if __name__ == "__main__":
load_dotenv()
cocoindex.init()
my_flow.update()Follow this structure:
@cocoindex.flow_def(name="DescriptiveName")
def flow_name(flow_builder: cocoindex.FlowBuilder, data_scope: cocoindex.DataScope):
# 1. Import source data
data_scope["source_name"] = flow_builder.add_source(
cocoindex.sources.SourceType(...)
)
# 2. Create collector(s) for outputs
collector = data_scope.add_collector()
# 3. Transform data (iterate through rows)
with data_scope["source_name"].row() as item:
# Apply transformations
item["new_field"] = item["existing_field"].transform(
cocoindex.functions.FunctionName(...)
)
...
# Nested iteration (e.g., chunks within documents)
with item["nested_table"].row() as nested_item:
# More transformations
nested_item["embedding"] = nested_item["text"].transform(...)
# Collect data for export
collector.collect(
field1=nested_item["field1"],
field2=item["field2"],
generated_id=cocoindex.GeneratedField.UUID
)
# 4. Export to target
collector.export(
"target_name",
cocoindex.targets.TargetType(...),
primary_key_fields=["field1"],
vector_indexes=[...] # If needed
)Key principles:
.row() to iterate through table dataitem["new_field"] = item["existing_field"].transform(...), NOT local variables like new_field = item["existing_field"].transform(...)Common mistakes to avoid:
❌ Wrong: Using local variables for transformations
with data_scope["files"].row() as file:
summary = file["content"].transform(...) # ❌ Local variable
summaries_collector.collect(filename=file["filename"], summary=summary)✅ Correct: Assigning to row fields
with data_scope["files"].row() as file:
file["summary"] = file["content"].transform(...) # ✅ Field assignment
summaries_collector.collect(filename=file["filename"], summary=file["summary"])❌ Wrong: Creating unnecessary dataclasses to mirror flow fields
from dataclasses import dataclass
@dataclass
class FileSummary: # ❌ Unnecessary - CocoIndex manages fields automatically
filename: str
summary: str
embedding: list[float]
# This dataclass is never used in the flow!IMPORTANT: The patterns listed below are common starting points, but you cannot exhaustively enumerate all possible scenarios. When user requirements don't match existing patterns:
Common starting patterns (use references for detailed examples):
For text embedding: Load references/flow_patterns.md and refer to "Pattern 1: Simple Text Embedding"
For code embedding: Load references/flow_patterns.md and refer to "Pattern 2: Code Embedding with Language Detection"
For LLM extraction + knowledge graph: Load references/flow_patterns.md and refer to "Pattern 3: LLM-based Extraction to Knowledge Graph"
For live updates: Load references/flow_patterns.md and refer to "Pattern 4: Live Updates with Refresh Interval"
For custom functions: Load references/flow_patterns.md and refer to "Pattern 5: Custom Transform Function"
For reusable query logic: Load references/flow_patterns.md and refer to "Pattern 6: Transform Flow for Reusable Logic"
For concurrency control: Load references/flow_patterns.md and refer to "Pattern 7: Concurrency Control"
Example of pattern composition:
If a user asks to "index images from S3, generate captions with a vision API, and store in Qdrant", combine:
No single pattern covers this exact scenario, but the building blocks are composable.
Guide user through testing:
# 1. Run with setup
cocoindex update --setup -f main # -f force setup without confirmation prompts
# 2. Start a server and redirect users to CocoInsight
cocoindex server -ci main
# Then open CocoInsight at https://cocoindex.io/cocoinsight
CocoIndex has a type system independent of programming languages. All data types are determined at flow definition time, making schemas clear and predictable.
IMPORTANT: When to define types:
Type annotation requirements:
Any, dict[str, Any], or omit annotations; engine already knows the typesWhy specific return types matter: Custom function return types let CocoIndex infer field types throughout the flow without processing real data. This enables creating proper target schemas (e.g., vector indexes with fixed dimensions).
Common type categories:
Primitive types: str, int, float, bool, bytes, datetime.date, datetime.datetime, uuid.UUID
Vector types (embeddings): Specify dimension in return type if you plan to export as vectors to targets, as most targets require a fixed vector dimension
cocoindex.Vector[cocoindex.Float32, typing.Literal[768]] - 768-dim float32 vector (recommended)list[float] without dimension also worksStruct types: Dataclass, NamedTuple, or Pydantic model
Person)dict[str, Any] or AnyTable types:
dict[K, V] where K = key type (primitive or frozen struct), V = Struct typelist[R] where R = Struct typedict[Any, Any] or list[Any]Json type: cocoindex.Json for unstructured/dynamic data
Optional types: T | None for nullable values
Examples:
from dataclasses import dataclass
from typing import Literal
import cocoindex
@dataclass
class Person:
name: str
age: int
# ✅ Vector with dimension (recommended for vector search)
@cocoindex.op.function(behavior_version=1)
def embed_text(text: str) -> cocoindex.Vector[cocoindex.Float32, Literal[768]]:
"""Generate 768-dim embedding - dimension needed for vector index."""
# ... embedding logic ...
return embedding # numpy array or list of 768 floats
# ✅ Struct return type, relaxed argument
@cocoindex.op.function(behavior_version=1)
def process_person(person: dict[str, Any]) -> Person:
"""Argument can be dict[str, Any], return must be specific Struct."""
return Person(name=person["name"], age=person["age"])
# ✅ LTable return type
@cocoindex.op.function(behavior_version=1)
def filter_people(people: list[Any]) -> list[Person]:
"""Return type specifies list of specific Struct."""
return [p for p in people if p.age >= 18]
# ❌ Wrong: dict[str, str] is not a valid specific CocoIndex type
# @cocoindex.op.function(...)
# def bad_example(person: Person) -> dict[str, str]:
# return {"name": person.name}For comprehensive data types documentation: https://cocoindex.io/docs/core/data_types
When users need custom transformation logic, create custom functions.
Use standalone function when:
Use spec+executor when:
@cocoindex.op.function(behavior_version=1)
def my_function(input_arg: str, optional_arg: int | None = None) -> dict:
"""
Function description.
Args:
input_arg: Description
optional_arg: Optional description
"""
# Transformation logic
return {"result": f"processed-{input_arg}"}Requirements:
@cocoindex.op.function()cache=True for expensive ops, behavior_version (required with cache)# 1. Define configuration spec
class MyFunction(cocoindex.op.FunctionSpec):
"""Configuration for MyFunction."""
model_name: str
threshold: float = 0.5
# 2. Define executor
@cocoindex.op.executor_class(cache=True, behavior_version=1)
class MyFunctionExecutor:
spec: MyFunction # Required: link to spec
model = None # Instance variables for state
def prepare(self) -> None:
"""Optional: run once before execution."""
# Load model, setup connections, etc.
self.model = load_model(self.spec.model_name)
def __call__(self, text: str) -> dict:
"""Required: execute for each data row."""
# Use self.spec for configuration
# Use self.model for loaded resources
result = self.model.process(text)
return {"result": result}When to enable cache:
Important: Increment behavior_version when function logic changes to invalidate cache.
For detailed examples and patterns, load references/custom_functions.md.
For more on custom functions: https://cocoindex.io/docs/custom_ops/custom_functions
Setup flow (create resources):
cocoindex setup mainOne-time update:
cocoindex update main
# With auto-setup
cocoindex update --setup main
# Force reset everything before setup and update
cocoindex update --reset mainLive update (continuous monitoring):
cocoindex update main.py -L
# Requires refresh_interval on source or source-specific change captureDrop flow (remove all resources):
cocoindex drop main.pyInspect flow:
cocoindex show main.py:FlowNameTest without side effects:
cocoindex evaluate main.py:FlowName --output-dir ./test_outputFor complete CLI reference, load references/cli_operations.md.
For CLI documentation: https://cocoindex.io/docs/core/cli
Basic setup:
from dotenv import load_dotenv
import cocoindex
load_dotenv()
cocoindex.init()
@cocoindex.flow_def(name="MyFlow")
def my_flow(flow_builder, data_scope):
# ... flow definition ...
passOne-time update:
stats = my_flow.update()
print(f"Processed {stats.total_rows} rows")
# Async
stats = await my_flow.update_async()Live update:
# As context manager
with cocoindex.FlowLiveUpdater(my_flow) as updater:
# Updater runs in background
# Your application logic here
pass
# Manual control
updater = cocoindex.FlowLiveUpdater(
my_flow,
cocoindex.FlowLiveUpdaterOptions(
live_mode=True,
print_stats=True
)
)
updater.start()
# ... application logic ...
updater.wait()Setup/drop:
my_flow.setup(report_to_stdout=True)
my_flow.drop(report_to_stdout=True)
cocoindex.setup_all_flows()
cocoindex.drop_all_flows()Query with transform flows:
@cocoindex.transform_flow()
def text_to_embedding(text: cocoindex.DataSlice[str]) -> cocoindex.DataSlice[list[float]]:
return text.transform(
cocoindex.functions.SentenceTransformerEmbed(model="...")
)
# Use in flow for indexing
doc["embedding"] = text_to_embedding(doc["content"])
# Use for querying
query_embedding = text_to_embedding.eval("search query")For complete API reference and patterns, load references/api_operations.md.
For API documentation: https://cocoindex.io/docs/core/flow_methods
SplitRecursively - Chunk text intelligently
doc["chunks"] = doc["content"].transform(
cocoindex.functions.SplitRecursively(),
language="markdown", # or "python", "javascript", etc.
chunk_size=2000,
chunk_overlap=500
)ParseJson - Parse JSON strings
data = json_string.transform(cocoindex.functions.ParseJson())DetectProgrammingLanguage - Detect language from filename
file["language"] = file["filename"].transform(
cocoindex.functions.DetectProgrammingLanguage()
)SentenceTransformerEmbed - Local embedding model
# Requires: cocoindex[embeddings]
chunk["embedding"] = chunk["text"].transform(
cocoindex.functions.SentenceTransformerEmbed(
model="sentence-transformers/all-MiniLM-L6-v2"
)
)EmbedText - LLM API embeddings
This is the recommended way to generate embeddings using LLM APIs (OpenAI, Voyage, etc.).
chunk["embedding"] = chunk["text"].transform(
cocoindex.functions.EmbedText(
api_type=cocoindex.LlmApiType.OPENAI,
model="text-embedding-3-small",
)
)ColPaliEmbedImage - Multimodal image embeddings
# Requires: cocoindex[colpali]
image["embedding"] = image["img_bytes"].transform(
cocoindex.functions.ColPaliEmbedImage(model="vidore/colpali-v1.2")
)ExtractByLlm - Extract structured data with LLM
This is the recommended way to use LLMs for extraction and summarization tasks. It supports both structured outputs (dataclasses, Pydantic models) and simple text outputs (str).
import dataclasses
# For structured extraction
@dataclasses.dataclass
class ProductInfo:
name: str
price: float
category: str
item["product_info"] = item["text"].transform(
cocoindex.functions.ExtractByLlm(
llm_spec=cocoindex.LlmSpec(
api_type=cocoindex.LlmApiType.OPENAI,
model="gpt-4o-mini"
),
output_type=ProductInfo,
instruction="Extract product information"
)
)
# For text summarization/generation
file["summary"] = file["content"].transform(
cocoindex.functions.ExtractByLlm(
llm_spec=cocoindex.LlmSpec(
api_type=cocoindex.LlmApiType.OPENAI,
model="gpt-4o-mini"
),
output_type=str,
instruction="Summarize this document in one paragraph"
)
)Browse all sources: https://cocoindex.io/docs/sources/ Browse all targets: https://cocoindex.io/docs/targets/
LocalFile:
cocoindex.sources.LocalFile(
path="documents",
included_patterns=["*.md", "*.txt"],
excluded_patterns=["**/.*", "node_modules"]
)AmazonS3:
cocoindex.sources.AmazonS3(
bucket="my-bucket",
prefix="documents/",
aws_access_key_id=cocoindex.add_transient_auth_entry("..."),
aws_secret_access_key=cocoindex.add_transient_auth_entry("...")
)Postgres:
cocoindex.sources.Postgres(
connection=cocoindex.add_auth_entry("conn", cocoindex.sources.PostgresConnection(...)),
query="SELECT id, content FROM documents"
)Postgres (with vector support):
collector.export(
"target_name",
cocoindex.targets.Postgres(),
primary_key_fields=["id"],
vector_indexes=[
cocoindex.VectorIndexDef(
field_name="embedding",
metric=cocoindex.VectorSimilarityMetric.COSINE_SIMILARITY
)
]
)Qdrant:
collector.export(
"target_name",
cocoindex.targets.Qdrant(collection_name="my_collection"),
primary_key_fields=["id"]
)LanceDB:
# Requires: cocoindex[lancedb]
collector.export(
"target_name",
cocoindex.targets.LanceDB(uri="lancedb_data", table_name="my_table"),
primary_key_fields=["id"]
)Neo4j (nodes):
collector.export(
"nodes",
cocoindex.targets.Neo4j(
connection=neo4j_conn,
mapping=cocoindex.targets.Nodes(label="Entity")
),
primary_key_fields=["id"]
)Neo4j (relationships):
collector.export(
"relationships",
cocoindex.targets.Neo4j(
connection=neo4j_conn,
mapping=cocoindex.targets.Relationships(
rel_type="RELATES_TO",
source=cocoindex.targets.NodeFromFields(
label="Entity",
fields=[cocoindex.targets.TargetFieldMapping(source="source_id", target="id")]
),
target=cocoindex.targets.NodeFromFields(
label="Entity",
fields=[cocoindex.targets.TargetFieldMapping(source="target_id", target="id")]
)
)
),
primary_key_fields=["id"]
)cocoindex show main.py--app-dir if not in project root.env has COCOINDEX_DATABASE_URLpsql $COCOINDEX_DATABASE_URL--env-file to specify custom locationcocoindex setup main.pycocoindex drop main.py && cocoindex setup main.pyrefresh_interval to sourcemax_inflight_rows, max_inflight_bytes.env: COCOINDEX_SOURCE_MAX_INFLIGHT_ROWSThis skill includes comprehensive reference documentation for common patterns and operations:
Load these references when users need:
For comprehensive documentation: https://cocoindex.io/docs/ Search specific topics: https://cocoindex.io/docs/search?q=url%20encoded%20keyword
© davila7, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in cli-tool/components/skills/development/cocoindex of davila7/claude-code-templates.
Open the folder on GitHubat commit 14680ec
We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in davila7/claude-code-templates, which our catalogue first saw on October 7, 2026.
Cocoindex next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Cocoindex this skilldavila7/claude-code-templates | 32k | 2 repos | ~6.4k | Automated safety check: Notes | MIT | |
| Pgvector Semantic Searchtimescale/pg-aiguide | 1.9k | 1 repos | ~3.8k | Automated safety check: Pass | Apache-2.0 | |
| Vector DBericrisco/rsc-harness | 167 | — | ~2.8k | Automated safety check: Pass | MIT | |
| Cognee Integrations Setuptopoteretes/cognee | 32k | — | ~1k | Automated safety check: Notes | Apache-2.0 | |
| I3brycewang-stanford/Auto-Empirical-Research-Skills | 4.5k | — | ~1.8k | Automated safety check: Pass | Custom licence | |
| Neo4j Vector Index Skillneo4j-contrib/neo4j-skills | 114 | — | ~5.6k | Automated safety check: Notes | MIT |
timescale/pg-aiguide
A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.
ericrisco/rsc-harness
A skill your agent uses when operating a vector store as a data layer — choosing or migrating between Pinecone, Qdrant, Weaviate and pgvector; designing a collection or index (distance metric…
topoteretes/cognee
Switches cognee's LLM, embedding, relational, vector and graph backends through environment variables, with the extras to install and the traps to avoid.
brycewang-stanford/Auto-Empirical-Research-Skills
RAG Builder with Parallel Document Processing Vector database construction with local embeddings (zero cost) Handles PDF download, text extraction, chunking, and vector database creation Absorbed B5…
neo4j-contrib/neo4j-skills
Create and manage Neo4j vector indexes, run vector similarity search (ANN/kNN), store embeddings on nodes or relationships, use SEARCH clause (Neo4j 2026.01+, preferred) or…
NVIDIA/skills
A skill your agent uses when operating PAIDF Curation and Retrieval or NVIDIA Cosmos Curator pipelines (split, filter, caption, embed, dedup, shard, image annotate) or PAIDF Data Mining…
davila7/claude-code-templates
Runs web-grounded searches through Perplexity's Sonar models over OpenRouter for current events, recent literature and cited facts beyond the model's training cutoff.
davila7/claude-code-templates
Analyzes Neuropixels recordings from SpikeGLX or Open Ephys through preprocessing, drift correction, Kilosort4 spike sorting, quality metrics and curation.
davila7/claude-code-templates
Supplies LaTeX templates and formatting rules for journals, conferences, posters, and grant proposals, then can check a draft against them.
davila7/claude-code-templates
Analyzes a brand's existing writing to lock in a consistent voice, then builds SEO blog posts and platform-specific social content around it.
davila7/claude-code-templates
Guides corrective and preventive action (CAPA) work in a quality management system, from initiation and root cause analysis through effectiveness verification.
davila7/claude-code-templates
Senior FDA consultant and specialist for medical device companies including HIPAA compliance and requirement management.
Works with
Comprehensive toolkit for developing with the CocoIndex library. Cocoindex is an agent skill from davila7/claude-code-templates. Comprehensive toolkit for developing with the CocoIndex library.
Cocoindex fits situations like: users need to create data transformation pipelines (flows); write custom functions; operate flows via CLI.
Run `npx skills add davila7/claude-code-templates --skill cocoindex -a claude-code`. Or copy the skill folder (cli-tool/components/skills/development/cocoindex in davila7/claude-code-templates) into .claude/skills/cocoindex in your project. Claude Code loads it when a task matches its description.
Run `npx skills add davila7/claude-code-templates --skill cocoindex -a codex`. Or copy the skill folder (cli-tool/components/skills/development/cocoindex in davila7/claude-code-templates) into .agents/skills/cocoindex in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davila7/claude-code-templates --skill cocoindex -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cocoindex, .gemini/skills/cocoindex, .github/skills/cocoindex and .opencode/skills/cocoindex in your project.
Going by SKILL.md and its folder, Cocoindex needs the command-line tools its instructions call (psql) and credentials named OPENAI_API_KEY, ANTHROPIC_API_KEY, GOOGLE_API_KEY and VOYAGE_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY; A credential in ANTHROPIC_API_KEY.
SKILL.md names 2 domains. In commands or code: cocoindex.io; the agent is likely to contact it when it follows the instructions. As links in the text: github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Cocoindex is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.4k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 12k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Cocoindex: Pgvector Semantic Search (timescale/pg-aiguide, 1.9k stars), Vector DB (ericrisco/rsc-harness, 167 stars), Cognee Integrations Setup (topoteretes/cognee, 32k stars) and I3 (brycewang-stanford/Auto-Empirical-Research-Skills, 4.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
davila7 (a GitHub user) maintains it in davila7/claude-code-templates, which has 32,463 GitHub stars. The repository holds 477 skills in this directory. The repository was last updated on October 8, 2026.
Source: davila7/claude-code-templates on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.