Agent skill

Embedding Optimization

by ancoleman in ancoleman/ai-design-components

Optimizing vector embeddings for RAG systems through model selection, chunking strategies, caching, and performance tuning.

MITAuto-check passedAI & LLM Engineering

Install Embedding Optimization

skills CLI
$ npx skills add ancoleman/ai-design-components --skill embedding-optimization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ancoleman/ai-design-components embedding-optimization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/embedding-optimization .claude/skills/embedding-optimization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
embedding-optimization
GitHub stars
526
Used in
1 other repo
Token cost
~2.1k tokens
SKILL.md length
746 words
Files
11 (incl. references)
Skills in repo
75
Repo updated
First seen
Licence
MIT

At a glance

Optimizing vector embeddings for RAG systems through model selection, chunking strategies, caching, and performance tuning.

  • Building semantic search
  • SKILL.md covers When to Use This Skill, Model Selection Framework, Chunking Strategies and Caching Implementation, plus 7 more sections
  • Runs Python scripts from its folder; calls python
  • Document retrieval systems that require cost-effective

What it does

Embedding Optimization is an agent skill from ancoleman/ai-design-components. Optimizing vector embeddings for RAG systems through model selection, chunking strategies, caching, and performance tuning. Use when building semantic search, RAG pipelines, or document retrieval systems that require cost-effective, high-quality embeddings.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including reference files (for example `examples/batch_processor.py`, `examples/benchmark_embeddings.py` and `examples/local_embedder.py`).

It sits in AI & LLM Engineering, covering Embeddings and Retrieval-augmented generation. The repository describes itself as: Comprehensive UI/UX and Backend component design skills for AI-assisted development with Claude. The licence is MIT.

When your agent uses it

  • Building semantic search
  • Document retrieval systems that require cost-effective
  • High-quality embeddings

Example prompts

  • “/embedding-optimization”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 76551b7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Embedding Optimization loads about 2.1k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 70 tokens; SKILL.md has 746 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~70
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~15k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ancoleman/ai-design-components at commit 76551b7, republished under its MIT licence (© ancoleman). 746 words, ~2,072 tokens.

Download SKILL.mdSave it as .claude/skills/embedding-optimization/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
embedding-optimization
description
Optimizing vector embeddings for RAG systems through model selection, chunking strategies, caching, and performance tuning. Use when building semantic search, RAG pipelines, or document retrieval systems that require cost-effective, high-quality embeddings.

Embedding Optimization

Optimize embedding generation for cost, performance, and quality in RAG and semantic search systems.

When to Use This Skill

Trigger this skill when:

  • Building RAG (Retrieval Augmented Generation) systems
  • Implementing semantic search or similarity detection
  • Optimizing embedding API costs (reducing by 70-90%)
  • Improving document retrieval quality through better chunking
  • Processing large document corpora (thousands to millions of documents)
  • Selecting between API-based vs. local embedding models

Model Selection Framework

Choose the optimal embedding model based on requirements:

Quick Recommendations:

  • Startup/MVP: all-MiniLM-L6-v2 (local, 384 dims, zero API costs)
  • Production: text-embedding-3-small (API, 1,536 dims, balanced quality/cost)
  • High Quality: text-embedding-3-large (API, 3,072 dims, premium)
  • Multilingual: multilingual-e5-base (local, 768 dims) or Cohere embed-multilingual-v3.0

For detailed decision frameworks including cost comparisons, quality benchmarks, and data privacy considerations, see references/model-selection-guide.md.

Model Comparison Summary:

ModelTypeDimensionsCost per 1M tokensBest For
all-MiniLM-L6-v2Local384$0 (compute only)High volume, tight budgets
BGE-base-en-v1.5Local768$0 (compute only)Quality + cost balance
text-embedding-3-smallAPI1,536$0.02General purpose production
text-embedding-3-largeAPI3,072$0.13Premium quality requirements
embed-multilingual-v3.0API1,024$0.10100+ language support

Chunking Strategies

Select chunking strategy based on content type and use case:

Content Type → Strategy Mapping:

  • Documentation: Recursive (heading-aware), 800 chars, 100 overlap
  • Code: Recursive (function-level), 1,000 chars, 100 overlap
  • Q&A/FAQ: Fixed-size, 500 chars, 50 overlap (precise retrieval)
  • Legal/Technical: Semantic (large), 1,500 chars, 200 overlap (context preservation)
  • Blog Posts: Semantic (paragraph), 1,000 chars, 100 overlap
  • Academic Papers: Recursive (section-aware), 1,200 chars, 150 overlap

For detailed chunking patterns, decision trees, and implementation guidance, see references/chunking-strategies.md.

Quick Start with CLI:

bash
python scripts/chunk_document.py \
  --input document.txt \
  --content-type markdown \
  --chunk-size 800 \
  --overlap 100 \
  --output chunks.jsonl

Caching Implementation

Achieve 80-90% cost reduction through content-addressable caching.

Caching Architecture by Query Volume:

  • <10K queries/month: In-memory cache (Python lru_cache)
  • 10K-100K queries/month: Redis (fast, TTL-based expiration)
  • 100K-1M queries/month: Redis (hot) + PostgreSQL (warm)
  • >1M queries/month: Multi-tier (Redis + PostgreSQL + S3)

Production Caching with Redis:

bash
# Embed documents with caching enabled
python scripts/cached_embedder.py \
  --model text-embedding-3-small \
  --input documents.jsonl \
  --output embeddings.npy \
  --cache-backend redis \
  --cache-ttl 2592000  # 30 days

Caching ROI Example:

  • 50,000 document chunks
  • 20% duplicate content
  • Without caching: $0.50 API cost
  • With caching (60% hit rate): $0.20 API cost
  • Savings: 60% ($0.30)

Dimensionality Trade-offs

Balance storage, search speed, and quality:

DimensionsStorage (1M vectors)Search Speed (p95)QualityUse Case
3841.5 GB10msGoodLarge-scale search
7683 GB15msHighGeneral purpose RAG
1,5366 GB25msVery HighHigh-quality retrieval
3,07212 GB40msHighestPremium applications

Key Insight: For most RAG applications, 768 dimensions (BGE-base-en-v1.5 local or equivalent) provides the best quality/cost/speed balance.

Batch Processing Optimization

Maximize throughput for large-scale ingestion:

OpenAI API:

  • Batch up to 2,048 inputs per request
  • Implement rate limiting (tier-dependent: 500-5,000 RPM)
  • Use parallel requests with backoff on rate limits

Local Models (sentence-transformers):

  • GPU acceleration (CUDA, MPS for Apple Silicon)
  • Batch size tuning (32-128 based on GPU memory)
  • Multi-GPU support for maximum throughput

Expected Throughput:

  • OpenAI API: 1,000-5,000 texts/minute (rate limit dependent)
  • Local GPU (RTX 3090): 5,000-10,000 texts/minute
  • Local CPU: 100-500 texts/minute
Show full SKILL.md (286 more words)Show less

Performance Monitoring

Track key metrics for optimization:

Critical Metrics:

  • Latency: Embedding generation time (p50, p95, p99)
  • Throughput: Embeddings per second/minute
  • Cost: API usage tracking (USD per 1K/1M tokens)
  • Cache Efficiency: Hit rate percentage

For detailed monitoring setup, metric collection patterns, and dashboarding, see references/performance-monitoring.md.

Monitor with Wrapper:

python
from scripts.performance_monitor import MonitoredEmbedder

monitored = MonitoredEmbedder(
    embedder=your_embedder,
    cost_per_1k_tokens=0.00002  # OpenAI pricing
)

embeddings = monitored.embed_batch(texts)
metrics = monitored.get_metrics()
print(f"Cache hit rate: {metrics['cache_hit_rate_pct']}%")
print(f"Total cost: ${metrics['total_cost_usd']}")

Working Examples

See examples/ directory for complete implementations:

Python Examples:

  • examples/openai_cached.py - OpenAI embeddings with Redis caching
  • examples/local_embedder.py - sentence-transformers local embedding
  • examples/smart_chunker.py - Content-aware recursive chunking
  • examples/performance_monitor.py - Pipeline performance tracking
  • examples/batch_processor.py - Large-scale document processing

All examples include:

  • Complete, runnable code
  • Dependency installation instructions
  • Error handling and retry logic
  • Configuration options

Integration Points

Upstream (This skill provides to):

  • Vector Databases: Embeddings flow to Pinecone, Weaviate, Qdrant, pgvector
  • RAG Systems: Optimized embeddings for retrieval pipelines
  • Semantic Search: Query and document embeddings for similarity search

Downstream (This skill uses from):

  • Document Processing: Chunk documents before embedding
  • Data Ingestion: Process documents from various sources

Related Skills:

  • For RAG architecture, see building-ai-chat skill
  • For vector database operations, see databases-vector skill
  • For data ingestion pipelines, see ingesting-data skill

Common Patterns

Pattern 1: RAG Pipeline

Document → Chunk → Embed → Store (vector DB) → Retrieve

Pattern 2: Semantic Search

Query → Embed → Search (vector DB) → Rank → Display

Pattern 3: Multi-Stage Retrieval (Cost Optimization)

Query → Cheap Embedding (384d) → Initial Search →
Expensive Embedding (1,536d) → Rerank Top-K → Return

Cost Savings: 70% reduction vs. single-stage with expensive embeddings

Quick Reference Checklist

Model Selection:

  • Identified data privacy requirements (local vs. API)
  • Calculated expected query volume
  • Determined quality requirements (good/high/highest)
  • Checked multilingual support needs

Chunking:

  • Analyzed content type (code, docs, legal, etc.)
  • Selected appropriate chunk size (500-1,500 chars)
  • Set overlap to prevent context loss (50-200 chars)
  • Validated chunks preserve semantic boundaries

Caching:

  • Implemented content-addressable hashing
  • Selected cache backend (Redis, PostgreSQL)
  • Set TTL based on content volatility
  • Monitoring cache hit rate (target: >60%)

Performance:

  • Tracking latency (embedding generation time)
  • Measuring throughput (embeddings/sec)
  • Monitoring costs (USD spent on API calls)
  • Optimizing batch sizes for maximum efficiency

© ancoleman, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (references) in skills/embedding-optimization of ancoleman/ai-design-components.

  • SKILL.md
  • examples/batch_processor.py
  • examples/benchmark_embeddings.py
  • examples/local_embedder.py
  • examples/openai_cached.py
  • examples/performance_monitor.py
  • examples/smart_chunker.py
  • outputs.yaml
  • references/chunking-strategies.md
  • references/model-selection-guide.md
  • references/performance-monitoring.md

Open the folder on GitHubat commit 76551b7

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in ancoleman/ai-design-components, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Embedding Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Embedding Optimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Embedding Optimization this skillancoleman/ai-design-components5261 repos~2.1kAutomated safety check: PassMIT
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k8 repos~2.3kAutomated safety check: PassMIT
Ms Agent Framework RAGshuyu-labs/WebCode278—~1.1kAutomated safety check: PassCustom licence
Pgvector Semantic Searchtimescale/pg-aiguide1.9k1 repos~3.8kAutomated safety check: PassApache-2.0
Evaluate RAGai-evals-course/evals-skills1.5k—~1.9kAutomated safety check: PassApache-2.0
RAG ArchitectJeffallan/claude-skills12k1 repos~2kAutomated safety check: PassMIT

Similar skills

  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Ms Agent Framework RAG

    shuyu-labs/WebCode

    Comprehensive guide for building Agentic RAG systems using Microsoft Agent Framework in C.

    278 GitHub stars~1.1k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Pgvector Semantic Search

    timescale/pg-aiguide

    A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.

    1.9k GitHub starsUsed in 1 repo~3.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Evaluate RAG

    ai-evals-course/evals-skills

    Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.

    1.5k GitHub stars~1.9k tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check passed
  • RAG Architect

    Jeffallan/claude-skills

    Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

    12k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • Memory Upgrade

    profbernardoj/everclaw-community-branches

    Diagnose and fix broken memory search in OpenClaw. An agent skill from profbernardoj/everclaw-community-branches.

    112 GitHub stars~574 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from ancoleman/ai-design-components

All 75 skills in this repo
  • Building AI Chat

    ancoleman/ai-design-components

    Builds AI chat interfaces and conversational UI with streaming responses, context management, and multi-modal support.

    526 GitHub starsUsed in 1 repo~3.4k tokens
    Auto-check passed
  • Building Forms

    ancoleman/ai-design-components

    Builds form components and data collection interfaces including contact forms, registration flows, checkout processes, surveys, and settings pages.

    526 GitHub stars~3.7k tokensUpdated 10 mo ago
    Auto-check passed
  • Building Tables

    ancoleman/ai-design-components

    Builds tables and data grids for displaying tabular information, from simple HTML tables to complex enterprise data grids.

    526 GitHub stars~1.8k tokensUpdated 10 mo ago
    Auto-check passed
  • Creating Dashboards

    ancoleman/ai-design-components

    Creates comprehensive dashboard and analytics interfaces that combine data visualization, KPI cards, real-time updates, and interactive layouts.

    526 GitHub stars~3.5k tokensUpdated 10 mo ago
    Auto-check passed
  • Designing Layouts

    ancoleman/ai-design-components

    Designs layout systems and responsive interfaces including grid systems, flexbox patterns, sidebar layouts, and responsive breakpoints.

    526 GitHub stars~1.7k tokensUpdated 10 mo ago
    Auto-check passed
  • Displaying Timelines

    ancoleman/ai-design-components

    Displays chronological events and activity through timelines, activity feeds, Gantt charts, and calendar interfaces.

    526 GitHub stars~2.7k tokensUpdated 10 mo ago
    Auto-check passed

Questions about Embedding Optimization

What does Embedding Optimization do?

Optimizing vector embeddings for RAG systems through model selection, chunking strategies, caching, and performance tuning. Embedding Optimization is an agent skill from ancoleman/ai-design-components. Optimizing vector embeddings for RAG systems through model selection, chunking strategies, caching, and performance tuning.

When should I use Embedding Optimization?

Embedding Optimization fits situations like: building semantic search; document retrieval systems that require cost-effective; high-quality embeddings.

How do I install Embedding Optimization in Claude Code?

Run `npx skills add ancoleman/ai-design-components --skill embedding-optimization -a claude-code`. Or copy the skill folder (skills/embedding-optimization in ancoleman/ai-design-components) into .claude/skills/embedding-optimization in your project. Claude Code loads it when a task matches its description.

How do I install Embedding Optimization in Codex?

Run `npx skills add ancoleman/ai-design-components --skill embedding-optimization -a codex`. Or copy the skill folder (skills/embedding-optimization in ancoleman/ai-design-components) into .agents/skills/embedding-optimization in your project. Codex loads it when a task matches its description.

Can I use Embedding Optimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ancoleman/ai-design-components --skill embedding-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/embedding-optimization, .gemini/skills/embedding-optimization, .github/skills/embedding-optimization and .opencode/skills/embedding-optimization in your project.

What does Embedding Optimization need to run?

Going by SKILL.md and its folder, Embedding Optimization needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Embedding Optimization access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Embedding Optimization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Embedding Optimization use?

Embedding Optimization is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Embedding Optimization use?

About 2.1k tokens (SKILL.md is roughly 8.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 12k tokens, read only when the agent opens those files.

What are the alternatives to Embedding Optimization?

Skills that share tags, products or a category with Embedding Optimization: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), Ms Agent Framework RAG (shuyu-labs/WebCode, 278 stars), Pgvector Semantic Search (timescale/pg-aiguide, 1.9k stars) and Evaluate RAG (ai-evals-course/evals-skills, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Embedding Optimization?

ancoleman (a GitHub user) maintains it in ancoleman/ai-design-components, which has 526 GitHub stars. The repository holds 75 skills in this directory. The repository was last updated on December 11, 2025.

Source: ancoleman/ai-design-components on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.