Agent skill

Clip Aware Embeddings

by curiositech in curiositech/some_claude_skills

Semantic image-text matching with CLIP and alternatives. An agent skill from curiositech/some_claude_skills.

MITAuto-check: notesAI & LLM Engineering

Install Clip Aware Embeddings

skills CLI
$ npx skills add curiositech/some_claude_skills --skill clip-aware-embeddings -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install curiositech/some_claude_skills clip-aware-embeddings --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/curiositech/some_claude_skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/clip-aware-embeddings .claude/skills/clip-aware-embeddings && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
clip-aware-embeddings
GitHub stars
243
Token cost
~2.3k tokens
SKILL.md length
538 words
Files
4 (incl. scripts)
Skills in repo
95
Repo updated
First seen
Licence
MIT

At a glance

Semantic image-text matching with CLIP and alternatives. An agent skill from curiositech/some_claude_skills.

  • Works in 4 steps: Is this a counting task? → Use object… → Fine-grained classification? → Use… → Spatial query? → Use spatial model → …
  • Zero-shot classification
  • SKILL.md covers MCP Integrations, Quick Decision Tree, When to Use This Skill and Installation, plus 9 more sections
  • Runs Python scripts from its folder; calls python and pip

What it does

Clip Aware Embeddings is an agent skill from curiositech/some_claude_skills. Semantic image-text matching with CLIP and alternatives. Use for image search, zero-shot classification, similarity matching. NOT for counting objects, fine-grained classification (celebrities, car models), spatial reasoning, or compositional queries. Activate on "CLIP", "embeddings", "image similarity", "semantic search", "zero-shot classification", "image-text matching".

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `.claude-plugin/plugin.json`, `CHANGELOG.md` and `scripts/validate_clip_usage.py`).

It sits in AI & LLM Engineering, covering Embeddings. The repository describes itself as: Claude skills that make my life easier. The licence is MIT.

When your agent uses it

  • Zero-shot classification
  • Similarity matching

Example prompts

  • “embeddings”
  • “image similarity”
  • “semantic search”
  • “/clip-aware-embeddings”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Is this a counting task? → Use object detection
  2. Fine-grained classification? → Use specialized model
  3. Spatial query? → Use spatial model
  4. Multiple objects with attributes? → Use compositional model

What it can do on your machine

Read from SKILL.md and the folder at commit 6713fc7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Clip Aware Embeddings loads about 2.3k tokens when it runs. Until then it costs about 99 tokens; SKILL.md has 538 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~99
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Write, Edit, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from curiositech/some_claude_skills at commit 6713fc7, republished under its MIT licence (© curiositech). 538 words, ~2,298 tokens.

Download SKILL.mdSave it as .claude/skills/clip-aware-embeddings/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
clip-aware-embeddings
description
Semantic image-text matching with CLIP and alternatives. Use for image search, zero-shot classification, similarity matching. NOT for counting objects, fine-grained classification (celebrities, car models), spatial reasoning, or compositional queries. Activate on "CLIP", "embeddings", "image similarity", "semantic search", "zero-shot classification", "image-text matching".
allowed-tools
Read, Write, Edit, Bash
metadata.category
AI & Machine Learning
metadata.tags
clip, embeddings, vision, similarity, zero-shot

CLIP-Aware Image Embeddings

Smart image-text matching that knows when CLIP works and when to use alternatives.

MCP Integrations

MCPPurpose
FirecrawlResearch latest CLIP alternatives and benchmarks
Hugging Face (if configured)Access model cards and documentation

Quick Decision Tree

Your task:
├─ Semantic search ("find beach images") → CLIP ✓
├─ Zero-shot classification (broad categories) → CLIP ✓
├─ Counting objects → DETR, Faster R-CNN ✗
├─ Fine-grained ID (celebrities, car models) → Specialized model ✗
├─ Spatial relations ("cat left of dog") → GQA, SWIG ✗
└─ Compositional ("red car AND blue truck") → DCSMs, PC-CLIP ✗

When to Use This Skill

✅ Use for:

  • Semantic image search
  • Broad category classification
  • Image similarity matching
  • Zero-shot tasks on new categories

❌ Do NOT use for:

  • Counting objects in images
  • Fine-grained classification
  • Spatial understanding
  • Attribute binding
  • Negation handling

Installation

bash
pip install transformers pillow torch sentence-transformers --break-system-packages

Validation: Run python scripts/validate_setup.py

Basic Usage

python
from transformers import CLIPProcessor, CLIPModel
from PIL import Image

model = CLIPModel.from_pretrained("openai/clip-vit-large-patch14")
processor = CLIPProcessor.from_pretrained("openai/clip-vit-large-patch14")

# Embed images
images = [Image.open(f"img{i}.jpg") for i in range(10)]
inputs = processor(images=images, return_tensors="pt")
image_features = model.get_image_features(**inputs)

# Search with text
text_inputs = processor(text=["a beach at sunset"], return_tensors="pt")
text_features = model.get_text_features(**text_inputs)

# Compute similarity
similarity = (image_features @ text_features.T).softmax(dim=0)

Common Anti-Patterns

Anti-Pattern 1: "CLIP for Everything"

❌ Wrong:

python
# Using CLIP to count cars in an image
prompt = "How many cars are in this image?"
# CLIP cannot count - it will give nonsense results

Why wrong: CLIP's architecture collapses spatial information into a single vector. It literally cannot count.

✓ Right:

python
from transformers import DetrImageProcessor, DetrForObjectDetection

processor = DetrImageProcessor.from_pretrained("facebook/detr-resnet-50")
model = DetrForObjectDetection.from_pretrained("facebook/detr-resnet-50")

# Detect objects
results = model(**processor(images=image, return_tensors="pt"))
# Filter for cars and count
car_detections = [d for d in results if d['label'] == 'car']
count = len(car_detections)

How to detect: If query contains "how many", "count", or numeric questions → Use object detection


Anti-Pattern 2: Fine-Grained Classification

❌ Wrong:

python
# Trying to identify specific celebrities with CLIP
prompts = ["Tom Hanks", "Brad Pitt", "Morgan Freeman"]
# CLIP will perform poorly - not trained for fine-grained face ID

Why wrong: CLIP trained on coarse categories. Fine-grained faces, car models, flower species require specialized models.

✓ Right:

python
# Use a fine-tuned face recognition model
from transformers import AutoFeatureExtractor, AutoModelForImageClassification

model = AutoModelForImageClassification.from_pretrained(
    "microsoft/resnet-50"  # Then fine-tune on celebrity dataset
)
# Or use dedicated face recognition: ArcFace, CosFace

How to detect: If query asks to distinguish between similar items in same category → Use specialized model


Anti-Pattern 3: Spatial Understanding

❌ Wrong:

python
# CLIP cannot understand spatial relationships
prompts = [
    "cat to the left of dog",
    "cat to the right of dog"
]
# Will give nearly identical scores

Why wrong: CLIP embeddings lose spatial topology. "Left" and "right" are treated as bag-of-words.

✓ Right:

python
# Use a spatial reasoning model
# Examples: GQA models, Visual Genome models, SWIG
from swig_model import SpatialRelationModel

model = SpatialRelationModel()
result = model.predict_relation(image, "cat", "dog")
# Returns: "left", "right", "above", "below", etc.

How to detect: If query contains directional words (left, right, above, under, next to) → Use spatial model


Anti-Pattern 4: Attribute Binding

❌ Wrong:

python
prompts = [
    "red car and blue truck",
    "blue car and red truck"
]
# CLIP often gives similar scores for both

Why wrong: CLIP cannot bind attributes to objects. It sees "red, blue, car, truck" as a bag of concepts.

✓ Right - Use PC-CLIP or DCSMs:

python
# PC-CLIP: Fine-tuned for pairwise comparisons
from pc_clip import PCCLIPModel

model = PCCLIPModel.from_pretrained("pc-clip-vit-l")
# Or use DCSMs (Dense Cosine Similarity Maps)

How to detect: If query has multiple objects with different attributes → Use compositional model


Evolution Timeline

2021: CLIP Released
  • Revolutionary: zero-shot, 400M image-text pairs
  • Widely adopted for everything
  • Limitations not yet understood
2022-2023: Limitations Discovered
  • Cannot count objects
  • Poor at fine-grained classification
  • Fails spatial reasoning
  • Can't bind attributes
2024: Alternatives Emerge
  • DCSMs: Preserve patch/token topology
  • PC-CLIP: Trained on pairwise comparisons
  • SpLiCE: Sparse interpretable embeddings
2025: Current Best Practices
  • Use CLIP for what it's good at
  • Task-specific models for limitations
  • Compositional models for complex queries

LLM Mistake: LLMs trained on 2021-2023 data will suggest CLIP for everything because limitations weren't widely known. This skill corrects that.


Show full SKILL.md (204 more words)Show less

Validation Script

Before using CLIP, check if it's appropriate:

bash
python scripts/validate_clip_usage.py \
    --query "your query here" \
    --check-all

Returns:

  • ✅ CLIP is appropriate
  • ❌ Use alternative (with suggestion)

Task-Specific Guidance

Image Search (CLIP ✓)
python
# Good use of CLIP
queries = ["beach", "mountain", "city skyline"]
# Works well for broad semantic concepts
Zero-Shot Classification (CLIP ✓)
python
# Good: Broad categories
categories = ["indoor", "outdoor", "nature", "urban"]
# CLIP excels at this
Object Counting (CLIP ✗)
python
# Use object detection instead
from transformers import DetrImageProcessor, DetrForObjectDetection
# See /references/object_detection.md
Fine-Grained Classification (CLIP ✗)
python
# Use specialized models
# See /references/fine_grained_models.md
Spatial Reasoning (CLIP ✗)
python
# Use spatial relation models
# See /references/spatial_models.md

Troubleshooting

Issue: CLIP gives unexpected results

Check:

  1. Is this a counting task? → Use object detection
  2. Fine-grained classification? → Use specialized model
  3. Spatial query? → Use spatial model
  4. Multiple objects with attributes? → Use compositional model

Validation:

bash
python scripts/diagnose_clip_issue.py --image path/to/image --query "your query"
Issue: Low similarity scores

Possible causes:

  1. Query too specific (CLIP works better with broad concepts)
  2. Fine-grained task (not CLIP's strength)
  3. Need to adjust threshold

Solution: Try broader query or use alternative model


Model Selection Guide

ModelBest ForAvoid For
CLIP ViT-L/14Semantic search, broad categoriesCounting, fine-grained, spatial
DETRObject detection, countingSemantic similarity
DINOv2Fine-grained featuresText-image matching
PC-CLIPAttribute binding, comparisonsGeneral embedding
DCSMsCompositional reasoningSimple similarity

Performance Notes

CLIP models:

  • ViT-B/32: Fast, lower quality
  • ViT-L/14: Balanced (recommended)
  • ViT-g-14: Highest quality, slower

Inference time (single image, CPU):

  • ViT-B/32: ~100ms
  • ViT-L/14: ~300ms
  • ViT-g-14: ~1000ms

Further Reading

  • /references/clip_limitations.md - Detailed analysis of CLIP's failures
  • /references/alternatives.md - When to use what model
  • /references/compositional_reasoning.md - DCSMs and PC-CLIP deep dive
  • /scripts/validate_clip_usage.py - Pre-flight validation tool
  • /scripts/diagnose_clip_issue.py - Debug unexpected results

See CHANGELOG.md for version history.

© curiositech, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in .claude/skills/clip-aware-embeddings of curiositech/some_claude_skills.

  • SKILL.md
  • .claude-plugin/plugin.json
  • CHANGELOG.md
  • scripts/validate_clip_usage.py

Open the folder on GitHubat commit 6713fc7

Compare with similar skills

Clip Aware Embeddings next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Clip Aware Embeddings compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Clip Aware Embeddings this skillcuriositech/some_claude_skills243—~2.3kAutomated safety check: NotesMIT
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k7 repos~2.3kAutomated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Codebase Managementgiancarloerra/SocratiCode3.3k1 repos~1.8kAutomated safety check: PassAGPL-3.0
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k7 repos~1.7kAutomated safety check: PassMIT
Sentence-Transformers Training Routerhuggingface/skills11k1 repos~2.6kAutomated safety check: PassApache-2.0

Similar skills

  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Codebase Management

    giancarloerra/SocratiCode

    Set up, index, and manage SocratiCode codebase indexing. An agent skill from giancarloerra/SocratiCode.

    3.3k GitHub starsUsed in 1 repo~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Mashup Mods

    rehan-remade/universal-modder

    Build cross-game mashups and total conversions, the "Minecraft inside Elden Ring" or "skateboarding in MW2" kind.

    5.8k GitHub stars~3.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from curiositech/some_claude_skills

All 95 skills in this repo
  • Crisis Detection Intervention AI

    curiositech/some_claude_skills

    Detect crisis signals in user content using NLP, mental health sentiment analysis, and safe intervention protocols.

    243 GitHub starsUsed in 2 repos~3.8k tokens
    Auto-check passed
  • Form Validation Architect

    curiositech/some_claude_skills

    End-to-end form handling with react-hook-form, Zod schemas, validation patterns, error messaging, field arrays, and multi-step wizards.

    243 GitHub stars~3.8k tokensUpdated 1 mo ago
    Auto-check passed
  • GitHub Actions Pipeline Builder

    curiositech/some_claude_skills

    Build production CI/CD pipelines with GitHub Actions. An agent skill from curiositech/some_claude_skills.

    243 GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check: notes
  • Background Job Orchestrator

    curiositech/some_claude_skills

    Expert in background job processing with Bull/BullMQ (Redis), Celery, and cloud queues.

    243 GitHub stars~3.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Competitive Cartographer

    curiositech/some_claude_skills

    Strategic analyst that maps competitive landscapes, identifies white space opportunities, and provides positioning recommendations.

    243 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Computer Vision Pipeline

    curiositech/some_claude_skills

    Build production computer vision pipelines for object detection, tracking, and video analysis.

    243 GitHub stars~4k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Clip Aware Embeddings

What does Clip Aware Embeddings do?

Semantic image-text matching with CLIP and alternatives. An agent skill from curiositech/some_claude_skills. Clip Aware Embeddings is an agent skill from curiositech/some_claude_skills. Semantic image-text matching with CLIP and alternatives.

When should I use Clip Aware Embeddings?

Clip Aware Embeddings fits situations like: zero-shot classification; similarity matching.

How do I install Clip Aware Embeddings in Claude Code?

Run `npx skills add curiositech/some_claude_skills --skill clip-aware-embeddings -a claude-code`. Or copy the skill folder (.claude/skills/clip-aware-embeddings in curiositech/some_claude_skills) into .claude/skills/clip-aware-embeddings in your project. Claude Code loads it when a task matches its description.

How do I install Clip Aware Embeddings in Codex?

Run `npx skills add curiositech/some_claude_skills --skill clip-aware-embeddings -a codex`. Or copy the skill folder (.claude/skills/clip-aware-embeddings in curiositech/some_claude_skills) into .agents/skills/clip-aware-embeddings in your project. Codex loads it when a task matches its description.

Can I use Clip Aware Embeddings in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add curiositech/some_claude_skills --skill clip-aware-embeddings -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/clip-aware-embeddings, .gemini/skills/clip-aware-embeddings, .github/skills/clip-aware-embeddings and .opencode/skills/clip-aware-embeddings in your project.

What does Clip Aware Embeddings need to run?

Going by SKILL.md and its folder, Clip Aware Embeddings needs Python for the scripts in its folder and the command-line tools its instructions call (python and pip). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash.

Does Clip Aware Embeddings access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Clip Aware Embeddings safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Clip Aware Embeddings use?

Clip Aware Embeddings is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Clip Aware Embeddings use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Clip Aware Embeddings?

Skills that share tags, products or a category with Clip Aware Embeddings: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars), Codebase Management (giancarloerra/SocratiCode, 3.3k stars) and CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Clip Aware Embeddings?

curiositech (a GitHub organization) maintains it in curiositech/some_claude_skills, which has 243 GitHub stars. The repository holds 95 skills in this directory. The repository was last updated on September 6, 2026.

Source: curiositech/some_claude_skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.