Agent skill

Semantic Grep

by oaustegard in oaustegard/claude-skills

In-process semantic search over text files or in-memory strings, using Gemini embeddings via the CF AI Gateway.

MITAuto-check passedAI & LLM Engineering

Install Semantic Grep

skills CLI
$ npx skills add oaustegard/claude-skills --skill semantic-grep -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install oaustegard/claude-skills semantic-grep --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/semantic-grep .claude/skills/semantic-grep && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
semantic-grep
GitHub stars
150
Token cost
~2.1k tokens
SKILL.md length
855 words
Files
4 (incl. scripts)
Skills in repo
66
Repo updated
First seen
Licence
MIT

At a glance

In-process semantic search over text files or in-memory strings, using Gemini embeddings via the CF AI Gateway.

  • User wants fuzzy/conceptual search where exact-keyword grep would miss — sessions discussing regulatory constraints
  • SKILL.md covers When Semantic Search Helps, Setup, Quick Start and Core API, plus 6 more sections
  • Runs Python scripts from its folder; needs GOOGLE_API_KEY and GEMINI_API_KEY
  • Code about retry logic

What it does

Semantic Grep is an agent skill from oaustegard/claude-skills. In-process semantic search over text files or in-memory strings, using Gemini embeddings via the CF AI Gateway. Use when user wants fuzzy/conceptual search where exact-keyword grep would miss — "sessions discussing regulatory constraints", "code about retry logic", "notes mentioning burnout even if the word isn't there". Complements searching-codebases (regex/AST) and extracting-keywords (YAKE). Do NOT use when an exact string/regex match is what's wanted — grep/rg wins on speed and precision there.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `CHANGELOG.md`, `README.md` and `scripts/semantic_grep.py`).

It sits in AI & LLM Engineering, covering Embeddings and Error handling. The repository describes itself as: My collection of Claude skills. The licence is MIT.

When your agent uses it

  • User wants fuzzy/conceptual search where exact-keyword grep would miss — sessions discussing regulatory constraints
  • Code about retry logic
  • Notes mentioning burnout even if the word isnt there
  • An exact string/regex match is whats wanted — grep/rg wins on speed and precision there

Example prompts

  • “sessions discussing regulatory constraints”
  • “code about retry logic”
  • “notes mentioning burnout even if the word isn”
  • “/semantic-grep”

Requirements

  • Python 3
  • A credential in CF_API_TOKEN
  • A credential in GOOGLE_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit cf49d47. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GOOGLE_API_KEY
    • GEMINI_API_KEY
    • CF_API_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Semantic Grep loads about 2.1k tokens when it runs. Until then it costs about 130 tokens; SKILL.md has 855 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~130
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from oaustegard/claude-skills at commit cf49d47, republished under its MIT licence (© oaustegard). 855 words, ~2,140 tokens.

Download SKILL.mdSave it as .claude/skills/semantic-grep/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
semantic-grep
description
In-process semantic search over text files or in-memory strings, using Gemini embeddings via the CF AI Gateway. Use when user wants fuzzy/conceptual search where exact-keyword grep would miss — "sessions discussing regulatory constraints", "code about retry logic", "notes mentioning burnout even if the word isn't there". Complements searching-codebases (regex/AST) and extracting-keywords (YAKE). Do NOT use when an exact string/regex match is what's wanted — grep/rg wins on speed and precision there.
metadata.version
0.2.1

Semantic Grep

jina-grep-style semantic search, done in-process via Python rather than as an external CLI. Embeds query + corpus chunks with gemini-embedding-2, ranks by cosine similarity, returns grep-format output.

When Semantic Search Helps

The core trade-off (lifted from jina-grep-cli's own docs and validated in testing):

TaskTool
Known exact string, filename, or regexgrep / rg / searching-codebases
"What files discuss concept X" when X may not appear verbatimsemantic-grep
Hybrid: prefilter with grep, rerank by conceptgrep → rerank_candidates()

Regression test result (workshop session corpus, 135 docs):

  • "handling regulatory constraints" → top hit "Engineering AI Systems Under Sovereignty Constraints" (0.67). ✓
  • "sessions about GEPA" → top hit "Gemma, DeepMind's Family of Open Models" (0.69). ✗ — false positive on phonetic neighbor. GEPA is mentioned verbatim in one session description; grep would find it correctly.

Rule: when the user query reads like a named entity or keyword, try grep first. Only reach for semantic-grep when paraphrase/concept matching is actually needed.

Setup

Credentials via proxy.env (Cloudflare AI Gateway w/ BYOK — same pattern as invoking-gemini):

CF_ACCOUNT_ID=...
CF_GATEWAY_ID=...
CF_API_TOKEN=...

Direct-API fallback: GOOGLE_API_KEY or GEMINI_API_KEY env var. No dependencies beyond requests + numpy.

Quick Start

python
import sys
sys.path.insert(0, '/mnt/skills/user/semantic-grep/scripts')
from semantic_grep import semantic_grep, format_grep

# Directory of .txt files
results = semantic_grep("error handling under load", "/path/to/notes",
                        top_k=5, granularity="paragraph")
print(format_grep(results))
# notes/incidents.txt:42:  When the queue depth exceeds... [0.71]
# notes/postmortem.txt:8:  Under sustained traffic we saw... [0.68]

Core API

semantic_grep(query, corpus, *, top_k=10, threshold=None, ...)

Main search function.

  • query (str) — the search query (embedded with RETRIEVAL_QUERY task type)
  • corpus (str | Path | list[Chunk]) — a file, directory, or pre-chunked list
  • top_k (int | None) — max results; None = all above threshold
  • threshold (float | None) — cosine similarity cutoff; None = no filter (top_k only)
  • granularity ("paragraph" | "line") — how to chunk files (default paragraph)
  • include (str) — filename-glob filter when corpus is a directory (default "*.txt"). Matches against Path.name only, not the full path — "*.md" works, "docs/*.md" does not.
  • model (str) — default "gemini-embedding-2". gemini-embedding-001 is retired (text-only) and warns if passed explicitly.
  • dim (int) — 128 / 768 / 1536 / 3072 (default 768; MRL-truncated + renormalized)
  • task ("text" | "code") — selects text vs code task types

Returns list[Match] where Match has path, line, text, score.

load_corpus(path, *, include="*.txt", granularity="paragraph") -> list[Chunk]

Load and chunk a file or directory without embedding. Useful for inspecting what gets embedded before paying for the API call.

embed_batch(texts, task_type, *, model, dim, group_size=100) -> np.ndarray

Lower-level: embed a list of strings directly via :batchEmbedContents. Returns (N, dim) float32 array, rows normalized when dim < 3072.

format_grep(matches, *, max_text_chars=200, show_score=True) -> str

Format matches as grep output: path:line: snippet [score].

Pipe-mode Rerank Pattern

The highest-leverage use isn't naive full-corpus semantic search — it's hybrid retrieval: fast coarse filter → semantic rerank.

python
import subprocess
from semantic_grep import Chunk, semantic_grep, format_grep

# Stage 1: fast exact/regex prefilter with rg
result = subprocess.run(
    ["rg", "-n", "--no-heading", "error|fail|timeout", "logs/"],
    capture_output=True, text=True,
)

# Parse `path:line:text` into Chunks
chunks = []
for raw in result.stdout.splitlines():
    path, line, text = raw.split(":", 2)
    chunks.append(Chunk(path=path, line=int(line), text=text))

# Stage 2: semantic rerank on the prefiltered subset
ranked = semantic_grep("intermittent queue saturation during peak traffic",
                       chunks, top_k=10)
print(format_grep(ranked))

This is how you scale past the "embed the whole corpus every call" limit without needing a vector DB. The exact-match stage cheaply cuts millions of lines to thousands; semantic reranks those.

Task Types (Gemini)

  • text mode (default): query → RETRIEVAL_QUERY, docs → RETRIEVAL_DOCUMENT. Asymmetric — documented to outperform symmetric encoding for retrieval.
  • code mode: query → CODE_RETRIEVAL_QUERY, docs → RETRIEVAL_DOCUMENT. Use when searching code with natural-language queries.

Use SEMANTIC_SIMILARITY (symmetric) only if you're doing pairwise sim, not retrieval. This module doesn't expose that path yet.

Show full SKILL.md (392 more words)Show less

Model Notes

gemini-embedding-2 (GA since 2026-04-22) — general-purpose and multimodal. Verified 2026-07-21 via the CF gateway: text, image and audio all embed to the same space at the requested dim, L2-normalized. The retired gemini-embedding-001 was text-only and rejected non-text input with HTTP 400:

  • 2,048 input token limit per text. Longer texts are truncated at ~8K chars (approximation).
  • Matryoshka (MRL) — 3072 native dims, safely truncatable to 1536/768/256/128.
  • 3072 is auto-normalized; lower dims need client-side renorm (handled here).
  • Pricing: $0.15 / 1M input tokens. 135 medium paragraphs ≈ 15K tokens ≈ $0.002 per query.

gemini-embedding-2-preview (March 2026) is multimodal and currently top of MTEB. Set model="gemini-embedding-2-preview" to opt in once the preview stabilizes.

Limitations

  • No persistent index. Every call re-embeds the corpus. Fine for <~1K chunks; prohibitive for real knowledge bases. Phase 2: cache embeddings by content hash.
  • Token budget is approximated by char count (×1.5). Conservative for mixed-script text; over-truncates English slightly. Real tokenizer would use the Gemini tokenizer endpoint but costs an extra call per embed.
  • Batch bulk-failure diagnostic. If one text in a group of 100 overflows or is rejected by safety filters, the whole batch fails and the 99 good ones are lost. No per-index fallback yet.
  • No memory ceiling on corpus size. semantic_grep pre-allocates (N, dim) float32; 1M chunks at dim=768 ≈ 3GB. Caller is responsible for sane chunk counts. load_corpus also follows symlinks via rglob — fine in a trusted single-user container, not for untrusted paths.
  • Sequential batch groups. group_size=100 per HTTP call; groups run serially. For >1K chunks, add asyncio — not needed yet.
  • No CLI shim. Called as a Python module, not a subprocess. Per design: "within an LLM rather than calling out to one."
  • Embedding function lives here, not in invoking-gemini. Should be factored up when invoking-gemini adds embedding support. Tracked as followup.
  • invoking-gemini — sibling; handles Gemini text + image generation through the same CF gateway. Shares credential pattern.
  • searching-codebases — regex/AST search. Use first when the query is a known pattern.
  • extracting-keywords — YAKE keyword extraction; orthogonal, but pairs well for building query terms from a long prompt.
  • exploring-codebases — for understanding repo structure. Semantic-grep doesn't replace AST-based navigation.

Attribution

Conceptually inspired by jina-grep-cli — we kept the retrieval shape (grep-compatible output, asymmetric query/doc embeddings, threshold + top-k) but swapped the MLX/Apple-Silicon backend for a portable Gemini API call. The original's pipe-mode rerank pattern is the most generalizable idea it contributes and is preserved here.

© oaustegard, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in semantic-grep of oaustegard/claude-skills.

  • SKILL.md
  • CHANGELOG.md
  • README.md
  • scripts/semantic_grep.py

Open the folder on GitHubat commit cf49d47

Compare with similar skills

Semantic Grep next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Semantic Grep compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Semantic Grep this skilloaustegard/claude-skills150—~2.1kAutomated safety check: PassMIT
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k7 repos~2.3kAutomated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Codebase Managementgiancarloerra/SocratiCode3.3k1 repos~1.8kAutomated safety check: PassAGPL-3.0
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k7 repos~1.7kAutomated safety check: PassMIT
Sentence-Transformers Training Routerhuggingface/skills11k1 repos~2.6kAutomated safety check: PassApache-2.0

Similar skills

  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Codebase Management

    giancarloerra/SocratiCode

    Set up, index, and manage SocratiCode codebase indexing. An agent skill from giancarloerra/SocratiCode.

    3.3k GitHub starsUsed in 1 repo~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Mashup Mods

    rehan-remade/universal-modder

    Build cross-game mashups and total conversions, the "Minecraft inside Elden Ring" or "skateboarding in MW2" kind.

    5.8k GitHub stars~3.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from oaustegard/claude-skills

All 66 skills in this repo
  • Vega-Lite Interactive Charts

    oaustegard/claude-skills

    Builds interactive Vega-Lite charts from uploaded data: analyzes the fields, picks five to ten fitting chart types, and produces a React artifact with the data embedded inline.

    150 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Single-File HTML Composer

    oaustegard/claude-skills

    Builds self-contained single-file HTML pages such as reports, decks, postmortems, flowcharts and prototypes from a small spec using a bundled Python composer and templates.

    150 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Declauding

    oaustegard/claude-skills

    Rewrites model-sounding prose into plain technical writing and checks that every claim survives, for PR text, docs, commit messages and similar drafts.

    150 GitHub stars~5.2k tokensUpdated today
    Auto-check passed
  • Preact Developer

    oaustegard/claude-skills

    Guides building standards-based Preact apps with native-first choices, HTM syntax, import maps and vendored ESM, from single-file demos to larger builds.

    150 GitHub stars~4.6k tokensUpdated today
    Auto-check passed
  • Bluesky Zeitgeist Sampler

    oaustegard/claude-skills

    Deprecated sampler that captures short windows of the Bluesky firehose, clusters trending terms and builds an HTML report; replaced by the browsing-bluesky skill.

    150 GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Adversarial Review Before Shipping

    oaustegard/claude-skills

    Has a fresh-context adversary attack a blog post, recommendation, analysis brief or piece of code before you ship it, using a profile suited to that kind of artifact.

    150 GitHub stars~3.4k tokensUpdated today
    Auto-check passed

Questions about Semantic Grep

What does Semantic Grep do?

In-process semantic search over text files or in-memory strings, using Gemini embeddings via the CF AI Gateway. Semantic Grep is an agent skill from oaustegard/claude-skills. In-process semantic search over text files or in-memory strings, using Gemini embeddings via the CF AI Gateway.

When should I use Semantic Grep?

Semantic Grep fits situations like: user wants fuzzy/conceptual search where exact-keyword grep would miss — sessions discussing regulatory constraints; code about retry logic; notes mentioning burnout even if the word isnt there; an exact string/regex match is whats wanted — grep/rg wins on speed and precision there.

How do I install Semantic Grep in Claude Code?

Run `npx skills add oaustegard/claude-skills --skill semantic-grep -a claude-code`. Or copy the skill folder (semantic-grep in oaustegard/claude-skills) into .claude/skills/semantic-grep in your project. Claude Code loads it when a task matches its description.

How do I install Semantic Grep in Codex?

Run `npx skills add oaustegard/claude-skills --skill semantic-grep -a codex`. Or copy the skill folder (semantic-grep in oaustegard/claude-skills) into .agents/skills/semantic-grep in your project. Codex loads it when a task matches its description.

Can I use Semantic Grep in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oaustegard/claude-skills --skill semantic-grep -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/semantic-grep, .gemini/skills/semantic-grep, .github/skills/semantic-grep and .opencode/skills/semantic-grep in your project.

What does Semantic Grep need to run?

Going by SKILL.md and its folder, Semantic Grep needs Python for the scripts in its folder and credentials named GOOGLE_API_KEY, GEMINI_API_KEY and CF_API_TOKEN. Our summary lists: Python 3; A credential in CF_API_TOKEN; A credential in GOOGLE_API_KEY.

Does Semantic Grep access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Semantic Grep safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Semantic Grep use?

Semantic Grep is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Semantic Grep use?

About 2.1k tokens (SKILL.md is roughly 8.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Semantic Grep?

Skills that share tags, products or a category with Semantic Grep: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars), Codebase Management (giancarloerra/SocratiCode, 3.3k stars) and CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Semantic Grep?

oaustegard (a GitHub user) maintains it in oaustegard/claude-skills, which has 150 GitHub stars. The repository holds 66 skills in this directory. The repository was last updated on October 8, 2026.

Source: oaustegard/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.