Agent skill

Cursor Codebase Indexing

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Set up and optimize Cursor codebase indexing for semantic code search and @Codebase queries.

MITAuto-check: notesAI & LLM Engineering

Install Cursor Codebase Indexing

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill cursor-codebase-indexing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace cursor-codebase-indexing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/cursor-codebase-indexing .claude/skills/cursor-codebase-indexing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cursor-codebase-indexing
GitHub stars
2.8k
Token cost
~2k tokens
SKILL.md length
736 words
Files
9 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Set up and optimize Cursor codebase indexing for semantic code search and @Codebase queries.

  • Works in 4 steps: Define the intended workspace scope and… → Verify required files can be found and… → Reindex only after a reviewed… → …
  • Codebase indexing
  • SKILL.md covers Overview, Prerequisites, Instructions and Output, plus 9 more sections
  • Calls cursor

What it does

Cursor Codebase Indexing is an agent skill from jeremylongshore/tons-of-skills-marketplace. Set up and optimize Cursor codebase indexing for semantic code search and @Codebase queries. Triggers on "cursor index", "codebase indexing", "index codebase", "cursor semantic search", "@codebase", "cursor embeddings".

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including reference files (for example `references/configuration.md`, `references/errors.md` and `references/examples.md`). Compatibility notes: Designed for Claude Code

It sits in AI & LLM Engineering, covering Embeddings and Codebase onboarding. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Codebase indexing
  • Cursor semantic search
  • Cursor embeddings

Example prompts

  • “cursor index”
  • “codebase indexing”
  • “index codebase”
  • “/cursor-codebase-indexing”

Requirements

  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash(cmd:*)

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Define the intended workspace scope and add exclusions before the first indexing pass.
  2. Verify required files can be found and excluded paths cannot be attached or discovered.
  3. Reindex only after a reviewed configuration/workspace change and capture a redacted result.
  4. Monitor index size/performance and narrow scope rather than weakening data controls.

What it can do on your machine

Read from SKILL.md and the folder at commit 23ea8d4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash(cmd:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • cursor

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.cursor.com
    • cursor.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Cursor Codebase Indexing loads about 2k tokens when it runs, and up to ~3.4k if it reads all its reference files. Until then it costs about 61 tokens; SKILL.md has 736 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~61
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:46
    d `docs/` for a service while excluding `.env*`, build output, vendor directories, and customer exports. Verify a saniti
  • NoteMentions a .env fileSKILL.md:128
    .env*
  • NoteRuns commands with sudoSKILL.md:229
    sudo sysctl fs.inotify.max_user_watches=524288
  • NoteRuns commands with sudoSKILL.md:232
    o "fs.inotify.max_user_watches=524288" | sudo tee -a /etc/sysctl.conf
  • NoteRuns commands with sudoSKILL.md:233
    sudo sysctl -p

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit 23ea8d4, republished under its MIT licence (© jeremylongshore). 736 words, ~2,037 tokens.

Download SKILL.mdSave it as .claude/skills/cursor-codebase-indexing/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
cursor-codebase-indexing
description
Set up and optimize Cursor codebase indexing for semantic code search and @Codebase queries. Triggers on "cursor index", "codebase indexing", "index codebase", "cursor semantic search", "@codebase", "cursor embeddings".
allowed-tools
Read, Write, Edit, Bash(cmd:*)
compatibility
Designed for Claude Code
version
1.19.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, cursor, cursor-codebase

Cursor Codebase Indexing

Overview

Configure codebase indexing so relevant source and documentation are available while sensitive, generated, oversized, and unrelated files remain excluded from AI context.

Prerequisites

  • Repository ownership, expected source/documentation paths, and reviewed .cursorignore/indexing exclusions.
  • A non-sensitive verification fixture and a change-review path for configuration updates.

Instructions

  1. Define the intended workspace scope and add exclusions before the first indexing pass.
  2. Verify required files can be found and excluded paths cannot be attached or discovered.
  3. Reindex only after a reviewed configuration/workspace change and capture a redacted result.
  4. Monitor index size/performance and narrow scope rather than weakening data controls.

Output

  • A verified index scope with explicit included/excluded paths and an owner for maintenance.

Examples

Index src/ and docs/ for a service while excluding .env*, build output, vendor directories, and customer exports. Verify a sanitized source file is available and an excluded placeholder is not before allowing AI work on the repository.

Set up and optimize Cursor's codebase indexing system. Indexing creates embeddings of your code, enabling @Codebase semantic search and improving AI context awareness across Chat, Composer, and Agent mode.

How Indexing Works

Your Code Files
      │
      ▼
  Syntax Chunking ─── splits files into meaningful code blocks
      │
      ▼
  Embedding Generation ─── converts chunks to vector representations
      │
      ▼
  Vector Storage (Turbopuffer) ─── cloud-hosted nearest-neighbor search
      │
      ▼
  @Codebase Query ─── your question → embedding → similarity search → relevant chunks
Key Architecture Details
  • Merkle tree for change detection: only modified files are re-indexed (every 10 minutes)
  • No plaintext storage: code is not stored server-side; only embeddings and obfuscated metadata
  • Privacy Mode compatible: with Privacy Mode on, embeddings are computed without retaining source code
  • Indexing runs in the background; small projects complete in seconds, large projects (50K+ files) may take hours initially

Initial Setup

  1. Open your project in Cursor
  2. Indexing starts automatically on first open
  3. Check status: look at the bottom status bar for "Indexing..." indicator
  4. View indexed files: Cursor Settings > Features > Codebase Indexing > View included files
Verify Indexing Status

The status bar shows:

  • "Indexing..." with progress indicator -- initial indexing in progress
  • "Indexed" -- indexing complete, @Codebase queries are available
  • No indicator -- indexing may be disabled or not started

Configuration

.cursorignore

Exclude files from indexing and AI features. Place in project root. Uses .gitignore syntax:

gitignore
# .cursorignore

# Build artifacts (large, not useful for AI context)
dist/
build/
out/
.next/
target/

# Dependencies
node_modules/
vendor/
venv/
.venv/

# Generated files
*.min.js
*.min.css
*.bundle.js
*.map
*.lock

# Large data files
*.csv
*.sql
*.sqlite
*.parquet
fixtures/
seed-data/

# Secrets (defense in depth -- also use .gitignore)
.env*
**/secrets/
**/credentials/
.cursorindexingignore

Exclude files from indexing only but keep them accessible to AI features when explicitly referenced:

gitignore
# .cursorindexingignore

# Large test fixtures -- don't index, but allow @Files reference
tests/fixtures/
e2e/recordings/

# Documentation build output
docs/.vitepress/dist/

Difference: .cursorignore hides files from both indexing and AI features. .cursorindexingignore only excludes from the index; files can still be referenced via @Files.

Default Exclusions

Cursor automatically excludes everything in .gitignore. You only need .cursorignore for files tracked by git that you want to exclude from AI.

Using the Index

@Codebase Queries

Ask semantic questions about your entire codebase:

@Codebase where is user authentication handled?

@Codebase show me all API endpoints that accept file uploads

@Codebase how does the payment processing flow work?

@Codebase find all places where we connect to Redis

@Codebase performs a nearest-neighbor search using your question's embedding. It returns the most semantically similar code chunks, even if they do not contain the exact keywords you used.

Show full SKILL.md (300 more words)Show less
MethodWhen to UseContext Cost
@CodebaseDiscovery -- you don't know which filesHigh (many chunks)
@FilesYou know exactly which fileLow (one file)
@FoldersYou know the directoryMedium-High
Ctrl+Shift+FExact text/regex matchN/A (editor search)

Use @Codebase for discovery, then switch to @Files once you know where the code lives.

Optimization for Large Projects

Monorepo Strategy

For monorepos with many packages, open the specific package directory instead of the root:

bash
# Instead of opening the entire monorepo:
cursor /path/to/monorepo           # Indexes everything -- slow

# Open the specific package:
cursor /path/to/monorepo/packages/api   # Indexes only this package -- fast

Or use .cursorignore at the root to exclude packages you are not actively working on:

gitignore
# .cursorignore -- monorepo, focus on api and shared
packages/web/
packages/mobile/
packages/admin/
# packages/api/    ← not listed, so it IS indexed
# packages/shared/ ← not listed, so it IS indexed
Re-Indexing

If search results are stale or indexing appears stuck:

  1. Cmd+Shift+P > Cursor: Resync Index
  2. Wait for status bar to show indexing progress
  3. If that fails, delete the local cache:
    • macOS: ~/Library/Application Support/Cursor/Cache/
    • Linux: ~/.config/Cursor/Cache/
    • Windows: %APPDATA%\Cursor\Cache\
  4. Restart Cursor and allow full re-index
File Watcher Limits (Linux)

On Linux, large projects may hit the file watcher limit:

bash
# Check current limit
cat /proc/sys/fs/inotify/max_user_watches

# Increase (temporary)
sudo sysctl fs.inotify.max_user_watches=524288

# Increase (permanent)
echo "fs.inotify.max_user_watches=524288" | sudo tee -a /etc/sysctl.conf
sudo sysctl -p

Enterprise Considerations

  • Data residency: Embeddings are stored in Turbopuffer (cloud). Obfuscated filenames and no plaintext code, but metadata exists
  • Privacy Mode: With Privacy Mode on, embeddings are computed with zero data retention at the provider
  • Air-gapped environments: Indexing requires network access to Cursor's embedding API. Not available offline
  • Indexing scope: Only files in the currently open workspace are indexed. Closing a project removes its index from active queries

Troubleshooting

SymptomCauseFix
@Codebase returns no resultsIndex not builtWait for "Indexed" in status bar
Search misses known filesFile in .gitignore or .cursorignoreCheck ignore files
Indexing stuck at N%Large project or network issueResync index via Command Palette
Stale results after refactorIndex not yet updatedWait 10 min or manual resync
High CPU during indexingInitial embedding computationNormal for first run; subsides

Resources

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (references) in skills/.curated/cursor-codebase-indexing of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/configuration.md
  • references/errors.md
  • references/examples.md
  • references/index-settings.md
  • references/maintaining-the-index.md
  • references/optimization-for-large-projects.md
  • references/troubleshooting.md
  • references/using-the-index.md

Open the folder on GitHubat commit 23ea8d4

Compare with similar skills

Cursor Codebase Indexing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cursor Codebase Indexing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cursor Codebase Indexing this skilljeremylongshore/tons-of-skills-marketplace2.8k—~2kAutomated safety check: NotesMIT
Codebase Explorationgiancarloerra/SocratiCode3.3k1 repos~1.5kAutomated safety check: PassAGPL-3.0
Symdex Code Searchhusnainpk/SymDex208—~2.3kAutomated safety check: PassMIT
Agentsop Repo Mapagentsope/SkillAlchemy459—~3.9kAutomated safety check: PassMIT
Trusty Searchbobmatnyc/claude-mpm155—~1.5kAutomated safety check: PassApache-2.0
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k8 repos~2.3kAutomated safety check: PassMIT

Similar skills

  • Codebase Exploration

    giancarloerra/SocratiCode

    Explore and understand codebases using SocratiCode semantic search, dependency graphs, and context artifacts.

    3.3k GitHub starsUsed in 1 repo~1.5k tokens
    DatabasesAuto-check passed
  • Symdex Code Search

    husnainpk/SymDex

    This skill should be used when finding, tracing, or understanding code in a repository with SymDex available.

    208 GitHub stars~2.3k tokensUpdated 5 mo ago
    DevelopmentAuto-check passed
  • Agentsop Repo Map

    agentsope/SkillAlchemy

    Symbol-level code context for LLM coder-agents: tree-sitter extracts symbols, PageRank ranks them over the cross-file reference graph, and the top class/function signatures are fed to the LLM as a…

    459 GitHub stars~3.9k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Trusty Search

    bobmatnyc/claude-mpm

    Hybrid code search (BM25 + vector + KG) with RRF fusion. An agent skill from bobmatnyc/claude-mpm.

    155 GitHub stars~1.5k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Analyzing Text With NLP

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to perform natural language processing and text analysis using the nlp-text-analyzer plugin.

    2.8k GitHub starsUsed in 1 repo~819 tokens
    Auto-check passed
  • Building Neural Networks

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill allows AI assistant to construct and configure neural network architectures using the neural-network-builder plugin.

    2.8k GitHub starsUsed in 1 repo~1k tokens
    Auto-check passed
  • Detecting Data Anomalies

    jeremylongshore/tons-of-skills-marketplace

    Process identify anomalies and outliers in datasets using machine learning algorithms.

    2.8k GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Explaining Machine Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill enables AI assistant to provide interpretability and explainability for machine learning models.

    2.8k GitHub starsUsed in 1 repo~1k tokens
    Auto-check passed
  • Optimizing Prompts

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill optimizes prompts for large language models (llms) to reduce token usage, lower costs, and improve performance.

    2.8k GitHub starsUsed in 1 repo~1k tokens
    Auto-check passed

Questions about Cursor Codebase Indexing

What does Cursor Codebase Indexing do?

Set up and optimize Cursor codebase indexing for semantic code search and @Codebase queries. Cursor Codebase Indexing is an agent skill from jeremylongshore/tons-of-skills-marketplace. Set up and optimize Cursor codebase indexing for semantic code search and @Codebase queries.

When should I use Cursor Codebase Indexing?

Cursor Codebase Indexing fits situations like: codebase indexing; Cursor semantic search; Cursor embeddings.

How do I install Cursor Codebase Indexing in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill cursor-codebase-indexing -a claude-code`. Or copy the skill folder (skills/.curated/cursor-codebase-indexing in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/cursor-codebase-indexing in your project. Claude Code loads it when a task matches its description.

How do I install Cursor Codebase Indexing in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill cursor-codebase-indexing -a codex`. Or copy the skill folder (skills/.curated/cursor-codebase-indexing in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/cursor-codebase-indexing in your project. Codex loads it when a task matches its description.

Can I use Cursor Codebase Indexing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill cursor-codebase-indexing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cursor-codebase-indexing, .gemini/skills/cursor-codebase-indexing, .github/skills/cursor-codebase-indexing and .opencode/skills/cursor-codebase-indexing in your project.

What does Cursor Codebase Indexing need to run?

Going by SKILL.md and its folder, Cursor Codebase Indexing needs the command-line tools its instructions call (cursor). Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(cmd:*). Compatibility (from SKILL.md): Designed for Claude Code.

Does Cursor Codebase Indexing access the network?

SKILL.md names 2 domains. As links in the text: docs.cursor.com and cursor.com. This is read from the text; nothing was executed.

Is Cursor Codebase Indexing safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file; runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Cursor Codebase Indexing use?

Cursor Codebase Indexing is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cursor Codebase Indexing use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.4k tokens, read only when the agent opens those files.

What are the alternatives to Cursor Codebase Indexing?

Skills that share tags, products or a category with Cursor Codebase Indexing: Codebase Exploration (giancarloerra/SocratiCode, 3.3k stars), Symdex Code Search (husnainpk/SymDex, 208 stars), Agentsop Repo Map (agentsope/SkillAlchemy, 459 stars) and Trusty Search (bobmatnyc/claude-mpm, 155 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cursor Codebase Indexing?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,821 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 8, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.