Langchain RAG
langchain-ai/langchain-skills
INVOKE THIS SKILL when building ANY retrieval-augmented generation (RAG) system.
Agent skill
by Orchestra-Research in Orchestra-Research/AI-Research-SKILLs
Generates text embeddings locally with the sentence-transformers library for RAG, semantic search, clustering and similarity, with model picks for general, multilingual and legal text.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill sentence-transformers -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs sentence-transformers --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/15-rag/sentence-transformers .claude/skills/sentence-transformers && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "sentence-transformers" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/15-rag/sentence-transformers into .claude/skills/sentence-transformers/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sentence-transformers", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/15-rag/sentence-transformersType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill sentence-transformers -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs sentence-transformers --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .agents/skills && cp -r skills-src/15-rag/sentence-transformers .agents/skills/sentence-transformers && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "sentence-transformers" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/15-rag/sentence-transformers into .agents/skills/sentence-transformers/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sentence-transformers", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill sentence-transformers -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs sentence-transformers --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/15-rag/sentence-transformers .cursor/skills/sentence-transformers && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "sentence-transformers" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/15-rag/sentence-transformers into .cursor/skills/sentence-transformers/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sentence-transformers", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Orchestra-Research/AI-Research-SKILLs.git --path 15-rag/sentence-transformers--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill sentence-transformers -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs sentence-transformers --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/15-rag/sentence-transformers .gemini/skills/sentence-transformers && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "sentence-transformers" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/15-rag/sentence-transformers into .gemini/skills/sentence-transformers/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sentence-transformers", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Orchestra-Research/AI-Research-SKILLs sentence-transformersInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill sentence-transformers -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .github/skills && cp -r skills-src/15-rag/sentence-transformers .github/skills/sentence-transformers && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "sentence-transformers" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/15-rag/sentence-transformers into .github/skills/sentence-transformers/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sentence-transformers", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill sentence-transformers -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs sentence-transformers --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/15-rag/sentence-transformers .opencode/skills/sentence-transformers && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "sentence-transformers" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/15-rag/sentence-transformers into .opencode/skills/sentence-transformers/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sentence-transformers", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
sentence-transformersGenerates text embeddings locally with the sentence-transformers library for RAG, semantic search, clustering and similarity, with model picks for general, multilingual and legal text.
Sentence Transformers is a Python framework that turns sentences and other text into embedding vectors without calling a hosted API. The skill covers installing it, loading a SentenceTransformer model, encoding text singly or in large batches, running semantic search with the util helpers, and scoring cosine similarity between embeddings. It names starter models such as all-MiniLM-L6-v2 for fast general use, a multilingual MiniLM variant, and a legal-domain BERT model.
It also shows fine-tuning with InputExample objects, losses and a PyTorch DataLoader, plus wrappers that plug the embeddings into LangChain and LlamaIndex. A model selection table compares dimensions, speed and quality. OpenAI embeddings, Instructor and Cohere Embed are listed as alternatives when you want an API or task-specific instructions. A reference file lists more models; the excerpt is cut off inside the selection table.
8 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comhuggingface.cosbert.netFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Sentence Transformers Embeddings loads about 1.6k tokens when it runs, and up to ~2.3k if it reads all its reference files. Until then it costs about 90 tokens; SKILL.md has 216 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 216 words, ~1,581 tokens.
.claude/skills/sentence-transformers/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Python framework for sentence and text embeddings using transformers.
Use when:
Metrics:
Use alternatives instead:
pip install sentence-transformersfrom sentence_transformers import SentenceTransformer
# Load model
model = SentenceTransformer('all-MiniLM-L6-v2')
# Generate embeddings
sentences = [
"This is an example sentence",
"Each sentence is converted to a vector"
]
embeddings = model.encode(sentences)
print(embeddings.shape) # (2, 384)
# Cosine similarity
from sentence_transformers.util import cos_sim
similarity = cos_sim(embeddings[0], embeddings[1])
print(f"Similarity: {similarity.item():.4f}")# Fast, good quality (384 dim)
model = SentenceTransformer('all-MiniLM-L6-v2')
# Better quality (768 dim)
model = SentenceTransformer('all-mpnet-base-v2')
# Best quality (1024 dim, slower)
model = SentenceTransformer('all-roberta-large-v1')# 50+ languages
model = SentenceTransformer('paraphrase-multilingual-MiniLM-L12-v2')
# 100+ languages
model = SentenceTransformer('paraphrase-multilingual-mpnet-base-v2')# Legal domain
model = SentenceTransformer('nlpaueb/legal-bert-base-uncased')
# Scientific papers
model = SentenceTransformer('allenai/specter')
# Code
model = SentenceTransformer('microsoft/codebert-base')from sentence_transformers import SentenceTransformer, util
model = SentenceTransformer('all-MiniLM-L6-v2')
# Corpus
corpus = [
"Python is a programming language",
"Machine learning uses algorithms",
"Neural networks are powerful"
]
# Encode corpus
corpus_embeddings = model.encode(corpus, convert_to_tensor=True)
# Query
query = "What is Python?"
query_embedding = model.encode(query, convert_to_tensor=True)
# Find most similar
hits = util.semantic_search(query_embedding, corpus_embeddings, top_k=3)
print(hits)# Cosine similarity
similarity = util.cos_sim(embedding1, embedding2)
# Dot product
similarity = util.dot_score(embedding1, embedding2)
# Pairwise cosine similarity
similarities = util.cos_sim(embeddings, embeddings)# Efficient batch processing
sentences = ["sentence 1", "sentence 2", ...] * 1000
embeddings = model.encode(
sentences,
batch_size=32,
show_progress_bar=True,
convert_to_tensor=False # or True for PyTorch tensors
)from sentence_transformers import InputExample, losses
from torch.utils.data import DataLoader
# Training data
train_examples = [
InputExample(texts=['sentence 1', 'sentence 2'], label=0.8),
InputExample(texts=['sentence 3', 'sentence 4'], label=0.3),
]
train_dataloader = DataLoader(train_examples, batch_size=16)
# Loss function
train_loss = losses.CosineSimilarityLoss(model)
# Train
model.fit(
train_objectives=[(train_dataloader, train_loss)],
epochs=10,
warmup_steps=100
)
# Save
model.save('my-finetuned-model')from langchain_community.embeddings import HuggingFaceEmbeddings
embeddings = HuggingFaceEmbeddings(
model_name="sentence-transformers/all-mpnet-base-v2"
)
# Use with vector stores
from langchain_chroma import Chroma
vectorstore = Chroma.from_documents(
documents=docs,
embedding=embeddings
)from llama_index.embeddings.huggingface import HuggingFaceEmbedding
embed_model = HuggingFaceEmbedding(
model_name="sentence-transformers/all-mpnet-base-v2"
)
from llama_index.core import Settings
Settings.embed_model = embed_model
# Use in index
index = VectorStoreIndex.from_documents(documents)| Model | Dimensions | Speed | Quality | Use Case |
|---|---|---|---|---|
| all-MiniLM-L6-v2 | 384 | Fast | Good | General, prototyping |
| all-mpnet-base-v2 | 768 | Medium | Better | Production RAG |
| all-roberta-large-v1 | 1024 | Slow | Best | High accuracy needed |
| paraphrase-multilingual | 768 | Medium | Good | Multilingual |
| Model | Speed (sentences/sec) | Memory | Dimension |
|---|---|---|---|
| MiniLM | ~2000 | 120MB | 384 |
| MPNet | ~600 | 420MB | 768 |
| RoBERTa | ~300 | 1.3GB | 1024 |
© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in 15-rag/sentence-transformers of Orchestra-Research/AI-Research-SKILLs.
Open the folder on GitHubat commit 773a529
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.
Sentence Transformers Embeddings next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Sentence Transformers Embeddings this skillOrchestra-Research/AI-Research-SKILLs | 13k | 2 repos | ~1.6k | Automated safety check: Pass | MIT | |
| Langchain RAGlangchain-ai/langchain-skills | 1.3k | — | ~3.9k | Automated safety check: Pass | MIT | |
| RAG Skillsllama-farm/llamafarm | 836 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Discover MLrand/cc-polymath | 181 | — | ~574 | Automated safety check: Pass | MIT | |
| Neo4j Graphrag Skillneo4j-contrib/neo4j-skills | 114 | — | ~4.2k | Automated safety check: Notes | MIT | |
| Hugging Face Transformers Usagedavila7/claude-code-templates | 33k | 11 repos | ~1.2k | Automated safety check: Pass | MIT |
langchain-ai/langchain-skills
INVOKE THIS SKILL when building ANY retrieval-augmented generation (RAG) system.
llama-farm/llamafarm
RAG-specific best practices for LlamaIndex, ChromaDB, and Celery workers.
rand/cc-polymath
Automatically discover machine learning and AI skills when working with machine learning, PyTorch, training, inference, RAG, embeddings, fine-tuning, LLM, DSPy, HuggingFace, or diffusion models.
neo4j-contrib/neo4j-skills
Build GraphRAG retrieval pipelines on Neo4j using the neo4j-graphrag Python package (v1.22.0+).
davila7/claude-code-templates
Loads pre-trained Hugging Face Transformers models for text, vision and audio tasks, runs inference with pipelines and fine-tunes on custom datasets.
Jeffallan/claude-skills
Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.
Orchestra-Research/AI-Research-SKILLs
Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
Orchestra-Research/AI-Research-SKILLs
Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.
Orchestra-Research/AI-Research-SKILLs
Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.
Categories
Generates text embeddings locally with the sentence-transformers library for RAG, semantic search, clustering and similarity, with model picks for general, multilingual and legal text. Sentence Transformers is a Python framework that turns sentences and other text into embedding vectors without calling a hosted API. The skill covers installing it, loading a SentenceTransformer model, encoding text singly or in large batches, running semantic search with the util helpers, and scoring cosine similarity between embeddings.
Sentence Transformers Embeddings fits situations like: embedding documents for a RAG system without paying for an API; finding semantically similar sentences or duplicate questions; clustering or classifying text by meaning; fine-tuning an embedding model on your own sentence pairs.
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill sentence-transformers -a claude-code`. Or copy the skill folder (15-rag/sentence-transformers in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/sentence-transformers in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill sentence-transformers -a codex`. Or copy the skill folder (15-rag/sentence-transformers in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/sentence-transformers in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill sentence-transformers -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sentence-transformers, .gemini/skills/sentence-transformers, .github/skills/sentence-transformers and .opencode/skills/sentence-transformers in your project.
Going by SKILL.md and its folder, Sentence Transformers Embeddings needs the command-line tools its instructions call (pip). Our summary lists: Python with the `sentence-transformers` package.
SKILL.md names 3 domains. As links in the text: github.com, huggingface.co and sbert.net. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Sentence Transformers Embeddings is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 732 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Sentence Transformers Embeddings: Langchain RAG (langchain-ai/langchain-skills, 1.3k stars), RAG Skills (llama-farm/llamafarm, 836 stars), Discover ML (rand/cc-polymath, 181 stars) and Neo4j Graphrag Skill (neo4j-contrib/neo4j-skills, 114 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,405 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.
Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.