Langchain RAG
langchain-ai/langchain-skills
INVOKE THIS SKILL when building ANY retrieval-augmented generation (RAG) system.
Load and chunk documents for LangChain 1.0 RAG pipelines correctly — language-aware splitters, table-safe PDF loaders, Cloudflare-compatible web loaders, chunk-boundary strategies that survive…
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-data-handling -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-data-handling --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/langchain-data-handling .claude/skills/langchain-data-handling && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "langchain-data-handling" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-data-handling into .claude/skills/langchain-data-handling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-data-handling", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-data-handlingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-data-handling -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-data-handling --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/.curated/langchain-data-handling .agents/skills/langchain-data-handling && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "langchain-data-handling" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-data-handling into .agents/skills/langchain-data-handling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-data-handling", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-data-handling -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-data-handling --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/.curated/langchain-data-handling .cursor/skills/langchain-data-handling && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "langchain-data-handling" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-data-handling into .cursor/skills/langchain-data-handling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-data-handling", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jeremylongshore/tons-of-skills-marketplace.git --path skills/.curated/langchain-data-handling--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-data-handling -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-data-handling --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/.curated/langchain-data-handling .gemini/skills/langchain-data-handling && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "langchain-data-handling" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-data-handling into .gemini/skills/langchain-data-handling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-data-handling", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-data-handlingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-data-handling -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/.curated/langchain-data-handling .github/skills/langchain-data-handling && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "langchain-data-handling" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-data-handling into .github/skills/langchain-data-handling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-data-handling", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-data-handling -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-data-handling --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/.curated/langchain-data-handling .opencode/skills/langchain-data-handling && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "langchain-data-handling" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-data-handling into .opencode/skills/langchain-data-handling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-data-handling", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
langchain-data-handlingLoad and chunk documents for LangChain 1.0 RAG pipelines correctly — language-aware splitters, table-safe PDF loaders, Cloudflare-compatible web loaders, chunk-boundary strategies that survive…
Langchain Data Handling is an agent skill from jeremylongshore/tons-of-skills-marketplace. Load and chunk documents for LangChain 1.0 RAG pipelines correctly — language-aware splitters, table-safe PDF loaders, Cloudflare-compatible web loaders, chunk-boundary strategies that survive real-world structure. Use when building a RAG pipeline, diagnosing why retrieval misquotes a table, or debugging a crawler returning blank content. Trigger with "langchain document loader", "text splitter", "chunking strategy", "pdf loader", "markdown splitter", "webbaseloader".
Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/crawler-hygiene.md`, `references/language-aware-splitters.md` and `references/loader-selection-matrix.md`). Compatibility notes: Designed for Claude Code
It sits in AI & LLM Engineering, covering Building AI agents, Retrieval-augmented generation and Web scraping. It works with LangChain, Cloudflare and Python. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadWriteEditBash(python:*)Bash(pip:*)From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
python.langchain.compymupdf.readthedocs.iounstructured.iodevelopers.cloudflare.comrfc-editor.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Designed for Claude Code
From compatibility in the SKILL.md frontmatter.
Langchain Data Handling loads about 4.1k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 124 tokens; SKILL.md has 1,382 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 1,382 words, ~4,057 tokens.
.claude/skills/langchain-data-handling/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.You have a RAG system over a Python docs site. A user asks "what does
trim_messages do?" and the retriever returns this chunk:
### `trim_messages(strategy="last", include_system=True)`
Trim a message history to fit a token budget. The newest messages are kept;
older messages are dropped. Pass `include_system=True` to preserve the system...and that's it. The chunk ends there. The code example showing the function body — the actual thing the user wanted — is in a different chunk, retrieved with a lower similarity score and dropped before the LLM sees it. The model then hallucinates the function's behavior from the signature alone.
This is pain-catalog entry P13. RecursiveCharacterTextSplitter's default
separators are ["\n\n", "\n", " ", ""]. It splits on any blank line — including
inside triple-backtick code fences in Markdown. The fix is a one-line swap
to RecursiveCharacterTextSplitter.from_language(Language.MARKDOWN), which
treats the fence as an atomic unit, but you have to know the bug exists.
The sibling failures this skill prevents:
PyPDFLoader splits by page. A 5-row financial table that spans
a page break gets torn in half; rows 1-3 go in one chunk, rows 4-5 in another
with no header. A RAG answer sourced from the second chunk misquotes the
numbers because the column meanings are in the first chunk. Fix: use
PyMuPDFLoader or UnstructuredPDFLoader, which detect tables and emit
them as distinct structured elements.WebBaseLoader's default User-Agent is python-requests/2.x.
Cloudflare-protected sites flag this as a bot and return a 403 interstitial
HTML page ("Checking your browser...") instead of real content. The crawler
indexes the challenge page. You notice weeks later when every retrieval from
that source returns the same Cloudflare text. Fix: set a realistic
header_template={"User-Agent": "Mozilla/5.0 ..."}, respect robots.txt,
and rate-limit per-host to 1 req/sec.Pinned versions: langchain-core 1.0.x, langchain-community 1.0.x,
langchain-text-splitters 1.0.x, pymupdf, unstructured.
Pain-catalog anchors: P13, P49, P50, P15.
This skill is the upstream half of the RAG pipeline — load and chunk.
For the downstream half (embedding, scoring, reranking) see the pair skill
langchain-embeddings-search, which covers score semantics (P12), dim guards
(P14), and reranker filtering (P15). Do not re-implement chunking there.
langchain-core >= 1.0, < 2.0 and langchain-community >= 1.0, < 2.0langchain-text-splitters >= 1.0, < 2.0pip install pymupdf unstructured[pdf]pip install beautifulsoup4 requestspip install datasketchLoader selection is the first decision — get it wrong and no amount of splitter tuning will recover. Use the decision table:
| Source | Use | NOT | Why |
|---|---|---|---|
| PDF with tables | PyMuPDFLoader or UnstructuredPDFLoader | PyPDFLoader | Tables torn by page splits (P49) |
| PDF text-only | PyPDFLoader | — | Simple, fast, OK when no tables |
| Web page | WebBaseLoader(header_template=...) | Default UA | Cloudflare 403 (P50) |
| Markdown docs | UnstructuredMarkdownLoader | Plain text read | Preserves heading structure |
| HTML long-form | WebBaseLoader + HTMLHeaderTextSplitter | Plain text | Keeps <h1>/<h2> context |
| Code repo | GenericLoader with language parser | DirectoryLoader as text | Language-aware chunking |
| Corpus (1000+ docs) | DirectoryLoader + glob filter | One-by-one | Parallel load, progress |
from langchain_community.document_loaders import (
PyMuPDFLoader, # table-aware PDF
WebBaseLoader, # web pages (set custom UA)
UnstructuredMarkdownLoader,
DirectoryLoader,
)
# PDF with tables — P49 fix
pdf_docs = PyMuPDFLoader("10-Q-filing.pdf").load()
# Web page — P50 fix
web_docs = WebBaseLoader(
"https://example.com/article",
header_template={
"User-Agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36"
},
).load()
# Markdown docs site
md_docs = UnstructuredMarkdownLoader("docs/guide.md").load()
# Corpus
corpus = DirectoryLoader(
"./docs", glob="**/*.md",
loader_cls=UnstructuredMarkdownLoader,
show_progress=True,
).load()Hard limit: keep single-PDF ingestion under 5 MB per call. Larger files
should be pre-split with pdftk / qpdf to avoid OOM on PyMuPDFLoader's
full-document parse.
See Loader Selection Matrix for the full per-format table with cost and accuracy notes.
| Content | Splitter | chunk_size | chunk_overlap | Why |
|---|---|---|---|---|
| Prose (docs, articles) | RecursiveCharacterTextSplitter.from_language(Language.MARKDOWN) | 1000 | 100 | Preserves code fences (P13) |
| Python source | RecursiveCharacterTextSplitter.from_language(Language.PYTHON) | 1500 | 150 | Splits at def/class |
| FAQ / Q&A | RecursiveCharacterTextSplitter with separators=["\n\n"] | 500 | 50 | One chunk per Q-A pair |
| HTML long-form | HTMLHeaderTextSplitter | — | — | Headers become metadata |
| Generic text | RecursiveCharacterTextSplitter | 1000 | 100 | Safe default |
from langchain_text_splitters import (
RecursiveCharacterTextSplitter,
Language,
HTMLHeaderTextSplitter,
)
# GOOD — P13 fix for Markdown
md_splitter = RecursiveCharacterTextSplitter.from_language(
Language.MARKDOWN, chunk_size=1000, chunk_overlap=100,
)
# GOOD — Python code
py_splitter = RecursiveCharacterTextSplitter.from_language(
Language.PYTHON, chunk_size=1500, chunk_overlap=150,
)
# GOOD — HTML long-form with heading-as-metadata
html_splitter = HTMLHeaderTextSplitter(
headers_to_split_on=[("h1", "Header 1"), ("h2", "Header 2")],
)
# BAD — breaks inside code fences (P13)
bad = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=100)See Language-Aware Splitters for the
full list of Language.* enum values, custom separator patterns, and the
code-fence-detection regex for when you need a custom splitter.
Defaults from the table work for most corpora. Tune when:
chunk_size (1000 → 1500) or
chunk_overlap (100 → 200). Overlap is what bridges a concept that crosses
chunk boundaries.chunk_size (1000 → 500).
Smaller chunks = more precise retrieval but more chunks to index.A 1% overlap-to-size ratio is too low (200/20000); 20% is the sweet spot for most prose. Code needs less overlap (10%) because function boundaries are natural splits.
Tables are not text. If your corpus has financial filings, product specs, or any tabular data, index tables as separate records with column metadata:
import fitz # pymupdf directly for table detection
def extract_tables_as_records(pdf_path: str) -> list[dict]:
"""Extract tables as one record per row."""
doc = fitz.open(pdf_path)
records = []
for page_num, page in enumerate(doc):
tables = page.find_tables()
for table in tables:
rows = table.extract()
if not rows:
continue
headers = rows[0]
for row_idx, row in enumerate(rows[1:], start=1):
record = {
"page": page_num,
"table_idx": tables.tables.index(table),
"row_idx": row_idx,
"content": " | ".join(f"{h}: {v}" for h, v in zip(headers, row)),
"metadata": dict(zip(headers, row)),
}
records.append(record)
return recordsNow a question like "what was Q3 revenue?" retrieves a single row with its column headers attached, not half a table missing the column meanings. See Table Preservation for the full pattern including hybrid retrieval (prose + table records).
The loader attaches metadata (source, page, heading); the splitter propagates
it. Front-matter in Markdown, PDF page numbers, and web URLs should all end
up in doc.metadata so retrieval results are citable:
for doc in md_docs:
# Markdown front-matter (if loader extracted it)
print(doc.metadata.get("title"), doc.metadata.get("date"))
# Splitter-preserved metadata
chunks = md_splitter.split_documents(md_docs)
assert chunks[0].metadata == md_docs[0].metadata # preservedCustom metadata (tenant_id, version, confidence) should be added before splitting so every chunk inherits it.
Web crawls and scraped docs often contain near-duplicate pages (nav chrome, footer boilerplate, syndicated posts). MinHash-based dedup at the chunk level keeps the index clean:
from datasketch import MinHash, MinHashLSH
lsh = MinHashLSH(threshold=0.9, num_perm=128)
kept = []
for i, chunk in enumerate(all_chunks):
mh = MinHash(num_perm=128)
for tok in chunk.page_content.lower().split():
mh.update(tok.encode())
if not list(lsh.query(mh)):
lsh.insert(str(i), mh)
kept.append(chunk)A threshold of 0.9 catches near-duplicates (minor wording differences) without eating legitimate paraphrases.
# Multi-stage: load → split → dedup → index
def build_rag_index(source_dir: str, store):
# 1. Load
docs = DirectoryLoader(
source_dir, glob="**/*.md",
loader_cls=UnstructuredMarkdownLoader,
).load()
# 2. Clean (empty-content filter)
docs = [d for d in docs if d.page_content.strip()]
# 3. Split (language-aware)
splitter = RecursiveCharacterTextSplitter.from_language(
Language.MARKDOWN, chunk_size=1000, chunk_overlap=100,
)
chunks = splitter.split_documents(docs)
# 4. Dedup (optional for noisy corpora)
# chunks = dedup_minhash(chunks, threshold=0.9)
# 5. Index — handoff to langchain-embeddings-search
store.add_documents(chunks)
return storeFor the embedding + indexing + retrieval steps, see langchain-embeddings-search.
robots.txt respect| Error / symptom | Cause | Fix |
|---|---|---|
| RAG retrieves function signature without body | RecursiveCharacterTextSplitter broke inside code fence (P13) | Use from_language(Language.MARKDOWN) or add "```" as first separator |
| Table rows misquoted in RAG answer | PyPDFLoader tore table by page (P49) | Switch to PyMuPDFLoader; index tables as structured records |
WebBaseLoader returns 403 / blank content | Default UA flagged by Cloudflare (P50) | Set header_template={"User-Agent": "Mozilla/5.0 ..."}; respect robots.txt |
ValueError: expected str, NoneType found during split | Empty page_content | Filter [d for d in docs if d.page_content.strip()] before splitting |
MemoryError loading PDF | PDF > 5 MB ingested in one call | Pre-split with pdftk / qpdf; process chunks separately |
| Chunks missing metadata after split | Custom metadata added after loading but before splitting was lost | Add metadata before split_documents(); verify chunks[0].metadata preserved |
| Retrieval quality low on FAQ corpus | Chunks too large, one chunk holds multiple Q-A pairs | Drop to chunk_size=500, chunk_overlap=50 with separators=["\n\n"] |
| Web crawl indexes Cloudflare challenge page | No check for HTTP status / response length | Assert len(doc.page_content) > 500 and reject pages containing "Checking your browser" |
| Duplicate chunks eat retrieval slots | Syndicated content, nav chrome not stripped | MinHash dedup at threshold 0.9 before indexing |
| Reranker scores inconsistent across chunks | Chunks of wildly different size change score distribution (P15) | Normalize chunk size within a corpus; target ±20% of chunk_size |
Markdown docs with Python code fences require Language.MARKDOWN to keep
fence boundaries intact. Chunk size 1000 with 100 overlap preserves one
function-sized example per chunk. Front-matter fields (title, date, author)
are attached as metadata for citation. See
Language-Aware Splitters.
10-Q filings have dozens of multi-row tables. Use PyMuPDFLoader for the prose
and a direct fitz.find_tables() pass to extract tables as structured records.
Index prose with chunk_size=1000 and tables as one-row-per-record with the
header row concatenated. Questions like "what was Q3 revenue?" hit a single row
with column meanings attached. See
Table Preservation.
Set a realistic User-Agent, fetch robots.txt first and respect Disallow
rules, rate-limit to 1 req/sec per host, and prefer the site's sitemap or RSS
feed when available. Assert response length > 500 chars and reject known
interstitial patterns. See Crawler Hygiene.
GenericLoader with LanguageParser(language=Language.PYTHON) preserves
function and class boundaries. Chunk size 1500 with 150 overlap gives enough
context for typical function-level queries. Imports and module docstrings
end up in their own chunks — tag them with metadata for higher precision
retrieval on "where is X imported from" queries.
page.find_tables()langchain-embeddings-search (embed/score/rerank — downstream of this skill)docs/pain-catalog.md (entries P13, P49, P50, P15)© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (references) in skills/.curated/langchain-data-handling of jeremylongshore/tons-of-skills-marketplace.
Open the folder on GitHubat commit cfae287
Langchain Data Handling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Langchain Data Handling this skilljeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~4.1k | Automated safety check: Pass | MIT | |
| Langchain RAGlangchain-ai/langchain-skills | 1.3k | — | ~3.9k | Automated safety check: Pass | MIT | |
| Neo4j Graphrag Skillneo4j-contrib/neo4j-skills | 114 | — | ~4.2k | Automated safety check: Notes | MIT | |
| Add Example AgentGetBindu/Bindu | 10k | — | ~1.1k | Automated safety check: Notes | Custom licence | |
| Failproof AI SDK IntegrationFailproofAI/failproofai | 5.3k | — | ~6k | Automated safety check: Pass | Custom licence | |
| SynalinksSynaLinks/synalinks-skills | 907 | — | ~4.8k | Automated safety check: Pass | Apache-2.0 |
langchain-ai/langchain-skills
INVOKE THIS SKILL when building ANY retrieval-augmented generation (RAG) system.
neo4j-contrib/neo4j-skills
Build GraphRAG retrieval pipelines on Neo4j using the neo4j-graphrag Python package (v1.22.0+).
GetBindu/Bindu
Add a new self-contained example agent under examples/. An agent skill from GetBindu/Bindu.
FailproofAI/failproofai
Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.
SynaLinks/synalinks-skills
A skill your agent uses for anything involving the Synalinks neuro-symbolic LM framework (Keras-inspired): DataModel/Field/Input, JSON operators (+ & | ^ ~), synalinks.ops…
andrewyng/context-hub
Guides building Tavily integrations for web search, URL extraction, site crawling and AI-assisted research in Python or JavaScript agent and RAG projects.
jeremylongshore/tons-of-skills-marketplace
Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.
jeremylongshore/tons-of-skills-marketplace
Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.
jeremylongshore/tons-of-skills-marketplace
Execute proactive auto-loading: automatically detects and loads agents.md files.
jeremylongshore/tons-of-skills-marketplace
Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.
jeremylongshore/tons-of-skills-marketplace
Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.
jeremylongshore/tons-of-skills-marketplace
Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.
Works with
Categories
Load and chunk documents for LangChain 1.0 RAG pipelines correctly — language-aware splitters, table-safe PDF loaders, Cloudflare-compatible web loaders, chunk-boundary strategies that survive…. Langchain Data Handling is an agent skill from jeremylongshore/tons-of-skills-marketplace.0 RAG pipelines correctly — language-aware splitters, table-safe PDF loaders, Cloudflare-compatible web loaders, chunk-boundary strategies that survive real-world structure.
Langchain Data Handling fits situations like: building a RAG pipeline; diagnosing why retrieval misquotes a table; debugging a crawler returning blank content; with langchain document loader.
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-data-handling -a claude-code`. Or copy the skill folder (skills/.curated/langchain-data-handling in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/langchain-data-handling in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-data-handling -a codex`. Or copy the skill folder (skills/.curated/langchain-data-handling in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/langchain-data-handling in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-data-handling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/langchain-data-handling, .gemini/skills/langchain-data-handling, .github/skills/langchain-data-handling and .opencode/skills/langchain-data-handling in your project.
Going by SKILL.md and its folder, Langchain Data Handling needs the command-line tools its instructions call (pip). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(python:*), Bash(pip:*). Compatibility (from SKILL.md): Designed for Claude Code.
SKILL.md names 5 domains. As links in the text: python.langchain.com, pymupdf.readthedocs.io, unstructured.io, developers.cloudflare.com and rfc-editor.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Langchain Data Handling is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.1k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Langchain Data Handling: Langchain RAG (langchain-ai/langchain-skills, 1.3k stars), Neo4j Graphrag Skill (neo4j-contrib/neo4j-skills, 114 stars), Add Example Agent (GetBindu/Bindu, 10k stars) and Failproof AI SDK Integration (FailproofAI/failproofai, 5.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.
Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.