Agent skill

Langchain Data Handling

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Load and chunk documents for LangChain 1.0 RAG pipelines correctly — language-aware splitters, table-safe PDF loaders, Cloudflare-compatible web loaders, chunk-boundary strategies that survive…

MITAuto-check passedAI & LLM Engineering

Install Langchain Data Handling

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-data-handling -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-data-handling --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/langchain-data-handling .claude/skills/langchain-data-handling && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
langchain-data-handling
GitHub stars
2.8k
Token cost
~4.1k tokens
SKILL.md length
1,382 words
Files
6 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Load and chunk documents for LangChain 1.0 RAG pipelines correctly — language-aware splitters, table-safe PDF loaders, Cloudflare-compatible web loaders, chunk-boundary strategies that survive…

  • Works in 7 steps: Choose a loader by source format → Pick a splitter by content type → Tune chunk_size and overlap → …
  • Building a RAG pipeline
  • SKILL.md covers Overview, Prerequisites, Instructions and Output, plus 3 more sections
  • Calls pip

What it does

Langchain Data Handling is an agent skill from jeremylongshore/tons-of-skills-marketplace. Load and chunk documents for LangChain 1.0 RAG pipelines correctly — language-aware splitters, table-safe PDF loaders, Cloudflare-compatible web loaders, chunk-boundary strategies that survive real-world structure. Use when building a RAG pipeline, diagnosing why retrieval misquotes a table, or debugging a crawler returning blank content. Trigger with "langchain document loader", "text splitter", "chunking strategy", "pdf loader", "markdown splitter", "webbaseloader".

Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/crawler-hygiene.md`, `references/language-aware-splitters.md` and `references/loader-selection-matrix.md`). Compatibility notes: Designed for Claude Code

It sits in AI & LLM Engineering, covering Building AI agents, Retrieval-augmented generation and Web scraping. It works with LangChain, Cloudflare and Python. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Building a RAG pipeline
  • Diagnosing why retrieval misquotes a table
  • Debugging a crawler returning blank content
  • With langchain document loader

Example prompts

  • “langchain document loader”
  • “text splitter”
  • “chunking strategy”
  • “/langchain-data-handling”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash(python:*), Bash(pip:*)

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Choose a loader by source format
  2. Pick a splitter by content type
  3. Tune chunk_size and overlap
  4. Detect and index tables as structured records
  5. Preserve metadata through the pipeline
  6. Deduplicate noisy corpora
  7. Compose the pipeline

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash(python:*)
    • Bash(pip:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • python.langchain.com
    • pymupdf.readthedocs.io
    • unstructured.io
    • developers.cloudflare.com
    • rfc-editor.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Langchain Data Handling loads about 4.1k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 124 tokens; SKILL.md has 1,382 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~124
When it runs · the whole SKILL.md, loaded when a task matches
~4.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~12k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 1,382 words, ~4,057 tokens.

Download SKILL.mdSave it as .claude/skills/langchain-data-handling/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
langchain-data-handling
description
Load and chunk documents for LangChain 1.0 RAG pipelines correctly — language-aware splitters, table-safe PDF loaders, Cloudflare-compatible web loaders, chunk-boundary strategies that survive real-world structure. Use when building a RAG pipeline, diagnosing why retrieval misquotes a table, or debugging a crawler returning blank content. Trigger with "langchain document loader", "text splitter", "chunking strategy", "pdf loader", "markdown splitter", "webbaseloader".
allowed-tools
Read, Write, Edit, Bash(python:*), Bash(pip:*)
compatibility
Designed for Claude Code
version
2.7.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, langchain, langgraph, python, langchain-1.0, document-loaders, text-splitters, rag

LangChain Data Handling — Loaders and Splitters (Python)

Overview

You have a RAG system over a Python docs site. A user asks "what does trim_messages do?" and the retriever returns this chunk:

### `trim_messages(strategy="last", include_system=True)`

Trim a message history to fit a token budget. The newest messages are kept;
older messages are dropped. Pass `include_system=True` to preserve the system

...and that's it. The chunk ends there. The code example showing the function body — the actual thing the user wanted — is in a different chunk, retrieved with a lower similarity score and dropped before the LLM sees it. The model then hallucinates the function's behavior from the signature alone.

This is pain-catalog entry P13. RecursiveCharacterTextSplitter's default separators are ["\n\n", "\n", " ", ""]. It splits on any blank line — including inside triple-backtick code fences in Markdown. The fix is a one-line swap to RecursiveCharacterTextSplitter.from_language(Language.MARKDOWN), which treats the fence as an atomic unit, but you have to know the bug exists.

The sibling failures this skill prevents:

  • P49 — PyPDFLoader splits by page. A 5-row financial table that spans a page break gets torn in half; rows 1-3 go in one chunk, rows 4-5 in another with no header. A RAG answer sourced from the second chunk misquotes the numbers because the column meanings are in the first chunk. Fix: use PyMuPDFLoader or UnstructuredPDFLoader, which detect tables and emit them as distinct structured elements.
  • P50 — WebBaseLoader's default User-Agent is python-requests/2.x. Cloudflare-protected sites flag this as a bot and return a 403 interstitial HTML page ("Checking your browser...") instead of real content. The crawler indexes the challenge page. You notice weeks later when every retrieval from that source returns the same Cloudflare text. Fix: set a realistic header_template={"User-Agent": "Mozilla/5.0 ..."}, respect robots.txt, and rate-limit per-host to 1 req/sec.

Pinned versions: langchain-core 1.0.x, langchain-community 1.0.x, langchain-text-splitters 1.0.x, pymupdf, unstructured. Pain-catalog anchors: P13, P49, P50, P15.

This skill is the upstream half of the RAG pipeline — load and chunk. For the downstream half (embedding, scoring, reranking) see the pair skill langchain-embeddings-search, which covers score semantics (P12), dim guards (P14), and reranker filtering (P15). Do not re-implement chunking there.

Prerequisites

  • Python 3.10+
  • langchain-core >= 1.0, < 2.0 and langchain-community >= 1.0, < 2.0
  • langchain-text-splitters >= 1.0, < 2.0
  • PDF support: pip install pymupdf unstructured[pdf]
  • Web loading: pip install beautifulsoup4 requests
  • For corpus dedup (optional): pip install datasketch

Instructions

Step 1 — Choose a loader by source format

Loader selection is the first decision — get it wrong and no amount of splitter tuning will recover. Use the decision table:

SourceUseNOTWhy
PDF with tablesPyMuPDFLoader or UnstructuredPDFLoaderPyPDFLoaderTables torn by page splits (P49)
PDF text-onlyPyPDFLoader—Simple, fast, OK when no tables
Web pageWebBaseLoader(header_template=...)Default UACloudflare 403 (P50)
Markdown docsUnstructuredMarkdownLoaderPlain text readPreserves heading structure
HTML long-formWebBaseLoader + HTMLHeaderTextSplitterPlain textKeeps <h1>/<h2> context
Code repoGenericLoader with language parserDirectoryLoader as textLanguage-aware chunking
Corpus (1000+ docs)DirectoryLoader + glob filterOne-by-oneParallel load, progress
python
from langchain_community.document_loaders import (
    PyMuPDFLoader,            # table-aware PDF
    WebBaseLoader,            # web pages (set custom UA)
    UnstructuredMarkdownLoader,
    DirectoryLoader,
)

# PDF with tables — P49 fix
pdf_docs = PyMuPDFLoader("10-Q-filing.pdf").load()

# Web page — P50 fix
web_docs = WebBaseLoader(
    "https://example.com/article",
    header_template={
        "User-Agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36"
    },
).load()

# Markdown docs site
md_docs = UnstructuredMarkdownLoader("docs/guide.md").load()

# Corpus
corpus = DirectoryLoader(
    "./docs", glob="**/*.md",
    loader_cls=UnstructuredMarkdownLoader,
    show_progress=True,
).load()

Hard limit: keep single-PDF ingestion under 5 MB per call. Larger files should be pre-split with pdftk / qpdf to avoid OOM on PyMuPDFLoader's full-document parse.

See Loader Selection Matrix for the full per-format table with cost and accuracy notes.

Step 2 — Pick a splitter by content type
ContentSplitterchunk_sizechunk_overlapWhy
Prose (docs, articles)RecursiveCharacterTextSplitter.from_language(Language.MARKDOWN)1000100Preserves code fences (P13)
Python sourceRecursiveCharacterTextSplitter.from_language(Language.PYTHON)1500150Splits at def/class
FAQ / Q&ARecursiveCharacterTextSplitter with separators=["\n\n"]50050One chunk per Q-A pair
HTML long-formHTMLHeaderTextSplitter——Headers become metadata
Generic textRecursiveCharacterTextSplitter1000100Safe default
python
from langchain_text_splitters import (
    RecursiveCharacterTextSplitter,
    Language,
    HTMLHeaderTextSplitter,
)

# GOOD — P13 fix for Markdown
md_splitter = RecursiveCharacterTextSplitter.from_language(
    Language.MARKDOWN, chunk_size=1000, chunk_overlap=100,
)

# GOOD — Python code
py_splitter = RecursiveCharacterTextSplitter.from_language(
    Language.PYTHON, chunk_size=1500, chunk_overlap=150,
)

# GOOD — HTML long-form with heading-as-metadata
html_splitter = HTMLHeaderTextSplitter(
    headers_to_split_on=[("h1", "Header 1"), ("h2", "Header 2")],
)

# BAD — breaks inside code fences (P13)
bad = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=100)

See Language-Aware Splitters for the full list of Language.* enum values, custom separator patterns, and the code-fence-detection regex for when you need a custom splitter.

Step 3 — Tune chunk_size and overlap

Defaults from the table work for most corpora. Tune when:

  • Retrieval misses context: increase chunk_size (1000 → 1500) or chunk_overlap (100 → 200). Overlap is what bridges a concept that crosses chunk boundaries.
  • Retrieval too broad, answers wander: decrease chunk_size (1000 → 500). Smaller chunks = more precise retrieval but more chunks to index.
  • Tables / structured data: do NOT tune — index them separately (step 4).

A 1% overlap-to-size ratio is too low (200/20000); 20% is the sweet spot for most prose. Code needs less overlap (10%) because function boundaries are natural splits.

Step 4 — Detect and index tables as structured records

Tables are not text. If your corpus has financial filings, product specs, or any tabular data, index tables as separate records with column metadata:

python
import fitz  # pymupdf directly for table detection

def extract_tables_as_records(pdf_path: str) -> list[dict]:
    """Extract tables as one record per row."""
    doc = fitz.open(pdf_path)
    records = []
    for page_num, page in enumerate(doc):
        tables = page.find_tables()
        for table in tables:
            rows = table.extract()
            if not rows:
                continue
            headers = rows[0]
            for row_idx, row in enumerate(rows[1:], start=1):
                record = {
                    "page": page_num,
                    "table_idx": tables.tables.index(table),
                    "row_idx": row_idx,
                    "content": " | ".join(f"{h}: {v}" for h, v in zip(headers, row)),
                    "metadata": dict(zip(headers, row)),
                }
                records.append(record)
    return records

Now a question like "what was Q3 revenue?" retrieves a single row with its column headers attached, not half a table missing the column meanings. See Table Preservation for the full pattern including hybrid retrieval (prose + table records).

Step 5 — Preserve metadata through the pipeline

The loader attaches metadata (source, page, heading); the splitter propagates it. Front-matter in Markdown, PDF page numbers, and web URLs should all end up in doc.metadata so retrieval results are citable:

python
for doc in md_docs:
    # Markdown front-matter (if loader extracted it)
    print(doc.metadata.get("title"), doc.metadata.get("date"))

# Splitter-preserved metadata
chunks = md_splitter.split_documents(md_docs)
assert chunks[0].metadata == md_docs[0].metadata  # preserved

Custom metadata (tenant_id, version, confidence) should be added before splitting so every chunk inherits it.

Step 6 — Deduplicate noisy corpora

Web crawls and scraped docs often contain near-duplicate pages (nav chrome, footer boilerplate, syndicated posts). MinHash-based dedup at the chunk level keeps the index clean:

python
from datasketch import MinHash, MinHashLSH

lsh = MinHashLSH(threshold=0.9, num_perm=128)
kept = []
for i, chunk in enumerate(all_chunks):
    mh = MinHash(num_perm=128)
    for tok in chunk.page_content.lower().split():
        mh.update(tok.encode())
    if not list(lsh.query(mh)):
        lsh.insert(str(i), mh)
        kept.append(chunk)

A threshold of 0.9 catches near-duplicates (minor wording differences) without eating legitimate paraphrases.

Show full SKILL.md (546 more words)Show less
Step 7 — Compose the pipeline
python
# Multi-stage: load → split → dedup → index
def build_rag_index(source_dir: str, store):
    # 1. Load
    docs = DirectoryLoader(
        source_dir, glob="**/*.md",
        loader_cls=UnstructuredMarkdownLoader,
    ).load()

    # 2. Clean (empty-content filter)
    docs = [d for d in docs if d.page_content.strip()]

    # 3. Split (language-aware)
    splitter = RecursiveCharacterTextSplitter.from_language(
        Language.MARKDOWN, chunk_size=1000, chunk_overlap=100,
    )
    chunks = splitter.split_documents(docs)

    # 4. Dedup (optional for noisy corpora)
    # chunks = dedup_minhash(chunks, threshold=0.9)

    # 5. Index — handoff to langchain-embeddings-search
    store.add_documents(chunks)
    return store

For the embedding + indexing + retrieval steps, see langchain-embeddings-search.

Output

  • Loader chosen from the selection matrix matching source format and table needs
  • Splitter chosen from the decision tree matching content type
  • Chunk size + overlap tuned from the defaults (1000/100 prose, 1500/150 code, 500/50 FAQ)
  • Tables extracted as structured records with column metadata (not text chunks)
  • Web loaders configured with realistic User-Agent and robots.txt respect
  • Metadata preserved through loader → splitter → index
  • Optional MinHash dedup (threshold 0.9) for noisy corpora

Error Handling

Error / symptomCauseFix
RAG retrieves function signature without bodyRecursiveCharacterTextSplitter broke inside code fence (P13)Use from_language(Language.MARKDOWN) or add "```" as first separator
Table rows misquoted in RAG answerPyPDFLoader tore table by page (P49)Switch to PyMuPDFLoader; index tables as structured records
WebBaseLoader returns 403 / blank contentDefault UA flagged by Cloudflare (P50)Set header_template={"User-Agent": "Mozilla/5.0 ..."}; respect robots.txt
ValueError: expected str, NoneType found during splitEmpty page_contentFilter [d for d in docs if d.page_content.strip()] before splitting
MemoryError loading PDFPDF > 5 MB ingested in one callPre-split with pdftk / qpdf; process chunks separately
Chunks missing metadata after splitCustom metadata added after loading but before splitting was lostAdd metadata before split_documents(); verify chunks[0].metadata preserved
Retrieval quality low on FAQ corpusChunks too large, one chunk holds multiple Q-A pairsDrop to chunk_size=500, chunk_overlap=50 with separators=["\n\n"]
Web crawl indexes Cloudflare challenge pageNo check for HTTP status / response lengthAssert len(doc.page_content) > 500 and reject pages containing "Checking your browser"
Duplicate chunks eat retrieval slotsSyndicated content, nav chrome not strippedMinHash dedup at threshold 0.9 before indexing
Reranker scores inconsistent across chunksChunks of wildly different size change score distribution (P15)Normalize chunk size within a corpus; target ±20% of chunk_size

Examples

Ingesting a Markdown docs site with code examples

Markdown docs with Python code fences require Language.MARKDOWN to keep fence boundaries intact. Chunk size 1000 with 100 overlap preserves one function-sized example per chunk. Front-matter fields (title, date, author) are attached as metadata for citation. See Language-Aware Splitters.

Ingesting a PDF filing with financial tables

10-Q filings have dozens of multi-row tables. Use PyMuPDFLoader for the prose and a direct fitz.find_tables() pass to extract tables as structured records. Index prose with chunk_size=1000 and tables as one-row-per-record with the header row concatenated. Questions like "what was Q3 revenue?" hit a single row with column meanings attached. See Table Preservation.

Crawling a documentation site behind Cloudflare

Set a realistic User-Agent, fetch robots.txt first and respect Disallow rules, rate-limit to 1 req/sec per host, and prefer the site's sitemap or RSS feed when available. Assert response length > 500 chars and reject known interstitial patterns. See Crawler Hygiene.

Ingesting a Python code repo for code RAG

GenericLoader with LanguageParser(language=Language.PYTHON) preserves function and class boundaries. Chunk size 1500 with 150 overlap gives enough context for typical function-level queries. Imports and module docstrings end up in their own chunks — tag them with metadata for higher precision retrieval on "where is X imported from" queries.

Resources

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/.curated/langchain-data-handling of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/crawler-hygiene.md
  • references/language-aware-splitters.md
  • references/loader-selection-matrix.md
  • references/one-pager.md
  • references/table-preservation.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Langchain Data Handling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Langchain Data Handling compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Langchain Data Handling this skilljeremylongshore/tons-of-skills-marketplace2.8k—~4.1kAutomated safety check: PassMIT
Langchain RAGlangchain-ai/langchain-skills1.3k—~3.9kAutomated safety check: PassMIT
Neo4j Graphrag Skillneo4j-contrib/neo4j-skills114—~4.2kAutomated safety check: NotesMIT
Add Example AgentGetBindu/Bindu10k—~1.1kAutomated safety check: NotesCustom licence
Failproof AI SDK IntegrationFailproofAI/failproofai5.3k—~6kAutomated safety check: PassCustom licence
SynalinksSynaLinks/synalinks-skills907—~4.8kAutomated safety check: PassApache-2.0

Similar skills

  • Langchain RAG

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when building ANY retrieval-augmented generation (RAG) system.

    1.3k GitHub stars~3.9k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Neo4j Graphrag Skill

    neo4j-contrib/neo4j-skills

    Build GraphRAG retrieval pipelines on Neo4j using the neo4j-graphrag Python package (v1.22.0+).

    114 GitHub stars~4.2k tokensUpdated yesterday
    Knowledge ManagementAuto-check: notes
  • Add Example Agent

    GetBindu/Bindu

    Add a new self-contained example agent under examples/. An agent skill from GetBindu/Bindu.

    10k GitHub stars~1.1k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Failproof AI SDK Integration

    FailproofAI/failproofai

    Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.

    5.3k GitHub stars~6k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Synalinks

    SynaLinks/synalinks-skills

    A skill your agent uses for anything involving the Synalinks neuro-symbolic LM framework (Keras-inspired): DataModel/Field/Input, JSON operators (+ & | ^ ~), synalinks.ops…

    907 GitHub stars~4.8k tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check passed
  • Tavily Search API Integration

    andrewyng/context-hub

    Guides building Tavily integrations for web search, URL extraction, site crawling and AI-assisted research in Python or JavaScript agent and RAG projects.

    14k GitHub stars~1.1k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Questions about Langchain Data Handling

What does Langchain Data Handling do?

Load and chunk documents for LangChain 1.0 RAG pipelines correctly — language-aware splitters, table-safe PDF loaders, Cloudflare-compatible web loaders, chunk-boundary strategies that survive…. Langchain Data Handling is an agent skill from jeremylongshore/tons-of-skills-marketplace.0 RAG pipelines correctly — language-aware splitters, table-safe PDF loaders, Cloudflare-compatible web loaders, chunk-boundary strategies that survive real-world structure.

When should I use Langchain Data Handling?

Langchain Data Handling fits situations like: building a RAG pipeline; diagnosing why retrieval misquotes a table; debugging a crawler returning blank content; with langchain document loader.

How do I install Langchain Data Handling in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-data-handling -a claude-code`. Or copy the skill folder (skills/.curated/langchain-data-handling in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/langchain-data-handling in your project. Claude Code loads it when a task matches its description.

How do I install Langchain Data Handling in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-data-handling -a codex`. Or copy the skill folder (skills/.curated/langchain-data-handling in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/langchain-data-handling in your project. Codex loads it when a task matches its description.

Can I use Langchain Data Handling in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-data-handling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/langchain-data-handling, .gemini/skills/langchain-data-handling, .github/skills/langchain-data-handling and .opencode/skills/langchain-data-handling in your project.

What does Langchain Data Handling need to run?

Going by SKILL.md and its folder, Langchain Data Handling needs the command-line tools its instructions call (pip). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(python:*), Bash(pip:*). Compatibility (from SKILL.md): Designed for Claude Code.

Does Langchain Data Handling access the network?

SKILL.md names 5 domains. As links in the text: python.langchain.com, pymupdf.readthedocs.io, unstructured.io, developers.cloudflare.com and rfc-editor.org. This is read from the text; nothing was executed.

Is Langchain Data Handling safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Langchain Data Handling use?

Langchain Data Handling is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Langchain Data Handling use?

About 4.1k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.9k tokens, read only when the agent opens those files.

What are the alternatives to Langchain Data Handling?

Skills that share tags, products or a category with Langchain Data Handling: Langchain RAG (langchain-ai/langchain-skills, 1.3k stars), Neo4j Graphrag Skill (neo4j-contrib/neo4j-skills, 114 stars), Add Example Agent (GetBindu/Bindu, 10k stars) and Failproof AI SDK Integration (FailproofAI/failproofai, 5.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Langchain Data Handling?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.