Langchain4j RAG Implementation Patterns
giuseppe-trisciuoglio/developer-kit
Provides Retrieval-Augmented Generation (RAG) implementation patterns with LangChain4j for Java.
Build a RAG (retrieval-augmented generation) pipeline or a custom search engine on top of Bright Data's Discover API — using intent-ranked web results + parsed page content as the…
$ npx skills add brightdata/skills --skill rag-pipeline -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install brightdata/skills rag-pipeline --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/brightdata/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/rag-pipeline .claude/skills/rag-pipeline && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "rag-pipeline" agent skill from https://github.com/brightdata/skills/tree/main/skills/rag-pipeline into .claude/skills/rag-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-pipeline", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/brightdata/skills/tree/main/skills/rag-pipelineType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add brightdata/skills --skill rag-pipeline -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install brightdata/skills rag-pipeline --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/brightdata/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/rag-pipeline .agents/skills/rag-pipeline && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "rag-pipeline" agent skill from https://github.com/brightdata/skills/tree/main/skills/rag-pipeline into .agents/skills/rag-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-pipeline", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add brightdata/skills --skill rag-pipeline -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install brightdata/skills rag-pipeline --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/brightdata/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/rag-pipeline .cursor/skills/rag-pipeline && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "rag-pipeline" agent skill from https://github.com/brightdata/skills/tree/main/skills/rag-pipeline into .cursor/skills/rag-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-pipeline", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/brightdata/skills.git --path skills/rag-pipeline--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add brightdata/skills --skill rag-pipeline -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install brightdata/skills rag-pipeline --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/brightdata/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/rag-pipeline .gemini/skills/rag-pipeline && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "rag-pipeline" agent skill from https://github.com/brightdata/skills/tree/main/skills/rag-pipeline into .gemini/skills/rag-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-pipeline", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install brightdata/skills rag-pipelineInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add brightdata/skills --skill rag-pipeline -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/brightdata/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/rag-pipeline .github/skills/rag-pipeline && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "rag-pipeline" agent skill from https://github.com/brightdata/skills/tree/main/skills/rag-pipeline into .github/skills/rag-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-pipeline", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add brightdata/skills --skill rag-pipeline -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install brightdata/skills rag-pipeline --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/brightdata/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/rag-pipeline .opencode/skills/rag-pipeline && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "rag-pipeline" agent skill from https://github.com/brightdata/skills/tree/main/skills/rag-pipeline into .opencode/skills/rag-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-pipeline", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
rag-pipelineBuild a RAG (retrieval-augmented generation) pipeline or a custom search engine on top of Bright Data's Discover API — using intent-ranked web results + parsed page content as the…
RAG Pipeline is an agent skill from brightdata/skills. Build a RAG (retrieval-augmented generation) pipeline or a custom search engine on top of Bright Data's Discover API — using intent-ranked web results + parsed page content as the retrieval/ingestion layer for an LLM or vector store. Use when the user wants to "build a RAG pipeline", "add web search to my LLM/agent", "ground my model in live web data", "build a search engine over the web", "ingest web content into a vector DB / knowledge base", or "give my chatbot retrieval". Covers both live retrieval (Discover…
Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/code.md`).
It sits in AI & LLM Engineering, covering Retrieval-augmented generation, Web search and Vector databases. It works with Bright Data. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 81f51af. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are javascript).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
BRIGHTDATA_API_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
RAG Pipeline loads about 1.9k tokens when it runs, and up to ~4.4k if it reads all its reference files. Until then it costs about 195 tokens; SKILL.md has 640 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from brightdata/skills at commit 81f51af, republished under its MIT licence (© brightdata). 640 words, ~1,850 tokens.
.claude/skills/rag-pipeline/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Use Discover as the retrieval layer for an LLM app or a custom search engine.
Discover already returns intent-ranked, relevance-scored results with parsed
page content, so it does the "search + fetch + clean" stage of RAG for you. This
is a code/architecture skill built on the discover-api skill — read that
for API mechanics (trigger/poll, modes, params, limits).
Pick the right neighbor: a written brief → live-research; markdown of specific
URLs you already have → scrape; structured platform records → data-feeds.
Does the corpus change every query, or is it a stable knowledge base?
├── Per-query, always-fresh ("ground each answer in live web data")
│ → LIVE RETRIEVAL: Discover(include_content) at query time → top-k → LLM
│ Pros: always current, no storage. Cons: per-query latency + cost.
│
└── Reused across many queries ("build a knowledge base / search engine")
→ INGESTION: Discover(include_content) → chunk → embed → vector store
then at query time: embed query → vector search → (rerank) → LLM
Pros: fast queries, cacheable. Cons: can go stale (re-ingest on a schedule).Many systems do both: an ingested base for breadth + a live Discover call for freshness, merged before the LLM.
Pattern: on each user question, run Discover with a sharp intent, take the
top-k by relevance_score, and pass their content as context to the LLM. The
LLM cites the links.
import { bdclient } from '@brightdata/sdk';
const client = new bdclient(); // BRIGHTDATA_API_TOKEN
async function retrieve(question, k = 6) {
const res = await client.discover(question, {
intent: `authoritative sources that directly answer: ${question}`,
includeContent: true,
numResults: Math.min(k * 2, 20), // over-fetch, then trim
});
// NOTE: the JS SDK returns a WRAPPER object, not a bare array:
// { success, data: [ {link,title,description,relevance_score,content?} ], totalResults, cost, taskId, ... }
// The result rows are in `.data` (CLI/REST use `.results` instead — see discover-api).
if (!res.success) throw new Error(`discover failed: ${res.error ?? 'unknown'}`);
return (res.data ?? [])
.filter(r => r.content && !/just a moment|captcha|access denied|not found/i.test(r.content) && r.content.length > 200)
.sort((a, b) => b.relevance_score - a.relevance_score)
.slice(0, k);
}
// → build a prompt from sources[].content, ask the LLM to answer WITH [n] citations to sources[].linkFull prompt-assembly + citation pattern: references/code.md.
Pattern: discover broadly (high volume — zeroRanking via REST is ideal here),
chunk each page's content, embed the chunks, upsert into a vector store with the
source URL as metadata. At query time: embed the query, vector-search, optionally
rerank, then feed to the LLM.
Stages: discover → dedup → chunk → embed → upsert (ingest), then
embed query → search → rerank → generate (serve). Provider-agnostic code for
both stages, including chunking and metadata, is in
references/code.md.
For bulk corpus building, prefer the raw REST "mode":"zeroRanking" flow (max raw
results, no ranking) from the discover-api skill — but note it ignores
num_results and does not support include_content, so you fetch content
separately (Discover standard/deep with content, or the scrape skill).
link (and ideally title +
relevance_score). RAG without citations is unverifiable.content before embedding. Skip block pages and empty bodies
(oversized PDFs return null content). Embedding garbage poisons retrieval.relevance_score. Discover's score is a strong prior
for top-k selection before (or instead of) a reranker.num_results ≤ 20 per call; dedup by normalized URL across
calls so one article via three aggregators isn't triple-weighted.content made it into the index — spot-check stored chunks.[n] the LLM emits maps to a real source link in the retrieved set.content without filtering block pages / nulls.num_results as unlimited (cap 20) or expecting include_content under zeroRanking.references/code.md — runnable JS + Python for both architectures: live retrieval with prompt+citation assembly, and the full ingestion pipeline (discover → dedup → chunk → embed → upsert → query), with a provider-agnostic embedder/vector-store interface.discover-api — the retrieval API (trigger/poll, modes, include_content, limits). Read first.live-research — one-off synthesized report instead of a standing system.scrape — fetch markdown for specific URLs you already have.js-sdk-best-practices / python-sdk-best-practices — client.discover() option details and batch patterns.© brightdata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in skills/rag-pipeline of brightdata/skills.
Open the folder on GitHubat commit 81f51af
RAG Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| RAG Pipeline this skillbrightdata/skills | 264 | — | ~1.9k | Automated safety check: Pass | MIT | |
| Langchain4j RAG Implementation Patternsgiuseppe-trisciuoglio/developer-kit | 355 | 1 repos | ~3.3k | Automated safety check: Notes | MIT | |
| Context Retrievalseb1n/awesome-ai-agent-skills | 206 | — | ~2.1k | Automated safety check: Pass | MIT | |
| Tavily Search API Integrationandrewyng/context-hub | 14k | — | ~1.1k | Automated safety check: Pass | MIT | |
| Sap AI Coresecondsky/sap-skills | 462 | — | ~3.3k | Automated safety check: Pass | GPL-3.0 | |
| AWS Cloudformation Bedrockgiuseppe-trisciuoglio/developer-kit | 355 | — | ~3.2k | Automated safety check: Notes | MIT |
giuseppe-trisciuoglio/developer-kit
Provides Retrieval-Augmented Generation (RAG) implementation patterns with LangChain4j for Java.
seb1n/awesome-ai-agent-skills
Retrieve relevant information from a knowledge base using semantic, keyword, or hybrid search to ground a query.
andrewyng/context-hub
Guides building Tavily integrations for web search, URL extraction, site crawling and AI-assisted research in Python or JavaScript agent and RAG projects.
secondsky/sap-skills
Guides development with SAP AI Core and SAP AI Launchpad for enterprise AI/ML workloads on SAP BTP.
giuseppe-trisciuoglio/developer-kit
Provides AWS CloudFormation patterns for Amazon Bedrock resources including agents, knowledge bases, data sources, guardrails, prompts, flows, and inference profiles.
davila7/claude-code-templates
Set up and run local web searches using Bright Data SERP API with the unfancy-search pipeline (query expansion, SERP retrieval, RRF reranking).
brightdata/skills
Replicate the visual style of any website and apply it to your existing codebase.
brightdata/skills
Generate working code that routes HTTP requests through Bright Data proxy networks (Datacenter, ISP, Residential, Mobile) and help users decide which network and IP pool type to use (shared pool…
brightdata/skills
Bright Data MCP handles ALL web data operations. An agent skill from brightdata/skills.
brightdata/skills
Produce a deep, multi-source, cited research brief on a topic from live web data using Bright Data's Discover API (intent-ranked web search + parsed page content).
brightdata/skills
Web data extraction and discovery using the Bright Data JavaScript/TypeScript SDK (@brightdata/sdk).
brightdata/skills
Extract structured data from 40+ supported platforms (Amazon, LinkedIn, Instagram, TikTok, Facebook, YouTube, Reddit, and more) via the Bright Data CLI (bdata pipelines).
Works with
Categories
Build a RAG (retrieval-augmented generation) pipeline or a custom search engine on top of Bright Data's Discover API — using intent-ranked web results + parsed page content as the…. RAG Pipeline is an agent skill from brightdata/skills. Build a RAG (retrieval-augmented generation) pipeline or a custom search engine on top of Bright Data's Discover API — using intent-ranked web results + parsed page content as the retrieval/ingestion layer for an LLM or vector store.
RAG Pipeline fits situations like: the user wants to build a RAG pipeline; add web search to my LLM/agent; ground my model in live web data; build a search engine over the web.
Run `npx skills add brightdata/skills --skill rag-pipeline -a claude-code`. Or copy the skill folder (skills/rag-pipeline in brightdata/skills) into .claude/skills/rag-pipeline in your project. Claude Code loads it when a task matches its description.
Run `npx skills add brightdata/skills --skill rag-pipeline -a codex`. Or copy the skill folder (skills/rag-pipeline in brightdata/skills) into .agents/skills/rag-pipeline in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brightdata/skills --skill rag-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag-pipeline, .gemini/skills/rag-pipeline, .github/skills/rag-pipeline and .opencode/skills/rag-pipeline in your project.
Going by SKILL.md and its folder, RAG Pipeline needs credentials named BRIGHTDATA_API_TOKEN. Our summary lists: Python 3; A credential in BRIGHTDATA_API_TOKEN.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
RAG Pipeline is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with RAG Pipeline: Langchain4j RAG Implementation Patterns (giuseppe-trisciuoglio/developer-kit, 355 stars), Context Retrieval (seb1n/awesome-ai-agent-skills, 206 stars), Tavily Search API Integration (andrewyng/context-hub, 14k stars) and Sap AI Core (secondsky/sap-skills, 462 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
brightdata (a GitHub organization) maintains it in brightdata/skills, which has 264 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 7, 2026.
Source: brightdata/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.