Agent skill

Haystack

by magnus919 in magnus919/agent-skills

Build production search and NLP pipelines with Haystack. An agent skill from magnus919/agent-skills.

MITAuto-check passedAI & LLM Engineering

Install Haystack

skills CLI
$ npx skills add magnus919/agent-skills --skill haystack -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install magnus919/agent-skills haystack --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/haystack .claude/skills/haystack && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
haystack
GitHub stars
111
Token cost
~1.6k tokens
SKILL.md length
480 words
Files
15 (incl. scripts, references)
Skills in repo
129
Repo updated
First seen
Licence
MIT

At a glance

Build production search and NLP pipelines with Haystack. An agent skill from magnus919/agent-skills.

  • Works in 5 steps: Pipelines are validated DAGs.… → Components are typed. Each component has… → PromptBuilder uses Jinja2. Templates are… → …
  • Building search pipelines
  • SKILL.md covers Core Paradigm, Core Principles, Where to Start and Quick Reference, plus 4 more sections
  • Runs Python scripts from its folder

What it does

Haystack is an agent skill from magnus919/agent-skills. Build production search and NLP pipelines with Haystack. Pipeline DAG composition, document stores, retrievers, PromptBuilder (Jinja2), generators, evaluation, Hayhooks deployment. Use when building search pipelines or comparing NLP application frameworks. Do not use this skill for unrelated requests; route to the nearest named specialist.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 18 other files, including scripts and reference files (for example `README.md`, `evals/evals.json` and `references/deployment.md`).

It sits in AI & LLM Engineering, covering Natural language processing and NoSQL databases. The repository describes itself as: Curated collection of AI agent skills for Hermes and other agent frameworks. The licence is MIT.

When your agent uses it

  • Building search pipelines
  • Comparing NLP application frameworks
  • Unrelated requests
  • Route to the nearest named specialist

Example prompts

  • “/haystack”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Pipelines are validated DAGs. add_component + connect. Pipeline validation catches errors BEFORE execution — leverage this during…
  2. Components are typed. Each component has input/output slots. Connections must match types. This prevents runtime errors.
  3. PromptBuilder uses Jinja2. Templates are Jinja2 strings, not f-strings. {{documents}}, {{query}}, {{question}} are variable placeholders.
  4. Indexing and query are separate pipelines. One pipeline loads/cleans/embeds/writes documents. Another retrieves/generates answers. They…
  5. Evaluation is a pipeline too. Add evaluator components to measure faithfulness, relevancy, or custom metrics.

What it can do on your machine

Read from SKILL.md and the folder at commit 96fbe07. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Haystack loads about 1.6k tokens when it runs, and up to ~5.5k if it reads all its reference files. Until then it costs about 88 tokens; SKILL.md has 480 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~88
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from magnus919/agent-skills at commit 96fbe07, republished under its MIT licence (© magnus919). 480 words, ~1,649 tokens.

Download SKILL.mdSave it as .claude/skills/haystack/SKILL.md (or your agent's skills folder). This skill also uses 14 other files; get the full folder from GitHub.
name
haystack
description
Build production search and NLP pipelines with Haystack. Pipeline DAG composition, document stores, retrievers, PromptBuilder (Jinja2), generators, evaluation, Hayhooks deployment. Use when building search pipelines or comparing NLP application frameworks. Do not use this skill for unrelated requests; route to the nearest named specialist.
license
MIT
metadata.author
Magnus Hedemark
metadata.version
1.1.0
metadata.source
https://docs.haystack.deepset.ai

Haystack Expert Skill

Haystack (by deepset) is a production-oriented framework for building search and NLP pipelines. Its core abstraction is the Pipeline — a directed acyclic graph of typed components with explicit connections. Unlike LangChain's LCEL (pipe operator) or LlamaIndex's query engines, Haystack pipelines are declared upfront with add_component and connect, giving validated, debuggable DAGs.

Core Paradigm

python
from haystack import Pipeline
from haystack.components.embedders import SentenceTransformersTextEmbedder
from haystack.components.retrievers.in_memory import InMemoryEmbeddingRetriever
from haystack.components.builders import PromptBuilder
from haystack.components.generators import OpenAIGenerator
from haystack.document_stores.in_memory import InMemoryDocumentStore

# Build a pipeline
document_store = InMemoryDocumentStore()
pipeline = Pipeline()
pipeline.add_component("embedder", SentenceTransformersTextEmbedder())
pipeline.add_component("retriever", InMemoryEmbeddingRetriever(document_store=document_store))
pipeline.add_component("prompt_builder", PromptBuilder(template="Answer using: {{documents}}\n\nQuestion: {{question}}"))
pipeline.add_component("generator", OpenAIGenerator())

# Connect components
pipeline.connect("embedder.embedding", "retriever.query_embedding")
pipeline.connect("retriever.documents", "prompt_builder.documents")
pipeline.connect("prompt_builder", "generator")

# Run
result = pipeline.run({"embedder": {"text": "What is Haystack?"}, "prompt_builder": {"question": "What is Haystack?"}})

Core Principles

  1. Pipelines are validated DAGs. add_component + connect. Pipeline validation catches errors BEFORE execution — leverage this during development.
  2. Components are typed. Each component has input/output slots. Connections must match types. This prevents runtime errors.
  3. PromptBuilder uses Jinja2. Templates are Jinja2 strings, not f-strings. {{documents}}, {{query}}, {{question}} are variable placeholders.
  4. Indexing and query are separate pipelines. One pipeline loads/cleans/embeds/writes documents. Another retrieves/generates answers. They share the DocumentStore.
  5. Evaluation is a pipeline too. Add evaluator components to measure faithfulness, relevancy, or custom metrics.

Where to Start

You already have...Start here
Nothing — exploring HaystackBuild a basic indexing + query pipeline
Documents to indexBuild an indexing pipeline (converters, splitter, embedder, writer)
A search use caseBuild a query pipeline (embedder, retriever, prompt, generator)
A production deploymentAdd Hayhooks + evaluation pipeline

Quick Reference

TaskApproachReference
Build indexing pipelineadd_component -> connect -> runreferences/pipeline-design.md
Build query pipelineretriever -> prompt_builder -> generatorreferences/pipeline-design.md
Choose document storeInMemory (dev), Elasticsearch/Pinecone (prod)references/document-stores.md
Embedding retrievalSentenceTransformersTextEmbedder + EmbeddingRetrieverreferences/retrievers.md
Hybrid retrievalBM25 + Embedding in parallel, DocumentJoinerreferences/retrievers.md
Prompt templatesJinja2 in PromptBuilderreferences/pipeline-design.md
EvaluationDeepEvalEvaluator, SASEvaluatorreferences/evaluation.md
DeployHayhooks REST APIreferences/deployment.md

Framework Routing Guide

ScenarioReach forWhy
Search / NLP pipelinesHaystackPipeline DAG model is most mature for retrieval-heavy workloads
Documents to query / RAGLlamaIndexData ingestion is the primary primitive
Chain/agent compositionLangChainLCEL pipe operator for general chain building
Compiled prompt programsDSPyAuto-optimizes prompts against a metric
Role-based multi-agentCrewAIHigher-level agent abstraction
Show full SKILL.md (180 more words)Show less

Reference Files

ReferenceLoad whenFile
Pipeline DesignBuilding indexing and query pipelinesreferences/pipeline-design.md
Document StoresStore selection and configurationreferences/document-stores.md
RetrieversEmbedding, BM25, hybrid retrievalreferences/retrievers.md
Validation AuditResearch validation of all API claimsreferences/validation-audit.md
File ConvertersMulti-format indexing, YAML serialization, component typesreferences/file-converters.md
EvaluationMetrics, evaluators, pipeline evaluationreferences/evaluation.md
DeploymentHayhooks, containerization, productionreferences/deployment.md
FAQ & TroubleshootingCommon errors and fixesreferences/faq-and-troubleshooting.md

Templates

TemplateWhen to useFile
Indexing PipelineLoad, split, embed, write to storetemplates/indexing-pipeline.py
Query PipelineRetrieve, prompt, generate answertemplates/query-pipeline.py
Hybrid RAGBM25 + embedding in paralleltemplates/hybrid-rag.py

Troubleshooting

SymptomLikely causeFixReference
Pipeline run errorsComponent connection mismatchCheck component input/output slot typesreferences/pipeline-design.md
No documents retrievedEmpty document storeRun indexing pipeline firstreferences/pipeline-design.md
Prompt not renderingWrong variable name in Jinja2 templateCheck {{variables}} match pipeline inputreferences/pipeline-design.md
Slow retrievalFull scan instead of ANNConfigure approximate nearest neighbor indexreferences/retrievers.md
Embedding mismatchDifferent models for indexing vs queryUse same model in both pipelinesreferences/retrievers.md
Hayhooks not startingPort conflict or missing configCheck port, run with --help for optionsreferences/deployment.md

© magnus919, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 14 other files (scripts, references) in haystack of magnus919/agent-skills.

  • SKILL.md
  • README.md
  • evals/evals.json
  • references/deployment.md
  • references/document-stores.md
  • references/evaluation.md
  • references/faq-and-troubleshooting.md
  • references/file-converters.md
  • references/pipeline-design.md
  • references/retrievers.md
  • references/validation-audit.md
  • scripts/check-setup.py
  • templates/hybrid-rag.py
  • templates/indexing-pipeline.py
  • templates/query-pipeline.py

Open the folder on GitHubat commit 96fbe07

Compare with similar skills

Haystack next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Haystack compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Haystack this skillmagnus919/agent-skills111—~1.6kAutomated safety check: PassMIT
Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassCustom licence
Data SpecialistProrise-cool/Claude-Code-Multi-Agent305—~1kAutomated safety check: PassNone
Hugging Face TokenizersOrchestra-Research/AI-Research-SKILLs13k7 repos~3.4kAutomated safety check: PassMIT
OpenMed Model Card Writermaziyarpanahi/openmed5.5k—~1.8kAutomated safety check: PassApache-2.0
Andrej KarpathyK-Dense-AI/mimeo282—~1.9kAutomated safety check: PassMIT

Similar skills

  • Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

    1.3k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Data Specialist

    Prorise-cool/Claude-Code-Multi-Agent

    提供数据库设计、优化、数据工程和数据分析能力。当需要处理数据库操作、数据管道或数据分析时使用. An agent skill from Prorise-cool/Claude-Code-Multi-Agent.

    305 GitHub stars~1k tokensUpdated 21 days ago
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Tokenizers

    Orchestra-Research/AI-Research-SKILLs

    Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking.

    13k GitHub starsUsed in 7 repos~3.4k tokens
    AI & LLM EngineeringAuto-check passed
  • OpenMed Model Card Writer

    maziyarpanahi/openmed

    Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.

    5.5k GitHub stars~1.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Andrej Karpathy

    K-Dense-AI/mimeo

    Applies the mental models and frameworks of Andrej Karpathy (deep learning, former Director of AI at Tesla, founding member of OpenAI, Eureka Labs).

    282 GitHub stars~1.9k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Compare

    taishi-i/awesome-japanese-nlp-resources

    Compare several Japanese NLP libraries, models, or datasets for a keyword (a specific tool name, or a function/task like '形態素解析') across a handful of criteria chosen for that comparison, rendered as…

    1k GitHub stars~4.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes

More from magnus919/agent-skills

All 129 skills in this repo
  • Artifact Pyramids

    magnus919/agent-skills

    Organize durable agent research outputs as summaries, analysis, and evidence dossiers.

    111 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check passed
  • Ascii City Engine

    magnus919/agent-skills

    Build portable, first-person colored ASCII city engines and small GIS-derived city packs.

    111 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Color Management

    magnus919/agent-skills

    Manage color workflows with ICC profiles, working spaces, gamut mapping, and color science.

    111 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check: notes
  • Data Scientist

    magnus919/agent-skills

    A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…

    111 GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed
  • Docker Compose

    magnus919/agent-skills

    Use Docker Compose to define, run, debug, and harden multi-container applications.

    111 GitHub stars~2k tokensUpdated yesterday
    Auto-check: notes
  • Fpga Development

    magnus919/agent-skills

    Design, review, simulate, and verify FPGA logic using explicit RTL contracts, clock and reset models, CDC analysis, timing constraints, and reproducible implementation evidence.

    111 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check passed

Questions about Haystack

What does Haystack do?

Build production search and NLP pipelines with Haystack. An agent skill from magnus919/agent-skills. Haystack is an agent skill from magnus919/agent-skills. Build production search and NLP pipelines with Haystack.

When should I use Haystack?

Haystack fits situations like: building search pipelines; comparing NLP application frameworks; unrelated requests; route to the nearest named specialist.

How do I install Haystack in Claude Code?

Run `npx skills add magnus919/agent-skills --skill haystack -a claude-code`. Or copy the skill folder (haystack in magnus919/agent-skills) into .claude/skills/haystack in your project. Claude Code loads it when a task matches its description.

How do I install Haystack in Codex?

Run `npx skills add magnus919/agent-skills --skill haystack -a codex`. Or copy the skill folder (haystack in magnus919/agent-skills) into .agents/skills/haystack in your project. Codex loads it when a task matches its description.

Can I use Haystack in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add magnus919/agent-skills --skill haystack -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/haystack, .gemini/skills/haystack, .github/skills/haystack and .opencode/skills/haystack in your project.

What does Haystack need to run?

Going by SKILL.md and its folder, Haystack needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Haystack access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Haystack safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Haystack use?

Haystack is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Haystack use?

About 1.6k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.8k tokens, read only when the agent opens those files.

What are the alternatives to Haystack?

Skills that share tags, products or a category with Haystack: Gptqmodel Tokenizer Normalization (ModelCloud/GPTQModel, 1.3k stars), Data Specialist (Prorise-cool/Claude-Code-Multi-Agent, 305 stars), Hugging Face Tokenizers (Orchestra-Research/AI-Research-SKILLs, 13k stars) and OpenMed Model Card Writer (maziyarpanahi/openmed, 5.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Haystack?

magnus919 (a GitHub user) maintains it in magnus919/agent-skills, which has 111 GitHub stars. The repository holds 129 skills in this directory. The repository was last updated on October 6, 2026.

Source: magnus919/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.