Postgres
timescale/pg-aiguide
A skill your agent uses for any PostgreSQL database work — table design, indexing, data types, constraints, extensions (pgvector, PostGIS, TimescaleDB), search, and migrations.
Builds and debugs retrieval with the gaik toolkit — PgVectorStore, Ranker, FinnishTextProcessor, RelevanceGate — as hybrid search: pgvector similarity plus Postgres full-text, fused by rank, and the…
$ npx skills add GAIK-project/gaik-toolkit --skill searching-documents -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GAIK-project/gaik-toolkit searching-documents --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GAIK-project/gaik-toolkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/implementation_layer/no-code-assets/agent-plugin/skills/searching-documents .claude/skills/searching-documents && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "searching-documents" agent skill from https://github.com/GAIK-project/gaik-toolkit/tree/main/implementation_layer/no-code-assets/agent-plugin/skills/searching-documents into .claude/skills/searching-documents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "searching-documents", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GAIK-project/gaik-toolkit/tree/main/implementation_layer/no-code-assets/agent-plugin/skills/searching-documentsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GAIK-project/gaik-toolkit --skill searching-documents -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GAIK-project/gaik-toolkit searching-documents --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GAIK-project/gaik-toolkit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/implementation_layer/no-code-assets/agent-plugin/skills/searching-documents .agents/skills/searching-documents && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "searching-documents" agent skill from https://github.com/GAIK-project/gaik-toolkit/tree/main/implementation_layer/no-code-assets/agent-plugin/skills/searching-documents into .agents/skills/searching-documents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "searching-documents", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GAIK-project/gaik-toolkit --skill searching-documents -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GAIK-project/gaik-toolkit searching-documents --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GAIK-project/gaik-toolkit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/implementation_layer/no-code-assets/agent-plugin/skills/searching-documents .cursor/skills/searching-documents && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "searching-documents" agent skill from https://github.com/GAIK-project/gaik-toolkit/tree/main/implementation_layer/no-code-assets/agent-plugin/skills/searching-documents into .cursor/skills/searching-documents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "searching-documents", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GAIK-project/gaik-toolkit.git --path implementation_layer/no-code-assets/agent-plugin/skills/searching-documents--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GAIK-project/gaik-toolkit --skill searching-documents -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GAIK-project/gaik-toolkit searching-documents --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GAIK-project/gaik-toolkit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/implementation_layer/no-code-assets/agent-plugin/skills/searching-documents .gemini/skills/searching-documents && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "searching-documents" agent skill from https://github.com/GAIK-project/gaik-toolkit/tree/main/implementation_layer/no-code-assets/agent-plugin/skills/searching-documents into .gemini/skills/searching-documents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "searching-documents", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GAIK-project/gaik-toolkit searching-documentsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GAIK-project/gaik-toolkit --skill searching-documents -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GAIK-project/gaik-toolkit.git skills-src && mkdir -p .github/skills && cp -r skills-src/implementation_layer/no-code-assets/agent-plugin/skills/searching-documents .github/skills/searching-documents && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "searching-documents" agent skill from https://github.com/GAIK-project/gaik-toolkit/tree/main/implementation_layer/no-code-assets/agent-plugin/skills/searching-documents into .github/skills/searching-documents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "searching-documents", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GAIK-project/gaik-toolkit --skill searching-documents -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GAIK-project/gaik-toolkit searching-documents --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GAIK-project/gaik-toolkit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/implementation_layer/no-code-assets/agent-plugin/skills/searching-documents .opencode/skills/searching-documents && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "searching-documents" agent skill from https://github.com/GAIK-project/gaik-toolkit/tree/main/implementation_layer/no-code-assets/agent-plugin/skills/searching-documents into .opencode/skills/searching-documents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "searching-documents", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
searching-documentsBuilds and debugs retrieval with the gaik toolkit — PgVectorStore, Ranker, FinnishTextProcessor, RelevanceGate — as hybrid search: pgvector similarity plus Postgres full-text, fused by rank, and the…
Searching Documents is an agent skill from GAIK-project/gaik-toolkit. Builds and debugs retrieval with the gaik toolkit — PgVectorStore, Ranker, FinnishTextProcessor, RelevanceGate — as hybrid search: pgvector similarity plus Postgres full-text, fused by rank, and the same design in plain SQL for a TypeScript app. Use whenever gaik is used for search or RAG retrieval, whenever Finnish text is indexed for full-text search, when adding semantic or hybrid search over document chunks, and when deciding whether a search found anything relevant at all. Also use when search misbehaves…
Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/evaluating-retrieval.md`, `references/finnish.md` and `references/postgres-without-gaik.md`).
It sits in AI & LLM Engineering, covering Retrieval-augmented generation, Search implementation and Translation. It works with pgvector, PostgreSQL, SQL and TypeScript. The repository describes itself as: Python toolkit providing reusable AI/ML utilities: schema extraction, structured outputs, and production-ready components. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit e516ece. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
AITTA_API_KEYAITTA_API_TOKENAITTA_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Searching Documents loads about 4.2k tokens when it runs, and up to ~9.7k if it reads all its reference files. Until then it costs about 219 tokens; SKILL.md has 1,950 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GAIK-project/gaik-toolkit at commit e516ece, republished under its MIT licence (© GAIK-project). 1,950 words, ~4,241 tokens.
.claude/skills/searching-documents/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.pip install "gaik[pg-vector-store,embedder,ranker]"
pip install "gaik[finnish-rag]" # Finnish lemmatizationHybrid search runs two arms that fail differently. The vector arm finds passages that
answer in other words; the keyword arm finds exact tokens an embedding blurs — a section
number, a standard's code, a product name. Reciprocal Rank Fusion (RRF) combines them by
rank position, so a cosine similarity and a ts_rank_cd never have to share a scale.
Building that takes an afternoon. What costs weeks is that either arm can stop working without an error, and the page still fills with plausible results because the other arm keeps answering. Most of this skill is about making that failure visible.
Do not assume hybrid wins. On one 92-question set, vector-only matched hybrid's recall from top-10 upwards with a slightly better MRR at an eighth of the latency; hybrid won only at top-5. That metric asks whether one chunk was found, so it cannot see the case hybrid exists for — an exact code or section number — and on other corpora the keyword arm was the only thing that found those at all.
So ship three modes behind one parameter — hybrid, vector, keyword — and judge hybrid
against vector-only on a query set that contains exact tokens. The modes double as the
cheapest health check there is (below).
from gaik.software_components.config import get_openai_config
from gaik.software_components.RAG.embedder import Embedder
from gaik.software_components.RAG.pg_vector_store import PgVectorStore
from gaik.software_components.RAG.ranker import Ranker
# model= is the deployment name on Azure; its output size must equal embedding_dim
embedder = Embedder(get_openai_config(use_azure=True), model="text-embedding-3-small")
store = PgVectorStore(
"postgresql://postgres:postgres@localhost:5432/search",
embedding_dim=1536,
fts_language="finnish", # "simple" when a text_processor supplies lemmas (below)
tsquery_mode="or", # a sentence query still matches (default from 0.8.4)
hnsw_ef_search=100,
)
store.setup() # table, HNSW + GIN indexes, SQL functions
embeddings, docs = embedder.embed(chunks) # chunks: list[langchain Document]
store.add(docs, embeddings)
query_vec = embedder.embed_query(query)
semantic = store.search_semantic(query_vec, top_k=50, threshold=0.0)
keyword = store.search_keyword(query, top_k=50)
hits = Ranker(expose_ranks=True).fuse(
semantic, keyword, names=("semantic", "keyword"), weights=(1.0, 1.0), top_k=10
)Embedder also accepts get_llm_config() from gaik.software_components.llm for
native google/vertex, aitta, openai_compatible, or optional litellm. Native
Google/Vertex needs gaik[llm-google]; LiteLLM needs gaik[llm-litellm] and an explicit
provider-prefixed embedding_model (or LITELLM_EMBEDDING_MODEL). Aitta requires
AITTA_API_KEY (also accepts AITTA_API_TOKEN or AITTA_TOKEN) and an explicit
AITTA_EMBEDDING_MODEL, or an embedding_model config override naming an embedding model
available on Aitta. A chat model is not an embedding model. Other compatible servers need their endpoint, API key,
and an explicit embedding model. Use the same embedding model for indexing and queries,
and match embedding_dim to its output. Native Anthropic does not provide embeddings.
For the Chroma-based end-to-end module, choose the three model stages independently:
from gaik.software_components.llm import get_llm_config
from gaik.software_modules.RAG_workflow import RAGWorkflow
workflow = RAGWorkflow(
parser_config=get_llm_config("openai", model="gpt-6-luna"),
embedding_config=get_llm_config("google", embedding_model="gemini-embedding-001"),
answer_config=get_llm_config("aitta"),
)Install gaik[rag-workflow,llm-google] for this example. The parser model must accept
images; answer generation only needs chat. Each omitted stage uses shared api_config
or the legacy OpenAI/Azure default. When all stages are explicit, no unrelated default
credentials are loaded. RAGWorkflow uses Chroma; the PostgreSQL recipe above remains
the path for PgVectorStore.
Fuse the two lists yourself instead of calling store.search_hybrid():
search_hybrid() returns only RRF scores;expose_ranks=True writes rank_semantic and rank_keyword into every hit's metadata,
so an arm that contributes nothing shows up per result;search_hybrid() ignores tsquery_mode and always parses in
websearch mode, which ANDs every term, and search_hybrid_weighted() fails on every
call with column reference "id" is ambiguous.From gaik 0.8.4 search_hybrid() also writes semantic_similarity (cosine) into every
hit's metadata, so one call is enough when the ranks it returns are all you need.
Tune the weights and the RRF constant, do not assume them. Where the keyword arm alone
reached 33.7% recall@15, equal weights at k=60 scored 8.7 points below vectors weighted 3×
at k=20 (Ranker(rrf_k=...) sets k).
embedding_dim must equal the model's output, and at most 2,000. setup() builds an
HNSW index on vector(N), which pgvector refuses above 2,000 dimensions. Without model=,
Embedder falls back to a default that depends on the config helper, and two of them
are too large: get_openai_config() gives text-embedding-3-large (3,072) and
get_llm_config("google") or "vertex" gives gemini-embedding-001 (3,072).
get_llm_config("openai") or "azure" gives EMBEDDING_MODEL, else
text-embedding-3-small (1,536). Pass model= explicitly, and for a model above 2,000
pass vector_type="halfvec", which indexes up to 4,000 (gaik 0.8.4+; before that, pick
a smaller model or build the schema yourself: references/postgres-without-gaik.md).
websearch_to_tsquery and plainto_tsquery conjoin
every term, so a nine-word question only matches a passage containing all nine stems —
on real prose, none. Rewrite & to | and let ts_rank_cd rank by how many terms
matched and how close together. In gaik that is tsquery_mode="or".GENERATED ALWAYS AS (…) STORED column the
migration declared. Every row was NULL for about six months, the GIN index indexed
nothing, and "hybrid" was vector search the whole time. Users noticed first.finnish index took the keyword arm's
recall@15 from 33.7% to 5.4%, and hybrid became vector-only. Translate the query into the
corpus language as a question, not a keyword list — a keyword list helped the keyword
arm but wrecked the vector arm and ended 14 points below not translating at all.A fourth, rarer: zero-width characters (U+200B) pasted in from a CMS glue onto words. Postgres does not treat them as whitespace, so the word stays unstemmed and never matches. Strip them at ingest and at query time.
Check the arms are alive, against the live database. A unit test cannot see a column
the migration declared correctly and the database never got. With gaik 0.8.4+,
report = store.health() does it in one call: report.ok, and str(report) names each
problem (a text_search column nothing fills, another fts_language than the index,
lemmas built with other settings, rows either arm cannot see, missing indexes, session
settings that did not stick). By hand, from a health endpoint or a startup check:
hybrid and in vector mode: identical result lists mean the keyword arm is
contributing nothing.Postgres' finnish snowball stemmer is inconsistent across a word family —
tilinpäätös stems to tilinpäätös, tilinpäätöksen to tilinpäätöks — so a search for
one form misses a passage containing the other, and nothing reports the miss. Lemmatize
both sides with the same analyser. Measured on sentences queried in a different
inflection than the document used: snowball matched 1 of 4, every prefix strategy 0–1 of
4, lemmas on both sides 4 of 4.
from gaik.software_components.RAG.finnish_text_processor import FinnishTextProcessor
# A named backend raises ImportError when it is missing; only "auto" falls back
# silently, to "simple", a tokenizer rather than a lemmatizer.
processor = FinnishTextProcessor(backend="pyvoikko", decompound=False)
store = PgVectorStore(dsn, embedding_dim=1536, fts_language="simple",
text_processor=processor, tsquery_mode="or")Read references/finnish.md before indexing Finnish. The rules that each cost a debugging
session:
backend="auto" picks native libvoikko where it is installed and
something else where it is not, so a laptop indexes different lemmas than the container
queries with, and the arm silently stops matching. pyvoikko is pure Python and runs
everywhere, including images where libvoikko is not packaged.decompound=False — gaik defaults to True. With both sides lemmatized,
whole compounds already match, and splitting adds noise: arvonlisävero yields arvo,
which reaches arvopaperi. Index and query must agree; after changing it,
store.relemmatize() (0.8.4+) rebuilds the lemmas without re-embedding. Through 0.8.3
backend="auto" ignored decompound=False and split anyway.unaccent. In Finnish, ä/a and ö/o are different letters: tähti/tahti,
sää/saa.The vector arm answers every query with its nearest neighbours: on one corpus, a gibberish string returned 24 results and a real question 12. What separates them is the distance to the closest chunk — 17 answerable queries scored 0.26–0.56 cosine distance, 16 unanswerable ones 0.64–0.87 — so a floor at 0.60 split them cleanly. That number belongs to one embedding model and one corpus; take your own reading.
from gaik.software_components.RAG.relevance_gate import RelevanceGate
# answerable_best / unanswerable_best: the top similarity of search_semantic(...) per query
reading = RelevanceGate.calibrate(answerable_best, unanswerable_best, lower_is_better=False)
print(reading) # says so plainly when the two populations overlap
gate = RelevanceGate(reading.floor, lower_is_better=False) if reading.separated else None
if gate and not gate.is_answerable(semantic, key=lambda hit: hit[1]):
... # "nothing in the library covers this"search_semantic() returns cosine similarity, higher is better, so pass
lower_is_better=False. Getting the direction backwards inverts the gate with no error.Search ranks chunks; the reader usually sees documents. Over-fetch chunks — about four times the documents wanted — collapse them onto their parent, and keep the closest distance per document: rows arrive in fused order, so a document's first chunk is its best-ranked, not necessarily its nearest.
Do not order documents by hit count. On one set that dropped document-level hit@1 from 0.79 to 0.63, because a long document that keeps mentioning a topic outranks the short one that is about it. Fusing the best-chunk rank with a length-normalised hit density raised it to 0.83 — and only for short topical queries; on long questions it moved nothing and pushed correct chunks down.
Read references/evaluating-retrieval.md before reporting a number or adopting a change.
The rules that matter most:
A cross-encoder reorders only the pool it is handed. On one set it raised hit@1 from 0.757
to 0.843 on full questions, did nothing for short domain terms, and made some of them
worse. Ranker().rerank(query, hits, on_error="fallback") returns the input order when
the model fails, but has no timeout — wrap it:
asyncio.wait_for(asyncio.to_thread(ranker.rerank, query, hits), timeout=4).
Hosted rerankers often ship with tight rate limits; check the quota before designing
around one.
Retriever(hybrid_search=True) is not a keyword arm. It BM25-rescores the vector
candidates (0.7 × vector + 0.3 × BM25), so a document the embedding missed cannot come
back. PgVectorStore.search_keyword() is the real one.search_semantic() defaults to threshold=0.7. Pass threshold=0.0 when fusing or
gating, or the list arrives pre-cut and the gate never sees the weak cases.hnsw.ef_search defaults to 40. Raising it to 100 measured recall@20 against an exact
scan at 96.2% → 99.2% for +0.7 ms on one 1,500-dimension corpus. On a few thousand rows
it changes nothing — the index already returns exact neighbours.PgVectorStore(..., hnsw_iterative_scan="relaxed_order") from gaik
0.8.4, or SET hnsw.iterative_scan = relaxed_order per connection or as a database
default.*_l2_ops and
queried with <=> is ignored, and every search scans the table.ts_rank_cd has no IDF: in a two-word query where one word is common, that word decides
the ranking. Dropping very common words from a short query's keyword text helps — check
it on both query sets.store.setup() creates the vector, pg_trgm and unaccent extensions. On a managed
database without that privilege, have them installed and call
setup(create_extensions=False).vector(N).© GAIK-project, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in implementation_layer/no-code-assets/agent-plugin/skills/searching-documents of GAIK-project/gaik-toolkit.
Open the folder on GitHubat commit e516ece
Searching Documents next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Searching Documents this skillGAIK-project/gaik-toolkit | 100 | — | ~4.2k | Automated safety check: Pass | MIT | |
| Postgrestimescale/pg-aiguide | 1.9k | — | ~941 | Automated safety check: Pass | Apache-2.0 | |
| Pgvector Semantic Searchtimescale/pg-aiguide | 1.9k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | |
| Postgres Hybrid Text Searchtimescale/pg-aiguide | 1.9k | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | |
| RAG Implementationwshobson/agents | 40k | 9 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Neon Postgresneondatabase/agent-skills | 100 | — | ~4.1k | Automated safety check: Notes | Apache-2.0 |
timescale/pg-aiguide
A skill your agent uses for any PostgreSQL database work — table design, indexing, data types, constraints, extensions (pgvector, PostGIS, TimescaleDB), search, and migrations.
timescale/pg-aiguide
A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.
timescale/pg-aiguide
A skill your agent uses to implement hybrid search combining BM25 keyword search with semantic vector search using Reciprocal Rank Fusion (RRF).
wshobson/agents
Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline.
neondatabase/agent-skills
Guides and best practices for working with Lakebase Postgres on Neon: connections, pooled vs direct, schema migrations, branching, autoscaling, scale-to-zero, instant restore, read replicas, IP…
wshobson/agents
Shows how to run vector and keyword search side by side and merge their results, so retrieval catches both meaning and exact terms in RAG and search systems.
GAIK-project/gaik-toolkit
Builds a visual, editable PowerPoint (.pptx) deck with speaker-ready notes, exact timing, citations and a layout-checked design from a topic, an audience and a length, using only the user's own…
GAIK-project/gaik-toolkit
GAIK toolkit overview and reference. An agent skill from GAIK-project/gaik-toolkit.
GAIK-project/gaik-toolkit
Extracts structured data — fields, tables, line items — out of documents into a validated schema using the gaik toolkit, and designs schemas that stay inside provider limits and produce checkable…
GAIK-project/gaik-toolkit
Converts PDFs, scans, and Word documents into text or markdown with the gaik toolkit's parsers, choosing the parser that will not silently destroy the structure the downstream task depends on.
GAIK-project/gaik-toolkit
Extracts structured data from Finnish construction site daily diary audio recordings (Työmaapäiväkirja) and creates a formatted Word document with extracted fields.
GAIK-project/gaik-toolkit
Adds or updates working code examples for GAIK toolkit components and pipelines in implementationlayer/examples/.
Works with
Categories
Builds and debugs retrieval with the gaik toolkit — PgVectorStore, Ranker, FinnishTextProcessor, RelevanceGate — as hybrid search: pgvector similarity plus Postgres full-text, fused by rank, and the…. Searching Documents is an agent skill from GAIK-project/gaik-toolkit. Builds and debugs retrieval with the gaik toolkit — PgVectorStore, Ranker, FinnishTextProcessor, RelevanceGate — as hybrid search: pgvector similarity plus Postgres full-text, fused by rank, and the same design in plain SQL for a TypeScript app.
Searching Documents fits situations like: gaik is used for search; whenever Finnish text is indexed for full-text search; adding semantic; hybrid search over document chunks.
Run `npx skills add GAIK-project/gaik-toolkit --skill searching-documents -a claude-code`. Or copy the skill folder (implementation_layer/no-code-assets/agent-plugin/skills/searching-documents in GAIK-project/gaik-toolkit) into .claude/skills/searching-documents in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GAIK-project/gaik-toolkit --skill searching-documents -a codex`. Or copy the skill folder (implementation_layer/no-code-assets/agent-plugin/skills/searching-documents in GAIK-project/gaik-toolkit) into .agents/skills/searching-documents in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GAIK-project/gaik-toolkit --skill searching-documents -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/searching-documents, .gemini/skills/searching-documents, .github/skills/searching-documents and .opencode/skills/searching-documents in your project.
Going by SKILL.md and its folder, Searching Documents needs the command-line tools its instructions call (pip) and credentials named AITTA_API_KEY, AITTA_API_TOKEN and AITTA_TOKEN. Our summary lists: Python 3; A credential in AITTA_API_KEY; A credential in AITTA_API_TOKEN.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Searching Documents is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.5k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Searching Documents: Postgres (timescale/pg-aiguide, 1.9k stars), Pgvector Semantic Search (timescale/pg-aiguide, 1.9k stars), Postgres Hybrid Text Search (timescale/pg-aiguide, 1.9k stars) and RAG Implementation (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GAIK-project (a GitHub organization) maintains it in GAIK-project/gaik-toolkit, which has 100 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 9, 2026.
Source: GAIK-project/gaik-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.