Agent skill

Searching Documents

by GAIK-project in GAIK-project/gaik-toolkit

Builds and debugs retrieval with the gaik toolkit — PgVectorStore, Ranker, FinnishTextProcessor, RelevanceGate — as hybrid search: pgvector similarity plus Postgres full-text, fused by rank, and the…

MITAuto-check passedAI & LLM Engineering

Install Searching Documents

skills CLI
$ npx skills add GAIK-project/gaik-toolkit --skill searching-documents -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GAIK-project/gaik-toolkit searching-documents --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GAIK-project/gaik-toolkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/implementation_layer/no-code-assets/agent-plugin/skills/searching-documents .claude/skills/searching-documents && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
searching-documents
GitHub stars
100
Token cost
~4.2k tokens
SKILL.md length
1,950 words
Files
4 (incl. references)
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Builds and debugs retrieval with the gaik toolkit — PgVectorStore, Ranker, FinnishTextProcessor, RelevanceGate — as hybrid search: pgvector similarity plus Postgres full-text, fused by rank, and the…

  • Works in 3 steps: The tsquery ANDs a sentence.… → The index column is empty. On one… → The query is in another language than…
  • Gaik is used for search
  • SKILL.md covers Keep vector-only as a mode and…, The gaik recipe, Three ways the keyword arm… and Finnish and other inflected…, plus 5 more sections
  • Calls pip; needs AITTA_API_KEY and AITTA_API_TOKEN

What it does

Searching Documents is an agent skill from GAIK-project/gaik-toolkit. Builds and debugs retrieval with the gaik toolkit — PgVectorStore, Ranker, FinnishTextProcessor, RelevanceGate — as hybrid search: pgvector similarity plus Postgres full-text, fused by rank, and the same design in plain SQL for a TypeScript app. Use whenever gaik is used for search or RAG retrieval, whenever Finnish text is indexed for full-text search, when adding semantic or hybrid search over document chunks, and when deciding whether a search found anything relevant at all. Also use when search misbehaves…

Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/evaluating-retrieval.md`, `references/finnish.md` and `references/postgres-without-gaik.md`).

It sits in AI & LLM Engineering, covering Retrieval-augmented generation, Search implementation and Translation. It works with pgvector, PostgreSQL, SQL and TypeScript. The repository describes itself as: Python toolkit providing reusable AI/ML utilities: schema extraction, structured outputs, and production-ready components. The licence is MIT.

When your agent uses it

  • Gaik is used for search
  • Whenever Finnish text is indexed for full-text search
  • Adding semantic
  • Hybrid search over document chunks

Example prompts

  • “Use the searching-documents skill to build and debugs retrieval with the gaik toolkit — PgVectorStore, Ranker, FinnishTextProcessor, RelevanceGate —…”
  • “/searching-documents”

Requirements

  • Python 3
  • A credential in AITTA_API_KEY
  • A credential in AITTA_API_TOKEN

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. The tsquery ANDs a sentence. websearch_to_tsquery and plainto_tsquery conjoin
  2. The index column is empty. On one production system the tsvector column existed as a
  3. The query is in another language than the index. Postgres full-text cannot cross

What it can do on your machine

Read from SKILL.md and the folder at commit e516ece. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • AITTA_API_KEY
    • AITTA_API_TOKEN
    • AITTA_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Searching Documents loads about 4.2k tokens when it runs, and up to ~9.7k if it reads all its reference files. Until then it costs about 219 tokens; SKILL.md has 1,950 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~219
When it runs · the whole SKILL.md, loaded when a task matches
~4.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GAIK-project/gaik-toolkit at commit e516ece, republished under its MIT licence (© GAIK-project). 1,950 words, ~4,241 tokens.

Download SKILL.mdSave it as .claude/skills/searching-documents/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
searching-documents
description
Builds and debugs retrieval with the gaik toolkit — PgVectorStore, Ranker, FinnishTextProcessor, RelevanceGate — as hybrid search: pgvector similarity plus Postgres full-text, fused by rank, and the same design in plain SQL for a TypeScript app. Use whenever gaik is used for search or RAG retrieval, whenever Finnish text is indexed for full-text search, when adding semantic or hybrid search over document chunks, and when deciding whether a search found anything relevant at all. Also use when search misbehaves without an error: a sentence query gets no keyword hits, a word visibly in the text is not found, hybrid returns what vector-only returns, gibberish fills a page of results, or a similarity threshold never filters anything; and when measuring retrieval with Hit@K or MRR, or judging whether a reranker, synonyms, or query translation helped.

Searching documents with gaik

bash
pip install "gaik[pg-vector-store,embedder,ranker]"
pip install "gaik[finnish-rag]"   # Finnish lemmatization

Hybrid search runs two arms that fail differently. The vector arm finds passages that answer in other words; the keyword arm finds exact tokens an embedding blurs — a section number, a standard's code, a product name. Reciprocal Rank Fusion (RRF) combines them by rank position, so a cosine similarity and a ts_rank_cd never have to share a scale.

Building that takes an afternoon. What costs weeks is that either arm can stop working without an error, and the page still fills with plausible results because the other arm keeps answering. Most of this skill is about making that failure visible.

Keep vector-only as a mode and make hybrid earn its place

Do not assume hybrid wins. On one 92-question set, vector-only matched hybrid's recall from top-10 upwards with a slightly better MRR at an eighth of the latency; hybrid won only at top-5. That metric asks whether one chunk was found, so it cannot see the case hybrid exists for — an exact code or section number — and on other corpora the keyword arm was the only thing that found those at all.

So ship three modes behind one parameter — hybrid, vector, keyword — and judge hybrid against vector-only on a query set that contains exact tokens. The modes double as the cheapest health check there is (below).

The gaik recipe

python
from gaik.software_components.config import get_openai_config
from gaik.software_components.RAG.embedder import Embedder
from gaik.software_components.RAG.pg_vector_store import PgVectorStore
from gaik.software_components.RAG.ranker import Ranker

# model= is the deployment name on Azure; its output size must equal embedding_dim
embedder = Embedder(get_openai_config(use_azure=True), model="text-embedding-3-small")

store = PgVectorStore(
    "postgresql://postgres:postgres@localhost:5432/search",
    embedding_dim=1536,
    fts_language="finnish",    # "simple" when a text_processor supplies lemmas (below)
    tsquery_mode="or",         # a sentence query still matches (default from 0.8.4)
    hnsw_ef_search=100,
)
store.setup()                                  # table, HNSW + GIN indexes, SQL functions
embeddings, docs = embedder.embed(chunks)      # chunks: list[langchain Document]
store.add(docs, embeddings)

query_vec = embedder.embed_query(query)
semantic = store.search_semantic(query_vec, top_k=50, threshold=0.0)
keyword = store.search_keyword(query, top_k=50)
hits = Ranker(expose_ranks=True).fuse(
    semantic, keyword, names=("semantic", "keyword"), weights=(1.0, 1.0), top_k=10
)

Embedder also accepts get_llm_config() from gaik.software_components.llm for native google/vertex, aitta, openai_compatible, or optional litellm. Native Google/Vertex needs gaik[llm-google]; LiteLLM needs gaik[llm-litellm] and an explicit provider-prefixed embedding_model (or LITELLM_EMBEDDING_MODEL). Aitta requires AITTA_API_KEY (also accepts AITTA_API_TOKEN or AITTA_TOKEN) and an explicit AITTA_EMBEDDING_MODEL, or an embedding_model config override naming an embedding model available on Aitta. A chat model is not an embedding model. Other compatible servers need their endpoint, API key, and an explicit embedding model. Use the same embedding model for indexing and queries, and match embedding_dim to its output. Native Anthropic does not provide embeddings.

For the Chroma-based end-to-end module, choose the three model stages independently:

python
from gaik.software_components.llm import get_llm_config
from gaik.software_modules.RAG_workflow import RAGWorkflow

workflow = RAGWorkflow(
    parser_config=get_llm_config("openai", model="gpt-6-luna"),
    embedding_config=get_llm_config("google", embedding_model="gemini-embedding-001"),
    answer_config=get_llm_config("aitta"),
)

Install gaik[rag-workflow,llm-google] for this example. The parser model must accept images; answer generation only needs chat. Each omitted stage uses shared api_config or the legacy OpenAI/Azure default. When all stages are explicit, no unrelated default credentials are loaded. RAGWorkflow uses Chroma; the PostgreSQL recipe above remains the path for PgVectorStore.

Fuse the two lists yourself instead of calling store.search_hybrid():

  • the semantic list keeps its cosine similarities, which the relevance gate needs — search_hybrid() returns only RRF scores;
  • expose_ranks=True writes rank_semantic and rank_keyword into every hit's metadata, so an arm that contributes nothing shows up per result;
  • through gaik 0.7.2 search_hybrid() ignores tsquery_mode and always parses in websearch mode, which ANDs every term, and search_hybrid_weighted() fails on every call with column reference "id" is ambiguous.

From gaik 0.8.4 search_hybrid() also writes semantic_similarity (cosine) into every hit's metadata, so one call is enough when the ranks it returns are all you need.

Tune the weights and the RRF constant, do not assume them. Where the keyword arm alone reached 33.7% recall@15, equal weights at k=60 scored 8.7 points below vectors weighted 3× at k=20 (Ranker(rrf_k=...) sets k).

embedding_dim must equal the model's output, and at most 2,000. setup() builds an HNSW index on vector(N), which pgvector refuses above 2,000 dimensions. Without model=, Embedder falls back to a default that depends on the config helper, and two of them are too large: get_openai_config() gives text-embedding-3-large (3,072) and get_llm_config("google") or "vertex" gives gemini-embedding-001 (3,072). get_llm_config("openai") or "azure" gives EMBEDDING_MODEL, else text-embedding-3-small (1,536). Pass model= explicitly, and for a model above 2,000 pass vector_type="halfvec", which indexes up to 4,000 (gaik 0.8.4+; before that, pick a smaller model or build the schema yourself: references/postgres-without-gaik.md).

Three ways the keyword arm returns nothing, silently

  1. The tsquery ANDs a sentence. websearch_to_tsquery and plainto_tsquery conjoin every term, so a nine-word question only matches a passage containing all nine stems — on real prose, none. Rewrite & to | and let ts_rank_cd rank by how many terms matched and how close together. In gaik that is tsquery_mode="or".
  2. The index column is empty. On one production system the tsvector column existed as a plain nullable column instead of the GENERATED ALWAYS AS (…) STORED column the migration declared. Every row was NULL for about six months, the GIN index indexed nothing, and "hybrid" was vector search the whole time. Users noticed first.
  3. The query is in another language than the index. Postgres full-text cannot cross languages: an English question against a finnish index took the keyword arm's recall@15 from 33.7% to 5.4%, and hybrid became vector-only. Translate the query into the corpus language as a question, not a keyword list — a keyword list helped the keyword arm but wrecked the vector arm and ended 14 points below not translating at all.

A fourth, rarer: zero-width characters (U+200B) pasted in from a CMS glue onto words. Postgres does not treat them as whitespace, so the word stays unstemmed and never matches. Strip them at ingest and at query time.

Check the arms are alive, against the live database. A unit test cannot see a column the migration declared correctly and the database never got. With gaik 0.8.4+, report = store.health() does it in one call: report.ok, and str(report) names each problem (a text_search column nothing fills, another fts_language than the index, lemmas built with other settings, rows either arm cannot see, missing indexes, session settings that did not stick). By hand, from a health endpoint or a startup check:

  • rows with a NULL tsvector: 0, or exactly the rows deliberately left unindexed;
  • a word copied out of a random stored chunk, searched in keyword mode, returns that chunk;
  • one query in hybrid and in vector mode: identical result lists mean the keyword arm is contributing nothing.

Finnish and other inflected languages

Postgres' finnish snowball stemmer is inconsistent across a word family — tilinpäätös stems to tilinpäätös, tilinpäätöksen to tilinpäätöks — so a search for one form misses a passage containing the other, and nothing reports the miss. Lemmatize both sides with the same analyser. Measured on sentences queried in a different inflection than the document used: snowball matched 1 of 4, every prefix strategy 0–1 of 4, lemmas on both sides 4 of 4.

python
from gaik.software_components.RAG.finnish_text_processor import FinnishTextProcessor

# A named backend raises ImportError when it is missing; only "auto" falls back
# silently, to "simple", a tokenizer rather than a lemmatizer.
processor = FinnishTextProcessor(backend="pyvoikko", decompound=False)

store = PgVectorStore(dsn, embedding_dim=1536, fts_language="simple",
                      text_processor=processor, tsquery_mode="or")

Read references/finnish.md before indexing Finnish. The rules that each cost a debugging session:

  • Pin the backend. backend="auto" picks native libvoikko where it is installed and something else where it is not, so a laptop indexes different lemmas than the container queries with, and the arm silently stops matching. pyvoikko is pure Python and runs everywhere, including images where libvoikko is not packaged.
  • Start with decompound=False — gaik defaults to True. With both sides lemmatized, whole compounds already match, and splitting adds noise: arvonlisävero yields arvo, which reaches arvopaperi. Index and query must agree; after changing it, store.relemmatize() (0.8.4+) rebuilds the lemmas without re-embedding. Through 0.8.3 backend="auto" ignored decompound=False and split anyway.
  • Long questions flood the lemma arm. A 25-word question became ~20 OR-ed lemmas that matched 78–79% of one corpus, and Hit@1 fell from 3/5 to 2/5 while short-term queries improved sharply. Split long questions into short topic queries before searching.
  • No unaccent. In Finnish, ä/a and ö/o are different letters: tähti/tahti, sää/saa.
Show full SKILL.md (764 more words)Show less

A result count is not evidence

The vector arm answers every query with its nearest neighbours: on one corpus, a gibberish string returned 24 results and a real question 12. What separates them is the distance to the closest chunk — 17 answerable queries scored 0.26–0.56 cosine distance, 16 unanswerable ones 0.64–0.87 — so a floor at 0.60 split them cleanly. That number belongs to one embedding model and one corpus; take your own reading.

python
from gaik.software_components.RAG.relevance_gate import RelevanceGate

# answerable_best / unanswerable_best: the top similarity of search_semantic(...) per query
reading = RelevanceGate.calibrate(answerable_best, unanswerable_best, lower_is_better=False)
print(reading)                   # says so plainly when the two populations overlap
gate = RelevanceGate(reading.floor, lower_is_better=False) if reading.separated else None

if gate and not gate.is_answerable(semantic, key=lambda hit: hit[1]):
    ...                          # "nothing in the library covers this"
  • search_semantic() returns cosine similarity, higher is better, so pass lower_is_better=False. Getting the direction backwards inverts the gate with no error.
  • Put ordinary questions about the wrong subject into the unanswerable set, not only gibberish. Gibberish separates easily; a sensible off-topic question is what users type.
  • Never gate on the RRF score. It is built from rank positions, so the top hit for gibberish scores exactly what the top hit for a real question does, and fused scores are compressed — rank 1 is 1/61 and rank 20 is 1/80, a ratio of 0.76 — so a relative floor such as "30% of the top score" can never fire.
  • Judge each result, not the page. A literal keyword match is evidence whatever its cosine; a vector-only hit has to clear the floor. Gating on the top hit alone lets one good result promote a page of unrelated ones.
  • Re-take the reading whenever the embedding model or the corpus changes.

From chunks to documents

Search ranks chunks; the reader usually sees documents. Over-fetch chunks — about four times the documents wanted — collapse them onto their parent, and keep the closest distance per document: rows arrive in fused order, so a document's first chunk is its best-ranked, not necessarily its nearest.

Do not order documents by hit count. On one set that dropped document-level hit@1 from 0.79 to 0.63, because a long document that keeps mentioning a topic outranks the short one that is about it. Fusing the best-chunk rank with a length-normalised hit density raised it to 0.83 — and only for short topical queries; on long questions it moved nothing and pushed correct chunks down.

Measuring

Read references/evaluating-retrieval.md before reporting a number or adopting a change. The rules that matter most:

  • Build two query sets, short domain terms and full questions. They disagree: dropping common words and adding a reranker each helped one set and hurt the other.
  • Measure end to end. An improvement to the keyword arm measured in isolation (hit@5 0.800 → 0.829) was an end-to-end loss (0.790 → 0.768) once fused.
  • Treat a recall of 0 as a bug until proven otherwise. One benchmark read 0% because the database returned bigint ids as strings and the comparison was strict.

Rerankers

A cross-encoder reorders only the pool it is handed. On one set it raised hit@1 from 0.757 to 0.843 on full questions, did nothing for short domain terms, and made some of them worse. Ranker().rerank(query, hits, on_error="fallback") returns the input order when the model fails, but has no timeout — wrap it: asyncio.wait_for(asyncio.to_thread(ranker.rerank, query, hits), timeout=4). Hosted rerankers often ship with tight rate limits; check the quota before designing around one.

Gotchas

  • Retriever(hybrid_search=True) is not a keyword arm. It BM25-rescores the vector candidates (0.7 × vector + 0.3 × BM25), so a document the embedding missed cannot come back. PgVectorStore.search_keyword() is the real one.
  • search_semantic() defaults to threshold=0.7. Pass threshold=0.0 when fusing or gating, or the list arrives pre-cut and the gate never sees the weak cases.
  • hnsw.ef_search defaults to 40. Raising it to 100 measured recall@20 against an exact scan at 96.2% → 99.2% for +0.7 ms on one 1,500-dimension corpus. On a few thousand rows it changes nothing — the index already returns exact neighbours.
  • A filtered vector query returns fewer rows than asked unless iterative scans are on (pgvector 0.8+): PgVectorStore(..., hnsw_iterative_scan="relaxed_order") from gaik 0.8.4, or SET hnsw.iterative_scan = relaxed_order per connection or as a database default.
  • The HNSW operator class must match the query operator. An index built with *_l2_ops and queried with <=> is ignored, and every search scans the table.
  • ts_rank_cd has no IDF: in a two-word query where one word is common, that word decides the ranking. Dropping very common words from a short query's keyword text helps — check it on both query sets.
  • store.setup() creates the vector, pg_trgm and unaccent extensions. On a managed database without that privilege, have them installed and call setup(create_extensions=False).
  • Changing the embedding model or its dimension means re-embedding everything; the column is vector(N).
  • Chunk size is a token limit, not a character one. Finnish runs about 2.5 characters per token, so an English-based estimate is twice too generous and an oversized chunk fails the whole document at the embedding endpoint.

© GAIK-project, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in implementation_layer/no-code-assets/agent-plugin/skills/searching-documents of GAIK-project/gaik-toolkit.

  • SKILL.md
  • references/evaluating-retrieval.md
  • references/finnish.md
  • references/postgres-without-gaik.md

Open the folder on GitHubat commit e516ece

Compare with similar skills

Searching Documents next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Searching Documents compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Searching Documents this skillGAIK-project/gaik-toolkit100—~4.2kAutomated safety check: PassMIT
Postgrestimescale/pg-aiguide1.9k—~941Automated safety check: PassApache-2.0
Pgvector Semantic Searchtimescale/pg-aiguide1.9k—~3.8kAutomated safety check: PassApache-2.0
Postgres Hybrid Text Searchtimescale/pg-aiguide1.9k—~3.1kAutomated safety check: PassApache-2.0
RAG Implementationwshobson/agents40k9 repos~1.1kAutomated safety check: PassMIT
Neon Postgresneondatabase/agent-skills100—~4.1kAutomated safety check: NotesApache-2.0

Similar skills

  • Postgres

    timescale/pg-aiguide

    A skill your agent uses for any PostgreSQL database work — table design, indexing, data types, constraints, extensions (pgvector, PostGIS, TimescaleDB), search, and migrations.

    1.9k GitHub stars~941 tokensUpdated yesterday
    DatabasesAuto-check passed
  • Pgvector Semantic Search

    timescale/pg-aiguide

    A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.

    1.9k GitHub stars~3.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Postgres Hybrid Text Search

    timescale/pg-aiguide

    A skill your agent uses to implement hybrid search combining BM25 keyword search with semantic vector search using Reciprocal Rank Fusion (RRF).

    1.9k GitHub stars~3.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • RAG Implementation

    wshobson/agents

    Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline.

    40k GitHub starsUsed in 9 repos~1.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Neon Postgres

    neondatabase/agent-skills

    Official

    Guides and best practices for working with Lakebase Postgres on Neon: connections, pooled vs direct, schema migrations, branching, autoscaling, scale-to-zero, instant restore, read replicas, IP…

    100 GitHub stars~4.1k tokensUpdated today
    DatabasesAuto-check: notes
  • Shows how to run vector and keyword search side by side and merge their results, so retrieval catches both meaning and exact terms in RAG and search systems.

    40k GitHub starsUsed in 9 repos~497 tokens
    AI & LLM EngineeringAuto-check passed

More from GAIK-project/gaik-toolkit

All 15 skills in this repo
  • Brief To Slides

    GAIK-project/gaik-toolkit

    Builds a visual, editable PowerPoint (.pptx) deck with speaker-ready notes, exact timing, citations and a layout-checked design from a topic, an audience and a length, using only the user's own…

    100 GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Gaik Toolkit

    GAIK-project/gaik-toolkit

    GAIK toolkit overview and reference. An agent skill from GAIK-project/gaik-toolkit.

    100 GitHub stars~5.7k tokensUpdated today
    Auto-check passed
  • Extracting Structured Data

    GAIK-project/gaik-toolkit

    Extracts structured data — fields, tables, line items — out of documents into a validated schema using the gaik toolkit, and designs schemas that stay inside provider limits and produce checkable…

    100 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Parsing Documents

    GAIK-project/gaik-toolkit

    Converts PDFs, scans, and Word documents into text or markdown with the gaik toolkit's parsers, choosing the parser that will not silently destroy the structure the downstream task depends on.

    100 GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Construction Diary Creation

    GAIK-project/gaik-toolkit

    Extracts structured data from Finnish construction site daily diary audio recordings (Työmaapäiväkirja) and creates a formatted Word document with extracted fields.

    100 GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Gaik Add Examples

    GAIK-project/gaik-toolkit

    Adds or updates working code examples for GAIK toolkit components and pipelines in implementationlayer/examples/.

    100 GitHub stars~2.4k tokensUpdated today
    Auto-check: notes

Questions about Searching Documents

What does Searching Documents do?

Builds and debugs retrieval with the gaik toolkit — PgVectorStore, Ranker, FinnishTextProcessor, RelevanceGate — as hybrid search: pgvector similarity plus Postgres full-text, fused by rank, and the…. Searching Documents is an agent skill from GAIK-project/gaik-toolkit. Builds and debugs retrieval with the gaik toolkit — PgVectorStore, Ranker, FinnishTextProcessor, RelevanceGate — as hybrid search: pgvector similarity plus Postgres full-text, fused by rank, and the same design in plain SQL for a TypeScript app.

When should I use Searching Documents?

Searching Documents fits situations like: gaik is used for search; whenever Finnish text is indexed for full-text search; adding semantic; hybrid search over document chunks.

How do I install Searching Documents in Claude Code?

Run `npx skills add GAIK-project/gaik-toolkit --skill searching-documents -a claude-code`. Or copy the skill folder (implementation_layer/no-code-assets/agent-plugin/skills/searching-documents in GAIK-project/gaik-toolkit) into .claude/skills/searching-documents in your project. Claude Code loads it when a task matches its description.

How do I install Searching Documents in Codex?

Run `npx skills add GAIK-project/gaik-toolkit --skill searching-documents -a codex`. Or copy the skill folder (implementation_layer/no-code-assets/agent-plugin/skills/searching-documents in GAIK-project/gaik-toolkit) into .agents/skills/searching-documents in your project. Codex loads it when a task matches its description.

Can I use Searching Documents in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GAIK-project/gaik-toolkit --skill searching-documents -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/searching-documents, .gemini/skills/searching-documents, .github/skills/searching-documents and .opencode/skills/searching-documents in your project.

What does Searching Documents need to run?

Going by SKILL.md and its folder, Searching Documents needs the command-line tools its instructions call (pip) and credentials named AITTA_API_KEY, AITTA_API_TOKEN and AITTA_TOKEN. Our summary lists: Python 3; A credential in AITTA_API_KEY; A credential in AITTA_API_TOKEN.

Does Searching Documents access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Searching Documents safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Searching Documents use?

Searching Documents is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Searching Documents use?

About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.5k tokens, read only when the agent opens those files.

What are the alternatives to Searching Documents?

Skills that share tags, products or a category with Searching Documents: Postgres (timescale/pg-aiguide, 1.9k stars), Pgvector Semantic Search (timescale/pg-aiguide, 1.9k stars), Postgres Hybrid Text Search (timescale/pg-aiguide, 1.9k stars) and RAG Implementation (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Searching Documents?

GAIK-project (a GitHub organization) maintains it in GAIK-project/gaik-toolkit, which has 100 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 9, 2026.

Source: GAIK-project/gaik-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.