Agent skill

Ketch

by 1broseidon in 1broseidon/ketch

Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but…

MITAuto-check passedData & Analytics

Install Ketch

skills CLI
$ npx skills add 1broseidon/ketch --skill ketch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install 1broseidon/ketch ketch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/1broseidon/ketch.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ketch .claude/skills/ketch && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ketch
GitHub stars
697
Used in
1 other repo
Token cost
~3.9k tokens
SKILL.md length
1,971 words
Files
4 (incl. references)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but…

  • Works in 4 steps: which ketch succeeds → the CLI is your… → Also check for ketch's MCP tools in your… → Both live → either transport serves… → …
  • A question needs live sources: research X
  • SKILL.md covers Transport: stateless CLI by…, Glossary, How to use this skill and Non-negotiable disciplines, plus 8 more sections
  • Calls brew and go

What it does

Ketch is an agent skill from 1broseidon/ketch. Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but the CLI is the primary interface. Use when a question needs live sources: 'research X', 'what are people saying about Y', 'find docs or real-world examples for Z', 'scrape/crawl this site' — or when installing or configuring ketch backends. Routes search vs code vs docs vs scrape vs crawl, keeps every fetch inside a…

Its SKILL.md is about 3.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/surfaces.md`, `references/verbs/ketch-research.md` and `references/verbs/setup.md`).

It sits in Data & Analytics, covering Web scraping. It works with Model Context Protocol, Visual Studio Code and Go. The repository describes itself as: Fast, stateless CLI for web search and scrape. Built for AI agents. The licence is MIT.

When your agent uses it

  • A question needs live sources: research X
  • What are people saying about Y
  • Real-world examples for Z
  • Scrape/crawl this site —

Example prompts

  • “research X”
  • “what are people saying about Y”
  • “find docs or real-world examples for Z”
  • “/ketch”

Requirements

  • Docker

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. which ketch succeeds → the CLI is your transport: --json on every call, exit codes as control flow.
  2. Also check for ketch's MCP tools in your tool list — search, code, docs, scrape, crawl and tag from a server named ketch (in Claude Code…
  3. Both live → either transport serves research calls, but know the tradeoff: a running MCP server holds the single-process page-cache lock…
  4. Neither CLI nor MCP tools → ketch is not installed. Offer brew install ketch or go install github.com/1broseidon/ketch@latest — an…

What it can do on your machine

Read from SKILL.md and the folder at commit d66bde7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • brew
    • go

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ketch loads about 3.9k tokens when it runs, and up to ~10k if it reads all its reference files. Until then it costs about 168 tokens; SKILL.md has 1,971 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~168
When it runs · the whole SKILL.md, loaded when a task matches
~3.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~10k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from 1broseidon/ketch at commit d66bde7, republished under its MIT licence (© 1broseidon). 1,971 words, ~3,919 tokens.

Download SKILL.mdSave it as .claude/skills/ketch/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
ketch
description
Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but the CLI is the primary interface. Use when a question needs live sources: 'research X', 'what are people saying about Y', 'find docs or real-world examples for Z', 'scrape/crawl this site' — or when installing or configuring ketch backends. Routes search vs code vs docs vs scrape vs crawl, keeps every fetch inside a token budget, turns error prefixes into control flow, and produces cited syntheses. Not for local codebase search, private repos, or pages behind auth.
version
0.1.0

Ketch

Route every live-source question to one of ketch's five research surfaces — search, code, docs, scrape, crawl — over the transport the operator gave you, with a token budget on every fetch and a source URL on every claim. ketch is one stateless binary — call, result, exit — with web search, OSS code grep, curated library docs, and page/site extraction together, so a complete research pipeline needs no other tool and no daemon.

Transport: stateless CLI by default, MCP when the operator wired it

The CLI is ketch's identity: call → result → exit, --json on every call, exit codes as control flow, zero daemon. That is the default transport and the zero-infrastructure path. The MCP server is a supported alternative for operators who want it — never a prerequisite.

Decide once per session, before the first call:

  1. which ketch succeeds → the CLI is your transport: --json on every call, exit codes as control flow.
  2. Also check for ketch's MCP tools in your tool list — search, code, docs, scrape, crawl and tag from a server named ketch (in Claude Code: mcp__ketch__search, …). Present → the operator wired them up on purpose, and using them for research calls is correct and good: structured output, per-URL errors, no shell round-trip. Do not shell out around tools the operator set up.
  3. Both live → either transport serves research calls, but know the tradeoff: a running MCP server holds the single-process page-cache lock, so concurrent CLI scrapes silently run cache-disabled.
  4. Neither CLI nor MCP tools → ketch is not installed. Offer brew install ketch or go install github.com/1broseidon/ketch@latest — an operator action: propose, wait for confirmation.

The rule: use the transport the operator gave you — when both are live, either is fine for research calls, and operator actions are always CLI.

tag is the one exception to the operator-action rule below: it changes local state but the agent is both writer and reader, so it is published over MCP as well as the CLI.

Config discovery is CLI regardless of transport: ketch config prints effective settings and available backends as JSON; there is no config tool over MCP. Operator actions — config set, cache, browser install, crawl --background/status/stop, doctor — are deliberately not in MCP. They are always CLI.

Glossary

Use only these terms in ketch output.

TermMeaning
surfaceOne of the five research operations: search, code, docs, scrape, crawl
tagA label filed across surfaces: search hits, code hits, docs chunks, scraped pages, crawled pages. Not a research surface — it answers "what did I already find?", not "what is out there?"
transportHow a surface is called: the CLI binary (default) or the optional MCP tools
backendThe provider behind a surface: auto (default; a fallback chain, not a provider)/brave/ddg/searxng/exa/firecrawl/keenable/tavily/parallel/serpbase/degoog/serply/youcom (search), grepapp/sourcegraph/github (code), context7 (docs)
operator actionA system-managing or diagnostic command — config set, cache, browser install, background crawls, doctor — CLI-only by design
error prefixThe stable class on every ketch error: CLI exit codes 2–6, mirrored as the bracketed prefix opening every MCP tool error — [validation], [not_found], [upstream], [precondition], [cancelled]
fan-outHow many queries are searched and URLs scraped under one plan
token budgetThe per-call output bound: max_chars/trim on scrapes, tokens on docs, limit/--minimal on lists
probeOne cheap read-only call that tests whether a surface is configured and reachable

How to use this skill

  • Default: answer one question with one or two routed calls. Use the surface routing table, token budgets, and error control flow below.
  • ketch research <question>: deep multi-source research — search fan-out → scrape top hits → optional code/docs corroboration → synthesized, cited answer. Read references/verbs/ketch-research.md.
  • ketch setup: configure backends with the operator — probe current state, propose exact commands, mutate only on confirmation. Read references/verbs/setup.md. Enter this verb whenever any call returns [precondition] / exit 5.

One question = one plan. Escalate a default run into ketch research when the first search shows the answer is contested, multi-part, or needs corroboration.

Non-negotiable disciplines

  1. Use the transport the operator gave you. The CLI is the default; MCP tools in your list mean the operator opted in — use them for research rather than shelling out around them. When both are live, either serves research calls (a running MCP server holds the page-cache lock, so concurrent CLI scrapes run uncached); operator actions — config, cache, browser, background crawls, doctor — are always CLI.
  2. Bound every fetch. max_chars 4000–8000 plus trim on any scrape of a page you have not seen — an unguarded page can cost ~25k tokens. Skipping the cap requires a stated one-line reason ("known ~200-word page").
  3. Cite every claim. A research synthesis without source URLs is not a deliverable.
  4. Error prefixes are control flow. Classify before reacting. Never retry [validation] or [not_found] unchanged.
  5. Propose, then mutate. config set, browser install, docker runs, installs — only after the operator confirms the exact command. Never touch a value that is already configured and working.
  6. The binary outranks this file. ketch config and --help are ground truth; where they disagree with a table here, trust the binary and flag the skill as drifted.

Gold decision trace

text
Request: "ketch research — do people actually use Go's iter.Seq in real projects, and what are the gotchas?"

Transport: operator wired mcp__ketch__* into this session → honor it; research calls go over MCP.
Plan: 2 queries · scrape top 3 · max_chars 6000 + trim · ≤8 calls

search {query: "Go iter.Seq real-world experience gotchas", limit: 5}
  → "[upstream] ddg rate limited" → an explicit backend failed; rotate to another
    usable provider from available_backends, retry once:
    search {query: ..., backend: "brave", limit: 5} → ok
search {query: "Go range-over-func adoption production", backend: "brave", limit: 5}
  → 10 results, 8 unique hosts → picked 3: official blog post, one experience
    report, one issue thread (primary sources over aggregators)

scrape {urls: [u1, u2, u3], max_chars: 6000, trim: true}
  → isError=false; checked results[] one by one: u1, u2 ok;
    u3.error = "[upstream] … 503" → dropped, will be named in synthesis

code {query: "iter.Seq", lang: "go", limit: 3}      # corroborate real usage
  → 3 repos with file/line URLs

Synthesis: five claims, each cited to its URL; u3 listed as unretrieved;
one conflict between u1 and u2 stated and attributed, not averaged.
Budget: 5 of 8 calls (the rate-limited attempt counts).

Surface routing

First match wins:

The question needsSurfaceNot
Current web pages, opinions, news, comparisonssearchdocs — that is curated library docs only
How real projects call an APIcodesearch — blogs talk about code; code greps public OSS repos via grep.app
A library's own documentation, version-awaredocsscrape of the docs site — docs is already extracted and token-budgeted
The content of a URL you already holdscrapesearch — never re-find a known URL
Many pages from one sitecrawllooped scrape — crawl dedupes, bounds, and streams
Anything you already found earlier in this projecttag (operation: show)search again — you already paid for these once

In reverse: search finds URLs; scrape reads them; crawl reads a site; code reads public source; docs reads library docs. search with scrape: true fuses the first two when you will want full content from every hit — budget it like a scrape.

Keeping a working set with tag

On work that spans more than one session or more than a handful of sources, pass --tag <name> (CLI) or tag (MCP) on the calls you will want again, naming the project or scope of work — remote-access, not results-3.

It works on every surface, and one tag holds them all: search, code, docs, scrape, crawl. So a project tag ends up holding the vendor's documentation, the code that calls it, and the write-up that explained the undocumented flag, in one list. Later, tag show returns that as an llms.txt-shaped index — titles, URLs, descriptions — limited to the newest 50, and you re-read one source instead of re-running the searches that found them. Tag the sources you actually used, not every hit.

Use --limit N / MCP limit for a smaller view; 0 explicitly requests all. entries counts the whole tag and shown counts returned pages. Read the index before searching again, but do not request all of a large tag by default.

The index is durable and outlives the cached page bodies, so it still answers days later. cached: false means no fresh body was confirmed, not a dead link. When cache_status is unavailable, the page cache could not be checked; bookmarks still work and sources can still be fetched. Entries from code, docs and unscraped search hits start that way by nature, since those calls return snippets rather than fetched pages; scrape the URL when you want the whole thing. Read the index first, then fetch only what you need: assembling the whole tag defeats the point.

Check bookmark diagnostics separately from research success: CLI warnings go to stderr (structured under --json), and MCP returns warnings. A failed write does not discard the useful research result. Durable bookmarks use a separate file; test labs must set KETCH_TAGS_PATH as well as isolating the page cache.

Show full SKILL.md (678 more words)Show less

Token budgets

CallBound withMeasured cost
search, limit 5limit~1.4 KB
code, limit 3limit~0.7 KB
docs, default budgettokens (default 4000)~3.3 KB
scrape, unknown pagemax_chars 4000–8000 + trimunguarded: up to ~100 KB (~25k tokens)
crawl (MCP)max_pages + per-page max_chars30 pages default, 100 cap, 3-min wall clock
tag show / MCP tag showlimit (default 50; 0 all)bounded source metadata; no page bodies
Any CLI list--minimalroughly halves output

Error control flow

Exit (CLI)Prefix (MCP)MeaningDo
2[validation]Bad inputFix the call; retrying unchanged can never succeed
3[not_found]Nothing matchedChange the query or selector; not an outage
4[upstream]Backend or network failureExplicit backend: rotate to another provider (available_backends in ketch config) or retry once. auto already fell through every usable provider: retry once, then report the outage
5[precondition]Operator config missingStop researching; enter ketch setup
6[cancelled]Cancelled or timed outRerun with smaller scope

Situations → class: unknown backend, regexp on github → [validation]. Selector matched nothing, a repo the code backend does not have → [not_found] (for repo, change the backend). ddg rate limit (it rate-limits readily under fan-out), DNS failure, grepapp's intermittent 504 → [upstream], rotate or retry once. Missing API key, docs backend local (planned, unimplemented), force_browser with no browser configured → [precondition]. One asymmetry: a CLI crawl interrupted by SIGINT exits 0 with partial results, by design.

Gotchas

Detail for each lives in references/surfaces.md.

  • Scraping a bare domain auto-probes /llms.txt and may silently return that instead of the homepage — the title field reveals the swap; no_llms_txt opts out.
  • docs is a two-step: resolve the name → vet the matches → fetch by library ID. Resolve never returns empty — garbage in gets confident fuzzy matches out, so check the name, not just the trust score.
  • Batch scrape reports per-URL failures inside a successful call: isError=false with results[].error set. Check every entry.
  • regexp works on grepapp and sourcegraph only; github rejects it with a pointer to those backends.
  • Background crawls (--background, status, stop) are CLI-only; the MCP crawl is synchronous and capped.
  • The page cache (bbolt, 72h default TTL) is single-process: a long-running MCP server holds the lock, so concurrent CLI scrapes silently run cache-disabled — ketch doctor reports the cache as locked by another process. Running the server degrades the CLI; prefer CLI-only when both would run long-term.

BAD/GOOD contrasts

BAD: scrape {url: "https://docs.example.com"} — no bound; you get llms.txt or ~25k tokens, whichever is worse. GOOD: scrape {url: "https://docs.example.com/quickstart", max_chars: 6000, trim: true} — plus no_llms_txt: true when you want the page itself, not the site's llms.txt.

BAD: Telling a user they must run an MCP server to use ketch with agents — the CLI plus a prompt block is the zero-infrastructure path, and a long-running server holds the page-cache lock against every CLI call. GOOD: CLI by default; MCP when the operator wired it — and when mcp__ketch__* tools are in your list, use them for research instead of shelling out around the operator's setup.

BAD: [upstream] ddg rate limited → retry the identical call three times. GOOD: Rotate — backend: "brave" (or another provider from available_backends; auto is a chain, not a rotation target) — retry once, and note the swap. When auto itself failed, it already tried every usable provider: retry once, then report the outage.

BAD: Fetch docs from resolve's first match because its trust score is high, even though its name is not the library you asked about. GOOD: Vet name + snippet count + trust; if no match names the intended library, say so instead of fetching junk docs.

Reference loading

  • ketch research … → read references/verbs/ketch-research.md before starting.
  • ketch setup, any [precondition]/exit 5, or an install → read references/verbs/setup.md.
  • Full flag/param tables, CLI↔MCP name mapping, backend/key matrix, or a surface behaving oddly → read references/surfaces.md.

Scope

In scope: the five research surfaces over both transports, the research and setup verbs, token budgets, error-prefix control flow, backend configuration. Out of scope: local or private codebase search (use repo tools), pages behind auth or paywalls, bulk archival crawling beyond the caps, browser automation beyond ketch's headless-rendering fallback.

Bound every fetch; cite every claim.

© 1broseidon, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/ketch of 1broseidon/ketch.

  • SKILL.md
  • references/surfaces.md
  • references/verbs/ketch-research.md
  • references/verbs/setup.md

Open the folder on GitHubat commit d66bde7

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in 1broseidon/ketch, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Ketch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ketch compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ketch this skill1broseidon/ketch6971 repos~3.9kAutomated safety check: PassMIT
AnakinscraperAnakin-Inc/anakin4.5k—~859Automated safety check: PassAGPL-3.0
Querying Indonesian Gov Datasuryast/indonesia-gov-apis172—~997Automated safety check: PassMIT
Media Crawlertsingyuai/growth-lab2k—~731Automated safety check: PassApache-2.0
Xquik MCPXquik-dev/x-twitter-scraper2091 repos~997Automated safety check: PassMIT
Erd Studio Setupliam-machine/erd-studio165—~8.5kAutomated safety check: PassCustom licence

Similar skills

  • Anakinscraper

    Anakin-Inc/anakin

    Scrape any website into clean markdown or structured JSON. An agent skill from Anakin-Inc/anakin.

    4.5k GitHub stars~859 tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Querying Indonesian Gov Data

    suryast/indonesia-gov-apis

    Query 57 Indonesian government APIs and data sources — BPJPH halal certification, BPOM food safety, OJK financial legality, BPS statistics, BMKG weather/earthquakes, Bank Indonesia exchange rates…

    172 GitHub stars~997 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Media Crawler

    tsingyuai/growth-lab

    Install, authenticate, configure, operate, and troubleshoot the external MediaCrawler client shared by Douyin, Kuaishou, Bilibili, Weibo, Tieba, and Zhihu collectors.

    2k GitHub stars~731 tokensUpdated 10 days ago
    Data & AnalyticsAuto-check passed
  • Xquik MCP

    Xquik-dev/x-twitter-scraper

    Connect, verify, and troubleshoot Xquik's remote MCP server.

    209 GitHub starsUsed in 1 repo~997 tokens
    Backend & APIsAuto-check passed
  • Erd Studio Setup

    liam-machine/erd-studio

    Friendly, step-by-step setup for ERD Studio in an existing dbt project, for people who may be new to dbt or data modelling.

    165 GitHub stars~8.5k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Apify Product Data Setup

    apify/awesome-skills

    Official

    Wire an AI agent to live e-commerce product data using Apify's E-commerce Scraping Tool over MCP, either as runtime tool calls or as a scheduled refresh into a vector store.

    264 GitHub starsUsed in 1 repo~2k tokens
    Data & AnalyticsAuto-check passed

Questions about Ketch

What does Ketch do?

Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but…. Ketch is an agent skill from 1broseidon/ketch. Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but the CLI is the primary interface.

When should I use Ketch?

Ketch fits situations like: A question needs live sources: research X; what are people saying about Y; real-world examples for Z; scrape/crawl this site —.

How do I install Ketch in Claude Code?

Run `npx skills add 1broseidon/ketch --skill ketch -a claude-code`. Or copy the skill folder (skills/ketch in 1broseidon/ketch) into .claude/skills/ketch in your project. Claude Code loads it when a task matches its description.

How do I install Ketch in Codex?

Run `npx skills add 1broseidon/ketch --skill ketch -a codex`. Or copy the skill folder (skills/ketch in 1broseidon/ketch) into .agents/skills/ketch in your project. Codex loads it when a task matches its description.

Can I use Ketch in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add 1broseidon/ketch --skill ketch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ketch, .gemini/skills/ketch, .github/skills/ketch and .opencode/skills/ketch in your project.

What does Ketch need to run?

Going by SKILL.md and its folder, Ketch needs the command-line tools its instructions call (brew and go). Our summary lists: Docker.

Does Ketch access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ketch safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ketch use?

Ketch is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ketch use?

About 3.9k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.4k tokens, read only when the agent opens those files.

What are the alternatives to Ketch?

Skills that share tags, products or a category with Ketch: Anakinscraper (Anakin-Inc/anakin, 4.5k stars), Querying Indonesian Gov Data (suryast/indonesia-gov-apis, 172 stars), Media Crawler (tsingyuai/growth-lab, 2k stars) and Xquik MCP (Xquik-dev/x-twitter-scraper, 209 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ketch?

1broseidon (a GitHub user) maintains it in 1broseidon/ketch, which has 697 GitHub stars. The repository was last updated on October 8, 2026.

Source: 1broseidon/ketch on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.