Agent skill

Specification Discovery Crawler

by NyxFoundation in NyxFoundation/speca

Crawls from a starting URL to find and list technical specification, whitepaper and architecture documents, writing the found URLs to a JSON file.

MITAuto-check passedResearch & Science

Install Specification Discovery Crawler

skills CLI
$ npx skills add NyxFoundation/speca --skill spec-discovery -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NyxFoundation/speca spec-discovery --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NyxFoundation/speca.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/spec-discovery .claude/skills/spec-discovery && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
spec-discovery
GitHub stars
459
Token cost
~669 tokens
SKILL.md length
255 words
Files
1
Skills in repo
2
Repo updated
First seen
Licence
MIT

At a glance

Crawls from a starting URL to find and list technical specification, whitepaper and architecture documents, writing the found URLs to a JSON file.

  • Works in 6 steps: Initial Fetch: Use mcpfetchfetch with… → Link Extraction: Parse the returned… → Recursive Fetch: For each discovered… → …
  • Finding the whitepaper and docs for a project starting from its website
  • SKILL.md covers Mindset, Goal, Tools and Input, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Given a seed URL as JSON, the agent acts as a web researcher hunting for specification documents. It fetches the page as Markdown with the MCP fetch tool, looks for links whose text suggests specifications, whitepapers, yellow papers, architecture, protocol, technical details or docs, and follows promising links recursively to a depth of two or three levels.

When the fetch tool fails, for example with a 403, an empty response or content that needs JavaScript, the agent falls back to browser navigation, scrolling and clicking for those URLs. Found pages and PDFs are deduplicated and written as a JSON object holding the start URL and the list of found specs to the path in the OUTPUT_FILE environment variable. Its repository is an auditing framework that turns specifications into checklists.

When your agent uses it

  • Finding the whitepaper and docs for a project starting from its website
  • Collecting specification URLs before an audit
  • Mapping which technical documents a project publishes

Example prompts

  • “Find all specification and whitepaper links starting from https://example.com/project.”
  • “Crawl this protocol's docs site and list its architecture documents.”
  • “Collect technical spec URLs for this project and write them to a JSON file.”

Requirements

  • The MCP fetch tool
  • Browser automation tools as a fallback
  • An OUTPUT_FILE environment variable for the result path
  • Pre-approved tools (allowed-tools): mcp__fetch__fetch, browser_navigate, browser_scroll, browser_click, browser_view, write

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Initial Fetch: Use mcpfetchfetch with the provided url to retrieve the page content as Markdown.
  2. Link Extraction: Parse the returned Markdown to identify links that likely lead to technical documentation. Keywords to look for include…
  3. Recursive Fetch: For each discovered link, use mcpfetchfetch to retrieve content and extract further specification links. Limit depth to…
  4. Fallback to Browser: If mcpfetchfetch fails (403, empty response, JavaScript-required content), fall back to browser tools…
  5. Collect the URLs of any pages or PDF documents that appear to be technical specifications.
  6. Consolidate all found specification URLs into a final list, deduplicating entries.

What it can do on your machine

Read from SKILL.md and the folder at commit d173893. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • mcp__fetch__fetch
    • browser_navigate
    • browser_scroll
    • browser_click
    • browser_view
    • write

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Specification Discovery Crawler loads about 669 tokens when it runs. Until then it costs about 19 tokens; SKILL.md has 255 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~19
When it runs · the whole SKILL.md, loaded when a task matches
~669

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NyxFoundation/speca at commit d173893, republished under its MIT licence (© NyxFoundation). 255 words, ~669 tokens.

Download SKILL.mdSave it as .claude/skills/spec-discovery/SKILL.md (or your agent's skills folder).
name
spec-discovery
description
Crawl and discover specification documents from a given URL.
allowed-tools
mcp__fetch__fetch, browser_navigate, browser_scroll, browser_click, browser_view, write
context
fork

SKILL: Specification Discovery

Mindset

You are a meticulous Web Researcher tasked with finding all relevant technical specification documents starting from a seed URL. Your goal is to be comprehensive and follow all promising links.

Goal

Given a starting URL, navigate the website to find and list all URLs pointing to technical specifications, whitepapers, or architectural documents. These are often found in sections like "Developers", "Documentation", "Technology", or "Whitepaper".

Tools

  • Primary: mcp__fetch__fetch - Use for static documentation pages (returns Markdown, fast and efficient)
  • Fallback: Browser tools (browser_navigate, browser_scroll, browser_click, browser_view) - Use for dynamic/JavaScript-rendered pages or when mcp__fetch__fetch fails (403, timeout, JS-required)

Input

A JSON object containing the starting URL:

json
{
  "url": "https://example.com/project"
}

Procedure

  1. Initial Fetch: Use mcp__fetch__fetch with the provided url to retrieve the page content as Markdown.
  2. Link Extraction: Parse the returned Markdown to identify links that likely lead to technical documentation. Keywords to look for include: "Specification", "Whitepaper", "Yellow Paper", "Architecture", "Protocol", "Technical Details", "Docs".
  3. Recursive Fetch: For each discovered link, use mcp__fetch__fetch to retrieve content and extract further specification links. Limit depth to 2-3 levels.
  4. Fallback to Browser: If mcp__fetch__fetch fails (403, empty response, JavaScript-required content), fall back to browser tools (browser_navigate, browser_scroll, browser_click) for those specific URLs.
  5. Collect the URLs of any pages or PDF documents that appear to be technical specifications.
  6. Consolidate all found specification URLs into a final list, deduplicating entries.

Output Format

Return a JSON object containing a list of found specification URLs. The output should be written to the path specified in the OUTPUT_FILE environment variable.

json
{
  "start_url": "https://example.com/project",
  "found_specs": [
    {
      "url": "https://example.com/project/docs/specification.md",
      "title": "Project Specification"
    },
    {
      "url": "https://example.com/project/whitepaper.pdf",
      "title": "Project Whitepaper"
    }
  ],
  "metadata": {
    "timestamp": "...",
    "urls_visited": []
  }
}

© NyxFoundation, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/spec-discovery of NyxFoundation/speca.

Open the folder on GitHubat commit d173893

Compare with similar skills

Specification Discovery Crawler next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Specification Discovery Crawler compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Specification Discovery Crawler this skillNyxFoundation/speca459—~669Automated safety check: PassMIT
Deep Researchaffaan-m/ECC276k2 repos~590Automated safety check: PassMIT
Browser SearchJohell1NS/browser-search529—~2.5kAutomated safety check: PassMIT
Deep Researchjordan-gibbs/hyperresearch3.8k—~1.2kAutomated safety check: PassMIT
Rival Search MCPdamionrashford/RivalSearchMCP132—~796Automated safety check: PassMIT
Deep Research MCP Guidepminervini/deep-research-mcp113—~5.8kAutomated safety check: PassMIT

Similar skills

  • Deep Research

    affaan-m/ECC

    使用firecrawl和exa MCPs进行多源深度研究。搜索网络、综合发现并交付带有来源引用的报告。适用于用户希望对任何主题进行有证据和引用的彻底研究时。

    276k GitHub starsUsed in 2 repos~590 tokens
    Research & ScienceAuto-check passed
  • Browser Search

    Johell1NS/browser-search

    Multi-engine web search (SearXNG) + browsing/scraping (Camofox, CloakBrowser).

    529 GitHub stars~2.5k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check passed
  • Deep Research

    jordan-gibbs/hyperresearch

    Deep research with hyperresearch, for Claude Code and OpenAI Codex.

    3.8k GitHub stars~1.2k tokensUpdated 7 days ago
    Research & ScienceAuto-check passed
  • Rival Search MCP

    damionrashford/RivalSearchMCP

    Deterministic deep research via RivalSearchMCP. An agent skill from damionrashford/RivalSearchMCP.

    132 GitHub stars~796 tokensUpdated today
    Research & ScienceAuto-check passed
  • Deep Research MCP Guide

    pminervini/deep-research-mcp

    Explains how to run, integrate and debug the deep-research-mcp project through its CLI, Python API or MCP server, with OpenAI, Gemini and DR-Tulu backends.

    113 GitHub stars~5.8k tokensUpdated 11 days ago
    Research & ScienceAuto-check passed
  • Interceptor Research

    Hacker-Valley-Media/Interceptor

    Deep web-research methodology for the interceptor browser surface — investigate a topic the way researchers, intelligence analysts, investigative journalists, private investigators, and OSINT…

    519 GitHub stars~3.8k tokensUpdated 7 days ago
    Research & ScienceAuto-check passed

More from NyxFoundation/speca

  • Turns one specification document into program graphs following Nielson and Nielson's definition, writing each graph as a Mermaid file, one per functional unit.

    459 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Specification Discovery Crawler

What does Specification Discovery Crawler do?

Crawls from a starting URL to find and list technical specification, whitepaper and architecture documents, writing the found URLs to a JSON file. Given a seed URL as JSON, the agent acts as a web researcher hunting for specification documents. It fetches the page as Markdown with the MCP fetch tool, looks for links whose text suggests specifications, whitepapers, yellow papers, architecture, protocol, technical details or docs, and follows promising links recursively to a depth of two or three levels.

When should I use Specification Discovery Crawler?

Specification Discovery Crawler fits situations like: finding the whitepaper and docs for a project starting from its website; collecting specification URLs before an audit; mapping which technical documents a project publishes.

How do I install Specification Discovery Crawler in Claude Code?

Run `npx skills add NyxFoundation/speca --skill spec-discovery -a claude-code`. Or copy the skill folder (.claude/skills/spec-discovery in NyxFoundation/speca) into .claude/skills/spec-discovery in your project. Claude Code loads it when a task matches its description.

How do I install Specification Discovery Crawler in Codex?

Run `npx skills add NyxFoundation/speca --skill spec-discovery -a codex`. Or copy the skill folder (.claude/skills/spec-discovery in NyxFoundation/speca) into .agents/skills/spec-discovery in your project. Codex loads it when a task matches its description.

Can I use Specification Discovery Crawler in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NyxFoundation/speca --skill spec-discovery -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spec-discovery, .gemini/skills/spec-discovery, .github/skills/spec-discovery and .opencode/skills/spec-discovery in your project.

What does Specification Discovery Crawler need to run?

SKILL.md names no scripts, command-line tools or credentials: Specification Discovery Crawler is instructions for the agent only. Our summary lists: The MCP fetch tool; Browser automation tools as a fallback; An OUTPUT_FILE environment variable for the result path. Its frontmatter pre-approves these tools: mcp__fetch__fetch, browser_navigate, browser_scroll, browser_click, browser_view, write.

Does Specification Discovery Crawler access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Specification Discovery Crawler safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Specification Discovery Crawler use?

Specification Discovery Crawler is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Specification Discovery Crawler use?

About 669 tokens (SKILL.md is roughly 2.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Specification Discovery Crawler?

Skills that share tags, products or a category with Specification Discovery Crawler: Deep Research (affaan-m/ECC, 276k stars), Browser Search (Johell1NS/browser-search, 529 stars), Deep Research (jordan-gibbs/hyperresearch, 3.8k stars) and Rival Search MCP (damionrashford/RivalSearchMCP, 132 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Specification Discovery Crawler?

NyxFoundation (a GitHub organization) maintains it in NyxFoundation/speca, which has 459 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on August 12, 2026.

Source: NyxFoundation/speca on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.