Agent skill

Web Archive Scraper

by gooseworks-ai in gooseworks-ai/goose-skills

Search the Wayback Machine for archived versions of websites.

MITAuto-check passedData & Analytics

Install Web Archive Scraper

skills CLI
$ npx skills add gooseworks-ai/goose-skills --skill web-archive-scraper -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install gooseworks-ai/goose-skills web-archive-scraper --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/gooseworks-ai/goose-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/research-tools/capabilities/web-archive-scraper .claude/skills/web-archive-scraper && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
web-archive-scraper
GitHub stars
1.2k
Used in
1 other repo
Token cost
~901 tokens
SKILL.md length
270 words
Files
3 (incl. scripts)
Skills in repo
273
Repo updated
First seen
Licence
MIT

At a glance

Search the Wayback Machine for archived versions of websites.

  • Works in 5 steps: CDX API search — Queries… → Filtering — Filters by date range, HTTP… → Dedup — Collapses to one snapshot per… → …
  • Tasks that involve Web scraping
  • SKILL.md covers Quick Start, How It Works, CLI Reference and Output Schema, plus 2 more sections
  • Runs Python scripts from its folder; calls python3; reaches botkeeper.com and web.archive.org

What it does

Web Archive Scraper is an agent skill from gooseworks-ai/goose-skills. Search the Wayback Machine for archived versions of websites. Extract cached pages, customer lists, testimonials, and partner directories from sites that have changed or gone offline. Uses the free CDX API — no API key needed.

Its SKILL.md is about 900 tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `scripts/search_archive.py` and `skill.meta.json`).

It sits in Data & Analytics, covering Web scraping. The repository describes itself as: Library of Growth & GTM skills + data APIs for Claude Code, Codex, Cursor to run ads, social, content, lead gen, seo and data scraping. The licence is MIT.

When your agent uses it

  • Tasks that involve Web scraping

Example prompts

  • “/web-archive-scraper”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. CDX API search — Queries web.archive.org/cdx/search/cdx for snapshots matching the URL
  2. Filtering — Filters by date range, HTTP status code, and MIME type
  3. Dedup — Collapses to one snapshot per day by default to avoid redundant results
  4. Content fetch — Optionally fetches the raw archived HTML (using id_ modifier to skip Wayback toolbar)
  5. Text extraction — Strips HTML tags for readable text output when fetching content

What it can do on your machine

Read from SKILL.md and the folder at commit c650c6d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • botkeeper.com
    • web.archive.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Web Archive Scraper loads about 901 tokens when it runs. Until then it costs about 62 tokens; SKILL.md has 270 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~62
When it runs · the whole SKILL.md, loaded when a task matches
~901

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from gooseworks-ai/goose-skills at commit c650c6d, republished under its MIT licence (© gooseworks-ai). 270 words, ~901 tokens.

Download SKILL.mdSave it as .claude/skills/web-archive-scraper/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
web-archive-scraper
description
Search the Wayback Machine for archived versions of websites. Extract cached pages, customer lists, testimonials, and partner directories from sites that have changed or gone offline. Uses the free CDX API — no API key needed.

Web Archive Scraper

Search the Wayback Machine (Internet Archive) for archived snapshots of websites. Fetch cached page content to find customer lists, testimonials, partner directories, and other information from sites that have changed or shut down.

Quick Start

Only dependency is requests. No API key needed.

bash
# Find all snapshots of a URL
python3 skills/web-archive-scraper/scripts/search_archive.py \
  --url "https://botkeeper.com/customers"

# Search with date range
python3 skills/web-archive-scraper/scripts/search_archive.py \
  --url "https://botkeeper.com" --from 2025-01-01 --to 2026-02-01

# Search all pages under a domain (prefix match)
python3 skills/web-archive-scraper/scripts/search_archive.py \
  --url "https://botkeeper.com" --match prefix --limit 50

# Fetch the actual archived page content
python3 skills/web-archive-scraper/scripts/search_archive.py \
  --url "https://botkeeper.com/customers" --fetch

# Output formats
python3 skills/web-archive-scraper/scripts/search_archive.py --url URL --output json
python3 skills/web-archive-scraper/scripts/search_archive.py --url URL --output csv
python3 skills/web-archive-scraper/scripts/search_archive.py --url URL --output summary

How It Works

  1. CDX API search — Queries web.archive.org/cdx/search/cdx for snapshots matching the URL
  2. Filtering — Filters by date range, HTTP status code, and MIME type
  3. Dedup — Collapses to one snapshot per day by default to avoid redundant results
  4. Content fetch — Optionally fetches the raw archived HTML (using id_ modifier to skip Wayback toolbar)
  5. Text extraction — Strips HTML tags for readable text output when fetching content

CLI Reference

FlagDefaultDescription
--urlrequiredTarget URL to search in the archive
--matchexactMatch type: exact, prefix, host, domain
--fromnoneStart date (YYYY-MM-DD)
--tononeEnd date (YYYY-MM-DD)
--limit25Max number of snapshots to return
--fetchfalseFetch and display the content of the most recent snapshot
--fetch-allfalseFetch content of ALL matched snapshots (use with small --limit)
--status200HTTP status filter (set to "any" to include all)
--outputjsonOutput format: json, csv, summary
--collapsedayDedup level: none, day, month, year

Output Schema

json
{
  "url": "https://botkeeper.com/customers",
  "timestamp": "20250915143022",
  "datetime": "2025-09-15T14:30:22",
  "status_code": "200",
  "mime_type": "text/html",
  "archive_url": "https://web.archive.org/web/20250915143022/https://botkeeper.com/customers",
  "raw_url": "https://web.archive.org/web/20250915143022id_/https://botkeeper.com/customers",
  "content": "..."
}

The content field is only populated when --fetch or --fetch-all is used.

Cost

Free. The Wayback Machine CDX API requires no authentication or API key. Rate limit is ~15 requests/minute.

Common Use Cases

  • Find customer lists from shut-down companies (e.g., botkeeper.com)
  • Recover testimonials/case studies before a site redesign
  • Track how a competitor's messaging changed over time
  • Find partner directories that have been removed

© gooseworks-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/research-tools/capabilities/web-archive-scraper of gooseworks-ai/goose-skills.

  • SKILL.md
  • scripts/search_archive.py
  • skill.meta.json

Open the folder on GitHubat commit c650c6d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in gooseworks-ai/goose-skills, which our catalogue first saw on October 9, 2026.

Compare with similar skills

Web Archive Scraper next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Web Archive Scraper compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Web Archive Scraper this skillgooseworks-ai/goose-skills1.2k1 repos~901Automated safety check: PassMIT
Tmuxtrpc-group/trpc-agent-go1.9k23 repos~868Automated safety check: PassApache-2.0
Ketch1broseidon/ketch7021 repos~3.9kAutomated safety check: PassMIT
Boss Zhipin Scrapereatmoreduck/boss-zhipin-scraper1.5k—~2.6kAutomated safety check: PassMIT
Crawl4AI Web Scrapingsmallnest/goclaw5991 repos~2.5kAutomated safety check: PassMIT
Axyusukebe/ax7191 repos~918Automated safety check: PassMIT

Similar skills

  • Tmux

    trpc-group/trpc-agent-go

    Remote-control tmux sessions for interactive CLIs by sending keystrokes and scraping pane output.

    1.9k GitHub starsUsed in 23 repos~868 tokens
    Data & AnalyticsAuto-check passed
  • Ketch

    1broseidon/ketch

    Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but…

    702 GitHub starsUsed in 1 repo~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Boss Zhipin Scraper

    eatmoreduck/boss-zhipin-scraper

    Scrape BOSS直聘 (job listing site) via Chrome CDP. An agent skill from eatmoreduck/boss-zhipin-scraper.

    1.5k GitHub stars~2.6k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Crawl4AI Web Scraping

    smallnest/goclaw

    Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.

    599 GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed
  • Ax

    yusukebe/ax

    Use the ax CLI instead of curl + throwaway parsing scripts whenever you fetch a URL, explore an unknown web page, or extract structured data from HTML.

    719 GitHub starsUsed in 1 repo~918 tokens
    Data & AnalyticsAuto-check passed
  • Anakinscraper

    Anakin-Inc/anakin

    Scrape any website into clean markdown or structured JSON. An agent skill from Anakin-Inc/anakin.

    4.5k GitHub stars~859 tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed

More from gooseworks-ai/goose-skills

All 273 skills in this repo
  • Reddit Post Finder

    gooseworks-ai/goose-skills

    Scrape and search Reddit posts using Apify. An agent skill from gooseworks-ai/goose-skills.

    1.2k GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Create Image Fal

    gooseworks-ai/goose-skills

    Generate or edit an image via any FAL image model (nano-banana edit, gpt-image, flux, ...), ROUTED THROUGH THE fal-proxy so it bills the Ads agent.

    1.2k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Render Hook Replacement

    gooseworks-ai/goose-skills

    Replace an existing video's opening with a supplied clip or free kinetic text hook while retaining and verifying every original body frame, audio, captions and ending.

    1.2k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Blog Feed Monitor

    gooseworks-ai/goose-skills

    Scrape blog posts via RSS feeds (free, no API key) with Apify fallback for JS-heavy sites.

    1.2k GitHub starsUsed in 1 repo~578 tokens
    Auto-check passed
  • Competitor Post Engagers

    gooseworks-ai/goose-skills

    Find leads by scraping engagers from a competitor's top LinkedIn posts.

    1.2k GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check: notes
  • Render Chatgpt Chat

    gooseworks-ai/goose-skills

    Assemble a ChatGPT chat-reveal video ad from a thread + timeline JSON — one continuous Playwright recording of a ChatGPT mobile chat (user types with the iOS keyboard up → taps send → keyboard…

    1.2k GitHub stars~2.3k tokensUpdated today
    Auto-check passed

Questions about Web Archive Scraper

What does Web Archive Scraper do?

Search the Wayback Machine for archived versions of websites. Web Archive Scraper is an agent skill from gooseworks-ai/goose-skills. Search the Wayback Machine for archived versions of websites.

When should I use Web Archive Scraper?

Web Archive Scraper fits situations like: tasks that involve Web scraping.

How do I install Web Archive Scraper in Claude Code?

Run `npx skills add gooseworks-ai/goose-skills --skill web-archive-scraper -a claude-code`. Or copy the skill folder (skills/research-tools/capabilities/web-archive-scraper in gooseworks-ai/goose-skills) into .claude/skills/web-archive-scraper in your project. Claude Code loads it when a task matches its description.

How do I install Web Archive Scraper in Codex?

Run `npx skills add gooseworks-ai/goose-skills --skill web-archive-scraper -a codex`. Or copy the skill folder (skills/research-tools/capabilities/web-archive-scraper in gooseworks-ai/goose-skills) into .agents/skills/web-archive-scraper in your project. Codex loads it when a task matches its description.

Can I use Web Archive Scraper in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add gooseworks-ai/goose-skills --skill web-archive-scraper -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/web-archive-scraper, .gemini/skills/web-archive-scraper, .github/skills/web-archive-scraper and .opencode/skills/web-archive-scraper in your project.

What does Web Archive Scraper need to run?

Going by SKILL.md and its folder, Web Archive Scraper needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Web Archive Scraper access the network?

SKILL.md names 2 domains. In commands or code: botkeeper.com and web.archive.org; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Web Archive Scraper safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Web Archive Scraper use?

Web Archive Scraper is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Web Archive Scraper use?

About 901 tokens (SKILL.md is roughly 3.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Web Archive Scraper?

Skills that share tags, products or a category with Web Archive Scraper: Tmux (trpc-group/trpc-agent-go, 1.9k stars), Ketch (1broseidon/ketch, 702 stars), Boss Zhipin Scraper (eatmoreduck/boss-zhipin-scraper, 1.5k stars) and Crawl4AI Web Scraping (smallnest/goclaw, 599 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Web Archive Scraper?

gooseworks-ai (a GitHub organization) maintains it in gooseworks-ai/goose-skills, which has 1,240 GitHub stars. The repository holds 273 skills in this directory. The repository was last updated on October 8, 2026.

Source: gooseworks-ai/goose-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.