Agent skill

Wayback Cdx Wildcard Pagination

by divinevideo in divinevideo/divine-mobile

Fix empty results when paginating Wayback Machine CDX API with wildcard URL queries.

MPL-2.0Auto-check passed

Install Wayback Cdx Wildcard Pagination

skills CLI
$ npx skills add divinevideo/divine-mobile --skill wayback-cdx-wildcard-pagination -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install divinevideo/divine-mobile wayback-cdx-wildcard-pagination --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/divinevideo/divine-mobile.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/wayback-cdx-wildcard-pagination .claude/skills/wayback-cdx-wildcard-pagination && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
wayback-cdx-wildcard-pagination
GitHub stars
266
Token cost
~1.1k tokens
SKILL.md length
268 words
Files
1
Skills in repo
103
Repo updated
First seen
Licence
MPL-2.0

At a glance

Fix empty results when paginating Wayback Machine CDX API with wildcard URL queries.

  • CDX query with url=foo/ and page=0 returns empty/0 bytes but works without page parameter
  • SKILL.md covers Problem, Context / Trigger Conditions, Solution and Verification, plus 3 more sections
  • Calls curl; reaches web.archive.org
  • CDX pagination returns no data for wildcard prefix searches

What it does

Wayback Cdx Wildcard Pagination is an agent skill from divinevideo/divine-mobile. Fix empty results when paginating Wayback Machine CDX API with wildcard URL queries. Use when: (1) CDX query with url=foo/ and page=0 returns empty/0 bytes but works without page parameter, (2) CDX pagination returns no data for wildcard prefix searches, (3) Need to paginate large CDX result sets using wildcard URL matching. Covers showResumeKey, offset, and page parameter incompatibilities.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The licence is MPL-2.0.

When your agent uses it

  • CDX query with url=foo/ and page=0 returns empty/0 bytes but works without page parameter
  • CDX pagination returns no data for wildcard prefix searches
  • Need to paginate large CDX result sets using wildcard URL matching

Example prompts

  • “/wayback-cdx-wildcard-pagination”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit c3d6f7e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • web.archive.org

    Also links to:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Wayback Cdx Wildcard Pagination loads about 1.1k tokens when it runs. Until then it costs about 107 tokens; SKILL.md has 268 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~107
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from divinevideo/divine-mobile at commit c3d6f7e, republished under its MPL-2.0 licence (© divinevideo). 268 words, ~1,072 tokens.

Download SKILL.mdSave it as .claude/skills/wayback-cdx-wildcard-pagination/SKILL.md (or your agent's skills folder).
name
wayback-cdx-wildcard-pagination
description
Fix empty results when paginating Wayback Machine CDX API with wildcard URL queries. Use when: (1) CDX query with url=foo/* and page=0 returns empty/0 bytes but works without page parameter, (2) CDX pagination returns no data for wildcard prefix searches, (3) Need to paginate large CDX result sets using wildcard URL matching. Covers showResumeKey, offset, and page parameter incompatibilities.
author
Claude Code
version
1.0.0
date
2026-02-20

Wayback CDX API: Wildcard Query Pagination

Problem

The Wayback Machine CDX API's page=N pagination parameter returns empty results when combined with wildcard URL queries (url=domain.com/path/*), even though the same query without page= returns data. This causes scripts to incorrectly report "no results found" when there are actually hundreds of thousands of results.

Context / Trigger Conditions

  • CDX query with url=*.example.com/* or url=example.com/path/* and page=0 returns 0 bytes
  • Same query without page= parameter returns expected data
  • Using web.archive.org/cdx/search/cdx API endpoint
  • matchType=prefix + collapse=urlkey + page=N combination also fails silently
  • Script works with limit=5 but fails when adding page=0

Solution

The most reliable pagination method. CDX appends a base64 resume token after a blank line separator in the response.

python
resume_key = None
while True:
    params = {
        'url': 'example.com/path/*',
        'output': 'text',
        'fl': 'original',
        'filter': 'statuscode:200',
        'limit': '10000',
        'showResumeKey': 'true',
    }
    if resume_key:
        params['resumeKey'] = resume_key

    data = fetch_cdx(params)
    if not data:
        break

    # Resume key is after the last blank line
    parts = data.rstrip().rsplit('\n\n', 1)
    if len(parts) == 2:
        data_lines = parts[0].strip().split('\n')
        resume_key = parts[1].strip()
    else:
        data_lines = parts[0].strip().split('\n')
        resume_key = None

    # Process data_lines...

    if not resume_key or len(data_lines) < limit:
        break
Option 2: Use offset=N

Works but less efficient for very large result sets.

python
offset = 0
limit = 10000
while True:
    params = {
        'url': 'example.com/path/*',
        'output': 'text',
        'fl': 'original',
        'limit': str(limit),
        'offset': str(offset),
    }
    data = fetch_cdx(params)
    lines = data.strip().split('\n')
    if not lines or not lines[0]:
        break
    offset += len(lines)
What NOT to do
python
# THIS RETURNS EMPTY for wildcard queries:
params = {
    'url': 'example.com/path/*',
    'page': '0',        # <-- INCOMPATIBLE with wildcard
    'limit': '10000',
}

# THIS ALSO FAILS from some IPs:
params = {
    'url': 'example.com/path/',
    'matchType': 'prefix',
    'collapse': 'urlkey',
    'page': '0',
}

Verification

  • Query with showResumeKey=true and no page= param returns data
  • Response ends with a blank line followed by a base64 token (the resume key)
  • Subsequent request with resumeKey=<token> returns the next page

Example

Fetching all vine.co/oembed/* URLs (690K+ captures, 312K unique vine IDs):

bash
# This works:
curl "https://web.archive.org/cdx/search/cdx?url=vine.co/oembed/*&output=text&fl=original&limit=10000&showResumeKey=true"

# This returns empty:
curl "https://web.archive.org/cdx/search/cdx?url=vine.co/oembed/*&output=text&fl=original&limit=10000&page=0"

Notes

  • page=N works fine for non-wildcard queries (e.g., exact URL lookups)
  • The showResumeKey approach is server-side cursor-based, more efficient than offset
  • Keep limit at 10000 or less to avoid timeouts, especially from cloud IPs
  • Always add time.sleep(3-5) between pages to be polite to the CDX server
  • The CDX API has no official documentation for this incompatibility

References

© divinevideo, MPL-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/wayback-cdx-wildcard-pagination of divinevideo/divine-mobile.

Open the folder on GitHubat commit c3d6f7e

Compare with similar skills

Wayback Cdx Wildcard Pagination next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Wayback Cdx Wildcard Pagination compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Wayback Cdx Wildcard Pagination this skilldivinevideo/divine-mobile266—~1.1kAutomated safety check: PassMPL-2.0
Paginationthedaviddias/Front-End-Checklist74k—~699Automated safety check: PassMIT
Pagination Accessibilitythedaviddias/Front-End-Checklist74k—~405Automated safety check: PassMIT
Empty Linksthedaviddias/Front-End-Checklist74k—~508Automated safety check: PassMIT
Empty Headingthedaviddias/Front-End-Checklist74k—~426Automated safety check: PassMIT
Building Product Empty StatesPostHog/posthog40k—~4.7kAutomated safety check: PassCustom licence

Similar skills

  • Pagination

    thedaviddias/Front-End-Checklist

    A skill your agent uses when auditing paginated content (blog archives, product category pages, search results).

    74k GitHub stars~699 tokensUpdated 4 days ago
    Auto-check passed
  • Pagination Accessibility

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing templates, rendered HTML, or shared components related to Make pagination accessible.

    74k GitHub stars~405 tokensUpdated 4 days ago
    Frontend & DesignAuto-check passed
  • Empty Links

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing rendered HTML, interactive components, or design-system patterns related to Fix empty and broken links.

    74k GitHub stars~508 tokensUpdated 4 days ago
    Frontend & DesignAuto-check passed
  • Empty Heading

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing rendered HTML, interactive components, or design-system patterns related to Ensure headings contain text.

    74k GitHub stars~426 tokensUpdated 4 days ago
    Frontend & DesignAuto-check passed
  • Official

    Guide for adding a product setup empty state — the skippable first-run screen a product scene shows until real data arrives, built on the shared ProductEmptyState component.

    40k GitHub stars~4.7k tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • Audit State Machine

    ben-manes/caffeine

    Audit explicit state machines (drain status, node lifecycle, async-value lifecycle) for illegal or missed transitions

    18k GitHub stars~1.2k tokensUpdated today
    Auto-check passed

More from divinevideo/divine-mobile

All 103 skills in this repo
  • Fix ArgoCD ExternalSecret deployment failing with "namespace X is not permitted in project Y".

    266 GitHub stars~931 tokensUpdated today
    Auto-check passed
  • Art Direct

    divinevideo/divine-mobile

    Art direction for any content — reads text, PDF, Word, HTML, PPT, then proposes 2-3 creative directions with photography style, mood, and visual language.

    266 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Async Await Null Race Condition

    divinevideo/divine-mobile

    Fix "Null check operator used on a null value" errors when an object is set to null during an async await.

    266 GitHub stars~881 tokensUpdated today
    Auto-check passed
  • AWS V4 Signing Custom Headers Gcs

    divinevideo/divine-mobile

    Add custom metadata headers (x-amz-meta-) to AWS v4 signed requests for GCS S3-compatible API.

    266 GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Bash Herestring Newline Secrets

    divinevideo/divine-mobile

    Fix password/secret authentication failures caused by trailing newlines when creating Google Cloud secrets (or similar) with bash here-strings.

    266 GitHub stars~791 tokensUpdated today
    Auto-check passed
  • Fix silent video/media processing failures caused by URL extraction code that filters on file extensions (.mp4, .webm, .webp).

    266 GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Questions about Wayback Cdx Wildcard Pagination

What does Wayback Cdx Wildcard Pagination do?

Fix empty results when paginating Wayback Machine CDX API with wildcard URL queries. Wayback Cdx Wildcard Pagination is an agent skill from divinevideo/divine-mobile. Fix empty results when paginating Wayback Machine CDX API with wildcard URL queries.

When should I use Wayback Cdx Wildcard Pagination?

Wayback Cdx Wildcard Pagination fits situations like: CDX query with url=foo/ and page=0 returns empty/0 bytes but works without page parameter; CDX pagination returns no data for wildcard prefix searches; need to paginate large CDX result sets using wildcard URL matching.

How do I install Wayback Cdx Wildcard Pagination in Claude Code?

Run `npx skills add divinevideo/divine-mobile --skill wayback-cdx-wildcard-pagination -a claude-code`. Or copy the skill folder (.agents/skills/wayback-cdx-wildcard-pagination in divinevideo/divine-mobile) into .claude/skills/wayback-cdx-wildcard-pagination in your project. Claude Code loads it when a task matches its description.

How do I install Wayback Cdx Wildcard Pagination in Codex?

Run `npx skills add divinevideo/divine-mobile --skill wayback-cdx-wildcard-pagination -a codex`. Or copy the skill folder (.agents/skills/wayback-cdx-wildcard-pagination in divinevideo/divine-mobile) into .agents/skills/wayback-cdx-wildcard-pagination in your project. Codex loads it when a task matches its description.

Can I use Wayback Cdx Wildcard Pagination in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add divinevideo/divine-mobile --skill wayback-cdx-wildcard-pagination -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/wayback-cdx-wildcard-pagination, .gemini/skills/wayback-cdx-wildcard-pagination, .github/skills/wayback-cdx-wildcard-pagination and .opencode/skills/wayback-cdx-wildcard-pagination in your project.

What does Wayback Cdx Wildcard Pagination need to run?

Going by SKILL.md and its folder, Wayback Cdx Wildcard Pagination needs the command-line tools its instructions call (curl). Our summary lists: Python 3.

Does Wayback Cdx Wildcard Pagination access the network?

SKILL.md names 2 domains. In commands or code: web.archive.org; the agent is likely to contact it when it follows the instructions. As links in the text: github.com. This is read from the text; nothing was executed.

Is Wayback Cdx Wildcard Pagination safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Wayback Cdx Wildcard Pagination use?

Wayback Cdx Wildcard Pagination is published under the MPL-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Wayback Cdx Wildcard Pagination use?

About 1.1k tokens (SKILL.md is roughly 4.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Wayback Cdx Wildcard Pagination?

Skills that share tags, products or a category with Wayback Cdx Wildcard Pagination: Pagination (thedaviddias/Front-End-Checklist, 74k stars), Pagination Accessibility (thedaviddias/Front-End-Checklist, 74k stars), Empty Links (thedaviddias/Front-End-Checklist, 74k stars) and Empty Heading (thedaviddias/Front-End-Checklist, 74k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Wayback Cdx Wildcard Pagination?

divinevideo (a GitHub organization) maintains it in divinevideo/divine-mobile, which has 266 GitHub stars. The repository holds 103 skills in this directory. The repository was last updated on October 10, 2026.

Source: divinevideo/divine-mobile on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.