Agent skill

Wayback Cdx Cloud Ip Workaround

by divinevideo in divinevideo/divine-mobile

Fix Wayback Machine CDX API returning empty results or timing out from cloud/datacenter IPs (Cloud Run, AWS Lambda, GCE, etc.) while working fine locally.

MPL-2.0Auto-check passedBackend & APIs

Install Wayback Cdx Cloud Ip Workaround

skills CLI
$ npx skills add divinevideo/divine-mobile --skill wayback-cdx-cloud-ip-workaround -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install divinevideo/divine-mobile wayback-cdx-cloud-ip-workaround --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/divinevideo/divine-mobile.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/wayback-cdx-cloud-ip-workaround .claude/skills/wayback-cdx-cloud-ip-workaround && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
wayback-cdx-cloud-ip-workaround
GitHub stars
266
Token cost
~1.4k tokens
SKILL.md length
420 words
Files
1
Skills in repo
103
Repo updated
First seen
Licence
MPL-2.0

At a glance

Fix Wayback Machine CDX API returning empty results or timing out from cloud/datacenter IPs (Cloud Run, AWS Lambda, GCE, etc.) while working fine locally.

  • Works in 3 steps: Fetch locally: Run CDX queries from your… → Upload to cloud storage: gsutil cp the… → Import on cloud: Have your Cloud Run job…
  • CDX queries timeout
  • SKILL.md covers Problem, Context / Trigger Conditions, Solution and Verification, plus 3 more sections
  • Calls gsutil and gcloud; reaches web.archive.org

What it does

Wayback Cdx Cloud Ip Workaround is an agent skill from divinevideo/divine-mobile. Fix Wayback Machine CDX API returning empty results or timing out from cloud/datacenter IPs (Cloud Run, AWS Lambda, GCE, etc.) while working fine locally. Use when: (1) CDX queries timeout or return 0 bytes from Cloud Run/cloud functions but work from local machine, (2) urllib.request.urlopen times out for web.archive.org from server, (3) Wayback Machine CDX silently returns empty from datacenter IPs with no HTTP error, (4) Need to run CDX-dependent scripts on Cloud Run or similar cloud compute. Workaround: fetch…

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Backend & APIs, covering Serverless. It works with Cloud Run and AWS Lambda. The licence is MPL-2.0.

When your agent uses it

  • CDX queries timeout
  • Return 0 bytes from Cloud Run/cloud functions but work from local machine
  • Urllib.request.urlopen times out for web.archive.org from server
  • Wayback Machine CDX silently returns empty from datacenter IPs with no HTTP error

Example prompts

  • “/wayback-cdx-cloud-ip-workaround”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Fetch locally: Run CDX queries from your local machine, save results to a file
  2. Upload to cloud storage: gsutil cp the file to GCS (or S3, etc.)
  3. Import on cloud: Have your Cloud Run job read from cloud storage instead of CDX

What it can do on your machine

Read from SKILL.md and the folder at commit 6487b05. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gsutil
    • gcloud

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • web.archive.org

    Also links to:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Wayback Cdx Cloud Ip Workaround loads about 1.4k tokens when it runs. Until then it costs about 152 tokens; SKILL.md has 420 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~152
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from divinevideo/divine-mobile at commit 6487b05, republished under its MPL-2.0 licence (© divinevideo). 420 words, ~1,421 tokens.

Download SKILL.mdSave it as .claude/skills/wayback-cdx-cloud-ip-workaround/SKILL.md (or your agent's skills folder).
name
wayback-cdx-cloud-ip-workaround
description
Fix Wayback Machine CDX API returning empty results or timing out from cloud/datacenter IPs (Cloud Run, AWS Lambda, GCE, etc.) while working fine locally. Use when: (1) CDX queries timeout or return 0 bytes from Cloud Run/cloud functions but work from local machine, (2) urllib.request.urlopen times out for web.archive.org from server, (3) Wayback Machine CDX silently returns empty from datacenter IPs with no HTTP error, (4) Need to run CDX-dependent scripts on Cloud Run or similar cloud compute. Workaround: fetch locally, upload to cloud storage, import on cloud compute.
author
Claude Code
version
1.0.0
date
2026-02-20

Wayback CDX API: Cloud/Datacenter IP Workaround

Problem

The Wayback Machine CDX API (web.archive.org/cdx/search/cdx) silently refuses to serve data to requests from cloud provider IP ranges (Google Cloud Run, AWS, GCE, etc.). There is no HTTP error code — requests either timeout or return 0 bytes. The same queries work perfectly from residential/local IPs.

Context / Trigger Conditions

  • CDX API returns empty (0 bytes) or times out from Cloud Run, Lambda, GCE, etc.
  • Same query via curl on local machine returns expected data
  • No HTTP error codes (not 429, not 403) — just empty or timeout
  • Other external APIs (Common Crawl, Arctic Shift, Arquivo.pt) work fine from the same cloud instance
  • Even with VPC connector + static IP configured, CDX still fails
  • Increasing timeouts and retries doesn't help

Solution

The Pattern: Local Fetch → Cloud Storage → Cloud Import

Since CDX won't respond to cloud IPs, split the workflow:

  1. Fetch locally: Run CDX queries from your local machine, save results to a file
  2. Upload to cloud storage: gsutil cp the file to GCS (or S3, etc.)
  3. Import on cloud: Have your Cloud Run job read from cloud storage instead of CDX
Implementation

Step 1: Local fetch script

python
# Run this locally — CDX works from residential IPs
import urllib.request, urllib.parse

CDX = 'https://web.archive.org/cdx/search/cdx'
results = set()
resume_key = None

while True:
    params = {
        'url': 'example.com/path/*',
        'output': 'text',
        'fl': 'original',
        'filter': 'statuscode:200',
        'limit': '10000',
        'showResumeKey': 'true',
    }
    if resume_key:
        params['resumeKey'] = resume_key

    url = f"{CDX}?{urllib.parse.urlencode(params)}"
    req = urllib.request.Request(url, headers={'User-Agent': 'MyBot/0.1'})
    with urllib.request.urlopen(req, timeout=120) as resp:
        raw = resp.read().decode('utf-8', errors='replace')

    parts = raw.rstrip().rsplit('\n\n', 1)
    data_lines = parts[0].strip().split('\n')
    resume_key = parts[1].strip() if len(parts) == 2 else None

    for line in data_lines:
        results.add(process_line(line))  # Extract what you need

    if not resume_key or len(data_lines) < 10000:
        break
    time.sleep(3)

# Save to file
with open('/tmp/results.txt', 'w') as f:
    for item in sorted(results):
        f.write(f'{item}\n')

Step 2: Upload to GCS

bash
gsutil cp /tmp/results.txt gs://my-bucket/imports/results.txt

Step 3: Cloud Run job with --from-gcs flag

python
def load_from_gcs(gcs_path: str) -> set[str]:
    """Load IDs from a GCS text file."""
    import subprocess
    result = subprocess.run(
        ["gsutil", "cp", gcs_path, "/tmp/results.txt"],
        capture_output=True, text=True, timeout=120
    )
    if result.returncode != 0:
        raise RuntimeError(f"gsutil failed: {result.stderr}")

    items = set()
    with open("/tmp/results.txt") as f:
        for line in f:
            items.add(line.strip())
    return items

# In main():
parser.add_argument("--from-gcs", type=str, default=None,
                    help="Import from GCS file instead of CDX")

if args.from_gcs:
    items = load_from_gcs(args.from_gcs)
else:
    items = fetch_from_cdx(delay=args.delay)

Step 4: Deploy with --from-gcs

bash
gcloud run jobs deploy my-job \
    --args="-m,my_module,--from-gcs,gs://my-bucket/imports/results.txt" \
    ...

Verification

  • Local fetch completes successfully with expected data volume
  • File uploads to GCS without errors
  • Cloud Run job reads from GCS and imports to database
  • gsutil ls -l gs://my-bucket/imports/results.txt shows expected file size
Show full SKILL.md (176 more words)Show less

Example

Real-world case: Fetching 312,427 vine IDs from Wayback CDX oEmbed captures.

  • CDX failed from Cloud Run (4 attempts over multiple deploys, all returning 0 bytes)
  • Local fetch completed in ~4 minutes with resumeKey pagination
  • Uploaded 3.6MB text file to GCS
  • Cloud Run imported 170,108 new IDs from GCS in ~10 minutes

Notes

  • This affects web.archive.org specifically. Other CDX servers (index.commoncrawl.org) work fine from cloud IPs
  • The Internet Archive may whitelist specific IPs on request — contact them if you have a legitimate archival use case
  • Even with a whitelisted static IP via VPC connector, the CDX API may still not respond (whitelisting may apply to download, not CDX)
  • collapse=urlkey and matchType=prefix are more likely to fail than simple wildcard queries, but even wildcards fail from cloud IPs
  • This workaround adds a manual step (local fetch + upload) but is reliable and only needs to be done once per data source
  • For frequently-changing data, consider scheduling the local fetch via cron and automating the GCS upload

References

© divinevideo, MPL-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/wayback-cdx-cloud-ip-workaround of divinevideo/divine-mobile.

Open the folder on GitHubat commit 6487b05

Compare with similar skills

Wayback Cdx Cloud Ip Workaround next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Wayback Cdx Cloud Ip Workaround compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Wayback Cdx Cloud Ip Workaround this skilldivinevideo/divine-mobile266—~1.4kAutomated safety check: PassMPL-2.0
Azure Cloud Migratemicrosoft/GitHub-Copilot-for-Azure2551 repos~1.1kAutomated safety check: PassMIT
AWS Serverless Edazxkane/aws-skills3674 repos~3.2kAutomated safety check: PassMIT
Polylith Base CreationDavidVujic/python-polylith553—~757Automated safety check: PassMIT
Serverless IntegrationsDataDog/dd-trace-js837—~1.1kAutomated safety check: PassCustom licence
Processing S3 Uploads With Step Functionsaws/agent-toolkit-for-aws2.8k—~4kAutomated safety check: PassApache-2.0

Similar skills

  • Azure Cloud Migrate

    microsoft/GitHub-Copilot-for-Azure

    Official

    Assess and migrate cross-cloud workloads to Azure with reports and code conversion.

    255 GitHub starsUsed in 1 repo~1.1k tokens
    DevOps & CloudAuto-check passed
  • AWS Serverless Eda

    zxkane/aws-skills

    AWS serverless and event-driven architecture expert based on Well-Architected Framework.

    367 GitHub starsUsed in 4 repos~3.2k tokens
    Backend & APIsAuto-check passed
  • Polylith Base Creation

    DavidVujic/python-polylith

    Create a Polylith base with poly create base — the entry point of a deployable application (HTTP API, CLI, message-queue consumer, AWS Lambda handler, GCP Cloud Function, scheduled job).

    553 GitHub stars~757 tokensUpdated 5 days ago
    Backend & APIsAuto-check passed
  • Serverless Integrations

    DataDog/dd-trace-js

    Official

    A skill your agent uses when adding, modifying, debugging, or reviewing dd-trace-js serverless platform integrations that create root invocation spans for AWS Lambda, Azure Functions, Google Cloud…

    837 GitHub stars~1.1k tokensUpdated today
    Backend & APIsAuto-check passed
  • Official

    Deploy an event-driven workflow that routes S3 uploads to either Lambda or Fargate via Step Functions based on file size.

    2.8k GitHub stars~4k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • AWS Serverless

    davila7/claude-code-templates

    Specialized skill for building production-ready serverless applications on AWS.

    32k GitHub starsUsed in 7 repos~2k tokens
    Backend & APIsAuto-check passed

More from divinevideo/divine-mobile

All 103 skills in this repo
  • Fix ArgoCD ExternalSecret deployment failing with "namespace X is not permitted in project Y".

    266 GitHub stars~931 tokensUpdated today
    Auto-check passed
  • Art Direct

    divinevideo/divine-mobile

    Art direction for any content — reads text, PDF, Word, HTML, PPT, then proposes 2-3 creative directions with photography style, mood, and visual language.

    266 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Async Await Null Race Condition

    divinevideo/divine-mobile

    Fix "Null check operator used on a null value" errors when an object is set to null during an async await.

    266 GitHub stars~881 tokensUpdated today
    Auto-check passed
  • AWS V4 Signing Custom Headers Gcs

    divinevideo/divine-mobile

    Add custom metadata headers (x-amz-meta-) to AWS v4 signed requests for GCS S3-compatible API.

    266 GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Bash Herestring Newline Secrets

    divinevideo/divine-mobile

    Fix password/secret authentication failures caused by trailing newlines when creating Google Cloud secrets (or similar) with bash here-strings.

    266 GitHub stars~791 tokensUpdated today
    Auto-check passed
  • Fix silent video/media processing failures caused by URL extraction code that filters on file extensions (.mp4, .webm, .webp).

    266 GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Categories

Questions about Wayback Cdx Cloud Ip Workaround

What does Wayback Cdx Cloud Ip Workaround do?

Fix Wayback Machine CDX API returning empty results or timing out from cloud/datacenter IPs (Cloud Run, AWS Lambda, GCE, etc.) while working fine locally. Wayback Cdx Cloud Ip Workaround is an agent skill from divinevideo/divine-mobile.) while working fine locally.

When should I use Wayback Cdx Cloud Ip Workaround?

Wayback Cdx Cloud Ip Workaround fits situations like: CDX queries timeout; return 0 bytes from Cloud Run/cloud functions but work from local machine; urllib.request.urlopen times out for web.archive.org from server; wayback Machine CDX silently returns empty from datacenter IPs with no HTTP error.

How do I install Wayback Cdx Cloud Ip Workaround in Claude Code?

Run `npx skills add divinevideo/divine-mobile --skill wayback-cdx-cloud-ip-workaround -a claude-code`. Or copy the skill folder (.agents/skills/wayback-cdx-cloud-ip-workaround in divinevideo/divine-mobile) into .claude/skills/wayback-cdx-cloud-ip-workaround in your project. Claude Code loads it when a task matches its description.

How do I install Wayback Cdx Cloud Ip Workaround in Codex?

Run `npx skills add divinevideo/divine-mobile --skill wayback-cdx-cloud-ip-workaround -a codex`. Or copy the skill folder (.agents/skills/wayback-cdx-cloud-ip-workaround in divinevideo/divine-mobile) into .agents/skills/wayback-cdx-cloud-ip-workaround in your project. Codex loads it when a task matches its description.

Can I use Wayback Cdx Cloud Ip Workaround in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add divinevideo/divine-mobile --skill wayback-cdx-cloud-ip-workaround -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/wayback-cdx-cloud-ip-workaround, .gemini/skills/wayback-cdx-cloud-ip-workaround, .github/skills/wayback-cdx-cloud-ip-workaround and .opencode/skills/wayback-cdx-cloud-ip-workaround in your project.

What does Wayback Cdx Cloud Ip Workaround need to run?

Going by SKILL.md and its folder, Wayback Cdx Cloud Ip Workaround needs the command-line tools its instructions call (gsutil and gcloud). Our summary lists: Python 3.

Does Wayback Cdx Cloud Ip Workaround access the network?

SKILL.md names 2 domains. In commands or code: web.archive.org; the agent is likely to contact it when it follows the instructions. As links in the text: github.com. This is read from the text; nothing was executed.

Is Wayback Cdx Cloud Ip Workaround safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Wayback Cdx Cloud Ip Workaround use?

Wayback Cdx Cloud Ip Workaround is published under the MPL-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Wayback Cdx Cloud Ip Workaround use?

About 1.4k tokens (SKILL.md is roughly 5.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Wayback Cdx Cloud Ip Workaround?

Skills that share tags, products or a category with Wayback Cdx Cloud Ip Workaround: Azure Cloud Migrate (microsoft/GitHub-Copilot-for-Azure, 255 stars), AWS Serverless Eda (zxkane/aws-skills, 367 stars), Polylith Base Creation (DavidVujic/python-polylith, 553 stars) and Serverless Integrations (DataDog/dd-trace-js, 837 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Wayback Cdx Cloud Ip Workaround?

divinevideo (a GitHub organization) maintains it in divinevideo/divine-mobile, which has 266 GitHub stars. The repository holds 103 skills in this directory. The repository was last updated on October 9, 2026.

Source: divinevideo/divine-mobile on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.