Agent skill

Wayback Indirect Asset Recovery

by divinevideo in divinevideo/divine-mobile

Recover assets (images, videos, data) that weren't directly archived by Wayback Machine by crawling archived pages that reference them.

MPL-2.0Auto-check passed

Install Wayback Indirect Asset Recovery

skills CLI
$ npx skills add divinevideo/divine-mobile --skill wayback-indirect-asset-recovery -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install divinevideo/divine-mobile wayback-indirect-asset-recovery --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/divinevideo/divine-mobile.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/wayback-indirect-asset-recovery .claude/skills/wayback-indirect-asset-recovery && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
wayback-indirect-asset-recovery
GitHub stars
266
Token cost
~1.5k tokens
SKILL.md length
287 words
Files
1
Skills in repo
103
Repo updated
First seen
Licence
MPL-2.0

At a glance

Recover assets (images, videos, data) that weren't directly archived by Wayback Machine by crawling archived pages that reference them.

  • Works in 4 steps: Fetch the Archived Page → Extract Asset URLs from HTML → Handle Wayback-Wrapped URLs → …
  • Direct asset URL returns 404/503 from Wayback
  • SKILL.md covers Problem, Context / Trigger Conditions, Solution and Verification, plus 3 more sections
  • Reaches web.archive.org

What it does

Wayback Indirect Asset Recovery is an agent skill from divinevideo/divine-mobile. Recover assets (images, videos, data) that weren't directly archived by Wayback Machine by crawling archived pages that reference them. Use when: (1) Direct asset URL returns 404/503 from Wayback, (2) Asset was hosted on CDN that Wayback didn't crawl, (3) You have the page URL but not the asset URL, (4) Recovering avatars, thumbnails, or embedded media from defunct services. The key insight: pages often got archived even when their assets didn't - extract URLs from archived HTML, then try alternative fetch methods.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The licence is MPL-2.0.

When your agent uses it

  • Direct asset URL returns 404/503 from Wayback
  • Asset was hosted on CDN that Wayback didnt crawl
  • You have the page URL but not the asset URL
  • Recovering avatars

Example prompts

  • “/wayback-indirect-asset-recovery”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Fetch the Archived Page
  2. Extract Asset URLs from HTML
  3. Handle Wayback-Wrapped URLs
  4. Fallback Methods if Wayback Fails

What it can do on your machine

Read from SKILL.md and the folder at commit c3d6f7e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • web.archive.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Wayback Indirect Asset Recovery loads about 1.5k tokens when it runs. Until then it costs about 138 tokens; SKILL.md has 287 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~138
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from divinevideo/divine-mobile at commit c3d6f7e, republished under its MPL-2.0 licence (© divinevideo). 287 words, ~1,540 tokens.

Download SKILL.mdSave it as .claude/skills/wayback-indirect-asset-recovery/SKILL.md (or your agent's skills folder).
name
wayback-indirect-asset-recovery
description
Recover assets (images, videos, data) that weren't directly archived by Wayback Machine by crawling archived pages that reference them. Use when: (1) Direct asset URL returns 404/503 from Wayback, (2) Asset was hosted on CDN that Wayback didn't crawl, (3) You have the page URL but not the asset URL, (4) Recovering avatars, thumbnails, or embedded media from defunct services. The key insight: pages often got archived even when their assets didn't - extract URLs from archived HTML, then try alternative fetch methods.
author
Claude Code
version
1.0.0
date
2026-01-28

Wayback Indirect Asset Recovery

Problem

When archiving content from defunct services, direct asset URLs (images, videos, avatars) often return 404/503 from Wayback Machine even when the pages referencing them were archived. The assets themselves may not have been crawled, but their URLs exist in archived HTML.

Context / Trigger Conditions

  • Direct Wayback URL for asset returns 404 or 503
  • You have a profile/page URL but not the asset URL
  • Recovering media from defunct services (Vine, Tumblr, defunct startups)
  • Wayback has the page but not the embedded assets
  • CDN-hosted assets that weren't in Wayback's crawl scope

Solution

Step 1: Fetch the Archived Page
python
import requests
import re

session = requests.Session()
session.headers['User-Agent'] = 'YourCrawler/0.1 (archival research)'

# Fetch archived profile/page
page_url = f"https://web.archive.org/web/20170110/https://example.com/user/{username}"
resp = session.get(page_url, timeout=20, allow_redirects=True)
Step 2: Extract Asset URLs from HTML
python
# Try multiple patterns - pages structure varies
patterns = [
    # Open Graph image
    (r'og:image["\s]+content="([^"]+)"', 'og:image'),
    # JSON data in page
    (r'"avatarUrl"\s*:\s*"([^"]+)"', 'JSON avatarUrl'),
    (r'"imageUrl"\s*:\s*"([^"]+)"', 'JSON imageUrl'),
    # CDN URLs
    (r'(https?://[^"\s]+cdn\.[^"\s]+\.(jpg|png|mp4))', 'CDN URL'),
    # S3 URLs
    (r'(https?://[^"\s]+\.s3\.amazonaws\.com/[^"\s]+)', 'S3 URL'),
]

for pattern, name in patterns:
    match = re.search(pattern, html)
    if match:
        asset_url = match.group(1).replace('&', '&')
        break
Step 3: Handle Wayback-Wrapped URLs

URLs extracted from archived pages are often already Wayback URLs:

python
# If URL is already wrapped, convert im_ to id_ for raw content
if 'web.archive.org/web/' in asset_url:
    # Replace /web/TIMESTAMP/ or /web/TIMESTAMPim_/ with /web/TIMESTAMPid_/
    download_url = re.sub(r'/web/(\d+)(im_)?/', r'/web/\1id_/', asset_url)
else:
    # Wrap raw URL in Wayback
    download_url = f"https://web.archive.org/web/{timestamp}id_/{asset_url}"
Step 4: Fallback Methods if Wayback Fails

If Wayback returns 404/503 for the asset, try alternatives:

python
# Extract original CDN URL
cdn_match = re.search(r'(https?://[^/]+cdn\.[^"\s]+)', asset_url)
if cdn_match:
    original_url = cdn_match.group(1)

    # Method 1: Try live CDN via IP (if DNS is dead but servers live)
    # See: dead-cdn-dns-bypass skill

    # Method 2: Try different Wayback timestamps
    for ts in ['20170110', '20160601', '20150601']:
        alt_url = f"https://web.archive.org/web/{ts}id_/{original_url}"
        resp = session.get(alt_url, timeout=30)
        if resp.status_code == 200:
            break

    # Method 3: Check if asset exists on successor platform
    # (e.g., user migrated to Twitter/TikTok with same username)

Verification

  1. Check HTTP 200 response from final download URL
  2. Validate content-type matches expected type
  3. Verify file magic bytes (JPEG: \xff\xd8\xff, PNG: \x89PNG)
  4. Confirm file size > minimum threshold

Example: Recovering Vine Avatars

python
def recover_avatar_from_profile(vanity_url: str, user_id: str, session) -> bool:
    """Recover avatar by crawling archived Vine profile."""
    timestamps = ['20170110', '20161201', '20160601', '20150601']

    for ts in timestamps:
        # Fetch archived profile page
        profile_url = f"https://web.archive.org/web/{ts}/https://vine.co/{vanity_url}"
        resp = session.get(profile_url, timeout=20)

        if resp.status_code != 200:
            continue

        # Extract avatar URL from page
        match = re.search(r'og:image["\s]+content="([^"]+)"', resp.text)
        if not match:
            match = re.search(r'"avatarUrl"\s*:\s*"([^"]+)"', resp.text)

        if match:
            avatar_url = match.group(1).replace('&', '&')

            # Convert to raw content URL
            if 'web.archive.org/web/' in avatar_url:
                download_url = re.sub(r'/web/(\d+)(im_)?/', r'/web/\1id_/', avatar_url)
            else:
                download_url = f"https://web.archive.org/web/{ts}id_/{avatar_url}"

            # Download
            img_resp = session.get(download_url, timeout=30)
            if img_resp.status_code == 200 and len(img_resp.content) > 100:
                # Validate and save
                if img_resp.content[:3] == b'\xff\xd8\xff':  # JPEG
                    save_avatar(user_id, img_resp.content, 'jpg')
                    return True

    return False

Result: Recovered avatars for 95% of top Vine creators (including Logan Paul, Nash Grier, Lele Pons) whose direct avatar URLs weren't archived.

Notes

  • Wayback im_ vs id_ modifiers: im_ returns image with Wayback toolbar, id_ returns raw bytes
  • Rate limiting: Respect Wayback's servers - add 1-2 second delays between requests
  • Multiple timestamps: Try several archive dates - some may have the asset, others may not
  • HTML entity decoding: Always decode & to & in extracted URLs
  • Redirect handling: Use allow_redirects=True - Wayback often redirects to nearest snapshot
  • dead-cdn-dns-bypass - For when CDN DNS is dead but servers still respond
  • wayback-api-archive-recovery - For discovering what was archived via CDX API
  • wayback-machine-raw-content-id-modifier - For fetching raw content without Wayback wrapper

© divinevideo, MPL-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/wayback-indirect-asset-recovery of divinevideo/divine-mobile.

Open the folder on GitHubat commit c3d6f7e

Compare with similar skills

Wayback Indirect Asset Recovery next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Wayback Indirect Asset Recovery compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Wayback Indirect Asset Recovery this skilldivinevideo/divine-mobile266—~1.5kAutomated safety check: PassMPL-2.0
Homepage Video Assetsremotion-dev/remotion63k—~763Automated safety check: PassCustom licence
Tasteforge Videoaffaan-m/ECC277k1 repos~3.7kAutomated safety check: PassMIT
AI Presenter VideoNousResearch/hermes-agent253k—~2.3kAutomated safety check: PassMIT
Videothedaviddias/Front-End-Checklist74k—~562Automated safety check: PassMIT
MoneyPrinterTurbo Video Generatorharry0703/MoneyPrinterTurbo130k—~2.1kAutomated safety check: WarnMIT

Similar skills

  • Homepage Video Assets

    remotion-dev/remotion

    Official

    A skill your agent uses when rendering or replacing Remotion homepage creator strip videos, especially transparent Chrome WebM and Safari MP4 assets copied into both promo-pages and docs.

    63k GitHub stars~763 tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Tasteforge Video

    affaan-m/ECC

    A skill your agent uses for file-driven multimodal image, video, and 3D-asset discovery; taste interviews; distill or apply workflows; style-pack validation; editable EDL/FCPXML export; provenance…

    277k GitHub starsUsed in 1 repo~3.7k tokens
    Game DevelopmentAuto-check passed
  • AI Presenter Video

    NousResearch/hermes-agent

    Produces a presenter-led video from a topic or script plus one authorized presenter image, with captions, lip-sync checks and acceptance reports.

    253k GitHub stars~2.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Video

    thedaviddias/Front-End-Checklist

    A skill your agent uses when applies to any page embedding or hosting video content (YouTube, Vimeo, self-hosted).

    74k GitHub stars~562 tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • MoneyPrinterTurbo Video Generator

    harry0703/MoneyPrinterTurbo

    Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.

    130k GitHub stars~2.1k tokensUpdated yesterday
    Media & CreativeAuto-check: warnings
  • Assets

    BuilderIO/agent-native

    Use Agent-Native Assets for image and video generation requests, brand-safe asset search/export, and human-in-the-loop asset selection through the hosted Assets MCP app.

    7.1k GitHub stars~1.2k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed

More from divinevideo/divine-mobile

All 103 skills in this repo
  • Fix ArgoCD ExternalSecret deployment failing with "namespace X is not permitted in project Y".

    266 GitHub stars~931 tokensUpdated today
    Auto-check passed
  • Art Direct

    divinevideo/divine-mobile

    Art direction for any content — reads text, PDF, Word, HTML, PPT, then proposes 2-3 creative directions with photography style, mood, and visual language.

    266 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Async Await Null Race Condition

    divinevideo/divine-mobile

    Fix "Null check operator used on a null value" errors when an object is set to null during an async await.

    266 GitHub stars~881 tokensUpdated today
    Auto-check passed
  • AWS V4 Signing Custom Headers Gcs

    divinevideo/divine-mobile

    Add custom metadata headers (x-amz-meta-) to AWS v4 signed requests for GCS S3-compatible API.

    266 GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Bash Herestring Newline Secrets

    divinevideo/divine-mobile

    Fix password/secret authentication failures caused by trailing newlines when creating Google Cloud secrets (or similar) with bash here-strings.

    266 GitHub stars~791 tokensUpdated today
    Auto-check passed
  • Fix silent video/media processing failures caused by URL extraction code that filters on file extensions (.mp4, .webm, .webp).

    266 GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Questions about Wayback Indirect Asset Recovery

What does Wayback Indirect Asset Recovery do?

Recover assets (images, videos, data) that weren't directly archived by Wayback Machine by crawling archived pages that reference them. Wayback Indirect Asset Recovery is an agent skill from divinevideo/divine-mobile. Recover assets (images, videos, data) that weren't directly archived by Wayback Machine by crawling archived pages that reference them.

When should I use Wayback Indirect Asset Recovery?

Wayback Indirect Asset Recovery fits situations like: direct asset URL returns 404/503 from Wayback; asset was hosted on CDN that Wayback didnt crawl; you have the page URL but not the asset URL; recovering avatars.

How do I install Wayback Indirect Asset Recovery in Claude Code?

Run `npx skills add divinevideo/divine-mobile --skill wayback-indirect-asset-recovery -a claude-code`. Or copy the skill folder (.agents/skills/wayback-indirect-asset-recovery in divinevideo/divine-mobile) into .claude/skills/wayback-indirect-asset-recovery in your project. Claude Code loads it when a task matches its description.

How do I install Wayback Indirect Asset Recovery in Codex?

Run `npx skills add divinevideo/divine-mobile --skill wayback-indirect-asset-recovery -a codex`. Or copy the skill folder (.agents/skills/wayback-indirect-asset-recovery in divinevideo/divine-mobile) into .agents/skills/wayback-indirect-asset-recovery in your project. Codex loads it when a task matches its description.

Can I use Wayback Indirect Asset Recovery in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add divinevideo/divine-mobile --skill wayback-indirect-asset-recovery -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/wayback-indirect-asset-recovery, .gemini/skills/wayback-indirect-asset-recovery, .github/skills/wayback-indirect-asset-recovery and .opencode/skills/wayback-indirect-asset-recovery in your project.

What does Wayback Indirect Asset Recovery need to run?

SKILL.md names no scripts, command-line tools or credentials: Wayback Indirect Asset Recovery is instructions for the agent only. Our summary lists: Python 3.

Does Wayback Indirect Asset Recovery access the network?

SKILL.md names 1 domain. In commands or code: web.archive.org; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Wayback Indirect Asset Recovery safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Wayback Indirect Asset Recovery use?

Wayback Indirect Asset Recovery is published under the MPL-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Wayback Indirect Asset Recovery use?

About 1.5k tokens (SKILL.md is roughly 6.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Wayback Indirect Asset Recovery?

Skills that share tags, products or a category with Wayback Indirect Asset Recovery: Homepage Video Assets (remotion-dev/remotion, 63k stars), Tasteforge Video (affaan-m/ECC, 277k stars), AI Presenter Video (NousResearch/hermes-agent, 253k stars) and Video (thedaviddias/Front-End-Checklist, 74k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Wayback Indirect Asset Recovery?

divinevideo (a GitHub organization) maintains it in divinevideo/divine-mobile, which has 266 GitHub stars. The repository holds 103 skills in this directory. The repository was last updated on October 10, 2026.

Source: divinevideo/divine-mobile on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.