Agent skill

Wp Static Clone

by jdevalk in jdevalk/skills

Clones a live WordPress (or other CMS-driven) site into a static HTML site deployable on any static host (Cloudflare Pages, Netlify, Vercel, S3+CloudFront, plain Apache/nginx).

MITAuto-check passedDevOps & Cloud

Install Wp Static Clone

skills CLI
$ npx skills add jdevalk/skills --skill wp-static-clone -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jdevalk/skills wp-static-clone --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jdevalk/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/wp-static-clone .claude/skills/wp-static-clone && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
wp-static-clone
GitHub stars
104
Token cost
~2.6k tokens
SKILL.md length
1,232 words
Files
7 (incl. scripts)
Skills in repo
9
Repo updated
First seen
Licence
MIT

At a glance

Clones a live WordPress (or other CMS-driven) site into a static HTML site deployable on any static host (Cloudflare Pages, Netlify, Vercel, S3+CloudFront, plain Apache/nginx).

  • Works in 10 steps: Confirm intent → Discover URLs and pull XML sitemaps → Scrape in one shot → …
  • The user wants to scrape
  • SKILL.md covers When to use, Workflow, Gotchas and Output structure
  • Runs Python scripts from its folder; calls python3, wget and git

What it does

Wp Static Clone is an agent skill from jdevalk/skills. Clones a live WordPress (or other CMS-driven) site into a static HTML site deployable on any static host (Cloudflare Pages, Netlify, Vercel, S3+CloudFront, plain Apache/nginx). Use when the user wants to "scrape", "freeze", "archive", "static-ify", or "move to [host]" a WordPress site, or asks to turn a sitemap into deployable static HTML. Pulls every URL from sitemapindex.xml, fetches all assets, rewrites paths to be root-relative, strips WP runtime markup, and outputs a flat directory ready to deploy with no…

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts (for example `AGENTS.md`, `README.md` and `scripts/insert-banner.py`).

It sits in DevOps & Cloud, covering Web scraping, Technical SEO and File uploads and storage. It works with WordPress, Netlify, Cloudflare Pages and Vercel. The repository describes itself as: Agent skills for GitHub repos and profiles, WordPress and EmDash plugins, Astro SEO, and content readability. The licence is MIT.

When your agent uses it

  • The user wants to scrape
  • Move to [host] a WordPress site
  • Asks to turn a sitemap into deployable static HTML

Example prompts

  • “scrape”
  • “freeze”
  • “archive”
  • “/wp-static-clone”

Requirements

  • Python 3

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. Confirm intent
  2. Discover URLs and pull XML sitemaps
  3. Scrape in one shot
  4. Pull assets the page-requisites pass missed
  5. Convert paths to root-relative
  6. Brand the static output
  7. Replace WP runtime hooks
  8. Copy robots.txt
  9. Verify locally
  10. Deploy

What it can do on your machine

Read from SKILL.md and the folder at commit 106fc68. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • wget
    • git
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use wget and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Wp Static Clone loads about 2.6k tokens when it runs. Until then it costs about 137 tokens; SKILL.md has 1,232 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~137
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from jdevalk/skills at commit 106fc68, republished under its MIT licence (© jdevalk). 1,232 words, ~2,576 tokens.

Download SKILL.mdSave it as .claude/skills/wp-static-clone/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
wp-static-clone
description
Clones a live WordPress (or other CMS-driven) site into a static HTML site deployable on any static host (Cloudflare Pages, Netlify, Vercel, S3+CloudFront, plain Apache/nginx). Use when the user wants to "scrape", "freeze", "archive", "static-ify", or "move to [host]" a WordPress site, or asks to turn a sitemap into deployable static HTML. Pulls every URL from sitemap_index.xml, fetches all assets, rewrites paths to be root-relative, strips WP runtime markup, and outputs a flat directory ready to deploy with no build command.

wp-static-clone

Turn a live WordPress site into a static HTML clone deployable on any static host. Driven by the site's XML sitemap. Handles the WordPress-specific gotchas — Cloudflare bot protection, mid-scrape link rewriting, proxied analytics, R2-offloaded uploads, comment-form runtime, Yoast attribution, Gravatar privacy — that a naïve wget run misses.

Recipes live in AGENTS.md; reusable scripts in scripts/. This file is the workflow, gotchas, and output structure.

When to use

Trigger on requests like:

  • "Scrape this WordPress site for [host]"
  • "Freeze [domain] as static HTML"
  • "Pull all the pages from this sitemap and turn them into static files"
  • "Move this WP site to [host] with no build step"

The broad shape (sitemap → wget → root-relative paths → static host) generalises to any CMS that emits a standard XML sitemap. The runtime cleanup (comment forms, Plausible proxy, Gravatar, Yoast) is WordPress-specific.

Workflow

Phase 0 — Confirm intent

Before scraping, confirm:

  • Source URL (the live site).
  • Target host — Cloudflare Pages, Netlify, Vercel, generic static. Drives Phase 9.
  • Same or different domain at the destination. Drives whether og:url, <link rel="canonical">, and JSON-LD @id stay absolute (same domain — correct SEO behaviour) or get rewritten (different domain).
  • What to do with analytics and forms. WP plugins for both can't run statically. Plausible gets replaced with the standard tracker (recipe in AGENTS.md); contact/search forms either get removed or wired through Pages Functions / Formspree / Netlify Forms — host-specific.
Phase 1 — Discover URLs and pull XML sitemaps

Fetch the sitemap index. Try <root>/sitemap_index.xml (Yoast convention) first, then <root>/sitemap.xml. If the index references sub-sitemaps (page-sitemap.xml, post-sitemap.xml, …), fetch each and concatenate <loc> values into urls.txt. Skip image-sitemap entries.

Also fetch the XML sitemaps themselves and the Yoast XSL stylesheet now (recipe in AGENTS.md) — they aren't linked from HTML, so wget -p won't find them later.

Phase 2 — Scrape in one shot

Critical: scrape every URL in a single wget invocation so its --convert-links pass sees all downloaded files and rewrites cross-page links correctly. Scraping URLs in separate runs leaves residual absolute links on whichever page was scraped first/last. Recipe in AGENTS.md.

Phase 3 — Pull assets the page-requisites pass missed

Some assets aren't -p-followed because they appear only in og:image, apple-touch-icon, JSON-LD image/logo, msapplication-TileImage, or <link rel="modulepreload">. Audit and fetch the long tail. Recipe in AGENTS.md covers all three asset roots (uploads, themes, plugins).

Phase 4 — Convert paths to root-relative

wget -k produces a mix of ../wp-content/... (depth-relative) and bare wp-content/... (homepage). Both work locally but break the moment a page moves. Convert to root-relative /wp-content/... everywhere:

sh
python3 scripts/rewrite-paths.py output/ urls.txt --source-domain example.com

The script derives the page-slug list from urls.txt, not from a directory walk — otherwise wget-grabbed archive directories like category/, feed/, author/, wp-json/ get wrongly classified as pages and their inter-page links get mis-rewritten.

The script defaults to WordPress asset roots (wp-content, wp-includes). For non-WP sources, override with --asset-roots: e.g. --asset-roots sites/default/files,sites/default/themes for Drupal, --asset-roots content/images for Ghost. The rest of the rewriter is CMS-agnostic.

Phase 5 — Brand the static output

So future-you (or anyone reading view-source) can tell at a glance that this is the static clone, not the live WP install:

sh
python3 scripts/insert-banner.py output/

Inserts an HTML comment after <!DOCTYPE html> on every page. Idempotent. Then replace the "Generated by Yoast SEO" attribution in wp-content/plugins/wordpress-seo/css/main-sitemap.xsl — recipe in AGENTS.md.

Phase 6 — Replace WP runtime hooks

Three categories of WP-only markup that breaks once the backend is gone:

1. Comment forms, reply links, and dead head tags. One pass:

sh
python3 scripts/strip-wp-runtime.py output/

Removes <div id="respond"> blocks (the comment form), comment-reply-link anchors in both block-theme and classic-theme variants, the comment-reply-js script tag and its underlying file, and dead <head> tags (REST API discovery, RSD, oEmbed alternates, RSS alternates, archive next links). Match-by-class throughout — no language assumptions about link text.

2. Plausible analytics proxy. The WP plugin proxies the script through /wp-content/uploads/<hash>/pa-XXX.js and posts events back to /wp-json/.... Both endpoints disappear. Replace the two-script block with the standard tracker — recipe in AGENTS.md.

3. Gravatar avatars. Self-host every distinct (hash, size) pair, drop the ?s=N&d=mm requests to a third party:

sh
python3 scripts/selfhost-gravatars.py output/

Saves under avatars/ and rewrites every reference. Detects extension from response bytes (PNG fallback vs JPEG real avatar), keeps size variants separate (?s=40 and ?s=80 are different files).

After the scripts, audit remaining absolute source-domain URLs (recipe in AGENTS.md) and triage by case: author archives → strip the <a> wrapper, server-rendered iframes → drop the wrapping <p>, Gravity Forms script blocks → strip on gform-mention, etc.

Phase 7 — Copy robots.txt

Not linked from HTML; fetch it explicitly. Adjust the Sitemap: reference if the deployed sitemap path differs from the source.

Show full SKILL.md (495 more words)Show less
Phase 8 — Verify locally

Serve from output/ with python3 -m http.server, then run the verify checklist in AGENTS.md:

  1. Every URL in urls.txt resolves to a file (no missed pages).
  2. No remaining https://<source-domain>/ outside the canonical / og:url / JSON-LD allow-list.
  3. No broken internal links from wget --spider.
  4. Spot-check the homepage and a deep page in a browser. Watch srcset images, sidebar widgets, and the header banner — those break silently if missed.
Phase 9 — Deploy

Host-specific recipes in AGENTS.md:

  • Cloudflare Pages — _redirects, _headers, "no build command, no output directory" defaults.
  • Netlify — same _redirects / _headers syntax, plus netlify.toml.
  • Vercel — vercel.json with redirects / headers.
  • Generic — nginx try_files, Apache Options +MultiViews.

Gotchas

These are the things that bit us. Don't repeat them.

  1. Cloudflare bot protection 403s the default Wget/1.x UA. Always set a real browser UA + Accept / Accept-Language headers (recipe). If you see 403 Forbidden after a burst of requests, that's it — back off, switch UA, retry.

  2. Cross-page link rewriting only works in a single wget invocation. wget's -k only rewrites to local paths it sees in the current run. If a page was downloaded in a separate invocation (e.g. to recover from a 403 on one URL), its links to the rest stay absolute. Solution: redo the full scrape once you have the right UA. Don't piecemeal it. If you're scraping at scale (10K+ URLs) and can't fit in one run, scrape in batches and re-run scripts/rewrite-paths.py afterwards as the canonical pass — -k's output is then redundant.

  3. Default publish directory by host. Cloudflare Pages serves the repo root when no build command is configured. Netlify and Vercel also default to root. If you scraped into output/, either move files to the repo root (git mv output/* .) or configure the host to publish from output/. Symptom of the wrong setup on Pages: every URL 404s with R2-style headers (access-control-allow-origin: *, cache-control: no-store) instead of a Pages-branded 404.

  4. WordPress Offload Media plugins route /wp-content/uploads/ to R2 / S3 buckets. wget may successfully fetch an image even when later direct access 404s (intermittent or partial bucket sync). Trust your local copy — that's why we scrape and self-host.

  5. Sitemaps and the Yoast XSL aren't linked from HTML. wget -p won't find them. Fetch explicitly in Phase 1.

  6. Filenames with ?ver=... query strings. wget keeps these as literal filenames; HTML uses %3F encoding. Standard servers (Pages, Netlify, Vercel, python -m http.server) URL-decode and serve correctly. Don't try to "clean these up" unless something actually breaks.

  7. og:url, canonical, JSON-LD stay absolute. They identify the canonical resource and are correct as-is when redeploying to the same domain. Only rewrite if changing domains.

  8. sed -i '' is macOS / BSD only. GNU sed needs sed -i (no empty-string argument). Recipes in AGENTS.md flag the macOS-isms; default to the Python scripts where there's a choice — they're portable.

Output structure

text
<repo-root>/
  index.html                 ← homepage
  <slug>/index.html          ← one per URL from sitemap
  wp-content/                ← assets (themes, uploads, plugins)
  wp-includes/               ← block library CSS, et al.
  avatars/                   ← self-hosted Gravatars (Phase 6)
  sitemap_index.xml
  page-sitemap.xml           ← + any other child sitemaps
  wp-content/plugins/wordpress-seo/css/main-sitemap.xsl
  robots.txt
  _redirects                 ← optional, host-specific
  _headers                   ← optional, host-specific

Push to a git host and connect to the static host with no build command and no build output directory — defaults work.

© jdevalk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts) in wp-static-clone of jdevalk/skills.

  • SKILL.md
  • AGENTS.md
  • README.md
  • scripts/insert-banner.py
  • scripts/rewrite-paths.py
  • scripts/selfhost-gravatars.py
  • scripts/strip-wp-runtime.py

Open the folder on GitHubat commit 106fc68

Compare with similar skills

Wp Static Clone next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Wp Static Clone compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Wp Static Clone this skilljdevalk/skills104—~2.6kAutomated safety check: PassMIT
Memstack Deployment Hetzner Setupcwinvestments/memstack423—~4.1kAutomated safety check: NotesProprietary
Nginx To Higress Migrationhigress-group/higress9.5k—~3.9kAutomated safety check: PassApache-2.0
NGINX Ingress Controller Feature Checklistsnginx/kubernetes-ingress5.1k—~1.4kAutomated safety check: PassApache-2.0
NGINX Ingress Policy CRD Guidenginx/kubernetes-ingress5.1k—~2kAutomated safety check: PassApache-2.0
Node Response Time Reviewblotcms/blot2k—~1.9kAutomated safety check: PassAGPL-3.0

Similar skills

  • Memstack Deployment Hetzner Setup

    cwinvestments/memstack

    A skill your agent uses when the user says 'Hetzner', 'VPS setup', 'server provisioning', 'deploy to VPS', 'hetzner-setup', 'cloud server', or needs to provision, harden, and deploy applications to…

    423 GitHub stars~4.1k tokensUpdated 11 days ago
    DevOps & CloudAuto-check: notes
  • Nginx To Higress Migration

    higress-group/higress

    Migrate from ingress-nginx to Higress in Kubernetes environments.

    9.5k GitHub stars~3.9k tokensUpdated 3 days ago
    DevOps & CloudAuto-check passed
  • Gives step-by-step checklists for adding Ingress annotations, VirtualServer fields and Helm values to the NGINX Kubernetes Ingress Controller, with common gotchas.

    5.1k GitHub stars~1.4k tokensUpdated today
    DevOps & CloudAuto-check passed
  • NGINX Ingress Policy CRD Guide

    nginx/kubernetes-ingress

    Step-by-step checklist for adding a new Policy CRD type to the NGINX Ingress Controller, from the Go types and validation to config generation and templates.

    5.1k GitHub stars~2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Analyze production Node.js app container response times to find slow-rendering sites, cross-checking against nginx queuing delay to rule out false positives (a site only looks slow because the event…

    2k GitHub stars~1.9k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Deploy Setup

    garrytan/gstack

    Detects where an app deploys, its production URL and health checks, then saves the deploy configuration in CLAUDE.md for /land-and-deploy.

    136k GitHub stars~11k tokensUpdated today
    DevOps & CloudAuto-check: notes

More from jdevalk/skills

All 9 skills in this repo
  • Astro SEO

    jdevalk/skills

    Audits and improves SEO for Astro sites. An agent skill from jdevalk/skills.

    104 GitHub stars~4.9k tokensUpdated 3 mo ago
    Auto-check passed
  • Content SEO

    jdevalk/skills

    Audits a blog post draft or page copy for content-level SEO: search intent fit, focus keyphrase placement, E-E-A-T signals (experience, expertise, authoritativeness, trustworthiness), helpfulness…

    104 GitHub stars~2.3k tokensUpdated 3 mo ago
    Auto-check passed
  • Metadata Check

    jdevalk/skills

    Reviews short high-value strings — page titles, meta descriptions, schema description fields, FAQ answers, GitHub repo taglines, profile bios, social-card copy, and other metadata where Flesch and…

    104 GitHub stars~1k tokensUpdated 3 mo ago
    Auto-check passed
  • Readability Check

    jdevalk/skills

    Runs a readability audit on a blog post draft or other multi-paragraph prose, calibrated for readers who read English as a second language.

    104 GitHub stars~3.3k tokensUpdated 3 mo ago
    Auto-check passed
  • Static SEO

    jdevalk/skills

    Audits and improves SEO for static HTML sites. An agent skill from jdevalk/skills.

    104 GitHub stars~4.8k tokensUpdated 3 mo ago
    Auto-check passed
  • GitHub Profile

    jdevalk/skills

    Audits and optimizes GitHub profile pages — profile README, metadata fields, pinned repositories, stats widgets, and contribution visibility.

    104 GitHub stars~1.9k tokensUpdated 3 mo ago
    Auto-check passed

Questions about Wp Static Clone

What does Wp Static Clone do?

Clones a live WordPress (or other CMS-driven) site into a static HTML site deployable on any static host (Cloudflare Pages, Netlify, Vercel, S3+CloudFront, plain Apache/nginx). Wp Static Clone is an agent skill from jdevalk/skills. Clones a live WordPress (or other CMS-driven) site into a static HTML site deployable on any static host (Cloudflare Pages, Netlify, Vercel, S3+CloudFront, plain Apache/nginx).

When should I use Wp Static Clone?

Wp Static Clone fits situations like: the user wants to scrape; move to [host] a WordPress site; asks to turn a sitemap into deployable static HTML.

How do I install Wp Static Clone in Claude Code?

Run `npx skills add jdevalk/skills --skill wp-static-clone -a claude-code`. Or copy the skill folder (wp-static-clone in jdevalk/skills) into .claude/skills/wp-static-clone in your project. Claude Code loads it when a task matches its description.

How do I install Wp Static Clone in Codex?

Run `npx skills add jdevalk/skills --skill wp-static-clone -a codex`. Or copy the skill folder (wp-static-clone in jdevalk/skills) into .agents/skills/wp-static-clone in your project. Codex loads it when a task matches its description.

Can I use Wp Static Clone in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jdevalk/skills --skill wp-static-clone -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/wp-static-clone, .gemini/skills/wp-static-clone, .github/skills/wp-static-clone and .opencode/skills/wp-static-clone in your project.

What does Wp Static Clone need to run?

Going by SKILL.md and its folder, Wp Static Clone needs Python for the scripts in its folder and the command-line tools its instructions call (python3, wget, git and python). Our summary lists: Python 3.

Does Wp Static Clone access the network?

SKILL.md contains no URLs. Its commands use wget and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Wp Static Clone safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Wp Static Clone use?

Wp Static Clone is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Wp Static Clone use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Wp Static Clone?

Skills that share tags, products or a category with Wp Static Clone: Memstack Deployment Hetzner Setup (cwinvestments/memstack, 423 stars), Nginx To Higress Migration (higress-group/higress, 9.5k stars), NGINX Ingress Controller Feature Checklists (nginx/kubernetes-ingress, 5.1k stars) and NGINX Ingress Policy CRD Guide (nginx/kubernetes-ingress, 5.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Wp Static Clone?

jdevalk (a GitHub user) maintains it in jdevalk/skills, which has 104 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on July 5, 2026.

Source: jdevalk/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.