Agent skill

Pp Archive Is

by mvanhorn in mvanhorn/printing-press-library

A skill your agent uses whenever the user wants to archive a URL, bypass a paywall, look up an existing archive, view a cached version of a webpage, pull article text from archive.today or the…

Apache-2.0Auto-check: notes

Install Pp Archive Is

skills CLI
$ npx skills add mvanhorn/printing-press-library --skill pp-archive-is -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mvanhorn/printing-press-library pp-archive-is --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mvanhorn/printing-press-library.git skills-src && mkdir -p .claude/skills && cp -r skills-src/cli-skills/pp-archive-is .claude/skills/pp-archive-is && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pp-archive-is
GitHub stars
2.1k
Token cost
~3k tokens
SKILL.md length
1,171 words
Files
1
Skills in repo
506
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses whenever the user wants to archive a URL, bypass a paywall, look up an existing archive, view a cached version of a webpage, pull article text from archive.today or the…

  • Works in 3 steps: Install via the Printing Press… → Verify: archive-is-pp-cli --version → Ensure the reported install directory is…
  • The user wants to archive a URL
  • SKILL.md covers Prerequisites: Install the CLI, When to Use This CLI, Unique Capabilities and Command Reference, plus 7 more sections
  • Calls go, npx and claude; reaches wsj.com and ft.com

What it does

Pp Archive Is is an agent skill from mvanhorn/printing-press-library. Use this skill whenever the user wants to archive a URL, bypass a paywall, look up an existing archive, view a cached version of a webpage, pull article text from archive.today or the Wayback Machine, or batch-archive a list of URLs. archive.today + Wayback Machine CLI with lookup-before-submit, automatic fallback when one backend is down, and agent-friendly output. No API key required. Triggers on phrasings like 'archive this article', 'bypass the paywall on this link', 'grab the cached text', 'save this url to…

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Official library of CLIs generated by the CLI Printing Press. Endorsed, tested, and community-contributed. The licence is Apache-2.0.

When your agent uses it

  • The user wants to archive a URL
  • Bypass a paywall
  • Look up an existing archive
  • View a cached version of a webpage

Example prompts

  • “archive this article”
  • “bypass the paywall on this link”
  • “grab the cached text”
  • “/pp-archive-is”

Requirements

  • Node.js
  • Pre-approved tools (allowed-tools): Read, Bash

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Install via the Printing Press installer. It defaults binaries to $HOME/.local/bin on macOS/Linux and…
  2. Verify: archive-is-pp-cli --version
  3. Ensure the reported install directory is on $PATH for the agent/runtime that will invoke this skill.

What it can do on your machine

Read from SKILL.md and the folder at commit d9a1696. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • go
    • npx
    • claude

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • wsj.com
    • ft.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pp Archive Is loads about 3k tokens when it runs. Until then it costs about 154 tokens; SKILL.md has 1,171 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~154
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mvanhorn/printing-press-library at commit d9a1696, republished under its Apache-2.0 licence (© mvanhorn). 1,171 words, ~3,028 tokens.

Download SKILL.mdSave it as .claude/skills/pp-archive-is/SKILL.md (or your agent's skills folder).
name
pp-archive-is
description
Use this skill whenever the user wants to archive a URL, bypass a paywall, look up an existing archive, view a cached version of a webpage, pull article text from archive.today or the Wayback Machine, or batch-archive a list of URLs. archive.today + Wayback Machine CLI with lookup-before-submit, automatic fallback when one backend is down, and agent-friendly output. No API key required. Triggers on phrasings like 'archive this article', 'bypass the paywall on this link', 'grab the cached text', 'save this url to archive.today', 'check if this was already archived', 'bulk archive these 20 URLs'.
allowed-tools
Read, Bash
author
Matt Van Horn
license
Apache-2.0
argument-hint
<command> [args] | install cli|mcp
<!-- GENERATED FILE — DO NOT EDIT.
     This file is a verbatim mirror of library/media-and-entertainment/archive-is/SKILL.md,
     regenerated post-merge by tools/generate-skills/. Hand-edits here are
     silently overwritten on the next regen. Edit the library/ source instead.
     See the repository agent guide, section "Generated artifacts: registry.json, cli-skills/". -->

archive.today — Printing Press CLI

Prerequisites: Install the CLI

This skill drives the archive-is-pp-cli binary. You must verify the CLI is installed before invoking any command from this skill. If it is missing, install it first:

  1. Install via the Printing Press installer. It defaults binaries to $HOME/.local/bin on macOS/Linux and %LOCALAPPDATA%\Programs\PrintingPress\bin on Windows:
    bash
    npx -y @mvanhorn/printing-press-library install archive-is --cli-only
  2. Verify: archive-is-pp-cli --version
  3. Ensure the reported install directory is on $PATH for the agent/runtime that will invoke this skill.

If the npx install fails (no Node, offline, etc.), fall back to a direct Go install (requires Go 1.26.6 or newer):

bash
go install github.com/mvanhorn/printing-press-library/library/media-and-entertainment/archive-is/cmd/archive-is-pp-cli@latest

If --version reports "command not found" after install, the runtime cannot see the binary directory on $PATH. Do not proceed with skill commands until verification succeeds.

When to Use This CLI

Reach for this whenever a user wants to archive a URL, read a paywalled article, check whether something was previously archived, or batch-capture a list of URLs for research. Specifically good when:

  • A user sends a paywalled link and asks "can you read this" → read fetches text via archive
  • They want to preserve a URL that might change → save forces a fresh capture
  • They want historical versions → history lists all known snapshots
  • They have 20+ URLs to archive → bulk runs rate-limited batch archival

Don't reach for this if the URL is trivially scrapeable without archive services (no paywall, robots-allowed, direct HTTP works), or if the user wants the original source rather than a cached version.

Unique Capabilities

The whole CLI is unique — archive.today has no official API. But within this CLI, certain commands are the differentiators.

The hero commands
  • read <url> — Find or create an archive for a URL. Looks up existing snapshots first (Memento timegate → CDX fallback); submits a fresh capture only if nothing exists. The "always do the right thing" command.

    This is how 90% of agent calls should start. It's idempotent — calling it twice on the same URL doesn't double-submit.

  • get <url> [--raw] / tldr <url> — Fetch article text, optionally LLM-summarized. Automatic Wayback fallback when archive.today serves a CAPTCHA (which happens daily to cloud IPs).

    tldr pipes the fetched text through a summarization step — useful for agent chains where you want a short take without shipping 20KB of HTML back.

Durability operations
  • save <url> — Force a fresh capture via /submit/?url=<x>&anyway=1. Use when read returns an existing snapshot that's too old or missing a paywall update.

  • history <url> — List all known snapshots via Memento timemap parsing. Shows every capture date across both archive.today and Wayback.

  • bulk [file] — Rate-limited batch archiving from a file or stdin. Reads URLs one per line, submits each with backoff, returns a report of successes / failures / pre-existing.

    grep -oE 'https?://[^ )]+' notes.md | archive-is-pp-cli bulk - archives every URL in a markdown file.

  • request <url> — Fire-and-forget submit with optional wait+poll. Useful for long captures where you want to come back later.

Observability
  • snapshots newest <url> — Just the newest snapshot URL for a target, useful in scripts.

  • captures — List your local capture index (post-sync).

  • feeds — archive.today's global recent-archives feed.

  • --backend archive-is,wayback — Every read/get accepts a backend preference. Defaults to archive-is with Wayback fallback; flip the order for Wayback-primary.

Command Reference

Archive + retrieve:

  • archive-is-pp-cli read <url> — Find or create (hero command)
  • archive-is-pp-cli get <url> — Fetch article text (with Wayback fallback)
  • archive-is-pp-cli tldr <url> — Fetch + summarize
  • archive-is-pp-cli save <url> — Force fresh capture
  • archive-is-pp-cli request <url> — Fire-and-forget submit
  • archive-is-pp-cli request check <url> — Check whether a submitted request has completed

Listing + history:

  • archive-is-pp-cli history <url> — All known snapshots
  • archive-is-pp-cli snapshots <url> — Best known snapshot for a URL
  • archive-is-pp-cli snapshots newest <url> — Newest snapshot URL
  • archive-is-pp-cli snapshots timemap <url> — Raw Memento timemap
  • archive-is-pp-cli captures — Local capture index
  • archive-is-pp-cli feeds — Global recent feed

Batch:

  • archive-is-pp-cli bulk [file] — Batch from file or stdin

Local store:

  • archive-is-pp-cli sync / export / import / workflow archive — Local SQLite ops

Auth + health:

  • archive-is-pp-cli auth — Config (no API key needed; auth is a no-op)
  • archive-is-pp-cli doctor — Verify backend reachability

Recipes

Read a paywalled article
bash
archive-is-pp-cli read "https://www.wsj.com/articles/..." --agent
# or: return just the text
archive-is-pp-cli get "https://www.wsj.com/articles/..." --agent

read returns the archive URL (finding existing or creating new). get returns extracted article text by default, falling back to Wayback if archive.today CAPTCHAs; add --raw only when you need the archived HTML.

Show full SKILL.md (492 more words)Show less
Preserve a URL before it changes
bash
archive-is-pp-cli save "https://example.com/important-page" --agent
archive-is-pp-cli history "https://example.com/important-page" --agent  # verify

Force capture, then check history to confirm the new snapshot registered.

Bulk archive a research batch
bash
grep -oE 'https?://[^ )]+' research-notes.md | archive-is-pp-cli bulk - --agent
# or from a file:
archive-is-pp-cli bulk urls.txt --agent

Reads URLs one per line, submits each with exponential backoff, returns per-URL status (archived, pre-existing, failed) as JSON.

Wayback-preferred for a reliable-read
bash
archive-is-pp-cli read "https://ft.com/content/xyz" --backend wayback,archive-is --agent

Use when the Wayback Machine snapshot is known to be cleaner or archive.today is rate-limiting.

Auth Setup

No API key required. Archive.today and Wayback Machine are both public. The auth subcommand exists for consistency but is a no-op — doctor reports "Auth: not required" which is the expected state.

Optional env:

  • ARCHIVE_IS_BASE_URL — override archive.today host (for mirrors)
  • WAYBACK_BASE_URL — override Wayback Machine host

Agent Mode

Add --agent to any command. Expands to --json --compact --no-input --no-color --yes --no-prompt. Every action command also prints structured next_actions hints on stderr when called non-interactively — the calling agent sees "tried X, got Y, consider Z" automatically.

Notable flags:

  • --submit-timeout <duration> — max wait for a fresh submit (default 10m; 0 = unbounded)
  • --backend archive-is,wayback — backend preference and fallback order
  • --raw — return raw HTML from get instead of extracted text
Filtering output

--select accepts dotted paths to descend into nested responses; arrays traverse element-wise:

bash
archive-is-pp-cli <command> --agent --select id,name
archive-is-pp-cli <command> --agent --select items.id,items.owner.name

Use this to narrow huge payloads to the fields you actually need — critical for deeply nested API responses.

Response envelope

Data-layer commands wrap output in {"meta": {...}, "results": <data>}. Parse .results for data and .meta.source to know whether it's live or local. The N results (live) summary is printed to stderr only when stdout is a TTY; piped/agent consumers see pure JSON on stdout.

Exit Codes

CodeMeaning
0Success
2Usage error
3Not found (no snapshot exists)
5API error (archive.today or Wayback down)
7Rate limited (too many submits)

Installation

bash
go install github.com/mvanhorn/printing-press-library/library/media-and-entertainment/archive-is/cmd/archive-is-pp-cli@latest
archive-is-pp-cli doctor
MCP Server
bash
go install github.com/mvanhorn/printing-press-library/library/media-and-entertainment/archive-is/cmd/archive-is-pp-mcp@latest
claude mcp add archive-is-pp-mcp -- archive-is-pp-mcp

Argument Parsing

Given $ARGUMENTS:

  1. Empty, help, or --help → run archive-is-pp-cli --help
  2. install → CLI; install mcp → MCP
  3. Anything that looks like a URL, or "archive <url>" / "bypass paywall on <url>" → read <url> --agent is the default — it's idempotent and covers the 90% case.
  4. "bulk archive" / "archive these" → bulk from stdin if URLs are pasted, else ask for the file path.
<!-- pr-218-features -->

Agent Workflow Features

This CLI exposes three shared agent-workflow capabilities patched in from cli-printing-press PR #218.

Named profiles

Persist a set of flags under a name and reuse them across invocations.

bash
# Save the current non-default flags as a named profile
archive-is-pp-cli profile save <name>

# Use a profile — overlays its values onto any flag you don't set explicitly
archive-is-pp-cli --profile <name> <command>

# List / inspect / remove
archive-is-pp-cli profile list
archive-is-pp-cli profile show <name>
archive-is-pp-cli profile delete <name> --yes

Flag precedence: explicit flag > env var > profile > default.

--deliver

Route command output to a sink other than stdout. Useful when an agent needs to hand a result to a file, a webhook, or another process without plumbing.

bash
archive-is-pp-cli <command> --deliver file:/path/to/out.json
archive-is-pp-cli <command> --deliver webhook:https://hooks.example/in

File sinks write atomically (tmp + rename). Webhook sinks POST application/json (or application/x-ndjson when --compact is set). Unknown schemes produce a structured refusal listing the supported set.

feedback

Record in-band feedback about this CLI from the agent side of the loop. Local-only by default; safe to call without configuration.

bash
archive-is-pp-cli feedback "what surprised you or tripped you up"
archive-is-pp-cli feedback list         # show local entries
archive-is-pp-cli feedback clear --yes  # wipe

Entries append to ~/.archive-is-pp-cli/feedback.jsonl as JSON lines. When ARCHIVE_IS_FEEDBACK_ENDPOINT is set and either --send is passed or ARCHIVE_IS_FEEDBACK_AUTO_SEND=true, the entry is also POSTed upstream (non-blocking — local write always succeeds).

© mvanhorn, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in cli-skills/pp-archive-is of mvanhorn/printing-press-library.

Open the folder on GitHubat commit d9a1696

Compare with similar skills

Pp Archive Is next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pp Archive Is compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pp Archive Is this skillmvanhorn/printing-press-library2.1k—~3kAutomated safety check: NotesApache-2.0
OpenSpec Bulk Change ArchiverFission-AI/OpenSpec72k2 repos~5.6kAutomated safety check: PassMIT
Hunt Auth Bypasssickn33/agentic-awesome-skills47k1 repos~6.6kAutomated safety check: PassMIT
CTF ZIP Archive Key Recoveryzhaoxuya520/reverse-skill41k1 repos~1.5kAutomated safety check: PassMIT
Conversation Archivegarrytan/gbrain31k—~5.6kAutomated safety check: PassMIT
Recipe Label And Archive Emailsgoogleworkspace/cli31k—~240Automated safety check: PassApache-2.0

Similar skills

  • Archives several completed OpenSpec changes in one operation, checking the codebase to resolve spec conflicts rather than archiving blindly.

    72k GitHub starsUsed in 2 repos~5.6k tokens
    DevelopmentAuto-check passed
  • Hunt Auth Bypass

    sickn33/agentic-awesome-skills

    Hunting skill for auth bypass vulnerabilities. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 1 repo~6.6k tokens
    Backend & APIsAuto-check passed
  • CTF ZIP Archive Key Recovery

    zhaoxuya520/reverse-skill

    Recovers internal keys from legacy ZipCrypto-protected ZIP archives in CTF challenges through a known-plaintext method, instead of brute force.

    41k GitHub starsUsed in 1 repo~1.5k tokens
    SecurityAuto-check passed
  • Conversation Archive

    garrytan/gbrain

    Import AI-assistant chat exports (ChatGPT, Claude, Perplexity) and agent session transcripts into the brain as one dated page per conversation under conversations/, validate each page against the…

    31k GitHub stars~5.6k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Apply Gmail labels to matching messages and archive them to keep your inbox clean.

    31k GitHub stars~240 tokensUpdated 3 days ago
    Productivity & AutomationAuto-check passed
  • Detecting Secure Boot Bypass

    mukul975/Anthropic-Cybersecurity-Skills

    Detect UEFI Secure Boot bypasses and bootkits such as BlackLotus and Bootkitty by verifying Secure Boot state, checking dbx revocation currency, and hashing EFI boot binaries against known-bad sets…

    34k GitHub stars~2.6k tokensUpdated 1 mo ago
    SecurityAuto-check: notes

More from mvanhorn/printing-press-library

All 506 skills in this repo
  • Agent Desktop

    mvanhorn/printing-press-library

    Desktop automation through the real Rust agent-desktop CLI, published in Printing Press through a small bridge.

    2.1k GitHub stars~2.3k tokensUpdated today
    Auto-check: notes
  • Gfonts

    mvanhorn/printing-press-library

    Search, browse, and download Google Fonts from the terminal via the gfonts CLI.

    2.1k GitHub stars~574 tokensUpdated today
    Auto-check passed
  • Pp 1688

    mvanhorn/printing-press-library

    The free, offline Trigger phrases: search 1688 for, find a factory on 1688 for, wholesale price on 1688 for, who is the cheapest supplier on 1688 for, compare 1688 suppliers for, use 1688, run 1688.

    2.1k GitHub stars~3k tokensUpdated today
    Auto-check: notes
  • Pp Activity Japan

    mvanhorn/printing-press-library

    Inspect known Activity Japan plan IDs or URLs, compare dated prices and sessions, check language-sitemap coverage, and hand off to canonical booking pages.

    2.1k GitHub stars~2k tokensUpdated today
    Auto-check: notes
  • Pp Adminbyrequest

    mvanhorn/printing-press-library

    Every Admin By Request portal action, plus a local SQLite mirror of audit, events, inventory and requests for ad-hoc...

    2.1k GitHub stars~3.3k tokensUpdated today
    Auto-check: notes
  • Pp Agent Capture

    mvanhorn/printing-press-library

    macOS screen capture, window recording, GIF conversion, and agent evidence bundles from the terminal.

    2.1k GitHub stars~1.6k tokensUpdated today
    Auto-check: notes

Questions about Pp Archive Is

What does Pp Archive Is do?

A skill your agent uses whenever the user wants to archive a URL, bypass a paywall, look up an existing archive, view a cached version of a webpage, pull article text from archive.today or the…. Pp Archive Is is an agent skill from mvanhorn/printing-press-library.today or the Wayback Machine, or batch-archive a list of URLs.

When should I use Pp Archive Is?

Pp Archive Is fits situations like: the user wants to archive a URL; bypass a paywall; look up an existing archive; view a cached version of a webpage.

How do I install Pp Archive Is in Claude Code?

Run `npx skills add mvanhorn/printing-press-library --skill pp-archive-is -a claude-code`. Or copy the skill folder (cli-skills/pp-archive-is in mvanhorn/printing-press-library) into .claude/skills/pp-archive-is in your project. Claude Code loads it when a task matches its description.

How do I install Pp Archive Is in Codex?

Run `npx skills add mvanhorn/printing-press-library --skill pp-archive-is -a codex`. Or copy the skill folder (cli-skills/pp-archive-is in mvanhorn/printing-press-library) into .agents/skills/pp-archive-is in your project. Codex loads it when a task matches its description.

Can I use Pp Archive Is in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mvanhorn/printing-press-library --skill pp-archive-is -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pp-archive-is, .gemini/skills/pp-archive-is, .github/skills/pp-archive-is and .opencode/skills/pp-archive-is in your project.

What does Pp Archive Is need to run?

Going by SKILL.md and its folder, Pp Archive Is needs the command-line tools its instructions call (go, npx and claude). Our summary lists: Node.js. Its frontmatter pre-approves these tools: Read, Bash.

Does Pp Archive Is access the network?

SKILL.md names 2 domains. In commands or code: wsj.com and ft.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Pp Archive Is safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Pp Archive Is use?

Pp Archive Is is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pp Archive Is use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Pp Archive Is?

Skills that share tags, products or a category with Pp Archive Is: OpenSpec Bulk Change Archiver (Fission-AI/OpenSpec, 72k stars), Hunt Auth Bypass (sickn33/agentic-awesome-skills, 47k stars), CTF ZIP Archive Key Recovery (zhaoxuya520/reverse-skill, 41k stars) and Conversation Archive (garrytan/gbrain, 31k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pp Archive Is?

mvanhorn (a GitHub user) maintains it in mvanhorn/printing-press-library, which has 2,056 GitHub stars. The repository holds 506 skills in this directory. The repository was last updated on October 9, 2026.

Source: mvanhorn/printing-press-library on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.