Agent skill

Glmv PDF To Web

by zai-org in zai-org/GLM-skills

Convert a PDF (research paper, technical report, or project document) into a beautiful single-page academic/project website with a structured outline JSON.

Apache-2.0Auto-check passedDocuments & Office

Install Glmv PDF To Web

skills CLI
$ npx skills add zai-org/GLM-skills --skill glmv-pdf-to-web -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install zai-org/GLM-skills glmv-pdf-to-web --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/zai-org/GLM-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/glmv-pdf-to-web .claude/skills/glmv-pdf-to-web && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
glmv-pdf-to-web
GitHub stars
476
Token cost
~3k tokens
SKILL.md length
1,001 words
Files
5 (incl. scripts)
Skills in repo
16
Repo updated
First seen
Licence
Apache-2.0

At a glance

Convert a PDF (research paper, technical report, or project document) into a beautiful single-page academic/project website with a structured outline JSON.

  • Works in 7 steps: Create Output Directory → Convert PDF Pages to Images (DPI 120) → Read All Pages in Order → …
  • This skill when the user wants to make a paper page
  • SKILL.md covers Dependencies, When to Use, Output Directory Convention and Input, plus 4 more sections
  • Runs Python scripts from its folder; calls python, pip and curl

What it does

Glmv PDF To Web is an agent skill from zai-org/GLM-skills. Convert a PDF (research paper, technical report, or project document) into a beautiful single-page academic/project website with a structured outline JSON. Trigger this skill when the user wants to make a paper page, project homepage, or academic website from a PDF — in Chinese or English.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `scripts/crop.py`, `scripts/generate_web.py` and `scripts/pdf_to_images.py`).

It sits in Documents & Office, covering PDF. The repository describes itself as: Official skills for the GLM family of models. The licence is Apache-2.0.

When your agent uses it

  • This skill when the user wants to make a paper page
  • Project homepage
  • Academic website from a PDF — in Chinese

Example prompts

  • “/glmv-pdf-to-web”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Create Output Directory
  2. Convert PDF Pages to Images (DPI 120)
  3. Read All Pages in Order
  4. Plan Sections & Save outline.json
  5. Crop Required Images (Grounding + Subagent)
  6. Measure Cropped Image Dimensions
  7. Generate the Single-Page HTML

What it can do on your machine

Read from SKILL.md and the folder at commit 2ecd31c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip
    • curl
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip and curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Glmv PDF To Web loads about 3k tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 1,001 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~77
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from zai-org/GLM-skills at commit 2ecd31c, republished under its Apache-2.0 licence (© zai-org). 1,001 words, ~2,965 tokens.

Download SKILL.mdSave it as .claude/skills/glmv-pdf-to-web/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
glmv-pdf-to-web
description
Convert a PDF (research paper, technical report, or project document) into a beautiful single-page academic/project website with a structured outline JSON. Trigger this skill when the user wants to make a paper page, project homepage, or academic website from a PDF — in Chinese or English.

PDF → Academic Project Website Skill

Convert a research paper or technical document PDF into a polished single-page project website — the kind used for NeurIPS/CVPR/ICLR paper releases. Pages are converted locally at DPI 120, a structured outline.json is saved, images are cropped locally, and the final page is saved with generate_web.py.

Scripts are in: {SKILL_DIR}/scripts/

Dependencies

Python packages (install once):

bash
pip install pymupdf pillow

System tools: curl (pre-installed on macOS/Linux).

When to Use

Trigger when the user asks to create a webpage or project page from a PDF — phrases like: "make a project page from a PDF", "create a paper website", "build an academic website for this paper", "论文主页", "做项目主页", "根据pdf做网页", "把论文做成主页", or any similar intent in Chinese or English.

Output Directory Convention

All output goes under {WORKSPACE}/web/<pdf_stem>_<timestamp>/:

web/
└── <pdf_stem>_<timestamp>/
    ├── outline.json        ← structured web plan (WebPlan schema)
    ├── crops/              ← locally-saved cropped images
    │   ├── fig_arch_crop.png
    │   ├── table_results_crop.png
    │   └── ...
    └── index.html          ← the website
  • <pdf_stem> = PDF filename without extension
  • <timestamp> = format YYYYMMDD_HHMMSS
  • HTML references images via relative path crops/<name>_crop.png

Input

$ARGUMENTS is the path to the PDF file (local) or an HTTP/HTTPS URL.

  • If user provides a URL: download with curl first, then convert
  • If user provides a local PDF path: convert directly

Workflow

Phase 0 — Create Output Directory
python
import os, datetime
pdf_stem = os.path.splitext(os.path.basename(pdf_path))[0]
timestamp = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
out_dir = os.path.join(workspace, "web", f"{pdf_stem}_{timestamp}")
bash
mkdir -p "<out_dir>/crops"

Phase 1 — Convert PDF Pages to Images (DPI 120)

If the input is a URL, download it first:

bash
pdf_stem=$(basename "$ARGUMENTS" .pdf)
curl -L -o "/tmp/${pdf_stem}.pdf" "$ARGUMENTS"

Then convert (pass either the downloaded path or the original local path):

bash
python {SKILL_DIR}/scripts/pdf_to_images.py "<pdf_path>" --dpi 120

Outputs JSON to stdout:

json
[{"page": 1, "path": "/abs/path/page_001.png"}, ...]

Parse and store the full page → path map.


Phase 2 — Read All Pages in Order

View all page images sequentially before planning. Goal: pure understanding of the document's content, figures, and structure.

While reading, note:

  • Title, authors, affiliations, venue, year
  • Abstract text (verbatim)
  • Key contributions
  • Paper/Code/Dataset links (arXiv, GitHub, etc.)
  • Figures, tables, diagrams — which pages, rough regions
  • Teaser/hero figure if present

Do NOT plan sections yet — read everything first.


Phase 3 — Plan Sections & Save outline.json

Plan the website sections. Standard structure for academic papers (adapt as needed):

section_idPurpose
heroTitle, authors, venue badge, link buttons
abstractFull abstract text
contributions3–5 key contribution cards
methodArchitecture figure + method explanation
resultsQuantitative table + qualitative figures
conclusionBrief conclusion
citationBibTeX block

For each section that needs an image, identify:

  • Which page it comes from (the local page path from Phase 1)
  • A description of what the visual shows and why it belongs in this section

Save as <out_dir>/outline.json using exactly this schema:

json
{
  "project_title": "Paper Title",
  "lang": "English",
  "authors": ["Author One", "Author Two"],
  "sections_plan": [
    {
      "section_index": 1,
      "section_id": "hero",
      "title": "Hero",
      "content": "Title, authors, venue, teaser figure description",
      "required_images": [
        {
          "url": "<local_page_path_from_phase1>",
          "visual_description": "Figure 1: teaser showing input-output examples",
          "usage_reason": "Hero section visual to immediately show the paper's output"
        }
      ]
    }
  ]
}

Field notes:

  • lang: "Chinese" or "English" — match the PDF language
  • required_images: empty array [] if section needs no images
  • url: the local file path of the source page (from Phase 1 path field)
  • For images that need cropping, note the approximate region — exact crop boxes are determined in Phase 4

Write outline.json using the Write tool to <out_dir>/outline.json.


Phase 4 — Crop Required Images (Grounding + Subagent)

IMPORTANT: You MUST delegate ALL cropping to a clean subagent using the Agent tool. By this phase your context is very long (all page images + outline), which degrades visual coordinate accuracy. A fresh subagent with only the target image produces much more precise coordinates.

IMPORTANT: You MUST use the provided {SKILL_DIR}/scripts/crop.py script for ALL image cropping. Do NOT write your own cropping code, do NOT use PIL/Pillow directly, do NOT use any other method.

Read outline.json. Collect all crops needed, then launch one subagent per source page (or one per crop if pages differ). The subagent uses grounding-style localization — it views the image, locates the target element, and outputs a precise bounding box in normalized 0–999 coordinates.

Use the Agent tool like this:

Agent tool call:
  description: "Grounding crop page N"
  prompt: |
    You are a visual grounding and cropping assistant. Your task is to precisely
    locate specified visual elements in a page image and crop them out.

    ## Grounding method

    Use visual grounding to locate each target:
    1. Read the source image using the Read tool to view it
    2. Identify the target element described below
    3. Determine its bounding box as normalized coordinates in the 0–999 range:
       - 0 = left/top edge of the image
       - 999 = right/bottom edge of the image
       - These are thousandths, NOT pixels, NOT percentages (0–100)
       - Format: [x1, y1, x2, y2] where (x1,y1) is top-left, (x2,y2) is bottom-right
       - Example: [0, 0, 500, 500] = top-left quarter of the image
    4. Be precise: tightly bound the target element with a small margin (~10–20 units)
       around it. Do NOT crop too wide or too narrow.

    ## Source image
    <page_image_path>

    ## Crops needed

    For each crop below, first do grounding (locate the element), then crop:

    1. Name: "<descriptive_name>"
       Target: "<visual_description from outline.json>"
       Context: "<usage_reason from outline.json>"

    ## Crop command

    After determining the bounding box [X1, Y1, X2, Y2] for each target, run:
    ```bash
    python <SKILL_DIR>/scripts/crop.py \
        --path "<page_image_path>" \
        --box X1 Y1 X2 Y2 \
        --name "<crop_name>" \
        --out-dir "<out_dir>/crops"
    ```

    ## Verification

    After each crop, READ the output image to visually verify the correct region
    was captured. If the crop missed the target or is too wide/narrow, adjust the
    coordinates and re-run crop.py.

    ## Output

    Report the final results as a list:
    - crop_name: <name>, file: <output_filename>, box: [X1, Y1, X2, Y2]

Replace <page_image_path>, <SKILL_DIR>, <out_dir>, and crop details with actual values from your context.

The crop.py script outputs JSON: {"path": "/abs/path/<name>_crop.png"}

Collect results from all subagents and build the mapping: section_id → [crop filename, ...] to reference in HTML.

Launch subagents for independent pages in parallel when possible. Wait for all to complete before proceeding.


Show full SKILL.md (388 more words)Show less
Phase 5 — Measure Cropped Image Dimensions
bash
python3 -c "
from PIL import Image; import os, json
d = '<out_dir>/crops'
sizes = {}
for f in sorted(os.listdir(d)):
    if f.endswith('.png'):
        w, h = Image.open(os.path.join(d, f)).size
        sizes[f] = {'width': w, 'height': h, 'aspect': round(w/h, 2)}
print(json.dumps(sizes, indent=2))
"
Aspect ratioLayout recommendation
< 0.7 (tall/narrow)max-width: 400–500px, centered
0.7 – 1.3 (square-ish)max-width: 600–700px
> 1.3 (wide)Full-width, max-width: 100%
> 2.0 (very wide, e.g. tables)Full-width with horizontal scroll fallback

Phase 6 — Generate the Single-Page HTML

Step A — Write HTML to /tmp/website.html

  • All <img src="..."> must use relative paths: crops/<name>_crop.png
  • Do NOT use absolute paths

Step B — Save:

bash
python {SKILL_DIR}/scripts/generate_web.py \
    --html-file /tmp/website.html \
    --title "<paper title>" \
    --out-dir "<out_dir>/"

HTML Spec

A single self-contained HTML file — embedded CSS, minimal vanilla JS only. No external JS frameworks. Google Fonts CDN is fine.

Page layout:

  • Max content width: 900px, centered, comfortable side padding
  • Sticky top nav with section anchor links + smooth scroll
  • Looks good at 1200px wide; readable at 768px

Typography:

  • Two Google Fonts: one for headings, one for body/UI
  • Body: 17–18px, line-height 1.7
  • Strong heading hierarchy (h1 >> h2 >> h3)

Visual style:

  • If the user specifies a style, follow it exactly
  • Otherwise, infer an appropriate aesthetic from the paper's domain and tone (e.g. CV/ML paper → clean modern academic; systems paper → dark technical; humanities → warm editorial serif)
  • Define colors and fonts as CSS variables; no fixed palette or font choices are required

Section guidelines:

hero:

  • Large title (2–3rem), authors list with affiliation superscripts, venue badge pill
  • Link buttons: [📄 Paper] [💻 Code] [🗄️ Dataset] — grey out if no URL
  • Teaser figure below (if found)

abstract:

  • Verbatim text with subtle left border accent

contributions:

  • Cards in a 2–3 column CSS grid, each with Unicode symbol + heading + description

method:

  • Full-width architecture figure (<figure><img><figcaption>) + prose explanation

results:

  • Quantitative table as real <table> — use actual numbers from the PDF, best numbers bolded
  • Qualitative figures in a grid (2–4 images with captions)

conclusion:

  • 2–3 paragraphs

citation:

  • <pre><code> BibTeX block reconstructed from PDF metadata
  • "Copy" button using navigator.clipboard vanilla JS

Images:

  • All <img> use relative paths: crops/<name>_crop.png
  • Add loading="lazy" and descriptive alt
  • Wrap in <figure> with <figcaption>

Animations (subtle only):

  • Fade-in on scroll via IntersectionObserver + CSS transitions
  • Hover states on buttons/cards


Quality Checklist

  • Output directory named <pdf_stem>_<timestamp>/
  • outline.json saved with valid WebPlan schema
  • All crops saved to crops/ (local only)
  • All metadata (title, authors, venue, year) from the PDF
  • Abstract is verbatim
  • Quantitative table has real numbers from the paper
  • All crop images referenced via crops/<name>_crop.png
  • BibTeX block accurate and copyable
  • Nav anchors scroll to correct sections
  • generate_web.py called and confirmed success

Language

Match the PDF language. English paper → English website. Chinese paper → Chinese. No mixing.

© zai-org, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts) in skills/glmv-pdf-to-web of zai-org/GLM-skills.

  • SKILL.md
  • .DS_Store
  • scripts/crop.py
  • scripts/generate_web.py
  • scripts/pdf_to_images.py

Open the folder on GitHubat commit 2ecd31c

Compare with similar skills

Glmv PDF To Web next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Glmv PDF To Web compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Glmv PDF To Web this skillzai-org/GLM-skills476—~3kAutomated safety check: PassApache-2.0
Split PDFscunning1975/MixtapeTools4742 repos~2.9kAutomated safety check: PassNone
Paper Interpretationdigoal/blog8.6k—~1.5kAutomated safety check: PassGPL-2.0
Paper2slidesQuZhan51496/paper2anything450—~3.8kAutomated safety check: NotesApache-2.0
Paper LensYSQ-boop/paper-lens101—~1.3kAutomated safety check: PassApache-2.0
Geng Academic Fraud Detectorwooly99/geng-academic-fraud-detector278—~970Automated safety check: PassNone

Similar skills

  • Split PDF

    scunning1975/MixtapeTools

    Download, split, and deeply read academic PDFs. An agent skill from scunning1975/MixtapeTools.

    474 GitHub starsUsed in 2 repos~2.9k tokens
    Documents & OfficeAuto-check passed
  • 从论文 PDF 文件或论文 PDF URL 生成通俗易懂、图文并茂、带批判性评估的中文 Markdown 解读,并保存到当前项目的 markdown 目录。Use when the user asks to interpret,精读,解读,summarize,explain,analyze, or write an article from an academic paper PDF…

    8.6k GitHub stars~1.5k tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Paper2slides

    QuZhan51496/paper2anything

    Turn an academic paper PDF into a presentation deck (.pptx) end-to-end.

    450 GitHub stars~3.8k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check: notes
  • Paper Lens

    YSQ-boop/paper-lens

    Read and critically analyze one academic paper from an arXiv URL/ID or a local PDF, producing a source-grounded Markdown report that can grow from a quick read into a reviewer-level deep review.

    101 GitHub stars~1.3k tokensUpdated 11 days ago
    Documents & OfficeAuto-check passed
  • Geng Academic Fraud Detector

    wooly99/geng-academic-fraud-detector

    学术论文打假检测器,致敬耿同学。分析学术论文 PDF,检测数据造假、图片复用/拼接、Western blot 操纵、统计异常等学术不端行为。当用户提供论文 PDF 要求"查重"、"打假"、"检测造假"、"论文分析"、"学术打假"时使用。

    278 GitHub stars~970 tokensUpdated 4 mo ago
    Documents & OfficeAuto-check passed
  • Paper2poster

    QuZhan51496/paper2anything

    Convert academic papers (PDF) into conference posters (HTML/PNG).

    450 GitHub stars~9.2k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check: notes

More from zai-org/GLM-skills

All 16 skills in this repo
  • Glmocr

    zai-org/GLM-skills

    Extract text from images using GLM-OCR API. An agent skill from zai-org/GLM-skills.

    476 GitHub stars~1.1k tokensUpdated 5 mo ago
    Auto-check passed
  • Glm Image Gen

    zai-org/GLM-skills

    Official skill for generating high-quality images from text prompts using ZhiPu GLM-Image API.

    476 GitHub stars~2.9k tokensUpdated 5 mo ago
    Auto-check passed
  • Glmocr Formula

    zai-org/GLM-skills

    Official skill for recognizing and extracting mathematical formulas from images and PDFs into LaTeX format using ZhiPu GLM-OCR API.

    476 GitHub stars~2.2k tokensUpdated 5 mo ago
    Auto-check passed
  • Glmocr Handwriting

    zai-org/GLM-skills

    Official skill for recognizing handwritten text from images using ZhiPu GLM-OCR API.

    476 GitHub stars~1.7k tokensUpdated 5 mo ago
    Auto-check passed
  • Glmocr Table

    zai-org/GLM-skills

    Official skill for recognizing and extracting tables from images and PDFs into Markdown format using ZhiPu GLM-OCR API.

    476 GitHub stars~1.7k tokensUpdated 5 mo ago
    Auto-check passed
  • Glmv Caption

    zai-org/GLM-skills

    Generate captions (descriptions) for images, videos, and documents using ZhiPu GLM-V multimodal model series.

    476 GitHub stars~2k tokensUpdated 5 mo ago
    Auto-check: notes

Questions about Glmv PDF To Web

What does Glmv PDF To Web do?

Convert a PDF (research paper, technical report, or project document) into a beautiful single-page academic/project website with a structured outline JSON. Glmv PDF To Web is an agent skill from zai-org/GLM-skills. Convert a PDF (research paper, technical report, or project document) into a beautiful single-page academic/project website with a structured outline JSON.

When should I use Glmv PDF To Web?

Glmv PDF To Web fits situations like: this skill when the user wants to make a paper page; project homepage; academic website from a PDF — in Chinese.

How do I install Glmv PDF To Web in Claude Code?

Run `npx skills add zai-org/GLM-skills --skill glmv-pdf-to-web -a claude-code`. Or copy the skill folder (skills/glmv-pdf-to-web in zai-org/GLM-skills) into .claude/skills/glmv-pdf-to-web in your project. Claude Code loads it when a task matches its description.

How do I install Glmv PDF To Web in Codex?

Run `npx skills add zai-org/GLM-skills --skill glmv-pdf-to-web -a codex`. Or copy the skill folder (skills/glmv-pdf-to-web in zai-org/GLM-skills) into .agents/skills/glmv-pdf-to-web in your project. Codex loads it when a task matches its description.

Can I use Glmv PDF To Web in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zai-org/GLM-skills --skill glmv-pdf-to-web -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/glmv-pdf-to-web, .gemini/skills/glmv-pdf-to-web, .github/skills/glmv-pdf-to-web and .opencode/skills/glmv-pdf-to-web in your project.

What does Glmv PDF To Web need to run?

Going by SKILL.md and its folder, Glmv PDF To Web needs Python for the scripts in its folder and the command-line tools its instructions call (python, pip, curl and python3). Our summary lists: Python 3.

Does Glmv PDF To Web access the network?

SKILL.md contains no URLs. Its commands use pip and curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Glmv PDF To Web safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Glmv PDF To Web use?

Glmv PDF To Web is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Glmv PDF To Web use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Glmv PDF To Web?

Skills that share tags, products or a category with Glmv PDF To Web: Split PDF (scunning1975/MixtapeTools, 474 stars), Paper Interpretation (digoal/blog, 8.6k stars), Paper2slides (QuZhan51496/paper2anything, 450 stars) and Paper Lens (YSQ-boop/paper-lens, 101 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Glmv PDF To Web?

zai-org (a GitHub organization) maintains it in zai-org/GLM-skills, which has 476 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on April 15, 2026.

Source: zai-org/GLM-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.