Agent skill

Browser4 Web Miner

by platonai in platonai/Browser4

Groups similar web pages together and produces an interactive HTML report with clusters of related pages, plus Excel spreadsheets for analysis.

Apache-2.0Auto-check passedDocuments & Office

Install Browser4 Web Miner

skills CLI
$ npx skills add platonai/Browser4 --skill browser4-web-miner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install platonai/Browser4 browser4-web-miner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/platonai/Browser4.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/browser4-web-miner .claude/skills/browser4-web-miner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
browser4-web-miner
GitHub stars
1.2k
Token cost
~2.5k tokens
SKILL.md length
963 words
Files
1
Skills in repo
17
Repo updated
First seen
Licence
Apache-2.0

At a glance

Groups similar web pages together and produces an interactive HTML report with clusters of related pages, plus Excel spreadsheets for analysis.

  • Works in 3 steps: Full pipeline on a folder of pages → Rebuild views from an existing run → Try it on the sample dataset
  • The user wants to cluster downloaded HTML files
  • SKILL.md covers Quick Start, When to Use, How It Works and Patterns, plus 7 more sections
  • Calls java; reaches github.com

What it does

Browser4 Web Miner is an agent skill from platonai/Browser4. Groups similar web pages together and produces an interactive HTML report with clusters of related pages, plus Excel spreadsheets for analysis. Use when the user wants to cluster downloaded HTML files, convert detail web pages into interactive views, or analyze a folder of web pages locally.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering HTML artifacts and Excel spreadsheets. It works with Java. The repository describes itself as: Browser4 — an AI-native browser engine for autonomous agents, intelligent extraction, and large-scale web automation. The licence is Apache-2.0.

When your agent uses it

  • The user wants to cluster downloaded HTML files
  • Convert detail web pages into interactive views
  • Analyze a folder of web pages locally

Example prompts

  • “Use the browser4-web-miner skill to group similar web pages together and produces an interactive HTML report with clusters of related pages, plus…”
  • “/browser4-web-miner”

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Full pipeline on a folder of pages
  2. Rebuild views from an existing run
  3. Try it on the sample dataset

What it can do on your machine

Read from SKILL.md and the folder at commit 0fdba82. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • java

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Browser4 Web Miner loads about 2.5k tokens when it runs. Until then it costs about 78 tokens; SKILL.md has 963 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from platonai/Browser4 at commit 0fdba82, republished under its Apache-2.0 licence (© platonai). 963 words, ~2,467 tokens.

Download SKILL.mdSave it as .claude/skills/browser4-web-miner/SKILL.md (or your agent's skills folder).
name
browser4-web-miner
description
Groups similar web pages together and produces an interactive HTML report with clusters of related pages, plus Excel spreadsheets for analysis. Use when the user wants to cluster downloaded HTML files, convert detail web pages into interactive views, or analyze a folder of web pages locally.
title
WebMiner — Convert Detail Web Pages into Interactive Views
tier
procedure

WebMiner — Convert Detail Web Pages into Interactive Views

Quick Start

bash
browser4-cli webminer install            # one-time install (Java 17+ auto-detected)
browser4-cli webminer all <html-dir>     # full pipeline: encode → cluster → views

WebMiner groups similar web pages together and produces an interactive HTML report with clusters of related pages — plus Excel spreadsheets for further analysis. Give it a folder of downloaded HTML files, and it handles the rest. Everything runs locally; no data leaves your machine.

When to Use

Use WebMiner when you have a folder of downloaded HTML pages and want to cluster them into interactive views and Excel reports — fully local, no LLM tokens. It complements rather than replaces browser4-cli crawl/swarm (which acquire pages): WebMiner analyzes pages you already have. Not for single-page extraction — use htmlsnapshot for that.

How It Works

WebMiner runs a three-stage local pipeline: encode converts each HTML page into a 69-dimension feature vector, cluster groups similar pages with SMILE KMeans (k auto-detected), and views renders an interactive HTML report plus Excel spreadsheets. Everything runs locally on your machine — no data leaves it, and no LLM tokens are consumed.

Patterns

1. Full pipeline on a folder of pages
bash
browser4-cli webminer all <html-dir>
2. Rebuild views from an existing run
bash
browser4-cli webminer views <result-dir>
3. Try it on the sample dataset
bash
browser4-cli webminer run-example

Flags

FlagApplies toDescription
--max-files <n>webminer allLimit the number of HTML files processed (default 40)
--output <dir>webminer allOverride the output directory
--resume [<project-id>]webminer allResume a previous run

Errors & Recovery

SymptomCauseFix
webminer install failsNo Java 17+ on PATHInstall JDK 17+ or point JAVA_HOME at it
webminer all finds no pagesDirectory has no .html filesCheck the input directory path and file extensions
Pipeline crashes on large corporaFree tier limit (< 1,000 pages)Reduce the corpus or use --max-files; see the commercial Spark tier for scale
Views land in an unexpected temp dirThe views stage uses the app task-output rootUse webminer views <result-dir> to rebuild beside the result dir

Using from the Browser4 CLI

WebMiner is a first-class Browser4 citizen: the browser4-cli webminer command installs, updates, and runs the tool natively (no PowerShell needed — the CLI locates a Java 17+ installation, preferring the JRE bundled with the Browser4 runtime, and launches scent-miner.jar directly). The JAR and its release metadata are installed to ~/.scent/webminer/.

bash
browser4-cli webminer install            # Download and install the latest release
browser4-cli webminer update             # Check for and install the latest release
browser4-cli webminer version            # Show installed and latest available versions
browser4-cli webminer uninstall          # Remove the installed release
browser4-cli webminer run-example        # Sample dataset + full pipeline (needs 7-Zip)
browser4-cli webminer all <html-dir>     # Full pipeline (encode → cluster → views)
browser4-cli webminer views <result-dir> # Rebuild views from an existing run
  • webminer all <dir> accepts the pipeline options directly (--max-files <n>, --output <dir>, --resume [<project-id>]).
  • Any other command is forwarded verbatim to scent-miner.jar, e.g. browser4-cli webminer encode <dir>.
  • Runs started through the CLI set -Dapp.name=webminer, so the views task-output root is %TEMP%\webminer-<user>\ml\tasks\... (<user> is the OS user name; see Output).
  • The bare webminer panel and webminer version keep the update check quiet: the GitHub → OSS-mirror fallback notices (rate limit, HTTP status, unreachable) are suppressed, and the Published line is omitted entirely when the release carries no published_at. webminer install / update still report the fallback.

Installing WebMiner

browser4-cli webminer install downloads, verifies, and installs the latest release (GitHub Releases with an Aliyun OSS mirror fallback; works on Windows, Linux, and macOS — no PowerShell needed):

bash
browser4-cli webminer install            # Download and install the latest release
browser4-cli webminer update             # Check for and install the latest release
browser4-cli webminer version            # Show installed and latest available versions
browser4-cli webminer uninstall          # Remove the installed release

Releases are installed to ~/.scent/webminer/ and checked against https://github.com/platonai/web-miner/releases. SHA-256 checksums are verified automatically on download.

You can also use the JAR directly if it's already available:

bash
java -jar scent-miner.jar <command> <args>

Converting Pages to Views

Running the Example

The run-example command downloads a pre-uploaded test dataset of real web pages, extracts it, and runs the full pipeline — no manual setup required beyond Java 17 and 7-Zip:

bash
browser4-cli webminer run-example

The dataset is cached at ~/.scent/test-data/amazon.com/ so subsequent runs skip the download.

Running on Your Own Pages
bash
# Full pipeline (one-shot)
browser4-cli webminer all /path/to/html/files

# Or with the JAR directly
java -jar scent-miner.jar all /path/to/html/files

The cluster count is always auto-detected from the data — this produces better results than guessing a number.

Show full SKILL.md (385 more words)Show less
Options
FlagDefaultPurpose
--max-files <n>40Maximum number of HTML files to process
--output <dir><html-dir>-ml-outputWhere to write the clustered results (CSV + clustering info; the views stage uses the app temp root — see Output)
--resume [<project-id>]—Pick up where a previous run left off. If no project ID is given, the most recent project is used.
Building Views from an Existing Run

If clustering has already completed and you just need to (re)build the views:

bash
java -jar scent-miner.jar views <html-dir>-ml-output/kmeans-result/p<timestamp>

Output

all produces two kinds of artifacts in two different places:

  1. Clustered results — written to <html-dir>-ml-output/kmeans-result/p<timestamp>/ (or wherever --output points): one result.csv per feature view (predictionAnd{Final,Minimal,Original}Features/result.csv) plus clusteringInfo.txt.
  2. Views (interactive HTML report + Excel + JSON) — the views stage of all writes them to the application's temp task-output root, NOT under <html-dir>-ml-output: %TEMP%\<app>-<user>\ml\tasks\unsupervised\result\p<timestamp>\predictionAndMinimalFeatures.views\ on Windows, and <java.io.tmpdir>/<app>-<user>/ml/tasks/unsupervised/result/p<timestamp>/predictionAndMinimalFeatures.views/ on Linux/macOS (/tmp/... on Linux, $TMPDIR on macOS) — the <app> prefix follows -Dapp.name (webminer when launched through browser4-cli webminer, pulsar for a direct java -jar run) and <user> is the OS user name. The end of the run prints the resolved absolute views path.

So after java -jar scent-miner.jar all ./html-pages/ the clustered results look like:

html-pages-ml-output/
  └── kmeans-result/
      └── p<timestamp>/
          ├── predictionAndFinalFeatures/result.csv
          ├── predictionAndMinimalFeatures/result.csv
          ├── predictionAndOriginalFeatures/result.csv
          └── clusteringInfo.txt

and the views (<project>.html, *.xlsx, *.json) live in the temp task-output directory printed by the run.

The real report is <project>.html (e.g. p<timestamp>.html), not index.html. The index.html inside the views directory is an auto-generated directory listing ("Index of predictionAndMinimalFeatures.views") — opening it shows a file list, not the interactive clustering report. Open <project>.html instead.

To place the views beside the clustered results (e.g. to archive them with the project), rebuild them from the result directory:

bash
browser4-cli webminer views <html-dir>-ml-output/kmeans-result/p<timestamp>
# (equivalent to: java -jar scent-miner.jar views <html-dir>-ml-output/kmeans-result/p<timestamp>)

This writes predictionAndMinimalFeatures.views/ inside the given result directory — the recommended way to locate artifacts, since the output path is explicit instead of an opaque temp path. Open the generated <project>.html in a browser to explore the clustering results. The .xlsx files can be opened in Excel for sorting, filtering, or further analysis.

Tips

  • Input files — only *.html and *.htm files are processed. Other files in the directory are ignored.
  • Resume interrupted runs — if a pipeline stops partway through, use --resume to continue from the last completed stage instead of starting over.
  • Offline only — WebMiner works with pre-downloaded HTML files. Use a browser, wget, or a crawler to fetch pages first.
  • Java 17 is required. Make sure java is on your PATH.

© platonai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/browser4-web-miner of platonai/Browser4.

Open the folder on GitHubat commit 0fdba82

Compare with similar skills

Browser4 Web Miner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Browser4 Web Miner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Browser4 Web Miner this skillplatonai/Browser41.2k—~2.5kAutomated safety check: PassApache-2.0
Docx4jplutext/docx4j2.4k—~2.5kAutomated safety check: PassNone
Eval Suite Plannermicrosoft/eval-guide138—~2.3kAutomated safety check: PassMIT
Dashboard BuilderTheCraigHewitt/skills157—~1.5kAutomated safety check: PassMIT
Frontend Slideszarazhangrui/frontend-slides30k16 repos~7kAutomated safety check: PassMIT
Single-File HTML Slide Decksop7418/guizang-ppt-skill27k1 repos~6.3kAutomated safety check: PassAGPL-3.0

Similar skills

  • Docx4j

    plutext/docx4j

    A skill your agent uses when writing Java code that creates, reads or edits Word (.docx), PowerPoint (.pptx) or Excel (.xlsx) files with docx4j — including generating documents, editing existing…

    2.4k GitHub stars~2.5k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Eval Suite Planner

    microsoft/eval-guide

    Official

    Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description.

    138 GitHub stars~2.3k tokensUpdated 3 mo ago
    Documents & OfficeAuto-check passed
  • Dashboard Builder

    TheCraigHewitt/skills

    Build an interactive HTML dashboard from any data source — a CSV file, a folder of files, a spreadsheet, or pasted numbers.

    157 GitHub stars~1.5k tokensUpdated 4 mo ago
    Documents & OfficeAuto-check passed
  • Frontend Slides

    zarazhangrui/frontend-slides

    Builds animated HTML slide decks that run in the browser with no dependencies, or converts PowerPoint files to the web, starting from visual style previews.

    30k GitHub starsUsed in 16 repos~7k tokens
    Documents & OfficeAuto-check passed
  • Single-File HTML Slide Decks

    op7418/guizang-ppt-skill

    Generates single-file HTML slide decks with horizontal paging, WebGL backgrounds and a presenter view, in an editorial-magazine style or a Swiss-style layout.

    27k GitHub starsUsed in 1 repo~6.3k tokens
    Documents & OfficeAuto-check passed
  • Openkb Deck Editorial

    VectifyAI/OpenKB

    A skill your agent uses when the user asks the openkb chat to make a deck / slide presentation / PPT / slides / 演示稿 / 幻灯片 from their compiled KB content.

    4.8k GitHub starsUsed in 1 repo~2.1k tokens
    Documents & OfficeAuto-check passed

More from platonai/Browser4

All 17 skills in this repo
  • Data Validation

    platonai/Browser4

    Validates data against common and custom rules (required fields, formats, ranges).

    1.2k GitHub stars~896 tokensUpdated 2 days ago
    Auto-check passed
  • Form Filling

    platonai/Browser4

    Automatically fills web forms using provided field data and can optionally submit the form.

    1.2k GitHub stars~1.1k tokensUpdated 2 days ago
    Auto-check passed
  • Web Scraping

    platonai/Browser4

    Extracts data from web pages using browser automation and CSS/JavaScript selectors.

    1.2k GitHub stars~1.1k tokensUpdated 2 days ago
    Auto-check passed
  • Organize Task Files

    platonai/Browser4

    Lists, pairs, deduplicates, and moves task files across the Coworker task state machine (0draft → 6git-pushed).

    1.2k GitHub stars~801 tokensUpdated 2 days ago
    Auto-check passed
  • Task Token Usage

    platonai/Browser4

    Analyzes Claude Code session traces (JSONL) and draft task files to report token usage per task.

    1.2k GitHub stars~701 tokensUpdated 2 days ago
    Auto-check passed
  • Weather

    platonai/Browser4

    Fetches current weather conditions and a 7-day forecast for a requested location using Open-Meteo geocoding and forecast APIs.

    1.2k GitHub stars~1k tokensUpdated 2 days ago
    Auto-check passed

Works with

Questions about Browser4 Web Miner

What does Browser4 Web Miner do?

Groups similar web pages together and produces an interactive HTML report with clusters of related pages, plus Excel spreadsheets for analysis. Browser4 Web Miner is an agent skill from platonai/Browser4. Groups similar web pages together and produces an interactive HTML report with clusters of related pages, plus Excel spreadsheets for analysis.

When should I use Browser4 Web Miner?

Browser4 Web Miner fits situations like: the user wants to cluster downloaded HTML files; convert detail web pages into interactive views; analyze a folder of web pages locally.

How do I install Browser4 Web Miner in Claude Code?

Run `npx skills add platonai/Browser4 --skill browser4-web-miner -a claude-code`. Or copy the skill folder (skills/browser4-web-miner in platonai/Browser4) into .claude/skills/browser4-web-miner in your project. Claude Code loads it when a task matches its description.

How do I install Browser4 Web Miner in Codex?

Run `npx skills add platonai/Browser4 --skill browser4-web-miner -a codex`. Or copy the skill folder (skills/browser4-web-miner in platonai/Browser4) into .agents/skills/browser4-web-miner in your project. Codex loads it when a task matches its description.

Can I use Browser4 Web Miner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add platonai/Browser4 --skill browser4-web-miner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/browser4-web-miner, .gemini/skills/browser4-web-miner, .github/skills/browser4-web-miner and .opencode/skills/browser4-web-miner in your project.

What does Browser4 Web Miner need to run?

Going by SKILL.md and its folder, Browser4 Web Miner needs the command-line tools its instructions call (java).

Does Browser4 Web Miner access the network?

SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Browser4 Web Miner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Browser4 Web Miner use?

Browser4 Web Miner is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Browser4 Web Miner use?

About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Browser4 Web Miner?

Skills that share tags, products or a category with Browser4 Web Miner: Docx4j (plutext/docx4j, 2.4k stars), Eval Suite Planner (microsoft/eval-guide, 138 stars), Dashboard Builder (TheCraigHewitt/skills, 157 stars) and Frontend Slides (zarazhangrui/frontend-slides, 30k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Browser4 Web Miner?

platonai (a GitHub user) maintains it in platonai/Browser4, which has 1,152 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on October 7, 2026.

Source: platonai/Browser4 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.