Docx4j
plutext/docx4j
A skill your agent uses when writing Java code that creates, reads or edits Word (.docx), PowerPoint (.pptx) or Excel (.xlsx) files with docx4j — including generating documents, editing existing…
Groups similar web pages together and produces an interactive HTML report with clusters of related pages, plus Excel spreadsheets for analysis.
$ npx skills add platonai/Browser4 --skill browser4-web-miner -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install platonai/Browser4 browser4-web-miner --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/platonai/Browser4.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/browser4-web-miner .claude/skills/browser4-web-miner && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "browser4-web-miner" agent skill from https://github.com/platonai/Browser4/tree/main/skills/browser4-web-miner into .claude/skills/browser4-web-miner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser4-web-miner", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/platonai/Browser4/tree/main/skills/browser4-web-minerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add platonai/Browser4 --skill browser4-web-miner -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install platonai/Browser4 browser4-web-miner --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/platonai/Browser4.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/browser4-web-miner .agents/skills/browser4-web-miner && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "browser4-web-miner" agent skill from https://github.com/platonai/Browser4/tree/main/skills/browser4-web-miner into .agents/skills/browser4-web-miner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser4-web-miner", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add platonai/Browser4 --skill browser4-web-miner -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install platonai/Browser4 browser4-web-miner --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/platonai/Browser4.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/browser4-web-miner .cursor/skills/browser4-web-miner && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "browser4-web-miner" agent skill from https://github.com/platonai/Browser4/tree/main/skills/browser4-web-miner into .cursor/skills/browser4-web-miner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser4-web-miner", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/platonai/Browser4.git --path skills/browser4-web-miner--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add platonai/Browser4 --skill browser4-web-miner -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install platonai/Browser4 browser4-web-miner --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/platonai/Browser4.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/browser4-web-miner .gemini/skills/browser4-web-miner && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "browser4-web-miner" agent skill from https://github.com/platonai/Browser4/tree/main/skills/browser4-web-miner into .gemini/skills/browser4-web-miner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser4-web-miner", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install platonai/Browser4 browser4-web-minerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add platonai/Browser4 --skill browser4-web-miner -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/platonai/Browser4.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/browser4-web-miner .github/skills/browser4-web-miner && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "browser4-web-miner" agent skill from https://github.com/platonai/Browser4/tree/main/skills/browser4-web-miner into .github/skills/browser4-web-miner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser4-web-miner", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add platonai/Browser4 --skill browser4-web-miner -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install platonai/Browser4 browser4-web-miner --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/platonai/Browser4.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/browser4-web-miner .opencode/skills/browser4-web-miner && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "browser4-web-miner" agent skill from https://github.com/platonai/Browser4/tree/main/skills/browser4-web-miner into .opencode/skills/browser4-web-miner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browser4-web-miner", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
browser4-web-minerGroups similar web pages together and produces an interactive HTML report with clusters of related pages, plus Excel spreadsheets for analysis.
Browser4 Web Miner is an agent skill from platonai/Browser4. Groups similar web pages together and produces an interactive HTML report with clusters of related pages, plus Excel spreadsheets for analysis. Use when the user wants to cluster downloaded HTML files, convert detail web pages into interactive views, or analyze a folder of web pages locally.
Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Documents & Office, covering HTML artifacts and Excel spreadsheets. It works with Java. The repository describes itself as: Browser4 — an AI-native browser engine for autonomous agents, intelligent extraction, and large-scale web automation. The licence is Apache-2.0.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 0fdba82. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
javaFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Browser4 Web Miner loads about 2.5k tokens when it runs. Until then it costs about 78 tokens; SKILL.md has 963 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from platonai/Browser4 at commit 0fdba82, republished under its Apache-2.0 licence (© platonai). 963 words, ~2,467 tokens.
.claude/skills/browser4-web-miner/SKILL.md (or your agent's skills folder).browser4-cli webminer install # one-time install (Java 17+ auto-detected)
browser4-cli webminer all <html-dir> # full pipeline: encode → cluster → viewsWebMiner groups similar web pages together and produces an interactive HTML report with clusters of related pages — plus Excel spreadsheets for further analysis. Give it a folder of downloaded HTML files, and it handles the rest. Everything runs locally; no data leaves your machine.
Use WebMiner when you have a folder of downloaded HTML pages and want to cluster them into interactive views and Excel reports — fully local, no LLM tokens. It complements rather than replaces browser4-cli crawl/swarm (which acquire pages): WebMiner analyzes pages you already have. Not for single-page extraction — use htmlsnapshot for that.
WebMiner runs a three-stage local pipeline: encode converts each HTML page into a 69-dimension feature vector, cluster groups similar pages with SMILE KMeans (k auto-detected), and views renders an interactive HTML report plus Excel spreadsheets. Everything runs locally on your machine — no data leaves it, and no LLM tokens are consumed.
browser4-cli webminer all <html-dir>browser4-cli webminer views <result-dir>browser4-cli webminer run-example| Flag | Applies to | Description |
|---|---|---|
--max-files <n> | webminer all | Limit the number of HTML files processed (default 40) |
--output <dir> | webminer all | Override the output directory |
--resume [<project-id>] | webminer all | Resume a previous run |
| Symptom | Cause | Fix |
|---|---|---|
webminer install fails | No Java 17+ on PATH | Install JDK 17+ or point JAVA_HOME at it |
webminer all finds no pages | Directory has no .html files | Check the input directory path and file extensions |
| Pipeline crashes on large corpora | Free tier limit (< 1,000 pages) | Reduce the corpus or use --max-files; see the commercial Spark tier for scale |
| Views land in an unexpected temp dir | The views stage uses the app task-output root | Use webminer views <result-dir> to rebuild beside the result dir |
WebMiner is a first-class Browser4 citizen: the browser4-cli webminer
command installs, updates, and runs the tool natively (no PowerShell needed —
the CLI locates a Java 17+ installation, preferring the JRE bundled with the
Browser4 runtime, and launches scent-miner.jar directly). The JAR and its
release metadata are installed to ~/.scent/webminer/.
browser4-cli webminer install # Download and install the latest release
browser4-cli webminer update # Check for and install the latest release
browser4-cli webminer version # Show installed and latest available versions
browser4-cli webminer uninstall # Remove the installed release
browser4-cli webminer run-example # Sample dataset + full pipeline (needs 7-Zip)
browser4-cli webminer all <html-dir> # Full pipeline (encode → cluster → views)
browser4-cli webminer views <result-dir> # Rebuild views from an existing runwebminer all <dir> accepts the pipeline options directly
(--max-files <n>, --output <dir>, --resume [<project-id>]).scent-miner.jar, e.g.
browser4-cli webminer encode <dir>.-Dapp.name=webminer, so the views
task-output root is %TEMP%\webminer-<user>\ml\tasks\... (<user> is the OS
user name; see Output).webminer panel and webminer version keep the update check quiet:
the GitHub → OSS-mirror fallback notices (rate limit, HTTP status, unreachable)
are suppressed, and the Published line is omitted entirely when the release
carries no published_at. webminer install / update still report the
fallback.browser4-cli webminer install downloads, verifies, and installs the latest
release (GitHub Releases with an Aliyun OSS mirror fallback; works on
Windows, Linux, and macOS — no PowerShell needed):
browser4-cli webminer install # Download and install the latest release
browser4-cli webminer update # Check for and install the latest release
browser4-cli webminer version # Show installed and latest available versions
browser4-cli webminer uninstall # Remove the installed releaseReleases are installed to ~/.scent/webminer/ and checked against
https://github.com/platonai/web-miner/releases. SHA-256 checksums are
verified automatically on download.
You can also use the JAR directly if it's already available:
java -jar scent-miner.jar <command> <args>The run-example command downloads a pre-uploaded test dataset of real web
pages, extracts it, and runs the full pipeline — no manual setup required
beyond Java 17 and 7-Zip:
browser4-cli webminer run-exampleThe dataset is cached at ~/.scent/test-data/amazon.com/ so subsequent runs
skip the download.
# Full pipeline (one-shot)
browser4-cli webminer all /path/to/html/files
# Or with the JAR directly
java -jar scent-miner.jar all /path/to/html/filesThe cluster count is always auto-detected from the data — this produces better results than guessing a number.
| Flag | Default | Purpose |
|---|---|---|
--max-files <n> | 40 | Maximum number of HTML files to process |
--output <dir> | <html-dir>-ml-output | Where to write the clustered results (CSV + clustering info; the views stage uses the app temp root — see Output) |
--resume [<project-id>] | — | Pick up where a previous run left off. If no project ID is given, the most recent project is used. |
If clustering has already completed and you just need to (re)build the views:
java -jar scent-miner.jar views <html-dir>-ml-output/kmeans-result/p<timestamp>all produces two kinds of artifacts in two different places:
<html-dir>-ml-output/kmeans-result/p<timestamp>/
(or wherever --output points): one result.csv per feature view
(predictionAnd{Final,Minimal,Original}Features/result.csv) plus
clusteringInfo.txt.views stage of
all writes them to the application's temp task-output root, NOT under
<html-dir>-ml-output:
%TEMP%\<app>-<user>\ml\tasks\unsupervised\result\p<timestamp>\predictionAndMinimalFeatures.views\
on Windows, and <java.io.tmpdir>/<app>-<user>/ml/tasks/unsupervised/result/p<timestamp>/predictionAndMinimalFeatures.views/
on Linux/macOS (/tmp/... on Linux, $TMPDIR on macOS) — the <app> prefix
follows -Dapp.name (webminer when launched through browser4-cli webminer,
pulsar for a direct java -jar run) and <user> is the OS user name. The
end of the run prints the resolved absolute views path.So after java -jar scent-miner.jar all ./html-pages/ the clustered results
look like:
html-pages-ml-output/
└── kmeans-result/
└── p<timestamp>/
├── predictionAndFinalFeatures/result.csv
├── predictionAndMinimalFeatures/result.csv
├── predictionAndOriginalFeatures/result.csv
└── clusteringInfo.txtand the views (<project>.html, *.xlsx, *.json) live in the temp
task-output directory printed by the run.
The real report is
<project>.html(e.g.p<timestamp>.html), notindex.html. Theindex.htmlinside the views directory is an auto-generated directory listing ("Index of predictionAndMinimalFeatures.views") — opening it shows a file list, not the interactive clustering report. Open<project>.htmlinstead.
To place the views beside the clustered results (e.g. to archive them with the project), rebuild them from the result directory:
browser4-cli webminer views <html-dir>-ml-output/kmeans-result/p<timestamp>
# (equivalent to: java -jar scent-miner.jar views <html-dir>-ml-output/kmeans-result/p<timestamp>)This writes predictionAndMinimalFeatures.views/ inside the given result
directory — the recommended way to locate artifacts, since the output path is
explicit instead of an opaque temp path. Open the generated <project>.html
in a browser to explore the clustering results. The .xlsx files can be
opened in Excel for sorting, filtering, or further analysis.
*.html and *.htm files are processed. Other files
in the directory are ignored.--resume to continue from the last completed stage instead of starting over.java is on your PATH.© platonai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/browser4-web-miner of platonai/Browser4.
Open the folder on GitHubat commit 0fdba82
Browser4 Web Miner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Browser4 Web Miner this skillplatonai/Browser4 | 1.2k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Docx4jplutext/docx4j | 2.4k | — | ~2.5k | Automated safety check: Pass | None | |
| Eval Suite Plannermicrosoft/eval-guide | 138 | — | ~2.3k | Automated safety check: Pass | MIT | |
| Dashboard BuilderTheCraigHewitt/skills | 157 | — | ~1.5k | Automated safety check: Pass | MIT | |
| Frontend Slideszarazhangrui/frontend-slides | 30k | 16 repos | ~7k | Automated safety check: Pass | MIT | |
| Single-File HTML Slide Decksop7418/guizang-ppt-skill | 27k | 1 repos | ~6.3k | Automated safety check: Pass | AGPL-3.0 |
plutext/docx4j
A skill your agent uses when writing Java code that creates, reads or edits Word (.docx), PowerPoint (.pptx) or Excel (.xlsx) files with docx4j — including generating documents, editing existing…
microsoft/eval-guide
Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description.
TheCraigHewitt/skills
Build an interactive HTML dashboard from any data source — a CSV file, a folder of files, a spreadsheet, or pasted numbers.
zarazhangrui/frontend-slides
Builds animated HTML slide decks that run in the browser with no dependencies, or converts PowerPoint files to the web, starting from visual style previews.
op7418/guizang-ppt-skill
Generates single-file HTML slide decks with horizontal paging, WebGL backgrounds and a presenter view, in an editorial-magazine style or a Swiss-style layout.
VectifyAI/OpenKB
A skill your agent uses when the user asks the openkb chat to make a deck / slide presentation / PPT / slides / 演示稿 / 幻灯片 from their compiled KB content.
platonai/Browser4
Validates data against common and custom rules (required fields, formats, ranges).
platonai/Browser4
Automatically fills web forms using provided field data and can optionally submit the form.
platonai/Browser4
Extracts data from web pages using browser automation and CSS/JavaScript selectors.
platonai/Browser4
Lists, pairs, deduplicates, and moves task files across the Coworker task state machine (0draft → 6git-pushed).
platonai/Browser4
Analyzes Claude Code session traces (JSONL) and draft task files to report token usage per task.
platonai/Browser4
Fetches current weather conditions and a 7-day forecast for a requested location using Open-Meteo geocoding and forecast APIs.
Works with
Categories
Groups similar web pages together and produces an interactive HTML report with clusters of related pages, plus Excel spreadsheets for analysis. Browser4 Web Miner is an agent skill from platonai/Browser4. Groups similar web pages together and produces an interactive HTML report with clusters of related pages, plus Excel spreadsheets for analysis.
Browser4 Web Miner fits situations like: the user wants to cluster downloaded HTML files; convert detail web pages into interactive views; analyze a folder of web pages locally.
Run `npx skills add platonai/Browser4 --skill browser4-web-miner -a claude-code`. Or copy the skill folder (skills/browser4-web-miner in platonai/Browser4) into .claude/skills/browser4-web-miner in your project. Claude Code loads it when a task matches its description.
Run `npx skills add platonai/Browser4 --skill browser4-web-miner -a codex`. Or copy the skill folder (skills/browser4-web-miner in platonai/Browser4) into .agents/skills/browser4-web-miner in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add platonai/Browser4 --skill browser4-web-miner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/browser4-web-miner, .gemini/skills/browser4-web-miner, .github/skills/browser4-web-miner and .opencode/skills/browser4-web-miner in your project.
Going by SKILL.md and its folder, Browser4 Web Miner needs the command-line tools its instructions call (java).
SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Browser4 Web Miner is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Browser4 Web Miner: Docx4j (plutext/docx4j, 2.4k stars), Eval Suite Planner (microsoft/eval-guide, 138 stars), Dashboard Builder (TheCraigHewitt/skills, 157 stars) and Frontend Slides (zarazhangrui/frontend-slides, 30k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
platonai (a GitHub user) maintains it in platonai/Browser4, which has 1,152 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on October 7, 2026.
Source: platonai/Browser4 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.