Agent skill

Page Agent

by Tommy-yw in Tommy-yw/RunbookHermes

Embed alibaba/page-agent into your own web application — a pure-JavaScript in-page GUI agent that ships as a single <script tag or npm package and lets end-users of your site drive the UI with…

MITAuto-check: notesProductivity & Automation

Install Page Agent

skills CLI
$ npx skills add Tommy-yw/RunbookHermes --skill page-agent -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Tommy-yw/RunbookHermes page-agent --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Tommy-yw/RunbookHermes.git skills-src && mkdir -p .claude/skills && cp -r skills-src/optional-skills/web-development/page-agent .claude/skills/page-agent && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
page-agent
GitHub stars
546
Used in
3 other repos
Token cost
~2.3k tokens
SKILL.md length
885 words
Files
1
Skills in repo
38
Repo updated
First seen
Licence
MIT

At a glance

Embed alibaba/page-agent into your own web application — a pure-JavaScript in-page GUI agent that ships as a single <script tag or npm package and lets end-users of your site drive the UI with…

  • Works in 4 steps: Open the page in a browser with devtools… → You should see a floating panel. If not,… → Type a simple instruction matching… → …
  • The user is a web developer who wants to add an AI copilot to their SaaS / admin panel / B2B tool
  • SKILL.md covers When to use this skill, When NOT to use this skill, Prerequisites and Path 1 — 30-second demo via…, plus 6 more sections
  • Calls npm, git and curl; reaches github.com and dashscope.aliyuncs.com; needs LLM_API_KEY

What it does

Page Agent is an agent skill from Tommy-yw/RunbookHermes. Embed alibaba/page-agent into your own web application — a pure-JavaScript in-page GUI agent that ships as a single <script tag or npm package and lets end-users of your site drive the UI with natural language ("click login, fill username as John"). No Python, no headless browser, no extension required. Use this skill when the user is a web developer who wants to add an AI copilot to their SaaS / admin panel / B2B tool, make a legacy web app accessible via natural language, or evaluate page-agent against a local…

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Productivity & Automation, covering Browser automation, LLM inference and serving and Model routing and gateways. It works with npm, OpenAI, Ollama and OpenRouter. The repository describes itself as: Hermes-native AIOps agent for evidence-driven incident response, approval-gated remediation, and runbook learning. The licence is MIT.

When your agent uses it

  • The user is a web developer who wants to add an AI copilot to their SaaS / admin panel / B2B tool
  • Make a legacy web app accessible via natural language
  • Evaluate page-agent against a local (Ollama)
  • Cloud (Qwen / OpenAI / OpenRouter) LLM

Example prompts

  • “click login, fill username as John”
  • “/page-agent”

Requirements

  • Node.js
  • A credential in LLM_API_KEY

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Open the page in a browser with devtools open
  2. You should see a floating panel. If not, check the console for errors (most common: CORS on the LLM endpoint, wrong baseURL, or a bad API…
  3. Type a simple instruction matching something visible on the page ("click the Login link")
  4. Watch the Network tab — you should see a request to your baseURL

What it can do on your machine

Read from SKILL.md and the folder at commit 7fd2b9a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm
    • git
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com
    • dashscope.aliyuncs.com
    • api.openai.com
    • cdn.jsdelivr.net
    • openrouter.ai

    Also links to:

    • alibaba.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • LLM_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Page Agent loads about 2.3k tokens when it runs. Until then it costs about 171 tokens; SKILL.md has 885 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~171
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:111
    Create `.env` in the repo root with an LLM endpoint. Example:
  • NoteMentions a .env fileSKILL.md:145
    **Warning:** your `.env` `LLM_API_KEY` is inlined into the IIFE bundle during dev builds. Don't share the bundle. Don't
  • NoteMentions a .env fileSKILL.md:181
    - **Restart dev server** after editing `.env` in Path 3 — Vite only reads env at startup.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Tommy-yw/RunbookHermes at commit 7fd2b9a, republished under its MIT licence (© Tommy-yw). 885 words, ~2,266 tokens.

Download SKILL.mdSave it as .claude/skills/page-agent/SKILL.md (or your agent's skills folder).
name
page-agent
description
Embed alibaba/page-agent into your own web application — a pure-JavaScript in-page GUI agent that ships as a single <script> tag or npm package and lets end-users of your site drive the UI with natural language ("click login, fill username as John"). No Python, no headless browser, no extension required. Use this skill when the user is a web developer who wants to add an AI copilot to their SaaS / admin panel / B2B tool, make a legacy web app accessible via natural language, or evaluate page-agent against a local (Ollama) or cloud (Qwen / OpenAI / OpenRouter) LLM. NOT for server-side browser automation — point those users to Hermes' built-in browser tool instead.
version
1.0.0
author
Hermes Agent
license
MIT

page-agent

alibaba/page-agent (https://github.com/alibaba/page-agent, 17k+ stars, MIT) is an in-page GUI agent written in TypeScript. It lives inside a webpage, reads the DOM as text (no screenshots, no multi-modal LLM), and executes natural-language instructions like "click the login button, then fill username as John" against the current page. Pure client-side — the host site just includes a script and passes an OpenAI-compatible LLM endpoint.

When to use this skill

Load this skill when a user wants to:

  • Ship an AI copilot inside their own web app (SaaS, admin panel, B2B tool, ERP, CRM) — "users on my dashboard should be able to type 'create invoice for Acme Corp and email it' instead of clicking through five screens"
  • Modernize a legacy web app without rewriting the frontend — page-agent drops on top of existing DOM
  • Add accessibility via natural language — voice / screen-reader users drive the UI by describing what they want
  • Demo or evaluate page-agent against a local (Ollama) or hosted (Qwen, OpenAI, OpenRouter) LLM
  • Build interactive training / product demos — let an AI walk a user through "how to submit an expense report" live in the real UI

When NOT to use this skill

  • User wants Hermes itself to drive a browser → use Hermes' built-in browser tool (Browserbase / Camofox). page-agent is the opposite direction.
  • User wants cross-tab automation without embedding → use Playwright, browser-use, or the page-agent Chrome extension
  • User needs visual grounding / screenshots → page-agent is text-DOM only; use a multimodal browser agent instead

Prerequisites

  • Node 22.13+ or 24+, npm 10+ (docs claim 11+ but 10.9 works fine)
  • An OpenAI-compatible LLM endpoint: Qwen (DashScope), OpenAI, Ollama, OpenRouter, or anything speaking /v1/chat/completions
  • Browser with devtools (for debugging)

Path 1 — 30-second demo via CDN (no install)

Fastest way to see it work. Uses alibaba's free testing LLM proxy — for evaluation only, subject to their terms.

Add to any HTML page (or paste into the devtools console as a bookmarklet):

html
<script src="https://cdn.jsdelivr.net/npm/page-agent@1.8.0/dist/iife/page-agent.demo.js" crossorigin="true"></script>

A panel appears. Type an instruction. Done.

Bookmarklet form (drop into bookmarks bar, click on any page):

javascript
javascript:(function(){var s=document.createElement('script');s.src='https://cdn.jsdelivr.net/npm/page-agent@1.8.0/dist/iife/page-agent.demo.js';document.head.appendChild(s);})();

Path 2 — npm install into your own web app (production use)

Inside an existing web project (React / Vue / Svelte / plain):

bash
npm install page-agent

Wire it up with your own LLM endpoint — never ship the demo CDN to real users:

javascript
import { PageAgent } from 'page-agent'

const agent = new PageAgent({
    model: 'qwen3.5-plus',
    baseURL: 'https://dashscope.aliyuncs.com/compatible-mode/v1',
    apiKey: process.env.LLM_API_KEY,   // never hardcode
    language: 'en-US',
})

// Show the panel for end users:
agent.panel.show()

// Or drive it programmatically:
await agent.execute('Click submit button, then fill username as John')

Provider examples (any OpenAI-compatible endpoint works):

ProviderbaseURLmodel
Qwen / DashScopehttps://dashscope.aliyuncs.com/compatible-mode/v1qwen3.5-plus
OpenAIhttps://api.openai.com/v1gpt-4o-mini
Ollama (local)http://localhost:11434/v1qwen3:14b
OpenRouterhttps://openrouter.ai/api/v1anthropic/claude-sonnet-4.6

Key config fields (passed to new PageAgent({...})):

  • model, baseURL, apiKey — LLM connection
  • language — UI language (en-US, zh-CN, etc.)
  • Allowlist and data-masking hooks exist for locking down what the agent can touch — see https://alibaba.github.io/page-agent/ for the full option list

Security. Don't put your apiKey in client-side code for a real deployment — proxy LLM calls through your backend and point baseURL at your proxy. The demo CDN exists because alibaba runs that proxy for evaluation.

Path 3 — clone the source repo (contributing, or hacking on it)

Use this when the user wants to modify page-agent itself, test it against arbitrary sites via a local IIFE bundle, or develop the browser extension.

bash
git clone https://github.com/alibaba/page-agent.git
cd page-agent
npm ci              # exact lockfile install (or `npm i` to allow updates)

Create .env in the repo root with an LLM endpoint. Example:

LLM_MODEL_NAME=gpt-4o-mini
LLM_API_KEY=sk-...
LLM_BASE_URL=https://api.openai.com/v1

Ollama flavor:

LLM_BASE_URL=http://localhost:11434/v1
LLM_API_KEY=NA
LLM_MODEL_NAME=qwen3:14b

Common commands:

bash
npm start           # docs/website dev server
npm run build       # build every package
npm run dev:demo    # serve IIFE bundle at http://localhost:5174/page-agent.demo.js
npm run dev:ext     # develop the browser extension (WXT + React)
npm run build:ext   # build the extension

Test on any website using the local IIFE bundle. Add this bookmarklet:

javascript
javascript:(function(){var s=document.createElement('script');s.src=`http://localhost:5174/page-agent.demo.js?t=${Math.random()}`;s.onload=()=>console.log('PageAgent ready!');document.head.appendChild(s);})();

Then: npm run dev:demo, click the bookmarklet on any page, and the local build injects. Auto-rebuilds on save.

Warning: your .env LLM_API_KEY is inlined into the IIFE bundle during dev builds. Don't share the bundle. Don't commit it. Don't paste the URL into Slack. (Verified: grepping the public dev bundle returns the literal values from .env.)

Show full SKILL.md (306 more words)Show less

Repo layout (Path 3)

Monorepo with npm workspaces. Key packages:

PackagePathPurpose
page-agentpackages/page-agent/Main entry with UI panel
@page-agent/corepackages/core/Core agent logic, no UI
@page-agent/mcppackages/mcp/MCP server (beta)
—packages/llms/LLM client
—packages/page-controller/DOM ops + visual feedback
—packages/ui/Panel + i18n
—packages/extension/Chrome/Firefox extension
—packages/website/Docs + landing site

Verifying it works

After Path 1 or Path 2:

  1. Open the page in a browser with devtools open
  2. You should see a floating panel. If not, check the console for errors (most common: CORS on the LLM endpoint, wrong baseURL, or a bad API key)
  3. Type a simple instruction matching something visible on the page ("click the Login link")
  4. Watch the Network tab — you should see a request to your baseURL

After Path 3:

  1. npm run dev:demo prints Accepting connections at http://localhost:5174
  2. curl -I http://localhost:5174/page-agent.demo.js returns HTTP/1.1 200 OK with Content-Type: application/javascript
  3. Click the bookmarklet on any site; panel appears

Pitfalls

  • Demo CDN in production — don't. It's rate-limited, uses alibaba's free proxy, and their terms forbid production use.
  • API key exposure — any key passed to new PageAgent({apiKey: ...}) ships in your JS bundle. Always proxy through your own backend for real deployments.
  • Non-OpenAI-compatible endpoints fail silently or with cryptic errors. If your provider needs native Anthropic/Gemini formatting, use an OpenAI-compatibility proxy (LiteLLM, OpenRouter) in front.
  • CSP blocks — sites with strict Content-Security-Policy may refuse to load the CDN script or disallow inline eval. In that case, self-host from your origin.
  • Restart dev server after editing .env in Path 3 — Vite only reads env at startup.
  • Node version — the repo declares ^22.13.0 || >=24. Node 20 will fail npm ci with engine errors.
  • npm 10 vs 11 — docs say npm 11+; npm 10.9 actually works fine.

Reference

© Tommy-yw, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in optional-skills/web-development/page-agent of Tommy-yw/RunbookHermes.

Open the folder on GitHubat commit 7fd2b9a

Used in 3 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in Tommy-yw/RunbookHermes, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Page Agent next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Page Agent compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Page Agent this skillTommy-yw/RunbookHermes5463 repos~2.3kAutomated safety check: NotesMIT
Claudish UsageMadAppGang/claudish1k—~9kAutomated safety check: PassNone
Configuring Visionoxbshw/watch-skill469—~509Automated safety check: NotesMIT
QuorumDetrol/quorum-cli119—~807Automated safety check: NotesCustom licence
Summarize Anythingswyxio/skills175—~6.3kAutomated safety check: PassMIT
Create SkillHyk260/PureChat5461 repos~823Automated safety check: PassMIT

Similar skills

  • Claudish Usage

    MadAppGang/claudish

    CRITICAL - Guide for using Claudish CLI ONLY through sub-agents to run Claude Code with any AI model (OpenRouter, Gemini, OpenAI, local models).

    1k GitHub stars~9k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed
  • Configuring Vision

    oxbshw/watch-skill

    The user wants to connect an LLM or vision provider, already has an API key, asks "can I use OpenAI/Anthropic/Gemini/OpenRouter", wants local Ollama, or needs different cheap and strong models.

    469 GitHub stars~509 tokensUpdated 25 days ago
    AI & LLM EngineeringAuto-check: notes
  • Quorum

    Detrol/quorum-cli

    Run a structured debate between agent CLIs (claude, codex, agy, grok) and the user's configured API or local models (OpenAI, Anthropic, Google, xAI, OpenRouter, Ollama and more) through the Quorum…

    119 GitHub stars~807 tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check: notes
  • Summarize Anything

    swyxio/skills

    Summarizes arbitrarily long text (1k-1M words) using recursive map-reduce with any LLM backend.

    175 GitHub stars~6.3k tokensUpdated 4 days ago
    Writing & ContentAuto-check passed
  • Create Skill

    Hyk260/PureChat

    Create a new skill in the current repository. An agent skill from Hyk260/PureChat.

    546 GitHub starsUsed in 1 repo~823 tokens
    DevelopmentAuto-check passed
  • New Provider

    finch-xu/cc-router

    用于在 cc-router 仓库新增一个 LLM provider(即在 src-tauri/providers/ 下添加 YAML 描述符并完成配套的同步改动)。当用户说「加 provider」「接入 XX 厂商」「新增订阅源」「provider YAML」「让 cc-router 支持 OpenRouter/Together/Groq/Ollama 之类」时必须触发本…

    272 GitHub stars~1.6k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed

More from Tommy-yw/RunbookHermes

All 38 skills in this repo
  • Fastmcp

    Tommy-yw/RunbookHermes

    Build, test, inspect, install, and deploy MCP servers with FastMCP in Python.

    546 GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Drug Discovery

    Tommy-yw/RunbookHermes

    Pharmaceutical research assistant for drug discovery workflows.

    546 GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Youtube Content

    Tommy-yw/RunbookHermes

    Fetch YouTube video transcripts and transform them into structured content (chapters, summaries, threads, blog posts).

    546 GitHub starsUsed in 1 repo~785 tokens
    Auto-check passed
  • Oss Forensics

    Tommy-yw/RunbookHermes

    Supply chain investigation, evidence recovery, and forensic analysis for GitHub repositories.

    546 GitHub starsUsed in 3 repos~5k tokens
    Auto-check passed
  • P5js

    Tommy-yw/RunbookHermes

    Production pipeline for interactive and generative visual art using p5.js.

    546 GitHub starsUsed in 1 repo~6.8k tokens
    Auto-check passed
  • Touchdesigner MCP

    Tommy-yw/RunbookHermes

    Control a running TouchDesigner instance via twozero MCP — create operators, set parameters, wire connections, execute Python, build real-time visuals.

    546 GitHub starsUsed in 2 repos~3.4k tokens
    Auto-check passed

Questions about Page Agent

What does Page Agent do?

Embed alibaba/page-agent into your own web application — a pure-JavaScript in-page GUI agent that ships as a single <script tag or npm package and lets end-users of your site drive the UI with…. Page Agent is an agent skill from Tommy-yw/RunbookHermes. Embed alibaba/page-agent into your own web application — a pure-JavaScript in-page GUI agent that ships as a single <script tag or npm package and lets end-users of your site drive the UI with natural language ("click login, fill username as John").

When should I use Page Agent?

Page Agent fits situations like: the user is a web developer who wants to add an AI copilot to their SaaS / admin panel / B2B tool; make a legacy web app accessible via natural language; evaluate page-agent against a local (Ollama); cloud (Qwen / OpenAI / OpenRouter) LLM.

How do I install Page Agent in Claude Code?

Run `npx skills add Tommy-yw/RunbookHermes --skill page-agent -a claude-code`. Or copy the skill folder (optional-skills/web-development/page-agent in Tommy-yw/RunbookHermes) into .claude/skills/page-agent in your project. Claude Code loads it when a task matches its description.

How do I install Page Agent in Codex?

Run `npx skills add Tommy-yw/RunbookHermes --skill page-agent -a codex`. Or copy the skill folder (optional-skills/web-development/page-agent in Tommy-yw/RunbookHermes) into .agents/skills/page-agent in your project. Codex loads it when a task matches its description.

Can I use Page Agent in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Tommy-yw/RunbookHermes --skill page-agent -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/page-agent, .gemini/skills/page-agent, .github/skills/page-agent and .opencode/skills/page-agent in your project.

What does Page Agent need to run?

Going by SKILL.md and its folder, Page Agent needs the command-line tools its instructions call (npm, git and curl) and credentials named LLM_API_KEY. Our summary lists: Node.js; A credential in LLM_API_KEY.

Does Page Agent access the network?

SKILL.md names 6 domains. In commands or code: github.com, dashscope.aliyuncs.com, api.openai.com, cdn.jsdelivr.net and openrouter.ai; the agent is likely to contact these when it follows the instructions. As links in the text: alibaba.github.io. This is read from the text; nothing was executed.

Is Page Agent safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Page Agent use?

Page Agent is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Page Agent use?

About 2.3k tokens (SKILL.md is roughly 9.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Page Agent?

Skills that share tags, products or a category with Page Agent: Claudish Usage (MadAppGang/claudish, 1k stars), Configuring Vision (oxbshw/watch-skill, 469 stars), Quorum (Detrol/quorum-cli, 119 stars) and Summarize Anything (swyxio/skills, 175 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Page Agent?

Tommy-yw (a GitHub user) maintains it in Tommy-yw/RunbookHermes, which has 546 GitHub stars. The repository holds 38 skills in this directory. The repository was last updated on May 18, 2026.

Source: Tommy-yw/RunbookHermes on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.