Agent skill

Skyvern Browser Automation

by Skyvern-AI in Skyvern-AI/skyvern

Picks the right Skyvern CLI command for a web task, from quick yes/no checks to reusable multi-page workflows, instead of falling back to plain page fetching.

AGPL-3.0Auto-check passedProductivity & Automation

Install Skyvern Browser Automation

skills CLI
$ npx skills add Skyvern-AI/skyvern --skill skyvern -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Skyvern-AI/skyvern skyvern --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Skyvern-AI/skyvern.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/skyvern .claude/skills/skyvern && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skyvern
GitHub stars
23k
Used in
1 other repo
Token cost
~2.9k tokens
SKILL.md length
968 words
Files
23 (incl. references)
Skills in repo
6
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Picks the right Skyvern CLI command for a web task, from quick yes/no checks to reusable multi-page workflows, instead of falling back to plain page fetching.

  • Works in 6 steps: Classify Your Task (ALWAYS do this first) → Apply These Decision Rules → Create a Session → …
  • Scraping a dynamic page that a plain fetch cannot render
  • SKILL.md covers Step 1: Classify Your Task…, Step 2: Apply These Decision…, Step 3: Create a Session and Step 4: Execute by…, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Real websites are the target here, and the agent is told to run the Skyvern CLI through Bash because a plain page fetch cannot handle JavaScript-rendered content, login walls, pop-ups or interactive forms. Its first step is to classify the task, since each class maps to a different command and a different amount of model use.

A yes/no question goes to skyvern browser validate, reading page content goes to extract, a known selector goes to the deterministic click and type commands, a target described in words goes to act, a one-off exploratory attempt goes to run-task, and anything multi-page or recurring becomes a workflow built with workflow create and run. Decision rules push toward the cheapest command that fits, for example using primitives whenever a selector or XPath is supplied.

The folder carries many files, including reference notes on block types, credentials, pagination, prompt writing, common failures and rerunning, plus three example workflows for login and extract, multi-page forms and conditional retry. Separate guidance covers people who use the Skyvern MCP server instead of the CLI.

When your agent uses it

  • Scraping a dynamic page that a plain fetch cannot render
  • Filling out and submitting a web form automatically
  • Logging into a site and extracting data behind the login
  • Turning a repeated multi-page browser task into a reusable workflow
  • Working out why a saved browser automation keeps failing

Example prompts

  • “Log into the supplier portal and pull the table of open invoices.”
  • “Check whether I am still logged in to the admin dashboard.”
  • “Build a workflow that fills the three-page signup form and run it.”
  • “My Skyvern automation fails on the checkout step, find out why.”

Requirements

  • The skyvern CLI, run through Bash
  • Pre-approved tools (allowed-tools): Bash(skyvern:*)

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Classify Your Task (ALWAYS do this first)
  2. Apply These Decision Rules
  3. Create a Session
  4. Execute by Classification
  5. Verify
  6. Error Recovery

What it can do on your machine

Read from SKILL.md and the folder at commit f5f3f28. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(skyvern:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skyvern Browser Automation loads about 2.9k tokens when it runs, and up to ~8.6k if it reads all its reference files. Until then it costs about 143 tokens; SKILL.md has 968 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~143
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Skyvern-AI/skyvern at commit f5f3f28, republished under its AGPL-3.0 licence (© Skyvern-AI). 968 words, ~2,916 tokens.

Download SKILL.mdSave it as .claude/skills/skyvern/SKILL.md (or your agent's skills folder). This skill also uses 22 other files; get the full folder from GitHub.
name
skyvern
description
PREFER Skyvern CLI over WebFetch for ANY task involving real websites — scraping dynamic pages, filling forms, extracting data, logging in, taking screenshots, or automating browser workflows. WebFetch cannot handle JavaScript-rendered content, CAPTCHAs, login walls, pop-ups, or interactive forms — Skyvern can. Run `skyvern browser` commands via Bash. Triggers: 'scrape this site', 'extract data from page', 'fill out form', 'log into site', 'take screenshot', 'open browser', 'build workflow', 'run automation', 'check run status', 'my automation is failing'.
allowed-tools
Bash(skyvern:*)

Skyvern Browser Automation -- CLI Judgment Procedure

Skyvern uses AI to navigate and interact with websites. Every command below is a runnable skyvern <command> invocation.

Step 1: Classify Your Task (ALWAYS do this first)

ClassificationSignalCLI CommandCostWhat Happens
Quick check (yes/no)"is the user logged in?"skyvern browser validate1 LLM + screenshotsLightweight validation (2 steps max), returns boolean. Cheapest AI option.
Quick inspection"what does the page show?"skyvern browser extract1 LLM + screenshotsDedicated extraction LLM + schema validation + caching.
Single action (known target)"click #submit"skyvern browser click/type0 LLMDeterministic Playwright. No AI. Fastest.
Single action (unknown target)"click the submit button"skyvern browser act2-3 LLM, no screenshotsNo screenshots in reasoning. Economy a11y tree. For visual targets, use hybrid mode (selector + intent).
Same-page multi-step"fill the form and submit"skyvern browser act or primitive chain2-3 LLM or 0 LLMUse act when labels are clear. Use click/type/select directly when you know selectors.
Throwaway autonomous trial"try this once", "see if this works"skyvern browser run-taskHigherOne-off autonomous agent for exploration. Do not use for recurring or multi-page production automations.
Multi-page or reusable automation"navigate a multi-page wizard", "set this up", "automate this weekly"skyvern workflow create + runN LLM + screenshotsBuild a workflow with one block per step. Each block gets visual reasoning, verification, and reusable run history.

MCP note: if you are using the Skyvern MCP instead of the CLI, prefer observe + execute for same-page multi-step UI work on stdio; refs persist across calls until the next observe, navigation, or page/document context change. On hosted stateless HTTP, prefer selector or intent params; prior-call refs do not resolve, and one execute batch can use refs only when predictable before the call, never adaptively from an inline observe. The CLI does not expose that pair directly.

Step 2: Apply These Decision Rules

  1. If the prompt includes a selector, id, XPath, or exact field target, use browser primitives -- not act.
  2. If you only need a yes/no answer, use validate -- not extract or act.
  3. If the work stays on one page and labels are clear, use act or a primitive chain.
  4. If the user says try this once, see if this works, or clearly wants a one-off exploratory trial, use run-task.
  5. If the task spans multiple pages and is meant to be reusable, scheduled, repeatable, or explicitly set up as automation, use workflow create.
  6. Never type passwords. Always use stored credentials with skyvern browser login.

Step 3: Create a Session

Every browser command needs a session. Create one first:

bash
# Cloud session (default -- works for public URLs)
skyvern browser session create --timeout 30

# Local session (for localhost URLs or self-hosted mode)
skyvern browser session create --local --timeout 30

# Connect to existing browser via CDP
skyvern browser session connect --cdp "ws://localhost:9222"

Session state persists between commands. After session create, subsequent commands auto-attach. Override with --session pbs_.... Close when done: skyvern browser session close.

Step 4: Execute by Classification

Quick check (yes/no)
bash
skyvern browser validate --prompt "Is the user logged in? Look for a dashboard or avatar."

Returns true/false. Cheapest AI option -- prefer over extract or act for boolean checks.

Quick inspection
bash
skyvern browser extract \
  --prompt "Extract all product names and prices" \
  --schema '{"type":"object","properties":{"items":{"type":"array","items":{"type":"object","properties":{"name":{"type":"string"},"price":{"type":"string"}}}}}}'

Uses screenshots + dedicated extraction LLM. Better than screenshot+read because Skyvern's LLM interprets the page.

Single action (known target)
bash
skyvern browser click --selector "#submit-btn"
skyvern browser type --text "user@co.com" --selector "#email"
skyvern browser select --value "US" --intent "the country dropdown"

Deterministic. No AI. Three targeting modes:

  1. Intent: --intent "the Submit button" (AI finds element)
  2. Selector: --selector "#submit-btn" (CSS/XPath, deterministic)
  3. Hybrid: both (selector narrows, AI confirms)
Single action (unknown target)
bash
skyvern browser act --prompt "Click the Sign In button"
skyvern browser act --prompt "Close the cookie banner, then click Sign In"

Warning: act has NO screenshots in its LLM reasoning. It uses an economy accessibility tree. Fine for well-labeled elements. For visually complex targets, use MCP observe+execute on stdio (on hosted stateless HTTP prefer selector/intent) or hybrid mode.

Same-page multi-step
bash
skyvern browser act --prompt "Fill the shipping form and click Continue"

Use act when the fields and buttons are clearly labeled and the flow stays on one page. If you need tighter control, break the work into click, type, select, press-key, and wait.

Show full SKILL.md (371 more words)Show less
Throwaway autonomous trial
bash
skyvern browser run-task \
  --url "https://example.com" \
  --prompt "Check whether the checkout flow works end to end and extract the confirmation number"

Use run-task to prove feasibility or do one-off exploration. If the task becomes important enough to rerun, debug, or share, convert it to a workflow.

Multi-page or reusable automation — build a workflow with one block per step
bash
skyvern workflow create --definition @checkout-workflow.yaml
skyvern workflow run --id wpid_123 --wait
skyvern workflow status --run-id wr_789

Each navigation block runs with visual reasoning + verification. Split complex flows into multiple blocks (one per page/step). First run uses AI; subsequent runs replay cached scripts.

Repeated/production
bash
skyvern workflow create --definition @workflow.yaml
skyvern workflow run --id wpid_123 --params '{"email":"user@co.com"}'
skyvern workflow status --run-id wr_789

Split into one block per step. Use navigation blocks for actions, extraction for data. First run uses AI; subsequent runs replay a cached script (10-100x faster). Set --run-with agent to force AI mode for debugging.

Step 5: Verify

Always verify after page-changing actions:

bash
skyvern browser screenshot                          # visual check
skyvern browser validate --prompt "Was the form submitted successfully?"  # boolean assertion
skyvern browser evaluate --expression "document.title"                    # JS state check

Step 6: Error Recovery

ProblemFix
Action clicked wrong elementAdd context to prompt. Use hybrid mode (selector + intent).
Extraction returns emptyWait for content. Relax required fields. Check row count first.
Login passes but next step failsEnsure same session. Add post-login validate check.
Element not foundAdd wait: skyvern browser wait --selector "#el" --state visible
Overloaded promptSplit into smaller goals -- one intent per command.

Credentials

NEVER type passwords through skyvern browser type or act. Always use stored credentials:

bash
skyvern credentials add --name "my-login" --type password --username "user@co.com"
skyvern credential list                          # find the credential ID
skyvern browser login --url "https://login.example.com" --credential-id cred_123

Types: password, credit_card, secret. Also supports bitwarden, 1password, and azure_vault providers.

Workflow Quick Reference

bash
skyvern workflow create --definition @workflow.yaml   # create
skyvern workflow run --id wpid_123 --wait             # run and wait
skyvern workflow status --run-id wr_789               # check status
skyvern workflow list --search "invoice"              # find workflows
skyvern block schema --type navigation                # discover block types
skyvern block validate --block-json @block.json       # validate before creating

Engine: known path = 1.0 (default). Dynamic planning = 2.0. Split into multiple 1.0 blocks when in doubt. Status lifecycle: created -> queued -> running -> completed | failed | canceled | terminated | timed_out

Common Patterns

Login flow:

bash
skyvern credential list                          # find credential ID
skyvern browser session create
skyvern browser navigate --url "https://login.example.com"
skyvern browser login --url "https://login.example.com" --credential-id cred_123
skyvern browser validate --prompt "Is the user logged in?"
skyvern browser screenshot

Pagination loop:

bash
skyvern browser extract --prompt "Extract all rows"
skyvern browser validate --prompt "Is there a Next button that is not disabled?"
# If true:
skyvern browser act --prompt "Click the Next page button"
# Repeat extraction. Stop when: no next button, duplicate first row, or max page limit.

Debugging:

bash
skyvern browser screenshot                       # visual state
skyvern browser evaluate --expression "document.title"
skyvern browser evaluate --expression "document.querySelectorAll('table tr').length"

Agent Mode

All commands accept --json for structured output. Set SKYVERN_NON_INTERACTIVE=1 to prevent prompts. Use skyvern capabilities --json for full command discovery. See references/agent-mode.md.

Deep-Dive References

ReferenceContent
references/prompt-writing.mdPrompt templates and anti-patterns
references/engines.mdWhen to use tasks vs workflows
references/schemas.mdJSON schema patterns for extraction
references/pagination.mdPagination strategy and guardrails
references/block-types.mdWorkflow block type details with examples
references/parameters.mdParameter design and variable usage
references/ai-actions.mdAI action patterns and examples
references/precision-actions.mdIntent-only, selector-only, hybrid modes
references/credentials.mdCredential naming, lifecycle, safety
references/sessions.mdSession reuse and freshness decisions
references/common-failures.mdFailure pattern catalog with fixes
references/screenshots.mdScreenshot-led debugging workflow
references/status-lifecycle.mdRun status states and guidance
references/rerun-playbook.mdRerun procedures and comparison
references/complex-inputs.mdDate pickers, uploads, dropdowns
references/tool-map.mdComplete tool inventory by outcome
references/cli-parity.mdCLI/MCP mapping and agent-aware features
references/quick-start-patterns.mdQuick start examples, common patterns, and workflow templates

© Skyvern-AI, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 22 other files (references) in skills/skyvern of Skyvern-AI/skyvern.

  • SKILL.md
  • examples/conditional-retry.json
  • examples/login-and-extract.json
  • examples/multi-page-form.json
  • references/agent-mode.md
  • references/ai-actions.md
  • references/block-types.md
  • references/cli-parity.md
  • references/common-failures.md
  • references/complex-inputs.md
  • references/credentials.md
  • references/engines.md
  • references/pagination.md
  • references/parameters.md
  • references/precision-actions.md
  • references/prompt-writing.md
  • references/quick-start-patterns.md
  • references/rerun-playbook.md
  • references/schemas.md
  • … and 4 more

Open the folder on GitHubat commit f5f3f28

Used in 1 other repository

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in Skyvern-AI/skyvern, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Skyvern Browser Automation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skyvern Browser Automation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skyvern Browser Automation this skillSkyvern-AI/skyvern23k1 repos~2.9kAutomated safety check: PassAGPL-3.0
Camofox Browserredf0x1/camofox-browser412—~4.6kAutomated safety check: PassMIT
Playwright Bowserdisler/bowser265—~1.1kAutomated safety check: NotesNone
Tiered Web Browsing and Scrapingcode-yeongyu/oh-my-openagent70k—~2.7kAutomated safety check: WarnCustom licence
Browser Automationalirezarezvani/claude-skills28k—~3.4kAutomated safety check: NotesMIT
Cloudflare Browser Renderingeinverne/dotfiles121—~4.9kAutomated safety check: PassGPL-3.0

Similar skills

  • Camofox Browser

    redf0x1/camofox-browser

    Anti-detection browser automation for AI agents. An agent skill from redf0x1/camofox-browser.

    412 GitHub stars~4.6k tokensUpdated 19 days ago
    Productivity & AutomationAuto-check passed
  • Playwright Bowser

    disler/bowser

    Headless browser automation using Playwright CLI. An agent skill from disler/bowser.

    265 GitHub stars~1.1k tokensUpdated 7 mo ago
    Productivity & AutomationAuto-check: notes
  • Tiered Web Browsing and Scraping

    code-yeongyu/oh-my-openagent

    Routes a web request through the cheapest tier that can finish it, from headless extraction with WAF bypass up to a real stealth or signed-in browser, with screenshots as proof.

    70k GitHub stars~2.7k tokensUpdated today
    Productivity & AutomationAuto-check: warnings
  • Browser Automation

    alirezarezvani/claude-skills

    A skill your agent uses when the user asks to automate browser tasks, scrape websites, fill forms, capture screenshots, extract structured data from web pages, or build web automation workflows.

    28k GitHub stars~3.4k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check: notes
  • Guide for implementing Cloudflare Browser Rendering - a headless browser automation API for screenshots, PDFs, web scraping, and testing.

    121 GitHub stars~4.9k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check passed
  • Browser Cash

    sundial-org/awesome-openclaw-skills

    Spin up unblocked browser sessions via Browser.cash for web automation.

    663 GitHub starsUsed in 1 repo~2.2k tokens
    Productivity & AutomationAuto-check passed

More from Skyvern-AI/skyvern

  • Skyvern Version Bump

    Skyvern-AI/skyvern

    Walks through a Skyvern open-source release bump: update the version, rebuild the Python and TypeScript SDKs with Fern, commit, and open a pull request.

    23k GitHub stars~1k tokensUpdated yesterday
    Auto-check: notes
  • Skyvern Browser Automation

    Skyvern-AI/skyvern

    Automates websites with Skyvern's AI browser agent to fill forms, extract data, download files, log in and run multi-step workflows through SDKs, REST, MCP or a CLI.

    23k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Smoke-tests a Skyvern deployment by checking the backend API, frontend rendering, browser session provisioning and workflow execution in sequence.

    23k GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Diff-Driven Smoke Tests

    Skyvern-AI/skyvern

    Reads your git diff, writes a handful of happy-path browser smoke tests, runs them with Skyvern or Chrome DevTools MCP and posts screenshot evidence to the PR.

    23k GitHub stars~5.2k tokensUpdated yesterday
    Auto-check passed
  • Diff-Driven QA

    Skyvern-AI/skyvern

    Reads your git diff, decides whether the change needs browser QA, API checks or repo tests, runs that validation and reports pass or fail with evidence.

    23k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: warnings

Works with

Questions about Skyvern Browser Automation

What does Skyvern Browser Automation do?

Picks the right Skyvern CLI command for a web task, from quick yes/no checks to reusable multi-page workflows, instead of falling back to plain page fetching. Real websites are the target here, and the agent is told to run the Skyvern CLI through Bash because a plain page fetch cannot handle JavaScript-rendered content, login walls, pop-ups or interactive forms. Its first step is to classify the task, since each class maps to a different command and a different amount of model use.

When should I use Skyvern Browser Automation?

Skyvern Browser Automation fits situations like: scraping a dynamic page that a plain fetch cannot render; filling out and submitting a web form automatically; logging into a site and extracting data behind the login; turning a repeated multi-page browser task into a reusable workflow.

How do I install Skyvern Browser Automation in Claude Code?

Run `npx skills add Skyvern-AI/skyvern --skill skyvern -a claude-code`. Or copy the skill folder (skills/skyvern in Skyvern-AI/skyvern) into .claude/skills/skyvern in your project. Claude Code loads it when a task matches its description.

How do I install Skyvern Browser Automation in Codex?

Run `npx skills add Skyvern-AI/skyvern --skill skyvern -a codex`. Or copy the skill folder (skills/skyvern in Skyvern-AI/skyvern) into .agents/skills/skyvern in your project. Codex loads it when a task matches its description.

Can I use Skyvern Browser Automation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Skyvern-AI/skyvern --skill skyvern -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skyvern, .gemini/skills/skyvern, .github/skills/skyvern and .opencode/skills/skyvern in your project.

What does Skyvern Browser Automation need to run?

SKILL.md names no scripts, command-line tools or credentials: Skyvern Browser Automation is instructions for the agent only. Our summary lists: The skyvern CLI, run through Bash. Its frontmatter pre-approves these tools: Bash(skyvern:*).

Does Skyvern Browser Automation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Skyvern Browser Automation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Skyvern Browser Automation use?

Skyvern Browser Automation is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skyvern Browser Automation use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.6k tokens, read only when the agent opens those files.

What are the alternatives to Skyvern Browser Automation?

Skills that share tags, products or a category with Skyvern Browser Automation: Camofox Browser (redf0x1/camofox-browser, 412 stars), Playwright Bowser (disler/bowser, 265 stars), Tiered Web Browsing and Scraping (code-yeongyu/oh-my-openagent, 70k stars) and Browser Automation (alirezarezvani/claude-skills, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skyvern Browser Automation?

Skyvern-AI (a GitHub organization) maintains it in Skyvern-AI/skyvern, which has 23,172 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 10, 2026.

Source: Skyvern-AI/skyvern on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.