Agent skill

Appium

by blokadaorg in blokadaorg/blokada

A skill your agent uses for dynamic inspection and navigation of the Blokada app through the repo-local Appium machine session.

MPL-2.0Auto-check passedMobile

Install Appium

skills CLI
$ npx skills add blokadaorg/blokada --skill appium -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install blokadaorg/blokada appium --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/blokadaorg/blokada.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/appium .claude/skills/appium && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
appium
GitHub stars
3.3k
Token cost
~3.5k tokens
SKILL.md length
1,593 words
Files
2
Skills in repo
3
Repo updated
First seen
Licence
MPL-2.0

At a glance

A skill your agent uses for dynamic inspection and navigation of the Blokada app through the repo-local Appium machine session.

  • Works in 12 steps: Request elevated access first so device… → Start with a fresh install by default so… → Send session.status to confirm the… → …
  • Dynamic inspection and navigation of the Blokada app through the repo-local Appium machine session
  • SKILL.md covers Start the session, Simulator mode (mocked builds), Drive the session and Default workflow, plus 5 more sections
  • Calls make

What it does

Appium is an agent skill from blokadaorg/blokada. Use for dynamic inspection and navigation of the Blokada app through the repo-local Appium machine session. Trigger when Codex needs to explore the live native UI on a real device, inspect labels or structure, tap through screens, capture screenshots or XML on demand, or verify interactive behavior without writing a static WDIO spec. The current workflow is implemented for iOS, but the skill name stays generic so it can later cover other Appium-backed platforms too.

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in Mobile, covering Mobile testing and debugging. It works with iOS and Android. The repository describes itself as: The official repo for Blokada apps. The licence is MPL-2.0.

When your agent uses it

  • Dynamic inspection and navigation of the Blokada app through the repo-local Appium machine session
  • Codex needs to explore the live native UI on a real device
  • Tap through screens
  • Capture screenshots

Example prompts

  • “/appium”

Workflow steps

12 steps, taken from the first numbered list in SKILL.md.

  1. Request elevated access first so device discovery, install, Appium, and WebDriverAgent can talk to the connected iPhone.
  2. Start with a fresh install by default so the app on device matches the current workspace.
  3. Send session.status to confirm the session and app state.
  4. If the app version/build on device is not clearly tied to the current workspace, stop and reinstall instead of trusting the existing…
  5. Use app.targets when you need to switch between six, family, and iOS Settings without restarting the session.
  6. Use app.activate with args.target, for example six, family, or settings.
  7. Use ui.summary as the fast default inspection command while navigating the current foreground app. It now reports the verified foreground…
  8. Use ui.inspect when you need bounded structure details. By default it includes labels, a tree summary, and a structured visible-element…
  9. Use ui.read to normalize common element attributes such as value, enabled, visible, and switch-like boolean state.
  10. Use ui.focusSearch, ui.search, ui.back, ui.swipe, and ui.scroll for dynamic system-app exploration without creating a temporary static spec.
  11. Use ui.screenshot or ui.source only when you need artifacts.
  12. Always finish with session.shutdown. This should end the active session without forcing a full WDA reset.

What it can do on your machine

Read from SKILL.md and the folder at commit 5279136. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Appium loads about 3.5k tokens when it runs. Until then it costs about 119 tokens; SKILL.md has 1,593 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~119
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from blokadaorg/blokada at commit 5279136, republished under its MPL-2.0 licence (© blokadaorg). 1,593 words, ~3,459 tokens.

Download SKILL.mdSave it as .claude/skills/appium/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
appium
description
Use for dynamic inspection and navigation of the Blokada app through the repo-local Appium machine session. Trigger when Codex needs to explore the live native UI on a real device, inspect labels or structure, tap through screens, capture screenshots or XML on demand, or verify interactive behavior without writing a static WDIO spec. The current workflow is implemented for iOS, but the skill name stays generic so it can later cover other Appium-backed platforms too.

Appium

Use the repo-local machine session only. Do not create temporary WDIO debug specs for ad hoc inspection.

Start the session

Prefer the root make target:

bash
make appium-explore-session IOS_DEVICE_NAME="<device-name>"

Important:

  • In Codex, request elevated access before running the Appium explorer, install targets, or related device-discovery commands. These flows depend on xcrun device services, WebDriverAgent/Appium startup, and real-device interaction that the sandbox cannot reliably access.
  • Do not assume the make target can auto-discover the device inside the sandbox. If you need the connected phone, run the command with elevated access from the start.
  • When invoking repo make targets, keep the command prefix stable as make <target> ... and put any make variable overrides after the target, for example make appium-explore-session APP_INSTALL=0 SHOW_XCODE_LOG=0. Avoid shell-style env prefixes like APP_INSTALL=0 make ..., because they create unnecessary sandbox permission prompt variants.

Notes:

  • This installs the current workspace build by default. Use this for reviews of current changes, startup behavior, cold-start behavior, notification handling, and other recent code changes.
  • Use APP_FLAVOR=family when the primary app under test should be Blokada Family. The default primary flavor is Blokada 6.
  • A session started with the default flavor only freshly installs Blokada 6. app.targets may still list family if Family is already on the device, but that is a reused install unless you also ran a fresh APP_FLAVOR=family session in the same turn.
  • Use APP_INSTALL=0 only when you already installed the current workspace build in the same turn, or when you have high confidence the installed app still matches the code you want to inspect.
  • APP_BUNDLE_ID may override the primary bundle id, but if it does not match the intended flavor, set APP_FLAVOR explicitly so the install target stays correct.
  • IOS_DEVICE_NAME is the normal device selector for the current iOS flow.
  • IOS_UDID is a low-level fallback only.
  • Interactive sessions must preserve the device's current Auto-Correction and Predictive settings. Do not add Settings navigation to the tests for this; keep the fix in the Appium/WDA layer.
  • Real-device sessions should normally reuse an existing WebDriverAgent if one is healthy. End the active session cleanly, but reserve a full device-side WDA kill for explicit recovery such as APPIUM_WDA_HARD_RESET=1.
  • If you do not expect to continue testing immediately, always shut the session down before you step away. Do not leave a device sitting in automation mode with the Automation Running banner visible.
  • Appium does not manage iOS auto-lock for us. Before a long session, set Auto-Lock to Never or the longest available value and keep the device unlocked before starting.
  • The process stays open and accepts JSONL commands on stdin.

Simulator mode (mocked builds)

The same explorer drives an iOS Simulator running a Mocked / FamilyMocked build, not just a physical device. Use this when verifying mocked-scheme work (make run-{six,family}-mocked) where there is no real tunnel and no connected phone. Opt in with IOS_USE_SIM=1; the harness then resolves the per-worktree sim from make -C ios sim-status, uses simulator capabilities (Appium-managed WDA, no signing identity), and skips the devicectl device-discovery the physical path needs. None of the "request elevated access" / connected-iPhone notes above apply in sim mode.

bash
# interactive JSONL session against the per-worktree mocked sim
IOS_USE_SIM=1 make appium-explore-session
# family flavor
IOS_USE_SIM=1 APP_FLAVOR=family make appium-explore-session
# reuse an already-installed/running mocked build, skip rebuild+reinstall
IOS_USE_SIM=1 APP_INSTALL=0 make appium-explore-session

Prerequisites and gotchas:

  • make run-{six,family}-mocked must have created the sim first; appium-explore-session does not provision it.
  • sim-status derives the sim name from SIM_BASE (default iPhone 17). If you created the sim with a non-default base, pass the same SIM_BASE here or the UDID will not resolve, e.g. IOS_USE_SIM=1 SIM_BASE="iPhone 15" make appium-explore-session.
  • Run from the repo root (the target cds into automation/appium/wdio itself); do not run it from inside wdio/.
  • See ios/SIMULATOR.md ("Appium targeting") for the full sim flow and a raw appium --udid ... fallback.

Drive the session

Send one JSON object per line to stdin. Wait for the terminal done or error event before sending the next command.

Request shape:

json
{"id":"1","command":"ui.summary","args":{}}

Response lifecycle:

  • ack
  • optional result
  • terminal done or error

Use these commands:

  • session.status
  • session.shutdown
  • app.targets
  • app.launch
  • app.activate
  • app.terminate
  • app.state
  • ui.summary
  • ui.inspect
  • ui.tree
  • ui.labels
  • ui.read
  • ui.tap
  • ui.type
  • ui.focusSearch
  • ui.search
  • ui.back
  • ui.swipe
  • ui.scroll
  • ui.wait
  • ui.exists
  • ui.attr
  • ui.source
  • ui.screenshot

Default workflow

Use this sequence unless the task requires something else:

  1. Request elevated access first so device discovery, install, Appium, and WebDriverAgent can talk to the connected iPhone.
  2. Start with a fresh install by default so the app on device matches the current workspace.
  3. Send session.status to confirm the session and app state.
  4. If the app version/build on device is not clearly tied to the current workspace, stop and reinstall instead of trusting the existing install. For example: if you start a fresh six session and then switch to family, treat that Family app as untrusted unless you separately ran make appium-explore-session APP_FLAVOR=family.
  5. Use app.targets when you need to switch between six, family, and iOS Settings without restarting the session.
  6. Use app.activate with args.target, for example six, family, or settings.
  7. Use ui.summary as the fast default inspection command while navigating the current foreground app. It now reports the verified foreground target and active app metadata instead of trusting app state alone.
  8. Use ui.inspect when you need bounded structure details. By default it includes labels, a tree summary, and a structured visible-element list.
  9. Use ui.read to normalize common element attributes such as value, enabled, visible, and switch-like boolean state.
  10. Use ui.focusSearch, ui.search, ui.back, ui.swipe, and ui.scroll for dynamic system-app exploration without creating a temporary static spec.
  11. Use ui.screenshot or ui.source only when you need artifacts.
  12. Always finish with session.shutdown. This should end the active session without forcing a full WDA reset.
  13. If you are done for now, do not keep the session open just to preserve reuse. Reuse matters between active test runs, but an idle phone should not be left showing Automation Running.

If the session drops unexpectedly, first suspect device auto-lock or lost foreground automation:

  • unlock the device
  • confirm whether the app is still installed / in foreground
  • restart the explorer session once before concluding the app itself broke

If Appium never reaches JSONL command handling after install and the server log shows repeated /status socket hangups, iProxy ... Unexpected data, or WebDriverAgent xcodebuild code 65, treat that as a harness/WDA problem first:

  • retry the explorer once
  • inspect automation/appium/output/appium-explore-server.log
  • do not treat the app under test as failed unless the same behavior also reproduces after WDA is healthy

When reporting findings from an interactive session, state whether the session used a freshly installed workspace build or a trusted reused install.

Show full SKILL.md (536 more words)Show less

Command guidance

Prefer ui.summary first:

json
{"id":"1","command":"ui.summary","args":{}}

Use ui.inspect for bounded structure:

json
{"id":"2","command":"ui.inspect","args":{"labels":true,"tree":true,"limit":40}}

Use generic navigation and reads for dynamic exploration:

json
{"id":"3","command":"ui.search","args":{"text":"keyboard"}}
{"id":"4","command":"ui.tap","args":{"selector":"~Keyboard"}}
{"id":"5","command":"ui.read","args":{"selector":"~Auto-Correction"}}
{"id":"6","command":"ui.back","args":{}}
{"id":"7","command":"ui.scroll","args":{"direction":"down"}}

Use raw WDIO/Appium selectors directly:

json
{"id":"8","command":"ui.tap","args":{"selector":"~Privacy Pulse"}}
{"id":"9","command":"ui.wait","args":{"selector":"~Avancerat","timeoutMs":10000}}
{"id":"10","command":"ui.attr","args":{"selector":"~automation.power_toggle","name":"value"}}

Switch between the built-in app targets without restarting Appium:

json
{"id":"11","command":"app.targets","args":{}}
{"id":"12","command":"app.activate","args":{"target":"family"}}
{"id":"13","command":"ui.summary","args":{}}
{"id":"14","command":"app.activate","args":{"target":"settings"}}
{"id":"15","command":"app.activate","args":{"target":"six"}}

For dynamic Settings work, prefer discovery over a maintained path catalog:

  • switch to settings
  • use ui.summary and ui.inspect to see the visible hierarchy
  • use ui.search if the screen exposes a search field
  • open rows and read switches dynamically with ui.tap and ui.read

Lightweight hint only: DNS is usually under General, then VPN & Device Management, then DNS, but use the live hierarchy instead of assuming labels or locale.

Practical DNS notes:

  • Treat DNS as a top-level system setting reached from Settings -> General -> the VPN/DNS management area. The exact row labels depend on locale, but it is not an app-specific setting under the Blokada app entry.
  • A failed Settings search for DNS is not enough to conclude the path is unavailable; prefer live hierarchy discovery.
  • If tapping the visible VPN/DNS management row is unreliable, try the accessibility id when present.
  • If the task is to put DNS "back", first verify the current selection on the DNS screen. If Automatic is already selected, leave it unchanged instead of toggling away and back.

Capture artifacts only when needed:

json
{"id":"16","command":"ui.screenshot","args":{"name":"settings-screen"}}
{"id":"17","command":"ui.source","args":{"name":"settings-screen"}}

System alerts (notification / location / etc.)

iOS system permission alerts live in SpringBoard, not the foreground app. ui.tap (selector or coordinate) cannot reach them — the alert is outside WDA's app-scoped accessibility tree, and a coordinate tap lands "underneath" the alert in the app's coordinate space.

Use ui.alert instead, which wraps WDA's mobile: alert extension and operates on the system alert layer:

json
{"id":"1","command":"ui.alert","args":{"action":"getButtons"}}
{"id":"2","command":"ui.alert","args":{"action":"accept","buttonLabel":"Tillåt"}}
{"id":"3","command":"ui.alert","args":{"action":"dismiss"}}
  • action: accept (positive button) | dismiss (negative button) | getButtons (return the button labels for diagnostics)
  • buttonLabel (optional): pick a specific button by visible label. Useful for localized alerts where "Allow" vs "Tillåt" vs "Erlauben" differs by sim locale.

If ui.alert returns "An attempt was made to operate on a modal dialog when one was not open", the alert hasn't materialized yet. Sleep briefly and retry, or send the user-action that triggers it first.

Flutter widget taps (when element.click() is a no-op)

Flutter often exposes interactive widgets as XCUIElementTypeStaticText (no Semantics(button: true)). ui.tap with a selector calls element.click(), which targets the accessibility element but doesn't dispatch the underlying GestureDetector callback. The tap reports success but nothing happens in the app.

Use ui.tapCenter for these — it resolves the element's on-screen rect and dispatches a real touch via mobile: tap:

json
{"id":"1","command":"ui.tapCenter","args":{"selector":"~Aktivera"}}

Quick diagnosis: if ui.tap ~Foo reports tapped but the app's log shows no trace of the expected callback, switch to ui.tapCenter with the same selector.

Cross-worktree sim targeting

By default the harness reads the per-worktree sim from make -C ios sim-status. To drive a sim that was provisioned by a sibling worktree (without reprovisioning a new sim in this worktree), set IOS_UDID (and optionally IOS_DEVICE_NAME):

bash
IOS_USE_SIM=1 \
IOS_UDID=83114086-0934-4944-A01C-7955DB91B0A8 \
IOS_DEVICE_NAME="iPhone 16 - other-worktree" \
APP_FLAVOR=family APP_INSTALL=0 \
make appium-explore-session

This is the path for: edit harness scripts in worktree A, drive the existing app/sim from worktree B without rebuilding.

Cleanup

Always send:

json
{"id":"999","command":"session.shutdown","args":{}}

The intended behavior is that shutdown removes active automation from the phone without discarding reusable WDA state. If the iPhone still shows Automation Running, treat it as a harness cleanup bug and use an explicit hard reset rather than normalizing full WDA teardown after every run.

© blokadaorg, MPL-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .agents/skills/appium of blokadaorg/blokada.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 5279136

Compare with similar skills

Appium next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Appium compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Appium this skillblokadaorg/blokada3.3k—~3.5kAutomated safety check: PassMPL-2.0
Mobilerun Docs Referencedroidrun/mobilerun9.6k—~943Automated safety check: PassMIT
Phoneagentrounak/PhoneAgent798—~2.2kAutomated safety check: PassMIT
Simvynpranshuchittora/simvyn437—~938Automated safety check: PassMIT
Mobile Automation with agent-devicenuclearpasta/react-native-drax714—~1.4kAutomated safety check: PassMIT
Screenmapaleqsio/screenmap239—~7.2kAutomated safety check: PassMIT

Similar skills

  • Mobilerun Docs Reference

    droidrun/mobilerun

    Answers questions about Mobilerun, the LLM-agent framework for automating Android and iOS devices, by pointing the agent to the right page of its v5 documentation.

    9.6k GitHub stars~943 tokensUpdated 2 days ago
    MobileAuto-check passed
  • Phoneagent

    rounak/PhoneAgent

    Control a connected iPhone, iOS simulator, Android emulator, or Android device from macOS through PhoneAgent's JSON-RPC bridge.

    798 GitHub stars~2.2k tokensUpdated 1 mo ago
    MobileAuto-check passed
  • Simvyn

    pranshuchittora/simvyn

    Operate iOS Simulators, Android Emulators, and connected mobile devices with Simvyn.

    437 GitHub stars~938 tokensUpdated 6 days ago
    MobileAuto-check passed
  • Mobile Automation with agent-device

    nuclearpasta/react-native-drax

    Drives iOS and Android devices and simulators from the command line: open apps, snapshot the UI tree, tap, type, scroll, take screenshots and read UI info.

    714 GitHub stars~1.4k tokensUpdated 3 days ago
    MobileAuto-check passed
  • Screenmap

    aleqsio/screenmap

    Generate a visual navigation map of an Expo / React Native or NativeScript app.

    239 GitHub stars~7.2k tokensUpdated 2 days ago
    MobileAuto-check passed
  • Simulator Audio E2E

    hyochan/react-native-nitro-sound

    Build and run repeatable react-native-nitro-sound recorder/player regression tests on an iOS Simulator or Android emulator, with explicit virtual-device selection, microphone permission, Maestro…

    961 GitHub stars~1.1k tokensUpdated 7 days ago
    MobileAuto-check passed

More from blokadaorg/blokada

  • Device Log

    blokadaorg/blokada

    A skill your agent uses for pulling recent Blokada app logs from a connected device, using the same share-log file exposed in Settings.

    3.3k GitHub stars~762 tokensUpdated today
    Auto-check passed
  • Dep Validate

    blokadaorg/blokada

    A skill your agent uses to validate risky dependency bumps end to end as a local or cloud-launched agent.

    3.3k GitHub stars~5.7k tokensUpdated today
    Auto-check: notes

Works with

Categories

Questions about Appium

What does Appium do?

A skill your agent uses for dynamic inspection and navigation of the Blokada app through the repo-local Appium machine session. Appium is an agent skill from blokadaorg/blokada. Use for dynamic inspection and navigation of the Blokada app through the repo-local Appium machine session.

When should I use Appium?

Appium fits situations like: dynamic inspection and navigation of the Blokada app through the repo-local Appium machine session; Codex needs to explore the live native UI on a real device; tap through screens; capture screenshots.

How do I install Appium in Claude Code?

Run `npx skills add blokadaorg/blokada --skill appium -a claude-code`. Or copy the skill folder (.agents/skills/appium in blokadaorg/blokada) into .claude/skills/appium in your project. Claude Code loads it when a task matches its description.

How do I install Appium in Codex?

Run `npx skills add blokadaorg/blokada --skill appium -a codex`. Or copy the skill folder (.agents/skills/appium in blokadaorg/blokada) into .agents/skills/appium in your project. Codex loads it when a task matches its description.

Can I use Appium in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add blokadaorg/blokada --skill appium -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/appium, .gemini/skills/appium, .github/skills/appium and .opencode/skills/appium in your project.

What does Appium need to run?

Going by SKILL.md and its folder, Appium needs the command-line tools its instructions call (make).

Does Appium access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Appium safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Appium use?

Appium is published under the MPL-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Appium use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Appium?

Skills that share tags, products or a category with Appium: Mobilerun Docs Reference (droidrun/mobilerun, 9.6k stars), Phoneagent (rounak/PhoneAgent, 798 stars), Simvyn (pranshuchittora/simvyn, 437 stars) and Mobile Automation with agent-device (nuclearpasta/react-native-drax, 714 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Appium?

blokadaorg (a GitHub organization) maintains it in blokadaorg/blokada, which has 3,265 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 7, 2026.

Source: blokadaorg/blokada on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.