Develop and verify HyperTTS GUI screens against a real, running Anki instance: launch Anki headlessly on a throwaway profile, inject notes with AnkiConnect, read the live Qt widget tree as text…
Install the "anki-gui-automation" agent skill from https://github.com/Vocab-Apps/anki-hyper-tts/tree/main/.claude/skills/anki-gui-automation into .claude/skills/anki-gui-automation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "anki-gui-automation", then confirm the skill loads.
Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Type this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
skills CLI
$ npx skills add Vocab-Apps/anki-hyper-tts --skill anki-gui-automation -a codex
Project install goes to .agents/skills/; add -g for ~/.codex/skills/.
Install the "anki-gui-automation" agent skill from https://github.com/Vocab-Apps/anki-hyper-tts/tree/main/.claude/skills/anki-gui-automation into .agents/skills/anki-gui-automation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "anki-gui-automation", then confirm the skill loads.
Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add Vocab-Apps/anki-hyper-tts --skill anki-gui-automation -a cursor
Project install goes to .agents/skills/; add -g for ~/.cursor/skills/.
Install the "anki-gui-automation" agent skill from https://github.com/Vocab-Apps/anki-hyper-tts/tree/main/.claude/skills/anki-gui-automation into .cursor/skills/anki-gui-automation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "anki-gui-automation", then confirm the skill loads.
Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
skills CLI
$ npx skills add Vocab-Apps/anki-hyper-tts --skill anki-gui-automation -a gemini-cli
Project install goes to .agents/skills/; add -g for ~/.gemini/skills/.
Install the "anki-gui-automation" agent skill from https://github.com/Vocab-Apps/anki-hyper-tts/tree/main/.claude/skills/anki-gui-automation into .gemini/skills/anki-gui-automation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "anki-gui-automation", then confirm the skill loads.
Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
skills CLI
$ npx skills add Vocab-Apps/anki-hyper-tts --skill anki-gui-automation -a github-copilot
Project install goes to .agents/skills/; add -g for ~/.copilot/skills/.
Install the "anki-gui-automation" agent skill from https://github.com/Vocab-Apps/anki-hyper-tts/tree/main/.claude/skills/anki-gui-automation into .github/skills/anki-gui-automation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "anki-gui-automation", then confirm the skill loads.
GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add Vocab-Apps/anki-hyper-tts --skill anki-gui-automation -a opencode
OpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
Install the "anki-gui-automation" agent skill from https://github.com/Vocab-Apps/anki-hyper-tts/tree/main/.claude/skills/anki-gui-automation into .opencode/skills/anki-gui-automation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "anki-gui-automation", then confirm the skill loads.
OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Facts
Skill name
anki-gui-automation
GitHub stars
284
Token cost
~2.8k tokens
SKILL.md length
1,161 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
GPL-3.0
At a glance
Develop and verify HyperTTS GUI screens against a real, running Anki instance: launch Anki headlessly on a throwaway profile, inject notes with AnkiConnect, read the live Qt widget tree as text…
Tasks that involve Study guides and flashcards
SKILL.md covers The loop, Widget naming convention (do…, scripts/gui_automation reference and Verifying a feature end to end, plus 3 more sections
Calls pytest and dnf
Tasks that involve Desktop control
What it does
Anki Gui Automation is an agent skill from Vocab-Apps/anki-hyper-tts. Develop and verify HyperTTS GUI screens against a real, running Anki instance: launch Anki headlessly on a throwaway profile, inject notes with AnkiConnect, read the live Qt widget tree as text, drive dialogs, screenshot them, and tear everything down
Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Education, covering Study guides and flashcards and Desktop control. The repository describes itself as: HyperTTS Addon for Anki. The licence is GPL-3.0.
When your agent uses it
Tasks that involve Study guides and flashcards
Tasks that involve Desktop control
Example prompts
“/anki-gui-automation”
Requirements
Python 3
What it can do on your machine
Read from SKILL.md and the folder at commit 05007b9. It shows what the files ask for, not the result of running them.
Tool permissions
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Runs code
Shell commands in SKILL.md call:
pytest
dnf
From the folder's file list and the shell code blocks in SKILL.md.
Network
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Credentials
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Context cost
Anki Gui Automation loads about 2.8k tokens when it runs. Until then it costs about 68 tokens; SKILL.md has 1,161 words of instructions outside code blocks.
Always· name and description, kept in context so the agent knows when to use it
~68
When it runs· the whole SKILL.md, loaded when a task matches
~2.8k
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
Safety
Auto-check: notes
The automated check noted patterns worth knowing about, such as sudo or a known installer.
NoteRuns commands with sudoSKILL.md:197
System packages (Fedora, installed with `sudo dnf install`; see
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
Download SKILL.mdSave it as .claude/skills/anki-gui-automation/SKILL.md (or your agent's skills folder).
name
anki-gui-automation
description
Develop and verify HyperTTS GUI screens against a real, running Anki instance: launch Anki headlessly on a throwaway profile, inject notes with AnkiConnect, read the live Qt widget tree as text, drive dialogs, screenshot them, and tear everything down
user_invocable
true
Agentic GUI development for HyperTTS
Use this whenever you add or change a HyperTTS dialog. pytest-qt tests prove the logic; this
harness proves the dialog is actually reachable, laid out correctly, and correct against a real
Anki collection (real notes, real update_note, real undo).
Everything runs on a virtual display against an isolated Anki base folder and profile, so the
developer's own Anki profile and collection are never touched. Both service ports are
deliberately non-default (AnkiConnect on 8766, not 8765) so the harness can never talk to a real
Anki session that happens to be open.
You must run scripts/gui_automation/teardown.sh when you are done. Leaving Xvfb, x11vnc,
websockify and a headless Anki running is the main failure mode of this workflow.
The loop
bash
cd scripts/gui_automation
./start_anki.sh # Xvfb + openbox + x11vnc + noVNC + Anki with HyperTTS
./ankiconnect.py seed --reset # deck + note type + 5 notes with assorted sound tags
./ankiconnect.py browse # open the browser on those notes
./gui_probe.py browser-select-all
./gui_probe.py actions --text Audio # what menu entries exist
./gui_probe.py trigger --text 'Remove Audio (Collection)...' # open the dialog
./gui_probe.py windows # confirm it opened, and is modal
./gui_probe.py tree --window hypertts_remove_audio_dialog # read the whole dialog as text
./gui_probe.py table --object-name hypertts_remove_audio_preview_table
./gui_probe.py screenshot --path /tmp/hypertts-gui-automation/artifacts/dialog.png \
--window hypertts_remove_audio_dialog # then Read the png
./teardown.sh # ALWAYS
After editing any HyperTTS python file, run ./start_anki.sh --restart. The add-on is
symlinked into the profile, but Anki only imports it at startup. --fresh additionally throws
away the collection.
Read text first, screenshot second. tree costs a few hundred tokens and tells you state
(enabled, checked, current index, combo items); a screenshot costs thousands but is the only way
to catch layout problems. Both matter — the two bugs found while building the Remove Audio dialog
(a QLabel rendering <b> as literal text, and preview columns truncated by an even stretch)
were invisible in the widget tree and obvious in the screenshot.
Name the dialog too (self.setObjectName('hypertts_remove_audio_dialog')) so --window can
target it. Without object names you have to address widgets by class + index path
(RemoveAudioDialog/QGroupBox[0]/QComboBox[0]), which breaks the moment the layout changes.
hypertts_addon/component_remove_audio.py is the reference implementation.
Older components (component_batch.py, component_voiceselection.py, …) have no object names
yet. For those, address widgets by --class + --text, or by the path printed by tree.
prepares the profile, launches Anki, waits for both ports
setup_profile.py
symlinks HyperTTS + installs AnkiConnect + anki_gui_probe, creates the profile
stop_anki.sh
stops only Anki, leaves the display up
teardown.sh
stops everything, cleans the X lock
status.sh
what is running, which ports answer, which windows are open
screenshot.sh [name]
full-screen grab into the artifacts dir (includes window decorations)
ankiconnect.py
inject/read collection data
gui_probe.py
inspect and drive the live GUI
Runtime state lives under /tmp/hypertts-gui-automation: anki_base/ (throwaway collection and
add-ons), logs/anki.log, logs/hypertts.log (HyperTTS debug logging is on), artifacts/
(screenshots), pids/.
gui_probe.py
Inspect:
windows — every top-level window: class, title, modal, active, geometry
tree [--window W] [--named-only] [--all] [--max-depth N] — indented widget tree
info --object-name X — one widget in detail
table --object-name X — dump a QTableView/QTreeView model as rows (QVariant unwrapped)
actions [--text substring] — every QAction, i.e. every menu entry
undo-status — the label of the next undoable operation
note-fields --note-id N — a note's fields straight from the collection
Drive:
click --object-name X [--no-wait] — also --class QPushButton --text Cancel
set-text --object-name X --value '...'
combo --object-name X --text '...' (or --index N)
check --object-name X [--off]
select-row --object-name X --row N
trigger --text 'Menu entry...' — opens dialogs; does not wait, by design
close --title '...' [--class ...], undo, browser-select-all, browser-search --query ...
screenshot --path P [--window W] — QWidget.grab() of one window
raw '{"action": "eval", "params": {"expression": "mw.pm.name"}}' — escape hatch, main-thread
eval with aqt, qt and mw in scope
Modal dialogs block the Qt main thread. Any action that opens one must not wait for the main
thread, or the request times out. trigger defaults to not waiting; pass --no-wait to click
when the click opens a dialog (including a QMessageBox such as HyperTTS's "Save changes to
current preset ?" confirmation) or closes the window it lives on. When a probe call returns
{"status": "pending"} that is not an error: a modal is open and holding the main thread — call
windows to see it and deal with it.
Anki's own dialogs are reachable the same way, which is how you dismiss things like the add-on
startup error box:
seed [--reset] — creates the deck HyperTTS Automation, the note type
HyperTTS Automation Note (fields Chinese / English / Sound / Sound English), stores dummy
media files, and adds 5 notes: a lone HyperTTS sound tag, text plus a HyperTTS sound tag, a
foreign sound tag (external-recording.mp3, i.e. audio HyperTTS did not generate), HyperTTS
audio in two fields, and a note with no audio. Extend seed_notes() for new cases.
notes [--query Q] — dump note fields, the cheapest way to assert what an operation did
browse [--query Q] — open the browser on a query
invoke <action> [--params JSON] — any AnkiConnect action
Two AnkiConnect gotchas, already handled in the helper but worth knowing:
AnkiConnect is unmaintained and still assigns the deck through the legacy note-type dict, which
modern Anki ignores, so addNotes lands cards in Default. seed moves them with changeDeck.
Anki search syntax wants the quotes around the whole term: "deck:HyperTTS Automation", not
deck:"HyperTTS Automation".
Verifying a feature end to end
The pattern that actually proves a collection-modifying dialog works — this is how the Remove
Audio dialog was validated:
bash
./ankiconnect.py notes # state before
./gui_probe.py undo-status # "" - nothing to undo yet
./gui_probe.py click --object-name hypertts_remove_audio_remove_button
./ankiconnect.py notes # state after: only the intended fields changed
./gui_probe.py undo-status # "HyperTTS: Remove Audio from Notes" <- undo support works
./gui_probe.py undo
./ankiconnect.py notes # back to the original state
Undo support in HyperTTS comes from running the mutation inside
anki_utils.run_in_background_collection_op(parent, update_fn, success_fn, undo_entry_name=...),
which wraps it in a custom undo entry. update_fn receives the collection and must call
collection.update_note(note); never call aqt.mw.col directly from a dialog.
pytest-qt tests are still required
The harness complements the test suite, it does not replace it. Every new dialog needs a
tests/test_component_<name>.py following the existing pattern:
build a mock instance with testing_utils.TestConfigGenerator().build_hypertts_instance_test_servicemanager('default')
component-level tests: gui_testing_utils.build_empty_dialog(), then component.draw(dialog.getLayout())
full-workflow tests: register a dialog_input_fn_map[constants.DIALOG_ID_<X>] callback and call
the create_component_* factory
one test_<name>_manual guarded by HYPERTTS_<X>_DIALOG_DEBUG=yes which calls dialog.exec(),
and a matching entry in scripts/openbox_menu_hypertts so the dialog can be eyeballed by hand
Run pytest -n auto before finishing. The tests/test_tts_services/ tests hit real TTS APIs and
a couple can fail for unrelated network/speech-recognition reasons.
Troubleshooting
"Add-on Startup Failed" — HyperTTS raised during import. The add-on folder must be named
anki-hyper-tts (it is what constants.CONFIG_ADDON_NAME looks up; any other name makes
getConfig() return None). Read the message with
./gui_probe.py raw '{"action":"eval","params":{"expression":"[t.toPlainText() for w in qt.QApplication.topLevelWidgets() for t in w.findChildren(qt.QTextBrowser)]"}}'
probe port not answering — tail -40 /tmp/hypertts-gui-automation/logs/anki.log
port already in use — a previous run was not torn down: ./teardown.sh
table dump shows QVariant objects — HyperTTS models return QVariant; the probe unwraps
them, so this means the probe is stale: ./start_anki.sh --restart
import/screenshot is black — nothing is mapped on the display yet, or Anki is still
starting; check ./status.sh
Requirements
System packages (Fedora, installed with sudo dnf install; see
docs/AI_GUI_AUTOMATION_SETUP.md):
xorg-x11-server-Xvfb openbox x11vnc novnc python3-websockify xdotool wmctrl ImageMagick
The helper scripts only use the python standard library, so there are no additions to
requirements.txt. xdotool/wmctrl are available for real X11 input events if a widget ever
resists QWidget.click().
The harness deliberately does not use the AT-SPI accessibility tree — the probe is cheaper and
more precise — and start_anki.sh exports QT_ACCESSIBILITY=0 / NO_AT_BRIDGE=1 so Qt does not
publish the widget tree over D-Bus. Do not install dbus-x11 / at-spi2-core for this workflow:
nothing here needs them.
Anki Gui Automation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
Anki Gui Automation compared with similar skills
Skill
Stars
Used in
Tokens
Auto-check
Licence
Repo updated
Anki Gui Automation this skillVocab-Apps/anki-hyper-tts
Comprehensive framework for deep analysis of articles, papers, and long-form content using 10+ thinking models (SCQA, 5W2H, critical thinking, inversion, mental models, first principles, systems…
A skill your agent uses when the user asks about Ni Haisha / 倪海厦 TCM course material, especially Shang Han Lun / 伤寒论, Jingui / 金匮要略, Zhongjing Xinfa / 仲景心法, clinical cases / 临床案例 / 倪师医案, Bagang…
Turns a named classical Chinese chapter, such as one from the Tao Te Ching or the Analects, into a single annotated PNG image with notes and commentary.
Review unresolved Sentry audio-request issues from the last 24 hours and report any whose exceptiontype looks mis-categorized against hyperttsaddon/errors.py
Develop and verify HyperTTS GUI screens against a real, running Anki instance: launch Anki headlessly on a throwaway profile, inject notes with AnkiConnect, read the live Qt widget tree as text…. Anki Gui Automation is an agent skill from Vocab-Apps/anki-hyper-tts.
When should I use Anki Gui Automation?
Anki Gui Automation fits situations like: tasks that involve Study guides and flashcards; tasks that involve Desktop control.
How do I install Anki Gui Automation in Claude Code?
Run `npx skills add Vocab-Apps/anki-hyper-tts --skill anki-gui-automation -a claude-code`. Or copy the skill folder (.claude/skills/anki-gui-automation in Vocab-Apps/anki-hyper-tts) into .claude/skills/anki-gui-automation in your project. Claude Code loads it when a task matches its description.
How do I install Anki Gui Automation in Codex?
Run `npx skills add Vocab-Apps/anki-hyper-tts --skill anki-gui-automation -a codex`. Or copy the skill folder (.claude/skills/anki-gui-automation in Vocab-Apps/anki-hyper-tts) into .agents/skills/anki-gui-automation in your project. Codex loads it when a task matches its description.
Can I use Anki Gui Automation in Cursor, Gemini CLI or GitHub Copilot?
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Vocab-Apps/anki-hyper-tts --skill anki-gui-automation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/anki-gui-automation, .gemini/skills/anki-gui-automation, .github/skills/anki-gui-automation and .opencode/skills/anki-gui-automation in your project.
What does Anki Gui Automation need to run?
Going by SKILL.md and its folder, Anki Gui Automation needs the command-line tools its instructions call (pytest and dnf). Our summary lists: Python 3.
Does Anki Gui Automation access the network?
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Is Anki Gui Automation safe to install?
Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
What licence does Anki Gui Automation use?
Anki Gui Automation is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
How many tokens does Anki Gui Automation use?
About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
What are the alternatives to Anki Gui Automation?
Skills that share tags, products or a category with Anki Gui Automation: Deep Reading Analyst (ginobefun/deep-reading-analyst-skill, 354 stars), Nihaisha (JuneYaooo/nihaisha-nishi-tcm, 2.2k stars), Claude Certification Tutor (rohitg00/ai-engineering-from-scratch, 67k stars) and StudyVault Quiz Tutor (bevibing/tutor-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Who maintains Anki Gui Automation?
Vocab-Apps (a GitHub organization) maintains it in Vocab-Apps/anki-hyper-tts, which has 284 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 13, 2026.
Source: Vocab-Apps/anki-hyper-tts on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.