Agent skill

Anki Gui Automation

by Vocab-Apps in Vocab-Apps/anki-hyper-tts

Develop and verify HyperTTS GUI screens against a real, running Anki instance: launch Anki headlessly on a throwaway profile, inject notes with AnkiConnect, read the live Qt widget tree as text…

GPL-3.0Auto-check: notesEducation

Install Anki Gui Automation

skills CLI
$ npx skills add Vocab-Apps/anki-hyper-tts --skill anki-gui-automation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Vocab-Apps/anki-hyper-tts anki-gui-automation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Vocab-Apps/anki-hyper-tts.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/anki-gui-automation .claude/skills/anki-gui-automation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
anki-gui-automation
GitHub stars
284
Token cost
~2.8k tokens
SKILL.md length
1,161 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
GPL-3.0

At a glance

Develop and verify HyperTTS GUI screens against a real, running Anki instance: launch Anki headlessly on a throwaway profile, inject notes with AnkiConnect, read the live Qt widget tree as text…

  • Tasks that involve Study guides and flashcards
  • SKILL.md covers The loop, Widget naming convention (do…, scripts/gui_automation reference and Verifying a feature end to end, plus 3 more sections
  • Calls pytest and dnf
  • Tasks that involve Desktop control

What it does

Anki Gui Automation is an agent skill from Vocab-Apps/anki-hyper-tts. Develop and verify HyperTTS GUI screens against a real, running Anki instance: launch Anki headlessly on a throwaway profile, inject notes with AnkiConnect, read the live Qt widget tree as text, drive dialogs, screenshot them, and tear everything down

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Education, covering Study guides and flashcards and Desktop control. The repository describes itself as: HyperTTS Addon for Anki. The licence is GPL-3.0.

When your agent uses it

  • Tasks that involve Study guides and flashcards
  • Tasks that involve Desktop control

Example prompts

  • “/anki-gui-automation”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 05007b9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pytest
    • dnf

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Anki Gui Automation loads about 2.8k tokens when it runs. Until then it costs about 68 tokens; SKILL.md has 1,161 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~68
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:197
    System packages (Fedora, installed with `sudo dnf install`; see

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Vocab-Apps/anki-hyper-tts at commit 05007b9, republished under its GPL-3.0 licence (© Vocab-Apps). 1,161 words, ~2,775 tokens.

Download SKILL.mdSave it as .claude/skills/anki-gui-automation/SKILL.md (or your agent's skills folder).
name
anki-gui-automation
description
Develop and verify HyperTTS GUI screens against a real, running Anki instance: launch Anki headlessly on a throwaway profile, inject notes with AnkiConnect, read the live Qt widget tree as text, drive dialogs, screenshot them, and tear everything down
user_invocable
true

Agentic GUI development for HyperTTS

Use this whenever you add or change a HyperTTS dialog. pytest-qt tests prove the logic; this harness proves the dialog is actually reachable, laid out correctly, and correct against a real Anki collection (real notes, real update_note, real undo).

Everything runs on a virtual display against an isolated Anki base folder and profile, so the developer's own Anki profile and collection are never touched. Both service ports are deliberately non-default (AnkiConnect on 8766, not 8765) so the harness can never talk to a real Anki session that happens to be open.

You must run scripts/gui_automation/teardown.sh when you are done. Leaving Xvfb, x11vnc, websockify and a headless Anki running is the main failure mode of this workflow.

The loop

bash
cd scripts/gui_automation

./start_anki.sh                 # Xvfb + openbox + x11vnc + noVNC + Anki with HyperTTS
./ankiconnect.py seed --reset   # deck + note type + 5 notes with assorted sound tags
./ankiconnect.py browse         # open the browser on those notes
./gui_probe.py browser-select-all

./gui_probe.py actions --text Audio                   # what menu entries exist
./gui_probe.py trigger --text 'Remove Audio (Collection)...'   # open the dialog
./gui_probe.py windows                                # confirm it opened, and is modal
./gui_probe.py tree --window hypertts_remove_audio_dialog      # read the whole dialog as text
./gui_probe.py table --object-name hypertts_remove_audio_preview_table
./gui_probe.py screenshot --path /tmp/hypertts-gui-automation/artifacts/dialog.png \
    --window hypertts_remove_audio_dialog             # then Read the png

./teardown.sh                   # ALWAYS

After editing any HyperTTS python file, run ./start_anki.sh --restart. The add-on is symlinked into the profile, but Anki only imports it at startup. --fresh additionally throws away the collection.

Read text first, screenshot second. tree costs a few hundred tokens and tells you state (enabled, checked, current index, combo items); a screenshot costs thousands but is the only way to catch layout problems. Both matter — the two bugs found while building the Remove Audio dialog (a QLabel rendering <b> as literal text, and preview columns truncated by an even stretch) were invisible in the widget tree and obvious in the screenshot.

Watch it live in a browser at http://localhost:6099/vnc.html while it runs.

Widget naming convention (do this first, it unblocks everything)

Give every widget you create a stable objectName, prefixed hypertts_<screen>_:

python
self.field_combobox = aqt.qt.QComboBox()
self.field_combobox.setObjectName('hypertts_remove_audio_field')

Name the dialog too (self.setObjectName('hypertts_remove_audio_dialog')) so --window can target it. Without object names you have to address widgets by class + index path (RemoveAudioDialog/QGroupBox[0]/QComboBox[0]), which breaks the moment the layout changes. hypertts_addon/component_remove_audio.py is the reference implementation.

Older components (component_batch.py, component_voiceselection.py, …) have no object names yet. For those, address widgets by --class + --text, or by the path printed by tree.

scripts/gui_automation reference

scriptwhat it does
start_display.shXvfb :99, openbox, x11vnc 5999, noVNC 6099. Idempotent.
start_anki.sh [--restart|--fresh]prepares the profile, launches Anki, waits for both ports
setup_profile.pysymlinks HyperTTS + installs AnkiConnect + anki_gui_probe, creates the profile
stop_anki.shstops only Anki, leaves the display up
teardown.shstops everything, cleans the X lock
status.shwhat is running, which ports answer, which windows are open
screenshot.sh [name]full-screen grab into the artifacts dir (includes window decorations)
ankiconnect.pyinject/read collection data
gui_probe.pyinspect and drive the live GUI

Runtime state lives under /tmp/hypertts-gui-automation: anki_base/ (throwaway collection and add-ons), logs/anki.log, logs/hypertts.log (HyperTTS debug logging is on), artifacts/ (screenshots), pids/.

gui_probe.py

Inspect:

  • windows — every top-level window: class, title, modal, active, geometry
  • tree [--window W] [--named-only] [--all] [--max-depth N] — indented widget tree
  • info --object-name X — one widget in detail
  • table --object-name X — dump a QTableView/QTreeView model as rows (QVariant unwrapped)
  • actions [--text substring] — every QAction, i.e. every menu entry
  • undo-status — the label of the next undoable operation
  • note-fields --note-id N — a note's fields straight from the collection

Drive:

  • click --object-name X [--no-wait] — also --class QPushButton --text Cancel
  • set-text --object-name X --value '...'
  • combo --object-name X --text '...' (or --index N)
  • check --object-name X [--off]
  • select-row --object-name X --row N
  • trigger --text 'Menu entry...' — opens dialogs; does not wait, by design
  • close --title '...' [--class ...], undo, browser-select-all, browser-search --query ...
  • screenshot --path P [--window W] — QWidget.grab() of one window
  • raw '{"action": "eval", "params": {"expression": "mw.pm.name"}}' — escape hatch, main-thread eval with aqt, qt and mw in scope

Modal dialogs block the Qt main thread. Any action that opens one must not wait for the main thread, or the request times out. trigger defaults to not waiting; pass --no-wait to click when the click opens a dialog (including a QMessageBox such as HyperTTS's "Save changes to current preset ?" confirmation) or closes the window it lives on. When a probe call returns {"status": "pending"} that is not an error: a modal is open and holding the main thread — call windows to see it and deal with it.

Anki's own dialogs are reachable the same way, which is how you dismiss things like the add-on startup error box:

bash
./gui_probe.py tree --window QMessageBox
./gui_probe.py click --class QPushButton --text '&No' --window QMessageBox --no-wait
Show full SKILL.md (498 more words)Show less
ankiconnect.py
  • seed [--reset] — creates the deck HyperTTS Automation, the note type HyperTTS Automation Note (fields Chinese / English / Sound / Sound English), stores dummy media files, and adds 5 notes: a lone HyperTTS sound tag, text plus a HyperTTS sound tag, a foreign sound tag (external-recording.mp3, i.e. audio HyperTTS did not generate), HyperTTS audio in two fields, and a note with no audio. Extend seed_notes() for new cases.
  • notes [--query Q] — dump note fields, the cheapest way to assert what an operation did
  • browse [--query Q] — open the browser on a query
  • invoke <action> [--params JSON] — any AnkiConnect action

Two AnkiConnect gotchas, already handled in the helper but worth knowing:

  • AnkiConnect is unmaintained and still assigns the deck through the legacy note-type dict, which modern Anki ignores, so addNotes lands cards in Default. seed moves them with changeDeck.
  • Anki search syntax wants the quotes around the whole term: "deck:HyperTTS Automation", not deck:"HyperTTS Automation".

Verifying a feature end to end

The pattern that actually proves a collection-modifying dialog works — this is how the Remove Audio dialog was validated:

bash
./ankiconnect.py notes            # state before
./gui_probe.py undo-status        # "" - nothing to undo yet
./gui_probe.py click --object-name hypertts_remove_audio_remove_button
./ankiconnect.py notes            # state after: only the intended fields changed
./gui_probe.py undo-status        # "HyperTTS: Remove Audio from Notes" <- undo support works
./gui_probe.py undo
./ankiconnect.py notes            # back to the original state

Undo support in HyperTTS comes from running the mutation inside anki_utils.run_in_background_collection_op(parent, update_fn, success_fn, undo_entry_name=...), which wraps it in a custom undo entry. update_fn receives the collection and must call collection.update_note(note); never call aqt.mw.col directly from a dialog.

pytest-qt tests are still required

The harness complements the test suite, it does not replace it. Every new dialog needs a tests/test_component_<name>.py following the existing pattern:

  • build a mock instance with testing_utils.TestConfigGenerator().build_hypertts_instance_test_servicemanager('default')
  • component-level tests: gui_testing_utils.build_empty_dialog(), then component.draw(dialog.getLayout())
  • full-workflow tests: register a dialog_input_fn_map[constants.DIALOG_ID_<X>] callback and call the create_component_* factory
  • one test_<name>_manual guarded by HYPERTTS_<X>_DIALOG_DEBUG=yes which calls dialog.exec(), and a matching entry in scripts/openbox_menu_hypertts so the dialog can be eyeballed by hand

Run pytest -n auto before finishing. The tests/test_tts_services/ tests hit real TTS APIs and a couple can fail for unrelated network/speech-recognition reasons.

Troubleshooting

  • "Add-on Startup Failed" — HyperTTS raised during import. The add-on folder must be named anki-hyper-tts (it is what constants.CONFIG_ADDON_NAME looks up; any other name makes getConfig() return None). Read the message with ./gui_probe.py raw '{"action":"eval","params":{"expression":"[t.toPlainText() for w in qt.QApplication.topLevelWidgets() for t in w.findChildren(qt.QTextBrowser)]"}}'
  • probe port not answering — tail -40 /tmp/hypertts-gui-automation/logs/anki.log
  • port already in use — a previous run was not torn down: ./teardown.sh
  • table dump shows QVariant objects — HyperTTS models return QVariant; the probe unwraps them, so this means the probe is stale: ./start_anki.sh --restart
  • import/screenshot is black — nothing is mapped on the display yet, or Anki is still starting; check ./status.sh

Requirements

System packages (Fedora, installed with sudo dnf install; see docs/AI_GUI_AUTOMATION_SETUP.md): xorg-x11-server-Xvfb openbox x11vnc novnc python3-websockify xdotool wmctrl ImageMagick

The helper scripts only use the python standard library, so there are no additions to requirements.txt. xdotool/wmctrl are available for real X11 input events if a widget ever resists QWidget.click().

The harness deliberately does not use the AT-SPI accessibility tree — the probe is cheaper and more precise — and start_anki.sh exports QT_ACCESSIBILITY=0 / NO_AT_BRIDGE=1 so Qt does not publish the widget tree over D-Bus. Do not install dbus-x11 / at-spi2-core for this workflow: nothing here needs them.

© Vocab-Apps, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/anki-gui-automation of Vocab-Apps/anki-hyper-tts.

Open the folder on GitHubat commit 05007b9

Compare with similar skills

Anki Gui Automation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Anki Gui Automation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Anki Gui Automation this skillVocab-Apps/anki-hyper-tts284—~2.8kAutomated safety check: NotesGPL-3.0
Deep Reading Analystginobefun/deep-reading-analyst-skill3544 repos~3.6kAutomated safety check: PassMIT
NihaishaJuneYaooo/nihaisha-nishi-tcm2.2k—~4kAutomated safety check: PassNone
Claude Certification Tutorrohitg00/ai-engineering-from-scratch67k—~3kAutomated safety check: PassMIT
StudyVault Quiz Tutorbevibing/tutor-skills1.3k—~1.4kAutomated safety check: PassMIT
Project Mastery Coachtudoumashu/ai-memory-skillpack456—~1.8kAutomated safety check: PassMIT

Similar skills

  • Deep Reading Analyst

    ginobefun/deep-reading-analyst-skill

    Comprehensive framework for deep analysis of articles, papers, and long-form content using 10+ thinking models (SCQA, 5W2H, critical thinking, inversion, mental models, first principles, systems…

    354 GitHub starsUsed in 4 repos~3.6k tokens
    EducationAuto-check passed
  • Nihaisha

    JuneYaooo/nihaisha-nishi-tcm

    A skill your agent uses when the user asks about Ni Haisha / 倪海厦 TCM course material, especially Shang Han Lun / 伤寒论, Jingui / 金匮要略, Zhongjing Xinfa / 仲景心法, clinical cases / 临床案例 / 倪师医案, Bagang…

    2.2k GitHub stars~4k tokensUpdated 25 days ago
    EducationAuto-check passed
  • Claude Certification Tutor

    rohitg00/ai-engineering-from-scratch

    Guides a learner through one of four independent Claude certification tracks with onboarding, lessons, practice labs, mock exams and remediation.

    67k GitHub stars~3k tokensUpdated today
    EducationAuto-check passed
  • StudyVault Quiz Tutor

    bevibing/tutor-skills

    Quizzes you on the notes in an Obsidian StudyVault, tracks proficiency per concept and drills weak areas in four-question rounds.

    1.3k GitHub stars~1.4k tokensUpdated 7 mo ago
    EducationAuto-check passed
  • Project Mastery Coach

    tudoumashu/ai-memory-skillpack

    Train strict project ownership from repo-local docs/ai memory and central LLM Wiki project entities.

    456 GitHub stars~1.8k tokensUpdated 1 mo ago
    EducationAuto-check passed
  • Turns a named classical Chinese chapter, such as one from the Tao Te Ching or the Analects, into a single annotated PNG image with notes and commentary.

    7.5k GitHub stars~551 tokensUpdated 2 days ago
    EducationAuto-check passed

More from Vocab-Apps/anki-hyper-tts

  • Sentry Issues

    Vocab-Apps/anki-hyper-tts

    Look up HyperTTS crash reports in Sentry — the project is language-tools/anki-hyper-tts, project ID 6170140.

    284 GitHub stars~1.9k tokensUpdated 28 days ago
    Auto-check passed
  • Sentry Audio Error Review

    Vocab-Apps/anki-hyper-tts

    Review unresolved Sentry audio-request issues from the last 24 hours and report any whose exceptiontype looks mis-categorized against hyperttsaddon/errors.py

    284 GitHub stars~1.8k tokensUpdated 28 days ago
    Auto-check passed

Questions about Anki Gui Automation

What does Anki Gui Automation do?

Develop and verify HyperTTS GUI screens against a real, running Anki instance: launch Anki headlessly on a throwaway profile, inject notes with AnkiConnect, read the live Qt widget tree as text…. Anki Gui Automation is an agent skill from Vocab-Apps/anki-hyper-tts.

When should I use Anki Gui Automation?

Anki Gui Automation fits situations like: tasks that involve Study guides and flashcards; tasks that involve Desktop control.

How do I install Anki Gui Automation in Claude Code?

Run `npx skills add Vocab-Apps/anki-hyper-tts --skill anki-gui-automation -a claude-code`. Or copy the skill folder (.claude/skills/anki-gui-automation in Vocab-Apps/anki-hyper-tts) into .claude/skills/anki-gui-automation in your project. Claude Code loads it when a task matches its description.

How do I install Anki Gui Automation in Codex?

Run `npx skills add Vocab-Apps/anki-hyper-tts --skill anki-gui-automation -a codex`. Or copy the skill folder (.claude/skills/anki-gui-automation in Vocab-Apps/anki-hyper-tts) into .agents/skills/anki-gui-automation in your project. Codex loads it when a task matches its description.

Can I use Anki Gui Automation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Vocab-Apps/anki-hyper-tts --skill anki-gui-automation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/anki-gui-automation, .gemini/skills/anki-gui-automation, .github/skills/anki-gui-automation and .opencode/skills/anki-gui-automation in your project.

What does Anki Gui Automation need to run?

Going by SKILL.md and its folder, Anki Gui Automation needs the command-line tools its instructions call (pytest and dnf). Our summary lists: Python 3.

Does Anki Gui Automation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Anki Gui Automation safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Anki Gui Automation use?

Anki Gui Automation is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Anki Gui Automation use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Anki Gui Automation?

Skills that share tags, products or a category with Anki Gui Automation: Deep Reading Analyst (ginobefun/deep-reading-analyst-skill, 354 stars), Nihaisha (JuneYaooo/nihaisha-nishi-tcm, 2.2k stars), Claude Certification Tutor (rohitg00/ai-engineering-from-scratch, 67k stars) and StudyVault Quiz Tutor (bevibing/tutor-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Anki Gui Automation?

Vocab-Apps (a GitHub organization) maintains it in Vocab-Apps/anki-hyper-tts, which has 284 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 13, 2026.

Source: Vocab-Apps/anki-hyper-tts on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.