Agent skill

CodexBar Live QA

by steipete in steipete/CodexBar

Runs live QA for the CodexBar app: provider usage matrix checks through its packaged CLI, config validation and menu checks, with 1Password-backed credentials handled safely.

MITAuto-check passedTesting & QA

Install CodexBar Live QA

skills CLI
$ npx skills add steipete/CodexBar --skill qa-test -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install steipete/CodexBar qa-test --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/steipete/CodexBar.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/qa-test .claude/skills/qa-test && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
qa-test
GitHub stars
22k
Token cost
~1.2k tokens
SKILL.md length
387 words
Files
4 (incl. scripts, references)
Skills in repo
3
Repo updated
First seen
Licence
MIT

At a glance

Runs live QA for the CodexBar app: provider usage matrix checks through its packaged CLI, config validation and menu checks, with 1Password-backed credentials handled safely.

  • Running a release smoke test across CodexBar's providers
  • SKILL.md covers Rules, CLI Matrix, Config QA and Live Menu QA, plus 3 more sections
  • Runs Shell scripts from its folder; calls jq, rg and make; needs OPENAI_API_KEY
  • Debugging a report that one provider stopped working

What it does

The skill is for live provider testing, release smoke tests, menu verification and debugging reports that a provider works or fails. The agent works from the CodexBar checkout and uses the packaged `CodexBarCLI` helper rather than the app binary, which can look like it hangs as a CLI. A bundled script, `scripts/live_provider_matrix.sh`, runs the checks in several modes: the providers enabled in the app's config, the app-facing default command, every registered provider (a discovery tool that is expected to fail for providers without sessions or keys) or a list you name. A config counts as green when the enabled and default runs are clean.

Config QA validates the config file, prints its shape with API keys, secrets and cookie headers redacted, and takes a timestamped backup before any edit. Menu checks use Peekaboo after the CLI passes. For API behavior the agent browses official provider docs only, with a reference file of API specs, and may use Browser Use for logged-in dashboards. Safety rules forbid dumping environment variables or secret patterns, keep 1Password commands inside one persistent tmux session with no raw secret output, and treat browser-cookie and keychain flows as likely to trigger prompts.

When your agent uses it

  • Running a release smoke test across CodexBar's providers
  • Debugging a report that one provider stopped working
  • Validating CodexBar's config file before and after a change
  • Verifying the menu bar output of the app

Example prompts

  • “Run the live provider matrix for the enabled providers and tell me which ones fail.”
  • “Validate my CodexBar config and show it with the secrets redacted.”
  • “Check the DeepSeek provider against the official API docs and run it live.”
  • “Use Peekaboo to list the CodexBar menu items after the CLI checks pass.”

Requirements

  • The CodexBar repository checkout with `CodexBar.app` and its packaged CLI
  • Peekaboo, for menu checks
  • 1Password and tmux, when credentials are needed

What it can do on your machine

Read from SKILL.md and the folder at commit dd7dca4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • jq
    • rg
    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

CodexBar Live QA loads about 1.2k tokens when it runs, and up to ~1.4k if it reads all its reference files. Until then it costs about 60 tokens; SKILL.md has 387 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from steipete/CodexBar at commit dd7dca4, republished under its MIT licence (© steipete). 387 words, ~1,231 tokens.

Download SKILL.mdSave it as .claude/skills/qa-test/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
qa-test
description
CodexBar live QA/e2e testing: run provider usage matrix checks, validate real app config, use Peekaboo for menu proof, use Browser Use/official docs for API spec or logged-in dashboard checks, and handle 1Password credentials safely.

CodexBar Live QA

Use for live provider testing, release smoke tests, menu verification, or debugging “provider works/fails” reports.

Rules

  • Work from the CodexBar repo checkout.
  • Use the packaged CLI first: CodexBar.app/Contents/Helpers/CodexBarCLI.
  • Do not use CodexBar.app/Contents/MacOS/codexbar; that is the app binary and may appear to hang as a CLI.
  • Never run broad env, set, or secret regex dumps.
  • Use $one-password for secrets: all op commands inside one persistent tmux session, service account first, no raw secret output.
  • Treat browser-cookie/keychain flows as prompt-risky. Prefer CLI/API-token checks and KeychainNoUIQuery-safe tests unless the user explicitly requested live UI.
  • For current API behavior, browse official provider docs only.

CLI Matrix

Run the bundled script:

bash
.agents/skills/qa-test/scripts/live_provider_matrix.sh --enabled

Useful modes:

bash
.agents/skills/qa-test/scripts/live_provider_matrix.sh --provider all
.agents/skills/qa-test/scripts/live_provider_matrix.sh --providers openai,zai,deepseek
.agents/skills/qa-test/scripts/live_provider_matrix.sh --default

Interpretation:

  • --enabled asks CodexBarCLI config providers for enabled providers, honoring CODEXBAR_CONFIG and default toggles.
  • --default runs the app-facing default command with no provider override.
  • --provider all forces every registered provider and is expected to fail for providers without sessions/keys.
  • A green app config needs --enabled and --default clean; --provider all is a discovery/triage tool.

Config QA

Validate config:

bash
CodexBar.app/Contents/Helpers/CodexBarCLI config validate
stat -f '%Lp %N' "$HOME/.codexbar/config.json"

Redact config shape:

bash
jq '(.providers // []) |= map(.apiKey = (if .apiKey then "<redacted>" else .apiKey end) |
  .secretKey = (if .secretKey then "<redacted>" else .secretKey end) |
  .cookieHeader = (if .cookieHeader then "<redacted>" else .cookieHeader end) |
  (if .id == "stepfun" and has("region") then .region = "<redacted>" else . end) |
  .tokenAccounts = (if .tokenAccounts then (.tokenAccounts | .accounts = (.accounts | map(.token = "<redacted>"))) else .tokenAccounts end))' \
  "$HOME/.codexbar/config.json"

Before editing config, make a backup:

bash
cp "$HOME/.codexbar/config.json" "$HOME/.codexbar/config.pre-qa-$(date +%Y%m%d%H%M%S).json"
chmod 600 "$HOME/.codexbar"/config.pre-qa-*.json

Live Menu QA

Use Peekaboo after CLI checks:

bash
pkill -x CodexBar || pkill -f 'CodexBar.app/Contents/MacOS/CodexBar' || true
open -n "$PWD/CodexBar.app"
peekaboo menu list-all --json | rg -i 'codexbar'
peekaboo menu click-extra --title codexbar-merged --json
screencapture -x /tmp/codexbar-live-menu.png

Crop top-right menu if needed:

bash
sips --cropToHeightWidth 900 340 --cropOffset 20 2650 /tmp/codexbar-live-menu.png \
  --out /tmp/codexbar-live-menu-crop.png >/dev/null

Verify visually with view_image. Confirm provider tabs/rows match enabled config and no failing provider dominates the first screen.

Browser Use

Use $browser-use only when a logged-in dashboard, API key page, or provider docs need browser/profile state.

Existing Chrome path:

bash
mcporter call chrome-devtools.list_pages --args '{}' --output text
mcporter call chrome-devtools.navigate_page --args '{"url":"https://provider.example"}' --output text
mcporter call chrome-devtools.take_snapshot --args '{}' --output text

If Browser Use is unavailable, say so and use web search for public official docs; do not substitute isolated Playwright for login/profile-dependent pages.

Show full SKILL.md (134 more words)Show less

Fix Triage

  • Missing auth/session: configure key/session if available; otherwise leave provider disabled or report blocked auth.
  • Wrong provider API/spec: inspect official docs, then patch fetcher/settings/tests.
  • Provider key exists but live API rejects it: keep key stored if useful, disable provider if the menu would show a persistent error.
  • User-facing behavior changes need CHANGELOG.md.
  • Code fixes need focused tests, make check, $autoreview, and live CLI proof before landing.

Known CodexBar QA Notes

  • OpenAI Admin API key is the useful usage provider key. Project OPENAI_API_KEY values can fail legacy credit-balance fallback with 403.
  • Deepgram usage requires a key/project with Management API permissions; transcription-only keys can return 403.
  • Groq usage uses the Prometheus metrics API, not ordinary inference endpoints.
  • MiniMax pay-as-you-go API keys and Token Plan/Coding Plan keys are different; wrong key kind can leave usage unavailable.

© steipete, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in .agents/skills/qa-test of steipete/CodexBar.

  • SKILL.md
  • agents/openai.yaml
  • references/api-specs.md
  • scripts/live_provider_matrix.sh

Open the folder on GitHubat commit dd7dca4

Compare with similar skills

CodexBar Live QA next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

CodexBar Live QA compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
CodexBar Live QA this skillsteipete/CodexBar22k—~1.2kAutomated safety check: PassMIT
Senpi Agent QA Harnesscode-yeongyu/senpi470—~2.7kAutomated safety check: NotesMIT
tmux Real User TestingQwenLM/qwen-code28k—~2.3kAutomated safety check: PassApache-2.0
Verify Omnigent End-to-Endomnigent-ai/omnigent11k—~1.6kAutomated safety check: PassApache-2.0
Tmux Manual QA Workercode-yeongyu/senpi470—~1.6kAutomated safety check: PassMIT
Acceptance Evidence for Deliverieslobehub/lobehub83k—~9.7kAutomated safety check: PassApache-2.0

Similar skills

  • Senpi Agent QA Harness

    code-yeongyu/senpi

    Checks changes to the senpi coding agent by driving the real CLI from source in an isolated sandbox, over RPC, terminal UI, mock model and CLI smoke channels.

    470 GitHub stars~2.7k tokensUpdated today
    Testing & QAAuto-check: notes
  • tmux Real User Testing

    QwenLM/qwen-code

    Drives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review.

    28k GitHub stars~2.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Verify Omnigent End-to-End

    omnigent-ai/omnigent

    Spins up an isolated Omnigent server, runner and mock model to prove a user-facing behavior or bug fix with recorded evidence instead of reasoning from code.

    11k GitHub stars~1.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Tmux Manual QA Worker

    code-yeongyu/senpi

    Runs one manual QA scenario for the todo continuation feature in the real ./pi-test.sh CLI inside tmux, captures scrollback and checks a deterministic count marker.

    470 GitHub stars~1.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Verifies a delivery end to end by driving the real product on a CLI, web, desktop or iOS Simulator surface, capturing evidence and publishing a round with the lh CLI.

    83k GitHub stars~9.7k tokensUpdated today
    Testing & QAAuto-check passed
  • Codex Plugin QA

    code-yeongyu/oh-my-openagent

    Tests the omo Codex plugin in an isolated CODEX_HOME with a local mock model, proving hooks fired through app-server notifications without touching ~/.codex.

    70k GitHub stars~1.9k tokensUpdated today
    Testing & QAAuto-check passed

More from steipete/CodexBar

  • CodexBar Usage Reader

    steipete/CodexBar

    CodexBar read. Provider usage, limits, credits, config health. JSON. No writes.

    22k GitHub stars~320 tokensUpdated today
    Auto-check passed
  • CodexBar macOS Release

    steipete/CodexBar

    Releases a signed, notarized CodexBar build: confirms the changelog, resolves signing credentials from 1Password, runs the release script in tmux, and updates the Sparkle appcast and Homebrew tap.

    22k GitHub stars~1.5k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about CodexBar Live QA

What does CodexBar Live QA do?

Runs live QA for the CodexBar app: provider usage matrix checks through its packaged CLI, config validation and menu checks, with 1Password-backed credentials handled safely. The skill is for live provider testing, release smoke tests, menu verification and debugging reports that a provider works or fails. The agent works from the CodexBar checkout and uses the packaged `CodexBarCLI` helper rather than the app binary, which can look like it hangs as a CLI.

When should I use CodexBar Live QA?

CodexBar Live QA fits situations like: running a release smoke test across CodexBar's providers; debugging a report that one provider stopped working; validating CodexBar's config file before and after a change; verifying the menu bar output of the app.

How do I install CodexBar Live QA in Claude Code?

Run `npx skills add steipete/CodexBar --skill qa-test -a claude-code`. Or copy the skill folder (.agents/skills/qa-test in steipete/CodexBar) into .claude/skills/qa-test in your project. Claude Code loads it when a task matches its description.

How do I install CodexBar Live QA in Codex?

Run `npx skills add steipete/CodexBar --skill qa-test -a codex`. Or copy the skill folder (.agents/skills/qa-test in steipete/CodexBar) into .agents/skills/qa-test in your project. Codex loads it when a task matches its description.

Can I use CodexBar Live QA in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add steipete/CodexBar --skill qa-test -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/qa-test, .gemini/skills/qa-test, .github/skills/qa-test and .opencode/skills/qa-test in your project.

What does CodexBar Live QA need to run?

Going by SKILL.md and its folder, CodexBar Live QA needs a shell for the scripts in its folder, the command-line tools its instructions call (jq, rg and make) and credentials named OPENAI_API_KEY. Our summary lists: The CodexBar repository checkout with `CodexBar.app` and its packaged CLI; Peekaboo, for menu checks; 1Password and tmux, when credentials are needed.

Does CodexBar Live QA access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is CodexBar Live QA safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does CodexBar Live QA use?

CodexBar Live QA is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does CodexBar Live QA use?

About 1.2k tokens (SKILL.md is roughly 4.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 189 tokens, read only when the agent opens those files.

What are the alternatives to CodexBar Live QA?

Skills that share tags, products or a category with CodexBar Live QA: Senpi Agent QA Harness (code-yeongyu/senpi, 470 stars), tmux Real User Testing (QwenLM/qwen-code, 28k stars), Verify Omnigent End-to-End (omnigent-ai/omnigent, 11k stars) and Tmux Manual QA Worker (code-yeongyu/senpi, 470 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains CodexBar Live QA?

steipete (a GitHub user) maintains it in steipete/CodexBar, which has 22,250 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 7, 2026.

Source: steipete/CodexBar on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.