Agent skill

QA Find Bugs MCP

by bex-co in bex-co/beancount-io

Hunt bugs in the Beancount.io remote MCP server by driving the real POST /api-gateway/mcp endpoint with JSON-RPC and real MCP clients, checking transport, discovery, credential boundaries, tool and…

MITAuto-check passedBackend & APIs

Install QA Find Bugs MCP

skills CLI
$ npx skills add bex-co/beancount-io --skill qa-find-bugs-mcp -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install bex-co/beancount-io qa-find-bugs-mcp --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/bex-co/beancount-io.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/qa-find-bugs-mcp .claude/skills/qa-find-bugs-mcp && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
qa-find-bugs-mcp
GitHub stars
297
Token cost
~3.1k tokens
SKILL.md length
1,546 words
Files
2
Skills in repo
27
Repo updated
First seen
Licence
MIT

At a glance

Hunt bugs in the Beancount.io remote MCP server by driving the real POST /api-gateway/mcp endpoint with JSON-RPC and real MCP clients, checking transport, discovery, credential boundaries, tool and…

  • An agent-endpoint bug hunt
  • SKILL.md covers Prepare targets and credentials, Sweep whole journeys and Reproduce, research, and hand…
  • Calls yarn and node; reaches beancount.io; needs QA_MCP_TOKEN and BEANCOUNT_MCP_TOKEN
  • Tasks that involve MCP servers

What it does

QA Find Bugs MCP is an agent skill from bex-co/beancount-io. Hunt bugs in the Beancount.io remote MCP server by driving the real POST /api-gateway/mcp endpoint with JSON-RPC and real MCP clients, checking transport, discovery, credential boundaries, tool and resource results, the result envelope and prompts against REST/GraphQL controls, reproducing failures, tracing root causes, and deduplicating findings. Use for MCP QA or an agent-endpoint bug hunt. Skip ordinary code review, bug implementation, browser, native-mobile, and CLI QA.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `evals/evals.json`).

It sits in Backend & APIs, covering MCP servers, Debugging and Microservices. It works with Model Context Protocol and GraphQL. The repository describes itself as: 💰 Double-entry bookkeeping made easy — plain-text accounting for humans and AI agents. Polished iOS & Android app built with React Native + Expo. The licence is MIT.

When your agent uses it

  • An agent-endpoint bug hunt
  • Tasks that involve MCP servers
  • Tasks that involve Debugging

Example prompts

  • “/qa-find-bugs-mcp”

Requirements

  • Docker
  • A credential in QA_MCP_TOKEN
  • A credential in BEANCOUNT_MCP_TOKEN

What it can do on your machine

Read from SKILL.md and the folder at commit 0227141. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • yarn
    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • beancount.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • QA_MCP_TOKEN
    • BEANCOUNT_MCP_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

QA Find Bugs MCP loads about 3.1k tokens when it runs. Until then it costs about 124 tokens; SKILL.md has 1,546 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~124
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from bex-co/beancount-io at commit 0227141, republished under its MIT licence (© bex-co). 1,546 words, ~3,075 tokens.

Download SKILL.mdSave it as .claude/skills/qa-find-bugs-mcp/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
qa-find-bugs-mcp
description
Hunt bugs in the Beancount.io remote MCP server by driving the real POST /api-gateway/mcp endpoint with JSON-RPC and real MCP clients, checking transport, discovery, credential boundaries, tool and resource results, the result envelope and prompts against REST/GraphQL controls, reproducing failures, tracing root causes, and deduplicating findings. Use for MCP QA or an agent-endpoint bug hunt. Skip ordinary code review, bug implementation, browser, native-mobile, and CLI QA.

MCP QA bug hunt

Read the shared QA contract first, then backend-cluster/backend-v2/AGENTS.md, backend-cluster/backend-v2/docs/mcp.md (the customer contract: tools, resources, prompts, envelope, failure codes), ADR 0007 (docs/adrs/ADR007-backend-v2-mcp-surface.md, the endpoint contract) and ADR 0008 (surface parity). Exercise the real endpoint over HTTP; unit tests and source inspection support a finding but do not replace a live journey.

Arguments: optional base URL, journey names, ledger owner/name, wN, SHIP=1, DRY_RUN=1. Default to https://beancount.io and report mode. A base URL changes the deployment under test; record the server name and version initialize reports and whether it matches local HEAD before calling a source fix missing from production. DRY_RUN=1 still allows disposable writes on a designated QA ledger needed to reproduce bugs.

Prepare targets and credentials

  • Record branch, HEAD, dirty files, Node version, target base URL, the initialize response's serverInfo and negotiated protocolVersion, the credential kind (API key or OAuth token), its scopes and pin, and the chosen synthetic ledger. Never record the credential itself.

  • The endpoint is POST {base}/api-gateway/mcp. Send Accept: application/json, text/event-stream, a JSON content type and MCP-Protocol-Version. Responses may be SSE-framed; parse data: lines rather than assuming one JSON document. The server is stateless: every POST is independent, initialize is not a prerequisite for a call, and there is no session id.

  • Make requests with the bundled helper so the credential never reaches argv, a transcript or a tool call:

    sh
    node .agents/skills/qa-shared/scripts/qa-mcp.mjs [--credentials-file FILE] [--url URL] <command>

    Commands: initialize, tools-list, resources-list, prompts-list, call <tool> [json], read <uri>, prompt <name> [json], rpc <method> [json], and raw --method GET|DELETE|POST [--accept TEXT] [--body TEXT] for transport checks. --anonymous omits the credential. It reads QA_MCP_TOKEN (falling back to BEANCOUNT_MCP_TOKEN) from the environment or --credentials-file (not --env-file, which Node itself consumes after the script path), prints status, an allowlist of headers and the parsed JSON-RPC messages, redacts bcio_ prefixes and the token value, exits 2 for usage or credential problems and 3 for network failures or hangs. Run it from the repo root; --url defaults to the production endpoint and accepts HTTP only on loopback.

  • Use a user-designated QA credential: a bcio_ key minted on a dedicated QA account for a synthetic ledger, ideally one read/write key pinned to the ledger and one ledger.read key on the same ledger, both short-lived. Ask for the variable names or environment-file path, never for the value. Do not mint keys on, or point tools at, a personal account or ledger. Browser cookies from qa-login.mjs are not MCP credentials; the endpoint refuses them by design.

  • Without a credential, continue the anonymous transport and discovery checks and mark every authenticated journey unverified.

  • Disposable writes, accounts and admin journeys belong on the already running local deploy/docker-mac stack (http://localhost:42601, see its README for ports and DEV_PREMIUM_USER_IDS for local key minting). Use it only when it is already up; .pm/DO_NOT_DO.md forbids provisioning a second stack for MCP testing. Label local results as local, with the backend image or HEAD they ran.

  • yarn mcp:conformance <base-url> [--token …] [--read-only-token …] inside backend-cluster/backend-v2/ runs ADR 0007's checklist and only observes; run it first against the target and treat a failed check as a candidate. yarn mcp:agent-eval drives real Claude Code and Codex sessions and is billed; run it only when the user asks for client journeys.

  • Store evidence under the gitignored repo-root .tmp/qa-<date>/: the exact helper command, sanitized request and response, status, headers and timings. Bound every request with a timeout; a hang is a result.

Sweep whole journeys

Use selected journey names to narrow this table; otherwise survey the surface and deepen the journeys that show failures. Report each skipped group and why.

JourneyObservable promise
transportAnonymous POST returns 401 with WWW-Authenticate: Bearer resource_metadata=…, and that URL returns 200 with an RFC 9728 document naming an authorization server that also resolves. Authenticated GET/DELETE return 405 with Allow: POST and complete; anonymous GET still returns 401. Missing Accept values or a wrong content type fail before tool execution. A notification returns 202; an unknown method returns a JSON-RPC error. /mcp behaves as docs/mcp.md states.
discoveryinitialize negotiates each supported protocol version and returns serverInfo and instructions. tools/list, resources/templates/list and prompts/list match the documented inventory: every tool has a title, description, inputSchema and an object outputSchema; names are unique; prompt arguments are declared; resources/list emptiness is documented. Compare with mcp-tools.ts, mcp-resources.ts, mcp-prompts.ts.
boundariesA ledger.read key gets isError: true with FORBIDDEN on every write and admin tool, never a success or an internal error. A pinned key naming another ledger is refused; an unpinned key without ledger is refused with a hint naming listLedgers; resource URIs obey the same pin. A revoked or expired credential fails on the next request with UNAUTHENTICATED or 401. A browser session cookie is refused with the discovery hint.
readsrunBqlQuery and runBqlQueryStructured agree with REST query and GraphQL queryShell on rows, columns, decimals and dates; structuredContent and the text block describe the same result. listLedgers paging, checkLedger, getLedgerContext, getEntryContext, listLedgerFiles and readLedgerFiles line ranges match ledger contents and their REST counterparts. Empty, missing and forbidden cases return the documented code, not an empty success.
resourcesVocabulary, analysis, journal, statement and file templates expand as documented, including reserved {+path}, query parameters and JSON-encoded array filters. MIME types and contents shapes match the doc; a failure is a JSON-RPC error whose data.code matches the table, with the message unprefixed. Each read matches the REST route with the same suffix.
writesdry_run previews leave the ledger unchanged in a fresh request and in REST. appendLedgerText inserts in date order and checkLedger stays clean; addLedgerEntries refuses unbalanced input with UNBALANCED and honors allowInvalid; editLedgerFiles batches create/update/delete into one commit; editEntrySource returns CONFLICT on a stale sha256sum; renameLedgerFile preserves content. Verify persistence through readLedgerFiles and REST, and confirm exactly one commit per operation.
promptsprompts/get returns each playbook; malformed month or ledger is refused; every tool and beancount:// URI the text cites exists in this deployment's lists; a pinned credential is told its ledger. Compare served text with mcp-prompts.ts and classify a difference as deployment lag.
envelopeEvery failure carries isError: true and { ok: false, error: { code, message, hint } }; success carries { ok: true, result }; retryAfter appears only with RATE_LIMITED. Missing required arguments are BAD_USER_INPUT, not INTERNAL_SERVER_ERROR. Unexpected errors are masked in production; a DomainError keeps its message. ok and isError never disagree.
limitsThe handshake methods are not charged; write budgets are smaller than read budgets; RATE_LIMITED names retryAfter. Probe at human pace with a few extra calls, never a flood, and never against production for write budgets.
accountmanageLedgerCollaborators, manageLedgers, setLedgerStar and account resources honor scopes and relationships. API-key and SSH-key management, account deletion, and plan tiers and usage are deliberately absent from MCP (ADR 019, 2026-10-09 amendment); their absence is not a bug, and their presence on a deployment is either deployment lag or a regression. Create or delete only qa-<yyyymmdd>- resources on the designated QA account; never change collaborators, banks or billing on a pre-existing account.
clientsOnly when requested: a real Claude Code or Codex session connects with a bearer key, lists the same tools, exposes the four prompts, and completes a read journey; yarn mcp:agent-eval is the deep, billed harness. Record client versions and cost.
Show full SKILL.md (364 more words)Show less

After a surprising response, capture the full sanitized exchange, then repeat it in a fresh request. Compare with the same operation on REST (/api-gateway/v1/…) or GraphQL using the same credential where the surface accepts it; a discrepancy is a parity candidate owned by backend-v2, not automatically an MCP bug. Distinguish BAD_USER_INPUT, NOT_FOUND, FORBIDDEN, UNAUTHENTICATED, PREMIUM_REQUIRED, RATE_LIMITED, SERVICE_UNAVAILABLE and INTERNAL_SERVER_ERROR before naming a defect. A Cloudflare or bot-protection page is not an application error.

Reproduce, research, and hand off

Reproduce from a fresh request with the same credential kind and target, and check a nearby working control (another tool, the REST route, or a second ledger). Documented behaviors are not bugs: 405 on authenticated GET, an empty resources/list, SSE framing, a loosened published outputSchema, PREMIUM_REQUIRED on a free account, and error text that states a limit without naming a plan. Deployment lag is not a bug either: when the served inventory or prompt text differs from HEAD, cite the commit that already changed it.

Trace the route in src/features/ai-agent/api/mcp-route.ts (identity, pin, method set), the registry and result wrapper in src/server/api/composition-root.ts, tool descriptors in mcp-tools.ts and src/features/ai-agent/tools/, resources in mcp-resources.ts and mcp-resource-template.ts, error translation in mcp-errors.ts and mcp-result-text.ts, prompts in mcp-prompts.ts, the op table in src/server/api/op-class.ts, the limiter in rate-limit.ts and the protected services the adapters delegate to. Existing suites to compare with: mcp-route-methods, mcp-errors, mcp-output-schema, mcp-prompt-list, mcp-tool-list, mcp-resources-list, surface-parity and the *-parity contract tests. A proposed fix names the owning package and, per the parity workflow, every eligible REST/GraphQL/MCP surface it must change together.

Apply the shared root-cause, caller search, board/history dedupe and optional pm/ship steps. In the finding record, replace route/device evidence with base URL, serverInfo, protocol version, credential kind, scopes and pin (never the value), the exact helper command or JSON-RPC body, status, relevant headers, the sanitized response and the REST/GraphQL control. Include a minimal synthetic reproducer in public filing text.

Finish with findings by severity, coverage/skips, dedupe/filing status, and cleanup: revoke keys minted by this run, delete only qa-<yyyymmdd>- resources, and report anything left behind. Do not implement fixes unless requested; if requested, add contract coverage through the adapters and run yarn typecheck, yarn test and yarn generate-v1-openapi inside backend-cluster/backend-v2/ before handoff.

© bex-co, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .agents/skills/qa-find-bugs-mcp of bex-co/beancount-io.

  • SKILL.md
  • evals/evals.json

Open the folder on GitHubat commit 0227141

Compare with similar skills

QA Find Bugs MCP next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

QA Find Bugs MCP compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
QA Find Bugs MCP this skillbex-co/beancount-io297—~3.1kAutomated safety check: PassMIT
Shopify AI Toolkit Wrapperjeremylongshore/tons-of-skills-marketplace2.8k—~1.5kAutomated safety check: PassMIT
Shopifyasgeirtj/system_prompts_leaks69k—~2kAutomated safety check: PassCC0-1.0
MondayLeoYeAI/openclaw-master-skills2.2k—~4.3kAutomated safety check: PassMIT
Build Error AdapterArcadeAI/arcade-mcp1k—~2kAutomated safety check: PassMIT
SpikardGoldziher/spikard124—~799Automated safety check: PassMIT

Similar skills

  • Shopify AI Toolkit Wrapper

    jeremylongshore/tons-of-skills-marketplace

    Integrate Shopify's AI Toolkit MCP server with Claude Code for GraphQL validation, Liquid linting, and documentation search.

    2.8k GitHub stars~1.5k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Shopify

    asgeirtj/system_prompts_leaks

    Set up and operate a Shopify store with Shopify's official MCP server (the installed shopify command, NOT the npm Shopify CLI).

    69k GitHub stars~2k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Monday

    LeoYeAI/openclaw-master-skills

    Manage monday.com boards, items, columns, groups, updates, and workflows via MCP server (preferred) and GraphQL API (fallback).

    2.2k GitHub stars~4.3k tokensUpdated 2 mo ago
    Backend & APIsAuto-check passed
  • Build Error Adapter

    ArcadeAI/arcade-mcp

    Build new Arcade error adapters from scratch using public Arcade TDK patterns.

    1k GitHub stars~2k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Spikard

    Goldziher/spikard

    Scaffold Spikard projects and generate code from OpenAPI, AsyncAPI, OpenRPC, GraphQL, and Protobuf schemas using the Spikard CLI or its MCP server.

    124 GitHub stars~799 tokensUpdated today
    Backend & APIsAuto-check passed
  • Flowstudio Power Automate Debug

    github/awesome-copilot

    Official

    Debug failing Power Automate cloud flows using the FlowStudio MCP server.

    40k GitHub starsUsed in 2 repos~5k tokens
    DevelopmentAuto-check passed

More from bex-co/beancount-io

All 27 skills in this repo
  • Beancount Close

    bex-co/beancount-io

    Close an accounting period in a Beancount ledger by reconciling each active account through beancount-reconcile, checking assertions and recurring gaps, reviewing flags, then proposing a commit with…

    297 GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Beancount Import

    bex-co/beancount-io

    Import a bank or card CSV, OFX/QFX, or QIF export into an existing Beancount ledger.

    297 GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Beancount Importer Author

    bex-co/beancount-io

    Write or repair a reusable Beangulp importer from a sample bank export, with reviewed golden files and a passing test harness.

    297 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Beancount Init

    bex-co/beancount-io

    Scaffold a new Beancount ledger with bea init and validation, with optional Fava browser setup when requested.

    297 GitHub stars~1.2k tokensUpdated today
    Auto-check: notes
  • Beancount Options

    bex-co/beancount-io

    Record a described options trade or lifecycle event as validated Beancount transactions through bea, after review and confirmation.

    297 GitHub stars~3.7k tokensUpdated today
    Auto-check passed
  • Beancount Reconcile

    bex-co/beancount-io

    Reconcile one Beancount account against a CSV statement or pasted PDF text.

    297 GitHub stars~4.1k tokensUpdated today
    Auto-check passed

Questions about QA Find Bugs MCP

What does QA Find Bugs MCP do?

Hunt bugs in the Beancount.io remote MCP server by driving the real POST /api-gateway/mcp endpoint with JSON-RPC and real MCP clients, checking transport, discovery, credential boundaries, tool and…. QA Find Bugs MCP is an agent skill from bex-co/beancount-io.io remote MCP server by driving the real POST /api-gateway/mcp endpoint with JSON-RPC and real MCP clients, checking transport, discovery, credential boundaries, tool and resource results, the result envelope and prompts against REST/GraphQL controls, reproducing failures, tracing root causes, and deduplicating findings.

When should I use QA Find Bugs MCP?

QA Find Bugs MCP fits situations like: an agent-endpoint bug hunt; tasks that involve MCP servers; tasks that involve Debugging.

How do I install QA Find Bugs MCP in Claude Code?

Run `npx skills add bex-co/beancount-io --skill qa-find-bugs-mcp -a claude-code`. Or copy the skill folder (.agents/skills/qa-find-bugs-mcp in bex-co/beancount-io) into .claude/skills/qa-find-bugs-mcp in your project. Claude Code loads it when a task matches its description.

How do I install QA Find Bugs MCP in Codex?

Run `npx skills add bex-co/beancount-io --skill qa-find-bugs-mcp -a codex`. Or copy the skill folder (.agents/skills/qa-find-bugs-mcp in bex-co/beancount-io) into .agents/skills/qa-find-bugs-mcp in your project. Codex loads it when a task matches its description.

Can I use QA Find Bugs MCP in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bex-co/beancount-io --skill qa-find-bugs-mcp -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/qa-find-bugs-mcp, .gemini/skills/qa-find-bugs-mcp, .github/skills/qa-find-bugs-mcp and .opencode/skills/qa-find-bugs-mcp in your project.

What does QA Find Bugs MCP need to run?

Going by SKILL.md and its folder, QA Find Bugs MCP needs the command-line tools its instructions call (yarn and node) and credentials named QA_MCP_TOKEN and BEANCOUNT_MCP_TOKEN. Our summary lists: Docker; A credential in QA_MCP_TOKEN; A credential in BEANCOUNT_MCP_TOKEN.

Does QA Find Bugs MCP access the network?

SKILL.md names 1 domain. In commands or code: beancount.io; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is QA Find Bugs MCP safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does QA Find Bugs MCP use?

QA Find Bugs MCP is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does QA Find Bugs MCP use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to QA Find Bugs MCP?

Skills that share tags, products or a category with QA Find Bugs MCP: Shopify AI Toolkit Wrapper (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Shopify (asgeirtj/system_prompts_leaks, 69k stars), Monday (LeoYeAI/openclaw-master-skills, 2.2k stars) and Build Error Adapter (ArcadeAI/arcade-mcp, 1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains QA Find Bugs MCP?

bex-co (a GitHub organization) maintains it in bex-co/beancount-io, which has 297 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on October 11, 2026.

Source: bex-co/beancount-io on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.