Black-box QA audit of slop-guard across MCP, CLI, fit, docs, agent workflows, and writing-effectiveness.

MITAuto-check passedAgent Workflows

Install QA

skills CLI
$ npx skills add eric-tramel/slop-guard --skill qa -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install eric-tramel/slop-guard qa --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/eric-tramel/slop-guard.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/qa .claude/skills/qa && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
qa
GitHub stars
164
Token cost
~2.5k tokens
SKILL.md length
1,377 words
Files
1
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Black-box QA audit of slop-guard across MCP, CLI, fit, docs, agent workflows, and writing-effectiveness.

  • Works in 6 steps: MCP Agent Interface (mcp) → CLI UX (cli) → Fit Workflow (fit) → …
  • Tasks that involve MCP servers
  • SKILL.md covers Setup, Testing Angles, Issue Filing and Summary Report
  • Calls uv, gh and uvx; reaches claude.com

What it does

QA is an agent skill from eric-tramel/slop-guard. Black-box QA audit of slop-guard across MCP, CLI, fit, docs, agent workflows, and writing-effectiveness. Files GitHub issues for real problems found.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering MCP servers. It works with Model Context Protocol and GitHub. The repository describes itself as: Slop Scoring to Stop Slop. The licence is MIT.

When your agent uses it

  • Tasks that involve MCP servers

Example prompts

  • “/qa”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. MCP Agent Interface (mcp)
  2. CLI UX (cli)
  3. Fit Workflow (fit)
  4. Docs and Onboarding (docs)
  5. End-to-End Workflows (workflows)
  6. Writing Effectiveness (effectiveness)

What it can do on your machine

Read from SKILL.md and the folder at commit 7ef2113. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • gh
    • uvx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • claude.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

QA loads about 2.5k tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 1,377 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~38
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from eric-tramel/slop-guard at commit 7ef2113, republished under its MIT licence (© eric-tramel). 1,377 words, ~2,479 tokens.

Download SKILL.mdSave it as .claude/skills/qa/SKILL.md (or your agent's skills folder).
name
qa
description
Black-box QA audit of slop-guard across MCP, CLI, fit, docs, agent workflows, and writing-effectiveness. Files GitHub issues for real problems found.
argument-hint
[focus-area or 'all']

Slop-Guard QA Audit

Black-box QA of slop-guard through its public interfaces only. Do not read source files or tests to find bugs — discover issues by using the product as an agent or user would.

Launch parallel subagents per testing angle. File genuine issues with gh issue create. If $ARGUMENTS specifies a focus area (mcp, cli, fit, docs, workflows, or effectiveness), run only that angle. Otherwise run all 6 in parallel.

Setup

  1. Read README.md for documented behavior.
  2. If mcp__slop_guard__* tools are available, fetch their schemas via tool search.
  3. Fetch open issues to avoid duplicates:
    bash
    gh issue list --repo eric-tramel/slop-guard --limit 100 --state open --json number,title,body,labels
  4. Prefer local QA fixtures if present. Otherwise create temp files under /tmp/slop-guard-qa. Do not write throwaway files into the repo.

Pass the open issues list into every subagent prompt.

Testing Angles

1. MCP Agent Interface (mcp)

Is the MCP surface clear, consistent, and useful for an agent?

  • Tool descriptions, parameter clarity, distinction between check_slop and check_slop_file
  • JSON output structure: score, band, violations, counts, advice, file
  • Edge cases: short text, empty strings, unicode, markdown, code fences, long inputs
  • Missing/unreadable file paths
  • Response size and parseability
  • Cross-check MCP vs CLI JSON output on the same input

Issue prefix: mcp:

2. CLI UX (cli)

Does uv run sg behave predictably?

  • All flags from sg --help: --json, --verbose, --quiet, --threshold, --score-only, --counts, --config, --version
  • Input modes: file paths, inline text, stdin (-), multiple inputs
  • Exit codes: 0 (pass), 1 (threshold fail), 2 (error)
  • Stdout/stderr separation for scripting
  • Path-vs-inline-text ambiguity
  • Missing files, bad config paths

Issue prefix: cli:

3. Fit Workflow (fit)

Can a user fit a custom rule config from docs alone?

  • Legacy shorthand: sg-fit TARGET_CORPUS OUTPUT
  • Multi-input mode with --output
  • All input formats: .jsonl, .txt, .md
  • --negative-dataset, --no-calibration, --init
  • Error cases: malformed JSONL, missing text field, invalid label, missing files, unsupported suffixes
  • Round-trip: is the fitted JSONL immediately usable with sg -c?

Issue prefix: fit:

4. Docs and Onboarding (docs)

Do install flows and docs match actual behavior?

  • Install paths: uvx slop-guard, uv tool install slop-guard, uv run sg
  • MCP setup snippets for Claude and Codex
  • Custom config examples
  • Whether documented commands, flags, and tool names match reality

Issue prefix: docs:

5. End-to-End Workflows (workflows)

Does slop-guard compose well for real agent work?

  1. Lint prose via MCP or CLI, assess whether advice is actionable enough to drive a rewrite
  2. Lint a real file via check_slop_file or sg README.md
  3. CI gating with sg -t 60 ... — is the output machine-parseable?
  4. Compare MCP and CLI results on the same input

Focus on cross-surface consistency, scoring coherence, and whether the product actually improves agent writing workflows.

Issue prefix: workflow:

6. Writing Effectiveness (effectiveness)

Does slop-guard's feedback actually make an agent produce better writing? This angle tests the quality of the feedback loop — not whether the tool runs correctly, but whether following its guidance leads to measurably improved prose.

Method: Act as an agent that writes, scores, interprets advice, rewrites, and rescores. Evaluate every step for friction, ambiguity, and actual improvement.

6a. Advice Actionability

For each test, write a paragraph in a specific register (technical docs, blog post, marketing copy, academic summary, casual explanation), score it, then attempt to follow every piece of advice literally.

  • Is each advice string specific enough to know what to change? ("Replace 'crucial' — what specifically do you mean?" is good; a vague suggestion would be bad)
  • Does the advice tell you how to fix it, or only that something is wrong?
  • When multiple violations overlap in the same sentence, does the combined advice make sense or contradict itself?
  • Are there violations with no corresponding advice entry? (The advice array should cover every fixable issue)
  • Does any advice lead you toward a different slop pattern? (e.g., replacing a slop word with another slop word)
6b. Feedback Loop Convergence

Run iterative rewrite cycles and track whether the score converges upward:

  1. Write 3-5 paragraphs of deliberately sloppy AI-style prose (heavy on slop words, uniform rhythm, bold-bullet structures, contrast pairs)
  2. Score via check_slop or sg --json
  3. Rewrite following only the advice array — no independent judgment
  4. Rescore the rewrite
  5. Repeat until score stabilizes or advice is empty

Evaluate:

  • Does the score improve monotonically with each cycle? If not, what caused a regression?
  • How many cycles to reach clean band? (Should be 1-2 for light, 2-3 for moderate)
  • Does the advice array shrink each cycle, or do new violations appear as old ones are fixed?
  • Is there a score plateau where advice remains but following it doesn't move the score? What's blocking?
  • At convergence, does the text actually read well — or has it been "optimized" into something sterile/awkward?
6c. Score Sensitivity and Proportionality

Test whether scores reflect actual writing quality differences:

  • Write two versions of the same content: one natural, one AI-sloppy. Do scores separate cleanly?
  • Take clean prose (human-written, score >80) and introduce one slop word. How much does the score drop? Is the penalty proportional?
  • Take a short text (<50 words) vs a long text (~500 words) with the same density of violations. Are scores comparable?
  • Does the concentration penalty feel right? Write text with 1 contrast pair (should be mild) vs 5 contrast pairs (should be harsh). Check that scoring reflects this.
Show full SKILL.md (516 more words)Show less
6d. Advice Interpretation by an Agent

Simulate how an LLM agent would parse and act on the JSON output:

  • Given the violations array with match and context fields, can you locate the exact position in the original text to edit? Is the 60-char context window sufficient?
  • When advice says "Replace X — what specifically do you mean?", does the agent have enough context from the surrounding violations to pick a good replacement?
  • Are the counts useful for triage? (e.g., "12 slop_words vs 1 rhythm issue" — should the agent focus on vocabulary first?)
  • If the agent can only make one edit, does the output help prioritize which fix has the highest impact? (Penalties vary: -10 for ai_disclosure vs -1 for a slop word)
  • Does the band label (light, moderate, etc.) help the agent decide whether to rewrite or accept?
6e. Register Sensitivity

Test whether slop-guard handles different writing registers fairly:

  • Technical documentation: Do legitimate technical terms get false-positived as slop? (e.g., "framework", "robust", "scalable" in a systems design context)
  • Persuasive/marketing copy: This register naturally uses more emphatic language. Does slop-guard penalize it unfairly, or does it correctly distinguish marketing voice from AI slop?
  • Academic writing: Hedging language ("arguably", "perhaps", "it seems") is standard in academic prose. Does the weasel rule over-penalize legitimate scholarly register?
  • Conversational/casual: Does informal writing score better simply because it avoids formal slop patterns, even if it's low-quality?

For each register, write a good example and a sloppy example. The tool should score good writing higher regardless of register.

6f. Rewrite Quality Assessment

After the agent completes a rewrite cycle, evaluate the output text:

  • Does the rewritten text preserve the original meaning and intent?
  • Has the rewrite introduced awkward phrasing or unnatural constructions to dodge rules?
  • Is the rewritten text something a human editor would accept, or does it feel "linted" — technically clean but lifeless?
  • Compare: original sloppy text vs rewritten text vs hand-written alternative. Where does the tool-guided rewrite land on that spectrum?

File issues for patterns where following advice consistently produces worse prose, where scores don't reflect quality, where the feedback loop stalls, or where the interface makes it hard for an agent to act on feedback.

Issue prefix: effectiveness:

Issue Filing

Before filing, check every open issue — skip if same root cause, same behavior, or strict subset. When in doubt, don't file. If related but distinct, add Related: #N.

Use existing repo labels (bug, enhancement, documentation, etc.). Each issue body must include:

  1. Summary (user/agent perspective)
  2. Reproduction steps (exact commands or tool calls)
  3. Expected vs observed behavior
  4. Severity: critical / high / medium / low
  5. Surface: mcp / cli / fit / docs / workflow / effectiveness
  6. Generated with [Claude Code](https://claude.com/claude-code)

For effectiveness issues: use the enhancement label unless the issue describes advice that actively misleads (then bug). Include the original text, the advice received, the rewrite attempt, and before/after scores. Concrete examples are mandatory — do not file vague "advice could be better" issues.

Summary Report

After all angles complete, compile results grouped by severity. Include: total issues filed, findings already tracked, what works well, what was not tested, and whether MCP tools were available.

© eric-tramel, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/qa of eric-tramel/slop-guard.

Open the folder on GitHubat commit 7ef2113

Compare with similar skills

QA next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

QA compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
QA this skilleric-tramel/slop-guard164—~2.5kAutomated safety check: PassMIT
Skill Seekers Builderyusufkaraaslan/Skill_Seekers15k—~760Automated safety check: PassMIT
MCP Apps Builderawslabs/cli-agent-orchestrator1.4k—~1.7kAutomated safety check: PassApache-2.0
Releasejgravelle/jcodemunch-mcp2.7k—~6.5kAutomated safety check: PassCustom licence
Open PRArcadeAI/arcade-mcp1k—~2.9kAutomated safety check: PassMIT
MCP Clientcoleam00/second-brain-skills832—~1kAutomated safety check: PassNone

Similar skills

  • Skill Seekers Builder

    yusufkaraaslan/Skill_Seekers

    Detects the type of a knowledge source and uses the Skill Seekers MCP tools to turn docs, repos, PDFs or videos into packaged AI skills.

    15k GitHub stars~760 tokensUpdated 7 days ago
    Agent WorkflowsAuto-check passed
  • MCP Apps Builder

    awslabs/cli-agent-orchestrator

    Official

    Load the official MCP Apps builder skills (create-mcp-app, migrate-oai-app, add-app-to-server, convert-web-app) from github.com/modelcontextprotocol/ext-apps.

    1.4k GitHub stars~1.7k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Release

    jgravelle/jcodemunch-mcp

    Publishing a jMunch release (jcodemunch-mcp, jdocmunch-mcp, jdatamunch-mcp, jragmunch-cli), reviewing/merging/closing PRs, and responding to the community.

    2.7k GitHub stars~6.5k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Open PR

    ArcadeAI/arcade-mcp

    Prepare arcade-mcp changes for review by verifying intended behavior, filling the repository PR template, and creating or updating the PR.

    1k GitHub stars~2.9k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • MCP Client

    coleam00/second-brain-skills

    Universal MCP client for connecting to any MCP server with progressive disclosure.

    832 GitHub stars~1k tokensUpdated 8 mo ago
    Agent WorkflowsAuto-check passed
  • Project Release

    swimmwatch/cloakbrowser-mcp

    Prepare, publish, verify, or recover a cloakbrowser-mcp release only when the user explicitly requests release work.

    161 GitHub stars~1.9k tokensUpdated 6 days ago
    Agent WorkflowsAuto-check passed

Categories

Questions about QA

What does QA do?

Black-box QA audit of slop-guard across MCP, CLI, fit, docs, agent workflows, and writing-effectiveness. QA is an agent skill from eric-tramel/slop-guard. Black-box QA audit of slop-guard across MCP, CLI, fit, docs, agent workflows, and writing-effectiveness.

When should I use QA?

QA fits situations like: tasks that involve MCP servers.

How do I install QA in Claude Code?

Run `npx skills add eric-tramel/slop-guard --skill qa -a claude-code`. Or copy the skill folder (.claude/skills/qa in eric-tramel/slop-guard) into .claude/skills/qa in your project. Claude Code loads it when a task matches its description.

How do I install QA in Codex?

Run `npx skills add eric-tramel/slop-guard --skill qa -a codex`. Or copy the skill folder (.claude/skills/qa in eric-tramel/slop-guard) into .agents/skills/qa in your project. Codex loads it when a task matches its description.

Can I use QA in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add eric-tramel/slop-guard --skill qa -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/qa, .gemini/skills/qa, .github/skills/qa and .opencode/skills/qa in your project.

What does QA need to run?

Going by SKILL.md and its folder, QA needs the command-line tools its instructions call (uv, gh and uvx).

Does QA access the network?

SKILL.md names 1 domain. In commands or code: claude.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is QA safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does QA use?

QA is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does QA use?

About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to QA?

Skills that share tags, products or a category with QA: Skill Seekers Builder (yusufkaraaslan/Skill_Seekers, 15k stars), MCP Apps Builder (awslabs/cli-agent-orchestrator, 1.4k stars), Release (jgravelle/jcodemunch-mcp, 2.7k stars) and Open PR (ArcadeAI/arcade-mcp, 1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains QA?

eric-tramel (a GitHub user) maintains it in eric-tramel/slop-guard, which has 164 GitHub stars. The repository was last updated on July 9, 2026.

Source: eric-tramel/slop-guard on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.