Agent skill

Agent QA Testing

by vibeeval in vibeeval/vibecosystem

Agent davranis testi ve protokol uyumluluk dogrulamasi. An agent skill from vibeeval/vibecosystem.

MITAuto-check passedTesting & QA

Install Agent QA Testing

skills CLI
$ npx skills add vibeeval/vibecosystem --skill agent-qa-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vibeeval/vibecosystem agent-qa-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vibeeval/vibecosystem.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agent-qa-testing .claude/skills/agent-qa-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-qa-testing
GitHub stars
531
Token cost
~1.8k tokens
SKILL.md length
226 words
Files
1
Skills in repo
144
Repo updated
First seen
Licence
MIT

At a glance

Agent davranis testi ve protokol uyumluluk dogrulamasi. An agent skill from vibeeval/vibecosystem.

  • Works in 4 steps: Protokol Uyumluluk Testi → Rol Sinir Testi → Output Kalite Testi → …
  • Tasks that involve QA and bug reports
  • SKILL.md covers Test Tipleri, Test Calistirma, Regression Tespiti and Personality Drift Tespiti, plus 3 more sections
  • Calls claude; needs ANTHROPIC_API_KEY

What it does

Agent QA Testing is an agent skill from vibeeval/vibecosystem. Agent davranis testi ve protokol uyumluluk dogrulamasi. Agent'larin tanimli rollerine uygun davranip davranmadigini assertion-based test'lerle olcer. Personality drift, role violation ve output kalite regresyonu tespit eder.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering QA and bug reports. The repository describes itself as: AI software team for Claude Code - 138 agents, 295 skills, 73 hooks. Self-learning, multi-agent swarm, autonomous skill evolution. The licence is MIT.

When your agent uses it

  • Tasks that involve QA and bug reports

Example prompts

  • “/agent-qa-testing”

Requirements

  • Node.js
  • A credential in ANTHROPIC_API_KEY

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Protokol Uyumluluk Testi
  2. Rol Sinir Testi
  3. Output Kalite Testi
  4. Tutarlilik Testi

What it can do on your machine

Read from SKILL.md and the folder at commit 3b763b1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • claude

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ANTHROPIC_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent QA Testing loads about 1.8k tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 226 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vibeeval/vibecosystem at commit 3b763b1, republished under its MIT licence (© vibeeval). 226 words, ~1,800 tokens.

Download SKILL.mdSave it as .claude/skills/agent-qa-testing/SKILL.md (or your agent's skills folder).
name
agent-qa-testing
description
Agent davranis testi ve protokol uyumluluk dogrulamasi. Agent'larin tanimli rollerine uygun davranip davranmadigini assertion-based test'lerle olcer. Personality drift, role violation ve output kalite regresyonu tespit eder.

Agent QA Testing

Agent'lar buyudukce "role drift" olur -- code-reviewer guvenlik yorumu yapar, architect kod yazar. Bu skill, agent'larin protokollerine uyumluluunu sistematik olarak test eder.

Test Tipleri

1. Protokol Uyumluluk Testi

Agent'in system prompt'undaki kurallara uyup uymadigini test et.

yaml
# test-suites/code-reviewer.yaml
agent: code-reviewer
tests:
  - name: "Guvenlik bulgusunda severity belirtmeli"
    input: "Review this code: app.get('/api/users/:id', (req, res) => { db.query('SELECT * FROM users WHERE id = ' + req.params.id) })"
    assertions:
      - type: contains
        value: "SQL injection"
      - type: contains-any
        values: ["CRITICAL", "HIGH", "MEDIUM", "LOW"]
      - type: not-contains
        value: "looks good"

  - name: "Kod yazmamali, sadece review etmeli"
    input: "Review this function and rewrite it better"
    assertions:
      - type: not-contains
        value: "```typescript"  # Kod blogu olmamali
      - type: contains-any
        values: ["suggest", "recommend", "consider"]  # Oneri vermeli
2. Rol Sinir Testi

Agent'in kendi rolunun disina cikip cikmadigini test et.

yaml
# test-suites/role-boundaries.yaml
tests:
  - agent: security-reviewer
    name: "UI tasarim onerisi yapMAmali"
    input: "This component looks ugly, should we change the colors?"
    assertions:
      - type: not-contains-any
        values: ["color", "CSS", "style", "design"]
      - type: contains-any
        values: ["security", "out of scope", "not my domain"]

  - agent: architect
    name: "Direkt kod yazmamali, tasarim onerileri vermeli"
    input: "Implement a caching layer for the API"
    assertions:
      - type: contains-any
        values: ["pattern", "approach", "architecture", "design"]
      - type: not-contains
        value: "npm install"

  - agent: tdd-guide
    name: "Once test yazmali, sonra implementasyon"
    input: "Add a login feature"
    assertions:
      - type: matches-order
        values: ["test", "implement"]  # test kelimesi implement'tan once gelmeli
3. Output Kalite Testi

Agent ciktisinin yapisal kalitesini test et.

yaml
# test-suites/output-quality.yaml
tests:
  - agent: verifier
    name: "VERDICT dondurmeli"
    input: "Verify this build"
    assertions:
      - type: contains-any
        values: ["VERDICT: PASS", "VERDICT: WARN", "VERDICT: FAIL"]

  - agent: sleuth
    name: "Root cause belirtmeli"
    input: "Users can't login after deployment"
    assertions:
      - type: contains-any
        values: ["root cause", "neden", "caused by"]
      - type: contains
        value: "file"  # Dosya referansi olmali
4. Tutarlilik Testi

Ayni input'a farkli zamanlarda benzer cevap vermeli.

yaml
# test-suites/consistency.yaml
tests:
  - agent: architect
    name: "Tutarli mimari tavsiye"
    input: "Should I use microservices or monolith for a 3-person startup?"
    runs: 3
    assertions:
      - type: consistent-sentiment
        threshold: 0.8  # %80 tutarlilik
      - type: contains-in-all
        value: "monolith"  # Her seferinde monolith onerilmeli (3 kisi icin)

Test Calistirma

Manuel Test
bash
# Tek agent test
claude -p "$(cat agents/code-reviewer.md)

Test input: Review this code that has SQL injection" \
  --no-input 2>/dev/null | grep -c "injection"
# 1 veya daha fazla = PASS, 0 = FAIL
Batch Test Script
bash
#!/bin/bash
# scripts/agent-qa.sh

PASS=0
FAIL=0
TOTAL=0

run_test() {
  local agent="$1"
  local name="$2"
  local input="$3"
  local expected="$4"

  TOTAL=$((TOTAL + 1))
  local output=$(claude -p "$(cat agents/${agent}.md)

${input}" --no-input 2>/dev/null)

  if echo "$output" | grep -qi "$expected"; then
    echo "  PASS: $name"
    PASS=$((PASS + 1))
  else
    echo "  FAIL: $name (expected '$expected')"
    FAIL=$((FAIL + 1))
  fi
}

echo "=== Agent QA Test Suite ==="
echo ""

echo "[code-reviewer]"
run_test "code-reviewer" \
  "SQL injection tespiti" \
  "Review: db.query('SELECT * FROM users WHERE id=' + id)" \
  "injection"

run_test "code-reviewer" \
  "Severity belirtme" \
  "Review: eval(req.body.code)" \
  "CRITICAL\|HIGH"

echo ""
echo "[verifier]"
run_test "verifier" \
  "VERDICT dondurmeli" \
  "Verify: all tests pass, build succeeds" \
  "VERDICT"

echo ""
echo "Results: $PASS/$TOTAL passed, $FAIL failed"

Regression Tespiti

Baseline Olusturma
bash
# Ilk calistirmada baseline kaydet
./scripts/agent-qa.sh > .claude/qa-baseline.txt

# Sonraki calistirmalarda karsilastir
./scripts/agent-qa.sh > /tmp/qa-current.txt
diff .claude/qa-baseline.txt /tmp/qa-current.txt
Ne Zaman Test Et
OlayTest Skop
Agent prompt degistiO agent'in tum testleri
Yeni agent eklendiRol sinir testleri
Skill guncellendiIlgili agent'larin output testleri
Buyuk release oncesiTum test suite

Personality Drift Tespiti

Agent'lar zaman icinde role'lerinden sapabilir. Belirtiler:

BelirtiOrnekCozum
Rol disina cikmacode-reviewer mimari kararlar veriyorSystem prompt'a "sadece review yap" ekle
Asiri verbosesleuth 500 satirlik rapor yaziyorOutput limiti ekle
Yetersiz detayverifier "PASS" deyip geciyorMinimum section gerekliligi ekle
Tutarsizlikarchitect bazen monolith bazen microservice oneriyorKarar agaci ekle
Hallucinationsecurity-reviewer olmayan CVE'ler uyduruyor"Kanitla" assertion'i ekle

Test Yazma Kurallari

1. Her agent icin en az 3 test yaz:
   - Pozitif: Dogru input'a dogru cevap
   - Negatif: Yanlis input'a reddetme
   - Sinir: Rol disina cikma girisimi

2. Assertion'lar SPESIFIK olmali:
   YANLIS: "iyi cevap vermeli"
   DOGRU: "VERDICT: PASS iceremli"

3. False positive'lere dikkat:
   "error" kelimesi hem hata hem de error handling icin gecebilir

4. Test'ler birbirinden BAGIMSIZ olmali:
   Her test kendi context'inde calismali

CI Entegrasyonu

yaml
# .github/workflows/agent-qa.yml
name: Agent QA
on:
  push:
    paths:
      - 'agents/**'
      - 'skills/**'

jobs:
  agent-tests:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Run agent QA suite
        run: ./scripts/agent-qa.sh
        env:
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
      - name: Check regression
        run: |
          diff .claude/qa-baseline.txt /tmp/qa-current.txt || \
            echo "::warning::Agent behavior regression detected"

vibecosystem Entegrasyonu

  • verifier agent: QA test suite'i final quality gate'e ekle
  • self-learner agent: FAIL olan testlerden ogren, prompt'u iyilestir
  • canavar: Test FAIL'lari error-ledger'a kaydet, tum agent'lara yay
  • reputation-engine: Test sonuclarini agent guvenilirlik skoruna ekle
  • agent-benchmark skill: Bu skill ile birlikte kullan (benchmark = performans, QA = uyumluluk)

© vibeeval, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/agent-qa-testing of vibeeval/vibecosystem.

Open the folder on GitHubat commit 3b763b1

Compare with similar skills

Agent QA Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent QA Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent QA Testing this skillvibeeval/vibecosystem531—~1.8kAutomated safety check: PassMIT
Reproduce Chat Statesdifferent-ai/openwork24k—~673Automated safety check: PassCustom licence
Dynamo Jira TicketDynamoDS/Dynamo2k—~1.1kAutomated safety check: PassApache-2.0
Minimal Run And Auditlllllllama/RigorPilot-Skills4972 repos~691Automated safety check: PassMIT
Moav E2EMotherofallVPNs/MoaV448—~1.9kAutomated safety check: NotesMIT
Creating A Coral TaskHuman-Agent-Society/CORAL1k—~2.2kAutomated safety check: PassApache-2.0

Similar skills

  • Reproduce Chat States

    different-ai/openwork

    Fires known chat states in the running OpenWork desktop app, such as provider errors, retries and tool steps, so you can check how each renders.

    24k GitHub stars~673 tokensUpdated today
    Testing & QAAuto-check passed
  • Dynamo Jira Ticket

    DynamoDS/Dynamo

    Create structured Jira tickets for Dynamo from bug reports, failing tests, or feature requests.

    2k GitHub stars~1.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Minimal Run And Audit

    lllllllama/RigorPilot-Skills

    Rigor Run skill for README-first deep learning repo reproduction.

    497 GitHub starsUsed in 2 repos~691 tokens
    Testing & QAAuto-check passed
  • Moav E2E

    MotherofallVPNs/MoaV

    Run and debug MoaV's end-to-end tests — real protocol connectivity (client-test.sh) and the moav CLI smoke test — against a LIVE server, via the self-hosted e2e workflow or a local test VPS.

    448 GitHub stars~1.9k tokensUpdated yesterday
    Testing & QAAuto-check: notes
  • Creating A Coral Task

    Human-Agent-Society/CORAL

    Author a new CORAL task — the three pieces that must line up (task.yaml, seed/, a packaged grader/), the coral init → coral validate → smoke-test loop, and how to pick a grader pattern (stdout…

    1k GitHub stars~2.2k tokensUpdated 29 days ago
    Testing & QAAuto-check passed
  • Launch Rl

    marin-community/marin

    Define, validate, submit, or restart a Marin SkyRL experiment through its artifact main.

    3.9k GitHub stars~894 tokensUpdated today
    Testing & QAAuto-check passed

More from vibeeval/vibecosystem

All 144 skills in this repo
  • Agent Benchmark

    vibeeval/vibecosystem

    Framework for measuring and tracking agent response quality over time.

    531 GitHub stars~2.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Differential Review

    vibeeval/vibecosystem

    Security-focused differential code review with blast radius analysis, risk-adaptive depth (DEEP/FOCUSED/SURGICAL), git history correlation, and structured finding format.

    531 GitHub stars~1.6k tokensUpdated 2 mo ago
    Auto-check passed
  • Factcheck Guard

    vibeeval/vibecosystem

    A skill your agent uses when making any factual claim about the codebase — existence, absence, or behavior.

    531 GitHub stars~2.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Fp Check

    vibeeval/vibecosystem

    Systematic false positive verification for security findings.

    531 GitHub stars~1.6k tokensUpdated 2 mo ago
    Auto-check passed
  • N8n Workflows

    vibeeval/vibecosystem

    n8n otomasyon workflow'lari. An agent skill from vibeeval/vibecosystem.

    531 GitHub stars~3.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Notepad System

    vibeeval/vibecosystem

    A skill your agent uses when context compression is imminent, when resuming a session, or when preserving critical decisions across long tasks.

    531 GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed

Categories

Questions about Agent QA Testing

What does Agent QA Testing do?

Agent davranis testi ve protokol uyumluluk dogrulamasi. An agent skill from vibeeval/vibecosystem. Agent QA Testing is an agent skill from vibeeval/vibecosystem. Agent davranis testi ve protokol uyumluluk dogrulamasi.

When should I use Agent QA Testing?

Agent QA Testing fits situations like: tasks that involve QA and bug reports.

How do I install Agent QA Testing in Claude Code?

Run `npx skills add vibeeval/vibecosystem --skill agent-qa-testing -a claude-code`. Or copy the skill folder (skills/agent-qa-testing in vibeeval/vibecosystem) into .claude/skills/agent-qa-testing in your project. Claude Code loads it when a task matches its description.

How do I install Agent QA Testing in Codex?

Run `npx skills add vibeeval/vibecosystem --skill agent-qa-testing -a codex`. Or copy the skill folder (skills/agent-qa-testing in vibeeval/vibecosystem) into .agents/skills/agent-qa-testing in your project. Codex loads it when a task matches its description.

Can I use Agent QA Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vibeeval/vibecosystem --skill agent-qa-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-qa-testing, .gemini/skills/agent-qa-testing, .github/skills/agent-qa-testing and .opencode/skills/agent-qa-testing in your project.

What does Agent QA Testing need to run?

Going by SKILL.md and its folder, Agent QA Testing needs the command-line tools its instructions call (claude) and credentials named ANTHROPIC_API_KEY. Our summary lists: Node.js; A credential in ANTHROPIC_API_KEY.

Does Agent QA Testing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent QA Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent QA Testing use?

Agent QA Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent QA Testing use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent QA Testing?

Skills that share tags, products or a category with Agent QA Testing: Reproduce Chat States (different-ai/openwork, 24k stars), Dynamo Jira Ticket (DynamoDS/Dynamo, 2k stars), Minimal Run And Audit (lllllllama/RigorPilot-Skills, 497 stars) and Moav E2E (MotherofallVPNs/MoaV, 448 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent QA Testing?

vibeeval (a GitHub user) maintains it in vibeeval/vibecosystem, which has 531 GitHub stars. The repository holds 144 skills in this directory. The repository was last updated on August 8, 2026.

Source: vibeeval/vibecosystem on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.