Agent skill

Red Teaming LLMs With Garak

by mukul975 in mukul975/Anthropic-Cybersecurity-Skills

Runs NVIDIA garak probe suites (jailbreak, prompt injection, data leakage, toxicity, and more) against an LLM endpoint - Hugging Face models, OpenAI-compatible APIs, or Bedrock - then interprets the…

Apache-2.0Auto-check: warningsSecurity

Install Red Teaming LLMs With Garak

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill red-teaming-llms-with-garak -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mukul975/Anthropic-Cybersecurity-Skills red-teaming-llms-with-garak --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mukul975/Anthropic-Cybersecurity-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/red-teaming-llms-with-garak .claude/skills/red-teaming-llms-with-garak && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
red-teaming-llms-with-garak
GitHub stars
34k
Token cost
~2.9k tokens
SKILL.md length
1,103 words
Files
5 (incl. scripts, references)
Skills in repo
644
Repo updated
First seen
Licence
Apache-2.0

At a glance

Runs NVIDIA garak probe suites (jailbreak, prompt injection, data leakage, toxicity, and more) against an LLM endpoint - Hugging Face models, OpenAI-compatible APIs, or Bedrock - then interprets the…

  • Works in 7 steps: Enumerate Probes and Detectors → Probe a Local Hugging Face Model → Probe an OpenAI-Compatible API → …
  • Baselining LLM security before/after deployment
  • SKILL.md covers Overview, When to Use, Prerequisites and Objectives, plus 5 more sections
  • Runs Python scripts from its folder; calls python and jq; reaches github.com; needs OPENAI_API_KEY and ENV_TOKEN

What it does

Red Teaming LLMs With Garak is an agent skill from mukul975/Anthropic-Cybersecurity-Skills. Runs NVIDIA garak probe suites (jailbreak, prompt injection, data leakage, toxicity, and more) against an LLM endpoint - Hugging Face models, OpenAI-compatible APIs, or Bedrock - then interprets the resulting hit-rate report for triage. Use when baselining LLM security before/after deployment, validating that a guardrail or fine-tune reduces jailbreak/injection success rates, or producing evidence for an AI risk assessment.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/api-reference.md`, `references/standards.md` and `scripts/agent.py`).

It sits in Security, covering Prompt injection and agent security, Red teaming and adversary simulation and LLM guardrails. It works with Hugging Face, OpenAI and NVIDIA AI Platform. The repository describes itself as: 817 structured cybersecurity skills for AI agents · Mapped to 6 frameworks: MITRE ATT&CK, NIST CSF 2.0, MITRE ATLAS, D3FEND, NIST AI RMF & MITRE F3 (Fight Fraud) · agentskills.io…. The licence is Apache-2.0.

When your agent uses it

  • Baselining LLM security before/after deployment
  • Validating that a guardrail
  • Fine-tune reduces jailbreak/injection success rates
  • Producing evidence for an AI risk assessment

Example prompts

  • “/red-teaming-llms-with-garak”

Requirements

  • Python 3
  • A credential in OPENAI_API_KEY
  • A credential in ENV_TOKEN

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Enumerate Probes and Detectors
  2. Probe a Local Hugging Face Model
  3. Probe an OpenAI-Compatible API
  4. Probe an Arbitrary REST Endpoint
  5. Run a Curated Config and Full Sweep
  6. Interpret the Hit-Rate Report
  7. Mitigate and Re-Test

What it can do on your machine

Read from SKILL.md and the folder at commit 54a7988. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    Also links to:

    • docs.garak.ai
    • reference.garak.ai
    • arxiv.org
    • genai.owasp.org
    • atlas.mitre.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY
    • ENV_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Red Teaming LLMs With Garak loads about 2.9k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 114 tokens; SKILL.md has 1,103 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~114
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningContains instruction-override wording (e.g. “without asking the user”)SKILL.md:81
    k's `dan` and related probes attempt to bypass safety guardrails so the model produces restricted content. |

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from mukul975/Anthropic-Cybersecurity-Skills at commit 54a7988, republished under its Apache-2.0 licence (© mukul975). 1,103 words, ~2,872 tokens.

Download SKILL.mdSave it as .claude/skills/red-teaming-llms-with-garak/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
red-teaming-llms-with-garak
description
Runs NVIDIA garak probe suites (jailbreak, prompt injection, data leakage, toxicity, and more) against an LLM endpoint - Hugging Face models, OpenAI-compatible APIs, or Bedrock - then interprets the resulting hit-rate report for triage. Use when baselining LLM security before/after deployment, validating that a guardrail or fine-tune reduces jailbreak/injection success rates, or producing evidence for an AI risk assessment.
domain
cybersecurity
subdomain
ai-security
tags
ai-security, llm-red-teaming, garak, prompt-injection, jailbreak, vulnerability-scanning, data-leakage, mitre-atlas
version
1.0
author
mahipal
license
Apache-2.0
nist_ai_rmf
MEASURE-2.7
atlas_techniques
AML.T0051, AML.T0054

Red-Teaming LLMs with garak

Legal and Authorized-Use Notice: This skill is for authorized AI security testing and educational purposes only. Probe only models, API keys, and endpoints you own or have explicit written permission to test. Automated probing of third-party LLM APIs may violate their terms of service and consume billable tokens. Unauthorized probing of systems you do not control may be illegal.

Overview

garak (Generative AI Red-teaming and Assessment Kit) is an open-source LLM vulnerability scanner maintained by NVIDIA. It plays the role that a network vulnerability scanner like Nessus plays for hosts, but for large language models: it sends thousands of adversarial prompts ("probes") at a target model, captures the generations, and runs automated "detectors" over the responses to decide whether each attempt succeeded. Probe families cover prompt injection (promptinject, latentinjection), jailbreaks (dan), training-data and system-prompt leakage (leakreplay), malware generation (malwaregen), cross-site-scripting payload emission (xss), encoding-based bypasses (encoding), toxicity, and more. garak is described in the paper "garak: A Framework for Security Probing Large Language Models" (arXiv:2406.11036) and is distributed from the NVIDIA/garak GitHub repository.

The scanner is generator-agnostic. It can target Hugging Face models loaded locally, OpenAI-compatible APIs, AWS Bedrock, Replicate, Cohere, NIM endpoints, GGUF/llama.cpp models, and arbitrary REST endpoints via a JSON generator spec. After a run, garak emits a .report.jsonl line-delimited log of every attempt and detector verdict, a human-readable .report.html, a garak.log debug log, and a hit log of confirmed vulnerabilities. The terminal output prints a per-probe, per-detector pass/fail summary with a hit rate (for example dan.Dan_11_0 jailbreak: FAIL ok on 38/40), which is the primary artifact you interpret.

This skill maps to the MITRE ATLAS techniques AML.T0051 (LLM Prompt Injection) and AML.T0054 (LLM Jailbreak) because garak operationalizes both: it crafts prompt-injection and jailbreak inputs at scale and measures whether the target's guardrails hold. It supports the NIST AI RMF MEASURE-2.7 subcategory by providing repeatable, quantitative security/resilience measurement of a deployed AI system.

When to Use

  • When you need a fast, repeatable baseline security assessment of an LLM before or after deployment.
  • When validating that a guardrail, system prompt, or safety fine-tune actually reduces jailbreak and injection success rates (run before/after and compare hit rates).
  • When producing evidence for an AI risk assessment or model card security section (NIST AI RMF MEASURE-2.7).
  • When triaging which OWASP LLM Top 10 risks (LLM01 prompt injection, LLM02 sensitive information disclosure, LLM07 system prompt leakage) actually manifest in your model.
  • When regression-testing an LLM endpoint in CI after model or prompt changes.

Prerequisites

  • Python 3.10+ (3.12 recommended) and a virtual environment.
  • Install garak from PyPI:
    bash
    python -m venv .venv && source .venv/bin/activate   # Windows: .venv\Scripts\activate
    python -m pip install -U garak
    garak --version
  • For the bleeding-edge version:
    bash
    python -m pip install -U git+https://github.com/NVIDIA/garak.git@main
  • An API key for the target if probing a hosted model (for example export OPENAI_API_KEY="sk-...").
  • Written authorization to test the target, and awareness of token cost (probes generate thousands of calls).

Objectives

  • Enumerate available probes and detectors and pick a scoped suite.
  • Configure garak against a local Hugging Face model, an OpenAI-compatible API, and an arbitrary REST endpoint.
  • Run prompt-injection, jailbreak, and leakage probe families and capture reports.
  • Interpret the per-probe hit-rate output and the .report.jsonl.
  • Re-run after applying a mitigation to demonstrate risk reduction.
  • Produce a defensible findings artifact for an AI risk assessment.

MITRE ATT&CK Mapping

This skill uses MITRE ATLAS (the adversarial-ML companion to ATT&CK) technique IDs.

IDTacticOfficial NameRelevance
AML.T0051ML Attack Staging / ImpactLLM Prompt Injectiongarak's promptinject and latentinjection probes craft malicious prompts that subvert intended model behavior.
AML.T0054Privilege Escalation / Defense EvasionLLM Jailbreakgarak's dan and related probes attempt to bypass safety guardrails so the model produces restricted content.

Workflow

Phase 1: Enumerate Probes and Detectors
  1. List every probe garak ships so you can scope the run:
    bash
    garak --list_probes
  2. List detectors (the modules that score whether a probe succeeded) and generators (target connectors):
    bash
    garak --list_detectors
    garak --list_generators
  3. Read the probe taxonomy. Key families:
    • promptinject — PromptInject-framework direct injection.
    • latentinjection — instructions hidden in documents/encoded text (indirect injection).
    • dan — "Do Anything Now" and related jailbreaks (e.g. dan.Dan_11_0).
    • leakreplay — coax the model into reproducing memorized/training or hidden-prompt text.
    • encoding — base64/ROT13/etc. injection bypasses.
    • xss — emit cross-site-scripting payloads (markdown/HTML exfil).
    • malwaregen — request AV-evading or malicious code.
Phase 2: Probe a Local Hugging Face Model
  1. Run a single jailbreak probe against a local model to validate setup:
    bash
    python -m garak --target_type huggingface --target_name gpt2 --probes dan.Dan_11_0
  2. Run a fuller suite against a chat model:
    bash
    python -m garak \
      --target_type huggingface \
      --target_name meta-llama/Llama-3.2-1B-Instruct \
      --probes promptinject,dan,leakreplay \
      --report_prefix llama32_baseline
Show full SKILL.md (490 more words)Show less
Phase 3: Probe an OpenAI-Compatible API
  1. Export the key and run injection + leakage probes:
    bash
    export OPENAI_API_KEY="sk-..."
    python -m garak \
      --target_type openai \
      --target_name gpt-4o-mini \
      --probes promptinject,latentinjection,leakreplay \
      --generations 5 \
      --parallel_attempts 8 \
      --report_prefix gpt4omini_injection
    • --generations controls how many completions per prompt (more = more statistical confidence, more cost).
    • --parallel_attempts raises throughput for remote APIs.
Phase 4: Probe an Arbitrary REST Endpoint
  1. garak can target any HTTP API via a JSON generator spec. Create rest.json:
    json
    {
      "rest": {
        "RestGenerator": {
          "name": "my-llm-gateway",
          "uri": "https://llm.internal.example/v1/chat",
          "method": "post",
          "headers": { "Authorization": "Bearer $ENV_TOKEN", "Content-Type": "application/json" },
          "req_template_json_object": { "model": "internal-bot", "prompt": "$INPUT" },
          "response_json": true,
          "response_json_field": "$.output"
        }
      }
    }
  2. Run garak against it:
    bash
    export ENV_TOKEN="..."
    python -m garak \
      --target_type rest \
      -G rest.json \
      --probes promptinject,dan \
      --report_prefix internal_gateway
Phase 5: Run a Curated Config and Full Sweep
  1. For repeatable assessments, pin everything in a YAML/JSON config and pass --config:
    bash
    python -m garak --config assessment.yaml
    yaml
    # assessment.yaml
    plugins:
      model_type: openai
      model_name: gpt-4o-mini
      probe_spec: promptinject,latentinjection,dan,leakreplay,xss,malwaregen
    run:
      generations: 5
      parallel_attempts: 8
    reporting:
      report_prefix: quarterly_llm_assessment
  2. For an exhaustive sweep (slow, expensive) run all probes by omitting --probes entirely.
Phase 6: Interpret the Hit-Rate Report
  1. Read the terminal summary. Each row is probe.Class detector: PASS|FAIL ok on N/M. A FAIL with a low ok fraction means the model frequently produced the unsafe behavior — a high-severity finding.
  2. Open the machine-readable report and aggregate failures:
    bash
    # Every attempt with detector verdicts is one JSON line
    jq -r 'select(.entry_type=="eval") | "\(.probe)\t\(.detector)\t\(.passed)/\(.total)"' \
      garak.<timestamp>.report.jsonl | sort
  3. Open the generated .report.html in a browser for the formatted scorecard and per-probe breakdown.
  4. Pull the actual successful attack strings from the hit log to use as proof-of-concept evidence.
Phase 7: Mitigate and Re-Test
  1. Apply a control (tighten the system prompt, add an input/output guardrail such as Llama Guard or LLM Guard, or change the model).
  2. Re-run the identical probe set with a new --report_prefix.
  3. Compare hit rates between runs to quantify risk reduction for the report.

Tools and Resources

ResourcePurposeLink
NVIDIA/garakSource, probe list, issueshttps://github.com/NVIDIA/garak
garak documentationCLI reference, generator configshttps://docs.garak.ai/ and https://reference.garak.ai/
garak paper (arXiv:2406.11036)Methodology and designhttps://arxiv.org/abs/2406.11036
OWASP Top 10 for LLM ApplicationsRisk taxonomy probes map tohttps://genai.owasp.org/
MITRE ATLASAML technique definitionshttps://atlas.mitre.org/

Probe Family Reference

Probe familyTargetsOWASP LLM mapping
promptinjectDirect prompt injectionLLM01
latentinjectionIndirect / hidden-context injectionLLM01
danJailbreak / guardrail bypassLLM01 / safety
leakreplayTraining-data & prompt leakageLLM02 / LLM07
encodingEncoding-based filter bypassLLM01
xssMarkdown/HTML exfiltration payloadsLLM02
malwaregenMalicious code generationmisuse

Validation Criteria

  • garak installed and garak --version succeeds.
  • Probes and detectors enumerated with --list_probes / --list_detectors.
  • At least one probe run completed against the target with a --report_prefix set.
  • .report.jsonl, .report.html, and garak.log produced and located.
  • Per-probe hit rates extracted and ranked by severity.
  • Successful attack strings captured from the hit log as evidence.
  • A mitigation applied and a comparison re-run completed showing changed hit rates.
  • Findings documented against OWASP LLM Top 10 and MITRE ATLAS for the risk assessment.

© mukul975, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in skills/red-teaming-llms-with-garak of mukul975/Anthropic-Cybersecurity-Skills.

  • SKILL.md
  • LICENSE
  • references/api-reference.md
  • references/standards.md
  • scripts/agent.py

Open the folder on GitHubat commit 54a7988

Compare with similar skills

Red Teaming LLMs With Garak next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Red Teaming LLMs With Garak compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Red Teaming LLMs With Garak this skillmukul975/Anthropic-Cybersecurity-Skills34k—~2.9kAutomated safety check: WarnApache-2.0
China AI Compliance AuditjnMetaCode/shellward140—~1.1kAutomated safety check: PassApache-2.0
LlamaGuard Content ModerationOrchestra-Research/AI-Research-SKILLs13k2 repos~2.3kAutomated safety check: PassMIT
Dataset Transformationawslabs/agent-plugins9161 repos~3.5kAutomated safety check: PassApache-2.0
9Router Speech-to-Textdecolua/9router31k—~914Automated safety check: PassMIT
Writing Eval Scenariosopen-bias/open-bias143—~1.5kAutomated safety check: PassApache-2.0

Similar skills

  • China AI Compliance Audit

    jnMetaCode/shellward

    按中国法规(网安法 / PIPL / 等保2.0 / 数据出境 / AI生成内容标识)审计一个 AI 项目的代码仓库,产出每条都带 文件:行 取证、经独立复核、经脚本校验的合规报告。当用户问「这个项目上线合不合规」「调用了 OpenAI/Claude 算不算数据出境」「要不要做 AI 标识」「帮我做合规自查/等保/PIPL 检查」时使用。Audit an AI project's…

    140 GitHub stars~1.1k tokensUpdated 12 days ago
    SecurityAuto-check passed
  • LlamaGuard Content Moderation

    Orchestra-Research/AI-Research-SKILLs

    Uses Meta's LlamaGuard moderation model to screen prompts and model replies against six safety categories, with vLLM, FastAPI and NeMo Guardrails setups.

    13k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Dataset Transformation

    awslabs/agent-plugins

    Official

    Generates code that transforms datasets between ML schemas for model training or evaluation.

    916 GitHub starsUsed in 1 repo~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • 9Router Speech-to-Text

    decolua/9router

    Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

    31k GitHub stars~914 tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Writing Eval Scenarios

    open-bias/open-bias

    Guide for writing eval conversation JSONs and running them through policy engines

    143 GitHub stars~1.5k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • AI News

    letta-ai/skills

    Fetch and summarize recent AI news from curated RSS feeds (Hugging Face, VentureBeat, The Verge, OpenAI, Anthropic, DeepMind, etc.) and YouTube channels (Yannic Kilcher, Two Minute Papers, AI…

    149 GitHub stars~582 tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check passed

More from mukul975/Anthropic-Cybersecurity-Skills

All 644 skills in this repo
  • Campaign Attribution Evidence Analysis

    mukul975/Anthropic-Cybersecurity-Skills

    Weighs infrastructure, TTP, malware code and timing evidence with the Diamond Model and competing hypotheses to reach a confidence-rated attribution.

    34k GitHub stars~2.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Go Malware Analysis in Ghidra

    mukul975/Anthropic-Cybersecurity-Skills

    Walks through reverse engineering Go-compiled malware in Ghidra: parsing buildinfo and pclntab, recovering stripped function names and extracting dependencies.

    34k GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed
  • LNK and Jump List Forensics

    mukul975/Anthropic-Cybersecurity-Skills

    Guides forensic analysis of Windows LNK shortcut files and Jump Lists with LECmd, JLECmd and manual parsing to show file access and program execution.

    34k GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Malware Persistence Analysis with Autoruns

    mukul975/Anthropic-Cybersecurity-Skills

    Hunts Windows malware persistence with Sysinternals Autoruns, covering run keys, services, scheduled tasks and drivers, with baseline comparison.

    34k GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed
  • NTFS MFT Deleted File Recovery

    mukul975/Anthropic-Cybersecurity-Skills

    Guides a Windows forensic examination of the NTFS Master File Table to recover deleted-file evidence, build timelines and spot timestomping.

    34k GitHub stars~2.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Network Covert Channel Analysis

    mukul975/Anthropic-Cybersecurity-Skills

    Detects DNS tunneling, ICMP exfiltration and HTTP-based covert channels in packet captures and DNS logs when hunting for hidden command-and-control traffic.

    34k GitHub stars~2k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Red Teaming LLMs With Garak

What does Red Teaming LLMs With Garak do?

Runs NVIDIA garak probe suites (jailbreak, prompt injection, data leakage, toxicity, and more) against an LLM endpoint - Hugging Face models, OpenAI-compatible APIs, or Bedrock - then interprets the…. Red Teaming LLMs With Garak is an agent skill from mukul975/Anthropic-Cybersecurity-Skills. Runs NVIDIA garak probe suites (jailbreak, prompt injection, data leakage, toxicity, and more) against an LLM endpoint - Hugging Face models, OpenAI-compatible APIs, or Bedrock - then interprets the resulting hit-rate report for triage.

When should I use Red Teaming LLMs With Garak?

Red Teaming LLMs With Garak fits situations like: baselining LLM security before/after deployment; validating that a guardrail; fine-tune reduces jailbreak/injection success rates; producing evidence for an AI risk assessment.

How do I install Red Teaming LLMs With Garak in Claude Code?

Run `npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill red-teaming-llms-with-garak -a claude-code`. Or copy the skill folder (skills/red-teaming-llms-with-garak in mukul975/Anthropic-Cybersecurity-Skills) into .claude/skills/red-teaming-llms-with-garak in your project. Claude Code loads it when a task matches its description.

How do I install Red Teaming LLMs With Garak in Codex?

Run `npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill red-teaming-llms-with-garak -a codex`. Or copy the skill folder (skills/red-teaming-llms-with-garak in mukul975/Anthropic-Cybersecurity-Skills) into .agents/skills/red-teaming-llms-with-garak in your project. Codex loads it when a task matches its description.

Can I use Red Teaming LLMs With Garak in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill red-teaming-llms-with-garak -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/red-teaming-llms-with-garak, .gemini/skills/red-teaming-llms-with-garak, .github/skills/red-teaming-llms-with-garak and .opencode/skills/red-teaming-llms-with-garak in your project.

What does Red Teaming LLMs With Garak need to run?

Going by SKILL.md and its folder, Red Teaming LLMs With Garak needs Python for the scripts in its folder, the command-line tools its instructions call (python and jq) and credentials named OPENAI_API_KEY and ENV_TOKEN. Our summary lists: Python 3; A credential in OPENAI_API_KEY; A credential in ENV_TOKEN.

Does Red Teaming LLMs With Garak access the network?

SKILL.md names 6 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: docs.garak.ai, reference.garak.ai, arxiv.org, genai.owasp.org and atlas.mitre.org. This is read from the text; nothing was executed.

Is Red Teaming LLMs With Garak safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): contains instruction-override wording (e.g. “without asking the user”). Read the flagged lines before installing; the check is not a guarantee either way. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Red Teaming LLMs With Garak use?

Red Teaming LLMs With Garak is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Red Teaming LLMs With Garak use?

About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.1k tokens, read only when the agent opens those files.

What are the alternatives to Red Teaming LLMs With Garak?

Skills that share tags, products or a category with Red Teaming LLMs With Garak: China AI Compliance Audit (jnMetaCode/shellward, 140 stars), LlamaGuard Content Moderation (Orchestra-Research/AI-Research-SKILLs, 13k stars), Dataset Transformation (awslabs/agent-plugins, 916 stars) and 9Router Speech-to-Text (decolua/9router, 31k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Red Teaming LLMs With Garak?

mukul975 (a GitHub user) maintains it in mukul975/Anthropic-Cybersecurity-Skills, which has 34,116 GitHub stars. The repository holds 644 skills in this directory. The repository was last updated on August 31, 2026.

Source: mukul975/Anthropic-Cybersecurity-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.