Search
Security · LLM guardrails
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Query AI Safety HOT news, research papers, incidents, hot topics, and daily/weekly/monthly reports through its public read-only MCP service. | wuyoscar/ | 827 | — | ~1.4k | Automated safety check: Pass | Unknown | today |
| 2 | Installs, tunes and enforces Sponsio contracts that block unsafe tool calls in LLM agents, covering setup, auditing, observe mode and flipping to enforce. | SponsioLabs/ | 454 | — | ~12k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 3 | 按中国法规(网安法 / PIPL / 等保2.0 / 数据出境 / AI生成内容标识)审计一个 AI 项目的代码仓库,产出每条都带 文件:行 取证、经独立复核、经脚本校验的合规报告。当用户问「这个项目上线合不合规」「调用了 OpenAI/Claude 算不算数据出境」「要不要做 AI 标识」「帮我做合规自查/等保/PIPL 检查」时使用。Audit an AI project's… | jnMetaCode/ | 140 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 4 | Installs hooks that check each agent action against security policies before it runs, blocking destructive commands and logging every decision. | pegasi-ai/ | 392 | — | ~1.4k | Automated safety check: Warn | Apache-2.0 | yesterday |
| 5 | Guide for writing eval conversation JSONs and running them through policy engines | open-bias/ | 143 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 6 | Uses Meta's LlamaGuard moderation model to screen prompts and model replies against six safety categories, with vLLM, FastAPI and NeMo Guardrails setups. | Orchestra-Research/ | 13k | 2 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 7 | Report/investigate RUNTIME ACTIVITY of AI agents (Agent 365 / Copilot Studio / M365 Copilot / Work IQ) — agents used, tools/connectors, channels, tokens, prompt/reply content, and Prompt Shield… | SCStelz/ | 249 | — | ~17k | Automated safety check: Pass | MIT | 2 days ago |
| 8 | OpenClaw 安全部署指南 / Security deployment guide — help users secure their OpenClaw installation | jnMetaCode/ | 140 | — | ~644 | Automated safety check: Warn | Apache-2.0 | 12 days ago |
| 9 | AI governance, EU AI Act compliance, OWASP LLM security, responsible AI practices for GitHub Copilot agents | Hack23/ | 239 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 10 | Guides designing a layered permission pipeline for agent tools that decides which calls are allowed, need confirmation or are denied, with scopes and hooks. | simbajigege/ | 183 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 11 | NVIDIA's runtime safety framework for LLM applications. An agent skill from Orchestra-Research/AI-Research-SKILLs. | Orchestra-Research/ | 13k | 2 repos | ~1.9k | Automated safety check: Warn | MIT | 3 mo ago |
| 12 | 12.Prompt Guard Meta's 86M prompt injection and jailbreak detector. An agent skill from Orchestra-Research/AI-Research-SKILLs. | Orchestra-Research/ | 13k | 1 repo | ~2.4k | Automated safety check: Warn | MIT | 3 mo ago |
| 13 | Authorized Android/iOS application reverse engineering and security testing: APK/IPA analysis, runtime instrumentation (Frida/Objection), SSL-pinning and jailbreak/root-detection bypass, per OWASP… | sickn33/ | 47k | 1 repo | ~1.5k | Automated safety check: Pass | MIT | 2 days ago |
| 14 | Wires Promptfoo and DeepTeam into CI/CD for automated, repeatable red-teaming of LLM apps against OWASP LLM Top 10, OWASP Agentic, and MITRE ATLAS presets, failing the build when jailbreak or… | mukul975/ | 34k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 15 | Application security defense knowledge for builders. An agent skill from telagod/code-abyss. | telagod/ | 244 | — | ~777 | Automated safety check: Pass | MIT | 2 mo ago |
| 16 | Defend AI systems against prompt injection and indirect prompt attacks using input controls, tool permissions, output validation, and isolation boundaries. | sickn33/ | 47k | 2 repos | ~4.2k | Automated safety check: Warn | MIT | 2 days ago |
| 17 | AI/LLM defensive security reference: prompt-injection defense, OWASP LLM Top 10 defensive mapping, MCP and agentic tool-call hardening, training-data poisoning detection, model-output validation and… | modu-ai/ | 1.2k | — | ~4.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 18 | Cloud posture security across AWS, Azure, and GCP — IAM least privilege, public exposure, encryption, logging coverage, landing-zone guardrails. | borghei/ | 891 | — | ~3.5k | Automated safety check: Pass | MIT | 4 days ago |
| 19 | Detect and defend against indirect prompt injection hidden in web pages, documents, and images consumed by an agent, via content extraction (HTML/PDF/OCR), normalization, and scanning with LLM… | mukul975/ | 34k | — | ~2.8k | Automated safety check: Warn | Apache-2.0 | 1 mo ago |
| 20 | Runs NVIDIA garak probe suites (jailbreak, prompt injection, data leakage, toxicity, and more) against an LLM endpoint - Hugging Face models, OpenAI-compatible APIs, or Bedrock - then interprets the… | mukul975/ | 34k | — | ~2.9k | Automated safety check: Warn | Apache-2.0 | 1 mo ago |
| 21 | Deploys Llama Guard 3 safety classification, NeMo Guardrails programmable dialogue rails, and LLM Guard input/output scanner pipelines as complementary runtime defenses that inspect and constrain… | mukul975/ | 34k | — | ~3.1k | Automated safety check: Warn | Apache-2.0 | 1 mo ago |
| 22 | 22.AI Security A skill your agent uses when assessing AI/ML systems for prompt injection, jailbreak vulnerabilities, model inversion risk, data poisoning exposure, or agent tool abuse. | alirezarezvani/ | 28k | — | ~4.5k | Automated safety check: Warn | MIT | 1 mo ago |
| 23 | Catalogue of prompt framings that determine whether an agent refuses or performs specification search, and the harness for probing them. | brycewang-stanford/ | 4.6k | — | ~1.4k | Automated safety check: Pass | Unknown | 5 days ago |
| 24 | 24.Polygraph Behavioral trust grades (A–F) for MCP servers. An agent skill from BankrBot/skills. | BankrBot/ | 1.2k | — | ~3.4k | Automated safety check: Pass | No licence | yesterday |
| 25 | Proactively harden a cloud account or organization before an incident — prioritizing IAM and identity risk over checkbox findings, closing the exposures that become attack paths (public storage… | trilwu/ | 157 | — | ~1.9k | Automated safety check: Pass | MIT | 1 mo ago |
| 26 | Enforce organizational governance for Supabase projects: shared RLS policy library with reusable templates, table and column naming conventions, migration review process with CI checks, cost alert… | jeremylongshore/ | 2.8k | — | ~2.8k | Automated safety check: Pass | MIT | yesterday |
| 27 | Apply Anthropic Claude API security best practices for key management, input validation, and prompt injection defense. | jeremylongshore/ | 2.8k | — | ~1.9k | Automated safety check: Notes | MIT | yesterday |
| 28 | Specify the safety and reliability guardrails for an LLM feature before it ships. | mohitagw15856/ | 1.4k | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |