Search

Security · LLM guardrails

28 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Query AI Safety HOT news, research papers, incidents, hot topics, and daily/weekly/monthly reports through its public read-only MCP service.

wuyoscar/AISafetyHot-Hub827—~1.4kAutomated safety check: PassUnknowntoday
2

Installs, tunes and enforces Sponsio contracts that block unsafe tool calls in LLM agents, covering setup, auditing, observe mode and flipping to enforce.

SponsioLabs/Sponsio454—~12kAutomated safety check: PassApache-2.0yesterday
3

按中国法规(网安法 / PIPL / 等保2.0 / 数据出境 / AI生成内容标识)审计一个 AI 项目的代码仓库,产出每条都带 文件:行 取证、经独立复核、经脚本校验的合规报告。当用户问「这个项目上线合不合规」「调用了 OpenAI/Claude 算不算数据出境」「要不要做 AI 标识」「帮我做合规自查/等保/PIPL 检查」时使用。Audit an AI project's…

jnMetaCode/shellward140—~1.1kAutomated safety check: PassApache-2.012 days ago
4

Installs hooks that check each agent action against security policies before it runs, blocking destructive commands and logging every decision.

pegasi-ai/reins392—~1.4kAutomated safety check: WarnApache-2.0yesterday
5

Guide for writing eval conversation JSONs and running them through policy engines

open-bias/open-bias143—~1.5kAutomated safety check: PassApache-2.04 days ago
6

Uses Meta's LlamaGuard moderation model to screen prompts and model replies against six safety categories, with vLLM, FastAPI and NeMo Guardrails setups.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.3kAutomated safety check: PassMIT3 mo ago
7

Report/investigate RUNTIME ACTIVITY of AI agents (Agent 365 / Copilot Studio / M365 Copilot / Work IQ) — agents used, tools/connectors, channels, tokens, prompt/reply content, and Prompt Shield…

SCStelz/security-investigator249—~17kAutomated safety check: PassMIT2 days ago
8

OpenClaw 安全部署指南 / Security deployment guide — help users secure their OpenClaw installation

jnMetaCode/shellward140—~644Automated safety check: WarnApache-2.012 days ago
9

AI governance, EU AI Act compliance, OWASP LLM security, responsible AI practices for GitHub Copilot agents

Hack23/cia239—~1.4kAutomated safety check: PassApache-2.0today
10

Guides designing a layered permission pipeline for agent tools that decides which calls are allowed, need confirmation or are denied, with scopes and hooks.

simbajigege/book2skills183—~2.1kAutomated safety check: PassApache-2.01 mo ago
11

NVIDIA's runtime safety framework for LLM applications. An agent skill from Orchestra-Research/AI-Research-SKILLs.

Orchestra-Research/AI-Research-SKILLs13k2 repos~1.9kAutomated safety check: WarnMIT3 mo ago
12

Meta's 86M prompt injection and jailbreak detector. An agent skill from Orchestra-Research/AI-Research-SKILLs.

Orchestra-Research/AI-Research-SKILLs13k1 repo~2.4kAutomated safety check: WarnMIT3 mo ago
13

Authorized Android/iOS application reverse engineering and security testing: APK/IPA analysis, runtime instrumentation (Frida/Objection), SSL-pinning and jailbreak/root-detection bypass, per OWASP…

sickn33/agentic-awesome-skills47k1 repo~1.5kAutomated safety check: PassMIT2 days ago
14

Wires Promptfoo and DeepTeam into CI/CD for automated, repeatable red-teaming of LLM apps against OWASP LLM Top 10, OWASP Agentic, and MITRE ATLAS presets, failing the build when jailbreak or…

mukul975/Anthropic-Cybersecurity-Skills34k—~2.5kAutomated safety check: PassApache-2.01 mo ago
15

Application security defense knowledge for builders. An agent skill from telagod/code-abyss.

telagod/code-abyss244—~777Automated safety check: PassMIT2 mo ago
16

Defend AI systems against prompt injection and indirect prompt attacks using input controls, tool permissions, output validation, and isolation boundaries.

sickn33/agentic-awesome-skills47k2 repos~4.2kAutomated safety check: WarnMIT2 days ago
17

AI/LLM defensive security reference: prompt-injection defense, OWASP LLM Top 10 defensive mapping, MCP and agentic tool-call hardening, training-data poisoning detection, model-output validation and…

modu-ai/moai-adk1.2k—~4.5kAutomated safety check: PassApache-2.0yesterday
18

Cloud posture security across AWS, Azure, and GCP — IAM least privilege, public exposure, encryption, logging coverage, landing-zone guardrails.

borghei/Claude-Skills891—~3.5kAutomated safety check: PassMIT4 days ago
19

Detect and defend against indirect prompt injection hidden in web pages, documents, and images consumed by an agent, via content extraction (HTML/PDF/OCR), normalization, and scanning with LLM…

mukul975/Anthropic-Cybersecurity-Skills34k—~2.8kAutomated safety check: WarnApache-2.01 mo ago
20

Runs NVIDIA garak probe suites (jailbreak, prompt injection, data leakage, toxicity, and more) against an LLM endpoint - Hugging Face models, OpenAI-compatible APIs, or Bedrock - then interprets the…

mukul975/Anthropic-Cybersecurity-Skills34k—~2.9kAutomated safety check: WarnApache-2.01 mo ago
21

Deploys Llama Guard 3 safety classification, NeMo Guardrails programmable dialogue rails, and LLM Guard input/output scanner pipelines as complementary runtime defenses that inspect and constrain…

mukul975/Anthropic-Cybersecurity-Skills34k—~3.1kAutomated safety check: WarnApache-2.01 mo ago
22

A skill your agent uses when assessing AI/ML systems for prompt injection, jailbreak vulnerabilities, model inversion risk, data poisoning exposure, or agent tool abuse.

alirezarezvani/claude-skills28k—~4.5kAutomated safety check: WarnMIT1 mo ago
23

Catalogue of prompt framings that determine whether an agent refuses or performs specification search, and the harness for probing them.

brycewang-stanford/Auto-Empirical-Research-Skills4.6k—~1.4kAutomated safety check: PassUnknown5 days ago
24

Behavioral trust grades (A–F) for MCP servers. An agent skill from BankrBot/skills.

BankrBot/skills1.2k—~3.4kAutomated safety check: PassNo licenceyesterday
25

Proactively harden a cloud account or organization before an incident — prioritizing IAM and identity risk over checkbox findings, closing the exposures that become attack paths (public storage…

trilwu/secskills157—~1.9kAutomated safety check: PassMIT1 mo ago
26

Enforce organizational governance for Supabase projects: shared RLS policy library with reusable templates, table and column naming conventions, migration review process with CI checks, cost alert…

jeremylongshore/tons-of-skills-marketplace2.8k—~2.8kAutomated safety check: PassMITyesterday
27

Apply Anthropic Claude API security best practices for key management, input validation, and prompt injection defense.

jeremylongshore/tons-of-skills-marketplace2.8k—~1.9kAutomated safety check: NotesMITyesterday
28

Specify the safety and reliability guardrails for an LLM feature before it ships.

mohitagw15856/pm-claude-skills1.4k—~1.1kAutomated safety check: PassMITyesterday