Search

Hugging Face · LLM guardrails

5 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Explains how to train a model to be harmless with self-critique, revision and AI-generated preference feedback, with Hugging Face and TRL code for each stage.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2kAutomated safety check: PassMIT3 mo ago
2

Uses Meta's LlamaGuard moderation model to screen prompts and model replies against six safety categories, with vLLM, FastAPI and NeMo Guardrails setups.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.3kAutomated safety check: PassMIT3 mo ago
3

Meta's 86M prompt injection and jailbreak detector. An agent skill from Orchestra-Research/AI-Research-SKILLs.

Orchestra-Research/AI-Research-SKILLs13k1 repo~2.4kAutomated safety check: WarnMIT3 mo ago
4

Detect and defend against indirect prompt injection hidden in web pages, documents, and images consumed by an agent, via content extraction (HTML/PDF/OCR), normalization, and scanning with LLM…

mukul975/Anthropic-Cybersecurity-Skills34k—~2.8kAutomated safety check: WarnApache-2.01 mo ago
5

Runs NVIDIA garak probe suites (jailbreak, prompt injection, data leakage, toxicity, and more) against an LLM endpoint - Hugging Face models, OpenAI-compatible APIs, or Bedrock - then interprets the…

mukul975/Anthropic-Cybersecurity-Skills34k—~2.9kAutomated safety check: WarnApache-2.01 mo ago