Search
Hugging Face · LLM guardrails
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Explains how to train a model to be harmless with self-critique, revision and AI-generated preference feedback, with Hugging Face and TRL code for each stage. | Orchestra-Research/ | 13k | 2 repos | ~2k | Automated safety check: Pass | MIT | 3 mo ago |
| 2 | Uses Meta's LlamaGuard moderation model to screen prompts and model replies against six safety categories, with vLLM, FastAPI and NeMo Guardrails setups. | Orchestra-Research/ | 13k | 2 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 3 | Meta's 86M prompt injection and jailbreak detector. An agent skill from Orchestra-Research/AI-Research-SKILLs. | Orchestra-Research/ | 13k | 1 repo | ~2.4k | Automated safety check: Warn | MIT | 3 mo ago |
| 4 | Detect and defend against indirect prompt injection hidden in web pages, documents, and images consumed by an agent, via content extraction (HTML/PDF/OCR), normalization, and scanning with LLM… | mukul975/ | 34k | — | ~2.8k | Automated safety check: Warn | Apache-2.0 | 1 mo ago |
| 5 | Runs NVIDIA garak probe suites (jailbreak, prompt injection, data leakage, toxicity, and more) against an LLM endpoint - Hugging Face models, OpenAI-compatible APIs, or Bedrock - then interprets the… | mukul975/ | 34k | — | ~2.9k | Automated safety check: Warn | Apache-2.0 | 1 mo ago |