Search
OpenAI · LLM evaluation
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | A skill your agent uses when designing, generating an MVP blueprint for, auditing, troubleshooting, refactoring, or explaining an agentic harness for any domain. | DenisSergeevitch/ | 2.4k | — | ~7.4k | Automated safety check: Pass | MIT | 6 days ago |
| 2 | A skill your agent uses for Codex models/pricing, scheduled tasks, skills, settings, setup, troubleshooting, customization, automations, and self-knowledge—including 'you,' 'your,' 'this app,' or… | theowenyoung/ | 115 | 1 repo | ~1.4k | Automated safety check: Pass | Apache-2.0 | 14 days ago |
| 3 | Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes. | microsoft/ | 3.1k | — | ~2.8k | Automated safety check: Pass | MIT | yesterday |
| 4 | Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server. | Orchestra-Research/ | 13k | 2 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 5 | Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic. | agentailor/ | 132 | — | ~5.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 6 | Test a pre-built afm binary at any path — runs pre-flight safety checks, then any combination of unit tests, assertions, smart analysis, promptfoo evals, batch validation, OpenAI compat, GPU… | scouzi1966/ | 346 | — | ~3.8k | Automated safety check: Pass | MIT | yesterday |
| 7 | Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment. | Jeffallan/ | 12k | — | ~1.7k | Automated safety check: Pass | MIT | 7 days ago |
| 8 | Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step. | Jeffallan/ | 12k | — | ~2k | Automated safety check: Pass | MIT | 7 days ago |
| 9 | Evaluate WooAIAssistant against a structured scenario suite with hard invariants + LLM-as-judge rubric scoring. | woocommerce/ | 358 | — | ~7.4k | Automated safety check: Notes | GPL-2.0 | yesterday |
| 10 | Configures and runs LLM evaluation using Promptfoo framework. | daymade/ | 1.4k | — | ~3k | Automated safety check: Pass | MIT | yesterday |
| 11 | Tests/benchmarks a third-party LLM endpoint (OpenAI- or Anthropic-compatible): availability, fidelity, speed, concurrency, protocol compliance, quality regression. | daymade/ | 1.4k | — | ~4.7k | Automated safety check: Pass | MIT | yesterday |
| 12 | A skill your agent uses when designing, auditing, refactoring, or explaining an agentic harness for any domain, especially when work must continue from a measured gap to verified completion. | AnastasiyaW/ | 154 | — | ~5.4k | Automated safety check: Pass | MIT | yesterday |