Search

OpenAI · LLM evaluation

12 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

A skill your agent uses when designing, generating an MVP blueprint for, auditing, troubleshooting, refactoring, or explaining an agentic harness for any domain.

DenisSergeevitch/agents-best-practices2.4k—~7.4kAutomated safety check: PassMIT6 days ago
2

A skill your agent uses for Codex models/pricing, scheduled tasks, skills, settings, setup, troubleshooting, customization, automations, and self-knowledge—including 'you,' 'your,' 'this app,' or…

theowenyoung/home1151 repo~1.4kAutomated safety check: PassApache-2.014 days ago
3

Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

microsoft/skills3.1k—~2.8kAutomated safety check: PassMITyesterday
4

Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT3 mo ago
5

Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic.

agentailor/fullstack-langgraph-nextjs-agent132—~5.3kAutomated safety check: PassMIT1 mo ago
6

Test a pre-built afm binary at any path — runs pre-flight safety checks, then any combination of unit tests, assertions, smart analysis, promptfoo evals, batch validation, OpenAI compat, GPU…

scouzi1966/maclocal-api346—~3.8kAutomated safety check: PassMITyesterday
7

Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

Jeffallan/claude-skills12k—~1.7kAutomated safety check: PassMIT7 days ago
8

Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

Jeffallan/claude-skills12k—~2kAutomated safety check: PassMIT7 days ago
9

Evaluate WooAIAssistant against a structured scenario suite with hard invariants + LLM-as-judge rubric scoring.

woocommerce/woocommerce-ios358—~7.4kAutomated safety check: NotesGPL-2.0yesterday
10

Configures and runs LLM evaluation using Promptfoo framework.

daymade/claude-code-skills1.4k—~3kAutomated safety check: PassMITyesterday
11

Tests/benchmarks a third-party LLM endpoint (OpenAI- or Anthropic-compatible): availability, fidelity, speed, concurrency, protocol compliance, quality regression.

daymade/claude-code-skills1.4k—~4.7kAutomated safety check: PassMITyesterday
12

A skill your agent uses when designing, auditing, refactoring, or explaining an agentic harness for any domain, especially when work must continue from a measured gap to verified completion.

AnastasiyaW/codex-claude-code-config154—~5.4kAutomated safety check: PassMITyesterday