Search
AI & LLM Engineering · By tikalk
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | A skill your agent uses when evaluation results need triage and loop-closing — spec failures route to deterministic checks or context rules, generalization failures to the evaluator backlog. | tikalk/ | 141 | — | ~1.2k | Automated safety check: Pass | MIT | yesterday |
| 2 | A skill your agent uses when draft eval criteria need refining, clustering, and acceptance into the published goldset with an isolated holdout split (goldset.md + goldset.json). | tikalk/ | 141 | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 3 | A skill your agent uses when a published goldset is ready to become executable graders and configs — Python graders and metrics with auto-run unit tests verifying grader correctness. | tikalk/ | 141 | — | ~1.4k | Automated safety check: Pass | MIT | yesterday |
| 4 | A skill your agent uses when standing up evals/{system}/ for the first time — scaffolds the EDD directory structure, picks PromptFoo or DeepEval by tech stack, and generates a security baseline. | tikalk/ | 141 | — | ~934 | Automated safety check: Pass | MIT | yesterday |
| 5 | A skill your agent uses when you want guided bottom-up error analysis, structured failure taxonomy discovery, or comprehensive trace coding before documenting eval criteria. | tikalk/ | 141 | — | ~915 | Automated safety check: Pass | MIT | yesterday |
| 6 | A skill your agent uses when a goldset with graders is ready to run — executes the evaluation pyramid and validates evaluator quality (SLA compliance, TPR/TNR, statistical accuracy). | tikalk/ | 141 | — | ~850 | Automated safety check: Pass | MIT | yesterday |
| 7 | A skill your agent uses when explicit manual re-discovery of team context modules is wanted beyond the injected CDR index (/team-discover) — produces a structured match table with relevance… | tikalk/ | 141 | — | ~568 | Automated safety check: Pass | MIT | yesterday |
| 8 | A skill your agent uses when coordinating continuous improvement loops (team-levelup + change + evals feedback + cleanup) targeting team-ai-directives — includes build-to-delete pruning and… | tikalk/ | 141 | — | ~1.5k | Automated safety check: Pass | MIT | yesterday |