Search

AI & LLM Engineering · By tikalk

8 skills found.
Product:
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

A skill your agent uses when evaluation results need triage and loop-closing — spec failures route to deterministic checks or context rules, generalization failures to the evaluator backlog.

tikalk/adlc-team-skills141—~1.2kAutomated safety check: PassMITyesterday
2

A skill your agent uses when draft eval criteria need refining, clustering, and acceptance into the published goldset with an isolated holdout split (goldset.md + goldset.json).

tikalk/adlc-team-skills141—~1.1kAutomated safety check: PassMITyesterday
3

A skill your agent uses when a published goldset is ready to become executable graders and configs — Python graders and metrics with auto-run unit tests verifying grader correctness.

tikalk/adlc-team-skills141—~1.4kAutomated safety check: PassMITyesterday
4

A skill your agent uses when standing up evals/{system}/ for the first time — scaffolds the EDD directory structure, picks PromptFoo or DeepEval by tech stack, and generates a security baseline.

tikalk/adlc-team-skills141—~934Automated safety check: PassMITyesterday
5

A skill your agent uses when you want guided bottom-up error analysis, structured failure taxonomy discovery, or comprehensive trace coding before documenting eval criteria.

tikalk/adlc-team-skills141—~915Automated safety check: PassMITyesterday
6

A skill your agent uses when a goldset with graders is ready to run — executes the evaluation pyramid and validates evaluator quality (SLA compliance, TPR/TNR, statistical accuracy).

tikalk/adlc-team-skills141—~850Automated safety check: PassMITyesterday
7

A skill your agent uses when explicit manual re-discovery of team context modules is wanted beyond the injected CDR index (/team-discover) — produces a structured match table with relevance…

tikalk/adlc-team-skills141—~568Automated safety check: PassMITyesterday
8

A skill your agent uses when coordinating continuous improvement loops (team-levelup + change + evals feedback + cleanup) targeting team-ai-directives — includes build-to-delete pruning and…

tikalk/adlc-team-skills141—~1.5kAutomated safety check: PassMITyesterday