Running Tests
brendanhasz/probflow
Run Python unit test suites strictly using the uv package manager and pytest.
A skill your agent uses when a published goldset is ready to become executable graders and configs — Python graders and metrics with auto-run unit tests verifying grader correctness.
$ npx skills add tikalk/adlc-team-skills --skill evals-implement -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install tikalk/adlc-team-skills evals-implement --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/tikalk/adlc-team-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evals/evals-implement .claude/skills/evals-implement && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "evals-implement" agent skill from https://github.com/tikalk/adlc-team-skills/tree/main/skills/evals/evals-implement into .claude/skills/evals-implement/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evals-implement", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/tikalk/adlc-team-skills/tree/main/skills/evals/evals-implementType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add tikalk/adlc-team-skills --skill evals-implement -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install tikalk/adlc-team-skills evals-implement --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tikalk/adlc-team-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/evals/evals-implement .agents/skills/evals-implement && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "evals-implement" agent skill from https://github.com/tikalk/adlc-team-skills/tree/main/skills/evals/evals-implement into .agents/skills/evals-implement/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evals-implement", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add tikalk/adlc-team-skills --skill evals-implement -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install tikalk/adlc-team-skills evals-implement --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tikalk/adlc-team-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/evals/evals-implement .cursor/skills/evals-implement && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "evals-implement" agent skill from https://github.com/tikalk/adlc-team-skills/tree/main/skills/evals/evals-implement into .cursor/skills/evals-implement/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evals-implement", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/tikalk/adlc-team-skills.git --path skills/evals/evals-implement--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add tikalk/adlc-team-skills --skill evals-implement -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install tikalk/adlc-team-skills evals-implement --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tikalk/adlc-team-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/evals/evals-implement .gemini/skills/evals-implement && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "evals-implement" agent skill from https://github.com/tikalk/adlc-team-skills/tree/main/skills/evals/evals-implement into .gemini/skills/evals-implement/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evals-implement", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install tikalk/adlc-team-skills evals-implementInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add tikalk/adlc-team-skills --skill evals-implement -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/tikalk/adlc-team-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/evals/evals-implement .github/skills/evals-implement && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "evals-implement" agent skill from https://github.com/tikalk/adlc-team-skills/tree/main/skills/evals/evals-implement into .github/skills/evals-implement/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evals-implement", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add tikalk/adlc-team-skills --skill evals-implement -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install tikalk/adlc-team-skills evals-implement --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tikalk/adlc-team-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/evals/evals-implement .opencode/skills/evals-implement && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "evals-implement" agent skill from https://github.com/tikalk/adlc-team-skills/tree/main/skills/evals/evals-implement into .opencode/skills/evals-implement/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evals-implement", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
evals-implementA skill your agent uses when a published goldset is ready to become executable graders and configs — Python graders and metrics with auto-run unit tests verifying grader correctness.
Evals Implement is an agent skill from tikalk/adlc-team-skills. Use when a published goldset is ready to become executable graders and configs — Python graders and metrics with auto-run unit tests verifying grader correctness.
Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `scripts/bash/setup-evals-implement.sh`).
It sits in Testing & QA, covering LLM evaluation and Unit testing. It works with Python. The repository describes itself as: Agent skills for the Agentic SDLC: team lifecycle (team-boot, team-learn, team-init, team-repair), software factory, evals, CDR lifecycle with confidence scoring, and… The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 2dbed36. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Shell and PowerShell), which the agent can run.
Shell commands in SKILL.md call:
pytestFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Evals Implement loads about 1.4k tokens when it runs. Until then it costs about 45 tokens; SKILL.md has 628 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from tikalk/adlc-team-skills at commit 2dbed36, republished under its MIT licence (© tikalk). 628 words, ~1,351 tokens.
.claude/skills/evals-implement/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Generates the complete executable evaluation implementation following EDD Principle VIII (Close Production Loop) from the published goldset, with automated unit testing to verify evaluator correctness.
Output:
evals/{system}/graders/BaseMetricevals/{system}/tests/test_check_*.py) that run the goldset pass/fail examples against the generated graders to ensure the evaluator itself is accurateconfig.js or config.py) with Tier 1 + Tier 2 evaluation structure/evals-validate to run validationKey EDD Principles Applied:
/evals-clarify: Convert accepted goldset criteria into executable code/evals-clarify to generate goldset.json first/evals-validate to run the suite against application outputs$ARGUMENTS--system SYSTEM — Override active evaluation framework (promptfoo or deepeval)--no-tests — Skip automated unit test generation for graders (not recommended)evals/{system}/goldset.json.pass_condition and fail_condition as the grader's core rubric.Root Cause Analysis and axial_coding notes as contextual prompt guidelines or regex patterns to catch exact failure manifestations.evals/{system}/graders/check_*.py) containing specialized, dynamic LLM-judge templates or regex checks compiled from these goldset inputs.BaseMetric compiled from these goldset inputs.1.0 or 0.0, with zero Likert scale leakage).evals/{system}/tests/test_check_*.py) for each grader.pytest evals/{system}/tests/) to verify evaluator accuracy.holdout.json) remains completely isolated and is never loaded or exposed to the self-tuning loop (to prevent overfitting).config.js or config.py).Trigger /evals-validate to run validation.
evals/{system}/graders/ contains Python grader scripts for each criterion compiled dynamically from goldset pass/fail examples and root-cause analysesevals/{system}/tests/ contains matching unit test filesconfig.js or config.py) successfully generatedholdout.json remained completely isolated and untouched during tuning)pytest evals/{system}/tests/)© tikalk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (scripts) in skills/evals/evals-implement of tikalk/adlc-team-skills.
Open the folder on GitHubat commit 2dbed36
Evals Implement next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Evals Implement this skilltikalk/adlc-team-skills | 141 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Running Testsbrendanhasz/probflow | 175 | — | ~657 | Automated safety check: Pass | MIT | |
| Eval Loopjacob-dietle/context-os | 111 | — | ~5.2k | Automated safety check: Pass | MIT | |
| Eval Driven Devgithub/awesome-copilot | 40k | 1 repos | ~4.4k | Automated safety check: Warn | MIT | |
| Adk Verify Snippetsgoogle/adk-python | 22k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| Hermetic Python Unit TestsdimensionalOS/dimos | 4.6k | — | ~1.4k | Automated safety check: Pass | Custom licence |
brendanhasz/probflow
Run Python unit test suites strictly using the uv package manager and pytest.
jacob-dietle/context-os
This skill should be used when a specific quality problem (UX, data, architecture, feature) needs systematic diagnosis and iterative fixing toward a defined target.
github/awesome-copilot
Improve AI application with evaluation-driven development. An agent skill from github/awesome-copilot.
google/adk-python
Checks that every Python code block in a Markdown file actually compiles and runs, by extracting each block to a temporary file, executing it in an isolated subprocess, and writing a pass/fail…
dimensionalOS/dimos
Rules for writing, fixing and reviewing pytest unit tests that are hermetic: behavior-focused, deterministic, isolated and cheap to run.
microsoft/onnxruntime
Runs and debugs ONNX Runtime tests: Google Test executables for C++ and unittest or pytest for Python, with filters and build-directory guidance.
tikalk/adlc-team-skills
A skill your agent uses when coordinating a multi-repo workspace — init the .adlc/ structure, discover and link child repos as submodules, or audit workspace health (branch, dirty, unpushed, SHA…
tikalk/adlc-team-skills
A skill your agent uses when a session starts or resumes after compaction (auto via the sessionstart and sessioncompact event hooks) and the team AI directives context — constitution, CDR index…
tikalk/adlc-team-skills
A skill your agent uses when ADRs need review, gaps need filling, or ADR status must be approved as Accepted before architecture generation.
tikalk/adlc-team-skills
A skill your agent uses when reviewing, accepting, rejecting, or deferring ChDRs mined by change-init, validating inferred decisions against their git and issue evidence before promotion to project…
tikalk/adlc-team-skills
A skill your agent uses when you want guided mining of git history, structured change-story clustering, or comprehensive rationale recovery before documenting.
tikalk/adlc-team-skills
A skill your agent uses when accepted ChDRs are ready for promotion from drafts to project memory at docs/adlc/memory/chdr/ and the boot-facing chdr.md index needs regenerating.
Works with
Categories
A skill your agent uses when a published goldset is ready to become executable graders and configs — Python graders and metrics with auto-run unit tests verifying grader correctness. Evals Implement is an agent skill from tikalk/adlc-team-skills. Use when a published goldset is ready to become executable graders and configs — Python graders and metrics with auto-run unit tests verifying grader correctness.
Evals Implement fits situations like: tasks that involve LLM evaluation; tasks that involve Unit testing.
Run `npx skills add tikalk/adlc-team-skills --skill evals-implement -a claude-code`. Or copy the skill folder (skills/evals/evals-implement in tikalk/adlc-team-skills) into .claude/skills/evals-implement in your project. Claude Code loads it when a task matches its description.
Run `npx skills add tikalk/adlc-team-skills --skill evals-implement -a codex`. Or copy the skill folder (skills/evals/evals-implement in tikalk/adlc-team-skills) into .agents/skills/evals-implement in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tikalk/adlc-team-skills --skill evals-implement -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evals-implement, .gemini/skills/evals-implement, .github/skills/evals-implement and .opencode/skills/evals-implement in your project.
Going by SKILL.md and its folder, Evals Implement needs a shell and PowerShell for the scripts in its folder and the command-line tools its instructions call (pytest). Our summary lists: Python 3; A Bash shell; PowerShell.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Evals Implement is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.4k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Evals Implement: Running Tests (brendanhasz/probflow, 175 stars), Eval Loop (jacob-dietle/context-os, 111 stars), Eval Driven Dev (github/awesome-copilot, 40k stars) and Adk Verify Snippets (google/adk-python, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
tikalk (a GitHub organization) maintains it in tikalk/adlc-team-skills, which has 141 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 6, 2026.
Source: tikalk/adlc-team-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.