Synthetic Coaching Session Generator
glebis/claude-skills
Writes fictional but realistic coaching or therapy session transcripts for evals, demos and few-shot examples, in several modalities and export formats.
Create, edit, evaluate, and package agent skills. An agent skill from smixs/skill-conductor.
$ npx skills add smixs/skill-conductor --skill skill-conductor -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install smixs/skill-conductor skill-conductor --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/smixs/skill-conductor.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/skill-conductor .claude/skills/skill-conductor && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "skill-conductor" agent skill from https://github.com/smixs/skill-conductor/tree/main/skills/skill-conductor into .claude/skills/skill-conductor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-conductor", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/smixs/skill-conductor/tree/main/skills/skill-conductorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add smixs/skill-conductor --skill skill-conductor -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install smixs/skill-conductor skill-conductor --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/smixs/skill-conductor.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/skill-conductor .agents/skills/skill-conductor && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "skill-conductor" agent skill from https://github.com/smixs/skill-conductor/tree/main/skills/skill-conductor into .agents/skills/skill-conductor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-conductor", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add smixs/skill-conductor --skill skill-conductor -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install smixs/skill-conductor skill-conductor --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/smixs/skill-conductor.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/skill-conductor .cursor/skills/skill-conductor && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "skill-conductor" agent skill from https://github.com/smixs/skill-conductor/tree/main/skills/skill-conductor into .cursor/skills/skill-conductor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-conductor", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/smixs/skill-conductor.git --path skills/skill-conductor--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add smixs/skill-conductor --skill skill-conductor -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install smixs/skill-conductor skill-conductor --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/smixs/skill-conductor.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/skill-conductor .gemini/skills/skill-conductor && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "skill-conductor" agent skill from https://github.com/smixs/skill-conductor/tree/main/skills/skill-conductor into .gemini/skills/skill-conductor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-conductor", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install smixs/skill-conductor skill-conductorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add smixs/skill-conductor --skill skill-conductor -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/smixs/skill-conductor.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/skill-conductor .github/skills/skill-conductor && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "skill-conductor" agent skill from https://github.com/smixs/skill-conductor/tree/main/skills/skill-conductor into .github/skills/skill-conductor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-conductor", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add smixs/skill-conductor --skill skill-conductor -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install smixs/skill-conductor skill-conductor --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/smixs/skill-conductor.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/skill-conductor .opencode/skills/skill-conductor && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "skill-conductor" agent skill from https://github.com/smixs/skill-conductor/tree/main/skills/skill-conductor into .opencode/skills/skill-conductor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-conductor", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skill-conductorCreate, edit, evaluate, and package agent skills. An agent skill from smixs/skill-conductor.
Skill Conductor is an agent skill from smixs/skill-conductor. Create, edit, evaluate, and package agent skills. Use when building a new skill from scratch, improving an existing skill, fixing a skill that never triggers or fires unreliably, running evals to test a skill, benchmarking skill performance, optimizing a skill's description, reviewing third-party skills for quality, or packaging skills for distribution — even if the user doesn't explicitly say "skill" (e.g. "teach Claude to do X", "make the agent always follow Y"). Not for using skills or general coding tasks.
Its SKILL.md is about 6.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 31 other files, including scripts, reference files and assets (for example `agents/analyzer.md`, `agents/bineval.md` and `agents/comparator.md`).
It sits in AI & LLM Engineering, covering LLM evaluation. It works with Obsidian. The repository describes itself as: Architecture-first skill lifecycle for AI agents. BinEval binary scoring with threshold-blind, cross-family-calibrated judges, gated self-update loop, pressure testing, 10… The licence is MIT.
12 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 3c21d2f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Skill Conductor loads about 6.6k tokens when it runs, and up to ~26k if it reads all its reference files. Until then it costs about 133 tokens; SKILL.md has 2,803 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from smixs/skill-conductor at commit 3c21d2f, republished under its MIT licence (© smixs). 2,803 words, ~6,572 tokens.
.claude/skills/skill-conductor/SKILL.md (or your agent's skills folder). This skill also uses 26 other files; get the full folder from GitHub.Full lifecycle management for agent skills: draft → test → review → improve → repeat.
One skill to rule them all — from architecture to packaging. The core loop is always the same: write something, test it, see what fails, fix it, test again.
Before any mode that touches scripts (CREATE, IMPROVE, VALIDATE, OPTIMIZE, PACKAGE), run the pre-flight block → references/runtime-setup.md (checks uv, sets UV_BIN/SKILL_CONDUCTOR_DIR, verifies LLM access). If uv is absent, stop and tell the user.
Read context cues. If the user is a skill author iterating on their own work, be direct and technical. If they're new to skills, explain the why behind each step — not just what to do, but why it matters. Default to conversational, not robotic.
Detect mode from context. If ambiguous, ask.
| Mode | When | What happens |
|---|---|---|
| 1. CREATE | "build a skill", "new skill for..." | Full lifecycle: intent → architecture → scaffold → write → test |
| 2. IMPROVE | "fix this skill", "it doesn't trigger" | Diagnose → eval loop → gated self-update → iterate |
| 3. VALIDATE | "test this skill", "run evals" | Structural checks + trigger testing + BinEval scoring |
| 4. REVIEW | "review this skill", third-party assessment | 11-point quality gate, quick and focused |
| 5. OPTIMIZE | "improve triggering", "description optimization" | Automated description optimization with train/test split |
| 6. PACKAGE | "package for distribution" | Validate + bundle into .skill file |
Before writing anything, extract 2–3 concrete scenarios.
Ask:
Don't move on until you have a clear picture of what the skill does, for whom, and when. This prevents the most common failure: a skill that does something but triggers for the wrong things.
Before writing the skill, verify the agent fails without it:
If the agent already handles it perfectly, the skill is unnecessary. This sounds obvious, but it's the most skipped step and the most valuable one.
Choose a primary pattern from references/patterns.md (can combine):
| Pattern | Use when |
|---|---|
| Sequential workflow | clear step-by-step process |
| Iterative refinement | output improves with cycles |
| Context-aware selection | same goal, different tools by context |
| Domain intelligence | specialized knowledge beyond tool access |
| Multi-MCP coordination | workflow spans multiple services |
Choose degrees of freedom — this determines how much control vs. flexibility the skill gives the agent:
| Freedom | When | Example |
|---|---|---|
| Low (scripts) | fragile, error-prone, must be exact | PDF rotation, API calls |
| Medium (pseudocode) | preferred pattern exists, some variation ok | data processing |
| High (text) | multiple valid approaches, judgment needed | design decisions |
Freedom test: ask "if the agent makes a mistake here, what is the consequence?" High consequence → low freedom (an exact script it must not modify). Low consequence → high freedom (prose, let it judge). Calibrate per step, not per skill — one skill can hold both.
Golden rule: read references/sop-practices.md before authoring or reviewing ANY skill. It holds the canonical 10 authoring principles (universal): pre-flight, no-process-in-description, MOC (SKILL.md = map, not prose), fresh-practitioner author, TWI "why", blind-agent test, inline checklists, one-term-per-concept, cut-the-fat (env/keys OUT of SKILL.md), match-the-form-to-the-failure. For procedural skills (business process with branching: request, quote, onboarding, escalation) the same file also has the deep SOP methodology — format selection, 7-step process, procedural checklist.
uv run scripts/init_skill.py <skill-name> --path <output-dir> [--resources scripts,references,assets]Or create manually:
skill-name/
├── SKILL.md # required — the brain
├── scripts/ # deterministic operations (executed, not loaded)
├── references/ # detailed docs (loaded on demand)
└── assets/ # templates, images for output (never loaded)---
name: kebab-case-name
description: >
[What it does]. Use when [4-5 phrasing variations users actually say] — even
if they don't explicitly say "[canonical term]". Do NOT use for [negatives].
---The description is the single most important line — it decides whether the skill triggers at all. The full formula, the pushy clause and worked GOOD/BAD examples live in references/sop-practices.md Principle #2. Read it before writing one.
name: lowercase, digits, hyphens only. No consecutive hyphens. Matches folder name. Max 64 charsdescription: max 1024 chars. No angle brackets. No process/workflow steps# Skill Name
## Overview
What this enables. 1-2 sentences. Core principle.
## [Main sections]
Step-by-step with numbered sequences.
Concrete templates over prose.
Imperative voice throughout.
## Common Mistakes
What goes wrong + how to fix.
## Troubleshooting (if applicable)
Error: [message] → Cause: [why] → Fix: [how]/home/<user>, /Users/<user>) — reference them, never inline (Principle 9a)references/sop-practices.md Principle 5, TWI)This is the critical step — most failures hide here. Treat it as three sub-phases.
Before the full loop, micro-test the wording of anything you just wrote (5+ fresh-context reps, always with a no-guidance control) → references/pressure-testing.md. For a discipline skill — one that makes the agent follow a rule it's tempted to break — a pressure scenario from that file is mandatory, not optional.
evals/evals.json exists with 3–5 prompts (see references/schemas.md)<skill-name>-workspace/iteration-1/eval-0) and eval_metadata.jsonuv and eval-viewer/generate_review.py are reachable from current working dirIf any item fails — fix before proceeding. A missing workspace dir mid-run loses outputs.
| What | Key move | Why |
|---|---|---|
| Spawn with-skill runs | One subagent per eval, skill active, save outputs to iteration-N/<eval-name>/with_skill/ | Parallel = same wall time as one run |
| Spawn baseline runs in the same turn | Same prompt, no skill (or old version snapshot for IMPROVE), save to without_skill/ or old_skill/ | If you wait, baselines drift in time and aren't comparable |
| Draft assertions while runs execute | Pull verifiable statements from eval prompts | Don't waste the 5–15 min of subagent time |
| Capture timing on each notification | Save total_tokens, duration_ms to timing.json immediately | Notification is the only source — process per-arrival, don't batch |
timing.json files written (one per run)grading.json with fields text, passed, evidence (not name/met)benchmark.json aggregated: uv run scripts/aggregate_benchmark.py <workspace>/iteration-N --skill-name <name>agents/analyzer.md for what to look for (non-discriminating assertions, high-variance evals, time/token tradeoffs)uv run eval-viewer/generate_review.py <workspace> --skill-name <name> --benchmark <path>--static <output.html> and send file to user--previous-workspace <previous-iteration-path>The last bullet is the trap. If you skip user review and "improve" based on your own reading of outputs, you optimize against your taste, not the user's.
If any fail → iterate. Find how the agent rationalizes around the skill, plug loopholes, re-verify.
Read the existing SKILL.md completely. Identify the problem class:
| Problem | Signal | Fix |
|---|---|---|
| Undertriggering | skill doesn't load | add keywords, trigger phrases, file types to description |
| Overtriggering | loads for unrelated queries | add negative triggers, be more specific |
| Skips body | follows description only | remove process/workflow from description |
| Inconsistent output | varies across sessions | add explicit templates, reduce freedom, add scripts |
| Too slow | large context | move detail to references/, cut body to <500 lines |
scripts/. Saves every future invocation from reinventing the wheelreferences/sop-practices.md — the 10 canonical principles (universal) map directly to skill failure modes: process leaking into description, SKILL.md bloated instead of a map, env/keys inlined, silent improvisation from missing "why", missed edge cases, agents skipping end-of-doc checklists, a rule whose form doesn't match its failure. For process skills (ticket, quote, escalation) also apply the deep SOP methodology in the same fileThe improvement cycle mirrors CREATE Step 6, but focused on the broken behavior. Micro-test each candidate wording before it enters the loop, and re-run the pressure scenarios if the skill enforces a rule → references/pressure-testing.md.
agents/grader.mduv run eval-viewer/generate_review.py <workspace>--static <output.html> instead of live serverDrive iteration off failing BinEval questions, not taste — and accept edits only against evidence the editor never saw. Full rules: references/bineval-method.md § Gated self-update loop.
uv run scripts/split_evals.py evals/evals.json --holdout 0.4 --write <workspace>/split.json — deterministic, stratified by the optional per-eval category. Never re-split after seeing resultsreferences/bineval-method.md) → collect failing[]agents/analyzer.md with train transcripts + gradings to produce generalized, deduped lessons. Held-out grading stays unopened until the gatetransitions block in benchmark.jsonfailing[] (or its critical subset) is empty, or after 3 iterations. Keep the best ACCEPTED version by held-out pass-rate, then train pass-rateWhen you have two meaningfully different versions:
agents/comparator.md — answers the SAME binary questions for outputs A and B without knowing which skill produced whichagents/analyzer.md — unblinds results, analyzes WHY the winner wonThis prevents bias. The comparator judges output quality, not skill design.
Three stages, run in order.
uv run scripts/eval_skill.py <skill-folder>Checks: frontmatter, naming, description quality, process leak detection, body size, structure, scripts. Target: 10/10, no warnings.
Generate 6 test prompts:
Run each in clean session. Target: 6/6 correct.
For automated trigger testing at scale, use:
uv run scripts/run_eval.py --eval-set <path> --skill-path <path> --runs-per-query 3Evaluate with atomic binary yes/no questions across 5 dimensions — each answered 1/0 after a written critique citing evidence. See references/bineval-method.md for the method, references/quality-questions.md for the question bank, and agents/bineval.md for the evaluator that emits bineval.json.
The 5 dimensions: Discovery, Clarity, Structure, Robustness, Completeness.
Questions for the skill-artifact come from two sources (question_source: "hybrid"):
scripts/eval_skill.py --json (the sole emitter), e.g. DET-STRUCT-SKILLMD-EXISTS, DET-DISCOVERY-DESC-PRESENT. Some are flagged critical.references/quality-questions.md; the judge answers them, never invents its own. (Generated per-task questions via the two-step meta-prompt belong to output grading in Modes 1–2, agents/grader.md — not to artifact scoring.)The judge only answers the questions. YOU aggregate: per-dimension dimension_scores S_d = mean of that dimension's answers; overall S = mean of all answers. Never put the bands or the GATE into a judge prompt — a judge that knows the bar is biased toward it (references/bineval-method.md).
Display bands: S≥0.90 production-ready · 0.70–0.89 solid · 0.50–0.69 needs-work · <0.50 rewrite.
GATE = every critical question (deterministic + critical bank questions) answered 1. The GATE is the pass criterion — not the scalar S.
Quick quality gate for third-party skills.
[ ] SKILL.md exists, exact case
[ ] Valid YAML frontmatter (name + description)
[ ] name: kebab-case, matches folder, ≤64 chars
[ ] description: ≤1024 chars, no angle brackets
[ ] description has triggers ("Use when...")
[ ] description has NO workflow/process steps
[ ] No README.md inside skill folder
[ ] SKILL.md < 500 lines
[ ] References max 1 level deep
[ ] Scripts tested and executable
[ ] No hardcoded paths/tokens/secretsThen run VALIDATE Stage 2 (discovery) on the description. Report score + checklist.
The deterministic subset of this checklist is emitted as binary BinEval question records by scripts/eval_skill.py --json (e.g. DET-STRUCT-SKILLMD-EXISTS, DET-DISCOVERY-DESC-PRESENT, DET-ROBUST-NO-SECRETS) — the sole emitter of those records.
The checklist exists because these are the failure modes that actually happen in practice — especially process-in-description, which causes the agent to skip the body entirely.
Automated description optimization. The description competes with other skills for Claude's attention — optimization finds the wording that triggers most accurately. The same train/held-out principle now gates body edits too — see Mode 2 Step 3.
Queries must be realistic — concrete, detailed, with file paths, context, abbreviations, typos. Not "Format this data" but "my boss sent Q4 sales final FINAL v2.xlsx, add profit margin % column, revenue is col C costs col D".
Should-trigger (10): Different phrasings of the same intent — formal, casual, implicit. Include cases where user doesn't name the skill but clearly needs it. Add competing-skill edge cases.
Should-NOT-trigger (10): Near-misses that share keywords but need something different. Adjacent domains, ambiguous phrasing. "Write fibonacci" as negative for PDF skill = useless — too easy. Make negatives genuinely tricky.
Triggering mechanics: Claude only consults skills for tasks it can't handle directly. Simple queries ("read this PDF") won't trigger skills regardless of description — Claude handles them with basic tools. Eval queries must be substantive enough that consulting a skill would help.
assets/eval_review.htmluv run scripts/run_loop.py \
--eval-set evals/eval_set.json \
--skill-path <skill-dir> \
--model <model-id> \
--max-iterations 5 \
--holdout 0.4 \
--verboseThe loop:
| Script | Purpose |
|---|---|
scripts/run_eval.py | Run trigger evaluation on a description |
scripts/improve_description.py | Claude proposes improved description |
scripts/generate_report.py | HTML visualization of optimization history |
scripts/aggregate_benchmark.py | Statistical aggregation of benchmark runs |
uv run scripts/quick_validate.py <skill-folder>uv run scripts/package_skill.py <skill-folder> [output-dir]Creates skill-name.skill (zip with .skill extension). Verify: unzip in temp dir, check structure intact.
references/sop-practices.md| Directory | Loaded? | Purpose |
|---|---|---|
| SKILL.md | on trigger | brain — instructions |
| references/ | on demand | detailed docs, schemas |
| scripts/ | executed, not loaded | deterministic operations |
| assets/ | never loaded | templates, images |
| Level | When loaded | Budget |
|---|---|---|
| Frontmatter | always (system prompt) | ~100 words |
| SKILL.md body | on trigger | <500 lines |
| Bundled resources | on demand | unlimited |
[What it does] + Use when [4-5 phrasings users actually say] + even if they don't explicitly say "<canonical term>" + Do NOT use for [negatives] — full rules and examples in references/sop-practices.md Principle #2.
Load on demand, at the point of use named in each mode — never wholesale. Load an agents/* file only at the step that spawns that agent; load references/schemas.md only when writing or reading a JSON artifact. Everything else stays unloaded.
| Path | What's inside |
|---|---|
agents/grader.md | Evidence-based assertion grading |
agents/comparator.md | Blind A/B output comparison |
agents/analyzer.md | Post-hoc analysis + benchmark notes |
agents/bineval.md | BinEval evaluator — emits bineval.json |
references/patterns.md | 5 architectural patterns + anti-patterns |
references/schemas.md | JSON schemas for evals, grading, benchmark |
references/bineval-method.md | BinEval method: dimensions, scoring, GATE |
references/quality-questions.md | BinEval question bank (deterministic + bank) |
references/pressure-testing.md | Micro-tests for wording + pressure scenarios for discipline skills |
references/sop-practices.md | Canon: 10 authoring principles (universal) + deep SOP methodology for procedural skills |
references/runtime-setup.md | Pre-flight: uv/env/path checks, LLM-access options |
eval-viewer/ | Interactive HTML viewer for eval results |
assets/eval_review.html | Trigger eval set editor |
scripts/eval_skill.py | Structural validation (10-point scoring) |
scripts/init_skill.py | Skill scaffolder |
scripts/run_eval.py | Trigger evaluation runner |
scripts/run_loop.py | Eval + improve optimization loop |
scripts/improve_description.py | Claude-powered description improvement |
scripts/aggregate_benchmark.py | Benchmark statistics aggregator |
scripts/generate_report.py | HTML report generator |
scripts/quick_validate.py | Quick validation for packager |
scripts/test_smoke.py | Smoke tests for all scripts (12 tests) |
scripts/package_skill.py | Skill → .skill packager |
scripts/utils.py | Shared utilities (parse_skill_md) |
© smixs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 26 other files (scripts, references, assets) in skills/skill-conductor of smixs/skill-conductor.
Open the folder on GitHubat commit 3c21d2f
Skill Conductor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Skill Conductor this skillsmixs/skill-conductor | 179 | — | ~6.6k | Automated safety check: Pass | MIT | |
| Synthetic Coaching Session Generatorglebis/claude-skills | 391 | — | ~2.9k | Automated safety check: Pass | MIT | |
| Agent Eval Engineeringlangchain-ai/langchain-skills | 1.3k | — | ~4k | Automated safety check: Pass | MIT | |
| Evaluate RAGai-evals-course/evals-skills | 1.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Synthetic Eval Data Generatorai-evals-course/evals-skills | 1.5k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| Code Model Evaluation HarnessOrchestra-Research/AI-Research-SKILLs | 13k | 4 repos | ~2.9k | Automated safety check: Pass | MIT |
glebis/claude-skills
Writes fictional but realistic coaching or therapy session transcripts for evals, demos and few-shot examples, in several modalities and export formats.
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
ai-evals-course/evals-skills
Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.
ai-evals-course/evals-skills
Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.
Orchestra-Research/AI-Research-SKILLs
Benchmarks code generation models with the BigCode Evaluation Harness across HumanEval, MBPP, MultiPL-E and other suites using pass@k metrics.
ai-evals-course/evals-skills
Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data.
Works with
Categories
Create, edit, evaluate, and package agent skills. An agent skill from smixs/skill-conductor. Skill Conductor is an agent skill from smixs/skill-conductor. Create, edit, evaluate, and package agent skills.
Skill Conductor fits situations like: building a new skill from scratch; improving an existing skill; fixing a skill that never triggers; fires unreliably.
Run `npx skills add smixs/skill-conductor --skill skill-conductor -a claude-code`. Or copy the skill folder (skills/skill-conductor in smixs/skill-conductor) into .claude/skills/skill-conductor in your project. Claude Code loads it when a task matches its description.
Run `npx skills add smixs/skill-conductor --skill skill-conductor -a codex`. Or copy the skill folder (skills/skill-conductor in smixs/skill-conductor) into .agents/skills/skill-conductor in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add smixs/skill-conductor --skill skill-conductor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-conductor, .gemini/skills/skill-conductor, .github/skills/skill-conductor and .opencode/skills/skill-conductor in your project.
Going by SKILL.md and its folder, Skill Conductor needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Skill Conductor is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.6k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 19k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Skill Conductor: Synthetic Coaching Session Generator (glebis/claude-skills, 391 stars), Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars), Evaluate RAG (ai-evals-course/evals-skills, 1.5k stars) and Synthetic Eval Data Generator (ai-evals-course/evals-skills, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
smixs (a GitHub user) maintains it in smixs/skill-conductor, which has 179 GitHub stars. The repository was last updated on August 3, 2026.
Source: smixs/skill-conductor on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.