Agent Eval Engineering
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
Writes fictional but realistic coaching or therapy session transcripts for evals, demos and few-shot examples, in several modalities and export formats.
$ npx skills add glebis/claude-skills --skill synthetic-session-generator -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install glebis/claude-skills synthetic-session-generator --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/synthetic-session-generator .claude/skills/synthetic-session-generator && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "synthetic-session-generator" agent skill from https://github.com/glebis/claude-skills/tree/main/synthetic-session-generator into .claude/skills/synthetic-session-generator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-session-generator", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/glebis/claude-skills/tree/main/synthetic-session-generatorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add glebis/claude-skills --skill synthetic-session-generator -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install glebis/claude-skills synthetic-session-generator --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/synthetic-session-generator .agents/skills/synthetic-session-generator && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "synthetic-session-generator" agent skill from https://github.com/glebis/claude-skills/tree/main/synthetic-session-generator into .agents/skills/synthetic-session-generator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-session-generator", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add glebis/claude-skills --skill synthetic-session-generator -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install glebis/claude-skills synthetic-session-generator --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/synthetic-session-generator .cursor/skills/synthetic-session-generator && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "synthetic-session-generator" agent skill from https://github.com/glebis/claude-skills/tree/main/synthetic-session-generator into .cursor/skills/synthetic-session-generator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-session-generator", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/glebis/claude-skills.git --path synthetic-session-generator--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add glebis/claude-skills --skill synthetic-session-generator -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install glebis/claude-skills synthetic-session-generator --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/synthetic-session-generator .gemini/skills/synthetic-session-generator && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "synthetic-session-generator" agent skill from https://github.com/glebis/claude-skills/tree/main/synthetic-session-generator into .gemini/skills/synthetic-session-generator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-session-generator", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install glebis/claude-skills synthetic-session-generatorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add glebis/claude-skills --skill synthetic-session-generator -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/synthetic-session-generator .github/skills/synthetic-session-generator && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "synthetic-session-generator" agent skill from https://github.com/glebis/claude-skills/tree/main/synthetic-session-generator into .github/skills/synthetic-session-generator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-session-generator", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add glebis/claude-skills --skill synthetic-session-generator -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install glebis/claude-skills synthetic-session-generator --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/synthetic-session-generator .opencode/skills/synthetic-session-generator && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "synthetic-session-generator" agent skill from https://github.com/glebis/claude-skills/tree/main/synthetic-session-generator into .opencode/skills/synthetic-session-generator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-session-generator", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
synthetic-session-generatorWrites fictional but realistic coaching or therapy session transcripts for evals, demos and few-shot examples, in several modalities and export formats.
The skill produces invented coach-and-client or therapist-and-client dialogue that reads like a recorded session while staying clearly synthetic; every output is watermarked. Realism rests on persona consistency, so a client keeps the same voice, history and presenting issues across a session arc, and on modality fidelity to ICF/GROW coaching, CBT, IFS parts work or ACT with motivational interviewing. Results can serve as eval data with ground-truth labels, demos without real client data, or few-shot material.
A setup mode stores defaults for language, modality and session length in `config.json` using `scripts/setup_config.py`, and later runs of `scaffold_session.py` inherit them unless flags override. Other scripts convert formats and make a card. Exports include Fathom or Granola transcript style, plain dialogue, structured JSON and Obsidian markdown, with templates, persona and modality references and a realism guide alongside. It is not for analyzing or summarizing a real transcript.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 3b88261. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 5 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Synthetic Coaching Session Generator loads about 2.9k tokens when it runs, and up to ~5.4k if it reads all its reference files. Until then it costs about 187 tokens; SKILL.md has 1,331 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from glebis/claude-skills at commit 3b88261, republished under its MIT licence (© glebis). 1,331 words, ~2,898 tokens.
.claude/skills/synthetic-session-generator/SKILL.md (or your agent's skills folder). This skill also uses 14 other files; get the full folder from GitHub.Generate fictional but believable coaching/therapy session transcripts that read like real recorded sessions, while remaining clearly synthetic. Outputs feed three jobs: eval datasets (with ground-truth labels to benchmark summarizers and analyzers), product demos (realistic sessions without exposing real client data), and training/prompt examples (few-shot material for a coaching or therapy assistant).
Realism comes from two disciplines: persona consistency (a client speaks the same way, carries the same history and presenting issues across a session arc) and modality fidelity (the practitioner uses the techniques, question forms, and pacing of the chosen framework). Every output is watermarked as synthetic so it can never be mistaken for a real clinical record.
Use when a user asks for fake/synthetic/mock/demo coaching or therapy transcripts, eval or test data
for session-analysis tools (e.g. the coaching-session-summarizer), few-shot dialogue examples, or
persona-consistent session series. Do not use to analyze or summarize a real transcript — that
is the job of coaching-session-summarizer or transcript-analyzer.
When the user wants to configure the skill ("setup", "set my defaults", "always use Russian / IFS / 50-minute sessions"), run setup mode. Offer the three choices via AskUserQuestion, then persist them:
en, ru, de, es, fr, pt, it, nl).icf-grow, cbt, ifs, act-mi).python3 scripts/setup_config.py --language ru --modality cbt --duration 50 --show
python3 scripts/setup_config.py --show # view current defaultsThis writes config.json in the skill directory. Later scaffold_session.py runs inherit these
defaults, so the user only specifies what differs (e.g. persona and session position). Per-run flags
always override the saved config.
Honour the setup-mode defaults (Step 0); only ask for parameters the user hasn't already fixed.
Collect (or infer sensible defaults for) these parameters. Ask only for what materially changes the output; default the rest.
icf-grow, cbt, ifs, or act-mi. See references/modalities.md for the
technique cheat-sheet, signature moves, and vocabulary of each.references/personas.md, or generate a new one and
persist it back into that file so a session series stays consistent. A persona = name,
demographics, presenting issue, history, speech register, defenses/resistances, goals.fathom, plain, json, or markdown (see Step 3). Markdown is always produced.--language. Author all dialogue, persona
voice, and the watermark-adjacent text in that language; keep eval tag keys in English.--duration <minutes> (preferred — maps to a turn budget) or the coarse
--length (short ~15 / standard ~30 / long ~50+).Run the scaffolding script to turn the spec into a structured skeleton (phases, beat list, turn budget, JSON shell, and the synthetic watermark):
python3 scripts/scaffold_session.py --modality cbt --persona maya --position mid-arc \
--length standard --format json --out /tmp/session_skeleton.jsonThen write the actual dialogue by hand (model-authored), filling each beat. The script provides
structure and guardrails; Claude provides the natural, non-templated language. Key realism rules
(full list in references/realism_guide.md):
Author once in the JSON turn structure, then convert. Always render the markdown format (it is the canonical, human-readable artifact); add any other formats the user asked for.
# markdown is always produced:
python3 scripts/convert_format.py --in /tmp/session.json --to markdown --auto-timestamps --out session.md
# plus any requested extras:
python3 scripts/convert_format.py --in /tmp/session.json --to fathom --auto-timestamps --out session.txtcoaching-session-summarizer, transcript-analyzer).Coach: / Client: turn-taking markdown.speaker, timestamp, text, and eval tags
(technique, emotion, phase); for evals, also the ground_truth block.Timestamps. Do not hand-invent timestamps. Pass --auto-timestamps so the converter emulates
them from each turn's word count (~150 wpm + a short inter-turn gap), keeping timing internally
consistent. Tune pace with --wpm. See assets/templates/ for a reference example of each format.
When the user wants a card summarizing the case (for demos, persona bibles, or eval context), build it from the same session JSON and pair it with a generated portrait:
python3 scripts/make_card.py --in /tmp/session.json --out /tmp/card.md # scaffold
python3 scripts/make_card.py --in /tmp/session.json --print-prompt # portrait promptmake_card.py to emit the card scaffold (modality-aware formulation skeleton + themes/goals
pulled from ground_truth + a watermark + a ready portrait prompt).<!-- FILL: ... --> blocks with the clinical formulation (model-authored).gpt-image-2 skill using the prompt from --print-prompt.
Keep it illustrative, not photoreal — a stylized image cannot be mistaken for a photo of a
real person. Then re-run with --image <path> (or edit the card) to embed it.tufte-reportWhen the user wants a shareable HTML page of the case card (portrait + conceptualization), hand
the filled card to the tufte-report skill, which produces a standalone Tufte-style HTML file.
tufte-report skill with the card's conceptualization as the narrative content and the
portrait as a figure. Map card sections to the report: Snapshot/Presenting issue → intro
narrative; Formulation → the main 2-column narrative+data section; Working themes and
Goals & experiments → a status/dashboard panel; Emotional arc → a sparkline or labelled
sequence. Pass the portrait path so it renders as the hero figure..html.The portrait must remain the illustrative, non-photoreal image from Step 4 — the HTML page is for demos and persona bibles, never presented as a real client record.
Always apply the synthetic watermark — this is non-negotiable. The scaffold script injects it; verify it survived format conversion. Each output must carry, in a location appropriate to its format (frontmatter, JSON metadata, or a header/footer comment):
⚠️ SYNTHETIC — AI-generated fictional session. Not a real person, not clinical advice.Confirm the save location before writing. Ask the user where to save and state the default —
the current working directory (.). Only fall back to /tmp/ for throwaway intermediate
scaffolds the user will not keep. Use clear filenames (e.g. <persona>_<modality>_<position>.md).
For eval batches, write one file per session into the chosen directory plus a manifest listing
personas, modalities, and label coverage.
session-anonymizer).coaching-session-summarizer and anonymization to session-anonymizer.© glebis, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 14 other files (scripts, references, assets) in synthetic-session-generator of glebis/claude-skills.
Open the folder on GitHubat commit 3b88261
Synthetic Coaching Session Generator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Synthetic Coaching Session Generator this skillglebis/claude-skills | 390 | — | ~2.9k | Automated safety check: Pass | MIT | |
| Agent Eval Engineeringlangchain-ai/langchain-skills | 1.3k | — | ~4k | Automated safety check: Pass | MIT | |
| Skill Conductorsmixs/skill-conductor | 179 | — | ~6.6k | Automated safety check: Pass | MIT | |
| Synthetic Eval Data Generatorai-evals-course/evals-skills | 1.5k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| Chatbox Session RAG Evalchatboxai/chatbox | 42k | — | ~758 | Automated safety check: Pass | GPL-3.0 | |
| Nexus TestingProfSynapse/nexus | 157 | — | ~994 | Automated safety check: Pass | MIT |
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
smixs/skill-conductor
Create, edit, evaluate, and package agent skills. An agent skill from smixs/skill-conductor.
ai-evals-course/evals-skills
Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.
chatboxai/chatbox
Runs and debugs evaluations of how Chatbox models answer questions about large attached files, using synthetic and real long-document fixtures.
ProfSynapse/nexus
Verify a Nexus change — pick a Jest lane, write a test that can actually fail, run the in-app Obsidian CLI loop, drive the eval harness, or fix a shipped-docs drift failure.
ai-evals-course/evals-skills
Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.
glebis/claude-skills
Runs a human-first workflow for labeling PII spans in a transcript, then scores inter-annotator agreement and drafts an adjudicated gold set.
glebis/claude-skills
Automates a dedicated, logged-in Chrome instance per profile without ever closing the user's own open tabs or browser windows.
glebis/claude-skills
This skill should be used when conducting comprehensive research on any topic using the OpenAI Deep Research API.
glebis/claude-skills
This skill should be used for elimination-style research where the user wants to choose from a shortlist of products, tools, services, vendors, or other options using explicit criteria, numeric…
glebis/claude-skills
Generates a self-contained HTML presentation with article and slides modes, ElevenLabs voiceover narration and optional GPT Image 2 illustrations.
glebis/claude-skills
Walks through Goldratt's Five Focusing Steps to find the real bottleneck in your work, then recommends one automation aimed at it and a list of what not to automate.
Works with
Categories
Writes fictional but realistic coaching or therapy session transcripts for evals, demos and few-shot examples, in several modalities and export formats. The skill produces invented coach-and-client or therapist-and-client dialogue that reads like a recorded session while staying clearly synthetic; every output is watermarked. Realism rests on persona consistency, so a client keeps the same voice, history and presenting issues across a session arc, and on modality fidelity to ICF/GROW coaching, CBT, IFS parts work or ACT with motivational interviewing.
Synthetic Coaching Session Generator fits situations like: creating test data for a session summarizer or other coaching analysis tool; building demo transcripts without exposing real client conversations; producing few-shot dialogue examples for a coaching or therapy assistant; generating a series of sessions that keep the same client persona.
Run `npx skills add glebis/claude-skills --skill synthetic-session-generator -a claude-code`. Or copy the skill folder (synthetic-session-generator in glebis/claude-skills) into .claude/skills/synthetic-session-generator in your project. Claude Code loads it when a task matches its description.
Run `npx skills add glebis/claude-skills --skill synthetic-session-generator -a codex`. Or copy the skill folder (synthetic-session-generator in glebis/claude-skills) into .agents/skills/synthetic-session-generator in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add glebis/claude-skills --skill synthetic-session-generator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/synthetic-session-generator, .gemini/skills/synthetic-session-generator, .github/skills/synthetic-session-generator and .opencode/skills/synthetic-session-generator in your project.
Going by SKILL.md and its folder, Synthetic Coaching Session Generator needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Synthetic Coaching Session Generator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.5k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Synthetic Coaching Session Generator: Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars), Skill Conductor (smixs/skill-conductor, 179 stars), Synthetic Eval Data Generator (ai-evals-course/evals-skills, 1.5k stars) and Chatbox Session RAG Eval (chatboxai/chatbox, 42k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
glebis (a GitHub user) maintains it in glebis/claude-skills, which has 390 GitHub stars. The repository holds 92 skills in this directory. The repository was last updated on October 8, 2026.
Source: glebis/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.