Experience Map
Owl-Listener/designer-skills
Map the full ecosystem of touchpoints, channels, and relationships across a service.
Run fast, progressive experiments to map tunable parameters and their effects.
$ npx skills add dzhng/skills --skill auto-research -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install dzhng/skills auto-research --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/dzhng/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/engineering/auto-research .claude/skills/auto-research && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "auto-research" agent skill from https://github.com/dzhng/skills/tree/main/skills/engineering/auto-research into .claude/skills/auto-research/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "auto-research", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/dzhng/skills/tree/main/skills/engineering/auto-researchType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add dzhng/skills --skill auto-research -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install dzhng/skills auto-research --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dzhng/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/engineering/auto-research .agents/skills/auto-research && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "auto-research" agent skill from https://github.com/dzhng/skills/tree/main/skills/engineering/auto-research into .agents/skills/auto-research/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "auto-research", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add dzhng/skills --skill auto-research -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install dzhng/skills auto-research --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dzhng/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/engineering/auto-research .cursor/skills/auto-research && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "auto-research" agent skill from https://github.com/dzhng/skills/tree/main/skills/engineering/auto-research into .cursor/skills/auto-research/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "auto-research", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/dzhng/skills.git --path skills/engineering/auto-research--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add dzhng/skills --skill auto-research -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install dzhng/skills auto-research --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dzhng/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/engineering/auto-research .gemini/skills/auto-research && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "auto-research" agent skill from https://github.com/dzhng/skills/tree/main/skills/engineering/auto-research into .gemini/skills/auto-research/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "auto-research", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install dzhng/skills auto-researchInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add dzhng/skills --skill auto-research -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/dzhng/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/engineering/auto-research .github/skills/auto-research && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "auto-research" agent skill from https://github.com/dzhng/skills/tree/main/skills/engineering/auto-research into .github/skills/auto-research/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "auto-research", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add dzhng/skills --skill auto-research -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install dzhng/skills auto-research --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dzhng/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/engineering/auto-research .opencode/skills/auto-research && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "auto-research" agent skill from https://github.com/dzhng/skills/tree/main/skills/engineering/auto-research into .opencode/skills/auto-research/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "auto-research", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
auto-researchRun fast, progressive experiments to map tunable parameters and their effects.
Auto Research is an agent skill from dzhng/skills. Run fast, progressive experiments to map tunable parameters and their effects. Use when optimizing code, prompts, configurations, or other artifacts against a goal or benchmark, or when a long research run needs a planned hypothesis queue, focused trials, and durable findings.
Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: Reusable AI agent skills for software factories: explore ideas, write specs, implement, review, and run autonomous research. Works with Claude Code, Codex, and other… The licence is MIT.
8 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d513228. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Auto Research loads about 2.7k tokens when it runs. Until then it costs about 73 tokens; SKILL.md has 1,453 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from dzhng/skills at commit d513228, republished under its MIT licence (© dzhng). 1,453 words, ~2,705 tokens.
.claude/skills/auto-research/SKILL.md (or your agent's skills folder).Optimize learning per unit of time. Start with one fast, revealing task, plan competing hypotheses, test them separately, combine verified winners, then expand coverage. The durable result is a parameter-effect map plus the best verified artifact. An experiment is a checkpoint, not a stopping point.
Recover the research state. Read the objective, user feedback, project instructions, and existing research records. Reconcile the handoff with actual artifacts and running jobs; collect a pending attempt before launching another. Identify the best verified artifact (the incumbent), active task, regression set, hypothesis queue, and remaining budget. Reuse existing record locations.
Clarify success before experimenting. Recover any already agreed criteria; do not ask the user to repeat them. If the evaluator, metric, comparison baseline, or required improvement is unclear, ask focused questions and wait for answers before editing candidates or launching evaluations. Inspect existing artifacts and propose concrete options to make answering easier; do not silently choose what success means. Continue clarification until the answers define a checkable criterion, including evaluation scope, aggregation, and non-regression gates. Restate that criterion in the research contract.
Distinguish passing correctness gates from achieving the optimization target. Specify whether the target is an absolute score, an absolute change, or a relative improvement against a named baseline. “Improve accuracy by 5%” is ambiguous: from 60%, five percentage points means 65%, while 5% relative means 63%. Clarify the units and direction; record the formula when needed. A baseline's value may be measured next, but its identity and the comparison rule must be settled first. Accept open-ended optimization only when that is the user's intent; do not invent a threshold or substitute any improvement for the requested improvement.
Then define the smallest useful loop. Record mutable parameters and the fixed evaluator and select one cheap task that exposes the behavior being improved. Measure its turnaround and remove unnecessary setup, models, cases, and implementation work from the inner loop. Prefer an architectural spike when the uncertainty is architectural. Available benchmark tasks are a pool, not a requirement to run them all; choose validation coverage to support the intended claim and user's scope.
Record the evaluation command, timeout, repair allowance, resource limits, promotion rule, and confirmation plan. Separate research spend from the cost or runtime being optimized. Done means the success criterion is resolved and the next trial is cheap to run with an unambiguous decision rule; do not build a large benchmark matrix before testing the first hypothesis.
Establish the active task's baseline. Run the unchanged artifact once, or reuse its compatible recorded result and traces. Verify the evaluator measures the intended outcome; use a known failing control if detection is unproven. Keep the baseline fixed for this task/model/harness; do not rerun it per variant. For noisy measurements, plan only the repeats needed to distinguish an effect, respecting any user instruction to reuse a single baseline and reporting that limitation. Save exact inputs, artifact identity, and raw output.
Plan a small hypothesis batch. Before editing, list distinct explanations of the current bottleneck and the parameters or strategies that could test them. Inspect actual failure traces and matched baseline behavior instead of guessing from aggregate scores. For each hypothesis, record:
Rank by information gained per unit of time. Test alternatives one by one against the same parent so their effects remain interpretable. Avoid exhaustive parameter grids. Keep the queue short and revise it after results; planning is part of every research cycle, not a one-time backlog or a reason to delay the first trial.
Test one hypothesis on the active task. Preserve the incumbent and isolate the candidate from unrelated work. Change one conceptual factor; if a coupled bundle is necessary, attribute the result to that bundle. Record the candidate identity and prediction before running. Use cheap checks first, then the focused task; do not run the broad suite for every variant. Save uniquely identified raw output and inspect relevant traces. Bound crashes and repairs; a repaired candidate gets a new identity. Missing results are unknown, not zero.
Compare observation with prediction, including regressions and null results.
Record keep, discard, inconclusive, or invalid and update the map.
Restore the parent between alternatives; retain reproducible snapshots or
patches and evidence outside rollback scope. Rejecting a candidate does not
disprove its entire strategy: diagnose a promising partial win and plan a
targeted follow-up when evidence supports one.
Combine and promote. Test compatible winners together against the strongest individual candidate. Do not assume gains add: check interactions, and remove one change at a time when attribution is unclear. A combination that loses to a simpler candidate does not become the incumbent.
A focused win is provisional. Before promotion, check the candidate against previously solved tasks in the accumulated regression set, using their saved baselines and the declared tolerances. Reject or repair regressions rather than hiding them in an average. Confirm noisy wins with the planned evidence, including all planned attempts and costs. Bank the verified candidate and update the parameter-effect map with interactions and task-specific limits.
Expand only after progress. Once the active task meets its local criterion and prior tasks remain green, add one task or a small batch that probes a new failure mode or an uncertain effect in the map. Establish missing baselines only for those tasks. If a new task fails, make it the active task, return to hypothesis planning, and retain the old tasks as regression checks. Do not repeatedly run the expanded suite while debugging that one failure.
This curriculum changes discovery coverage, not the overall success criteria. Reserve broader evaluation for coverage-expansion checkpoints and final confirmation. A single-task win cannot complete a broader objective.
Replan from what was learned. After a hypothesis batch, a new failure, or several non-improving trials, consolidate the map and choose the next most informative batch. Change the suspected mechanism or experiment design when the evidence contradicts it; do not endlessly retune one family or rerun the same tests. Revisit a rejected idea only with a changed premise. Continue without asking whether to proceed after each experiment.
Use the project's existing format; keep these views compact and linked to evidence:
Continue within the authorized session and limits until the target passes its required validation, the user stops the work, a budget is exhausted, or a genuine external blocker prevents useful progress. A plateau calls for replanning, not an invented claim of optimality. Preserve research state across continuations; this skill does not itself schedule future work.
At a stopping boundary, leave the incumbent recoverable and account for pending jobs. Deliver the parameter-effect map, baseline versus best verified results, coverage and limitations, budget consumed, stop reason, and exact next experiment if unfinished. Meeting a score does not replace recording what the experiments established; an unfinished run still leaves a useful map of what is known.
© dzhng, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/engineering/auto-research of dzhng/skills.
Open the folder on GitHubat commit d513228
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in dzhng/skills, which our catalogue first saw on October 7, 2026.
Auto Research next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Auto Research this skilldzhng/skills | 1k | — | ~2.7k | Automated safety check: Pass | MIT | |
| Experience MapOwl-Listener/designer-skills | 2.9k | 1 repos | ~351 | Automated safety check: Pass | MIT | |
| Parametersthedaviddias/Front-End-Checklist | 74k | — | ~638 | Automated safety check: Pass | MIT | |
| Token Mapnexu-io/open-design | 100k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| Finding ExperimentsPostHog/posthog | 40k | — | ~826 | Automated safety check: Pass | Custom licence | |
| ExperimentsArize-ai/phoenix | 12k | — | ~1.8k | Automated safety check: Pass | Custom licence |
Owl-Listener/designer-skills
Map the full ecosystem of touchpoints, channels, and relationships across a service.
thedaviddias/Front-End-Checklist
A skill your agent uses when auditing URL structure or configuring search engine handling of filtered, sorted, or tracked URLs.
nexu-io/open-design
Map an extracted Figma / source-code token bag onto the active OD design system, producing a deterministic mapping the generate stage can consume.
PostHog/posthog
Resolves a PostHog experiment reference from natural language to a concrete experiment ID by browsing experiment-list (not feature-flag tools), with disambiguation when multiple experiments match.
Arize-ai/phoenix
Run, read, and compare dataset-backed experiments to find evidence that a prompt or pipeline is improving.
asgeirtj/system_prompts_leaks
Accurate maps from real geo data — use for any map, or whenever geography would make a good graphic for a deliverable
dzhng/skills
Compare screenshots against the intended design, distinguishing approved references from historical baselines.
dzhng/skills
Use Claude Code as an independent claude -p subagent when the user explicitly asks for Claude, wants a second-agent opinion from Claude, or asks to delegate a well-scoped task to Claude.
dzhng/skills
Refactor cleanly instead of layering sediment. An agent skill from dzhng/skills.
dzhng/skills
Create or revise agent skills. An agent skill from dzhng/skills.
dzhng/skills
Use the local Codex CLI as an independent second agent. An agent skill from dzhng/skills.
dzhng/skills
Audit or rewrite AGENTS.md so it holds only lasting principles.
Run fast, progressive experiments to map tunable parameters and their effects. Auto Research is an agent skill from dzhng/skills. Run fast, progressive experiments to map tunable parameters and their effects.
Auto Research fits situations like: optimizing code; other artifacts against a goal; A long research run needs a planned hypothesis queue; durable findings.
Run `npx skills add dzhng/skills --skill auto-research -a claude-code`. Or copy the skill folder (skills/engineering/auto-research in dzhng/skills) into .claude/skills/auto-research in your project. Claude Code loads it when a task matches its description.
Run `npx skills add dzhng/skills --skill auto-research -a codex`. Or copy the skill folder (skills/engineering/auto-research in dzhng/skills) into .agents/skills/auto-research in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dzhng/skills --skill auto-research -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/auto-research, .gemini/skills/auto-research, .github/skills/auto-research and .opencode/skills/auto-research in your project.
SKILL.md names no scripts, command-line tools or credentials: Auto Research is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Auto Research is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Auto Research: Experience Map (Owl-Listener/designer-skills, 2.9k stars), Parameters (thedaviddias/Front-End-Checklist, 74k stars), Token Map (nexu-io/open-design, 100k stars) and Finding Experiments (PostHog/posthog, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
dzhng (a GitHub user) maintains it in dzhng/skills, which has 1,020 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on October 5, 2026.
Source: dzhng/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.