Show Me Your Work Decision Log
cursor/plugins
Keeps a TSV decision log for long or unattended agent runs, one row per decision with what, why, evidence and result, so a reviewer can check the work later.
Human-In-The-Loop gate that presents the action plan with full context, collects an informed approval/modification/rejection decision, and records the outcome.
$ npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-6-hitl -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install kayba-ai/agentic-context-engine kayba-stage-6-hitl --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/kayba-ai/agentic-context-engine.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/kayba-pipeline/stage-6-hitl .claude/skills/kayba-stage-6-hitl && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "kayba-stage-6-hitl" agent skill from https://github.com/kayba-ai/agentic-context-engine/tree/main/.claude/skills/kayba-pipeline/stage-6-hitl into .claude/skills/kayba-stage-6-hitl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kayba-stage-6-hitl", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/kayba-ai/agentic-context-engine/tree/main/.claude/skills/kayba-pipeline/stage-6-hitlType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-6-hitl -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install kayba-ai/agentic-context-engine kayba-stage-6-hitl --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kayba-ai/agentic-context-engine.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/kayba-pipeline/stage-6-hitl .agents/skills/kayba-stage-6-hitl && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "kayba-stage-6-hitl" agent skill from https://github.com/kayba-ai/agentic-context-engine/tree/main/.claude/skills/kayba-pipeline/stage-6-hitl into .agents/skills/kayba-stage-6-hitl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kayba-stage-6-hitl", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-6-hitl -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install kayba-ai/agentic-context-engine kayba-stage-6-hitl --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kayba-ai/agentic-context-engine.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/kayba-pipeline/stage-6-hitl .cursor/skills/kayba-stage-6-hitl && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "kayba-stage-6-hitl" agent skill from https://github.com/kayba-ai/agentic-context-engine/tree/main/.claude/skills/kayba-pipeline/stage-6-hitl into .cursor/skills/kayba-stage-6-hitl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kayba-stage-6-hitl", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/kayba-ai/agentic-context-engine.git --path .claude/skills/kayba-pipeline/stage-6-hitl--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-6-hitl -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install kayba-ai/agentic-context-engine kayba-stage-6-hitl --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kayba-ai/agentic-context-engine.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/kayba-pipeline/stage-6-hitl .gemini/skills/kayba-stage-6-hitl && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "kayba-stage-6-hitl" agent skill from https://github.com/kayba-ai/agentic-context-engine/tree/main/.claude/skills/kayba-pipeline/stage-6-hitl into .gemini/skills/kayba-stage-6-hitl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kayba-stage-6-hitl", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install kayba-ai/agentic-context-engine kayba-stage-6-hitlInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-6-hitl -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/kayba-ai/agentic-context-engine.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/kayba-pipeline/stage-6-hitl .github/skills/kayba-stage-6-hitl && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "kayba-stage-6-hitl" agent skill from https://github.com/kayba-ai/agentic-context-engine/tree/main/.claude/skills/kayba-pipeline/stage-6-hitl into .github/skills/kayba-stage-6-hitl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kayba-stage-6-hitl", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-6-hitl -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install kayba-ai/agentic-context-engine kayba-stage-6-hitl --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kayba-ai/agentic-context-engine.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/kayba-pipeline/stage-6-hitl .opencode/skills/kayba-stage-6-hitl && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "kayba-stage-6-hitl" agent skill from https://github.com/kayba-ai/agentic-context-engine/tree/main/.claude/skills/kayba-pipeline/stage-6-hitl into .opencode/skills/kayba-stage-6-hitl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kayba-stage-6-hitl", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
kayba-stage-6-hitlHuman-In-The-Loop gate that presents the action plan with full context, collects an informed approval/modification/rejection decision, and records the outcome.
Kayba Stage 6 Hitl is an agent skill from kayba-ai/agentic-context-engine. Human-In-The-Loop gate that presents the action plan with full context, collects an informed approval/modification/rejection decision, and records the outcome. Trigger when the user says "run stage 6", "HITL review", "approve action plan", or when invoked by the kayba-pipeline orchestrator. Requires eval/actionplan.md and eval/baselinemetrics.md to exist.
Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Agent Workflows, covering Human-in-the-loop approvals. The repository describes itself as: 🧠 Make your agents learn from experience. Now available as a hosted solution at kayba.ai. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 3a31983. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Kayba Stage 6 Hitl loads about 2.6k tokens when it runs. Until then it costs about 95 tokens; SKILL.md has 877 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from kayba-ai/agentic-context-engine at commit 3a31983, republished under its Apache-2.0 licence (© kayba-ai). 877 words, ~2,621 tokens.
.claude/skills/kayba-stage-6-hitl/SKILL.md (or your agent's skills folder).Present the action plan with enough context for an informed decision, collect the user's approval, and record the outcome.
The goal is not rubber-stamping. The user must receive enough information to genuinely evaluate, modify, or reject the plan -- even if they have not seen Stages 1-5.
eval/action_plan.md -- the prioritized action plan from Stage 5eval/baseline_metrics.md -- the evaluation rubric with baseline valueseval/baseline_metrics.json -- raw metric data (for exact numerator/denominator counts)eval/stage1_insights_summary.md -- original insights (for trace evidence references)Read all four files before starting.
Compute and present the following counts from the action plan:
Format:
EXECUTIVE SUMMARY
-----------------
Insights analyzed: 19 (raw) -> 12 distinct after dedup
Actionable: 9 (8 prompt fixes, 1 code fix)
Discarded: 3 (reasons listed below)
Discards:
- 5ac7f4ce (Upfront Info Collection): conflicts with higher-priority turn discipline
- fe2d51cb (Proactive Reservation Lookup): already default behavior, no failure evidence
- 1fa1b826 (Cancellation Denial Enumeration): subsumed into cancellation checklistFor each of the top 3 fixes by priority, present:
Before/after behavior -- use concrete examples from actual traces referenced in the insights. Quote the specific agent behavior that was wrong (before) and describe what the agent should do instead (after). Reference the trace task ID.
Target metric delta -- which metric(s) this fix targets, the current baseline value, and the expected direction. Do not fabricate precise target numbers. Use the format: "M1: 41.4% -> higher (target: 90%+)" only when the action plan provides a target; otherwise use "M1: 41.4% -> up".
Risk rating -- assess each fix:
Low -- additive prompt instruction, no behavioral side effects expectedMedium -- changes existing behavior, could affect adjacent workflowsHigh -- modifies code/infrastructure, or could degrade a metric while improving anotherFormat each as a numbered block:
#1: Turn Discipline (covers 55c00c40, d9683144)
Type: prompt fix
Metrics: M1 (41.4% -> up), M2 (20.7% -> up)
Risk: Low
BEFORE (task_1, task_5, task_7, ...):
Agent batches 2-3 tool calls per turn (e.g., get_reservation + get_flight_status
in a single response). Also includes user-facing text alongside tool calls.
AFTER:
Exactly one tool call per response. No user-facing content in tool-call turns.
Agent processes each result before making the next call.Display all non-discarded fixes in a table:
| Priority | Fix Name | Type | Target Metrics | Risk | Effort |
|----------|-----------------------------------|------------|-----------------|--------|--------|
| 1 | Turn Discipline | prompt fix | M1, M2 | Low | Low |
| 2 | Post-Confirmation Execution | prompt fix | M3 | Low | Low |
| 3 | Cancellation Checklist | prompt fix | M5 | Low | Low |
| ... | ... | ... | ... | ... | ... |Effort ratings:
Low -- single prompt addition, under 5 linesMedium -- multiple prompt additions or minor code changeHigh -- significant code changes, new metric implementation, or architectural changesList every discarded insight with:
This section exists so the user can override a discard if they disagree.
Any metric with denominator < 5 must be explicitly called out:
LOW-CONFIDENCE METRICS (small sample size):
- M5 (Cancellation Policy Compliance): based on 2 observations -- directional only
- M6 (Compensation Execution Rate): based on 1 observation -- directional only
Fixes targeting these metrics (Cancellation Checklist, Compensation Rules) are
still recommended because the policy violations are clear from trace evidence,
but the measured improvement may not be statistically meaningful until the
trace corpus grows.Also flag any fix where the action plan notes uncertainty or partial evidence.
For each fix, present the chain: insight -> metric -> fix -> expected improvement. This can be a compact list or a table. The purpose is to let the user verify that nothing was lost or invented between stages.
TRACEABILITY:
55c00c40 (Tool Call Discipline) -> M1, M2 -> Skill 1 (Turn Discipline) -> M1 up, M2 up
6ea141e1 (Execution Discipline) -> M3 -> Skill 2 (Post-Confirmation) -> M3 up
0f4a952b + 6ce88ebb (Cancellation) -> M5 -> Skill 3 (Cancellation Checklist) -> M5 up
...Present exactly three options:
OPTIONS:
[A] Approve all -- implement all 9 fixes as described
[B] Approve with modifications -- review each fix individually
[C] Reject -- return to Stage 5 with feedbackUse the appropriate mechanism to collect the user's choice (direct question or AskUserQuestion if available).
Record the decision and proceed. No further interaction needed.
Walk through each fix individually, in priority order. For each fix, present:
Then ask: "Approve / Skip / Modify?"
After walking through all fixes, present a summary of changes:
Ask for final confirmation: "Proceed with this modified plan?"
Then update eval/action_plan.md:
Ask the user for specific feedback:
Record the feedback in eval/stage6_decision.md and signal that Stage 5 should be re-run with the user's feedback incorporated.
Write this file regardless of which option was selected.
# Stage 6: HITL Decision Record
## Date
[timestamp]
## Decision
[Approve all | Approve with modifications | Reject]
## What was presented
- Total insights: N (M distinct after dedup)
- Actionable fixes: X (Y prompt, Z code)
- Discarded: W
- Metrics: [list metric IDs and baselines]
- Low-confidence flags: [list metrics with small denominators]
## Top 3 changes presented
1. [fix name] -- [type] -- targets [metrics] -- risk [rating]
2. ...
3. ...
## Decision details
### If Approve all:
User approved all N fixes without modification.
Reasoning: [any reasoning the user provided, or "No additional reasoning provided"]
### If Approve with modifications:
| Fix | Original Status | Decision | Reason |
|-----|----------------|----------|--------|
| Turn Discipline | Priority 1 | Approved | -- |
| Compensation Rules | Priority 5 | Modified | User changed wording to... |
| Cabin Change Rules | Priority 8 | Skipped | User considers low priority |
Modifications detail:
- [Fix name]: Original: "..." -> Modified: "..." -- User rationale: "..."
### If Reject:
User feedback: [verbatim feedback]
Specific concerns: [list]
Re-run instructions for Stage 5: [what to change]
## Traceability snapshot
[Copy of the traceability chain from step 6, so the decision record is self-contained]If the user selected [B] and made changes:
eval/action_plan.md unless the user explicitly requests modifications.eval/stage6_decision.md -- full record of what was presented, decided, and whyeval/action_plan.md -- updated only if the user selected "Approve with modifications"© kayba-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/kayba-pipeline/stage-6-hitl of kayba-ai/agentic-context-engine.
Open the folder on GitHubat commit 3a31983
Kayba Stage 6 Hitl next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Kayba Stage 6 Hitl this skillkayba-ai/agentic-context-engine | 2.6k | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | |
| Show Me Your Work Decision Logcursor/plugins | 11k | 8 repos | ~1.6k | Automated safety check: Pass | None | |
| Darwin Skill Optimizeralchaincyf/darwin-skill | 6.2k | 1 repos | ~4.7k | Automated safety check: Pass | MIT | |
| Loop Constraints Enforcercobusgreyling/loop-engineering | 11k | 1 repos | ~475 | Automated safety check: Notes | MIT | |
| Ask User QuestionMemTensor/MemOS | 12k | — | ~1k | Automated safety check: Pass | Apache-2.0 | |
| PUA High-Agency Governancetanweai/pua | 20k | — | ~502 | Automated safety check: Pass | MIT |
cursor/plugins
Keeps a TSV decision log for long or unattended agent runs, one row per decision with what, why, evidence and result, so a reviewer can check the work later.
alchaincyf/darwin-skill
Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.
cobusgreyling/loop-engineering
Loads a project's loop-constraints.md before any other action and blocks pushes, edits or merges that violate the rules it defines.
MemTensor/MemOS
Shows a question as a modal in the interface to clarify a task, collect a preference or get approval, since the user cannot see terminal output.
tanweai/pua
Pushes an agent to keep verifying and changing approach after repeated failures, using a diagnosis line, evidence-based completion and confirmation before risky edits.
rohitg00/agentmemory
Deletes chosen memories from agentmemory only after showing the matches and getting an explicit yes, for privacy requests and cleanup of outdated notes.
kayba-ai/agentic-context-engine
End-to-end agent evaluation and improvement pipeline. An agent skill from kayba-ai/agentic-context-engine.
kayba-ai/agentic-context-engine
Fetch pre-computed insights from the Kayba API and build a structured summary.
kayba-ai/agentic-context-engine
Gather domain context about the repository and agent — system prompt, tool definitions, domain docs, and behavior patterns from traces.
kayba-ai/agentic-context-engine
Define metrics from Kayba insights, implement them as Python measurement code, run against traces, and iterate until the metrics are clean and meaningful.
kayba-ai/agentic-context-engine
Organize computed metrics into a tiered evaluation rubric with leading, lagging, and quality indicators.
kayba-ai/agentic-context-engine
Triage each insight into discard/code-fix/prompt-fix and produce a prioritized action plan with specific recommendations.
Categories
Human-In-The-Loop gate that presents the action plan with full context, collects an informed approval/modification/rejection decision, and records the outcome. Kayba Stage 6 Hitl is an agent skill from kayba-ai/agentic-context-engine. Human-In-The-Loop gate that presents the action plan with full context, collects an informed approval/modification/rejection decision, and records the outcome.
Kayba Stage 6 Hitl fits situations like: the user says run stage 6; approve action plan; invoked by the kayba-pipeline orchestrator.
Run `npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-6-hitl -a claude-code`. Or copy the skill folder (.claude/skills/kayba-pipeline/stage-6-hitl in kayba-ai/agentic-context-engine) into .claude/skills/kayba-stage-6-hitl in your project. Claude Code loads it when a task matches its description.
Run `npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-6-hitl -a codex`. Or copy the skill folder (.claude/skills/kayba-pipeline/stage-6-hitl in kayba-ai/agentic-context-engine) into .agents/skills/kayba-stage-6-hitl in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-6-hitl -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kayba-stage-6-hitl, .gemini/skills/kayba-stage-6-hitl, .github/skills/kayba-stage-6-hitl and .opencode/skills/kayba-stage-6-hitl in your project.
SKILL.md names no scripts, command-line tools or credentials: Kayba Stage 6 Hitl is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Kayba Stage 6 Hitl is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Kayba Stage 6 Hitl: Show Me Your Work Decision Log (cursor/plugins, 11k stars), Darwin Skill Optimizer (alchaincyf/darwin-skill, 6.2k stars), Loop Constraints Enforcer (cobusgreyling/loop-engineering, 11k stars) and Ask User Question (MemTensor/MemOS, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
kayba-ai (a GitHub organization) maintains it in kayba-ai/agentic-context-engine, which has 2,590 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on September 24, 2026.
Source: kayba-ai/agentic-context-engine on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.