Dimensions
thedaviddias/Front-End-Checklist
A skill your agent uses when reviewing image assets, markup, and CDN or build transforms related to Set explicit width and height on images.
Assesses and rates quality 0-10 across multiple dimensions (correctness, maintainability, security, performance, testability, simplicity) with pros/cons analysis.
$ npx skills add yonatangross/orchestkit --skill assess -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install yonatangross/orchestkit assess --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/skills/assess .claude/skills/assess && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "assess" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/assess into .claude/skills/assess/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "assess", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/yonatangross/orchestkit/tree/main/src/skills/assessType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add yonatangross/orchestkit --skill assess -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install yonatangross/orchestkit assess --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/src/skills/assess .agents/skills/assess && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "assess" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/assess into .agents/skills/assess/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "assess", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add yonatangross/orchestkit --skill assess -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install yonatangross/orchestkit assess --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/src/skills/assess .cursor/skills/assess && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "assess" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/assess into .cursor/skills/assess/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "assess", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/yonatangross/orchestkit.git --path src/skills/assess--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add yonatangross/orchestkit --skill assess -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install yonatangross/orchestkit assess --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/src/skills/assess .gemini/skills/assess && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "assess" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/assess into .gemini/skills/assess/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "assess", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install yonatangross/orchestkit assessInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add yonatangross/orchestkit --skill assess -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .github/skills && cp -r skills-src/src/skills/assess .github/skills/assess && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "assess" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/assess into .github/skills/assess/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "assess", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add yonatangross/orchestkit --skill assess -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install yonatangross/orchestkit assess --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/src/skills/assess .opencode/skills/assess && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "assess" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/assess into .opencode/skills/assess/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "assess", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
assessAssesses and rates quality 0-10 across multiple dimensions (correctness, maintainability, security, performance, testability, simplicity) with pros/cons analysis.
Assess is an agent skill from yonatangross/orchestkit. Assesses and rates quality 0-10 across multiple dimensions (correctness, maintainability, security, performance, testability, simplicity) with pros/cons analysis. Compares against project conventions and prior decisions from memory. Produces structured evaluation reports with actionable improvement suggestions. Use when evaluating code, designs, architectures, or comparing alternative approaches.
Its SKILL.md is about 6.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 34 other files, including scripts, reference files and assets (for example `assets/assessment-report.md`, `assets/comparison-table.md` and `checklists/assessment-checklist.md`). Compatibility notes: Claude Code 2.1.277+. Requires memory MCP server.
The repository describes itself as: The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install ork for stable (v9.x), or ork-alpha for the v10 line, which ships daily. The licence is MIT.
Read from SKILL.md and the folder at commit 02bbf9a. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
AskUserQuestionReadWriteGrepGlobAgentWorkflowTaskCreateTaskUpdateTaskList…and 3 more on the same allowed-tools line.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
Shell commands in SKILL.md call:
nodepython3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Claude Code 2.1.277+. Requires memory MCP server.
From compatibility in the SKILL.md frontmatter.
Assess loads about 6.6k tokens when it runs, and up to ~16k if it reads all its reference files. Until then it costs about 102 tokens; SKILL.md has 2,317 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: AskUserQuestion, Read, Write, Grep, Glob, Agent, Workflow, TaskCreate, TaskUpdate, TaskList, ToolSeaAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from yonatangross/orchestkit at commit 02bbf9a, republished under its MIT licence (© yonatangross). 2,317 words, ~6,564 tokens.
.claude/skills/assess/SKILL.md (or your agent's skills folder). This skill also uses 30 other files; get the full folder from GitHub.Host-neutral workflow. Invoke by skill name (assess). Claude Code slash routing, YAML hook loaders, and .claude/chain live in references/claude-code.md.
Comprehensive assessment skill for answering "is this good?" with structured evaluation, scoring, and actionable recommendations.
assess backend/app/services/auth.py
assess our caching strategy
assess --model=opus the current database schema
assess frontend/src/components/Dashboardxhigh)| Effort | Behavior |
|---|---|
low / medium | Subset of dimensions, faster turnaround |
high (default) | All six dimensions with pros/cons |
xhigh | All six dimensions + one additional assessor pass focused on uncertainty/caveats; emits confidence per dimension |
xhighsilently falls back tohighon a model that does not implement it: no error, no log line.doctorCategory 14 reports this, and only when it can positively prove the active model lacks the tier.
$ARGUMENTS is often not a path. For a bare pronoun or deictic (them, this, that,
these, they, same, the above, the last one, what we just did) or an empty target
after flags are stripped, the subject is in the conversation. Read back for the NEAREST
concrete one (a file just discussed, a diff or PR just opened, a component just investigated)
and announce the resolution in one line, so a wrong guess costs a correction rather than a
turn: "Reading 'them' as the 3 pretool guards we just probed; say otherwise and I'll switch."
Refusing is the bug, not the safe option. Asking "what does this refer to?" when the
previous turn named the subject burns a round-trip re-deriving what is already on screen.
Measured 2026-08-28: the operator sent assess them throguhly one message after "bug in
orchestkit hooks", mid-investigation of pretool/bash/dangerous-command-blocker, and this
skill replied that "them" had "no antecedent anywhere in this conversation". It had two.
Ask only when the conversation is genuinely empty (a fresh session opening with a bare pronoun). Every other case: resolve and announce.
Not unique to this skill:
verify,cover,fix-issue,review-prandimplementall read$ARGUMENTSas a literal path or topic, and no skill mentions resolving a reference. Tracked separately; this one fixes its own door.
TARGET = "$ARGUMENTS" # Full argument string, e.g., "backend/app/services/auth.py"
# $ARGUMENTS[0] is the first token (CC 2.1.59 indexed access)
# Model override detection (CC 2.1.72)
MODEL_OVERRIDE = None
for token in "$ARGUMENTS".split():
if token.startswith("--model="):
MODEL_OVERRIDE = token.split("=", 1)[1] # "opus", "sonnet", "haiku", "fable"
TARGET = TARGET.replace(token, "").strip()Pass MODEL_OVERRIDE to all Agent() calls via model=MODEL_OVERRIDE when set. Accepts symbolic names (opus, sonnet, haiku, fable on harnesses whose Agent tool lists it; note fable is premium API spend after 2026-07-12) or full IDs (claude-opus-5-5) per CC 2.1.74.
Switching to Opus via
/model(CC 2.1.144+):/modelnow changes the model for the current session only, so picking Opus for an assess run no longer persists past it. Pressdin the picker only to set a default for new sessions.
$CLAUDE_EFFORT is the primary signal. CC 2.1.120 sets this env var from /effort or the model picker. --effort= token in $ARGUMENTS is the explicit override fallback (also covers older CC).
# Read env first (CC 2.1.120+), then check explicit override
EFFORT = os.environ.get("CLAUDE_EFFORT") # "low" | "medium" | "high" | "xhigh" | None
for token in "$ARGUMENTS".split():
if token.startswith("--effort="):
EFFORT = token.split("=", 1)[1] # explicit override wins
TARGET = TARGET.replace(token, "").strip()
EFFORT = EFFORT or "high" # default when CC < 2.1.120 and no flagUse EFFORT to gate dimension count, agent count, and the optional xhigh uncertainty pass — see "Effort levels" table above. On CC < 2.1.120 the env var is unset; the explicit --effort= override is the only path. doctor Category 14 reports a provably unsupported xhigh request.
Load:
Read("../chain-patterns/references/mcp-detection.md")
# 1. Probe MCP servers (once at skill start)
# memory is alwaysLoad in .mcp.json (CC 2.1.121+, #1541) — probe below kept as fallback for older CC:
ToolSearch(query="select:mcp__memory__search_nodes")
# 2. Store capabilities
Write(".claude/chain/capabilities.json", {
"memory": probe_memory.found,
"skill": "assess",
"timestamp": now()
})
# 3. Check for resume
state = Read(".claude/chain/state.json") # may not exist
if state.skill == "assess" and state.status == "in_progress":
last_handoff = Read(f".claude/chain/{state.last_handoff}")| Phase | Handoff File | Contents |
|---|---|---|
| 0 | 00-intent.json | Dimensions, target, mode |
| 1 | 01-baseline.json | Initial codebase scan results |
| 2 | 02-evaluation.json | Per-dimension scores + evidence |
| 3 | 03-report.json | Final report, grade, recommendations |
BEFORE creating tasks, clarify assessment dimensions:
AskUserQuestion(
questions=[{
"question": "What dimensions to assess?",
"header": "Dimensions",
"options": [
{"label": "Full assessment (Recommended)", "description": "All dimensions: quality, maintainability, security, performance"},
{"label": "Code quality only", "description": "Readability, complexity, best practices"},
{"label": "Security focus", "description": "Vulnerabilities, attack surface, compliance"},
{"label": "Quick score", "description": "Just give me a 0-10 score with brief notes"}
],
"multiSelect": false
}]
)Based on answer, adjust workflow (passed to Phase 2 as focus):
full): All 7 phases, parallel agents, dimension subset scaled by effortquality): Skip security and performance phasessecurity): Prioritize security-auditor agentquick): Single pass, brief outputLoad details: Read("references/orchestration-mode.md") for env var check logic, Agent Teams vs Agent Tool comparison, and mode selection rules.
# 1. Create main task IMMEDIATELY
TaskCreate(
subject="Assess: {target}",
description="Comprehensive evaluation with quality scores and recommendations",
activeForm="Assessing {target}"
)
# 2. Create subtasks for each assessment phase
TaskCreate(subject="Understand target and gather context", activeForm="Understanding target") # id=2
TaskCreate(subject="Discover scope and build file list", activeForm="Discovering scope") # id=3
TaskCreate(subject="Rate quality across 6 dimensions", activeForm="Rating quality") # id=4
TaskCreate(subject="Analyze pros and cons", activeForm="Analyzing pros/cons") # id=5
TaskCreate(subject="Compare alternatives", activeForm="Comparing alternatives") # id=6
TaskCreate(subject="Generate improvement suggestions", activeForm="Generating suggestions") # id=7
TaskCreate(subject="Compile assessment report", activeForm="Compiling report") # id=8
# 3. Set dependencies for sequential phases
TaskUpdate(taskId="3", addBlockedBy=["2"]) # Scope needs target understanding
TaskUpdate(taskId="4", addBlockedBy=["3"]) # Rating needs scoped file list
TaskUpdate(taskId="5", addBlockedBy=["4"]) # Pros/cons needs quality scores
TaskUpdate(taskId="6", addBlockedBy=["4"]) # Alternatives need quality scores
TaskUpdate(taskId="7", addBlockedBy=["5", "6"]) # Suggestions need analysis
TaskUpdate(taskId="8", addBlockedBy=["7"]) # Report needs suggestions
# 4. Update status as you progress
TaskUpdate(taskId="2", status="in_progress") # When starting
TaskUpdate(taskId="2", status="completed") # When done — repeat for each subtask| Phase | Activities | Output |
|---|---|---|
| 1. Target Understanding | Read code/design, identify scope | Context summary |
| 1.5. Scope Discovery | Build bounded file list | Scoped file list |
| 2. Quality Rating | 6-dimension scoring (0-10) | Scores with reasoning |
| 3. Pros/Cons Analysis | Strengths and weaknesses | Balanced evaluation |
| 4. Alternative Comparison | Score alternatives | Comparison matrix |
| 5. Improvement Suggestions | Actionable recommendations | Prioritized list |
| 6. Effort Estimation | Time and complexity estimates | Effort breakdown |
| 7. Assessment Report | Compile findings | Final report |
Identify what's being assessed and gather context. TARGET here is the value Step 0 already
resolved, which is not necessarily what the user typed.
# PARALLEL - Gather context
Read(file_path=TARGET) # only when TARGET is a path
Grep(pattern=TARGET, output_mode="files_with_matches") # topic or symbol
mcp__memory__search_nodes(query=TARGET) # past decisionsRead failing is NOT a reason to stop. A target resolved from the conversation is usually a
subject rather than a filename ("the three pretool guards", "today's hook fixes"), so the Read
misses and the Grep plus the conversation carry the context. Treat a failed Read as "this is a
topic, not a path" and continue to Phase 1.5, which discovers the real file list anyway.
Load Read("references/scope-discovery.md") for the full file discovery, limit application (MAX 30 files), and sampling priority logic. Always include the scoped file list in every agent prompt.
Output results incrementally as each evaluation phase completes:
| After Phase | Show User |
|---|---|
| 1. Target Understanding | Scope summary, file list, context |
| 1.5. Scope Discovery | Bounded file list (max 30 files) |
| 2. Quality Rating | Each dimension's score as the evaluating agent returns |
| 3. Pros/Cons | Balanced evaluation summary |
The Phase 2 workflow returns once, so show every dimension's score from its result and lead with priorityConcerns (any dimension below 4/10) as a concern needing user attention. Per-agent streaming applies only on the Agent tool fallback.
Rate each dimension 0-10 with weighted composite score. Load Read("../quality-gates/references/unified-scoring-framework.md") for dimensions, weights, grade interpretation, and per-dimension criteria. Load Read("references/quality-model.md") for assess-specific overrides.
Do NOT hand-roll the assessors. Run the executor, which owns Phases 2 and 2.5:
result = Workflow(
scriptPath="${CLAUDE_SKILL_DIR}/workflows/assess-fanout.js",
args={"target": TARGET, "effort": EFFORT, "focus": FOCUS, # FOCUS from STEP 0
"mode": "comparison" if COMPARING else "default", # quality-model.md
"domain": "frontend" or "backend", # picks the performance engineer
"scopeFiles": SCOPE_FILES, # Phase 1.5 list
"projectContext": MEMORY_CONTEXT, # Phase 1 memory search
"rubric": Read("rubric.json"), "modelOverride": MODEL_OVERRIDE,
"feature": FEATURE})
Write(".claude/chain/02-evaluation.json", result)The script owns the mechanics, not the prose. It picks the assessors from focus and effort (security first), gives each a score schema that demands file:line evidence, sends every decision-bearing score to blind refuters (Phase 2.5), and computes the weighted composite, grade, rubric verdict and blockers. A score with no file:line evidence counts as unscored, a repeated dimension keeps only its first entry, and a selected dimension with a min_blocker that nobody scored is a blocker. It returns composite, grade, verdict, blockers (producer basis), postRefutation, chainVerdict, chainVerdictIfConfirmed, revisions, confirmationNeeded, manualReview, advisory, priorityConcerns, quickWins, unscored, rejectedDimensions, unassessed, dimensions, ledger and reasons. It never asks and never writes: those stay in this shell.
Fallback (no Workflow tool, ORCHESTKIT_FORCE_TASK_TOOL=1, or the cross-model lane below): Read("references/agent-spawn-definitions.md") for Agent tool and Agent Teams spawns, then run Phase 2.5 by hand.
Composite Score: Weighted average of the scored dimensions (see quality-model.md).
The assessor that scores a dimension is also its only judge, a self-preferential bias. A separate blind refuter forms its own band for each decision-bearing score. Effort gate: low/medium skip it; high runs up to 4 single advisory refuters (no auto-swing); xhigh runs a 3-refuter majority that revises to the near band edge. On the Workflow path the script already ran it; the shell finishes it:
ledger to .claude/chain/02b-refutation.json (engine section 10).file:line in revisions (engine section 3). A citation that does not hold reverts that dimension to its producer score; recompute the composite with the returned weights.confirmationNeeded is non-empty, AskUserQuestion before using chainVerdictIfConfirmed; otherwise use chainVerdict. Refutation alone never raises a score or flips fail to pass (engine section 7).manualReview dimensions as "not independently refuted", and surface every advisory overturn at high effort.Protocol and assess bindings: Read("references/adversarial-refutation.md") (loads the shared engine ../../shared/rules/adversarial-refutation.md). Producer findings must first pass the evidence-replay gate before entering any score or verdict: Read("../../shared/rules/evidence-replay.md").
When ORK_ALT_MODEL_CMD is configured and effort is high/xhigh, one quorum slot per high-weight or boundary-adjacent dimension score can route to a non-Claude model (Codex/GPT) for diverse failure modes. Off by default; substitutes one same-model slot, stamps refuter_model for provenance, cannot silently raise the grade (engine §7), owns no credentials/egress (shells out via ORK_ALT_MODEL_CMD, matches the egress guard #2533), and degrades to same-model on an absent command. Shares the review-pr operational doc: Read("../review-pr/references/cross-model-refuter.md").
The workflow does not run this lane (a script cannot shell out). When the user wants it, choose the Agent tool fallback before Phase 2 and run Phases 2 and 2.5 there.
Refuters are ALWAYS isolated spawns with no team_name, fed only the dimension and the scoped files: no producer score, identity, or prose. Keep the producer-basis score AND a labeled post-refutation score.
Load Read("references/phase-templates.md") for output templates for pros/cons, alternatives, improvements, effort, and the final report.
See also: Read("references/alternative-analysis.md") | Read("references/improvement-prioritization.md")
Parse --render= from $ARGUMENTS. Default is both.
| Mode | Behavior |
|---|---|
markdown | Current behavior — markdown assessment report only. No spec emitted. |
json-render | Emit .claude/chain/assess-dashboard.json only. Skip markdown report. |
both | Emit spec and markdown. Default — human reads the report, downstream skills parse the spec. |
When emitting a spec:
Read("references/dashboard-spec.md"). Example: references/dashboard-example.json.Card, StatGrid, DataTable, StatusBadge, BarMeter, Markdown. Top-level fields composite (number) and grade (string) are required for assess specs.BarMeter per dimension scored. The verdict element is a StatusBadge with status success/warning/error mapped from grade (A/B → success, C → warning, D/F → error)..claude/chain/assess-dashboard.json with compact JSON.node "${CLAUDE_SKILL_DIR}/scripts/render-spec.mjs" .claude/chain/assess-dashboard.json --checkIf validation fails, fall back to markdown-only and surface the error. Never write a partial spec.
--render=both, render the markdown view from the spec:node "${CLAUDE_SKILL_DIR}/scripts/render-spec.mjs" .claude/chain/assess-dashboard.jsonThis guarantees JSON spec and markdown report stay in sync.
xhigh effort: when effort=xhigh is active, add a sibling Markdown element per dimension containing confidence and caveats from the uncertainty pass. Reference list it in the dimensions Card's children alongside the BarMeter. See references/dashboard-spec.md for the exact pattern.
Downstream consumption: implement reads .claude/chain/assess-dashboard.json and pulls the lowest-scoring dimension and high-priority improvements (effort ≤ 2 AND impact ≥ 4) without parsing markdown tables. Measured: assess spec ≈ 830 tokens vs ~3500 token markdown for the same content.
When the assessment lands with a composite score, optionally persist scores + summary to the memory MCP knowledge graph as a typed entity. Future memory queries can then surface assessment lineage (which decisions did this codebase score 9/10 on testability? when did security regress below 7.0?).
python3 ${CLAUDE_SKILL_DIR}/scripts/memory_writeback.py "<assessment-dir>"<assessment-dir> is the dir containing assessment.json (typically the session's .claude/chain/). The script writes a memory-writeback.json handoff alongside it.
Auto-skip conditions (all exit 0, all WARN-logged):
| Skip reason | Trigger |
|---|---|
no composite score | assessment.json has no top-level composite numeric field |
yg-mcp-core not importable | yg-mcp-core>=0.3.0 not installed (orchestkit is public; yg-mcp-core lives on private pypi.yonyon.ai — HQ-only) |
memory MCP unreachable | memory MCP server down OR .mcp.json doesn't define memory |
The created entity has:
name: <slug-or-dir>@<timestamp> (stable across re-runs — re-runs create new entities)entityType: assessment (override with --entity-type <type>)observations: composite=X.XX, one <dim>=X.XX per scored dimension, optional summary: ... and topic: ...Mirrors Yonatan-HQ/hq-ext-plugin#194 (audio_podcast handler) and orchestkit#1886 (post-synthesis podcast) pattern. Unblocked by Yonatan-HQ/core#993 (yg-mcp-core 0.3.0).
After the composite and grade are final (post-refutation, Phase 2.5), ALWAYS write the machine-readable verdict: this is the stop-gate implement reads before Phase 1. On the Workflow path it is the returned chainVerdict (or chainVerdictIfConfirmed after a yes in Phase 2.5). Mirror the Phase 7b spec-emit pattern: write compact JSON, never a partial file.
// .claude/chain/assess-verdict.json
{
"rubric": "ork-rubric/1.0",
"skill": "assess",
"verdict": "fail",
"composite": 5.1,
"dimension_scores": {"correctness": 7.0, "maintainability": 6.5, "performance": 5.5, "security": 3.2, "scalability": 6.0, "testability": 4.8, "compliance": 6.2},
"blockers": [
{"dimension": "security", "score": 3.2, "reason": "Unparameterized SQL in auth path (src/api/auth.ts:42)"}
],
"feature": "<assessment topic, e.g. first non-flag token of $ARGUMENTS>"
}Verdict rules — thresholds come from rubric.json (schema: ../../shared/rubric.schema.json):
verdict = "fail" when composite < min_pass (5.5) OR any dimension scores below its min_blocker. Otherwise "pass".min_blocker gets a blockers[] entry — dimension, score, one evidence-backed reason. blockers is [] on pass.Consumers: implement Step -0.5 blocks Phase 1 on verdict == "fail" (user must fix-first or explicitly override); Phase 7c memory writeback persists the verdict + dimension scores to the memory graph (add a verdict=pass|fail observation) for cross-session learning.
xhigh effort)Current-generation models report their own limits far better than older tiers did. When xhigh effort is active, enrich each dimension's rating with a confidence level and a list of caveats — things the model couldn't verify, assumptions it relied on, or cases it didn't test.
Output schema per dimension (JSON):
{
"dimension": "security",
"score": 7.2,
"confidence": "medium", // "low" | "medium" | "high"
"caveats": [
"Didn't execute the SQL queries against a real DB to confirm parameterization",
"Assumed NODE_ENV=production in deployment; didn't verify CI config",
"Reviewed 12 of 15 handlers; remaining 3 deferred by scope filter"
],
"evidence": ["src/api/auth.ts:42", "src/middleware/guard.ts:88"]
}Rules:
confidence as an auto-gate. It's a signal for the human reader, not a pass/fail threshold.caveats must be specific. "Didn't check X" with file paths beats "uncertainty about security".score only — not weighted by confidence — to keep the number comparable across runs.Load Read("../quality-gates/references/unified-scoring-framework.md") for grade thresholds and scoring criteria.
| Decision | Choice | Rationale |
|---|---|---|
| 6 dimensions | Comprehensive coverage | All quality aspects without overwhelming |
| 0-10 scale | Industry standard | Easy to understand and compare |
| Parallel assessment | workflows/assess-fanout.js, up to 4 assessors | Scores, refutation and verdict held in code, not prose |
| Effort/Impact scoring | 1-5 scale | Simple prioritization math |
| Rule | Impact | What It Covers |
|---|---|---|
complexity-metrics (load rules/complexity-metrics.md) | HIGH | 7-criterion scoring (1-5), complexity levels, thresholds |
complexity-breakdown (load rules/complexity-breakdown.md) | HIGH | Task decomposition strategies, risk assessment |
Done means all of these hold:
high/xhigh effort, decision-bearing scores passed the adversarial refutation lane before entering the composite.claude/chain/assess-verdict.json written with verdict pass/fail and a blockers[] entry for every dimension below its min_blockerrender-spec.mjs --check and carries the required composite + grade fieldsork:verify - Post-implementation verificationork:code-review-playbook - Code review patternsork:quality-gates - Task complexity assessment, gate patternsVersion: 1.9.0 (September 2026): Phases 2 and 2.5 run as a Workflow script (workflows/assess-fanout.js)
© yonatangross, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 30 other files (scripts, references, assets) in src/skills/assess of yonatangross/orchestkit.
Open the folder on GitHubat commit 02bbf9a
Assess next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Assess this skillyonatangross/orchestkit | 290 | — | ~6.6k | Automated safety check: Notes | MIT | |
| Dimensionsthedaviddias/Front-End-Checklist | 74k | — | ~926 | Automated safety check: Pass | MIT | |
| Web3 Rate Limiting Circuit Breakersickn33/agentic-awesome-skills | 47k | 1 repos | ~1.4k | Automated safety check: Pass | MIT | |
| Performing API Rate Limiting Bypassmukul975/Anthropic-Cybersecurity-Skills | 34k | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | |
| Acreadiness Assessgithub/awesome-copilot | 40k | 1 repos | ~839 | Automated safety check: Pass | MIT | |
| Relsa Severity AssessmentK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~5.2k | Automated safety check: Notes | MIT |
thedaviddias/Front-End-Checklist
A skill your agent uses when reviewing image assets, markup, and CDN or build transforms related to Set explicit width and height on images.
sickn33/agentic-awesome-skills
On-chain and relayer rate-limiting circuit breaker register: throughput thresholds, emergency pause triggers, and multi-sig recovery.
mukul975/Anthropic-Cybersecurity-Skills
Tests API rate limiting for bypass vulnerabilities using Python (requests/aiohttp) and Burp Suite Turbo Intruder to manipulate headers (e.g.
github/awesome-copilot
Run the AgentRC readiness assessment on the current repository and produce a static HTML dashboard at reports/index.html.
K-Dense-AI/scientific-agent-skills
Supports multivariate severity assessment and exploratory endpoint-time score forecasting for laboratory animal studies using the RELSA (RELative Severity Assessment) score and ARIMA-based foRcast…
mukul975/Anthropic-Cybersecurity-Skills
Implements API abuse detection using token bucket, sliding window, and fixed window rate-limiting algorithms backed by Redis, including adaptive limits that tighten during detected attacks and relax…
yonatangross/orchestkit
API contract design for REST and GraphQL, covering resource shape, URL and header versioning with deprecation windows, RFC 9457 Problem Details error handling, and OpenAPI specs.
yonatangross/orchestkit
ADR templates in the Nygard format with context, decision, consequences, and alternatives.
yonatangross/orchestkit
Single-pass codebase analysis leveraging a 1M-token context window for comprehensive security scanning, architecture review, and dependency auditing.
yonatangross/orchestkit
Structured review processes, conventional comments, language-specific checklists, and feedback templates.
yonatangross/orchestkit
Creates GitHub pull requests with pre-flight validation, conventional title formatting, and structured summary generation.
yonatangross/orchestkit
Multi-angle codebase exploration spawning 3-5 parallel agents for code structure, data flow, architecture patterns, and health assessment.
Assesses and rates quality 0-10 across multiple dimensions (correctness, maintainability, security, performance, testability, simplicity) with pros/cons analysis. Assess is an agent skill from yonatangross/orchestkit. Assesses and rates quality 0-10 across multiple dimensions (correctness, maintainability, security, performance, testability, simplicity) with pros/cons analysis.
Assess fits situations like: evaluating code; comparing alternative approaches.
Run `npx skills add yonatangross/orchestkit --skill assess -a claude-code`. Or copy the skill folder (src/skills/assess in yonatangross/orchestkit) into .claude/skills/assess in your project. Claude Code loads it when a task matches its description.
Run `npx skills add yonatangross/orchestkit --skill assess -a codex`. Or copy the skill folder (src/skills/assess in yonatangross/orchestkit) into .agents/skills/assess in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yonatangross/orchestkit --skill assess -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/assess, .gemini/skills/assess, .github/skills/assess and .opencode/skills/assess in your project.
Going by SKILL.md and its folder, Assess needs the command-line tools its instructions call (node and python3). Our summary lists: Python 3. Its frontmatter pre-approves these tools: AskUserQuestion, Read, Write, Grep, Glob, Agent, Workflow, TaskCreate, TaskUpdate, TaskList, ToolSearch, mcp__memory__search_nodes, Bash. Compatibility (from SKILL.md): Claude Code 2.1.277+. Requires memory MCP server..
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Assess is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.6k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.1k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Assess: Dimensions (thedaviddias/Front-End-Checklist, 74k stars), Web3 Rate Limiting Circuit Breaker (sickn33/agentic-awesome-skills, 47k stars), Performing API Rate Limiting Bypass (mukul975/Anthropic-Cybersecurity-Skills, 34k stars) and Acreadiness Assess (github/awesome-copilot, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
yonatangross (a GitHub user) maintains it in yonatangross/orchestkit, which has 290 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 9, 2026.
Source: yonatangross/orchestkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.