Agent skill

Evaluator

by alecs5am in alecs5am/ralphy

Quality evaluation of rendered UGC mp4s — scene segmentation, audio loudness / dead-air, caption density, and per-scene visual analysis.

Apache-2.0Auto-check passedTesting & QA

Install Evaluator

skills CLI
$ npx skills add alecs5am/ralphy --skill evaluator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install alecs5am/ralphy evaluator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/alecs5am/ralphy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/evaluator .claude/skills/evaluator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluator
GitHub stars
138
Token cost
~3.6k tokens
SKILL.md length
1,699 words
Files
2 (incl. references)
Skills in repo
28
Repo updated
First seen
Licence
Apache-2.0

At a glance

Quality evaluation of rendered UGC mp4s — scene segmentation, audio loudness / dead-air, caption density, and per-scene visual analysis.

  • Works in 3 steps: The user said "validate against… → The user shows you a scrape-profile… → The project has a discoverable…
  • The user asks to evaluate / score / grade / review / QA / check quality of a rendered video
  • SKILL.md covers Trigger refinements, Hard invariants, What this skill is not and The single command, plus 5 more sections
  • Needs OPENROUTER_API_KEY

What it does

Evaluator is an agent skill from alecs5am/ralphy. Quality evaluation of rendered UGC mp4s — scene segmentation, audio loudness / dead-air, caption density, and per-scene visual analysis. Produces an actionable report (eval.json + eval-report.md) sized for a downstream fixer agent. USE WHEN the user asks to "evaluate / score / grade / review / QA / check quality of" a rendered video, asks "is this video good?", drops an mp4 path with no other instruction, mentions "find issues / problems / artifacts", asks for retention or scroll-stop assessment, or has just…

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/report-schema.md`).

It sits in Testing & QA, covering Influencer and creator marketing and Quality gates. The repository describes itself as: Open-source desktop app for content creation, with an agent runtime and standalone CLI. The licence is Apache-2.0.

When your agent uses it

  • The user asks to evaluate / score / grade / review / QA / check quality of a rendered video
  • Asks is this video good?
  • Drops an mp4 path with no other instruction
  • Mentions find issues / problems / artifacts

Example prompts

  • “evaluate / score / grade / review / QA / check quality of”
  • “is this video good?”
  • “find issues / problems / artifacts”
  • “/evaluator”

Requirements

  • A credential in OPENROUTER_API_KEY

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. The user said "validate against [creator]" / "evaluate against my style" / "is this on-brand for [niche]".
  2. The user shows you a scrape-profile style-sheet path and then drops an mp4.
  3. The project has a discoverable STYLE_LOCK.md / style-sheet.md (auto-discovered by walking up from the mp4 path). Eval auto-upgrades the…

What it can do on your machine

Read from SKILL.md and the folder at commit 8d139f0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENROUTER_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluator loads about 3.6k tokens when it runs, and up to ~5.1k if it reads all its reference files. Until then it costs about 231 tokens; SKILL.md has 1,699 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~231
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from alecs5am/ralphy at commit 8d139f0, republished under its Apache-2.0 licence (© alecs5am). 1,699 words, ~3,582 tokens.

Download SKILL.mdSave it as .claude/skills/evaluator/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
evaluator
description
Quality evaluation of rendered UGC mp4s — scene segmentation, audio loudness / dead-air, caption density, and per-scene visual analysis. Produces an actionable report (eval.json + eval-report.md) sized for a downstream fixer agent. USE WHEN the user asks to "evaluate / score / grade / review / QA / check quality of" a rendered video, asks "is this video good?", drops an mp4 path with no other instruction, mentions "find issues / problems / artifacts", asks for retention or scroll-stop assessment, or has just rendered something and wants verification before publishing. TRIGGER (EN): "evaluate this video", "score the render", "grade the mp4", "review the final cut", "QA this video", "is this ready to ship", "what's wrong with this video", "find issues in <path.mp4>", "audit the video", "scene-by-scene breakdown", "retention check", "quality gate". See body for ALSO FIRE / DO NOT FIRE / HARD INVARIANTS.
namespace
user

evaluator

Trigger refinements

ALSO FIRE when the user just dropped a path that ends in .mp4 from .ralphy/workspaces/<ws>/projects/<id>/render/ with no other instructions, or when an editor handed off and the user asks "and now?" (any language).

DO NOT FIRE for unrendered projects (handback to editor for ralphy render), for raw research downloads (those go through researcher's analyze-video flow, not eval), or for source media that hasn't been composed yet.

Hard invariants

  • Every model call (vision pass) routes through cli/lib/providers/llm.ts → callLLM() via the CLI. No direct OpenAI / fal calls.
  • Findings are deterministic outputs of cli/lib/eval/* — don't paraphrase them; pass through verbatim to the fixer agent.
  • Keyframe slicing is a cheap diagnostic, NOT a ship gate (#411). A keyframe-only (or structure-only) report can NEVER mark a Unit ship-ready — report.gate.shipReady is hard-false on it. The final gate before forming/publishing a Unit is the native-video pass (full mp4 → model), or deep-style when a STYLE_LOCK/brief exists. Screenshot slicing misses temporal continuity, audio-picture alignment, pacing, and caption sync — exactly the failures that hallucinate when the model only sees stills.

Where this sits in the Unit lifecycle. Eval is phase 14 of the canonical Unit lifecycle. Its native-video gate is the one that flips polished to true in ralphy project status <id> --contract; the quality-gate-failed and native-gate-required stop conditions are derived from eval.json's gate / scoring.verdict. A block verdict feeds the repair loop (phase 15) and the optional polish council (phase 16).


You evaluate rendered UGC videos and produce a report that another agent (the fixer) can act on without reading the video itself. The contract is: the report is the handoff.

What this skill is not

  • Not a researcher tool. For "analyze this TikTok/Reel from a creator I want to imitate", route to /researcher.
  • Not a fixer. The findings list is meant to be read by a separate agent (or the editor / art-director / scenarist) that will execute the fixes. Don't try to fix issues from inside this skill — that's a different role and would skip the user's chance to triage.
  • Not a publisher / scheduler. Verdict is informational, not a publish gate.

The single command

bash
ralphy eval video <path-to-mp4>

Auto-detects the project ID when the mp4 lives at .ralphy/workspaces/<ws>/projects/<id>/render/... (the current layout — fixed in #411; the legacy workspace/projects/<id>/ shape still resolves as a fallback). If detected, the report incorporates scenario.json, captions.json, BRIEF.md, STYLE_LOCK.md, and the template name from the project — these unlock the declared-vs-actual findings (duration drift, hook-zone-thin-vo, intent-drift, etc.) that are otherwise unavailable.

Validation modes (--mode, #411)

Eval has four explicit modes, cheapest → most thorough. Choose by what you're doing: a quick smoke check vs. the final gate before a Unit ships.

ModeWhat it runsModel spendCan mark a Unit ship-ready?
structureDeterministic only: ffprobe, scene durations, loudness, dead-air, caption density. No model call.$0No
keyframestructure + the cheap per-scene keyframe vision pass (one still/scene, gemini-flash). A smoke check for blank/garbled frames.~$0.01No
native-videostructure + a full-mp4 model pass (gemini-3.1-pro-preview sees every frame at native temporal resolution) for temporal continuity, audio-picture alignment, pacing, caption sync, format fit. No style sheet required.model on full mp4Yes (when verdict passes)
deep-stylenative-video PLUS style-lock / brief / reference conformance scoring.model on full mp4Yes (when verdict passes)
bash
ralphy eval video <mp4> --mode native-video      # the final gate before forming/publishing a Unit
ralphy eval video <mp4> --mode keyframe           # cheap diagnostic only — does NOT approve a polished Unit

Default (no --mode) = the final gate. When you omit --mode, eval runs the native-video gate automatically if a model provider is configured (OPENROUTER_API_KEY), and upgrades to deep-style when a project STYLE_LOCK.md / BRIEF.md is discoverable. With NO credentials it falls back to structure and explicitly marks the report not ship-ready (a eval.mode-downgrade info finding records why).

Why keyframe is not enough for a polished Unit: a still never reveals a continuity jump between cuts, a caption that lags the VO, a music hit on the wrong frame, or a draggy hold. Those are exactly the failures the native-video pass catches and the keyframe pass hallucinates around. Use keyframe to triage fast and free; use native-video (or deep-style) as the gate before ralphy unit / publish. The report's gate.shipReady boolean is the authoritative signal — it is hard-false on any non-native report.

Legacy flags (still work, mapped to modes):

  • --no-vision ⇒ --mode structure.
  • --no-deep-vision ⇒ caps at --mode keyframe (never escalates to the full-mp4 pass).
  • --style-sheet / --brief ⇒ implies --mode deep-style.
Deep-style pass (project-specific, anti-generic findings)

When the user asks "validate against my niche / style / creator reference" — or the project carries a style-sheet (typically from ralphy research scrape-profile or ralphy project style-lock) — the deep-style mode scores the full mp4 against every rule in the style sheet's "Vibe & visual register" and "What this creator NEVER does" sections. Trigger it explicitly with --mode deep-style, or just pass --style-sheet (which implies it):

bash
ralphy eval video <mp4> --style-sheet <style-sheet.md> [--brief <BRIEF.md>] [--reference-urls <url> <url> ...]

Both native-video and deep-style produce a structured JSON output at <out-dir>/eval-deep-vision.json (the repair loop, #409, consumes its what_to_redo). It carries:

  • overall_verdict — holistic pass/warn/fail
  • register_match — declared vs observed cinematographic register, with severity if mismatched
  • rule_conformance[] — per-rule pass/warn/fail with verbatim style-sheet quotes and specific timestamp evidence from the rendered video
  • brief_conformance[] — same shape, scoring against BRIEF.md intent
  • uncanny_mechanism_check — whether the render delivers the style sheet's proprietary aesthetic mechanism or just mimics the surface
  • pacing_and_timing — hook / body / closer evaluation
  • ai_artifacts[] — concrete timestamp-tagged artifacts the model spotted
  • what_works — be honest, what the render did right
  • what_to_redo — prioritized 1-6 item fix list with target (start-frame / end-frame / i2v / audio / scene-prompt / model-swap / regen-entire)

Each rule violation also flows into the main findings[] array under style.register-mismatch, style.rule-violation, brief.intent-drift, style.aesthetic-mechanism-missing, or style.timing-* categories so the unified scoring + downstream fixer pipeline pick them up.

When to use deep-style over native-video:

  1. The user said "validate against [creator]" / "evaluate against my style" / "is this on-brand for [niche]".
  2. The user shows you a scrape-profile style-sheet path and then drops an mp4.
  3. The project has a discoverable STYLE_LOCK.md / style-sheet.md (auto-discovered by walking up from the mp4 path). Eval auto-upgrades the default gate to deep-style in that case.

When native-video is the right gate (no style scoring): a generic Unit-readiness check with no creator-style reference — "is this ready to ship", "QA the final cut". This is the default. It still catches temporal/audio/pacing/caption/format failures; it just doesn't score against a specific creator's rules.

When NOT to run the full-mp4 pass at all:

  • A fast triage pass mid-iteration — use --mode keyframe (cheap, free-ish) but remember it can't approve a Unit.
  • The mp4 is over 40 MB — the model rejects on body size. Re-encode at lower bitrate first.
Show full SKILL.md (656 more words)Show less
Standard flags
  • --mode <structure|keyframe|native-video|deep-style> — the explicit validation mode (see the table above). Omit for the default final gate (native-video, or deep-style when a STYLE_LOCK/brief is discoverable).
  • --no-vision — legacy alias for --mode structure (deterministic only, $0). Use for a quick structure/audio sanity pass; not a ship gate.
  • --no-deep-vision — legacy: cap the mode at keyframe (never run the full-mp4 native pass even if a --style-sheet / --brief / project BRIEF.md is present).
  • --deep-vision-model <id> — override the full-mp4 model. Default google/gemini-3.1-pro-preview. For cheaper smoke tests, swap to google/gemini-2.5-pro.
  • --project <id> — force project context when the mp4 was moved out of the project tree.
  • --no-project — explicitly evaluate as a standalone video (skips scenario.json-derived findings).
  • --out-dir <path> — override where eval.json + eval-report.md + eval-deep-vision.json land. Default: project dir, or the mp4's parent for standalone.

The command returns JSON with verdict, score, mode, shipReady, gateReason, findings (count), and the output paths. shipReady is the gate to honor before forming/publishing a Unit — never form a Unit off a shipReady: false report unless the user explicitly accepted a cheap-mode result.

How to read the report

Two files written:

  • eval.json — machine contract. The fixer agent reads this. Schema in references/report-schema.md.
  • eval-report.md — same data flattened for humans. Show the user this one.

The shape that matters: report.findings[] is the actionable list. Each finding has:

  • id (F1, F2, …) — stable ref to call out in chat
  • category — taxonomy like audio.loudness, vision.text, structure.duration-drift
  • severity — info | warn | fail
  • sceneIndex + timestampSec — where in the video, when applicable
  • message — what's wrong (specific, not generic)
  • fixHint — what kind of fix, conceptually
  • fixCommand — a copy-pasteable ralphy / ffmpeg command if one applies

scoring.verdict is pass, warn, or fail — a quality summary. report.gate is the readiness signal: gate.mode (which mode ran), gate.nativeVideo (was it a full-mp4 pass), and gate.shipReady (the single boolean a Unit-forming step gates on). A pass verdict from a keyframe gate still has shipReady: false — keyframe slicing cannot approve a polished Unit. The user always decides whether to ship, but do not present a non-native report as ship-ready approval.

Workflow

  1. Confirm the path. If the user gave a project id instead of an mp4 path, resolve to .ralphy/workspaces/<ws>/projects/<id>/render/final.mp4 (or whatever the project's render output is — check composition-props.json if the path isn't obvious).
  2. Run ralphy eval video <path>. Omit --mode for the final gate — it runs native-video automatically (or deep-style when a STYLE_LOCK/brief is discoverable). Use --mode keyframe only when the user explicitly wants a fast/cheap triage and accepts it isn't a ship gate.
  3. Show the markdown report to the user, highlighting gate.shipReady, the verdict, and the top 3-5 findings by severity. If shipReady is false because the run was a cheap mode, say so and offer to re-run native-video.
  4. Hand off if the user wants fixes. The fixer agent reads eval.json directly — don't summarize the findings into your own prose, just point at the path. Suggested handoffs by finding category:
    • vision.text, vision.composition, vision.ai-artifacts, vision.quality → /ralph-art-director (regen affected keyframes / tweak prompts).
    • structure.duration-drift, structure.hook-zone-* → /ralph-scenarist (re-time / re-script).
    • audio.*, format.* → /ralph-editor (loudnorm / re-render / re-cut).
    • captions.* → /ralph-editor (regenerate captions or tighten the script).

When findings are clearly false-positives

The eval pipeline is tuned for the common UGC cases. Some templates legitimately violate "rules" — the brainrot-ai-meme top-half is often a single static image for the whole clip, which fires structure.hook-zone-static. Don't suppress in code; instead, in the chat handoff, mark such findings as expected-for-template so the fixer agent skips them.

Handoff to a fixer agent

When the user says "fix the issues" or similar, a downstream agent will read eval.json. The minimum it needs from you:

  • Path to the report
  • Path to the original mp4
  • Project id (if any)
  • Optional: which finding ids to skip (template false-positives)

Do not try to fix from inside evaluator. The skill ends at the report.

References

  • references/report-schema.md — full JSON schema of eval.json
  • cli/lib/eval/findings.ts — rule taxonomy + thresholds (the source of truth for category and severity ladders)
  • MODELS.md — vision model used (google/gemini-2.5-flash via OpenRouter)
  • docs/green-zone.md (when added) — the safe-zone geometry the vision prompt references

© alecs5am, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in .agents/skills/evaluator of alecs5am/ralphy.

  • SKILL.md
  • references/report-schema.md

Open the folder on GitHubat commit 8d139f0

Compare with similar skills

Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluator this skillalecs5am/ralphy138—~3.6kAutomated safety check: PassApache-2.0
Purple Cow AuditAffitor/affiliate-skills700—~2.5kAutomated safety check: PassMIT
Feature Plannerserendipity1004/cc-feature-implementer176—~2.4kAutomated safety check: PassNone
Ccg Workflowfengshao1227/ccg-workflow5.9k—~2.3kAutomated safety check: PassMIT
Conducty Checkpointrobertbarclayy/conducty176—~1.5kAutomated safety check: PassMIT
Mission Plannerjdforsythe/forge151—~3.5kAutomated safety check: PassMIT

Similar skills

  • Purple Cow Audit

    Affitor/affiliate-skills

    Score product remarkability 1-10 to decide if it's worth promoting.

    700 GitHub stars~2.5k tokensUpdated 24 days ago
    Testing & QAAuto-check passed
  • Feature Planner

    serendipity1004/cc-feature-implementer

    Creates phase-based feature plans with quality gates and incremental delivery structure.

    176 GitHub stars~2.4k tokensUpdated 9 mo ago
    Testing & QAAuto-check passed
  • Ccg Workflow

    fengshao1227/ccg-workflow

    How to run a non-trivial change end to end with the CCG role tools (ccganalyze / ccgdesign / ccgbuild / ccgdebug / ccgoptimize / ccgreview / ccgtest) and the verify- quality gates.

    5.9k GitHub stars~2.3k tokensUpdated 23 days ago
    Testing & QAAuto-check passed
  • Conducty Checkpoint

    robertbarclayy/conducty

    Quality gate between parallelization groups. An agent skill from robertbarclayy/conducty.

    176 GitHub stars~1.5k tokensUpdated 3 mo ago
    Testing & QAAuto-check passed
  • Mission Planner

    jdforsythe/forge

    Decomposes goals into team blueprints using evidence-based scaling laws, topology selection, and role design.

    151 GitHub stars~3.5k tokensUpdated 3 mo ago
    Testing & QAAuto-check passed
  • Quality Gate

    0xNyk/lacp

    Production quality gate for agent sessions. An agent skill from 0xNyk/lacp.

    305 GitHub stars~382 tokensUpdated 17 days ago
    Testing & QAAuto-check passed

More from alecs5am/ralphy

All 28 skills in this repo
  • Gsap

    alecs5am/ralphy

    GSAP animation reference for HyperFrames. An agent skill from alecs5am/ralphy.

    138 GitHub stars~2.4k tokensUpdated 16 days ago
    Auto-check passed
  • Researcher

    alecs5am/ralphy

    Deep-research workflow for UGC reference material — turns one or more URLs / handles / trend queries into a single cited research report (report.md + sources.json) that a scenarist or art-director…

    138 GitHub stars~1.9k tokensUpdated 16 days ago
    Auto-check passed
  • Editor

    alecs5am/ralphy

    Composition and render craft — assembles scenario.json plus asset-manifest.json into a HyperFrames HTML composition and renders the mp4.

    138 GitHub stars~3.6k tokensUpdated 16 days ago
    Auto-check passed
  • Producer

    alecs5am/ralphy

    End-to-end orchestration — the wrapper that drives the whole production contract across roles, plus batch production.

    138 GitHub stars~4.7k tokensUpdated 16 days ago
    Auto-check passed
  • Scenarist

    alecs5am/ralphy

    Scenario and script craft — writes and reworks the scene-by-scene scenario.json: hook, beat structure, per-scene VO, on-screen text, pacing, and the language/aspect pre-flight.

    138 GitHub stars~3.2k tokensUpdated 16 days ago
    Auto-check passed
  • Troubleshooting

    alecs5am/ralphy

    Ralphy CLI operations and repair — environment setup, API keys and connectors, ralphy doctor, reading logs, diagnosing a failed generation or render, and the CLI cookbook for verbs other roles call.

    138 GitHub stars~1.4k tokensUpdated 16 days ago
    Auto-check passed

Categories

Questions about Evaluator

What does Evaluator do?

Quality evaluation of rendered UGC mp4s — scene segmentation, audio loudness / dead-air, caption density, and per-scene visual analysis. Evaluator is an agent skill from alecs5am/ralphy. Quality evaluation of rendered UGC mp4s — scene segmentation, audio loudness / dead-air, caption density, and per-scene visual analysis.

When should I use Evaluator?

Evaluator fits situations like: the user asks to evaluate / score / grade / review / QA / check quality of a rendered video; asks is this video good?; drops an mp4 path with no other instruction; mentions find issues / problems / artifacts.

How do I install Evaluator in Claude Code?

Run `npx skills add alecs5am/ralphy --skill evaluator -a claude-code`. Or copy the skill folder (.agents/skills/evaluator in alecs5am/ralphy) into .claude/skills/evaluator in your project. Claude Code loads it when a task matches its description.

How do I install Evaluator in Codex?

Run `npx skills add alecs5am/ralphy --skill evaluator -a codex`. Or copy the skill folder (.agents/skills/evaluator in alecs5am/ralphy) into .agents/skills/evaluator in your project. Codex loads it when a task matches its description.

Can I use Evaluator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alecs5am/ralphy --skill evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluator, .gemini/skills/evaluator, .github/skills/evaluator and .opencode/skills/evaluator in your project.

What does Evaluator need to run?

Going by SKILL.md and its folder, Evaluator needs credentials named OPENROUTER_API_KEY. Our summary lists: A credential in OPENROUTER_API_KEY.

Does Evaluator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evaluator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evaluator use?

Evaluator is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluator use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.5k tokens, read only when the agent opens those files.

What are the alternatives to Evaluator?

Skills that share tags, products or a category with Evaluator: Purple Cow Audit (Affitor/affiliate-skills, 700 stars), Feature Planner (serendipity1004/cc-feature-implementer, 176 stars), Ccg Workflow (fengshao1227/ccg-workflow, 5.9k stars) and Conducty Checkpoint (robertbarclayy/conducty, 176 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluator?

alecs5am (a GitHub user) maintains it in alecs5am/ralphy, which has 138 GitHub stars. The repository holds 28 skills in this directory. The repository was last updated on September 22, 2026.

Source: alecs5am/ralphy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.