Iterative rubric-based evaluation and self-improvement loop.

MITAuto-check: notesEducation

Install Rulph

skills CLI
$ npx skills add team-attention/hoyeon --skill rulph -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install team-attention/hoyeon rulph --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/team-attention/hoyeon.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/rulph .claude/skills/rulph && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rulph
GitHub stars
173
Token cost
~4.3k tokens
SKILL.md length
1,253 words
Files
1
Skills in repo
36
Repo updated
First seen
Licence
MIT

At a glance

Iterative rubric-based evaluation and self-improvement loop.

  • Works in 4 steps: Rubric Building → Multi-Model Evaluation → Improvement Loop → …
  • Tasks that involve Quizzes and assessments
  • SKILL.md covers Phase 1: Rubric Building, Phase 2: Multi-Model Evaluation, Phase 3: Improvement Loop and Phase 4: Completion, plus 1 more section
  • Calls gemini and codex

What it does

Rulph is an agent skill from team-attention/hoyeon. Iterative rubric-based evaluation and self-improvement loop. Builds a scoring rubric interactively, evaluates an artifact with multiple models in parallel (Codex, Gemini, Claude), then autonomously improves the artifact one criterion at a time until a score threshold is met or circuit breaker fires. "/rulph", "rubric evaluate", "rubric score", "multi-model evaluate", "score and improve", "evaluate and iterate", "grade this", "루브릭 루프", "채점 루프", "자율 개선", "개선 루프", "루브릭 평가"

Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Education, covering Quizzes and assessments. The repository describes itself as: Requirements-first Harness — derive, verify, execute. The licence is MIT.

When your agent uses it

  • Tasks that involve Quizzes and assessments

Example prompts

  • “/rulph”
  • “rubric evaluate”
  • “rubric score”
  • “/rulph”

Requirements

  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Bash, Write, AskUserQuestion, Agent

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Rubric Building
  2. Multi-Model Evaluation
  3. Improvement Loop
  4. Completion

What it can do on your machine

Read from SKILL.md and the folder at commit 7cff032. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Bash
    • Write
    • AskUserQuestion
    • Agent

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gemini
    • codex

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Rulph loads about 4.3k tokens when it runs. Until then it costs about 120 tokens; SKILL.md has 1,253 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~120
When it runs · the whole SKILL.md, loaded when a task matches
~4.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Grep, Glob, Bash, Write, AskUserQuestion, Agent

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from team-attention/hoyeon at commit 7cff032, republished under its MIT licence (© team-attention). 1,253 words, ~4,283 tokens.

Download SKILL.mdSave it as .claude/skills/rulph/SKILL.md (or your agent's skills folder).
name
rulph
description
Iterative rubric-based evaluation and self-improvement loop. Builds a scoring rubric interactively, evaluates an artifact with multiple models in parallel (Codex, Gemini, Claude), then autonomously improves the artifact one criterion at a time until a score threshold is met or circuit breaker fires. "/rulph", "rubric evaluate", "rubric score", "multi-model evaluate", "score and improve", "evaluate and iterate", "grade this", "루브릭 루프", "채점 루프", "자율 개선", "개선 루프", "루브릭 평가"
allowed-tools
Read, Grep, Glob, Bash, Write, AskUserQuestion, Agent
validate_prompt
Must contain all 4 Phases (Rubric Building, Evaluation, Improvement Loop, Completion). Must include 3-step rubric building interaction with per-criterion…

rulph

Iterative self-improvement skill driven by a user-defined rubric. Builds a scoring rubric interactively, evaluates an artifact with multiple models in parallel, then loops autonomously — improving one criterion at a time — until the score meets the threshold or the circuit breaker fires. No user interaction after Phase 1.


Phase 1: Rubric Building

Build an evaluation rubric through a 3-step interactive process before any scoring begins.

User interaction: Use the AskUserQuestion tool for all user-facing questions in this skill. This ensures the UI renders properly and waits for real user input.

Step 1 — Criteria Collection

Use AskUserQuestion to ask what they are evaluating and what criteria matter. Suggest common categories (code quality, writing quality, system design) but let them describe freely.

After the user responds, parse:

  • Target: the artifact or output being evaluated (file path, text block, or description)
  • Criteria: the named dimensions to score (extract from free text or selection)

Require a minimum of 2 criteria. If fewer than 2 are given, prompt again:

"Please provide at least 2 criteria so we can triangulate quality. What else matters?"

Rubric Validation — before proceeding, check each criterion:

  • Warn on criteria that are not LLM-evaluable (e.g., "is it beautiful", "feels right", "gut check")

    "Warning: '[criterion]' is hard to score objectively. Consider rewording to something measurable, e.g., 'visual hierarchy is clear and consistent'."

  • Block on purely subjective criteria only if the user cannot clarify after one prompt.
Step 2 — Rubric Draft Presentation

Generate a rubric draft based on the collected criteria. Assign equal weights by default.

Checklist Decomposition (default): Break each criterion into 5–10 yes/no sub-items. Score is computed as (checked / total) × 100. This eliminates evaluator interpretation variance.

Present the draft as a table with sub-items:

## Draft Rubric

| # | Criterion       | Weight | Sub-items (yes/no each)                                    |
|---|-----------------|--------|------------------------------------------------------------|
| 1 | [criterion]     | 25%    | □ [sub-item-1] · □ [sub-item-2] · □ [sub-item-3]          |
|   |                 |        | □ [sub-item-4] · □ [sub-item-5]                            |
|   |                 |        | score = (checked / 5) × 100                                |
| 2 | [criterion]     | 25%    | □ [sub-item-1] · □ [sub-item-2] · □ [sub-item-3]          |
|   |                 |        | □ [sub-item-4] · □ [sub-item-5] · □ [sub-item-6]          |
|   |                 |        | score = (checked / 6) × 100                                |

Sub-item design rules:

  • Each sub-item must be binary observable — answerable with yes or no by reading the artifact
  • Avoid subjective sub-items ("code is clean") — reword to observable ("all functions have explicit return types")
  • Order sub-items from basic (easy to pass) to advanced (hard to pass)
  • Unchecked sub-items automatically become improvement targets in Phase 3

Qualitative fallback: If a criterion genuinely cannot be decomposed into sub-items (e.g., "writing tone"), use level-based anchors instead:

| # | Criterion       | Weight | Scoring Guidance (0–100)                                   |
|---|-----------------|--------|------------------------------------------------------------|
| N | [criterion]     | 25%    | 0=absent · 25=minimal · 50=partial · 75=good · 100=full   |

Level-based anchors must have 5 levels (0/25/50/75/100) with one concrete observable indicator per level. 3-level anchors (0/50/80) are too coarse.

Each criterion gets:

  • A 0–100 scoring range
  • Either checklist sub-items (preferred) or 5-level anchors with observable indicators

Then use AskUserQuestion to confirm or modify (accept, adjust weights, edit criteria, or start over). Loop until the user accepts.

Weight validation: After any adjustment, verify sum(weights) == 100% (±1% tolerance for rounding). If invalid, prompt:

"Weights must sum to 100%. Current sum: [X]%. Please redistribute." Re-present the rubric table until weights are valid.

Step 3 — Threshold & Floor Setting

Use AskUserQuestion to ask two things:

  1. Overall threshold (0–100): what overall score the artifact should reach before stopping. Suggest 70/80/90 as options. Default is 70 if the user doesn't specify.

  2. Per-criterion floor (0–100): the minimum score that EACH individual criterion must meet, regardless of overall score. Suggest 50/60 as options. Default is 60 if the user doesn't specify. Set to 0 to disable.

Why floor matters: Without a floor, strong criteria can mask weak ones (e.g., overall 80 passes threshold 70, but one criterion scores 50). The floor ensures every dimension meets a minimum bar.

Rubric Summary (Evaluation Contract)

Display the final rubric before Phase 2 begins:

## Evaluation Contract

**Target**: [artifact description or path]
**Threshold**: [threshold]/100
**Per-criterion floor**: [floor]/100
**Max rounds**: 5
**Scoring method**: Checklist Decomposition

| # | Criterion   | Weight | Sub-items                              | Formula              |
|---|-------------|--------|----------------------------------------|----------------------|
| 1 | [criterion] | [W]%   | □ A · □ B · □ C · □ D · □ E           | (checked/5) × 100   |
| 2 | [criterion] | [W]%   | □ A · □ B · □ C · □ D                 | (checked/4) × 100   |
...

Pass condition: overall >= [threshold] AND every criterion >= [floor]
Rubric locked. Starting evaluation.

State init — write the loop state so the Stop hook can track progress. The state file is session-scoped to prevent cross-session interference:

Bash: SESSION_ID="[session ID from UserPromptSubmit hook]" && hoyeon-cli session set --sid $SESSION_ID --json '{"rulph": {"round": 0, "max_rounds": 5, "score": 0, "threshold": [threshold], "status": "active", "iteration": 0, "max_iterations": 15}}'

Replace [threshold] with the actual threshold value. The state is stored under the .rulph key in the session-scoped state.json. This file is read by the Stop hook to decide whether the loop should continue. The iteration/max_iterations fields are the Stop hook's safety counter — always preserve them in subsequent state updates.


Phase 2: Multi-Model Evaluation

Score the artifact independently using up to 3 models in parallel.

CLI Availability Check

Before scoring, check which CLIs are available:

Bash: command -v codex && command -v gemini

Model states: AVAILABLE (CLI found) / SKIPPED (not found) / DEGRADED (found but call failed).

Note: The 3rd evaluator (Claude) runs as a subagent — no CLI check needed.

Parallel Scoring

Score isolation rule: Pass only the current artifact content to each model. Do NOT include previous round scores, improvement history, or prior evaluation feedback.

Each evaluator receives the same prompt template with the rubric, artifact content, and required JSON output format:

## Rubric Evaluation Task

You are a strict evaluator. Score the artifact below using the provided rubric.
For each criterion, check every sub-item (yes/no) and compute: score = (checked / total) × 100.
Return ONLY a JSON object — no prose before or after.

## Rubric
[criterion list with weights and sub-items checklist]

## Artifact
[Full artifact content — read the file]

## Required Output Format
{
  "scores": { "[criterion]": <0-100>, ... },
  "checklist": { "[criterion]": { "[sub-item-1]": true/false, "[sub-item-2]": true/false, ... }, ... },
  "suggestions": { "[criterion]": "<one concrete action targeting an unchecked sub-item>", ... }
}

Launch all 3 evaluators in a single message using run_in_background: true:

# All 3 in ONE message — true parallel execution
Agent(subagent_type="general-purpose", run_in_background=true,
      description="Codex evaluator",
      prompt="Run: codex exec <<'PROMPT'\n[evaluation prompt with rubric + artifact]\nPROMPT")

Agent(subagent_type="general-purpose", run_in_background=true,
      description="Gemini evaluator",
      prompt="Run: gemini -p \"$(cat <<'PROMPT'\n[evaluation prompt with rubric + artifact]\nPROMPT)\"\n")

Agent(subagent_type="general-purpose", run_in_background=true,
      description="Claude evaluator",
      prompt="[evaluation prompt with rubric + artifact — subagent evaluates directly]")

After launching, wait for all 3 to complete (check TaskOutput for each background agent). Then proceed to Score Aggregation.

Show full SKILL.md (505 more words)Show less
Score Aggregation

After all models complete (or fail):

Minimum model guarantee: If all 3 CLIs fail, fall back to main agent self-evaluation as a last resort. Score aggregation is guaranteed to have at least one model result.

Low confidence flag: If only 1 model is AVAILABLE, flag the round as LOW CONFIDENCE in the inline display. Single-model scores lack cross-validation.

  1. For each criterion, compute the average score across AVAILABLE models only.
  2. Compute the overall weighted average:
    overall = sum(criterion_avg[i] * weight[i]) for all i
  3. Record per-model status: AVAILABLE / SKIPPED / DEGRADED.

Inline display:

📊 Score: XX/100 (Codex: XX | Gemini: XX | Claude: XX) — Threshold: [threshold] · Floor: [floor]
   [criterion_1]: XX  (Codex: XX, Gemini: XX, Claude: XX)
   [criterion_2]: XX  (Codex: XX, Gemini: XX, Claude: XX)  ⚠️ BELOW FLOOR
   ...
   Model status: Codex=AVAILABLE · Gemini=SKIPPED · Claude=AVAILABLE
   Floor violations: [list of criteria below floor, or "None"]

Convergence / Divergence Analysis:

If any two models differ by more than 20 points on the same criterion:

"Warning: Model disagreement on '[criterion]' (gap: XX pts). Scores may reflect differing interpretations of the rubric. Consider clarifying the scoring anchor for this dimension."

Improvement Suggestion Synthesis:

Collect suggestions from all AVAILABLE models. Prioritize the criterion with the lowest average score. Present the top suggestion per criterion, labeled by source model.

State update — after every scoring round, update the session-scoped state file (preserve iteration/max_iterations for the Stop hook's safety counter):

Bash: SESSION_ID="[session ID from UserPromptSubmit hook]" && hoyeon-cli session set --sid $SESSION_ID --json '{"rulph": {"round": [round], "score": [overall], "threshold": [threshold], "status": "active", "iteration": 0}}'

Replace [round], [overall], etc. with actual values. Note: iteration resets to 0 here — the Stop hook increments it each time it fires within a round, providing a per-round safety net.


Phase 3: Improvement Loop

Iteratively improve the artifact one criterion at a time until the threshold is met or the circuit breaker fires. No user interaction in this phase — the loop runs autonomously.

Initialize: round = 1, max_rounds = 5, score_history = []

Loop Structure

The initial Phase 2 scoring produces baseline scores. Phase 3 then runs this loop:

LOOP:
  1. Pass Check → if overall >= threshold AND all criteria >= floor → Phase 4 (PASSED)
  2. Circuit Breaker → if round > max_rounds → Phase 4 (CIRCUIT BREAKER)
  3. Improvement Dispatch (improve lowest criterion — floor violations first)
  4. Re-score (return to Phase 2)
  5. Append to score_history, round += 1
  6. Repeat from 1
Pass Check (Threshold + Floor)
below_floor = [c for c in criteria if c.score < floor]

if overall >= threshold AND len(below_floor) == 0:
  → Proceed to Phase 4 immediately (PASSED)

if len(below_floor) > 0:
  → Log: "Floor violation: [criterion] at [score] < floor [floor]. Auto-targeting for improvement."
  → Improvement target = lowest below-floor criterion (not lowest overall)

if overall < threshold AND len(below_floor) == 0:
  → Improvement target = lowest criterion (original behavior)

Floor priority: Floor violations take precedence over overall threshold. Even if overall >= threshold, a below-floor criterion blocks PASSED and triggers improvement.

Circuit Breaker Check
if round > max_rounds:
  → Proceed to Phase 4 immediately (result: CIRCUIT BREAKER)
Improvement Dispatch

Select the single lowest-scoring criterion (prevents scope creep). If multiple criteria tie for the lowest score, pick the one with the higher weight (greater impact on overall score).

Dispatch a worker agent:

Agent(subagent_type="worker",
     prompt="## Improvement Task — Round [round]

## Artifact
Location: [artifact file path or content block]

## Target Criterion
[criterion name]: current score [score]/100
Weight: [W]%

## Unchecked Sub-items (fix these)
[List each unchecked sub-item from the checklist — these are the specific gaps to close]

## Improvement Instructions
[Synthesized suggestions from all AVAILABLE models for this criterion]

## Constraint
Improve ONLY this criterion. Focus on the unchecked sub-items listed above.
Do not restructure or rewrite unrelated sections.
Return the improved artifact to the same location.")

After the worker completes:

  1. Return to Phase 2 for re-scoring (which updates state file automatically)
  2. Append to score history: score_history.append({ round, overall, per_criterion_scores, model_states })
  3. Increment round counter: round += 1
  4. Return to top of loop (Threshold Check)

Phase 4: Completion

State update — mark as completed so the Stop hook allows exit:

Bash: SESSION_ID="[session ID from UserPromptSubmit hook]" && hoyeon-cli session set --sid $SESSION_ID --json '{"rulph": {"status": "completed"}}'
Final Report

Display the complete evaluation summary:

## Rulph Final Report

**Artifact**: [artifact description or path]
**Rubric**: [N] criteria · threshold [threshold]/100 · floor [floor]/100
**Result**: [PASSED / CIRCUIT BREAKER]

### Score History

| Round | Overall | [C1] | [C2] | ... | Models Used         |
|-------|---------|------|------|-----|---------------------|
| 1     | XX      | XX   | XX   | ... | Codex, Claude       |
| 2     | XX      | XX   | XX   | ... | Codex, Claude       |
| ...   |         |      |      |     |                     |
| N     | XX      | XX   | XX   | ... | Codex, Claude       |

### Final Scores (Round [N])

| Criterion   | Weight | Score | Top Suggestion                        |
|-------------|--------|-------|---------------------------------------|
| [criterion] | [W]%   | XX    | [best suggestion from last round]     |
| ...         |        |       |                                       |

**Overall: [final_score]/100**
[PASSED threshold of [threshold] ✓ / Did not reach threshold — stopped at round N]
Auto-Save Report

Always save the rubric and scores automatically. Include the full report in the saved file.

SESSION_ID="[session ID from UserPromptSubmit hook]"
REPORT_DIR="$HOME/.hoyeon/$SESSION_ID/tmp/rulph"
Bash: mkdir -p "$REPORT_DIR"

Write to $REPORT_DIR/$(date +%Y-%m-%d-%H%M%S)-report.md:
  [Full rubric definition]
  [Score history table]
  [Final scores table]
  [Model availability log per round]

Close with:

"Finished! Final score: [final_score]/100 after [N] round(s). Report saved to session tmp."


Prompt Hardening

  • Never interpolate user input directly into CLI parameters. Always wrap artifact content and rubric text in a heredoc (<<'PROMPT' ... PROMPT). For Gemini, use gemini -p "$(cat <<'PROMPT' ... PROMPT)" to prevent shell injection. The Claude evaluator runs as a subagent so no CLI escaping is needed.
  • Isolate artifact content from evaluator prompt. Rubric definition and artifact content must appear in separate labeled blocks.
  • Score isolation. When re-evaluating after improvement, pass only the current artifact state. Strip prior scores, history, and suggestions from the evaluator prompt.

© team-attention, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/rulph of team-attention/hoyeon.

Open the folder on GitHubat commit 7cff032

Compare with similar skills

Rulph next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Rulph compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Rulph this skillteam-attention/hoyeon173—~4.3kAutomated safety check: NotesMIT
DeepTutor CLIHKUDS/DeepTutor41k—~2.8kAutomated safety check: PassApache-2.0
AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch66k—~2kAutomated safety check: PassMIT
Codebase to Coursezarazhangrui/codebase-to-course5.7k—~4.4kAutomated safety check: PassNone
AI Engineering Phase Quizrohitg00/ai-engineering-from-scratch66k—~2.1kAutomated safety check: PassMIT
Scholar EvaluationK-Dense-AI/claude-scientific-writer2.4k2 repos~2.9kAutomated safety check: NotesMIT

Similar skills

  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.8k tokensUpdated yesterday
    EducationAuto-check passed
  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    66k GitHub stars~2k tokensUpdated today
    EducationAuto-check passed
  • Codebase to Course

    zarazhangrui/codebase-to-course

    Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.

    5.7k GitHub stars~4.4k tokensUpdated 6 mo ago
    EducationAuto-check passed
  • AI Engineering Phase Quiz

    rohitg00/ai-engineering-from-scratch

    Quizzes you on a completed phase of the AI Engineering from Scratch course, taking a phase number or name and mapping it to that phase's directory.

    66k GitHub stars~2.1k tokensUpdated today
    EducationAuto-check passed
  • Scholar Evaluation

    K-Dense-AI/claude-scientific-writer

    Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls.

    2.4k GitHub starsUsed in 2 repos~2.9k tokens
    EducationAuto-check: notes
  • Evaluation

    guanyang/open-agent-hub

    This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…

    977 GitHub starsUsed in 2 repos~4.2k tokens
    EducationAuto-check passed

More from team-attention/hoyeon

All 36 skills in this repo
  • Skill Session Analyzer

    team-attention/hoyeon

    This skill should be used when the user asks to "analyze session", "evaluate skill execution", "check session logs", provides a session ID with a skill path, or wants to verify that a skill executed…

    173 GitHub stars~1.9k tokensUpdated 4 mo ago
    Auto-check: notes
  • Browser Work

    team-attention/hoyeon

    Recon-first browser automation. An agent skill from team-attention/hoyeon.

    173 GitHub stars~1.9k tokensUpdated 4 mo ago
    Auto-check passed
  • Check

    team-attention/hoyeon

    This skill should be used when the user wants to verify their changes before pushing, or update the project's rule checklists.

    173 GitHub stars~1.8k tokensUpdated 4 mo ago
    Auto-check: notes
  • Compound

    team-attention/hoyeon

    This skill should be used when the user says "/compound", "compound this", "document learnings", "save what we learned", or after completing a PR.

    173 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check: notes
  • QA

    team-attention/hoyeon

    Systematically QA test any application — web apps, native macOS apps, Electron apps, CLI tools, interactive REPLs, or anything on screen.

    173 GitHub stars~2.6k tokensUpdated 4 mo ago
    Auto-check: notes
  • Tech Decision

    team-attention/hoyeon

    This skill should be used when the user asks about "technical decision", "what to use", "A vs B", "comparison analysis", "library selection", "architecture decision", "which one to use"…

    173 GitHub stars~1.4k tokensUpdated 4 mo ago
    Auto-check passed

Categories

Questions about Rulph

What does Rulph do?

Iterative rubric-based evaluation and self-improvement loop. Rulph is an agent skill from team-attention/hoyeon. Iterative rubric-based evaluation and self-improvement loop.

When should I use Rulph?

Rulph fits situations like: tasks that involve Quizzes and assessments.

How do I install Rulph in Claude Code?

Run `npx skills add team-attention/hoyeon --skill rulph -a claude-code`. Or copy the skill folder (skills/rulph in team-attention/hoyeon) into .claude/skills/rulph in your project. Claude Code loads it when a task matches its description.

How do I install Rulph in Codex?

Run `npx skills add team-attention/hoyeon --skill rulph -a codex`. Or copy the skill folder (skills/rulph in team-attention/hoyeon) into .agents/skills/rulph in your project. Codex loads it when a task matches its description.

Can I use Rulph in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add team-attention/hoyeon --skill rulph -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rulph, .gemini/skills/rulph, .github/skills/rulph and .opencode/skills/rulph in your project.

What does Rulph need to run?

Going by SKILL.md and its folder, Rulph needs the command-line tools its instructions call (gemini and codex). Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash, Write, AskUserQuestion, Agent.

Does Rulph access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Rulph safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Rulph use?

Rulph is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Rulph use?

About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Rulph?

Skills that share tags, products or a category with Rulph: DeepTutor CLI (HKUDS/DeepTutor, 41k stars), AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 66k stars), Codebase to Course (zarazhangrui/codebase-to-course, 5.7k stars) and AI Engineering Phase Quiz (rohitg00/ai-engineering-from-scratch, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Rulph?

team-attention (a GitHub organization) maintains it in team-attention/hoyeon, which has 173 GitHub stars. The repository holds 36 skills in this directory. The repository was last updated on May 21, 2026.

Source: team-attention/hoyeon on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.