Agent skill

Comment Judge

by fmflurry in fmflurry/settings-opencode

LLM-as-a-judge rubric for code comments (forbidden, false, stale, narration, noise, keep).

MITAuto-check passedEducation

Install Comment Judge

skills CLI
$ npx skills add fmflurry/settings-opencode --skill comment-judge -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install fmflurry/settings-opencode comment-judge --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/fmflurry/settings-opencode.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/comment-judge .claude/skills/comment-judge && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
comment-judge
GitHub stars
171
Token cost
~2.5k tokens
SKILL.md length
1,209 words
Files
1
Skills in repo
20
Repo updated
First seen
Licence
MIT

At a glance

LLM-as-a-judge rubric for code comments (forbidden, false, stale, narration, noise, keep).

  • Works in 3 steps: Apply rubric categories literally in… → Do not skip a category with a charitable… → Example: a comment with history noise…
  • Judging whether comments are true
  • SKILL.md covers Modes, Grouping, Asymmetric cost (REVIEW vs… and Rubric, plus 4 more sections
  • Calls bash and git

What it does

Comment Judge is an agent skill from fmflurry/settings-opencode. LLM-as-a-judge rubric for code comments (forbidden, false, stale, narration, noise, keep). Loaded in-process by code-reviewer, dotnet-cop and angular-cop to judge every added or changed comment in a review; used by the comment-judge agent for per-module comment purges. Use when judging whether comments are true, useful, or should be deleted/rewritten.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Education, covering Quizzes and assessments, Code review and Technical documentation. It works with Angular and .NET. The repository describes itself as: Custom OpenCode settings. The licence is MIT.

When your agent uses it

  • Judging whether comments are true
  • Should be deleted/rewritten

Example prompts

  • “/comment-judge”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Apply rubric categories literally in order: forbidden → false → stale → narration → noise → keep.
  2. Do not skip a category with a charitable reading unless the comment text genuinely supports multiple meanings (ambiguous wording), in…
  3. Example: a comment with history noise AND stale behavior → first check forbidden (no), then false (no), then stale (yes) → STALE wins. Do…

What it can do on your machine

Read from SKILL.md and the folder at commit 0e6c33c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bash
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Comment Judge loads about 2.5k tokens when it runs. Until then it costs about 92 tokens; SKILL.md has 1,209 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~92
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from fmflurry/settings-opencode at commit 0e6c33c, republished under its MIT licence (© fmflurry). 1,209 words, ~2,543 tokens.

Download SKILL.mdSave it as .claude/skills/comment-judge/SKILL.md (or your agent's skills folder).
name
comment-judge
description
LLM-as-a-judge rubric for code comments (forbidden, false, stale, narration, noise, keep). Loaded in-process by code-reviewer, dotnet-cop and angular-cop to judge every added or changed comment in a review; used by the comment-judge agent for per-module comment purges. Use when judging whether comments are true, useful, or should be deleted/rewritten.

comment-judge

LLM-as-a-judge for code comments. Applies rubric-based verdict, harvests blocks, emits structured findings. Loaded in-process by code-reviewer/dotnet-cop/angular-cop during review; used by the comment-judge agent for module-scoped purges.

Modes

REVIEW: Candidates are added/changed comments in the diff. Source: bash scripts/check-added-comments.sh <base> if the repo provides that script; otherwise git diff -U0 <base>...HEAD | grep -nE '^\+.*(//|/\*|#|<!--)', merged with pre-existing comments within ±10 lines of changed hunks (git diff -U10 <base>...HEAD | grep -nE '^\+.*(//|/\*|#|<!--)'). Goal: verify each comment truthfully describes code behavior.

PURGE: Scans a directory for all comments (added or pre-existing). Source: bash scripts/check-added-comments.sh --path <dir> if available; otherwise grep -rnE '(//|/\*|#|<!--)' <dir> (directory-scoped only; repo-wide is prohibitively slow). Goal: delete stale, false, forbidden, or noisy comments in bulk.

EXPLICIT: Caller supplies a list of path:start-end line ranges (used for manual calibration and edge-case handling).

Grouping

Merge consecutive path:line hits in the same file into blocks: start_line..end_line. One block = one contiguous region. Emit verdict and reasoning once per block, not per line.

Asymmetric cost (REVIEW vs PURGE)

REVIEW mode: Missing a false comment is costly (reviewers trust the code). A wrong flag is cheap (human ignores). Strategy: flag suspected false/stale/forbidden comments even with partial evidence. Tier as ❓ (low) when evidence is incomplete.

PURGE mode: Verdicts are bulk-applied to many files. Delete only at high confidence (🔴 or 🟡). Any doubt → keep + ❓ marker for follow-up.

Never auto-apply false: Comment may be correct and code may be wrong = possible bug. Always require human review.

Rubric

Read the code the comment describes: enclosing symbol, or ±15 lines, or ≤60 lines below a doc block. Never read files >400 lines whole. For stale checks, use codememory_definitions or one targeted Grep.

Ask in order; first failure wins:

a) FORBIDDEN class (per the project's code-comments rule — ~/.claude/rules/common/code-comments.md / ~/.config/opencode/instructions/code-comments.md, or the repo's own copy if it has one)

  • Narration: restates what code does; rename the code instead (tier: 🟡).
  • adr-restate: Restating, paraphrasing, or quoting design-doc/RFC/rule sections, points, phases, or citing project convention files (AGENTS.md, rules, CLAUDE.md) content. Citations must be bare pointers only: (ADR-1042) or (RFC 7636 §4.2) for external specs (section allowed for RFCs only). Tier: 🟡.
  • task-ref (extended): Requirement or work-item IDs: US\d+, EX\d+, SEC\d+, R\d{1,2}, F\d[a-z]?, #\d{5,}, WI \d+, TC #\d+, Phase \d, sub-phase, P\d(\.\d)?, Pin #\d+. Fix: strip the ID; keep the sentence if a why remains. Tier: 🟡.
  • Suppression/TODO missing reason or ticket number (tier: 🟡).
  • History/changelog/task reference (tier: 🟡).
  • Quality claim or verification assertion (tier: 🟡).
  • Section banners or #region markers (tier: 🟡).
  • Verdict: delete or rewrite (reason rule).

b) FALSE or STALE (tier: 🔴)

  • FALSE: Comment claims behavior the code does NOT perform (e.g., "thread-safe" without locks). The contradicting evidence is IN the attached code itself (enclosing symbol, ±15 lines). Must quote contradicting path:line + exact line.
  • STALE: Comment describes behavior, a call site, a symbol, or an architecture that has moved or been removed. The contradicting evidence lives OUTSIDE the attached code (another file, a removed/renamed symbol, a deleted method call). Even though the claim is technically untrue now, the comment was once correct and was never updated when the code changed. Verdict: rewrite if the attached code still needs a non-obvious why (supply corrected ≤1-line reason); else delete. Evidence: quote removal or cite codememory_definitions result.
  • Mixed (noise + stale): When history/ticket noise is attached to stale code, classify STALE (🔴 outranks 🟡); remove noise + supply corrected reason if needed.
  • Verdict: false (REVIEW mode only; PURGE requires high confidence). If low confidence → ❓.

c) NARRATION (restates next lines or method signature)

  • Example: // Creates a new user above public User Create(…) {}.
  • Verdict: delete.

d) NOISE (task/phase/agent history, verbose doc blocks, overstated impl detail on private members)

  • Example: // As discussed in ticket 1234, // Agent wrote this.
  • Verdict: delete, or rewrite to one line if non-obvious why is buried.

e) KEEP (non-obvious why, invariant with bare pointer, public/cross-module contract)

  • Example: // Sort by created_at DESC; matches UI mockup (ADR-1042).
  • Example: // Retry budget is capped at 3; matches upstream rate-limit window (ADR-2031).
  • A reference is allowed ONLY as a bare trailing pointer after a one-line why. Multi-line design-doc paraphrases must be rewritten to one line (why + bare pointer) or deleted if no local why remains.
  • Verdict: keep.

f) UNSURE (evidence partial, contradiction unclear, or context ambiguous)

  • Verdict: ❓ (low tier; flag for human).
Show full SKILL.md (521 more words)Show less

Decision Rules

DELETE vs REWRITE (when verdict is forbidden, narration, or noise): After removing the offending part (ticket ids, history, narration, stale claim), does a non-obvious why, invariant, or constraint remain that a reader of the attached code would otherwise miss?

  • YES: rewrite to ≤1 line, keeping only the non-obvious why. Example: comment says "avoid O(n²) — used in hot path", remove "Ticket 123: ", keep "avoid O(n²) — used in hot path".
  • NO (nothing of value left): delete.
  • Pure narration (comment merely restates code) is always delete.

External evidence for FALSE: When a comment cites external evidence (RFC, spec, ADR section, standard name) as proof of its claim:

  • If you can fetch and verify the cited source → apply it as evidence; if it contradicts the code, verdict is false with URL/path + section quoted.
  • If you cannot fetch/verify the cited source → ❓ (unsure), not false. Example: "// RFC 7232 says ETag must be…" — must verify RFC; do not guess.

Rubric order (tie-breaker for equally valid verdicts): When two verdicts seem equally valid:

  1. Apply rubric categories literally in order: forbidden → false → stale → narration → noise → keep.
  2. Do not skip a category with a charitable reading unless the comment text genuinely supports multiple meanings (ambiguous wording), in which case → ❓.
  3. Example: a comment with history noise AND stale behavior → first check forbidden (no), then false (no), then stale (yes) → STALE wins. Do not downgrade to noise just because history is also present.

Output contract

One finding per judged block, JSON:

json
{
  "path": "src/file.ts",
  "start_line": 42,
  "end_line": 45,
  "category": "comment",
  "severity": "high",
  "tier": "🔴",
  "rule_ref": "code-comments.md#forbidden-class",
  "source": "llm",
  "existing_code": "// This is thread-safe.\nfunction update(cache, key, value) {",
  "verdict": "false",
  "suggestion_code": "",
  "reason": "Code mutates cache[key] directly without synchronization; comment claims thread-safety but no locking present."
}

Optional verdict field (if omitted, finding is recorded but not auto-applied):

  • keep — no change needed.
  • delete — remove the comment (suggestion_code = "").
  • rewrite — replace with suggestion_code (≤1 line, English, no narration).
  • false — assertion proven false; must quote evidence.

Tier (severity mapping):

  • 🔴 false/stale: severity critical, tier 🔴.
  • 🟡 forbidden/narration/noise: severity medium, tier 🟡.
  • ❓ unsure: severity low, tier ❓.

Summary line format:

Comments: judged N · keep K · delete D · rewrite R · false F · ❓ Q

Add caveat: "Calibrated to REVIEW/PURGE mode; refactor-cleaner may auto-apply delete+rewrite; human must approve false verdicts."

Budget

Per-file limit: ≤40 blocks per batch. Total dispatch limit: ≤200 blocks. Over cap: return ## Continuation: next=<path> with the next file to process.

Trust limits (calibration baseline)

Blind calibration on real cases showed: 65–70% inter-run agreement, catches roughly a quarter of false comments, produces about one wrong delete per run, and can fabricate external evidence (e.g., misattributing which section of an RFC defines a term).

Reliable: forbidden, noise, narration verdicts (consistent across independent runs).

Unreliable: false, stale verdicts (low recall; fabrication risk with external evidence).

Application:

  1. REVIEW mode (code-reviewer / cops include judge findings): Judge false/stale verdicts are advisory flags only (tier ❓). They never block by themselves. Block only when the reviewer confirms by quoting a contradicting line from the repo (not memory or external recall). Deterministic forbidden hits still block as before.

  2. PURGE mode (refactor-cleaner applies verdicts): May auto-apply only delete/rewrite findings whose verdict is forbidden, noise, or narration AND two independent judge runs agree on (same verdict, same block). Require explicit human review ("human skimmed: yes" in brief) before applying; false, stale, and ❓ always go to human. Never apply verdicts whose evidence cites external sources without human confirmation of the source.

  3. Re-calibrate: widen these limits only after a substantial batch of human-approved cases accumulates.

© fmflurry, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/comment-judge of fmflurry/settings-opencode.

Open the folder on GitHubat commit 0e6c33c

Compare with similar skills

Comment Judge next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Comment Judge compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Comment Judge this skillfmflurry/settings-opencode171—~2.5kAutomated safety check: PassMIT
Auto Improvecrimeacs/auto-improve135—~651Automated safety check: PassMIT
Create Skill Testdotnet/skills5.6k1 repos~6.1kAutomated safety check: PassMIT
Malloy Reviewmalloydata/publisher116—~2.5kAutomated safety check: PassMIT
PR ReviewNVIDIA/Megatron-LM18k—~839Automated safety check: PassApache-2.0
Copilot PR Autopilotgithub/awesome-copilot40k—~3.4kAutomated safety check: PassMIT

Similar skills

  • Auto Improve

    crimeacs/auto-improve

    GAN-style iterative improvement loop for any text artifact. An agent skill from crimeacs/auto-improve.

    135 GitHub stars~651 tokensUpdated 2 mo ago
    EducationAuto-check passed
  • Create Skill Test

    dotnet/skills

    Official

    Scaffolds eval.yaml evaluation specs for skills, custom agents, and redistributable gh-aw workflow packages in the dotnet/skills repository.

    5.6k GitHub starsUsed in 1 repo~6.1k tokens
    EducationAuto-check passed
  • Malloy Review

    malloydata/publisher

    Malloy semantic-model code review. An agent skill from malloydata/publisher.

    116 GitHub stars~2.5k tokensUpdated today
    EducationAuto-check passed
  • PR Review

    NVIDIA/Megatron-LM

    Official

    Review rubric for the /review pull-request command. An agent skill from NVIDIA/Megatron-LM.

    18k GitHub stars~839 tokensUpdated today
    DevelopmentAuto-check passed
  • Copilot PR Autopilot

    github/awesome-copilot

    Official

    Copilot left 14 review comments on your PR — half are nits. An agent skill from github/awesome-copilot.

    40k GitHub stars~3.4k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Review a FHIRPath implementation change in Pathling against a correctness rubric covering collection semantics, empty propagation, column cardinality, type coercion, error-vs-empty behaviour, spec…

    137 GitHub stars~2k tokensUpdated 2 days ago
    EducationAuto-check passed

More from fmflurry/settings-opencode

All 20 skills in this repo
  • Show Your Work

    fmflurry/settings-opencode

    Keep a reviewable decision trail for long-running or unattended work: a TSV log with one row per decision (what, why, evidence, result).

    171 GitHub stars~1.8k tokensUpdated 3 days ago
    Auto-check passed
  • Playwright E2E Authoring

    fmflurry/settings-opencode

    Scaffold and extend Playwright E2E tests for the gc.platform suite (tests/playwright), wiring every artifact to the real frontend (localhost:4200) + real .NET backend — never mocks.

    171 GitHub stars~2.4k tokensUpdated 3 days ago
    Auto-check passed
  • Why

    fmflurry/settings-opencode

    A skill your agent uses for 'why does X work this way', 'why we picked Y', design rationale, regressions, postmortems, or data-backed thresholds.

    171 GitHub stars~2k tokensUpdated 3 days ago
    Auto-check passed
  • Angular Accessibility

    fmflurry/settings-opencode

    Audit and fix common accessibility issues in Angular templates and Angular Material components.

    171 GitHub stars~2.7k tokensUpdated 3 days ago
    Auto-check passed
  • Angular Clean Architecture

    fmflurry/settings-opencode

    Scaffolds and extends Angular standalone feature MODULES under src/app/modules/{name} using Clean Architecture layering (presentation/application/core/infrastructure), a self-registering module…

    171 GitHub stars~3.6k tokensUpdated 3 days ago
    Auto-check passed
  • Angular Cop

    fmflurry/settings-opencode

    Pre-merge code review for Angular + TypeScript pull requests.

    171 GitHub stars~1.4k tokensUpdated 3 days ago
    Auto-check passed

Works with

Questions about Comment Judge

What does Comment Judge do?

LLM-as-a-judge rubric for code comments (forbidden, false, stale, narration, noise, keep). Comment Judge is an agent skill from fmflurry/settings-opencode. LLM-as-a-judge rubric for code comments (forbidden, false, stale, narration, noise, keep).

When should I use Comment Judge?

Comment Judge fits situations like: judging whether comments are true; should be deleted/rewritten.

How do I install Comment Judge in Claude Code?

Run `npx skills add fmflurry/settings-opencode --skill comment-judge -a claude-code`. Or copy the skill folder (skills/comment-judge in fmflurry/settings-opencode) into .claude/skills/comment-judge in your project. Claude Code loads it when a task matches its description.

How do I install Comment Judge in Codex?

Run `npx skills add fmflurry/settings-opencode --skill comment-judge -a codex`. Or copy the skill folder (skills/comment-judge in fmflurry/settings-opencode) into .agents/skills/comment-judge in your project. Codex loads it when a task matches its description.

Can I use Comment Judge in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add fmflurry/settings-opencode --skill comment-judge -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/comment-judge, .gemini/skills/comment-judge, .github/skills/comment-judge and .opencode/skills/comment-judge in your project.

What does Comment Judge need to run?

Going by SKILL.md and its folder, Comment Judge needs the command-line tools its instructions call (bash and git).

Does Comment Judge access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Comment Judge safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Comment Judge use?

Comment Judge is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Comment Judge use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Comment Judge?

Skills that share tags, products or a category with Comment Judge: Auto Improve (crimeacs/auto-improve, 135 stars), Create Skill Test (dotnet/skills, 5.6k stars), Malloy Review (malloydata/publisher, 116 stars) and PR Review (NVIDIA/Megatron-LM, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Comment Judge?

fmflurry (a GitHub user) maintains it in fmflurry/settings-opencode, which has 171 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on October 7, 2026.

Source: fmflurry/settings-opencode on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.