Agent skill

Judge Sim

by MAGA2010 in MAGA2010/hackathon-run

Scores a hackathon project 0 to 5 across seven judging dimensions and produces a prioritized fix list for the last hour.

MITAuto-check passed

Install Judge Sim

skills CLI
$ npx skills add MAGA2010/hackathon-run --skill judge-sim -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install MAGA2010/hackathon-run judge-sim --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/MAGA2010/hackathon-run.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/judge-sim .claude/skills/judge-sim && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
judge-sim
GitHub stars
411
Token cost
~970 tokens
SKILL.md length
314 words
Files
2 (incl. scripts)
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Scores a hackathon project 0 to 5 across seven judging dimensions and produces a prioritized fix list for the last hour.

  • Works in 4 steps: Score seven dimensions → For each dimension, output → Compute fix priorities → …
  • Pre-submission self-review to surface the weak points judges will attack
  • SKILL.md covers Input contract, Execution, Output contract and Acceptance criteria, plus 2 more sections
  • Runs Python scripts from its folder

What it does

Judge Sim is an agent skill from MAGA2010/hackathon-run. Scores a hackathon project 0 to 5 across seven judging dimensions and produces a prioritized fix list for the last hour. Use for pre-submission self-review to surface the weak points judges will attack.

Its SKILL.md is about 970 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/score.py`).

The licence is MIT.

When your agent uses it

  • Pre-submission self-review to surface the weak points judges will attack

Example prompts

  • “Use the judge-sim skill to score a hackathon project 0 to 5 across seven judging dimensions and produces a prioritized fix list for the last hour”
  • “/judge-sim”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Score seven dimensions
  2. For each dimension, output
  3. Compute fix priorities
  4. Hard rule

What it can do on your machine

Read from SKILL.md and the folder at commit de3bcfb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Judge Sim loads about 970 tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 314 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~53
When it runs · the whole SKILL.md, loaded when a task matches
~970

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from MAGA2010/hackathon-run at commit de3bcfb, republished under its MIT licence (© MAGA2010). 314 words, ~970 tokens.

Download SKILL.mdSave it as .claude/skills/judge-sim/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
judge-sim
description
Scores a hackathon project 0 to 5 across seven judging dimensions and produces a prioritized fix list for the last hour. Use for pre-submission self-review to surface the weak points judges will attack.
when_to_use
Trigger when the user is about to submit, wants feedback before the final hour, or asks "how would judges score this". Do not invoke before the demo path…
version
1
category
judging
tags
scoring, 7-dimensions, feedback
dependencies
demo-coach
side_effects
review
triggers
how would judges score, pre-submission review, judge this, what would judges say, review my submission
capabilities
fs_read, fs_write, net, env

judge-sim

Input contract

Required:

  • repo_root: project root
  • .hackathon/state/plan.json (preferred; uses demo_path and KEEP list)
  • .hackathon/state/demo.json (preferred; uses one_liner)

Optional:

  • time_remaining_minutes: influences the fix priority list
  • HACKATHON_JUDGE_BACKEND: optional HTTP endpoint for an LLM judge; the script POSTs the state inputs and uses the returned scores. Any failure falls back to the heuristic scorer.
  • HACKATHON_JUDGE_TIMEOUT_SECONDS: request timeout for the backend (default 3).

Execution

1. Score seven dimensions

Each dimension: 0 (catastrophic) to 5 (excellent).

DimensionWhat judges look for
problem_clarityIs the pain obvious in 10 seconds?
originalityIs this novel vs. existing solutions?
completenessDoes the demo path actually work?
technical_depthIs the implementation non-trivial?
demo_qualityIs the pitch tight and rehearsed?
business_valueWould someone pay / use this?
submission_readinessREADME, run steps, secret hygiene?
2. For each dimension, output
  • score: 0..5
  • deduction_reason: one line why not higher
  • judge_questions: 2–3 likely questions
  • improvements: 1–3 concrete actions
3. Compute fix priorities

Three buckets:

  • FIX_NOW: changes that take < 30 min and improve any score
  • FIX_LAST_10MIN: cosmetic / typo-level changes only
  • DO_NOT_TOUCH: things that look fixable but risk breaking the demo
4. Hard rule

If .hackathon/state/verify.json last status is fail, any dimension scoring above 3 is invalid. Cap them at 3 and explain.

Output contract

Files written:

  • .hackathon/state/review.json (matches src/state/schemas/review.schema.json)
  • .hackathon/artifacts/judge-feedback.md (human-readable review)

Acceptance criteria

  • Provides per-dimension score (0-5).
  • Provides deduction_reason per dimension.
  • Provides improvement suggestions per dimension.
  • Distinguishes FIX_NOW / FIX_LAST_10MIN / DO_NOT_TOUCH.
  • Caps dimensions at 3 when verify status is fail.
  • Outputs a single overall score (mean).

Failure modes

ModeBehavior
No demo.jsonRefuse; ask to run demo-coach first
plan.json missingContinue, but note "demo not derived from a known plan"
All 7 scores = 5Refuse to rubber-stamp; ask the human to challenge one score
All 7 scores = 0Refuse to dunk; ask the human to confirm

Trigger phrases

  • "how will judges score us"
  • "what will the judges ask"
  • "is this competitive"
  • "final review before submit"
  • "what should I fix"

© MAGA2010, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in skills/judge-sim of MAGA2010/hackathon-run.

  • SKILL.md
  • scripts/score.py

Open the folder on GitHubat commit de3bcfb

Compare with similar skills

Judge Sim next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Judge Sim compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Judge Sim this skillMAGA2010/hackathon-run411—~970Automated safety check: PassMIT
Claw Scoreopenclaw/openclaw392k—~2.5kAutomated safety check: PassMIT
Harness Scoreruvnet/ruflo74k—~605Automated safety check: NotesMIT
Sim Helmsimstudioai/sim30k—~2.2kAutomated safety check: PassApache-2.0
Score Evalsickn33/agentic-awesome-skills47k1 repos~304Automated safety check: PassMIT
UI Scoresickn33/agentic-awesome-skills47k1 repos~1.8kAutomated safety check: PassMIT

Similar skills

  • Claw Score

    openclaw/openclaw

    Audit or refresh OpenClaw maturity scorecard docs from root taxonomy, maturity scores, and QA evidence artifacts without using maintainer discrawl data or committed inventory reports.

    392k GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Harness Score

    ruvnet/ruflo

    5-dimension harness readiness scorecard from metaharness score {path}.

    74k GitHub stars~605 tokensUpdated today
    DevelopmentAuto-check: notes
  • Sim Helm

    simstudioai/sim

    Install, upgrade, and operate the Sim Helm chart on Kubernetes.

    30k GitHub stars~2.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Score Eval

    sickn33/agentic-awesome-skills

    Imported skill score-eval from upstream source. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 1 repo~304 tokens
    EducationAuto-check passed
  • UI Score

    sickn33/agentic-awesome-skills

    Score a UI file's design quality 0-100 against StyleSeed's design language — per-category breakdown, the worst offenders, and a prioritized fix list.

    47k GitHub starsUsed in 1 repo~1.8k tokens
    Testing & QAAuto-check passed
  • Judge

    NeoLabHQ/context-engineering-kit

    Launch a meta-judge then a judge sub-agent to evaluate results produced in the current conversation

    1.8k GitHub stars~2k tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed

More from MAGA2010/hackathon-run

All 15 skills in this repo
  • Decision Log

    MAGA2010/hackathon-run

    Writes every KEEP/CUT/DEFER/PIVOT scope decision into an append-only team log with rationale, author, and timestamp.

    411 GitHub stars~847 tokensUpdated 2 days ago
    Auto-check passed
  • Demo Coach

    MAGA2010/hackathon-run

    Drafts a 30, 60, or 90-second pitch script for the demo, structured as opening, pain, product, action, result, close.

    411 GitHub stars~895 tokensUpdated 2 days ago
    Auto-check passed
  • Scope Knife

    MAGA2010/hackathon-run

    Forces a KEEP, CUT, or DEFER decision on every feature when scope is too large, no MVP consensus exists, or time is running out.

    411 GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed
  • Time Box

    MAGA2010/hackathon-run

    Allocates time across the hackathon lifecycle (idea - scope - build - verify - demo - ship) and warns before each deadline slips.

    411 GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed
  • Fast Verify

    MAGA2010/hackathon-run

    Verifies the demo path runs end-to-end by executing each step in order, recording the actual outcome, and stopping at the first failure with a diagnostic.

    411 GitHub stars~893 tokensUpdated 2 days ago
    Auto-check passed
  • Pivot

    MAGA2010/hackathon-run

    Detects a mid-build scope change and re-runs scope-knife against the new direction without losing prior progress.

    411 GitHub stars~981 tokensUpdated 2 days ago
    Auto-check passed

Questions about Judge Sim

What does Judge Sim do?

Scores a hackathon project 0 to 5 across seven judging dimensions and produces a prioritized fix list for the last hour. Judge Sim is an agent skill from MAGA2010/hackathon-run. Scores a hackathon project 0 to 5 across seven judging dimensions and produces a prioritized fix list for the last hour.

When should I use Judge Sim?

Judge Sim fits situations like: pre-submission self-review to surface the weak points judges will attack.

How do I install Judge Sim in Claude Code?

Run `npx skills add MAGA2010/hackathon-run --skill judge-sim -a claude-code`. Or copy the skill folder (skills/judge-sim in MAGA2010/hackathon-run) into .claude/skills/judge-sim in your project. Claude Code loads it when a task matches its description.

How do I install Judge Sim in Codex?

Run `npx skills add MAGA2010/hackathon-run --skill judge-sim -a codex`. Or copy the skill folder (skills/judge-sim in MAGA2010/hackathon-run) into .agents/skills/judge-sim in your project. Codex loads it when a task matches its description.

Can I use Judge Sim in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add MAGA2010/hackathon-run --skill judge-sim -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/judge-sim, .gemini/skills/judge-sim, .github/skills/judge-sim and .opencode/skills/judge-sim in your project.

What does Judge Sim need to run?

Going by SKILL.md and its folder, Judge Sim needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Judge Sim access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Judge Sim safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Judge Sim use?

Judge Sim is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Judge Sim use?

About 970 tokens (SKILL.md is roughly 3.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Judge Sim?

Skills that share tags, products or a category with Judge Sim: Claw Score (openclaw/openclaw, 392k stars), Harness Score (ruvnet/ruflo, 74k stars), Sim Helm (simstudioai/sim, 30k stars) and Score Eval (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Judge Sim?

MAGA2010 (a GitHub user) maintains it in MAGA2010/hackathon-run, which has 411 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 7, 2026.

Source: MAGA2010/hackathon-run on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.