Agent skill

Slop Eval

by fabricioctelles in fabricioctelles/skills

Objectively evaluate a UI/web design against the pols.dev anti-slop design law: detect catalogued slop tells with cited evidence, score 8 weighted axes (color, type, components, layout, motion…

Apache-2.0Auto-check passedFrontend & Design

Install Slop Eval

skills CLI
$ npx skills add fabricioctelles/skills --skill slop-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install fabricioctelles/skills slop-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/fabricioctelles/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/slop-eval .claude/skills/slop-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
slop-eval
GitHub stars
106
Token cost
~5.4k tokens
SKILL.md length
2,551 words
Files
8 (incl. scripts, references)
Skills in repo
15
Repo updated
First seen
Licence
Apache-2.0

At a glance

Objectively evaluate a UI/web design against the pols.dev anti-slop design law: detect catalogued slop tells with cited evidence, score 8 weighted axes (color, type, components, layout, motion…

  • Works in 6 steps: Load at desktop (1440×900) and mobile… → Full-page screenshot of both viewports… → Scroll pass top to bottom, then a second… → …
  • The user asks to evaluate design slop
  • SKILL.md covers Source, Parameters, Evaluation modes and Evidence channels, plus 10 more sections
  • Runs Python scripts from its folder; calls npx

What it does

Slop Eval is an agent skill from fabricioctelles/skills. Objectively evaluate a UI/web design against the pols.dev anti-slop design law: detect catalogued slop tells with cited evidence, score 8 weighted axes (color, type, components, layout, motion, execution, signature, cohesion), and emit a Slop Report with a 0–100 Slop Index and grade. Use when the user asks to "evaluate design slop", "slop report", "is this design AI slop", "audit this landing page design", "de-slop review", or wants an objective score of how generic/machine-made a design looks. To fix text (not…

Its SKILL.md is about 5.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `references/contexts.md`, `references/jev-integration.md` and `references/output-template.md`).

It sits in Frontend & Design, covering Accessibility, Humanizing AI text and Landing pages. The repository describes itself as: A collection of skills for AI agents (Kiro, Cursor, Windsurf, Claude Code, and others). Each skill is a reusable module that teaches the agent to perform complex tasks with… The licence is Apache-2.0.

When your agent uses it

  • The user asks to evaluate design slop
  • Is this design AI slop
  • Audit this landing page design
  • Wants an objective score of how generic/machine-made a design looks

Example prompts

  • “evaluate design slop”
  • “slop report”
  • “is this design AI slop”
  • “/slop-eval”

Requirements

  • Python 3
  • Node.js

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Load at desktop (1440×900) and mobile (390×844); wait for network idle.
  2. Full-page screenshot of both viewports **immediately after load,
  3. Scroll pass top to bottom, then a second full-page capture; diff the
  4. Interaction pass: hover the primary CTA, one card, one nav link
  5. Zoom crops at 2x of: anything near a clipped edge (X2), circled/tiled
  6. Pull the rendered sources for the code-channel greps: font names from

What it can do on your machine

Read from SKILL.md and the folder at commit f1de632. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • pols.dev
    • docs.typesafe.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Slop Eval loads about 5.4k tokens when it runs, and up to ~21k if it reads all its reference files. Until then it costs about 144 tokens; SKILL.md has 2,551 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~144
When it runs · the whole SKILL.md, loaded when a task matches
~5.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~21k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from fabricioctelles/skills at commit f1de632, republished under its Apache-2.0 licence (© fabricioctelles). 2,551 words, ~5,438 tokens.

Download SKILL.mdSave it as .claude/skills/slop-eval/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
slop-eval
description
Objectively evaluate a UI/web design against the pols.dev anti-slop design law: detect catalogued slop tells with cited evidence, score 8 weighted axes (color, type, components, layout, motion, execution, signature, cohesion), and emit a Slop Report with a 0–100 Slop Index and grade. Use when the user asks to "evaluate design slop", "slop report", "is this design AI slop", "audit this landing page design", "de-slop review", or wants an objective score of how generic/machine-made a design looks. To fix text (not design), use human-ai or humanizar skills instead.
metadata.author
https://ft.ia.br
metadata.version
1.1.0
metadata.date
2026-08-24
metadata.repository
https://github.com/fabricioctelles/skills
metadata.license
Apache-2.0
metadata.category
code-quality-and-review

Slop Eval

Evaluate a design the way skill-evaluation evaluates a skill: every finding cites concrete evidence, every axis gets a 0–100 score, arithmetic runs through a script, and the output is a structured report — never a vibe check.

The tell catalog lives in references/tells.md; read it before sweeping. The positive rubric (signature formula, cohesion checks, slop→premium pairs, and the "Adding Soul" guide) lives in references/premium-markers.md; read it before scoring Axes 7–8 and when writing fix prescriptions. The design context guide lives in references/contexts.md; read it to adjust priorities and tolerances based on the type of design being evaluated.

Source

Parameters

ParameterDescriptionDefault
targetWhat to evaluate: live URL, screenshot(s), code path, or Figma exportAsk user
briefBrand brief or explicit user directions the design followedNone
contextDesign type: landing, saas, editorial, ecommerce, or autoauto
outputPath to write the report./SLOP-REPORT.md
comparePath to a previous report for tracking mode (temporal evolution)None

Write the report in the language the user is speaking; keep tell IDs and names in English so they stay greppable against the catalog.

Evaluation modes

Standard mode (default)

Single evaluation of a design. Produces a Slop Report with scores, tells, section ledger, and prioritized fixes.

Comparison mode

Side-by-side evaluation of two different designs (e.g., competitor analysis, A/B variants). Add --compare pointing to another target or existing report.

Tracking mode

Evaluate the same design over time to measure improvement. Use when:

  • Running weekly/sprint design reviews
  • Measuring progress after a redesign
  • Validating that fixes actually moved the score

Usage:

bash
# First evaluation — establishes baseline
slop-eval --target https://site.com --output ./reports/baseline.md

# Later evaluation — tracks evolution
slop-eval --target https://site.com --output ./reports/week-2.md \
  --compare ./reports/baseline.md

Tracking mode adds to the report:

  • Score progression table with trends (✅ improved / ⚠️ regressed)
  • Tells resolved (what got fixed)
  • New tells (what got introduced)
  • Regressions (axes/sections that got worse)
  • Section ledger evolution
  • Velocity metrics (tells resolved per week, score improvement rate)
  • Recommendations for next iteration

See references/output-template.md for the full tracking output format.

Evidence channels

What you can verify depends on what you were given. Never score a check you could not observe — mark it Unverifiable and exclude it (like N/A in skill-evaluation).

ChannelCan verifyCannot verify
Code (CSS/JSX/HTML)Fonts, hex values, gradients, shadows, radii, opacity:0 gating, icon imports, layout skeletonsOptical centering, rendered contrast, seams, whether controls respond
Screenshot(s)Everything visual: palette, type, layout, alignment, centering, clipping, contrast, seamsHover/scroll motion, dead controls, invisible-content trap, responsive behavior
Live URL (browse + screenshot)All of the above plus interactions, motion, fold ownershipOnly what you didn't exercise

With code, grep before you stare: fonts.googleapis|next/font, lucide-react, linear-gradient, box-shadow, border-radius: *9999, backdrop-filter, opacity: *0, initial={{ *opacity: *0, overflow: *hidden, clip-path, position: *fixed. Each hit is a lead, not a verdict — confirm against the catalog entry before recording it.

Evidence acquisition SOP

Route by what the target is; always end with an evidence inventory (what was captured, what is Unverifiable) — it feeds the report header.

Live URL — the richest channel; prefer it whenever reachable. Use whatever browser automation this session has (a browser MCP such as Playwright or Chrome DevTools, or npx playwright screenshot as the no-MCP fallback) and capture, saving every artifact to the scratchpad so findings can cite file + region:

  1. Load at desktop (1440×900) and mobile (390×844); wait for network idle.
  2. Full-page screenshot of both viewports immediately after load, before any scrolling — sections sitting at opacity:0 waiting for a scroll reveal show up blank here (M1 evidence).
  3. Scroll pass top to bottom, then a second full-page capture; diff the two mentally for reveal-gated content, seams (C11, X13), and fold ownership (L16).
  4. Interaction pass: hover the primary CTA, one card, one nav link (M2–M4); click every tab, accordion, toggle, and button (M8); Tab through the page and confirm a visible focus ring (X14).
  5. Zoom crops at 2x of: anything near a clipped edge (X2), circled/tiled numbers and icons (X1), pricing columns side by side (X3), button labels (X5).
  6. Pull the rendered sources for the code-channel greps: font names from the network panel or <link>/@font-face, computed hex values from the stylesheets.

No browser automation available → fetch the HTML/CSS (curl) and run the code channel on it, ask the user for full-page desktop + mobile prints, and mark every visual-only and interaction check Unverifiable until the prints arrive. Never score a visual check from raw HTML.

Screenshots — Read each image. If only partial crops were provided, ask for full-page desktop + mobile before sweeping (a hero-only print cannot support L11, L15, or the cohesion axis). All interaction checks (M1, M8, X14, hover tells) are Unverifiable.

Code path — run the greps, read every file they hit, plus the layout/ page components and global styles. If the project runs locally, start its dev server and continue under the Live URL SOP — code plus a live render is the only combination that can verify everything.

Figma export — treat as Screenshots for visual tells; additionally fonts, hex values, and spacing are exact from the file. Motion and interaction axes are Unverifiable (score NA for Axis 5 unless prototypes were shared).

Axes and weights

#AxisWeightScored from
1Color & Light2xTells C1–C15
2Typography & Copy2xTells T1–T10, W1–W3
3Components & Ornament1xTells K1–K27
4Layout & Composition2xTells L1–L21
5Motion & Interaction1xTells M1–M8
6Execution & Craft2xTells X1–X14
7Signature & Uniqueness3x7-element formula (positive rubric)
8Cohesion2x4 checks (positive rubric)

Axis 7 carries the heaviest weight on purpose: the law's deepest rule is that dodging the tell list is still slop — a page with zero tells and no signature is unfinished work wearing restraint as an alibi.

Scoring

Axes 1–6 (tell-counted). Count confirmed tells on the axis by severity, then: score = max(0, 100 − 30·critical − 15·major − 5·minor). Run scripts/score.py axis CRIT MAJOR MINOR — don't do it by hand. One tell, one count: a pattern repeated across sections is still one tell (note the repetition in the evidence; repetition may upgrade minor → major where the catalog says so).

Axis 7 (Signature). Score each of the 7 formula elements 0 (absent), 50 (attempted, weak), or 100 (strong) per the rubric in premium-markers.md; the axis is their mean.

Axis 8 (Cohesion). Same 0/50/100 on the 4 cohesion checks; mean.

Compounding rule. Three or more major layout tells on one page cap Axis 4 at 40 — a page assembled from known skeletons is slop no matter how clean each block is.

Gates (pass as --cap to the overall run):

  • Signature gate: Axis 7 < 40 caps the overall at 59 (grade C max). No amount of clean spacing rescues a page with no signature.
  • Absolute-rule gate: any confirmed critical tell caps the overall at 69 (no grade A with broken execution).

Overall & Slop Index.

overall    = sum(axis_score × weight) / sum(weight)   # capped by gates
Slop Index = 100 − overall

Run scripts/score.py overall 1:80:2 2:65:2 ... [--cap 59] [--cap 69]. Unverifiable axes score NA and drop out of both sums. --fail-below N exits non-zero for CI gating, e.g. gating a PR on its preview deploy:

yaml
# .github/workflows/slop-gate.yml (step excerpt)
- name: Slop gate
  run: |
    # run slop-eval against $PREVIEW_URL, export each axis score, then:
    python3 skills/slop-eval/scripts/score.py overall \
      1:$A1:2 2:$A2:2 3:$A3:1 4:$A4:2 5:$A5:1 6:$A6:2 7:$A7:3 8:$A8:2 \
      --fail-below 40

Grade scale

GradeOverallSlop IndexVerdict
A80–1000–20Premium — deliberate, signed, executed
B60–7921–40Considered — mostly deliberate, some defaults
C40–5941–60Generic — clean but templated or unsigned
D20–3961–80Slop — assembled from presets
F0–1981–100Pure slop

Absolute rules check

Six execution laws, each pass/fail/unverifiable, reported in their own table. Any fail is a critical tell (counts on its axis AND triggers the absolute-rule gate):

  1. Content visible by default — nothing gated on an entrance animation (opacity:0 + reveal) (M1)
  2. Clear the cut — no text/control sliced by clip, notch, overflow, or fixed height (X2, X11)
  3. Parallel alignment — comparable columns share baselines; buttons anchored (X3)
  4. Real centering — everything meant to be centered is, mathematically and optically (X1)
  5. Legible contrast — every text clears its background by a real value gap (X5)
  6. Controls work — every interactive-looking control responds (M8)

Workflow

  1. Gather evidence — route the target through the Evidence acquisition SOP above. Done when the evidence inventory states what was captured and what is Unverifiable.
  2. Read references/tells.md — the catalog you sweep against.
  3. Sweep axes 1–6 — walk the catalog group by group. Cite-or-cut: a tell is only recorded with concrete evidence (hex value, font name, file:line, or screenshot region); no evidence, no tell. Check each candidate against its premium-pair note — the crafted version of a pattern is not the tell. Done when every catalog group has been swept and every recorded tell carries a citation.
  4. Run the absolute rules check — all six, pass/fail/unverifiable with evidence.
  5. Score Axes 7–8 — read references/premium-markers.md, score the 7 signature elements and 4 cohesion checks with one-line justifications each. Done when all 11 items carry a score and a justification.
  6. Compute — score.py axis per tell-counted axis, then score.py overall with weights and any triggered --cap. Never hand-compute.
  7. Write the report — read references/output-template.md and emit exactly that structure to output, ending with the 3–5 prioritized fixes that would move the score most (biggest weighted deltas first; a missing signature usually outranks any single tell).
Show full SKILL.md (1,053 more words)Show less

Gotchas

  • The brief overrides the law. If the user or brand explicitly directed a choice (a color, a layout, an effect), it is not a tell — the law itself says the user's word wins 100%. Ask for the brief when the design clearly follows one; note excluded tells in the report with proper justification tags (see Exclusion system below).
  • Context flips a tell. Mono on real data is correct; a populated, real-feeling product window is a signature, not the fake-window tell; a tight micro-grid with texture is premium, a full-page graph paper is slop. Always check the premium pair before recording.
  • Don't reward the clean miss. Zero tells with a weak signature is the most common failure of designs that tried to avoid slop. The signature gate exists for this — apply it without mercy.
  • Severity discipline. Critical is reserved for broken (the six absolute rules). A blue-purple gradient is loud but not broken: major.
  • One-axis bleed. Some tells could sit on two axes (cut-off glow is color and execution). The catalog assigns each tell to exactly one axis — count it only there.
  • Portfolio tells. L19 (recycling your own house style) needs prior work from the same author to verify; without it, mark Unverifiable rather than guessing.

Exclusion system

Every excluded tell MUST have a justification tag. A tell without a tag counts — no exceptions. This creates an audit trail and prevents lazy exclusions.

Justification tags
TagWhen to useExample
// BRIEF:Client/stakeholder explicitly directed this choice// BRIEF: client requested blue-purple gradient as brand identity
// DESIGN DECISION:Documented design decision with concrete reasoning// DESIGN DECISION: countdown is real — sale ends 2026-08-01
// CONTEXT:Design context makes this pattern acceptable// CONTEXT: mono typeface is appropriate for code snippets in SaaS docs
// PREMIUM PAIR:This is the crafted version, not the slop version// PREMIUM PAIR: glass effect has proper refraction, edge dispersion, tuned shadows
Valid vs invalid exclusions

Valid exclusions:

markdown
| C1 | Blue→purple gradient | `// BRIEF: brand guidelines v2.3 specify #6366f1→#8b5cf6` |
| K14 | Countdown timer | `// DESIGN DECISION: real sale ends 2026-12-31, verified in CMS` |
| T4 | Mono as house voice | `// CONTEXT: SaaS product with code-heavy documentation` |
| K25 | Glass effect | `// PREMIUM PAIR: proper backdrop blur, chromatic dispersion, directional light` |

Invalid exclusions (tell still counts):

markdown
| C1 | Blue→purple gradient | "we liked it" | ❌ Not a justification
| K9 | Default CTA pair | "it's our style" | ❌ Too vague
| L1 | Default hero stack | "approved by team" | ❌ Who? When? Why?
| K6 | Kitchen-sink card | "industry standard" | ❌ Slop IS the industry standard
Exclusion limits
  • >5 exclusions → Review each one. Mass exclusions suggest the brief wasn't followed or the evaluator is being too lenient.
  • >10 exclusions → Something is wrong. Either the brief allows nearly everything (in which case, why evaluate?) or exclusions are being used to inflate the score.
  • Excluding signature elements → Almost never valid. If S1–S7 are excluded, the design has no signature by definition.
Exclusion documentation in report

In the Excluded tells table, format as:

markdown
## Excluded tells

| ID | Tell | Exclusion reason |
|----|------|------------------|
| C1 | Blue→purple gradient | `// BRIEF: brand guidelines v2.3 specify #6366f1→#8b5cf6` |
| K14 | Countdown timer | `// DESIGN DECISION: real sale ends 2026-12-31, verified in CMS` |

**Exclusion summary:** 2 tells excluded (1 BRIEF, 1 DESIGN DECISION)
Challenging exclusions

When reviewing someone else's slop report, check exclusions first:

  1. Is the tag present? No tag = tell counts.
  2. Is the tag appropriate? // BRIEF: needs an actual brief reference.
  3. Is the reasoning concrete? Vague reasoning = tell counts.
  4. Is the exclusion count reasonable? >5 warrants scrutiny.

Evaluation with Jev (Optional)

When the harness has access to TypeSafe Jev, the subjective scoring steps (Axes 7-8) and Quality Checklist verification can use Jev for calibrated assessment.

Where Jev is used:

  • Axis 7 (Signature) — 7 Score questions (S1-S7) with 0/50/100 rubric
  • Axis 8 (Cohesion) — 4 Score questions (H1-H4) with 0/50/100 rubric
  • Quality Checklist — 27 Noul questions for binary verification

Where Jev is NOT used:

  • Axes 1-6 — Tell detection is factual (cite-or-cut); score.py handles arithmetic
  • Gates & Caps — Deterministic rules applied by score.py
Discovery Protocol
1. MCP Tool `jev_eval` configured in harness → use it
2. Model `typesafe/jev-latest` via OpenRouter → request it
3. Auxiliary slot (Hermes/Devin/Codex) with Jev → delegate
4. Fallback → inline scoring via current LLM using premium-markers.md rubric
Integration Files
FileDescription
scripts/jev_questions.json38 typed questions (11 Score + 27 Noul)
references/jev-integration.mdFull protocol, request/response formats

Full documentation: See references/jev-integration.md for discovery details, harness-specific instructions, and request/response structures.


Quality checklist

Final gate before delivering. Run through every item — a single failure means the report is not ready. This is the self-evaluation rubric; treat it as a hard gate, not a suggestion.

Pre-sweep checks
  • Evidence inventory complete — documented what was captured (code, screenshots, live URL) and what is Unverifiable
  • Brief documented — if provided, summarized in report header; if not provided, noted as "no brief"
  • Design context identified — what type of design is this? (landing page, SaaS dashboard, editorial, e-commerce). Read references/contexts.md to adjust priorities and tolerances
  • All reference files read — tells.md, premium-markers.md, and contexts.md loaded before starting the sweep
During-sweep checks
  • Cite-or-cut enforced — every recorded tell has ID + severity + concrete citation (hex value, font name, file:line, or screenshot region)
  • Premium pair checked — before recording any tell, verified it's not the crafted premium version of the pattern
  • Portability test applied — for borderline cases, asked: "Could this element be moved to another site without alteration?" If yes → tell. If no (it's specific to this brand) → not a tell
  • Defense test applied — for borderline cases, asked: "Could the designer defend this choice with concrete reasoning if asked?" If no → tell. Slop cannot be defended; deliberate choices can.
  • Section attribution — every tell assigned to a specific section (Hero, Features, Pricing, Footer, etc.) for the Section Ledger
  • Severity discipline — critical reserved for absolute-rule violations only; no severity inflation
Exclusion checks
  • Exclusions documented — every excluded tell has a // BRIEF: or // DESIGN DECISION: justification
  • Exclusions are genuine — "we liked it" or "it looked good" are NOT valid exclusion reasons. Only explicit brief direction or documented design decisions with concrete reasoning qualify.
  • Exclusion count reasonable — if >5 tells excluded, double-check each one. Mass exclusions suggest the brief wasn't followed, not that the tells don't apply.
Post-sweep checks
  • All unverifiable checks marked — not silently passed or skipped
  • All 6 absolute rules reported — pass/fail/unverifiable with evidence
  • All 11 signature/cohesion items scored — 0/50/100 with one-line justification each
  • Section Ledger complete — every major section has a verdict (CLEAN/SUSPICIOUS/INFLATED/CRITICAL) with tell count and action
  • Gates applied correctly:
    • Signature gate: if Axis 7 < 40, overall capped at 59
    • Absolute-rule gate: if any crit, overall capped at 69
    • Compounding cap: if ≥3 major layout tells, Axis 4 capped at 40
  • Math from script only — all scoring via score.py, never hand-computed
Report checks
  • Template followed exactly — structure matches output-template.md
  • Fixes ranked by weighted impact — signature issues (3x weight) typically outrank single tells
  • Language correct — report in user's language, tell IDs in English
Final self-audit

Before delivering, ask yourself:

  • "What still looks like obvious slop that I didn't flag?" — if something visually screams slop but isn't in your findings, either find the tell that covers it or note it as a gap in the catalog.
  • "Did I over-correct?" — a sparse report on a clearly-slop design suggests missed tells. A bloated report on a premium design suggests false positives.
  • "Would I trust this report if someone else wrote it?" — read the report as if reviewing a colleague's work. Does every claim hold up?

© fabricioctelles, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references) in skills/slop-eval of fabricioctelles/skills.

  • SKILL.md
  • references/contexts.md
  • references/jev-integration.md
  • references/output-template.md
  • references/premium-markers.md
  • references/tells.md
  • scripts/jev_questions.json
  • scripts/score.py

Open the folder on GitHubat commit f1de632

Compare with similar skills

Slop Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Slop Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Slop Eval this skillfabricioctelles/skills106—~5.4kAutomated safety check: PassApache-2.0
Auteuragiwhitelist/auteur1k—~5kAutomated safety check: PassMIT
Taste Skillyc-software/qm15k—~1.2kAutomated safety check: PassMIT
Design Anti SlopKartikLabhshetwar/better-shot2.4k—~2.5kAutomated safety check: PassCustom licence
Human Gatealirezarezvani/claude-skills28k—~1.5kAutomated safety check: PassMIT
Anti Slop FrontendBlackBeltTechnology/pi-agent-dashboard315—~4.8kAutomated safety check: PassMIT

Similar skills

  • Auteur

    agiwhitelist/auteur

    Design and build complete web experiences from scratch — award-level product and marketing pages, cinematic scroll-directed sites where the page is directed like a film, and multi-screen products…

    1k GitHub stars~5k tokensUpdated 2 mo ago
    Frontend & DesignAuto-check passed
  • Taste Skill

    yc-software/qm

    Design taste and process for anything a person will look at in a browser — landing page, dashboard, prototype, deck.

    15k GitHub stars~1.2k tokensUpdated today
    Frontend & DesignAuto-check passed
  • Design Anti Slop

    KartikLabhshetwar/better-shot

    Detect and fix AI design slop — the convergent look that shows up across AI-generated landing pages and dashboards (purple/indigo gradients, rounded-2xl everywhere, three-box feature grids, bento…

    2.4k GitHub stars~2.5k tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • Human Gate

    alirezarezvani/claude-skills

    Runs the human-verification lane of an agent loop, and proves review happened before work is called done.

    28k GitHub stars~1.5k tokensUpdated 1 mo ago
    Frontend & DesignAuto-check passed
  • Anti Slop Frontend

    BlackBeltTechnology/pi-agent-dashboard

    A mechanical, countable anti-slop checklist for AI-generated frontend.

    315 GitHub stars~4.8k tokensUpdated today
    Frontend & DesignAuto-check passed
  • Design Taste Frontend

    caorushizi/oss-client

    Anti-slop frontend skill for landing pages, portfolios, and redesigns.

    113 GitHub starsUsed in 22 repos~22k tokens
    Frontend & DesignAuto-check passed

More from fabricioctelles/skills

All 15 skills in this repo
  • Motion Ad

    fabricioctelles/skills

    Produce a short motion-graphics video ad — a 15s Facebook/Instagram/TikTok spot — as a rendered MP4.

    106 GitHub stars~4.1k tokensUpdated 6 days ago
    Auto-check passed
  • Agent Plugin Eval

    fabricioctelles/skills

    Audit, score, and compare repositories containing portable Agent Plugins against the official Agent Plugins specification.

    106 GitHub stars~2.1k tokensUpdated 6 days ago
    Auto-check passed
  • Loop Architect

    fabricioctelles/skills

    Design well-structured agent loops with best-practice coaching and cross-model review gates before you run them.

    106 GitHub stars~2.1k tokensUpdated 6 days ago
    Auto-check: notes
  • Pier Cloud

    fabricioctelles/skills

    This skill should be used when the user needs to consume the Pier Cloud (Lighthouse) API for cloud cost management — including JWT authentication, listing contexts, workspaces, workspace groups, and…

    106 GitHub stars~1.1k tokensUpdated 6 days ago
    Auto-check: notes
  • Ralph Loop Kiro Specs

    fabricioctelles/skills

    Automated iterative agent runner for spec-based development in Kiro.

    106 GitHub stars~2.6k tokensUpdated 6 days ago
    Auto-check passed
  • Security Specialist

    fabricioctelles/skills

    Runs security audits on codebases — full scans, diff reviews, threat models, vulnerability triage, remediation guidance, and finding tracking.

    106 GitHub stars~2.8k tokensUpdated 6 days ago
    Auto-check passed

Questions about Slop Eval

What does Slop Eval do?

Objectively evaluate a UI/web design against the pols.dev anti-slop design law: detect catalogued slop tells with cited evidence, score 8 weighted axes (color, type, components, layout, motion…. Slop Eval is an agent skill from fabricioctelles/skills.dev anti-slop design law: detect catalogued slop tells with cited evidence, score 8 weighted axes (color, type, components, layout, motion, execution, signature, cohesion), and emit a Slop Report with a 0–100 Slop Index and grade.

When should I use Slop Eval?

Slop Eval fits situations like: the user asks to evaluate design slop; is this design AI slop; audit this landing page design; wants an objective score of how generic/machine-made a design looks.

How do I install Slop Eval in Claude Code?

Run `npx skills add fabricioctelles/skills --skill slop-eval -a claude-code`. Or copy the skill folder (skills/slop-eval in fabricioctelles/skills) into .claude/skills/slop-eval in your project. Claude Code loads it when a task matches its description.

How do I install Slop Eval in Codex?

Run `npx skills add fabricioctelles/skills --skill slop-eval -a codex`. Or copy the skill folder (skills/slop-eval in fabricioctelles/skills) into .agents/skills/slop-eval in your project. Codex loads it when a task matches its description.

Can I use Slop Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add fabricioctelles/skills --skill slop-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/slop-eval, .gemini/skills/slop-eval, .github/skills/slop-eval and .opencode/skills/slop-eval in your project.

What does Slop Eval need to run?

Going by SKILL.md and its folder, Slop Eval needs Python for the scripts in its folder and the command-line tools its instructions call (npx). Our summary lists: Python 3; Node.js.

Does Slop Eval access the network?

SKILL.md names 2 domains. As links in the text: pols.dev and docs.typesafe.ai. This is read from the text; nothing was executed.

Is Slop Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Slop Eval use?

Slop Eval is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Slop Eval use?

About 5.4k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 16k tokens, read only when the agent opens those files.

What are the alternatives to Slop Eval?

Skills that share tags, products or a category with Slop Eval: Auteur (agiwhitelist/auteur, 1k stars), Taste Skill (yc-software/qm, 15k stars), Design Anti Slop (KartikLabhshetwar/better-shot, 2.4k stars) and Human Gate (alirezarezvani/claude-skills, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Slop Eval?

fabricioctelles (a GitHub user) maintains it in fabricioctelles/skills, which has 106 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 4, 2026.

Source: fabricioctelles/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.