Evaluate a batch of content in one run: ranked quality report, systemic issues.

MITAuto-check passed

Install Eval Suite

skills CLI
$ npx skills add indranilbanerjee/digital-marketing-pro --skill eval-suite -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install indranilbanerjee/digital-marketing-pro eval-suite --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/indranilbanerjee/digital-marketing-pro.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-suite .claude/skills/eval-suite && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-suite
GitHub stars
862
Used in
1 other repo
Token cost
~2.2k tokens
SKILL.md length
1,092 words
Files
1
Skills in repo
162
Repo updated
First seen
Licence
MIT

At a glance

Evaluate a batch of content in one run: ranked quality report, systemic issues.

  • Works in 9 steps: Load brand context: Read… → Enumerate all content items: Resolve the… → Evaluate each content item: For each… → …
  • SKILL.md covers Purpose, Input Required, Process and Output, plus 1 more section
  • Calls python

What it does

Eval Suite is an agent skill from indranilbanerjee/digital-marketing-pro. Evaluate a batch of content in one run: ranked quality report, systemic issues. "score our whole content library"

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: An open-source AI marketing operating system for strategy, SEO, AEO/GEO, paid media, content, CRM, and analytics - grounded in brand context, human approval, and verifiable… The licence is MIT.

Example prompts

  • “score our whole content library”
  • “/eval-suite”

Requirements

  • Python 3

Workflow steps

9 steps, taken from the first numbered list in SKILL.md.

  1. Load brand context: Read ~/.claude-marketing/brands/_active-brand.json for the active slug, then load…
  2. Enumerate all content items: Resolve the provided sources into a flat list of content items. For directory paths, scan for text-based…
  3. Evaluate each content item: For each item in the set, run python "${CLAUDE_PLUGIN_ROOT}/scripts/eval-runner.py" --brand {slug} --action…
  4. Log each evaluation: For every evaluated item, run python "${CLAUDE_PLUGIN_ROOT}/scripts/quality-tracker.py" --brand {slug} --action…
  5. Aggregate results: Compute portfolio-level statistics
  6. Rank all content pieces: Sort items from highest to lowest composite score. Present the full ranked list with scores, grades, and content…
  7. Identify common issues: Analyze the per-dimension scores across all items to find patterns — e.g., "7 of 12 items score below 70 on…
  8. Generate prioritized revision list: Sort items that need revision by potential impact. Prioritize items that are (a) below the auto-reject…
  9. Compare against baseline (if provided): If the user provided a previous suite run ID, retrieve both the current and baseline suite scores…

What it can do on your machine

Read from SKILL.md and the folder at commit 9e949f3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Suite loads about 2.2k tokens when it runs. Until then it costs about 31 tokens; SKILL.md has 1,092 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~31
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from indranilbanerjee/digital-marketing-pro at commit 9e949f3, republished under its MIT licence (© indranilbanerjee). 1,092 words, ~2,182 tokens.

Download SKILL.mdSave it as .claude/skills/eval-suite/SKILL.md (or your agent's skills folder).
name
eval-suite
description
Evaluate a batch of content in one run: ranked quality report, systemic issues. "score our whole content library"

/digital-marketing-pro:eval-suite

Script location. If your host does not set ${CLAUDE_PLUGIN_ROOT}, the scripts are in this plugin's scripts/ folder, next to skills/.

Purpose

Batch evaluation across multiple content pieces to produce a portfolio-level quality assessment. Evaluate an entire content library, all assets in a campaign, or a set of deliverables in one run. Instead of evaluating content one piece at a time, this command processes everything together and delivers a holistic view of content quality.

The output includes content rankings, per-dimension analysis, overall quality distribution, common issues across the set, and a prioritized revision list. This is the command to use before a campaign launch (to catch weak assets before they go live), during a content audit (to assess library health), or after a production sprint (to quality-check all deliverables at once). Every evaluation is logged to the quality tracker for longitudinal trend analysis.

Input Required

The user must provide (or will be prompted for):

  • Content sources: One or more of the following:
    • A list of file paths (e.g., "evaluate these 5 files: email-v1.txt, email-v2.txt, landing-page.html, ad-copy-fb.txt, ad-copy-google.txt")
    • A directory path (e.g., "evaluate everything in /campaign-q1-assets/") — all text-based files in the directory will be included
    • Multiple inline content blocks with labels (e.g., "Evaluate these: [Label: Homepage Hero] content... [Label: Email Subject] content...")
  • Content type: Optional — applied globally (e.g., "these are all email subject lines") or specified per item. If omitted, the evaluator will infer type from content characteristics
  • Evidence file: Optional — shared context document (brief, strategy doc, audience research) applied across all evaluations for more relevant scoring
  • Evaluation depth: Optional — quick (default, faster per-item evaluation) or full (comprehensive evaluation with detailed per-dimension commentary per item). Quick is recommended for sets larger than 10 items; full for critical campaign assets
  • Auto-reject threshold: Optional — composite score below which content is flagged as needing mandatory revision (default: 60)
  • Comparison baseline: Optional — a previous eval-suite run ID to compare against, showing improvement or regression per piece

Process

  1. Load brand context: Read ~/.claude-marketing/brands/_active-brand.json for the active slug, then load ~/.claude-marketing/brands/{slug}/profile.json. Apply brand voice, compliance rules for target markets (skills/context-engine/compliance-rules.md), and industry context. Check for guidelines at ~/.claude-marketing/brands/{slug}/guidelines/_manifest.json — if present, load restrictions and relevant category files (voice-and-tone, messaging, channel styles). Check for custom templates at ~/.claude-marketing/brands/{slug}/templates/. Check for agency SOPs at ~/.claude-marketing/sops/. If no brand exists, ask: "Set up a brand first (/digital-marketing-pro:brand-setup)?" — or proceed with defaults.
  2. Enumerate all content items: Resolve the provided sources into a flat list of content items. For directory paths, scan for text-based files (.txt, .md, .html, .csv rows). For inline content, parse labels and content blocks. Assign a label to each item (filename, provided label, or auto-generated index). Report the total item count to the user before proceeding and confirm if the set is larger than 25 items (to set expectations on processing time).
  3. Evaluate each content item: For each item in the set, run python "${CLAUDE_PLUGIN_ROOT}/scripts/eval-runner.py" --brand {slug} --action run-quick --file "{path}" --content-type "{type}" for file items (use --text "{content}" instead of --file for inline content blocks; use --action run-full if the user requested comprehensive depth). Pass --evidence "{evidence_path}" if an evidence file was provided. Collect the per-dimension scores (content_quality, brand_voice, hallucination_risk, claim_verification, output_structure, readability) and composite score for each item.
  4. Log each evaluation: For every evaluated item, run python "${CLAUDE_PLUGIN_ROOT}/scripts/quality-tracker.py" --brand {slug} --action log-eval --content-type "{type}" --data '{"label": "{label}", "scores": {scores_json}, "suite_id": "{suite_run_id}"}' to persist results for longitudinal tracking. The suite-id groups all items from this batch together.
  5. Aggregate results: Compute portfolio-level statistics:
    • Average composite score across all items
    • Score distribution — count of items in each grade band (90+: Excellent, 80-89: Strong, 70-79: Good, 60-69: Needs Work, <60: Auto-reject)
    • Per-dimension portfolio averages — identify which quality dimensions are consistently strong or weak across the entire set
    • Standard deviation to assess consistency (high deviation means uneven quality)
  6. Rank all content pieces: Sort items from highest to lowest composite score. Present the full ranked list with scores, grades, and content type labels.
  7. Identify common issues: Analyze the per-dimension scores across all items to find patterns — e.g., "7 of 12 items score below 70 on claim_verification" or "hallucination_risk scores are consistently 15+ points below content_quality scores." These systemic patterns indicate process or template issues rather than individual content problems.
  8. Generate prioritized revision list: Sort items that need revision by potential impact. Prioritize items that are (a) below the auto-reject threshold, (b) high-visibility content types (landing pages, ads) with below-average scores, or (c) items where a single dimension drags down an otherwise strong composite. For each item on the revision list, specify which dimension(s) to focus on and what kind of improvement is needed.
  9. Compare against baseline (if provided): If the user provided a previous suite run ID, retrieve both the current and baseline suite scores from the quality tracker using python "${CLAUDE_PLUGIN_ROOT}/scripts/quality-tracker.py" --brand {slug} --action get-summary for each suite period. Then compute per-item and portfolio-level deltas yourself by matching items across the two runs by label/content-type and calculating score differences. Present results as improved, regressed, or unchanged per item and overall.
Show full SKILL.md (255 more words)Show less

Output

A structured portfolio quality assessment containing:

  • Portfolio summary: Total piece count, average composite score, grade distribution (Excellent/Strong/Good/Needs Work/Auto-reject counts), overall portfolio grade, consistency score (based on standard deviation)
  • Ranked content list: All items sorted best to worst — each with label, content type, composite score, grade, and a one-line quality summary
  • Top performers: The 3 highest-scoring items with specific notes on what makes them strong — useful as internal benchmarks or templates
  • Per-dimension portfolio analysis: Average score per dimension across the full set (content_quality, brand_voice, hallucination_risk, claim_verification, output_structure, readability), identifying the strongest and weakest dimensions with specific observations (e.g., "brand_voice averages 88 across the set — voice guidelines are being followed well. claim_verification averages 62 — sources and supporting evidence are frequently missing.")
  • Common issues report: Systemic patterns found across multiple items — these indicate process-level problems worth fixing at the template or brief stage rather than per-item revision
  • Prioritized revision list: Items most in need of revision, sorted by impact, with specific guidance on which dimensions to improve and what kind of changes are needed
  • Auto-reject list: Items scoring below the threshold with specific reasons and mandatory revision flags
  • Baseline comparison (if applicable): Per-item deltas and portfolio-level improvement/regression metrics
  • Recommendations: Actionable next steps — which items to revise first, which process improvements would lift the entire portfolio, and whether any content types consistently underperform (suggesting brief or template issues)

Agents Used

  • quality-assurance -- Evaluates each content piece across all quality dimensions, maintains scoring consistency across the batch, identifies systemic quality patterns, generates portfolio-level insights, and produces the prioritized revision recommendations

© indranilbanerjee, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/eval-suite of indranilbanerjee/digital-marketing-pro.

Open the folder on GitHubat commit 9e949f3

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in indranilbanerjee/digital-marketing-pro, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Eval Suite next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Suite compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Suite this skillindranilbanerjee/digital-marketing-pro8621 repos~2.2kAutomated safety check: PassMIT
Evals Create Suiteelastic/kibana21k—~1.7kAutomated safety check: PassCustom licence
Evalalirezarezvani/claude-skills28k—~618Automated safety check: PassMIT
Eval-Driven Development Harnessaffaan-m/ECC276k—~1.5kAutomated safety check: PassMIT
Eval Harnessaffaan-m/ECC276k—~2.2kAutomated safety check: PassMIT
OmniRoute CLI Evalsdiegosouzapw/OmniRoute75k—~1.3kAutomated safety check: PassMIT

Similar skills

  • Evals Create Suite

    elastic/kibana

    Official

    Scaffold a new LLM evaluation suite package with Playwright config, evaluate fixture, and package files.

    21k GitHub stars~1.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Eval

    alirezarezvani/claude-skills

    Evaluate and rank agent results by metric or LLM judge for an AgentHub session.

    28k GitHub stars~618 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.

    276k GitHub stars~1.5k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Eval Harness

    affaan-m/ECC

    Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k…

    276k GitHub stars~2.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • OmniRoute CLI Evals

    diegosouzapw/OmniRoute

    Creates and runs LLM evaluation suites from the omniroute CLI, follows live runs, shows scorecards, compares models and ties eval runs into CI.

    75k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    affaan-m/ECC

    Eval-driven development (EDD) ilkelerini uygulayan Claude Code oturumları için formal değerlendirme çerçevesi

    276k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed

More from indranilbanerjee/digital-marketing-pro

All 162 skills in this repo
  • Import Template

    indranilbanerjee/digital-marketing-pro

    Import a deliverable template as a reusable placeholder template per brand.

    862 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed
  • Ab Test Plan

    indranilbanerjee/digital-marketing-pro

    Plan an A/B test by script: sample size per variant, days to run, stopping rules.

    862 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Aeo Audit

    indranilbanerjee/digital-marketing-pro

    Run a one-time AEO audit of six AI answer engines, scored per surface.

    862 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Agent Readiness Audit

    indranilbanerjee/digital-marketing-pro

    Audit agent readiness by script: AI-crawler rules, product schema, no-JS HTML, feeds.

    862 GitHub starsUsed in 1 repo~3.7k tokens
    Auto-check passed
  • Backlink Gap

    indranilbanerjee/digital-marketing-pro

    Find backlink gap domains linking to competitors, not you, scored by script.

    862 GitHub starsUsed in 1 repo~2.6k tokens
    Auto-check passed
  • C2pa Metadata

    indranilbanerjee/digital-marketing-pro

    Embed C2PA provenance in AI-generated images, video or PDF by script.

    862 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed

Questions about Eval Suite

What does Eval Suite do?

Evaluate a batch of content in one run: ranked quality report, systemic issues. Eval Suite is an agent skill from indranilbanerjee/digital-marketing-pro. Evaluate a batch of content in one run: ranked quality report, systemic issues.

How do I install Eval Suite in Claude Code?

Run `npx skills add indranilbanerjee/digital-marketing-pro --skill eval-suite -a claude-code`. Or copy the skill folder (skills/eval-suite in indranilbanerjee/digital-marketing-pro) into .claude/skills/eval-suite in your project. Claude Code loads it when a task matches its description.

How do I install Eval Suite in Codex?

Run `npx skills add indranilbanerjee/digital-marketing-pro --skill eval-suite -a codex`. Or copy the skill folder (skills/eval-suite in indranilbanerjee/digital-marketing-pro) into .agents/skills/eval-suite in your project. Codex loads it when a task matches its description.

Can I use Eval Suite in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add indranilbanerjee/digital-marketing-pro --skill eval-suite -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-suite, .gemini/skills/eval-suite, .github/skills/eval-suite and .opencode/skills/eval-suite in your project.

What does Eval Suite need to run?

Going by SKILL.md and its folder, Eval Suite needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Eval Suite access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Suite safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Suite use?

Eval Suite is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Suite use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Suite?

Skills that share tags, products or a category with Eval Suite: Evals Create Suite (elastic/kibana, 21k stars), Eval (alirezarezvani/claude-skills, 28k stars), Eval-Driven Development Harness (affaan-m/ECC, 276k stars) and Eval Harness (affaan-m/ECC, 276k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Suite?

indranilbanerjee (a GitHub user) maintains it in indranilbanerjee/digital-marketing-pro, which has 862 GitHub stars. The repository holds 162 skills in this directory. The repository was last updated on October 9, 2026.

Source: indranilbanerjee/digital-marketing-pro on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.