Agent skill

Evaluate Skill

by every-app in every-app/open-seo

Test a candidate OpenSEO skill end to end by running fresh, isolated Codex sessions against the local backend and scoring the reports they save.

MITAuto-check: notesMarketing & SEO

Install Evaluate Skill

skills CLI
$ npx skills add every-app/open-seo --skill evaluate-skill -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install every-app/open-seo evaluate-skill --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/every-app/open-seo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/evaluate-skill .claude/skills/evaluate-skill && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluate-skill
GitHub stars
23k
Token cost
~1.8k tokens
SKILL.md length
901 words
Files
3 (incl. scripts)
Skills in repo
19
Repo updated
First seen
Licence
MIT

At a glance

Test a candidate OpenSEO skill end to end by running fresh, isolated Codex sessions against the local backend and scoring the reports they save.

  • Works in 5 steps: Copies the candidate skill and… → Creates a new local project seeded only… → Starts a gateway in front of the local… → …
  • Editing a product skill (seo-audit
  • SKILL.md covers Prerequisites, Run, Isolation rules and Compare, plus 2 more sections
  • Runs JavaScript scripts from its folder; calls node, pnpm and codex

What it does

Evaluate Skill is an agent skill from every-app/open-seo. Test a candidate OpenSEO skill end to end by running fresh, isolated Codex sessions against the local backend and scoring the reports they save. Use when editing a product skill (seo-audit, keyword-research, ...) and you need evidence that the new instructions produce better output, not just a cleaner file.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts.

It sits in Marketing & SEO, covering SEO audit and Keyword research. It works with Model Context Protocol and pnpm. The repository describes itself as: Open source alternative to Semrush and Ahrefs. The licence is MIT.

When your agent uses it

  • Editing a product skill (seo-audit
  • Keyword-research
  • ...) and you need evidence that the new instructions produce better output
  • Not just a cleaner file

Example prompts

  • “/evaluate-skill”

Requirements

  • Node.js

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Copies the candidate skill and seo-report into a fresh temp folder. Edits during the run cannot change its instructions; the manifest…
  2. Creates a new local project seeded only with the brief.
  3. Starts a gateway in front of the local MCP that allows a fixed tool list and rejects any call for another project id, then probes that…
  4. Runs codex exec ephemeral, ignoring user config, memories, apps, plugins, and multi-agent, with the skill's self-review as the only…
  5. Captures the saved report HTML, every MCP call and response, the event stream, and the final message under the output folder.

What it can do on your machine

Read from SKILL.md and the folder at commit deb4491. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • node
    • pnpm
    • codex

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pnpm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluate Skill loads about 1.8k tokens when it runs. Until then it costs about 81 tokens; SKILL.md has 901 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~81
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:14
    -auth mode: `AUTH_MODE=local_noauth` in `.env.local`, then `pnpm db:migrate:local` and `pnpm dev` (or `pnpm dev:agents`)

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from every-app/open-seo at commit deb4491, republished under its MIT licence (© every-app). 901 words, ~1,796 tokens.

Download SKILL.mdSave it as .claude/skills/evaluate-skill/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
evaluate-skill
description
Test a candidate OpenSEO skill end to end by running fresh, isolated Codex sessions against the local backend and scoring the reports they save. Use when editing a product skill (seo-audit, keyword-research, ...) and you need evidence that the new instructions produce better output, not just a cleaner file.
metadata.internal
true

Evaluate a skill

A skill edit is only as good as the reports it produces. This workflow runs the candidate skill in sessions that know nothing except the skill file, a one-line business brief, and a locked-down local MCP endpoint, then compares what came back.

Prerequisites

  • Local server in no-auth mode: AUTH_MODE=local_noauth in .env.local, then pnpm db:migrate:local and pnpm dev (or pnpm dev:agents). Pass the URL it prints, with /mcp appended, as --endpoint. Default local D1, never remote bindings or remote Postgres. Keep it private; it has no login.
  • codex CLI on the path. Runs use its configured default model at high reasoning; record the resolved model from the events log if you need to name it.
  • Research credits: the local server still calls DataForSEO. An audit run makes 20 to 40 tool calls. Do not run more sessions than the comparison needs.

Run

sh
node .agents/skills/evaluate-skill/scripts/run.mjs \
  --skill seo-audit \
  --site example.com --site holdout.example \
  --brief "Evaluation brief: Acme at example.com sells <what>. The goal is relevant organic visitors who could become customers. No first-party analytics are connected." \
  --brief "Evaluation brief: <the holdout business>. ..." \
  --endpoint http://<local-dev-host>/mcp

Runs take about 10 to 14 minutes and execute in parallel. Launch the script detached, because a tool-call timeout that kills the shell kills the sessions too:

sh
(nohup node .agents/skills/evaluate-skill/scripts/run.mjs ... > .logs/eval.log 2>&1 &)

What the script does, per site:

  1. Copies the candidate skill and seo-report into a fresh temp folder. Edits during the run cannot change its instructions; the manifest records the skill hashes.
  2. Creates a new local project seeded only with the brief.
  3. Starts a gateway in front of the local MCP that allows a fixed tool list and rejects any call for another project id, then probes that rejection before starting.
  4. Runs codex exec ephemeral, ignoring user config, memories, apps, plugins, and multi-agent, with the skill's self-review as the only reviewer.
  5. Captures the saved report HTML, every MCP call and response, the event stream, and the final message under the output folder.

Read the reports with node .agents/skills/evaluate-skill/scripts/report-text.mjs <out>/*-report.html. Each session's working notes (for seo-audit, opportunities.md and evidence/) stay in its temp folder, named in manifest.json.

Isolation rules

  • Instructions: frozen copies only. Never point a session at the working tree.
  • Conversation: a new codex exec per run. Never resume or fork.
  • Data: a new project per run, seeded with the user-provided brief and nothing else. Never write evaluator notes, prior reports, or a reference answer into any project the sessions can read.
  • Prompt: the same neutral prompt for every variant. Do not name expected findings, target keywords, or the rubric.
  • Evaluator material (rubric, reference audits, prior reports) lives outside the session folders and outside the repository. Raw artifacts belong in the ignored .logs/.

This is practical isolation, not a security boundary: sessions share the local account, provider caches, and organization limits.

Show full SKILL.md (475 more words)Show less

Compare

Always run at least two sessions of the site you are tuning on and one holdout site with a different shape, so an instruction that fits one company's answer is caught. Keep every run, including failures; do not pick the best of several and call the skill fixed.

Score each report before opening its tool trace, then read the trace to classify any miss as discovery, tool, reasoning, or reporting. For an audit report:

DimensionWeightA strong report...
Discovery and coverage25%Reads commercial pages and their siblings, tools, and technical paths; broadens when a signal appears
Diagnosis and evidence20%Checks live pages and live results; separates observations from causes; records source, date, and settings
Business value and scale20%Ties each action to a real buyer; compares competing opportunities; never invents conversion or revenue
Prioritization15%Leads with a bounded change to a demonstrated problem for likely buyers; keeps maintenance and broad programs back
Calibration10%Makes useful calls without analytics; no penalty, causation, or trend claims the evidence cannot carry
Communication and actionability10%Names the page, the change, and the reason; material findings survive compression

A material factual error overrides the score. Ask a second model to recompute every live position from the run's own MCP log before trusting a report's numbers.

Coverage checks that generalize beyond any one site:

  • Sibling pages in an important family (comparison, pricing, template, location) were read side by side, and a page that only swaps a name into a shared answer was noticed.
  • A page already at the top for its query was protected, not rewritten.
  • Every serious candidate ended as a recommendation or as a table row with a real reason. A vague "later" is a miss.
  • Search volume, visits, and revenue were kept distinct; US demand was not multiplied into a global figure.

What past evaluations taught

  • Instruction length hurts. A 3,600-word skill of cautions performed worse than a 2,200-word skill with a fixed process. Add a step, not a warning.
  • Force the comparison. Sessions that wrote a shortlist before drafting kept their diagnoses; sessions that drafted straight from research ranked by raw query volume.
  • Show the rejected candidates. A "What else we checked" table in the report is what stops material findings from disappearing when the report is shortened.
  • One snapshot lies. Positions around 8 to 12 change between two checks minutes apart; the skill re-runs decisive queries and reports both.
  • Labels beat notes. Readers rejected "10th of 17 organic rows" and arrows; tables now say "#10 (page 1)" with market and date in the header.

Clean up

Each run's project, report, and audit can be removed through MCP with delete_report, delete_site_audit, and update_project_context, or left in the local database. A new project per run is simpler than proving an old one is clean. Stop the local server when finished.

© every-app, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in .agents/skills/evaluate-skill of every-app/open-seo.

  • SKILL.md
  • scripts/report-text.mjs
  • scripts/run.mjs

Open the folder on GitHubat commit deb4491

Compare with similar skills

Evaluate Skill next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluate Skill compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluate Skill this skillevery-app/open-seo23k—~1.8kAutomated safety check: NotesMIT
SEONexus-JPF/note-companion869—~2.2kAutomated safety check: PassMIT
FLOW SEO FrameworkAgriciDaniel/claude-seo18k2 repos~1.4kAutomated safety check: PassMIT
SEO DataforseoAgriciDaniel/codex-seo7872 repos~4.6kAutomated safety check: PassMIT
SEO Audit with Search Consolenowork-studio/notfair-plugin3.9k1 repos~15kAutomated safety check: WarnMIT
Competitor GapRyze-AI-Adgent/open-seo-mcp-skills4.4k—~560Automated safety check: PassMIT

Similar skills

  • SEO

    Nexus-JPF/note-companion

    Use and read this skill immediately if the user request is in any way related to SEO or a site's organic search or AI search presence.

    869 GitHub stars~2.2k tokensUpdated 4 days ago
    Marketing & SEOAuto-check passed
  • FLOW SEO Framework

    AgriciDaniel/claude-seo

    Brings the FLOW framework's stage-specific SEO prompts into the agent, from keyword discovery through backlinks, on-page work and conversion to local SEO, loaded on demand.

    18k GitHub starsUsed in 2 repos~1.4k tokens
    Marketing & SEOAuto-check passed
  • SEO Dataforseo

    AgriciDaniel/codex-seo

    Live SEO data via DataForSEO MCP server. An agent skill from AgriciDaniel/codex-seo.

    787 GitHub starsUsed in 2 repos~4.6k tokens
    Marketing & SEOAuto-check passed
  • SEO Audit with Search Console

    nowork-studio/notfair-plugin

    Runs a full SEO audit that combines Google Search Console, URL Inspection, PageSpeed Insights and a technical crawl, then ranks quick wins and a 30-day plan.

    3.9k GitHub starsUsed in 1 repo~15k tokens
    Marketing & SEOAuto-check: warnings
  • Competitor Gap

    Ryze-AI-Adgent/open-seo-mcp-skills

    Find keywords a competitor ranks for that the user's site doesn't — the content gap, prioritized by volume and winnability.

    4.4k GitHub stars~560 tokensUpdated 13 days ago
    Marketing & SEOAuto-check passed
  • 90 Day SEO Sprint

    Bomx/distribb-skill

    Run the Distribb 90-Day SEO Sprint - a founder-built 13-week playbook for shipping pre-launch SEO, core pages, a content engine, and a backlink starter stack on the way to compounding organic traffic.

    198 GitHub stars~3.6k tokensUpdated 6 days ago
    Marketing & SEOAuto-check passed

More from every-app/open-seo

All 19 skills in this repo
  • Papercuts

    every-app/open-seo

    Log genuine, recurring repository friction to .agents/PAPERCUTS.md — confusing setup, a flaky repo command or script, a misleading in-repo error, stale generated files, or a non-obvious gotcha that…

    23k GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Simple Issue Description

    every-app/open-seo

    Turn a rough bug report, feature request, support note, or pull request into a short, plain-language issue focused on the problem and desired behavior.

    23k GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Create Repo Skill

    every-app/open-seo

    Create or update a skill in this repository the right way — canonical home in .agents/skills, internal-vs-public marking, symlink mirroring into .claude/skills, and public docs registration for…

    23k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Deslop

    every-app/open-seo

    Remove AI writing patterns from prose so it reads like a person wrote it.

    23k GitHub stars~602 tokensUpdated today
    Auto-check passed
  • Observability Triage

    every-app/open-seo

    Triage OpenSEO production errors in Cloudflare Workers Observability — verified query recipes, counting gotchas, and a known-noise filter list applied automatically.

    23k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • SEO Report

    every-app/open-seo

    Write and save an OpenSEO report as one self-contained HTML page.

    23k GitHub stars~5.4k tokensUpdated today
    Auto-check passed

Categories

Questions about Evaluate Skill

What does Evaluate Skill do?

Test a candidate OpenSEO skill end to end by running fresh, isolated Codex sessions against the local backend and scoring the reports they save. Evaluate Skill is an agent skill from every-app/open-seo. Test a candidate OpenSEO skill end to end by running fresh, isolated Codex sessions against the local backend and scoring the reports they save.

When should I use Evaluate Skill?

Evaluate Skill fits situations like: editing a product skill (seo-audit; keyword-research; ...) and you need evidence that the new instructions produce better output; not just a cleaner file.

How do I install Evaluate Skill in Claude Code?

Run `npx skills add every-app/open-seo --skill evaluate-skill -a claude-code`. Or copy the skill folder (.agents/skills/evaluate-skill in every-app/open-seo) into .claude/skills/evaluate-skill in your project. Claude Code loads it when a task matches its description.

How do I install Evaluate Skill in Codex?

Run `npx skills add every-app/open-seo --skill evaluate-skill -a codex`. Or copy the skill folder (.agents/skills/evaluate-skill in every-app/open-seo) into .agents/skills/evaluate-skill in your project. Codex loads it when a task matches its description.

Can I use Evaluate Skill in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add every-app/open-seo --skill evaluate-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluate-skill, .gemini/skills/evaluate-skill, .github/skills/evaluate-skill and .opencode/skills/evaluate-skill in your project.

What does Evaluate Skill need to run?

Going by SKILL.md and its folder, Evaluate Skill needs JavaScript for the scripts in its folder and the command-line tools its instructions call (node, pnpm and codex). Our summary lists: Node.js.

Does Evaluate Skill access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evaluate Skill safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Evaluate Skill use?

Evaluate Skill is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluate Skill use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evaluate Skill?

Skills that share tags, products or a category with Evaluate Skill: SEO (Nexus-JPF/note-companion, 869 stars), FLOW SEO Framework (AgriciDaniel/claude-seo, 18k stars), SEO Dataforseo (AgriciDaniel/codex-seo, 787 stars) and SEO Audit with Search Console (nowork-studio/notfair-plugin, 3.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluate Skill?

every-app (a GitHub organization) maintains it in every-app/open-seo, which has 22,587 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 6, 2026.

Source: every-app/open-seo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.