Agent skill

Hermes Self Evaluation

by AtlasOmnia in AtlasOmnia/donna-starter

hermes-self-evaluation — Use when the user asks to evaluate, audit, or optimize Hermes itself — analyzing session history, skill library, costs, and architecture to identify improvements, automation…

MITAuto-check: notes

Install Hermes Self Evaluation

skills CLI
$ npx skills add AtlasOmnia/donna-starter --skill hermes-self-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AtlasOmnia/donna-starter hermes-self-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AtlasOmnia/donna-starter.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/hermes/hermes-self-evaluation .claude/skills/hermes-self-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hermes-self-evaluation
GitHub stars
126
Token cost
~3k tokens
SKILL.md length
1,282 words
Files
4 (incl. references)
Skills in repo
11
Repo updated
First seen
Licence
MIT

At a glance

hermes-self-evaluation — Use when the user asks to evaluate, audit, or optimize Hermes itself — analyzing session history, skill library, costs, and architecture to identify improvements, automation…

  • Works in 6 steps: Map the Session Store → Map the Skill Library → Gather Usage Statistics → …
  • The user asks to evaluate
  • SKILL.md covers Overview, When to Use, Workflow and Pitfalls, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Hermes Self Evaluation is an agent skill from AtlasOmnia/donna-starter. hermes-self-evaluation — Use when the user asks to evaluate, audit, or optimize Hermes itself — analyzing session history, skill library, costs, and architecture to identify improvements, automation opportunities, and system optimizations. Covers generating structured analyst prompts for external (stronger) models to review Hermes's own performance.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/analyst-prompt-template.md`, `references/evidence-first-self-check-validation.md` and `references/runaway-session-diagnostics.md`).

The repository describes itself as: Donna — a starter Hermes Agent profile: opinionated persona, 73 curated skills, guided first-run orientation, optional Token Router. MIT. The licence is MIT.

When your agent uses it

  • The user asks to evaluate
  • Optimize Hermes itself — analyzing session history
  • Architecture to identify improvements
  • Automation opportunities

Example prompts

  • “/hermes-self-evaluation”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Map the Session Store
  2. Map the Skill Library
  3. Gather Usage Statistics
  4. Compose the Analyst Prompt
  5. Write the Prompt File
  6. Deliver

What it can do on your machine

Read from SKILL.md and the folder at commit a3710bd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are sql).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hermes Self Evaluation loads about 3k tokens when it runs, and up to ~7.7k if it reads all its reference files. Until then it costs about 94 tokens; SKILL.md has 1,282 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:133
    el/provider setup (check config.yaml and .env)

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AtlasOmnia/donna-starter at commit a3710bd, republished under its MIT licence (© AtlasOmnia). 1,282 words, ~3,002 tokens.

Download SKILL.mdSave it as .claude/skills/hermes-self-evaluation/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
hermes-self-evaluation
description
hermes-self-evaluation — Use when the user asks to evaluate, audit, or optimize Hermes itself — analyzing session history, skill library, costs, and architecture to identify improvements, automation opportunities, and system optimizations. Covers generating structured analyst prompts for external (stronger) models to review Hermes's own performance.
version
1.0.0
license
MIT
tags
hermes, evaluation, audit, optimization, meta-analysis, systems-review
metadata.tags
hermes, evaluation, audit, optimization, meta-analysis, systems-review
metadata.related_skills
skill-auditor - session-artifact-indexing

Hermes Self-Evaluation

Use this skill when the user asks to audit, review, or optimize Hermes's own performance — analyzing session data, skills, configuration, costs, and usage patterns to identify improvements, automation opportunities, and system optimizations.

Don't use for: skill content quality grading (use skill-auditor instead), one-off task analysis, or session artifact indexing after a work session (use session-artifact-indexing).

Overview

The core workflow: gather live evidence about Hermes's state and history → validate each finding against the subsystem's actual semantics and control surface → either produce a direct evidence-backed review or compose a structured analyst prompt for an independent model → verify recommendations before implementation.

External models are useful anomaly detectors and critics, but they are not automatically reliable root-cause analysts. Separate the observed symptom, supported interpretation, confirmed producer/root cause, and proposed change. Before acting on any self-check finding,

When to Use

Triggers:

  • "How can we improve Hermes?"
  • "Analyze my sessions and tell me what to optimize"
  • "I want to have another model evaluate X"
  • "Where do sessions/skills live so I can analyze them?"
  • "Do an audit of the system"
  • "What's the token cost breakdown?"

Workflow

Fast Path: Evaluate a Single Runaway Session

When the user names a specific session and says it was “working” too long, hit a tool-call guardrail, ignored “stop,” or needs the problem evaluated, do not build a broad system-audit prompt first. Diagnose the named session directly.

Use for the exact SQL/Python checks. Minimum evidence to collect:

  1. Session metadata from ~/.hermes/state.db: source, title, model, start/end times, message count, tool-call count, token totals, end reason.
  2. Role counts and top tool counts.
  3. User-message timeline and non-tool assistant replies, especially around compaction/restore and the latest steering instruction.
  4. Repeated assistant tool_call IDs. Exact repeated call_ids are a strong sign of stale tool-call replay after context compaction or gateway restore.
  5. Log markers for the session ID: max_iterations_reached, Preflight compression, Pre-API compression, gateway shutdown, Operation interrupted, tool-call guardrail, idempotent_no_progress, and transport retry loops.
  6. If relevant, check whether any live process from the runaway task is still active before saying it is safe to abandon.

Reporting rule: separate root cause from secondary symptoms. For example, browser/CUA failures may explain retries, but repeated historical tool-call IDs point to restore/replay contamination. If the user says “stop and evaluate,” stop the operational task immediately and evaluate; do not continue trying to finish the stale task.

Step 1: Map the Session Store

The canonical session database is ~/.hermes/state.db. Query its current size and schema live; never hardcode historical counts or gigabytes.

Key tables:

sql
-- sessions: id, source (cli/cron/telegram/tui/api_server/subagent/discord/bluebubbles/speech-bridge),
-- model, input_tokens, output_tokens, reasoning_tokens, cache_read_tokens,
-- cache_write_tokens, message_count, title, started_at, ended_at,
-- estimated_cost_usd, handoff_state, git_branch
CREATE TABLE sessions (...)

-- messages: session_id, role (user/assistant/tool), content (full text),
-- tool_calls, token_count, timestamp, reasoning, finish_reason
CREATE TABLE messages (...)
-- Has FTS5 + trigram full-text search indexes on messages.content

Lifecycle semantics: sessions.ended_at IS NULL means an open DB row. Total retained rows are not active sessions. Gateway routing files map resumable platform conversations; they do not prove a process is currently executing. Review detached sources by age and treat long-lived messaging rows separately. Use the queries above and the interpretation rules in this section.

Other session data:

  • ~/.hermes/session-log/*.md — daily Markdown activity logs when the session-log plugin is enabled
  • ~/.hermes/sessions/ — gateway routing/session artifacts and optional transcript snapshots; inspect current contents rather than assuming a format or count
Step 2: Map the Skill Library
  • Default profile: ~/.hermes/skills/
  • Profile-specific overrides: ~/.hermes/profiles/<profile>/skills/
  • Each skill is a directory with SKILL.md plus optional references/, templates/, and scripts/ directories.

List skills and inspect the live category tree. Never hardcode skill counts or total size: installs, curator actions, and profile changes make those values stale quickly.

Step 3: Gather Usage Statistics

Run aggregation queries against state.db to get the profile picture:

sql
-- Sessions by source
SELECT source, COUNT(*) as sessions, SUM(input_tokens + output_tokens) as total_tokens,
 SUM(message_count) as total_msgs, ROUND(SUM(estimated_cost_usd), 2) as cost
FROM sessions GROUP BY source ORDER BY total_tokens DESC;

-- Token cost by model
SELECT model, SUM(input_tokens), SUM(output_tokens),
 SUM(input_tokens + output_tokens) as total
FROM sessions WHERE model IS NOT NULL GROUP BY model ORDER BY total DESC;

-- Most active hours / patterns
SELECT strftime('%H', datetime(started_at, 'unixepoch')) as hour, COUNT(*) as sessions
FROM sessions GROUP BY hour ORDER BY sessions DESC;

-- Longest sessions (high message count)
SELECT id, source, message_count, input_tokens, output_tokens, started_at
FROM sessions ORDER BY message_count DESC LIMIT 20;

-- Repeated session titles (recurring task types)
SELECT title, COUNT(*) as freq FROM sessions
WHERE title IS NOT NULL AND title != '' GROUP BY title ORDER BY freq DESC LIMIT 30;

-- Cron session energy (most expensive cron jobs)
SELECT s.id, s.title, s.message_count, s.input_tokens, s.output_tokens
FROM sessions s WHERE s.source = 'cron'
ORDER BY s.input_tokens + s.output_tokens DESC LIMIT 20;

Also collect:

  • Total session count and total tokens across all sources
  • Current model/provider setup (check config.yaml and .env)
  • Cron job list (cronjob action='list')
  • Multi-profile architecture notes (default → domain-specific profiles delegation pattern)
Step 4: Compose the Analyst Prompt

Use the analyst-prompt template structure below as the starting point.

The prompt must be self-contained — the receiving model has no knowledge of this conversation. Include:

  1. Data locations — exact paths the model would use to query or reference
  2. Usage profile — session counts, token costs, message volumes per source
  3. Skill library structure — categories, key skills, per-profile overrides
  4. Business context — the user's businesses and workflows
  5. Architecture summary — profiles, model setup, multi-machine layout
  6. Analysis dimensions — what you want the model to evaluate (automation, skill quality, costs, architecture, UX, system health)
  7. Output format — priority-ranked findings with evidence and expected impact
  8. Starter queries — SQL the model can use to deep-dive
Step 5: Write the Prompt File

Save the composed prompt to a known location so the user can:

  • Feed it to another model (Claude, GPT, local Heretic)
  • Review and modify it before analyzing
  • Re-run the same evaluation later with updated data

Standard location: Notes/hermes-[audit|evaluation]-prompt-<date>.md

Show full SKILL.md (518 more words)Show less
Step 6: Deliver

Tell the user:

  • Where the prompt lives
  • What data it includes (dates, session counts, what was gathered)
  • Key stats from the usage profile (most expensive sessions, patterns found)
  • Offer to feed it to an external model directly if he wants

Pitfalls

  1. Prompt too long for the target model. DeepSeek Flash v4 has a 1M token context window but weaker reasoning. A local 27B Heretic has 128K. If the target model can't handle the full prompt, truncate the oldest data or summarize low-value noise sessions. Know the target model's limits before composing.

  2. Session DB queries are expensive. 160K messages in a 5.4 GB database means some aggregations take seconds. Don't run heavy queries in a loop — batch them into a single SQL multi-query or collect stats once.

  3. Cost data may be incomplete. The estimated_cost_usd column in sessions only has values for sessions where the billing provider was reachable. Many local-model sessions show $0.00 cost. Note this in the prompt as a caveat rather than asserting "these sessions cost nothing."

  4. Skill library is ~26 items. Don't enumerate every skill in the prompt body if the target model has a small context window — reference the category tree and let the model query specific categories of interest.

  5. Profiles may share a session store. Named profiles can write to the same session database; the source column distinguishes them. Don't assert clean profile isolation in session data.

  6. The prompt template can get stale. After major Hermes upgrades or session-schema changes, update the template. Query the live CLI help, config, schema, and docs instead of preserving version-specific assumptions.

  7. A real symptom does not validate the proposed fix. Verify that the named setting exists, is enabled, and governs the affected subsystem. Examples: checkpoint retention does not control session rows; a retention value does nothing when auto-prune is disabled; staged profile-only plugin enablement is not the same as a globally broken plugin; and an invented convenience key is not a configuration experiment.

  8. Read-only self-checks must stay read-only. Do not restart, prune, vacuum, consolidate memory, change delivery targets, or enable plugins globally from an unattended audit. Surface exact evidence and the narrow next action.

  9. CLI failures require syntax verification, not folklore. Run the current command's --help and a redacted dry-run. For JSONL session export, include an output target such as -: hermes sessions export - --session-id <ID> --dry-run --redact.

Verification

  • state.db exists and queries return results
  • Usage statistics collected and written into the prompt
  • Skill library structure summarized (categories, counts)
  • Prompt references are current (check dates on data samples)
  • Output file written to a known location
  • Reported to the user with key findings summary

Reference Files

  • references/analyst-prompt-template.md — the full structured analyst prompt template used as the base for composing evaluation prompts. Update this when the Hermes session schema or architecture changes meaningfully.
  • references/runaway-session-diagnostics.md — SQL/Python snippets and interpretation notes for diagnosing a specific runaway/restored session, repeated tool-call IDs, compaction contamination, and stop/steer handling.
  • references/evidence-first-self-check-validation.md — class-level rubric and exact checks for validating session, memory, retention, delivery, plugin-rollout, and CLI-syntax findings before implementing self-check recommendations.

Public support files

  • references/analyst-prompt-template.md
  • references/evidence-first-self-check-validation.md
  • references/runaway-session-diagnostics.md

© AtlasOmnia, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/hermes/hermes-self-evaluation of AtlasOmnia/donna-starter.

  • SKILL.md
  • references/analyst-prompt-template.md
  • references/evidence-first-self-check-validation.md
  • references/runaway-session-diagnostics.md

Open the folder on GitHubat commit a3710bd

Compare with similar skills

Hermes Self Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hermes Self Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hermes Self Evaluation this skillAtlasOmnia/donna-starter126—~3kAutomated safety check: NotesMIT
Arize Evaluatorgithub/awesome-copilot40k1 repos~8.1kAutomated safety check: NotesMIT
Hermes Importsaffaan-m/ECC276k—~324Automated safety check: PassMIT
LLM Evaluationdavila7/claude-code-templates32k12 repos~3.5kAutomated safety check: PassMIT
Agent Evaluationsickn33/agentic-awesome-skills47k1 repos~2kAutomated safety check: PassMIT
Hermes Importsaffaan-m/ECC276k1 repos~752Automated safety check: PassMIT

Similar skills

  • Arize Evaluator

    github/awesome-copilot

    Official

    Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

    40k GitHub starsUsed in 1 repo~8.1k tokens
    AI & LLM EngineeringAuto-check: notes
  • Hermes Imports

    affaan-m/ECC

    将本地 Hermes 操作员工作流转换为经过清理的 ECC 技能和发布包工件。在准备将 Hermes 工作流用于公共 ECC 重用而不泄露私有工作区状态、凭据或仅本地路径时使用。

    276k GitHub stars~324 tokensUpdated 4 days ago
    Auto-check passed
  • LLM Evaluation

    davila7/claude-code-templates

    Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

    32k GitHub starsUsed in 12 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Evaluation

    sickn33/agentic-awesome-skills

    Evaluate agent behavior with versioned cases and explicit verifiers.

    47k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed
  • Hermes Imports

    affaan-m/ECC

    Convert local Hermes operator workflows into sanitized ECC skills and release-pack artifacts.

    276k GitHub starsUsed in 1 repo~752 tokens
    Auto-check passed
  • Inspecting Hermes Desktop Dom

    NousResearch/hermes-agent

    Read the live Hermes desktop DOM/CSS over CDP. An agent skill from NousResearch/hermes-agent.

    252k GitHub starsUsed in 1 repo~1.6k tokens
    Frontend & DesignAuto-check passed

More from AtlasOmnia/donna-starter

All 11 skills in this repo
  • macOS Storage Management

    AtlasOmnia/donna-starter

    macos-storage-management — Use when freeing Mac storage or moving files to SSDs.

    126 GitHub stars~4.2k tokensUpdated 20 days ago
    Auto-check passed
  • Marketing Collateral Design

    AtlasOmnia/donna-starter

    marketing-collateral-design — Use when designing, recreating, critiquing, or exporting static marketing collateral such as flyers, social graphics, postcards, brochures, business cards, print ads…

    126 GitHub stars~4.2k tokensUpdated 20 days ago
    Auto-check passed
  • Skill Auditor

    AtlasOmnia/donna-starter

    skill-auditor — Use when auditing, reviewing, or grading Hermes skills for quality.

    126 GitHub stars~3.8k tokensUpdated 20 days ago
    Auto-check passed
  • Local Discovery

    AtlasOmnia/donna-starter

    local-discovery — Find local events, venues, and activities — ad-hoc web discovery when the user asks 'what's happening' or 'what should I do this weekend'.

    126 GitHub stars~3.1k tokensUpdated 20 days ago
    Auto-check passed
  • Cross Browser Typography QA

    AtlasOmnia/donna-starter

    cross-browser-typography-qa — Diagnose and verify web typography rendering defects across Chromium, WebKit, and native Safari, including clipped glyphs, broken descenders, wrapping, font metrics…

    126 GitHub stars~2.3k tokensUpdated 20 days ago
    Auto-check passed
  • Hermes Mnemosyne

    AtlasOmnia/donna-starter

    hermes-mnemosyne — Configure, troubleshoot, and operate the Mnemosyne memory provider for Hermes Agent.

    126 GitHub stars~3.9k tokensUpdated 20 days ago
    Auto-check: notes

Questions about Hermes Self Evaluation

What does Hermes Self Evaluation do?

hermes-self-evaluation — Use when the user asks to evaluate, audit, or optimize Hermes itself — analyzing session history, skill library, costs, and architecture to identify improvements, automation…. Hermes Self Evaluation is an agent skill from AtlasOmnia/donna-starter. hermes-self-evaluation — Use when the user asks to evaluate, audit, or optimize Hermes itself — analyzing session history, skill library, costs, and architecture to identify improvements, automation opportunities, and system optimizations.

When should I use Hermes Self Evaluation?

Hermes Self Evaluation fits situations like: the user asks to evaluate; optimize Hermes itself — analyzing session history; architecture to identify improvements; automation opportunities.

How do I install Hermes Self Evaluation in Claude Code?

Run `npx skills add AtlasOmnia/donna-starter --skill hermes-self-evaluation -a claude-code`. Or copy the skill folder (skills/hermes/hermes-self-evaluation in AtlasOmnia/donna-starter) into .claude/skills/hermes-self-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Hermes Self Evaluation in Codex?

Run `npx skills add AtlasOmnia/donna-starter --skill hermes-self-evaluation -a codex`. Or copy the skill folder (skills/hermes/hermes-self-evaluation in AtlasOmnia/donna-starter) into .agents/skills/hermes-self-evaluation in your project. Codex loads it when a task matches its description.

Can I use Hermes Self Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AtlasOmnia/donna-starter --skill hermes-self-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hermes-self-evaluation, .gemini/skills/hermes-self-evaluation, .github/skills/hermes-self-evaluation and .opencode/skills/hermes-self-evaluation in your project.

What does Hermes Self Evaluation need to run?

SKILL.md names no scripts, command-line tools or credentials: Hermes Self Evaluation is instructions for the agent only. Our summary lists: Python 3.

Does Hermes Self Evaluation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Hermes Self Evaluation safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Hermes Self Evaluation use?

Hermes Self Evaluation is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hermes Self Evaluation use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.7k tokens, read only when the agent opens those files.

What are the alternatives to Hermes Self Evaluation?

Skills that share tags, products or a category with Hermes Self Evaluation: Arize Evaluator (github/awesome-copilot, 40k stars), Hermes Imports (affaan-m/ECC, 276k stars), LLM Evaluation (davila7/claude-code-templates, 32k stars) and Agent Evaluation (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hermes Self Evaluation?

AtlasOmnia (a GitHub user) maintains it in AtlasOmnia/donna-starter, which has 126 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on September 19, 2026.

Source: AtlasOmnia/donna-starter on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.