Agent skill

Evaluate Agent Native

by haoruilee in haoruilee/awesome-agent-native-services

Evaluate catalog candidates against either the standard five agent-native criteria or the narrow operator-surface track, and check URL Onboarding.

CC0-1.0Auto-check passedAgent Workflows

Install Evaluate Agent Native

skills CLI
$ npx skills add haoruilee/awesome-agent-native-services --skill evaluate-agent-native -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install haoruilee/awesome-agent-native-services evaluate-agent-native --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/haoruilee/awesome-agent-native-services.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.skills/evaluate-agent-native .claude/skills/evaluate-agent-native && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluate-agent-native
GitHub stars
476
Token cost
~2.8k tokens
SKILL.md length
1,068 words
Files
1
Skills in repo
4
Repo updated
First seen
Licence
CC0-1.0

At a glance

Evaluate catalog candidates against either the standard five agent-native criteria or the narrow operator-surface track, and check URL Onboarding.

  • Works in 5 steps: Agent-operations-first: Official… → Agent-specific live state: It… → Session attribution: Displayed state… → …
  • Asked whether a service
  • SKILL.md covers The gold standard: URL…, When to activate, Choose the admission track and Standard track: five hard…, plus 5 more sections
  • Calls npx; reaches moltbook.com and ensue.dev

What it does

Evaluate Agent Native is an agent skill from haoruilee/awesome-agent-native-services. Evaluate catalog candidates against either the standard five agent-native criteria or the narrow operator-surface track, and check URL Onboarding. Use when asked whether a service, harness, HUD, status line, or control surface belongs in the catalog.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Works with agents that can inspect official websites, repositories, and protocol documentation.

It sits in Agent Workflows, covering Agent evaluation and testing. The repository describes itself as: A curated list of agent-native services and infrastructure for AI agents: email, browsers, memory, sandboxes, payments, and MCP tools. Includes selection criteria and onboarding… The licence is CC0-1.0.

When your agent uses it

  • Asked whether a service
  • Control surface belongs in the catalog

Example prompts

  • “/evaluate-agent-native”

Requirements

  • Node.js
  • Compatibility (from SKILL.md): Works with agents that can inspect official websites, repositories, and protocol documentation.
  • Pre-approved tools (allowed-tools): WebSearch, Read

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Agent-operations-first: Official positioning is explicitly about running or supervising AI agents.
  2. Agent-specific live state: It continuously interprets state such as rollout/session identity, context pressure, tools, subagents…
  3. Session attribution: Displayed state maps to a concrete agent session, rollout, or typed agent path.
  4. Dedicated operational surface: It exposes a CLI, wrapper, daemon, hook, or status-line protocol purpose-built for agent operations.
  5. Honest boundary: It states when autonomy, machine APIs, delegated credentials, approval enforcement, Skills, or MCP are absent. Read-only…

What it can do on your machine

Read from SKILL.md and the folder at commit b1e6126. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • WebSearch
    • Read

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • moltbook.com
    • ensue.dev
    • api.ensue-network.ai
    • raw.githubusercontent.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Works with agents that can inspect official websites, repositories, and protocol documentation.

    From compatibility in the SKILL.md frontmatter.

Context cost

Evaluate Agent Native loads about 2.8k tokens when it runs. Until then it costs about 68 tokens; SKILL.md has 1,068 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~68
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from haoruilee/awesome-agent-native-services at commit b1e6126, republished under its CC0-1.0 licence (© haoruilee). 1,068 words, ~2,828 tokens.

Download SKILL.mdSave it as .claude/skills/evaluate-agent-native/SKILL.md (or your agent's skills folder).
name
evaluate-agent-native
description
Evaluate catalog candidates against either the standard five agent-native criteria or the narrow operator-surface track, and check URL Onboarding. Use when asked whether a service, harness, HUD, status line, or control surface belongs in the catalog.
allowed-tools
WebSearch, Read
compatibility
Works with agents that can inspect official websites, repositories, and protocol documentation.
license
CC0-1.0
metadata.repo
https://github.com/haoruilee/awesome-agent-native-services
metadata.catalog-version
2026-10-05

Skill: evaluate-agent-native

Use this skill to select the correct admission track, collect primary-source evidence, and classify a candidate. Most services use the five-criterion standard. Purpose-built surfaces for operating live agents may use the narrow operator-surface track.

The gold standard: URL Onboarding

Before applying the five criteria, ask the highest-level question:

Can an agent join and start using this service by reading a single URL?

Services that answer YES are exhibiting the strongest possible form of agent-nativeness. They have internalized the agent as first-class user so deeply that the onboarding flow itself is machine-readable:

# The full agent onboarding in one instruction:
Read <url> and follow the instructions.

Examples:

  • Moltbook: Read https://www.moltbook.com/skill.md — complete registration, heartbeat, posting, DM protocol
  • Ensue: Read https://ensue.dev/docs and call POST https://api.ensue-network.ai/auth/agent-register — shared-memory agent registration
  • autoresearch@home: Read https://raw.githubusercontent.com/mutable-state-inc/autoresearch-at-home/master/collab.md — complete swarm joining, claiming, publishing protocol
  • db9 / mem9 / mails.dev / MailboxKit / Agents Mail: domain-hosted skill.md URL onboarding for database, memory, and email workflows

This is qualitatively different from:

  • An SDK that a human developer installs (requires human coding time)
  • An MCP server that a human adds to a config file (requires human config edit)
  • A REST API that requires API key setup (requires human account creation)

URL Onboarding means the agent itself handles all of this — reading, understanding, and executing the join sequence autonomously.

Mark URL Onboarding as a strong bonus signal and highlight it prominently in the evaluation report.


When to activate

Activate when the user asks:

  • "Is [service] agent-native?"
  • "Does [service] qualify for the awesome list?"
  • "I want to add [service] — does it meet the criteria?"
  • "What's the difference between agent-native and agent-adapted?"
  • "Why isn't [service] on the list?"
  • "Does [service] have URL Onboarding?"

Choose the admission track

Use the standard track for infrastructure an agent consumes or invokes. Use the operator-surface track only when the product's original and primary purpose is operating live AI agents.

Do not use the exception for a generic dashboard, IDE skin, terminal theme, or process monitor that merely adds agent labels.

Standard track: five hard criteria

A standard-track service must pass all five. Evaluate each one explicitly.

Criterion 1 — Agent-First Positioning

Test: Does the official homepage or documentation explicitly identify AI agents as the primary consumer?

Evidence to look for:

  • Homepage headline naming AI agents
  • Documentation framing agents as the core user
  • Product name or tagline that only makes sense for agents

Red flags:

  • "Now with AI agent support" (agents are an add-on)
  • "Build apps, workflows, and agents" (agents are one of many outputs)
Criterion 2 — Agent-Specific Primitives

Test: Does the API expose at least one primitive with no meaningful human-facing equivalent?

Questions to ask:

  • What is the core API object? Agent inbox? KYA token? Claim? Heartbeat? Or generic inbox/token/task?
  • Would this primitive exist if agents didn't exist?
  • Is the output format optimized for LLM consumption or human reading?

Pass examples: agent inbox, KYA identity token, approval gate with context-window injection, claim_experiment(), heartbeat protocol, publish_hypothesis().

Fail examples: a REST API that sends emails (humans use it too), a webhook any server can receive.

Criterion 3 — Autonomy-Compatible Control Plane

Test: Can an agent complete a full task loop without a human clicking anything?

Questions to ask:

  • Can the agent provision its own credentials?
  • Can the agent initiate, execute, and complete the action without a human redirect?
  • Does the service provide agent-appropriate constraint mechanisms?
Criterion 4 — Machine-to-Machine Integration Surface

Test: Is the primary interface an SDK, REST API, MCP server, webhook, or machine-readable URL?

Questions to ask:

  • Can an agent use this service without a human ever opening a browser?
  • Is there a URL, SDK, REST API, or MCP server documented as the primary integration path?

Note: A service that exposes a machine-readable skill.md or protocol URL (URL Onboarding) passes this criterion with exceptional strength.

Criterion 5 — Agent Identity / Delegation Semantics

Test: Does the service distinguish (a) agent's own identity, (b) delegated user permissions, (c) audit trail?


Show full SKILL.md (440 more words)Show less

Operator-surface track

An operator surface must pass all five track requirements:

  1. Agent-operations-first: Official positioning is explicitly about running or supervising AI agents.
  2. Agent-specific live state: It continuously interprets state such as rollout/session identity, context pressure, tools, subagents, permissions, Skills, MCP configuration, or quota windows.
  3. Session attribution: Displayed state maps to a concrete agent session, rollout, or typed agent path.
  4. Dedicated operational surface: It exposes a CLI, wrapper, daemon, hook, or status-line protocol purpose-built for agent operations.
  5. Honest boundary: It states when autonomy, machine APIs, delegated credentials, approval enforcement, Skills, or MCP are absent. Read-only observation must not be presented as orchestration or authorization.

Session attach/list/stop controls strengthen the case but are not mandatory. An operator surface can be human-facing and read-only; that is the point of this narrow track.


Bonus signals (check all that apply)

SignalWeightEvidence to look for
URL Onboarding ⭐⭐⭐HighestService hosts a machine-readable skill.md / protocol doc an agent reads and follows to self-register
Dedicated agent identity modelHighAgent gets its own credential/wallet/token
MCP server publishedMediumOfficial MCP server with documented tools
Agent Skills (SKILL.md) publishedMediumnpx skills add org/repo works
Per-agent state / memory / sessionMediumState isolated by agent instance
Audit / trajectory artifactsMediumMachine-readable evidence of agent actions

How to test for URL Onboarding:

  1. Look for a skill.md, SKILL.md, collab.md, or similar machine-readable protocol file hosted at the service's domain or GitHub.
  2. Ask: could an agent read that URL and complete the full registration/onboarding sequence autonomously?
  3. Try the instruction: Read <url> and follow the instructions — does it work?

Classification decision tree

Is this infrastructure agents consume or invoke?
├── YES → apply all five standard criteria
│   └── PASS → agent-native (standard) ✅
└── NO → was it purpose-built to operate live AI agents?
    ├── YES → apply all five operator-surface requirements
    │   └── PASS → agent-native (operator surface) ✅
    └── NO → agent-adapted, agent-builder, or out of scope

For either qualifying track, add ⭐ when URL Onboarding is real.

Evaluation output format

## Evaluation: {Service Name}
**Website:** {url}
**Admission track:** Standard / Operator surface

### URL Onboarding Check ⭐
**Has URL Onboarding:** YES / NO
**Onboarding instruction (if YES):** Read {url} and follow the instructions to {join/register/participate}
**Notes:** {what the agent gets by reading that URL}

---

### Criterion 1 — Agent-First Positioning
**Result:** PASS / FAIL / PARTIAL
**Evidence:** "{exact quote}" — {source URL}

### Criterion 2 — Agent-Specific Primitives
**Result:** PASS / FAIL / PARTIAL
**Evidence:** {primitive name and description}
**No human equivalent because:** {explanation}

### Criterion 3 — Autonomy-Compatible Control Plane
**Result:** PASS / FAIL / PARTIAL
**Evidence:** {how agents operate without human confirmation}

### Criterion 4 — Machine-to-Machine Integration Surface
**Result:** PASS / FAIL / PARTIAL
**Evidence:** {URL, SDK, API, MCP details}

### Criterion 5 — Agent Identity / Delegation Semantics
**Result:** PASS / FAIL / PARTIAL / N/A
**Evidence:** {identity model details}

### Operator-Surface Track (complete instead of standard criteria when selected)
1. Agent-operations-first — PASS / FAIL — {evidence}
2. Agent-specific live state — PASS / FAIL — {evidence}
3. Session attribution — PASS / FAIL — {evidence}
4. Dedicated operational surface — PASS / FAIL — {evidence}
5. Honest boundary — PASS / FAIL — {documented limitations}

---

### Bonus signals
- [ ] URL Onboarding ⭐⭐⭐ — agent joins by reading one URL
- [ ] Dedicated agent identity model
- [ ] MCP server published
- [ ] Agent Skills (SKILL.md) published
- [ ] Per-agent state/memory/session
- [ ] Audit/trajectory/replay artifacts

---

### Overall verdict
**Classification:** agent-native ⭐ / agent-native (standard) / agent-native (operator surface) / agent-adapted / agent-builder / out of scope
**Recommendation:** Add to main list / Add to Excluded section / Do not add
**Confidence:** High / Medium / Low
**Reasoning:** {one paragraph summary}

### Next steps
{If agent-native with URL Onboarding: highlight this in the issue and service file prominently}
{If agent-native without: link to issue template}
{If agent-adapted: explain what would need to change}

Common borderline cases

"The product added an MCP server — does that make it agent-native?"

No. MCP support is a bonus signal, not a criterion. The core question is whether the service was designed from inception for agents. A human email provider that adds an MCP server is still agent-adapted.

"The service has URL Onboarding but other criteria are weak."

URL Onboarding is the strongest bonus signal but cannot substitute for the selected admission track. It is an amplifier, not a replacement.

"The service says 'for AI agents' in marketing."

Check the actual primitives. URL Onboarding is a reliable signal because it requires genuine design effort — you can't fake it with a marketing blog post.

"It is a human-facing Codex HUD. Is that automatically excluded?"

No. Apply the operator-surface track. A HUD qualifies only when it was purpose-built for live agents, interprets agent-specific runtime state, attributes it to concrete sessions, exposes a dedicated operational surface, and discloses its lack of control or delegated authority. A generic terminal dashboard still fails.

© haoruilee, CC0-1.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .skills/evaluate-agent-native of haoruilee/awesome-agent-native-services.

Open the folder on GitHubat commit b1e6126

Compare with similar skills

Evaluate Agent Native next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluate Agent Native compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluate Agent Native this skillhaoruilee/awesome-agent-native-services476—~2.8kAutomated safety check: PassCC0-1.0
MCP Server Builderanthropics/skills180k62 repos~2.3kAutomated safety check: PassApache-2.0
Diagnosing Superpowers Sessionsobra/superpowers296k3 repos~1.7kAutomated safety check: PassMIT
Darwin Skill Optimizeralchaincyf/darwin-skill6.2k1 repos~4.7kAutomated safety check: PassMIT
Skill Release Gaterohitg00/ai-engineering-from-scratch65k—~1kAutomated safety check: PassMIT
CodeGraph Agent Evalcolbymchenry/codegraph73k—~950Automated safety check: PassMIT

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 62 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers.

    296k GitHub starsUsed in 3 repos~1.7k tokens
    Agent WorkflowsAuto-check passed
  • Darwin Skill Optimizer

    alchaincyf/darwin-skill

    Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.

    6.2k GitHub starsUsed in 1 repo~4.7k tokens
    Agent WorkflowsAuto-check passed
  • Skill Release Gate

    rohitg00/ai-engineering-from-scratch

    Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability.

    65k GitHub stars~1k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • CodeGraph Agent Eval

    colbymchenry/codegraph

    Benchmarks how much CodeGraph helps a coding agent on a real repository, comparing runs with and without it for a chosen local or published version.

    73k GitHub stars~950 tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.

    23k GitHub stars~3.4k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed

More from haoruilee/awesome-agent-native-services

  • Add To Awesome List

    haoruilee/awesome-agent-native-services

    Guide a contributor through the full process of adding a new service to the catalog: selecting the standard or operator-surface admission track, checking URL Onboarding, opening an issue, writing…

    476 GitHub stars~2.8k tokensUpdated 2 days ago
    Auto-check: notes
  • Find Agent Service

    haoruilee/awesome-agent-native-services

    Given a task an AI agent needs to perform, find the right agent-native service from the awesome-agent-native-services catalog.

    476 GitHub stars~2.3k tokensUpdated 2 days ago
    Auto-check passed
  • Install Agent Service

    haoruilee/awesome-agent-native-services

    Use the Awesome Agent-Native Services catalog as an installation entry point.

    476 GitHub stars~1.7k tokensUpdated 2 days ago
    Auto-check: notes

Categories

Questions about Evaluate Agent Native

What does Evaluate Agent Native do?

Evaluate catalog candidates against either the standard five agent-native criteria or the narrow operator-surface track, and check URL Onboarding. Evaluate Agent Native is an agent skill from haoruilee/awesome-agent-native-services. Evaluate catalog candidates against either the standard five agent-native criteria or the narrow operator-surface track, and check URL Onboarding.

When should I use Evaluate Agent Native?

Evaluate Agent Native fits situations like: asked whether a service; control surface belongs in the catalog.

How do I install Evaluate Agent Native in Claude Code?

Run `npx skills add haoruilee/awesome-agent-native-services --skill evaluate-agent-native -a claude-code`. Or copy the skill folder (.skills/evaluate-agent-native in haoruilee/awesome-agent-native-services) into .claude/skills/evaluate-agent-native in your project. Claude Code loads it when a task matches its description.

How do I install Evaluate Agent Native in Codex?

Run `npx skills add haoruilee/awesome-agent-native-services --skill evaluate-agent-native -a codex`. Or copy the skill folder (.skills/evaluate-agent-native in haoruilee/awesome-agent-native-services) into .agents/skills/evaluate-agent-native in your project. Codex loads it when a task matches its description.

Can I use Evaluate Agent Native in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add haoruilee/awesome-agent-native-services --skill evaluate-agent-native -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluate-agent-native, .gemini/skills/evaluate-agent-native, .github/skills/evaluate-agent-native and .opencode/skills/evaluate-agent-native in your project.

What does Evaluate Agent Native need to run?

Going by SKILL.md and its folder, Evaluate Agent Native needs the command-line tools its instructions call (npx). Our summary lists: Node.js. Its frontmatter pre-approves these tools: WebSearch, Read. Compatibility (from SKILL.md): Works with agents that can inspect official websites, repositories, and protocol documentation..

Does Evaluate Agent Native access the network?

SKILL.md names 4 domains. In commands or code: moltbook.com, ensue.dev, api.ensue-network.ai and raw.githubusercontent.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Evaluate Agent Native safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evaluate Agent Native use?

Evaluate Agent Native is published under the CC0-1.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluate Agent Native use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evaluate Agent Native?

Skills that share tags, products or a category with Evaluate Agent Native: MCP Server Builder (anthropics/skills, 180k stars), Diagnosing Superpowers Sessions (obra/superpowers, 296k stars), Darwin Skill Optimizer (alchaincyf/darwin-skill, 6.2k stars) and Skill Release Gate (rohitg00/ai-engineering-from-scratch, 65k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluate Agent Native?

haoruilee (a GitHub user) maintains it in haoruilee/awesome-agent-native-services, which has 476 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 5, 2026.

Source: haoruilee/awesome-agent-native-services on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.