Agent skill

Conventions MCP

by stella in stella/stella

Apply when adding or changing an MCP tool, a capability the CLI generates, a tool input or output schema, a tool description, an error envelope, or an agent-facing reference resource.

Apache-2.0Auto-check passedAgent Workflows

Install Conventions MCP

skills CLI
$ npx skills add stella/stella --skill conventions-mcp -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install stella/stella conventions-mcp --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/stella/stella.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/conventions-mcp .claude/skills/conventions-mcp && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
conventions-mcp
GitHub stars
258
Token cost
~2.7k tokens
SKILL.md length
1,567 words
Files
1
Skills in repo
24
Repo updated
First seen
Licence
Apache-2.0

At a glance

Apply when adding or changing an MCP tool, a capability the CLI generates, a tool input or output schema, a tool description, an error envelope, or an agent-facing reference resource.

  • Works in 3 steps: An eval is the acceptance test. Every… → Keep the rejected call, in the eval… → Parse through the schema the handler…
  • Tasks that involve MCP servers
  • SKILL.md covers Target Property, Measure Before Believing, Shape Tools For How Models… and Every Error Is A Next Step, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Conventions MCP is an agent skill from stella/stella. Apply when adding or changing an MCP tool, a capability the CLI generates, a tool input or output schema, a tool description, an error envelope, or an agent-facing reference resource. Enforces contracts that language models can actually drive, measured by evals rather than by schema soundness.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering MCP servers and LLM evaluation. It works with Model Context Protocol. The repository describes itself as: Open-source legal workspace. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve MCP servers
  • Tasks that involve LLM evaluation

Example prompts

  • “/conventions-mcp”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. An eval is the acceptance test. Every agent-facing workflow has an eval that
  2. Keep the rejected call, in the eval only. The eval records the raw input
  3. Parse through the schema the handler uses. The eval, the tests, and the

What it can do on your machine

Read from SKILL.md and the folder at commit b225fd8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Conventions MCP loads about 2.7k tokens when it runs. Until then it costs about 78 tokens; SKILL.md has 1,567 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from stella/stella at commit b225fd8, republished under its Apache-2.0 licence (© stella). 1,567 words, ~2,733 tokens.

Download SKILL.mdSave it as .claude/skills/conventions-mcp/SKILL.md (or your agent's skills folder).
name
conventions-mcp
description
Apply when adding or changing an MCP tool, a capability the CLI generates, a tool input or output schema, a tool description, an error envelope, or an agent-facing reference resource. Enforces contracts that language models can actually drive, measured by evals rather than by schema soundness.

Agent-Facing Contract Conventions

An MCP tool or CLI capability is a user interface whose user is a language model. A schema can be type-sound, validated, and documented and still be undrivable: models copy examples, fill every property they see, retry with the same call, and cannot inspect bytes. Design for that behaviour and measure it.

Target Property

A capable model, given only what tools/list and the reference resources expose, completes the workflow on the first or second attempt, and every rejection it receives names the next call to make. The authoring eval, not code review, decides whether a contract change helped.

Measure Before Believing

  1. An eval is the acceptance test. Every agent-facing workflow has an eval that drives the real tools with the real schemas and scores each step separately (authored, created, configured, filled, or the workflow's equivalents). Run it before and after a contract change and put both tables in the PR. A change that is "obviously better" without numbers is a hypothesis.
  2. Keep the rejected call, in the eval only. The eval records the raw input of every call the schema rejected before the handler ran; a pass rate without the rejected payloads cannot say whether the model or the contract failed. Eval fixtures are synthetic, so the trace may hold them whole. Production telemetry for a rejected call records the tool name, the issue codes, and the issue paths, never the payload: tool arguments carry matter content and personal data.
  3. Parse through the schema the handler uses. The eval, the tests, and the handler must validate the same object; a normalisation that lives only in one handler's pre-step is invisible to the eval and drifts.

Shape Tools For How Models Behave

  1. One tool, one intent. Split create, configure, and update into separate tools rather than one tool whose meaning depends on which optional arguments are present. Cross-field "provide A or B, not both" checks are a sign the tool has more than one job.
  2. Few optionals, one discriminator. Replace a set of mutually exclusive optional keys with one discriminated union (source: { type: "ai" | "lookup" | ... }). A model that fills every property cannot produce a contradictory union member; it can produce six contradictory optionals.
  3. Best effort over all-or-nothing for collections. When a call carries a list of entries, validate each entry independently, apply the valid ones, and return the rest as per-entry issues[]. One bad property must not sink a call that carried seven good entries.
  4. Echo the next call. Return, from the call that discovers state, the exact payload the next call accepts (a configure skeleton, canonical path spelling, allowed values). Copying beats inferring. Read-back must round-trip: the shape a describe tool returns is the shape the configure tool accepts, byte for byte, and a test pins that fixed point.
  5. Idempotent and re-entrant. Models retry with the same call and resend everything they know. Accept the resend: upsert on a stable id, idempotency keys that cover every argument, durable receipts replayed without re-executing.
  6. Keep artefacts out of the model. Bytes are the hardest step in any loop. Prefer references (a host file reference, an id of something already stored) over inline base64; when inline is the only path, bound it by the request frame, derive the limit from the shared constant, and state it in the description. Never let an error hint invite the model to shrink or rewrite a document.

Every Error Is A Next Step

  1. Structured envelope, closed codes. Failures return { code, message, hint, issues[] } with codes from a closed set. hint names the corrective action, the tool to call, and where to go (a deep link when the fix is in the UI). "Disabled" without "enable it here" is a dead end.
  2. Never leak the substrate. A malformed id is a validation issue at the boundary, never a database cast error reported as internal_error. Every id input is declared with the shared id schema; a registry-wide test enforces it.
  3. Warn about what the model probably meant. Discovery and save return warnings[] with closed codes for the known traps (unprefixed loop item, unknown directive, split marker, and so on). A census test keeps the code list and the reference in step.
  4. Read each value kind leniently, in one place. A model spells a value the way its training data did: null for an unset optional, 4 000 for a number, 1. 10. 2026 for a date, ano for a boolean, cs_CZ for a locale. Every kind the wire accepts gets ONE owner that auto-normalizes the spellings carrying a single meaning and returns ONE ask-for-a-fix shape (received, expected, hint) when a spelling carries two: 01/02/2026 and a bare 1,234 are asked about with both readings named, never guessed, because guessing wrong is a wrong date or a factor of a thousand on an instrument. Null and the placeholder encodings are that rule for "absent" and live in the tool factory; the value kinds live in packages/agent-input/: number (a page size clamps with a note), boolean, date (a range bound reads 2020 as its first or last day and 0001-01-01/9999-12-31 as no bound), uuid (the all-zero and example ids are placeholders), filter (" ", "all", "-" mean "not filtering"), vocabulary (a data-owned value such as a court, read through case, diacritics, abbreviations and English names), string list (a lone string is a one-item list; only constrained tokens split on commas), ELI, country, locale, enum. A repaired value returns a note ("Read X as Y") that reaches the caller beside the result. A placeholder in an optional property is absence on a read and an ask on a write, where an optional id can switch update into create. An optional filter read against a data-owned vocabulary (court) never empties a search: a value that names no stored value, or several, is dropped with a warning naming the stored ones. A filter with no such vocabulary yet reads only its placeholders, and an empty page under it carries no_hits_filtered naming the values that would have matched; do not advertise resolution a filter does not do. An opaque token the server issues extends "absent" one step: a client filling every declared property invents a first-call cursor (" ", "0", "start"), so cursorInput reads a value outside the class its encoders emit as no cursor, while a value inside that class still reaches the decoder and fails through invalidCursorResult, which names the restart; restarting a damaged or invented cursor at page one silently would repeat a page the caller already read. Kinds a property's name implies (limit, date_from, date_to) are bound by the factory, so a new tool inherits them. Cover each kind with a property test over its whole spelling class rather than the examples someone happened to write down, and add a guard (an ownership row, a census test) so a new call site cannot parse the kind itself. Never per-tool tolerance code, and never a second reader: two lenient readers of one kind are worse than one strict one, because they disagree.
Show full SKILL.md (396 more words)Show less

References Are Read Verbatim

  1. Every sentence must be true of the code. Reference resources and tool descriptions are copied into model context as the complete contract. A promised behaviour the code does not have (a preserved extension, a size limit that is not the enforced one) is a bug, not a docs nit. Render numbers from the constants; type tool names against the registry; keep drift tests.
  2. One workflow resource per workflow. Publish the ordered steps (create, read back, configure, preview, persist) as a resource pointed to from the MCP instructions, and keep the quiz or eval that checks a model understood it.
  3. Decide grammar leniency by evidence. When two independent models write the same form the grammar rejects, that is a product signal: either accept the form or make the reference unambiguous, then re-measure. Do not do both at once.

Name Capabilities By Resource

A capability id is its handler path under apps/api/src/handlers/: <domain>[.<resource>…].<action>. The action is ONE word, either a canonical verb (list, get, create, update, delete) or an entry in the closed DOMAIN_ACTION_VERBS list (apps/api/scripts/lib/capability-catalog.ts; docs/capability-coverage.md renders the current list). A compound action is a nested resource directory: clauses/categories/create.ts, not clauses/categories-create.ts; entities/versions/list.ts, not entities/read-versions.ts. A new domain verb is a reviewed addition to that list; the exporter fails on an unlisted verb and on a listed verb no capability uses, and the capability-domain-action-verbs ratchet holds hyphenated entries at 0.

Guards Make It Stick

  • Registry-wide tests over every tool definition: id schemas, null-optional tolerance, description and reference drift, and a zero-diff capability export.
  • Total companion maps keyed by tool name (as const satisfies Record<ToolName, ...>) for policy, consent, projection, and CLI disposition, so a new tool cannot land without each decision.
  • Compile-time gates in the tool factory (a schema that is not wrapped, an id that is not the shared schema) rather than review discipline.
  • The CLI is generated from the same catalog: a capability whose transport the generator cannot invoke is not advertised anywhere, and generated artifacts are regenerated in order (capability export, then CLI codegen) until the diff is zero.

Existing References

  • Tool factory and boundary normalisation: apps/api/src/mcp/valibot-tool-definition.ts, apps/api/src/mcp/tool-utils.ts
  • One reader per value kind, and its census: packages/agent-input/src/, apps/api/src/lib/agent-input-owner.test.ts
  • Registry-wide guards: apps/api/src/mcp/uuid-id-inputs.test.ts, apps/api/src/mcp/null-optional-inputs.test.ts
  • Warnings census and references: apps/api/src/lib/docx/template-warnings.ts, apps/api/src/mcp/template-workflow-reference.ts
  • Authoring eval: apps/api/evals/template-authoring.ts

These are examples of the mechanisms, not proof that a new tool meets the bar: run the eval.

© stella, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/conventions-mcp of stella/stella.

Open the folder on GitHubat commit b225fd8

Compare with similar skills

Conventions MCP next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Conventions MCP compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Conventions MCP this skillstella/stella258—~2.7kAutomated safety check: PassApache-2.0
Improving MCP ToolsPostHog/posthog40k—~1.5kAutomated safety check: PassCustom licence
Opik Online Evalcomet-ml/opik-mcp220—~3kAutomated safety check: NotesApache-2.0
Skill Creatorcuriositech/some_claude_skills243—~7.2kAutomated safety check: PassApache-2.0
Agent Eval Casesagentailor/fullstack-langgraph-nextjs-agent132—~5.3kAutomated safety check: PassMIT
Managed Deep Agentslangchain-ai/langchain-skills1.3k—~8.7kAutomated safety check: NotesMIT

Similar skills

  • Improving MCP Tools

    PostHog/posthog

    Official

    Run an improve-my-MCP campaign: an autoresearch-style loop that measures the MCP agent experience with the eval harness, picks the highest-impact tool problem from production data, makes one bounded…

    40k GitHub stars~1.5k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Opik Online Eval

    comet-ml/opik-mcp

    Take a judge live on production traffic — create an Opik online evaluation rule (LLM-as-judge or Python metric) on a project with sampling, filters, variable mapping, and a cost cap, then confirm…

    220 GitHub stars~3k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Skill Creator

    curiositech/some_claude_skills

    A skill your agent uses when creating a new Claude skill from scratch, editing or improving an existing skill, or measuring skill performance with evals and benchmarks.

    243 GitHub stars~7.2k tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed
  • Agent Eval Cases

    agentailor/fullstack-langgraph-nextjs-agent

    Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic.

    132 GitHub stars~5.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Managed Deep Agents

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when building, testing, or deploying Managed Deep Agents in LangSmith.

    1.3k GitHub stars~8.7k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Agent Observability Eval Bootstrap

    datadog-labs/agent-skills

    Bootstrap evaluators from production traces — by default propose online LLM-judge evaluators and, after you confirm, create them in Datadog as disabled drafts (never auto-enabled); on request emit…

    177 GitHub stars~25k tokensUpdated today
    DevOps & CloudAuto-check passed

More from stella/stella

All 24 skills in this repo
  • Plan

    stella/stella

    Create a concise, evidence-backed implementation plan in the repository planning area when the user explicitly asks for a plan.

    258 GitHub stars~917 tokensUpdated today
    Auto-check passed
  • Answer From Sources

    stella/stella

    Answers data-protection (GDPR) questions grounded in the regulation and supervisory guidance, with a citation for every claim.

    258 GitHub stars~735 tokensUpdated today
    Auto-check passed
  • Check Against Rules

    stella/stella

    Reviews a non-disclosure agreement against the firm's NDA checklist and reports findings with citations.

    258 GitHub stars~856 tokensUpdated today
    Auto-check passed
  • Intake To Draft

    stella/stella

    Collects the facts of an unpaid invoice, then drafts a payment demand letter.

    258 GitHub stars~537 tokensUpdated today
    Auto-check passed
  • Conventions Perf

    stella/stella

    Apply when a performance-guard check (network baseline, bundle baseline, DB query count, loader-prefetch lint, RC bailouts) fails or when touching a hot route/endpoint.

    258 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Apply when writing or reviewing React effects in apps/web. An agent skill from stella/stella.

    258 GitHub stars~2.8k tokensUpdated today
    Auto-check passed

Questions about Conventions MCP

What does Conventions MCP do?

Apply when adding or changing an MCP tool, a capability the CLI generates, a tool input or output schema, a tool description, an error envelope, or an agent-facing reference resource. Conventions MCP is an agent skill from stella/stella. Apply when adding or changing an MCP tool, a capability the CLI generates, a tool input or output schema, a tool description, an error envelope, or an agent-facing reference resource.

When should I use Conventions MCP?

Conventions MCP fits situations like: tasks that involve MCP servers; tasks that involve LLM evaluation.

How do I install Conventions MCP in Claude Code?

Run `npx skills add stella/stella --skill conventions-mcp -a claude-code`. Or copy the skill folder (.agents/skills/conventions-mcp in stella/stella) into .claude/skills/conventions-mcp in your project. Claude Code loads it when a task matches its description.

How do I install Conventions MCP in Codex?

Run `npx skills add stella/stella --skill conventions-mcp -a codex`. Or copy the skill folder (.agents/skills/conventions-mcp in stella/stella) into .agents/skills/conventions-mcp in your project. Codex loads it when a task matches its description.

Can I use Conventions MCP in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add stella/stella --skill conventions-mcp -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/conventions-mcp, .gemini/skills/conventions-mcp, .github/skills/conventions-mcp and .opencode/skills/conventions-mcp in your project.

What does Conventions MCP need to run?

SKILL.md names no scripts, command-line tools or credentials: Conventions MCP is instructions for the agent only.

Does Conventions MCP access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Conventions MCP safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Conventions MCP use?

Conventions MCP is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Conventions MCP use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Conventions MCP?

Skills that share tags, products or a category with Conventions MCP: Improving MCP Tools (PostHog/posthog, 40k stars), Opik Online Eval (comet-ml/opik-mcp, 220 stars), Skill Creator (curiositech/some_claude_skills, 243 stars) and Agent Eval Cases (agentailor/fullstack-langgraph-nextjs-agent, 132 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Conventions MCP?

stella (a GitHub organization) maintains it in stella/stella, which has 258 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 9, 2026.

Source: stella/stella on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.