Agent skill

Validate Model

by simstudioai in simstudioai/sim

Validate a model entry (or every model in a provider) in apps/sim/providers/models.ts against the provider's live API docs (no hallucination — reports what cannot be verified)

Apache-2.0Auto-check passedDevelopment

Install Validate Model

skills CLI
$ npx skills add simstudioai/sim --skill validate-model -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install simstudioai/sim validate-model --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/simstudioai/sim.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/validate-model .claude/skills/validate-model && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
validate-model
GitHub stars
30k
Token cost
~2.5k tokens
SKILL.md length
1,172 words
Files
2
Skills in repo
40
Repo updated
First seen
Licence
Apache-2.0

At a glance

Validate a model entry (or every model in a provider) in apps/sim/providers/models.ts against the provider's live API docs (no hallucination — reports what cannot be verified)

  • Works in 6 steps: Read entries from models.ts → Live-fetch authoritative sources → Build the consumption map for this… → …
  • Tasks that involve Technical documentation
  • SKILL.md covers Hard rules (do not skip), Your Task, Step 1: Read entries from… and Step 2: Live-fetch…, plus 7 more sections
  • Calls bun; reaches docs.x.ai and cloudprice.net

What it does

Validate Model is an agent skill from simstudioai/sim. Validate a model entry (or every model in a provider) in apps/sim/providers/models.ts against the provider's live API docs (no hallucination — reports what cannot be verified)

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in Development, covering Technical documentation. The repository describes itself as: Sim is the collaborative workspace to build, deploy, and monitor AI agents and workflows. Used by 100,000+ builders. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Technical documentation

Example prompts

  • “/validate-model”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Read entries from models.ts
  2. Live-fetch authoritative sources
  3. Build the consumption map for this provider
  4. Run the checklist
  5. Report (mandatory format)
  6. Offer to fix

What it can do on your machine

Read from SKILL.md and the folder at commit 546d4e7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bun

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • docs.x.ai
    • cloudprice.net
    • openrouter.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Validate Model loads about 2.5k tokens when it runs. Until then it costs about 48 tokens; SKILL.md has 1,172 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~48
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from simstudioai/sim at commit 546d4e7, republished under its Apache-2.0 licence (© simstudioai). 1,172 words, ~2,511 tokens.

Download SKILL.mdSave it as .claude/skills/validate-model/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
validate-model
description
Validate a model entry (or every model in a provider) in apps/sim/providers/models.ts against the provider's live API docs (no hallucination — reports what cannot be verified)
argument-hint
<provider> [model-id]

Validate Model Skill

You audit one or more model entries in apps/sim/providers/models.ts against the provider's official live API docs. Hallucinated pricing and capabilities are the #1 failure mode in this file. Every numeric and capability claim must be re-derived from a live web fetch in this session — not from memory, not from training data, not from the user's marketing email.

Hard rules (do not skip)

  1. Live-fetch or report unverified. Each field must be backed by a live WebFetch in this session. If you cannot reach an authoritative URL for a field, mark it UNVERIFIED in the report — do not silently confirm it from memory.
  2. Cite every fact. Every value in the report must show the source URL it was checked against. No URL → mark UNVERIFIED.
  3. Two-source rule for pricing. Cross-check input/output/cached against at least one secondary source (OpenRouter, Artificial Analysis, CloudPrice). If sources disagree, the provider's own docs win — flag the disagreement.
  4. Inspect provider implementation before flagging capability mismatches. A capability flag in models.ts is dead unless the provider's code under apps/sim/providers/{provider}/ consumes it (see Consumption Matrix below). Setting a flag the provider ignores is a warning, not a critical.
  5. Never auto-fix without printing the diff. Show the user the proposed diff before applying. Get confirmation.

Your Task

When invoked as /validate-model <provider> [model-id]:

  1. Read the target entries from models.ts
  2. Live-fetch the provider's official models, pricing, and capability/reasoning pages + at least one secondary source for pricing
  3. Inspect the provider implementation to know which flags are actually consumed
  4. Run the checklist below per model
  5. Report findings (critical / warning / suggestion / unverified) with every cell linked to its source URL
  6. Offer to fix; on confirm, edit models.ts in a single pass and re-lint

If model-id is omitted, validate every model in the provider.

Step 1: Read entries from models.ts

Capture per model: id, full pricing, full capabilities, contextWindow, releaseDate, recommended, speedOptimized, deprecated.

Step 2: Live-fetch authoritative sources

Use the canonical provider URL table in the add-model skill (.agents/skills/add-model/SKILL.md), Step 1, as the single source of truth — fetch the models index, pricing, and reasoning/parameter caveats pages listed there for the target provider.

Secondary cross-check (use at least one): OpenRouter, Artificial Analysis, CloudPrice.

If a fetch fails (404, timeout, paywall), record the URL attempted and mark dependent fields UNVERIFIED.

Step 3: Build the consumption map for this provider

Use the Consumption Matrix in .agents/skills/add-model/SKILL.md Step 2 and run its re-grep commands for the target provider before relying on it. A flag set in models.ts that the provider's code does not read = warning: dead flag.

Step 4: Run the checklist

For each model, evaluate every row. Statuses: ✓ matches docs, ✗ disagrees, ⚠️ single-source, ❓ UNVERIFIED (could not fetch).

Identity
  • id exactly matches provider's API model identifier (case, dots, dashes, prefix for resellers)
  • releaseDate matches launch announcement
  • deprecated: true set if provider has announced retirement (or removed from active list)
Pricing (per 1M tokens, USD)
  • pricing.input matches provider pricing page
  • pricing.output matches provider pricing page
  • pricing.cachedInput matches provider's documented cached/prompt-cache rate (or is correctly omitted if no caching offered)
  • pricing.updatedAt is recent — warn if older than 60 days
Context & output limits
  • contextWindow matches docs (in tokens)
  • capabilities.maxOutputTokens matches documented output cap (or is correctly omitted if "no output limit")
Capabilities (each must be DOCUMENTED-AS-SUPPORTED and CONSUMED-BY-PROVIDER-CODE)
  • temperature — provider accepts it for this model (reasoning-always-on models often reject)
  • reasoningEffort.values — list matches docs; omitted for always-reasoning models that reject the parameter (e.g., grok-4.3, where xAI docs explicitly state reasoning_effort is not supported). Verify per model — some always-reasoning models (e.g., OpenAI's o-series) DO accept reasoning_effort and should keep the flag.
  • verbosity.values — only on OpenAI gpt-5.x family; values match docs
  • thinking.levels + thinking.default — only on providers whose code reads thinking/thinkingLevel (see the add-model Consumption Matrix); values match docs
  • thinking.streamed — REQUIRED on Anthropic-family thinking models. Current Claude thinking models use 'summary' (Sim requests display: 'summarized'); use 'full' only if official API docs explicitly guarantee raw thinking deltas. Verify against the provider's thinking-display docs. After any change, run bun run agent-stream-docs:generate so the Agent block docs table stays in sync (CI diffs it)
  • nativeStructuredOutputs — only on providers whose code consumes it (see the Consumption Matrix); provider must document Structured Outputs / JSON-mode for this model
  • toolUsageControl — provider supports tool_choice semantics
  • computerUse — provider implements computer-use loop AND model is a computer-use SKU
  • deepResearch — only on actual deep-research SKUs
  • memory: false — only when the model genuinely cannot maintain conversation history
Show full SKILL.md (441 more words)Show less
Flags
  • recommended: true — at most one or two per provider; should be current flagship
  • speedOptimized: true — only on smallest/fastest tier (nano / flash-lite / haiku class)
Hosting / billing
  • If getHostedModels() includes the model ID (providers/models.ts expands whole providers — more than openai/anthropic/google — plus the static Fireworks catalog), the model is served with Sim's rotating key and billed via shouldBillModelUsage(). Confirm that is the intent (a BYOK-only model parked under a hosted provider is a billing bug — warning).
  • If the model is hosted, the deployment is expected to have its {PREFIX}_COUNT / {PREFIX}_1..N env vars set (ops concern; note if it looks unset for a model claiming hosted support).

Step 5: Report (mandatory format)

For each model, emit a table with one row per checklist item. Every row that claims ✓ must have a URL.

markdown
### Validation — <model-id>

| Field | Repo | Live docs | Source URL | Status |
|---|---|---|---|---|
| `input` | $1.25/M | $1.25/M | https://docs.x.ai/... | ✓ |
| `cachedInput` | $0.50/M | $0.20/M | https://cloudprice.net/... | ✗ stale (price cut not picked up) |
| `reasoningEffort` | low/medium/high | rejected by API | https://docs.x.ai/.../reasoning | ✗ inert — selecting silently no-ops |
| `contextWindow` | 1,000,000 | 1,000,000 | https://docs.x.ai/... + https://openrouter.ai/... | ✓ (2 sources) |
| `releaseDate` | 2026-04-30 | not found in scraped pages | _attempted: docs.x.ai, x.ai/news_ | ❓ UNVERIFIED |

**Findings**
- 🔴 critical — `cachedInput` is wrong: docs say $0.20/M, repo has $0.50/M
- 🟡 warning — `reasoningEffort` is set but provider rejects it for this model (xAI docs explicitly: "reasoning_effort is not supported by grok-4.3")
- 🔵 suggestion — `pricing.updatedAt` is 90 days old; refresh
- ❓ unverified — `releaseDate` could not be confirmed from any fetched page; ask user

**Disagreements between sources**
- _none_ OR _OpenRouter says $X, provider docs say $Y — went with provider docs_

End each multi-model run with a summary count: N models checked · X critical · Y warnings · Z suggestions · W unverified.

Step 6: Offer to fix

After reporting, ask: "Want me to fix the critical and warning items? I'll print the diff first." On yes:

  1. Print the proposed diff (do not apply yet)
  2. Get user confirmation
  3. Edit models.ts in a single pass
  4. Run bun run lint
  5. Re-run only the failed rows of the checklist on the new state

Severity definitions

  • 🔴 critical — wrong number or wrong identifier that misleads users about cost or breaks API calls. Examples: incorrect pricing, wrong model id, wrong context window, capability the API rejects.
  • 🟡 warning — dead code or internal inconsistency. Examples: capability flag the provider ignores, multiple recommended: true per provider, pricing.updatedAt >60 days old, missing deprecated: true on retired model.
  • 🔵 suggestion — style/consistency. Examples: field order, missing speedOptimized on a clearly smallest-tier model.
  • ❓ unverified — could not fetch an authoritative source for this field. Surface it; never silently confirm.

Common drift

Pricing changes after provider price cuts; reasoningEffort/thinking/verbosity set on a model whose provider code or API does not accept them; stale pricing.updatedAt; wrong context window; more than one recommended after a flagship swap; missing deprecated: true after a provider retirement announcement.

What "I cannot verify this" looks like

If, after fetching the documented sources, a field cannot be confirmed:

  • Mark the row ❓ UNVERIFIED with the URL(s) attempted
  • Surface it in the Findings section with severity ❓
  • Do NOT mark the validation as passed
  • Ask the user for a docs URL or guidance before changing anything

The skill is allowed to say "I could not verify the cached input price for grok-4.3 from the official xAI docs in this session — I attempted [URLs] without finding the value. Third-party sources [URL1, URL2] both report $0.20/M. Confirm before I update." That is correct behavior. Hallucinating a number is not.

© simstudioai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .agents/skills/validate-model of simstudioai/sim.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 546d4e7

Compare with similar skills

Validate Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Validate Model compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Validate Model this skillsimstudioai/sim30k—~2.5kAutomated safety check: PassApache-2.0
Diagram Designcathrynlavery/diagram-design44k1 repos~7.5kAutomated safety check: PassMIT
Simple Englishmoeru-ai/airi50k2 repos~4.6kAutomated safety check: PassMIT
Get API Docs with chubandrewyng/context-hub14k2 repos~775Automated safety check: PassMIT
Doc SyncJetBrains/ideavim10k2 repos~2.6kAutomated safety check: PassMIT
Mailspring App ScreenshotsFoundry376/Mailspring18k—~1.4kAutomated safety check: PassGPL-3.0

Similar skills

  • Diagram Design

    cathrynlavery/diagram-design

    Creates branded diagrams, from architecture, flowchart and sequence to charts and maps, as self-contained HTML with inline SVG, with import from draw.io, Mermaid and Excalidraw.

    44k GitHub starsUsed in 1 repo~7.5k tokens
    DevelopmentAuto-check passed
  • Simple English

    moeru-ai/airi

    Write or rewrite technical text with the rules of ASD-STE100 Simplified Technical English so it is clear, unambiguous, and free of AI slop.

    50k GitHub starsUsed in 2 repos~4.6k tokens
    DevelopmentAuto-check passed
  • Get API Docs with chub

    andrewyng/context-hub

    Fetches current documentation for third-party APIs and SDKs with the chub CLI before the agent writes code against them, instead of relying on remembered API shapes.

    14k GitHub starsUsed in 2 repos~775 tokens
    DevelopmentAuto-check passed
  • Doc Sync

    JetBrains/ideavim

    Official

    Keeps IdeaVim documentation in sync with code changes. An agent skill from JetBrains/ideavim.

    10k GitHub starsUsed in 2 repos~2.6k tokens
    DevelopmentAuto-check passed
  • Mailspring App Screenshots

    Foundry376/Mailspring

    Captures screenshots of the running Mailspring dev app for docs, PRs or visual checks by launching it with a debugging port, driving the UI and clipping to an element.

    18k GitHub stars~1.4k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Draw.io Diagram Studio

    Agents365-ai/drawio-skill

    Creates and edits editable draw.io diagrams from descriptions, code, infrastructure files, SQL and API schemas, with sync, review, test and export tools.

    10k GitHub stars~2.4k tokensUpdated 5 days ago
    DevelopmentAuto-check: notes

More from simstudioai/sim

All 40 skills in this repo
  • Sim Helm

    simstudioai/sim

    Install, upgrade, and operate the Sim Helm chart on Kubernetes.

    30k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Add Column Type

    simstudioai/sim

    Add a new table column type to Sim — registry entry, icon, storage shape, coercion, and the behavioral hooks the grid and API read.

    30k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Add Enrichment

    simstudioai/sim

    Add a code-defined table enrichment (registry entry) under apps/sim/enrichments/ backed by an ordered provider cascade, ensuring every provider tool it calls has hosted-key support.

    30k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Add Hosted Key

    simstudioai/sim

    Add hosted API key support to a tool so Sim provides the key (metered and billed to the workspace) when a user has not brought their own.

    30k GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Add Managed CLI

    simstudioai/sim

    Add or upgrade a curated, immutable managed CLI for Sim Function sandboxes, including client-safe catalog metadata, a pinned server-only installation recipe, checksum and executable verification…

    30k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Add Selector

    simstudioai/sim

    Add or update a Sim dynamic selector using the shared manifest, server attachment, and selectors.execute path.

    30k GitHub stars~1.7k tokensUpdated today
    Auto-check passed

Categories

Questions about Validate Model

What does Validate Model do?

Validate a model entry (or every model in a provider) in apps/sim/providers/models.ts against the provider's live API docs (no hallucination — reports what cannot be verified). Validate Model is an agent skill from simstudioai/sim.

When should I use Validate Model?

Validate Model fits situations like: tasks that involve Technical documentation.

How do I install Validate Model in Claude Code?

Run `npx skills add simstudioai/sim --skill validate-model -a claude-code`. Or copy the skill folder (.agents/skills/validate-model in simstudioai/sim) into .claude/skills/validate-model in your project. Claude Code loads it when a task matches its description.

How do I install Validate Model in Codex?

Run `npx skills add simstudioai/sim --skill validate-model -a codex`. Or copy the skill folder (.agents/skills/validate-model in simstudioai/sim) into .agents/skills/validate-model in your project. Codex loads it when a task matches its description.

Can I use Validate Model in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add simstudioai/sim --skill validate-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/validate-model, .gemini/skills/validate-model, .github/skills/validate-model and .opencode/skills/validate-model in your project.

What does Validate Model need to run?

Going by SKILL.md and its folder, Validate Model needs the command-line tools its instructions call (bun).

Does Validate Model access the network?

SKILL.md names 3 domains. In commands or code: docs.x.ai, cloudprice.net and openrouter.ai; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Validate Model safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Validate Model use?

Validate Model is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Validate Model use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Validate Model?

Skills that share tags, products or a category with Validate Model: Diagram Design (cathrynlavery/diagram-design, 44k stars), Simple English (moeru-ai/airi, 50k stars), Get API Docs with chub (andrewyng/context-hub, 14k stars) and Doc Sync (JetBrains/ideavim, 10k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Validate Model?

simstudioai (a GitHub organization) maintains it in simstudioai/sim, which has 29,785 GitHub stars. The repository holds 40 skills in this directory. The repository was last updated on October 7, 2026.

Source: simstudioai/sim on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.