Agent skill

Cost Watchdog

by evloghq in evloghq/evlog

Weekly review of evlog's model cost and performance. An agent skill from evloghq/evlog.

MITAuto-check passedFrontend & Design

Install Cost Watchdog

skills CLI
$ npx skills add evloghq/evlog --skill cost-watchdog -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install evloghq/evlog cost-watchdog --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/evloghq/evlog.git skills-src && mkdir -p .claude/skills && cp -r skills-src/apps/evi/agent/skills/cost-watchdog .claude/skills/cost-watchdog && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cost-watchdog
GitHub stars
1.9k
Token cost
~2.2k tokens
SKILL.md length
1,304 words
Files
1
Skills in repo
20
Repo updated
First seen
Licence
MIT

At a glance

Weekly review of evlog's model cost and performance. An agent skill from evloghq/evlog.

  • Works in 6 steps: Define the window → Pull the numbers → Build the candidate set → …
  • Tasks that involve Static sites and blogs
  • SKILL.md covers What the report gives you, Resources, Steps and Deliver, plus 2 more sections
  • Reaches vercel.com and artificialanalysis.ai

What it does

Cost Watchdog is an agent skill from evloghq/evlog. Weekly review of evlog's model cost and performance. Load this when the cost-watchdog schedule fires, or when Hugo asks for a cost check, a model review, a per-surface model analysis, or a spend/drift report for the gateway.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Frontend & Design, covering Static sites and blogs and Journaling and reflection. The repository describes itself as: Digging through logs is not observability. It's hope — wide events, structured errors, TypeScript-first, every runtime. The licence is MIT.

When your agent uses it

  • Tasks that involve Static sites and blogs
  • Tasks that involve Journaling and reflection

Example prompts

  • “/cost-watchdog”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Define the window
  2. Pull the numbers
  3. Build the candidate set
  4. Compare candidates
  5. Flag drift
  6. Propose per-surface model adjustments

What it can do on your machine

Read from SKILL.md and the folder at commit 1b6e1b9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • vercel.com
    • artificialanalysis.ai
    • ai-gateway.vercel.sh
    • arena.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cost Watchdog loads about 2.2k tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 1,304 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from evloghq/evlog at commit 1b6e1b9, republished under its MIT licence (© evloghq). 1,304 words, ~2,246 tokens.

Download SKILL.mdSave it as .claude/skills/cost-watchdog/SKILL.md (or your agent's skills folder).
name
cost-watchdog
description
Weekly review of evlog's model cost and performance. Load this when the cost-watchdog schedule fires, or when Hugo asks for a cost check, a model review, a per-surface model analysis, or a spend/drift report for the gateway.

Cost and model watchdog

A recurring read of how evlog spends its model budget and whether the models in use are still the right ones. Run it weekly, on the last full week. Grounded in the AI Gateway report and the current model landscape, never in your memory of prices.

The core question: for every surface, is the model it runs still a sensible buy? The honest answer is often "yes, no change." A quiet week is a real result.

What the report gives you

ai_gateway__report with groupBy: 'tag' returns one row per tag value scoped to the current environment: the evi:env:* row (the total) plus one evi:surface:* row per surface. Each row carries total_cost, market_cost, input_tokens, output_tokens, cached_input_tokens, reasoning_tokens and request_count. groupBy: 'model' returns one row per model.

The surface list is whatever evi:surface:* rows the report actually returns. Do not assume the set; read it from the data.

Resources

Open these directly instead of searching; they are the stable home for everything the model-landscape step needs.

  • AI Gateway model catalog, as JSON (pricing and capabilities for every model in one fetch): https://ai-gateway.vercel.sh/v1/models
  • AI Gateway models browser (human-readable, filter by provider, pricing, latency, throughput): https://vercel.com/ai-gateway/models
  • AI Gateway docs, models & providers: https://vercel.com/docs/ai-gateway/models-and-providers
  • Model quality leaderboard (ex-LMArena, blind A/B human preference Elo): https://arena.ai/leaderboard
  • Independent cost-efficiency and benchmarks (Intelligence Index, cost per task, tokens per task, time per task, per-benchmark scores): https://artificialanalysis.ai/. Head-to-head pages live at https://artificialanalysis.ai/models/comparisons/<a>-vs-<b>, using AA slugs (glm-5-3-flash, gpt-6-luna-medium).

Use web_search/web_fetch only for what these do not cover, such as a candidate model's fit for a specific surface. Every figure cited still needs a source and a recency, and a benchmark figure comes from Artificial Analysis or the leaderboard directly, never from a blog or aggregator quoting them.

Steps

1. Define the window

Run Monday morning. Cover the last 7 full days ending yesterday, and pull the 7 days before that as the comparison window, so every drift figure is period-over-period.

2. Pull the numbers
  • ai_gateway__report for both windows, groupBy: 'tag'. That is the spend and token picture per surface and the total.
  • ai_gateway__report for both windows, groupBy: 'model'. The per-model mix (today this is usually one model everywhere).
  • For any surface worth a closer look, ai_gateway__report scoped to that surface (tags: ['evi:env:<env>', 'evi:surface:<name>']) with groupBy: 'model' to see what it runs and at what cost.

Use the eval environment tag to keep benchmark and eval traffic out of the production read when the report lets you.

3. Build the candidate set

Derive it from the catalog every run; do not pick alternatives from memory. From /v1/models, keep every model tagged both tool-use and reasoning that takes image input (the base model reads images natively, see docs/vision.md), and whose input and output prices are each within 3x of the current model's. Then add the top three of that set by Artificial Analysis cost per task. List the full candidate set in the report, with the reason each one was dropped.

4. Compare candidates

Score the current model and every surviving candidate on the same four axes, in one table:

  • Agentic quality. Terminal-Bench, AutomationBench, tau-bench and the hallucination rate, the benchmarks closest to what Evi's surfaces do. The headline Intelligence Index is context, not the verdict.
  • Cost per task. Artificial Analysis cost per task and output tokens per task. Two models at the same per-token price can differ tenfold per task, and a model that spends fewer tokens is cheaper even when it is less capable.
  • Speed. Time per task and output tokens per second. This matters most on interactive surfaces (slack, photon), where a person is waiting.
  • Projected cost on Evi's real traffic. Reprice last week's token mix from the report (uncached input, cached input, output, reasoning) at the candidate's catalog rates, including cache-write pricing when the candidate has one. The advertised price is a floor: gatewayRouting sends zeroDataRetention, which can drop the cheapest deployments, so the price Evi pays for the current model comes from a call's provider_metadata.gateway, not from the catalog.

Where quality and cost disagree, also compute cost per solved task (cost per task divided by score) on the benchmark closest to the surface.

Never discard a candidate for being less capable alone. A model that is materially cheaper or faster and trails on quality is a tradeoff to report, not a non-starter.

5. Flag drift

Compare the two windows and call out what moved, with a reason where one is visible:

  • Total or per-surface cost up or down, as a percent and a dollar figure.
  • Model mix change: a model appearing, disappearing, or shifting share.
  • Token shape change (input, output, cached, reasoning) that hints at a behavior or prompt drift, not just volume.
  • A surface whose cost is out of proportion to its request_count.
Show full SKILL.md (515 more words)Show less
6. Propose per-surface model adjustments

For each surface with nontrivial spend, give each candidate one of three verdicts, with the projected weekly cost and speed effect:

  • Swap. It wins on cost or speed and does not trail on agentic quality, or it wins on quality at a cost the spend justifies.
  • Run evals. It wins clearly on cost or speed but trails on quality, or the published numbers disagree. Benchmarks cannot settle this; Evi's own suite can. Recommend a manual run of the evi-evals workflow with the model input set to the candidate's gateway id, then a comparison of cost, latency and pass rate in PostHog (evi_eval_run, broken down by model).
  • Keep. It loses on the axes that matter for that surface. Say which ones.

A candidate listed under Settled decisions below gets its line in the table and the recorded reason, and is not recommended again unless its price or benchmarks have moved since the decision.

One constraint the report does not show: today the agent runs a single model everywhere, set by EVI_MODEL in agent/lib/model.ts (see agent/lib/gateway.ts for tagging). If a per-surface recommendation implies different models per surface, say that routing is currently global and the swap is one of two things: changing the global model, or adding surface-scoped routing as a follow-up decision. Never present a per-surface swap as a one-line config change when routing does not exist yet.

Deliver

The full report is a Linear document on the evlog team, titled Cost/model watchdog: YYYY-MM-DD, with markdown sections: spend and model mix per surface, drift, the candidate set and comparison table with sources, and the per-surface recommendations (or the explicit "nothing to improve").

The thread get two or three lines: the single most attention-worthy number or finding, and the document link.

A material, decision-worthy recommendation becomes a Linear issue on the evlog team via linear__save_issue. Search first (linear__list_issues) for a covering issue, including your own from earlier runs; update rather than duplicate. File the strongest one or two, never a report's worth. A model change is Hugo's call, and the issue is where he makes it.

If linear__save_document is unavailable or fails, fall back to posting the full report in the thread and say why.

When nothing is warranted

One line. Spend flat, no drift, and the models in use still the sane choice means the report says so and stops. Never invent a drift or a swap to make the week look busy.

Settled decisions

  • openai/gpt-6-luna (medium), decided 2026-09-24: keep zai/glm-5.3-flash. Luna is cheaper per task and much faster, but trails badly on agentic work (Terminal-Bench 4.0 at 2.5% against 32.8%) and hallucinates far more (85% against 28%). Revisit if a new Luna release closes the agentic gap.
  • anthropic/claude-haiku-5.5, decided 2026-10-09: replaces zai/glm-5.3-flash as the base model. It runs on the team's Anthropic BYOK key, which the gateway tries first, so its rows report total_cost near zero while market_cost carries the list price. Read groupBy: 'credential_type' before calling a drop in spend real: once the key's monthly credit runs out, the same traffic falls back to gateway credentials and bills there.

© evloghq, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in apps/evi/agent/skills/cost-watchdog of evloghq/evlog.

Open the folder on GitHubat commit 1b6e1b9

Compare with similar skills

Cost Watchdog next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cost Watchdog compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cost Watchdog this skillevloghq/evlog1.9k—~2.2kAutomated safety check: PassMIT
Tabler Astro Dev Servertabler/tabler42k—~1.2kAutomated safety check: PassMIT
Create Docsvictorgarciaesgi/nuxt-typed-router4132 repos~2.8kAutomated safety check: PassMIT
Tabler Astro Component Scriptstabler/tabler42k—~2.1kAutomated safety check: PassMIT
Paperclip Pagepaperclipai/paperclip100k—~1kAutomated safety check: PassMIT
Kill AI Slopyetone/kill-ai-slop1.3k—~1.4kAutomated safety check: PassApache-2.0

Similar skills

  • Starts the right Tabler dev server, keeps it from clashing with builds and verifies changes in the browser before a page or component is handed back.

    42k GitHub stars~1.2k tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • Create Docs

    victorgarciaesgi/nuxt-typed-router

    Create complete documentation sites for projects. An agent skill from victorgarciaesgi/nuxt-typed-router.

    413 GitHub starsUsed in 2 repos~2.8k tokens
    Frontend & DesignAuto-check passed
  • Rules for adding or fixing client-side scripts in Tabler's Astro components so the copied preview HTML stays readable, self-contained and runs in the right order.

    42k GitHub stars~2.1k tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • Paperclip Page

    paperclipai/paperclip

    Publish static HTML pages and asset folders to the Paperclip S3/CloudFront page host.

    100k GitHub stars~1k tokensUpdated today
    Frontend & DesignAuto-check passed
  • Kill AI Slop

    yetone/kill-ai-slop

    Find and remove AI slop — the generic, machine-default visual and copy tics of vibe-coded products — from a web project.

    1.3k GitHub stars~1.4k tokensUpdated 26 days ago
    Frontend & DesignAuto-check passed
  • Official

    Audit a documentation site for agent-friendliness: discovery, markdown delivery, crawlability, semantic structure, machine-readable surfaces, and content legibility.

    4.7k GitHub stars~1.8k tokensUpdated yesterday
    Frontend & DesignAuto-check passed

More from evloghq/evlog

All 20 skills in this repo
  • Walks through adding a new built-in evlog drain adapter for an observability platform: source, build config, exports, tests, docs and PR scope.

    1.9k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Guides adding a new built-in enricher to the evlog package, covering the source, tests, docs, README, a related skill and a changeset.

    1.9k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Walks a contributor through adding a new HTTP framework integration to the evlog logging package: middleware source, build entry, exports, tests, example app and docs.

    1.9k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Walks through adding a new rule or framework adapter to `evlog map` in @evlog/cli, from the rule source and registry to types, tests, docs and the published skill.

    1.9k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Rules for writing and reviewing evlog docs, blog posts, READMEs, skills and AGENTS.md files, with separate review and rewrite roles, a house voice and a catalog of AI-sounding tells.

    1.9k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Before After

    evloghq/evlog

    Produce a before/after visual comparison of an evlog surface (landing, docs, telemetry, playgrounds) and share it as public Blob URLs.

    1.9k GitHub stars~2.7k tokensUpdated today
    Auto-check passed

Questions about Cost Watchdog

What does Cost Watchdog do?

Weekly review of evlog's model cost and performance. An agent skill from evloghq/evlog. Cost Watchdog is an agent skill from evloghq/evlog. Weekly review of evlog's model cost and performance.

When should I use Cost Watchdog?

Cost Watchdog fits situations like: tasks that involve Static sites and blogs; tasks that involve Journaling and reflection.

How do I install Cost Watchdog in Claude Code?

Run `npx skills add evloghq/evlog --skill cost-watchdog -a claude-code`. Or copy the skill folder (apps/evi/agent/skills/cost-watchdog in evloghq/evlog) into .claude/skills/cost-watchdog in your project. Claude Code loads it when a task matches its description.

How do I install Cost Watchdog in Codex?

Run `npx skills add evloghq/evlog --skill cost-watchdog -a codex`. Or copy the skill folder (apps/evi/agent/skills/cost-watchdog in evloghq/evlog) into .agents/skills/cost-watchdog in your project. Codex loads it when a task matches its description.

Can I use Cost Watchdog in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add evloghq/evlog --skill cost-watchdog -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cost-watchdog, .gemini/skills/cost-watchdog, .github/skills/cost-watchdog and .opencode/skills/cost-watchdog in your project.

What does Cost Watchdog need to run?

SKILL.md names no scripts, command-line tools or credentials: Cost Watchdog is instructions for the agent only.

Does Cost Watchdog access the network?

SKILL.md names 4 domains. In commands or code: vercel.com, artificialanalysis.ai, ai-gateway.vercel.sh and arena.ai; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Cost Watchdog safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cost Watchdog use?

Cost Watchdog is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cost Watchdog use?

About 2.2k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cost Watchdog?

Skills that share tags, products or a category with Cost Watchdog: Tabler Astro Dev Server (tabler/tabler, 42k stars), Create Docs (victorgarciaesgi/nuxt-typed-router, 413 stars), Tabler Astro Component Scripts (tabler/tabler, 42k stars) and Paperclip Page (paperclipai/paperclip, 100k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cost Watchdog?

evloghq (a GitHub organization) maintains it in evloghq/evlog, which has 1,889 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on October 9, 2026.

Source: evloghq/evlog on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.