Agent skill

Evals Context

by zgsm-ai in zgsm-ai/costrict

Provides context about the CoStrict evals system structure in this monorepo.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Evals Context

skills CLI
$ npx skills add zgsm-ai/costrict --skill evals-context -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install zgsm-ai/costrict evals-context --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/zgsm-ai/costrict.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.roo/skills/evals-context .claude/skills/evals-context && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evals-context
GitHub stars
4.4k
Used in
1 other repo
Token cost
~1.9k tokens
SKILL.md length
380 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
Apache-2.0

At a glance

Provides context about the CoStrict evals system structure in this monorepo.

  • Works in 2 steps: Add exercise to Roo-Code-Evals repo… → See packages/evals/ADDING-EVALS.md for…
  • Tasks mention evals
  • SKILL.md covers When to Use This Skill, When NOT to Use This Skill, Key Disambiguation: Two… and Directory Structure Reference, plus 5 more sections
  • Calls pnpm and npx

What it does

Evals Context is an agent skill from zgsm-ai/costrict. Provides context about the CoStrict evals system structure in this monorepo. Use when tasks mention "evals", "evaluation", "eval runs", "eval exercises", or working with the evals infrastructure. Helps distinguish between the evals execution system (packages/evals, apps/web-evals) and the public website evals display page (apps/web-roo-code/src/app/evals).

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation and Monorepo tooling. It works with Visual Studio Code, DeepSeek, Google Gemini and MiniMax. The repository describes itself as: Costrict - strict AI coder for enterprises, quality first, including AI Agent, AI CodeReview, AI Completion. The licence is Apache-2.0.

When your agent uses it

  • Tasks mention evals
  • Working with the evals infrastructure

Example prompts

  • “evaluation”
  • “eval runs”
  • “eval exercises”
  • “/evals-context”

Requirements

  • Node.js
  • Docker

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Add exercise to Roo-Code-Evals repo (external)
  2. See packages/evals/ADDING-EVALS.md for structure

What it can do on your machine

Read from SKILL.md and the folder at commit dd38f54. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pnpm
    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evals Context loads about 1.9k tokens when it runs. Until then it costs about 93 tokens; SKILL.md has 380 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~93
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from zgsm-ai/costrict at commit dd38f54, republished under its Apache-2.0 licence (© zgsm-ai). 380 words, ~1,881 tokens.

Download SKILL.mdSave it as .claude/skills/evals-context/SKILL.md (or your agent's skills folder).
name
evals-context
description
Provides context about the CoStrict evals system structure in this monorepo. Use when tasks mention "evals", "evaluation", "eval runs", "eval exercises", or working with the evals infrastructure. Helps distinguish between the evals execution system (packages/evals, apps/web-evals) and the public website evals display page (apps/web-roo-code/src/app/evals).

Evals Codebase Context

When to Use This Skill

Use this skill when the task involves:

  • Modifying or debugging the evals execution infrastructure
  • Adding new eval exercises or languages
  • Working with the evals web interface (apps/web-evals)
  • Modifying the public evals display page on roocode.com
  • Understanding where evals code lives in this monorepo

When NOT to Use This Skill

Do NOT use this skill when:

  • Working on unrelated parts of the codebase (extension, webview-ui, etc.)
  • The task is purely about the VS Code extension's core functionality
  • Working on the main website pages that don't involve evals

Key Disambiguation: Two "Evals" Locations

This monorepo has two distinct evals-related locations that can cause confusion:

ComponentPathPurpose
Evals Execution Systempackages/evals/Core eval infrastructure: CLI, DB schema, Docker configs
Evals Management UIapps/web-evals/Next.js app for creating/monitoring eval runs (localhost:3446)
Website Evals Pageapps/web-roo-code/src/app/evals/Public roocode.com page displaying eval results
External Exercises RepoRoo-Code-EvalsActual coding exercises (NOT in this monorepo)

Directory Structure Reference

packages/evals/ - Core Evals Package
packages/evals/
├── ARCHITECTURE.md          # Detailed architecture documentation
├── ADDING-EVALS.md          # Guide for adding new exercises/languages
├── README.md                # Setup and running instructions
├── docker-compose.yml       # Container orchestration
├── Dockerfile.runner        # Runner container definition
├── Dockerfile.web           # Web app container
├── drizzle.config.ts        # Database ORM config
├── src/
│   ├── index.ts             # Package exports
│   ├── cli/                 # CLI commands for running evals
│   │   ├── runEvals.ts      # Orchestrates complete eval runs
│   │   ├── runTask.ts       # Executes individual tasks in containers
│   │   ├── runUnitTest.ts   # Validates task completion via tests
│   │   └── redis.ts         # Redis pub/sub integration
│   ├── db/
│   │   ├── schema.ts        # Database schema (runs, tasks)
│   │   ├── queries/         # Database query functions
│   │   └── migrations/      # SQL migrations
│   └── exercises/
│       └── index.ts         # Exercise loading utilities
└── scripts/
    └── setup.sh             # Local macOS setup script
apps/web-evals/ - Evals Management Web App
apps/web-evals/
├── src/
│   ├── app/
│   │   ├── page.tsx         # Home page (runs list)
│   │   ├── runs/
│   │   │   ├── new/         # Create new eval run
│   │   │   └── [id]/        # View specific run status
│   │   └── api/runs/        # SSE streaming endpoint
│   ├── actions/             # Server actions
│   │   ├── runs.ts          # Run CRUD operations
│   │   ├── tasks.ts         # Task queries
│   │   ├── exercises.ts     # Exercise listing
│   │   └── heartbeat.ts     # Controller health checks
│   ├── hooks/               # React hooks (SSE, models, etc.)
│   └── lib/                 # Utilities and schemas
apps/web-roo-code/src/app/evals/ - Public Website Evals Page
apps/web-roo-code/src/app/evals/
├── page.tsx      # Fetches and displays public eval results
├── evals.tsx     # Main evals display component
├── plot.tsx      # Visualization component
└── types.ts      # EvalRun type (extends packages/evals types)

This page displays eval results on the public roocode.com website. It imports types from @roo-code/evals but does NOT run evals.

Architecture Overview

The evals system is a distributed evaluation platform that runs AI coding tasks in isolated VS Code environments:

┌─────────────────────────────────────────────────────────────┐
│  Web App (apps/web-evals)  ──────────────────────────────── │
│        │                                                    │
│        ▼                                                    │
│  PostgreSQL ◄────► Controller Container                     │
│        │               │                                    │
│        ▼               ▼                                    │
│     Redis ◄───► Runner Containers (1-25 parallel)           │
└─────────────────────────────────────────────────────────────┘

Key components:

  • Controller: Orchestrates eval runs, spawns runners, manages task queue (p-queue)
  • Runner: Isolated Docker container with VS Code + CoStrict extension + language runtimes
  • Redis: Pub/sub for real-time events (NOT task queuing)
  • PostgreSQL: Stores runs, tasks, metrics
Show full SKILL.md (130 more words)Show less

Common Tasks Quick Reference

Adding a New Eval Exercise
  1. Add exercise to Roo-Code-Evals repo (external)
  2. See packages/evals/ADDING-EVALS.md for structure
Modifying Eval CLI Behavior

Edit files in packages/evals/src/cli/:

Modifying the Evals Web Interface

Edit files in apps/web-evals/src/:

Modifying the Public Evals Display Page

Edit files in apps/web-roo-code/src/app/evals/:

Database Schema Changes
  1. Edit packages/evals/src/db/schema.ts
  2. Generate migration: cd packages/evals && pnpm drizzle-kit generate
  3. Apply migration: pnpm drizzle-kit migrate

Running Evals Locally

bash
# From repo root
pnpm evals

# Opens web UI at http://localhost:3446

Ports (defaults):

  • PostgreSQL: 5433
  • Redis: 6380
  • Web: 3446

Testing

bash
# packages/evals tests
cd packages/evals && npx vitest run

# apps/web-evals tests
cd apps/web-evals && npx vitest run

Key Types/Exports from @roo-code/evals

The package exports are defined in packages/evals/src/index.ts:

  • Database queries: getRuns, getTasks, getTaskMetrics, etc.
  • Schema types: Run, Task, TaskMetrics
  • Used by both apps/web-evals and apps/web-roo-code

© zgsm-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .roo/skills/evals-context of zgsm-ai/costrict.

Open the folder on GitHubat commit dd38f54

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in zgsm-ai/costrict, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Evals Context next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evals Context compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evals Context this skillzgsm-ai/costrict4.4k1 repos~1.9kAutomated safety check: PassApache-2.0
OmniRoute Chat CLIdiegosouzapw/OmniRoute74k—~345Automated safety check: PassMIT
ModLens Image Vision Bridgeliustack/modlens4.2k—~1.3kAutomated safety check: NotesMIT
Quality FlywheelGoogleCloudPlatform/vertex-ai-samples792—~2kAutomated safety check: PassApache-2.0
Cross-Model Benchmarkgarrytan/gstack136k—~4kAutomated safety check: NotesMIT
Frontierharness Evalfrontier-harness-eval/eval301—~8kAutomated safety check: PassNone

Similar skills

  • OmniRoute Chat CLI

    diegosouzapw/OmniRoute

    Sends chat completions, streams responses, and opens an interactive REPL against any OmniRoute-routed model provider.

    74k GitHub stars~345 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Gives text-only models sight by running the modlens CLI on an image path or URL and returning structured JSON evidence with transcribed text, layout and semantics.

    4.2k GitHub stars~1.3k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check: notes
  • Quality Flywheel

    GoogleCloudPlatform/vertex-ai-samples

    Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

    792 GitHub stars~2k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Cross-Model Benchmark

    garrytan/gstack

    Sends one prompt to Claude, GPT through the Codex CLI and Gemini, then tabulates response time, token use and cost, with an optional judged quality score.

    136k GitHub stars~4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Frontierharness Eval

    frontier-harness-eval/eval

    Benchmark a third-party coding-agent harness against FrontierHarness Eval using Runta runtimes.

    301 GitHub stars~8k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Bridgic LLMs

    bitsky-tech/bridgic

    LLM provider initialization for bridgic projects. An agent skill from bitsky-tech/bridgic.

    155 GitHub stars~839 tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes

More from zgsm-ai/costrict

  • Roo Conflict Resolution

    zgsm-ai/costrict

    Provides comprehensive guidelines for resolving merge conflicts intelligently using git history and commit context.

    4.4k GitHub stars~2.3k tokensUpdated 9 days ago
    Auto-check passed
  • Roo Translation

    zgsm-ai/costrict

    Provides comprehensive guidelines for translating and localizing CoStrict extension strings.

    4.4k GitHub stars~1.8k tokensUpdated 9 days ago
    Auto-check passed

Questions about Evals Context

What does Evals Context do?

Provides context about the CoStrict evals system structure in this monorepo. Evals Context is an agent skill from zgsm-ai/costrict. Provides context about the CoStrict evals system structure in this monorepo.

When should I use Evals Context?

Evals Context fits situations like: tasks mention evals; working with the evals infrastructure.

How do I install Evals Context in Claude Code?

Run `npx skills add zgsm-ai/costrict --skill evals-context -a claude-code`. Or copy the skill folder (.roo/skills/evals-context in zgsm-ai/costrict) into .claude/skills/evals-context in your project. Claude Code loads it when a task matches its description.

How do I install Evals Context in Codex?

Run `npx skills add zgsm-ai/costrict --skill evals-context -a codex`. Or copy the skill folder (.roo/skills/evals-context in zgsm-ai/costrict) into .agents/skills/evals-context in your project. Codex loads it when a task matches its description.

Can I use Evals Context in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zgsm-ai/costrict --skill evals-context -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evals-context, .gemini/skills/evals-context, .github/skills/evals-context and .opencode/skills/evals-context in your project.

What does Evals Context need to run?

Going by SKILL.md and its folder, Evals Context needs the command-line tools its instructions call (pnpm and npx). Our summary lists: Node.js; Docker.

Does Evals Context access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Evals Context safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evals Context use?

Evals Context is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evals Context use?

About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evals Context?

Skills that share tags, products or a category with Evals Context: OmniRoute Chat CLI (diegosouzapw/OmniRoute, 74k stars), ModLens Image Vision Bridge (liustack/modlens, 4.2k stars), Quality Flywheel (GoogleCloudPlatform/vertex-ai-samples, 792 stars) and Cross-Model Benchmark (garrytan/gstack, 136k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evals Context?

zgsm-ai (a GitHub organization) maintains it in zgsm-ai/costrict, which has 4,448 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 30, 2026.

Source: zgsm-ai/costrict on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.