Agent skill

Benchmark Translate

by shapeshift in shapeshift/web

Run a quality benchmark of the /translate skill by selecting stratified test keys, capturing ground truth, translating, judging with sub-agents, and compiling a regression report.

MITAuto-check passedWriting & Content

Install Benchmark Translate

skills CLI
$ npx skills add shapeshift/web --skill benchmark-translate -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install shapeshift/web benchmark-translate --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/shapeshift/web.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/benchmark-translate .claude/skills/benchmark-translate && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark-translate
GitHub stars
206
Token cost
~1.6k tokens
SKILL.md length
526 words
Files
5 (incl. scripts)
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Run a quality benchmark of the /translate skill by selecting stratified test keys, capturing ground truth, translating, judging with sub-agents, and compiling a regression report.

  • Works in 7 steps: Select Keys → Setup → Translate → …
  • Tasks that involve Translation
  • SKILL.md covers Data Artifacts and Pipeline (7 Steps)
  • Runs JavaScript scripts from its folder; calls node and git

What it does

Benchmark Translate is an agent skill from shapeshift/web. Run a quality benchmark of the /translate skill by selecting stratified test keys, capturing ground truth, translating, judging with sub-agents, and compiling a regression report. Invoke with /benchmark-translate.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `scripts/compile-report.js`, `scripts/restore.js` and `scripts/select-keys.js`).

It sits in Writing & Content, covering Translation, Subagents and LLM evaluation. The licence is MIT.

When your agent uses it

  • Tasks that involve Translation
  • Tasks that involve Subagents
  • Tasks that involve LLM evaluation

Example prompts

  • “/benchmark-translate”

Requirements

  • Node.js
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Grep, Glob, Bash(node *), Bash(git checkout*), Bash(git diff*), Bash(git status*), Bash(git rev-parse*), Task, Skill, AskUserQuestion

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Select Keys
  2. Setup
  3. Translate
  4. Judge (Sub-Agents)
  5. Compile Report
  6. Restore
  7. Present Results

What it can do on your machine

Read from SKILL.md and the folder at commit 45096d2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Grep
    • Glob
    • Bash(node *)
    • Bash(git checkout*)
    • Bash(git diff*)
    • Bash(git status*)
    • Bash(git rev-parse*)

    …and 3 more on the same allowed-tools line.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • node
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark Translate loads about 1.6k tokens when it runs. Until then it costs about 58 tokens; SKILL.md has 526 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~58
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from shapeshift/web at commit 45096d2, republished under its MIT licence (© shapeshift). 526 words, ~1,550 tokens.

Download SKILL.mdSave it as .claude/skills/benchmark-translate/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
benchmark-translate
description
Run a quality benchmark of the /translate skill by selecting stratified test keys, capturing ground truth, translating, judging with sub-agents, and compiling a regression report. Invoke with /benchmark-translate.
allowed-tools
Read, Write, Edit, Grep, Glob, Bash(node *), Bash(git checkout*), Bash(git diff*), Bash(git status*), Bash(git rev-parse*), Task, Skill, AskUserQuestion

Translation Quality Benchmark

Measures the quality of the /translate skill by comparing its output against existing human translations. Uses stratified key selection with a fixed/rotating split, LLM judges, and programmatic validation to produce a comprehensive quality report with regression tracking across all 9 supported locales.

Data Artifacts

All benchmark data lives in scripts/translations/benchmark/ (gitignored):

FilePurpose
testKeys.jsonSelected test keys with categories and fixed flag
coreKeys.jsonPersistent core key set (stable across runs)
ground-truth.jsonCaptured human translations before removal
report.jsonLatest benchmark report (becomes baseline on next run)
baseline.jsonPrevious report (auto-copied by setup.js)

Pipeline (7 Steps)

Step 1: Select Keys
bash
node .claude/skills/benchmark-translate/scripts/select-keys.js [--count N] [--core N]

Selects N keys (default 150) stratified across 6 categories: glossary-term, financial-error, single-word, interpolation, defi-jargon, general. Validates all selected keys exist in en + all 9 locales.

Fixed/rotating split:

  • --core N (default 100): Number of fixed core keys for stable regression tracking
  • --count N (default 150): Total keys (core + rotating)
  • If coreKeys.json exists: loads it, validates keys still exist in all locales, tops up if needed
  • If coreKeys.json doesn't exist: selects core keys via stratified sampling and saves them
  • Remaining keys (default 50) are randomly selected as rotating keys from the non-core pool
  • Each entry in testKeys.json has "fixed": true (core) or "fixed": false (rotating)

Outputs scripts/translations/benchmark/testKeys.json.

Step 2: Setup
bash
node .claude/skills/benchmark-translate/scripts/setup.js
  • If report.json exists from a previous run, copies it to baseline.json
  • Reads testKeys.json, captures ground truth translations for all 9 locales
  • Writes ground-truth.json
  • Removes test keys from locale files so /translate can regenerate them
Step 3: Translate

Invoke the /translate skill using the Skill tool. This regenerates the removed keys through the full translate-review-refine pipeline.

Show full SKILL.md (262 more words)Show less
Step 4: Judge (Sub-Agents)

Launch 9 sub-agents in 3 waves of 3 (matching /translate's wave structure) using the Task tool. Each sub-agent receives the locale info, all key triplets, and glossary terms.

Wave 1: de, es, fr Wave 2: pt, ru, tr Wave 3: ja, uk, zh

For each locale, use this prompt:

You are an expert multilingual localization quality assessor for a cryptocurrency/DeFi application.
Rate translations from English into {LANGUAGE_NAME} on a 1-5 scale.

1 = Wrong/misleading meaning
2 = Significant issues (wrong register, missing nuance)
3 = Acceptable but could be more natural
4 = Good, natural, accurate
5 = Excellent, indistinguishable from professional native translation

Check: meaning preservation, naturalness, register ({REGISTER}), UI conciseness,
glossary compliance (these stay English: {NEVER_TRANSLATE_TERMS}),
placeholder integrity (%{...} preserved), DeFi terminology conventions.

Rate each translation INDEPENDENTLY. Community translations can contain errors.

Input: JSON array of {key, english, human, skill}
{ITEMS_JSON}

Output: Return ONLY a JSON array of objects with these exact fields:
{key, humanScore, skillScore, humanJustification, skillJustification, preferenceNote}

Scores must be integers 1-5. Justifications should be 1-2 sentences. preferenceNote should say which is better and why, or "tie" if equal.

Locale info for prompt substitution:

LocaleLanguageRegister
deGermanFormal (Sie)
esSpanishInformal (tú)
frFrenchFormal (vous)
jaJapanesePolite (です/ます)
ptPortugueseInformal (você)
ruRussianFormal (вы)
trTurkishFormal (siz)
ukUkrainianFormal (ви)
zhChinese (Simplified)Neutral/formal

Building the items array for each locale:

  1. Read scripts/translations/benchmark/ground-truth.json
  2. Read the current (post-translate) src/assets/translations/{locale}/main.json
  3. For each test key, build: { key: dottedPath, english: groundTruth.english[key], human: groundTruth.groundTruth[locale][key], skill: getValueFromLocaleFile(key) }

Getting never-translate terms: Read src/assets/translations/glossary.json, collect all keys where value is null (excluding _meta).

Each sub-agent must write its output to /tmp/{locale}-judge-scores.json. Parse the JSON array from the sub-agent's response and write it to that path.

Step 5: Compile Report
bash
node .claude/skills/benchmark-translate/scripts/compile-report.js

Loads judge scores from /tmp/{locale}-judge-scores.json, runs programmatic validation (including Cyrillic script check for ru/uk), computes summary stats, and writes scripts/translations/benchmark/report.json. If baseline.json exists, includes regression deltas. Report includes coreSummary and rotatingSummary alongside the overall summary.

Step 6: Restore
bash
node .claude/skills/benchmark-translate/scripts/restore.js

Restores locale files via git checkout --, verifies no diff remains.

Step 7: Present Results

Read the compile output (printed to stdout) and present to the user:

  • Overall score summary with baseline regression (if available)
  • Core vs rotating stats (divergence suggests overfitting)
  • Notable improvements/regressions
  • Per-locale and per-category highlights
  • Any items needing attention (low scores, validation failures, glossary issues)

© shapeshift, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts) in .claude/skills/benchmark-translate of shapeshift/web.

  • SKILL.md
  • scripts/compile-report.js
  • scripts/restore.js
  • scripts/select-keys.js
  • scripts/setup.js

Open the folder on GitHubat commit 45096d2

Compare with similar skills

Benchmark Translate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark Translate compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark Translate this skillshapeshift/web206—~1.6kAutomated safety check: PassMIT
Po Translatenatsukium/dotfiles106—~1.4kAutomated safety check: PassCC0-1.0
Translateforthecraft/drf-auth-kit121—~2.4kAutomated safety check: NotesMIT
Docs Leadlablup/backend.ai-webui133—~2.9kAutomated safety check: PassLGPL-3.0
Canvas HumanizerX-isdoingreat/canvas-pilot125—~10kAutomated safety check: NotesAGPL-3.0
Yao Meta Skillyaojingang/yao-meta-skill2.7k—~768Automated safety check: PassMIT

Similar skills

  • Po Translate

    natsukium/dotfiles

    Orchestrate English→Japanese translation of po/ja.po — classify, delegate translation/review to subagents, iterate until clean

    106 GitHub stars~1.4k tokensUpdated yesterday
    Writing & ContentAuto-check passed
  • Translate

    forthecraft/drf-auth-kit

    Run Django translation workflow - extract untranslated strings, translate them to 59 languages using parallel translator subagents, and apply back to PO files.

    121 GitHub stars~2.4k tokensUpdated 1 mo ago
    Writing & ContentAuto-check: notes
  • Docs Lead

    lablup/backend.ai-webui

    A skill your agent uses whenever the user mentions docs, the manual, documentation, terminology, translations, or screenshots — including indirect mentions like "이 PR 문서 영향 봐줘", "문서 점검", "용어 통일"…

    133 GitHub stars~2.9k tokensUpdated yesterday
    Writing & ContentAuto-check passed
  • Canvas Humanizer

    X-isdoingreat/canvas-pilot

    Reduces AI-detection signals in drafted text by routing every non-locked sentence through round-trip translation (English → intermediate language → English) and selecting the candidate that…

    125 GitHub stars~10k tokensUpdated 2 mo ago
    Writing & ContentAuto-check: notes
  • Yao Meta Skill

    yaojingang/yao-meta-skill

    Create, improve, or evaluate an existing skill from workflows, prompts, SOPs, scripts.

    2.7k GitHub stars~768 tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed
  • Translate Book

    deusyu/translate-book

    Translate books (PDF/DOCX/EPUB) into any language using parallel sub-agents.

    2.1k GitHub stars~5.5k tokensUpdated 14 days ago
    Documents & OfficeAuto-check: notes

More from shapeshift/web

  • React Best Practices

    shapeshift/web

    Comprehensive React and Next.js performance optimization guide with 40+ rules for eliminating waterfalls, optimizing bundles, and improving rendering.

    206 GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Qabot Fixture

    shapeshift/web

    Create a new qabot E2E test fixture interactively. An agent skill from shapeshift/web.

    206 GitHub stars~838 tokensUpdated yesterday
    Auto-check: notes
  • Translate

    shapeshift/web

    Translate new/changed English UI strings into all supported languages using a translate-review-refine pipeline.

    206 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check passed
  • Chain Integration

    shapeshift/web

    Integrate a new blockchain as a second-class citizen in ShapeShift Web.

    206 GitHub stars~20k tokensUpdated yesterday
    Auto-check: notes
  • Qabot

    shapeshift/web

    Run QA tests using agent-browser and post results to the qabot dashboard.

    206 GitHub stars~8.2k tokensUpdated yesterday
    Auto-check: notes
  • Swapper Integration

    shapeshift/web

    Integrate new DEX aggregators, swappers, or bridge protocols (like Bebop, Portals, Jupiter, 0x, 1inch, etc.) into ShapeShift Web.

    206 GitHub stars~11k tokensUpdated yesterday
    Auto-check: notes

Questions about Benchmark Translate

What does Benchmark Translate do?

Run a quality benchmark of the /translate skill by selecting stratified test keys, capturing ground truth, translating, judging with sub-agents, and compiling a regression report. Benchmark Translate is an agent skill from shapeshift/web. Run a quality benchmark of the /translate skill by selecting stratified test keys, capturing ground truth, translating, judging with sub-agents, and compiling a regression report.

When should I use Benchmark Translate?

Benchmark Translate fits situations like: tasks that involve Translation; tasks that involve Subagents; tasks that involve LLM evaluation.

How do I install Benchmark Translate in Claude Code?

Run `npx skills add shapeshift/web --skill benchmark-translate -a claude-code`. Or copy the skill folder (.claude/skills/benchmark-translate in shapeshift/web) into .claude/skills/benchmark-translate in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark Translate in Codex?

Run `npx skills add shapeshift/web --skill benchmark-translate -a codex`. Or copy the skill folder (.claude/skills/benchmark-translate in shapeshift/web) into .agents/skills/benchmark-translate in your project. Codex loads it when a task matches its description.

Can I use Benchmark Translate in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add shapeshift/web --skill benchmark-translate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-translate, .gemini/skills/benchmark-translate, .github/skills/benchmark-translate and .opencode/skills/benchmark-translate in your project.

What does Benchmark Translate need to run?

Going by SKILL.md and its folder, Benchmark Translate needs JavaScript for the scripts in its folder and the command-line tools its instructions call (node and git). Our summary lists: Node.js. Its frontmatter pre-approves these tools: Read, Write, Edit, Grep, Glob, Bash(node *), Bash(git checkout*), Bash(git diff*), Bash(git status*), Bash(git rev-parse*), Task, Skill, AskUserQuestion.

Does Benchmark Translate access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Benchmark Translate safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Benchmark Translate use?

Benchmark Translate is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmark Translate use?

About 1.6k tokens (SKILL.md is roughly 6.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmark Translate?

Skills that share tags, products or a category with Benchmark Translate: Po Translate (natsukium/dotfiles, 106 stars), Translate (forthecraft/drf-auth-kit, 121 stars), Docs Lead (lablup/backend.ai-webui, 133 stars) and Canvas Humanizer (X-isdoingreat/canvas-pilot, 125 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark Translate?

shapeshift (a GitHub organization) maintains it in shapeshift/web, which has 206 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 7, 2026.

Source: shapeshift/web on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.