Agent skill

Unicode Data

by cometkim in cometkim/unicode-segmenter

Regenerate the generated Unicode data modules (src/graphemedata.js, generaldata.js, emojidata.js, test/unicodetestdata.js) and verify correctness.

MITAuto-check passed

Install Unicode Data

skills CLI
$ npx skills add cometkim/unicode-segmenter --skill unicode-data -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cometkim/unicode-segmenter unicode-data --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cometkim/unicode-segmenter.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/unicode-data .claude/skills/unicode-data && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
unicode-data
GitHub stars
112
Token cost
~1.1k tokens
SKILL.md length
508 words
Files
1
Skills in repo
4
Repo updated
First seen
Licence
MIT

At a glance

Regenerate the generated Unicode data modules (src/graphemedata.js, generaldata.js, emojidata.js, test/unicodetestdata.js) and verify correctness.

  • Works in 4 steps: yarn test — includes the official UCD… → For rule changes, deep-fuzz in the… → yarn bundle-stats:grapheme — data… → …
  • Bumping the Unicode version
  • SKILL.md covers Encoding invariants (breaking…, Pair table semantics, Inlined fast paths and Verification checklist after…
  • Calls yarn and node

What it does

Unicode Data is an agent skill from cometkim/unicode-segmenter. Regenerate the generated Unicode data modules (src/graphemedata.js, generaldata.js, emojidata.js, test/unicodetestdata.js) and verify correctness. Use when bumping the Unicode version, editing scripts/unicode.js or scripts/lib/encoding.js, or changing grapheme break rules or the pair table.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: A lightweight implementation of the Unicode Text Segmentation (UAX 29). The licence is MIT.

When your agent uses it

  • Bumping the Unicode version
  • Editing scripts/unicode.js
  • Scripts/lib/encoding.js
  • Changing grapheme break rules

Example prompts

  • “/unicode-data”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. yarn test — includes the official UCD GraphemeBreakTest suite, fast-check property tests against Intl.Segmenter (100k runs each), and the…
  2. For rule changes, deep-fuzz in the scratchpad: fast-check comparing graphemeSegments against Intl.Segmenter at ≥1M runs over full-Unicode…
  3. yarn bundle-stats:grapheme — data changes move bundle size; report the delta.
  4. Add a changeset (yarn changeset) for anything user-visible; Unicode version bumps are user-visible.

What it can do on your machine

Read from SKILL.md and the folder at commit 5d3c738. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • yarn
    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use yarn, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Unicode Data loads about 1.1k tokens when it runs. Until then it costs about 78 tokens; SKILL.md has 508 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from cometkim/unicode-segmenter at commit 5d3c738, republished under its MIT licence (© cometkim). 508 words, ~1,071 tokens.

Download SKILL.mdSave it as .claude/skills/unicode-data/SKILL.md (or your agent's skills folder).
name
unicode-data
description
Regenerate the generated Unicode data modules (src/_grapheme_data.js, _general_data.js, _emoji_data.js, test/_unicode_testdata.js) and verify correctness. Use when bumping the Unicode version, editing scripts/unicode.js or scripts/lib/encoding.js, or changing grapheme break rules or the pair table.

Unicode data pipeline

node scripts/unicode.js downloads UCD files for the pinned UNICODE_VERSION (a const near the top of the script) into scripts/unicode_data/ (cached — delete a file to re-fetch) and regenerates:

  • src/_grapheme_data.js — grapheme_data (encoded ranges), grapheme_cats (one base36 digit per range), grapheme_pairs (256-digit 16×16 break-decision table), GraphemeCategory enum
  • src/_general_data.js, src/_emoji_data.js — flat membership tables
  • test/_unicode_testdata.js — official GraphemeBreakTest cases consumed by yarn test

Never hand-edit generated files; fix the generator and re-run.

Encoding invariants (breaking any of these corrupts data silently)

  • Ranges are delta + base36-VLQ encoded: 4 payload bits + 1 continuation bit per character, alphabet [0-9a-v], decoded by native parseInt(ch, 36) in src/core.js (decodeUnicodeData). Encoder lives in scripts/lib/encoding.js. base36 was chosen over base64 deliberately: no codec needed in the bundle, and it compresses better under gzip/brotli.
  • Ranges must be sorted and non-overlapping (delta gaps are non-negative by construction).
  • Flat tables pack end << 5 | category into Uint32 (findUnicodeRangeCategory unpacks with >>> 5 / & 31). Category max is 31; the largest real value is 23 (word break). A future category beyond 31 requires a format change in both core.js and the generator.
  • Grapheme category 15 (InCB_Consonant) is internal-only — it exists so InCB=Consonant data folds into the main grapheme table instead of a separate module. It must never leak through the public API: graphemeSegments masks it to 0 in _catBegin/_catEnd.

Pair table semantics

grapheme_pairs[prevCat * 16 + curCat] values: 0 = break, 1 = no break, 2 = GB12/13 (Regional Indicator parity), 3 = GB11 (ExtPic + ZWJ sequence), 4 = GB9c (InCB linker). The stateful cases (2–4) are resolved by the packed sequence state in src/grapheme.js (nextState + the gate mask). Rule changes belong in buildGraphemePairTable() in scripts/unicode.js, not in hand edits to the emitted string.

Show full SKILL.md (238 more words)Show less

Inlined fast paths

src/grapheme.js hardcodes hot regions (Latin, CJK, Hangul LV/LVT, PUA, variation selectors, E0000 tags, …) as computed checks plus three Uint8Array windows (T0: 0x0000–0x2FFF, T1: 0xA000–0xABFF, T2: 0x1F000–0x1FAFF); only the remainder reaches the binary-search tail. scripts/unicode.js asserts every inlined assumption against the raw UCD tables at codegen time, plus a tail-size cap (≤320 ranges). If regeneration throws an assertion, a new Unicode version changed a region that grapheme.js inlines — update the inline logic in src/grapheme.js to match reality; never weaken the assertion.

Verification checklist after any regen or rule change

  1. yarn test — includes the official UCD GraphemeBreakTest suite, fast-check property tests against Intl.Segmenter (100k runs each), and the real-world corpus suite (test/corpora.js over committed test/_corpora/ fixtures — natural texts caught bugs both of the others missed; see issue #124).
  2. For rule changes, deep-fuzz in the scratchpad: fast-check comparing graphemeSegments against Intl.Segmenter at ≥1M runs over full-Unicode strings. Divergences are not automatically bugs — the host ICU may lag or lead the pinned UNICODE_VERSION; adjudicate against the UAX #29 rules for the pinned version before "fixing" anything.
  3. yarn bundle-stats:grapheme — data changes move bundle size; report the delta.
  4. Add a changeset (yarn changeset) for anything user-visible; Unicode version bumps are user-visible.

On a Unicode version bump, also update UNICODE_VERSION_STRING in scripts/corpora.js (it pins the emoji-test.txt URL) and re-run node scripts/corpora.js to refresh the emoji corpus fixture. Downloads cache under scripts/corpora_data/ (gitignored); the fixtures in test/_corpora/ are committed.

© cometkim, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/unicode-data of cometkim/unicode-segmenter.

Open the folder on GitHubat commit 5d3c738

Compare with similar skills

Unicode Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Unicode Data compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Unicode Data this skillcometkim/unicode-segmenter112—~1.1kAutomated safety check: PassMIT
Es Modulesthedaviddias/Front-End-Checklist74k—~482Automated safety check: PassMIT
Write UI Module Readmeremix-run/remix33k—~1.1kAutomated safety check: PassMIT
Splitting Oversized ModulesPostHog/posthog40k—~2.2kAutomated safety check: PassCustom licence
Abp Moduleabpframework/abp14k—~1.6kAutomated safety check: PassLGPL-3.0
Orchardcore Nswag RegenerateOrchardCMS/OrchardCore8.2k—~2kAutomated safety check: PassBSD-3-Clause

Similar skills

  • Es Modules

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing scripts, client components, bundles, or runtime behavior related to Use ES modules (import/export).

    74k GitHub stars~482 tokensUpdated yesterday
    Auto-check passed
  • Write UI Module Readme

    remix-run/remix

    Write concise module README files for packages/ui/src/lib/ primitives.

    33k GitHub stars~1.1k tokensUpdated today
    DevelopmentAuto-check passed
  • Official

    Split an oversized Python module (a thousand-plus-line logic.py, models.py, api.py, or its test file) into a package of one module per concern, mechanically and provably without changing behavior.

    40k GitHub stars~2.2k tokensUpdated today
    Frontend & DesignAuto-check passed
  • Abp Module

    abpframework/abp

    ABP reusable Module solution template - EF Core + MongoDB dual support, virtual methods for extensibility, DbTablePrefix, module options pattern, entity extension, separate connection string.

    14k GitHub stars~1.6k tokensUpdated today
    DatabasesAuto-check passed
  • Orchardcore Nswag Regenerate

    OrchardCMS/OrchardCore

    Regenerate the OrchardCore.OpenApi module's NSwag-generated C/TypeScript API clients, and verify the regeneration produces a stable (non-reshuffled) diff.

    8.2k GitHub stars~2k tokensUpdated today
    Backend & APIsAuto-check passed
  • Extract Module

    JetBrains/intellij-community

    Official

    Extract optional plugin dependencies into IntelliJ content modules.

    21k GitHub stars~5.2k tokensUpdated today
    MobileAuto-check passed

More from cometkim/unicode-segmenter

  • Codspeed Debug

    cometkim/unicode-segmenter

    Investigate CodSpeed CI performance reports on PRs — regressions or improbable improvements that don't reproduce locally.

    112 GitHub stars~814 tokensUpdated 2 mo ago
    Auto-check passed
  • Bench Records

    cometkim/unicode-segmenter

    Archive benchmark runs into benchmark/grapheme/records and regenerate the HTML report.

    112 GitHub stars~723 tokensUpdated 2 mo ago
    Auto-check passed
  • Benchmark

    cometkim/unicode-segmenter

    Measure runtime perf, bundle size, and memory impact of unicode-segmenter changes.

    112 GitHub stars~1.3k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Unicode Data

What does Unicode Data do?

Regenerate the generated Unicode data modules (src/graphemedata.js, generaldata.js, emojidata.js, test/unicodetestdata.js) and verify correctness. Unicode Data is an agent skill from cometkim/unicode-segmenter.js) and verify correctness.

When should I use Unicode Data?

Unicode Data fits situations like: bumping the Unicode version; editing scripts/unicode.js; scripts/lib/encoding.js; changing grapheme break rules.

How do I install Unicode Data in Claude Code?

Run `npx skills add cometkim/unicode-segmenter --skill unicode-data -a claude-code`. Or copy the skill folder (.claude/skills/unicode-data in cometkim/unicode-segmenter) into .claude/skills/unicode-data in your project. Claude Code loads it when a task matches its description.

How do I install Unicode Data in Codex?

Run `npx skills add cometkim/unicode-segmenter --skill unicode-data -a codex`. Or copy the skill folder (.claude/skills/unicode-data in cometkim/unicode-segmenter) into .agents/skills/unicode-data in your project. Codex loads it when a task matches its description.

Can I use Unicode Data in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cometkim/unicode-segmenter --skill unicode-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/unicode-data, .gemini/skills/unicode-data, .github/skills/unicode-data and .opencode/skills/unicode-data in your project.

What does Unicode Data need to run?

Going by SKILL.md and its folder, Unicode Data needs the command-line tools its instructions call (yarn and node).

Does Unicode Data access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Unicode Data safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Unicode Data use?

Unicode Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Unicode Data use?

About 1.1k tokens (SKILL.md is roughly 4.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Unicode Data?

Skills that share tags, products or a category with Unicode Data: Es Modules (thedaviddias/Front-End-Checklist, 74k stars), Write UI Module Readme (remix-run/remix, 33k stars), Splitting Oversized Modules (PostHog/posthog, 40k stars) and Abp Module (abpframework/abp, 14k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Unicode Data?

cometkim (a GitHub user) maintains it in cometkim/unicode-segmenter, which has 112 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on July 29, 2026.

Source: cometkim/unicode-segmenter on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.