Es Modules
thedaviddias/Front-End-Checklist
A skill your agent uses when reviewing scripts, client components, bundles, or runtime behavior related to Use ES modules (import/export).
Regenerate the generated Unicode data modules (src/graphemedata.js, generaldata.js, emojidata.js, test/unicodetestdata.js) and verify correctness.
$ npx skills add cometkim/unicode-segmenter --skill unicode-data -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install cometkim/unicode-segmenter unicode-data --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/cometkim/unicode-segmenter.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/unicode-data .claude/skills/unicode-data && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "unicode-data" agent skill from https://github.com/cometkim/unicode-segmenter/tree/main/.claude/skills/unicode-data into .claude/skills/unicode-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "unicode-data", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/cometkim/unicode-segmenter/tree/main/.claude/skills/unicode-dataType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add cometkim/unicode-segmenter --skill unicode-data -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install cometkim/unicode-segmenter unicode-data --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cometkim/unicode-segmenter.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/unicode-data .agents/skills/unicode-data && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "unicode-data" agent skill from https://github.com/cometkim/unicode-segmenter/tree/main/.claude/skills/unicode-data into .agents/skills/unicode-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "unicode-data", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cometkim/unicode-segmenter --skill unicode-data -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install cometkim/unicode-segmenter unicode-data --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cometkim/unicode-segmenter.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/unicode-data .cursor/skills/unicode-data && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "unicode-data" agent skill from https://github.com/cometkim/unicode-segmenter/tree/main/.claude/skills/unicode-data into .cursor/skills/unicode-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "unicode-data", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/cometkim/unicode-segmenter.git --path .claude/skills/unicode-data--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add cometkim/unicode-segmenter --skill unicode-data -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install cometkim/unicode-segmenter unicode-data --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cometkim/unicode-segmenter.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/unicode-data .gemini/skills/unicode-data && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "unicode-data" agent skill from https://github.com/cometkim/unicode-segmenter/tree/main/.claude/skills/unicode-data into .gemini/skills/unicode-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "unicode-data", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install cometkim/unicode-segmenter unicode-dataInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add cometkim/unicode-segmenter --skill unicode-data -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/cometkim/unicode-segmenter.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/unicode-data .github/skills/unicode-data && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "unicode-data" agent skill from https://github.com/cometkim/unicode-segmenter/tree/main/.claude/skills/unicode-data into .github/skills/unicode-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "unicode-data", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cometkim/unicode-segmenter --skill unicode-data -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install cometkim/unicode-segmenter unicode-data --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cometkim/unicode-segmenter.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/unicode-data .opencode/skills/unicode-data && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "unicode-data" agent skill from https://github.com/cometkim/unicode-segmenter/tree/main/.claude/skills/unicode-data into .opencode/skills/unicode-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "unicode-data", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
unicode-dataRegenerate the generated Unicode data modules (src/graphemedata.js, generaldata.js, emojidata.js, test/unicodetestdata.js) and verify correctness.
Unicode Data is an agent skill from cometkim/unicode-segmenter. Regenerate the generated Unicode data modules (src/graphemedata.js, generaldata.js, emojidata.js, test/unicodetestdata.js) and verify correctness. Use when bumping the Unicode version, editing scripts/unicode.js or scripts/lib/encoding.js, or changing grapheme break rules or the pair table.
Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: A lightweight implementation of the Unicode Text Segmentation (UAX 29). The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 5d3c738. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
yarnnodeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use yarn, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Unicode Data loads about 1.1k tokens when it runs. Until then it costs about 78 tokens; SKILL.md has 508 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from cometkim/unicode-segmenter at commit 5d3c738, republished under its MIT licence (© cometkim). 508 words, ~1,071 tokens.
.claude/skills/unicode-data/SKILL.md (or your agent's skills folder).node scripts/unicode.js downloads UCD files for the pinned UNICODE_VERSION (a const near the top of the script) into scripts/unicode_data/ (cached — delete a file to re-fetch) and regenerates:
src/_grapheme_data.js — grapheme_data (encoded ranges), grapheme_cats (one base36 digit per range), grapheme_pairs (256-digit 16×16 break-decision table), GraphemeCategory enumsrc/_general_data.js, src/_emoji_data.js — flat membership tablestest/_unicode_testdata.js — official GraphemeBreakTest cases consumed by yarn testNever hand-edit generated files; fix the generator and re-run.
[0-9a-v], decoded by native parseInt(ch, 36) in src/core.js (decodeUnicodeData). Encoder lives in scripts/lib/encoding.js. base36 was chosen over base64 deliberately: no codec needed in the bundle, and it compresses better under gzip/brotli.end << 5 | category into Uint32 (findUnicodeRangeCategory unpacks with >>> 5 / & 31). Category max is 31; the largest real value is 23 (word break). A future category beyond 31 requires a format change in both core.js and the generator.InCB_Consonant) is internal-only — it exists so InCB=Consonant data folds into the main grapheme table instead of a separate module. It must never leak through the public API: graphemeSegments masks it to 0 in _catBegin/_catEnd.grapheme_pairs[prevCat * 16 + curCat] values: 0 = break, 1 = no break, 2 = GB12/13 (Regional Indicator parity), 3 = GB11 (ExtPic + ZWJ sequence), 4 = GB9c (InCB linker). The stateful cases (2–4) are resolved by the packed sequence state in src/grapheme.js (nextState + the gate mask). Rule changes belong in buildGraphemePairTable() in scripts/unicode.js, not in hand edits to the emitted string.
src/grapheme.js hardcodes hot regions (Latin, CJK, Hangul LV/LVT, PUA, variation selectors, E0000 tags, …) as computed checks plus three Uint8Array windows (T0: 0x0000–0x2FFF, T1: 0xA000–0xABFF, T2: 0x1F000–0x1FAFF); only the remainder reaches the binary-search tail. scripts/unicode.js asserts every inlined assumption against the raw UCD tables at codegen time, plus a tail-size cap (≤320 ranges). If regeneration throws an assertion, a new Unicode version changed a region that grapheme.js inlines — update the inline logic in src/grapheme.js to match reality; never weaken the assertion.
yarn test — includes the official UCD GraphemeBreakTest suite, fast-check property tests against Intl.Segmenter (100k runs each), and the real-world corpus suite (test/corpora.js over committed test/_corpora/ fixtures — natural texts caught bugs both of the others missed; see issue #124).graphemeSegments against Intl.Segmenter at ≥1M runs over full-Unicode strings. Divergences are not automatically bugs — the host ICU may lag or lead the pinned UNICODE_VERSION; adjudicate against the UAX #29 rules for the pinned version before "fixing" anything.yarn bundle-stats:grapheme — data changes move bundle size; report the delta.yarn changeset) for anything user-visible; Unicode version bumps are user-visible.On a Unicode version bump, also update UNICODE_VERSION_STRING in scripts/corpora.js (it pins the emoji-test.txt URL) and re-run node scripts/corpora.js to refresh the emoji corpus fixture. Downloads cache under scripts/corpora_data/ (gitignored); the fixtures in test/_corpora/ are committed.
© cometkim, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/unicode-data of cometkim/unicode-segmenter.
Open the folder on GitHubat commit 5d3c738
Unicode Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Unicode Data this skillcometkim/unicode-segmenter | 112 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Es Modulesthedaviddias/Front-End-Checklist | 74k | — | ~482 | Automated safety check: Pass | MIT | |
| Write UI Module Readmeremix-run/remix | 33k | — | ~1.1k | Automated safety check: Pass | MIT | |
| Splitting Oversized ModulesPostHog/posthog | 40k | — | ~2.2k | Automated safety check: Pass | Custom licence | |
| Abp Moduleabpframework/abp | 14k | — | ~1.6k | Automated safety check: Pass | LGPL-3.0 | |
| Orchardcore Nswag RegenerateOrchardCMS/OrchardCore | 8.2k | — | ~2k | Automated safety check: Pass | BSD-3-Clause |
thedaviddias/Front-End-Checklist
A skill your agent uses when reviewing scripts, client components, bundles, or runtime behavior related to Use ES modules (import/export).
remix-run/remix
Write concise module README files for packages/ui/src/lib/ primitives.
PostHog/posthog
Split an oversized Python module (a thousand-plus-line logic.py, models.py, api.py, or its test file) into a package of one module per concern, mechanically and provably without changing behavior.
abpframework/abp
ABP reusable Module solution template - EF Core + MongoDB dual support, virtual methods for extensibility, DbTablePrefix, module options pattern, entity extension, separate connection string.
OrchardCMS/OrchardCore
Regenerate the OrchardCore.OpenApi module's NSwag-generated C/TypeScript API clients, and verify the regeneration produces a stable (non-reshuffled) diff.
JetBrains/intellij-community
Extract optional plugin dependencies into IntelliJ content modules.
cometkim/unicode-segmenter
Investigate CodSpeed CI performance reports on PRs — regressions or improbable improvements that don't reproduce locally.
cometkim/unicode-segmenter
Archive benchmark runs into benchmark/grapheme/records and regenerate the HTML report.
cometkim/unicode-segmenter
Measure runtime perf, bundle size, and memory impact of unicode-segmenter changes.
Regenerate the generated Unicode data modules (src/graphemedata.js, generaldata.js, emojidata.js, test/unicodetestdata.js) and verify correctness. Unicode Data is an agent skill from cometkim/unicode-segmenter.js) and verify correctness.
Unicode Data fits situations like: bumping the Unicode version; editing scripts/unicode.js; scripts/lib/encoding.js; changing grapheme break rules.
Run `npx skills add cometkim/unicode-segmenter --skill unicode-data -a claude-code`. Or copy the skill folder (.claude/skills/unicode-data in cometkim/unicode-segmenter) into .claude/skills/unicode-data in your project. Claude Code loads it when a task matches its description.
Run `npx skills add cometkim/unicode-segmenter --skill unicode-data -a codex`. Or copy the skill folder (.claude/skills/unicode-data in cometkim/unicode-segmenter) into .agents/skills/unicode-data in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cometkim/unicode-segmenter --skill unicode-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/unicode-data, .gemini/skills/unicode-data, .github/skills/unicode-data and .opencode/skills/unicode-data in your project.
Going by SKILL.md and its folder, Unicode Data needs the command-line tools its instructions call (yarn and node).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Unicode Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.1k tokens (SKILL.md is roughly 4.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Unicode Data: Es Modules (thedaviddias/Front-End-Checklist, 74k stars), Write UI Module Readme (remix-run/remix, 33k stars), Splitting Oversized Modules (PostHog/posthog, 40k stars) and Abp Module (abpframework/abp, 14k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
cometkim (a GitHub user) maintains it in cometkim/unicode-segmenter, which has 112 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on July 29, 2026.
Source: cometkim/unicode-segmenter on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.