Hugging Face Tokenizers
Orchestra-Research/AI-Research-SKILLs
Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking.
Search local videos by meaning, exact words or reference image, analyze scenes, and export clips using the cerul CLI.
$ npx skills add cerul-ai/cerul --skill cerul -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install cerul-ai/cerul cerul --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/cerul-ai/cerul.git skills-src && mkdir -p .claude/skills && cp -r skills-src/prompts .claude/skills/cerul && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "cerul" agent skill from https://github.com/cerul-ai/cerul/tree/main/prompts into .claude/skills/cerul/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cerul", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/cerul-ai/cerul/tree/main/promptsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add cerul-ai/cerul --skill cerul -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install cerul-ai/cerul cerul --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cerul-ai/cerul.git skills-src && mkdir -p .agents/skills && cp -r skills-src/prompts .agents/skills/cerul && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "cerul" agent skill from https://github.com/cerul-ai/cerul/tree/main/prompts into .agents/skills/cerul/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cerul", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cerul-ai/cerul --skill cerul -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install cerul-ai/cerul cerul --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cerul-ai/cerul.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/prompts .cursor/skills/cerul && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "cerul" agent skill from https://github.com/cerul-ai/cerul/tree/main/prompts into .cursor/skills/cerul/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cerul", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/cerul-ai/cerul.git --path prompts--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add cerul-ai/cerul --skill cerul -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install cerul-ai/cerul cerul --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cerul-ai/cerul.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/prompts .gemini/skills/cerul && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "cerul" agent skill from https://github.com/cerul-ai/cerul/tree/main/prompts into .gemini/skills/cerul/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cerul", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install cerul-ai/cerul cerulInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add cerul-ai/cerul --skill cerul -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/cerul-ai/cerul.git skills-src && mkdir -p .github/skills && cp -r skills-src/prompts .github/skills/cerul && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "cerul" agent skill from https://github.com/cerul-ai/cerul/tree/main/prompts into .github/skills/cerul/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cerul", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cerul-ai/cerul --skill cerul -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install cerul-ai/cerul cerul --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cerul-ai/cerul.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/prompts .opencode/skills/cerul && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "cerul" agent skill from https://github.com/cerul-ai/cerul/tree/main/prompts into .opencode/skills/cerul/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cerul", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
cerulSearch local videos by meaning, exact words or reference image, analyze scenes, and export clips using the cerul CLI.
Cerul is an agent skill from cerul-ai/cerul. Search local videos by meaning, exact words or reference image, analyze scenes, and export clips using the cerul CLI.
Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering. It works with Rust. The repository describes itself as: Open-source video processing core and CLI. Search and annotate local videos with your own model endpoints. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 3698190. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Cerul loads about 1.8k tokens when it runs. Until then it costs about 31 tokens; SKILL.md has 988 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
commit, or read an unrelated project's `.env`. When no keyAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from cerul-ai/cerul at commit 3698190, republished under its Apache-2.0 licence (© cerul-ai). 988 words, ~1,793 tokens.
.claude/skills/cerul/SKILL.md (or your agent's skills folder).Cerul turns local video into searchable, annotated data using the user's own model endpoints. There is no account and no server to start: run the binary and read its output. Sidecar files beside each video are authoritative; search indexes are caches that rebuild from them without model calls.
Run cerul --version. If it is missing, follow the installation runbook at
https://github.com/cerul-ai/cerul/blob/main/docs/agent-setup.md. Do not install
Rust, Python, Ollama, or system FFmpeg for this workflow, and do not switch to a
source build when a download fails; report the error instead.
cerul --json upgrade reports the newest published release and installs
nothing. It replaces the program only with --yes, so ask the user before
running cerul upgrade --yes: it changes which build answers every later
command. A build too old for a flag you need is a reason to offer the upgrade,
not to work around it.
Run cerul --json auth to see whether a key is available. It reports whether a
key is saved or exported, never the value. Never echo a key, print part of it,
write it to a log or a commit, or read an unrelated project's .env. When no key
is available, ask the user to run cerul auth set in their own terminal, or to
export the key themselves.
--json is the machine contract, and it is the only mode to use:
invalid_arguments, not a
question.Exit codes: 0 success, 2 arguments or configuration, 3 unavailable
dependency or capability, 4 execution failure, 5 cancelled, 6 partial
success. Read the code as well as the output. Exit 6 is not success: part of
the work finished and the rest did not.
Events on stderr:
progress counts units of work for one station.annotation_progress reports annotation work units, cached units and phase;
percentage is completed planned work, not model-internal progress.checkpoint marks a window that is saved and will be reused after an
interruption, so a rerun resumes rather than repeating it.published marks a validated annotation file that now exists, with its record
count and path. Only a published file is safe to read or train on.log carries human-readable notices.Commands can return exit 6 for partial results. Repeat the same command to reuse
completed work; do not add --recompute unless fresh processing is intended.
Preview any expensive command with --dry-run first. It writes nothing, calls no
model, and asks for no credential.
Search a video. Indexing is required for search, and it sends media to the configured endpoints.
cerul --json --dry-run index ./video.mp4
cerul --json index ./video.mp4
cerul --json search "a person picking up a cup" --in ./video.mp4
cerul --json search --text "ERROR 500"
cerul --json search "a person picking up a cup" --save ./clipsResult numbers are global, so cerul open 1 plays the first moment in the user's
video player and cerul open 3 plays the third. score is a ranking similarity, not a probability that the
moment is the right one; do not present it as a confidence or an accuracy.
Indexing builds video embeddings, OCR, and available speech/text search data.
It does not call the vision model or generate scene descriptions, sections, or
summaries. Use cerul analyze ./video.mp4 for scenes and an overview. Existing analysis and cached
description vectors remain readable and searchable; indexing does not refresh
them. Search suggestions reuse current evidence without model generation.
Do not treat suggested queries as verified retrieval results.
cerul --json status ./video.mp4 --timeline --type summary
cerul --json status ./video.mp4 --timeline --type sceneVisual descriptions, spoken words, and screen text have distinct provenance. Use scene evidence for what is visible, transcript evidence for what was said, and OCR for visible words. Sparse visual samples cannot establish exact motion boundaries, success, intent, or camera trajectories. Scene descriptions do not replace the task and action annotations below.
Inspect performance. cerul --json diagnostics reads the latest index/analyze
stage timings, request latency and cache counts. Stage wall times overlap; use
invocation elapsed time for total throughput. This makes no model calls.
Analyze a video. No indexing step is needed.
cerul --json analyze ./video.mp4
cerul --json --dry-run analyze ./videos/Without options, returns scenes, chapters and an overview. Add --prompt "Question"
for a focused answer, repeat --image ./reference.png for comparison images, and
use --from 00:30 --to 01:10 for a half-open episode time range. No embeddings
are generated. Focused answers use at most 120 sampled frames and valid cached
in-range text; sparse samples cannot establish continuous motion or absence.
Results include the actual sample timestamps, reference hashes and limitations.
Reference images are not evidence of occurrence in the video.
--stream --json emits provisional analysis_delta NDJSON on stderr; stdout
contains only the final structured report. Never treat a delta as validated
evidence. Exit 6 and per-stream errors mean incomplete analysis. Cached answers
emit one delta marked cached: true. Question-specific results preserve full
video scene/overview records; --recompute refreshes the selected request.
The fixed response schema is published; arbitrary user JSON schemas are not
accepted. Preserve errors and coverage when interpreting results.
Robotics workflows. Embodied labels, human-hand tracking, LeRobot processing and review-video rendering belong to the separate cerul-robotics CLI and skill: https://github.com/cerul-ai/cerul-robotics. Existing records remain readable here.
Read what was produced. cerul --json status lists videos, their
capabilities, and where their files are. cerul --json status PATH --timeline
returns published annotation records in time order, with --type and --limit.
status --providers unless asked; it makes real model requests.--workspace for a whole task. Do not switch it to escape a lock.Generated from this build's argument definitions.
© cerul-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in prompts of cerul-ai/cerul.
Open the folder on GitHubat commit 3698190
Cerul next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Cerul this skillcerul-ai/cerul | 160 | — | ~1.8k | Automated safety check: Notes | Apache-2.0 | |
| Hugging Face TokenizersOrchestra-Research/AI-Research-SKILLs | 13k | 6 repos | ~3.4k | Automated safety check: Pass | MIT | |
| Flowflow Spacesmirkobozzetto/flowflow | 171 | — | ~1k | Automated safety check: Pass | EUPL-1.2 | |
| Celestiacelestiaorg/docs | 183 | — | ~2.8k | Automated safety check: Pass | None | |
| Golem Create Agent Instance Rustgolemcloud/golem | 1.5k | — | ~983 | Automated safety check: Pass | Custom licence | |
| Evalhashgraph-online/awesome-codex-plugins | 1.3k | — | ~2k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking.
mirkobozzetto/flowflow
Use FlowFlow spaces safely through their scoped MCP server. An agent skill from mirkobozzetto/flowflow.
celestiaorg/docs
Route Celestia requests to the correct repo and apply canonical blob submit/retrieve guidance (Go, Rust, and Node RPC) with docs guardrails.
golemcloud/golem
Creating a new Rust Golem agent instance with golem agent new.
hashgraph-online/awesome-codex-plugins
Quality and performance evaluation with baseline comparison.
facebook/pyrefly
A skill your agent uses when Pyrefly computes a wrong tensor shape (or is missing one that can't be expressed in a stub signature) and you need to add or fix a shape-DSL rule.
Works with
Categories
Search local videos by meaning, exact words or reference image, analyze scenes, and export clips using the cerul CLI. Cerul is an agent skill from cerul-ai/cerul. Search local videos by meaning, exact words or reference image, analyze scenes, and export clips using the cerul CLI.
Cerul fits situations like: AI & LLM Engineering work in your project.
Run `npx skills add cerul-ai/cerul --skill cerul -a claude-code`. Or copy the skill folder (prompts in cerul-ai/cerul) into .claude/skills/cerul in your project. Claude Code loads it when a task matches its description.
Run `npx skills add cerul-ai/cerul --skill cerul -a codex`. Or copy the skill folder (prompts in cerul-ai/cerul) into .agents/skills/cerul in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cerul-ai/cerul --skill cerul -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cerul, .gemini/skills/cerul, .github/skills/cerul and .opencode/skills/cerul in your project.
SKILL.md names no scripts, command-line tools or credentials: Cerul is instructions for the agent only.
SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Cerul is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Cerul: Hugging Face Tokenizers (Orchestra-Research/AI-Research-SKILLs, 13k stars), Flowflow Spaces (mirkobozzetto/flowflow, 171 stars), Celestia (celestiaorg/docs, 183 stars) and Golem Create Agent Instance Rust (golemcloud/golem, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
cerul-ai (a GitHub organization) maintains it in cerul-ai/cerul, which has 160 GitHub stars. The repository was last updated on October 8, 2026.
Source: cerul-ai/cerul on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.