Benchmark is an agent skill from serejaris/personal-corp-os. Use when choosing between engines or models for an agent pipeline and the answer must come from measurement on your own data, not from marketing pages: speech recognition for meeting recordings, the model that turns a transcript into notes, the critic model that checks it. Runs every variant in an isolated container, measures time, cost per hour of input (USD and RUB, with price source and date) and quality (WER/CER and course terms for ASR; code checks plus a judge of a different model for LLM steps), writes…
Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including scripts (for example `README.md`, `README.ru.md` and `examples/adapter_example.py`).
It sits in AI & LLM Engineering, covering Speech recognition and synthesis and LLM evaluation. The repository describes itself as: Personal Corp OS — управление личной компанией через AI-агентов: задачи вне головы, отделы вместо памяти, недельное ретро. Открытые скиллы для Claude Code и Codex. The licence is MIT.