Install the "one-eval" agent skill from https://github.com/OpenDCAI/One-Eval/tree/main/one-eval-skill into .claude/skills/one-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "one-eval", then confirm the skill loads.
Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Type this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
skills CLI
$ npx skills add OpenDCAI/One-Eval --skill one-eval -a codex
Project install goes to .agents/skills/; add -g for ~/.codex/skills/.
Install the "one-eval" agent skill from https://github.com/OpenDCAI/One-Eval/tree/main/one-eval-skill into .agents/skills/one-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "one-eval", then confirm the skill loads.
Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add OpenDCAI/One-Eval --skill one-eval -a cursor
Project install goes to .agents/skills/; add -g for ~/.cursor/skills/.
Install the "one-eval" agent skill from https://github.com/OpenDCAI/One-Eval/tree/main/one-eval-skill into .cursor/skills/one-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "one-eval", then confirm the skill loads.
Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
skills CLI
$ npx skills add OpenDCAI/One-Eval --skill one-eval -a gemini-cli
Project install goes to .agents/skills/; add -g for ~/.gemini/skills/.
Install the "one-eval" agent skill from https://github.com/OpenDCAI/One-Eval/tree/main/one-eval-skill into .gemini/skills/one-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "one-eval", then confirm the skill loads.
Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
GitHub CLI
$ gh skill install OpenDCAI/One-Eval one-eval
Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
skills CLI
$ npx skills add OpenDCAI/One-Eval --skill one-eval -a github-copilot
Project install goes to .agents/skills/; add -g for ~/.copilot/skills/.
Install the "one-eval" agent skill from https://github.com/OpenDCAI/One-Eval/tree/main/one-eval-skill into .github/skills/one-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "one-eval", then confirm the skill loads.
GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add OpenDCAI/One-Eval --skill one-eval -a opencode
OpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
Install the "one-eval" agent skill from https://github.com/OpenDCAI/One-Eval/tree/main/one-eval-skill into .opencode/skills/one-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "one-eval", then confirm the skill loads.
OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Works in 9 steps: 先确认运行环境(优先复用已有环境,能不装就不装) → 选模型 + 测连通性(强制门槛) → 选 benchmark → …
Tasks that involve Agent evaluation and testing
SKILL.md covers 前置环境(首次使用必读), 标准流程(按序执行,不要跳步), 文件地图 and 安全 & 边界, plus 1 more section
Runs Python scripts from its folder; calls python, pip and uv
What it does
One Eval is an agent skill from OpenDCAI/One-Eval. 驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。
Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 27 other files, including scripts, reference files and assets (for example `DEV_NOTES.md`, `assets/custom_metric.template.py` and `assets/evalspec.template.yaml`).
It sits in AI & LLM Engineering, covering Agent evaluation and testing and LLM inference and serving. It works with vLLM, Python and SGLang. The repository describes itself as: Automated system for LLM evaluation via agents. Doc as below:. The licence is Apache-2.0.
When your agent uses it
Tasks that involve Agent evaluation and testing
Tasks that involve LLM inference and serving
Example prompts
“/one-eval”
Requirements
Python 3
Workflow steps
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 8aba20a. It shows what the files ask for, not the result of running them.
Tool permissions
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Runs code
Ships 3 files in scripts/ (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
python
pip
uv
From the folder's file list and the shell code blocks in SKILL.md.
Network
No URLs in SKILL.md. Its commands use pip and uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Credentials
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Context cost
One Eval loads about 2.4k tokens when it runs, and up to ~27k if it reads all its reference files. Until then it costs about 35 tokens; SKILL.md has 634 words of instructions outside code blocks.
Always· name and description, kept in context so the agent knows when to use it
~35
When it runs· the whole SKILL.md, loaded when a task matches
~2.4k
With references· SKILL.md plus every file in references/, read only if the agent opens them
~27k
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
Safety
Auto-check passed
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
Download SKILL.mdSave it as .claude/skills/one-eval/SKILL.md (or your agent's skills folder). This skill also uses 23 other files; get the full folder from GitHub.
One Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.
Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs and reports a validated gain, using the Hyperloom multi-agent optimizer.
Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA.
驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。. One Eval is an agent skill from OpenDCAI/One-Eval.
When should I use One Eval?
One Eval fits situations like: tasks that involve Agent evaluation and testing; tasks that involve LLM inference and serving.
How do I install One Eval in Claude Code?
Run `npx skills add OpenDCAI/One-Eval --skill one-eval -a claude-code`. Or copy the skill folder (one-eval-skill in OpenDCAI/One-Eval) into .claude/skills/one-eval in your project. Claude Code loads it when a task matches its description.
How do I install One Eval in Codex?
Run `npx skills add OpenDCAI/One-Eval --skill one-eval -a codex`. Or copy the skill folder (one-eval-skill in OpenDCAI/One-Eval) into .agents/skills/one-eval in your project. Codex loads it when a task matches its description.
Can I use One Eval in Cursor, Gemini CLI or GitHub Copilot?
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OpenDCAI/One-Eval --skill one-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/one-eval, .gemini/skills/one-eval, .github/skills/one-eval and .opencode/skills/one-eval in your project.
What does One Eval need to run?
Going by SKILL.md and its folder, One Eval needs Python for the scripts in its folder and the command-line tools its instructions call (python, pip and uv). Our summary lists: Python 3.
Does One Eval access the network?
SKILL.md contains no URLs. Its commands use pip and uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Is One Eval safe to install?
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
What licence does One Eval use?
One Eval is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
How many tokens does One Eval use?
About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 24k tokens, read only when the agent opens those files.
What are the alternatives to One Eval?
Skills that share tags, products or a category with One Eval: Dstack Prototyping (dstackai/dstack, 2.3k stars), SGLang Structured Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hyperloom Workload Optimizer (amd/skills, 398 stars) and LLM Serving Framework Benchmark (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Who maintains One Eval?
OpenDCAI (a GitHub organization) maintains it in OpenDCAI/One-Eval, which has 165 GitHub stars. The repository was last updated on August 31, 2026.
Source: OpenDCAI/One-Eval on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.