Evals Context
zgsm-ai/costrict
Provides context about the CoStrict evals system structure in this monorepo.
Benchmark a third-party coding-agent harness against FrontierHarness Eval using Runta runtimes.
$ npx skills add frontier-harness-eval/eval --skill frontierharness-eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install frontier-harness-eval/eval frontierharness-eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/frontier-harness-eval/eval.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/frontierharness-eval .claude/skills/frontierharness-eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "frontierharness-eval" agent skill from https://github.com/frontier-harness-eval/eval/tree/main/skills/frontierharness-eval into .claude/skills/frontierharness-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "frontierharness-eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/frontier-harness-eval/eval/tree/main/skills/frontierharness-evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add frontier-harness-eval/eval --skill frontierharness-eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install frontier-harness-eval/eval frontierharness-eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/frontier-harness-eval/eval.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/frontierharness-eval .agents/skills/frontierharness-eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "frontierharness-eval" agent skill from https://github.com/frontier-harness-eval/eval/tree/main/skills/frontierharness-eval into .agents/skills/frontierharness-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "frontierharness-eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add frontier-harness-eval/eval --skill frontierharness-eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install frontier-harness-eval/eval frontierharness-eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/frontier-harness-eval/eval.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/frontierharness-eval .cursor/skills/frontierharness-eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "frontierharness-eval" agent skill from https://github.com/frontier-harness-eval/eval/tree/main/skills/frontierharness-eval into .cursor/skills/frontierharness-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "frontierharness-eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/frontier-harness-eval/eval.git --path skills/frontierharness-eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add frontier-harness-eval/eval --skill frontierharness-eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install frontier-harness-eval/eval frontierharness-eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/frontier-harness-eval/eval.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/frontierharness-eval .gemini/skills/frontierharness-eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "frontierharness-eval" agent skill from https://github.com/frontier-harness-eval/eval/tree/main/skills/frontierharness-eval into .gemini/skills/frontierharness-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "frontierharness-eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install frontier-harness-eval/eval frontierharness-evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add frontier-harness-eval/eval --skill frontierharness-eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/frontier-harness-eval/eval.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/frontierharness-eval .github/skills/frontierharness-eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "frontierharness-eval" agent skill from https://github.com/frontier-harness-eval/eval/tree/main/skills/frontierharness-eval into .github/skills/frontierharness-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "frontierharness-eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add frontier-harness-eval/eval --skill frontierharness-eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install frontier-harness-eval/eval frontierharness-eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/frontier-harness-eval/eval.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/frontierharness-eval .opencode/skills/frontierharness-eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "frontierharness-eval" agent skill from https://github.com/frontier-harness-eval/eval/tree/main/skills/frontierharness-eval into .opencode/skills/frontierharness-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "frontierharness-eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
frontierharness-evalBenchmark a third-party coding-agent harness against FrontierHarness Eval using Runta runtimes.
Frontierharness Eval is an agent skill from frontier-harness-eval/eval. Benchmark a third-party coding-agent harness against FrontierHarness Eval using Runta runtimes. Provisions a clean runtime bound to the GitHub repo under evaluation, installs the Harbor and Pier stacks needed for Terminal-Bench and DeepSWE tasks, freezes a golden checkpoint, runs tasks from identical fresh restores while saving trajectories as evidence, generates a comparison diagram, and builds a shareable report. Use when evaluating, benchmarking, scoring, or comparing a coding agent harness, or when the user…
Its SKILL.md is about 8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 22 other files, including scripts (for example `PROMPT.md`, `isolated-evaluation-repairs.md` and `reference.md`).
It sits in AI & LLM Engineering, covering LLM evaluation. It works with GitHub and DeepSeek. The repository describes itself as: Public results and task definitions for FrontierHarness Eval.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit e837a70. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 16 files in scripts/ (Shell, Python and JavaScript, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
nodegitnpxbashjqghFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
runta.comgithub.comAlso links to:
frontierharness.orgFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
FIREWORKS_API_KEYMOONSHOT_API_KEYOPENROUTER_API_KEYTOGETHER_API_KEYRUNTA_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Frontierharness Eval loads about 8k tokens when it runs. Until then it costs about 162 tokens; SKILL.md has 3,962 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
Without a licence we can't republish the file, so here is its outline and opening line. It has 3,962 words (~8,006 tokens).
“Score a harness that is not in the published FrontierHarness v1.0 set, on the same tasks, runtime, and cost accounting, so the result can be placed next to the twelve baseline configurations in results/eval-data.json.”
SKILL.md and 21 other files (scripts) in skills/frontierharness-eval of frontier-harness-eval/eval.
Open the folder on GitHubat commit e837a70
Frontierharness Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Frontierharness Eval this skillfrontier-harness-eval/eval | 301 | — | ~8k | Automated safety check: Pass | None | |
| Evals Contextzgsm-ai/costrict | 4.4k | 1 repos | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Octocode Benchmark Runnerbgauryy/octocode | 949 | — | ~2.1k | Automated safety check: Pass | MIT | |
| AI Project Copilotsun461941-hub/ai-project-copilot | 100 | — | ~3k | Automated safety check: Pass | MIT | |
| SWE Benchmark Task Adderory/lumen | 307 | — | ~497 | Automated safety check: Pass | Custom licence | |
| Managed Deep Agentslangchain-ai/langchain-skills | 1.3k | — | ~8.7k | Automated safety check: Notes | MIT |
zgsm-ai/costrict
Provides context about the CoStrict evals system structure in this monorepo.
bgauryy/octocode
Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.
sun461941-hub/ai-project-copilot
A skill your agent uses to turn an AI idea or existing repository into a credible open-source product and to run evidence-first repository engineering across codebase discovery, context-efficient…
ory/lumen
Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch.
langchain-ai/langchain-skills
INVOKE THIS SKILL when building, testing, or deploying Managed Deep Agents in LangSmith.
ascend-ai-coding/awesome-ascend-skills
Track daily PRs and Issues from vllm-project/vllm and vllm-project/vllm-ascend, filter by model (DeepSeek/Qwen/GLM/MiniMax/Kimi) and tech topics (PD disaggregation, MTP, quantization, graph mode…
Categories
Benchmark a third-party coding-agent harness against FrontierHarness Eval using Runta runtimes. Frontierharness Eval is an agent skill from frontier-harness-eval/eval. Benchmark a third-party coding-agent harness against FrontierHarness Eval using Runta runtimes.
Frontierharness Eval fits situations like: comparing a coding agent harness; the user mentions FrontierHarness; golden checkpoints; harness trajectories.
Run `npx skills add frontier-harness-eval/eval --skill frontierharness-eval -a claude-code`. Or copy the skill folder (skills/frontierharness-eval in frontier-harness-eval/eval) into .claude/skills/frontierharness-eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add frontier-harness-eval/eval --skill frontierharness-eval -a codex`. Or copy the skill folder (skills/frontierharness-eval in frontier-harness-eval/eval) into .agents/skills/frontierharness-eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add frontier-harness-eval/eval --skill frontierharness-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/frontierharness-eval, .gemini/skills/frontierharness-eval, .github/skills/frontierharness-eval and .opencode/skills/frontierharness-eval in your project.
Going by SKILL.md and its folder, Frontierharness Eval needs a shell, Python and JavaScript for the scripts in its folder, the command-line tools its instructions call (node, git, npx, bash, jq and gh) and credentials named FIREWORKS_API_KEY, MOONSHOT_API_KEY, OPENROUTER_API_KEY and TOGETHER_API_KEY. Our summary lists: Python 3; Node.js; A Bash shell; A credential in FIREWORKS_API_KEY; A credential in MOONSHOT_API_KEY.
SKILL.md names 3 domains. In commands or code: runta.com and github.com; the agent is likely to contact these when it follows the instructions. As links in the text: frontierharness.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
No licence was found for Frontierharness Eval or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.
About 8k tokens (SKILL.md is roughly 32k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Frontierharness Eval: Evals Context (zgsm-ai/costrict, 4.4k stars), Octocode Benchmark Runner (bgauryy/octocode, 949 stars), AI Project Copilot (sun461941-hub/ai-project-copilot, 100 stars) and SWE Benchmark Task Adder (ory/lumen, 307 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
frontier-harness-eval (a GitHub organization) maintains it in frontier-harness-eval/eval, which has 301 GitHub stars. The repository was last updated on September 8, 2026.
Source: frontier-harness-eval/eval on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.