Code Model Evaluation Harness
Orchestra-Research/AI-Research-SKILLs
Benchmarks code generation models with the BigCode Evaluation Harness across HumanEval, MBPP, MultiPL-E and other suites using pass@k metrics.
Inspect a mathematical-modeling workspace, evaluate lean or submission gates per subquestion, update machine-readable manifests, classify change impact, and route one next action without duplicating…
$ npx skills add zhnnky329/MathModeling-skills --skill workflow-orchestrator -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install zhnnky329/MathModeling-skills workflow-orchestrator --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/zhnnky329/MathModeling-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.codex/skills/workflow-orchestrator .claude/skills/workflow-orchestrator && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "workflow-orchestrator" agent skill from https://github.com/zhnnky329/MathModeling-skills/tree/main/.codex/skills/workflow-orchestrator into .claude/skills/workflow-orchestrator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "workflow-orchestrator", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/zhnnky329/MathModeling-skills/tree/main/.codex/skills/workflow-orchestratorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add zhnnky329/MathModeling-skills --skill workflow-orchestrator -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install zhnnky329/MathModeling-skills workflow-orchestrator --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/zhnnky329/MathModeling-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.codex/skills/workflow-orchestrator .agents/skills/workflow-orchestrator && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "workflow-orchestrator" agent skill from https://github.com/zhnnky329/MathModeling-skills/tree/main/.codex/skills/workflow-orchestrator into .agents/skills/workflow-orchestrator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "workflow-orchestrator", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add zhnnky329/MathModeling-skills --skill workflow-orchestrator -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install zhnnky329/MathModeling-skills workflow-orchestrator --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/zhnnky329/MathModeling-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.codex/skills/workflow-orchestrator .cursor/skills/workflow-orchestrator && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "workflow-orchestrator" agent skill from https://github.com/zhnnky329/MathModeling-skills/tree/main/.codex/skills/workflow-orchestrator into .cursor/skills/workflow-orchestrator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "workflow-orchestrator", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/zhnnky329/MathModeling-skills.git --path .codex/skills/workflow-orchestrator--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add zhnnky329/MathModeling-skills --skill workflow-orchestrator -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install zhnnky329/MathModeling-skills workflow-orchestrator --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/zhnnky329/MathModeling-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.codex/skills/workflow-orchestrator .gemini/skills/workflow-orchestrator && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "workflow-orchestrator" agent skill from https://github.com/zhnnky329/MathModeling-skills/tree/main/.codex/skills/workflow-orchestrator into .gemini/skills/workflow-orchestrator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "workflow-orchestrator", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install zhnnky329/MathModeling-skills workflow-orchestratorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add zhnnky329/MathModeling-skills --skill workflow-orchestrator -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/zhnnky329/MathModeling-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/.codex/skills/workflow-orchestrator .github/skills/workflow-orchestrator && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "workflow-orchestrator" agent skill from https://github.com/zhnnky329/MathModeling-skills/tree/main/.codex/skills/workflow-orchestrator into .github/skills/workflow-orchestrator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "workflow-orchestrator", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add zhnnky329/MathModeling-skills --skill workflow-orchestrator -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install zhnnky329/MathModeling-skills workflow-orchestrator --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/zhnnky329/MathModeling-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.codex/skills/workflow-orchestrator .opencode/skills/workflow-orchestrator && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "workflow-orchestrator" agent skill from https://github.com/zhnnky329/MathModeling-skills/tree/main/.codex/skills/workflow-orchestrator into .opencode/skills/workflow-orchestrator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "workflow-orchestrator", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
workflow-orchestratorInspect a mathematical-modeling workspace, evaluate lean or submission gates per subquestion, update machine-readable manifests, classify change impact, and route one next action without duplicating…
Workflow Orchestrator is an agent skill from zhnnky329/MathModeling-skills. Inspect a mathematical-modeling workspace, evaluate lean or submission gates per subquestion, update machine-readable manifests, classify change impact, and route one next action without duplicating downstream work.
Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: 面向数学建模竞赛的 Claude Code / Codex Skills ,支持分阶段建模流程与 Python、MATLAB/北太天元代码分支。 The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 0b46e9c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Workflow Orchestrator loads about 1.8k tokens when it runs. Until then it costs about 59 tokens; SKILL.md has 745 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from zhnnky329/MathModeling-skills at commit 0b46e9c, republished under its MIT licence (© zhnnky329). 745 words, ~1,751 tokens.
.claude/skills/workflow-orchestrator/SKILL.md (or your agent's skills folder).Act as the gate-driven scheduler and state reader. Do not solve models, write model code, or draft paper sections.
../../AGENTS.md is the packaged policy source. Prefer a project-root AGENTS.md when one exists; otherwise read the packaged copy relative to this SKILL.md. Apply that policy without reproducing large reports or dashboards.
Before orchestration in a new workspace:
git status --short;planning/session_config.json, accepting legacy mode.Report warnings concisely. Do not create the full project skeleton unless the user is initializing a project.
Prefer, in order:
planning/manifests/Qx.jsonNever trust a dashboard over newer canonical artifacts.
Maintain one compact JSON manifest per subquestion:
{
"schema_version": 1,
"question_id": "Q1",
"rigor_profile": "lean",
"current_gate": "G2",
"status": "method_screened_waiting_human",
"artifacts": {
"method_card": "methods/Q1/q1_method_card.md",
"decision_ledger": "methods/Q1/q1_decisions.jsonl",
"risk_probe": "methods/Q1/probes/risk_probe_summary.json",
"latest_run": null
},
"allowed": {
"code_generation": false,
"freeze": false,
"paper_writing": false,
"final_assembly": false
},
"blockers": [],
"next_action": {
"owner": "human",
"skill": "decision-prompt-builder",
"reason": "method choice not recorded"
},
"updated_at": "ISO-8601"
}Update only fields affected by the current state change. Generate a human dashboard on request or at a milestone; otherwise derive status directly from manifests.
Evaluate each Qx independently.
Pass when parse, classification, data inventory, success criteria, and human framing exist. A placeholder in a human-owned field blocks the gate.
Pass when:
qx_method_card.md defines a main candidate and usable baseline;risk_probe_summary.json covers applicable checks, including output degeneracy;PASS or justified CONDITIONAL;Do not require a fixed number of candidates, universal PoCs, or a source-line limit.
Pass when qx_decisions.jsonl contains a human DECIDED method choice citing probe evidence. While blocked, allow data preparation but not model code generation.
Pass when:
run_summary.json is complete;Accept legacy Markdown review artifacts during migration, but prefer JSON for new work.
In lean, pass the result-judgment subgate when final-result and stability decisions cite computed evidence. Continue iterating without freezing when the human selects adjust or fallback.
In submission, additionally require:
frozen_numbers.json.Require the three writer rules, frozen-number sourcing, human-confirmed interpretation/claim scope, and verified figures.
Evaluate only in submission. Require passing consistency, completeness, and QA artifacts. Never infer that one auditor covers another.
Choose one primary next action:
data-auditor-cleaner;method-selector;decision-prompt-builder;model-code-analyzer;robustness-checker;Do not invoke several judgment-bearing skills speculatively.
Classify changes before scheduling checks:
NONE: scratch, formatting, comments, non-semantic docs.LOCAL: exploratory code or method-card updates before freeze.CANONICAL: schema/units, symbols, equations, parameters, official values, figure paths.FROZEN: changes affecting frozen values or paper claims.Route checks:
NONE: none.LOCAL: local tests/review.CANONICAL: scoped consistency for affected Qx.FROZEN: thaw log, rerun affected work, re-freeze, scoped consistency.Never schedule a full-workspace consistency audit solely because more than one file changed.
In lean:
In submission:
Read legacy artifacts when new ones are absent:
planning/progress_dashboard.mdqx_method_candidates.mdqx_method_iteration_log.mdqx_decision_log.mddecisions/*_modeler_decision.mdMark them legacy_source in the manifest and recommend migration at the next material edit. Do not regenerate legacy files for new work.
Return a compact state report:
Do not paste a full dashboard or large JSON structure unless the user asks.
© zhnnky329, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .codex/skills/workflow-orchestrator of zhnnky329/MathModeling-skills.
Open the folder on GitHubat commit 0b46e9c
Workflow Orchestrator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Workflow Orchestrator this skillzhnnky329/MathModeling-skills | 1.1k | — | ~1.8k | Automated safety check: Pass | MIT | |
| Code Model Evaluation HarnessOrchestra-Research/AI-Research-SKILLs | 13k | 4 repos | ~2.9k | Automated safety check: Pass | MIT | |
| Model Evaluationawslabs/agent-plugins | 916 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Model Evaluation Metricsjeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~578 | Automated safety check: Pass | MIT | |
| OmniRoute Model Catalogdiegosouzapw/OmniRoute | 75k | — | ~589 | Automated safety check: Pass | MIT | |
| Evaluating Machine Learning Modelsforyourhealth111-pixel/Vibe-Skills | 3.6k | — | ~390 | Automated safety check: Pass | MIT |
Orchestra-Research/AI-Research-SKILLs
Benchmarks code generation models with the BigCode Evaluation Harness across HumanEval, MBPP, MultiPL-E and other suites using pass@k metrics.
awslabs/agent-plugins
Generates python code that evaluates SageMaker models. An agent skill from awslabs/agent-plugins.
jeremylongshore/tons-of-skills-marketplace
Build model evaluation metrics operations. An agent skill from jeremylongshore/tons-of-skills-marketplace.
diegosouzapw/OmniRoute
Looks up which AI models an OmniRoute gateway can reach, creates or updates model aliases and tests whether individual models respond.
foryourhealth111-pixel/Vibe-Skills
Evaluate trained machine learning models with the right metrics and comparison logic.
github/awesome-copilot
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…
zhnnky329/MathModeling-skills
Build and risk-screen a compact role-based method shortlist for a mathematical-modeling subquestion.
zhnnky329/MathModeling-skills
Classify each parsed mathematical-modeling subquestion by required output and structure, surface ambiguous framing trade-offs for human choice, and record primary/secondary task types without…
zhnnky329/MathModeling-skills
Map contest attachments to subquestions, audit and clean raw data, and emit one reusable data profile with quality, coverage, imbalance, concentration, and method-readiness evidence for downstream…
zhnnky329/MathModeling-skills
Build one compact choice card at a genuine mathematical-modeling judgment point.
zhnnky329/MathModeling-skills
Generate and run minimal reproducible MATLAB or Beita Tianyuan compatible code for the human-approved main method and usable baseline, with compact experiment artifacts and a canonical run summary.
zhnnky329/MathModeling-skills
Translate a human-approved main method and usable baseline into a minimal language-neutral implementation and experiment contract.
Inspect a mathematical-modeling workspace, evaluate lean or submission gates per subquestion, update machine-readable manifests, classify change impact, and route one next action without duplicating…. Workflow Orchestrator is an agent skill from zhnnky329/MathModeling-skills. Inspect a mathematical-modeling workspace, evaluate lean or submission gates per subquestion, update machine-readable manifests, classify change impact, and route one next action without duplicating downstream work.
Run `npx skills add zhnnky329/MathModeling-skills --skill workflow-orchestrator -a claude-code`. Or copy the skill folder (.codex/skills/workflow-orchestrator in zhnnky329/MathModeling-skills) into .claude/skills/workflow-orchestrator in your project. Claude Code loads it when a task matches its description.
Run `npx skills add zhnnky329/MathModeling-skills --skill workflow-orchestrator -a codex`. Or copy the skill folder (.codex/skills/workflow-orchestrator in zhnnky329/MathModeling-skills) into .agents/skills/workflow-orchestrator in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zhnnky329/MathModeling-skills --skill workflow-orchestrator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/workflow-orchestrator, .gemini/skills/workflow-orchestrator, .github/skills/workflow-orchestrator and .opencode/skills/workflow-orchestrator in your project.
Going by SKILL.md and its folder, Workflow Orchestrator needs the command-line tools its instructions call (git).
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Workflow Orchestrator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.8k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Workflow Orchestrator: Code Model Evaluation Harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Model Evaluation (awslabs/agent-plugins, 916 stars), Model Evaluation Metrics (jeremylongshore/tons-of-skills-marketplace, 2.8k stars) and OmniRoute Model Catalog (diegosouzapw/OmniRoute, 75k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
zhnnky329 (a GitHub user) maintains it in zhnnky329/MathModeling-skills, which has 1,060 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on September 24, 2026.
Source: zhnnky329/MathModeling-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.