Hugging Face Local Model Evals
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
Agent skill
by Orchestra-Research in Orchestra-Research/AI-Research-SKILLs
Benchmarks code generation models with the BigCode Evaluation Harness across HumanEval, MBPP, MultiPL-E and other suites using pass@k metrics.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill evaluating-code-models -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs evaluating-code-models --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/11-evaluation/bigcode-evaluation-harness .claude/skills/evaluating-code-models && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "evaluating-code-models" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/11-evaluation/bigcode-evaluation-harness into .claude/skills/evaluating-code-models/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluating-code-models", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/11-evaluation/bigcode-evaluation-harnessType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill evaluating-code-models -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs evaluating-code-models --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .agents/skills && cp -r skills-src/11-evaluation/bigcode-evaluation-harness .agents/skills/evaluating-code-models && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "evaluating-code-models" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/11-evaluation/bigcode-evaluation-harness into .agents/skills/evaluating-code-models/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluating-code-models", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill evaluating-code-models -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs evaluating-code-models --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/11-evaluation/bigcode-evaluation-harness .cursor/skills/evaluating-code-models && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "evaluating-code-models" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/11-evaluation/bigcode-evaluation-harness into .cursor/skills/evaluating-code-models/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluating-code-models", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Orchestra-Research/AI-Research-SKILLs.git --path 11-evaluation/bigcode-evaluation-harness--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill evaluating-code-models -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs evaluating-code-models --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/11-evaluation/bigcode-evaluation-harness .gemini/skills/evaluating-code-models && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "evaluating-code-models" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/11-evaluation/bigcode-evaluation-harness into .gemini/skills/evaluating-code-models/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluating-code-models", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Orchestra-Research/AI-Research-SKILLs evaluating-code-modelsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill evaluating-code-models -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .github/skills && cp -r skills-src/11-evaluation/bigcode-evaluation-harness .github/skills/evaluating-code-models && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "evaluating-code-models" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/11-evaluation/bigcode-evaluation-harness into .github/skills/evaluating-code-models/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluating-code-models", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill evaluating-code-models -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs evaluating-code-models --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/11-evaluation/bigcode-evaluation-harness .opencode/skills/evaluating-code-models && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "evaluating-code-models" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/11-evaluation/bigcode-evaluation-harness into .opencode/skills/evaluating-code-models/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluating-code-models", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
evaluating-code-modelsBenchmarks code generation models with the BigCode Evaluation Harness across HumanEval, MBPP, MultiPL-E and other suites using pass@k metrics.
The skill covers the BigCode Evaluation Harness, which evaluates code generation models on 15+ benchmarks. It lists the suites to choose from: HumanEval with 164 handwritten problems and HumanEval+ with 80× more tests, MBPP with 500 crowd-sourced problems and MBPP+ with 399 curated ones, MultiPL-E, which translates HumanEval and MBPP into 18 languages, APPS with 10,000 problems, and DS-1000 with 1,000 data science problems across 7 libraries.
Installation is a git clone followed by pip install -e ., and runs go through accelerate launch main.py with a model such as bigcode/starcoder2-7b and a task name. The first workflow evaluates the standard benchmarks with pass@k estimation at several values of k and writes results to a JSON file. The second, for multi-language runs, generates solutions on the host without executing them and then runs the evaluation in Docker so untrusted code executes safely. Reference files cover benchmarks, custom tasks and known issues.
Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
dockergitpippythonFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comAlso links to:
huggingface.coFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Code Model Evaluation Harness loads about 2.9k tokens when it runs, and up to ~9.8k if it reads all its reference files. Until then it costs about 84 tokens; SKILL.md has 551 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 551 words, ~2,923 tokens.
.claude/skills/evaluating-code-models/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.BigCode Evaluation Harness evaluates code generation models across 15+ benchmarks including HumanEval, MBPP, and MultiPL-E (18 languages).
Installation:
git clone https://github.com/bigcode-project/bigcode-evaluation-harness.git
cd bigcode-evaluation-harness
pip install -e .
accelerate configEvaluate on HumanEval:
accelerate launch main.py \
--model bigcode/starcoder2-7b \
--tasks humaneval \
--max_length_generation 512 \
--temperature 0.2 \
--n_samples 20 \
--batch_size 10 \
--allow_code_execution \
--save_generationsView available tasks:
python -c "from bigcode_eval.tasks import ALL_TASKS; print(ALL_TASKS)"Evaluate model on core code benchmarks (HumanEval, MBPP, HumanEval+).
Checklist:
Code Benchmark Evaluation:
- [ ] Step 1: Choose benchmark suite
- [ ] Step 2: Configure model and generation
- [ ] Step 3: Run evaluation with code execution
- [ ] Step 4: Analyze pass@k resultsStep 1: Choose benchmark suite
Python code generation (most common):
Multi-language (18 languages):
Advanced:
Step 2: Configure model and generation
# Standard HuggingFace model
accelerate launch main.py \
--model bigcode/starcoder2-7b \
--tasks humaneval \
--max_length_generation 512 \
--temperature 0.2 \
--do_sample True \
--n_samples 200 \
--batch_size 50 \
--allow_code_execution
# Quantized model (4-bit)
accelerate launch main.py \
--model codellama/CodeLlama-34b-hf \
--tasks humaneval \
--load_in_4bit \
--max_length_generation 512 \
--allow_code_execution
# Custom/private model
accelerate launch main.py \
--model /path/to/my-code-model \
--tasks humaneval \
--trust_remote_code \
--use_auth_token \
--allow_code_executionStep 3: Run evaluation
# Full evaluation with pass@k estimation (k=1,10,100)
accelerate launch main.py \
--model bigcode/starcoder2-7b \
--tasks humaneval \
--temperature 0.8 \
--n_samples 200 \
--batch_size 50 \
--allow_code_execution \
--save_generations \
--metric_output_path results/starcoder2-humaneval.jsonStep 4: Analyze results
Results in results/starcoder2-humaneval.json:
{
"humaneval": {
"pass@1": 0.354,
"pass@10": 0.521,
"pass@100": 0.689
},
"config": {
"model": "bigcode/starcoder2-7b",
"temperature": 0.8,
"n_samples": 200
}
}Evaluate code generation across 18 programming languages.
Checklist:
Multi-Language Evaluation:
- [ ] Step 1: Generate solutions (host machine)
- [ ] Step 2: Run evaluation in Docker (safe execution)
- [ ] Step 3: Compare across languagesStep 1: Generate solutions on host
# Generate without execution (safe)
accelerate launch main.py \
--model bigcode/starcoder2-7b \
--tasks multiple-py,multiple-js,multiple-java,multiple-cpp \
--max_length_generation 650 \
--temperature 0.8 \
--n_samples 50 \
--batch_size 50 \
--generation_only \
--save_generations \
--save_generations_path generations_multi.jsonStep 2: Evaluate in Docker container
# Pull the MultiPL-E Docker image
docker pull ghcr.io/bigcode-project/evaluation-harness-multiple
# Run evaluation inside container
docker run -v $(pwd)/generations_multi.json:/app/generations.json:ro \
-it evaluation-harness-multiple python3 main.py \
--model bigcode/starcoder2-7b \
--tasks multiple-py,multiple-js,multiple-java,multiple-cpp \
--load_generations_path /app/generations.json \
--allow_code_execution \
--n_samples 50Supported languages: Python, JavaScript, Java, C++, Go, Rust, TypeScript, C#, PHP, Ruby, Swift, Kotlin, Scala, Perl, Julia, Lua, R, Racket
Evaluate chat/instruction models with proper formatting.
Checklist:
Instruction Model Evaluation:
- [ ] Step 1: Use instruction-tuned tasks
- [ ] Step 2: Configure instruction tokens
- [ ] Step 3: Run evaluationStep 1: Choose instruction tasks
Step 2: Configure instruction tokens
# For models with chat templates (e.g., CodeLlama-Instruct)
accelerate launch main.py \
--model codellama/CodeLlama-7b-Instruct-hf \
--tasks instruct-humaneval \
--instruction_tokens "<s>[INST],</s>,[/INST]" \
--max_length_generation 512 \
--allow_code_executionStep 3: HumanEvalPack for instruction models
# Test code synthesis across 6 languages
accelerate launch main.py \
--model codellama/CodeLlama-7b-Instruct-hf \
--tasks humanevalsynthesize-python,humanevalsynthesize-js \
--prompt instruct \
--max_length_generation 512 \
--allow_code_executionBenchmark suite for model comparison.
Step 1: Create evaluation script
#!/bin/bash
# eval_models.sh
MODELS=(
"bigcode/starcoder2-7b"
"codellama/CodeLlama-7b-hf"
"deepseek-ai/deepseek-coder-6.7b-base"
)
TASKS="humaneval,mbpp"
for model in "${MODELS[@]}"; do
model_name=$(echo $model | tr '/' '-')
echo "Evaluating $model"
accelerate launch main.py \
--model $model \
--tasks $TASKS \
--temperature 0.2 \
--n_samples 20 \
--batch_size 20 \
--allow_code_execution \
--metric_output_path results/${model_name}.json
doneStep 2: Generate comparison table
import json
import pandas as pd
models = ["bigcode-starcoder2-7b", "codellama-CodeLlama-7b-hf", "deepseek-ai-deepseek-coder-6.7b-base"]
results = []
for model in models:
with open(f"results/{model}.json") as f:
data = json.load(f)
results.append({
"Model": model,
"HumanEval pass@1": f"{data['humaneval']['pass@1']:.3f}",
"MBPP pass@1": f"{data['mbpp']['pass@1']:.3f}"
})
df = pd.DataFrame(results)
print(df.to_markdown(index=False))Use BigCode Evaluation Harness when:
Use alternatives instead:
| Benchmark | Problems | Languages | Metric | Use Case |
|---|---|---|---|---|
| HumanEval | 164 | Python | pass@k | Standard code completion |
| HumanEval+ | 164 | Python | pass@k | Stricter evaluation (80× tests) |
| MBPP | 500 | Python | pass@k | Entry-level problems |
| MBPP+ | 399 | Python | pass@k | Stricter evaluation (35× tests) |
| MultiPL-E | 164×18 | 18 languages | pass@k | Multi-language evaluation |
| APPS | 10,000 | Python | pass@k | Competition-level |
| DS-1000 | 1,000 | Python | pass@k | Data science (pandas, numpy, etc.) |
| HumanEvalPack | 164×3×6 | 6 languages | pass@k | Synthesis/fix/explain |
| Mercury | 1,889 | Python | Efficiency | Computational efficiency |
Issue: Different results than reported in papers
Check these factors:
# 1. Verify n_samples (need 200 for accurate pass@k)
--n_samples 200
# 2. Check temperature (0.2 for greedy-ish, 0.8 for sampling)
--temperature 0.8
# 3. Verify task name matches exactly
--tasks humaneval # Not "human_eval" or "HumanEval"
# 4. Check max_length_generation
--max_length_generation 512 # Increase for longer problemsIssue: CUDA out of memory
# Use quantization
--load_in_8bit
# OR
--load_in_4bit
# Reduce batch size
--batch_size 1
# Set memory limit
--max_memory_per_gpu "20GiB"Issue: Code execution hangs or times out
Use Docker for safe execution:
# Generate on host (no execution)
--generation_only --save_generations
# Evaluate in Docker
docker run ... --allow_code_execution --load_generations_path ...Issue: Low scores on instruction models
Ensure proper instruction formatting:
# Use instruction-specific tasks
--tasks instruct-humaneval
# Set instruction tokens for your model
--instruction_tokens "<s>[INST],</s>,[/INST]"Issue: MultiPL-E language failures
Use the dedicated Docker image:
docker pull ghcr.io/bigcode-project/evaluation-harness-multiple| Argument | Default | Description |
|---|---|---|
--model | - | HuggingFace model ID or local path |
--tasks | - | Comma-separated task names |
--n_samples | 1 | Samples per problem (200 for pass@k) |
--temperature | 0.2 | Sampling temperature |
--max_length_generation | 512 | Max tokens (prompt + generation) |
--batch_size | 1 | Batch size per GPU |
--allow_code_execution | False | Enable code execution (required) |
--generation_only | False | Generate without evaluation |
--load_generations_path | - | Load pre-generated solutions |
--save_generations | False | Save generated code |
--metric_output_path | results.json | Output file for metrics |
--load_in_8bit | False | 8-bit quantization |
--load_in_4bit | False | 4-bit quantization |
--trust_remote_code | False | Allow custom model code |
--precision | fp32 | Model precision (fp32/fp16/bf16) |
| Model Size | VRAM (fp16) | VRAM (4-bit) | Time (HumanEval, n=200) |
|---|---|---|---|
| 7B | 14GB | 6GB | ~30 min (A100) |
| 13B | 26GB | 10GB | ~1 hour (A100) |
| 34B | 68GB | 20GB | ~2 hours (A100) |
© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in 11-evaluation/bigcode-evaluation-harness of Orchestra-Research/AI-Research-SKILLs.
Open the folder on GitHubat commit 773a529
We found 7 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 5 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.
Code Model Evaluation Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Code Model Evaluation Harness this skillOrchestra-Research/AI-Research-SKILLs | 13k | 5 repos | ~2.9k | Automated safety check: Pass | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Best Model Finderhuggingface/skills | 11k | 2 repos | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Generate Openenv Envadithya-s-k/FineEnvs | 421 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Huggingface Spaceshuggingface/skills | 11k | 1 repos | ~4.4k | Automated safety check: Pass | Apache-2.0 | |
| Deploy To Hf Spacesadithya-s-k/FineEnvs | 421 | — | ~864 | Automated safety check: Pass | Apache-2.0 |
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
huggingface/skills
Finds top models for a task from official Hugging Face benchmark leaderboards, filters them to what fits your hardware, and returns a comparison table with scores.
adithya-s-k/FineEnvs
Builds an OpenEnv (Hugging Face) variant of an RL environment.
huggingface/skills
Build, deploy, and maintain applications on Hugging Face Spaces — Gradio / Docker / Static SDKs, ZeroGPU and dedicated hardware, model loading, debugging, buckets, inference providers, community…
adithya-s-k/FineEnvs
Deploy the article to a Hugging Face Space. An agent skill from adithya-s-k/FineEnvs.
sickn33/agentic-awesome-skills
Run evaluations for Hugging Face Hub models using inspect-ai and lighteval on local hardware.
Orchestra-Research/AI-Research-SKILLs
Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.
Orchestra-Research/AI-Research-SKILLs
Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
Orchestra-Research/AI-Research-SKILLs
Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Works with
Categories
Benchmarks code generation models with the BigCode Evaluation Harness across HumanEval, MBPP, MultiPL-E and other suites using pass@k metrics. The skill covers the BigCode Evaluation Harness, which evaluates code generation models on 15+ benchmarks. It lists the suites to choose from: HumanEval with 164 handwritten problems and HumanEval+ with 80× more tests, MBPP with 500 crowd-sourced problems and MBPP+ with 399 curated ones, MultiPL-E, which translates HumanEval and MBPP into 18 languages, APPS with 10,000 problems, and DS-1000 with 1,000 data science problems across 7 libraries.
Code Model Evaluation Harness fits situations like: benchmarking a code generation model on HumanEval or MBPP; comparing the coding ability of several models with pass@k; testing multi-language code generation with MultiPL-E; adding a custom evaluation task.
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill evaluating-code-models -a claude-code`. Or copy the skill folder (11-evaluation/bigcode-evaluation-harness in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/evaluating-code-models in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill evaluating-code-models -a codex`. Or copy the skill folder (11-evaluation/bigcode-evaluation-harness in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/evaluating-code-models in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill evaluating-code-models -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluating-code-models, .gemini/skills/evaluating-code-models, .github/skills/evaluating-code-models and .opencode/skills/evaluating-code-models in your project.
Going by SKILL.md and its folder, Code Model Evaluation Harness needs the command-line tools its instructions call (docker, git, pip and python). Our summary lists: The bigcode-evaluation-harness repository installed with pip install -e .; A model to evaluate, such as a Hugging Face checkpoint; Docker, for safe execution of multi-language runs.
SKILL.md names 2 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: huggingface.co. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Code Model Evaluation Harness is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Code Model Evaluation Harness: Hugging Face Local Model Evals (huggingface/skills, 11k stars), Hugging Face Best Model Finder (huggingface/skills, 11k stars), Generate Openenv Env (adithya-s-k/FineEnvs, 421 stars) and Huggingface Spaces (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,313 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.
Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.