Agent skill

Code Model Evaluation Harness

by Orchestra-Research in Orchestra-Research/AI-Research-SKILLs

Benchmarks code generation models with the BigCode Evaluation Harness across HumanEval, MBPP, MultiPL-E and other suites using pass@k metrics.

MITAuto-check passedAI & LLM Engineering

Install Code Model Evaluation Harness

skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill evaluating-code-models -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orchestra-Research/AI-Research-SKILLs evaluating-code-models --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/11-evaluation/bigcode-evaluation-harness .claude/skills/evaluating-code-models && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluating-code-models
GitHub stars
13k
Used in
5 other repos
Token cost
~2.9k tokens
SKILL.md length
551 words
Files
4 (incl. references)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

Benchmarks code generation models with the BigCode Evaluation Harness across HumanEval, MBPP, MultiPL-E and other suites using pass@k metrics.

  • Benchmarking a code generation model on HumanEval or MBPP
  • SKILL.md covers Quick Start, Common Workflows, When to Use vs Alternatives and Supported Benchmarks, plus 4 more sections
  • Calls docker, git and pip; reaches github.com
  • Comparing the coding ability of several models with pass@k

What it does

The skill covers the BigCode Evaluation Harness, which evaluates code generation models on 15+ benchmarks. It lists the suites to choose from: HumanEval with 164 handwritten problems and HumanEval+ with 80× more tests, MBPP with 500 crowd-sourced problems and MBPP+ with 399 curated ones, MultiPL-E, which translates HumanEval and MBPP into 18 languages, APPS with 10,000 problems, and DS-1000 with 1,000 data science problems across 7 libraries.

Installation is a git clone followed by pip install -e ., and runs go through accelerate launch main.py with a model such as bigcode/starcoder2-7b and a task name. The first workflow evaluates the standard benchmarks with pass@k estimation at several values of k and writes results to a JSON file. The second, for multi-language runs, generates solutions on the host without executing them and then runs the evaluation in Docker so untrusted code executes safely. Reference files cover benchmarks, custom tasks and known issues.

When your agent uses it

  • Benchmarking a code generation model on HumanEval or MBPP
  • Comparing the coding ability of several models with pass@k
  • Testing multi-language code generation with MultiPL-E
  • Adding a custom evaluation task

Example prompts

  • “Evaluate bigcode/starcoder2-7b on HumanEval and report pass@1.”
  • “Run MultiPL-E for Rust and Go, generating on the host and executing inside Docker.”
  • “Compare my fine-tuned checkpoint against the base model on MBPP and HumanEval+.”

Requirements

  • The bigcode-evaluation-harness repository installed with pip install -e .
  • A model to evaluate, such as a Hugging Face checkpoint
  • Docker, for safe execution of multi-language runs

What it can do on your machine

Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker
    • git
    • pip
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    Also links to:

    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Code Model Evaluation Harness loads about 2.9k tokens when it runs, and up to ~9.8k if it reads all its reference files. Until then it costs about 84 tokens; SKILL.md has 551 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 551 words, ~2,923 tokens.

Download SKILL.mdSave it as .claude/skills/evaluating-code-models/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
evaluating-code-models
description
Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities, testing multi-language support, or measuring code generation quality. Industry standard from BigCode Project used by HuggingFace leaderboards.
version
1.0.0
author
Orchestra Research
license
MIT
tags
Evaluation, Code Generation, HumanEval, MBPP, MultiPL-E, Pass@k, BigCode, Benchmarking, Code Models
dependencies
bigcode-evaluation-harness, transformers>=4.25.1, accelerate>=0.13.2, datasets>=2.6.1

BigCode Evaluation Harness - Code Model Benchmarking

Quick Start

BigCode Evaluation Harness evaluates code generation models across 15+ benchmarks including HumanEval, MBPP, and MultiPL-E (18 languages).

Installation:

bash
git clone https://github.com/bigcode-project/bigcode-evaluation-harness.git
cd bigcode-evaluation-harness
pip install -e .
accelerate config

Evaluate on HumanEval:

bash
accelerate launch main.py \
  --model bigcode/starcoder2-7b \
  --tasks humaneval \
  --max_length_generation 512 \
  --temperature 0.2 \
  --n_samples 20 \
  --batch_size 10 \
  --allow_code_execution \
  --save_generations

View available tasks:

bash
python -c "from bigcode_eval.tasks import ALL_TASKS; print(ALL_TASKS)"

Common Workflows

Workflow 1: Standard Code Benchmark Evaluation

Evaluate model on core code benchmarks (HumanEval, MBPP, HumanEval+).

Checklist:

Code Benchmark Evaluation:
- [ ] Step 1: Choose benchmark suite
- [ ] Step 2: Configure model and generation
- [ ] Step 3: Run evaluation with code execution
- [ ] Step 4: Analyze pass@k results

Step 1: Choose benchmark suite

Python code generation (most common):

  • HumanEval: 164 handwritten problems, function completion
  • HumanEval+: Same 164 problems with 80× more tests (stricter)
  • MBPP: 500 crowd-sourced problems, entry-level difficulty
  • MBPP+: 399 curated problems with 35× more tests

Multi-language (18 languages):

  • MultiPL-E: HumanEval/MBPP translated to C++, Java, JavaScript, Go, Rust, etc.

Advanced:

  • APPS: 10,000 problems (introductory/interview/competition)
  • DS-1000: 1,000 data science problems across 7 libraries

Step 2: Configure model and generation

bash
# Standard HuggingFace model
accelerate launch main.py \
  --model bigcode/starcoder2-7b \
  --tasks humaneval \
  --max_length_generation 512 \
  --temperature 0.2 \
  --do_sample True \
  --n_samples 200 \
  --batch_size 50 \
  --allow_code_execution

# Quantized model (4-bit)
accelerate launch main.py \
  --model codellama/CodeLlama-34b-hf \
  --tasks humaneval \
  --load_in_4bit \
  --max_length_generation 512 \
  --allow_code_execution

# Custom/private model
accelerate launch main.py \
  --model /path/to/my-code-model \
  --tasks humaneval \
  --trust_remote_code \
  --use_auth_token \
  --allow_code_execution

Step 3: Run evaluation

bash
# Full evaluation with pass@k estimation (k=1,10,100)
accelerate launch main.py \
  --model bigcode/starcoder2-7b \
  --tasks humaneval \
  --temperature 0.8 \
  --n_samples 200 \
  --batch_size 50 \
  --allow_code_execution \
  --save_generations \
  --metric_output_path results/starcoder2-humaneval.json

Step 4: Analyze results

Results in results/starcoder2-humaneval.json:

json
{
  "humaneval": {
    "pass@1": 0.354,
    "pass@10": 0.521,
    "pass@100": 0.689
  },
  "config": {
    "model": "bigcode/starcoder2-7b",
    "temperature": 0.8,
    "n_samples": 200
  }
}
Workflow 2: Multi-Language Evaluation (MultiPL-E)

Evaluate code generation across 18 programming languages.

Checklist:

Multi-Language Evaluation:
- [ ] Step 1: Generate solutions (host machine)
- [ ] Step 2: Run evaluation in Docker (safe execution)
- [ ] Step 3: Compare across languages

Step 1: Generate solutions on host

bash
# Generate without execution (safe)
accelerate launch main.py \
  --model bigcode/starcoder2-7b \
  --tasks multiple-py,multiple-js,multiple-java,multiple-cpp \
  --max_length_generation 650 \
  --temperature 0.8 \
  --n_samples 50 \
  --batch_size 50 \
  --generation_only \
  --save_generations \
  --save_generations_path generations_multi.json

Step 2: Evaluate in Docker container

bash
# Pull the MultiPL-E Docker image
docker pull ghcr.io/bigcode-project/evaluation-harness-multiple

# Run evaluation inside container
docker run -v $(pwd)/generations_multi.json:/app/generations.json:ro \
  -it evaluation-harness-multiple python3 main.py \
  --model bigcode/starcoder2-7b \
  --tasks multiple-py,multiple-js,multiple-java,multiple-cpp \
  --load_generations_path /app/generations.json \
  --allow_code_execution \
  --n_samples 50

Supported languages: Python, JavaScript, Java, C++, Go, Rust, TypeScript, C#, PHP, Ruby, Swift, Kotlin, Scala, Perl, Julia, Lua, R, Racket

Workflow 3: Instruction-Tuned Model Evaluation

Evaluate chat/instruction models with proper formatting.

Checklist:

Instruction Model Evaluation:
- [ ] Step 1: Use instruction-tuned tasks
- [ ] Step 2: Configure instruction tokens
- [ ] Step 3: Run evaluation

Step 1: Choose instruction tasks

  • instruct-humaneval: HumanEval with instruction prompts
  • humanevalsynthesize-{lang}: HumanEvalPack synthesis tasks

Step 2: Configure instruction tokens

bash
# For models with chat templates (e.g., CodeLlama-Instruct)
accelerate launch main.py \
  --model codellama/CodeLlama-7b-Instruct-hf \
  --tasks instruct-humaneval \
  --instruction_tokens "<s>[INST],</s>,[/INST]" \
  --max_length_generation 512 \
  --allow_code_execution

Step 3: HumanEvalPack for instruction models

bash
# Test code synthesis across 6 languages
accelerate launch main.py \
  --model codellama/CodeLlama-7b-Instruct-hf \
  --tasks humanevalsynthesize-python,humanevalsynthesize-js \
  --prompt instruct \
  --max_length_generation 512 \
  --allow_code_execution
Workflow 4: Compare Multiple Models

Benchmark suite for model comparison.

Step 1: Create evaluation script

bash
#!/bin/bash
# eval_models.sh

MODELS=(
  "bigcode/starcoder2-7b"
  "codellama/CodeLlama-7b-hf"
  "deepseek-ai/deepseek-coder-6.7b-base"
)
TASKS="humaneval,mbpp"

for model in "${MODELS[@]}"; do
  model_name=$(echo $model | tr '/' '-')
  echo "Evaluating $model"

  accelerate launch main.py \
    --model $model \
    --tasks $TASKS \
    --temperature 0.2 \
    --n_samples 20 \
    --batch_size 20 \
    --allow_code_execution \
    --metric_output_path results/${model_name}.json
done

Step 2: Generate comparison table

python
import json
import pandas as pd

models = ["bigcode-starcoder2-7b", "codellama-CodeLlama-7b-hf", "deepseek-ai-deepseek-coder-6.7b-base"]
results = []

for model in models:
    with open(f"results/{model}.json") as f:
        data = json.load(f)
        results.append({
            "Model": model,
            "HumanEval pass@1": f"{data['humaneval']['pass@1']:.3f}",
            "MBPP pass@1": f"{data['mbpp']['pass@1']:.3f}"
        })

df = pd.DataFrame(results)
print(df.to_markdown(index=False))

When to Use vs Alternatives

Use BigCode Evaluation Harness when:

  • Evaluating code generation models specifically
  • Need multi-language evaluation (18 languages via MultiPL-E)
  • Testing functional correctness with unit tests (pass@k)
  • Benchmarking for BigCode/HuggingFace leaderboards
  • Evaluating fill-in-the-middle (FIM) capabilities

Use alternatives instead:

  • lm-evaluation-harness: General LLM benchmarks (MMLU, GSM8K, HellaSwag)
  • EvalPlus: Stricter HumanEval+/MBPP+ with more test cases
  • SWE-bench: Real-world GitHub issue resolution
  • LiveCodeBench: Contamination-free, continuously updated problems
  • CodeXGLUE: Code understanding tasks (clone detection, defect prediction)
Show full SKILL.md (245 more words)Show less

Supported Benchmarks

BenchmarkProblemsLanguagesMetricUse Case
HumanEval164Pythonpass@kStandard code completion
HumanEval+164Pythonpass@kStricter evaluation (80× tests)
MBPP500Pythonpass@kEntry-level problems
MBPP+399Pythonpass@kStricter evaluation (35× tests)
MultiPL-E164×1818 languagespass@kMulti-language evaluation
APPS10,000Pythonpass@kCompetition-level
DS-10001,000Pythonpass@kData science (pandas, numpy, etc.)
HumanEvalPack164×3×66 languagespass@kSynthesis/fix/explain
Mercury1,889PythonEfficiencyComputational efficiency

Common Issues

Issue: Different results than reported in papers

Check these factors:

bash
# 1. Verify n_samples (need 200 for accurate pass@k)
--n_samples 200

# 2. Check temperature (0.2 for greedy-ish, 0.8 for sampling)
--temperature 0.8

# 3. Verify task name matches exactly
--tasks humaneval  # Not "human_eval" or "HumanEval"

# 4. Check max_length_generation
--max_length_generation 512  # Increase for longer problems

Issue: CUDA out of memory

bash
# Use quantization
--load_in_8bit
# OR
--load_in_4bit

# Reduce batch size
--batch_size 1

# Set memory limit
--max_memory_per_gpu "20GiB"

Issue: Code execution hangs or times out

Use Docker for safe execution:

bash
# Generate on host (no execution)
--generation_only --save_generations

# Evaluate in Docker
docker run ... --allow_code_execution --load_generations_path ...

Issue: Low scores on instruction models

Ensure proper instruction formatting:

bash
# Use instruction-specific tasks
--tasks instruct-humaneval

# Set instruction tokens for your model
--instruction_tokens "<s>[INST],</s>,[/INST]"

Issue: MultiPL-E language failures

Use the dedicated Docker image:

bash
docker pull ghcr.io/bigcode-project/evaluation-harness-multiple

Command Reference

ArgumentDefaultDescription
--model-HuggingFace model ID or local path
--tasks-Comma-separated task names
--n_samples1Samples per problem (200 for pass@k)
--temperature0.2Sampling temperature
--max_length_generation512Max tokens (prompt + generation)
--batch_size1Batch size per GPU
--allow_code_executionFalseEnable code execution (required)
--generation_onlyFalseGenerate without evaluation
--load_generations_path-Load pre-generated solutions
--save_generationsFalseSave generated code
--metric_output_pathresults.jsonOutput file for metrics
--load_in_8bitFalse8-bit quantization
--load_in_4bitFalse4-bit quantization
--trust_remote_codeFalseAllow custom model code
--precisionfp32Model precision (fp32/fp16/bf16)

Hardware Requirements

Model SizeVRAM (fp16)VRAM (4-bit)Time (HumanEval, n=200)
7B14GB6GB~30 min (A100)
13B26GB10GB~1 hour (A100)
34B68GB20GB~2 hours (A100)

Resources

© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in 11-evaluation/bigcode-evaluation-harness of Orchestra-Research/AI-Research-SKILLs.

  • SKILL.md
  • references/benchmarks.md
  • references/custom-tasks.md
  • references/issues.md

Open the folder on GitHubat commit 773a529

Used in 5 other repositories

We found 7 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 5 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Code Model Evaluation Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Code Model Evaluation Harness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Code Model Evaluation Harness this skillOrchestra-Research/AI-Research-SKILLs13k5 repos~2.9kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Hugging Face Best Model Finderhuggingface/skills11k2 repos~1.5kAutomated safety check: PassApache-2.0
Generate Openenv Envadithya-s-k/FineEnvs421—~2.4kAutomated safety check: PassApache-2.0
Huggingface Spaceshuggingface/skills11k1 repos~4.4kAutomated safety check: PassApache-2.0
Deploy To Hf Spacesadithya-s-k/FineEnvs421—~864Automated safety check: PassApache-2.0

Similar skills

  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Finds top models for a task from official Hugging Face benchmark leaderboards, filters them to what fits your hardware, and returns a comparison table with scores.

    11k GitHub starsUsed in 2 repos~1.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Generate Openenv Env

    adithya-s-k/FineEnvs

    Builds an OpenEnv (Hugging Face) variant of an RL environment.

    421 GitHub stars~2.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Huggingface Spaces

    huggingface/skills

    Official

    Build, deploy, and maintain applications on Hugging Face Spaces — Gradio / Docker / Static SDKs, ZeroGPU and dedicated hardware, model loading, debugging, buckets, inference providers, community…

    11k GitHub starsUsed in 1 repo~4.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Deploy To Hf Spaces

    adithya-s-k/FineEnvs

    Deploy the article to a Hugging Face Space. An agent skill from adithya-s-k/FineEnvs.

    421 GitHub stars~864 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Community Evals

    sickn33/agentic-awesome-skills

    Run evaluations for Hugging Face Hub models using inspect-ai and lighteval on local hardware.

    47k GitHub starsUsed in 1 repo~1.9k tokens
    AI & LLM EngineeringAuto-check passed

More from Orchestra-Research/AI-Research-SKILLs

All 96 skills in this repo
  • AudioCraft Audio Generation

    Orchestra-Research/AI-Research-SKILLs

    Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.

    13k GitHub starsUsed in 9 repos~3.9k tokens
    Auto-check passed
  • Peft Fine Tuning

    Orchestra-Research/AI-Research-SKILLs

    Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.

    13k GitHub starsUsed in 9 repos~3.1k tokens
    Auto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 9 repos~3.3k tokens
    Auto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    Auto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 8 repos~1.7k tokens
    Auto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    Auto-check passed

Questions about Code Model Evaluation Harness

What does Code Model Evaluation Harness do?

Benchmarks code generation models with the BigCode Evaluation Harness across HumanEval, MBPP, MultiPL-E and other suites using pass@k metrics. The skill covers the BigCode Evaluation Harness, which evaluates code generation models on 15+ benchmarks. It lists the suites to choose from: HumanEval with 164 handwritten problems and HumanEval+ with 80× more tests, MBPP with 500 crowd-sourced problems and MBPP+ with 399 curated ones, MultiPL-E, which translates HumanEval and MBPP into 18 languages, APPS with 10,000 problems, and DS-1000 with 1,000 data science problems across 7 libraries.

When should I use Code Model Evaluation Harness?

Code Model Evaluation Harness fits situations like: benchmarking a code generation model on HumanEval or MBPP; comparing the coding ability of several models with pass@k; testing multi-language code generation with MultiPL-E; adding a custom evaluation task.

How do I install Code Model Evaluation Harness in Claude Code?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill evaluating-code-models -a claude-code`. Or copy the skill folder (11-evaluation/bigcode-evaluation-harness in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/evaluating-code-models in your project. Claude Code loads it when a task matches its description.

How do I install Code Model Evaluation Harness in Codex?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill evaluating-code-models -a codex`. Or copy the skill folder (11-evaluation/bigcode-evaluation-harness in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/evaluating-code-models in your project. Codex loads it when a task matches its description.

Can I use Code Model Evaluation Harness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill evaluating-code-models -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluating-code-models, .gemini/skills/evaluating-code-models, .github/skills/evaluating-code-models and .opencode/skills/evaluating-code-models in your project.

What does Code Model Evaluation Harness need to run?

Going by SKILL.md and its folder, Code Model Evaluation Harness needs the command-line tools its instructions call (docker, git, pip and python). Our summary lists: The bigcode-evaluation-harness repository installed with pip install -e .; A model to evaluate, such as a Hugging Face checkpoint; Docker, for safe execution of multi-language runs.

Does Code Model Evaluation Harness access the network?

SKILL.md names 2 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: huggingface.co. This is read from the text; nothing was executed.

Is Code Model Evaluation Harness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Code Model Evaluation Harness use?

Code Model Evaluation Harness is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Code Model Evaluation Harness use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.9k tokens, read only when the agent opens those files.

What are the alternatives to Code Model Evaluation Harness?

Skills that share tags, products or a category with Code Model Evaluation Harness: Hugging Face Local Model Evals (huggingface/skills, 11k stars), Hugging Face Best Model Finder (huggingface/skills, 11k stars), Generate Openenv Env (adithya-s-k/FineEnvs, 421 stars) and Huggingface Spaces (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Code Model Evaluation Harness?

Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,313 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.

Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.