LLM Benchmarking with lm-evaluation-harness
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Step-by-step guide for adding a new built-in test evaluator provider to Agent Kernel (beyond DeepEval, Opik and JEV).
$ npx skills add yaalalabs/agent-kernel --skill ak-dev-new-evaluator-provider -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install yaalalabs/agent-kernel ak-dev-new-evaluator-provider --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/yaalalabs/agent-kernel.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/ak-dev-new-evaluator-provider .claude/skills/ak-dev-new-evaluator-provider && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ak-dev-new-evaluator-provider" agent skill from https://github.com/yaalalabs/agent-kernel/tree/develop/.agents/skills/ak-dev-new-evaluator-provider into .claude/skills/ak-dev-new-evaluator-provider/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ak-dev-new-evaluator-provider", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/yaalalabs/agent-kernel/tree/develop/.agents/skills/ak-dev-new-evaluator-providerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add yaalalabs/agent-kernel --skill ak-dev-new-evaluator-provider -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install yaalalabs/agent-kernel ak-dev-new-evaluator-provider --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yaalalabs/agent-kernel.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/ak-dev-new-evaluator-provider .agents/skills/ak-dev-new-evaluator-provider && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ak-dev-new-evaluator-provider" agent skill from https://github.com/yaalalabs/agent-kernel/tree/develop/.agents/skills/ak-dev-new-evaluator-provider into .agents/skills/ak-dev-new-evaluator-provider/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ak-dev-new-evaluator-provider", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add yaalalabs/agent-kernel --skill ak-dev-new-evaluator-provider -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install yaalalabs/agent-kernel ak-dev-new-evaluator-provider --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yaalalabs/agent-kernel.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/ak-dev-new-evaluator-provider .cursor/skills/ak-dev-new-evaluator-provider && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ak-dev-new-evaluator-provider" agent skill from https://github.com/yaalalabs/agent-kernel/tree/develop/.agents/skills/ak-dev-new-evaluator-provider into .cursor/skills/ak-dev-new-evaluator-provider/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ak-dev-new-evaluator-provider", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/yaalalabs/agent-kernel.git --path .agents/skills/ak-dev-new-evaluator-provider--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add yaalalabs/agent-kernel --skill ak-dev-new-evaluator-provider -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install yaalalabs/agent-kernel ak-dev-new-evaluator-provider --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yaalalabs/agent-kernel.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/ak-dev-new-evaluator-provider .gemini/skills/ak-dev-new-evaluator-provider && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ak-dev-new-evaluator-provider" agent skill from https://github.com/yaalalabs/agent-kernel/tree/develop/.agents/skills/ak-dev-new-evaluator-provider into .gemini/skills/ak-dev-new-evaluator-provider/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ak-dev-new-evaluator-provider", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install yaalalabs/agent-kernel ak-dev-new-evaluator-providerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add yaalalabs/agent-kernel --skill ak-dev-new-evaluator-provider -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/yaalalabs/agent-kernel.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/ak-dev-new-evaluator-provider .github/skills/ak-dev-new-evaluator-provider && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ak-dev-new-evaluator-provider" agent skill from https://github.com/yaalalabs/agent-kernel/tree/develop/.agents/skills/ak-dev-new-evaluator-provider into .github/skills/ak-dev-new-evaluator-provider/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ak-dev-new-evaluator-provider", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add yaalalabs/agent-kernel --skill ak-dev-new-evaluator-provider -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install yaalalabs/agent-kernel ak-dev-new-evaluator-provider --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yaalalabs/agent-kernel.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/ak-dev-new-evaluator-provider .opencode/skills/ak-dev-new-evaluator-provider && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ak-dev-new-evaluator-provider" agent skill from https://github.com/yaalalabs/agent-kernel/tree/develop/.agents/skills/ak-dev-new-evaluator-provider into .opencode/skills/ak-dev-new-evaluator-provider/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ak-dev-new-evaluator-provider", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ak-dev-new-evaluator-providerStep-by-step guide for adding a new built-in test evaluator provider to Agent Kernel (beyond DeepEval, Opik and JEV).
Ak Dev New Evaluator Provider is an agent skill from yaalalabs/agent-kernel. Step-by-step guide for adding a new built-in test evaluator provider to Agent Kernel (beyond DeepEval, Opik and JEV). Use this skill when you need to give the test framework's pluggable AKEvaluator interface a new first-party scoring/judge backend addressable by a short config name (e.g. "trulens"), not a one-off bring-your-own evaluator. Covers implementing score-based and LLM-as-judge evaluation, factory registration, configuration, optional dependencies, and testing.
Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: The Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or… The licence is Apache-2.0.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 97fa8d9. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python, toml and yaml).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ak Dev New Evaluator Provider loads about 3.4k tokens when it runs. Until then it costs about 126 tokens; SKILL.md has 1,054 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from yaalalabs/agent-kernel at commit 97fa8d9, republished under its Apache-2.0 licence (© yaalalabs). 1,054 words, ~3,431 tokens.
.claude/skills/ak-dev-new-evaluator-provider/SKILL.md (or your agent's skills folder).This guide walks through adding a new built-in evaluator provider to Agent Kernel's test
framework. Use the existing DeepEval implementation
(ak-py/src/agentkernel/test/core/evaluator/deepeval.py) as reference.
Before starting, check whether you actually need this skill: if the evaluator only needs to exist
for your own project (not addressable by every AK user via a short built-in name), you don't
need any of the steps below — just subclass AKEvaluator anywhere importable and point
test-config.yaml's evaluator: at its dotted path. That's the "bring your own evaluator"
path described in the user-facing ak-test skill and
docs/docs/testing/cli-testing.md;
examples/cli/custom-evaluator/ is a complete worked example of it. This skill is only for adding
a first-party, in-repo provider that ships with AK and gets its own short type name.
| Provider | Short name | Scoring mode | LLM-judge mode | Extra |
|---|---|---|---|---|
| DeepEval | deepeval | Scorer.quasi_exact_match_score (whole-string, normalised) | GEval LLM-as-judge metric | agentkernel[test] |
| Opik | opik | LevenshteinRatio (fuzzy string similarity) | GEval LLM-as-judge metric | agentkernel[opik] |
| JEV | jev | — (AKMetricNotSupported) | TypeSafe Noul (yes/no probability) | agentkernel[jev] |
AKEvaluator (ak-py/src/agentkernel/test/core/evaluator/base.py) is the abstract base
every evaluator — built-in or bring-your-own — implements. It has exactly two abstract methods:evaluate_by_score(case: AKEvaluationCase) -> AKEvaluationResult — deterministic scoring,
no LLM callevaluate_by_llm(case: AKEvaluationCase) -> AKEvaluationResult — LLM-as-judge scoringAKEvaluationCase carries the comparison inputs (user_input, actual, expected,
threshold, context, criteria); AKEvaluationResult carries the outcome (score,
passed, metric, evaluator, reason, cost, attempts, metadata).AKMissingInput when a field the requested metric needs (e.g. case.expected) wasn't
supplied.AKMetricNotSupported from whichever of the two methods your backend structurally
cannot implement (e.g. a pure LLM-judge service has no offline scoring mode).AKEvaluationError when a configured backend fails to produce a score (missing
credentials, transport error, unparseable judge output). Never raise AssertionError and
never silently return a 0.0 to stand in for a failure — 0.0 must only ever mean "scored
zero", not "couldn't be scored". Test.compare is the only place that decides pass/fail
fatality; evaluators only ever set result.passed.Test._resolve_evaluator_class in ak-py/src/agentkernel/test/test.py is the factory. It
shares the same pluggable-backend shape as guardrails, sandbox providers, and trace backends
(core/util/factory.py's resolve_dotted/require_extra/AKConfigError): an if-per-built-in
branch with the SDK import wrapped in require_extra (actionable ImportError naming the pip
extra if missing), then a dotted-path bring-your-own fallback for anything else.Test._evaluator (keyed by the configured value),
guarded by Test._evaluator_lock — construction happens once per distinct evaluator: config
value, not once per Test.compare call.Create ak-py/src/agentkernel/test/core/evaluator/<provider>.py. Keep the provider's SDK imports
inside this file only — test/core/evaluator/__init__.py and base.py stay pure Python with no
optional-dependency imports at module level, so importing the AKEvaluator interface never
requires your provider's SDK to be installed.
# ak-py/src/agentkernel/test/core/evaluator/<provider>.py
from agentkernel.test.config import AKTestConfig
from .base import AKEvaluationCase, AKEvaluationError, AKEvaluationResult, AKEvaluator, AKMissingInput
class <Provider>AKEvaluator(AKEvaluator):
def __init__(self, config: AKTestConfig) -> None:
super().__init__(config)
# Lazy-init any client/model here only if evaluate_by_score never needs it
# (mirrors DeepevalAKEvaluator's lazy LiteLLMModel, built only on first evaluate_by_llm call).
def evaluate_by_score(self, case: AKEvaluationCase) -> AKEvaluationResult:
if not case.expected:
raise AKMissingInput("evaluate_by_score requires AKEvaluationCase.expected")
# Deterministic, offline scoring logic here.
score = ... # float
return AKEvaluationResult(
metric="<metric_name>",
evaluator="<provider>",
score=score,
passed=score >= case.threshold,
)
def evaluate_by_llm(self, case: AKEvaluationCase) -> AKEvaluationResult:
if not case.expected:
raise AKMissingInput("evaluate_by_llm requires AKEvaluationCase.expected")
try:
score = ... # call the judge
except Exception as exc:
raise AKEvaluationError(f"<provider> llm-based evaluation failed: {exc}") from exc
return AKEvaluationResult(
metric="<metric_name>",
evaluator="<provider>",
score=score,
reason=..., # judge's explanation, if the backend provides one
passed=score is not None and score >= case.threshold,
)If a mode genuinely doesn't apply to your backend (e.g. a provider that is LLM-judge-only), raise
AKMetricNotSupported from that method instead of faking a result. Test.compare does not catch
it — in fallback mode it propagates out of evaluate_by_score before evaluate_by_llm runs — so
document that users of your provider must set the matching mode (e.g. JEV requires mode: llm).
Add the short name to _BUILTIN_EVALUATORS and a branch in Test._resolve_evaluator_class, both
in ak-py/src/agentkernel/test/test.py:
_BUILTIN_EVALUATORS = ["deepeval", "opik", "jev", "<provider>"] # ADD THIS
class Test:
...
@classmethod
def _resolve_evaluator_class(cls, configured: str) -> type[AKEvaluator]:
if configured == "deepeval":
with require_extra("test", "evaluator: deepeval"):
from .core.evaluator.deepeval import DeepevalAKEvaluator
return DeepevalAKEvaluator
if configured == "opik":
with require_extra("opik", "evaluator: opik"):
from .core.evaluator.opik import OpikAKEvaluator
return OpikAKEvaluator
if configured == "jev":
with require_extra("jev", "evaluator: jev"):
from .core.evaluator.jev import JevAKEvaluator
return JevAKEvaluator
if configured == "<provider>": # ADD THIS
with require_extra("<provider>", "evaluator: <provider>"):
from .core.evaluator.<provider> import <Provider>AKEvaluator
return <Provider>AKEvaluator
if "." not in configured:
raise AKConfigError(
f"unknown evaluator '{configured}'; expected one of {_BUILTIN_EVALUATORS} or a dotted path to an AKEvaluator subclass"
)
return resolve_dotted(configured, base=AKEvaluator)A dotted evaluator: value (e.g. myorg.evaluators.CustomEvaluator) resolves via resolve_dotted
without any factory edit at all — only add an if branch here for a first-party, in-repo provider
you want addressable by a short name.
Add a new extras group to ak-py/pyproject.toml for the provider's SDK — don't fold it into the
existing test extra (that one stays DeepEval's, since every test user already needs it for the
framework itself). Follow the pattern of the opik extra, the first provider added on top of the
original DeepEval-only test extra:
[project.optional-dependencies]
<provider> = [
"provider-sdk>=x.y.z",
]evaluator: in test-config.yaml is already a free-form string on AKTestConfig (built-in short
name or dotted path) — no config schema change is needed for a new built-in, since it's just a new
value the same field accepts:
mode: fallback
evaluator: <provider>If your provider needs extra config fields (e.g. an API key env var name, a judge model override),
read them from AKTestConfig the same way DeepevalAKEvaluator reads self._config.llm — don't
invent a parallel config path.
Add ak-py/tests/test_evaluator_<provider>.py, following the shape of
ak-py/tests/test_evaluator_deepeval.py: exercise evaluate_by_score for real (offline, no
network) where possible, and mock the judge call in evaluate_by_llm so the suite stays
network-free. At minimum cover:
evaluate_by_score: exact/mismatch cases, threshold boundary, AKMissingInput when expected
is absentevaluate_by_llm: success, failure wrapped as AKEvaluationError, AKMissingInput when
expected is absentTest._resolve_evaluator_class("<provider>") resolves to your class, and
(if the SDK is optional) the require_extra ImportError path when it's missing — see
test_resolve_evaluator_class_deepeval_missing_extra_raises_import_error in
ak-py/tests/test_cli_tester.py for the pattern (patching builtins.__import__, since a
cached submodule import can otherwise mask the missing dependency).Add examples/cli/<provider>-evaluator/, following the shape of examples/cli/opik-evaluator/
(a minimal agent, a demo_test.py exercising the new evaluator, and a test-config.yaml pointing
evaluator: at the new short name). Register it in .github/test-config.yaml's e2e matrix so it
runs in CI, the way every other examples/cli/* entry does.
Neither doc page carries a literal "evaluator backend table" — both describe the built-ins in
prose next to the score/llm/fallback mode explanations. Update every prose mention that
enumerates the built-ins by name, not just one page:
docs/docs/core-concepts/configuration.md
and docs/docs/testing/cli-testing.md — the
evaluator: field description and the score/llm mode explanations.docs/docs/testing/automated-testing.md and
docs/docs/testing/overview.md — same prose pattern,
duplicated across these pages.docs/docs/agent-skills.md — the skill directory rows for
this skill and for ak-dev-testing-conventions..agents/skills/ak-dev-testing-conventions/SKILL.md — the evaluator config/mode section.ak-py/README.md — the Test Configuration reference (evaluator field) and the test-config
walkthrough section.ak-test skill (ak-py/src/agentkernel/skills/ak-test/SKILL.md) and its
evals/evals.json.docs/src/components/*/data.tsx): a tile in the Observability,
safety & testing row of IntegrationsMarquee/data.tsx (role Evaluator, href to the
automated testing page, logo or react-icons/si glyph), and the provider in the Pluggable
Evaluators card's tags and description under the Observe tab in
FeatureExplorer/data.tsx. Logo sourcing and the build check are in
ak-dev-sync-docs-from-branch, Docs-Site Landing and Features Pages.docs/src/pages/features.tsx): the approaches entry for Pluggable
Evaluators under Testing & Evaluation names every built-in.ak-py/src/agentkernel/test/core/evaluator/<provider>.py implementing AKEvaluatorTest._resolve_evaluator_class (ak-py/src/agentkernel/test/test.py)
and _BUILTIN_EVALUATORSak-py/pyproject.tomlak-py/tests/test_evaluator_<provider>.pyexamples/cli/<provider>-evaluator/, registered in .github/test-config.yamldocs/docs/core-concepts/configuration.md,
docs/docs/testing/cli-testing.md, docs/docs/testing/automated-testing.md,
docs/docs/testing/overview.md, docs/docs/agent-skills.md,
.agents/skills/ak-dev-testing-conventions/SKILL.md, ak-py/README.md, the ak-test skill
and its evals/evals.jsonIntegrationsMarquee/data.tsx), Pluggable Evaluators card
tags (FeatureExplorer/data.tsx); the Pluggable Evaluators approaches entry in features.tsx© yaalalabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/ak-dev-new-evaluator-provider of yaalalabs/agent-kernel.
Open the folder on GitHubat commit 97fa8d9
Ak Dev New Evaluator Provider next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ak Dev New Evaluator Provider this skillyaalalabs/agent-kernel | 192 | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Looperksimback/looper | 710 | — | ~2.7k | Automated safety check: Notes | MIT | |
| Agent Eval Engineeringlangchain-ai/langchain-skills | 1.3k | — | ~4k | Automated safety check: Pass | MIT | |
| Quality FlywheelGoogleCloudPlatform/vertex-ai-samples | 792 | — | ~2k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
ksimback/looper
Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
GoogleCloudPlatform/vertex-ai-samples
Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.
cloudnative-co/claude-code-starter-kit
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.
yaalalabs/agent-kernel
Code quality standards, formatting, Python style rules (classes over script-style functions, configuration-field rules), commit conventions, and PR workflow for Agent Kernel development.
yaalalabs/agent-kernel
Step-by-step guide for adding a new guardrail provider to Agent Kernel.
yaalalabs/agent-kernel
Step-by-step guide for adding a new knowledge base backend to Agent Kernel.
yaalalabs/agent-kernel
Step-by-step guide for adding a new messaging platform integration to Agent Kernel.
yaalalabs/agent-kernel
Step-by-step guide for adding a new multimodal attachment storage backend to Agent Kernel.
yaalalabs/agent-kernel
Step-by-step guide for adding a new queue transport to Agent Kernel's execution pipeline.
Categories
Step-by-step guide for adding a new built-in test evaluator provider to Agent Kernel (beyond DeepEval, Opik and JEV). Ak Dev New Evaluator Provider is an agent skill from yaalalabs/agent-kernel. Step-by-step guide for adding a new built-in test evaluator provider to Agent Kernel (beyond DeepEval, Opik and JEV).
Ak Dev New Evaluator Provider fits situations like: tasks that involve LLM evaluation.
Run `npx skills add yaalalabs/agent-kernel --skill ak-dev-new-evaluator-provider -a claude-code`. Or copy the skill folder (.agents/skills/ak-dev-new-evaluator-provider in yaalalabs/agent-kernel) into .claude/skills/ak-dev-new-evaluator-provider in your project. Claude Code loads it when a task matches its description.
Run `npx skills add yaalalabs/agent-kernel --skill ak-dev-new-evaluator-provider -a codex`. Or copy the skill folder (.agents/skills/ak-dev-new-evaluator-provider in yaalalabs/agent-kernel) into .agents/skills/ak-dev-new-evaluator-provider in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yaalalabs/agent-kernel --skill ak-dev-new-evaluator-provider -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ak-dev-new-evaluator-provider, .gemini/skills/ak-dev-new-evaluator-provider, .github/skills/ak-dev-new-evaluator-provider and .opencode/skills/ak-dev-new-evaluator-provider in your project.
SKILL.md names no scripts, command-line tools or credentials: Ak Dev New Evaluator Provider is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Ak Dev New Evaluator Provider is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Ak Dev New Evaluator Provider: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Looper (ksimback/looper, 710 stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
yaalalabs (a GitHub organization) maintains it in yaalalabs/agent-kernel, which has 192 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 9, 2026.
Source: yaalalabs/agent-kernel on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.