Testing framework for evaluating Databricks skills. An agent skill from databricks-solutions/ai-dev-kit.

Custom licenceAuto-check passedTesting & QA

Install Skill Test

skills CLI
$ npx skills add databricks-solutions/ai-dev-kit --skill skill-test -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install databricks-solutions/ai-dev-kit skill-test --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/databricks-solutions/ai-dev-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.test .claude/skills/skill-test && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skill-test
GitHub stars
1.9k
Token cost
~1.9k tokens
SKILL.md length
487 words
Files
201 (incl. scripts, references)
Skills in repo
5
Repo updated
First seen
Licence
Custom licence

At a glance

Testing framework for evaluating Databricks skills. An agent skill from databricks-solutions/ai-dev-kit.

  • Works in 4 steps: Read the skill's SKILL.md to understand… → Create manifest.yaml with appropriate… → Create empty ground_truth.yaml and… → …
  • Building test cases for skills
  • SKILL.md covers Quick References, /skill-test Command, Execution Instructions and Command Handler, plus 3 more sections
  • Calls uv

What it does

Skill Test is an agent skill from databricks-solutions/ai-dev-kit. Testing framework for evaluating Databricks skills. Use when building test cases for skills, running skill evaluations, comparing skill versions, or creating ground truth datasets with the Generate-Review-Promote (GRP) pipeline. Triggers include "test skill", "evaluate skill", "skill regression", "ground truth", "GRP pipeline", "skill quality", and "skill metrics".

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 209 other files, including scripts and reference files (for example `CUSTOMIZATION_GUIDE.md`, `README.md` and `TECHNICAL.md`).

It sits in Testing & QA, covering Test generation and Agent evaluation and testing. It works with Databricks and MLflow. The repository describes itself as: Databricks Toolkit for Coding Agents provided by Field Engineering.

When your agent uses it

  • Building test cases for skills
  • Running skill evaluations
  • Comparing skill versions
  • Creating ground truth datasets with the Generate-Review-Promote (GRP) pipeline

Example prompts

  • “test skill”
  • “evaluate skill”
  • “skill regression”
  • “/skill-test”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Read the skill's SKILL.md to understand its purpose
  2. Create manifest.yaml with appropriate scorers and trace_expectations
  3. Create empty ground_truth.yaml and candidates.yaml templates
  4. Recommend test prompts based on documentation examples

What it can do on your machine

Read from SKILL.md and the folder at commit b059fd0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skill Test loads about 1.9k tokens when it runs, and up to ~8.3k if it reads all its reference files. Until then it costs about 95 tokens; SKILL.md has 487 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~95
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 487 words (~1,938 tokens).

“Offline YAML-first evaluation with human-in-the-loop review and interactive skill improvement.”

— opening of SKILL.md by databricks-solutions, Custom licence
name
skill-test
command
skill-test
arguments
[skill-name] [subcommand]

Read the full SKILL.md on GitHub

Files

SKILL.md and 200 other files (scripts, references) in .test of databricks-solutions/ai-dev-kit.

  • SKILL.md
  • CUSTOMIZATION_GUIDE.md
  • README.md
  • TECHNICAL.md
  • baselines/databricks-agent-bricks/baseline.yaml
  • baselines/databricks-aibi-dashboards/baseline.yaml
  • baselines/databricks-apps-python/baseline.yaml
  • baselines/databricks-bundles/baseline.yaml
  • baselines/databricks-genie/baseline.yaml
  • baselines/databricks-model-serving/baseline.yaml
  • baselines/databricks-python-sdk/baseline.yaml
  • baselines/databricks-spark-declarative-pipelines/baseline.yaml
  • … and 189 more

Open the folder on GitHubat commit b059fd0

Compare with similar skills

Skill Test next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skill Test compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skill Test this skilldatabricks-solutions/ai-dev-kit1.9k—~1.9kAutomated safety check: PassCustom licence
Databricks Mlflow Evaluationdatabricks/databricks-agent-skills345—~2.7kAutomated safety check: PassCustom licence
Eval Triage And Improvementmicrosoft/eval-guide138—~5.9kAutomated safety check: PassMIT
Eval Guidemicrosoft/eval-guide138—~22kAutomated safety check: WarnMIT
Testing Livekit Agentslivekit-examples/agent-starter-python2641 repos~1.9kAutomated safety check: PassMIT
Evalmikeyobrien/rho372—~9.7kAutomated safety check: PassMIT

Similar skills

  • Databricks Mlflow Evaluation

    databricks/databricks-agent-skills

    Official

    MLflow 3 GenAI agent evaluation. An agent skill from databricks/databricks-agent-skills.

    345 GitHub stars~2.7k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Eval Triage And Improvement

    microsoft/eval-guide

    Official

    A skill your agent uses when the user's Copilot Studio agent evaluations have come back and they need to interpret scores, diagnose root causes of underperforming test cases, find remediation steps…

    138 GitHub stars~5.9k tokensUpdated 3 mo ago
    Testing & QAAuto-check passed
  • Eval Guide

    microsoft/eval-guide

    Official

    Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately.

    138 GitHub stars~22k tokensUpdated 3 mo ago
    Testing & QAAuto-check: warnings
  • Testing Livekit Agents

    livekit-examples/agent-starter-python

    Writes turn-level tests for a LiveKit agent in the user's normal test suite: pytest (Python) or Vitest (Node.js).

    264 GitHub starsUsed in 1 repo~1.9k tokens
    Testing & QAAuto-check passed
  • Eval

    mikeyobrien/rho

    Plan and run conversational AI agent evaluations with test generation and analysis.

    372 GitHub stars~9.7k tokensUpdated 7 days ago
    Testing & QAAuto-check passed
  • Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.

    16k GitHub starsUsed in 1 repo~2.1k tokens
    Testing & QAAuto-check: notes

More from databricks-solutions/ai-dev-kit

  • Python Dev

    databricks-solutions/ai-dev-kit

    Python development guidance with code quality standards, error handling, testing practices, and environment management.

    1.9k GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Tool Selection

    databricks-solutions/ai-dev-kit

    Evaluates whether the agent selected appropriate MCP tools instead of shell workarounds.

    1.9k GitHub stars~519 tokensUpdated 1 mo ago
    Auto-check passed
  • SQL Correctness

    databricks-solutions/ai-dev-kit

    SQL evaluation criteria for Databricks. An agent skill from databricks-solutions/ai-dev-kit.

    1.9k GitHub stars~438 tokensUpdated 1 mo ago
    Auto-check passed
  • General Quality

    databricks-solutions/ai-dev-kit

    General response quality evaluation. An agent skill from databricks-solutions/ai-dev-kit.

    1.9k GitHub stars~358 tokensUpdated 1 mo ago
    Auto-check passed

Questions about Skill Test

What does Skill Test do?

Testing framework for evaluating Databricks skills. An agent skill from databricks-solutions/ai-dev-kit. Skill Test is an agent skill from databricks-solutions/ai-dev-kit. Testing framework for evaluating Databricks skills.

When should I use Skill Test?

Skill Test fits situations like: building test cases for skills; running skill evaluations; comparing skill versions; creating ground truth datasets with the Generate-Review-Promote (GRP) pipeline.

How do I install Skill Test in Claude Code?

Run `npx skills add databricks-solutions/ai-dev-kit --skill skill-test -a claude-code`. Or copy the skill folder (.test in databricks-solutions/ai-dev-kit) into .claude/skills/skill-test in your project. Claude Code loads it when a task matches its description.

How do I install Skill Test in Codex?

Run `npx skills add databricks-solutions/ai-dev-kit --skill skill-test -a codex`. Or copy the skill folder (.test in databricks-solutions/ai-dev-kit) into .agents/skills/skill-test in your project. Codex loads it when a task matches its description.

Can I use Skill Test in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add databricks-solutions/ai-dev-kit --skill skill-test -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-test, .gemini/skills/skill-test, .github/skills/skill-test and .opencode/skills/skill-test in your project.

What does Skill Test need to run?

Going by SKILL.md and its folder, Skill Test needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Skill Test access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Skill Test safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Skill Test use?

Skill Test has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Skill Test use?

About 1.9k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.3k tokens, read only when the agent opens those files.

What are the alternatives to Skill Test?

Skills that share tags, products or a category with Skill Test: Databricks Mlflow Evaluation (databricks/databricks-agent-skills, 345 stars), Eval Triage And Improvement (microsoft/eval-guide, 138 stars), Eval Guide (microsoft/eval-guide, 138 stars) and Testing Livekit Agents (livekit-examples/agent-starter-python, 264 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skill Test?

databricks-solutions (a GitHub organization) maintains it in databricks-solutions/ai-dev-kit, which has 1,939 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on August 13, 2026.

Source: databricks-solutions/ai-dev-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.