Official agent skill

Databricks Mlflow Evaluation

by databricks in databricks/databricks-agent-skills

MLflow 3 GenAI agent evaluation. An agent skill from databricks/databricks-agent-skills.

OfficialCustom licenceAuto-check passedAgent Workflows

Install Databricks Mlflow Evaluation

skills CLI
$ npx skills add databricks/databricks-agent-skills --skill databricks-mlflow-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install databricks/databricks-agent-skills databricks-mlflow-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/databricks/databricks-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/databricks-mlflow-evaluation .claude/skills/databricks-mlflow-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
databricks-mlflow-evaluation
GitHub stars
345
Token cost
~2.7k tokens
SKILL.md length
1,046 words
Files
15 (incl. references, assets)
Skills in repo
32
Repo updated
First seen
Licence
Custom licence

At a glance

MLflow 3 GenAI agent evaluation. An agent skill from databricks/databricks-agent-skills.

  • Works in 2 steps: Read GOTCHAS.md - 15+ common mistakes… → Read CRITICAL-interfaces.md - Exact API…
  • Writing mlflow.genai.evaluate() code
  • SKILL.md covers Scope vs upstream mlflow/skills, Before Writing Any Code, End-to-End Workflows and Reference Files Quick Lookup, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Databricks Mlflow Evaluation is an agent skill from databricks/databricks-agent-skills, published by the product's own GitHub organization. MLflow 3 GenAI agent evaluation. Use when writing mlflow.genai.evaluate() code, creating @scorer functions, using built-in scorers (Guidelines, Correctness, Safety, RetrievalGroundedness), building eval datasets from traces, setting up trace ingestion and production monitoring, aligning judges with MemAlign from domain expert feedback, or running optimizeprompts() with GEPA for automated prompt improvement.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 17 other files, including reference files and assets (for example `agents/openai.yaml`, `references/CRITICAL-interfaces.md` and `references/GOTCHAS.md`). Compatibility notes: Requires databricks CLI (= v1.0.0)

It sits in Agent Workflows, covering Agent evaluation and testing. It works with MLflow and Databricks. The repository describes itself as: Databricks AI Tools: skills and plugins for building on Databricks with Claude Code, Cursor, Codex, GitHub Copilot, and other AI coding agents.

When your agent uses it

  • Writing mlflow.genai.evaluate() code
  • Creating @scorer functions
  • Using built-in scorers (Guidelines
  • RetrievalGroundedness)

Example prompts

  • “/databricks-mlflow-evaluation”

Requirements

  • Compatibility (from SKILL.md): Requires databricks CLI (>= v1.0.0)

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Read GOTCHAS.md - 15+ common mistakes that cause failures
  2. Read CRITICAL-interfaces.md - Exact API signatures and data schemas

What it can do on your machine

Read from SKILL.md and the folder at commit f4fcec5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires databricks CLI (>= v1.0.0)

    From compatibility in the SKILL.md frontmatter.

Context cost

Databricks Mlflow Evaluation loads about 2.7k tokens when it runs, and up to ~54k if it reads all its reference files. Until then it costs about 110 tokens; SKILL.md has 1,046 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~110
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~54k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 1,046 words (~2,685 tokens).

“The OSS mlflow/skills repo ships agent-evaluation and related skills (instrumenting-with-mlflow-tracing, analyze-mlflow-trace, retrieving-mlflow-traces, querying-mlflow-metrics) that cover the generic MLflow GenAI evaluation workflow — mlflow.genai.evaluate(), scorers/judges, datasets, tracing setup, and the 5-step evaluation loop.”

— opening of SKILL.md by databricks, Custom licence
name
databricks-mlflow-evaluation
compatibility
Requires databricks CLI (>= v1.0.0)
metadata.version
0.1.0
parent
databricks-core

Read the full SKILL.md on GitHub

Files

SKILL.md and 14 other files (references, assets) in skills/databricks-mlflow-evaluation of databricks/databricks-agent-skills.

  • SKILL.md
  • agents/openai.yaml
  • assets/databricks.png
  • assets/databricks.svg
  • references/CRITICAL-interfaces.md
  • references/GOTCHAS.md
  • references/patterns-context-optimization.md
  • references/patterns-datasets.md
  • references/patterns-evaluation.md
  • references/patterns-judge-alignment.md
  • references/patterns-prompt-optimization.md
  • references/patterns-scorers.md
  • references/patterns-trace-analysis.md
  • references/patterns-trace-ingestion.md
  • references/user-journeys.md

Open the folder on GitHubat commit f4fcec5

Compare with similar skills

Databricks Mlflow Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Databricks Mlflow Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Databricks Mlflow Evaluation this skilldatabricks/databricks-agent-skills345—~2.7kAutomated safety check: PassCustom licence
Skill Testdatabricks-solutions/ai-dev-kit1.9k—~1.9kAutomated safety check: PassCustom licence
Azure Machine LearningMicrosoftDocs/Agent-Skills7751 repos~19kAutomated safety check: PassCC-BY-4.0
Azure Data Science VmMicrosoftDocs/Agent-Skills775—~1.8kAutomated safety check: PassCC-BY-4.0
MCP Server Builderanthropics/skills180k62 repos~2.3kAutomated safety check: PassApache-2.0
Diagnosing Superpowers Sessionsobra/superpowers296k3 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Skill Test

    databricks-solutions/ai-dev-kit

    Testing framework for evaluating Databricks skills. An agent skill from databricks-solutions/ai-dev-kit.

    1.9k GitHub stars~1.9k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Azure Machine Learning

    MicrosoftDocs/Agent-Skills

    Official

    Expert knowledge for Azure Machine Learning development including troubleshooting, best practices, decision making, architecture & design patterns, limits & quotas, security, configuration…

    775 GitHub starsUsed in 1 repo~19k tokens
    DevelopmentAuto-check passed
  • Azure Data Science Vm

    MicrosoftDocs/Agent-Skills

    Official

    Expert knowledge for Azure Data Science Virtual Machines development including troubleshooting, decision making, architecture & design patterns, security, configuration, integrations & coding…

    775 GitHub stars~1.8k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 62 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers.

    296k GitHub starsUsed in 3 repos~1.7k tokens
    Agent WorkflowsAuto-check passed
  • Darwin Skill Optimizer

    alchaincyf/darwin-skill

    Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.

    6.2k GitHub starsUsed in 1 repo~4.7k tokens
    Agent WorkflowsAuto-check passed

More from databricks/databricks-agent-skills

All 32 skills in this repo
  • Databricks Dbsql

    databricks/databricks-agent-skills

    Official

    Databricks SQL (DBSQL) advanced features and SQL warehouse capabilities.

    345 GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed
  • Databricks Synthetic Data Gen

    databricks/databricks-agent-skills

    Official

    Generate realistic synthetic data using Spark + Faker (strongly recommended).

    345 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Databricks Model Serving

    databricks/databricks-agent-skills

    Official

    Databricks Model Serving endpoint lifecycle and ops. An agent skill from databricks/databricks-agent-skills.

    345 GitHub starsUsed in 1 repo~3.3k tokens
    Auto-check passed
  • Databricks Python SDK

    databricks/databricks-agent-skills

    Official

    Databricks development guidance including Python SDK, Databricks Connect, CLI, and REST API.

    345 GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed
  • Databricks Spark Structured Streaming

    databricks/databricks-agent-skills

    Official

    Comprehensive guide to Spark Structured Streaming for production workloads.

    345 GitHub starsUsed in 1 repo~956 tokens
    Auto-check passed
  • Databricks App Design

    databricks/databricks-agent-skills

    Official

    Design the UX of custom-code Databricks Apps (AppKit/React) data screens — KPI/overview pages, reports, charts, tables, and Genie/chat data assistants — mapped to concrete AppKit components.

    345 GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Databricks Mlflow Evaluation

What does Databricks Mlflow Evaluation do?

MLflow 3 GenAI agent evaluation. An agent skill from databricks/databricks-agent-skills. Databricks Mlflow Evaluation is an agent skill from databricks/databricks-agent-skills, published by the product's own GitHub organization. MLflow 3 GenAI agent evaluation.

When should I use Databricks Mlflow Evaluation?

Databricks Mlflow Evaluation fits situations like: writing mlflow.genai.evaluate() code; creating @scorer functions; using built-in scorers (Guidelines; retrievalGroundedness).

How do I install Databricks Mlflow Evaluation in Claude Code?

Run `npx skills add databricks/databricks-agent-skills --skill databricks-mlflow-evaluation -a claude-code`. Or copy the skill folder (skills/databricks-mlflow-evaluation in databricks/databricks-agent-skills) into .claude/skills/databricks-mlflow-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Databricks Mlflow Evaluation in Codex?

Run `npx skills add databricks/databricks-agent-skills --skill databricks-mlflow-evaluation -a codex`. Or copy the skill folder (skills/databricks-mlflow-evaluation in databricks/databricks-agent-skills) into .agents/skills/databricks-mlflow-evaluation in your project. Codex loads it when a task matches its description.

Can I use Databricks Mlflow Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add databricks/databricks-agent-skills --skill databricks-mlflow-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/databricks-mlflow-evaluation, .gemini/skills/databricks-mlflow-evaluation, .github/skills/databricks-mlflow-evaluation and .opencode/skills/databricks-mlflow-evaluation in your project.

What does Databricks Mlflow Evaluation need to run?

SKILL.md names no scripts, command-line tools or credentials: Databricks Mlflow Evaluation is instructions for the agent only. Compatibility (from SKILL.md): Requires databricks CLI (>= v1.0.0).

Does Databricks Mlflow Evaluation access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Databricks Mlflow Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Databricks Mlflow Evaluation use?

Databricks Mlflow Evaluation has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Databricks Mlflow Evaluation use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 52k tokens, read only when the agent opens those files.

What are the alternatives to Databricks Mlflow Evaluation?

Skills that share tags, products or a category with Databricks Mlflow Evaluation: Skill Test (databricks-solutions/ai-dev-kit, 1.9k stars), Azure Machine Learning (MicrosoftDocs/Agent-Skills, 775 stars), Azure Data Science Vm (MicrosoftDocs/Agent-Skills, 775 stars) and MCP Server Builder (anthropics/skills, 180k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Databricks Mlflow Evaluation?

databricks (a GitHub organization, an official publisher) maintains it in databricks/databricks-agent-skills, which has 345 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 6, 2026.

Source: databricks/databricks-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.