Official agent skill

Logfire Evals

by pydantic in pydantic/skills

Run offline Python (pydanticevals) or Node.js (logfire/evals) evaluations and review them in Logfire.

OfficialMITAuto-check passedAI & LLM Engineering

Install Logfire Evals

skills CLI
$ npx skills add pydantic/skills --skill logfire-evals -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pydantic/skills logfire-evals --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pydantic/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/logfire-evals .claude/skills/logfire-evals && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
logfire-evals
GitHub stars
140
Token cost
~3.6k tokens
SKILL.md length
1,592 words
Files
1
Skills in repo
9
Repo updated
First seen
Licence
MIT

At a glance

Run offline Python (pydanticevals) or Node.js (logfire/evals) evaluations and review them in Logfire.

  • Works in 5 steps: Check for an Existing Braintrust Suite… → Authenticate When the Run Needs Logfire → Detect What to Evaluate → …
  • Braintrust migration
  • SKILL.md covers How This Works, Step 1: Check for an Existing…, Step 2: Authenticate When the… and Step 3: Detect What to Evaluate, plus 2 more sections
  • Calls uv, poetry and pnpm; reaches logfire-us.pydantic.dev; needs BRAINTRUST_API_KEY

What it does

Logfire Evals is an agent skill from pydantic/skills, published by the product's own GitHub organization. Run offline Python (pydanticevals) or Node.js (logfire/evals) evaluations and review them in Logfire. Also redirect existing Braintrust Eval() suites. Use for eval setup, test datasets, AI scoring, agent checks, LLM judges, Braintrust migration, or Logfire Datasets & Experiments. Not for live traffic or infrastructure monitoring.

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation. It works with Pydantic, Python, Node.js and OpenTelemetry. The licence is MIT.

When your agent uses it

  • Braintrust migration
  • Logfire Datasets & Experiments

Example prompts

  • “/logfire-evals”

Requirements

  • Python 3
  • Node.js
  • A credential in BRAINTRUST_API_KEY

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Check for an Existing Braintrust Suite First
  2. Authenticate When the Run Needs Logfire
  3. Detect What to Evaluate
  4. Define the Dataset and Run It
  5. Verify

What it can do on your machine

Read from SKILL.md and the folder at commit 238d971. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • poetry
    • pnpm
    • yarn
    • bun
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • logfire-us.pydantic.dev

    Also links to:

    • pydantic.dev

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • BRAINTRUST_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Logfire Evals loads about 3.6k tokens when it runs. Until then it costs about 88 tokens; SKILL.md has 1,592 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~88
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from pydantic/skills at commit 238d971, republished under its MIT licence (© pydantic). 1,592 words, ~3,551 tokens.

Download SKILL.mdSave it as .claude/skills/logfire-evals/SKILL.md (or your agent's skills folder).
name
logfire-evals
description
Run offline Python (`pydantic_evals`) or Node.js (`logfire/evals`) evaluations and review them in Logfire. Also redirect existing Braintrust `Eval()` suites. Use for eval setup, test datasets, AI scoring, agent checks, LLM judges, Braintrust migration, or Logfire Datasets & Experiments. Not for live traffic or infrastructure monitoring.

Evaluate AI code with Logfire

How This Works

Python's pydantic_evals and Node.js's logfire/evals run a task against cases, apply evaluators, and return a report. An active Logfire or OpenTelemetry provider may export inputs and outputs even if this skill did not configure it. Intentional upload needs logfire.configure() in Python or a configured Node.js exporter.

Span-based evaluators inspect the task's OpenTelemetry span tree. Without working Logfire instrumentation, Python reports "No span tree available"; Node.js HasMatchingSpan can produce no evaluator result at all. Treat either signal as a setup failure, not evidence about the agent.

Step 1: Check for an Existing Braintrust Suite First

Cheap check, before anything else: does this repo already have an existing Braintrust suite — actual Eval(...) calls or from braintrust import Eval in source, not just a braintrust dependency listed without any real usage? This path needs no CLI auth at all — don't run Step 2 for it. The compatibility endpoint is documented only for Logfire Cloud US and EU; for any other supplied origin, use the native evaluation path instead.

Keep the existing Eval() code (Python braintrust>=0.30.1 / TypeScript braintrust>=3.24.0 — verified versions) and redirect its next run to Logfire by changing environment variables only, no pydantic_evals involved:

bash
export BRAINTRUST_APP_URL="https://logfire-us.pydantic.dev/v1/braintrust"  # EU: logfire-eu.pydantic.dev
export BRAINTRUST_API_KEY="<logfire-project-api-key>"                     # Settings -> API Keys
unset BRAINTRUST_API_URL BRAINTRUST_PROXY_URL  # these override the endpoint above if set — the #1 "it still hit Braintrust" cause

The destination project's API key needs project:write_otlp and project:read_datasets: the SDK writes the run, then reads experiment metadata for its summary. An ingest-only write token fails with 403.

This compatibility preview covers inline/callable data, local tasks and scorers, multiple scores, one label per name, and normal summary finalization. It excludes Braintrust-hosted resources, BTQL, the model proxy, server-side scoring, and post-finalization feedback. Rust, summarize_scores=False, and manual flush() without comparison do not request the summary needed for projection. See the full coverage and concept mapping.

Skip straight to Step 5 (Verify) — the SDK's own printed result URL also opens directly in Logfire, and nothing else here (auth, dataset definition) applies to this path.

No existing Braintrust suite? Continue to Step 2 now, before the more detailed identification in Step 3 — nothing past this point requires knowing the function/agent or dataset shape yet.

Step 2: Authenticate When the Run Needs Logfire

For an explicitly local-only run without span evaluators, use a fresh process that neither preloads nor imports application telemetry; omit Python's logfire.configure() and any Node.js exporter bootstrap. If the task configures an exporter and has no documented disable switch, stop rather than claiming local-only. Uploading, hosted datasets, and span evaluators require Logfire authentication to the exact project.

For a Logfire-backed run, use Authenticate and Select the Exact Project to derive the CLI target from the supplied Logfire URL and run its target-aware whoami check with a verified CLI path — for JS/TS projects without uv, use the external-prefix npm fallback instead of plain npx, which can execute a repository-local binary. Skip to Step 3 if that already reports the right project and resolved --region or --base-url target; otherwise, continue through the full authentication and project-selection sequence there. This CLI flow is for logfire.configure(); Step 3's hosted-dataset operations use a separate API key with different scopes.

Step 3: Detect What to Evaluate

Identify the real task and any dataset. Follow repository package, test, and dependency conventions. For a first eval, prefer 3-5 cases from tests, schemas, examples, or synthetic fixtures, with deterministic checks for defined behavior. Do not copy the example unless it fits, replace an eval framework, or refactor unrelated code. Without a runnable task or safe expected behavior, ask one focused question instead of inventing either.

  • In-code dataset: a Python module using pydantic_evals, or a Node.js module using logfire/evals. This is the default for an agent-driven workflow.
  • Hosted/managed dataset: cases live in the Logfire UI, edited by non-engineers, pulled/pushed via a separate LogfireAPIClient (from logfire.experimental.api_client import LogfireAPIClient). Hosted inputs and outputs must be JSON objects; scalar roots accepted by code-defined datasets are rejected. client.get_dataset(name) with no type arguments returns a raw dict, not something push_dataset or .evaluate_sync() can take — pass the input/output (and metadata, if used) types to get back a real pydantic_evals.Dataset: client.get_dataset(name, MyInputType, MyOutputType). If the stored dataset contains custom evaluators, also pass their classes with custom_evaluator_types=[MyEvaluator] (and custom report evaluators with custom_report_evaluator_types=[...]) so they can be deserialized. Push with client.push_dataset(dataset). This needs its own API key from Settings → API Keys (scoped project:read_datasets/project:write_datasets), not Step 2's CLI auth flow. Only relevant if the user specifically wants case editing outside code.

Step 4: Define the Dataset and Run It

Use the repository's existing package manager and lockfile. Install only the missing integration for its language.

Python

Add pydantic-evals[logfire] with the detected Python manager: uv add, poetry add, or pdm add. For a pip/requirements project, update its declared requirements and install from that file; do not introduce a second manager or lockfile.

python
import logfire
from pydantic_evals import Case, Dataset
from pydantic_evals.evaluators import EqualsExpected, IsInstance

logfire.configure()  # omit only in the isolated local-only process described above

def classify_sentiment(text: str) -> str:
    return 'positive' if 'love' in text else 'negative'


dataset = Dataset[str, str, None](
    name='sentiment-eval',
    cases=[
        Case(name='positive', inputs='I love this', expected_output='positive'),
        Case(name='negative', inputs='This is terrible', expected_output='negative'),
    ],
    evaluators=[EqualsExpected(), IsInstance(type_name='str')],
)

report = dataset.evaluate_sync(classify_sentiment)  # or `await dataset.evaluate(...)`
report.print(include_input=True, include_output=True)

Use logfire[datasets] instead only when the task specifically needs the hosted-dataset API from Step 3.

JavaScript or TypeScript on Node.js

Do not apply this section to Deno, Bun, browsers, or workers; their exporter setup is not validated by this skill.

Add logfire and @pydantic/logfire-node with the manager selected by the existing lockfile: pnpm add, yarn add, bun add, or npm install. Do not introduce a second lockfile.

Configure Logfire before loading the task. Reuse an existing instrumentation entry point rather than configuring it twice.

ts
import * as logfire from '@pydantic/logfire-node'
import { Case, Dataset, EqualsExpected, renderReport } from 'logfire/evals'

logfire.configure()

const classifySentiment = (text: string) => (text.includes('love') ? 'positive' : 'negative')
const dataset = new Dataset<string, string>({
  name: 'sentiment-eval',
  cases: [new Case({ name: 'positive', inputs: 'I love this', expectedOutput: 'positive' })],
  evaluators: [new EqualsExpected()],
})

await dataset.evaluate(classifySentiment).then((report) => {
  console.log(renderReport(report, { includeInput: true, includeOutput: true }))
}).finally(() => logfire.shutdown({ timeoutMillis: 5000 }))

Other built-ins include Equals, Contains, IsInstance, MaxDuration, HasMatchingSpan, and LLMJudge. Node.js custom evaluators extend Evaluator, and LLMJudge needs a judge callback. Use @pydantic/logfire-node/datasets only for hosted datasets.

Show full SKILL.md (702 more words)Show less
Smoke test before a paid or full run

Before running the full dataset, run a smoke test on 2-3 cases if the dataset is large or uses LLMJudge or any evaluator that makes billed model calls. This catches setup errors before they multiply cost across the dataset.

Python:

python
smoke = Dataset(
    name=dataset.name,
    cases=dataset.cases[:3],
    evaluators=dataset.evaluators,
    report_evaluators=dataset.report_evaluators,
)
smoke_report = smoke.evaluate_sync(classify_sentiment)
smoke_report.print(include_input=True, include_output=True)

Node.js:

ts
const smoke = new Dataset({
  name: dataset.name,
  cases: dataset.cases.slice(0, 3),
  evaluators: dataset.evaluators,
  reportEvaluators: dataset.reportEvaluators,
})
await smoke.evaluate(classifySentiment).then((report) => {
  console.log(renderReport(report, { includeInput: true, includeOutput: true }))
}).finally(() => logfire.shutdown({ timeoutMillis: 5000 }))

Confirm the smoke run has zero unexpected errors and the assertions that should pass do. Then, if the full dataset is large or uses paid model calls, tell the user the case count and which evaluators will make model calls, and get explicit confirmation before running the full dataset — don't run an expensive full pass on the strength of a clean smoke test alone without saying so.

The remaining details in this section are Python-specific. Custom evaluators inherit Evaluator and implement evaluate; use @dataclass for configurable fields and portable serialization. Case names must be unique within a dataset. The evaluators reached for most:

EvaluatorChecks
Equals(value) / EqualsExpected()Exact match against a literal / expected_output (no-op if expected_output is unset — don't rely on it silently catching that)
IsInstance(type_name)Output's type matches by name
LLMJudge(rubric, model=None, score=False)Subjective or rubric-based judgment; makes billed model requests, so validate the rubric against human-reviewed examples before treating it as a quality gate
ToolCorrectness(expected_tools, ...)Which tools an agent called — reads the span tree, so needs Step 2's logfire.configure() to work at all, not just to upload

Also available: Contains, MaxDuration, TrajectoryMatch, ArgumentCorrectness, MaxToolCalls, MaxModelRequests — same span-tree dependency as ToolCorrectness for the tool/trajectory ones; see pydantic_evals.evaluators for the full set. These five agentic (span-based) evaluators need pydantic-evals>=2.4.0 — on an older pin, check pyproject.toml/uv.lock and upgrade before reaching for them, since the import itself is what fails, not a silent no-op.

The Python evaluator (arbitrary code execution) was removed for security reasons — don't reach for it even if an older example references it.

If editing a hosted dataset: client.push_dataset(dataset) overwrites server-side evaluators on every push, including removing ones you deleted locally — don't push a stale local copy over a dataset others have edited in the UI.

Step 5: Verify

A report printing to the terminal isn't proof it reached Logfire — confirm the run actually landed. Never report a case as passed, a score, or a run as complete without having actually checked it in this session — if a run fails, cancels, or produces no scores, report that failure plainly; never substitute an invented score or a manual guess at what the result "should" be.

Came from the Step 1 Braintrust path (Step 2 skipped)? There's no whoami-resolved project to look up here — use the SDK's own printed result URL instead, which already opens directly in the right Logfire project. Confirm the same things below (completion, pass mix, case detail) from that page rather than searching by name.

  1. Query for the run directly, if a Logfire MCP server or API is connected — the root span for a run is named evaluate {name} and carries gen_ai.operation.name = 'experiment', dataset_name, and task_name attributes; find the most recent one matching your dataset's name and confirm logfire.experiment.metadata shows the case count and pass rate you expect. Otherwise, open AI Evaluations → Datasets & Experiments → Experiments in Logfire for the exact project from Step 2, and find the run by name/timestamp.
  2. Read the Overview tab (or the queried metadata) first: completion count, assertion pass mix, task errors, average duration. If completion says "Not reported," the run sent case data but never signaled it finished — treat that as a broken run, not a passing one.
  3. Open the Cases tab, starting from Needs Review / Failed / Errors, not the full list.
  4. Drill into a failing case's trace in Live view for the actual evidence, rather than trusting the summary score alone.
  5. Fix and re-run until the cases that should pass do, and any tool-call/trajectory checks show real span data, not "No span tree available."

Close with a final report built from what you just confirmed — the run name, exact case count and pass rate you queried, and which evaluators ran — not a template. Include the direct link to this experiment (the SDK's own printed result URL, or the Datasets & Experiments page you opened it from), so the user can see the run without having to ask where to look.

© pydantic, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/logfire-evals of pydantic/skills.

Open the folder on GitHubat commit 238d971

Compare with similar skills

Logfire Evals next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Logfire Evals compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Logfire Evals this skillpydantic/skills140—~3.6kAutomated safety check: PassMIT
Phoenix Release PleaseArize-ai/phoenix12k—~708Automated safety check: PassApache-2.0
Gorm ExpertLeoYeAI/openclaw-master-skills2.2k—~3.4kAutomated safety check: PassMIT
Building Pydantic AI Agentsdocling-project/docling69k—~2.8kAutomated safety check: PassMIT
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Failproof AI SDK IntegrationFailproofAI/failproofai5.3k—~6kAutomated safety check: PassCustom licence

Similar skills

  • Phoenix Release Please

    Arize-ai/phoenix

    Bump the next release-please version for a Phoenix Python package (arize-phoenix, arize-phoenix-client, arize-phoenix-evals, arize-phoenix-otel) by opening a PR with a Release-As commit footer.

    12k GitHub stars~708 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Gorm Expert

    LeoYeAI/openclaw-master-skills

    GORM v2 最佳实践与性能优化。适用于:代码审查、慢查询优化、N+1、连接池、 事务管理、分库分表、Prometheus/OTel监控、Session安全、Clause/Upsert、 缓存集成、BaseModel脚手架、SQL→struct生成、多租户隔离。

    2.2k GitHub stars~3.4k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • Building Pydantic AI Agents

    docling-project/docling

    Patterns and tested examples for building agents with Pydantic AI: tools, capabilities, structured output, dependency injection, hooks, YAML specs, streaming and testing.

    69k GitHub stars~2.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Failproof AI SDK Integration

    FailproofAI/failproofai

    Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.

    5.3k GitHub stars~6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.

    40k GitHub stars~1.3k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed

More from pydantic/skills

All 9 skills in this repo
  • Logfire Infrastructure

    pydantic/skills

    Official

    Monitor hosts, Docker containers, Kubernetes clusters, database/queue/cache servers, and cloud-provider metrics with Pydantic Logfire — no application code required.

    140 GitHub stars~1.8k tokensUpdated 7 days ago
    Auto-check passed
  • Logfire Query

    pydantic/skills

    Official

    Query and analyze Logfire telemetry data — traces, logs, spans, metrics, summaries, and SQL results.

    140 GitHub stars~2.2k tokensUpdated 7 days ago
    Auto-check passed
  • Pydantic AI Harness

    pydantic/skills

    Official

    Extend Pydantic AI agents with batteries-included capabilities from pydantic-ai-harness -- Code Mode (collapse many tool calls into one sandboxed Python execution), a filesystem and shell…

    140 GitHub stars~1.9k tokensUpdated 7 days ago
    Auto-check passed
  • Official

    Build AI agents with Pydantic AI — tools, capabilities (including on-demand loading), structured output, streaming, testing, and multi-agent patterns.

    140 GitHub stars~5.4k tokensUpdated 7 days ago
    Auto-check passed
  • Official

    Add Pydantic Logfire observability to application code — traces, logs, metrics, and AI/agent spans.

    140 GitHub stars~6.1k tokensUpdated 7 days ago
    Auto-check passed
  • Logfire UI

    pydantic/skills

    Official

    Open or return Logfire project pages, live views, trace links, and Explore pages in the Codex browser without querying telemetry first.

    140 GitHub stars~2.9k tokensUpdated 7 days ago
    Auto-check passed

Questions about Logfire Evals

What does Logfire Evals do?

Run offline Python (pydanticevals) or Node.js (logfire/evals) evaluations and review them in Logfire. Logfire Evals is an agent skill from pydantic/skills, published by the product's own GitHub organization.js (logfire/evals) evaluations and review them in Logfire.

When should I use Logfire Evals?

Logfire Evals fits situations like: braintrust migration; logfire Datasets & Experiments.

How do I install Logfire Evals in Claude Code?

Run `npx skills add pydantic/skills --skill logfire-evals -a claude-code`. Or copy the skill folder (skills/logfire-evals in pydantic/skills) into .claude/skills/logfire-evals in your project. Claude Code loads it when a task matches its description.

How do I install Logfire Evals in Codex?

Run `npx skills add pydantic/skills --skill logfire-evals -a codex`. Or copy the skill folder (skills/logfire-evals in pydantic/skills) into .agents/skills/logfire-evals in your project. Codex loads it when a task matches its description.

Can I use Logfire Evals in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pydantic/skills --skill logfire-evals -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/logfire-evals, .gemini/skills/logfire-evals, .github/skills/logfire-evals and .opencode/skills/logfire-evals in your project.

What does Logfire Evals need to run?

Going by SKILL.md and its folder, Logfire Evals needs the command-line tools its instructions call (uv, poetry, pnpm, yarn, bun and npm) and credentials named BRAINTRUST_API_KEY. Our summary lists: Python 3; Node.js; A credential in BRAINTRUST_API_KEY.

Does Logfire Evals access the network?

SKILL.md names 2 domains. In commands or code: logfire-us.pydantic.dev; the agent is likely to contact it when it follows the instructions. As links in the text: pydantic.dev. This is read from the text; nothing was executed.

Is Logfire Evals safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Logfire Evals use?

Logfire Evals is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Logfire Evals use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Logfire Evals?

Skills that share tags, products or a category with Logfire Evals: Phoenix Release Please (Arize-ai/phoenix, 12k stars), Gorm Expert (LeoYeAI/openclaw-master-skills, 2.2k stars), Building Pydantic AI Agents (docling-project/docling, 69k stars) and Azure AI Projects Python SDK (microsoft/skills, 3.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Logfire Evals?

pydantic (a GitHub organization, an official publisher) maintains it in pydantic/skills, which has 140 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 1, 2026.

Source: pydantic/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.