Agent skill

Braintrust Agent Evals

by githits-com in githits-com/githits-cli

Inspect, query, compare, or explicitly export GitHits agent-eval history in Braintrust using the repository's verified workflow.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Braintrust Agent Evals

skills CLI
$ npx skills add githits-com/githits-cli --skill braintrust-agent-evals -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install githits-com/githits-cli braintrust-agent-evals --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/githits-com/githits-cli.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/braintrust-agent-evals .claude/skills/braintrust-agent-evals && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
braintrust-agent-evals
GitHub stars
114
Token cost
~3.1k tokens
SKILL.md length
1,347 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Inspect, query, compare, or explicitly export GitHits agent-eval history in Braintrust using the repository's verified workflow.

  • Tasks that involve LLM evaluation
  • SKILL.md covers Safety and interpretation, Inspect experiments and Validate or explicitly export
  • Calls bun; needs BRAINTRUST_API_KEY and OPENROUTER_API_KEY

What it does

Braintrust Agent Evals is an agent skill from githits-com/githits-cli. Inspect, query, compare, or explicitly export GitHits agent-eval history in Braintrust using the repository's verified workflow.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation. It works with Model Context Protocol. The repository describes itself as: CLI & MCP for GitHits - The Code Context Layer for AI Coding Agents. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve LLM evaluation

Example prompts

  • “/braintrust-agent-evals”

Requirements

  • A credential in BRAINTRUST_API_KEY
  • A credential in OPENROUTER_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 7449018. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bun

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • braintrust.dev

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • BRAINTRUST_API_KEY
    • OPENROUTER_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Braintrust Agent Evals loads about 3.1k tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 1,347 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~38
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from githits-com/githits-cli at commit 7449018, republished under its Apache-2.0 licence (© githits-com). 1,347 words, ~3,102 tokens.

Download SKILL.mdSave it as .claude/skills/braintrust-agent-evals/SKILL.md (or your agent's skills folder).
name
braintrust-agent-evals
description
Inspect, query, compare, or explicitly export GitHits agent-eval history in Braintrust using the repository's verified workflow.
metadata.internal
true

Braintrust agent evals

Use this skill for read-only inspection of the GitHits agent-eval history in the Braintrust project githits-cli-agent-evals, or when the user explicitly asks to export a validated local suite. The normalized exporter stores one top-level eval span per scenario/workload cell plus structural tool children; see docs/implementation/agentic-eval-metrics.md for the field contract.

The persistence unit is one exporter invocation = one experiment, one scenario/workload cell = one eval row, and one normalized logical tool call = one structural tool child. Current exporter-owned experiment names are main-r<run-id>-a<attempt>, pr-<pr-number>-r<run-id>-a<attempt>, and local-<branch-slug>-<UTC-timestamp-with-milliseconds>-<short-sha>. Historical github-* experiments predate this identity contract and should be treated as historical evidence, not as current names or baseline candidates.

Safety and interpretation

  • Read-only is the default. Never delete experiments or upload raw stdout, stderr, environment/configuration, provider events, or arbitrary artifacts.
  • Never read or print BRAINTRUST_API_KEY, .bt/, Keychain contents, or any credential/environment value. CI scopes the key only to its exporter step.
  • Do not treat an agent's self-reported confidence as result quality. This phase has no scorer or quality score.
  • A failed or partial cell can be valid history when its normalized evidence is complete; distinguish that from a rejected suite or failed preflight.

Inspect experiments

Use the exercised project-option placement for list/view:

bash
bt experiments --json --project githits-cli-agent-evals list
bt experiments --json --project githits-cli-agent-evals view <experiment-name>

Use the experiment ID returned by the view result for a bounded field query:

bash
bt sql --json --non-interactive "SELECT input, output, metrics, metadata, tags FROM experiment('<experiment-id>') WHERE span_attributes.type = 'eval' LIMIT 100"
bt sql --json --non-interactive "SELECT name, span_attributes.type, metrics, metadata FROM experiment('<experiment-id>') WHERE span_attributes.type = 'tool' LIMIT 100"

The eval-root query is the verified path for prompts, neutral answers, hashes, statuses, native token/cost/duration metrics, and root metadata. Query tool children separately for native tool counts/errors and exact lifecycle timing. An unfiltered count(*) includes both eval roots and tool children, so it is not the workload-row count. A local proof experiment poc-native-tool-spans-v2-20260831 (ID e8480301-6622-4a06-a37b-0ebd0e42bb64, https://www.braintrust.dev/app/GitHits/p/githits-cli-agent-evals/experiments/poc-native-tool-spans-v2-20260831) read back two eval roots and 10 tool children. Native comparison reported tool_calls average 5.0 and tool_errors 0; child durations totaled 30.970 seconds and ranged from 0.006 to 10.400 seconds. Native token and cost fields remain populated. Open the experiment permalink when row-level UI inspection is useful.

The labeled CI path is proven by run 33424857668 at code SHA 7195ccc56b9ac9288dfb3d8de854f2f0e7ae7cf0. Its experiment is github-33424857668-1 (ID 182ee9db-0df3-40f4-8987-6eeb6d91a89b), source github, exporter/schema 2, metrics schema 3: 23 eval spans and 116 tool children, exactly matching 116 MCP calls, with zero CLI calls and zero failed tool spans. Totals were 513.911 seconds eval duration, 126.458999872 seconds tool duration, 2,686,094 prompt tokens, 20,172 completion tokens, 2,706,266 total tokens, and estimated cost $0.22819038. Compare averages were duration 22.343956532685652, estimated cost $0.009921320869565216, tool calls 5.043478260869565, tool errors 0, and total tokens 117663.73913043478. The first stable default-branch bootstrap is run 33477846273 at SHA 40796bd0eabaf87afec5ea0e4460ff47e7448603. Experiment main-r33477846273-a1 (ID 6f3847fc-3816-4b32-b1f6-65019c2757b7) read back 23 eval roots and 112 tool children, zero CLI calls, 3,024,404 tokens, 445.728 seconds cumulative agent duration, and estimated cost $0.24188221. Its null base is the expected one-time bootstrap result. Main pushes now temporarily run the same matrix, in addition to the daily/manual/label paths, to collect variance and workload-optimization evidence.

The current agent-eval-openrouter label runs the shared main matrix on trusted same-repository PRs: canary discovery plus stable-full intent and full guidance. eval/agentic/suites.json owns the workload counts. The trial PR must commit credential-free eval/agentic/openrouter.toml selecting its exact candidate model; the repository's blank-model example deliberately selects none. Keep the active config out of main and use OPENROUTER_API_KEY as its provider env_key, the only wired provider execution credential. The shared .github/workflows/agent-evals.yml retains Codex 0.154.0/prompt-json and execution-only OpenRouter auth; other triggers retain Luna/low/schema. The old DeepSeek label no longer starts a run. The dedicated canary workflow is removed. Each trial exports all three scenarios into one PR experiment with actual model/report-format metadata and linked main Luna baseline. Configuration support does not prove compatibility or quality of an untried model. Account for every cell and verify actual linked base/stable inputs; a single preset comparison is not a quality or consistency score. The following DeepSeek runs remain historical measured evidence.

The full comparison is live-proven by pr-401-r35099796991-a1 (ID a6313674-e0cd-45b0-8d5b-037d885f1876): exactly 50 eval roots and 495 tool children. Its persisted base is 13590571-39c1-4a33-831d-db144fb1fc7a (main-r35085880981-a1), with all 50 stable inputs matching. DeepSeek validates 48 reports versus main Luna's 50, takes 2950.371 versus 787.752 cumulative seconds, and uses 495 versus 205 MCP calls. Two malformed JSON finals cause the summary to fail while complete failed-cell export succeeds; neither timed out. Do not repair their finals or confuse successful export with successful cells. See the permanent comparison for scenario metrics, failed cells, source-path differences and interpretation.

The OpenRouter DeepSeek two-workload canary is proven by run 35093150512 on draft PR #401 at SHA e3fe68c40b80ac74d0c9fa59b0009b28c0841660. Experiment pr-401-r35093150512-a1 (ID cf6ec867-e67a-4adb-86bf-ace617b30dc0) read back two eval spans and 25 tool children, matching 25 completed MCP calls (package 5, router 20), zero failed calls and validated JSON finals. Metadata is DeepSeek/high/prompt-json, channel PR, exporter/schema 3. Actual base is main-r35085880981-a1 (ID 13590571-39c1-4a33-831d-db144fb1fc7a), sampled as Luna/low. This verifies PR linkage and integration. Cost remains unknown without a verified DeepSeek rate card; quality is ungraded and a single canary does not prove repeat consistency.

For current comparisons, inspect experiment-level metadata.channel and baseExperiment in the safe exporter result or CI summary. A current main baseline has a main-r...-a... name and channel: main; PR and local exports resolve the newest such main experiment before initialization. The exporter reports the actual linked base {id, name} after fetchBaseExperiment(). Validate-only reports the base as unresolved/not queried and performs no discovery. The first main run is a one-time bootstrap; PR and default-local exports fail before initialization when no main baseline exists. Explicit local --base-experiment takes precedence and skips discovery. Live readback has proven the first main bootstrap and the PR linkage recorded below. For exports, use the returned experiment name from the SDK readback; it can differ from a reused explicit local name if Braintrust de-duplicates it. Validate-only reports the requested or generated name.

Show full SKILL.md (406 more words)Show less

The exercised comparison syntax is:

bash
bt experiments --json --project githits-cli-agent-evals compare <experiment-a> <experiment-b>

For custom cross-experiment SQL analysis, join eval rows by metadata.cellId and verify identical stable input values (including promptSha256), rather than joining only by metadata.workloadId: the same workload can appear in multiple scenarios. Braintrust's built-in experiment comparison already matches the stable row inputs and avoids this ambiguity.

The prior custom-only experiments succeeded but reported only generic Braintrust trace metrics, which were zero and did not expose their custom eval telemetry. Treat that only as historical evidence about the older rows. The preceding native-root experiment is also historical: it set root tool_calls=119 and tool_errors=2, so comparison reported zero before structural children were implemented. Use bounded SQL and the experiment UI for GitHits-specific/custom telemetry; the current exporter uses exact harness-observed lifecycle boundaries and never fabricates timing.

Validate or explicitly export

Credential-free validation maps complete suite artifacts without initializing Braintrust:

bash
bun run agent:e2e:braintrust \
  --suite discovery=.agent-eval/suites/<discovery>/suite.json \
  --suite intent=.agent-eval/suites/<intent>/suite.json \
  --project githits-cli-agent-evals \
  --validate-only

An authenticated local subscription export uses the saved bt profile to run the same official entrypoint. The exporter, not bt, owns the experiment options and safe result file:

bash
bt eval --runner bun --no-auto-instrumentation scripts/agent-eval-braintrust.ts -- \
  --suite discovery=.agent-eval/suites/<discovery>/suite.json \
  --suite intent=.agent-eval/suites/<intent>/suite.json \
  --project githits-cli-agent-evals \
  --source local \
  --result-out .agent-eval/braintrust-result.json

This default local export lets the exporter derive its stable name and resolve the latest main baseline. Add --branch <branch> only when the evaluated suite is detached or has no branch; add --base-experiment <main-r...-a...> to use an explicit local main override. --experiment <name> is also a local-only override. GitHub workflow exports supply their channel, branch, PR number, run identity, and URL through environment-bound arguments and never pass --experiment.

The suite preflight rejects dry-run suites, suites with no workload cells, duplicate cells, mixed identity or schema contracts, and missing/unsafe child evidence before network setup. It does not reject a failed cell that retains complete report, metrics, workload, and contained prompt evidence. The result file is nonsecret and uses result-file schemaVersion: 2; it contains only mode, project, experiment, row count, suite summaries, an export URL when applicable, and baseExperiment. In validate-only mode baseExperiment: null means unresolved/not queried; in export mode null means the required Braintrust readback returned no actual linked base. Experiment metadata records exporter schema/version 3, including model, reasoning effort and Codex report format identity. Historical exporter/schema-2 experiments retain their recorded version. It never contains row bodies, prompts, answers, artifact paths, or credentials. Terminal tool-bearing rows lacking complete/valid observed lifecycle timing are rejected because they cannot produce accurate structural children; an observed started-only call remains an open child. Zero-tool legacy rows remain exportable. Do not create or upload a new experiment unless the user explicitly requests that export.

© githits-com, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/braintrust-agent-evals of githits-com/githits-cli.

Open the folder on GitHubat commit 7449018

Compare with similar skills

Braintrust Agent Evals next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Braintrust Agent Evals compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Braintrust Agent Evals this skillgithits-com/githits-cli114—~3.1kAutomated safety check: PassApache-2.0
LLM Eval Pipeline Auditai-evals-course/evals-skills1.5k—~2.5kAutomated safety check: PassApache-2.0
Opik Evaluatecomet-ml/opik-mcp219—~2.5kAutomated safety check: NotesApache-2.0
Compliance Drift Evalsucsandman/DashClaw310—~1.8kAutomated safety check: PassMIT
Caveman Experiment ManagerJuliusBrussee/caveman110k1 repos~975Automated safety check: PassApache-2.0
Agents Best PracticesDenisSergeevitch/agents-best-practices2.4k—~7.4kAutomated safety check: PassMIT

Similar skills

  • LLM Eval Pipeline Audit

    ai-evals-course/evals-skills

    Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes.

    1.5k GitHub stars~2.5k tokensUpdated 14 days ago
    AI & LLM EngineeringAuto-check passed
  • Opik Evaluate

    comet-ml/opik-mcp

    Build an LLM evaluation and run it against the app, returning an Opik experiment with scores and its link.

    219 GitHub stars~2.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Compliance Drift Evals

    ucsandman/DashClaw

    Set up compliance exports, drift detection, evaluations, scoring, and learning analytics

    310 GitHub stars~1.8k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Caveman Experiment Manager

    JuliusBrussee/caveman

    Reads the state and results of Caveman Cloud experiments and reports one recommendation or a block, without changing an experiment's lifecycle itself.

    110k GitHub starsUsed in 1 repo~975 tokens
    AI & LLM EngineeringAuto-check passed
  • Agents Best Practices

    DenisSergeevitch/agents-best-practices

    A skill your agent uses when designing, generating an MVP blueprint for, auditing, troubleshooting, refactoring, or explaining an agentic harness for any domain.

    2.4k GitHub stars~7.4k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.

    949 GitHub stars~2.1k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed

More from githits-com/githits-cli

  • Githits Code

    githits-com/githits-cli

    A skill your agent uses whenever invoking the GitHits CLI for public OSS source, documentation, or example evidence, including code search/grep, file navigation, source verification, docs lookup, or…

    114 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Githits Onboarding

    githits-com/githits-cli

    A skill your agent uses when the user asks to install, connect, configure, sign in to, sign up for, or start using GitHits.

    114 GitHub stars~3.5k tokensUpdated yesterday
    Auto-check passed
  • Githits Package

    githits-com/githits-cli

    A skill your agent uses whenever invoking the GitHits CLI for public package or dependency evidence, including metadata, versions, licenses, vulnerabilities, dependency graphs, changelogs, release…

    114 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Githits Plugin Maintenance

    githits-com/githits-cli

    Internal repository-maintenance skill for GitHits cross-host plugin and Agent Skill surfaces.

    114 GitHub stars~1k tokensUpdated yesterday
    Auto-check passed
  • Githits Release

    githits-com/githits-cli

    A skill your agent uses when maintaining the GitHits changelog or preparing, reviewing, or executing a GitHits release.

    114 GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed
  • Githits MCP

    githits-com/githits-cli

    Route public OSS code, documentation, examples, and package questions to GitHits tools.

    114 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed

Questions about Braintrust Agent Evals

What does Braintrust Agent Evals do?

Inspect, query, compare, or explicitly export GitHits agent-eval history in Braintrust using the repository's verified workflow. Braintrust Agent Evals is an agent skill from githits-com/githits-cli. Inspect, query, compare, or explicitly export GitHits agent-eval history in Braintrust using the repository's verified workflow.

When should I use Braintrust Agent Evals?

Braintrust Agent Evals fits situations like: tasks that involve LLM evaluation.

How do I install Braintrust Agent Evals in Claude Code?

Run `npx skills add githits-com/githits-cli --skill braintrust-agent-evals -a claude-code`. Or copy the skill folder (.agents/skills/braintrust-agent-evals in githits-com/githits-cli) into .claude/skills/braintrust-agent-evals in your project. Claude Code loads it when a task matches its description.

How do I install Braintrust Agent Evals in Codex?

Run `npx skills add githits-com/githits-cli --skill braintrust-agent-evals -a codex`. Or copy the skill folder (.agents/skills/braintrust-agent-evals in githits-com/githits-cli) into .agents/skills/braintrust-agent-evals in your project. Codex loads it when a task matches its description.

Can I use Braintrust Agent Evals in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add githits-com/githits-cli --skill braintrust-agent-evals -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/braintrust-agent-evals, .gemini/skills/braintrust-agent-evals, .github/skills/braintrust-agent-evals and .opencode/skills/braintrust-agent-evals in your project.

What does Braintrust Agent Evals need to run?

Going by SKILL.md and its folder, Braintrust Agent Evals needs the command-line tools its instructions call (bun) and credentials named BRAINTRUST_API_KEY and OPENROUTER_API_KEY. Our summary lists: A credential in BRAINTRUST_API_KEY; A credential in OPENROUTER_API_KEY.

Does Braintrust Agent Evals access the network?

SKILL.md names 1 domain. As links in the text: braintrust.dev. This is read from the text; nothing was executed.

Is Braintrust Agent Evals safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Braintrust Agent Evals use?

Braintrust Agent Evals is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Braintrust Agent Evals use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Braintrust Agent Evals?

Skills that share tags, products or a category with Braintrust Agent Evals: LLM Eval Pipeline Audit (ai-evals-course/evals-skills, 1.5k stars), Opik Evaluate (comet-ml/opik-mcp, 219 stars), Compliance Drift Evals (ucsandman/DashClaw, 310 stars) and Caveman Experiment Manager (JuliusBrussee/caveman, 110k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Braintrust Agent Evals?

githits-com (a GitHub organization) maintains it in githits-com/githits-cli, which has 114 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 7, 2026.

Source: githits-com/githits-cli on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.