LLM Eval Pipeline Audit
ai-evals-course/evals-skills
Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes.
Inspect, query, compare, or explicitly export GitHits agent-eval history in Braintrust using the repository's verified workflow.
$ npx skills add githits-com/githits-cli --skill braintrust-agent-evals -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install githits-com/githits-cli braintrust-agent-evals --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/githits-com/githits-cli.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/braintrust-agent-evals .claude/skills/braintrust-agent-evals && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "braintrust-agent-evals" agent skill from https://github.com/githits-com/githits-cli/tree/main/.agents/skills/braintrust-agent-evals into .claude/skills/braintrust-agent-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "braintrust-agent-evals", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/githits-com/githits-cli/tree/main/.agents/skills/braintrust-agent-evalsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add githits-com/githits-cli --skill braintrust-agent-evals -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install githits-com/githits-cli braintrust-agent-evals --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/githits-com/githits-cli.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/braintrust-agent-evals .agents/skills/braintrust-agent-evals && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "braintrust-agent-evals" agent skill from https://github.com/githits-com/githits-cli/tree/main/.agents/skills/braintrust-agent-evals into .agents/skills/braintrust-agent-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "braintrust-agent-evals", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add githits-com/githits-cli --skill braintrust-agent-evals -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install githits-com/githits-cli braintrust-agent-evals --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/githits-com/githits-cli.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/braintrust-agent-evals .cursor/skills/braintrust-agent-evals && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "braintrust-agent-evals" agent skill from https://github.com/githits-com/githits-cli/tree/main/.agents/skills/braintrust-agent-evals into .cursor/skills/braintrust-agent-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "braintrust-agent-evals", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/githits-com/githits-cli.git --path .agents/skills/braintrust-agent-evals--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add githits-com/githits-cli --skill braintrust-agent-evals -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install githits-com/githits-cli braintrust-agent-evals --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/githits-com/githits-cli.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/braintrust-agent-evals .gemini/skills/braintrust-agent-evals && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "braintrust-agent-evals" agent skill from https://github.com/githits-com/githits-cli/tree/main/.agents/skills/braintrust-agent-evals into .gemini/skills/braintrust-agent-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "braintrust-agent-evals", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install githits-com/githits-cli braintrust-agent-evalsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add githits-com/githits-cli --skill braintrust-agent-evals -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/githits-com/githits-cli.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/braintrust-agent-evals .github/skills/braintrust-agent-evals && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "braintrust-agent-evals" agent skill from https://github.com/githits-com/githits-cli/tree/main/.agents/skills/braintrust-agent-evals into .github/skills/braintrust-agent-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "braintrust-agent-evals", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add githits-com/githits-cli --skill braintrust-agent-evals -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install githits-com/githits-cli braintrust-agent-evals --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/githits-com/githits-cli.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/braintrust-agent-evals .opencode/skills/braintrust-agent-evals && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "braintrust-agent-evals" agent skill from https://github.com/githits-com/githits-cli/tree/main/.agents/skills/braintrust-agent-evals into .opencode/skills/braintrust-agent-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "braintrust-agent-evals", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
braintrust-agent-evalsInspect, query, compare, or explicitly export GitHits agent-eval history in Braintrust using the repository's verified workflow.
Braintrust Agent Evals is an agent skill from githits-com/githits-cli. Inspect, query, compare, or explicitly export GitHits agent-eval history in Braintrust using the repository's verified workflow.
Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM evaluation. It works with Model Context Protocol. The repository describes itself as: CLI & MCP for GitHits - The Code Context Layer for AI Coding Agents. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 7449018. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
bunFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
braintrust.devFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
BRAINTRUST_API_KEYOPENROUTER_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Braintrust Agent Evals loads about 3.1k tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 1,347 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from githits-com/githits-cli at commit 7449018, republished under its Apache-2.0 licence (© githits-com). 1,347 words, ~3,102 tokens.
.claude/skills/braintrust-agent-evals/SKILL.md (or your agent's skills folder).Use this skill for read-only inspection of the GitHits agent-eval history in the
Braintrust project githits-cli-agent-evals, or when the user explicitly asks
to export a validated local suite. The normalized exporter stores one top-level
eval span per scenario/workload cell plus structural tool children; see
docs/implementation/agentic-eval-metrics.md
for the field contract.
The persistence unit is one exporter invocation = one experiment, one
scenario/workload cell = one eval row, and one normalized logical tool call =
one structural tool child. Current exporter-owned experiment names are
main-r<run-id>-a<attempt>, pr-<pr-number>-r<run-id>-a<attempt>, and
local-<branch-slug>-<UTC-timestamp-with-milliseconds>-<short-sha>. Historical
github-* experiments predate this identity contract and should be treated as
historical evidence, not as current names or baseline candidates.
BRAINTRUST_API_KEY, .bt/, Keychain contents, or any
credential/environment value. CI scopes the key only to its exporter step.Use the exercised project-option placement for list/view:
bt experiments --json --project githits-cli-agent-evals list
bt experiments --json --project githits-cli-agent-evals view <experiment-name>Use the experiment ID returned by the view result for a bounded field query:
bt sql --json --non-interactive "SELECT input, output, metrics, metadata, tags FROM experiment('<experiment-id>') WHERE span_attributes.type = 'eval' LIMIT 100"
bt sql --json --non-interactive "SELECT name, span_attributes.type, metrics, metadata FROM experiment('<experiment-id>') WHERE span_attributes.type = 'tool' LIMIT 100"The eval-root query is the verified path for prompts, neutral answers, hashes,
statuses, native token/cost/duration metrics, and root metadata. Query tool
children separately for native tool counts/errors and exact lifecycle timing.
An unfiltered count(*) includes both eval roots and tool children, so it is not
the workload-row count. A local proof experiment
poc-native-tool-spans-v2-20260831 (ID
e8480301-6622-4a06-a37b-0ebd0e42bb64,
https://www.braintrust.dev/app/GitHits/p/githits-cli-agent-evals/experiments/poc-native-tool-spans-v2-20260831)
read back two eval roots and 10 tool children. Native comparison reported
tool_calls average 5.0 and
tool_errors 0; child durations totaled 30.970 seconds and ranged from
0.006 to 10.400 seconds. Native token and cost fields remain populated. Open
the experiment permalink when row-level UI inspection is useful.
The labeled CI path is proven by run
33424857668
at code SHA 7195ccc56b9ac9288dfb3d8de854f2f0e7ae7cf0. Its experiment is
github-33424857668-1 (ID 182ee9db-0df3-40f4-8987-6eeb6d91a89b), source
github, exporter/schema 2, metrics schema 3: 23 eval spans and 116 tool
children, exactly matching 116 MCP calls, with zero CLI calls and zero failed
tool spans. Totals were 513.911 seconds eval duration, 126.458999872 seconds tool
duration, 2,686,094 prompt tokens, 20,172 completion tokens, 2,706,266 total
tokens, and estimated cost $0.22819038. Compare averages were duration
22.343956532685652, estimated cost $0.009921320869565216, tool calls
5.043478260869565, tool errors 0, and total tokens 117663.73913043478.
The first stable default-branch bootstrap is run
33477846273
at SHA 40796bd0eabaf87afec5ea0e4460ff47e7448603. Experiment
main-r33477846273-a1 (ID 6f3847fc-3816-4b32-b1f6-65019c2757b7) read back
23 eval roots and 112 tool children, zero CLI calls, 3,024,404 tokens,
445.728 seconds cumulative agent duration, and estimated cost $0.24188221.
Its null base is the expected one-time bootstrap result. Main pushes now
temporarily run the same matrix, in addition to the daily/manual/label paths,
to collect variance and workload-optimization evidence.
The current agent-eval-openrouter label runs the shared main matrix on
trusted same-repository PRs: canary discovery plus stable-full intent and
full guidance. eval/agentic/suites.json owns the workload counts.
The trial PR must commit credential-free eval/agentic/openrouter.toml selecting
its exact candidate model; the repository's blank-model example deliberately
selects none. Keep the active config out of main and use OPENROUTER_API_KEY as
its provider env_key, the only wired provider execution credential. The shared
.github/workflows/agent-evals.yml retains Codex 0.154.0/prompt-json and
execution-only OpenRouter auth; other triggers retain Luna/low/schema. The old
DeepSeek label no longer starts a run. The dedicated canary workflow is removed.
Each trial exports all three scenarios into one PR experiment with actual
model/report-format metadata and linked main Luna baseline. Configuration
support does not prove compatibility or quality of an untried model. Account
for every cell and verify actual linked base/stable inputs; a single preset
comparison is not a quality or consistency score. The following DeepSeek runs
remain historical measured evidence.
The full comparison is live-proven by
pr-401-r35099796991-a1
(ID a6313674-e0cd-45b0-8d5b-037d885f1876): exactly 50 eval roots and 495 tool
children. Its persisted base is 13590571-39c1-4a33-831d-db144fb1fc7a
(main-r35085880981-a1), with all 50 stable inputs matching. DeepSeek validates
48 reports versus main Luna's 50, takes 2950.371 versus 787.752 cumulative
seconds, and uses 495 versus 205 MCP calls. Two malformed JSON finals cause the
summary to fail while complete failed-cell export succeeds; neither timed out.
Do not repair their finals or confuse successful export with successful cells.
See the permanent comparison
for scenario metrics, failed cells, source-path differences and interpretation.
The OpenRouter DeepSeek two-workload canary is proven by run
35093150512
on draft PR #401 at SHA e3fe68c40b80ac74d0c9fa59b0009b28c0841660. Experiment
pr-401-r35093150512-a1
(ID cf6ec867-e67a-4adb-86bf-ace617b30dc0) read back two eval spans and 25 tool
children, matching 25 completed MCP calls (package 5, router 20), zero failed
calls and validated JSON finals. Metadata is DeepSeek/high/prompt-json, channel
PR, exporter/schema 3. Actual base is main-r35085880981-a1
(ID 13590571-39c1-4a33-831d-db144fb1fc7a), sampled as Luna/low. This verifies PR
linkage and integration. Cost remains unknown without a verified DeepSeek rate
card; quality is ungraded and a single canary does not prove repeat consistency.
For current comparisons, inspect experiment-level metadata.channel and
baseExperiment in the safe exporter result or CI summary. A current main
baseline has a main-r...-a... name and channel: main; PR and local exports
resolve the newest such main experiment before initialization. The exporter
reports the actual linked base {id, name} after fetchBaseExperiment().
Validate-only reports the base as unresolved/not queried and performs no
discovery. The first main run is a one-time bootstrap; PR and default-local
exports fail before initialization when no main baseline exists. Explicit
local --base-experiment takes precedence and skips discovery. Live
readback has proven the first main bootstrap and the PR linkage recorded below.
For exports, use the returned experiment name from the SDK readback; it can
differ from a reused explicit local name if Braintrust de-duplicates it.
Validate-only reports the requested or generated name.
The exercised comparison syntax is:
bt experiments --json --project githits-cli-agent-evals compare <experiment-a> <experiment-b>For custom cross-experiment SQL analysis, join eval rows by
metadata.cellId and verify identical stable input values (including
promptSha256), rather than joining only by metadata.workloadId: the same workload can appear in
multiple scenarios. Braintrust's built-in experiment comparison already
matches the stable row inputs and avoids this ambiguity.
The prior custom-only experiments succeeded but reported only generic
Braintrust trace metrics, which were zero and did not expose their custom eval
telemetry. Treat that only as historical evidence about the older rows. The
preceding native-root experiment is also historical: it set root
tool_calls=119 and tool_errors=2, so comparison reported zero before
structural children were implemented. Use bounded SQL and the experiment UI for
GitHits-specific/custom telemetry; the current exporter uses exact
harness-observed lifecycle boundaries and never fabricates timing.
Credential-free validation maps complete suite artifacts without initializing Braintrust:
bun run agent:e2e:braintrust \
--suite discovery=.agent-eval/suites/<discovery>/suite.json \
--suite intent=.agent-eval/suites/<intent>/suite.json \
--project githits-cli-agent-evals \
--validate-onlyAn authenticated local subscription export uses the saved bt profile to run
the same official entrypoint. The exporter, not bt, owns the experiment
options and safe result file:
bt eval --runner bun --no-auto-instrumentation scripts/agent-eval-braintrust.ts -- \
--suite discovery=.agent-eval/suites/<discovery>/suite.json \
--suite intent=.agent-eval/suites/<intent>/suite.json \
--project githits-cli-agent-evals \
--source local \
--result-out .agent-eval/braintrust-result.jsonThis default local export lets the exporter derive its stable name and resolve
the latest main baseline. Add --branch <branch> only when the evaluated suite
is detached or has no branch; add --base-experiment <main-r...-a...> to use an
explicit local main override. --experiment <name> is also a local-only
override. GitHub workflow exports supply their channel, branch, PR number, run
identity, and URL through environment-bound arguments and never pass
--experiment.
The suite preflight rejects dry-run suites, suites with no workload cells,
duplicate cells, mixed identity or schema contracts, and missing/unsafe child
evidence before network setup. It does not reject a failed cell that retains
complete report, metrics, workload, and contained prompt evidence. The result
file is nonsecret and uses result-file schemaVersion: 2; it contains only
mode, project, experiment, row count, suite summaries, an export URL when
applicable, and baseExperiment. In validate-only mode baseExperiment: null
means unresolved/not queried; in export mode null means the required
Braintrust readback returned no actual linked base. Experiment metadata records
exporter schema/version 3, including model, reasoning effort and Codex report
format identity. Historical exporter/schema-2 experiments retain their recorded
version. It never contains row bodies, prompts, answers,
artifact paths, or credentials.
Terminal tool-bearing rows lacking complete/valid observed lifecycle timing are
rejected because they cannot produce accurate structural children; an observed
started-only call remains an open child. Zero-tool legacy rows remain
exportable. Do not create or upload a new experiment unless the user explicitly
requests that export.
© githits-com, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/braintrust-agent-evals of githits-com/githits-cli.
Open the folder on GitHubat commit 7449018
Braintrust Agent Evals next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Braintrust Agent Evals this skillgithits-com/githits-cli | 114 | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | |
| LLM Eval Pipeline Auditai-evals-course/evals-skills | 1.5k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Opik Evaluatecomet-ml/opik-mcp | 219 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | |
| Compliance Drift Evalsucsandman/DashClaw | 310 | — | ~1.8k | Automated safety check: Pass | MIT | |
| Caveman Experiment ManagerJuliusBrussee/caveman | 110k | 1 repos | ~975 | Automated safety check: Pass | Apache-2.0 | |
| Agents Best PracticesDenisSergeevitch/agents-best-practices | 2.4k | — | ~7.4k | Automated safety check: Pass | MIT |
ai-evals-course/evals-skills
Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes.
comet-ml/opik-mcp
Build an LLM evaluation and run it against the app, returning an Opik experiment with scores and its link.
ucsandman/DashClaw
Set up compliance exports, drift detection, evaluations, scoring, and learning analytics
JuliusBrussee/caveman
Reads the state and results of Caveman Cloud experiments and reports one recommendation or a block, without changing an experiment's lifecycle itself.
DenisSergeevitch/agents-best-practices
A skill your agent uses when designing, generating an MVP blueprint for, auditing, troubleshooting, refactoring, or explaining an agentic harness for any domain.
bgauryy/octocode
Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.
githits-com/githits-cli
A skill your agent uses whenever invoking the GitHits CLI for public OSS source, documentation, or example evidence, including code search/grep, file navigation, source verification, docs lookup, or…
githits-com/githits-cli
A skill your agent uses when the user asks to install, connect, configure, sign in to, sign up for, or start using GitHits.
githits-com/githits-cli
A skill your agent uses whenever invoking the GitHits CLI for public package or dependency evidence, including metadata, versions, licenses, vulnerabilities, dependency graphs, changelogs, release…
githits-com/githits-cli
Internal repository-maintenance skill for GitHits cross-host plugin and Agent Skill surfaces.
githits-com/githits-cli
A skill your agent uses when maintaining the GitHits changelog or preparing, reviewing, or executing a GitHits release.
githits-com/githits-cli
Route public OSS code, documentation, examples, and package questions to GitHits tools.
Works with
Categories
Inspect, query, compare, or explicitly export GitHits agent-eval history in Braintrust using the repository's verified workflow. Braintrust Agent Evals is an agent skill from githits-com/githits-cli. Inspect, query, compare, or explicitly export GitHits agent-eval history in Braintrust using the repository's verified workflow.
Braintrust Agent Evals fits situations like: tasks that involve LLM evaluation.
Run `npx skills add githits-com/githits-cli --skill braintrust-agent-evals -a claude-code`. Or copy the skill folder (.agents/skills/braintrust-agent-evals in githits-com/githits-cli) into .claude/skills/braintrust-agent-evals in your project. Claude Code loads it when a task matches its description.
Run `npx skills add githits-com/githits-cli --skill braintrust-agent-evals -a codex`. Or copy the skill folder (.agents/skills/braintrust-agent-evals in githits-com/githits-cli) into .agents/skills/braintrust-agent-evals in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add githits-com/githits-cli --skill braintrust-agent-evals -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/braintrust-agent-evals, .gemini/skills/braintrust-agent-evals, .github/skills/braintrust-agent-evals and .opencode/skills/braintrust-agent-evals in your project.
Going by SKILL.md and its folder, Braintrust Agent Evals needs the command-line tools its instructions call (bun) and credentials named BRAINTRUST_API_KEY and OPENROUTER_API_KEY. Our summary lists: A credential in BRAINTRUST_API_KEY; A credential in OPENROUTER_API_KEY.
SKILL.md names 1 domain. As links in the text: braintrust.dev. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Braintrust Agent Evals is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Braintrust Agent Evals: LLM Eval Pipeline Audit (ai-evals-course/evals-skills, 1.5k stars), Opik Evaluate (comet-ml/opik-mcp, 219 stars), Compliance Drift Evals (ucsandman/DashClaw, 310 stars) and Caveman Experiment Manager (JuliusBrussee/caveman, 110k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
githits-com (a GitHub organization) maintains it in githits-com/githits-cli, which has 114 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 7, 2026.
Source: githits-com/githits-cli on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.