SDK AI Bot Run Evaluation
Azure/azure-sdk-tools
Run Azure SDK QA bot evaluations on curated datasets locally, including a single test case.
Turn a vague task into a testable definition of done and generate an executable goal-test script for it, optionally with a bounded retry loop around a headless agent.
$ npx skills add undefined-ui/second-brain-os --skill goal-test -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install undefined-ui/second-brain-os goal-test --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/undefined-ui/second-brain-os.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/agents-course/skills/goal-test .claude/skills/goal-test && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "goal-test" agent skill from https://github.com/undefined-ui/second-brain-os/tree/main/plugins/agents-course/skills/goal-test into .claude/skills/goal-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "goal-test", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/undefined-ui/second-brain-os/tree/main/plugins/agents-course/skills/goal-testType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add undefined-ui/second-brain-os --skill goal-test -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install undefined-ui/second-brain-os goal-test --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/undefined-ui/second-brain-os.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/agents-course/skills/goal-test .agents/skills/goal-test && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "goal-test" agent skill from https://github.com/undefined-ui/second-brain-os/tree/main/plugins/agents-course/skills/goal-test into .agents/skills/goal-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "goal-test", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add undefined-ui/second-brain-os --skill goal-test -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install undefined-ui/second-brain-os goal-test --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/undefined-ui/second-brain-os.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/agents-course/skills/goal-test .cursor/skills/goal-test && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "goal-test" agent skill from https://github.com/undefined-ui/second-brain-os/tree/main/plugins/agents-course/skills/goal-test into .cursor/skills/goal-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "goal-test", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/undefined-ui/second-brain-os.git --path plugins/agents-course/skills/goal-test--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add undefined-ui/second-brain-os --skill goal-test -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install undefined-ui/second-brain-os goal-test --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/undefined-ui/second-brain-os.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/agents-course/skills/goal-test .gemini/skills/goal-test && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "goal-test" agent skill from https://github.com/undefined-ui/second-brain-os/tree/main/plugins/agents-course/skills/goal-test into .gemini/skills/goal-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "goal-test", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install undefined-ui/second-brain-os goal-testInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add undefined-ui/second-brain-os --skill goal-test -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/undefined-ui/second-brain-os.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/agents-course/skills/goal-test .github/skills/goal-test && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "goal-test" agent skill from https://github.com/undefined-ui/second-brain-os/tree/main/plugins/agents-course/skills/goal-test into .github/skills/goal-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "goal-test", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add undefined-ui/second-brain-os --skill goal-test -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install undefined-ui/second-brain-os goal-test --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/undefined-ui/second-brain-os.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/agents-course/skills/goal-test .opencode/skills/goal-test && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "goal-test" agent skill from https://github.com/undefined-ui/second-brain-os/tree/main/plugins/agents-course/skills/goal-test into .opencode/skills/goal-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "goal-test", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
goal-testTurn a vague task into a testable definition of done and generate an executable goal-test script for it, optionally with a bounded retry loop around a headless agent.
Goal Test is an agent skill from undefined-ui/second-brain-os. Turn a vague task into a testable definition of done and generate an executable goal-test script for it, optionally with a bounded retry loop around a headless agent. Use when the user wants to run an agent in a loop, asks how to know when an agent task is finished, or says a goal like "improve X" needs to become checkable. Do NOT use for building eval suites over many cases — that is evals-bootstrap.
Its SKILL.md is about 750 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Knowledge Management, covering LLM evaluation. The repository describes itself as: An AI second brain that maintains itself. Full guide, starter vault, agent skills and scripts for a self-organizing knowledge base in Claude Code and Obsidian. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit c7fa35b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
claudemakeFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
secondbrainos.devFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Goal Test loads about 754 tokens when it runs. Until then it costs about 104 tokens; SKILL.md has 326 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from undefined-ui/second-brain-os at commit c7fa35b, republished under its MIT licence (© undefined-ui). 326 words, ~754 tokens.
.claude/skills/goal-test/SKILL.md (or your agent's skills folder).Theory: The four parts of a loop and the goal test build page. "Improve the error handling" cannot terminate a loop, because nothing can ever say it is finished. The goal must be phrased so that a program — not a person, not the model — returns true or false against it.
make test is
better than one that reimplements it.goal-test.sh (or .py if the checks are easier there) in the
project root: runs every check, prints one line per check with pass/fail,
exits 0 only when all pass. Keep it under ~40 lines; it must run in
seconds and be safe to run repeatedly.#!/usr/bin/env bash
# goal-loop.sh — run a headless agent until the goal test passes or attempts run out.
set -u
MAX_ATTEMPTS=5
for i in $(seq 1 "$MAX_ATTEMPTS"); do
FEEDBACK=$(./goal-test.sh 2>&1) && { echo "done in $i attempt(s)"; exit 0; }
claude -p "Goal: <the goal>. The goal test currently fails with:
$FEEDBACK
Fix the code so the goal test passes." \
--permission-mode acceptEdits --output-format json > ".attempt-$i.json"
done
echo "goal test still failing after $MAX_ATTEMPTS attempts — falling back to a human"
exit 1Adjust the agent command to whatever CLI the user runs. Keep the three
brakes visible and named: the checker outside the model (goal-test.sh),
the stop rule (test passes), the budget (MAX_ATTEMPTS, plus a dollar cap
read from the JSON output if they want one).
© undefined-ui, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in plugins/agents-course/skills/goal-test of undefined-ui/second-brain-os.
Open the folder on GitHubat commit c7fa35b
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in undefined-ui/second-brain-os, which our catalogue first saw on October 7, 2026.
Goal Test next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Goal Test this skillundefined-ui/second-brain-os | 1k | 1 repos | ~754 | Automated safety check: Pass | MIT | |
| SDK AI Bot Run EvaluationAzure/azure-sdk-tools | 134 | — | ~1.1k | Automated safety check: Notes | MIT | |
| Qmdalsk1992/CloddsBot | 2.9k | 3 repos | ~1.2k | Automated safety check: Pass | MIT | |
| Gitnexus CLIaws-samples/sample-kolya-br-proxy | 106 | 10 repos | ~822 | Automated safety check: Pass | MIT-0 | |
| Ragu BuildRaguTeam/RAGU | 135 | — | ~3.8k | Automated safety check: Notes | MIT | |
| Analyze Id Eval Rankingopen-thoughts/OpenThoughts-Agent | 301 | — | ~3.1k | Automated safety check: Pass | Apache-2.0 |
Azure/azure-sdk-tools
Run Azure SDK QA bot evaluations on curated datasets locally, including a single test case.
alsk1992/CloddsBot
Local hybrid search for markdown notes and docs. An agent skill from alsk1992/CloddsBot.
aws-samples/sample-kolya-br-proxy
A skill your agent uses when the user needs to run GitNexus CLI commands like analyze/index a repo, check status, clean the index, generate a wiki, or list indexed repos.
RaguTeam/RAGU
Interview the user about their RAGU use case, select an appropriate RAGU pipeline, and generate a validated ragubuild.yaml plus a runnable build<name.py script.
open-thoughts/OpenThoughts-Agent
Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…
Health-Yang/MineEcho
技能系统导航。当用户询问"你能做什么"、"有什么功能"、"有什么技能"或不确定如何完成某个任务时,使用此技能列出所有可用技能并建议最合适的技能。此技能是每个对话开始时默认加载的,用于技能发现。
undefined-ui/second-brain-os
Audit an agent's context layout against the four places: system prompt, tools, history, tail.
undefined-ui/second-brain-os
Scaffold a first eval suite for an agent: mine real failures into cases, write behavioural checks over traces, and generate the runner.
undefined-ui/second-brain-os
Find the decisions in a pipeline that do not need the expensive model and propose the gate for each: a rule, a classic classifier, or a small model, with fail-closed routing.
undefined-ui/second-brain-os
Audit an agent's harness against the four rings: containment, guides, sensors, permissions.
undefined-ui/second-brain-os
Triage an exported chat history into what to ingest, what to archive and what to delete, with a privacy pass first.
undefined-ui/second-brain-os
Analyse the vault's link graph and report on its shape: orphan rate, average degree, components, hubs, bridges and clusters, with what each number means for retrieval.
Categories
Turn a vague task into a testable definition of done and generate an executable goal-test script for it, optionally with a bounded retry loop around a headless agent. Goal Test is an agent skill from undefined-ui/second-brain-os. Turn a vague task into a testable definition of done and generate an executable goal-test script for it, optionally with a bounded retry loop around a headless agent.
Goal Test fits situations like: the user wants to run an agent in a loop; asks how to know when an agent task is finished; says a goal like improve X needs to become checkable; building eval suites over many cases — that is evals-bootstrap.
Run `npx skills add undefined-ui/second-brain-os --skill goal-test -a claude-code`. Or copy the skill folder (plugins/agents-course/skills/goal-test in undefined-ui/second-brain-os) into .claude/skills/goal-test in your project. Claude Code loads it when a task matches its description.
Run `npx skills add undefined-ui/second-brain-os --skill goal-test -a codex`. Or copy the skill folder (plugins/agents-course/skills/goal-test in undefined-ui/second-brain-os) into .agents/skills/goal-test in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add undefined-ui/second-brain-os --skill goal-test -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/goal-test, .gemini/skills/goal-test, .github/skills/goal-test and .opencode/skills/goal-test in your project.
Going by SKILL.md and its folder, Goal Test needs the command-line tools its instructions call (claude and make).
SKILL.md names 1 domain. As links in the text: secondbrainos.dev. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Goal Test is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 754 tokens (SKILL.md is roughly 3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Goal Test: SDK AI Bot Run Evaluation (Azure/azure-sdk-tools, 134 stars), Qmd (alsk1992/CloddsBot, 2.9k stars), Gitnexus CLI (aws-samples/sample-kolya-br-proxy, 106 stars) and Ragu Build (RaguTeam/RAGU, 135 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
undefined-ui (a GitHub user) maintains it in undefined-ui/second-brain-os, which has 1,009 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 7, 2026.
Source: undefined-ui/second-brain-os on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.