Agent skill

Goal Test

by undefined-ui in undefined-ui/second-brain-os

Turn a vague task into a testable definition of done and generate an executable goal-test script for it, optionally with a bounded retry loop around a headless agent.

MITAuto-check passedKnowledge Management

Install Goal Test

skills CLI
$ npx skills add undefined-ui/second-brain-os --skill goal-test -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install undefined-ui/second-brain-os goal-test --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/undefined-ui/second-brain-os.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/agents-course/skills/goal-test .claude/skills/goal-test && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
goal-test
GitHub stars
1k
Used in
1 other repo
Token cost
~754 tokens
SKILL.md length
326 words
Files
1
Skills in repo
23
Repo updated
First seen
Licence
MIT

At a glance

Turn a vague task into a testable definition of done and generate an executable goal-test script for it, optionally with a bounded retry loop around a headless agent.

  • Works in 5 steps: Extract the claim. Ask what the user… → Prefer checks that already exist. The… → Write goal-test.sh (or .py if the checks… → …
  • The user wants to run an agent in a loop
  • SKILL.md covers Workflow and Rules
  • Calls claude and make

What it does

Goal Test is an agent skill from undefined-ui/second-brain-os. Turn a vague task into a testable definition of done and generate an executable goal-test script for it, optionally with a bounded retry loop around a headless agent. Use when the user wants to run an agent in a loop, asks how to know when an agent task is finished, or says a goal like "improve X" needs to become checkable. Do NOT use for building eval suites over many cases — that is evals-bootstrap.

Its SKILL.md is about 750 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Knowledge Management, covering LLM evaluation. The repository describes itself as: An AI second brain that maintains itself. Full guide, starter vault, agent skills and scripts for a self-organizing knowledge base in Claude Code and Obsidian. The licence is MIT.

When your agent uses it

  • The user wants to run an agent in a loop
  • Asks how to know when an agent task is finished
  • Says a goal like improve X needs to become checkable
  • Building eval suites over many cases — that is evals-bootstrap

Example prompts

  • “improve X”
  • “/goal-test”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Extract the claim. Ask what the user wants true at the end, then
  2. Prefer checks that already exist. The test suite, the linter, the
  3. Write goal-test.sh (or .py if the checks are easier there) in the
  4. Run it now. It should fail before the work is done — a goal test that
  5. Offer the loop. If the user wants the agent driven until done, wrap it

What it can do on your machine

Read from SKILL.md and the folder at commit c7fa35b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • claude
    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • secondbrainos.dev

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Goal Test loads about 754 tokens when it runs. Until then it costs about 104 tokens; SKILL.md has 326 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~104
When it runs · the whole SKILL.md, loaded when a task matches
~754

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from undefined-ui/second-brain-os at commit c7fa35b, republished under its MIT licence (© undefined-ui). 326 words, ~754 tokens.

Download SKILL.mdSave it as .claude/skills/goal-test/SKILL.md (or your agent's skills folder).
name
goal-test
description
Turn a vague task into a testable definition of done and generate an executable goal-test script for it, optionally with a bounded retry loop around a headless agent. Use when the user wants to run an agent in a loop, asks how to know when an agent task is finished, or says a goal like "improve X" needs to become checkable. Do NOT use for building eval suites over many cases — that is evals-bootstrap.

Write done as a script

Theory: The four parts of a loop and the goal test build page. "Improve the error handling" cannot terminate a loop, because nothing can ever say it is finished. The goal must be phrased so that a program — not a person, not the model — returns true or false against it.

Workflow

  1. Extract the claim. Ask what the user wants true at the end, then restate it as verifiable facts: which command exits 0, which file exists, which string appears, which number crosses which threshold. If a part cannot be checked by a program, say so and negotiate it down to what can.
  2. Prefer checks that already exist. The test suite, the linter, the build, the type checker. A goal test that shells out to make test is better than one that reimplements it.
  3. Write goal-test.sh (or .py if the checks are easier there) in the project root: runs every check, prints one line per check with pass/fail, exits 0 only when all pass. Keep it under ~40 lines; it must run in seconds and be safe to run repeatedly.
  4. Run it now. It should fail before the work is done — a goal test that passes on the current state is testing nothing. Show the failing output.
  5. Offer the loop. If the user wants the agent driven until done, wrap it:
bash
#!/usr/bin/env bash
# goal-loop.sh — run a headless agent until the goal test passes or attempts run out.
set -u
MAX_ATTEMPTS=5
for i in $(seq 1 "$MAX_ATTEMPTS"); do
  FEEDBACK=$(./goal-test.sh 2>&1) && { echo "done in $i attempt(s)"; exit 0; }
  claude -p "Goal: <the goal>. The goal test currently fails with:
$FEEDBACK
Fix the code so the goal test passes." \
    --permission-mode acceptEdits --output-format json > ".attempt-$i.json"
done
echo "goal test still failing after $MAX_ATTEMPTS attempts — falling back to a human"
exit 1

Adjust the agent command to whatever CLI the user runs. Keep the three brakes visible and named: the checker outside the model (goal-test.sh), the stop rule (test passes), the budget (MAX_ATTEMPTS, plus a dollar cap read from the JSON output if they want one).

Rules

  • The checker never asks the model's opinion; self-review is not a checker.
  • Feed the checker's output back verbatim — the failure text is the prompt.
  • If the user cannot state done as a check, they do not have a loopable task yet; tell them that plainly and suggest an interactive session instead.

© undefined-ui, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/agents-course/skills/goal-test of undefined-ui/second-brain-os.

Open the folder on GitHubat commit c7fa35b

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in undefined-ui/second-brain-os, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Goal Test next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Goal Test compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Goal Test this skillundefined-ui/second-brain-os1k1 repos~754Automated safety check: PassMIT
SDK AI Bot Run EvaluationAzure/azure-sdk-tools134—~1.1kAutomated safety check: NotesMIT
Qmdalsk1992/CloddsBot2.9k3 repos~1.2kAutomated safety check: PassMIT
Gitnexus CLIaws-samples/sample-kolya-br-proxy10610 repos~822Automated safety check: PassMIT-0
Ragu BuildRaguTeam/RAGU135—~3.8kAutomated safety check: NotesMIT
Analyze Id Eval Rankingopen-thoughts/OpenThoughts-Agent301—~3.1kAutomated safety check: PassApache-2.0

Similar skills

  • SDK AI Bot Run Evaluation

    Azure/azure-sdk-tools

    Official

    Run Azure SDK QA bot evaluations on curated datasets locally, including a single test case.

    134 GitHub stars~1.1k tokensUpdated today
    Knowledge ManagementAuto-check: notes
  • Qmd

    alsk1992/CloddsBot

    Local hybrid search for markdown notes and docs. An agent skill from alsk1992/CloddsBot.

    2.9k GitHub starsUsed in 3 repos~1.2k tokens
    Knowledge ManagementAuto-check passed
  • Gitnexus CLI

    aws-samples/sample-kolya-br-proxy

    Official

    A skill your agent uses when the user needs to run GitNexus CLI commands like analyze/index a repo, check status, clean the index, generate a wiki, or list indexed repos.

    106 GitHub starsUsed in 10 repos~822 tokens
    Knowledge ManagementAuto-check passed
  • Ragu Build

    RaguTeam/RAGU

    Interview the user about their RAGU use case, select an appropriate RAGU pipeline, and generate a validated ragubuild.yaml plus a runnable build<name.py script.

    135 GitHub stars~3.8k tokensUpdated 3 days ago
    Knowledge ManagementAuto-check: notes
  • Analyze Id Eval Ranking

    open-thoughts/OpenThoughts-Agent

    Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…

    301 GitHub stars~3.1k tokensUpdated 9 days ago
    Research & ScienceAuto-check passed
  • Using Mineecho Skills

    Health-Yang/MineEcho

    技能系统导航。当用户询问"你能做什么"、"有什么功能"、"有什么技能"或不确定如何完成某个任务时,使用此技能列出所有可用技能并建议最合适的技能。此技能是每个对话开始时默认加载的,用于技能发现。

    220 GitHub stars~124 tokensUpdated 4 mo ago
    Knowledge ManagementAuto-check passed

More from undefined-ui/second-brain-os

All 23 skills in this repo
  • Context Audit

    undefined-ui/second-brain-os

    Audit an agent's context layout against the four places: system prompt, tools, history, tail.

    1k GitHub stars~802 tokensUpdated today
    Auto-check passed
  • Evals Bootstrap

    undefined-ui/second-brain-os

    Scaffold a first eval suite for an agent: mine real failures into cases, write behavioural checks over traces, and generate the runner.

    1k GitHub stars~786 tokensUpdated today
    Auto-check passed
  • Gate Check

    undefined-ui/second-brain-os

    Find the decisions in a pipeline that do not need the expensive model and propose the gate for each: a rule, a classic classifier, or a small model, with fail-closed routing.

    1k GitHub stars~854 tokensUpdated today
    Auto-check passed
  • Harness Audit

    undefined-ui/second-brain-os

    Audit an agent's harness against the four rings: containment, guides, sensors, permissions.

    1k GitHub stars~855 tokensUpdated today
    Auto-check passed
  • Second Brain Chat Import

    undefined-ui/second-brain-os

    Triage an exported chat history into what to ingest, what to archive and what to delete, with a privacy pass first.

    1k GitHub stars~527 tokensUpdated today
    Auto-check passed
  • Second Brain Graph

    undefined-ui/second-brain-os

    Analyse the vault's link graph and report on its shape: orphan rate, average degree, components, hubs, bridges and clusters, with what each number means for retrieval.

    1k GitHub stars~533 tokensUpdated today
    Auto-check passed

Questions about Goal Test

What does Goal Test do?

Turn a vague task into a testable definition of done and generate an executable goal-test script for it, optionally with a bounded retry loop around a headless agent. Goal Test is an agent skill from undefined-ui/second-brain-os. Turn a vague task into a testable definition of done and generate an executable goal-test script for it, optionally with a bounded retry loop around a headless agent.

When should I use Goal Test?

Goal Test fits situations like: the user wants to run an agent in a loop; asks how to know when an agent task is finished; says a goal like improve X needs to become checkable; building eval suites over many cases — that is evals-bootstrap.

How do I install Goal Test in Claude Code?

Run `npx skills add undefined-ui/second-brain-os --skill goal-test -a claude-code`. Or copy the skill folder (plugins/agents-course/skills/goal-test in undefined-ui/second-brain-os) into .claude/skills/goal-test in your project. Claude Code loads it when a task matches its description.

How do I install Goal Test in Codex?

Run `npx skills add undefined-ui/second-brain-os --skill goal-test -a codex`. Or copy the skill folder (plugins/agents-course/skills/goal-test in undefined-ui/second-brain-os) into .agents/skills/goal-test in your project. Codex loads it when a task matches its description.

Can I use Goal Test in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add undefined-ui/second-brain-os --skill goal-test -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/goal-test, .gemini/skills/goal-test, .github/skills/goal-test and .opencode/skills/goal-test in your project.

What does Goal Test need to run?

Going by SKILL.md and its folder, Goal Test needs the command-line tools its instructions call (claude and make).

Does Goal Test access the network?

SKILL.md names 1 domain. As links in the text: secondbrainos.dev. This is read from the text; nothing was executed.

Is Goal Test safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Goal Test use?

Goal Test is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Goal Test use?

About 754 tokens (SKILL.md is roughly 3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Goal Test?

Skills that share tags, products or a category with Goal Test: SDK AI Bot Run Evaluation (Azure/azure-sdk-tools, 134 stars), Qmd (alsk1992/CloddsBot, 2.9k stars), Gitnexus CLI (aws-samples/sample-kolya-br-proxy, 106 stars) and Ragu Build (RaguTeam/RAGU, 135 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Goal Test?

undefined-ui (a GitHub user) maintains it in undefined-ui/second-brain-os, which has 1,009 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 7, 2026.

Source: undefined-ui/second-brain-os on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.