Google Antigravity SDK
google-antigravity/antigravity-sdk-python
Design, implement, and debug autonomous AI agents and multi-agent systems using the Google Antigravity (AGY) SDK.
Benchmark validation procedure for one assigned benchmark-related target.
$ npx skills add Intelligent-Internet/zenith --skill benchmark-validator -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Intelligent-Internet/zenith benchmark-validator --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .claude/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/benchmark-validator .claude/skills/benchmark-validator && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "benchmark-validator" agent skill from https://github.com/Intelligent-Internet/zenith/tree/main/zenith/src/zenith_harness/bundled/skills/benchmark-validator into .claude/skills/benchmark-validator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-validator", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Intelligent-Internet/zenith/tree/main/zenith/src/zenith_harness/bundled/skills/benchmark-validatorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Intelligent-Internet/zenith --skill benchmark-validator -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Intelligent-Internet/zenith benchmark-validator --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .agents/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/benchmark-validator .agents/skills/benchmark-validator && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "benchmark-validator" agent skill from https://github.com/Intelligent-Internet/zenith/tree/main/zenith/src/zenith_harness/bundled/skills/benchmark-validator into .agents/skills/benchmark-validator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-validator", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Intelligent-Internet/zenith --skill benchmark-validator -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Intelligent-Internet/zenith benchmark-validator --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/benchmark-validator .cursor/skills/benchmark-validator && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "benchmark-validator" agent skill from https://github.com/Intelligent-Internet/zenith/tree/main/zenith/src/zenith_harness/bundled/skills/benchmark-validator into .cursor/skills/benchmark-validator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-validator", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Intelligent-Internet/zenith.git --path zenith/src/zenith_harness/bundled/skills/benchmark-validator--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Intelligent-Internet/zenith --skill benchmark-validator -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Intelligent-Internet/zenith benchmark-validator --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/benchmark-validator .gemini/skills/benchmark-validator && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "benchmark-validator" agent skill from https://github.com/Intelligent-Internet/zenith/tree/main/zenith/src/zenith_harness/bundled/skills/benchmark-validator into .gemini/skills/benchmark-validator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-validator", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Intelligent-Internet/zenith benchmark-validatorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Intelligent-Internet/zenith --skill benchmark-validator -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .github/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/benchmark-validator .github/skills/benchmark-validator && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "benchmark-validator" agent skill from https://github.com/Intelligent-Internet/zenith/tree/main/zenith/src/zenith_harness/bundled/skills/benchmark-validator into .github/skills/benchmark-validator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-validator", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Intelligent-Internet/zenith --skill benchmark-validator -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Intelligent-Internet/zenith benchmark-validator --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/benchmark-validator .opencode/skills/benchmark-validator && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "benchmark-validator" agent skill from https://github.com/Intelligent-Internet/zenith/tree/main/zenith/src/zenith_harness/bundled/skills/benchmark-validator into .opencode/skills/benchmark-validator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-validator", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
benchmark-validatorBenchmark validation procedure for one assigned benchmark-related target.
Benchmark Validator is an agent skill from Intelligent-Internet/zenith. Benchmark validation procedure for one assigned benchmark-related target. For optimization EXP- targets, independently classify candidate outcome. For engineering VAL- or legacy engineering targets, prove or disprove the required benchmark/performance assertion.
Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Agent Workflows. It works with Model Context Protocol. The repository describes itself as: Zenith: a continuous-improvement harness for long-running agent tasks. Turns Claude Code, Codex, or Hermes into a multi-agent mission orchestrator via MCP/ACP. The licence is Apache-2.0.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit a8d9b57. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Benchmark Validator loads about 1.8k tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 720 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Intelligent-Internet/zenith at commit a8d9b57, republished under its Apache-2.0 licence (© Intelligent-Internet). 720 words, ~1,788 tokens.
.claude/skills/benchmark-validator/SKILL.md (or your agent's skills folder).Use this skill for a validation assignment targeting exactly one benchmark-related assertion.
For optimization EXP-* targets, you are not selecting the winner or optimizing
further. You independently decide whether the experiment produced a trustworthy
outcome under its contract.
For optimization EXP-* targets, act as an adversarial tester for the selected
candidate. The contract is the minimum bar, not the whole test plan. Before
promotion, define and run a compact but comprehensive validation plan that
attacks the candidate's likely correctness and performance failure modes.
For engineering VAL-* or legacy engineering targets, do not use optimization
outcome semantics. Pass only when the assigned benchmark, performance,
correctness, and guardrail behavior required by the assignment and contract is
proven with fresh evidence.
Read:
AGENTS.md.If the assignment targets more than one benchmark-related assertion, fail the assignment as too broad and request attention.
Identify artifacts
Audit evidence integrity
invalid; for an engineering target, mark the item
passed=false.Define adversarial validation plan
EXP-*, do this before remeasurement. Do not only run the
contract's benchmark command.invalid rather than promoting from aggregate
speed alone.Remeasure independently
Run correctness and guardrails
EXP-*, run the adversarial validation plan as well as
the contract-required checks.rejected, not promoted.invalid.Classify outcome
EXP-*
target or an engineering/legacy target.promoted: credible integrity, setup complete, correctness/guardrails pass, metric meets promotion rule.rejected: credible integrity and setup, but correctness/guardrails fail or metric does not win.budget_exhausted: declared budget prevented completion and the contract allows that honest outcome.invalid: missing ledger/ref/checks, broken candidate binding, compromised evidence, or unverifiable setup.For optimization EXP-* targets, use passed=true for promoted, rejected,
or legitimate budget_exhausted; use passed=false for invalid.
For engineering VAL-* or legacy engineering targets, use normal validation
semantics: passed=true only when the assigned benchmark/performance assertion
and required guardrails pass with fresh evidence. A rejected candidate, failed
guardrail, missing evidence, or budget-exhausted result is passed=false unless
the contract explicitly defines that outcome as the required behavior.
passed=false, write <regressions_dir>/<item_id>.md with setup,
command, expected credible measurement or required behavior, observed
invalidating issue, and evidence artifact paths.## Experiment
- Item ID: <id>
- Parent artifact: <ref/path>
- Candidate artifact: <ref/path>
- Ledger: <path>
## Evidence Integrity
- Changed protected/evidence-sensitive paths: <list>
- Verdict: <credible | invalid>
## Measurements
- Parent: <runs, aggregate, spread, raw paths>
- Candidate: <runs, aggregate, spread, raw paths>
- Improvement: <calculation and threshold>
## Adversarial Validation Plan
- Risk axes: <candidate-specific axes attacked>
- Cases: <case -> expected evidence -> failure boundary>
- Limitations: <untested risk axes and why>
## Correctness And Guardrails
- `<command>` -> exit <code>, <result>
- Per-case adversarial results: <case -> passed/rejected/invalid evidence>
## Outcome
- <promoted | rejected | budget_exhausted | invalid>
- Reason: <contract-tied explanation>
## Limitations
- <setup or measurement caveats>Call end_node with exactly one item for the assigned target id, then exit immediately.
© Intelligent-Internet, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in zenith/src/zenith_harness/bundled/skills/benchmark-validator of Intelligent-Internet/zenith.
Open the folder on GitHubat commit a8d9b57
Benchmark Validator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Benchmark Validator this skillIntelligent-Internet/zenith | 334 | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| Google Antigravity SDKgoogle-antigravity/antigravity-sdk-python | 3.6k | — | ~2.1k | Automated safety check: Notes | Apache-2.0 | |
| Chatgpt AppsHaohao-end/openagent | 807 | 1 repos | ~4.9k | Automated safety check: Pass | Apache-2.0 | |
| Autocontext for Hermesgreyhaven-ai/autocontext | 1.3k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Workflow Schema Tuningbreaking-brake/cc-wf-studio | 5.4k | — | ~1.3k | Automated safety check: Pass | Custom licence | |
| Documentation Serverandrea9293/mcp-documentation-server | 343 | — | ~2.3k | Automated safety check: Pass | MIT |
google-antigravity/antigravity-sdk-python
Design, implement, and debug autonomous AI agents and multi-agent systems using the Google Antigravity (AGY) SDK.
Haohao-end/openagent
Build, scaffold, refactor, and troubleshoot ChatGPT Apps SDK applications that combine an MCP server and widget UI.
greyhaven-ai/autocontext
Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI.
breaking-brake/cc-wf-studio
Guides edits to cc-wf-studio's workflow schema so AI agents generate better workflows, treating schema text as prompt engineering rather than validation.
andrea9293/mcp-documentation-server
A skill your agent uses when you need to store, retrieve, search, or manage documents in a local knowledge base with semantic search and hybrid (vector + full-text) retrieval.
yoloshii/ClawMem
ClawMem operational reference for agents at query time — the 3-rule escalation gate, MCP tool routing, the 4 query-optimization levers, pipeline behavior (query vs intentsearch), composite scoring…
Intelligent-Internet/zenith
A skill your agent uses when planning or replanning engineering missions that create, change, port, migrate, integrate, or preserve durable codebase behavior across UI, API, CLI, background jobs…
Intelligent-Internet/zenith
Adversarial scrutiny procedure for engineering validation assignments.
Intelligent-Internet/zenith
Real-surface validation coordinator for engineering validation assignments.
Intelligent-Internet/zenith
Automates browser and Electron app interactions for user-flow validation.
Intelligent-Internet/zenith
Domain playbook for optimization missions — any task whose goal is to move a metric: performance, latency, throughput, memory, cost, score, quality, compression, ranking, solver, model/eval, and…
Works with
Categories
Benchmark validation procedure for one assigned benchmark-related target. Benchmark Validator is an agent skill from Intelligent-Internet/zenith. Benchmark validation procedure for one assigned benchmark-related target.
Benchmark Validator fits situations like: agent Workflows work in your project.
Run `npx skills add Intelligent-Internet/zenith --skill benchmark-validator -a claude-code`. Or copy the skill folder (zenith/src/zenith_harness/bundled/skills/benchmark-validator in Intelligent-Internet/zenith) into .claude/skills/benchmark-validator in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Intelligent-Internet/zenith --skill benchmark-validator -a codex`. Or copy the skill folder (zenith/src/zenith_harness/bundled/skills/benchmark-validator in Intelligent-Internet/zenith) into .agents/skills/benchmark-validator in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Intelligent-Internet/zenith --skill benchmark-validator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-validator, .gemini/skills/benchmark-validator, .github/skills/benchmark-validator and .opencode/skills/benchmark-validator in your project.
SKILL.md names no scripts, command-line tools or credentials: Benchmark Validator is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Benchmark Validator is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Benchmark Validator: Google Antigravity SDK (google-antigravity/antigravity-sdk-python, 3.6k stars), Chatgpt Apps (Haohao-end/openagent, 807 stars), Autocontext for Hermes (greyhaven-ai/autocontext, 1.3k stars) and Workflow Schema Tuning (breaking-brake/cc-wf-studio, 5.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Intelligent-Internet (a GitHub organization) maintains it in Intelligent-Internet/zenith, which has 334 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on September 6, 2026.
Source: Intelligent-Internet/zenith on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.