Agent skill

Terminal Agent Benchmark

by nowledge-co in nowledge-co/con-terminal

Run and maintain Con's terminal-agent benchmark against a live app session.

MITAuto-check passedAgent Workflows

Install Terminal Agent Benchmark

skills CLI
$ npx skills add nowledge-co/con-terminal --skill terminal-agent-benchmark -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install nowledge-co/con-terminal terminal-agent-benchmark --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/nowledge-co/con-terminal.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/terminal-agent-benchmark .claude/skills/terminal-agent-benchmark && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
terminal-agent-benchmark
GitHub stars
625
Token cost
~822 tokens
SKILL.md length
313 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Run and maintain Con's terminal-agent benchmark against a live app session.

  • Works in 9 steps: Confirm a live app session and socket… → Run the strict benchmark → If provider setup is present, run → …
  • Validating con-cli
  • SKILL.md covers Default workflow and Rules
  • Calls python3

What it does

Terminal Agent Benchmark is an agent skill from nowledge-co/con-terminal. Run and maintain Con's terminal-agent benchmark against a live app session. Use when validating con-cli, SSH workspace reuse, tmux awareness, agent-target preparation, or when collecting benchmark evidence for regressions and release notes.

Its SKILL.md is about 820 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Agent evaluation and testing and Changelog and release notes. It works with tmux and Python. The repository describes itself as: The Native Terminal Emulator with a builtin AI Harness. The licence is MIT.

When your agent uses it

  • Validating con-cli
  • SSH workspace reuse
  • Agent-target preparation
  • Collecting benchmark evidence for regressions and release notes

Example prompts

  • “/terminal-agent-benchmark”

Requirements

  • Python 3

Workflow steps

9 steps, taken from the first numbered list in SKILL.md.

  1. Confirm a live app session and socket exist.
  2. Run the strict benchmark
  3. If provider setup is present, run
  4. Prefer a built-in profile when one matches the workflow
  5. Use starter profiles for quick regression checks and operator profiles for richer coding, SSH maintenance, or tmux dev-loop evaluation.
  6. For SSH/tmux changes, run the relevant playbook under benchmarks/terminal-agent/playbooks/.
  7. Save the JSON record under .context/benchmarks/ and cite it in your report.
  8. If the run is an operator benchmark, score it with
  9. Generate a trend report when comparing many runs

What it can do on your machine

Read from SKILL.md and the folder at commit a371ecb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Terminal Agent Benchmark loads about 822 tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 313 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~822

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from nowledge-co/con-terminal at commit a371ecb, republished under its MIT licence (© nowledge-co). 313 words, ~822 tokens.

Download SKILL.mdSave it as .claude/skills/terminal-agent-benchmark/SKILL.md (or your agent's skills folder).
name
terminal-agent-benchmark
description
Run and maintain Con's terminal-agent benchmark against a live app session. Use when validating con-cli, SSH workspace reuse, tmux awareness, agent-target preparation, or when collecting benchmark evidence for regressions and release notes.

Terminal Agent Benchmark

Use this skill when you need to evaluate Con as a terminal-native agent, not just compile it.

Primary references:

Default workflow

  1. Confirm a live app session and socket exist.
  2. Run the strict benchmark:
    • python3 benchmarks/terminal-agent/run.py --suite strict
  3. If provider setup is present, run:
    • CON_BENCH_ENABLE_AGENT=1 python3 benchmarks/terminal-agent/run.py --suite all
  4. Prefer a built-in profile when one matches the workflow:
    • python3 benchmarks/terminal-agent/run.py --list-profiles
    • python3 benchmarks/terminal-agent/run.py --profile basic-local-shell
    • CON_BENCH_ENABLE_AGENT=1 python3 benchmarks/terminal-agent/run.py --profile basic-local-codex --suite all
  5. Use starter profiles for quick regression checks and operator profiles for richer coding, SSH maintenance, or tmux dev-loop evaluation.
    • python3 benchmarks/terminal-agent/run.py --profile operator-local-codex-devloop --suite operator
    • python3 benchmarks/terminal-agent/run.py --profile operator-local-claude-devloop --suite operator
    • python3 benchmarks/terminal-agent/run.py --profile operator-local-opencode-devloop --suite operator
    • python3 benchmarks/terminal-agent/run.py --profile operator-ssh-dual-host-maintenance --suite operator
    • python3 benchmarks/terminal-agent/run.py --profile operator-ssh-tmux-devloop --suite operator
  6. For SSH/tmux changes, run the relevant playbook under benchmarks/terminal-agent/playbooks/.
  7. Save the JSON record under .context/benchmarks/ and cite it in your report.
  8. If the run is an operator benchmark, score it with:
    • python3 benchmarks/terminal-agent/score.py --profile ... --record ... --score ...
  9. Generate a trend report when comparing many runs:
    • python3 benchmarks/terminal-agent/report.py

Rules

  • Prefer pane_id over pane_index when following up on benchmark findings.
  • Keep benchmark control on the existing panes.* API unless the benchmark explicitly targets pane-local surfaces. Surface support is additive and should not change the built-in agent's pane/tool contract.
  • Do not treat playbook observations as strict pass/fail evidence unless the behavior is actually deterministic.
  • If a scenario depends on host setup, say so explicitly.
  • Keep operator playbooks safe-by-default. Prefer read-only checks first, and treat destructive or privileged steps as explicit branches.
  • If a benchmark reveals a product limit, document the limit instead of hiding it behind a softer assertion.
  • Operator suites intentionally serialize agent ask turns. If a tab already has a pending agent request, let the runner wait and reuse the same tab instead of opening a parallel benchmark against it.

© nowledge-co, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/terminal-agent-benchmark of nowledge-co/con-terminal.

Open the folder on GitHubat commit a371ecb

Compare with similar skills

Terminal Agent Benchmark next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Terminal Agent Benchmark compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Terminal Agent Benchmark this skillnowledge-co/con-terminal625—~822Automated safety check: PassMIT
Skills Constitutionjiabaobei/skills-constitution232—~5.6kAutomated safety check: PassMIT
MCP Server Builderanthropics/skills180k63 repos~2.3kAutomated safety check: PassApache-2.0
CodeGraph Agent Evalcolbymchenry/codegraph74k—~950Automated safety check: PassMIT
Agent Wiki Compare OutcomesAgentToolkit/altk-evolve124—~1.7kAutomated safety check: PassApache-2.0
TelegramAIOSAI/AIPass288—~4.9kAutomated safety check: PassMIT

Similar skills

  • Skills Constitution

    jiabaobei/skills-constitution

    当 Agent 接到专业任务(编码/爬虫/文件操作/API调用/数据分析/文档/部署/推送等)时,强制先查记忆层和技能索引,有匹配必用、无匹配必搜、答复时自动推荐(排除已装)。用于防止 Agent 跳过技能直接硬扛通用能力。跨平台通用(WorkBuddy/Claude/ChatGPT/Cursor/Gemini 等 20+ 框架)。完整版本史见 CHANGELOG.md。

    232 GitHub stars~5.6k tokensUpdated 6 days ago
    Agent WorkflowsAuto-check passed
  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 63 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • CodeGraph Agent Eval

    colbymchenry/codegraph

    Benchmarks how much CodeGraph helps a coding agent on a real repository, comparing runs with and without it for a chosen local or published version.

    74k GitHub stars~950 tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Agent Wiki Compare Outcomes

    AgentToolkit/altk-evolve

    Contrasts successful and failed agent runs of the same task and derives guidelines backed by evidence from transcripts, tool calls and outcome judgments.

    124 GitHub stars~1.7k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Telegram

    AIOSAI/AIPass

    Multi-bot Telegram bridge — routes messages between Telegram and Claude tmux sessions

    288 GitHub stars~4.9k tokensUpdated 4 days ago
    Agent WorkflowsAuto-check passed
  • Os Release

    CronusL-1141/AI-company

    发布 AI Team OS 新版本的完整清单——预检、版本七处锁步、中英双语 CHANGELOG、双份 dist 构建、私有术语扫描、commit/tag、双仓推送、建 GitHub Release 条目并核对 latest 徽章、事后核对。当准备发版、补建漏掉的 Release 条目、或核对已发版本的线上状态时使用。

    371 GitHub stars~2.3k tokensUpdated 9 days ago
    DevelopmentAuto-check passed

More from nowledge-co/con-terminal

  • Con CLI E2E

    nowledge-co/con-terminal

    Validate Con's local socket control plane against a real running app session, and write/run con-test integration tests.

    625 GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Changelog Release Notes

    nowledge-co/con-terminal

    Maintain Con's CHANGELOG.md and release notes. An agent skill from nowledge-co/con-terminal.

    625 GitHub stars~686 tokensUpdated today
    Auto-check passed
  • Release

    nowledge-co/con-terminal

    Run and verify Con beta, dev, and stable releases without publishing incomplete artifacts.

    625 GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Terminal Agent Improvement Loop

    nowledge-co/con-terminal

    Run a benchmark-driven improvement loop for Con's terminal agent.

    625 GitHub stars~681 tokensUpdated today
    Auto-check passed

Works with

Questions about Terminal Agent Benchmark

What does Terminal Agent Benchmark do?

Run and maintain Con's terminal-agent benchmark against a live app session. Terminal Agent Benchmark is an agent skill from nowledge-co/con-terminal. Run and maintain Con's terminal-agent benchmark against a live app session.

When should I use Terminal Agent Benchmark?

Terminal Agent Benchmark fits situations like: validating con-cli; SSH workspace reuse; agent-target preparation; collecting benchmark evidence for regressions and release notes.

How do I install Terminal Agent Benchmark in Claude Code?

Run `npx skills add nowledge-co/con-terminal --skill terminal-agent-benchmark -a claude-code`. Or copy the skill folder (skills/terminal-agent-benchmark in nowledge-co/con-terminal) into .claude/skills/terminal-agent-benchmark in your project. Claude Code loads it when a task matches its description.

How do I install Terminal Agent Benchmark in Codex?

Run `npx skills add nowledge-co/con-terminal --skill terminal-agent-benchmark -a codex`. Or copy the skill folder (skills/terminal-agent-benchmark in nowledge-co/con-terminal) into .agents/skills/terminal-agent-benchmark in your project. Codex loads it when a task matches its description.

Can I use Terminal Agent Benchmark in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add nowledge-co/con-terminal --skill terminal-agent-benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/terminal-agent-benchmark, .gemini/skills/terminal-agent-benchmark, .github/skills/terminal-agent-benchmark and .opencode/skills/terminal-agent-benchmark in your project.

What does Terminal Agent Benchmark need to run?

Going by SKILL.md and its folder, Terminal Agent Benchmark needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Terminal Agent Benchmark access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Terminal Agent Benchmark safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Terminal Agent Benchmark use?

Terminal Agent Benchmark is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Terminal Agent Benchmark use?

About 822 tokens (SKILL.md is roughly 3.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Terminal Agent Benchmark?

Skills that share tags, products or a category with Terminal Agent Benchmark: Skills Constitution (jiabaobei/skills-constitution, 232 stars), MCP Server Builder (anthropics/skills, 180k stars), CodeGraph Agent Eval (colbymchenry/codegraph, 74k stars) and Agent Wiki Compare Outcomes (AgentToolkit/altk-evolve, 124 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Terminal Agent Benchmark?

nowledge-co (a GitHub organization) maintains it in nowledge-co/con-terminal, which has 625 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 9, 2026.

Source: nowledge-co/con-terminal on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.