Agent skill

Heterogeneous Agent Compatibility Test

by lobehub in lobehub/lobehub

A user-invoked check that every official model the server advertises works through six external CLI agents, from Claude Code and Codex to TRAE, in a live test matrix.

Custom licenceAuto-check passedTesting & QA

Install Heterogeneous Agent Compatibility Test

skills CLI
$ npx skills add lobehub/lobehub --skill testing-heterogeneous-agents -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install lobehub/lobehub testing-heterogeneous-agents --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/lobehub/lobehub.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/testing-heterogeneous-agents .claude/skills/testing-heterogeneous-agents && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
testing-heterogeneous-agents
GitHub stars
83k
Token cost
~1.8k tokens
SKILL.md length
803 words
Files
4 (incl. scripts)
Skills in repo
50
Repo updated
First seen
Licence
Custom licence

At a glance

A user-invoked check that every official model the server advertises works through six external CLI agents, from Claude Code and Codex to TRAE, in a live test matrix.

  • Works in 7 steps: Load the parent contract and project… → Discover before spending → Establish the Electron surface → …
  • Verifying that all official models work through each external CLI agent
  • SKILL.md covers Manual Invocation Only, Scope, Harness and Run Workflow, plus 2 more sections
  • Runs JavaScript and Shell scripts from its folder; calls node and bash

What it does

This project skill extends the acceptance skill with one scenario: proving that every server-advertised official model completes through each supported external CLI agent and LobeHub's server-default provider binding. It is invoked only by the user, with /testing-heterogeneous-agents in Claude Code or $testing-heterogeneous-agents in Codex, never automatically or on a schedule, and it still needs scope and cost approval before live requests. The path tested runs from the Desktop renderer through IPC and the CLI to the official relay, using a unique marker round trip.

A harness script, scripts/official-smoke.mjs, can discover the live matrix and installed CLI versions without making model calls. Exit codes are 0 when every selected cell passed, 1 when at least one failed, 2 when none failed but one was blocked, and 3 when Electron, authentication, topic or capability preflight failed. User-defined providers, user-managed API keys and general model quality are out of scope, and acceptance planning, evidence and publishing stay with the acceptance skill.

When your agent uses it

  • Verifying that all official models work through each external CLI agent
  • Listing the live compatibility matrix and installed CLI versions before a run
  • Interpreting the harness exit codes after a failed or blocked run

Example prompts

  • “Discover the live compatibility matrix without making any model calls.”
  • “Run the official provider matrix for Codex only, after I approve the scope and cost.”
  • “The harness exited with code 2. Which cells were blocked and why?”

Requirements

  • A LobeHub Desktop (Electron) instance that the harness can attach to
  • The external CLI agents under test installed
  • Node.js to run the harness script

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Load the parent contract and project adapter
  2. Discover before spending
  3. Establish the Electron surface
  4. Confirm scope, cost, and history
  5. Execute once
  6. Inspect and publish the round
  7. Diagnose failures at the owning layer

What it can do on your machine

Read from SKILL.md and the folder at commit 6aaeeee. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (JavaScript and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • node
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Heterogeneous Agent Compatibility Test loads about 1.8k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 803 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 803 words (~1,753 tokens).

“This project skill extends acceptance with one scenario: proving that every server-advertised official model completes through each supported external CLI agent and LobeHub's server-default provider binding.”

— opening of SKILL.md by lobehub, Custom licence
name
testing-heterogeneous-agents
disable-model-invocation
true

Read the full SKILL.md on GitHub

Files

SKILL.md and 3 other files (scripts) in .agents/skills/testing-heterogeneous-agents of lobehub/lobehub.

  • SKILL.md
  • agents/openai.yaml
  • scripts/official-smoke.mjs
  • scripts/official-smoke.test.sh

Open the folder on GitHubat commit 6aaeeee

Compare with similar skills

Heterogeneous Agent Compatibility Test next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Heterogeneous Agent Compatibility Test compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Heterogeneous Agent Compatibility Test this skilllobehub/lobehub83k—~1.8kAutomated safety check: PassCustom licence
OpenHarness End-to-End EvalsHKUDS/OpenHarness16k1 repos~2.1kAutomated safety check: NotesMIT
E2E Testinglangflow-ai/langflow156k—~3.3kAutomated safety check: PassMIT
MongoDB Source Connector E2E Harnessairbytehq/airbyte22k—~1.9kAutomated safety check: PassCustom licence
Java SDK E2E Test with Replay Snapshotgithub/copilot-sdk11k—~1.8kAutomated safety check: PassMIT
Go Redis Client Test Runnerredis/go-redis22k—~786Automated safety check: PassBSD-2-Clause

Similar skills

  • Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.

    16k GitHub starsUsed in 1 repo~2.1k tokens
    Testing & QAAuto-check: notes
  • E2E Testing

    langflow-ai/langflow

    Write and review Playwright E2E tests for Langflow. An agent skill from langflow-ai/langflow.

    156k GitHub stars~3.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Official

    Starts a throwaway MongoDB 7.0 replica set and runs the Airbyte spec, check, discover and read commands against source-mongodb-v2 images for local end-to-end testing.

    22k GitHub stars~1.9k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Official

    Creates a Java SDK end-to-end test for the Copilot SDK that runs against a recorded YAML snapshot through a replay proxy, so CI needs no real authentication.

    11k GitHub stars~1.8k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Official

    Explains how to run go-redis tests: the Docker Compose stack, make targets, focusing a single Ginkgo spec, the e2e suite and the version environment variables.

    22k GitHub stars~786 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Official

    Starts a local PostgreSQL 16 container, loads SQL fixtures and runs the Airbyte spec, check, discover and read commands against a chosen source-postgres image.

    22k GitHub stars~2.5k tokensUpdated yesterday
    Testing & QAAuto-check passed

More from lobehub/lobehub

All 50 skills in this repo
  • Builds single-file interactive HTML prototypes rendered with the real LobeHub UI components and written as production-style React, so they can later be split into files.

    83k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Verifies a delivery end to end by driving the real product on a CLI, web, desktop or iOS Simulator surface, capturing evidence and publishing a round with the lh CLI.

    83k GitHub stars~9.7k tokensUpdated today
    Auto-check passed
  • Git Worktree Cleanup

    lobehub/lobehub

    Audits stale Git worktrees and branches with a bundled script, classifies each one, and deletes only after you approve the exact candidates.

    83k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Maintains LobeHub's model-backed alint rule set: writing rules, removing false positives against real code, deciding warn versus error and tracking token cost.

    83k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Guides building LobeHub builtin agent tools, from the manifest and execution runtime to executors, chat UI renders and registry wiring.

    83k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Explains how LobeHub client code fetches data through services, SWR store hooks and cache keys, and when to avoid useEffect fetching or duplicated state.

    83k GitHub stars~1.7k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Heterogeneous Agent Compatibility Test

What does Heterogeneous Agent Compatibility Test do?

A user-invoked check that every official model the server advertises works through six external CLI agents, from Claude Code and Codex to TRAE, in a live test matrix. This project skill extends the acceptance skill with one scenario: proving that every server-advertised official model completes through each supported external CLI agent and LobeHub's server-default provider binding. It is invoked only by the user, with /testing-heterogeneous-agents in Claude Code or $testing-heterogeneous-agents in Codex, never automatically or on a schedule, and it still needs scope and cost approval before live requests.

When should I use Heterogeneous Agent Compatibility Test?

Heterogeneous Agent Compatibility Test fits situations like: verifying that all official models work through each external CLI agent; listing the live compatibility matrix and installed CLI versions before a run; interpreting the harness exit codes after a failed or blocked run.

How do I install Heterogeneous Agent Compatibility Test in Claude Code?

Run `npx skills add lobehub/lobehub --skill testing-heterogeneous-agents -a claude-code`. Or copy the skill folder (.agents/skills/testing-heterogeneous-agents in lobehub/lobehub) into .claude/skills/testing-heterogeneous-agents in your project. Claude Code loads it when a task matches its description.

How do I install Heterogeneous Agent Compatibility Test in Codex?

Run `npx skills add lobehub/lobehub --skill testing-heterogeneous-agents -a codex`. Or copy the skill folder (.agents/skills/testing-heterogeneous-agents in lobehub/lobehub) into .agents/skills/testing-heterogeneous-agents in your project. Codex loads it when a task matches its description.

Can I use Heterogeneous Agent Compatibility Test in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add lobehub/lobehub --skill testing-heterogeneous-agents -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/testing-heterogeneous-agents, .gemini/skills/testing-heterogeneous-agents, .github/skills/testing-heterogeneous-agents and .opencode/skills/testing-heterogeneous-agents in your project.

What does Heterogeneous Agent Compatibility Test need to run?

Going by SKILL.md and its folder, Heterogeneous Agent Compatibility Test needs JavaScript and a shell for the scripts in its folder and the command-line tools its instructions call (node and bash). Our summary lists: A LobeHub Desktop (Electron) instance that the harness can attach to; The external CLI agents under test installed; Node.js to run the harness script.

Does Heterogeneous Agent Compatibility Test access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Heterogeneous Agent Compatibility Test safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Heterogeneous Agent Compatibility Test use?

Heterogeneous Agent Compatibility Test has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Heterogeneous Agent Compatibility Test use?

About 1.8k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Heterogeneous Agent Compatibility Test?

Skills that share tags, products or a category with Heterogeneous Agent Compatibility Test: OpenHarness End-to-End Evals (HKUDS/OpenHarness, 16k stars), E2E Testing (langflow-ai/langflow, 156k stars), MongoDB Source Connector E2E Harness (airbytehq/airbyte, 22k stars) and Java SDK E2E Test with Replay Snapshot (github/copilot-sdk, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Heterogeneous Agent Compatibility Test?

lobehub (a GitHub organization) maintains it in lobehub/lobehub, which has 83,044 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on October 8, 2026.

Source: lobehub/lobehub on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.