Agent skill

Copilot SDK E2E Dev

by omnigent-ai in omnigent-ai/omnigent

Spin up a live local Omnigent server and exercise the GitHub Copilot SDK harness end-to-end — build copilot agents, run real turns, smoke-test, and bug-bash.

Apache-2.0Auto-check passedTesting & QA

Install Copilot SDK E2E Dev

skills CLI
$ npx skills add omnigent-ai/omnigent --skill copilot-sdk-e2e-dev -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install omnigent-ai/omnigent copilot-sdk-e2e-dev --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/omnigent-ai/omnigent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/copilot-sdk-e2e-dev .claude/skills/copilot-sdk-e2e-dev && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
copilot-sdk-e2e-dev
GitHub stars
11k
Token cost
~2.7k tokens
SKILL.md length
1,149 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

Spin up a live local Omnigent server and exercise the GitHub Copilot SDK harness end-to-end — build copilot agents, run real turns, smoke-test, and bug-bash.

  • Works in 3 steps: start a local server → build a copilot agent bundle → run a turn (and smoke-test)
  • Tasks that involve End-to-end testing
  • SKILL.md covers Prerequisites (check these…, Step 1 — start a local server, Step 2 — build a copilot agent… and Step 3 — run a turn (and…, plus 7 more sections
  • Calls python, uv and gh; needs GH_TOKEN and COPILOT_GITHUB_TOKEN

What it does

Copilot SDK E2E Dev is an agent skill from omnigent-ai/omnigent. Spin up a live local Omnigent server and exercise the GitHub Copilot SDK harness end-to-end — build copilot agents, run real turns, smoke-test, and bug-bash. Load when developing, testing, or debugging the copilot harness (omnigent/inner/copilotexecutor.py, copilotharness.py, omnigent/onboarding/copilotauth.py) or its auth / model / tool-bridge behavior.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering End-to-end testing and QA and bug reports. It works with Bash, GitHub and Python. The repository describes itself as: Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve End-to-end testing
  • Tasks that involve QA and bug reports

Example prompts

  • “/copilot-sdk-e2e-dev”

Requirements

  • Python 3
  • A credential in COPILOT_GITHUB_TOKEN
  • A credential in GITHUB_TOKEN

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. start a local server
  2. build a copilot agent bundle
  3. run a turn (and smoke-test)

What it can do on your machine

Read from SKILL.md and the folder at commit fa1dbe6. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • uv
    • gh
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, gh and curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GH_TOKEN
    • COPILOT_GITHUB_TOKEN
    • GITHUB_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Copilot SDK E2E Dev loads about 2.7k tokens when it runs. Until then it costs about 95 tokens; SKILL.md has 1,149 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~95
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from omnigent-ai/omnigent at commit fa1dbe6, republished under its Apache-2.0 licence (© omnigent-ai). 1,149 words, ~2,738 tokens.

Download SKILL.mdSave it as .claude/skills/copilot-sdk-e2e-dev/SKILL.md (or your agent's skills folder).
name
copilot-sdk-e2e-dev
description
Spin up a live local Omnigent server and exercise the GitHub Copilot SDK harness end-to-end — build copilot agents, run real turns, smoke-test, and bug-bash. Load when developing, testing, or debugging the copilot harness (omnigent/inner/copilot_executor.py, copilot_harness.py, omnigent/onboarding/copilot_auth.py) or its auth / model / tool-bridge behavior.

Copilot SDK harness: end-to-end dev & testing

The copilot harness drives the GitHub Copilot SDK (github-copilot-sdk, imported as copilot) — a persistent CopilotClient + CopilotSession per Omnigent conversation — and bridges Omnigent's sys_* tools into Copilot as SDK Tools. The Python SDK bundles the Copilot CLI binary it drives as a backing server, so there is no separate @github/copilot install. This skill is the proven recipe for running it for real against a live local server — not just the unit tests.

The harness runs as a local runner from your current checkout, so omni run <bundle> --server <url> exercises exactly the code you're on.

Prerequisites (check these first)

  1. You're on the branch you want to test. The copilot harness is an optional extra — install it (without disturbing other extras) with uv sync --frozen --group test --extra copilot. NB: a bare uv run --frozen --group test re-syncs the venv and prunes the copilot SDK; for live testing call .venv/bin/omni / .venv/bin/python directly and avoid uv run mid-session.
  2. The SDK is installed: .venv/bin/python -c "import copilot; print(copilot.__file__)".
  3. A GitHub token with Copilot access is configured. Copilot needs a fine-grained PAT with the "Copilot Requests" permission, or an OAuth token from the GitHub CLI / Copilot CLI app (classic ghp_ PATs are rejected). Verify (booleans only — never print the token):
    bash
    .venv/bin/python -c "from omnigent.onboarding.copilot_auth import copilot_github_token_configured; import os; print('config:', copilot_github_token_configured(), 'env:', bool(os.environ.get('GH_TOKEN') or os.environ.get('COPILOT_GITHUB_TOKEN')))"
    If both are False, run omni setup and register a Copilot token, or export GH_TOKEN=$(gh auth token) (when gh is logged into an account with Copilot). Check the account's entitlement with gh api /copilot_internal/user (look for chat_enabled/cli_enabled).
  4. Network egress to GitHub's Copilot backend. A turn that hangs or fails to connect on a locked-down host is usually egress, not a harness bug.

Step 1 — start a local server

bash
cd /path/to/omnigent
.venv/bin/omni server --port 7788 --no-open    # foreground; or `omni server --background` for detached
curl -s http://127.0.0.1:7788/health           # {"status":"ok"}

Use the URL below as $SERVER.

Step 2 — build a copilot agent bundle

A spec with spec_version must be a directory containing config.yaml — not a single .yaml file. Minimal copilot agent:

bash
mkdir -p /tmp/copilot-dev
cat > /tmp/copilot-dev/config.yaml <<'YAML'
spec_version: 1
name: copilot-dev
description: Copilot SDK dev/test agent.
executor:
  type: omnigent
  config:
    harness: copilot
    # model: gpt-5-mini      # optional; omit for Copilot auto-select
prompt: |
  You are a terse test agent. Answer in as few words as possible.
YAML

For sub-agents, tools, guardrails/policies, copy the field shapes from examples/polly/config.yaml and examples/debby/config.yaml. (Declare policies under guardrails.policies: — a top-level policies: key is silently dropped on the spec_version + config.yaml path.)

Step 3 — run a turn (and smoke-test)

bash
SERVER=http://127.0.0.1:7788
timeout 280 .venv/bin/omni run /tmp/copilot-dev \
  -p "Reply with exactly the single word: PONG" \
  --server "$SERVER" 2>&1

A healthy run prints connection lines then the reply (PONG). If that works, the full stack is good: token, egress, bundled CLI, harness.

  • Shell / file tools: add --tools coding.
  • Specific model: add --model gpt-5-mini (or claude-haiku-4.5, auto).

Targeted scenarios

GoalHow
Native tools (shell/edit/read)--tools coding, prompt to create→read→edit a file; confirm it actually touches disk
Bridged sys_* / sub-agent dispatchdeclare a sub-agent (harness copilot so auth is satisfied), prompt the parent to delegate — exercises the SDK Tool async-handler bridge into _tool_executor
Model routingrun the same bundle with several --model values; an unknown id fails loud, a databricks-* id is dropped to auto with a warning
LLM-phase policyadd a guardrail that denies a keyword; confirm PHASE_LLM_REQUEST/PHASE_LLM_RESPONSE blocks it
Concurrency / leaksfire several omni run … & at once; then pgrep -af "copilot/bin/copilot" to check for orphaned bundled-CLI subprocesses

Running polly (or any orchestrator) on a copilot brain

The copilot harness can serve as an async orchestrator brain (polly / debby), not just a standalone agent — it dispatches to sub-agents via the bridged sys_* tools and synthesizes their results. Two ways to exercise it:

1. Committed regression guard (brain smoke). tests/e2e/test_polly_copilot_e2e.py boots a local server from your checkout and runs examples/polly with --harness copilot --model auto, asserting the brain boots and replies. It is skipped unless a Copilot token is configured (so CI without one skips it). Run it with:

bash
.venv/bin/python -m pytest -o addopts="" tests/e2e/test_polly_copilot_e2e.py -v

2. Full orchestration (dispatch → collect → synthesize). Use the polly-e2e-dev driver (in the internal agent-framework clone) — it boots a local server, polls the AP API, auto-answers elicitations, and asserts the fan-out. Drive the brain on copilot with --brain-harness copilot, and always pass a Copilot-catalog --brain-model (auto, claude-haiku-4.5, gpt-5-mini): the driver's default --brain-model is a Claude id that Copilot (no Databricks gateway) can't route. From the agent-framework clone:

bash
.venv/bin/python .claude/skills/polly-e2e-dev/polly_driver.py \
  --local --code-dir <this-worktree> \
  --cuj smoke --brain-harness copilot --brain-model auto      # brain only
# --cuj fanout  …  and  --cuj review-pr --repo omnigent-ai/omnigent --pr <n>  …
#   exercise real sub-agent dispatch (claude_code + codex) under a copilot brain.

All three CUJs (smoke / fanout / review-pr) pass on a copilot brain (verified live: fanout dispatched 8 sub-agents, 8/8 OK + a synthesis). Note omni run -p exits after the dispatch turn (the brain parks until woken), so a sub-agent's final answer lands server-side — read it over the AP API (GET /v1/sessions/{id}/items, child sessions), not just stdout.

Show full SKILL.md (451 more words)Show less

Gotchas (these cost real time)

  1. config.yaml's server: defaults to a remote server. Omitting --server sends your turn to that remote deploy — which may be stale and reject the copilot harness with executor.config.harness: must be one of […]. Always pass --server http://127.0.0.1:<port>. (If a local server rejects copilot, it's running stale code — restart it from your checkout.)
  2. A spec with spec_version must be a directory + config.yaml, never a single .yaml file.
  3. Copilot needs a GitHub token (fine-grained PAT w/ Copilot Requests, or a gh/Copilot-CLI OAuth token). Resolution precedence: spec executor.auth (api_key) > stored copilot: config block (omni setup) > ambient COPILOT_GITHUB_TOKEN / GH_TOKEN / GITHUB_TOKEN. Classic ghp_ rejected.
  4. No Databricks gateway. Copilot talks only to GitHub's backend, so a databricks-* model is silently resolved to Copilot's auto-select — it will not route through the AI Gateway like claude-sdk/codex/pi.
  5. Use a model id from the account's catalog. free_limited offers auto, claude-haiku-4.5, gpt-5-mini. Run .venv/bin/python + client.list_models() to discover the live set; an unknown id fails loud (server-side failed session).
  6. Turns take 30–90s — always wrap in timeout 280.
  7. Never print/echo the GitHub token in logs or commands.

Code & tests

  • Executor (SDK bridge): omnigent/inner/copilot_executor.py
  • Wrap (HARNESS_COPILOT_ env → executor):* omnigent/inner/copilot_harness.py
  • Auth / token resolution: omnigent/onboarding/copilot_auth.py
  • Spawn env: _build_copilot_spawn_env in omnigent/runtime/workflow.py
bash
uv run --frozen --group test python -m pytest \
  tests/inner/test_copilot_executor.py \
  tests/inner/test_copilot_harness.py \
  tests/runtime/test_copilot_spawn_env.py \
  tests/onboarding/test_copilot_auth.py -q

Bug-bash (fan out)

To stress the harness, run several scenario probes in parallel — each builds a bundle and runs real turns against the same $SERVER, then reports what broke. Highest-value targets: the Tool async-handler bridge (hangs / lost tool results / errors reported as success), model routing, policy enforcement, streamed-output rendering, and orphaned bundled-CLI processes after teardown. Cross-check the AP API (GET /v1/sessions/{id}/items) — a start failure can exit 0 with empty stdout while the server records a failed session.

Known sharp edges (found via live bug-bash — "as of this writing")

  • Native tools bypass on:[tool_call] policies and aren't recorded. Copilot's built-in create/view/edit/bash run inside the SDK, so an on:[tool_call] DENY guardrail (e.g. blast_radius) never sees them, and they leave no function_call item in the transcript (only streamed narration). Bridged sys_* tools ARE gated and recorded. Gate Copilot's built-ins at the LLM phase (PHASE_LLM_REQUEST/RESPONSE, which fire) or via the OS-env sandbox — not on:[tool_call]. (Same shape as the cursor harness.)
  • Copilot fails loud (unlike cursor's swallowed start failures). Bad token, empty/invalid model, and unknown model ids all exit non-zero with a clear error AND a server-side failed session + error item — verified, not swallowed.
  • omni run -p against an async orchestrator exits after the dispatch turn, so a delegated sub-agent's final answer is persisted server-side but may not reach stdout in one-shot mode. Read the session over the AP API to see it.
  • Non-graceful exit can orphan the bundled CLI. Graceful teardown reaps it (client.stop()); after a SIGKILL/hard-exit, sweep pgrep -af "copilot/bin/copilot".

Cleanup

bash
.venv/bin/omni server stop        # or kill the foreground `omni server`
rm -rf /tmp/copilot-dev           # remove scratch bundles
pgrep -af "copilot/bin/copilot"   # confirm no orphaned bundled-CLI subprocesses linger

© omnigent-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/copilot-sdk-e2e-dev of omnigent-ai/omnigent.

Open the folder on GitHubat commit fa1dbe6

Compare with similar skills

Copilot SDK E2E Dev next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Copilot SDK E2E Dev compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Copilot SDK E2E Dev this skillomnigent-ai/omnigent11k—~2.7kAutomated safety check: PassApache-2.0
E2Ewp-media/wp-rocket767—~1.2kAutomated safety check: PassGPL-2.0
Specx Testsmaksimzayats/specx202—~1.9kAutomated safety check: PassMIT
OpenROAD Issue TriageThe-OpenROAD-Project/OpenROAD3.2k—~842Automated safety check: PassBSD-3-Clause
Beava PR Reviewbeava-dev/beava138—~2.1kAutomated safety check: PassApache-2.0
Oss Bounty Findertinyfish-io/tinyfish-cookbook2.2k—~4.6kAutomated safety check: PassMIT

Similar skills

  • E2E

    wp-media/wp-rocket

    Run a basic E2E behavioral probe — one primary scenario smoke test for the grooming step.

    767 GitHub stars~1.2k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Specx Tests

    maksimzayats/specx

    Add or refine tests for specx Python services. An agent skill from maksimzayats/specx.

    202 GitHub stars~1.9k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • OpenROAD Issue Triage

    The-OpenROAD-Project/OpenROAD

    Reproduces an OpenROAD GitHub bug from an attached tarball and shrinks the failing design with whittle.py so maintainers get a minimal test case.

    3.2k GitHub stars~842 tokensUpdated today
    DevelopmentAuto-check passed
  • Beava PR Review

    beava-dev/beava

    Reviews a beava PR diff for real bugs, beava-specific architectural invariants, and AI-generated "slop" patterns (hollow code, phantom imports, inflated comments, disconnected pipelines).

    138 GitHub stars~2.1k tokensUpdated 4 mo ago
    DevelopmentAuto-check passed
  • Oss Bounty Finder

    tinyfish-io/tinyfish-cookbook

    Find paid open-source work, OSS bounties, open source grants, or ways to get paid contributing to open source.

    2.2k GitHub stars~4.6k tokensUpdated 5 days ago
    Data & AnalyticsAuto-check passed
  • Web Application Testing

    anthropics/skills

    Official

    Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

    180k GitHub starsUsed in 51 repos~966 tokens
    Testing & QAAuto-check passed

More from omnigent-ai/omnigent

All 19 skills in this repo
  • Omnigent Docker Compose Deploy

    omnigent-ai/omnigent

    Brings up the Omnigent server and Postgres as a Docker compose stack on any Docker host, and covers the Dockerfile's runtime and host build targets for extending it to a new platform.

    11k GitHub stars~1.3k tokensUpdated today
    Auto-check: notes
  • Omnigent Framework Detection

    omnigent-ai/omnigent

    Scans Python agent code for framework imports and recommends the matching Omnigent executor type, or says when the framework is not natively supported yet.

    11k GitHub stars~610 tokensUpdated today
    Auto-check passed
  • Omnigent Load Test Runner

    omnigent-ai/omnigent

    Runs the Omnigent load test with real hosts and multi-turn sessions against a mocked LLM, then explains the latency results from summary.md.

    11k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Verify Omnigent End-to-End

    omnigent-ai/omnigent

    Spins up an isolated Omnigent server, runner and mock model to prove a user-facing behavior or bug fix with recorded evidence instead of reasoning from code.

    11k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Spins up a local Omnigent server and exercises the Antigravity (Gemini) SDK harness end to end: building agents, running real turns, smoke tests and bug-bashing.

    11k GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Omnigent Agent Builder

    omnigent-ai/omnigent

    Gives patterns for generating a minimal, valid Omnigent agent directory: the config.yaml fields, the right executor type, and the files each agent needs.

    11k GitHub stars~2.1k tokensUpdated today
    Auto-check passed

Categories

Questions about Copilot SDK E2E Dev

What does Copilot SDK E2E Dev do?

Spin up a live local Omnigent server and exercise the GitHub Copilot SDK harness end-to-end — build copilot agents, run real turns, smoke-test, and bug-bash. Copilot SDK E2E Dev is an agent skill from omnigent-ai/omnigent. Spin up a live local Omnigent server and exercise the GitHub Copilot SDK harness end-to-end — build copilot agents, run real turns, smoke-test, and bug-bash.

When should I use Copilot SDK E2E Dev?

Copilot SDK E2E Dev fits situations like: tasks that involve End-to-end testing; tasks that involve QA and bug reports.

How do I install Copilot SDK E2E Dev in Claude Code?

Run `npx skills add omnigent-ai/omnigent --skill copilot-sdk-e2e-dev -a claude-code`. Or copy the skill folder (.claude/skills/copilot-sdk-e2e-dev in omnigent-ai/omnigent) into .claude/skills/copilot-sdk-e2e-dev in your project. Claude Code loads it when a task matches its description.

How do I install Copilot SDK E2E Dev in Codex?

Run `npx skills add omnigent-ai/omnigent --skill copilot-sdk-e2e-dev -a codex`. Or copy the skill folder (.claude/skills/copilot-sdk-e2e-dev in omnigent-ai/omnigent) into .agents/skills/copilot-sdk-e2e-dev in your project. Codex loads it when a task matches its description.

Can I use Copilot SDK E2E Dev in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add omnigent-ai/omnigent --skill copilot-sdk-e2e-dev -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/copilot-sdk-e2e-dev, .gemini/skills/copilot-sdk-e2e-dev, .github/skills/copilot-sdk-e2e-dev and .opencode/skills/copilot-sdk-e2e-dev in your project.

What does Copilot SDK E2E Dev need to run?

Going by SKILL.md and its folder, Copilot SDK E2E Dev needs the command-line tools its instructions call (python, uv, gh and curl) and credentials named GH_TOKEN, COPILOT_GITHUB_TOKEN and GITHUB_TOKEN. Our summary lists: Python 3; A credential in COPILOT_GITHUB_TOKEN; A credential in GITHUB_TOKEN.

Does Copilot SDK E2E Dev access the network?

SKILL.md contains no URLs. Its commands use uv, gh and curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Copilot SDK E2E Dev safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Copilot SDK E2E Dev use?

Copilot SDK E2E Dev is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Copilot SDK E2E Dev use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Copilot SDK E2E Dev?

Skills that share tags, products or a category with Copilot SDK E2E Dev: E2E (wp-media/wp-rocket, 767 stars), Specx Tests (maksimzayats/specx, 202 stars), OpenROAD Issue Triage (The-OpenROAD-Project/OpenROAD, 3.2k stars) and Beava PR Review (beava-dev/beava, 138 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Copilot SDK E2E Dev?

omnigent-ai (a GitHub organization) maintains it in omnigent-ai/omnigent, which has 10,633 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 7, 2026.

Source: omnigent-ai/omnigent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.