Agent skill

Tech Debt CI Review

by ZaxbyHub in ZaxbyHub/opencode-swarm

Deep technical debt and CI stability audit for identifying test theater, missing or mis-scoped tests, actual and potential test failures, flaky-test risk, dependency/toolchain brittleness, and…

MITAuto-check passedDevelopment

Install Tech Debt CI Review

skills CLI
$ npx skills add ZaxbyHub/opencode-swarm --skill tech-debt-ci-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ZaxbyHub/opencode-swarm tech-debt-ci-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ZaxbyHub/opencode-swarm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/tech-debt-ci-review .claude/skills/tech-debt-ci-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tech-debt-ci-review
GitHub stars
494
Token cost
~2.6k tokens
SKILL.md length
1,431 words
Files
1
Skills in repo
91
Repo updated
First seen
Licence
MIT

At a glance

Deep technical debt and CI stability audit for identifying test theater, missing or mis-scoped tests, actual and potential test failures, flaky-test risk, dependency/toolchain brittleness, and…

  • Works in 6 steps: Inventory first → Parallel exploration → Candidate validation in a fresh reviewer… → …
  • Tasks that involve Technical debt
  • SKILL.md covers Mission, Operating stance, Quality-over-speed directive and Required workflow, plus 12 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Tech Debt CI Review is an agent skill from ZaxbyHub/opencode-swarm. Deep technical debt and CI stability audit for identifying test theater, missing or mis-scoped tests, actual and potential test failures, flaky-test risk, dependency/toolchain brittleness, and structural debt that prevents PRs from going green safely.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Technical debt and Failing and flaky tests. The repository describes itself as: Architect-centric agentic swarm plugin for OpenCode. Hub-and-spoke orchestration with SME consultation, code generation, and QA review. The licence is MIT.

When your agent uses it

  • Tasks that involve Technical debt
  • Tasks that involve Failing and flaky tests

Example prompts

  • “/tech-debt-ci-review”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Inventory first
  2. Parallel exploration
  3. Candidate validation in a fresh reviewer context
  4. Audit checklist
  5. Runtime-aware validation
  6. Final output

What it can do on your machine

Read from SKILL.md and the folder at commit b63a4bd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tech Debt CI Review loads about 2.6k tokens when it runs. Until then it costs about 68 tokens; SKILL.md has 1,431 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~68
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ZaxbyHub/opencode-swarm at commit b63a4bd, republished under its MIT licence (© ZaxbyHub). 1,431 words, ~2,649 tokens.

Download SKILL.mdSave it as .claude/skills/tech-debt-ci-review/SKILL.md (or your agent's skills folder).
name
tech-debt-ci-review
description
Deep technical debt and CI stability audit for identifying test theater, missing or mis-scoped tests, actual and potential test failures, flaky-test risk, dependency/toolchain brittleness, and structural debt that prevents PRs from going green safely.
audience
swarm-plugin
disable-model-invocation
true

/tech-debt-ci-review

Run a deep technical debt and CI stability audit of the current repository.

Mission

Identify every meaningful source of:

  • technical debt with real CI impact
  • test theater
  • missing or mis-scoped tests
  • actual and potential test failures
  • CI instability
  • flaky-test risk
  • dependency/build/toolchain brittleness
  • verification gaps that prevent the repository from reaching and staying green in pull requests

Do not build features. Do not do opportunistic cleanup for its own sake. Do not preserve noisy tests or workflows just because they make dashboards look busy.

Operating stance

  • Treat green-looking CI, high coverage, test names, comments, docs, release notes, and examples as claims or hints, not proof.
  • Treat code and tests as plausible until verified, not correct until disproven.
  • No finding is valid without exact file path and line evidence, or exact workflow/job/command evidence for CI-level issues.
  • No approval-like conclusion is valid without positive evidence of what was checked.
  • If a finding depends on runtime behavior, framework behavior, timing, sequencing, state, or exploitability, do not over-claim from static code alone if safe validation is available.

Quality-over-speed directive

Confidence in the test and CI signal is the main success metric. There is no time pressure. There is no reward for finishing in fewer passes. Do not batch more aggressively, skip validation, or stop early because the repository is large or the audit feels expensive. Large codebases require more disciplined verification, not less.

Required workflow

Phase 0 — Inventory first

In this repository, start from the authoritative contracts: AGENTS.md (especially invariant 6 — never validate the whole repo through broad test_runner scopes; use the per-file isolation loops from TESTING.md and contributing.md) and .github/workflows/ci.yml for what CI actually enforces.

Read enough of the repo to build a quality and CI map:

  • dependency manifests and lockfiles
  • CI workflows and job configs first
  • top-level README/docs
  • existing QA reports, CI handoffs, or known-failure notes if present
  • test runner configuration
  • any task brief / PR text / diff context if relevant

Build a map of:

  • tech stack
  • CI surface
  • test surface
  • quality surface
  • debt surface
  • public/risky surfaces
Phase 1 — Parallel exploration

Use subagents for breadth so the main context stays clean. Split the repo into disjoint scopes and/or audit families. Prefer small, focused scopes over giant repo-wide sweeps.

Recommended exploration lanes:

  • CI workflow and job graph correctness
  • test theater and falsifiability
  • flaky-test risk and nondeterminism
  • required-vs-optional test necessity
  • debt hotspots and maintainability drag
  • dependency/build/toolchain brittleness
  • behavioral regression exposure in risky modules

Explorer-style subagents should:

  • read every file in their assigned scope
  • stay focused on one scope or issue family
  • return candidate findings only, not final truth
  • cite exact paths and lines
  • call out where deeper validation is needed
Phase 2 — Candidate validation in a fresh reviewer context

Treat exploration output as a hypothesis engine, not a final verdict. Use one or more fresh reviewer contexts to validate candidate findings.

For each candidate finding, classify exactly one:

  • CONFIRMED
  • DISPROVED
  • UNVERIFIED
  • PRE_EXISTING

Reviewer must:

  1. Re-open the exact files and lines referenced by the explorer
  2. Read enough surrounding context to judge correctly
  3. Check callers, callees, tests, config, CI jobs, manifests, scripts, docs, and runtime assumptions as needed
  4. Check whether a mitigating control invalidates the candidate
  5. Run the smallest safe validation loop available when the claim depends on runtime behavior, timing, sequencing, or environment state
  6. Reclassify severity if the explorer overclaimed it
  7. Record the reason for disproof when rejecting a candidate

If the issue cannot be proven due to ambiguity or missing context, mark it UNVERIFIED, not DISPROVED.

Phase 3 — Audit checklist

Apply this checklist across the repo:

  1. Actual CI failure risk
  • broken workflow YAML or invalid job wiring
  • incorrect needs/dependency graph
  • CI commands that do not match real repo tooling
  • matrix values or OS assumptions likely to fail
  • caches restoring incompatible artifacts
  • missing setup/bootstrap steps
  • local-only assumptions in CI
  • required checks whose producer/consumer contracts have drifted
  1. Test theater
  • tests asserting only existence, truthiness, or snapshots with no behavioral meaning
  • tests that would pass if the implementation were removed
  • tests validating mocks instead of behavior
  • edge-case-named tests that only exercise happy path
  • coverage theater that ignores critical paths
  1. Need for tests
  • changed behavior with no regression tests
  • public API/schema changes with no protective tests
  • critical paths lacking focused tests
  • expensive edge cases not covered
  • places where tests are unnecessary or low ROI
  • opportunities to replace brittle E2E with cheaper lower-level tests
  1. Actual and potential test failures
  • currently failing tests
  • tests likely to fail in CI despite passing locally
  • environment-sensitive, order-sensitive, timing-sensitive, or race-prone tests
  • hidden shared state between tests
  • improper cleanup/teardown
  • dependence on wall clock, timezone, locale, random seeds, network, filesystem, temp paths, or machine performance
  • brittle selectors or snapshot churn
  • hidden credential or unavailable service assumptions
  1. Flaky-test risk
  • intermittent assertions
  • retries masking failures instead of fixing them
  • sleeps, polling hacks, fixed delays
  • nondeterministic data generation without seeding
  • parallel execution hazards
  • flaky fixture setup or external dependency reliance
  • quarantine/retry practices that keep CI green while hiding root causes
  1. Mutation-minded test quality Use mutation thinking even if no mutation tool exists. For important tests, ask:
  • would this fail if the function returned early?
  • would it fail if a boolean were inverted?
  • would it fail if an error path were swallowed?
  • would it fail if a result were hardcoded?
  • would it fail if the dependency call were skipped? If not, the test is likely theater or too weak.
  1. Technical debt with CI impact
  • duplicated logic increasing bug-fix surface area
  • high-complexity modules hard to test safely
  • hidden coupling across modules or fixtures
  • stale abstractions and wrappers with no value
  • outdated or partially migrated tooling
  • dead code or dead tests raising maintenance cost
  • brittle build/bootstrap flows
  • lack of ownership signals for flaky tests or red suites
  1. Dependency, build, and toolchain brittleness
  • phantom or suspicious dependencies
  • undeclared runtime tools
  • local and CI install/build drift
  • lockfile or version drift
  • cross-platform dependency issues
  • build scripts relying on hidden global tools
  • optional tools treated as required without guardrails
  1. Intended-vs-actual verification for CI and testing For every important workflow claim, verify whether the repo actually does what it claims, for example:
  • PRs are green by default
  • tests catch regressions
  • lint/typecheck/build are enforced
  • reviewers can trust CI
  • this suite protects behavior X
Show full SKILL.md (400 more words)Show less
Phase 4 — Runtime-aware validation

Use runtime validation selectively, not blindly. If a candidate depends on actual behavior, timing, workflow sequencing, role/state transitions, or environment interaction, use the smallest safe validation loop available. Do not use retries as proof of correctness. Retries can help separate flaky from deterministic failures, but they are not a fix.

Phase 5 — Final output

Write the final report to .claude/reports/tech-debt-report.md (create the directory if needed). Do not write it to the repository root, and do not commit it unless the user asks — it is working output, not a deliverable.

Use this structure:

Technical Debt and CI Stability Report

Generated: [timestamp] Scope: [scopes and artifacts covered] Files reviewed: [count] Parallel exploration used: YES | NO Independent reviewer validation used: YES | NO Runtime validation used: YES | NO

Executive Summary

[2-4 sentences]

Current PR/CI Blockers

[only issues actively breaking or very likely to break PR CI]

Critical and High Findings

[full detail for CONFIRMED and PRE_EXISTING only]

Test Theater Findings

[tests that create confidence without protection]

Missing or Mis-Scoped Tests

[where tests are needed, where they are not needed, and where the testing pyramid is wrong]

Flaky-Test Risks

[confirmed or likely nondeterminism sources]

Structural Debt with CI Impact

[coupling, brittle setup, outdated tooling, duplicated logic, etc.]

Dependency and Toolchain Risks

[build/install/runtime brittleness]

Coverage Notes

[what remained unverified and why]

Validation Notes

  • candidate findings generated: N
  • reviewer confirmed: N
  • reviewer disproved: N
  • reviewer unverified: N
  • reviewer pre_existing: N

Green-PR Remediation Order

  1. current red-build blockers
  2. flaky tests and nondeterministic CI failures
  3. test theater in critical paths
  4. missing tests for expensive regressions
  5. structural debt directly causing CI and change fragility
  6. lower-value cleanup only after the above

When presenting the report, include:

  1. the current PR/CI blockers in one-line form
  2. the most important test theater patterns
  3. the highest-value missing-test gaps
  4. the top flaky-test risks
  5. the minimal remediation order required to get PR CI trustworthy and green

Final rules

  • No finding without exact evidence.
  • No approval-like conclusion without positive evidence of what was checked.
  • If a candidate cannot be proven, mark it UNVERIFIED rather than CONFIRMED.
  • Do not recommend tests that do not materially improve defect detection.
  • Do not preserve test theater just because it makes dashboards look good.
  • The point is not to maximize test count or finding count.
  • The point is to maximize confidence that PR CI can go green without unsafe bypasses.

© ZaxbyHub, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/tech-debt-ci-review of ZaxbyHub/opencode-swarm.

Open the folder on GitHubat commit b63a4bd

Compare with similar skills

Tech Debt CI Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tech Debt CI Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tech Debt CI Review this skillZaxbyHub/opencode-swarm494—~2.6kAutomated safety check: PassMIT
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Iterate PRmeshery/meshery-operator1517 repos~2.2kAutomated safety check: PassApache-2.0
React Router Bug Fix Workflowremix-run/react-router57k—~1.3kAutomated safety check: PassMIT
Bug InvestigatorMageByte-Zero/spec-superflow8461 repos~1.6kAutomated safety check: PassMIT
Runtime Debugvercel/next.js143k1 repos~618Automated safety check: PassMIT

Similar skills

  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Iterate PR

    meshery/meshery-operator

    Iterate on a PR until CI passes. An agent skill from meshery/meshery-operator.

    151 GitHub starsUsed in 7 repos~2.2k tokens
    DevelopmentAuto-check passed
  • React Router Bug Fix Workflow

    remix-run/react-router

    Fixes a React Router bug reported in a GitHub issue end to end: fetching the issue, validating the reproduction, writing a failing test and implementing the fix on a new branch.

    57k GitHub stars~1.3k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Bug Investigator

    MageByte-Zero/spec-superflow

    A skill your agent uses when encountering any bug, test failure, or unexpected behavior during spec-superflow execution, before proposing fixes.

    846 GitHub starsUsed in 1 repo~1.6k tokens
    DevelopmentAuto-check passed
  • Runtime Debug

    vercel/next.js

    Official

    Debug and verification workflow for runtime-bundle and module-resolution regressions.

    143k GitHub starsUsed in 1 repo~618 tokens
    DevelopmentAuto-check passed
  • PlotJuggler Ship Check

    PlotJuggler/PlotJuggler

    Runs a gated finish-line checklist before committing a PlotJuggler PJ4 change: build proof, red-test triage, hooks, docs freshness and a diff self-review.

    6.2k GitHub stars~1.3k tokensUpdated 10 days ago
    DevelopmentAuto-check passed

More from ZaxbyHub/opencode-swarm

All 91 skills in this repo
  • Codebase Review Swarm

    ZaxbyHub/opencode-swarm

    Runs an evidence-gated, quote-grounded audit of a codebase for security, QA, accessibility, performance and more, and writes a verified report without changing source files.

    496 GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Issue Tracer

    ZaxbyHub/opencode-swarm

    Drives a bug report from validation and root-cause tracing through a critic-reviewed plan, an approved minimal fix and a PR-ready closure, never merging without recorded human approval.

    496 GitHub stars~4.4k tokensUpdated today
    Auto-check passed
  • Commit and PR Publishing for Codex

    ZaxbyHub/opencode-swarm

    Codex adapter for opencode-swarm that governs commits, pushes, draft PRs, PR body updates and CI closeout, deferring to the repo's canonical commit-pr protocol.

    496 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Durable Session State

    ZaxbyHub/opencode-swarm

    Keeps plans, decisions, evidence and reviewer verdicts in small files so long multi-phase tasks survive context compaction and session resumes.

    496 GitHub stars~896 tokensUpdated today
    Auto-check passed
  • Swarm PR Feedback Closer

    ZaxbyHub/opencode-swarm

    Ingests existing pull request feedback such as review comments and CI failures, verifies each claim, fixes confirmed issues and reports closure status for every item.

    496 GitHub stars~14k tokensUpdated today
    Auto-check passed
  • Swarm PR Subscribe

    ZaxbyHub/opencode-swarm

    Monitor a pull request after creation and act autonomously on pushed PR activity.

    496 GitHub stars~2.2k tokensUpdated today
    Auto-check passed

Questions about Tech Debt CI Review

What does Tech Debt CI Review do?

Deep technical debt and CI stability audit for identifying test theater, missing or mis-scoped tests, actual and potential test failures, flaky-test risk, dependency/toolchain brittleness, and…. Tech Debt CI Review is an agent skill from ZaxbyHub/opencode-swarm. Deep technical debt and CI stability audit for identifying test theater, missing or mis-scoped tests, actual and potential test failures, flaky-test risk, dependency/toolchain brittleness, and structural debt that prevents PRs from going green safely.

When should I use Tech Debt CI Review?

Tech Debt CI Review fits situations like: tasks that involve Technical debt; tasks that involve Failing and flaky tests.

How do I install Tech Debt CI Review in Claude Code?

Run `npx skills add ZaxbyHub/opencode-swarm --skill tech-debt-ci-review -a claude-code`. Or copy the skill folder (.claude/skills/tech-debt-ci-review in ZaxbyHub/opencode-swarm) into .claude/skills/tech-debt-ci-review in your project. Claude Code loads it when a task matches its description.

How do I install Tech Debt CI Review in Codex?

Run `npx skills add ZaxbyHub/opencode-swarm --skill tech-debt-ci-review -a codex`. Or copy the skill folder (.claude/skills/tech-debt-ci-review in ZaxbyHub/opencode-swarm) into .agents/skills/tech-debt-ci-review in your project. Codex loads it when a task matches its description.

Can I use Tech Debt CI Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ZaxbyHub/opencode-swarm --skill tech-debt-ci-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tech-debt-ci-review, .gemini/skills/tech-debt-ci-review, .github/skills/tech-debt-ci-review and .opencode/skills/tech-debt-ci-review in your project.

What does Tech Debt CI Review need to run?

SKILL.md names no scripts, command-line tools or credentials: Tech Debt CI Review is instructions for the agent only.

Does Tech Debt CI Review access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Tech Debt CI Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Tech Debt CI Review use?

Tech Debt CI Review is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tech Debt CI Review use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Tech Debt CI Review?

Skills that share tags, products or a category with Tech Debt CI Review: PR Babysitter (openinterpreter/openinterpreter, 69k stars), Iterate PR (meshery/meshery-operator, 151 stars), React Router Bug Fix Workflow (remix-run/react-router, 57k stars) and Bug Investigator (MageByte-Zero/spec-superflow, 846 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tech Debt CI Review?

ZaxbyHub (a GitHub organization) maintains it in ZaxbyHub/opencode-swarm, which has 494 GitHub stars. The repository holds 91 skills in this directory. The repository was last updated on October 10, 2026.

Source: ZaxbyHub/opencode-swarm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.