CI/CD Failure Troubleshooting
ruby-git/ruby-git
Diagnoses and fixes failing GitHub Actions runs by identifying the failure, fetching only the relevant logs, finding the root cause and reproducing it locally.
Debugs failing GitHub Actions CI runs for PostHog PRs, commits, and branches, and answers broad CI-health questions ("is CI red?", "is master green today?", "what's broken right now?").
$ npx skills add PostHog/posthog-foss --skill debugging-ci-failures -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install PostHog/posthog-foss debugging-ci-failures --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/PostHog/posthog-foss.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/debugging-ci-failures .claude/skills/debugging-ci-failures && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "debugging-ci-failures" agent skill from https://github.com/PostHog/posthog-foss/tree/master/.agents/skills/debugging-ci-failures into .claude/skills/debugging-ci-failures/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debugging-ci-failures", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/PostHog/posthog-foss/tree/master/.agents/skills/debugging-ci-failuresType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add PostHog/posthog-foss --skill debugging-ci-failures -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install PostHog/posthog-foss debugging-ci-failures --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PostHog/posthog-foss.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/debugging-ci-failures .agents/skills/debugging-ci-failures && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "debugging-ci-failures" agent skill from https://github.com/PostHog/posthog-foss/tree/master/.agents/skills/debugging-ci-failures into .agents/skills/debugging-ci-failures/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debugging-ci-failures", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add PostHog/posthog-foss --skill debugging-ci-failures -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install PostHog/posthog-foss debugging-ci-failures --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PostHog/posthog-foss.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/debugging-ci-failures .cursor/skills/debugging-ci-failures && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "debugging-ci-failures" agent skill from https://github.com/PostHog/posthog-foss/tree/master/.agents/skills/debugging-ci-failures into .cursor/skills/debugging-ci-failures/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debugging-ci-failures", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/PostHog/posthog-foss.git --path .agents/skills/debugging-ci-failures--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add PostHog/posthog-foss --skill debugging-ci-failures -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install PostHog/posthog-foss debugging-ci-failures --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PostHog/posthog-foss.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/debugging-ci-failures .gemini/skills/debugging-ci-failures && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "debugging-ci-failures" agent skill from https://github.com/PostHog/posthog-foss/tree/master/.agents/skills/debugging-ci-failures into .gemini/skills/debugging-ci-failures/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debugging-ci-failures", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install PostHog/posthog-foss debugging-ci-failuresInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add PostHog/posthog-foss --skill debugging-ci-failures -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/PostHog/posthog-foss.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/debugging-ci-failures .github/skills/debugging-ci-failures && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "debugging-ci-failures" agent skill from https://github.com/PostHog/posthog-foss/tree/master/.agents/skills/debugging-ci-failures into .github/skills/debugging-ci-failures/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debugging-ci-failures", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add PostHog/posthog-foss --skill debugging-ci-failures -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install PostHog/posthog-foss debugging-ci-failures --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PostHog/posthog-foss.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/debugging-ci-failures .opencode/skills/debugging-ci-failures && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "debugging-ci-failures" agent skill from https://github.com/PostHog/posthog-foss/tree/master/.agents/skills/debugging-ci-failures into .opencode/skills/debugging-ci-failures/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debugging-ci-failures", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
debugging-ci-failuresDebugs failing GitHub Actions CI runs for PostHog PRs, commits, and branches, and answers broad CI-health questions ("is CI red?", "is master green today?", "what's broken right now?").
Debugging CI Failures is an agent skill from PostHog/posthog-foss, published by the product's own GitHub organization. Debugs failing GitHub Actions CI runs for PostHog PRs, commits, and branches, and answers broad CI-health questions ("is CI red?", "is master green today?", "what's broken right now?"). Use when the user asks why CI is red, asks for the current CI or master status, or mentions a failing check, GitHub Actions run, Depot runner, Depot CI run, workflow, job, shard, merge queue kick, flaky test, lint failure, typecheck failure, snapshot diff, migration check, generated types drift, or skills build failure…
Its SKILL.md is about 5.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/master-red-incident.md` and `references/master-red-workflow-setup.md`).
It sits in Testing & QA, covering Failing and flaky tests, CI/CD and Debugging. It works with PostHog, GitHub Actions and GitHub. The repository describes itself as: PostHog FOSS is a read-only mirror of PostHog, with all proprietary code removed. NOTE: This repo is synced automatically from the main PostHog repo. Please raise any issues and… The licence is MIT.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 2c48221. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
ghpnpmuvFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comAlso links to:
githubstatus.comstatus.depot.devFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
TRUNK_API_TOKENDEPOT_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Debugging CI Failures loads about 5.7k tokens when it runs, and up to ~9.6k if it reads all its reference files. Until then it costs about 234 tokens; SKILL.md has 2,852 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from PostHog/posthog-foss at commit 2c48221, republished under its MIT licence (© PostHog). 2,852 words, ~5,717 tokens.
.claude/skills/debugging-ci-failures/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Before you propose a change to CI, check things already tried for the idea. It records what was measured, and why some good-sounding changes were reverted or rejected.
Find the first meaningful failure, classify it, reproduce the smallest useful case locally when appropriate, and report the result. Avoid public-visible or irreversible actions unless the user explicitly asks.
For an interactive investigation, start with the hogli ci:insights digest.
It aggregates across runs and branches, which gh cannot do cheaply, and tells
you whether a failure is likely trunk-borne, gate-only, or isolated to a small
set of branches. gh is authoritative for one run's current state and
attribution. Use the digest to decide what to inspect; use gh to confirm
whose failure it is and exactly what failed in a given run.
The unattended Master-red workflow is the exception. Follow its reference, which defines a first step that works in the task sandbox.
This skill triages and classifies. Once a failure is confirmed flaky, hand off
to the fixing-flaky-tests skill, which owns local reproduction, root-cause
fixing, and N-run validation. For "who broke master" — the culprit commit and
the commit that fixed it — hand off to the investigating-ci-failures skill,
which owns the green/red boundary analysis. For aggregate pipeline health (is CI
getting slower, which workflow is the long pole, how long PRs take to merge),
read diagnosing-ci-and-merge-bottlenecks. Both are product skills under
products/engineering_analytics/skills/, not invocable here: read the
SKILL.md at that path.
GitHub Actions goes down often enough that it belongs before any log reading, and
a platform incident makes every other signal a symptom. It costs two page loads:
https://www.githubstatus.com/ for GitHub, https://status.depot.dev/ for the
runners. Check them whenever failures are broad — several workflows at once, a
burst of runs failing together, jobs dying before Checkout, or anything red
across unrelated PRs.
Report an outage as an outage, name the component, and stop recommending reruns.
Do not do any of these without explicit approval in the current conversation:
.github/workflows/ files (CI infra changes need human review).Read-only gh calls and read-only GitHub tools are fine. If you need to
change local Git state, make sure it is necessary for the task and does not
overwrite unrelated work.
hogli ci:insights reads PostHog's own engineering analytics — the cross-run
failure history a single run can't show. Consult it before any raw gh log
archaeology.
hogli ci:insights # digest for the current repo + branch
hogli ci:insights search "<error or test name>" # match a specific failure
hogli ci:insights view <ref> # one failure in full
hogli ci:insights view <ref> --logs # ...plus the failing log linessearch "<error>" to match it before reading logs, then
view <ref> on a row the digest or search printed. --logs prints the thinned
failing lines from that failure's latest run, which is usually enough to
classify without touching gh.Read each row's state as a triage ranking:
| State | Means |
|---|---|
breaking_master | failing on the default branch and that job's latest run is still red |
blocking_merge_queue | failed only on merge-queue gate branches in the window |
novel_burst | new within a day, already spreading across branches, not on trunk yet |
potentially_resolved | hit trunk but that job's latest run is green again |
flaky | sporadic across two or more branches over more than a day |
pr_only | limited branch spread; job status may be missing or behind the failure lines |
potentially_resolved is a hint, not a conclusion: confirm from run data before
reporting a failure as already fixed.
blocking_merge_queue proves a gate failure happened in the window, not that it
still blocks landings. Check the current queue run with gh before reporting it
as active. Likewise, confirm pr_only from the current run before assigning the
failure to a PR; it is also the fallback when job status is missing or stale.
Caveats to carry into whatever you report:
workflow_run webhook settles it, so
during a live incident confirm a specific run against gh.If nobody has signed in on this machine, hogli ci:insights exits 78. Treat
exit 78 as "no CI insights available" and fall back to the gh-based
inspection below — then tell the user they can run hogli posthog:login
once, which opens a browser and needs no API key. Do not run it yourself: it
waits on a consent screen you cannot see. Surface what you find per the Safety
rules — do not auto-apply a fix.
Determine the target in this order:
gh pr view --json number,headRefName,statusCheckRollup.A PR kicked from the merge queue is the exception: its own checks are the
wrong target. Trunk tests each queued PR on a trunk-merge/pr-<n>/<uuid>
branch holding master plus every PR queued ahead of it whose impacted targets
overlap its own. The lane script over-reports targets on purpose, so in
practice that is most of the queue. So:
trunk merge status <n> (/merging-prs step 4), not gh pr checks. Trunk publishes no check run here.
The branch is ephemeral; the run and its logs stay on GitHub, and the
warehouse keeps its jobs under that head_branch (query 8 in the
investigating-ci-failures references).
Backend CI is the exception when CI_BACKEND_DEPOT_MERGE_QUEUE_PERCENT hands the batch to Depot CI: the backend tests then run on Depot, and the Django Tests Pass check links that run.…-bisection branch is the exception to "the other merge commits are the
first suspects": Trunk creates it after a batch fails, and it carries only
master plus the one PR. The test PR's body names the master SHA it is based
on; if master's run at that SHA passed the same test, the failure is an
environment-dependent race (fixing-flaky-tests step 1), not this PR's.blocking_merge_queue state.Before reading logs, read the shared CI report comment. It collects independent CI signals and advisories in many sections.
Read the full raw comment, not GitHub's collapsed view. This command finds it even when the PR has many comments:
gh api --paginate "repos/<owner>/<repo>/issues/<pr>/comments?per_page=100" \
--jq '.[] | select(.user.login == "github-actions[bot]" and (.body | startswith("<!-- posthog-ci-report -->"))) | .body'Inspect every ci-report:section block, including unknown sections. Record its
title, status, summary, and links. Do not stop at the first fail section or
treat the comment heading as an overall verdict.
fail as a lead, and match it to the current job.alert and warn as non-blocking findings. Review their details and
follow their stated action, but do not call either a failed job unless its
job failed.ok and info are not test results.The report is a summary. A section without a head SHA or run link can be stale. Confirm it against the current job. If it disagrees with the current logs, report the mismatch and use the logs for the cause.
Inspect read-only:
gh pr checks <pr>
gh pr view <pr> --json statusCheckRollup
gh run view <run-id> --json jobs,conclusion,name,workflowName,url
gh run view <run-id> --log-failedUse the full job log only when --log-failed lacks the failing command or
enough surrounding output:
gh run view <run-id> --log --job <job-id>Given a run id, the engineering-analytics-run-failure-logs MCP tool returns
every failed job's error region with original line numbers, already thinned.
One call instead of a jobs listing plus a log download, and it works when the
job died before any test ran. It is bounded by Logs retention, so fall back to
gh for older runs. Given a PR number instead of a run id,
engineering-analytics-ci-failure-logs does the same across every run that PR
has pushed, so an earlier push's failure is still there.
Extract these before classifying:
.github/workflows/ci-backend.yml.backend-tests (4/10).Run pytest.When scanning logs, search for FAIL, Error, error:, assert,
Traceback, exit code, and ##[error]. Stop at the first failing step that
explains the run's conclusion. Keep excerpts under 40 lines.
For test-job failures, the trunk MCP server's investigate-ci-failure tool
is a shortcut past log scanning: give it the run URL
(https://github.com/PostHog/posthog/actions/runs/<run-id>) and it returns
structured failing-test details — names, error messages, stdout/stderr — from
the results CI uploaded to Trunk Flaky Tests, with quarantined known-flaky
tests already filtered out. It only covers what ran and uploaded: for jobs
that died before tests (build, setup, lint), use
engineering-analytics-run-failure-logs or gh run view --log.
Authenticate once via /mcp → trunk (browser OAuth); headless environments
instead add an Authorization: Bearer header with a TRUNK_API_TOKEN org
token to the server entry in .mcp.json. To dig into one test's flakiness
history, hand off to fixing-flaky-tests, which covers the search-test and
fix-flaky-test tools.
A PR can route its backend tests to Depot CI.
Then the GitHub Actions run holds only the relay: Django Tests Pass fails with "Backend tests on Depot CI concluded failure", and the GitHub run logs show nothing more.
A re-run of that job is not read-only: it retries the failed Depot jobs and reports the new verdict, so the rule against re-running CI applies to it.
The log of the Relay the Depot verdict step prints the Depot run URL, each failed step with its log lines, and the label that sends the PR back to GitHub Actions.
Read more of the failure from Depot with the depot CLI, which the flox environment installs.
Run depot ci diagnose --org <org> --workflow <workflow> with the two ids from the Depot run URL in the relay log.
For logs, artifacts, and test results, load the depot-ci skill.
Its references/posthog-check-run-semantics.md gets Depot ids from a commit's check runs when you have no relay log.
If a depot command fails to authenticate, check DEPOT_TOKEN first: it overrides the saved login, so a wrong one fails every command.
Without it, ask the user to run ! depot login --clear.
Do not run it yourself: it waits on a browser sign-in.
If the user has no access to the Depot org, report the relay log's no-Depot retry options instead.
| Signal in the log | Class | First action |
|---|---|---|
AssertionError, test diff, FAILED test_... in a committed test file | code regression | reproduce with hogli test <path>::<test> |
Test failed here, passed on master or on rerun in the same PR | flaky test | confirm against master history; to fix, use fixing-flaky-tests |
ruff, oxlint, stylelint, markdownlint, prettier errors | lint | hogli lint:python:fix or hogli format on touched files |
mypy, pyright, tsc, typescript:check errors | typecheck | run the same checker locally, not the full suite |
| Chromatic / Storybook / Playwright visual diff, snapshot mismatch | snapshot / visual | surface the diff URL; do NOT auto-accept snapshots |
manage.py migrate error, migrations:check failure, missing migration | migration / schema | hogli migrations:check locally |
| OpenAPI schema diff, generated API types out of sync | codegen drift | hogli build:openapi |
Cannot connect, ECONNREFUSED, address already in use, OOM, runner killed, setup step timeout | infra / runner | get the base rate before calling it transient (below) |
startup_failure conclusion, a job with zero recorded steps, or a log blob that 404s | infra / runner | no log to read; check https://www.githubstatus.com/ and the runs around it |
apt-get, uv sync, pnpm install, docker pull, setup action failures | environment / setup | diff .nvmrc, pyproject.toml, package.json, Dockerfiles |
hogli lint:skills, hogli build:skills failure | skills build | run the same hogli command locally |
SDK compat check, ci-survey-sdk-check, cross-version failure | SDK compatibility | check SDK version matrix for the affected package |
If multiple signals match, choose the most specific class. For example, prefer codegen drift over lint, migration over typecheck, and snapshot / visual over a generic Playwright test failure.
"Transient" is a claim about how often the job fails, so do not assert it from
one run. A job that dies before its tests run leaves no test-level evidence: no
FAILED line, so no fingerprint and no span, so it never appears as a row in
broken_tests or in the flaky-tests tool. It is still visible as a job
conclusion, which is what the digest's master-failures section groups by, and
engineering-analytics-run-failure-logs still returns its failing lines because
it reads by run, not by test. Start there, not from a test.
Unlike the span-derived test reads, this one can give you a real rate: the
warehouse records every job attempt, greens included, so the denominator is
honest. Query 7 in
products/engineering_analytics/skills/investigating-ci-failures/references/investigation-queries.md
is copy-ready; that skill also owns the wider investigation.
Read the result as:
trunk merge yourself needs approval, per the Safety rules above.startup_failure before they
start and so leaves no log at all.Run only the narrowest command that exercises the failure. If the command shape
is unclear, read .agents/skills/hogli/SKILL.md and hogli <command> --help.
| Class | Repro guidance |
|---|---|
| code regression | hogli test path/to/test.py::TestClass::test_method or hogli test <file.test.ts> |
| flaky test | Hand off to the fixing-flaky-tests skill. |
| lint | Use the failing formatter/linter on touched files, e.g. hogli format:python. |
| typecheck | Run the failing checker, e.g. pnpm --filter=@posthog/frontend typescript:check. |
| snapshot / visual | Run the specific Playwright or Storybook workflow; read playwright-test if needed. |
| migration / schema | hogli migrations:check; run migrations only if the user agrees. |
| codegen drift | hogli build:openapi. |
| infra / runner | No local repro. Get the base rate (above), report, and stop. |
| environment / setup | Reproduce the setup step only if cheap and relevant to changed files. |
| skills build | hogli lint:skills; if that passes, hogli build:skills. |
Do NOT run hogli test with no arguments. Do NOT run hogli nuke or
hogli dev:reset as a shortcut. Do NOT bypass hooks with --no-verify.
depot-ubuntu-latest or depot-ubuntu-latest-16.
Depot runs surface logs through the GitHub Actions UI / gh run view just
like standard GitHub-hosted runners, so read them there first.status.depot.dev covers the case
where Depot itself is the outage, and the depot-github-runners skill owns
runner troubleshooting beyond triage.Checkout completes (no app code ran), classify as
infra / runner. Do not propose code fixes.backend-tests (3/10) style). Reproduce from the specific failing test
path, not the shard index.Keep the response short. Include one likely-cause sentence and avoid deeper speculation.
Target: PR #<num> - run <run-id> (<workflow file>)
Failing job: <job name>
Failing step: <step name>
Command: <failing command>
Excerpt:
<up to 40 lines, trimmed around the failure>
Classification: <class from the table>
Shadow run: <yes | no>
Likely cause: <one sentence>
Local repro: <exact command, or "none">
Next action (needs your approval):
- <push fix | rerun job | update snapshot | none>If the classification is infra / runner or a shadow run, say so and stop;
do not propose a code change. For infra / runner, include the base rate and
whether a retry is warranted.
© PostHog, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in .agents/skills/debugging-ci-failures of PostHog/posthog-foss.
Open the folder on GitHubat commit 2c48221
Debugging CI Failures next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Debugging CI Failures this skillPostHog/posthog-foss | 721 | — | ~5.7k | Automated safety check: Pass | MIT | |
| CI/CD Failure Troubleshootingruby-git/ruby-git | 1.8k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Babysit PRZenUml/web-sequence | 150 | — | ~871 | Automated safety check: Pass | MIT | |
| GreptimeDB Fuzz CI Failure InvestigationGreptimeTeam/greptimedb | 6.7k | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | |
| CI Workflow Guidesgl-project/sglang | 37k | 2 repos | ~5.5k | Automated safety check: Pass | Apache-2.0 | |
| Analysing CI Failuresgolemcloud/golem | 1.5k | — | ~862 | Automated safety check: Pass | Custom licence |
ruby-git/ruby-git
Diagnoses and fixes failing GitHub Actions runs by identifying the failure, fetching only the relevant logs, finding the root cause and reproducing it locally.
ZenUml/web-sequence
Monitor and diagnose GitHub Actions checks on ZenUML web-sequence PRs, fixing code-caused CI failures when appropriate.
GreptimeTeam/greptimedb
Diagnoses a failed GreptimeDB fuzz CI job by pulling its GitHub Actions logs and fuzz artifacts, then matching the evidence to the local source code.
sgl-project/sglang
Guide to SGLang CI workflow orchestration — stage ordering, fail-fast, gating, partitioning, execution modes, and debugging CI failures.
golemcloud/golem
Analysing GitHub Actions CI failures from a run URL. An agent skill from golemcloud/golem.
latitude-dev/latitude-llm
Continuously monitor GitHub PR CI checks and automatically fix failures until all checks pass.
PostHog/posthog-foss
Author useful, low-noise log alerts on services in a PostHog project.
PostHog/posthog-foss
Operating procedure for the conflict-autoresolver agent: sweep open PostHog/posthog PRs that conflict with master, resolve the trivial conflicts (generated artifacts deterministically, source…
PostHog/posthog-foss
Help users debug PostHog Error Tracking stack-trace symbolication for any supported platform — JavaScript/TypeScript web, React Native (Hermes), Android (Proguard / R8), or iOS / macOS (dSYM).
PostHog/posthog-foss
Investigates distributed application performance using PostHog APM (OpenTelemetry span) data via MCP.
PostHog/posthog-foss
Debug and inspect LLM/AI agent traces using PostHog's MCP tools.
PostHog/posthog-foss
Diagnose why a product metric changed (dropped, spiked, or plateaued) by orchestrating breakdowns, actors, paths, lifecycle, retention, and annotations queries.
Works with
Categories
Debugs failing GitHub Actions CI runs for PostHog PRs, commits, and branches, and answers broad CI-health questions ("is CI red?", "is master green today?", "what's broken right now?"). Debugging CI Failures is an agent skill from PostHog/posthog-foss, published by the product's own GitHub organization.").
Debugging CI Failures fits situations like: the user asks why CI is red; asks for the current CI; mentions a failing check; GitHub Actions run.
Run `npx skills add PostHog/posthog-foss --skill debugging-ci-failures -a claude-code`. Or copy the skill folder (.agents/skills/debugging-ci-failures in PostHog/posthog-foss) into .claude/skills/debugging-ci-failures in your project. Claude Code loads it when a task matches its description.
Run `npx skills add PostHog/posthog-foss --skill debugging-ci-failures -a codex`. Or copy the skill folder (.agents/skills/debugging-ci-failures in PostHog/posthog-foss) into .agents/skills/debugging-ci-failures in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PostHog/posthog-foss --skill debugging-ci-failures -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debugging-ci-failures, .gemini/skills/debugging-ci-failures, .github/skills/debugging-ci-failures and .opencode/skills/debugging-ci-failures in your project.
Going by SKILL.md and its folder, Debugging CI Failures needs the command-line tools its instructions call (gh, pnpm and uv) and credentials named TRUNK_API_TOKEN and DEPOT_TOKEN.
SKILL.md names 3 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: githubstatus.com and status.depot.dev. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Debugging CI Failures is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.7k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Debugging CI Failures: CI/CD Failure Troubleshooting (ruby-git/ruby-git, 1.8k stars), Babysit PR (ZenUml/web-sequence, 150 stars), GreptimeDB Fuzz CI Failure Investigation (GreptimeTeam/greptimedb, 6.7k stars) and CI Workflow Guide (sgl-project/sglang, 37k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
PostHog (a GitHub organization, an official publisher) maintains it in PostHog/posthog-foss, which has 721 GitHub stars. The repository holds 213 skills in this directory. The repository was last updated on October 7, 2026.
Source: PostHog/posthog-foss on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.