GreptimeDB Fuzz CI Failure Investigation
GreptimeTeam/greptimedb
Diagnoses a failed GreptimeDB fuzz CI job by pulling its GitHub Actions logs and fuzz artifacts, then matching the evidence to the local source code.
Investigates a failing GitHub Actions run or job for Megatron-LM, finds the root cause plus the PR and test author involved, and files a structured bug issue.
$ npx skills add NVIDIA/Megatron-LM --skill mcore-create-issue -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/Megatron-LM mcore-create-issue --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/mcore-create-issue .claude/skills/mcore-create-issue && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "mcore-create-issue" agent skill from https://github.com/NVIDIA/Megatron-LM/tree/main/skills/mcore-create-issue into .claude/skills/mcore-create-issue/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mcore-create-issue", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/Megatron-LM/tree/main/skills/mcore-create-issueType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/Megatron-LM --skill mcore-create-issue -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/Megatron-LM mcore-create-issue --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/mcore-create-issue .agents/skills/mcore-create-issue && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "mcore-create-issue" agent skill from https://github.com/NVIDIA/Megatron-LM/tree/main/skills/mcore-create-issue into .agents/skills/mcore-create-issue/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mcore-create-issue", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/Megatron-LM --skill mcore-create-issue -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/Megatron-LM mcore-create-issue --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/mcore-create-issue .cursor/skills/mcore-create-issue && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "mcore-create-issue" agent skill from https://github.com/NVIDIA/Megatron-LM/tree/main/skills/mcore-create-issue into .cursor/skills/mcore-create-issue/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mcore-create-issue", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/Megatron-LM.git --path skills/mcore-create-issue--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/Megatron-LM --skill mcore-create-issue -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/Megatron-LM mcore-create-issue --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/mcore-create-issue .gemini/skills/mcore-create-issue && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "mcore-create-issue" agent skill from https://github.com/NVIDIA/Megatron-LM/tree/main/skills/mcore-create-issue into .gemini/skills/mcore-create-issue/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mcore-create-issue", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/Megatron-LM mcore-create-issueInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/Megatron-LM --skill mcore-create-issue -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/mcore-create-issue .github/skills/mcore-create-issue && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "mcore-create-issue" agent skill from https://github.com/NVIDIA/Megatron-LM/tree/main/skills/mcore-create-issue into .github/skills/mcore-create-issue/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mcore-create-issue", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/Megatron-LM --skill mcore-create-issue -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/Megatron-LM mcore-create-issue --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/mcore-create-issue .opencode/skills/mcore-create-issue && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "mcore-create-issue" agent skill from https://github.com/NVIDIA/Megatron-LM/tree/main/skills/mcore-create-issue into .opencode/skills/mcore-create-issue/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mcore-create-issue", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
mcore-create-issueInvestigates a failing GitHub Actions run or job for Megatron-LM, finds the root cause plus the PR and test author involved, and files a structured bug issue.
You pass a GitHub Actions job or run URL. The agent pulls out the run and job ids, lists the failed jobs when only a run was given (asking which to triage if several failed), and fetches the logs with the gh CLI, narrowing them to error lines such as FAILED, Traceback and assertion failures. If the filtered output is thin, it downloads the full logs and looks for the pytest failures section or the last non-zero exit.
It then resolves the triggering pull request from the run's head branch, which follows a pull-request/number pattern, and finds who last touched the failing test file by checking the PR's base branch, falling back to the PR's commits when the PR itself introduced the file. The result is a bug issue against NVIDIA/Megatron-LM. The skill ships with an evals file, a benchmark note and a skill card.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 486a126. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
ghpytestFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Megatron-LM CI Failure Triage loads about 1.6k tokens when it runs. Until then it costs about 27 tokens; SKILL.md has 396 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/Megatron-LM at commit 486a126, republished under its Apache-2.0 licence (© NVIDIA). 396 words, ~1,588 tokens.
.claude/skills/mcore-create-issue/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.Investigate a failing GitHub Actions job, extract the root cause, and file a
well-structured bug issue against NVIDIA/Megatron-LM.
The argument is a GitHub Actions URL. It will be one of:
https://github.com/<owner>/<repo>/actions/runs/<run_id>/job/<job_id>https://github.com/<owner>/<repo>/actions/runs/<run_id>Extract run_id and, if present, job_id.
If a job_id was provided, use that job directly.
If only a run_id was provided, list all failed jobs in the run:
gh run view <run_id> --repo NVIDIA/Megatron-LM --json jobs \
--jq '[.jobs[] | select(.conclusion == "failure") | {id: .databaseId, name: .name, url: .url}]'If multiple jobs failed, ask the user which one to triage, or triage all of them if they say so.
For each failed job, retrieve the logs and narrow them down to the failure:
# Pull the raw log and keep only error-bearing lines
gh api repos/NVIDIA/Megatron-LM/actions/jobs/<job_id>/logs 2>&1 \
| grep -E "(FAILED|ERROR|\bError\b|assert|Traceback|Exception|##\[error\])" \
| head -200Also capture the full job name:
gh run view --job <job_id> --repo NVIDIA/Megatron-LM --json name --jq .nameIf the grep output is sparse, download the full logs and look for the pytest
FAILURES section or the last non-zero exit signal.
Triggering PR: the run's head branch follows the pattern pull-request/<number>.
Extract it and resolve the PR:
gh run view <run_id> --repo NVIDIA/Megatron-LM --json headBranch --jq .headBranch
# → e.g. "pull-request/4332"
# Extract PR number and fetch metadata:
gh pr view <pr_number> --repo NVIDIA/Megatron-LM --json number,title,url \
--jq '{number: .number, title: .title, url: .url}'Test file author: find the GitHub login of whoever last touched the failing
test file. The file may not exist on main — first determine the PR's base
branch, then search from there:
# 1. Get the PR's base branch (e.g. "main", "dev", "release/X.Y")
gh pr view <pr_number> --repo NVIDIA/Megatron-LM --json baseRefName --jq .baseRefName
# 2. Search commits on that base branch
gh api "repos/NVIDIA/Megatron-LM/commits?path=<test-file-path>&sha=<base-branch>&per_page=1" \
--jq '.[0] | {login: .author.login, name: .commit.author.name, sha: .sha}'If the result is empty (file was introduced by the PR itself), query the PR's commits instead:
gh api "repos/NVIDIA/Megatron-LM/pulls/<pr_number>/commits" \
--jq '[.[] | select(.files? // [] | any(.filename == "<test-file-path>"))] | .[0].author.login'As a last resort, list the PR commits and pick the author of the commit whose message most closely relates to the failing test file.
From the logs, identify:
FAILED tests/...::... give the exact pytest node IDs.tests/unit_tests/transformer/moe/**/*.py - latest).Search for open issues that already cover the same test:
gh issue list --repo NVIDIA/Megatron-LM \
--state open \
--search "<failed-test-filename>" \
--json number,title,url \
--limit 10Pass --assignee <test-author-login> to assign the issue to the test file's
author. Include the triggering PR URL in the issue body.
gh issue create \
--repo NVIDIA/Megatron-LM \
--title "🐛 CI failure: <failed-test-node-id>" \
--label "bug" \
--assignee "<test-author-login>" \
--body "..."Use the bug-report template body structure:
**Describe the bug**
CI test `<failed-test-node-id>` failed in job [`<job-name>`](<job-url>).
Tag @NVIDIA/mcore-oncall to get oncall's attention to this issue.
**Failing run**
| Field | Value |
|-------|-------|
| PR | [#<pr_number>: <pr_title>](<pr_url>) |
| Run | [<run_id>](<run_url>) |
| Job | [<job_name>](<job_url>) |
**Error**
<core error message / traceback — 30 lines max>
**Steps/Code to reproduce bug**
Re-run the failing CI job linked above, or locally inside the dev container:
```bash
pytest <failed-test-node-id>Additional context
Triaged automatically via /triage-issue.
If multiple tests failed in the same job, list each one as a separate bullet
under "Describe the bug" and include the combined error snippets. Assign the
issue to the author of whichever test file appears first in the failure list.
### 8. Report back to the user
Print the URL of the newly created issue (or the duplicate, if found) so the
user can review or share it.
## Important guidelines
- Never create an issue if a duplicate already exists — link the existing one instead.
- Always include the triggering PR link in the issue body.
- Always assign the issue to the test file's most recent author. If the author
lookup fails (e.g. the commit was made by a bot or the login is unavailable),
skip `--assignee` and note it in the "Additional context" section.
- Keep the error snippet concise (≤30 lines). Truncate long tracebacks and note that the full log is available via the job URL.
- Do not guess the root cause — quote the actual log output verbatim.
- If the job is still in progress or the logs are unavailable, say so and ask the user to retry once the run completes.© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files in skills/mcore-create-issue of NVIDIA/Megatron-LM.
Open the folder on GitHubat commit 486a126
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in NVIDIA/Megatron-LM, which our catalogue first saw on October 7, 2026.
Megatron-LM CI Failure Triage next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Megatron-LM CI Failure Triage this skillNVIDIA/Megatron-LM | 18k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| GreptimeDB Fuzz CI Failure InvestigationGreptimeTeam/greptimedb | 6.7k | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | |
| CI Failure Triage and RepairChachamaru127/claude-code-harness | 3.2k | 1 repos | ~1.1k | Automated safety check: Notes | MIT | |
| CI Watchdoglatitude-dev/latitude-llm | 4.7k | — | ~1.6k | Automated safety check: Pass | MIT | |
| Parallel CI Triagespencerpauly/awesome-cursor-skills | 841 | — | ~793 | Automated safety check: Pass | CC0-1.0 | |
| Gh Fix CIJetBrains/skills | 363 | 7 repos | ~953 | Automated safety check: Pass | Apache-2.0 |
GreptimeTeam/greptimedb
Diagnoses a failed GreptimeDB fuzz CI job by pulling its GitHub Actions logs and fuzz artifacts, then matching the evidence to the local source code.
Chachamaru127/claude-code-harness
Diagnoses failing CI pipelines and tests, deciding first whether the test or the implementation is at fault, and hands hard cases to a dedicated fixer subagent.
latitude-dev/latitude-llm
Continuously monitor GitHub PR CI checks and automatically fix failures until all checks pass.
spencerpauly/awesome-cursor-skills
When GitHub Actions fails, fetch failing job logs and assign each failing job to a separate subagent that fixes its slice of the problem in parallel.
JetBrains/skills
A skill your agent uses when a user asks to debug or fix failing GitHub PR checks that run in GitHub Actions; use gh to inspect checks and logs, summarize failure context, draft a fix plan, and…
trailofbits/skills-curated
A skill your agent uses when a user asks to debug or fix failing GitHub PR checks that run in GitHub Actions; use gh to inspect checks and logs, summarize failure context, draft a fix plan, and…
NVIDIA/Megatron-LM
Refreshes stored golden values from a GitHub Actions run, reports signed percentage changes per model, and writes a summary ready for a pull request description.
NVIDIA/Megatron-LM
Walks an agent through working inside the Megatron-LM CI container and changing dependencies with uv, so lock files resolve the same locally and in CI.
NVIDIA/Megatron-LM
Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up.
NVIDIA/Megatron-LM
Explains Megatron-LM's CI pipeline, PR scope labels, triggering the internal GitLab CI with a dry run first, and investigating CI failures.
NVIDIA/Megatron-LM
Guides moving Megatron Core GPTModel checkpoints, configs, training commands and launch scripts to HybridModel, following the repository's migration document.
NVIDIA/Megatron-LM
Creates one-node GitHub merge-request variants of existing two-node GB200 functional tests in Megatron-LM, adjusting parallelism settings to fit four GPUs.
Works with
Categories
Investigates a failing GitHub Actions run or job for Megatron-LM, finds the root cause plus the PR and test author involved, and files a structured bug issue. You pass a GitHub Actions job or run URL. The agent pulls out the run and job ids, lists the failed jobs when only a run was given (asking which to triage if several failed), and fetches the logs with the gh CLI, narrowing them to error lines such as FAILED, Traceback and assertion failures.
Megatron-LM CI Failure Triage fits situations like: A Megatron-LM CI job failed and needs a bug issue; tracing a failed Actions run back to its pull request and test author; triaging several failed jobs from one workflow run.
Run `npx skills add NVIDIA/Megatron-LM --skill mcore-create-issue -a claude-code`. Or copy the skill folder (skills/mcore-create-issue in NVIDIA/Megatron-LM) into .claude/skills/mcore-create-issue in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/Megatron-LM --skill mcore-create-issue -a codex`. Or copy the skill folder (skills/mcore-create-issue in NVIDIA/Megatron-LM) into .agents/skills/mcore-create-issue in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/Megatron-LM --skill mcore-create-issue -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/mcore-create-issue, .gemini/skills/mcore-create-issue, .github/skills/mcore-create-issue and .opencode/skills/mcore-create-issue in your project.
Going by SKILL.md and its folder, Megatron-LM CI Failure Triage needs the command-line tools its instructions call (gh and pytest). Our summary lists: The gh CLI authenticated with access to NVIDIA/Megatron-LM.
SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Megatron-LM CI Failure Triage is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Megatron-LM CI Failure Triage: GreptimeDB Fuzz CI Failure Investigation (GreptimeTeam/greptimedb, 6.7k stars), CI Failure Triage and Repair (Chachamaru127/claude-code-harness, 3.2k stars), CI Watchdog (latitude-dev/latitude-llm, 4.7k stars) and Parallel CI Triage (spencerpauly/awesome-cursor-skills, 841 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/Megatron-LM, which has 18,078 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 7, 2026.
Source: NVIDIA/Megatron-LM on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.