Search
NVIDIA/Megatron-LM
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Walks an agent through working inside the Megatron-LM CI container and changing dependencies with uv, so lock files resolve the same locally and in CI. | NVIDIA/ | 18k | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up. | NVIDIA/ | 18k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | today |
| 3 | Explains Megatron-LM's CI pipeline, PR scope labels, triggering the internal GitLab CI with a dry run first, and investigating CI failures. | NVIDIA/ | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 4 | Investigates a failing GitHub Actions run or job for Megatron-LM, finds the root cause plus the PR and test author involved, and files a structured bug issue. | NVIDIA/ | 18k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 5 | Guides moving Megatron Core GPTModel checkpoints, configs, training commands and launch scripts to HybridModel, following the repository's migration document. | NVIDIA/ | 18k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 6 | Creates one-node GitHub merge-request variants of existing two-node GB200 functional tests in Megatron-LM, adjusting parallelism settings to fit four GPUs. | NVIDIA/ | 18k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 7 | Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis. | NVIDIA/ | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 8 | Guide to the Megatron-LM test system: layout, recipe YAML, running and adding unit and functional tests, golden values, marker filters and CI parity. | NVIDIA/ | 18k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | today |
| 9 | Refreshes stored golden values from a GitHub Actions run, reports signed percentage changes per model, and writes a summary ready for a pull request description. | NVIDIA/ | 18k | — | ~2.8k | Automated safety check: Pass | Unknown | today |
| 10 | Runs the Megatron-LM autoformat script and its linting tools before a pull request, and keeps Python imports in order with isort. | NVIDIA/ | 18k | — | ~316 | Automated safety check: Pass | Apache-2.0 | today |
| 11 | Domain knowledge for the nightly workflow that merges Megatron-LM's main branch into dev, covering conflict handling, CI iteration, failure investigation and known issues. | NVIDIA/ | 18k | — | ~8.5k | Automated safety check: Pass | Unknown | today |
| 12 | Split a PR into multiple PRs to reduce the number of required CODEOWNERS reviewer groups. | NVIDIA/ | 18k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 13 | Research and draft a response to a GitHub issue or question from an external contributor. | NVIDIA/ | 18k | — | ~915 | Automated safety check: Pass | Unknown | today |
| 14 | Review rubric for the /review pull-request command. An agent skill from NVIDIA/Megatron-LM. | NVIDIA/ | 18k | — | ~839 | Automated safety check: Pass | Apache-2.0 | today |