Megatron-LM Base Image Bump
NVIDIA/Megatron-LM
Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up.
Monitor the current PR's GitLab pipeline to completion, then report success, auto-fix, or investigate a failure.
$ npx skills add DataDog/datadog-agent --skill follow-pr -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install DataDog/datadog-agent follow-pr --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/DataDog/datadog-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/follow-pr .claude/skills/follow-pr && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "follow-pr" agent skill from https://github.com/DataDog/datadog-agent/tree/main/.agents/skills/follow-pr into .claude/skills/follow-pr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "follow-pr", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/DataDog/datadog-agent/tree/main/.agents/skills/follow-prType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add DataDog/datadog-agent --skill follow-pr -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install DataDog/datadog-agent follow-pr --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/DataDog/datadog-agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/follow-pr .agents/skills/follow-pr && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "follow-pr" agent skill from https://github.com/DataDog/datadog-agent/tree/main/.agents/skills/follow-pr into .agents/skills/follow-pr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "follow-pr", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add DataDog/datadog-agent --skill follow-pr -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install DataDog/datadog-agent follow-pr --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/DataDog/datadog-agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/follow-pr .cursor/skills/follow-pr && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "follow-pr" agent skill from https://github.com/DataDog/datadog-agent/tree/main/.agents/skills/follow-pr into .cursor/skills/follow-pr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "follow-pr", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/DataDog/datadog-agent.git --path .agents/skills/follow-pr--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add DataDog/datadog-agent --skill follow-pr -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install DataDog/datadog-agent follow-pr --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/DataDog/datadog-agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/follow-pr .gemini/skills/follow-pr && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "follow-pr" agent skill from https://github.com/DataDog/datadog-agent/tree/main/.agents/skills/follow-pr into .gemini/skills/follow-pr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "follow-pr", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install DataDog/datadog-agent follow-prInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add DataDog/datadog-agent --skill follow-pr -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/DataDog/datadog-agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/follow-pr .github/skills/follow-pr && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "follow-pr" agent skill from https://github.com/DataDog/datadog-agent/tree/main/.agents/skills/follow-pr into .github/skills/follow-pr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "follow-pr", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add DataDog/datadog-agent --skill follow-pr -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install DataDog/datadog-agent follow-pr --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/DataDog/datadog-agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/follow-pr .opencode/skills/follow-pr && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "follow-pr" agent skill from https://github.com/DataDog/datadog-agent/tree/main/.agents/skills/follow-pr into .opencode/skills/follow-pr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "follow-pr", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
follow-prMonitor the current PR's GitLab pipeline to completion, then report success, auto-fix, or investigate a failure.
Follow PR is an agent skill from DataDog/datadog-agent, published by the product's own GitHub organization. Monitor the current PR's GitLab pipeline to completion, then report success, auto-fix, or investigate a failure. Use when the user asks to follow, babysit, watch, or wait on a PR/pipeline, or just after pushing to / creating a PR.
Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `scripts/config.py` and `scripts/create_devenv.sh`).
It sits in DevOps & Cloud. It works with GitLab and Datadog. The repository describes itself as: Main repository for Datadog Agent. The licence is Apache-2.0.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit a706f1a. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python and Shell), which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Follow PR loads about 3.2k tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 1,660 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from DataDog/datadog-agent at commit a706f1a, republished under its Apache-2.0 licence (© DataDog). 1,660 words, ~3,177 tokens.
.claude/skills/follow-pr/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Watch the latest Gitlab CI pipeline for the current PR to a terminal state, report the outcome, and — for failures reliably caused by this PR — autonomously fix and push bounded ones under the resolved autonomy policy while investigating the rest locally.
Owning team: @DataDog/agent-devx
The appropriate tool for this usecase is ddgl, and more specifically ddgl attach.
Check if ddgl is available - which ddgl. If so, move to Step 1. Otherwise, use a dev env as specified below.
First, check if you are running in a dev env: test -f /.started will exit 0 if so. If you are in an outdated devenv without ddgl, stop and notify the user to recreate his dev env.
Otherwise, check for the existence of a dev env by using dda env dev show.
If there are existing dev envs:
repos and extra_(mount|volume)_specs fields)If the environment is already started and contains the right repo, move to the next step.
Otherwise, create one by using ./scripts/create_devenv.sh, then use the environment ID printed by the script in subsequent commands.
To run commands inside a dev env, use the following template:
dda env dev run --id <dev-env-id> -- [command]Watch out for space-splitting. For example:
dda env dev run --id follow-pr-attach-7C2C42F6 -- ddgl attach --detail=normal --follow --plainIf the user gave a ref, branch, or pipeline ID, pass it through (--ref <ref> or --pipeline <id>).
Otherwise omit both — ddgl attach resolves the pipeline for the current branch on its own.
Before starting the monitoring loop, resolve how this run should handle failures caused by the PR's own code:
python3 .agents/skills/follow-pr/scripts/config.py resolve \
[--mode <from --fix-mode arg>] [--max-fix-cycles <from arg>] [--policy <from arg>]Pass through any --fix-mode, --max-fix-cycles, or policy text given in this invocation; otherwise the script falls back to environment variables, then worktree-local config, then global config, then its own default (autofix).
If the config ended up resolving using only the default values, suggest the user create a global/local config file with the settings he just chose for future invocations. Check the script contents to get the appropriate paths and formats.
autofix or no-autofix: report the resolved mode, cycle budget, and whether a custom policy is active, then continue to Step 3.ask, or the script errors: ask the user directly, before monitoring starts, whether PR-caused failures this run should be fixed and pushed (autofix) or only investigated locally (no-autofix). Offer to persist the answer (worktree-local or global config) if they don't want to be asked again; otherwise use it for this run only.Keep the resolved mode, cycle budget (default 2), and policy text in context — you'll pass them straight through as --mode/--max-fix-cycles/--policy to subskills that might need it.
All pipeline discovery, polling, follow/rebind, and timeout handling is covered by the internals of ddgl attach.
Do not implement a second polling loop or persist monitoring state of your own.
Check whether you have a long-lived monitoring tool available, one that can run a command in the background and forward each stdout line as it arrives, without a timeout of its own (e.g. Claude Code's Monitor tool).
With such a tool: start it on
ddgl attach --plain --follow --detail=full [--ref <ref> | --pipeline <id>]and wait for a [FINAL] line — no --timeout needed.
Without one: run it in the foreground, bounded so the invocation cannot outlive your own harness timeout:
ddgl attach --plain --follow --detail=full --timeout 600 [--ref <ref> | --pipeline <id>]If the [FINAL] line reports a timeout (not a pipeline outcome), start an identical invocation again.
This is safe: attach is stateless and each invocation begins with a fresh snapshot of the pipeline.
NOTE: If the pipeline is already terminal or does not exist when you start monitoring, the user might have just pushed and the pipeline is still waiting to be created. In this case, wait for 60 seconds and then re-attempt monitoring. The
--followargument will make sureddgl attachalways monitors the latest pipeline for the ref.
You may see:
[POLL] - rollup summary after a changed poll tick (jobs done/total, stage,
failure count). Informational only.[INFO] - an informational log from ddgl itself.[PIPE] - a change in the pipeline status.[JOB] - a job finished running and changed state.[FINAL] - the terminal, authoritative outcome. Treat this line as the
source of truth regardless of the command's exit code — it names the pipeline id, terminal status, and, on failure, the failed job names.datadog-agent CI retry once automatically on any failure (.gitlab-ci.yml's
default retry: max: 1, when: always), so a first-attempt failure alone isn't yet evidence
of anything — GitLab will retry it once. A minority of jobs (most e2e, Windows, macOS —
.retry_only_infra_failure) retry only on GitLab's infra-flavoured failure_reason values,
so a script_failure there gets no automatic retry at all and is worth a closer look sooner.
Unit test, linter, and build failures are less likely to be flakes regardless of policy.
If you're unsure which policy a job is on: grep -rn '<job-name>' .gitlab/ .gitlab-ci.yml.
Otherwise, ask the user whether to continue monitoring, or if this job failure is already a
problem. In the latter case, move to Step 6.[FINAL]: re-invoke ddgl attach as in Step 3; this is not a true terminal outcome.ddgl itself, or from the monitoring tool): report what happened. Do not attempt a recovery action.Invoke /triage-ci-failure on the pipeline id from the [FINAL] line; it returns one CI triage result block per failed job.
Route each block by its Blame field — never by whether Incident happens to be none, since that value alone doesn't tell you the job wasn't PR-caused:
upstream: if Incident is active and still breaking, continue to Step 7 and wait it out. If stable or resolved, tell the user it's safe to rebase onto main and re-run — say plainly that stable is a weaker signal than resolved (the fix may still be in progress). If no incident is declared at all, say CI looks broken on main with nothing declared for it — worth surfacing loudly. Investigation ends here for that job.
infra or flake: report the verdict and its suggested action (typically a retry, citing the evidence /triage-ci-failure gave you). Investigation ends here for that job.
inconclusive: report the evidence and the two most likely readings. Investigation ends here for that job.
pr-code: collect every pr-code block from this pipeline. Before invoking the handler, check for repeats against every push made earlier in this same run:
Failure signature matches one from an earlier push, that fix didn't work — tell /handle-pr-ci-failure explicitly that this failure already survived one autonomous attempt, so it treats that root cause as blocked rather than reattempting it as safe./handle-pr-ci-failure it has no push budget left this run, so it should investigate every root cause without committing or pushing, same as no-autofix.Either way, invoke /handle-pr-ci-failure once with all of this pipeline's pr-code blocks together, passing --mode/--max-fix-cycles/--policy set to the values resolved in Step 2. Continue to Step 8 with its result.
Only entered when /triage-ci-failure reported an incident that's still active and breaking for a failed job — this is the other half of watching a PR through:
the pipeline is red because of something outside the PR, and it will stay red until that something changes.
Poll the incident on an interval (a few minutes is reasonable; don't busy-loop):
.agents/skills/triage-ci-failure/scripts/incidents.py timeline <IR-nnnnn>Watch for a state transition off active — to stable (a rollback or workaround has likely landed; rebasing is probably safe even if the root cause isn't fully fixed yet) or resolved/completed (the stronger signal).
Once either happens, confirm recovery before telling the user to act — check that the job is passing again on main:
pup cicd events aggregate \
--query='ci_level:job @ci.pipeline.name:DataDog/datadog-agent @git.branch:main @ci.job.name:"<job name>"' \
--compute=count --group-by='@ci.status' --from='2h'Once main is clean, tell the user it's time to rebase onto main and re-run.
Read /handle-pr-ci-failure's result block:
Outcome: pushed: record its Failure signatures and increment the cycle count — Step 6 needs both for the repeat/budget checks on the next pipeline. Go back to Step 3 to watch the replacement pipeline at the Pushed SHA regardless of whether the budget is now exhausted; you still need to confirm this fix actually worked before you can stop. Return here through Step 6 once it finishes.Outcome: committed-not-pushed: report the local commit and the remaining complex root cause(s) blocking a push using the final report template, then stop and let the user decide.Outcome: needs-user or blocked: report the evidence and the specific question /handle-pr-ci-failure asked for using the final report template, then stop.Every trip back through this loop re-runs /triage-ci-failure from scratch on the new pipeline — never reuse an earlier verdict for a different pipeline.
Whenever you stop for good — pipeline succeeded, an upstream/infra/flake verdict ended investigation, or /handle-pr-ci-failure returned anything other than a fresh pushed, close with this block instead of ad hoc prose:
Follow-PR report
Pipeline: <final pipeline id> (<SHA>)
Status: succeeded | failed | needs-user
Fixes pushed: <count> — <commit SHAs, or none>
Cycles consumed: <N>/<max_fix_cycles>
Remaining issues: <summary, or none>
User decision needed: <question, or none>
End follow-PR report/triage-ci-failure returns one pr-code block; /handle-pr-ci-failure classifies it safe, fixes it, verifies with dda inv linter.go, commits, and pushes. Step 8 sees Outcome: pushed (cycle 1 of 2), goes back to Step 3, and the replacement pipeline goes green.-race. /handle-pr-ci-failure classifies it complex, reproduces it locally, tries two distinct hypotheses, and stops with Outcome: needs-user and an uncommitted candidate diff — nothing is pushed.Failure signature. Step 6 recognizes the repeat before invoking the handler, tells it not to reattempt a fix, and it investigates instead — no second push happens.© DataDog, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (scripts) in .agents/skills/follow-pr of DataDog/datadog-agent.
Open the folder on GitHubat commit a706f1a
Follow PR next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Follow PR this skillDataDog/datadog-agent | 3.8k | — | ~3.2k | Automated safety check: Pass | Apache-2.0 | |
| Megatron-LM Base Image BumpNVIDIA/Megatron-LM | 18k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Megatron-LM CI/CD GuideNVIDIA/Megatron-LM | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| Setup Osworldxlang-ai/OSWorld-V2 | 355 | 1 repos | ~2.6k | Automated safety check: Notes | Apache-2.0 | |
| Apm IntegrationsDataDog/dd-trace-js | 837 | — | ~3k | Automated safety check: Pass | Custom licence | |
| Dd IdpDataDog/pup | 1k | — | ~2k | Automated safety check: Pass | Apache-2.0 |
NVIDIA/Megatron-LM
Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up.
NVIDIA/Megatron-LM
Explains Megatron-LM's CI pipeline, PR scope labels, triggering the internal GitLab CI with a dry run first, and investigating CI failures.
xlang-ai/OSWorld-V2
Provision and verify an OSWorld-V2 checkout after clone. An agent skill from xlang-ai/OSWorld-V2.
DataDog/dd-trace-js
A skill your agent uses when adding, debugging, fixing, or modifying instrumentation and plugins for third-party libraries in dd-trace-js.
DataDog/pup
Find, filter, count, and connect software, teams, engineering work and delivery, infrastructure, and operational or security records through Pup's read-only Datadog entity graph.
DataDog/terraform-provider-datadog
Generates a Datadog Terraform provider data source from an OpenAPI operation with tfgen and opens a review-ready GitHub PR with a risk scan and testing guide.
DataDog/datadog-agent
Classify a failed CI as either caused by an active incident, flakiness, or a true code regression.
DataDog/datadog-agent
Run a structured discovery session to build an Allium specification through conversation.
DataDog/datadog-agent
A skill your agent uses when an engineer or manager asks to recap, summarize, or post an update on a Jira Epic — a progress update for an in-progress Epic (how far along it is, what's shipped so…
DataDog/datadog-agent
Explains a lading.yaml config file from the regression test suite, using the lading Rust source as ground truth for field meanings and defaults.
DataDog/datadog-agent
Extract an Allium specification from an existing codebase. An agent skill from DataDog/datadog-agent.
DataDog/datadog-agent
Run one already-written new-e2e test locally and triage the setup failures that stop it — "run the containers e2e tests", "my e2e run fails before any test starts".
Categories
Monitor the current PR's GitLab pipeline to completion, then report success, auto-fix, or investigate a failure. Follow PR is an agent skill from DataDog/datadog-agent, published by the product's own GitHub organization. Monitor the current PR's GitLab pipeline to completion, then report success, auto-fix, or investigate a failure.
Follow PR fits situations like: the user asks to follow; wait on a PR/pipeline; just after pushing to / creating a PR.
Run `npx skills add DataDog/datadog-agent --skill follow-pr -a claude-code`. Or copy the skill folder (.agents/skills/follow-pr in DataDog/datadog-agent) into .claude/skills/follow-pr in your project. Claude Code loads it when a task matches its description.
Run `npx skills add DataDog/datadog-agent --skill follow-pr -a codex`. Or copy the skill folder (.agents/skills/follow-pr in DataDog/datadog-agent) into .agents/skills/follow-pr in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add DataDog/datadog-agent --skill follow-pr -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/follow-pr, .gemini/skills/follow-pr, .github/skills/follow-pr and .opencode/skills/follow-pr in your project.
Going by SKILL.md and its folder, Follow PR needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3; A Bash shell.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Follow PR is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Follow PR: Megatron-LM Base Image Bump (NVIDIA/Megatron-LM, 18k stars), Megatron-LM CI/CD Guide (NVIDIA/Megatron-LM, 18k stars), Setup Osworld (xlang-ai/OSWorld-V2, 355 stars) and Apm Integrations (DataDog/dd-trace-js, 837 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
DataDog (a GitHub organization, an official publisher) maintains it in DataDog/datadog-agent, which has 3,759 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 9, 2026.
Source: DataDog/datadog-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.