Hugging Face Local Model Evals
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
Agent skill
by jeremylongshore in jeremylongshore/tons-of-skills-marketplace
Diagnose the most expensive silent failure on a CoreWeave multi-node GPU job: GPUDirect RDMA falling back from InfiniBand to TCP.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-fabric-diagnostics -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace coreweave-fabric-diagnostics --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/coreweave-fabric-diagnostics .claude/skills/coreweave-fabric-diagnostics && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "coreweave-fabric-diagnostics" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/coreweave-fabric-diagnostics into .claude/skills/coreweave-fabric-diagnostics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "coreweave-fabric-diagnostics", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/coreweave-fabric-diagnosticsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-fabric-diagnostics -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace coreweave-fabric-diagnostics --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/.curated/coreweave-fabric-diagnostics .agents/skills/coreweave-fabric-diagnostics && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "coreweave-fabric-diagnostics" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/coreweave-fabric-diagnostics into .agents/skills/coreweave-fabric-diagnostics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "coreweave-fabric-diagnostics", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-fabric-diagnostics -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace coreweave-fabric-diagnostics --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/.curated/coreweave-fabric-diagnostics .cursor/skills/coreweave-fabric-diagnostics && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "coreweave-fabric-diagnostics" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/coreweave-fabric-diagnostics into .cursor/skills/coreweave-fabric-diagnostics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "coreweave-fabric-diagnostics", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jeremylongshore/tons-of-skills-marketplace.git --path skills/.curated/coreweave-fabric-diagnostics--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-fabric-diagnostics -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace coreweave-fabric-diagnostics --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/.curated/coreweave-fabric-diagnostics .gemini/skills/coreweave-fabric-diagnostics && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "coreweave-fabric-diagnostics" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/coreweave-fabric-diagnostics into .gemini/skills/coreweave-fabric-diagnostics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "coreweave-fabric-diagnostics", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jeremylongshore/tons-of-skills-marketplace coreweave-fabric-diagnosticsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-fabric-diagnostics -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/.curated/coreweave-fabric-diagnostics .github/skills/coreweave-fabric-diagnostics && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "coreweave-fabric-diagnostics" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/coreweave-fabric-diagnostics into .github/skills/coreweave-fabric-diagnostics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "coreweave-fabric-diagnostics", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-fabric-diagnostics -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace coreweave-fabric-diagnostics --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/.curated/coreweave-fabric-diagnostics .opencode/skills/coreweave-fabric-diagnostics && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "coreweave-fabric-diagnostics" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/coreweave-fabric-diagnostics into .opencode/skills/coreweave-fabric-diagnostics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "coreweave-fabric-diagnostics", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
coreweave-fabric-diagnosticsDiagnose the most expensive silent failure on a CoreWeave multi-node GPU job: GPUDirect RDMA falling back from InfiniBand to TCP.
Coreweave Fabric Diagnostics is an agent skill from jeremylongshore/tons-of-skills-marketplace. Diagnose the most expensive silent failure on a CoreWeave multi-node GPU job: GPUDirect RDMA falling back from InfiniBand to TCP. When NCCL drops from NET/IB to NET/Socket, collectives keep running with NO error but throughput collapses (commonly 5-20x slower) while every GPU still bills at full rate — 5x the GPU bill for the same work, invisibly. Paste an NCCLDEBUG=INFO log (and/or a pod-spec, ibstat, or allreduceperf output) and the bundled deterministic script verdicts whether RDMA is actually engaged, which…
Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `ARD.md`, `PRD.md` and `eval-spec.yaml`). Compatibility notes: Designed for Claude Code
It sits in AI & LLM Engineering, covering GPU and accelerator computing. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadWriteEditGlobBash(kubectl get:*)Bash(python3:*)From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
kubectlpython3From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.nvidia.comdocs.coreweave.comgithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Designed for Claude Code
From compatibility in the SKILL.md frontmatter.
Coreweave Fabric Diagnostics loads about 3.4k tokens when it runs, and up to ~6.3k if it reads all its reference files. Until then it costs about 225 tokens; SKILL.md has 1,457 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
sudo systemctl stop nvidia-fabricmanagersudo nvidia-smi -r # GPU resetsudo systemctl start nvidia-fabricmanagerAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 1,457 words, ~3,364 tokens.
.claude/skills/coreweave-fabric-diagnostics/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.Community-contributed. Not affiliated with, endorsed by, or sponsored by CoreWeave, Inc. CoreWeave is a registered trademark of CoreWeave, Inc.
Detects when a CoreWeave multi-node GPU job has silently fallen off the InfiniBand fabric onto TCP — the failure that makes distributed training run at a fraction of the hardware's speed while every GPU keeps billing at the full rate — and gives the exact fix.
On a CoreWeave multi-node job, NCCL should carry collectives over InfiniBand with
GPUDirect RDMA (NET/IB). If any one of three conditions is missing, NCCL silently
falls back to TCP sockets (NET/Socket): the job still runs, still converges, and
raises no error — but all-reduce throughput collapses (commonly cited as 5-20x
slower [nccl]) because it now crosses the Ethernet control plane instead of the
400 Gb/s-class fabric. You keep paying full GPU rate for a multi-node run that performs
like a badly-connected one. This is the single highest-dollar invisible failure on the
platform, and nothing in the default output flags it.
The diagnosis is deterministic: the bundled scripts/fabric-check.py greps the pasted
NCCL_DEBUG=INFO log for the decisive Using network line (and NET/IB vs NET/Socket),
parses the pod-spec's resources block for the RDMA device request, reads ibstat port
state, and echoes any all_reduce_perf bus bandwidth — then emits a VERDICT with the
rdma_engaged/transport call, the missing conditions, and the fix. The LLM never
eyeballs which transport is in use; the script decides. Deep grounding lives in
references/, loaded only when a leg of the diagnosis needs it.
NCCL_DEBUG=INFO log from the actual run — the primary signal. Re-run the job
(or one rank) with NCCL_DEBUG=INFO set and capture stderr. The decisive line is
Using network IB (good) vs Using network Socket (the fallback). This is the one input
the skill really needs; everything else corroborates.kubectl get pod NAME -o yaml) to
check the RDMA device request; ibstat output from the node for port health; and
all_reduce_perf results from CoreWeave's nccl-tests to measure bus bandwidth.python3 to run the deterministic checker (stdlib only).kubectl (read-only) if corroborating the live pod spec / node cordon state.Authentication. Nothing secret is read. If the pod spec is pulled live, kubectl uses
the existing $KUBECONFIG; the skill only ever runs kubectl get (read-only) — it never
cordons, drains, or applies.
resources.requests AND resources.limits
(rdma/ib: 1). If it is in only one — or absent — the device plugin does not inject the
IB device into the pod and NCCL never sees a HCA. [unverified — the exact resource key
(e.g. rdma/ib) depends on the installed RDMA device-plugin config; confirm with
kubectl describe node / kubectl get node -o yaml.]NCCL_IB_HCA=ibp and NCCL_SOCKET_IFNAME=eth0 are set (CoreWeave's documented
values [cw]) — unless you launch via the MPI Operator, which manages this network
config for you [nt].NCCL_DEBUG=INFO then confirms NET/IB (ideally a GPU Direct RDMA Enabled line).
If it shows NET/Socket / Using network Socket, RDMA is not engaged.Full checklist with verification commands: references/rdma-engagement-checklist.md.
The pipeline is gather → verdict → fix → confirm. The script does the transport call;
references/ carry the grounding:
NCCL_DEBUG=INFO log (required) plus any pod-spec / ibstat / all_reduce_perf
output you have. Concatenate them into one paste — the checker keys on each signal
independently.The log is the load-bearing input. If the user has not run with NCCL_DEBUG=INFO, tell
them to — without it, transport selection is unknowable. To pull the live pod spec:
kubectl get pod "$POD" -o yaml > pod.yamlPipe everything you gathered to fabric-check.py. It greps for the decisive Using network line, the resources block, ibstat state, and any Avg bus bandwidth:
cat nccl-debug.log pod.yaml ibstat.txt allreduce.txt 2>/dev/null | \
python3 scripts/fabric-check.pyThe verdict names rdma_engaged (yes/no/partial/unknown), the transport in use, the
missing conditions, and the fix. Use --json to capture the structured result for further
processing. Reading the log by eye is what this step exists to prevent — see
references/nccl-debug-reading.md for what each line
means.
Use Glob to gather multiple pasted log files when a run spans several ranks, Write the
verdict report to the working directory, and Edit it to refine the fix as the user
iterates on the manifest.
This is the money case. Fix in order (the checker prints the same list):
rdma/ib: 1 to both resources.requests and resources.limits.NCCL_IB_HCA=ibp and NCCL_SOCKET_IFNAME=eth0 (or launch via the MPI Operator).NCCL_DEBUG=INFO and confirm the log now shows NET/IB +
GPU Direct RDMA Enabled, not NET/Socket.If the log shows NCCL_IB_DISABLE=1, that alone forces sockets — set it to 0 (RoCE and
IB both need the IB verbs transport enabled [env]).
RDMA can be engaged yet slow. Two corroborating checks:
ibstat — every port must read State: Active / Physical state: LinkUp. A port
Down/Polling, or a link that flaps, drags the whole collective; CoreWeave
auto-cordons flapping links, so a shrinking node count mid-run is a fabric symptom.all_reduce_perf bus bandwidth — compare the reported busbw against CoreWeave's
published nccl-tests manifest baseline for your GPU count + NCCL version [nt]. Do
not compare against a fixed number: the baseline moves with GPU type, node count, NCCL
version, and SHARP. The checker echoes the observed figure tagged [unverified vs baseline] precisely so nobody reads it as a hard pass/fail.Details + the busbw-vs-algbw distinction: references/allreduce-baseline.md.
On NVSwitch/NVLink systems, a wedged fabric shows up as NVLink/NVSwitch errors rather than IB fallback. The safe reset order is stop Fabric Manager → reset the GPUs → start Fabric Manager, never the reverse:
sudo systemctl stop nvidia-fabricmanager
sudo nvidia-smi -r # GPU reset
sudo systemctl start nvidia-fabricmanager[unverified — service unit name and reset support vary by image/driver; on managed CoreWeave nodes prefer opening a support ticket / cordoning over an in-place reset.]
rdma/ib-in-requests-AND-limits change, the env vars, and the
re-verify step.busbw (tagged [unverified vs baseline]).| Error | Cause | Solution |
|---|---|---|
Verdict is unknown | No NET/IB / NET/Socket / Using network line in the paste | Re-run the job with NCCL_DEBUG=INFO and capture stderr; without it transport is unknowable. |
Verdict Socket but the pod "has RDMA" | rdma/ib in limits only (or only requests) | Add it to BOTH blocks; the device plugin injects the IB device only when the resource is requested. |
NET/IB present yet training still slow | GDR not actually enabled; nvidia-peermem unloaded → traffic stages through host memory | Confirm a GPU Direct RDMA Enabled line; verify nvidia-peermem is loaded on the node [nccl]. |
busbw "looks low" | Compared against a wrong/guessed baseline | Compare only against CoreWeave's nccl-tests manifest baseline for your GPU count + NCCL version; the number is workload/version-dependent. |
| Nodes drop out mid-run | Flapping IB link → CoreWeave auto-cordon | Check ibstat for Physical state != LinkUp; the cordoned node's link is the cause, not your job. |
rdma/ib resource not schedulable | Wrong resource key for the installed device plugin | Confirm the exact key with kubectl describe node (search the Allocatable list) and substitute it. |
The user pastes an NCCL_DEBUG=INFO excerpt plus the pod spec. The checker finds Using network Socket and rdma/ib only in requests, and verdicts:
### VERDICT: RDMA is NOT engaged -- NCCL fell back to TCP (NET/Socket). Multi-node collectives are running over the Ethernet control plane, commonly 5-20x slower for the same GPU-hours -- you pay full GPU rate for a fraction of the throughput, and NCCL raised no error.
- RDMA engaged: **no**
- Transport in use: **Socket**
- Missing conditions (each one alone forces a silent TCP fallback):
- `rdma/ib` missing from resources.limits
- `NCCL_IB_HCA` not set (e.g. `ibp`) -- unless the MPI Operator manages it
**The fix (in order):**
1. Request the RDMA device in BOTH requests AND limits: `rdma/ib: 1` (if it is in only one, the device plugin will not inject the IB device).
2. Set `NCCL_IB_HCA=ibp` and `NCCL_SOCKET_IFNAME=eth0` (CoreWeave values), or let the MPI Operator manage them.
3. Re-run with `NCCL_DEBUG=INFO` and confirm the log now shows `NET/IB` and `GPU Direct RDMA Enabled` -- not `NET/Socket` / `Using network Socket`.
4. Confirm each IB port is `State: Active` / `Physical state: LinkUp` via `ibstat`; a flapping link gets auto-cordoned by CoreWeave.The log shows NET/IB and GPU Direct RDMA Enabled, so the checker returns
rdma_engaged: yes. It then surfaces the ibstat port that reads Physical state: Polling as a degraded signal and echoes the observed busbw tagged [unverified vs baseline], directing the user to compare against CoreWeave's nccl-tests manifest for their
GPU count + NCCL version rather than a guessed number.
references/rdma-engagement-checklist.md — the three required conditions + how to verify each, cited.references/nccl-debug-reading.md — reading NCCL_DEBUG=INFO: NET/IB vs NET/Socket, the decisive Using network line, GDR.references/allreduce-baseline.md — all_reduce_perf busbw/algbw and why the baseline is never hardcoded.coreweave-gpu-cost-leak-hunter dollarizes idle/right-sizing spend; this skill finds the throughput leak (fabric fallback) that a cost report cannot see.© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 7 other files (scripts, references) in skills/.curated/coreweave-fabric-diagnostics of jeremylongshore/tons-of-skills-marketplace.
Open the folder on GitHubat commit cfae287
Coreweave Fabric Diagnostics next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Coreweave Fabric Diagnostics this skilljeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~3.4k | Automated safety check: Notes | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Liger Kernel Perflinkedin/Liger-Kernel | 6.7k | — | ~1.5k | Automated safety check: Pass | BSD-2-Clause | |
| Hugging Face LLM Trainerhuggingface/skills | 11k | 1 repos | ~7.2k | Automated safety check: Pass | Apache-2.0 | |
| MUSA GPU Training Optimizeropen-infra-skills/infra-skills | 141 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Cuda Kernel OptimizerKernelFlow-ops/cuda-optimized-skill | 214 | — | ~4.3k | Automated safety check: Pass | MIT |
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
linkedin/Liger-Kernel
Optimizes the performance of existing Liger Kernel Triton kernels.
huggingface/skills
Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.
open-infra-skills/infra-skills
Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.
KernelFlow-ops/cuda-optimized-skill
Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.
inclusionAI/AReno
Diagnose failed, hung, slow, OOM, NaN, illegal-memory-access, NCCL, compilation, rollout, or training runs in AReno.
jeremylongshore/tons-of-skills-marketplace
Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.
jeremylongshore/tons-of-skills-marketplace
Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.
jeremylongshore/tons-of-skills-marketplace
Execute proactive auto-loading: automatically detects and loads agents.md files.
jeremylongshore/tons-of-skills-marketplace
Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.
jeremylongshore/tons-of-skills-marketplace
Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.
jeremylongshore/tons-of-skills-marketplace
Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.
Categories
Diagnose the most expensive silent failure on a CoreWeave multi-node GPU job: GPUDirect RDMA falling back from InfiniBand to TCP. Coreweave Fabric Diagnostics is an agent skill from jeremylongshore/tons-of-skills-marketplace. Diagnose the most expensive silent failure on a CoreWeave multi-node GPU job: GPUDirect RDMA falling back from InfiniBand to TCP.
Coreweave Fabric Diagnostics fits situations like: multi-node training is slow; checking whether RDMA/InfiniBand is engaged; all-reduce bandwidth looks low; with coreweave slow training.
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-fabric-diagnostics -a claude-code`. Or copy the skill folder (skills/.curated/coreweave-fabric-diagnostics in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/coreweave-fabric-diagnostics in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-fabric-diagnostics -a codex`. Or copy the skill folder (skills/.curated/coreweave-fabric-diagnostics in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/coreweave-fabric-diagnostics in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-fabric-diagnostics -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/coreweave-fabric-diagnostics, .gemini/skills/coreweave-fabric-diagnostics, .github/skills/coreweave-fabric-diagnostics and .opencode/skills/coreweave-fabric-diagnostics in your project.
Going by SKILL.md and its folder, Coreweave Fabric Diagnostics needs Python for the scripts in its folder and the command-line tools its instructions call (kubectl and python3). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Glob, Bash(kubectl get:*), Bash(python3:*). Compatibility (from SKILL.md): Designed for Claude Code.
SKILL.md names 3 domains. As links in the text: docs.nvidia.com, docs.coreweave.com and github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Coreweave Fabric Diagnostics is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.4k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Coreweave Fabric Diagnostics: Hugging Face Local Model Evals (huggingface/skills, 11k stars), Liger Kernel Perf (linkedin/Liger-Kernel, 6.7k stars), Hugging Face LLM Trainer (huggingface/skills, 11k stars) and MUSA GPU Training Optimizer (open-infra-skills/infra-skills, 141 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.
Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.