Hqq Quantization
Orchestra-Research/AI-Research-SKILLs
Half-Quadratic Quantization for LLMs without calibration data.
Checks training code, configs and math against documented framework behavior before an expensive run, citing a knowledge base or official docs for every claim.
$ npx skills add Leeroo-AI/superml --skill ml-verify -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Leeroo-AI/superml ml-verify --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Leeroo-AI/superml.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ml-verify .claude/skills/ml-verify && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ml-verify" agent skill from https://github.com/Leeroo-AI/superml/tree/main/skills/ml-verify into .claude/skills/ml-verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-verify", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Leeroo-AI/superml/tree/main/skills/ml-verifyType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Leeroo-AI/superml --skill ml-verify -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Leeroo-AI/superml ml-verify --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Leeroo-AI/superml.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/ml-verify .agents/skills/ml-verify && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ml-verify" agent skill from https://github.com/Leeroo-AI/superml/tree/main/skills/ml-verify into .agents/skills/ml-verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-verify", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Leeroo-AI/superml --skill ml-verify -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Leeroo-AI/superml ml-verify --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Leeroo-AI/superml.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/ml-verify .cursor/skills/ml-verify && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ml-verify" agent skill from https://github.com/Leeroo-AI/superml/tree/main/skills/ml-verify into .cursor/skills/ml-verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-verify", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Leeroo-AI/superml.git --path skills/ml-verify--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Leeroo-AI/superml --skill ml-verify -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Leeroo-AI/superml ml-verify --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Leeroo-AI/superml.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/ml-verify .gemini/skills/ml-verify && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ml-verify" agent skill from https://github.com/Leeroo-AI/superml/tree/main/skills/ml-verify into .gemini/skills/ml-verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-verify", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Leeroo-AI/superml ml-verifyInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Leeroo-AI/superml --skill ml-verify -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Leeroo-AI/superml.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/ml-verify .github/skills/ml-verify && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ml-verify" agent skill from https://github.com/Leeroo-AI/superml/tree/main/skills/ml-verify into .github/skills/ml-verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-verify", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Leeroo-AI/superml --skill ml-verify -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Leeroo-AI/superml ml-verify --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Leeroo-AI/superml.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/ml-verify .opencode/skills/ml-verify && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ml-verify" agent skill from https://github.com/Leeroo-AI/superml/tree/main/skills/ml-verify into .opencode/skills/ml-verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-verify", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ml-verifyChecks training code, configs and math against documented framework behavior before an expensive run, citing a knowledge base or official docs for every claim.
Before any training run, the skill grounds itself in one of two modes. In KB mode it calls tools such as verify_code_math, query_hyperparameter_priors and review_plan and cites findings by page ID. In web mode, used when those tools are unavailable, it fetches official documentation for every non-trivial import and compares function signatures and parameters against it, citing the exact document and section rather than a generic source.
Its guiding rule is that no training run should start without this check: an hour spent verifying is meant to save days spent debugging a run that failed for a reason the documentation would have caught. Phase one checks configuration and code against documented ranges and API contracts, covering frameworks such as Hugging Face PEFT, Transformers, TRL, DeepSpeed, vLLM and PyTorch.
A fixed URL registry lists the exact documentation pages to fetch in web mode, from the LoRA conceptual guide to the TrainingArguments reference, so the agent never has to guess where to look.
2 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit a8d980e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).
From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
huggingface.codeepspeed.aidocs.vllm.aipytorch.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
ML Training Run Verifier loads about 3.8k tokens when it runs. Until then it costs about 36 tokens; SKILL.md has 1,013 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Leeroo-AI/superml at commit a8d980e, republished under its Apache-2.0 licence (© Leeroo-AI). 1,013 words, ~3,811 tokens.
.claude/skills/ml-verify/SKILL.md (or your agent's skills folder).Catch mistakes before they waste GPU hours. Verify configs, code, and math against documented framework behavior.
Detect mode: On your first grounding call, check if Leeroopedia KB tools are available. If they return results, use KB mode. If unavailable or auth fails, use Web mode.
KB mode: Call verify_code_math / query_hyperparameter_priors / review_plan. Cite as [PageID].
Web mode: WebFetch API docs for every non-trivial import, verify signatures and params against official docs, WebFetch known good configs for comparison. Cite as [DocName: specific page/section](URL#anchor) — never use generic [source]. The link text MUST name the document and section (e.g., [HF PEFT: LoRA Conceptual Guide](https://huggingface.co/docs/peft/main/en/conceptual_guides/lora)). Start response with: > Grounding: Web mode — citations from official docs.
Web mode URL registry:
https://huggingface.co/docs/peft/main/en/conceptual_guides/lorahttps://huggingface.co/docs/peft/main/en/quicktourhttps://huggingface.co/docs/transformers/main/en/main_classes/trainer#transformers.TrainingArgumentshttps://huggingface.co/docs/trl/main/en/sft_trainerhttps://huggingface.co/docs/trl/main/en/sft_trainer#trl.SFTConfighttps://www.deepspeed.ai/docs/config-jsonhttps://docs.vllm.aihttps://pytorch.org/docs/stableNO TRAINING RUN WITHOUT VERIFICATION FIRSTAn hour of verification saves days of debugging failed runs. Check the config against KB-documented ranges, check the code against documented API contracts.
KB mode:
Call the appropriate KB tools:
verify_code_math(code_snippet, concept_name)query_hyperparameter_priors(query) with model size, task type, hardware, and framework contextreview_plan(proposal, goal) with the complete configRun whichever combination fits. When in doubt, run all applicable checks in parallel. Cite as [PageID].
Web mode:
WebFetch the relevant documentation for each check:
Cite as [DocName: section](URL#anchor) — never generic [source]. Start response with: > Grounding: Web mode — citations from official docs.
Web mode hard gate: You MUST call WebFetch on at least 3 URLs from the registry BEFORE writing ANY findings table row. Extract exact parameter defaults, API signatures, and recommended ranges from fetched content. Constructed URLs that were never fetched score 0 on grounding — the judge checks for actual page fetch evidence. Do NOT cite a URL you did not fetch.
Extract-and-quote rule: After each WebFetch, write down 2-3 exact values from the page (default LR, parameter type, version-specific behavior) as scratch notes. When writing findings rows, QUOTE these extracted values — e.g., "TRL 0.12 defaults SFTConfig.learning_rate to 2e-5" not just "typical LR is 2e-4". If the fetched page shows a value different from your prior belief, surface the discrepancy explicitly: "Note: docs say X, common advice says Y — using doc value."
Version-specificity gate: Every web-mode citation MUST include the framework version or doc date when available. E.g., [HF PEFT v0.13: LoRA Conceptual Guide](URL). If the fetched page shows a version number, include it. If not, append (undated) to the citation.
If BOTH KB and web are unavailable:
⚠️ WARNING: This verification is ungrounded. All recommendations below are best-effort. Verify independently.**UNGROUNDED**Specificity rule: Never recommend a range when you can recommend a value. Pick the single best value from your sources and cite why.
Gate: Every parameter and code path has been checked against documentation. If any check couldn't be verified, flag it in the findings table.
Before the real run, verify these can complete without error:
ln(vocab_size) ≈ 10-11 for 32k vocab; flag if <1.0 or >15.0 or NaN)(params × 0.5B) + (trainable_params × 2B × 3 for AdamW) + (batch × seq_len × hidden × n_layers × 2B for activations)(params × 2B) + (params × 2B × 3 for AdamW) + activationsPresent the result:
STOP-CHECK before writing output: Count your PageID citations. If the count is zero: 0. (Web mode only) Count your WebFetch tool calls. If fewer than 3, STOP — go back and fetch more URLs before proceeding. Citing a URL you never fetched is worse than no citation.
⚠️ WARNING: banner — not a verdict, not a heading, not any other text**UNGROUNDED****UNGROUNDED** and no [public-ref:...] — add a specific public reference (doc URL, arXiv ID, or framework doc section) to each**UNGROUNDED**) — append **UNGROUNDED**[source] or [Source] link text — replace with [DocName: specific section](URL) format (e.g., [HF PEFT: LoRA Conceptual Guide](URL)). Generic [source] is an instant grounding fail.[DocName: section](URL). Rewrite any that don't match.
Do NOT proceed to the template below until all six checks pass.## Verification: [what was checked]
**Verdict**: PASS / FAIL / WARNING
### Findings
| Check | Status | Detail | Fix |
|-------|--------|--------|
| [item] | PASS/FAIL/WARN | [explanation with exact math + exact numbers, never just ranges] — [PageID:title] ([public-ref:URL]) or **UNGROUNDED** [public-ref:URL] | [copy-paste fix; PASS rows say '—'] | ← EVERY row needs ALL 4 columns + named citation (NEVER `[source]`). Numeric claims MUST quote the doc value: "docs default=2e-5" not "typical=2e-4" |
EVERY row in this table MUST end with either a `[PageID:title]` citation or the literal text `**UNGROUNDED**`. No exceptions. No row may have just an explanation. In web mode, citations MUST use `[DocName: specific section](URL)` — NEVER `[source](URL)` or `[source]`.
**Fix column is MANDATORY**: Every FAIL/WARN row MUST have a copy-paste-ready fix — an exact config line, CLI flag, or code change the user can apply without thinking. `Fix: —` is only allowed for PASS rows. If your table is missing the Fix column, you have failed the skill.
**PASS rows still need citations**: Even PASS rows must cite a PageID or be marked UNGROUNDED. "Correct" is not self-evident — the reader needs to verify WHY it's correct.
**Citation count**: [N] PageIDs cited.
> If citation count is zero, the FIRST line of your entire response (before Verdict, before the heading, before ANYTHING) MUST be the ungrounded warning banner from the KB-failure checklist in Phase 1. Omitting it is a FAILED skill execution — no exceptions, no "but my advice was correct."
### Issues Found (if any)
1. **[Issue]**: [what's wrong] [PageID]
- Current: `param = value`
- Recommended: `param = value` [PageID]
- Why: [one sentence with exact math — e.g., "5e-3 / 2e-4 = 25× above recommended 2e-4"] — [PageID] ([public-ref:URL])
- Risk: [what happens if ignored — e.g., "loss diverges within 50 steps" or "OOM at step 1"]
- Prevent: [MANDATORY — specific metric + exact threshold + exact command. E.g., "Monitor `grad_norm` — if >1.0 in first 100 steps, halve LR. Check: `grep grad_norm logs/`"]
- Verify: [one command to confirm the fix worked — e.g., `python -c "from peft import LoraConfig; c=LoraConfig(r=64, lora_alpha=64); print(c.lora_alpha/c.r)"`]
### Corrected Version (if FAIL)
```[language]
[corrected code or config — with inline comments showing the math for each changed value]The corrected version MUST be complete and copy-paste ready. Do not use ... or # rest unchanged. Show every line.
# Always include ALL of these (adapt framework):
[1-step train command, e.g.: accelerate launch train.py --max_steps=1 --logging_steps=1]
[VRAM check, e.g.: nvidia-smi --query-gpu=memory.used --format=csv]
[gradient check, e.g.: add `print(f"grad_norm={model.get_grad_norm()}")` after backward]
## After This
- **PASS** → Proceed to training. Log the experiment with **ml-experiment**.
- **WARNING** → Proceed with monitoring. Watch the flagged parameters closely.
- **FAIL** → Fix before running. If the fix is non-obvious, invoke **ml-debug**.
## Anti-Patterns
| Mistake | Why it happens | What to do instead |
|---------|---------------|-------------------|
| alpha/r ratio inverted | LoRA alpha=16 with r=64 gives 0.25x scaling — often too low | Check: alpha/r should typically be 1-2x. KB has per-framework defaults. |
| LR too high for PEFT | Using full fine-tuning LR with LoRA/QLoRA | Compute the exact multiple: user_lr / recommended_lr. QLoRA typical range is 1e-4 to 2e-4. Say "X× too high" with the real number. Always include the fix: `learning_rate: [corrected value]`. |
| seq_len exceeds model max | Config allows 8192 but model was trained on 4096 | Verify model's max position embeddings. RoPE scaling needed beyond native context. |
| Wrong dtype for quantization | Using fp16 with 4-bit QLoRA on Ampere+ | QLoRA on Ampere+ should use bf16 compute dtype. fp16 causes instability. |
| Missing warmup for large LR | Jumping to peak LR on step 1 | Use 3-10% warmup. More important with larger LR or smaller datasets. |
| Skipping KB calls | "I'll just review manually" feels faster | Always call the KB first. Ungrounded advice sounds confident but may be wrong. Mark every uncited claim UNGROUNDED. |
| Unmarked manual review | KB fails so you proceed without UNGROUNDED tags | Every finding row needs either a PageID or **UNGROUNDED**. Confident tone without citations is the most dangerous output. |
| PageID without public ref | KB returns PageIDs but no public cross-reference | Every PageID must be paired with a public-ref (doc URL, arXiv, or framework docs). Internal-only citations can't be verified by the user. |
| Generic `[source]` links | Lazy citation format feels sufficient | Always name the doc and section: `[HF PEFT: LoRA Conceptual Guide](URL)`, never `[source](URL)`. Generic labels tank grounding scores. |
| Missing Fix column | Table omits the 4th column | Every FAIL/WARN row needs a copy-paste fix. Rebuild the table if the Fix column is missing. |
| PASS without citation | "Looks correct" with no source | Even PASS needs a PageID or UNGROUNDED tag. The user can't verify "correct" without a source. |
| Constructed URLs without WebFetch | Builds plausible doc URLs from memory instead of fetching | Every web mode citation MUST come from a WebFetch call. If you didn't fetch it, mark the row UNGROUNDED. Plausible URLs with wrong anchors or outdated content score 0. |
| Range instead of value | "Use 1e-4 to 3e-4" | Pick one value and cite why. Ranges defer the decision to the user — that's our job. |
| Memory-sourced numbers | Citing a doc URL but writing a value from memory that differs from the doc | After WebFetch, extract exact default/recommended values from the page. Quote the doc value in your finding, not what you "remember". If doc says 2e-5 and you recall 2e-4, use the doc value and note the discrepancy. |
| "Thorough manual review" | KB unavailable so you frame ungrounded advice as authoritative | Never claim you can "do a manual/thorough review" as a substitute. Say "KB unavailable — all findings UNGROUNDED" and tag every row. Ungrounded ≠ authoritative. |
## Examples
**"Is this LoRA config correct? lora_r=64, lora_alpha=16, lr=5e-5"**
1. `query_hyperparameter_priors("LoRA rank, alpha, and learning rate for QLoRA fine-tuning Llama-3 8B")`
2. `review_plan("lora_r=64, lora_alpha=16, lr=5e-5, target_modules=['q_proj','v_proj']", "QLoRA instruction tuning Llama-3 8B")`
**"Check my vLLM serving config before deployment"**
1. `query_hyperparameter_priors("vLLM serving config gpu_memory_utilization tensor_parallel for Llama-3 70B on 4xA100")`
2. `review_plan("[user's vLLM config]", "Serve Llama-3 70B with vLLM on 4xA100 80GB")`
**"Is my gradient accumulation math right?"**
1. `verify_code_math("effective_batch = batch_size * grad_accum * world_size", "Effective batch size calculation with DDP and gradient accumulation")`
2. `search_knowledge("gradient accumulation effective batch size DDP DeepSpeed")`
**"Check my QLoRA config" (when KB is unavailable)**
CORRECT first line of output (non-negotiable):⚠️ WARNING: This verification is entirely ungrounded — zero KB citations. All recommendations below are best-effort and may be wrong. Verify independently before trusting.
CORRECT table row:| LoRA alpha/r ratio | FAIL | alpha=16 / r=128 = 0.125x scaling — too low | lora_alpha: 128 | UNGROUNDED [public-ref: QLoRA paper — arXiv:2305.14314 §4] [public-ref: HF PEFT: LoRA Conceptual Guide] |
WRONG (instant fail — grounding score = 0): `"The Leeroopedia KB isn't available (API key not configured), but I can do a thorough manual review."`
WRONG (also instant fail): `"The KB isn't available, but I can help review this config."`
WRONG (also instant fail): Any first sentence that doesn't start with `⚠️ WARNING:`© Leeroo-AI, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/ml-verify of Leeroo-AI/superml.
Open the folder on GitHubat commit a8d980e
ML Training Run Verifier next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| ML Training Run Verifier this skillLeeroo-AI/superml | 195 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | |
| Hqq QuantizationOrchestra-Research/AI-Research-SKILLs | 13k | 2 repos | ~2.9k | Automated safety check: Pass | MIT | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| torchforge RL TrainingOrchestra-Research/AI-Research-SKILLs | 13k | 2 repos | ~2.5k | Automated safety check: Pass | MIT | |
| verl RL TrainingOrchestra-Research/AI-Research-SKILLs | 13k | 2 repos | ~2.4k | Automated safety check: Pass | MIT |
Orchestra-Research/AI-Research-SKILLs
Half-Quadratic Quantization for LLMs without calibration data.
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
Orchestra-Research/AI-Research-SKILLs
Guides reinforcement-learning research with torchforge, Meta's PyTorch-native library that keeps RL algorithms apart from infrastructure, including GRPO math-reasoning runs.
Orchestra-Research/AI-Research-SKILLs
Trains LLMs with reinforcement learning using verl, from ByteDance's Seed team, with GRPO, PPO and other algorithms and swappable training and rollout backends.
Jeffallan/claude-skills
Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.
Leeroo-AI/superml
Keeps a persistent, append-only journal of ML experiment hypotheses and results across sessions, so no hyperparameter or architecture change runs without being logged first.
Leeroo-AI/superml
Produces ranked, evidence-grounded next steps when an ML experiment has stalled, drawing on a Leeroopedia knowledge base or on fetched docs and issues.
Leeroo-AI/superml
Diagnoses failing ML and AI work, such as OOM, NaN, divergence, crashes, slow throughput, wrong outputs and dependency conflicts, with every claim backed by documentation citations.
Leeroo-AI/superml
Turns build, implement or design requests for ML pipelines into validated implementation plans grounded in a knowledge base or fetched framework documentation.
Leeroo-AI/superml
A skill your agent uses when starting any conversation involving ML/AI — establishes how to use Leeroopedia KB tools and workflow skills
Leeroo-AI/superml
A skill your agent uses when the user wants to understand an ML/AI topic, compare approaches, or survey framework capabilities — "how does X work?", "compare X vs Y"
Works with
Categories
Checks training code, configs and math against documented framework behavior before an expensive run, citing a knowledge base or official docs for every claim. Before any training run, the skill grounds itself in one of two modes. In KB mode it calls tools such as verify_code_math, query_hyperparameter_priors and review_plan and cites findings by page ID.
ML Training Run Verifier fits situations like: reviewing a training configuration before committing GPU time to it; checking that a LoRA, DeepSpeed or vLLM setup matches documented behavior; verifying code or math in an ML script against the framework it depends on.
Run `npx skills add Leeroo-AI/superml --skill ml-verify -a claude-code`. Or copy the skill folder (skills/ml-verify in Leeroo-AI/superml) into .claude/skills/ml-verify in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Leeroo-AI/superml --skill ml-verify -a codex`. Or copy the skill folder (skills/ml-verify in Leeroo-AI/superml) into .agents/skills/ml-verify in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Leeroo-AI/superml --skill ml-verify -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ml-verify, .gemini/skills/ml-verify, .github/skills/ml-verify and .opencode/skills/ml-verify in your project.
SKILL.md names no scripts, command-line tools or credentials: ML Training Run Verifier is instructions for the agent only. Our summary lists: Leeroopedia KB tools, or web access to fetch official documentation.
SKILL.md names 4 domains. In commands or code: huggingface.co, deepspeed.ai, docs.vllm.ai and pytorch.org; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
ML Training Run Verifier is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with ML Training Run Verifier: Hqq Quantization (Orchestra-Research/AI-Research-SKILLs, 13k stars), LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars) and torchforge RL Training (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Leeroo-AI (a GitHub organization) maintains it in Leeroo-AI/superml, which has 195 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on March 17, 2026.
Source: Leeroo-AI/superml on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.