Megatron-LM on SLURM
NVIDIA/Megatron-LM
Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.
Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives.
The automated check flagged lines worth reading first. See the safety section below.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill lambda-labs-gpu-cloud -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs lambda-labs-gpu-cloud --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/09-infrastructure/lambda-labs .claude/skills/lambda-labs-gpu-cloud && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "lambda-labs-gpu-cloud" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/09-infrastructure/lambda-labs into .claude/skills/lambda-labs-gpu-cloud/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lambda-labs-gpu-cloud", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/09-infrastructure/lambda-labsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill lambda-labs-gpu-cloud -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs lambda-labs-gpu-cloud --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .agents/skills && cp -r skills-src/09-infrastructure/lambda-labs .agents/skills/lambda-labs-gpu-cloud && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "lambda-labs-gpu-cloud" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/09-infrastructure/lambda-labs into .agents/skills/lambda-labs-gpu-cloud/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lambda-labs-gpu-cloud", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill lambda-labs-gpu-cloud -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs lambda-labs-gpu-cloud --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/09-infrastructure/lambda-labs .cursor/skills/lambda-labs-gpu-cloud && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "lambda-labs-gpu-cloud" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/09-infrastructure/lambda-labs into .cursor/skills/lambda-labs-gpu-cloud/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lambda-labs-gpu-cloud", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Orchestra-Research/AI-Research-SKILLs.git --path 09-infrastructure/lambda-labs--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill lambda-labs-gpu-cloud -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs lambda-labs-gpu-cloud --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/09-infrastructure/lambda-labs .gemini/skills/lambda-labs-gpu-cloud && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "lambda-labs-gpu-cloud" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/09-infrastructure/lambda-labs into .gemini/skills/lambda-labs-gpu-cloud/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lambda-labs-gpu-cloud", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Orchestra-Research/AI-Research-SKILLs lambda-labs-gpu-cloudInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill lambda-labs-gpu-cloud -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .github/skills && cp -r skills-src/09-infrastructure/lambda-labs .github/skills/lambda-labs-gpu-cloud && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "lambda-labs-gpu-cloud" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/09-infrastructure/lambda-labs into .github/skills/lambda-labs-gpu-cloud/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lambda-labs-gpu-cloud", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill lambda-labs-gpu-cloud -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs lambda-labs-gpu-cloud --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/09-infrastructure/lambda-labs .opencode/skills/lambda-labs-gpu-cloud && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "lambda-labs-gpu-cloud" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/09-infrastructure/lambda-labs into .opencode/skills/lambda-labs-gpu-cloud/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lambda-labs-gpu-cloud", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
lambda-labs-gpu-cloudGuide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives.
This skill is a guide to renting GPUs from Lambda Labs for machine-learning work, either as on-demand instances or as 1-Click Clusters. It lists when Lambda fits: dedicated instances with full SSH access, long training runs, persistent storage, no egress fees, multi-node clusters of 16 to 512 GPUs, and a pre-installed Lambda Stack with PyTorch, CUDA and NCCL. It names Modal, SkyPilot, RunPod and Vast.ai as alternatives for serverless, multi-cloud orchestration, cheaper spot capacity and marketplace pricing.
The quick start covers account setup, which needs a payment method, an API key and an SSH key added before any launch, then launching through the console, choosing a GPU type and region, optionally attaching a persistent filesystem and connecting over SSH as the ubuntu user. It tabulates GPU models from B200 to V100 with memory and intended use, explains 8x, 4x and 2x instance layouts for distributed training and gives launch times. Reference files hold advanced usage and troubleshooting.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
sshpythonpipcurljqjupytergitFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
cloud.lambdalabs.comgithub.comAlso links to:
lambda.aicloud.lambda.aidocs.lambda.aisupport.lambdalabs.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
LAMBDA_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Lambda Labs GPU Cloud loads about 3k tokens when it runs, and up to ~9.7k if it reads all its reference files. Until then it costs about 63 tokens; SKILL.md has 671 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found patterns that need a careful read before installing.
ssh -i ~/.ssh/lambda_key ubuntu@<INSTANCE-IP>ssh-keygen -t ed25519 -f ~/.ssh/lambda_keyecho 'ssh-rsa AAAA...' >> ~/.ssh/authorized_keysAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 671 words, ~3,031 tokens.
.claude/skills/lambda-labs-gpu-cloud/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Comprehensive guide to running ML workloads on Lambda Labs GPU cloud with on-demand instances and 1-Click Clusters.
Use Lambda Labs when:
Key features:
Use alternatives instead:
# Get instance IP from console
ssh ubuntu@<INSTANCE-IP>
# Or with specific key
ssh -i ~/.ssh/lambda_key ubuntu@<INSTANCE-IP>| GPU | VRAM | Price/GPU/hr | Best For |
|---|---|---|---|
| B200 SXM6 | 180 GB | $4.99 | Largest models, fastest training |
| H100 SXM | 80 GB | $2.99-3.29 | Large model training |
| H100 PCIe | 80 GB | $2.49 | Cost-effective H100 |
| GH200 | 96 GB | $1.49 | Single-GPU large models |
| A100 80GB | 80 GB | $1.79 | Production training |
| A100 40GB | 40 GB | $1.29 | Standard training |
| A10 | 24 GB | $0.75 | Inference, fine-tuning |
| A6000 | 48 GB | $0.80 | Good VRAM/price ratio |
| V100 | 16 GB | $0.55 | Budget training |
8x GPU: Best for distributed training (DDP, FSDP)
4x GPU: Large models, multi-GPU training
2x GPU: Medium workloads
1x GPU: Fine-tuning, inference, developmentAll instances come with Lambda Stack pre-installed:
# Included software
- Ubuntu 22.04 LTS
- NVIDIA drivers (latest)
- CUDA 12.x
- cuDNN 8.x
- NCCL (for multi-GPU)
- PyTorch (latest)
- TensorFlow (latest)
- JAX
- JupyterLab# Check GPU
nvidia-smi
# Check PyTorch
python -c "import torch; print(torch.cuda.is_available())"
# Check CUDA version
nvcc --versionpip install lambda-cloud-clientimport os
import lambda_cloud_client
# Configure with API key
configuration = lambda_cloud_client.Configuration(
host="https://cloud.lambdalabs.com/api/v1",
access_token=os.environ["LAMBDA_API_KEY"]
)with lambda_cloud_client.ApiClient(configuration) as api_client:
api = lambda_cloud_client.DefaultApi(api_client)
# Get available instance types
types = api.instance_types()
for name, info in types.data.items():
print(f"{name}: {info.instance_type.description}")from lambda_cloud_client.models import LaunchInstanceRequest
request = LaunchInstanceRequest(
region_name="us-west-1",
instance_type_name="gpu_1x_h100_sxm5",
ssh_key_names=["my-ssh-key"],
file_system_names=["my-filesystem"], # Optional
name="training-job"
)
response = api.launch_instance(request)
instance_id = response.data.instance_ids[0]
print(f"Launched: {instance_id}")instances = api.list_instances()
for instance in instances.data:
print(f"{instance.name}: {instance.ip} ({instance.status})")from lambda_cloud_client.models import TerminateInstanceRequest
request = TerminateInstanceRequest(
instance_ids=[instance_id]
)
api.terminate_instance(request)from lambda_cloud_client.models import AddSshKeyRequest
# Add SSH key
request = AddSshKeyRequest(
name="my-key",
public_key="ssh-rsa AAAA..."
)
api.add_ssh_key(request)
# List keys
keys = api.list_ssh_keys()
# Delete key
api.delete_ssh_key(key_id)curl -u $LAMBDA_API_KEY: \
https://cloud.lambdalabs.com/api/v1/instance-types | jqcurl -u $LAMBDA_API_KEY: \
-X POST https://cloud.lambdalabs.com/api/v1/instance-operations/launch \
-H "Content-Type: application/json" \
-d '{
"region_name": "us-west-1",
"instance_type_name": "gpu_1x_h100_sxm5",
"ssh_key_names": ["my-key"]
}' | jqcurl -u $LAMBDA_API_KEY: \
-X POST https://cloud.lambdalabs.com/api/v1/instance-operations/terminate \
-H "Content-Type: application/json" \
-d '{"instance_ids": ["<INSTANCE-ID>"]}' | jqFilesystems persist data across instance restarts:
# Mount location
/lambda/nfs/<FILESYSTEM_NAME>
# Example: save checkpoints
python train.py --checkpoint-dir /lambda/nfs/my-storage/checkpointsFilesystems must be attached at instance launch time:
file_system_names in launch request# Store on filesystem (persists)
/lambda/nfs/storage/
├── datasets/
├── checkpoints/
├── models/
└── outputs/
# Local SSD (faster, ephemeral)
/home/ubuntu/
└── working/ # Temporary files# Generate key locally
ssh-keygen -t ed25519 -f ~/.ssh/lambda_key
# Add public key to Lambda console
# Or via API# On instance, add more keys
echo 'ssh-rsa AAAA...' >> ~/.ssh/authorized_keys# On instance
ssh-import-id gh:username# Forward Jupyter
ssh -L 8888:localhost:8888 ubuntu@<IP>
# Forward TensorBoard
ssh -L 6006:localhost:6006 ubuntu@<IP>
# Multiple ports
ssh -L 8888:localhost:8888 -L 6006:localhost:6006 ubuntu@<IP># On instance
jupyter lab --ip=0.0.0.0 --port=8888
# From local machine with tunnel
ssh -L 8888:localhost:8888 ubuntu@<IP>
# Open http://localhost:8888# SSH to instance
ssh ubuntu@<IP>
# Clone repo
git clone https://github.com/user/project
cd project
# Install dependencies
pip install -r requirements.txt
# Train
python train.py --epochs 100 --checkpoint-dir /lambda/nfs/storage/checkpoints# train_ddp.py
import torch
import torch.distributed as dist
from torch.nn.parallel import DistributedDataParallel as DDP
def main():
dist.init_process_group("nccl")
rank = dist.get_rank()
device = rank % torch.cuda.device_count()
model = MyModel().to(device)
model = DDP(model, device_ids=[device])
# Training loop...
if __name__ == "__main__":
main()# Launch with torchrun (8 GPUs)
torchrun --nproc_per_node=8 train_ddp.pyimport os
checkpoint_dir = "/lambda/nfs/my-storage/checkpoints"
os.makedirs(checkpoint_dir, exist_ok=True)
# Save checkpoint
torch.save({
'epoch': epoch,
'model_state_dict': model.state_dict(),
'optimizer_state_dict': optimizer.state_dict(),
'loss': loss,
}, f"{checkpoint_dir}/checkpoint_{epoch}.pt")High-performance Slurm clusters with:
# On Slurm cluster
srun --nodes=4 --ntasks-per-node=8 --gpus-per-node=8 \
torchrun --nnodes=4 --nproc_per_node=8 \
--rdzv_backend=c10d --rdzv_endpoint=$MASTER_ADDR:29500 \
train.py# Find private IP
ip addr show | grep 'inet '# 1. Launch 8x H100 instance with filesystem
# 2. SSH and setup
ssh ubuntu@<IP>
pip install transformers accelerate peft
# 3. Download model to filesystem
python -c "
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained('meta-llama/Llama-2-7b-hf')
model.save_pretrained('/lambda/nfs/storage/models/llama-2-7b')
"
# 4. Fine-tune with checkpoints on filesystem
accelerate launch --num_processes 8 train.py \
--model_path /lambda/nfs/storage/models/llama-2-7b \
--output_dir /lambda/nfs/storage/outputs \
--checkpoint_dir /lambda/nfs/storage/checkpoints# 1. Launch A10 instance (cost-effective for inference)
# 2. Run inference
python inference.py \
--model /lambda/nfs/storage/models/fine-tuned \
--input /lambda/nfs/storage/data/inputs.jsonl \
--output /lambda/nfs/storage/data/outputs.jsonl| Task | Recommended GPU |
|---|---|
| LLM fine-tuning (7B) | A100 40GB |
| LLM fine-tuning (70B) | 8x H100 |
| Inference | A10, A6000 |
| Development | V100, A10 |
| Maximum performance | B200 |
| Issue | Solution |
|---|---|
| Instance won't launch | Check region availability, try different GPU |
| SSH connection refused | Wait for instance to initialize (3-15 min) |
| Data lost after terminate | Use persistent filesystems |
| Slow data transfer | Use filesystem in same region |
| GPU not detected | Reboot instance, check drivers |
© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in 09-infrastructure/lambda-labs of Orchestra-Research/AI-Research-SKILLs.
Open the folder on GitHubat commit 773a529
We found 13 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 5 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.
Lambda Labs GPU Cloud next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Lambda Labs GPU Cloud this skillOrchestra-Research/AI-Research-SKILLs | 13k | 5 repos | ~3k | Automated safety check: Warn | MIT | |
| Megatron-LM on SLURMNVIDIA/Megatron-LM | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| AI ML Engineertheneoai/awesome-skills | 183 | — | ~2.9k | Automated safety check: Pass | MIT | |
| MUSA GPU Training Optimizeropen-infra-skills/infra-skills | 141 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| GPU OptimizerMathews-Tom/armory | 327 | — | ~3.5k | Automated safety check: Notes | MIT | |
| Cuda Index Widthpytorch/pytorch | 104k | — | ~1.6k | Automated safety check: Pass | Custom licence |
NVIDIA/Megatron-LM
Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.
theneoai/awesome-skills
Expert AI/ML Engineer with deep MLOps expertise. An agent skill from theneoai/awesome-skills.
open-infra-skills/infra-skills
Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.
Mathews-Tom/armory
GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.
pytorch/pytorch
Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
Orchestra-Research/AI-Research-SKILLs
Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.
Orchestra-Research/AI-Research-SKILLs
Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
Orchestra-Research/AI-Research-SKILLs
Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Works with
Categories
Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives. This skill is a guide to renting GPUs from Lambda Labs for machine-learning work, either as on-demand instances or as 1-Click Clusters. It lists when Lambda fits: dedicated instances with full SSH access, long training runs, persistent storage, no egress fees, multi-node clusters of 16 to 512 GPUs, and a pre-installed Lambda Stack with PyTorch, CUDA and NCCL.
Lambda Labs GPU Cloud fits situations like: choosing a GPU cloud for a long training job that needs SSH access; launching an on-demand Lambda instance and connecting over SSH; planning a multi-node cluster for large-scale training; keeping datasets on a persistent filesystem across instance restarts.
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill lambda-labs-gpu-cloud -a claude-code`. Or copy the skill folder (09-infrastructure/lambda-labs in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/lambda-labs-gpu-cloud in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill lambda-labs-gpu-cloud -a codex`. Or copy the skill folder (09-infrastructure/lambda-labs in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/lambda-labs-gpu-cloud in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill lambda-labs-gpu-cloud -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/lambda-labs-gpu-cloud, .gemini/skills/lambda-labs-gpu-cloud, .github/skills/lambda-labs-gpu-cloud and .opencode/skills/lambda-labs-gpu-cloud in your project.
Going by SKILL.md and its folder, Lambda Labs GPU Cloud needs the command-line tools its instructions call (ssh, python, pip, curl, jq and jupyter) and credentials named LAMBDA_API_KEY. Our summary lists: A Lambda account with a payment method and an API key; An SSH key added before launching instances.
SKILL.md names 6 domains. In commands or code: cloud.lambdalabs.com and github.com; the agent is likely to contact these when it follows the instructions. As links in the text: lambda.ai, cloud.lambda.ai, docs.lambda.ai and support.lambdalabs.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md flagged 3 warning(s): mentions a credentials file (ssh keys, cloud or package-manager tokens). Read the flagged lines before installing; the check is not a guarantee either way.
Lambda Labs GPU Cloud is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.7k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Lambda Labs GPU Cloud: Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars), AI ML Engineer (theneoai/awesome-skills, 183 stars), MUSA GPU Training Optimizer (open-infra-skills/infra-skills, 141 stars) and GPU Optimizer (Mathews-Tom/armory, 327 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,313 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.
Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.