Agent skill

Modal Serverless GPU

by Orchestra-Research in Orchestra-Research/AI-Research-SKILLs

Serverless GPU cloud platform for running ML workloads. An agent skill from Orchestra-Research/AI-Research-SKILLs.

MITAuto-check passedBackend & APIs

Install Modal Serverless GPU

skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill modal-serverless-gpu -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orchestra-Research/AI-Research-SKILLs modal-serverless-gpu --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/09-infrastructure/modal .claude/skills/modal-serverless-gpu && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
modal-serverless-gpu
GitHub stars
13k
Used in
5 other repos
Token cost
~2.1k tokens
SKILL.md length
391 words
Files
3 (incl. references)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

Serverless GPU cloud platform for running ML workloads. An agent skill from Orchestra-Research/AI-Research-SKILLs.

  • You need on-demand GPU access without infrastructure management
  • SKILL.md covers When to use Modal, Quick start, Core concepts and GPU configuration, plus 13 more sections
  • Calls modal and pip; needs HF_TOKEN
  • Deploying ML models as APIs

What it does

Modal Serverless GPU is an agent skill from Orchestra-Research/AI-Research-SKILLs. Serverless GPU cloud platform for running ML workloads. Use when you need on-demand GPU access without infrastructure management, deploying ML models as APIs, or running batch jobs with automatic scaling.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/advanced-usage.md` and `references/troubleshooting.md`).

It sits in Backend & APIs, covering Serverless, GPU and accelerator computing and Background jobs. The repository describes itself as: Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent… The licence is MIT.

When your agent uses it

  • You need on-demand GPU access without infrastructure management
  • Deploying ML models as APIs
  • Running batch jobs with automatic scaling

Example prompts

  • “/modal-serverless-gpu”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • modal
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • modal.com
    • github.com
    • discord.gg

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Modal Serverless GPU loads about 2.1k tokens when it runs, and up to ~7.5k if it reads all its reference files. Until then it costs about 56 tokens; SKILL.md has 391 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~56
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 391 words, ~2,139 tokens.

Download SKILL.mdSave it as .claude/skills/modal-serverless-gpu/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
modal-serverless-gpu
description
Serverless GPU cloud platform for running ML workloads. Use when you need on-demand GPU access without infrastructure management, deploying ML models as APIs, or running batch jobs with automatic scaling.
version
1.0.0
author
Orchestra Research
license
MIT
tags
Infrastructure, Serverless, GPU, Cloud, Deployment, Modal
dependencies
modal>=0.64.0

Modal Serverless GPU

Comprehensive guide to running ML workloads on Modal's serverless GPU cloud platform.

When to use Modal

Use Modal when:

  • Running GPU-intensive ML workloads without managing infrastructure
  • Deploying ML models as auto-scaling APIs
  • Running batch processing jobs (training, inference, data processing)
  • Need pay-per-second GPU pricing without idle costs
  • Prototyping ML applications quickly
  • Running scheduled jobs (cron-like workloads)

Key features:

  • Serverless GPUs: T4, L4, A10G, L40S, A100, H100, H200, B200 on-demand
  • Python-native: Define infrastructure in Python code, no YAML
  • Auto-scaling: Scale to zero, scale to 100+ GPUs instantly
  • Sub-second cold starts: Rust-based infrastructure for fast container launches
  • Container caching: Image layers cached for rapid iteration
  • Web endpoints: Deploy functions as REST APIs with zero-downtime updates

Use alternatives instead:

  • RunPod: For longer-running pods with persistent state
  • Lambda Labs: For reserved GPU instances
  • SkyPilot: For multi-cloud orchestration and cost optimization
  • Kubernetes: For complex multi-service architectures

Quick start

Installation
bash
pip install modal
modal setup  # Opens browser for authentication
Hello World with GPU
python
import modal

app = modal.App("hello-gpu")

@app.function(gpu="T4")
def gpu_info():
    import subprocess
    return subprocess.run(["nvidia-smi"], capture_output=True, text=True).stdout

@app.local_entrypoint()
def main():
    print(gpu_info.remote())

Run: modal run hello_gpu.py

Basic inference endpoint
python
import modal

app = modal.App("text-generation")
image = modal.Image.debian_slim().pip_install("transformers", "torch", "accelerate")

@app.cls(gpu="A10G", image=image)
class TextGenerator:
    @modal.enter()
    def load_model(self):
        from transformers import pipeline
        self.pipe = pipeline("text-generation", model="gpt2", device=0)

    @modal.method()
    def generate(self, prompt: str) -> str:
        return self.pipe(prompt, max_length=100)[0]["generated_text"]

@app.local_entrypoint()
def main():
    print(TextGenerator().generate.remote("Hello, world"))

Core concepts

Key components
ComponentPurpose
AppContainer for functions and resources
FunctionServerless function with compute specs
ClsClass-based functions with lifecycle hooks
ImageContainer image definition
VolumePersistent storage for models/data
SecretSecure credential storage
Execution modes
CommandDescription
modal run script.pyExecute and exit
modal serve script.pyDevelopment with live reload
modal deploy script.pyPersistent cloud deployment

GPU configuration

Show full SKILL.md (171 more words)Show less
Available GPUs
GPUVRAMBest For
T416GBBudget inference, small models
L424GBInference, Ada Lovelace arch
A10G24GBTraining/inference, 3.3x faster than T4
L40S48GBRecommended for inference (best cost/perf)
A100-40GB40GBLarge model training
A100-80GB80GBVery large models
H10080GBFastest, FP8 + Transformer Engine
H200141GBAuto-upgrade from H100, 4.8TB/s bandwidth
B200LatestBlackwell architecture
GPU specification patterns
python
# Single GPU
@app.function(gpu="A100")

# Specific memory variant
@app.function(gpu="A100-80GB")

# Multiple GPUs (up to 8)
@app.function(gpu="H100:4")

# GPU with fallbacks
@app.function(gpu=["H100", "A100", "L40S"])

# Any available GPU
@app.function(gpu="any")

Container images

python
# Basic image with pip
image = modal.Image.debian_slim(python_version="3.11").pip_install(
    "torch==2.1.0", "transformers==4.36.0", "accelerate"
)

# From CUDA base
image = modal.Image.from_registry(
    "nvidia/cuda:12.1.0-cudnn8-devel-ubuntu22.04",
    add_python="3.11"
).pip_install("torch", "transformers")

# With system packages
image = modal.Image.debian_slim().apt_install("git", "ffmpeg").pip_install("whisper")

Persistent storage

python
volume = modal.Volume.from_name("model-cache", create_if_missing=True)

@app.function(gpu="A10G", volumes={"/models": volume})
def load_model():
    import os
    model_path = "/models/llama-7b"
    if not os.path.exists(model_path):
        model = download_model()
        model.save_pretrained(model_path)
        volume.commit()  # Persist changes
    return load_from_path(model_path)

Web endpoints

FastAPI endpoint decorator
python
@app.function()
@modal.fastapi_endpoint(method="POST")
def predict(text: str) -> dict:
    return {"result": model.predict(text)}
Full ASGI app
python
from fastapi import FastAPI
web_app = FastAPI()

@web_app.post("/predict")
async def predict(text: str):
    return {"result": await model.predict.remote.aio(text)}

@app.function()
@modal.asgi_app()
def fastapi_app():
    return web_app
Web endpoint types
DecoratorUse Case
@modal.fastapi_endpoint()Simple function → API
@modal.asgi_app()Full FastAPI/Starlette apps
@modal.wsgi_app()Django/Flask apps
@modal.web_server(port)Arbitrary HTTP servers

Dynamic batching

python
@app.function()
@modal.batched(max_batch_size=32, wait_ms=100)
async def batch_predict(inputs: list[str]) -> list[dict]:
    # Inputs automatically batched
    return model.batch_predict(inputs)

Secrets management

bash
# Create secret
modal secret create huggingface HF_TOKEN=hf_xxx
python
@app.function(secrets=[modal.Secret.from_name("huggingface")])
def download_model():
    import os
    token = os.environ["HF_TOKEN"]

Scheduling

python
@app.function(schedule=modal.Cron("0 0 * * *"))  # Daily midnight
def daily_job():
    pass

@app.function(schedule=modal.Period(hours=1))
def hourly_job():
    pass

Performance optimization

Cold start mitigation
python
@app.function(
    container_idle_timeout=300,  # Keep warm 5 min
    allow_concurrent_inputs=10,  # Handle concurrent requests
)
def inference():
    pass
Model loading best practices
python
@app.cls(gpu="A100")
class Model:
    @modal.enter()  # Run once at container start
    def load(self):
        self.model = load_model()  # Load during warm-up

    @modal.method()
    def predict(self, x):
        return self.model(x)

Parallel processing

python
@app.function()
def process_item(item):
    return expensive_computation(item)

@app.function()
def run_parallel():
    items = list(range(1000))
    # Fan out to parallel containers
    results = list(process_item.map(items))
    return results

Common configuration

python
@app.function(
    gpu="A100",
    memory=32768,              # 32GB RAM
    cpu=4,                     # 4 CPU cores
    timeout=3600,              # 1 hour max
    container_idle_timeout=120,# Keep warm 2 min
    retries=3,                 # Retry on failure
    concurrency_limit=10,      # Max concurrent containers
)
def my_function():
    pass

Debugging

python
# Test locally
if __name__ == "__main__":
    result = my_function.local()

# View logs
# modal app logs my-app

Common issues

IssueSolution
Cold start latencyIncrease container_idle_timeout, use @modal.enter()
GPU OOMUse larger GPU (A100-80GB), enable gradient checkpointing
Image build failsPin dependency versions, check CUDA compatibility
Timeout errorsIncrease timeout, add checkpointing

References

Resources

© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in 09-infrastructure/modal of Orchestra-Research/AI-Research-SKILLs.

  • SKILL.md
  • references/advanced-usage.md
  • references/troubleshooting.md

Open the folder on GitHubat commit 773a529

Used in 5 other repositories

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 5 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Modal Serverless GPU next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Modal Serverless GPU compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Modal Serverless GPU this skillOrchestra-Research/AI-Research-SKILLs13k5 repos~2.1kAutomated safety check: PassMIT
ModalK-Dense-AI/scientific-agent-skills48k1 repos~4.5kAutomated safety check: NotesApache-2.0
ModalBioTender-max/awesome-bio-agent-skills197—~3.1kAutomated safety check: NotesApache-2.0
AI Model NodejsTencentCloudBase/CloudBase-AI-Toolkit1.1k3 repos~5kAutomated safety check: PassMIT
NubaseOtterMind/Nubase624—~2.2kAutomated safety check: NotesApache-2.0
Polylith Base CreationDavidVujic/python-polylith553—~757Automated safety check: PassMIT

Similar skills

  • Modal

    K-Dense-AI/scientific-agent-skills

    Modal is a serverless cloud platform for running Python on demand, including on-demand GPUs.

    48k GitHub starsUsed in 1 repo~4.5k tokens
    Backend & APIsAuto-check: notes
  • Modal

    BioTender-max/awesome-bio-agent-skills

    Cloud computing platform for running Python on GPUs and serverless infrastructure.

    197 GitHub stars~3.1k tokensUpdated 3 mo ago
    Backend & APIsAuto-check: notes
  • AI Model Nodejs

    TencentCloudBase/CloudBase-AI-Toolkit

    A skill your agent uses for Node.js backend AI via @cloudbase/node-sdk (=3.16.0) — cloud functions, CloudRun, Express/Koa/NestJS, serverless APIs, scheduled jobs, LLM proxies, agent orchestration.

    1.1k GitHub starsUsed in 3 repos~5k tokens
    Backend & APIsAuto-check passed
  • Nubase

    OtterMind/Nubase

    A skill your agent uses when the user mentions Nubase broadly, wants a backend for an AI-generated app, or needs to deploy/publish generated code online — across Database, Auth, Storage, Assets…

    624 GitHub stars~2.2k tokensUpdated 9 days ago
    Backend & APIsAuto-check: notes
  • Polylith Base Creation

    DavidVujic/python-polylith

    Create a Polylith base with poly create base — the entry point of a deployable application (HTTP API, CLI, message-queue consumer, AWS Lambda handler, GCP Cloud Function, scheduled job).

    553 GitHub stars~757 tokensUpdated 2 days ago
    Backend & APIsAuto-check passed
  • Neon

    xai-org/plugin-marketplace

    Overview of the Neon platform for apps and agents, spanning Postgres, Auth, Data API, and the new services: Object Storage, Compute Functions, and AI Gateway.

    281 GitHub stars~3.9k tokensUpdated today
    Backend & APIsAuto-check: notes

More from Orchestra-Research/AI-Research-SKILLs

All 96 skills in this repo
  • AudioCraft Audio Generation

    Orchestra-Research/AI-Research-SKILLs

    Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.

    13k GitHub starsUsed in 9 repos~3.9k tokens
    Auto-check passed
  • Peft Fine Tuning

    Orchestra-Research/AI-Research-SKILLs

    Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.

    13k GitHub starsUsed in 9 repos~3.1k tokens
    Auto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 9 repos~3.3k tokens
    Auto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    Auto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 8 repos~1.7k tokens
    Auto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    Auto-check passed

Categories

Questions about Modal Serverless GPU

What does Modal Serverless GPU do?

Serverless GPU cloud platform for running ML workloads. An agent skill from Orchestra-Research/AI-Research-SKILLs. Modal Serverless GPU is an agent skill from Orchestra-Research/AI-Research-SKILLs. Serverless GPU cloud platform for running ML workloads.

When should I use Modal Serverless GPU?

Modal Serverless GPU fits situations like: you need on-demand GPU access without infrastructure management; deploying ML models as APIs; running batch jobs with automatic scaling.

How do I install Modal Serverless GPU in Claude Code?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill modal-serverless-gpu -a claude-code`. Or copy the skill folder (09-infrastructure/modal in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/modal-serverless-gpu in your project. Claude Code loads it when a task matches its description.

How do I install Modal Serverless GPU in Codex?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill modal-serverless-gpu -a codex`. Or copy the skill folder (09-infrastructure/modal in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/modal-serverless-gpu in your project. Codex loads it when a task matches its description.

Can I use Modal Serverless GPU in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill modal-serverless-gpu -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/modal-serverless-gpu, .gemini/skills/modal-serverless-gpu, .github/skills/modal-serverless-gpu and .opencode/skills/modal-serverless-gpu in your project.

What does Modal Serverless GPU need to run?

Going by SKILL.md and its folder, Modal Serverless GPU needs the command-line tools its instructions call (modal and pip) and credentials named HF_TOKEN. Our summary lists: Python 3.

Does Modal Serverless GPU access the network?

SKILL.md names 3 domains. As links in the text: modal.com, github.com and discord.gg. This is read from the text; nothing was executed.

Is Modal Serverless GPU safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Modal Serverless GPU use?

Modal Serverless GPU is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Modal Serverless GPU use?

About 2.1k tokens (SKILL.md is roughly 8.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.4k tokens, read only when the agent opens those files.

What are the alternatives to Modal Serverless GPU?

Skills that share tags, products or a category with Modal Serverless GPU: Modal (K-Dense-AI/scientific-agent-skills, 48k stars), Modal (BioTender-max/awesome-bio-agent-skills, 197 stars), AI Model Nodejs (TencentCloudBase/CloudBase-AI-Toolkit, 1.1k stars) and Nubase (OtterMind/Nubase, 624 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Modal Serverless GPU?

Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,313 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.

Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.