Agent skill

Runtime Skills

by llama-farm in llama-farm/llamafarm

Universal Runtime best practices for PyTorch inference, Transformers models, and FastAPI serving.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Runtime Skills

skills CLI
$ npx skills add llama-farm/llamafarm --skill runtime-skills -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install llama-farm/llamafarm runtime-skills --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/llama-farm/llamafarm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/runtime-skills .claude/skills/runtime-skills && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
runtime-skills
GitHub stars
836
Token cost
~1.3k tokens
SKILL.md length
210 words
Files
5
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

Universal Runtime best practices for PyTorch inference, Transformers models, and FastAPI serving.

  • Works in 4 steps: Model Loading with Double-Checked Locking → Device-Aware Tensor Operations → TTL-Based Model Caching → …
  • Tasks that involve Deep learning
  • SKILL.md covers Overview, Links to Shared Skills, Runtime-Specific Checklists and Architecture, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Runtime Skills is an agent skill from llama-farm/llamafarm. Universal Runtime best practices for PyTorch inference, Transformers models, and FastAPI serving. Covers device management, model loading, memory optimization, and performance tuning.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `fastapi.md`, `performance.md` and `pytorch.md`).

It sits in AI & LLM Engineering, covering Deep learning and Backend development. It works with FastAPI, PyTorch and Python. The repository describes itself as: Deploy any AI model, agent, database, RAG, and pipeline locally or remotely in minutes. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Deep learning
  • Tasks that involve Backend development

Example prompts

  • “/runtime-skills”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Grep, Glob

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Model Loading with Double-Checked Locking
  2. Device-Aware Tensor Operations
  3. TTL-Based Model Caching
  4. Async Generation with Thread Pools

What it can do on your machine

Read from SKILL.md and the folder at commit 6244d46. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Runtime Skills loads about 1.3k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 210 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from llama-farm/llamafarm at commit 6244d46, republished under its Apache-2.0 licence (© llama-farm). 210 words, ~1,278 tokens.

Download SKILL.mdSave it as .claude/skills/runtime-skills/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
runtime-skills
description
Universal Runtime best practices for PyTorch inference, Transformers models, and FastAPI serving. Covers device management, model loading, memory optimization, and performance tuning.
allowed-tools
Read, Grep, Glob
user-invocable
false

Universal Runtime Skills

Best practices and code review checklists for the Universal Runtime - LlamaFarm's local ML inference server.

Overview

The Universal Runtime provides OpenAI-compatible endpoints for HuggingFace models:

  • Text generation (Causal LMs: GPT, Llama, Mistral, Qwen)
  • Text embeddings (BERT, sentence-transformers, ModernBERT)
  • Classification, NER, and reranking
  • OCR and document understanding
  • Anomaly detection

Directory: runtimes/universal/ Python: 3.11+ Key Dependencies: PyTorch, Transformers, FastAPI, llama-cpp-python

This skill extends the shared Python practices. Always apply these first:

Runtime-Specific Checklists

TopicFileKey Points
PyTorchpytorch.mdDevice management, dtype, memory cleanup
Transformerstransformers.mdModel loading, tokenization, inference
FastAPIfastapi.mdAPI design, streaming, lifespan
Performanceperformance.mdBatching, caching, optimizations

Architecture

runtimes/universal/
├── server.py              # FastAPI app, model caching, endpoints
├── core/
│   └── logging.py         # UniversalRuntimeLogger (structlog)
├── models/
│   ├── base.py            # BaseModel ABC with device management
│   ├── language_model.py  # Transformers text generation
│   ├── gguf_language_model.py  # llama-cpp-python for GGUF
│   ├── encoder_model.py   # Embeddings, classification, NER, reranking
│   └── ...                # OCR, anomaly, document models
├── routers/
│   └── chat_completions/  # Chat completions with streaming
├── utils/
│   ├── device.py          # Device detection (CUDA/MPS/CPU)
│   ├── model_cache.py     # TTL-based model caching
│   ├── model_format.py    # GGUF vs transformers detection
│   └── context_calculator.py  # GGUF context size computation
└── tests/

Key Patterns

1. Model Loading with Double-Checked Locking
python
_model_load_lock = asyncio.Lock()

async def load_encoder(model_id: str, task: str = "embedding"):
    cache_key = f"encoder:{task}:{model_id}"
    if cache_key not in _models:
        async with _model_load_lock:
            # Double-check after acquiring lock
            if cache_key not in _models:
                model = EncoderModel(model_id, device, task=task)
                await model.load()
                _models[cache_key] = model
    return _models.get(cache_key)
2. Device-Aware Tensor Operations
python
class BaseModel(ABC):
    def get_dtype(self, force_float32: bool = False):
        if force_float32:
            return torch.float32
        if self.device in ("cuda", "mps"):
            return torch.float16
        return torch.float32

    def to_device(self, tensor: torch.Tensor, dtype=None):
        # Don't change dtype for integer tensors
        if tensor.dtype in (torch.int32, torch.int64, torch.long):
            return tensor.to(device=self.device)
        dtype = dtype or self.get_dtype()
        return tensor.to(device=self.device, dtype=dtype)
3. TTL-Based Model Caching
python
_models: ModelCache[BaseModel] = ModelCache(ttl=300)  # 5 min TTL

async def _cleanup_idle_models():
    while True:
        await asyncio.sleep(CLEANUP_CHECK_INTERVAL)
        for cache_key, model in _models.pop_expired():
            await model.unload()
4. Async Generation with Thread Pools
python
# GGUF models use blocking llama-cpp, run in executor
self._executor = ThreadPoolExecutor(max_workers=1)

async def generate(self, messages, max_tokens=512, ...):
    loop = asyncio.get_running_loop()
    return await loop.run_in_executor(self._executor, self._generate_sync)

Review Priority

When reviewing Universal Runtime code:

  1. Critical - Security

    • Path traversal prevention in file endpoints
    • Input sanitization for model IDs
  2. High - Memory & Device

    • Proper CUDA/MPS cache clearing on unload
    • torch.no_grad() for inference
    • Correct dtype for device
  3. Medium - Performance

    • Model caching patterns
    • Batch processing where applicable
    • Streaming implementation
  4. Low - Code Style

    • Consistent with patterns.md
    • Proper type hints

© llama-farm, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in .claude/skills/runtime-skills of llama-farm/llamafarm.

  • SKILL.md
  • fastapi.md
  • performance.md
  • pytorch.md
  • transformers.md

Open the folder on GitHubat commit 6244d46

Compare with similar skills

Runtime Skills next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Runtime Skills compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Runtime Skills this skillllama-farm/llamafarm836—~1.3kAutomated safety check: PassApache-2.0
Benchmark Pyreflyfacebook/pyrefly7.1k—~1.8kAutomated safety check: PassMIT
Document Public APIspytorch/pytorch104k—~4.2kAutomated safety check: PassCustom licence
ExecuTorch Cortex-M Backendpytorch/executorch5.1k—~872Automated safety check: PassCustom licence
Cloudbase Agent PythonTencentCloudBase/CloudBase-AI-Toolkit1.1k2 repos~2.9kAutomated safety check: NotesMIT
Homepage Generatorwanshuiyin/ARIS-in-AI-Offer583—~4.8kAutomated safety check: NotesMIT

Similar skills

  • Benchmark Pyrefly

    facebook/pyrefly

    Official

    Run Pyrefly benchmarks locally via Buck or Cargo, including PyTorch real-world LSP benchmarks.

    7.1k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Document Public APIs

    pytorch/pytorch

    Document undocumented public APIs in PyTorch by removing functions from coverageignorefunctions and coverageignoreclasses in docs/source/conf.py, running Sphinx coverage, and adding the appropriate…

    104k GitHub stars~4.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • ExecuTorch Cortex-M Backend

    pytorch/executorch

    Developer guide for the Cortex-M (CMSIS-NN) backend in ExecuTorch: quantization pipeline, pass manager, tests and adding new ops.

    5.1k GitHub stars~872 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Cloudbase Agent Python

    TencentCloudBase/CloudBase-AI-Toolkit

    Build production-ready AI agent backends using the CloudBase Agent Python SDK — create agents with LangGraph/CrewAI/LlamaIndex, serve them via FastAPI with AG-UI protocol streaming +…

    1.1k GitHub starsUsed in 2 repos~2.9k tokens
    AI & LLM EngineeringAuto-check: notes
  • Homepage Generator

    wanshuiyin/ARIS-in-AI-Offer

    Generate a fact-checked academic personal homepage from a CV, optionally augmented by an existing manual homepage and an assets directory.

    583 GitHub stars~4.8k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check: notes
  • Ako4all

    TongmingLAIC/AKO4ALL

    Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.

    369 GitHub stars~4k tokensUpdated 24 days ago
    AI & LLM EngineeringAuto-check passed

More from llama-farm/llamafarm

All 19 skills in this repo
  • Reflect

    llama-farm/llamafarm

    Analyze the current session and propose improvements to skills.

    836 GitHub stars~1.5k tokensUpdated 4 mo ago
    Auto-check: notes
  • Temp Files

    llama-farm/llamafarm

    Guidelines for creating temporary files in system temp directory.

    836 GitHub stars~515 tokensUpdated 4 mo ago
    Auto-check: notes
  • CLI Skills

    llama-farm/llamafarm

    CLI best practices for LlamaFarm. An agent skill from llama-farm/llamafarm.

    836 GitHub stars~1.2k tokensUpdated 4 mo ago
    Auto-check passed
  • Code Review

    llama-farm/llamafarm

    Comprehensive code review for diffs. An agent skill from llama-farm/llamafarm.

    836 GitHub stars~2.3k tokensUpdated 4 mo ago
    Auto-check: notes
  • Commit Push PR

    llama-farm/llamafarm

    Commit changes, push to GitHub, and open a PR. An agent skill from llama-farm/llamafarm.

    836 GitHub stars~2.3k tokensUpdated 4 mo ago
    Auto-check: notes
  • Common Skills

    llama-farm/llamafarm

    Best practices for the Common utilities package in LlamaFarm.

    836 GitHub stars~885 tokensUpdated 4 mo ago
    Auto-check passed

Questions about Runtime Skills

What does Runtime Skills do?

Universal Runtime best practices for PyTorch inference, Transformers models, and FastAPI serving. Runtime Skills is an agent skill from llama-farm/llamafarm. Universal Runtime best practices for PyTorch inference, Transformers models, and FastAPI serving.

When should I use Runtime Skills?

Runtime Skills fits situations like: tasks that involve Deep learning; tasks that involve Backend development.

How do I install Runtime Skills in Claude Code?

Run `npx skills add llama-farm/llamafarm --skill runtime-skills -a claude-code`. Or copy the skill folder (.claude/skills/runtime-skills in llama-farm/llamafarm) into .claude/skills/runtime-skills in your project. Claude Code loads it when a task matches its description.

How do I install Runtime Skills in Codex?

Run `npx skills add llama-farm/llamafarm --skill runtime-skills -a codex`. Or copy the skill folder (.claude/skills/runtime-skills in llama-farm/llamafarm) into .agents/skills/runtime-skills in your project. Codex loads it when a task matches its description.

Can I use Runtime Skills in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add llama-farm/llamafarm --skill runtime-skills -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/runtime-skills, .gemini/skills/runtime-skills, .github/skills/runtime-skills and .opencode/skills/runtime-skills in your project.

What does Runtime Skills need to run?

SKILL.md names no scripts, command-line tools or credentials: Runtime Skills is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Grep, Glob.

Does Runtime Skills access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Runtime Skills safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Runtime Skills use?

Runtime Skills is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Runtime Skills use?

About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Runtime Skills?

Skills that share tags, products or a category with Runtime Skills: Benchmark Pyrefly (facebook/pyrefly, 7.1k stars), Document Public APIs (pytorch/pytorch, 104k stars), ExecuTorch Cortex-M Backend (pytorch/executorch, 5.1k stars) and Cloudbase Agent Python (TencentCloudBase/CloudBase-AI-Toolkit, 1.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Runtime Skills?

llama-farm (a GitHub organization) maintains it in llama-farm/llamafarm, which has 836 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on June 10, 2026.

Source: llama-farm/llamafarm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.