Benchmark Pyrefly
facebook/pyrefly
Run Pyrefly benchmarks locally via Buck or Cargo, including PyTorch real-world LSP benchmarks.
Universal Runtime best practices for PyTorch inference, Transformers models, and FastAPI serving.
$ npx skills add llama-farm/llamafarm --skill runtime-skills -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install llama-farm/llamafarm runtime-skills --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/llama-farm/llamafarm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/runtime-skills .claude/skills/runtime-skills && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "runtime-skills" agent skill from https://github.com/llama-farm/llamafarm/tree/main/.claude/skills/runtime-skills into .claude/skills/runtime-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "runtime-skills", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/llama-farm/llamafarm/tree/main/.claude/skills/runtime-skillsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add llama-farm/llamafarm --skill runtime-skills -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install llama-farm/llamafarm runtime-skills --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/llama-farm/llamafarm.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/runtime-skills .agents/skills/runtime-skills && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "runtime-skills" agent skill from https://github.com/llama-farm/llamafarm/tree/main/.claude/skills/runtime-skills into .agents/skills/runtime-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "runtime-skills", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add llama-farm/llamafarm --skill runtime-skills -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install llama-farm/llamafarm runtime-skills --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/llama-farm/llamafarm.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/runtime-skills .cursor/skills/runtime-skills && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "runtime-skills" agent skill from https://github.com/llama-farm/llamafarm/tree/main/.claude/skills/runtime-skills into .cursor/skills/runtime-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "runtime-skills", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/llama-farm/llamafarm.git --path .claude/skills/runtime-skills--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add llama-farm/llamafarm --skill runtime-skills -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install llama-farm/llamafarm runtime-skills --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/llama-farm/llamafarm.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/runtime-skills .gemini/skills/runtime-skills && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "runtime-skills" agent skill from https://github.com/llama-farm/llamafarm/tree/main/.claude/skills/runtime-skills into .gemini/skills/runtime-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "runtime-skills", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install llama-farm/llamafarm runtime-skillsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add llama-farm/llamafarm --skill runtime-skills -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/llama-farm/llamafarm.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/runtime-skills .github/skills/runtime-skills && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "runtime-skills" agent skill from https://github.com/llama-farm/llamafarm/tree/main/.claude/skills/runtime-skills into .github/skills/runtime-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "runtime-skills", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add llama-farm/llamafarm --skill runtime-skills -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install llama-farm/llamafarm runtime-skills --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/llama-farm/llamafarm.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/runtime-skills .opencode/skills/runtime-skills && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "runtime-skills" agent skill from https://github.com/llama-farm/llamafarm/tree/main/.claude/skills/runtime-skills into .opencode/skills/runtime-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "runtime-skills", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
runtime-skillsUniversal Runtime best practices for PyTorch inference, Transformers models, and FastAPI serving.
Runtime Skills is an agent skill from llama-farm/llamafarm. Universal Runtime best practices for PyTorch inference, Transformers models, and FastAPI serving. Covers device management, model loading, memory optimization, and performance tuning.
Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `fastapi.md`, `performance.md` and `pytorch.md`).
It sits in AI & LLM Engineering, covering Deep learning and Backend development. It works with FastAPI, PyTorch and Python. The repository describes itself as: Deploy any AI model, agent, database, RAG, and pipeline locally or remotely in minutes. The licence is Apache-2.0.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6244d46. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadGrepGlobFrom allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Runtime Skills loads about 1.3k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 210 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from llama-farm/llamafarm at commit 6244d46, republished under its Apache-2.0 licence (© llama-farm). 210 words, ~1,278 tokens.
.claude/skills/runtime-skills/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.Best practices and code review checklists for the Universal Runtime - LlamaFarm's local ML inference server.
The Universal Runtime provides OpenAI-compatible endpoints for HuggingFace models:
Directory: runtimes/universal/
Python: 3.11+
Key Dependencies: PyTorch, Transformers, FastAPI, llama-cpp-python
This skill extends the shared Python practices. Always apply these first:
| Topic | File | Priority |
|---|---|---|
| Patterns | python-skills/patterns.md | Medium |
| Async | python-skills/async.md | High |
| Typing | python-skills/typing.md | Medium |
| Testing | python-skills/testing.md | Medium |
| Errors | python-skills/error-handling.md | High |
| Security | python-skills/security.md | Critical |
| Topic | File | Key Points |
|---|---|---|
| PyTorch | pytorch.md | Device management, dtype, memory cleanup |
| Transformers | transformers.md | Model loading, tokenization, inference |
| FastAPI | fastapi.md | API design, streaming, lifespan |
| Performance | performance.md | Batching, caching, optimizations |
runtimes/universal/
├── server.py # FastAPI app, model caching, endpoints
├── core/
│ └── logging.py # UniversalRuntimeLogger (structlog)
├── models/
│ ├── base.py # BaseModel ABC with device management
│ ├── language_model.py # Transformers text generation
│ ├── gguf_language_model.py # llama-cpp-python for GGUF
│ ├── encoder_model.py # Embeddings, classification, NER, reranking
│ └── ... # OCR, anomaly, document models
├── routers/
│ └── chat_completions/ # Chat completions with streaming
├── utils/
│ ├── device.py # Device detection (CUDA/MPS/CPU)
│ ├── model_cache.py # TTL-based model caching
│ ├── model_format.py # GGUF vs transformers detection
│ └── context_calculator.py # GGUF context size computation
└── tests/_model_load_lock = asyncio.Lock()
async def load_encoder(model_id: str, task: str = "embedding"):
cache_key = f"encoder:{task}:{model_id}"
if cache_key not in _models:
async with _model_load_lock:
# Double-check after acquiring lock
if cache_key not in _models:
model = EncoderModel(model_id, device, task=task)
await model.load()
_models[cache_key] = model
return _models.get(cache_key)class BaseModel(ABC):
def get_dtype(self, force_float32: bool = False):
if force_float32:
return torch.float32
if self.device in ("cuda", "mps"):
return torch.float16
return torch.float32
def to_device(self, tensor: torch.Tensor, dtype=None):
# Don't change dtype for integer tensors
if tensor.dtype in (torch.int32, torch.int64, torch.long):
return tensor.to(device=self.device)
dtype = dtype or self.get_dtype()
return tensor.to(device=self.device, dtype=dtype)_models: ModelCache[BaseModel] = ModelCache(ttl=300) # 5 min TTL
async def _cleanup_idle_models():
while True:
await asyncio.sleep(CLEANUP_CHECK_INTERVAL)
for cache_key, model in _models.pop_expired():
await model.unload()# GGUF models use blocking llama-cpp, run in executor
self._executor = ThreadPoolExecutor(max_workers=1)
async def generate(self, messages, max_tokens=512, ...):
loop = asyncio.get_running_loop()
return await loop.run_in_executor(self._executor, self._generate_sync)When reviewing Universal Runtime code:
Critical - Security
High - Memory & Device
Medium - Performance
Low - Code Style
© llama-farm, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files in .claude/skills/runtime-skills of llama-farm/llamafarm.
Open the folder on GitHubat commit 6244d46
Runtime Skills next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Runtime Skills this skillllama-farm/llamafarm | 836 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Benchmark Pyreflyfacebook/pyrefly | 7.1k | — | ~1.8k | Automated safety check: Pass | MIT | |
| Document Public APIspytorch/pytorch | 104k | — | ~4.2k | Automated safety check: Pass | Custom licence | |
| ExecuTorch Cortex-M Backendpytorch/executorch | 5.1k | — | ~872 | Automated safety check: Pass | Custom licence | |
| Cloudbase Agent PythonTencentCloudBase/CloudBase-AI-Toolkit | 1.1k | 2 repos | ~2.9k | Automated safety check: Notes | MIT | |
| Homepage Generatorwanshuiyin/ARIS-in-AI-Offer | 583 | — | ~4.8k | Automated safety check: Notes | MIT |
facebook/pyrefly
Run Pyrefly benchmarks locally via Buck or Cargo, including PyTorch real-world LSP benchmarks.
pytorch/pytorch
Document undocumented public APIs in PyTorch by removing functions from coverageignorefunctions and coverageignoreclasses in docs/source/conf.py, running Sphinx coverage, and adding the appropriate…
pytorch/executorch
Developer guide for the Cortex-M (CMSIS-NN) backend in ExecuTorch: quantization pipeline, pass manager, tests and adding new ops.
TencentCloudBase/CloudBase-AI-Toolkit
Build production-ready AI agent backends using the CloudBase Agent Python SDK — create agents with LangGraph/CrewAI/LlamaIndex, serve them via FastAPI with AG-UI protocol streaming +…
wanshuiyin/ARIS-in-AI-Offer
Generate a fact-checked academic personal homepage from a CV, optionally augmented by an existing manual homepage and an assets directory.
TongmingLAIC/AKO4ALL
Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.
llama-farm/llamafarm
Analyze the current session and propose improvements to skills.
llama-farm/llamafarm
Guidelines for creating temporary files in system temp directory.
llama-farm/llamafarm
CLI best practices for LlamaFarm. An agent skill from llama-farm/llamafarm.
llama-farm/llamafarm
Comprehensive code review for diffs. An agent skill from llama-farm/llamafarm.
llama-farm/llamafarm
Commit changes, push to GitHub, and open a PR. An agent skill from llama-farm/llamafarm.
llama-farm/llamafarm
Best practices for the Common utilities package in LlamaFarm.
Categories
Universal Runtime best practices for PyTorch inference, Transformers models, and FastAPI serving. Runtime Skills is an agent skill from llama-farm/llamafarm. Universal Runtime best practices for PyTorch inference, Transformers models, and FastAPI serving.
Runtime Skills fits situations like: tasks that involve Deep learning; tasks that involve Backend development.
Run `npx skills add llama-farm/llamafarm --skill runtime-skills -a claude-code`. Or copy the skill folder (.claude/skills/runtime-skills in llama-farm/llamafarm) into .claude/skills/runtime-skills in your project. Claude Code loads it when a task matches its description.
Run `npx skills add llama-farm/llamafarm --skill runtime-skills -a codex`. Or copy the skill folder (.claude/skills/runtime-skills in llama-farm/llamafarm) into .agents/skills/runtime-skills in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add llama-farm/llamafarm --skill runtime-skills -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/runtime-skills, .gemini/skills/runtime-skills, .github/skills/runtime-skills and .opencode/skills/runtime-skills in your project.
SKILL.md names no scripts, command-line tools or credentials: Runtime Skills is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Grep, Glob.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Runtime Skills is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Runtime Skills: Benchmark Pyrefly (facebook/pyrefly, 7.1k stars), Document Public APIs (pytorch/pytorch, 104k stars), ExecuTorch Cortex-M Backend (pytorch/executorch, 5.1k stars) and Cloudbase Agent Python (TencentCloudBase/CloudBase-AI-Toolkit, 1.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
llama-farm (a GitHub organization) maintains it in llama-farm/llamafarm, which has 836 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on June 10, 2026.
Source: llama-farm/llamafarm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.