Agent skill

Multi Tenant LLM Hosting

by sickn33 in sickn33/agentic-awesome-skills

Design secure, multi-tenant LLM hosting platforms with tenant isolation, quotas, billing attribution, noisy-neighbor protection, and per-tenant policy controls.

MITAuto-check passedBackend & APIs

Install Multi Tenant LLM Hosting

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill multi-tenant-llm-hosting -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills multi-tenant-llm-hosting --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/multi-tenant-llm-hosting .claude/skills/multi-tenant-llm-hosting && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
multi-tenant-llm-hosting
GitHub stars
47k
Used in
2 other repos
Token cost
~3.1k tokens
SKILL.md length
180 words
Files
2 (incl. references)
Skills in repo
1,497
Repo updated
First seen
Licence
MIT

At a glance

Design secure, multi-tenant LLM hosting platforms with tenant isolation, quotas, billing attribution, noisy-neighbor protection, and per-tenant policy controls.

  • Tasks that involve Multi-tenancy
  • SKILL.md covers Prerequisites, Isolation Model, vLLM Multi-Model Serving and Per-Tenant Quota Configuration, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Multi Tenant LLM Hosting is an agent skill from sickn33/agentic-awesome-skills. Design secure, multi-tenant LLM hosting platforms with tenant isolation, quotas, billing attribution, noisy-neighbor protection, and per-tenant policy controls.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/details.md`). Compatibility notes: Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.

It sits in Backend & APIs, covering Multi-tenancy. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Tasks that involve Multi-tenancy

Example prompts

  • “/multi-tenant-llm-hosting”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.

What it can do on your machine

Read from SKILL.md and the folder at commit 1c7bdea. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.

    From compatibility in the SKILL.md frontmatter.

Context cost

Multi Tenant LLM Hosting loads about 3.1k tokens when it runs, and up to ~4.6k if it reads all its reference files. Until then it costs about 46 tokens; SKILL.md has 180 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit 1c7bdea, republished under its MIT licence (© sickn33). 180 words, ~3,091 tokens.

Download SKILL.mdSave it as .claude/skills/multi-tenant-llm-hosting/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
multi-tenant-llm-hosting
description
Design secure, multi-tenant LLM hosting platforms with tenant isolation, quotas, billing attribution, noisy-neighbor protection, and per-tenant policy controls.
compatibility
Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.
category
devops
risk
critical
source
https://github.com/BagelHole/DevOps-Security-Agent-Skills
source_repo
BagelHole/DevOps-Security-Agent-Skills
source_type
community
date_added
2026-09-20
license
MIT
license_source
https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
metadata.author
devops-skills
metadata.version
1.0

Multi-Tenant LLM Hosting

Host many teams/customers on shared inference infrastructure without sacrificing security, performance, or cost governance.

Prerequisites

  • Kubernetes cluster with GPU node pools
  • API gateway or LLM gateway (LiteLLM, Envoy, Kong)
  • Prometheus + Grafana for per-tenant observability
  • Redis or equivalent for rate limiting state
  • Billing system or cost attribution database

Isolation Model

  • Strong tenant identity on every request
  • Per-tenant API keys and scoped model access
  • Namespace or workload isolation for high-risk tenants
  • Strict data retention and log partitioning controls

vLLM Multi-Model Serving

yaml
# vllm-deployment.yaml - Multi-model serving with vLLM
apiVersion: apps/v1
kind: Deployment
metadata:
  name: vllm-gpt4o-equivalent
  namespace: llm-serving
  labels:
    app: vllm
    model-tier: premium
spec:
  replicas: 3
  selector:
    matchLabels:
      app: vllm
      model-tier: premium
  template:
    metadata:
      labels:
        app: vllm
        model-tier: premium
      annotations:
        prometheus.io/scrape: "true"
        prometheus.io/port: "8080"
    spec:
      containers:
        - name: vllm
          image: vllm/vllm-openai:v0.4.1
          args:
            - "--model=/models/llama-3.1-70b"
            - "--tensor-parallel-size=2"
            - "--max-model-len=8192"
            - "--gpu-memory-utilization=0.90"
            - "--max-num-seqs=128"
            - "--enable-prefix-caching"
          ports:
            - containerPort: 8000
              name: inference
            - containerPort: 8080
              name: metrics
          resources:
            requests:
              nvidia.com/gpu: 2
              cpu: "8"
              memory: "64Gi"
            limits:
              nvidia.com/gpu: 2
              cpu: "16"
              memory: "128Gi"
          volumeMounts:
            - name: model-weights
              mountPath: /models
              readOnly: true
      volumes:
        - name: model-weights
          persistentVolumeClaim:
            claimName: premium-model-weights
      tolerations:
        - key: nvidia.com/gpu
          operator: Exists
          effect: NoSchedule
      nodeSelector:
        gpu-type: a100
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: vllm-economy
  namespace: llm-serving
  labels:
    app: vllm
    model-tier: economy
spec:
  replicas: 2
  selector:
    matchLabels:
      app: vllm
      model-tier: economy
  template:
    metadata:
      labels:
        app: vllm
        model-tier: economy
    spec:
      containers:
        - name: vllm
          image: vllm/vllm-openai:v0.4.1
          args:
            - "--model=/models/llama-3.1-8b"
            - "--max-model-len=4096"
            - "--gpu-memory-utilization=0.85"
            - "--max-num-seqs=256"
            - "--enable-prefix-caching"
          ports:
            - containerPort: 8000
              name: inference
            - containerPort: 8080
              name: metrics
          resources:
            requests:
              nvidia.com/gpu: 1
              cpu: "4"
              memory: "32Gi"
            limits:
              nvidia.com/gpu: 1
              cpu: "8"
              memory: "64Gi"
          volumeMounts:
            - name: model-weights
              mountPath: /models
              readOnly: true
      volumes:
        - name: model-weights
          persistentVolumeClaim:
            claimName: economy-model-weights
      tolerations:
        - key: nvidia.com/gpu
          operator: Exists
          effect: NoSchedule

Per-Tenant Quota Configuration

yaml
# tenant-quotas-configmap.yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: tenant-quotas
  namespace: llm-serving
data:
  quotas.yaml: |
    tenants:
      acme-corp:
        tier: enterprise
        models_allowed:
          - llama-3.1-70b
          - llama-3.1-8b
          - nomic-embed-text
        rate_limits:
          requests_per_minute: 300
          tokens_per_minute: 500000
          concurrent_requests: 50
        budget:
          daily_limit_usd: 500.00
          monthly_limit_usd: 10000.00
          alert_threshold_percent: 80
        priority: high

      startup-xyz:
        tier: standard
        models_allowed:
          - llama-3.1-8b
          - nomic-embed-text
        rate_limits:
          requests_per_minute: 60
          tokens_per_minute: 100000
          concurrent_requests: 10
        budget:
          daily_limit_usd: 50.00
          monthly_limit_usd: 1000.00
          alert_threshold_percent: 80
        priority: medium

      internal-dev:
        tier: free
        models_allowed:
          - llama-3.1-8b
        rate_limits:
          requests_per_minute: 20
          tokens_per_minute: 50000
          concurrent_requests: 5
        budget:
          daily_limit_usd: 10.00
          monthly_limit_usd: 200.00
          alert_threshold_percent: 90
        priority: low

Namespace Isolation for High-Risk Tenants

yaml
# tenant-namespace.yaml
apiVersion: v1
kind: Namespace
metadata:
  name: tenant-acme-corp
  labels:
    tenant: acme-corp
    isolation: strict
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: tenant-isolation
  namespace: tenant-acme-corp
spec:
  podSelector: {}
  policyTypes:
    - Ingress
    - Egress
  ingress:
    - from:
        - namespaceSelector:
            matchLabels:
              name: llm-gateway
  egress:
    - to:
        - namespaceSelector:
            matchLabels:
              name: llm-serving
      ports:
        - port: 8000
          protocol: TCP
    - to:
        - namespaceSelector:
            matchLabels:
              name: kube-dns
      ports:
        - port: 53
          protocol: UDP
---
apiVersion: v1
kind: ResourceQuota
metadata:
  name: tenant-quota
  namespace: tenant-acme-corp
spec:
  hard:
    requests.cpu: "16"
    requests.memory: "64Gi"
    limits.cpu: "32"
    limits.memory: "128Gi"
    requests.nvidia.com/gpu: "4"
    pods: "20"

Request Routing and Rate Limiting

python
# gateway_router.py
"""Multi-tenant request router with rate limiting and model routing."""
import time
import json
import redis
from fastapi import FastAPI, HTTPException, Header, Request
from typing import Optional
import httpx
import yaml

app = FastAPI()
redis_client = redis.Redis(host="redis", port=6379, decode_responses=True)

# Load tenant config
with open("/etc/config/quotas.yaml") as f:
    TENANT_CONFIG = yaml.safe_load(f)["tenants"]

MODEL_ENDPOINTS = {
    "llama-3.1-70b": "http://vllm-gpt4o-equivalent:8000",
    "llama-3.1-8b": "http://vllm-economy:8000",
    "nomic-embed-text": "http://embedding-service:8000",
}

def check_rate_limit(tenant_id: str, config: dict) -> bool:
    """Check and update rate limit for a tenant."""
    key = f"ratelimit:{tenant_id}:{int(time.time() // 60)}"
    current = redis_client.incr(key)
    if current == 1:
        redis_client.expire(key, 120)
    return current <= config["rate_limits"]["requests_per_minute"]

def check_concurrent(tenant_id: str, config: dict) -> bool:
    """Check concurrent request limit."""
    key = f"concurrent:{tenant_id}"
    current = int(redis_client.get(key) or 0)
    return current < config["rate_limits"]["concurrent_requests"]

def check_budget(tenant_id: str, config: dict) -> bool:
    """Check if tenant is within daily budget."""
    key = f"spend:{tenant_id}:{time.strftime('%Y-%m-%d')}"
    current_spend = float(redis_client.get(key) or 0)
    return current_spend < config["budget"]["daily_limit_usd"]

def record_usage(tenant_id: str, model: str, prompt_tokens: int, completion_tokens: int):
    """Record token usage and cost for billing."""
    # Cost rates per 1K tokens
    rates = {
        "llama-3.1-70b": {"prompt": 0.004, "completion": 0.012},
        "llama-3.1-8b": {"prompt": 0.0005, "completion": 0.0015},
        "nomic-embed-text": {"prompt": 0.0001, "completion": 0.0},
    }
    rate = rates.get(model, {"prompt": 0.001, "completion": 0.003})
    cost = (prompt_tokens * rate["prompt"] + completion_tokens * rate["completion"]) / 1000

    # Update daily spend
    spend_key = f"spend:{tenant_id}:{time.strftime('%Y-%m-%d')}"
    redis_client.incrbyfloat(spend_key, cost)
    redis_client.expire(spend_key, 172800)

    # Record for billing export
    billing_key = f"billing:{tenant_id}:{time.strftime('%Y-%m')}"
    redis_client.rpush(billing_key, json.dumps({
        "timestamp": time.time(),
        "model": model,
        "prompt_tokens": prompt_tokens,
        "completion_tokens": completion_tokens,
        "cost_usd": cost,
    }))

@app.post("/v1/chat/completions")
async def chat_completions(
    request: Request,
    x_tenant_id: str = Header(...),
    x_api_key: str = Header(...),
):
    """Route chat completion request with tenant controls."""
    if x_tenant_id not in TENANT_CONFIG:
        raise HTTPException(status_code=403, detail="Unknown tenant")

    config = TENANT_CONFIG[x_tenant_id]
    body = await request.json()
    model = body.get("model", "llama-3.1-8b")

    # Check model access
    if model not in config["models_allowed"]:
        raise HTTPException(status_code=403, detail=f"Model {model} not allowed for tenant")

    # Check rate limit
    if not check_rate_limit(x_tenant_id, config):
        raise HTTPException(status_code=429, detail="Rate limit exceeded")

    # Check concurrent requests
    if not check_concurrent(x_tenant_id, config):
        raise HTTPException(status_code=429, detail="Concurrent request limit exceeded")

    # Check budget
    if not check_budget(x_tenant_id, config):
        raise HTTPException(status_code=402, detail="Daily budget exceeded")

    # Route to model endpoint
    endpoint = MODEL_ENDPOINTS.get(model)
    if not endpoint:
        raise HTTPException(status_code=404, detail=f"Model {model} not available")

    # Track concurrent requests
    concurrent_key = f"concurrent:{x_tenant_id}"
    redis_client.incr(concurrent_key)

    try:
        async with httpx.AsyncClient(timeout=120.0) as client:
            response = await client.post(
                f"{endpoint}/v1/chat/completions",
                json=body,
                headers={"Content-Type": "application/json"},
            )
            result = response.json()

            # Record usage
            usage = result.get("usage", {})
            record_usage(
                x_tenant_id, model,
                usage.get("prompt_tokens", 0),
                usage.get("completion_tokens", 0),
            )

            return result
    finally:
        redis_client.decr(concurrent_key)

Contents

When to Use This Skill

  • Building an internal LLM platform shared by multiple teams
  • Hosting LLM inference for external customers with isolation requirements
  • Implementing per-tenant quotas, billing, and rate limiting
  • Designing request routing for multi-model, multi-tenant environments
  • Preventing noisy-neighbor issues on shared GPU infrastructure

Limitations

  • Infrastructure commands can disrupt services: confirm target host/scope and have backups/snapshots before mutating state.
  • Docs-only import: upstream scripts and templates not bundled.

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/multi-tenant-llm-hosting of sickn33/agentic-awesome-skills.

  • SKILL.md
  • references/details.md

Open the folder on GitHubat commit 1c7bdea

Used in 2 other repositories

We found 6 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 9, 2026.

Compare with similar skills

Multi Tenant LLM Hosting next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Multi Tenant LLM Hosting compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Multi Tenant LLM Hosting this skillsickn33/agentic-awesome-skills47k2 repos~3.1kAutomated safety check: PassMIT
PR Review Provideryansongda/pay5.4k—~2.4kAutomated safety check: PassMIT
Abp Authorizationabpframework/abp14k—~1.3kAutomated safety check: PassLGPL-3.0
Django Access Reviewgetsentry/skills1k3 repos~2.6kAutomated safety check: NotesApache-2.0
Backend Dev Guidelineslitefuse/litefuse1011 repos~5.8kAutomated safety check: PassCustom licence
Ecto Persistence Patternsgeorgeguimaraes/elixir-agent-tools184—~1.2kAutomated safety check: PassApache-2.0

Similar skills

  • PR Review Provider

    yansongda/pay

    A skill your agent uses when reviewing PRs that add or modify a payment Provider in yansongda/pay - covers plugin pipeline, multi-tenant safety, signature verification, docs, and naming conventions.

    5.4k GitHub stars~2.4k tokensUpdated 2 days ago
    Backend & APIsAuto-check passed
  • Abp Authorization

    abpframework/abp

    ABP permission system - PermissionDefinitionProvider, [Authorize] attribute, CheckPolicyAsync, IsGrantedAsync, ICurrentUser, IPermissionManager, multi-tenancy side.

    14k GitHub stars~1.3k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Django Access Review

    getsentry/skills

    Official

    Django access control and IDOR security review. An agent skill from getsentry/skills.

    1k GitHub starsUsed in 3 repos~2.6k tokens
    Backend & APIsAuto-check: notes
  • Backend Dev Guidelines

    litefuse/litefuse

    Comprehensive backend development guide for Litefuse's Next.js 14/tRPC/Express/TypeScript monorepo.

    101 GitHub starsUsed in 1 repo~5.8k tokens
    Backend & APIsAuto-check passed
  • Ecto Persistence Patterns

    georgeguimaraes/elixir-agent-tools

    Designs and debugs Elixir persistence with Ecto: schemas, changesets, queries, preloads, migrations and multi-tenancy, kept within application contexts.

    184 GitHub stars~1.2k tokensUpdated 21 days ago
    Backend & APIsAuto-check passed
  • Rails Security Multitenancy

    marckohlbrugge/37signals-skills

    Apply Rails security and multi-tenant safety practices including scoped queries, SSRF defenses, rate limiting, and tenant-scoped realtime updates.

    724 GitHub stars~1.7k tokensUpdated 4 mo ago
    Backend & APIsAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,497 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Categories

Questions about Multi Tenant LLM Hosting

What does Multi Tenant LLM Hosting do?

Design secure, multi-tenant LLM hosting platforms with tenant isolation, quotas, billing attribution, noisy-neighbor protection, and per-tenant policy controls. Multi Tenant LLM Hosting is an agent skill from sickn33/agentic-awesome-skills. Design secure, multi-tenant LLM hosting platforms with tenant isolation, quotas, billing attribution, noisy-neighbor protection, and per-tenant policy controls.

When should I use Multi Tenant LLM Hosting?

Multi Tenant LLM Hosting fits situations like: tasks that involve Multi-tenancy.

How do I install Multi Tenant LLM Hosting in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill multi-tenant-llm-hosting -a claude-code`. Or copy the skill folder (skills/multi-tenant-llm-hosting in sickn33/agentic-awesome-skills) into .claude/skills/multi-tenant-llm-hosting in your project. Claude Code loads it when a task matches its description.

How do I install Multi Tenant LLM Hosting in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill multi-tenant-llm-hosting -a codex`. Or copy the skill folder (skills/multi-tenant-llm-hosting in sickn33/agentic-awesome-skills) into .agents/skills/multi-tenant-llm-hosting in your project. Codex loads it when a task matches its description.

Can I use Multi Tenant LLM Hosting in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill multi-tenant-llm-hosting -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/multi-tenant-llm-hosting, .gemini/skills/multi-tenant-llm-hosting, .github/skills/multi-tenant-llm-hosting and .opencode/skills/multi-tenant-llm-hosting in your project.

What does Multi Tenant LLM Hosting need to run?

SKILL.md names no scripts, command-line tools or credentials: Multi Tenant LLM Hosting is instructions for the agent only. Our summary lists: Python 3. Compatibility (from SKILL.md): Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled..

Does Multi Tenant LLM Hosting access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Multi Tenant LLM Hosting safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Multi Tenant LLM Hosting use?

Multi Tenant LLM Hosting is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Multi Tenant LLM Hosting use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.5k tokens, read only when the agent opens those files.

What are the alternatives to Multi Tenant LLM Hosting?

Skills that share tags, products or a category with Multi Tenant LLM Hosting: PR Review Provider (yansongda/pay, 5.4k stars), Abp Authorization (abpframework/abp, 14k stars), Django Access Review (getsentry/skills, 1k stars) and Backend Dev Guidelines (litefuse/litefuse, 101 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Multi Tenant LLM Hosting?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,443 GitHub stars. The repository holds 1,497 skills in this directory. The repository was last updated on October 10, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.