Agent skill

Turbofit

by SouthpawIN in SouthpawIN/turbofit

Operate the Turbofit adaptive local inference plugin. An agent skill from SouthpawIN/turbofit.

MITAuto-check passed

Install Turbofit

skills CLI
$ npx skills add SouthpawIN/turbofit --skill turbofit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install SouthpawIN/turbofit turbofit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
turbofit
GitHub stars
107
Token cost
~1.4k tokens
SKILL.md length
693 words
Files
616 (incl. scripts, references, assets)
Skills in repo
2
Repo updated
First seen
Licence
MIT

At a glance

Operate the Turbofit adaptive local inference plugin. An agent skill from SouthpawIN/turbofit.

  • Works in 8 steps: Call turbofit_status to inspect provider… → Call turbofit_configure with profile:… → Set primary: true to use custom:turbofit… → …
  • SKILL.md covers 2.3 model authority, Operator workflow, Intelligence benchmarks and Portable memory allocation, plus 1 more section
  • Runs Python scripts from its folder

What it does

Turbofit is an agent skill from SouthpawIN/turbofit. Operate the Turbofit adaptive local inference plugin.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 618 other files, including scripts, reference files and assets (for example `.github/workflows/model-intelligence.yml`, `CHANGELOG.md` and `README.md`).

The repository describes itself as: Adaptive local inference for Hermes — TurboFit Check, evidence-built TurboFit List, repaired discovery/benchmark campaigns, Qwen 3.8 DFlash2, Bonsai DSpark, FreeToken, and Sirvir. The licence is MIT.

Example prompts

  • “/turbofit”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Call turbofit_status to inspect provider registration, gateway health, selected hardware profile, active rung, and stable routes.
  2. Call turbofit_configure with profile: auto for hardware selection. Manual hardware-*gb profiles are accepted only when physical topology…
  3. Set primary: true to use custom:turbofit with model auto as the main Hermes provider.
  4. Set fallback: true to append Turbofit to the canonical fallback_providers chain; set it false to remove only Turbofit while preserving…
  5. Set publish_tailnet: true to create private Tailscale Serve routes for the provider and dashboard; the returned HTTPS provider URL is…
  6. Set install_sirvir: true (alias install_turbosouth: true) to install or update the canonical SouthpawIN/turbosouth GitHub-current profile…
  7. Set install_freetoken: true only on Linux x86_64 + NVIDIA driver 580+ + CUDA toolkit 13+ to install pinned FreeToken 0.1.2 as a text-only…
  8. Start a new Hermes session after provider changes.

What it can do on your machine

Read from SKILL.md and the folder at commit 42fc223. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python, from the files we listed), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Turbofit loads about 1.4k tokens when it runs, and up to ~651k if it reads all its reference files. Until then it costs about 16 tokens; SKILL.md has 693 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~16
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~651k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from SouthpawIN/turbofit at commit 42fc223, republished under its MIT licence (© SouthpawIN). 693 words, ~1,398 tokens.

Download SKILL.mdSave it as .claude/skills/turbofit/SKILL.md (or your agent's skills folder). This skill also uses 615 other files; get the full folder from GitHub.
name
turbofit
description
Operate the Turbofit adaptive local inference plugin.
version
2.3.1
author
SouthpawIN + Nous Girl
license
MIT
tags
hermes-agent, plugin, llama-cpp, adaptive-runtime

Turbofit

2.3 model authority

Fit List keeps dedicated VRAM separate from integrated/RAM-only total memory. Dedicated: Maple Preview TQ2_0 at 8GB, Qwen 3.8 27B Unleashed UD-IQ3_XXS at 16GB, UD-Q3_K_XL at 24–95GB, and Qwen 3.8 27B 16-bit at 96GB+ until an Unleashed FP16 GGUF is published. Shared total memory: Maple at 8–15GB, Ornith 1.5 35A3B at 16–23GB, and Unleashed UD-Q3_K_XL at 24GB+. An 8GB GPU may also use Ornith when host RAM can hold offloaded experts. Maple remains Auto on dedicated 8GB at native 64K/128K. Auxiliary is Ornith, optional Carwin Nano, or auto.

Use this bundled plugin skill when configuring or inspecting Turbofit for Hermes Agent.

Operator workflow

  1. Call turbofit_status to inspect provider registration, gateway health, selected hardware profile, active rung, and stable routes.
  2. Call turbofit_configure with profile: auto for hardware selection. Manual hardware-*gb profiles are accepted only when physical topology fits.
  3. Set primary: true to use custom:turbofit with model auto as the main Hermes provider.
  4. Set fallback: true to append Turbofit to the canonical fallback_providers chain; set it false to remove only Turbofit while preserving other fallbacks.
  5. Set publish_tailnet: true to create private Tailscale Serve routes for the provider and dashboard; the returned HTTPS provider URL is registered automatically.
  6. Set install_sirvir: true (alias install_turbosouth: true) to install or update the canonical SouthpawIN/turbosouth GitHub-current profile without replacing its memories or user state — TurboSouth sends tested pull requests upstream to TurboFit.
  7. Set install_freetoken: true only on Linux x86_64 + NVIDIA driver 580+ + CUDA toolkit 13+ to install pinned FreeToken 0.1.2 as a text-only MoE candidate. It never changes Auto until exact on-box campaigns promote a supported model recipe.
  8. Start a new Hermes session after provider changes.

The same controls are available in Hermes Desktop under Turbofit and through /turbofit status|update|shift|serve|tiers|setup.

/turbofit setup refreshes Hermes Desktop. A refused http://127.0.0.1:8091/v1/models means the Turbofit runtime is down — not the Hermes messaging gateway. Diagnose that from Sirvir or Desktop with turbofit_status. Setup downloads recommended models if they are missing.

Intelligence benchmarks

  • scripts/turbofit-catalog-campaign proves native runtime fit and TPS; it does not produce intelligence scores.
  • scripts/turbofit-intelligence-campaign runs the exact successful quantized production recipe through pinned DeepSWE and the Turbofit agentic main/auxiliary pair harness.
  • Use status, run-one, or run --limit N; state is resumable in references/intelligence-campaign-state.json.
  • Scores require both benchmark suites and immutable raw evidence. Never replace missing scores with catalog tiers, parameter counts, or vendor benchmark claims.
  • /turbofit tiers and scripts/turbofit-hardware-tiers show every 8/16/24/32/48/64/96/128/192/256/384 GB class with pending versus measured intelligence and TPS.
  • scripts/turbofit-intelligence-campaign benchmarks only the current machine's TurboFit List tournament candidates. rebuild-scores recomputes derived composites from raw suite counts; zero-call/token trials remain invalid infrastructure.
  • scripts/turbofit-promote-list-winner promotes only an exact-tier candidate with current physical evidence, positive intelligence/TPS/balanced values, and matching recipe hashes. scripts/turbofit-list renders the global evidence-only List.
  • Qwen 3.8 DFlash2 is a separate candidate runtime/artifact pair (dflash2-llama.cpp, Qwen3.8-27B-DFlash2-Q4_K_M.gguf). Never attach that drafter to Bonsai. Bonsai uses its own released DSpark sidecar and Prism runtime until a dedicated Bonsai DFlash checkpoint exists.
Show full SKILL.md (201 more words)Show less

Portable memory allocation

  • Hardware fingerprints classify memory as dedicated, unified, or cpu and reserve 5% of host RAM, bounded to 1–8 GiB.
  • Dedicated systems may combine accelerator VRAM with host RAM through llama.cpp offload; contexts beyond the model's native window place KV cache in host RAM when at least 32 GiB is usable.
  • Unified-memory systems count RAM once and suppress discrete multi-GPU split flags.
  • CPU-only systems set model and draft GPU layers to zero.
  • Native backend order is CUDA, ROCm, Vulkan, then CPU on Linux/Windows, and Metal on macOS. Use scripts/install-native-runtimes --backend <backend> for an explicit build.

Invariants

  • Stable model IDs are auto, active:main, and active:aux.
  • FreeToken support is candidate-only: no active Qwen 3.8/Ornith replacement, no source TPS inheritance, and no Auto promotion without exact hardware/model evidence.
  • External GPU processes are read-only pressure signals and are never terminated or signaled.
  • The hardware recommendation remains the healing ceiling; transient pressure changes only the effective rung.
  • Runtime activation and model lifecycle remain owned by NativeRuntimeBackend, which signals only PID-verified children it launched.
  • Plain HTTP provider endpoints are limited to loopback or Tailscale addresses; all other endpoints require HTTPS.
  • Every native llama.cpp command includes --jinja; DSpark variants include target, draft, projector, and draft-attention arguments.

© SouthpawIN, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 615 other files (scripts, references, assets) in the repository root of SouthpawIN/turbofit.

  • SKILL.md
  • .gitattributes
  • .github/workflows/model-intelligence.yml
  • .gitignore
  • CHANGELOG.md
  • LICENSE
  • README.md
  • __init__.py
  • after-install.md
  • assets/hermes-desktop-turbofit-settings.png
  • assets/provider-integration.png
  • assets/scaling-ladder.png
  • assets/turbofit-2.2-model-ladder.png
  • assets/turbofit-2.2-model-ladder.svg
  • assets/turbofit-2.2-multimodal.png
  • assets/turbofit-2.2-multimodal.svg
  • assets/turbofit-2.2-qwen-lineup.png
  • assets/turbofit-2.2-qwen-lineup.svg
  • … and 598 more

Open the folder on GitHubat commit 42fc223

Compare with similar skills

Turbofit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Turbofit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Turbofit this skillSouthpawIN/turbofit107—~1.4kAutomated safety check: PassMIT
Ito Inferenceaffaan-m/ECC277k1 repos~1.5kAutomated safety check: PassMIT
It Operationsdavila7/claude-code-templates33k1 repos~3.7kAutomated safety check: PassMIT
Gke Inferencegoogle/skills21k—~2kAutomated safety check: PassApache-2.0
LLM Inference Scalingsickn33/agentic-awesome-skills47k1 repos~2.1kAutomated safety check: PassMIT
Debug InferenceNVIDIA/OpenShell16k—~1.9kAutomated safety check: PassApache-2.0

Similar skills

  • Ito Inference

    affaan-m/ECC

    Inspect the availability of model serving on a completed Itô compute booking and, when the canonical backend becomes available, hand off an explicitly confirmed serving manifest.

    277k GitHub starsUsed in 1 repo~1.5k tokens
    AI & LLM EngineeringAuto-check passed
  • It Operations

    davila7/claude-code-templates

    Manages IT infrastructure, monitoring, incident response, and service reliability.

    33k GitHub starsUsed in 1 repo~3.7k tokens
    DevOps & CloudAuto-check passed
  • Gke Inference

    google/skills

    Official

    Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers.

    21k GitHub stars~2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • LLM Inference Scaling

    sickn33/agentic-awesome-skills

    Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

    47k GitHub starsUsed in 1 repo~2.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Debug Inference

    NVIDIA/OpenShell

    Official

    Debug inference clients that use an attached provider and its native endpoint, including hosted APIs and host-local Ollama, vLLM, SGLang, TRT-LLM, LM Studio, or NIM.

    16k GitHub stars~1.9k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Agent skill for adaptive-coordinator - invoke with $agent-adaptive-coordinator

    74k GitHub starsUsed in 2 repos~4k tokens
    Data & AnalyticsAuto-check passed

More from SouthpawIN/turbofit

  • Turbofit

    SouthpawIN/turbofit

    Hardware-aware adaptive Hermes runtime using portable Turbofiles, total usable memory, owned native llama.cpp residency, stable auto/active:main/active:aux routes, and evidence-backed promotion.

    107 GitHub stars~1.9k tokensUpdated 2 days ago
    Auto-check passed

Questions about Turbofit

What does Turbofit do?

Operate the Turbofit adaptive local inference plugin. An agent skill from SouthpawIN/turbofit. Turbofit is an agent skill from SouthpawIN/turbofit. Operate the Turbofit adaptive local inference plugin.

How do I install Turbofit in Claude Code?

Run `npx skills add SouthpawIN/turbofit --skill turbofit -a claude-code`. Or copy the skill folder (the SouthpawIN/turbofit repository) into .claude/skills/turbofit in your project. Claude Code loads it when a task matches its description.

How do I install Turbofit in Codex?

Run `npx skills add SouthpawIN/turbofit --skill turbofit -a codex`. Or copy the skill folder (the SouthpawIN/turbofit repository) into .agents/skills/turbofit in your project. Codex loads it when a task matches its description.

Can I use Turbofit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add SouthpawIN/turbofit --skill turbofit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/turbofit, .gemini/skills/turbofit, .github/skills/turbofit and .opencode/skills/turbofit in your project.

What does Turbofit need to run?

Going by SKILL.md and its folder, Turbofit needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Turbofit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Turbofit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Turbofit use?

Turbofit is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Turbofit use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 650k tokens, read only when the agent opens those files.

What are the alternatives to Turbofit?

Skills that share tags, products or a category with Turbofit: Ito Inference (affaan-m/ECC, 277k stars), It Operations (davila7/claude-code-templates, 33k stars), Gke Inference (google/skills, 21k stars) and LLM Inference Scaling (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Turbofit?

SouthpawIN (a GitHub user) maintains it in SouthpawIN/turbofit, which has 107 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 8, 2026.

Source: SouthpawIN/turbofit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.