Agent skill

Turbofit

by SouthpawIN in SouthpawIN/turbofit

Hardware-aware adaptive Hermes runtime using portable Turbofiles, total usable memory, owned native llama.cpp residency, stable auto/active:main/active:aux routes, and evidence-backed promotion.

MITAuto-check passedAI & LLM Engineering

Install Turbofit

skills CLI
$ npx skills add SouthpawIN/turbofit --skill turbofit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install SouthpawIN/turbofit turbofit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/SouthpawIN/turbofit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/turbofit .claude/skills/turbofit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
turbofit
GitHub stars
107
Token cost
~1.9k tokens
SKILL.md length
819 words
Files
32 (incl. scripts, references)
Skills in repo
2
Repo updated
First seen
Licence
MIT

At a glance

Hardware-aware adaptive Hermes runtime using portable Turbofiles, total usable memory, owned native llama.cpp residency, stable auto/active:main/active:aux routes, and evidence-backed promotion.

  • Works in 7 steps: Turbofile: portable recommendation and… → Hardware fingerprint: physical topology,… → Pressure snapshot: ownership-aware… → …
  • Troubleshooting Turbofit runtimes
  • SKILL.md covers 2.4 model authority, Use when, Canonical workflow and Runtime authorities, plus 7 more sections
  • Calls python3 and curl

What it does

Turbofit is an agent skill from SouthpawIN/turbofit. Hardware-aware adaptive Hermes runtime using portable Turbofiles, total usable memory, owned native llama.cpp residency, stable auto/active:main/active:aux routes, and evidence-backed promotion. Use for recommending, activating, inspecting, testing, or troubleshooting Turbofit runtimes.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 32 other files, including scripts and reference files (for example `README.md`, `distribution.yaml` and `references/SOUL.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with llama.cpp and Qwen. The repository describes itself as: Adaptive local inference for Hermes — TurboFit Check, evidence-built TurboFit List, repaired discovery/benchmark campaigns, Qwen 3.8 DFlash2, Bonsai DSpark, FreeToken, and Sirvir. The licence is MIT.

When your agent uses it

  • Troubleshooting Turbofit runtimes
  • Tasks that involve LLM inference and serving

Example prompts

  • “/turbofit”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Turbofile: portable recommendation and ordered rung policy.
  2. Hardware fingerprint: physical topology, total usable memory, and per-device capacity; never transient free memory.
  3. Pressure snapshot: ownership-aware transient capacity.
  4. Pure policy: dwell/hysteresis/cooldown/flap decision.
  5. Native runtime backend: sole local residency authority for owned processes.
  6. Reconciler: drain, activate, verify, publish, rollback.
  7. Gateway route state: backing targets for stable IDs.

What it can do on your machine

Read from SKILL.md and the folder at commit 42fc223. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Turbofit loads about 1.9k tokens when it runs, and up to ~37k if it reads all its reference files. Until then it costs about 74 tokens; SKILL.md has 819 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~37k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from SouthpawIN/turbofit at commit 42fc223, republished under its MIT licence (© SouthpawIN). 819 words, ~1,944 tokens.

Download SKILL.mdSave it as .claude/skills/turbofit/SKILL.md (or your agent's skills folder). This skill also uses 31 other files; get the full folder from GitHub.
name
turbofit
description
Hardware-aware adaptive Hermes runtime using portable Turbofiles, total usable memory, owned native llama.cpp residency, stable auto/active:main/active:aux routes, and evidence-backed promotion. Use for recommending, activating, inspecting, testing, or troubleshooting Turbofit runtimes.
version
2.3.1
author
SouthpawIN + Nous Girl
license
MIT
tags
hermes-agent, llama-cpp, llm, accelerator, cpu, adaptive-runtime, turbofile

Turbofit

2.4 model authority

Fit List keeps dedicated VRAM separate from integrated/RAM-only total memory. Dedicated: Maple Preview TQ2_0 at 8GB, Qwen 3.8 27B Unleashed UD-IQ3_XXS at 16GB, UD-Q3_K_XL at 24–95GB, and Qwen 3.8 27B 16-bit at 96GB+ until an Unleashed FP16 GGUF is published. Shared total memory: Maple at 8–15GB, Ornith 1.5 35A3B at 16–23GB, and Unleashed UD-Q3_K_XL at 24GB+. An 8GB GPU may also use Ornith when host RAM can hold offloaded experts. Maple remains Auto on dedicated 8GB at native 64K/128K. Auxiliary is Ornith, optional Carwin Nano, or auto.

Use when

  • Selecting an evidence-backed main/aux runtime for physical hardware
  • Activating or inspecting a Turbofile profile
  • Diagnosing pressure, contraction, expansion, routing, or native residency
  • Benchmarking or promoting a model pair
  • Updating candidate intelligence or generated wiki views

Canonical workflow

Work from the Git repository, not an installed copy.

bash
scripts/turbofit-runtime list
scripts/turbofit-runtime set auto
scripts/turbofit-runtime set <profile-id>
scripts/turbofit-runtime status
scripts/turbofit-controller --once
curl -fsS http://127.0.0.1:8091/v1/models

set auto chooses a canonical profile from immutable physical topology. set <profile-id> validates a measured, natively resolvable manual combination. Both begin at API safety and use the same adaptive controller to contract and heal; manual selection changes only the healing ceiling.

Use only stable provider IDs: auto, active:main, and active:aux.

FreeToken 0.1.2 at revision 0ab982f10905fa775962a4eddcb44caa50065251 is an optional NVIDIA/CUDA-13 text-only MoE candidate. Install or probe with scripts/install-freetoken-runtime; never expose it as an Auto rung, inherit its published TPS, or replace active Qwen 3.8/Ornith authority until an exact supported model recipe passes the full physical and intelligence campaigns.

Runtime authorities

  1. Turbofile: portable recommendation and ordered rung policy.
  2. Hardware fingerprint: physical topology, total usable memory, and per-device capacity; never transient free memory.
  3. Pressure snapshot: ownership-aware transient capacity.
  4. Pure policy: dwell/hysteresis/cooldown/flap decision.
  5. Native runtime backend: sole local residency authority for owned processes.
  6. Reconciler: drain, activate, verify, publish, rollback.
  7. Gateway route state: backing targets for stable IDs.

Legacy serve, direct launchers, and scaling watcher are compatibility tools, not adaptive authorities.

Portable memory allocation

Hardware fingerprints classify memory as dedicated, unified, or cpu. Turbofit reserves 5% of host RAM, bounded to 1–8 GiB, and never double-counts unified memory. Dedicated systems can combine VRAM and host RAM through llama.cpp offload; contexts beyond the model's native window move KV cache pressure to host RAM when at least 32 GiB is usable. Unified-memory systems suppress discrete split flags. CPU-only systems use pinned CPU runtimes with both model and draft GPU layers set to zero.

Backend order is CUDA → ROCm → Vulkan → CPU on Linux/Windows and Metal on macOS. Build or verify the current machine's pinned backend with scripts/install-native-runtimes --backend cuda|rocm|metal|vulkan|cpu.

Non-negotiable safety

  • Never kill or signal external accelerator/model processes.
  • Signal only PID-and-command-verified processes owned by Turbofit.
  • Count external memory as unavailable and managed residency as reclaimable.
  • A temporary auxiliary admission redirect may precede drain; never publish a new target rung before verification.
  • Restore and verify the previous rung after any failed transition.
  • Never place paths, secrets, credentials, provider keys, or device indices in Turbofiles.
  • Never treat research candidates or generated wiki text as production authority.
  • Never mark benchmark success without a canonical promotion record.
Show full SKILL.md (318 more words)Show less

Profile/recommendation checks

bash
PYTHONPATH=src python3 scripts/turbofit-runtime-recommend --fit-only --json
PYTHONPATH=src:. python3 -m pytest tests/test_runtime_profile.py tests/test_profile_io.py tests/test_hardware.py tests/test_recommend.py -q -o 'addopts='

Topology matters: 1x48 and 2x24 are different classes. Unmeasured classes keep API as the Auto safety rung while setup may expose separately labeled portable-fit local candidates for on-box validation.

TurboFit Check means the system scan-to-configuration process. TurboFit List means only the exact physical hardware-level winners generated by scripts/turbofit-list. Intelligence campaigns run only the current tier's tournament candidates; zero-call/token DeepSWE trials are invalid, and suite composites are rebuilt from real hash-bound pass counts rather than stale stored zeros.

Qwen 3.8 DFlash2 is a separate candidate using the pinned Inco Q4_K_M drafter and pinned z-lab llama.cpp PR runtime. Never reuse it for Bonsai. Bonsai keeps its own Prism DSpark sidecar/runtime until a dedicated Bonsai DFlash checkpoint exists.

Pressure and adaptation checks

bash
PYTHONPATH=src:. python3 -m pytest tests/test_pressure.py tests/test_pressure_probe.py tests/test_policy.py tests/test_reconciler.py tests/test_controller.py tests/test_runtime_service.py tests/integration -q -o 'addopts='

Expected contraction:

text
dedicated aux → shared-main → smaller context/model → terminal API

Expected recovery walks one rung at a time toward the recommendation after margin and dwell.

Release gates

bash
scripts/release-check
scripts/release-check --real

The first command validates syntax, tests, profiles, links, and simulated transitions. The second additionally requires working accelerator telemetry, stable live routes, and controlled real pressure/recovery evidence. Do not claim release readiness if --real is blocked.

Acceptance evidence: references/results/adaptive-runtime-acceptance.json.

Candidate intelligence

Collectors write only research/candidates.json:

bash
PYTHONPATH=. python3 research/discover_huggingface.py
PYTHONPATH=. python3 research/discover_model_news.py --url <public-feed>
PYTHONPATH=. python3 research/discover_api_models.py --provider <name> --url <public-model-list>

No collector may modify runtime profiles, routes, or credentials. Live cron schedules/delivery require explicit user approval.

Troubleshooting order

  1. scripts/turbofit-runtime status and the hardware fingerprint in Dashboard/Desktop
  2. The platform's available native inventory probe (CUDA, ROCm, Metal, Vulkan, or CPU)
  3. Native runtime /health, /v1/models, and /metrics
  4. Gateway /v1/models
  5. Route-state freshness and stable IDs
  6. Acceptance record blockers
  7. Focused tests, then full scripts/release-check

If http://127.0.0.1:8091/v1/models is connection-refused, the Turbofit provider gateway is not running. That is setup missing. Do not restart the Hermes messaging gateway, do not start with a firewall hunt when nothing listens, and do not invoke Sirvir until the endpoint answers.

If the platform reports a driver/runtime mismatch, stop the real pressure test. Do not attempt blind driver reloads or disruptive accelerator work.

Full architecture and schema: README.md.

© SouthpawIN, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 31 other files (scripts, references) in skills/turbofit of SouthpawIN/turbofit.

  • SKILL.md
  • .skillignore
  • README.md
  • distribution.yaml
  • references/SOUL.md
  • references/api-model-rankings.md
  • references/api-pairing-matrix.md
  • references/api-tier-rankings.md
  • references/auto-update.yaml
  • references/benchmark-analysis.md
  • references/benchmark-results.json
  • references/binary-selection.md
  • references/curated-lineup.md
  • references/fleet-benchmarks.json
  • references/model-database.yaml
  • references/model-pricing.json
  • references/optimization-matrix.yaml
  • references/optimization-results.json
  • references/research-report.md
  • references/scaling-ladder.md
  • … and 12 more

Open the folder on GitHubat commit 42fc223

Compare with similar skills

Turbofit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Turbofit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Turbofit this skillSouthpawIN/turbofit107—~1.9kAutomated safety check: PassMIT
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT
Add Modelguoqingbao/xinfer334—~4.2kAutomated safety check: NotesMIT
Resolvealexziskind1/model-shelf130—~792Automated safety check: PassMIT
Test Modelguoqingbao/xinfer334—~2.6kAutomated safety check: PassMIT
Add New ModelJakeATX/llamAmpere166—~4.1kAutomated safety check: PassMIT

Similar skills

  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    334 GitHub stars~4.2k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Test Model

    guoqingbao/xinfer

    Test LLM models served by xinfer for correctness, output quality, and performance.

    334 GitHub stars~2.6k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Add New Model

    JakeATX/llamAmpere

    Guided workflow for adding a new model architecture to llama.cpp.

    166 GitHub stars~4.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Code Review

    JakeATX/llamAmpere

    Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR.

    166 GitHub stars~5.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from SouthpawIN/turbofit

  • Turbofit

    SouthpawIN/turbofit

    Operate the Turbofit adaptive local inference plugin. An agent skill from SouthpawIN/turbofit.

    107 GitHub stars~1.4k tokensUpdated 2 days ago
    Auto-check passed

Works with

Questions about Turbofit

What does Turbofit do?

Hardware-aware adaptive Hermes runtime using portable Turbofiles, total usable memory, owned native llama.cpp residency, stable auto/active:main/active:aux routes, and evidence-backed promotion. Turbofit is an agent skill from SouthpawIN/turbofit.cpp residency, stable auto/active:main/active:aux routes, and evidence-backed promotion.

When should I use Turbofit?

Turbofit fits situations like: troubleshooting Turbofit runtimes; tasks that involve LLM inference and serving.

How do I install Turbofit in Claude Code?

Run `npx skills add SouthpawIN/turbofit --skill turbofit -a claude-code`. Or copy the skill folder (skills/turbofit in SouthpawIN/turbofit) into .claude/skills/turbofit in your project. Claude Code loads it when a task matches its description.

How do I install Turbofit in Codex?

Run `npx skills add SouthpawIN/turbofit --skill turbofit -a codex`. Or copy the skill folder (skills/turbofit in SouthpawIN/turbofit) into .agents/skills/turbofit in your project. Codex loads it when a task matches its description.

Can I use Turbofit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add SouthpawIN/turbofit --skill turbofit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/turbofit, .gemini/skills/turbofit, .github/skills/turbofit and .opencode/skills/turbofit in your project.

What does Turbofit need to run?

Going by SKILL.md and its folder, Turbofit needs the command-line tools its instructions call (python3 and curl). Our summary lists: Python 3.

Does Turbofit access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Turbofit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Turbofit use?

Turbofit is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Turbofit use?

About 1.9k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 35k tokens, read only when the agent opens those files.

What are the alternatives to Turbofit?

Skills that share tags, products or a category with Turbofit: Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), Add Model (guoqingbao/xinfer, 334 stars), Resolve (alexziskind1/model-shelf, 130 stars) and Test Model (guoqingbao/xinfer, 334 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Turbofit?

SouthpawIN (a GitHub user) maintains it in SouthpawIN/turbofit, which has 107 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 8, 2026.

Source: SouthpawIN/turbofit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.