Post-training 4-bit quantization for LLMs with minimal accuracy loss.

MITAuto-check passedAI & LLM Engineering

Install Gptq

skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill gptq -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orchestra-Research/AI-Research-SKILLs gptq --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/10-optimization/gptq .claude/skills/gptq && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gptq
GitHub stars
13k
Used in
2 other repos
Token cost
~2.9k tokens
SKILL.md length
539 words
Files
4 (incl. references)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

Post-training 4-bit quantization for LLMs with minimal accuracy loss.

  • Works in 4 steps: Group weights: Divide each weight matrix… → Quantize per-group: Each group has its… → Minimize error: Uses Hessian information… → …
  • Deploying large models (70B
  • SKILL.md covers When to use GPTQ, Quick start, Group-wise quantization and Quantization configurations, plus 8 more sections
  • Calls pip; reaches huggingface.co

What it does

Gptq is an agent skill from Orchestra-Research/AI-Research-SKILLs. Post-training 4-bit quantization for LLMs with minimal accuracy loss. Use for deploying large models (70B, 405B) on consumer GPUs, when you need 4× memory reduction with <2% perplexity degradation, or for faster inference (3-4× speedup) vs FP16. Integrates with transformers and PEFT for QLoRA fine-tuning.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/calibration.md`, `references/integration.md` and `references/troubleshooting.md`).

It sits in AI & LLM Engineering, covering Fine-tuning, LLM inference and serving and Web search. It works with Perplexity. The repository describes itself as: Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent… The licence is MIT.

When your agent uses it

  • Deploying large models (70B
  • 405B) on consumer GPUs
  • You need 4× memory reduction with <2% perplexity degradation
  • For faster inference (3-4× speedup) vs FP16

Example prompts

  • “/gptq”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Group weights: Divide each weight matrix into groups (typically 128 elements)
  2. Quantize per-group: Each group has its own scale/zero-point
  3. Minimize error: Uses Hessian information to minimize quantization error
  4. Result: 4-bit weights with near-FP16 accuracy

What it can do on your machine

Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • huggingface.co

    Also links to:

    • github.com
    • discord.gg

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gptq loads about 2.9k tokens when it runs, and up to ~6.1k if it reads all its reference files. Until then it costs about 78 tokens; SKILL.md has 539 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 539 words, ~2,891 tokens.

Download SKILL.mdSave it as .claude/skills/gptq/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
gptq
description
Post-training 4-bit quantization for LLMs with minimal accuracy loss. Use for deploying large models (70B, 405B) on consumer GPUs, when you need 4× memory reduction with <2% perplexity degradation, or for faster inference (3-4× speedup) vs FP16. Integrates with transformers and PEFT for QLoRA fine-tuning.
version
1.0.0
author
Orchestra Research
license
MIT
tags
Optimization, GPTQ, Quantization, 4-Bit, Post-Training, Memory Optimization, Consumer GPUs, Fast Inference, QLoRA, Group-Wise Quantization
dependencies
auto-gptq, transformers, optimum, peft

GPTQ (Generative Pre-trained Transformer Quantization)

Post-training quantization method that compresses LLMs to 4-bit with minimal accuracy loss using group-wise quantization.

When to use GPTQ

Use GPTQ when:

  • Need to fit large models (70B+) on limited GPU memory
  • Want 4× memory reduction with <2% accuracy loss
  • Deploying on consumer GPUs (RTX 4090, 3090)
  • Need faster inference (3-4× speedup vs FP16)

Use AWQ instead when:

  • Need slightly better accuracy (<1% loss)
  • Have newer GPUs (Ampere, Ada)
  • Want Marlin kernel support (2× faster on some GPUs)

Use bitsandbytes instead when:

  • Need simple integration with transformers
  • Want 8-bit quantization (less compression, better quality)
  • Don't need pre-quantized model files

Quick start

Installation
bash
# Install AutoGPTQ
pip install auto-gptq

# With Triton (Linux only, faster)
pip install auto-gptq[triton]

# With CUDA extensions (faster)
pip install auto-gptq --no-build-isolation

# Full installation
pip install auto-gptq transformers accelerate
Load pre-quantized model
python
from transformers import AutoTokenizer
from auto_gptq import AutoGPTQForCausalLM

# Load quantized model from HuggingFace
model_name = "TheBloke/Llama-2-7B-Chat-GPTQ"

model = AutoGPTQForCausalLM.from_quantized(
    model_name,
    device="cuda:0",
    use_triton=False  # Set True on Linux for speed
)

tokenizer = AutoTokenizer.from_pretrained(model_name)

# Generate
prompt = "Explain quantum computing"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda:0")
outputs = model.generate(**inputs, max_new_tokens=200)
print(tokenizer.decode(outputs[0]))
Quantize your own model
python
from transformers import AutoTokenizer
from auto_gptq import AutoGPTQForCausalLM, BaseQuantizeConfig
from datasets import load_dataset

# Load model
model_name = "meta-llama/Llama-2-7b-chat-hf"
tokenizer = AutoTokenizer.from_pretrained(model_name)

# Quantization config
quantize_config = BaseQuantizeConfig(
    bits=4,              # 4-bit quantization
    group_size=128,      # Group size (recommended: 128)
    desc_act=False,      # Activation order (False for CUDA kernel)
    damp_percent=0.01    # Dampening factor
)

# Load model for quantization
model = AutoGPTQForCausalLM.from_pretrained(
    model_name,
    quantize_config=quantize_config
)

# Prepare calibration data
dataset = load_dataset("c4", split="train", streaming=True)
calibration_data = [
    tokenizer(example["text"])["input_ids"][:512]
    for example in dataset.take(128)
]

# Quantize
model.quantize(calibration_data)

# Save quantized model
model.save_quantized("llama-2-7b-gptq")
tokenizer.save_pretrained("llama-2-7b-gptq")

# Push to HuggingFace
model.push_to_hub("username/llama-2-7b-gptq")

Group-wise quantization

How GPTQ works:

  1. Group weights: Divide each weight matrix into groups (typically 128 elements)
  2. Quantize per-group: Each group has its own scale/zero-point
  3. Minimize error: Uses Hessian information to minimize quantization error
  4. Result: 4-bit weights with near-FP16 accuracy

Group size trade-off:

Group SizeModel SizeAccuracySpeedRecommendation
-1 (per-column)SmallestBestSlowestResearch only
32SmallerBetterSlowerHigh accuracy needed
128MediumGoodFastRecommended default
256LargerLowerFasterSpeed critical
1024LargestLowestFastestNot recommended

Example:

Weight matrix: [1024, 4096] = 4.2M elements

Group size = 128:
- Groups: 4.2M / 128 = 32,768 groups
- Each group: own 4-bit scale + zero-point
- Result: Better granularity → better accuracy

Quantization configurations

python
from auto_gptq import BaseQuantizeConfig

config = BaseQuantizeConfig(
    bits=4,              # 4-bit quantization
    group_size=128,      # Standard group size
    desc_act=False,      # Faster CUDA kernel
    damp_percent=0.01    # Dampening factor
)

Performance:

  • Memory: 4× reduction (70B model: 140GB → 35GB)
  • Accuracy: ~1.5% perplexity increase
  • Speed: 3-4× faster than FP16
High accuracy (3-bit with larger groups)
python
config = BaseQuantizeConfig(
    bits=3,              # 3-bit (more compression)
    group_size=128,      # Keep standard group size
    desc_act=True,       # Better accuracy (slower)
    damp_percent=0.01
)

Trade-off:

  • Memory: 5× reduction
  • Accuracy: ~3% perplexity increase
  • Speed: 5× faster (but less accurate)
Maximum accuracy (4-bit with small groups)
python
config = BaseQuantizeConfig(
    bits=4,
    group_size=32,       # Smaller groups (better accuracy)
    desc_act=True,       # Activation reordering
    damp_percent=0.005   # Lower dampening
)

Trade-off:

  • Memory: 3.5× reduction (slightly larger)
  • Accuracy: ~0.8% perplexity increase (best)
  • Speed: 2-3× faster (kernel overhead)

Kernel backends

ExLlamaV2 (default, fastest)
python
model = AutoGPTQForCausalLM.from_quantized(
    model_name,
    device="cuda:0",
    use_exllama=True,      # Use ExLlamaV2
    exllama_config={"version": 2}
)

Performance: 1.5-2× faster than Triton

Marlin (Ampere+ GPUs)
python
# Quantize with Marlin format
config = BaseQuantizeConfig(
    bits=4,
    group_size=128,
    desc_act=False  # Required for Marlin
)

model.quantize(calibration_data, use_marlin=True)

# Load with Marlin
model = AutoGPTQForCausalLM.from_quantized(
    model_name,
    device="cuda:0",
    use_marlin=True  # 2× faster on A100/H100
)

Requirements:

  • NVIDIA Ampere or newer (A100, H100, RTX 40xx)
  • Compute capability ≥ 8.0
Triton (Linux only)
python
model = AutoGPTQForCausalLM.from_quantized(
    model_name,
    device="cuda:0",
    use_triton=True  # Linux only
)

Performance: 1.2-1.5× faster than CUDA backend

Integration with transformers

Direct transformers usage
python
from transformers import AutoModelForCausalLM, AutoTokenizer

# Load quantized model (transformers auto-detects GPTQ)
model = AutoModelForCausalLM.from_pretrained(
    "TheBloke/Llama-2-13B-Chat-GPTQ",
    device_map="auto",
    trust_remote_code=False
)

tokenizer = AutoTokenizer.from_pretrained("TheBloke/Llama-2-13B-Chat-GPTQ")

# Use like any transformers model
inputs = tokenizer("Hello", return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=100)
QLoRA fine-tuning (GPTQ + LoRA)
python
from transformers import AutoModelForCausalLM
from peft import prepare_model_for_kbit_training, LoraConfig, get_peft_model

# Load GPTQ model
model = AutoModelForCausalLM.from_pretrained(
    "TheBloke/Llama-2-7B-GPTQ",
    device_map="auto"
)

# Prepare for LoRA training
model = prepare_model_for_kbit_training(model)

# LoRA config
lora_config = LoraConfig(
    r=16,
    lora_alpha=32,
    target_modules=["q_proj", "v_proj"],
    lora_dropout=0.05,
    bias="none",
    task_type="CAUSAL_LM"
)

# Add LoRA adapters
model = get_peft_model(model, lora_config)

# Fine-tune (memory efficient!)
# 70B model trainable on single A100 80GB

Performance benchmarks

Show full SKILL.md (229 more words)Show less
Memory reduction
ModelFP16GPTQ 4-bitReduction
Llama 2-7B14 GB3.5 GB4×
Llama 2-13B26 GB6.5 GB4×
Llama 2-70B140 GB35 GB4×
Llama 3-405B810 GB203 GB4×

Enables:

  • 70B on single A100 80GB (vs 2× A100 needed for FP16)
  • 405B on 3× A100 80GB (vs 11× A100 needed for FP16)
  • 13B on RTX 4090 24GB (vs OOM with FP16)
Inference speed (Llama 2-7B, A100)
PrecisionTokens/secvs FP16
FP1625 tok/s1×
GPTQ 4-bit (CUDA)85 tok/s3.4×
GPTQ 4-bit (ExLlama)105 tok/s4.2×
GPTQ 4-bit (Marlin)120 tok/s4.8×
Accuracy (perplexity on WikiText-2)
ModelFP16GPTQ 4-bit (g=128)Degradation
Llama 2-7B5.475.55+1.5%
Llama 2-13B4.884.95+1.4%
Llama 2-70B3.323.38+1.8%

Excellent quality preservation - less than 2% degradation!

Common patterns

Multi-GPU deployment
python
# Automatic device mapping
model = AutoGPTQForCausalLM.from_quantized(
    "TheBloke/Llama-2-70B-GPTQ",
    device_map="auto",  # Automatically split across GPUs
    max_memory={0: "40GB", 1: "40GB"}  # Limit per GPU
)

# Manual device mapping
device_map = {
    "model.embed_tokens": 0,
    "model.layers.0-39": 0,  # First 40 layers on GPU 0
    "model.layers.40-79": 1,  # Last 40 layers on GPU 1
    "model.norm": 1,
    "lm_head": 1
}

model = AutoGPTQForCausalLM.from_quantized(
    model_name,
    device_map=device_map
)
CPU offloading
python
# Offload some layers to CPU (for very large models)
model = AutoGPTQForCausalLM.from_quantized(
    "TheBloke/Llama-2-405B-GPTQ",
    device_map="auto",
    max_memory={
        0: "80GB",  # GPU 0
        1: "80GB",  # GPU 1
        2: "80GB",  # GPU 2
        "cpu": "200GB"  # Offload overflow to CPU
    }
)
Batch inference
python
# Process multiple prompts efficiently
prompts = [
    "Explain AI",
    "Explain ML",
    "Explain DL"
]

inputs = tokenizer(prompts, return_tensors="pt", padding=True).to("cuda")

outputs = model.generate(
    **inputs,
    max_new_tokens=100,
    pad_token_id=tokenizer.eos_token_id
)

for i, output in enumerate(outputs):
    print(f"Prompt {i}: {tokenizer.decode(output)}")

Finding pre-quantized models

TheBloke on HuggingFace:

Search:

bash
# Find GPTQ models on HuggingFace
https://huggingface.co/models?library=gptq

Download:

python
from auto_gptq import AutoGPTQForCausalLM

# Automatically downloads from HuggingFace
model = AutoGPTQForCausalLM.from_quantized(
    "TheBloke/Llama-2-70B-Chat-GPTQ",
    device="cuda:0"
)

Supported models

  • LLaMA family: Llama 2, Llama 3, Code Llama
  • Mistral: Mistral 7B, Mixtral 8x7B, 8x22B
  • Qwen: Qwen, Qwen2, QwQ
  • DeepSeek: V2, V3
  • Phi: Phi-2, Phi-3
  • Yi, Falcon, BLOOM, OPT
  • 100+ models on HuggingFace

References

Resources

© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in 10-optimization/gptq of Orchestra-Research/AI-Research-SKILLs.

  • SKILL.md
  • references/calibration.md
  • references/integration.md
  • references/troubleshooting.md

Open the folder on GitHubat commit 773a529

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Gptq next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gptq compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gptq this skillOrchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT
Quark Torch LLM Evalamd/Quark182—~6.2kAutomated safety check: PassMIT
Nemotron Super3NVIDIA-NeMo/Nemotron2.1k—~2.4kAutomated safety check: PassApache-2.0
Nemotron UltraNVIDIA-NeMo/Nemotron2.1k—~1.8kAutomated safety check: PassApache-2.0
Fine-Tuning ExpertJeffallan/claude-skills12k—~1.7kAutomated safety check: PassMIT
Quantized Exportwshobson/agents40k—~2kAutomated safety check: PassMIT

Similar skills

  • End-to-end LLM accuracy evaluation on AMD ROCm (ROCm-only) — container setup, vLLM/SGLang/ATOM serving, lm-eval / lighteval / evalscope benchmarks.

    182 GitHub stars~6.2k tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed
  • Nemotron Super3

    NVIDIA-NeMo/Nemotron

    Reference desk for NVIDIA Nemotron 3 Super — architecture, training data, recipes (pretrain/SFT/RL/eval/quantization), and deployment notes.

    2.1k GitHub stars~2.4k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Nemotron Ultra

    NVIDIA-NeMo/Nemotron

    Reference desk for NVIDIA Nemotron 3 Ultra (550B-A55B) — architecture, NVFP4 pretraining, SFT, MOPD (multi-teacher on-policy distillation), MTP boosting, quantization, inference.

    2.1k GitHub stars~1.8k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub stars~1.7k tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed
  • Quantized Export

    wshobson/agents

    Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8.

    40k GitHub stars~2k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • ML Research Lab

    AnastasiyaW/codex-claude-code-config

    Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability.

    154 GitHub stars~794 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from Orchestra-Research/AI-Research-SKILLs

All 96 skills in this repo
  • AudioCraft Audio Generation

    Orchestra-Research/AI-Research-SKILLs

    Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.

    13k GitHub starsUsed in 8 repos~3.9k tokens
    Auto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    Auto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    Auto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    Auto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    Auto-check passed
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    Auto-check: notes

Works with

Questions about Gptq

What does Gptq do?

Post-training 4-bit quantization for LLMs with minimal accuracy loss. Gptq is an agent skill from Orchestra-Research/AI-Research-SKILLs. Post-training 4-bit quantization for LLMs with minimal accuracy loss.

When should I use Gptq?

Gptq fits situations like: deploying large models (70B; 405B) on consumer GPUs; you need 4× memory reduction with <2% perplexity degradation; for faster inference (3-4× speedup) vs FP16.

How do I install Gptq in Claude Code?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill gptq -a claude-code`. Or copy the skill folder (10-optimization/gptq in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/gptq in your project. Claude Code loads it when a task matches its description.

How do I install Gptq in Codex?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill gptq -a codex`. Or copy the skill folder (10-optimization/gptq in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/gptq in your project. Codex loads it when a task matches its description.

Can I use Gptq in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill gptq -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gptq, .gemini/skills/gptq, .github/skills/gptq and .opencode/skills/gptq in your project.

What does Gptq need to run?

Going by SKILL.md and its folder, Gptq needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Gptq access the network?

SKILL.md names 3 domains. In commands or code: huggingface.co; the agent is likely to contact it when it follows the instructions. As links in the text: github.com and discord.gg. This is read from the text; nothing was executed.

Is Gptq safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gptq use?

Gptq is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gptq use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.2k tokens, read only when the agent opens those files.

What are the alternatives to Gptq?

Skills that share tags, products or a category with Gptq: Quark Torch LLM Eval (amd/Quark, 182 stars), Nemotron Super3 (NVIDIA-NeMo/Nemotron, 2.1k stars), Nemotron Ultra (NVIDIA-NeMo/Nemotron, 2.1k stars) and Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gptq?

Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,405 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.

Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.