Agent skill

Ptq Workflow Integration

by vipshop in vipshop/cache-dit

A skill your agent uses when integrating a new PTQ workflow into cache-dit; designing quantize/load API shape, backend-specific config validation, save/load manifests, benchmark and regression…

Apache-2.0Auto-check passedTesting & QA

Install Ptq Workflow Integration

skills CLI
$ npx skills add vipshop/cache-dit --skill ptq-workflow-integration -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vipshop/cache-dit ptq-workflow-integration --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/ptq-workflow-integration .claude/skills/ptq-workflow-integration && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ptq-workflow-integration
GitHub stars
1.3k
Token cost
~2.8k tokens
SKILL.md length
1,372 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when integrating a new PTQ workflow into cache-dit; designing quantize/load API shape, backend-specific config validation, save/load manifests, benchmark and regression…

  • Works in 5 steps: Public API symmetry matters → Keep backend-specific knobs grouped and… → Hide internal PTQ machinery → …
  • Integrating a new PTQ workflow into cache-dit
  • SKILL.md covers Goal, Core Rule, When to Use and Reference Style Rule, plus 11 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Ptq Workflow Integration is an agent skill from vipshop/cache-dit. Use when integrating a new PTQ workflow into cache-dit; designing quantize/load API shape, backend-specific config validation, save/load manifests, benchmark and regression tests, or reviewing a PTQ integration plan. Uses the SVDQ PTQ integration only as a style and coverage reference. Do not copy the SVDQ implementation mechanically.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA. The repository describes itself as: A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs. The licence is Apache-2.0.

When your agent uses it

  • Integrating a new PTQ workflow into cache-dit
  • Designing quantize/load API shape
  • Backend-specific config validation
  • Save/load manifests

Example prompts

  • “/ptq-workflow-integration”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Public API symmetry matters
  2. Keep backend-specific knobs grouped and validated
  3. Hide internal PTQ machinery
  4. Save/load should be ergonomic and deterministic
  5. Slow validation must be opt-in

What it can do on your machine

Read from SKILL.md and the folder at commit a7898aa. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ptq Workflow Integration loads about 2.8k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 1,372 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vipshop/cache-dit at commit a7898aa, republished under its Apache-2.0 licence (© vipshop). 1,372 words, ~2,780 tokens.

Download SKILL.mdSave it as .claude/skills/ptq-workflow-integration/SKILL.md (or your agent's skills folder).
name
ptq-workflow-integration
description
Use when integrating a new PTQ workflow into cache-dit; designing quantize/load API shape, backend-specific config validation, save/load manifests, benchmark and regression tests, or reviewing a PTQ integration plan. Uses the SVDQ PTQ integration only as a style and coverage reference. Do not copy the SVDQ implementation mechanically.
argument-hint
Describe the PTQ algorithm or backend, target public API, calibration flow, serialization requirements, target models, and required validation layers.
user-invocable
true

PTQ Workflow Integration for cache-dit

Goal

Integrate one PTQ workflow in a way that feels native to cache-dit:

  • public API stays simple
  • backend-specific logic stays localized
  • save/load UX is predictable
  • fast tests and slow tests cover different risks
  • docs show only the public workflow

This skill is based on lessons from the SVDQ PTQ integration, but SVDQ PTQ is a reference only.

Core Rule

Do not mechanically copy SVDQ PTQ files.

Use the SVDQ PTQ integration to learn:

  • what belongs in public API vs private backend code
  • where config validation should live
  • where backend serialization and load logic should live
  • how tests should be split by cost and purpose
  • what documentation shape is acceptable for cache-dit

Do not reuse SVDQ PTQ by blind copy-paste.

Specifically do not copy without redesigning first:

  • private helper structure
  • class names or helper names
  • logging layout
  • exact file decomposition
  • prompts, thresholds, or benchmark constants
  • test bodies that only happen to fit SVDQ

Treat SVDQ as a style reference, not a template to replay.

When to Use

Use this skill when you need to:

  • add a new PTQ backend or algorithm into cache-dit
  • extend an existing PTQ backend with save/load support
  • decide where PTQ files and tests should live
  • review whether a PTQ integration follows cache-dit API style
  • plan coverage for a PTQ feature before coding

Do not use this skill for:

  • generic quantization work that does not involve PTQ workflow design
  • blind upstream porting
  • one-off benchmark scripts with no repository integration

Reference Style Rule

Use repo-relative references only.

  • For cache-dit files, use paths like src/cache_dit/quantization/config.py.
  • For docs, use paths like docs/user_guide/QUANTIZATION.md.
  • For tests, use paths like tests/quantization/test_svdquant_ptq.py.
  • Do not write machine-local absolute paths into the skill.

Design Principles to Keep

1. Public API symmetry matters

Prefer a user-facing flow like:

  • cache_dit.quantize(...)
  • cache_dit.load(...)
  • QuantizeConfig(...)

If save/load is part of the workflow, keep quantize and load at the same API layer unless there is a very strong reason not to.

2. Keep backend-specific knobs grouped and validated

Prefer validated backend-specific kwargs or a clearly scoped backend config section over many new top-level config fields.

The pattern to follow is:

  • generic config contract in src/cache_dit/quantization/config.py
  • backend-specific validation still triggered from that shared config layer
  • backend math and orchestration remain under the backend package
3. Hide internal PTQ machinery

Private PTQ context objects, calibrators, observers, or loaders should stay internal unless users truly need them.

Tests and docs should generally use public APIs only.

4. Save/load should be ergonomic and deterministic

If the PTQ workflow serializes checkpoints:

  • normalize output to a deterministic file name
  • keep a machine-readable manifest next to the checkpoint when directory loading is supported
  • resolve config, file path, and directory path through one internal load path resolver
  • validate metadata before mutating the target module
5. Slow validation must be opt-in

Fast regression coverage should run by default when feasible.

Pipeline-scale validation, compile validation, and large-model comparisons should be environment-gated.

File Placement Guidelines

Only add new files when they correspond to a real boundary.

Usually edit existing shared files for
  • generic config schema: src/cache_dit/quantization/config.py
  • public quantize/load routing: src/cache_dit/quantization/dispatch.py
  • package exports if public API changes: src/cache_dit/__init__.py
  • optional quantization package exports: src/cache_dit/quantization/__init__.py
  • user docs: docs/user_guide/QUANTIZATION.md
Usually add backend-local files for
  • PTQ orchestration and serialization: src/cache_dit/quantization/<backend>/ptq.py
  • backend math or accumulation logic: src/cache_dit/quantization/<backend>/quantizer.py
  • backend module wrappers or load helpers: src/cache_dit/quantization/<backend>/...
Usually add tests in
  • backend public-workflow tests: tests/kernels/test_<backend>_ptq.py
  • backend math or lower-level tests: tests/kernels/test_<backend>_quantizer.py

Add a separate shared test utility file only if multiple test files genuinely reuse the same helpers.

What the SVDQ PTQ Integration Teaches About API Design

Use these as design lessons, not copy targets.

Shared config layer

src/cache_dit/quantization/config.py is the right place for:

  • quant type parsing and normalization
  • backend auto-resolution
  • validation of PTQ-specific unsupported combinations
  • normalization of serialize_to
  • validation of backend-specific kwargs

Do not push these checks into only the backend implementation file if the public config object can reject them earlier.

Backend PTQ implementation layer

src/cache_dit/quantization/svdquant/ptq.py demonstrates the right kind of responsibilities for a backend PTQ file:

  • run calibration through the public config callback
  • quantize target modules
  • serialize checkpoint artifacts
  • write a lightweight JSON manifest for directory load UX
  • resolve load inputs and validate metadata
  • attach runtime metadata back onto the loaded module

This is the right layer for backend save/load orchestration.

Public docs and tests layer

docs/user_guide/QUANTIZATION.md and tests/quantization/test_svdquant_ptq.py demonstrate the preferred user story:

  • user interacts with QuantizeConfig
  • user calls cache_dit.quantize
  • user calls cache_dit.load
  • user does not need private PTQ classes

Keep new PTQ integrations aligned with that style unless the backend genuinely requires a different experience.

Serialization and Load Checklist

If the PTQ workflow needs saved checkpoints, check all of the following:

  1. The serialized checkpoint file name is deterministic.
  2. A colocated manifest exists when directory loading is supported.
  3. The load path accepts ergonomic inputs only when they can be resolved unambiguously.
  4. Metadata validation happens before module mutation.
  5. Quant type mismatches fail clearly.
  6. Missing manifest or malformed metadata fail clearly.
  7. Round-trip tests prove loaded output matches quantized output.
Show full SKILL.md (542 more words)Show less

Test Strategy

Do not ship a PTQ integration with only one slow end-to-end test.

Fast tests should cover
  • public API quantization replaces the expected layers
  • save/load roundtrip restores the quantized module
  • config validation rejects unsupported combinations
  • directory load, checkpoint load, and config-driven load if all are supported
  • invalid metadata and incomplete checkpoint failure cases
  • exclusion or filtering behavior
  • backend-specific calibration or buffering behavior when applicable
Slow tests should cover
  • at least one real pipeline or model integration path
  • serialization plus reload closure
  • one primary quality gate
  • additional metrics reported separately from the hard gate
  • latency, transformer memory, or peak memory when those are part of the PTQ value proposition
Optional compile tests should cover
  • loading an already quantized module
  • enabling compile configs if the repo expects them
  • torch.compile(...)
  • one warmup run
  • one actual inference run

Compile validation should be behind a separate environment variable, not mixed into the default slow test path.

Test Style Rules to Preserve

Follow these rules unless there is a strong reason to violate them:

  • integration tests should use public APIs, not private PTQ classes
  • slow tests should self-skip with clear environment-variable guidance
  • deterministic generators or seeds should be used for pipeline tests
  • large test artifacts should go under repo-local .tmp/tests/...
  • benchmark tables and visuals are reports, not pass/fail criteria unless explicitly required
  • hard accuracy gates should stay minimal and explainable

Suggested Coverage Map

When integrating a new PTQ workflow, think in layers:

  1. config/schema layer
  2. backend quantizer/math layer
  3. serialization/load layer
  4. public API layer
  5. model or pipeline integration layer
  6. optional compile layer

If one of these layers is intentionally out of scope, say so explicitly in the PR or plan.

  1. Survey existing public quantization API and decide whether the new PTQ flow fits it.
  2. Add or update shared config validation in src/cache_dit/quantization/config.py.
  3. Implement backend PTQ orchestration in src/cache_dit/quantization/<backend>/ptq.py.
  4. Add backend helper logic only where a real separation exists.
  5. Wire dispatch and exports only after backend behavior is stable.
  6. Add fast tests for public API, roundtrip, validation, and failure cases.
  7. Add env-gated slow tests for real model integration.
  8. Add optional compile tests only if compile compatibility matters.
  9. Update docs/user_guide/QUANTIZATION.md with public API examples only.

Review Questions

Before merging a PTQ integration, ask:

  1. Does the user-facing flow still look like cache-dit?
  2. Are backend-specific options validated centrally?
  3. Are save/load artifacts deterministic and discoverable?
  4. Can the feature be tested quickly without a giant model?
  5. Are slow tests opt-in and clearly scoped?
  6. Do docs and tests avoid private PTQ symbols?
  7. Does this integration borrow SVDQ lessons without copying SVDQ internals?

Common Mistakes

  • exposing private PTQ context classes in docs or integration tests
  • adding too many top-level config fields instead of validated backend-specific kwargs
  • implementing save/load UX only through one exact file path
  • skipping malformed metadata tests
  • writing only slow model tests and no fast public-API tests
  • making compile validation part of the default slow path
  • hardcoding local machine paths into docs or instructions
  • copying SVDQ helper structure line-for-line

Reference Touchpoints

These are the main SVDQ PTQ reference files for style and coverage shape:

  • src/cache_dit/quantization/config.py
  • src/cache_dit/quantization/svdquant/ptq.py
  • tests/quantization/test_svdquant_ptq.py
  • tests/quantization/test_svdquant_quantizer.py
  • docs/user_guide/QUANTIZATION.md

Use them to understand cache-dit conventions.

Do not reproduce them mechanically.

© vipshop, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .github/skills/ptq-workflow-integration of vipshop/cache-dit.

Open the folder on GitHubat commit a7898aa

Compare with similar skills

Ptq Workflow Integration next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ptq Workflow Integration compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ptq Workflow Integration this skillvipshop/cache-dit1.3k—~2.8kAutomated safety check: PassApache-2.0
Create Taskmelandlabs/openloomi1k—~4.6kAutomated safety check: PassApache-2.0
Early Experience DataOSU-NLP-Group/EarlyExperience102—~4.1kAutomated safety check: PassMIT
EvaluationPrimeIntellect-ai/prime-envs130—~4.6kAutomated safety check: PassApache-2.0
5 Persona Advisory Boardharryvondiesel-web/5-persona-advisory-board131—~2.8kAutomated safety check: PassMIT
Veomni Patchgen ModelByteDance-Seed/VeOmni2.2k—~9.6kAutomated safety check: PassApache-2.0

Similar skills

  • Create Task

    melandlabs/openloomi

    Create an end-to-end Continual Learning Bench task. An agent skill from melandlabs/openloomi.

    1k GitHub stars~4.6k tokensUpdated 15 days ago
    Testing & QAAuto-check passed
  • Early Experience Data

    OSU-NLP-Group/EarlyExperience

    A skill your agent uses whenever the user asks to generate, collect, inspect, or prepare early-experience training data (Implicit World Modeling or Self-Reflection, in the sense of arXiv:2510.08558)…

    102 GitHub stars~4.1k tokensUpdated 3 mo ago
    Testing & QAAuto-check passed
  • Evaluation

    PrimeIntellect-ai/prime-envs

    Install and run a verifiers environment — smoke testing during development and full benchmark evals.

    130 GitHub stars~4.6k tokensUpdated today
    Testing & QAAuto-check passed
  • 5 Persona Advisory Board

    harryvondiesel-web/5-persona-advisory-board

    Run a 5 Persona Advisory Board review, board review, strategic decision stress-test, offer critique, risk check, or pricing/timing/positioning decision review.

    131 GitHub stars~2.8k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Veomni Patchgen Model

    ByteDance-Seed/VeOmni

    Author or refresh a VeOmni model's patchgen-generated modeling under generated/ — GPU and/or NPU config, dense or MoE, text / VLM / Omni.

    2.2k GitHub stars~9.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Rsibench Data Factory

    evolvent-ai/RSIBench-Data

    Use inside RSIBench-Data when testing whether an automation agent can improve a target model on a configured benchmark through synthetic Tinker SFT data, Tinker sampling, and E2B-based Harbor…

    171 GitHub stars~640 tokensUpdated 1 mo ago
    Testing & QAAuto-check: notes

More from vipshop/cache-dit

  • High-level guide for integrating a new DiT model into cache-dit: Cache (BlockAdapter/ForwardPattern), Context Parallelism, Tensor Parallelism, Text Encoder Parallelism (TE-P), VAE Parallelism…

    1.3k GitHub stars~11k tokensUpdated 10 days ago
    Auto-check passed
  • Cuda Cpp Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

    1.3k GitHub stars~2.3k tokensUpdated 10 days ago
    Auto-check passed
  • Cute Dsl Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or…

    1.3k GitHub stars~2.8k tokensUpdated 10 days ago
    Auto-check passed
  • Cutlass Cpp Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM…

    1.3k GitHub stars~2.2k tokensUpdated 10 days ago
    Auto-check passed
  • Operator Migration

    vipshop/cache-dit

    A skill your agent uses when doing operator migration or kernel migration for CUDA, Triton, or custom ops in cache-dit; porting kernels from nunchaku, deepcompressor, or other repos; designing…

    1.3k GitHub stars~3.8k tokensUpdated 10 days ago
    Auto-check passed
  • Triton Kernel

    vipshop/cache-dit

    Write optimized Triton GPU kernels for deep learning operations.

    1.3k GitHub stars~1.1k tokensUpdated 10 days ago
    Auto-check passed

Questions about Ptq Workflow Integration

What does Ptq Workflow Integration do?

A skill your agent uses when integrating a new PTQ workflow into cache-dit; designing quantize/load API shape, backend-specific config validation, save/load manifests, benchmark and regression…. Ptq Workflow Integration is an agent skill from vipshop/cache-dit. Use when integrating a new PTQ workflow into cache-dit; designing quantize/load API shape, backend-specific config validation, save/load manifests, benchmark and regression tests, or reviewing a PTQ integration plan.

When should I use Ptq Workflow Integration?

Ptq Workflow Integration fits situations like: integrating a new PTQ workflow into cache-dit; designing quantize/load API shape; backend-specific config validation; save/load manifests.

How do I install Ptq Workflow Integration in Claude Code?

Run `npx skills add vipshop/cache-dit --skill ptq-workflow-integration -a claude-code`. Or copy the skill folder (.github/skills/ptq-workflow-integration in vipshop/cache-dit) into .claude/skills/ptq-workflow-integration in your project. Claude Code loads it when a task matches its description.

How do I install Ptq Workflow Integration in Codex?

Run `npx skills add vipshop/cache-dit --skill ptq-workflow-integration -a codex`. Or copy the skill folder (.github/skills/ptq-workflow-integration in vipshop/cache-dit) into .agents/skills/ptq-workflow-integration in your project. Codex loads it when a task matches its description.

Can I use Ptq Workflow Integration in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vipshop/cache-dit --skill ptq-workflow-integration -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ptq-workflow-integration, .gemini/skills/ptq-workflow-integration, .github/skills/ptq-workflow-integration and .opencode/skills/ptq-workflow-integration in your project.

What does Ptq Workflow Integration need to run?

SKILL.md names no scripts, command-line tools or credentials: Ptq Workflow Integration is instructions for the agent only.

Does Ptq Workflow Integration access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ptq Workflow Integration safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ptq Workflow Integration use?

Ptq Workflow Integration is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ptq Workflow Integration use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ptq Workflow Integration?

Skills that share tags, products or a category with Ptq Workflow Integration: Create Task (melandlabs/openloomi, 1k stars), Early Experience Data (OSU-NLP-Group/EarlyExperience, 102 stars), Evaluation (PrimeIntellect-ai/prime-envs, 130 stars) and 5 Persona Advisory Board (harryvondiesel-web/5-persona-advisory-board, 131 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ptq Workflow Integration?

vipshop (a GitHub organization) maintains it in vipshop/cache-dit, which has 1,289 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 29, 2026.

Source: vipshop/cache-dit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.