Agent skill

Phyai Model Implement

by mingti-org in mingti-org/phyai

A skill your agent uses when implementing, porting, integrating, reproducing, or debugging support for a model in PHYAI.

MITAuto-check passedAgent Workflows

Install Phyai Model Implement

skills CLI
$ npx skills add mingti-org/phyai --skill phyai-model-implement -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mingti-org/phyai phyai-model-implement --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mingti-org/phyai.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/phyai-model-implement .claude/skills/phyai-model-implement && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
phyai-model-implement
GitHub stars
129
Token cost
~3.7k tokens
SKILL.md length
1,629 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when implementing, porting, integrating, reproducing, or debugging support for a model in PHYAI.

  • Works in 7 steps: Configuration (configuration_.py) → Modeling (modeling_.py) → Weight Loading → …
  • Debugging support for a model in PHYAI
  • SKILL.md covers Before You Implement, Core Three-Layer Architecture, File Naming Conventions and Implementation Steps, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Phyai Model Implement is an agent skill from mingti-org/phyai. Use this skill when implementing, porting, integrating, reproducing, or debugging support for a model in PHYAI. This includes translating architecture research into PHYAI code, adding model configuration, modeling modules, runners, schedulers, weight loading, layer reuse, focused tests, validation scripts, and implementation plans while respecting PHYAI model-development constraints.

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Planning and Translation. The repository describes itself as: PhyAI is a high-performance framework for running Physical AI models (VLA, WAM, and beyond), supporting both cloud-based serving and on-device deployment. The licence is MIT.

When your agent uses it

  • Debugging support for a model in PHYAI
  • Tasks that involve Planning
  • Tasks that involve Translation

Example prompts

  • “/phyai-model-implement”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Configuration (configuration_.py)
  2. Modeling (modeling_.py)
  3. Weight Loading
  4. Model Runner (model_runner_.py)
  5. Scheduler (scheduler_ws1_.py)
  6. Engine Plugin (main_.py)
  7. Processor (phyai-utils-tools package)

What it can do on your machine

Read from SKILL.md and the folder at commit 36a46bf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Phyai Model Implement loads about 3.7k tokens when it runs. Until then it costs about 102 tokens; SKILL.md has 1,629 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~102
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mingti-org/phyai at commit 36a46bf, republished under its MIT licence (© mingti-org). 1,629 words, ~3,743 tokens.

Download SKILL.mdSave it as .claude/skills/phyai-model-implement/SKILL.md (or your agent's skills folder).
name
phyai-model-implement
description
Use this skill when implementing, porting, integrating, reproducing, or debugging support for a model in PHYAI. This includes translating architecture research into PHYAI code, adding model configuration, modeling modules, runners, schedulers, weight loading, layer reuse, focused tests, validation scripts, and implementation plans while respecting PHYAI model-development constraints.

PHYAI Model Implementation Guide

Before You Implement

First, carefully study the existing repositories and articles. If the user has not provided references, ask them for the relevant references before proceeding.

Core Three-Layer Architecture

PHYAI uses a strict three-layer separation for every model:

text
modeling_xxx.py     - Pure architecture definition: stateless, no KV cache, no runtime state
model_runner_xxx.py - Runtime state management: KV cache pool, condition cache, prefill/decode switching
scheduler_xxx.py    - Orchestration loop: denoise loop, sampler, CFG, multi-step orchestration,
                      multi-GPU and multi-stage coordination
vae.py              - Extra models or processing files, such as VAE, and auxiliary models

Never:

  • Never put KV cache inside modeling classes. A modeling class performs one forward pass only.
  • Never introduce external modeling files from libraries such as diffusers, transformers, or wan_x.
  • Never put sampling or noise-schedule logic in the modeling layer. That is the scheduler's responsibility.
  • Never put matrix multiplication or attention in the scheduler layer unless there is no viable alternative. Those belong in the modeling layer.

Call relationship between layers:

text
Engine Plugin (main_xxx.py)
  +-- Scheduler (scheduler_ws1_xxx.py)
        +-- ModelRunner (model_runner_xxx.py)
              +-- Model (modeling_xxx.py)

Each layer calls only the layer immediately below it. Do not skip layers. A scheduler must not call an nn.Module.forward directly; it must go through the runner.

File Naming Conventions

text
phyai/src/phyai/models/<model_name>/
+-- __init__.py                      # Export all public symbols
+-- configuration_<model>.py         # Frozen dataclass configuration
+-- modeling_<model>.py              # Pure network architecture
+-- model_runner_<model>.py          # Runtime state wrapper
+-- scheduler_ws1_<model>.py         # Single-GPU scheduler; ws = world_size
+-- main_<model>.py                  # Engine plugin entry point; Entry subclass
+-- (optional) sampler_*.py          # Diffusion or ODE sampler
     NOTE: name it XxxSampler, not XxxScheduler,
           to avoid conflict with phyai.runtime.schedule.Scheduler.

Naming rules:

  • Use CamelCase for class names and snake_case for file names.
  • Weight remap function: <model>_weight_remap.
  • Engine plugin name: a short lowercase name, such as "pi05" or "cosmos3_policy".

Implementation Steps

Follow these steps in order.

1. Configuration (configuration_<model>.py)
  • Define the model configuration with @dataclass(frozen=True).
  • Support loading from checkpoint config.json with load_config(path, XxxConfig).
  • If checkpoint JSON keys do not match dataclass fields, use the nested_sources class variable to define the mapping.
  • Validate the configuration in __post_init__, such as GQA divisibility and even head_dim.
  • Give every numeric parameter a reasonable default so XxxConfig() can be constructed without arguments in tests.
2. Modeling (modeling_<model>.py)

Principles:

  • Keep modeling stateless. Each forward pass depends only on its input arguments, not on internal mutable state.
  • Use shared layers from phyai.layers, such as Linear, RMSNorm, Attention, and RotaryEmbedding.
  • Do not implement a custom attention kernel. Use phyai.layers.attention.
  • Do not implement RoPE yourself. Use phyai.layers.rotary_embedding.
  • Do not implement normalization yourself. Use RMSNorm or LayerNorm from phyai.layers.
  • The weight mapping function xxx_weight_remap(name) returns the checkpoint key to PHYAI key mapping.
  • If a new general-purpose layer is needed, add it to phyai.layers rather than to the model directory.

Typical forward signature:

python
class XxxModel(nn.Module):
    def __init__(self, config: XxxConfig, *, params_dtype=torch.bfloat16, device=None):
        ...

    def forward(self, inputs, ...) -> output:
        """Run one forward pass without storing intermediate state."""
3. Weight Loading
  • Use phyai.weights.loader.load_pretrained(module, path, remap=xxx_weight_remap, strict=False).
  • Define the remap function as def xxx_weight_remap(name: str) -> str | None; returning None means the weight is skipped.
  • Support both single-file safetensors and sharded indexes; automatic detection is already available.
  • Validate with assert len(report.missing) == 0.
4. Model Runner (model_runner_<model>.py)
  • Inherit from phyai.runtime.model_runner.ModelRunner, whose abstract methods are setup() and forward().
  • Responsibilities:
    • Condition caching: run encode_condition once and reuse the result in later steps.
    • CUDA graph capture, optionally controlled by a use_cuda_graph flag.
    • Optional torch.compile, applied to submodules in setup().
    • reset(), which clears per-request caches and is called at the start of each new request.
  • The runner holds a reference to the model but does not own the weights. Weights are loaded at the plugin layer and then passed in.
5. Scheduler (scheduler_ws1_<model>.py)
  • Inherit from phyai.runtime.schedule.Scheduler, whose abstract methods are setup() and step(request).
  • Responsibilities:
    • Inference loop, such as repeatedly calling runner.forward for a diffusion denoise loop or an autoregressive decode loop.
    • Multi-branch forward passes, such as CFG dual branches plus guidance interpolation.
    • Sampler control, including timestep schedules and ODE solvers.
    • Condition re-imposition, such as writing conditioned regions back at every step.
    • Noise initialization: seed to generator to randn.
  • Define requests with @dataclass; include all inference parameters in the request.
  • step() should behave functionally: given a request, return a result, and do not keep state across requests.
6. Engine Plugin (main_<model>.py)
  • Register the plugin with the @Engine.register decorator.
  • Implement setup(args), step(request), and close().
  • setup is responsible for loading the configuration, building the model, loading weights, building the runner and scheduler, and warming up.
  • step delegates directly to the scheduler.
  • close releases all GPU resources.
7. Processor (phyai-utils-tools package)
  • Inherit from BaseModelProcessor and implement build_preprocessor() and build_postprocessor().
  • Preprocessing: raw image/text/state to model-ready tensors.
  • Postprocessing: model output to a user-friendly format.
  • Do not import phyai. phyai-utils-tools is an independent leaf package and must not depend on the main library.
  • If preprocessing requires a model forward, such as VAE encode, keep it in the scheduler or plugin layer rather than in the processor.

Critical Constraints

Code Organization
  1. Do not modify general components such as phyai.layers or phyai.runtime unless a required layer or system component is missing. If a change is needed, tell the user first.
  2. Use one directory per model. Put all model-specific code under phyai/models/<model>/; do not scatter it elsewhere.
  3. Do not import across models. Model A must not import code from Model B. Promote shared logic to phyai.layers.
Attention and Computation
  1. Use flashinfer or SDPA for attention by default. Do not use eager attention except for debug.
  2. Precision: default to bf16. If a submodule needs fp32, such as a timestep MLP or ViT, control it through params_dtype during construction.
  3. Do not hardcode devices in model code. Pass the device as an argument, or infer it from existing tensors.
Testing
  1. Layer-level tests belong in phyai/tests/; the suite requires CUDA and runs real kernels.
  2. Model-level tests that require full weights belong in .cache/, because CI does not have enough resources for them.
  3. conftest.py aborts collection on a machine without CUDA; layers construct on the engine default device.target = "cuda".
Style
  1. Logging: get the module logger with get_logger(__name__) from phyai.utils; never logging.getLogger. Use logger.info_rank0(...) for rank-0-only lines, plain logger.info(...) when every rank should log (the [rank R/W] label comes from the formatter), and logger.warning_once(...) on per-request / per-layer paths. Do not pass a logger or a level as an argument, and do not call print directly.
  2. Comments: write all comments in English.
  3. Naming: public by default. Do not add leading underscores casually. Expose singletons through get_*() getters.
  4. Do not add type: ignore. Fix the type instead of suppressing the warning.
  5. Import order: stdlib, third-party, phyai, then local imports. Ruff will sort imports automatically.
Show full SKILL.md (671 more words)Show less

Validation Workflow

Validation must compare PHYAI against the original reference implementation, not only against shape checks or smoke tests. Keep the reference repository available locally when possible, run the same checkpoint and deterministic inputs through both implementations, and save enough intermediate tensors to identify the first layer that diverges.

Before judging final quality, align the basics with the reference repository:

  1. Use the same checkpoint, tokenizer or processor, preprocessing rules, dtype policy, random seed, timestep schedule, sampler settings, guidance settings, and device placement.
  2. Verify config parity: every architecture field that affects tensor shapes, attention layout, normalization, MLP width, RoPE, patching, channel order, or action/video dimensions must match the reference.
  3. Verify weight parity: map each checkpoint tensor to the intended PHYAI parameter, check missing and unexpected keys, and spot-check representative tensor values after loading.
  4. Verify input parity: feed identical model-ready tensors to both implementations. If processors differ, dump the processed tensors and compare them before running the model.
  5. Compare progressively: embedding/projection outputs, each block or major submodule, final model outputs, and finally the scheduler or end-to-end result.
Single-Step Velocity Parity
python
# Same input: PHYAI forward vs. reference forward.
# Expected cosine > 0.99, allowing for bf16 accumulation error.
cosine = F.cosine_similarity(phyai_out.flatten(), ref_out.flatten(), dim=0)

For a diffusion or flow model, compare the predicted velocity/noise/action for one fixed step against the reference repository first. The final single-step output cosine similarity must be greater than 0.99. If it is below 0.99, do not treat the port as validated; find the earliest diverging intermediate tensor and fix the corresponding config, weight mapping, layout, precision, or preprocessing issue.

End-to-End Inference Validation
python
# Determinism: same seed -> exactly identical output, cosine = 1.0.
# Convergence: after the denoise loop, output std should be far below noise std.

After single-step parity passes, run the full PHYAI scheduler against the reference repository with the same request and deterministic seed. The final end-to-end result should also reach cosine similarity greater than 0.99 against the reference output, unless the reference uses a documented non-deterministic kernel. If it cannot meet this threshold, document the exact source of drift and whether it comes from precision, sampler implementation, preprocessing, or an intentional algorithmic deviation.

Weight Loading Validation
python
report = load_pretrained(model, path, remap=remap, strict=False)
assert len(report.missing) == 0  # All weights have been loaded.

References

Existing model implementations live under phyai/src/phyai/models/. Before implementing a new model, read one complete existing model implementation as a reference.

Basic Performance Considerations

Write code that remains friendly to later performance optimization, including but not limited to torch.compile and CUDA graphs.

torch.compile Friendly
  1. Avoid data-dependent control flow. Do not branch on runtime tensor values, such as if tensor.item() > 0, because it causes graph breaks unless the algorithm truly requires it.
  2. Avoid dynamic shapes. Keep tensor shapes inferable at compile time whenever possible. If a shape must be dynamic, mark it with torch._dynamo.mark_dynamic().
  3. Do not create a tensor in forward and immediately call .to(device). Preallocate buffers in __init__ or setup().
  4. Avoid mixing Python list/dict operations with tensors. [t1, t2, t3] followed by torch.stack() is fine, but appending tensors to a list in a loop and then calling cat may trigger a graph break.
  5. Prefer torch.nn.functional over custom Python loops. For example, use F.scaled_dot_product_attention rather than a hand-written softmax plus matmul.
CUDA Graph Friendly
  1. Use fixed tensor shapes. All input and output shapes must be fixed during CUDA graph capture. Shape changes require re-capture.
  2. Do not cause CPU-GPU synchronization in forward. tensor.item(), tensor.cpu(), and print(tensor) all break graph execution.
  3. Do not allocate memory in forward. Preallocate torch.empty, torch.zeros, and torch.randn buffers, then fill them in-place in forward.
  4. Avoid Python side effects. CUDA graph replay does not re-run Python code; it replays only the captured CUDA kernel sequence.
  5. Use static flags for conditional branches instead of runtime tensors. if self.use_xxx: with a Python bool is acceptable; if tensor > 0: is not.
General Principles
  1. Reduce kernel launches. Fuse consecutive small operations where possible, such as RMSNorm plus Linear.
  2. Avoid unnecessary contiguous() calls. Call it only when a kernel truly requires contiguous memory.
  3. Do not write attention yourself. Use the unified interface in phyai.layers.attention; it chooses flashinfer, flash_attn, or SDPA automatically.
  4. Prefer reshape/view over permute plus contiguous for large tensors. The latter triggers an extra memory copy.
  5. Do not call torch.cuda.synchronize() in forward. It serializes all streams unless needed for debugging.

© mingti-org, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/phyai-model-implement of mingti-org/phyai.

Open the folder on GitHubat commit 36a46bf

Compare with similar skills

Phyai Model Implement next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Phyai Model Implement compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Phyai Model Implement this skillmingti-org/phyai129—~3.7kAutomated safety check: PassMIT
OpenSpec Guided OnboardingFission-AI/OpenSpec71k1 repos~4.5kAutomated safety check: PassMIT
Paseo Committeegetpaseo/paseo20k1 repos~496Automated safety check: PassCustom licence
Improvefossasia/eventyay-interpretation1.6k10 repos~3.7kAutomated safety check: WarnMIT
Dsh Web Documentationzhu1090093659/dsh-web8.5k—~479Automated safety check: PassApache-2.0
Implementation Plan Creatortailcallhq/forgecode7.6k1 repos~1.1kAutomated safety check: PassApache-2.0

Similar skills

  • OpenSpec Guided Onboarding

    Fission-AI/OpenSpec

    Walks you through a complete OpenSpec workflow cycle with narration while doing real work in your codebase.

    71k GitHub starsUsed in 1 repo~4.5k tokens
    Agent WorkflowsAuto-check passed
  • Paseo Committee

    getpaseo/paseo

    Forms a two-agent committee with contrasting profiles to analyze a stuck problem in parallel, reconcile their views and return a consensus plan without editing files.

    20k GitHub starsUsed in 1 repo~496 tokens
    Agent WorkflowsAuto-check passed
  • Improve

    fossasia/eventyay-interpretation

    Survey any codebase as a senior advisor and produce prioritized, self-contained implementation plans for OTHER models/agents to execute.

    1.6k GitHub starsUsed in 10 repos~3.7k tokens
    Agent WorkflowsAuto-check: warnings
  • Dsh Web Documentation

    zhu1090093659/dsh-web

    A skill your agent uses when adding or editing dsh-web README files, docs, AGENTS.md instructions, user-facing configuration text, or bilingual documentation pairs.

    8.5k GitHub stars~479 tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Implementation Plan Creator

    tailcallhq/forgecode

    Writes a structured Markdown implementation plan with checkbox tasks, verification criteria and risks, then checks it with a validation script; no code changes.

    7.6k GitHub starsUsed in 1 repo~1.1k tokens
    Agent WorkflowsAuto-check passed
  • Plan Preview

    u-ichi/reviewable-html-workbench

    Plan Mode の <proposedplan を出す直前に、計画の段階・依存関係・検証観点を一時HTMLで視覚確認したい時に使う agent-internal skill。Use this agent-internal skill to create a temporary HTML preview for a plan just before presenting…

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Agent WorkflowsAuto-check passed

More from mingti-org/phyai

  • Phyai Local Env Report

    mingti-org/phyai

    Generate a local phyai environment report for debugging system, Python, CUDA/GPU, dependency, workspace package, git, and PHYAI configuration issues.

    129 GitHub stars~590 tokensUpdated yesterday
    Auto-check passed
  • A skill your agent uses when the user provides a paper, arXiv link, technical report, model card, checkpoint name, GitHub repository, or local codebase and asks to research, explain, compare, or…

    129 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Phyai Solve PR Comments

    mingti-org/phyai

    Triage and resolve review comments on a GitHub PR — fetch all comment surfaces (issue / inline / review), validate each suggestion against upstream source rather than trusting blindly, present…

    129 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Analyze a phyai .memory file or directory supplied by the user: parse what task/session it records, locate and inspect the referenced code repository when available, verify claims against local code…

    129 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed

Questions about Phyai Model Implement

What does Phyai Model Implement do?

A skill your agent uses when implementing, porting, integrating, reproducing, or debugging support for a model in PHYAI. Phyai Model Implement is an agent skill from mingti-org/phyai. Use this skill when implementing, porting, integrating, reproducing, or debugging support for a model in PHYAI.

When should I use Phyai Model Implement?

Phyai Model Implement fits situations like: debugging support for a model in PHYAI; tasks that involve Planning; tasks that involve Translation.

How do I install Phyai Model Implement in Claude Code?

Run `npx skills add mingti-org/phyai --skill phyai-model-implement -a claude-code`. Or copy the skill folder (.claude/skills/phyai-model-implement in mingti-org/phyai) into .claude/skills/phyai-model-implement in your project. Claude Code loads it when a task matches its description.

How do I install Phyai Model Implement in Codex?

Run `npx skills add mingti-org/phyai --skill phyai-model-implement -a codex`. Or copy the skill folder (.claude/skills/phyai-model-implement in mingti-org/phyai) into .agents/skills/phyai-model-implement in your project. Codex loads it when a task matches its description.

Can I use Phyai Model Implement in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mingti-org/phyai --skill phyai-model-implement -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/phyai-model-implement, .gemini/skills/phyai-model-implement, .github/skills/phyai-model-implement and .opencode/skills/phyai-model-implement in your project.

What does Phyai Model Implement need to run?

SKILL.md names no scripts, command-line tools or credentials: Phyai Model Implement is instructions for the agent only. Our summary lists: Python 3.

Does Phyai Model Implement access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Phyai Model Implement safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Phyai Model Implement use?

Phyai Model Implement is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Phyai Model Implement use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Phyai Model Implement?

Skills that share tags, products or a category with Phyai Model Implement: OpenSpec Guided Onboarding (Fission-AI/OpenSpec, 71k stars), Paseo Committee (getpaseo/paseo, 20k stars), Improve (fossasia/eventyay-interpretation, 1.6k stars) and Dsh Web Documentation (zhu1090093659/dsh-web, 8.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Phyai Model Implement?

mingti-org (a GitHub organization) maintains it in mingti-org/phyai, which has 129 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 6, 2026.

Source: mingti-org/phyai on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.