Agent skill

Add Archon Model

by areal-project in areal-project/AReaL

Guide for adding a new model to the Archon engine. An agent skill from areal-project/AReaL.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Add Archon Model

skills CLI
$ npx skills add areal-project/AReaL --skill add-archon-model -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install areal-project/AReaL add-archon-model --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/areal-project/AReaL.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/add-archon-model .claude/skills/add-archon-model && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-archon-model
GitHub stars
5.8k
Token cost
~4.9k tokens
SKILL.md length
1,273 words
Files
1
Skills in repo
11
Repo updated
First seen
Licence
Apache-2.0

At a glance

Guide for adding a new model to the Archon engine. An agent skill from areal-project/AReaL.

  • Works in 10 steps: Analyze the Target Model Architecture → Select the Reference Model → Implement args.py → …
  • User wants to add support for a new HuggingFace model architecture in ArchonEngine
  • SKILL.md covers When to Use, Prerequisites, Step-by-Step Guide and Reference Implementations, plus 5 more sections
  • Calls hf

What it does

Add Archon Model is an agent skill from areal-project/AReaL. Guide for adding a new model to the Archon engine. Use when user wants to add support for a new HuggingFace model architecture in ArchonEngine.

Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Model hubs and datasets. It works with Hugging Face. The repository describes itself as: The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible. The licence is Apache-2.0.

When your agent uses it

  • User wants to add support for a new HuggingFace model architecture in ArchonEngine
  • Tasks that involve Model hubs and datasets

Example prompts

  • “/add-archon-model”

Requirements

  • Python 3

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. Analyze the Target Model Architecture
  2. Select the Reference Model
  3. Implement args.py
  4. Implement model.py
  5. Implement rope.py
  6. Implement state_dict_adapter.py
  7. Implement parallelize.py
  8. Create spec.py and Register
  9. Register in init.py
  10. Verify and Test

What it can do on your machine

Read from SKILL.md and the folder at commit 298412a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • hf

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add Archon Model loads about 4.9k tokens when it runs. Until then it costs about 40 tokens; SKILL.md has 1,273 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~40
When it runs · the whole SKILL.md, loaded when a task matches
~4.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from areal-project/AReaL at commit 298412a, republished under its Apache-2.0 licence (© areal-project). 1,273 words, ~4,912 tokens.

Download SKILL.mdSave it as .claude/skills/add-archon-model/SKILL.md (or your agent's skills folder).
name
add-archon-model
description
Guide for adding a new model to the Archon engine. Use when user wants to add support for a new HuggingFace model architecture in ArchonEngine.

Add Archon Model

Add support for a new HuggingFace model architecture in the Archon training engine.

When to Use

This skill is triggered when:

  • User asks "how do I add a model to Archon?"
  • User wants to support a new model family (e.g., Llama, Mistral, DeepSeek) in ArchonEngine
  • User mentions adding a new ModelSpec or model type for Archon

Prerequisites

Before starting, ensure:

  • The target model is available on HuggingFace (has config.json with model_type)
  • You know the HuggingFace model ID (e.g., meta-llama/Llama-3-8B)
  • The model uses a standard transformer architecture (decoder-only)

Step-by-Step Guide

Step 1: Analyze the Target Model Architecture

Read the HuggingFace model's source code to extract key architecture information.

Action: Fetch and analyze the model's HuggingFace configuration and modeling files.

  1. Read the model's config.json (via AutoConfig.from_pretrained) to identify:

    • model_type string (this is the key used for registry lookup)
    • All architecture hyperparameters (hidden_size, num_layers, etc.)
    • Any model-specific fields (e.g., qk_norm, attention_bias, MoE fields)
  2. Read the HuggingFace modeling_*.py source to identify:

    • Attention variant: Does it have Q/K norm? Attention bias? Sliding window? Multi-latent attention?
    • FFN variant: SwiGLU (gate_proj + up_proj + down_proj)? GeGLU? Standard MLP?
    • MoE support: Does it have MoE layers? What router type? Shared experts?
    • RoPE variant: Standard RoPE? YaRN? NTK-aware scaling? What is the inv_freq formula?
    • Normalization: RMSNorm or LayerNorm? Pre-norm or post-norm? Elementwise affine?
    • Weight tying: Does tie_word_embeddings appear in config?
    • State dict key names: What are the HF weight key naming conventions?
  3. Summarize findings in a checklist like:

Target model: <name>
HF model_type: "<model_type>" (and variants like "<model_type>_moe" if applicable)
Attention: [standard GQA / with QK norm / with bias / sliding window / ...]
FFN: [SwiGLU / GeGLU / standard MLP / ...]
MoE: [no / yes - num_experts, top_k, shared_experts]
RoPE: [standard / YaRN / NTK-aware / ...]
Norm: [RMSNorm / LayerNorm] with [pre-norm / post-norm]
Weight tying: [yes / no]
Step 2: Select the Reference Model

Choose the closest existing implementation as a starting point:

Target characteristicsReferenceWhy
Dense-only, standard GQA, no QK normqwen2Simplest baseline, pure dense
Has QK norm, or has MoE supportqwen3Supports QK norm + MoE + shared experts

Action: Copy the reference model directory as the starting point:

areal/experimental/models/archon/<model>/
  __init__.py
  spec.py
  model/
    args.py
    model.py
    rope.py
    state_dict_adapter.py
  infra/
    parallelize.py
Step 3: Implement args.py

Adapt <Model>ModelArgs to match the target model's HuggingFace config fields.

Key changes from reference:

  1. Update the @dataclass fields to match the target model's hyperparameters:

    • Field names should use Archon conventions (dim, n_layers, n_heads, n_kv_heads, vocab_size, head_dim, hidden_dim, norm_eps, rope_theta, etc.)
    • Default values should match the smallest variant of the target model
    • Add model-specific fields (e.g., attention_bias, qk_norm, sliding_window)
  2. Update from_hf_config() to correctly map HuggingFace config attributes:

    • Use getattr(hf_config, "field_name", default) for optional fields
    • Handle variant-specific fields (e.g., MoE fields only present in MoE variants)
    • The method must return an instance of the model args class

Critical: Verify every field mapping against the HF model's config.json. Incorrect mappings here cause silent errors downstream.

Base class contract (BaseModelArgs):

python
@dataclass
class <Model>ModelArgs(BaseModelArgs):
    # ... model-specific fields ...

    @classmethod
    def from_hf_config(
        cls,
        hf_config: PretrainedConfig,
        is_critic: bool = False,
        **kwargs,
    ) -> <Model>ModelArgs:
        # Map HF config fields to Archon model args
        ...
Step 4: Implement model.py

Adapt the model architecture to match the target model.

Key components to adapt:

  1. Normalization (RMSNorm or similar):

    • Check if elementwise_affine is configurable
    • Check the epsilon default value
    • If the model uses LayerNorm, implement accordingly
  2. Attention module:

    • Q/K/V projection: Check bias presence (nn.Linear(..., bias=True/False))
    • QK norm: Add q_norm/k_norm if the model has them, remove if it doesn't
    • GQA: n_kv_heads < n_heads for grouped-query attention
    • Ulysses SP: Keep the set_cp_group / _sp_enabled pattern from the reference
    • Output projection: Check bias presence
  3. FeedForward module:

    • SwiGLU: w2(silu(w1(x)) * w3(x)) -- most common for modern LLMs
    • Check bias in linear layers
    • For MoE models: MoE module replaces FeedForward on designated layers
  4. TransformerBlock: Pre-norm (most modern LLMs) vs post-norm

    • MoE layer detection via _is_moe_layer() if applicable
  5. Top-level Model (<Model>Model(BaseArchonModel)):

    • tok_embeddings, layers (as ModuleDict), norm, output/score
    • init_weights(): Match initialization scheme from HF
    • init_buffers(): RoPE cache + MoE buffers
    • forward(): Must follow BaseArchonModel signature: (tokens, positions, cu_seqlens, max_seqlen, tree_attn_meta=None) -> Tensor

Base class contract (BaseArchonModel):

python
class <Model>Model(BaseArchonModel):
    def forward(self, tokens, positions, cu_seqlens, max_seqlen, tree_attn_meta=None) -> torch.Tensor: ...
    def init_weights(self) -> None: ...
    def init_buffers(self, buffer_device) -> None: ...
Step 5: Implement rope.py

Handle the rotary position embedding variant.

Options:

  1. Standard RoPE (same as qwen2/qwen3): Re-export from qwen2:

    python
    from areal.experimental.models.archon.qwen2.model.rope import (
        apply_rotary_emb,
        precompute_rope_cache,
        repeat_kv,
        reshape_for_broadcast,
        rotate_half,
    )
  2. Custom RoPE (YaRN, NTK-aware, etc.): Implement custom precompute_rope_cache() and apply_rotary_emb() functions. The key difference is usually in how inv_freq is computed (scaling factors, interpolation, etc.).

Step 6: Implement state_dict_adapter.py

Map between HuggingFace and Archon weight key names.

This is the most error-prone step. The adapter must correctly handle:

  1. Key name mapping (from_hf_map dict):

    • Embedding: model.embed_tokens.weight -> tok_embeddings.weight
    • Attention: model.layers.{}.self_attn.q_proj.weight -> layers.{}.attention.wq.weight
    • FFN: model.layers.{}.mlp.gate_proj.weight -> layers.{}.feed_forward.w1.weight
    • Norms: model.layers.{}.input_layernorm.weight -> layers.{}.attention_norm.weight
    • Output: lm_head.weight -> output.weight
    • Skip keys (set to None): rotary_emb.inv_freq (computed at runtime)
    • Model-specific keys: bias terms, QK norm weights, etc.
  2. Reverse mapping (to_hf_map): Auto-generated from from_hf_map

  3. MoE expert weights (if applicable): 3D<->2D conversion for expert weights. Copy the MoE handling from qwen3 if the model has MoE.

  4. Weight tying: Skip output.weight during to_hf() if tie_word_embeddings=True

Verification approach: After implementation, the adapter should satisfy:

python
# Roundtrip: archon -> hf -> archon preserves all keys
hf_sd = adapter.to_hf(archon_sd)
roundtrip_sd = adapter.from_hf(hf_sd)
assert set(roundtrip_sd.keys()) == set(archon_sd.keys())

Base class contract (BaseStateDictAdapter):

python
class <Model>StateDictAdapter(BaseStateDictAdapter):
    def from_hf(self, hf_state_dict) -> dict[str, Any]: ...
    def to_hf(self, archon_state_dict) -> dict[str, Any]: ...
    def convert_single_to_hf(self, name, tensor) -> list[tuple[str, torch.Tensor]]: ...
Step 7: Implement parallelize.py

Define the parallelization strategy for the model.

The parallelize function applies parallelism in this order:

  1. TP (Tensor Parallelism) -- shard attention/FFN across devices
  2. EP (Expert Parallelism) -- for MoE models only
  3. CP (Context Parallelism / Ulysses SP) -- sequence parallelism
  4. AC (Activation Checkpointing) -- memory optimization
  5. torch.compile -- compilation optimization
  6. FSDP (Fully Sharded Data Parallelism) -- data parallelism

Key adaptations by model architecture:

  • Attention with QK norm: wq/wk use use_local_output=False (DTensor output for norm), add SequenceParallel(sequence_dim=2) for q_norm/k_norm
  • Attention without QK norm: wq/wk/wv all use use_local_output=True
  • Attention with bias: Bias terms follow the same parallel plan as their weights
  • MoE layers: Separate TP plan for MoE input/output, router gate, and expert weights. Copy from qwen3's apply_moe_ep_tp() and apply_non_moe_tp()
  • Dense-only models: Simpler plan without MoE handling. Copy from qwen2

Function signature (must match ParallelizeFn protocol):

python
def parallelize_<model>(
    model: nn.Module,
    parallel_dims: ArchonParallelDims,
    param_dtype: torch.dtype = torch.bfloat16,
    reduce_dtype: torch.dtype = torch.float32,
    loss_parallel: bool = True,
    cpu_offload: bool = False,
    reshard_after_forward_policy: str = "default",
    ac_config: ActivationCheckpointConfig | None = None,
    enable_compile: bool = True,
) -> nn.Module:
Step 8: Create spec.py and Register

Assemble the ModelSpec and register it.

python
from areal.experimental.models.archon.model_spec import ModelSpec, register_model_spec
from areal.experimental.models.archon.pipeline_parallel import pipeline_llm
from areal.experimental.models.archon.<model>.infra.parallelize import parallelize_<model>
from areal.experimental.models.archon.<model>.model.args import <Model>ModelArgs
from areal.experimental.models.archon.<model>.model.model import <Model>Model
from areal.experimental.models.archon.<model>.model.state_dict_adapter import (
    <Model>StateDictAdapter,
)

<MODEL>_SPEC = ModelSpec(
    name="<Model>",
    model_class=<Model>Model,
    model_args_class=<Model>ModelArgs,
    state_dict_adapter_class=<Model>StateDictAdapter,
    parallelize_fn=parallelize_<model>,
    supported_model_types=frozenset({"<model_type>"}),  # From HF config.json
    pipelining_fn=pipeline_llm,
)

# Auto-register when module is imported
register_model_spec(<MODEL>_SPEC)

__all__ = ["<MODEL>_SPEC"]

Note: supported_model_types should include all HF model_type strings that this implementation handles (e.g., {"qwen3", "qwen3_moe"} for Qwen3).

Show full SKILL.md (599 more words)Show less
Step 9: Register in __init__.py

Add the import to areal/experimental/models/archon/__init__.py:

python
from areal.experimental.models.archon.<model> import spec as <model>_spec  # noqa: F401

This triggers auto-registration when the module is imported.

Step 10: Verify and Test

Verification should be done in stages, adapting based on available hardware and the test patterns in tests/experimental/archon/.

Before writing tests, examine the existing test files to understand current patterns:

tests/experimental/archon/
  conftest.py             -- Pytest configuration (version checks)
  utils.py                -- Shared utilities (model loading, comparison)
  test_qwen3_args.py      -- Args unit tests (CPU-only)
  test_state_dict_adapter.py  -- State dict roundtrip tests
  test_weight_sync.py     -- Weight completeness tests (meta device)
  test_forward.py         -- Forward precision comparison (single GPU)
  ...

Test stages (write tests appropriate for the model's complexity):

Stage 1: Args Tests (CPU-only, always write these)

Test from_hf_config() with mock HuggingFace configs:

python
# Pattern: Create mock PretrainedConfig, verify args mapping
from unittest.mock import MagicMock

def test_args_from_hf_config():
    hf_config = MagicMock()
    hf_config.hidden_size = 4096
    hf_config.num_hidden_layers = 32
    # ... set all required fields
    args = <Model>ModelArgs.from_hf_config(hf_config)
    assert args.dim == 4096
    assert args.n_layers == 32
Stage 2: State Dict Adapter Tests (CPU-only)

Test key mapping roundtrip:

python
def test_state_dict_roundtrip():
    # Create adapter with mock config
    adapter = <Model>StateDictAdapter(mock_config)
    # Create fake archon state dict with expected keys
    archon_sd = {"tok_embeddings.weight": torch.randn(vocab, dim), ...}
    # Roundtrip
    hf_sd = adapter.to_hf(archon_sd)
    roundtrip = adapter.from_hf(hf_sd)
    assert set(roundtrip.keys()) == set(archon_sd.keys())
Stage 3: Weight Completeness (meta device, CPU-only)

Verify all model parameters have HF mappings:

python
def test_weight_completeness():
    # Create model on meta device
    with torch.device("meta"):
        model = <Model>Model(args)
    adapter = <Model>StateDictAdapter(hf_config)
    # Check every archon param has a HF mapping
    for name, _ in model.named_parameters():
        hf_pairs = adapter.convert_single_to_hf(name, torch.empty(0))
        assert len(hf_pairs) > 0, f"No HF mapping for {name}"
Stage 4: Forward Precision (single GPU, if available)

Compare Archon model output against HuggingFace reference:

python
@pytest.mark.skipif(not torch.cuda.is_available(), reason="Requires CUDA")
def test_forward_matches_hf():
    # Load both HF and Archon models
    # Run forward on same input
    # Compare logits within tolerance

Important: Do NOT hardcode the test categories. Inspect the existing test files in tests/experimental/archon/ and follow the same patterns, fixtures, and markers. Adapt test scope to the model's specific features (e.g., add MoE-specific tests only if the model has MoE).

Reference Implementations

ModelDirectoryFeatures
Qwen2areal/experimental/models/archon/qwen2/Dense, attention bias, no QK norm
Qwen3areal/experimental/models/archon/qwen3/Dense + MoE, QK norm, no attention bias, shared experts

Architecture Decision Map

Featureqwen2qwen3What to check in target model
Attention biasYesNoattention_bias in HF config
QK normNoYesqk_norm in HF config or QKNorm module in modeling file
MoENoYesnum_experts/num_local_experts in HF config
Shared expertsNoYesnum_shared_experts in HF config
Decoder sparse stepNoYesdecoder_sparse_step in HF config
Weight tyingBothBothtie_word_embeddings in HF config
RoPEStandardStandard (re-export qwen2)Check inv_freq formula in HF modeling code

Common Mistakes

  • Not mapping all HF keys in state_dict_adapter.py (causes silent weight drops)
  • Wrong from_hf_config() field mapping (uses wrong HF config attribute name)
  • Forgetting to handle None keys in from_hf_map (keys to skip like rotary_emb.inv_freq)
  • Missing MoE expert weight 3D<->2D conversion when model has MoE
  • Wrong TP plan for attention with/without QK norm (use_local_output must match)
  • Forgetting to add import line in areal/experimental/models/archon/__init__.py
  • Not including all model_type variants in supported_model_types frozenset
  • Using print instead of areal.utils.logging.getLogger()

File Checklist

After completion, verify all files exist and are consistent:

  • areal/experimental/models/archon/<model>/__init__.py
  • areal/experimental/models/archon/<model>/spec.py -- ModelSpec + register
  • areal/experimental/models/archon/<model>/model/args.py -- ModelArgs + from_hf_config
  • areal/experimental/models/archon/<model>/model/model.py -- Model + Attention + FFN
  • areal/experimental/models/archon/<model>/model/rope.py -- RoPE (or re-export)
  • areal/experimental/models/archon/<model>/model/state_dict_adapter.py -- Key mapping
  • areal/experimental/models/archon/<model>/infra/parallelize.py -- Parallel strategy
  • areal/experimental/models/archon/__init__.py -- Import line added
  • tests/experimental/archon/test_<model>_*.py -- Tests

<!--
================================================================================
                            MAINTAINER GUIDE
================================================================================

Canonical location: .agents/skills/add-archon-model/SKILL.md
Mirrors: .opencode/skills/add-archon-model/SKILL.md, .claude/skills/add-archon-model/SKILL.md
Invocation: $add-archon-model (Codex) / /add-archon-model (OpenCode, Claude Code)

## Purpose

Semi-automated guide for adding new model architectures to the Archon training engine.
Unlike simpler skills (add-reward, add-dataset), this skill actively guides the agent to:
1. Analyze HuggingFace source code to extract architecture details
2. Select the closest reference implementation (qwen2 or qwen3)
3. Generate code skeletons adapted to the target architecture
4. Create appropriate tests based on existing test patterns

## How to Update

### When New Reference Models Are Added
1. Add to "Reference Implementations" table
2. Update "Architecture Decision Map" with new feature columns
3. Update Step 2 (reference selection) with new options

### When Base Classes Change
1. Update contract signatures in Steps 3, 4, 6, 7
2. Update file checklist if new files are required

### When ModelSpec Changes
1. Update Step 8 with new ModelSpec fields
2. Update spec.py template

### When Test Patterns Change
1. Update Step 10 with new test patterns
2. Do NOT hardcode test categories -- keep it flexible

### Important Design Decisions
- This skill is SEMI-AUTOMATED: Claude should read HF source and generate code,
  not just provide templates for the user to fill in manually
- The skill references existing test files rather than hardcoding test categories,
  ensuring it stays current as the test suite evolves
- Reference model selection (qwen2 vs qwen3) is based on MoE and QK norm presence

================================================================================
-->

© areal-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/add-archon-model of areal-project/AReaL.

Open the folder on GitHubat commit 298412a

Compare with similar skills

Add Archon Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add Archon Model compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add Archon Model this skillareal-project/AReaL5.8k—~4.9kAutomated safety check: PassApache-2.0
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Upload Post Imagehuggingface/blog3.5k—~1.1kAutomated safety check: PassNone
Esmfold2JimLiu/science-skills2284 repos~2.5kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Upload Post Image

    huggingface/blog

    Official

    A skill your agent uses when adding or migrating non-thumbnail images for a Hugging Face Blog post.

    3.5k GitHub stars~1.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    228 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed

More from areal-project/AReaL

All 11 skills in this repo
  • Add Dataset

    areal-project/AReaL

    Guide for adding a new dataset loader to AReaL. An agent skill from areal-project/AReaL.

    5.8k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Add Reward

    areal-project/AReaL

    Guide for adding a new reward function to AReaL. An agent skill from areal-project/AReaL.

    5.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Add Unit Tests

    areal-project/AReaL

    Guide for adding unit tests to AReaL. An agent skill from areal-project/AReaL.

    5.8k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Add Workflow

    areal-project/AReaL

    Guide for adding a new RolloutWorkflow to AReaL. An agent skill from areal-project/AReaL.

    5.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Debug Distributed

    areal-project/AReaL

    Guide for debugging distributed training issues in AReaL. An agent skill from areal-project/AReaL.

    5.8k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Review PR

    areal-project/AReaL

    Read-only pull request review workflow with risk analysis, targeted checklists, and Codex subagent consultation.

    5.8k GitHub stars~704 tokensUpdated today
    Auto-check passed

Works with

Questions about Add Archon Model

What does Add Archon Model do?

Guide for adding a new model to the Archon engine. An agent skill from areal-project/AReaL. Add Archon Model is an agent skill from areal-project/AReaL. Guide for adding a new model to the Archon engine.

When should I use Add Archon Model?

Add Archon Model fits situations like: user wants to add support for a new HuggingFace model architecture in ArchonEngine; tasks that involve Model hubs and datasets.

How do I install Add Archon Model in Claude Code?

Run `npx skills add areal-project/AReaL --skill add-archon-model -a claude-code`. Or copy the skill folder (.agents/skills/add-archon-model in areal-project/AReaL) into .claude/skills/add-archon-model in your project. Claude Code loads it when a task matches its description.

How do I install Add Archon Model in Codex?

Run `npx skills add areal-project/AReaL --skill add-archon-model -a codex`. Or copy the skill folder (.agents/skills/add-archon-model in areal-project/AReaL) into .agents/skills/add-archon-model in your project. Codex loads it when a task matches its description.

Can I use Add Archon Model in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add areal-project/AReaL --skill add-archon-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-archon-model, .gemini/skills/add-archon-model, .github/skills/add-archon-model and .opencode/skills/add-archon-model in your project.

What does Add Archon Model need to run?

Going by SKILL.md and its folder, Add Archon Model needs the command-line tools its instructions call (hf). Our summary lists: Python 3.

Does Add Archon Model access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Add Archon Model safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Add Archon Model use?

Add Archon Model is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add Archon Model use?

About 4.9k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Add Archon Model?

Skills that share tags, products or a category with Add Archon Model: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars) and Upload Post Image (huggingface/blog, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add Archon Model?

areal-project (a GitHub organization) maintains it in areal-project/AReaL, which has 5,824 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 10, 2026.

Source: areal-project/AReaL on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.