LLM Benchmarking with lm-evaluation-harness
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Guide for adding a new model to the Archon engine. An agent skill from areal-project/AReaL.
$ npx skills add areal-project/AReaL --skill add-archon-model -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install areal-project/AReaL add-archon-model --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/areal-project/AReaL.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/add-archon-model .claude/skills/add-archon-model && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "add-archon-model" agent skill from https://github.com/areal-project/AReaL/tree/main/.agents/skills/add-archon-model into .claude/skills/add-archon-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-archon-model", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/areal-project/AReaL/tree/main/.agents/skills/add-archon-modelType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add areal-project/AReaL --skill add-archon-model -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install areal-project/AReaL add-archon-model --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/areal-project/AReaL.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/add-archon-model .agents/skills/add-archon-model && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "add-archon-model" agent skill from https://github.com/areal-project/AReaL/tree/main/.agents/skills/add-archon-model into .agents/skills/add-archon-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-archon-model", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add areal-project/AReaL --skill add-archon-model -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install areal-project/AReaL add-archon-model --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/areal-project/AReaL.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/add-archon-model .cursor/skills/add-archon-model && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "add-archon-model" agent skill from https://github.com/areal-project/AReaL/tree/main/.agents/skills/add-archon-model into .cursor/skills/add-archon-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-archon-model", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/areal-project/AReaL.git --path .agents/skills/add-archon-model--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add areal-project/AReaL --skill add-archon-model -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install areal-project/AReaL add-archon-model --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/areal-project/AReaL.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/add-archon-model .gemini/skills/add-archon-model && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "add-archon-model" agent skill from https://github.com/areal-project/AReaL/tree/main/.agents/skills/add-archon-model into .gemini/skills/add-archon-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-archon-model", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install areal-project/AReaL add-archon-modelInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add areal-project/AReaL --skill add-archon-model -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/areal-project/AReaL.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/add-archon-model .github/skills/add-archon-model && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "add-archon-model" agent skill from https://github.com/areal-project/AReaL/tree/main/.agents/skills/add-archon-model into .github/skills/add-archon-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-archon-model", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add areal-project/AReaL --skill add-archon-model -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install areal-project/AReaL add-archon-model --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/areal-project/AReaL.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/add-archon-model .opencode/skills/add-archon-model && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "add-archon-model" agent skill from https://github.com/areal-project/AReaL/tree/main/.agents/skills/add-archon-model into .opencode/skills/add-archon-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-archon-model", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
add-archon-modelGuide for adding a new model to the Archon engine. An agent skill from areal-project/AReaL.
Add Archon Model is an agent skill from areal-project/AReaL. Guide for adding a new model to the Archon engine. Use when user wants to add support for a new HuggingFace model architecture in ArchonEngine.
Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Model hubs and datasets. It works with Hugging Face. The repository describes itself as: The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible. The licence is Apache-2.0.
10 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 298412a. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
hfFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Add Archon Model loads about 4.9k tokens when it runs. Until then it costs about 40 tokens; SKILL.md has 1,273 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from areal-project/AReaL at commit 298412a, republished under its Apache-2.0 licence (© areal-project). 1,273 words, ~4,912 tokens.
.claude/skills/add-archon-model/SKILL.md (or your agent's skills folder).Add support for a new HuggingFace model architecture in the Archon training engine.
This skill is triggered when:
ModelSpec or model type for ArchonBefore starting, ensure:
config.json with model_type)meta-llama/Llama-3-8B)Read the HuggingFace model's source code to extract key architecture information.
Action: Fetch and analyze the model's HuggingFace configuration and modeling files.
Read the model's config.json (via AutoConfig.from_pretrained) to identify:
model_type string (this is the key used for registry lookup)qk_norm, attention_bias, MoE fields)Read the HuggingFace modeling_*.py source to identify:
tie_word_embeddings appear in config?Summarize findings in a checklist like:
Target model: <name>
HF model_type: "<model_type>" (and variants like "<model_type>_moe" if applicable)
Attention: [standard GQA / with QK norm / with bias / sliding window / ...]
FFN: [SwiGLU / GeGLU / standard MLP / ...]
MoE: [no / yes - num_experts, top_k, shared_experts]
RoPE: [standard / YaRN / NTK-aware / ...]
Norm: [RMSNorm / LayerNorm] with [pre-norm / post-norm]
Weight tying: [yes / no]Choose the closest existing implementation as a starting point:
| Target characteristics | Reference | Why |
|---|---|---|
| Dense-only, standard GQA, no QK norm | qwen2 | Simplest baseline, pure dense |
| Has QK norm, or has MoE support | qwen3 | Supports QK norm + MoE + shared experts |
Action: Copy the reference model directory as the starting point:
areal/experimental/models/archon/<model>/
__init__.py
spec.py
model/
args.py
model.py
rope.py
state_dict_adapter.py
infra/
parallelize.pyargs.pyAdapt <Model>ModelArgs to match the target model's HuggingFace config fields.
Key changes from reference:
Update the @dataclass fields to match the target model's hyperparameters:
dim, n_layers, n_heads,
n_kv_heads, vocab_size, head_dim, hidden_dim, norm_eps, rope_theta,
etc.)attention_bias, qk_norm, sliding_window)Update from_hf_config() to correctly map HuggingFace config attributes:
getattr(hf_config, "field_name", default) for optional fieldsCritical: Verify every field mapping against the HF model's config.json. Incorrect
mappings here cause silent errors downstream.
Base class contract (BaseModelArgs):
@dataclass
class <Model>ModelArgs(BaseModelArgs):
# ... model-specific fields ...
@classmethod
def from_hf_config(
cls,
hf_config: PretrainedConfig,
is_critic: bool = False,
**kwargs,
) -> <Model>ModelArgs:
# Map HF config fields to Archon model args
...model.pyAdapt the model architecture to match the target model.
Key components to adapt:
Normalization (RMSNorm or similar):
elementwise_affine is configurableLayerNorm, implement accordinglyAttention module:
nn.Linear(..., bias=True/False))q_norm/k_norm if the model has them, remove if it doesn'tn_kv_heads < n_heads for grouped-query attentionset_cp_group / _sp_enabled pattern from the referenceFeedForward module:
w2(silu(w1(x)) * w3(x)) -- most common for modern LLMsMoE module replaces FeedForward on designated layersTransformerBlock: Pre-norm (most modern LLMs) vs post-norm
_is_moe_layer() if applicableTop-level Model (<Model>Model(BaseArchonModel)):
tok_embeddings, layers (as ModuleDict), norm, output/scoreinit_weights(): Match initialization scheme from HFinit_buffers(): RoPE cache + MoE buffersforward(): Must follow BaseArchonModel signature:
(tokens, positions, cu_seqlens, max_seqlen, tree_attn_meta=None) -> TensorBase class contract (BaseArchonModel):
class <Model>Model(BaseArchonModel):
def forward(self, tokens, positions, cu_seqlens, max_seqlen, tree_attn_meta=None) -> torch.Tensor: ...
def init_weights(self) -> None: ...
def init_buffers(self, buffer_device) -> None: ...rope.pyHandle the rotary position embedding variant.
Options:
Standard RoPE (same as qwen2/qwen3): Re-export from qwen2:
from areal.experimental.models.archon.qwen2.model.rope import (
apply_rotary_emb,
precompute_rope_cache,
repeat_kv,
reshape_for_broadcast,
rotate_half,
)Custom RoPE (YaRN, NTK-aware, etc.): Implement custom precompute_rope_cache()
and apply_rotary_emb() functions. The key difference is usually in how inv_freq
is computed (scaling factors, interpolation, etc.).
state_dict_adapter.pyMap between HuggingFace and Archon weight key names.
This is the most error-prone step. The adapter must correctly handle:
Key name mapping (from_hf_map dict):
model.embed_tokens.weight -> tok_embeddings.weightmodel.layers.{}.self_attn.q_proj.weight ->
layers.{}.attention.wq.weightmodel.layers.{}.mlp.gate_proj.weight -> layers.{}.feed_forward.w1.weightmodel.layers.{}.input_layernorm.weight ->
layers.{}.attention_norm.weightlm_head.weight -> output.weightNone): rotary_emb.inv_freq (computed at runtime)Reverse mapping (to_hf_map): Auto-generated from from_hf_map
MoE expert weights (if applicable): 3D<->2D conversion for expert weights. Copy the MoE handling from qwen3 if the model has MoE.
Weight tying: Skip output.weight during to_hf() if tie_word_embeddings=True
Verification approach: After implementation, the adapter should satisfy:
# Roundtrip: archon -> hf -> archon preserves all keys
hf_sd = adapter.to_hf(archon_sd)
roundtrip_sd = adapter.from_hf(hf_sd)
assert set(roundtrip_sd.keys()) == set(archon_sd.keys())Base class contract (BaseStateDictAdapter):
class <Model>StateDictAdapter(BaseStateDictAdapter):
def from_hf(self, hf_state_dict) -> dict[str, Any]: ...
def to_hf(self, archon_state_dict) -> dict[str, Any]: ...
def convert_single_to_hf(self, name, tensor) -> list[tuple[str, torch.Tensor]]: ...parallelize.pyDefine the parallelization strategy for the model.
The parallelize function applies parallelism in this order:
Key adaptations by model architecture:
use_local_output=False (DTensor output for
norm), add SequenceParallel(sequence_dim=2) for q_norm/k_normuse_local_output=Trueapply_moe_ep_tp() and apply_non_moe_tp()Function signature (must match ParallelizeFn protocol):
def parallelize_<model>(
model: nn.Module,
parallel_dims: ArchonParallelDims,
param_dtype: torch.dtype = torch.bfloat16,
reduce_dtype: torch.dtype = torch.float32,
loss_parallel: bool = True,
cpu_offload: bool = False,
reshard_after_forward_policy: str = "default",
ac_config: ActivationCheckpointConfig | None = None,
enable_compile: bool = True,
) -> nn.Module:spec.py and RegisterAssemble the ModelSpec and register it.
from areal.experimental.models.archon.model_spec import ModelSpec, register_model_spec
from areal.experimental.models.archon.pipeline_parallel import pipeline_llm
from areal.experimental.models.archon.<model>.infra.parallelize import parallelize_<model>
from areal.experimental.models.archon.<model>.model.args import <Model>ModelArgs
from areal.experimental.models.archon.<model>.model.model import <Model>Model
from areal.experimental.models.archon.<model>.model.state_dict_adapter import (
<Model>StateDictAdapter,
)
<MODEL>_SPEC = ModelSpec(
name="<Model>",
model_class=<Model>Model,
model_args_class=<Model>ModelArgs,
state_dict_adapter_class=<Model>StateDictAdapter,
parallelize_fn=parallelize_<model>,
supported_model_types=frozenset({"<model_type>"}), # From HF config.json
pipelining_fn=pipeline_llm,
)
# Auto-register when module is imported
register_model_spec(<MODEL>_SPEC)
__all__ = ["<MODEL>_SPEC"]Note: supported_model_types should include all HF model_type strings that this
implementation handles (e.g., {"qwen3", "qwen3_moe"} for Qwen3).
__init__.pyAdd the import to areal/experimental/models/archon/__init__.py:
from areal.experimental.models.archon.<model> import spec as <model>_spec # noqa: F401This triggers auto-registration when the module is imported.
Verification should be done in stages, adapting based on available hardware and the test
patterns in tests/experimental/archon/.
Before writing tests, examine the existing test files to understand current patterns:
tests/experimental/archon/
conftest.py -- Pytest configuration (version checks)
utils.py -- Shared utilities (model loading, comparison)
test_qwen3_args.py -- Args unit tests (CPU-only)
test_state_dict_adapter.py -- State dict roundtrip tests
test_weight_sync.py -- Weight completeness tests (meta device)
test_forward.py -- Forward precision comparison (single GPU)
...Test stages (write tests appropriate for the model's complexity):
Test from_hf_config() with mock HuggingFace configs:
# Pattern: Create mock PretrainedConfig, verify args mapping
from unittest.mock import MagicMock
def test_args_from_hf_config():
hf_config = MagicMock()
hf_config.hidden_size = 4096
hf_config.num_hidden_layers = 32
# ... set all required fields
args = <Model>ModelArgs.from_hf_config(hf_config)
assert args.dim == 4096
assert args.n_layers == 32Test key mapping roundtrip:
def test_state_dict_roundtrip():
# Create adapter with mock config
adapter = <Model>StateDictAdapter(mock_config)
# Create fake archon state dict with expected keys
archon_sd = {"tok_embeddings.weight": torch.randn(vocab, dim), ...}
# Roundtrip
hf_sd = adapter.to_hf(archon_sd)
roundtrip = adapter.from_hf(hf_sd)
assert set(roundtrip.keys()) == set(archon_sd.keys())Verify all model parameters have HF mappings:
def test_weight_completeness():
# Create model on meta device
with torch.device("meta"):
model = <Model>Model(args)
adapter = <Model>StateDictAdapter(hf_config)
# Check every archon param has a HF mapping
for name, _ in model.named_parameters():
hf_pairs = adapter.convert_single_to_hf(name, torch.empty(0))
assert len(hf_pairs) > 0, f"No HF mapping for {name}"Compare Archon model output against HuggingFace reference:
@pytest.mark.skipif(not torch.cuda.is_available(), reason="Requires CUDA")
def test_forward_matches_hf():
# Load both HF and Archon models
# Run forward on same input
# Compare logits within toleranceImportant: Do NOT hardcode the test categories. Inspect the existing test files in
tests/experimental/archon/ and follow the same patterns, fixtures, and markers. Adapt
test scope to the model's specific features (e.g., add MoE-specific tests only if the
model has MoE).
| Model | Directory | Features |
|---|---|---|
| Qwen2 | areal/experimental/models/archon/qwen2/ | Dense, attention bias, no QK norm |
| Qwen3 | areal/experimental/models/archon/qwen3/ | Dense + MoE, QK norm, no attention bias, shared experts |
| Feature | qwen2 | qwen3 | What to check in target model |
|---|---|---|---|
| Attention bias | Yes | No | attention_bias in HF config |
| QK norm | No | Yes | qk_norm in HF config or QKNorm module in modeling file |
| MoE | No | Yes | num_experts/num_local_experts in HF config |
| Shared experts | No | Yes | num_shared_experts in HF config |
| Decoder sparse step | No | Yes | decoder_sparse_step in HF config |
| Weight tying | Both | Both | tie_word_embeddings in HF config |
| RoPE | Standard | Standard (re-export qwen2) | Check inv_freq formula in HF modeling code |
state_dict_adapter.py (causes silent weight drops)from_hf_config() field mapping (uses wrong HF config attribute name)None keys in from_hf_map (keys to skip like
rotary_emb.inv_freq)use_local_output must match)areal/experimental/models/archon/__init__.pymodel_type variants in supported_model_types frozensetprint instead of areal.utils.logging.getLogger()After completion, verify all files exist and are consistent:
areal/experimental/models/archon/<model>/__init__.pyareal/experimental/models/archon/<model>/spec.py -- ModelSpec + registerareal/experimental/models/archon/<model>/model/args.py -- ModelArgs +
from_hf_configareal/experimental/models/archon/<model>/model/model.py -- Model + Attention +
FFNareal/experimental/models/archon/<model>/model/rope.py -- RoPE (or re-export)areal/experimental/models/archon/<model>/model/state_dict_adapter.py -- Key
mappingareal/experimental/models/archon/<model>/infra/parallelize.py -- Parallel
strategyareal/experimental/models/archon/__init__.py -- Import line addedtests/experimental/archon/test_<model>_*.py -- Tests<!--
================================================================================
MAINTAINER GUIDE
================================================================================
Canonical location: .agents/skills/add-archon-model/SKILL.md
Mirrors: .opencode/skills/add-archon-model/SKILL.md, .claude/skills/add-archon-model/SKILL.md
Invocation: $add-archon-model (Codex) / /add-archon-model (OpenCode, Claude Code)
## Purpose
Semi-automated guide for adding new model architectures to the Archon training engine.
Unlike simpler skills (add-reward, add-dataset), this skill actively guides the agent to:
1. Analyze HuggingFace source code to extract architecture details
2. Select the closest reference implementation (qwen2 or qwen3)
3. Generate code skeletons adapted to the target architecture
4. Create appropriate tests based on existing test patterns
## How to Update
### When New Reference Models Are Added
1. Add to "Reference Implementations" table
2. Update "Architecture Decision Map" with new feature columns
3. Update Step 2 (reference selection) with new options
### When Base Classes Change
1. Update contract signatures in Steps 3, 4, 6, 7
2. Update file checklist if new files are required
### When ModelSpec Changes
1. Update Step 8 with new ModelSpec fields
2. Update spec.py template
### When Test Patterns Change
1. Update Step 10 with new test patterns
2. Do NOT hardcode test categories -- keep it flexible
### Important Design Decisions
- This skill is SEMI-AUTOMATED: Claude should read HF source and generate code,
not just provide templates for the user to fill in manually
- The skill references existing test files rather than hardcoding test categories,
ensuring it stays current as the test suite evolves
- Reference model selection (qwen2 vs qwen3) is based on MoE and QK norm presence
================================================================================
-->
© areal-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/add-archon-model of areal-project/AReaL.
Open the folder on GitHubat commit 298412a
Add Archon Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Add Archon Model this skillareal-project/AReaL | 5.8k | — | ~4.9k | Automated safety check: Pass | Apache-2.0 | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Upload Post Imagehuggingface/blog | 3.5k | — | ~1.1k | Automated safety check: Pass | None | |
| Esmfold2JimLiu/science-skills | 228 | 4 repos | ~2.5k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
huggingface/blog
A skill your agent uses when adding or migrating non-thumbnail images for a Hugging Face Blog post.
JimLiu/science-skills
Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.
R6410418/Jackrong-llm-finetuning-guide
Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.
areal-project/AReaL
Guide for adding a new dataset loader to AReaL. An agent skill from areal-project/AReaL.
areal-project/AReaL
Guide for adding a new reward function to AReaL. An agent skill from areal-project/AReaL.
areal-project/AReaL
Guide for adding unit tests to AReaL. An agent skill from areal-project/AReaL.
areal-project/AReaL
Guide for adding a new RolloutWorkflow to AReaL. An agent skill from areal-project/AReaL.
areal-project/AReaL
Guide for debugging distributed training issues in AReaL. An agent skill from areal-project/AReaL.
areal-project/AReaL
Read-only pull request review workflow with risk analysis, targeted checklists, and Codex subagent consultation.
Works with
Categories
Guide for adding a new model to the Archon engine. An agent skill from areal-project/AReaL. Add Archon Model is an agent skill from areal-project/AReaL. Guide for adding a new model to the Archon engine.
Add Archon Model fits situations like: user wants to add support for a new HuggingFace model architecture in ArchonEngine; tasks that involve Model hubs and datasets.
Run `npx skills add areal-project/AReaL --skill add-archon-model -a claude-code`. Or copy the skill folder (.agents/skills/add-archon-model in areal-project/AReaL) into .claude/skills/add-archon-model in your project. Claude Code loads it when a task matches its description.
Run `npx skills add areal-project/AReaL --skill add-archon-model -a codex`. Or copy the skill folder (.agents/skills/add-archon-model in areal-project/AReaL) into .agents/skills/add-archon-model in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add areal-project/AReaL --skill add-archon-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-archon-model, .gemini/skills/add-archon-model, .github/skills/add-archon-model and .opencode/skills/add-archon-model in your project.
Going by SKILL.md and its folder, Add Archon Model needs the command-line tools its instructions call (hf). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Add Archon Model is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.9k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Add Archon Model: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars) and Upload Post Image (huggingface/blog, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
areal-project (a GitHub organization) maintains it in areal-project/AReaL, which has 5,824 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 10, 2026.
Source: areal-project/AReaL on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.