Agent skill

Wally Architecture

by RunanywhereAI in RunanywhereAI/wally

Where Wally logic belongs — command layering, proto as SOT, kit vs CLI ownership, Apple MLX host vs wally-cxx.

MITAuto-check passed

Install Wally Architecture

skills CLI
$ npx skills add RunanywhereAI/wally --skill wally-architecture -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install RunanywhereAI/wally wally-architecture --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/RunanywhereAI/wally.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/wally-architecture .claude/skills/wally-architecture && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
wally-architecture
GitHub stars
1.6k
Token cost
~992 tokens
SKILL.md length
428 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Where Wally logic belongs — command layering, proto as SOT, kit vs CLI ownership, Apple MLX host vs wally-cxx.

  • Adding a command
  • SKILL.md covers Ownership, Layering rules, Command surface and Apple MLX host, plus 1 more section
  • Calls cmake and bash
  • Moving inference logic

What it does

Wally Architecture is an agent skill from RunanywhereAI/wally. Where Wally logic belongs — command layering, proto as SOT, kit vs CLI ownership, Apple MLX host vs wally-cxx. Use when adding a command, moving inference logic, or deciding whether a bug is SDK or CLI.

Its SKILL.md is about 990 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It works with C++. The repository describes itself as: Get up and running with GLM-5.3-flash, DeepSeek, Gemma and other open source frontier models. The licence is MIT.

When your agent uses it

  • Adding a command
  • Moving inference logic
  • Deciding whether a bug is SDK

Example prompts

  • “/wally-architecture”

What it can do on your machine

Read from SKILL.md and the folder at commit 3819f46. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • cmake
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Wally Architecture loads about 992 tokens when it runs. Until then it costs about 55 tokens; SKILL.md has 428 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~55
When it runs · the whole SKILL.md, loaded when a task matches
~992

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from RunanywhereAI/wally at commit 3819f46, republished under its MIT licence (© RunanywhereAI). 428 words, ~992 tokens.

Download SKILL.mdSave it as .claude/skills/wally-architecture/SKILL.md (or your agent's skills folder).
name
wally-architecture
description
Where Wally logic belongs — command layering, proto as SOT, kit vs CLI ownership, Apple MLX host vs wally-cxx. Use when adding a command, moving inference logic, or deciding whether a bug is SDK or CLI.

Wally architecture

Repo: RunanywhereAI/wally. Product CLI named wally. The application is a Rust crate (Cargo.toml); it consumes a packaged C++ desktop kit via CMake's find_package(RunAnywhere). CMake stays the build entry and the owner of the kit, and runs cargo for the application code. It does not add_subdirectory or FetchContent the SDK, and it does not compile llama.cpp / Sherpa / ONNX / MLX from source.

Product version (project(wally VERSION …) in CMakeLists.txt) is independent of the SDK kit pin in cmake/sdk-pin.cmake.

Ownership

text
argv / flags / env
  -> src/commands/cmd_*.rs      thin: parse → bootstrap() → one rac_* → render
  -> C++ desktop kit            catalog, download, lifecycle, generate, serve
  -> engines (in the kit)       llama.cpp, Sherpa, ONNX, MLX (Apple host);
                                NeuRT / QHexRT only when the private overlay
                                was applied at configure time

The kit owns truth: models, backends, proto contracts, download, inference. The CLI renders and interacts. If a command is composing a multi-step bootstrap, hardcoding an engine name, or post-processing model output, that is a bug in the SDK — fix it there, then consume a new kit (wally-kit-pin).

Layering rules

  • Command modules stay thin. Business rules do not live in command callbacks (src/cli/mod.rs's CLI11-compatible builder), Swift, or the REPL.
  • Proto is the SOT. crate::io::proto::v1::* (prost) is generated at build time straight from the kit's own .proto files (build.rs, via protox), after checking the kit's SCHEMA_LOCK hash against the pin in versions.toml. Wally never runs protoc, and nothing generated is committed. Parse rac_* byte buffers with io::proto into v1::*.
  • No parallel hand-written enums for values that exist in idl/*.proto.
  • Structured errors. Machine-readable codes from the ABI; human text on stderr. Results on stdout. --json prints exactly one document on stdout.
  • Never log API keys, tokens, or Authorization headers.

RunAnywhere::commons applies google=runanywhere_internal when the kit was built with namespace isolation. Never find_package(Protobuf) against Homebrew.

Show full SKILL.md (179 more words)Show less

Command surface

Dual grammar: spec namespaces (llm generate, models download) plus terminal aliases (run, pull, stt). One configure_* wires both (src/commands/mod.rs).

Do not reintroduce FetchContent of the SDK, a second inference backend tree, or a retired MetalRT / hardcoded catalog.

Apple MLX host

On Apple Silicon, cmake --build produces build/wally (Swift host wrapping wally_run_main, the crate's exported entry point in src/lib.rs). Users never run wally-cxx; that name exists only so CMake cannot overwrite the product binary. Independent clones set WALLY_SDK_SWIFT_PATH to a runanywhere-sdks checkout (CI does this). Nested EXTERNAL/Wally finds ../../Package.swift automatically. Disable with -DWALLY_APPLE_MLX_HOST=OFF only for a fast loop that skips the Swift build.

NeuRT image gen is #[cfg(wally_has_neurt)] in src/commands/cmd_image.rs (the WALLY_HAS_NEURT capability flag CMake writes to wally-build.env becomes that cfg in build.rs), true only when the NeuRT overlay is applied. Public bottles stay OSS. --engine qhexrt / qnn / npu / hexagon map to INFERENCE_FRAMEWORK_QHEXRT. Local HNPU trees are inferred from v75/v79/ v81, context.bin, or *_HNPU directory names.

Skills trees

Canonical: .claude/skills/. Mirror: .agents/skills/. Edit canonical, then bash scripts/ci/check-agents-sync.sh --fix. Never hand-edit the mirror. CLAUDE.md is a symlink to AGENTS.md.

© RunanywhereAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/wally-architecture of RunanywhereAI/wally.

Open the folder on GitHubat commit 3819f46

Compare with similar skills

Wally Architecture next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Wally Architecture compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Wally Architecture this skillRunanywhereAI/wally1.6k—~992Automated safety check: PassMIT
Paddle BuildPaddlePaddle/Paddle24k—~1kAutomated safety check: PassApache-2.0
Fory Releaseapache/fory4.6k—~2.9kAutomated safety check: PassApache-2.0
ONNX Runtime Shape Inference Safety Auditmicrosoft/onnxruntime22k—~3.3kAutomated safety check: PassMIT
Code Audit3stoneBrother/code-audit8921 repos~2.7kAutomated safety check: PassNone
Leetcuda Tex To Readthedocsxlite-dev/LeetCUDA12k—~902Automated safety check: PassGPL-3.0

Similar skills

  • Paddle Build

    PaddlePaddle/Paddle

    A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.

    24k GitHub stars~1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Fory Release

    apache/fory

    Prepare an Apache Fory release candidate from a clean release branch, including the version bump, RC tag, JVM staging, ASF source artifacts, SVN upload, and vote email.

    4.6k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Finds and fixes out-of-range output writes in ONNX Runtime operator shape-inference functions where a getNumOutputs guard admits too few outputs.

    22k GitHub stars~3.3k tokensUpdated today
    SecurityAuto-check passed
  • Code Audit

    3stoneBrother/code-audit

    Professional code security audit skill covering 55+ vulnerability types.

    892 GitHub starsUsed in 1 repo~2.7k tokens
    SecurityAuto-check passed
  • Leetcuda Tex To Readthedocs

    xlite-dev/LeetCUDA

    LeetCUDA 书稿 LaTeX 到 Read the Docs 站点的转换管线维护 skill(站点目录 LeetCUDA/docs/readthedocs/)。当任务涉及:改完书稿后让站点同步、改转换器 convert/、本地构建与预览 build.sh、内容核对 convert.verify、浏览器验收…

    12k GitHub stars~902 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Qt C++ Code Review

    x-tools-author/x-tools

    Read-only review of Qt6 C++ code that combines a deterministic lint script with six parallel analysis agents and reports only high-confidence issues.

    1.1k GitHub starsUsed in 2 repos~4.3k tokens
    DevelopmentAuto-check passed

More from RunanywhereAI/wally

  • Wally Device E2E

    RunanywhereAI/wally

    Run wally's LLM e2e on Apple Neural Engine (NeuRT) and Snapdragon Hexagon NPU (QHexRT) devices.

    1.6k GitHub stars~845 tokensUpdated today
    Auto-check passed
  • Wally Release

    RunanywhereAI/wally

    Cut an Wally product release (independent of SDK version) — version bump, release:patch label, merge, auto-tag, bottles.

    1.6k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Wally E2E

    RunanywhereAI/wally

    Verify a built wally binary against a pinned C++ desktop kit on macOS and Windows.

    1.6k GitHub stars~3.8k tokensUpdated today
    Auto-check passed
  • Wally Kit Pin

    RunanywhereAI/wally

    Bump cmake/sdk-pin.cmake to a new published SDK C++ desktop kit (version + SHA-256 + IDL lock).

    1.6k GitHub stars~1k tokensUpdated today
    Auto-check passed

Works with

Questions about Wally Architecture

What does Wally Architecture do?

Where Wally logic belongs — command layering, proto as SOT, kit vs CLI ownership, Apple MLX host vs wally-cxx. Wally Architecture is an agent skill from RunanywhereAI/wally. Where Wally logic belongs — command layering, proto as SOT, kit vs CLI ownership, Apple MLX host vs wally-cxx.

When should I use Wally Architecture?

Wally Architecture fits situations like: adding a command; moving inference logic; deciding whether a bug is SDK.

How do I install Wally Architecture in Claude Code?

Run `npx skills add RunanywhereAI/wally --skill wally-architecture -a claude-code`. Or copy the skill folder (.agents/skills/wally-architecture in RunanywhereAI/wally) into .claude/skills/wally-architecture in your project. Claude Code loads it when a task matches its description.

How do I install Wally Architecture in Codex?

Run `npx skills add RunanywhereAI/wally --skill wally-architecture -a codex`. Or copy the skill folder (.agents/skills/wally-architecture in RunanywhereAI/wally) into .agents/skills/wally-architecture in your project. Codex loads it when a task matches its description.

Can I use Wally Architecture in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add RunanywhereAI/wally --skill wally-architecture -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/wally-architecture, .gemini/skills/wally-architecture, .github/skills/wally-architecture and .opencode/skills/wally-architecture in your project.

What does Wally Architecture need to run?

Going by SKILL.md and its folder, Wally Architecture needs the command-line tools its instructions call (cmake and bash).

Does Wally Architecture access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Wally Architecture safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Wally Architecture use?

Wally Architecture is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Wally Architecture use?

About 992 tokens (SKILL.md is roughly 4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Wally Architecture?

Skills that share tags, products or a category with Wally Architecture: Paddle Build (PaddlePaddle/Paddle, 24k stars), Fory Release (apache/fory, 4.6k stars), ONNX Runtime Shape Inference Safety Audit (microsoft/onnxruntime, 22k stars) and Code Audit (3stoneBrother/code-audit, 892 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Wally Architecture?

RunanywhereAI (a GitHub organization) maintains it in RunanywhereAI/wally, which has 1,562 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 10, 2026.

Source: RunanywhereAI/wally on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.