Agent skill

Remote Compute Ops

by AnastasiyaW in AnastasiyaW/codex-claude-code-config

Operate GPU and remote compute across RunPod (Pods and Serverless), Massed Compute VMs, and owned or virtual remote servers through existing bridges, SSH sessions, MCP/API adapters, bounded polling…

MITAuto-check passedBackend & APIs

Install Remote Compute Ops

skills CLI
$ npx skills add AnastasiyaW/codex-claude-code-config --skill remote-compute-ops -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AnastasiyaW/codex-claude-code-config remote-compute-ops --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AnastasiyaW/codex-claude-code-config.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/operational/remote-compute-ops .claude/skills/remote-compute-ops && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
remote-compute-ops
GitHub stars
154
Token cost
~3k tokens
SKILL.md length
1,584 words
Files
5 (incl. references)
Skills in repo
50
Repo updated
First seen
Licence
MIT

At a glance

Operate GPU and remote compute across RunPod (Pods and Serverless), Massed Compute VMs, and owned or virtual remote servers through existing bridges, SSH sessions, MCP/API adapters, bounded polling…

  • Works in 7 steps: Freeze the target and read the… → Classify the provider mode before… → Reconcile read-only state through the… → …
  • The user mentions RunPod
  • SKILL.md covers Non-negotiable transport rule, Workflow, Gotchas and Troubleshooting
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Remote Compute Ops is an agent skill from AnastasiyaW/codex-claude-code-config. Operate GPU and remote compute across RunPod (Pods and Serverless), Massed Compute VMs, and owned or virtual remote servers through existing bridges, SSH sessions, MCP/API adapters, bounded polling, cost controls, and resumable lifecycle checks. Use when the user mentions RunPod, Massed Compute, a remote GPU/server/VM, SSH bridge/tunnel/bastion/Tailscale, training or inference on rented compute, GPU inventory, billing, or asks to minimize API/SSH connections and avoid rate limits. Do not use for generic cloud…

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `agents/openai.yaml`, `references/massed-compute-recipes.md` and `references/provider-matrix.md`).

It sits in Backend & APIs, covering Rate limiting, Budgeting and forecasting and Serverless. It works with Model Context Protocol. The repository describes itself as: Claude Code, Codex, and multi-agent configuration system: principles, hooks, skills, and workflow patterns for AI-assisted development. The licence is MIT.

When your agent uses it

  • The user mentions RunPod
  • A remote GPU/server/VM
  • SSH bridge/tunnel/bastion/Tailscale
  • Inference on rented compute

Example prompts

  • “/remote-compute-ops”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Freeze the target and read the architecture
  2. Classify the provider mode before choosing a channel
  3. Reconcile read-only state through the cheapest valid path
  4. Select the provider adapter
  5. Mutate only the named resource
  6. Reconcile instead of duplicating
  7. Close the loop

What it can do on your machine

Read from SKILL.md and the folder at commit 67709af. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Remote Compute Ops loads about 3k tokens when it runs, and up to ~6.3k if it reads all its reference files. Until then it costs about 156 tokens; SKILL.md has 1,584 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~156
When it runs · the whole SKILL.md, loaded when a task matches
~3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AnastasiyaW/codex-claude-code-config at commit 67709af, republished under its MIT licence (© AnastasiyaW). 1,584 words, ~2,964 tokens.

Download SKILL.mdSave it as .claude/skills/remote-compute-ops/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
remote-compute-ops
description
Operate GPU and remote compute across RunPod (Pods and Serverless), Massed Compute VMs, and owned or virtual remote servers through existing bridges, SSH sessions, MCP/API adapters, bounded polling, cost controls, and resumable lifecycle checks. Use when the user mentions RunPod, Massed Compute, a remote GPU/server/VM, SSH bridge/tunnel/bastion/Tailscale, training or inference on rented compute, GPU inventory, billing, or asks to minimize API/SSH connections and avoid rate limits. Do not use for generic cloud architecture, local-only GPU work, or application code with no remote-resource operation.

Remote compute operations

Use this skill as the provider-neutral workflow for remote GPU and server work. The provider is an adapter, not the name of the skill: RunPod, Massed Compute, and an owned/virtual server must all follow the same evidence, transport, lifecycle, and handoff rules.

When auditing this skill or designing a plan offline, do not call a provider, SSH, or bridge at all; state that live mode/identity is unproven. The read-only lookup below applies only when the user explicitly asks to inspect live remote state.

Non-negotiable transport rule

Reuse the already-created bridge or live connection before opening a new one. The purpose is reliable, rate-limit-compliant operation and avoiding unnecessary authentication attempts; never disguise traffic, evade a provider limit, rotate identities, or bypass a ban.

  • Inspect the local connection registry/bridge health when one exists. A registry is a coordination hint, not proof that a tunnel is alive. Resolve the existing helper through %USERPROFILE%\\.claude\\scripts\\conn_registry.py on Windows or $HOME/.claude/scripts/conn_registry.py on POSIX when that file exists; if it is not discoverable, report “registry unavailable” instead of guessing a path.
  • In a read-only investigation, do not create or re-register a bridge. If a live route is required, allow at most one explicitly authorized health probe. Record the host alias/route, registry entry age, session owner, local PID/service or control-socket metadata when available, target identity, last probe result, and whether a probe was permitted; never record credentials.
  • Use one persistent provider client/session per task phase. Group compatible read-only queries and reuse keep-alive connections; do not create a client or authenticate once per command.
  • The bridge probe has the stricter budget: local registry/config inspection is network-free, but SSH/tunnel health is at most one attempt total per target and phase, with no SSH retry after a timeout or connection error. The API-read retry budget in transport-safety.md does not apply to that probe.
  • Batch related remote shell checks into one SSH invocation. Use ControlMaster/ControlPersist only after the exact route has passed a health check. If multiplexing fails on the platform or bridge, do not retry it blindly: use one batched command over the known working bridge.
  • Do not fan out API or SSH calls merely to reduce wall-clock time. Parallelism is allowed only when the provider documents it, the connection budget allows it, and the calls cannot duplicate a mutation.
  • For 429, 503, connection resets, or transport timeouts, stop increasing the request rate. Honor Retry-After, use bounded exponential backoff with jitter, and record the retry budget. See transport-safety.md.

Workflow

1. Freeze the target and read the architecture

Before a remote mutation, read the repository AGENTS.md, provider runbook, and the relevant reference. Establish:

  • provider and mode (RunPod Pod, RunPod Serverless, Massed VM, or owned server);
  • exact target ID/name, region, image/template, job ID, and intended outcome;
  • traffic path: existing bridge, bastion, VPN/Tailscale, SSH host alias, proxy, or provider API endpoint;
  • current checkout, deployment/source revision, process/job state, storage and checkpoint path, and cost/burn boundary.

Do not infer a live state from a stale handoff, old dashboard, or a command that only proves that a process exists locally.

2. Classify the provider mode before choosing a channel

Use existing target metadata first. If the mode is unknown, make one read-only control-plane lookup and classify it before touching SSH:

  • RunPod Serverless: endpoint/job ID, /run//status//health, webhook, or stream. Do not try to SSH to a Serverless job.
  • RunPod Pod: exact Pod ID plus a documented SSH/TCP/HTTP connection route and exact bridge/host alias. Use SSH only when the target is an actual Pod and the route is already verified; never infer Pod identity from a generic job name.
  • Massed VM: instance UUID and the provider-returned SSH target, plus Massed MCP for account/instance state.
  • Owned/virtual server: documented host alias and existing bridge/tunnel.

If the provider or mode remains ambiguous after that one lookup, stop and report the missing identity instead of opening a second kind of connection.

3. Reconcile read-only state through the cheapest valid path

Prefer this order:

  1. existing bridge/session health and the shared connection registry;
  2. one batched remote probe for host, GPU, process, disk, and durable logs;
  3. one provider client session for exact inventory, billing, target, or job state;
  4. a second provider call only when the first result is incomplete or ambiguous.

For long jobs, prefer durable logs, checkpoints, job events, or a webhook over a tight status loop. If polling is the only supported observation path, use one job-specific timer with a minimum interval, a maximum deadline, a request budget, and terminal-state exit. Never poll every target independently from several agents.

4. Select the provider adapter

Read provider-matrix.md and then the provider's existing detailed skill/runbook when available.

  • RunPod: use the local runpod-gpu-ops skill if it is installed; otherwise use provider-matrix.md and the linked official RunPod docs for account-specific images, volumes, and lifecycle. Serverless is the default for scale-to-zero inference; a Pod is a persistent billed resource and needs an explicit reason plus a cleanup owner. Use the returned endpoint/job/pod ID as the identity for all later calls.
  • Massed Compute: use the Massed MCP tools and read massed-compute-recipes.md for provider-specific recipes. Prefer read-only tools first; destructive tools may be absent from a read-only key by design. Keep the MCP session and reconcile after any timeout before considering a retry.
  • Owned or virtual server: do not invent a cloud API. Reuse the verified SSH or tunnel route, batch probes, inspect the actual service/process/GPU/log state, and use the host's runbook for restart or shutdown decisions.
5. Mutate only the named resource

Launching or restarting affects cost and capacity. State the chosen target, image, quantity, region, expected hourly/per-job burn, checkpoint/storage path, and stop condition before executing within the user's request.

Termination, deletion, key removal, volume destruction, and any action that can lose unrecoverable work require exact target disclosure, explicit confirmation, the smallest possible scope, and post-action verification. A vague label such as "the idle pod" is not an exact target.

Show full SKILL.md (608 more words)Show less
6. Reconcile instead of duplicating

Every mutation must have an identity and a durable observation record. If a launch or restart times out, assume it may have succeeded: list/get by exact ID, name, client idempotency key, or a narrow creation-time filter before retrying. Do not send a second launch because the first response was lost.

For each state transition, record provider, target ID, bridge/session used, source revision, last observation timestamp, state, job/checkpoint marker, and next allowed action. Do not record tokens, passwords, private keys, or full response bodies containing credentials.

7. Close the loop

After launch/restart/deploy, verify the actual user-facing or job outcome, not only that a VM is listed as running:

  • SSH/bridge reaches the intended host;
  • GPU and process are the expected ones;
  • service/job health is ready and the first safe probe succeeds;
  • output/checkpoint/log marker advances;
  • cost and cleanup owner are known.

If a remote action is left running, write the handoff/journal entry and state the exact next observation. Do not leave a paid resource without a shutdown rule.

Gotchas

  • Provider rate limits are not interchangeable. RunPod publishes limits per endpoint and operation; Massed Compute documents a generic 429 recovery path. Always re-check the current provider page and response headers.
  • An SSH control socket can be unsupported or broken on a particular Windows or ProxyCommand route. The safe fallback is one batched connection, not a storm of short SSH calls and not an unverified direct route.
  • A “running” Pod can still be starting a service; a Serverless /health result is not the same as a completed job. Check the service/job marker and logs.
  • A timeout is an ambiguous mutation result. Reconcile by exact identity before retrying; never rely on a fresh list alone when multiple jobs have similar names.
  • A shared connection registry can contain stale entries. Heartbeat expiry narrows the candidates but cannot replace an external health probe.
  • API keys and VM passwords stay in the approved local secret store or provider UI. Never copy them into this skill, a handoff, Git, or a command transcript.
  • Do not terminate a GPU merely because it is idle for one observation. Compare the active task, owner, checkpoint/output progress, and declared stop condition.

Troubleshooting

  • Several agents keep opening SSH/API sessions -> inspect the shared connection registry and active bridge, nominate one owner for the connection, batch the remaining checks, and make other agents consume the durable log/heartbeat.
  • 429 or 503 -> stop fan-out, honor Retry-After, back off with jitter, reduce polling frequency, and retry only idempotent reads. For an ambiguous mutation, reconcile first.
  • ControlMaster reports a socket or getsockname error -> mark multiplexing unavailable for that route, use the verified bridge with one batched command, and preserve the local guard/reminder that prevents repeated calls.
  • RunPod shows a healthy endpoint but the job is stuck -> inspect queue/worker state and job status, then worker logs and the actual output marker. Do not create a diagnostic Pod by default.
  • Massed tools are missing -> check the MCP entry and token scope; a read-only key intentionally hides launch/restart/terminate/key-management tools. Do not compensate with ad-hoc REST calls unless the provider runbook explicitly allows it.
  • A launch command timed out -> list/get the exact target and check billing, inventory, and capacity before any retry. Treat the first request as possibly successful.
  • A bridge is recorded but unreachable -> do not re-register or reconnect in a read-only investigation. Report the route/alias, registry age, owner/session, local PID/service or control-socket metadata when available, last health result, target identity, and whether one authorized probe was allowed. Perform only the local checklist: registry heartbeat/TTL, bridge process/socket presence, SSH alias and ProxyCommand mapping, and bridge-owner confirmation. Never reclaim a live tunnel based only on a stale timestamp.

© AnastasiyaW, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in skills/operational/remote-compute-ops of AnastasiyaW/codex-claude-code-config.

  • SKILL.md
  • agents/openai.yaml
  • references/massed-compute-recipes.md
  • references/provider-matrix.md
  • references/transport-safety.md

Open the folder on GitHubat commit 67709af

Compare with similar skills

Remote Compute Ops next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Remote Compute Ops compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Remote Compute Ops this skillAnastasiyaW/codex-claude-code-config154—~3kAutomated safety check: PassMIT
Neon Functionsneondatabase/agent-skills100—~12kAutomated safety check: NotesApache-2.0
Azure Aigatewaymicrosoft/GitHub-Copilot-for-Azure2551 repos~1.3kAutomated safety check: PassMIT
AWS Serverless Edazxkane/aws-skills3674 repos~3.2kAutomated safety check: PassMIT
NubaseOtterMind/Nubase622—~2.2kAutomated safety check: NotesApache-2.0
AWS Solution Architectalirezarezvani/claude-skills28k1 repos~2.5kAutomated safety check: PassMIT

Similar skills

  • Neon Functions

    neondatabase/agent-skills

    Official

    Long-running, serverless Node.js HTTP functions deployed onto your Neon branch, with DATABASEURL injected automatically and compute that runs next to your data.

    100 GitHub stars~12k tokensUpdated yesterday
    Backend & APIsAuto-check: notes
  • Azure Aigateway

    microsoft/GitHub-Copilot-for-Azure

    Official

    Configure Azure API Management as an AI Gateway for AI models, MCP tools, and agents.

    255 GitHub starsUsed in 1 repo~1.3k tokens
    Backend & APIsAuto-check passed
  • AWS Serverless Eda

    zxkane/aws-skills

    AWS serverless and event-driven architecture expert based on Well-Architected Framework.

    367 GitHub starsUsed in 4 repos~3.2k tokens
    Backend & APIsAuto-check passed
  • Nubase

    OtterMind/Nubase

    A skill your agent uses when the user mentions Nubase broadly, wants a backend for an AI-generated app, or needs to deploy/publish generated code online — across Database, Auth, Storage, Assets…

    622 GitHub stars~2.2k tokensUpdated 12 days ago
    Backend & APIsAuto-check: notes
  • AWS Solution Architect

    alirezarezvani/claude-skills

    Design AWS architectures for startups using serverless patterns and IaC templates.

    28k GitHub starsUsed in 1 repo~2.5k tokens
    Backend & APIsAuto-check passed
  • Testing Livepeer

    daydreamlive/scope

    Test Scope locally in Livepeer mode end to end using a prebuilt go-livepeer artifact from the ja/serverless PR, uv run --extra livepeer livepeer-runner, and Scope.

    452 GitHub stars~2.5k tokensUpdated 2 mo ago
    Backend & APIsAuto-check passed

More from AnastasiyaW/codex-claude-code-config

All 50 skills in this repo
  • Bug Reproducer

    AnastasiyaW/codex-claude-code-config

    Find likely software bugs in a codebase, rank concrete bug candidates, and prove or reject them with focused regression tests before proposing a fix.

    154 GitHub stars~4.1k tokensUpdated today
    Auto-check passed
  • Motion Framer

    AnastasiyaW/codex-claude-code-config

    A skill your agent uses when implementing Motion or Framer Motion in React/JavaScript: interactive UI components, micro-interactions, gestures, layout or page transitions, and scroll-based animation.

    154 GitHub starsUsed in 1 repo~5.2k tokens
    Auto-check passed
  • Proof Verify

    AnastasiyaW/codex-claude-code-config

    Plan-based verification - freeze acceptance criteria before building, then verify after with an independent fresh-context agent (the builder must not verify their own work).

    154 GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Workflow Orchestration

    AnastasiyaW/codex-claude-code-config

    Написание и запуск Claude Code dynamic workflows (JS-оркестратор субагентов).

    154 GitHub stars~3.8k tokensUpdated today
    Auto-check passed
  • Notebooklm Grounded Research

    AnastasiyaW/codex-claude-code-config

    A skill your agent uses when: NotebookLM, notebooklm MCP, large documentation sets, courses, books, papers, or citation-backed research are mentioned.

    154 GitHub stars~2.4k tokensUpdated today
    Auto-check: warnings
  • Deepseek Provider Contract

    AnastasiyaW/codex-claude-code-config

    Validate a proposed DeepSeek API integration before any key or project context is sent: check thinking-mode tool-call history, strict-schema assumptions, bounded output, and provider data boundaries.

    154 GitHub stars~1.2k tokensUpdated today
    Auto-check passed

Categories

Questions about Remote Compute Ops

What does Remote Compute Ops do?

Operate GPU and remote compute across RunPod (Pods and Serverless), Massed Compute VMs, and owned or virtual remote servers through existing bridges, SSH sessions, MCP/API adapters, bounded polling…. Remote Compute Ops is an agent skill from AnastasiyaW/codex-claude-code-config. Operate GPU and remote compute across RunPod (Pods and Serverless), Massed Compute VMs, and owned or virtual remote servers through existing bridges, SSH sessions, MCP/API adapters, bounded polling, cost controls, and resumable lifecycle checks.

When should I use Remote Compute Ops?

Remote Compute Ops fits situations like: the user mentions RunPod; A remote GPU/server/VM; SSH bridge/tunnel/bastion/Tailscale; inference on rented compute.

How do I install Remote Compute Ops in Claude Code?

Run `npx skills add AnastasiyaW/codex-claude-code-config --skill remote-compute-ops -a claude-code`. Or copy the skill folder (skills/operational/remote-compute-ops in AnastasiyaW/codex-claude-code-config) into .claude/skills/remote-compute-ops in your project. Claude Code loads it when a task matches its description.

How do I install Remote Compute Ops in Codex?

Run `npx skills add AnastasiyaW/codex-claude-code-config --skill remote-compute-ops -a codex`. Or copy the skill folder (skills/operational/remote-compute-ops in AnastasiyaW/codex-claude-code-config) into .agents/skills/remote-compute-ops in your project. Codex loads it when a task matches its description.

Can I use Remote Compute Ops in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AnastasiyaW/codex-claude-code-config --skill remote-compute-ops -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/remote-compute-ops, .gemini/skills/remote-compute-ops, .github/skills/remote-compute-ops and .opencode/skills/remote-compute-ops in your project.

What does Remote Compute Ops need to run?

SKILL.md names no scripts, command-line tools or credentials: Remote Compute Ops is instructions for the agent only.

Does Remote Compute Ops access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Remote Compute Ops safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Remote Compute Ops use?

Remote Compute Ops is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Remote Compute Ops use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.3k tokens, read only when the agent opens those files.

What are the alternatives to Remote Compute Ops?

Skills that share tags, products or a category with Remote Compute Ops: Neon Functions (neondatabase/agent-skills, 100 stars), Azure Aigateway (microsoft/GitHub-Copilot-for-Azure, 255 stars), AWS Serverless Eda (zxkane/aws-skills, 367 stars) and Nubase (OtterMind/Nubase, 622 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Remote Compute Ops?

AnastasiyaW (a GitHub user) maintains it in AnastasiyaW/codex-claude-code-config, which has 154 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on October 9, 2026.

Source: AnastasiyaW/codex-claude-code-config on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.