Diagnose runtime infrastructure issues — cold starts, timeouts, scaling problems, network failures.

MITAuto-check: notesDevOps & Cloud

Install Forge Diagnose

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill forge-diagnose -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace forge-diagnose --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ai-agency/tonone/skills/forge-diagnose .claude/skills/forge-diagnose && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
forge-diagnose
GitHub stars
2.8k
Token cost
~1.1k tokens
SKILL.md length
340 words
Files
2
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Diagnose runtime infrastructure issues — cold starts, timeouts, scaling problems, network failures.

  • Works in 5 steps: Detect Environment → Identify the Symptom → Gather Diagnostic Data → …
  • Asked about infra is slow
  • SKILL.md covers Steps and Delivery
  • Calls gcloud, aws and kubectl

What it does

Forge Diagnose is an agent skill from jeremylongshore/tons-of-skills-marketplace. Diagnose runtime infrastructure issues — cold starts, timeouts, scaling problems, network failures. Use when asked about "infra is slow", "cold starts", "network issues", "why is this timing out", "scaling problem", "latency spikes", or "service is down".

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `.claude-plugin/plugin.json`).

It sits in DevOps & Cloud. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Asked about infra is slow
  • Why is this timing out
  • Scaling problem
  • Service is down

Example prompts

  • “infra is slow”
  • “cold starts”
  • “network issues”
  • “/forge-diagnose”

Requirements

  • Docker
  • Pre-approved tools (allowed-tools): Read, Bash, Glob, Grep, WebFetch, WebSearch, AskUserQuestion

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Detect Environment
  2. Identify the Symptom
  3. Gather Diagnostic Data
  4. Analyze and Diagnose
  5. Propose Fix

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash
    • Glob
    • Grep
    • WebFetch
    • WebSearch
    • AskUserQuestion

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gcloud
    • aws
    • kubectl
    • fly
    • wrangler

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gcloud, aws, kubectl and wrangler, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Forge Diagnose loads about 1.1k tokens when it runs. Until then it costs about 68 tokens; SKILL.md has 340 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~68
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash, Glob, Grep, WebFetch, WebSearch, AskUserQuestion

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 340 words, ~1,057 tokens.

Download SKILL.mdSave it as .claude/skills/forge-diagnose/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
forge-diagnose
description
Diagnose runtime infrastructure issues — cold starts, timeouts, scaling problems, network failures. Use when asked about "infra is slow", "cold starts", "network issues", "why is this timing out", "scaling problem", "latency spikes", or "service is down".
allowed-tools
Read, Bash, Glob, Grep, WebFetch, WebSearch, AskUserQuestion
version
0.6.4
author
tonone-ai <hello@tonone.ai>
license
MIT

Diagnose Runtime Infrastructure Issues

You are Forge — the infrastructure engineer on the Engineering Team.

Follow the output format defined in docs/output-kit.md — 40-line CLI max, box-drawing skeleton, unified severity indicators, compressed prose.

Steps

Step 0: Detect Environment

Scan the project to determine the platform and available diagnostic tools:

bash
# Check for cloud CLI configs
gcloud config get-value project 2>/dev/null
aws sts get-caller-identity 2>/dev/null
cat wrangler.toml 2>/dev/null
cat fly.toml 2>/dev/null

# Check for IaC to understand the architecture
find . -name '*.tf' -not -path './.terraform/*' 2>/dev/null
ls docker-compose.yml fly.toml wrangler.toml vercel.json render.yaml 2>/dev/null

# Check available CLI tools
which gcloud aws flyctl wrangler kubectl docker 2>/dev/null
Step 1: Identify the Symptom

Classify what the user is experiencing:

  • Latency — slow responses, high p99
  • Cold starts — first request after idle is slow
  • Timeouts — requests failing after N seconds
  • Scaling — can't handle load, 429s or 503s
  • Network — connection refused, DNS failures, TLS errors
  • Resource exhaustion — OOM kills, CPU throttling, disk full
  • Intermittent failures — works sometimes, fails sometimes
Step 2: Gather Diagnostic Data

Based on the symptom, run targeted diagnostics:

For GCP/Cloud Run:

bash
gcloud run services describe SERVICE --region REGION --format yaml
gcloud run revisions list --service SERVICE --region REGION
gcloud logging read "resource.type=cloud_run_revision AND resource.labels.service_name=SERVICE" --limit 50 --format json

For AWS/ECS:

bash
aws ecs describe-services --cluster CLUSTER --services SERVICE
aws logs get-log-events --log-group-name LOG_GROUP --limit 50
aws cloudwatch get-metric-statistics --namespace AWS/ECS --metric-name CPUUtilization --period 300 --statistics Average --start-time START --end-time END

For Fly.io:

bash
fly status -a APP
fly logs -a APP --limit 50
fly scale show -a APP

For Cloudflare Workers:

bash
wrangler tail --format json 2>/dev/null

For Kubernetes:

bash
kubectl get pods -l app=APP
kubectl describe pod POD
kubectl top pods -l app=APP
kubectl logs -l app=APP --tail=50

Read all IaC files to understand the intended configuration vs what's actually running.

Step 3: Analyze and Diagnose

Check for common root causes:

  • Undersized instances — CPU/memory too low for the workload
  • Cold start patterns — min instances set to 0, no keep-warm strategy
  • Network misconfiguration — wrong VPC connector, missing firewall rules, DNS propagation
  • Scaling limits — max instances too low, concurrency too high per instance
  • Resource contention — noisy neighbors, shared database connections, connection pool exhaustion
  • Timeout mismatches — load balancer timeout < app startup time, or request timeout < downstream call
  • Missing health checks — traffic routed to unhealthy instances
  • Disk/memory leaks — gradual degradation over time
Step 4: Propose Fix

For each identified issue:

  1. What's wrong — specific misconfiguration or bottleneck
  2. Why it causes the symptom — the causal chain
  3. The fix — exact config change, IaC update, or CLI command
  4. Verification — how to confirm the fix worked

Implement the fix in IaC if possible. If it requires a CLI command (e.g., emergency scaling), provide it but also update the IaC so it doesn't drift back.

Delivery

If output exceeds the 40-line CLI budget, invoke /atlas-report with the full findings. The HTML report is the output. CLI is the receipt — box header, one-line verdict, top 3 findings, and the report path. Never dump analysis to CLI.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in plugins/ai-agency/tonone/skills/forge-diagnose of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • .claude-plugin/plugin.json

Open the folder on GitHubat commit cfae287

Compare with similar skills

Forge Diagnose next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Forge Diagnose compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Forge Diagnose this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.1kAutomated safety check: NotesMIT
Monitor CInrwl/nx29k6 repos~4.7kAutomated safety check: PassMIT
Terraform and OpenTofu Guideagentscope-ai/QwenPaw36k6 repos~4.2kAutomated safety check: PassApache-2.0
Vercel Optimize Auditvercel-labs/agent-skills32k8 repos~4.3kAutomated safety check: PassNone
Analyze GitHub Action Logswithastro/astro63k1 repos~1.3kAutomated safety check: PassCustom licence
Openclaw Live Updateropenclaw/openclaw392k—~3.7kAutomated safety check: PassMIT

Similar skills

  • Monitor CI

    nrwl/nx

    Monitor Nx Cloud CI pipeline and handle self-healing fixes. An agent skill from nrwl/nx.

    29k GitHub starsUsed in 6 repos~4.7k tokens
    DevOps & CloudAuto-check passed
  • Terraform and OpenTofu Guide

    agentscope-ai/QwenPaw

    Guidance for writing and testing Terraform and OpenTofu code: module structure, naming, test approaches, CI/CD workflows, state handling and security scanning.

    36k GitHub starsUsed in 6 repos~4.2k tokens
    DevOps & CloudAuto-check passed
  • Vercel Optimize Audit

    vercel-labs/agent-skills

    Official

    Runs a metrics-first audit of a deployed Vercel project, gating investigations on real signals to produce ranked, citation-backed cost and performance recommendations.

    32k GitHub starsUsed in 8 repos~4.3k tokens
    DevOps & CloudAuto-check passed
  • Official

    Analyze recent GitHub Actions workflow runs to identify patterns, mistakes, and improvements.

    63k GitHub starsUsed in 1 repo~1.3k tokens
    DevOps & CloudAuto-check passed
  • Openclaw Live Updater

    openclaw/openclaw

    Maintain the canonical live OpenClaw main checkout, macOS LaunchAgent-managed Gateway, local macOS app, exact-head main CI, and recurring full release validation.

    392k GitHub stars~3.7k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Docs Learn PR Preview

    netdata/netdata

    Use only when the user explicitly asks to build, run, preview, inspect, or validate learn.netdata.cloud locally using the contents of a PR or documentation branch before merge.

    81k GitHub stars~2k tokensUpdated today
    DevOps & CloudAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Categories

Questions about Forge Diagnose

What does Forge Diagnose do?

Diagnose runtime infrastructure issues — cold starts, timeouts, scaling problems, network failures. Forge Diagnose is an agent skill from jeremylongshore/tons-of-skills-marketplace. Diagnose runtime infrastructure issues — cold starts, timeouts, scaling problems, network failures.

When should I use Forge Diagnose?

Forge Diagnose fits situations like: asked about infra is slow; why is this timing out; scaling problem; service is down.

How do I install Forge Diagnose in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill forge-diagnose -a claude-code`. Or copy the skill folder (plugins/ai-agency/tonone/skills/forge-diagnose in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/forge-diagnose in your project. Claude Code loads it when a task matches its description.

How do I install Forge Diagnose in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill forge-diagnose -a codex`. Or copy the skill folder (plugins/ai-agency/tonone/skills/forge-diagnose in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/forge-diagnose in your project. Codex loads it when a task matches its description.

Can I use Forge Diagnose in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill forge-diagnose -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/forge-diagnose, .gemini/skills/forge-diagnose, .github/skills/forge-diagnose and .opencode/skills/forge-diagnose in your project.

What does Forge Diagnose need to run?

Going by SKILL.md and its folder, Forge Diagnose needs the command-line tools its instructions call (gcloud, aws, kubectl, fly and wrangler). Our summary lists: Docker. Its frontmatter pre-approves these tools: Read, Bash, Glob, Grep, WebFetch, WebSearch, AskUserQuestion.

Does Forge Diagnose access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Forge Diagnose safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Forge Diagnose use?

Forge Diagnose is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Forge Diagnose use?

About 1.1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Forge Diagnose?

Skills that share tags, products or a category with Forge Diagnose: Monitor CI (nrwl/nx, 29k stars), Terraform and OpenTofu Guide (agentscope-ai/QwenPaw, 36k stars), Vercel Optimize Audit (vercel-labs/agent-skills, 32k stars) and Analyze GitHub Action Logs (withastro/astro, 63k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Forge Diagnose?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.