Agent skill

Groq Incident Runbook

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Execute Groq incident response: triage, mitigation, fallback, and postmortem.

MITAuto-check passedDevOps & Cloud

Install Groq Incident Runbook

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill groq-incident-runbook -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace groq-incident-runbook --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/groq-incident-runbook .claude/skills/groq-incident-runbook && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
groq-incident-runbook
GitHub stars
2.8k
Token cost
~1.3k tokens
SKILL.md length
511 words
Files
4 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Execute Groq incident response: triage, mitigation, fallback, and postmortem.

  • Works in 5 steps: Classify severity. Match user impact to… → Triage. Run the Quick Triage script… → Decide. Walk the decision tree in → …
  • Responding to Groq-related outages
  • SKILL.md covers Overview, Prerequisites, Instructions and Output, plus 4 more sections
  • Calls curl; reaches api.groq.com; needs GROQ_API_KEY

What it does

Groq Incident Runbook is an agent skill from jeremylongshore/tons-of-skills-marketplace. Execute Groq incident response: triage, mitigation, fallback, and postmortem. Use when responding to Groq-related outages, investigating errors, or running post-incident reviews for Groq integration failures. Trigger with phrases like "groq incident", "groq outage", "groq down", "groq on-call", "groq emergency", "groq broken".

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/communication-and-postmortem.md`, `references/mitigations.md` and `references/triage-and-diagnostics.md`). Compatibility notes: Designed for Claude Code

It sits in DevOps & Cloud, covering Incident response and Runbooks and postmortems. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Responding to Groq-related outages
  • Investigating errors
  • Running post-incident reviews for Groq integration failures
  • With phrases like groq incident

Example prompts

  • “groq incident”
  • “groq outage”
  • “groq down”
  • “/groq-incident-runbook”

Requirements

  • A credential in GROQ_API_KEY
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Bash(kubectl:*), Bash(curl:*)

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Classify severity. Match user impact to the P1–P4 table in
  2. Triage. Run the Quick Triage script (status reachability, auth, per-model availability, rate-limit headers). The one-line probe that…
  3. Decide. Walk the decision tree in
  4. Mitigate. Apply the matching fix from mitigations.md: fallback-model routing for 5xx on one model, wait-or-reroute for 429, key rotation…
  5. Communicate & close. Post the internal alert and status-page update, then after resolution collect evidence and write the postmortem — all…

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash(kubectl:*)
    • Bash(curl:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.groq.com

    Also links to:

    • console.groq.com
    • status.groq.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GROQ_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Groq Incident Runbook loads about 1.3k tokens when it runs, and up to ~3k if it reads all its reference files. Until then it costs about 88 tokens; SKILL.md has 511 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~88
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 511 words, ~1,306 tokens.

Download SKILL.mdSave it as .claude/skills/groq-incident-runbook/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
groq-incident-runbook
description
Execute Groq incident response: triage, mitigation, fallback, and postmortem. Use when responding to Groq-related outages, investigating errors, or running post-incident reviews for Groq integration failures. Trigger with phrases like "groq incident", "groq outage", "groq down", "groq on-call", "groq emergency", "groq broken".
allowed-tools
Read, Bash(kubectl:*), Bash(curl:*)
compatibility
Designed for Claude Code
version
1.11.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, groq, incident-response

Groq Incident Runbook

Overview

Rapid incident response procedures for Groq API failures. Groq is a third-party inference provider -- when it goes down, your mitigation options are: wait, fall back to a different model, or fall back to a different provider.

This SKILL.md is the high-level flow. Deep, copy-paste-ready material lives in references/:

Prerequisites

  • GROQ_API_KEY exported in the environment you run the triage commands from.
  • curl for API probes; kubectl only if you collect logs from a Kubernetes deployment.
  • Access to console.groq.com to rotate keys or upgrade the plan.
  • A configured fallback provider (e.g. OpenAI) if you need to fail away from Groq entirely.

Authentication: every Groq API call in this runbook authenticates with a bearer token — Authorization: Bearer $GROQ_API_KEY. Keep the key in a secret manager, never inline; the evidence-collection step in communication-and-postmortem.md redacts gsk_ tokens from logs before archiving.

Instructions

Work the incident in five phases. Each phase points to the reference file with the exact commands.

  1. Classify severity. Match user impact to the P1–P4 table in triage-and-diagnostics.md — this sets your response-time budget (P1 < 15 min, P4 next business day).

  2. Triage. Run the Quick Triage script (status reachability, auth, per-model availability, rate-limit headers). The one-line probe that starts most incidents:

    bash
    curl -s -o /dev/null -w "%{http_code}\n" \
      https://api.groq.com/openai/v1/models \
      -H "Authorization: Bearer $GROQ_API_KEY"
  3. Decide. Walk the decision tree in triage-and-diagnostics.md to turn the HTTP code (timeout / 401 / 429 / 5xx / slow) into an action path.

  4. Mitigate. Apply the matching fix from mitigations.md: fallback-model routing for 5xx on one model, wait-or-reroute for 429, key rotation for 401, enable the fallback provider for a Groq-wide outage.

  5. Communicate & close. Post the internal alert and status-page update, then after resolution collect evidence and write the postmortem — all in communication-and-postmortem.md.

Show full SKILL.md (205 more words)Show less

Output

Running this runbook produces:

  • A triage verdict — the HTTP status per model and whether the fault is Groq-side or ours.
  • An applied mitigation — traffic routed to a healthy model or provider, or a rotated key.
  • A communication trail — internal alert + external status-page message.
  • An evidence bundle — groq-incident-TIMESTAMP.tar.gz containing models.json and redacted app-logs.txt.
  • A postmortem document — timeline, root cause, and dated action items.

Error Handling

IssueCauseSolution
Can't reach status.groq.comNetwork issueUse mobile or different network
All models failingGroq-wide outageEnable fallback provider (OpenAI, etc.)
Key rotation failsNo admin accessEscalate to team lead with console access
Fallback provider also downMulti-provider outageDegrade gracefully, show cached content

Examples

Example — 429 on the primary model. Triage shows llama-3.3-70b-versatile: HTTP 429 while llama-3.1-8b-instant: HTTP 200. The decision tree routes "one model 429 → route to a different model," so you switch traffic to the 8B model per mitigations.md, post a P3 internal alert, and file an action item to add fallback routing. The fallback-routing function lives in mitigations.md; the alert and postmortem templates are in communication-and-postmortem.md.

Resources

Next Steps

For data-handling and compliance procedures after an incident, see the groq-data-handling skill in this pack.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/.curated/groq-incident-runbook of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/communication-and-postmortem.md
  • references/mitigations.md
  • references/triage-and-diagnostics.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Groq Incident Runbook next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Groq Incident Runbook compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Groq Incident Runbook this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.3kAutomated safety check: PassMIT
Oncallpigweed-project/pigweed548—~963Automated safety check: PassApache-2.0
Activation Governance Chaos RolloutAli-Marandi/DataSense107—~1.9kAutomated safety check: PassMIT
Incident Response686f6c61/alfred-dev117—~1.1kAutomated safety check: PassMIT
Superset Incident Triagesuperset-sh/superset15k—~1kAutomated safety check: PassCustom licence
Post-Incident DebriefVeryGoodOpenSource/vgv-wingspan109—~1.9kAutomated safety check: PassMIT

Similar skills

  • Oncall

    pigweed-project/pigweed

    Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).

    548 GitHub stars~963 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Design, validate, and govern fail-closed customer-activation automations that use an Outbox/worker pattern.

    107 GitHub stars~1.9k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Incident Response

    686f6c61/alfred-dev

    Protocolo de respuesta ante incidentes en produccion: triaje, mitigacion, causa raiz y postmortem.

    117 GitHub stars~1.1k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Superset Incident Triage

    superset-sh/superset

    Does a read-only first pass on a possible production incident: gathers deploy, Sentry and health-check signals, proposes a severity and status message, then stops for human approval.

    15k GitHub stars~1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Post-Incident Debrief

    VeryGoodOpenSource/vgv-wingspan

    Produces a blameless post-incident debrief with timeline, root cause and follow-up actions after an outage, failed release or significant bug, while details are fresh.

    109 GitHub stars~1.9k tokensUpdated 4 days ago
    DevOps & CloudAuto-check passed
  • SRE Engineer

    Jeffallan/claude-skills

    Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.

    12k GitHub stars~1.7k tokensUpdated 7 days ago
    DevOps & CloudAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Categories

Questions about Groq Incident Runbook

What does Groq Incident Runbook do?

Execute Groq incident response: triage, mitigation, fallback, and postmortem. Groq Incident Runbook is an agent skill from jeremylongshore/tons-of-skills-marketplace. Execute Groq incident response: triage, mitigation, fallback, and postmortem.

When should I use Groq Incident Runbook?

Groq Incident Runbook fits situations like: responding to Groq-related outages; investigating errors; running post-incident reviews for Groq integration failures; with phrases like groq incident.

How do I install Groq Incident Runbook in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill groq-incident-runbook -a claude-code`. Or copy the skill folder (skills/.curated/groq-incident-runbook in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/groq-incident-runbook in your project. Claude Code loads it when a task matches its description.

How do I install Groq Incident Runbook in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill groq-incident-runbook -a codex`. Or copy the skill folder (skills/.curated/groq-incident-runbook in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/groq-incident-runbook in your project. Codex loads it when a task matches its description.

Can I use Groq Incident Runbook in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill groq-incident-runbook -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/groq-incident-runbook, .gemini/skills/groq-incident-runbook, .github/skills/groq-incident-runbook and .opencode/skills/groq-incident-runbook in your project.

What does Groq Incident Runbook need to run?

Going by SKILL.md and its folder, Groq Incident Runbook needs the command-line tools its instructions call (curl) and credentials named GROQ_API_KEY. Our summary lists: A credential in GROQ_API_KEY. Its frontmatter pre-approves these tools: Read, Bash(kubectl:*), Bash(curl:*). Compatibility (from SKILL.md): Designed for Claude Code.

Does Groq Incident Runbook access the network?

SKILL.md names 3 domains. In commands or code: api.groq.com; the agent is likely to contact it when it follows the instructions. As links in the text: console.groq.com and status.groq.com. This is read from the text; nothing was executed.

Is Groq Incident Runbook safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Groq Incident Runbook use?

Groq Incident Runbook is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Groq Incident Runbook use?

About 1.3k tokens (SKILL.md is roughly 5.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.7k tokens, read only when the agent opens those files.

What are the alternatives to Groq Incident Runbook?

Skills that share tags, products or a category with Groq Incident Runbook: Oncall (pigweed-project/pigweed, 548 stars), Activation Governance Chaos Rollout (Ali-Marandi/DataSense, 107 stars), Incident Response (686f6c61/alfred-dev, 117 stars) and Superset Incident Triage (superset-sh/superset, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Groq Incident Runbook?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.