Set up observability for Groq integrations: latency histograms, token throughput, rate limit gauges, cost tracking, and Prometheus alerts.

MITAuto-check passedDevOps & Cloud

Install Groq Observability

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill groq-observability -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace groq-observability --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/groq-observability .claude/skills/groq-observability && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
groq-observability
GitHub stars
2.8k
Token cost
~1.6k tokens
SKILL.md length
515 words
Files
3 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Set up observability for Groq integrations: latency histograms, token throughput, rate limit gauges, cost tracking, and Prometheus alerts.

  • Works in 6 steps: Instrumented client — wrap… → Prometheus metrics — register a… → Rate limit header tracking — parse… → …
  • Instrumenting Groq API calls
  • SKILL.md covers Overview, Prerequisites, Key Metrics to Track and Instructions, plus 4 more sections
  • Calls npm; needs GROQ_API_KEY

What it does

Groq Observability is an agent skill from jeremylongshore/tons-of-skills-marketplace. Set up observability for Groq integrations: latency histograms, token throughput, rate limit gauges, cost tracking, and Prometheus alerts. Use when instrumenting Groq API calls, building a metrics dashboard, or wiring latency/cost/rate-limit alerts. Trigger with phrases like "groq monitoring", "groq metrics", "groq observability", "monitor groq", "groq alerts", "groq dashboard".

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/examples.md` and `references/implementation.md`). Compatibility notes: Designed for Claude Code

It sits in DevOps & Cloud, covering Observability, Monitoring and alerting and Rate limiting. It works with Prometheus. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Instrumenting Groq API calls
  • Building a metrics dashboard
  • Wiring latency/cost/rate-limit alerts
  • With phrases like groq monitoring

Example prompts

  • “groq monitoring”
  • “groq metrics”
  • “groq observability”
  • “/groq-observability”

Requirements

  • Node.js
  • A credential in GROQ_API_KEY
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Instrumented client — wrap groq.chat.completions.create so latency, tokens, queue time, and estimated cost are captured on the same path…
  2. Prometheus metrics — register a histogram (latency), counters (tokens, cost, errors), and gauges (throughput, rate-limit remaining), then…
  3. Rate limit header tracking — parse x-ratelimit-remaining-* off every response into a gauge so you alert before a 429, not after.
  4. Prometheus alert rules — ship latency/rate-limit/throughput/error/cost alerts tuned to Groq's sub-200ms, 280+ tok/s baseline.
  5. Structured request logging — emit one JSON line per request for log aggregation, preserving per-request detail metrics roll up.
  6. Dashboard panels — TTFT distribution, tokens/sec, rate-limit utilization, request volume, error rate, cost, and queue time.

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • console.groq.com
    • npmjs.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GROQ_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Groq Observability loads about 1.6k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 100 tokens; SKILL.md has 515 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~100
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 515 words, ~1,573 tokens.

Download SKILL.mdSave it as .claude/skills/groq-observability/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
groq-observability
description
Set up observability for Groq integrations: latency histograms, token throughput, rate limit gauges, cost tracking, and Prometheus alerts. Use when instrumenting Groq API calls, building a metrics dashboard, or wiring latency/cost/rate-limit alerts. Trigger with phrases like "groq monitoring", "groq metrics", "groq observability", "monitor groq", "groq alerts", "groq dashboard".
allowed-tools
Read, Write, Edit
compatibility
Designed for Claude Code
version
1.11.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, groq, monitoring, observability, dashboard

Groq Observability

Overview

Monitor Groq LPU inference for latency, token throughput, rate limit utilization, and cost. Groq's defining advantage is speed (280-560 tok/s), so latency degradation is the highest-priority signal. The API returns rich timing metadata (queue_time, prompt_time, completion_time) and rate limit headers on every response.

Prerequisites

  • A Groq account with an API key exported as the GROQ_API_KEY environment variable — the groq-sdk client reads it automatically (new Groq()).
  • Node.js with groq-sdk and prom-client installed (npm install groq-sdk prom-client).
  • A Prometheus scrape target and (optionally) Grafana for the dashboard panels.

Key Metrics to Track

MetricTypeSourceWhy
TTFT (time to first token)HistogramClient-side timingGroq's main value prop
Tokens/secondGaugeusage.completion_timeThroughput degradation
Total latencyHistogramClient-side timingEnd-to-end performance
Rate limit remainingGaugex-ratelimit-remaining-* headersPrevent 429s
Token usageCounterusage.total_tokensCost attribution
Error rate by codeCounterError handlerAvailability
Estimated costCounterTokens * model priceBudget tracking

Instructions

Apply these six steps in order. Steps 1-2 are the core instrumentation loop — wrap the client, then feed a Prometheus instrument set from each call. Steps 3-6 add rate-limit tracking, alerting, structured logs, and dashboards on top. The lean client skeleton is below; the full code for every step lives in references/implementation.md.

  1. Instrumented client — wrap groq.chat.completions.create so latency, tokens, queue time, and estimated cost are captured on the same path as the request (trackedCompletion).
  2. Prometheus metrics — register a histogram (latency), counters (tokens, cost, errors), and gauges (throughput, rate-limit remaining), then feed them from emitMetrics.
  3. Rate limit header tracking — parse x-ratelimit-remaining-* off every response into a gauge so you alert before a 429, not after.
  4. Prometheus alert rules — ship latency/rate-limit/throughput/error/cost alerts tuned to Groq's sub-200ms, 280+ tok/s baseline.
  5. Structured request logging — emit one JSON line per request for log aggregation, preserving per-request detail metrics roll up.
  6. Dashboard panels — TTFT distribution, tokens/sec, rate-limit utilization, request volume, error rate, cost, and queue time.
typescript
import Groq from "groq-sdk";

const groq = new Groq(); // reads GROQ_API_KEY

async function trackedCompletion(model: string, messages: any[]) {
  const start = performance.now();
  const result = await groq.chat.completions.create({ model, messages });
  const latencyMs = performance.now() - start;
  const usage = result.usage!;
  const metrics = {
    model,
    latencyMs: Math.round(latencyMs),
    tokensPerSec: Math.round(usage.completion_tokens / ((usage as any).completion_time || latencyMs / 1000)),
    totalTokens: usage.total_tokens,
  };
  emitMetrics(metrics); // -> Prometheus (Step 2)
  return { result, metrics };
}

See references/implementation.md for the complete GroqMetrics shape, pricing table, Prometheus instruments, rate-limit tracking, alert rules, structured logging, and dashboard panel list.

Show full SKILL.md (176 more words)Show less

Output

Applying the workflow produces:

  • A trackedCompletion wrapper that returns { result, metrics }, where metrics is a GroqMetrics object (latency, TTFT, tokens/sec, token counts, queue time, estimated cost).
  • A Prometheus metric set — groq_latency_ms (histogram), groq_tokens_total / groq_cost_usd / groq_errors_total (counters), and groq_tokens_per_second / groq_ratelimit_remaining (gauges).
  • Five alert rules (GroqLatencyHigh, GroqRateLimitCritical, GroqThroughputDrop, GroqErrorRateHigh, GroqCostSpike).
  • A structured JSON log line per request and a 7-panel dashboard spec.

Examples

Instrument a single completion and emit a structured log line:

typescript
const { result, metrics } = await trackedCompletion(
  "llama-3.3-70b-versatile",
  [{ role: "user", content: "Summarize this incident report in two sentences." }]
);
logGroqRequest(metrics, result.id);
// metrics.tokensPerSec -> 310, metrics.estimatedCostUsd -> 0.000404

For a 429-guard using rate-limit headers and a dashboard health-reading table, see references/examples.md.

Error Handling

IssueCauseSolution
429 with high retry-afterRPM or TPM exhaustedImplement request queuing
Latency spike > 2sModel overloaded or large promptReduce prompt size or switch to lighter model
503 Service UnavailableGroq capacity issueEnable fallback to alternative provider
Tokens/sec dropStreaming disabled or large promptsEnable streaming for better perceived performance

Resources

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/.curated/groq-observability of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/examples.md
  • references/implementation.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Groq Observability next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Groq Observability compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Groq Observability this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.6kAutomated safety check: PassMIT
Cost Exportruvnet/ruflo74k—~687Automated safety check: NotesMIT
Monitoring Observabilityyonatangross/orchestkit292—~2.2kAutomated safety check: PassMIT
Happy Infra Metrics and Grafanaslopus/happy24k—~2kAutomated safety check: NotesMIT
WizTelemetry Platform Servicekubesphere/kubesphere17k—~1.8kAutomated safety check: PassCustom licence
Redis Observabilityredis/agent-skills1662 repos~911Automated safety check: PassMIT

Similar skills

  • Cost Export

    ruvnet/ruflo

    Export cost-tracking telemetry in Prometheus textfile or webhook JSON formats — for external observability (Grafana, Datadog, custom dashboards)

    74k GitHub stars~687 tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Monitoring Observability

    yonatangross/orchestkit

    Monitoring and observability patterns for Prometheus metrics, Grafana dashboards, Langfuse v4 LLM tracing (astype, scorecurrentspan, shouldexportspan, LangfuseMedia), and drift detection.

    292 GitHub stars~2.2k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API.

    24k GitHub stars~2k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • WizTelemetry Platform Service

    kubesphere/kubesphere

    Installs and configures the WizTelemetry Platform Service extension for KubeSphere, the shared API server behind its observability extensions.

    17k GitHub stars~1.8k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • Redis Observability

    redis/agent-skills

    Official

    Redis observability guidance — which metrics to monitor (memory, connections, hit ratio, ops/sec, rejected connections), which built-in commands to reach for during incident triage (SLOWLOG, INFO…

    166 GitHub starsUsed in 2 repos~911 tokens
    DevOps & CloudAuto-check passed
  • 当需要为 funboost 创建 Consumer 或 Publisher 的 Mixin 扩展类时使用。触发场景:添加监控、熔断、限流、链路追踪等横切关注点,编写自定义前置/后置处理钩子。关键词:mixin, consumeroverridecls, publisheroverridecls, ConsumerMixin, 自定义消费者, hook, 拦截器, 熔断器, 监控…

    895 GitHub stars~2.1k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Groq Observability

What does Groq Observability do?

Set up observability for Groq integrations: latency histograms, token throughput, rate limit gauges, cost tracking, and Prometheus alerts. Groq Observability is an agent skill from jeremylongshore/tons-of-skills-marketplace. Set up observability for Groq integrations: latency histograms, token throughput, rate limit gauges, cost tracking, and Prometheus alerts.

When should I use Groq Observability?

Groq Observability fits situations like: instrumenting Groq API calls; building a metrics dashboard; wiring latency/cost/rate-limit alerts; with phrases like groq monitoring.

How do I install Groq Observability in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill groq-observability -a claude-code`. Or copy the skill folder (skills/.curated/groq-observability in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/groq-observability in your project. Claude Code loads it when a task matches its description.

How do I install Groq Observability in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill groq-observability -a codex`. Or copy the skill folder (skills/.curated/groq-observability in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/groq-observability in your project. Codex loads it when a task matches its description.

Can I use Groq Observability in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill groq-observability -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/groq-observability, .gemini/skills/groq-observability, .github/skills/groq-observability and .opencode/skills/groq-observability in your project.

What does Groq Observability need to run?

Going by SKILL.md and its folder, Groq Observability needs the command-line tools its instructions call (npm) and credentials named GROQ_API_KEY. Our summary lists: Node.js; A credential in GROQ_API_KEY. Its frontmatter pre-approves these tools: Read, Write, Edit. Compatibility (from SKILL.md): Designed for Claude Code.

Does Groq Observability access the network?

SKILL.md names 2 domains. As links in the text: console.groq.com and npmjs.com. This is read from the text; nothing was executed.

Is Groq Observability safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Groq Observability use?

Groq Observability is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Groq Observability use?

About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.4k tokens, read only when the agent opens those files.

What are the alternatives to Groq Observability?

Skills that share tags, products or a category with Groq Observability: Cost Export (ruvnet/ruflo, 74k stars), Monitoring Observability (yonatangross/orchestkit, 292 stars), Happy Infra Metrics and Grafana (slopus/happy, 24k stars) and WizTelemetry Platform Service (kubesphere/kubesphere, 17k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Groq Observability?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.