Agent skill

Groq Performance Tuning

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Optimize Groq API performance with model selection, caching, streaming, and parallel requests.

MITAuto-check passedBackend & APIs

Install Groq Performance Tuning

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill groq-performance-tuning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace groq-performance-tuning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/groq-performance-tuning .claude/skills/groq-performance-tuning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
groq-performance-tuning
GitHub stars
2.8k
Token cost
~1.9k tokens
SKILL.md length
711 words
Files
3 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Optimize Groq API performance with model selection, caching, streaming, and parallel requests.

  • Works in 6 steps: Choose the right model for speed. Map… → Minimize token count. Trim verbose… → Stream for perceived performance. For… → …
  • Experiencing slow responses
  • SKILL.md covers Overview, Prerequisites, Groq Speed Benchmarks and Instructions, plus 5 more sections
  • Calls npm; needs GROQ_API_KEY

What it does

Groq Performance Tuning is an agent skill from jeremylongshore/tons-of-skills-marketplace. Optimize Groq API performance with model selection, caching, streaming, and parallel requests. Use when experiencing slow responses, implementing caching strategies, or optimizing request throughput for Groq integrations. Trigger with phrases like "groq performance", "optimize groq", "groq latency", "groq caching", "groq slow", "groq speed".

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/examples.md` and `references/implementation.md`). Compatibility notes: Designed for Claude Code

It sits in Backend & APIs, covering Caching. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Experiencing slow responses
  • Implementing caching strategies
  • Optimizing request throughput for Groq integrations
  • With phrases like groq performance

Example prompts

  • “groq performance”
  • “optimize groq”
  • “groq latency”
  • “/groq-performance-tuning”

Requirements

  • Node.js
  • A credential in GROQ_API_KEY
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Choose the right model for speed. Map each call site to a speed tier: llama-3.1-8b-instant for latency-critical paths…
  2. Minimize token count. Trim verbose system prompts to their essence and set max_tokens to the expected output size, not a safe-looking…
  3. Stream for perceived performance. For any output the user watches arrive, stream chunks and surface live TTFT / tokens-per-second metrics…
  4. Cache deterministic responses. Hash {messages, model} and serve repeat temperature: 0 requests from an LRU cache with a short TTL…
  5. Parallelize under a rate-limit-aware queue. Fan out bulk work with p-queue, capping concurrency and per-minute volume so you saturate…
  6. Benchmark before you commit. Measure the candidate models against your real prompt shape and pick the fastest that clears your quality bar.

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • console.groq.com
    • npmjs.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GROQ_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Groq Performance Tuning loads about 1.9k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 92 tokens; SKILL.md has 711 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~92
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 711 words, ~1,950 tokens.

Download SKILL.mdSave it as .claude/skills/groq-performance-tuning/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
groq-performance-tuning
description
Optimize Groq API performance with model selection, caching, streaming, and parallel requests. Use when experiencing slow responses, implementing caching strategies, or optimizing request throughput for Groq integrations. Trigger with phrases like "groq performance", "optimize groq", "groq latency", "groq caching", "groq slow", "groq speed".
allowed-tools
Read, Write, Edit
compatibility
Designed for Claude Code
version
1.11.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, groq, api, performance

Groq Performance Tuning

Overview

Maximize Groq's LPU inference speed advantage. Groq already delivers extreme throughput (280-560 tok/s) and low latency (<200ms TTFT), but client-side optimization -- model selection, prompt size, streaming, caching, and parallelism -- determines whether your application fully exploits that speed.

This skill walks through six tuning levers at a high level; the complete, copy-pasteable code for each lives in references/implementation.md, and end-to-end worked scenarios live in references/examples.md.

Prerequisites

  • Groq API key — set GROQ_API_KEY in the environment. The groq-sdk client (new Groq()) reads it automatically; never hardcode the key.
  • Node.js 18+ with the groq-sdk package installed (npm install groq-sdk).
  • Optional packages for the caching and parallelism steps: lru-cache and p-queue (npm install lru-cache p-queue).
  • A baseline latency measurement of your current integration so you can confirm the tuning actually helps.

Groq Speed Benchmarks

ModelTTFTThroughputContext
llama-3.1-8b-instant~50ms~560 tok/s128K
llama-3.3-70b-versatile~150ms~280 tok/s128K
llama-3.3-70b-specdec~100ms~400 tok/s128K
meta-llama/llama-4-scout-17b-16e-instruct~80ms~460 tok/s128K

TTFT = Time to First Token. Actual values depend on prompt size and server load.

Instructions

Apply these six levers in order. Each is a small, independent change — start with the ones that match your bottleneck (model choice and caching give the biggest wins on most workloads). The full code for every step is in references/implementation.md.

  1. Choose the right model for speed. Map each call site to a speed tier: llama-3.1-8b-instant for latency-critical paths, llama-3.3-70b-versatile for quality-sensitive paths, llama-3.3-70b-specdec for 70b quality at higher throughput. Set temperature: 0 so responses are deterministic (and cacheable).
  2. Minimize token count. Trim verbose system prompts to their essence and set max_tokens to the expected output size, not a safe-looking ceiling. Fewer tokens means faster responses and less TPM-quota pressure.
  3. Stream for perceived performance. For any output the user watches arrive, stream chunks and surface live TTFT / tokens-per-second metrics. Streaming hides TTFT even when total wall-clock is unchanged.
  4. Cache deterministic responses. Hash {messages, model} and serve repeat temperature: 0 requests from an LRU cache with a short TTL — turning a repeated call into a ~0ms hit.
  5. Parallelize under a rate-limit-aware queue. Fan out bulk work with p-queue, capping concurrency and per-minute volume so you saturate throughput without tripping 429s.
  6. Benchmark before you commit. Measure the candidate models against your real prompt shape and pick the fastest that clears your quality bar.

The essential skeleton — a tiered client every other step builds on:

typescript
import Groq from "groq-sdk";

const groq = new Groq();  // reads GROQ_API_KEY from the environment

const SPEED_MAP = {
  instant: "llama-3.1-8b-instant",      // <100ms TTFT — latency-critical
  balanced: "llama-3.3-70b-versatile",  // <200ms TTFT — quality-sensitive
  fast70b: "llama-3.3-70b-specdec",     // 70b quality, faster throughput
} as const;

async function tieredCompletion(prompt: string, tier: keyof typeof SPEED_MAP = "instant") {
  return groq.chat.completions.create({
    model: SPEED_MAP[tier],
    messages: [{ role: "user", content: prompt }],
    temperature: 0,   // deterministic = cacheable
    max_tokens: 256,  // request only what you need
  });
}

See references/implementation.md for the streaming, caching, parallel-queue, and benchmarking functions in full.

Show full SKILL.md (300 more words)Show less

Output

Applying these levers to a Groq integration produces:

  • A tiered model map (SPEED_MAP) so each call site uses the fastest model that meets its quality bar.
  • A streaming helper that returns { content, ttftMs, totalMs, tokPerSec } for live latency instrumentation.
  • A deterministic prompt cache (LRU + SHA-256 key) that collapses repeated requests to ~0ms.
  • A rate-limit-aware parallel executor that maximizes throughput without hitting 429s.
  • A benchmark report printing average latency and tokens/sec per model, e.g.:
text
llama-3.1-8b-instant     |  61ms avg | 548 tok/s avg
llama-3.3-70b-versatile  | 148ms avg | 279 tok/s avg
llama-3.3-70b-specdec    | 103ms avg | 401 tok/s avg

Performance Decision Matrix

ScenarioModelmax_tokensstreamcache
Classification8b-instant5NoYes
Chat response70b-versatile1024YesNo
Data extraction8b-instant200NoYes
Code generation70b-versatile2048YesNo
Bulk processing8b-instant256NoYes

Examples

Common scenarios mapped to the levers above. Full code for each is in references/examples.md.

  • Latency-critical classification — 8b-instant + one-word prompt + max_tokens: 5 + cache. First call ~50ms TTFT; identical repeats return from cache at ~0ms.
  • Interactive chat — 70b-versatile streamed with streamWithMetrics, printing tokens as they arrive plus a [TTFT | tok/s] footer.
  • Bulk processing (500 records) — parallelCompletions wraps each call in a rate-limit-aware p-queue and reuses the cache for duplicate rows.
  • Empirical model choice — run benchmarkModels against your real prompt, then hardcode the fastest tier that clears your quality bar.
typescript
// Latency-critical classification, cached
const label = await cachedCompletion(
  [
    { role: "system", content: "Classify as positive/negative/neutral. One word only." },
    { role: "user", content: "This product exceeded every expectation." },
  ],
  "llama-3.1-8b-instant"
);
// => "positive"

See references/examples.md for the streaming, bulk, and benchmarking walkthroughs.

Error Handling

IssueCauseSolution
High TTFTUsing 70b for simple tasksSwitch to llama-3.1-8b-instant
Rate limit (429)Over RPM or TPMUse queue with interval limiting
Stream disconnectNetwork timeoutImplement reconnection with partial content
Token overflowmax_tokens too highSet to expected output size
Cache miss rate highUnique promptsNormalize prompts, use template patterns

Resources

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/.curated/groq-performance-tuning of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/examples.md
  • references/implementation.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Groq Performance Tuning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Groq Performance Tuning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Groq Performance Tuning this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.9kAutomated safety check: PassMIT
Stripe Projectsfossasia/eventyay1.7k5 repos~2kAutomated safety check: NotesApache-2.0
FoundatioFoundatioFx/Foundatio2.1k—~3.9kAutomated safety check: PassApache-2.0
Wp Block Themesgambitph/Stackable3513 repos~985Automated safety check: PassGPL-3.0
Wp Performancegambitph/Stackable3513 repos~1.5kAutomated safety check: PassGPL-3.0
Effect Portable Patternsmillionco/expect3.6k—~3.7kAutomated safety check: PassCustom licence

Similar skills

  • Stripe Projects

    fossasia/eventyay

    A skill your agent uses when the user wants to provision infrastructure or third-party services using Stripe Projects.

    1.7k GitHub starsUsed in 5 repos~2k tokens
    Backend & APIsAuto-check: notes
  • Foundatio

    FoundatioFx/Foundatio

    A skill your agent uses when working with Foundatio infrastructure abstractions for .NET -- caching, queuing, messaging, file storage, distributed locking, or background jobs.

    2.1k GitHub stars~3.9k tokensUpdated 2 days ago
    Backend & APIsAuto-check passed
  • Wp Block Themes

    gambitph/Stackable

    A skill your agent uses when developing WordPress block themes: theme.json (global settings/styles), templates and template parts, patterns, style variations, and Site Editor troubleshooting (style…

    351 GitHub starsUsed in 3 repos~985 tokens
    Backend & APIsAuto-check passed
  • Wp Performance

    gambitph/Stackable

    A skill your agent uses when investigating or improving WordPress performance (backend-only agent): profiling and measurement (WP-CLI profile/doctor, Server-Timing, Query Monitor via REST headers)…

    351 GitHub starsUsed in 3 repos~1.5k tokens
    Backend & APIsAuto-check passed
  • Effect Portable Patterns

    millionco/expect

    Portable Effect patterns for robust promise execution. An agent skill from millionco/expect.

    3.6k GitHub stars~3.7k tokensUpdated 5 mo ago
    Backend & APIsAuto-check passed
  • FastAPI-Redis SDK Development

    redis/fastapi-redis-sdk

    Official

    Guides development on the fastapi-redis-sdk library itself - its connection lifecycle, dependency-injected caching, and async/sync bridging.

    405 GitHub stars~2.5k tokensUpdated 10 days ago
    Backend & APIsAuto-check: notes

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Categories

Questions about Groq Performance Tuning

What does Groq Performance Tuning do?

Optimize Groq API performance with model selection, caching, streaming, and parallel requests. Groq Performance Tuning is an agent skill from jeremylongshore/tons-of-skills-marketplace. Optimize Groq API performance with model selection, caching, streaming, and parallel requests.

When should I use Groq Performance Tuning?

Groq Performance Tuning fits situations like: experiencing slow responses; implementing caching strategies; optimizing request throughput for Groq integrations; with phrases like groq performance.

How do I install Groq Performance Tuning in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill groq-performance-tuning -a claude-code`. Or copy the skill folder (skills/.curated/groq-performance-tuning in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/groq-performance-tuning in your project. Claude Code loads it when a task matches its description.

How do I install Groq Performance Tuning in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill groq-performance-tuning -a codex`. Or copy the skill folder (skills/.curated/groq-performance-tuning in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/groq-performance-tuning in your project. Codex loads it when a task matches its description.

Can I use Groq Performance Tuning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill groq-performance-tuning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/groq-performance-tuning, .gemini/skills/groq-performance-tuning, .github/skills/groq-performance-tuning and .opencode/skills/groq-performance-tuning in your project.

What does Groq Performance Tuning need to run?

Going by SKILL.md and its folder, Groq Performance Tuning needs the command-line tools its instructions call (npm) and credentials named GROQ_API_KEY. Our summary lists: Node.js; A credential in GROQ_API_KEY. Its frontmatter pre-approves these tools: Read, Write, Edit. Compatibility (from SKILL.md): Designed for Claude Code.

Does Groq Performance Tuning access the network?

SKILL.md names 2 domains. As links in the text: console.groq.com and npmjs.com. This is read from the text; nothing was executed.

Is Groq Performance Tuning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Groq Performance Tuning use?

Groq Performance Tuning is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Groq Performance Tuning use?

About 1.9k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2k tokens, read only when the agent opens those files.

What are the alternatives to Groq Performance Tuning?

Skills that share tags, products or a category with Groq Performance Tuning: Stripe Projects (fossasia/eventyay, 1.7k stars), Foundatio (FoundatioFx/Foundatio, 2.1k stars), Wp Block Themes (gambitph/Stackable, 351 stars) and Wp Performance (gambitph/Stackable, 351 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Groq Performance Tuning?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.