Agent skill

AI Engineering

by swyxio in swyxio/skills

Diagnose or improve reliability of a structured, multi-request, rate-limited, or cost-sensitive AI workflow.

MITAuto-check passedBackend & APIs

Install AI Engineering

skills CLI
$ npx skills add swyxio/skills --skill ai-engineering -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install swyxio/skills ai-engineering --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/ai-engineering .claude/skills/ai-engineering && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ai-engineering
GitHub stars
175
Token cost
~2.1k tokens
SKILL.md length
1,082 words
Files
5 (incl. references)
Skills in repo
89
Repo updated
First seen
Licence
MIT

At a glance

Diagnose or improve reliability of a structured, multi-request, rate-limited, or cost-sensitive AI workflow.

  • Works in 6 steps: Inspect the boundary → Shape the request before scaling it → Scale deliberately → …
  • Truncated outputs
  • SKILL.md covers Useful defaults, Workflow and References
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

AI Engineering is an agent skill from swyxio/skills. Diagnose or improve reliability of a structured, multi-request, rate-limited, or cost-sensitive AI workflow. Use for malformed or truncated outputs, retry/rate-limit failures, unreliable fan-out, cache/resume bugs, or missing run telemetry. Do not use for ordinary prompt edits or simple one-shot model calls.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `agents/openai.yaml`, `references/failure-matrix.md` and `references/image-video-fal.md`).

It sits in Backend & APIs, covering Rate limiting. The repository describes itself as: Agent skills for Claude Code and other AI agents. The licence is MIT.

When your agent uses it

  • Truncated outputs
  • Retry/rate-limit failures
  • Unreliable fan-out
  • Cache/resume bugs

Example prompts

  • “/ai-engineering”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Inspect the boundary
  2. Shape the request before scaling it
  3. Scale deliberately
  4. Repair by failure class
  5. Persist proportionately
  6. Verify the relevant unhappy paths

What it can do on your machine

Read from SKILL.md and the folder at commit 038ef34. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AI Engineering loads about 2.1k tokens when it runs, and up to ~6.6k if it reads all its reference files. Until then it costs about 81 tokens; SKILL.md has 1,082 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~81
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from swyxio/skills at commit 038ef34, republished under its MIT licence (© swyxio). 1,082 words, ~2,146 tokens.

Download SKILL.mdSave it as .claude/skills/ai-engineering/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
ai-engineering
description
Diagnose or improve reliability of a structured, multi-request, rate-limited, or cost-sensitive AI workflow. Use for malformed or truncated outputs, retry/rate-limit failures, unreliable fan-out, cache/resume bugs, or missing run telemetry. Do not use for ordinary prompt edits or simple one-shot model calls.

AI Engineering

Use enough structure to make a model workflow explainable and recoverable without turning every prototype into an operations project. This skill owns request validation, retry/cache behavior, rate-aware fan-out, and measurement. Pair it with live-ai-pipelines only when a run needs live progress or durable resume.

For provider-specific API behavior, read provider operation notes when the adapter is OpenAI, Anthropic, or OpenRouter, then verify volatile details in current provider docs. For image/video generation or FAL model fan-outs, read image, video, and FAL operation notes.

Useful defaults

  • Improve the established workflow before introducing another path. Do not turn implementation choices or temporary exceptions into user requirements.
  • Do not publish a malformed, partial, or schema-invalid result as complete.
  • For a machine-consumed structured artifact, use the selected model provider's documented native structured-output API with an explicit schema. Do not treat a free-form completion prompted with “return JSON” plus local JSON.parse as structured output.
  • Prefer the official provider API for schema-critical work. A router is acceptable only when its exact pinned endpoint advertises native structured-output support and a canary has verified the complete request/stream/validation path; otherwise call the provider directly.
  • Prompted JSON is acceptable only as a human-readable/debug artifact. It is not an input contract for a canonical graph, database write, workflow transition, or published page.
  • Keep observed source data, generated prose, inferred claims, and repair output distinct when downstream consumers need that distinction.
  • Treat workers as an in-flight limit, not as a request-rate setting.
  • When a fallback changes coverage or semantics, record that difference rather than silently treating it as the richer result.
  • Keep provider delivery separate from domain or human acceptance; a valid artifact URL does not establish that an image, video, or other subjective output is usable.
  • Match validation to the deliverable. A machine-consumed article envelope may need a schema; illustrative code inside it is reader-facing text, not software to compile or certify unless the user explicitly requests that.
  • Keep enough redacted telemetry to answer what happened, what it cost, and what may safely resume.

Workflow

1. Inspect the boundary

Find the first boundary that changed or rejected the artifact: provider delivery, local redaction, parsing, validation or rendering. Do not assume malformed local output means the model failed. Inspect the relevant wrapper and retained evidence before changing prompts or retrying; preserve valid work and keep sensitive source text out of logs.

2. Shape the request before scaling it

Normalize source material into a usable representation before model admission, preserving originals and required coverage. Repair unreadable or oversized inputs locally rather than silently truncating them.

Budget each request so the model can finish; split or continue long work while preserving required coverage. Do not turn token budgets into arbitrary section/item ceilings that omit substantive source material. Validate schema, source references and domain invariants; native structured output does not establish semantic completeness.

Before broad admission, run cheap deterministic checks across all selected packets: identity, required fields, canonical links, source locators and alias consistency. Do not spend model calls discovering a malformed packet. Calibrate unfamiliar request classes on a small sparse/typical/dense sample; reuse comparable successful calibration when its relevant inputs have not changed. Use the result to tune per-class input and output budgets. Long-form requests benefit from a bounded evidence packet: deduplicate, rank representative support, preserve conflicts, and keep stable locators. Global synthesis should receive normalized IDs and compact summaries rather than the raw corpus plus every intermediate artifact.

For heterogeneous model fan-outs, define endpoint-specific capability and payload profiles. Include reference topology and ordering, supported parameters, safety-control fields, output schema, and fallback policy in the effective request. Do not send a universal parameter bundle or guessed provider controls.

3. Scale deliberately

For an unmeasured request class, start with modest in-flight concurrency. Preserve a measured healthy setting on resume rather than restarting its ramp. Pace requests and tokens separately, honor Retry-After, and use provider headers when available. Raise concurrency only after measuring throughput, latency, retries, 429s, context size, and remaining headroom; high latency can be a context or generation bottleneck rather than a rate-limit problem.

Show full SKILL.md (412 more words)Show less
4. Repair by failure class

After a repair, rerun only work whose inputs or acceptance were affected; reuse valid independent results.

Recover a completed response before considering another request. Retry genuinely incomplete transient failures through the same limiter; distinguish provider failure from host interruption, controller deadline and explicit cancellation. Fix validator/redactor/renderer mistakes locally, without asking the model to satisfy a broken check. For genuine truncation, compact intermediate detail or continue without dropping required coverage. For a content repair, request only the affected fields or blocks and assemble them against the saved base; do not regenerate the whole artifact for a small edit. Record any fallback's coverage loss.

See failure matrix for practical actions and lightweight attempt fields.

5. Persist proportionately

For a significant run, write complete artifacts atomically and maintain a status snapshot plus attempt history. A success cache can resume results but cannot explain failed attempts, waits, headers, or interruption. Record the logical item, effective request, outcome, latency, usage/cost when available, and redacted provider identifiers.

For subjective artifacts, persist provider completion and human adjudication independently. Support blind evaluation when requested: store artifact references and operational metadata without fetching, opening, classifying, or scoring the artifact, and leave acceptance to the named reviewer.

If workers can outlive their caller or another runner may resume work, add explicit ownership, heartbeats, and cancellation behavior. A simple in-process job does not need a lease protocol.

For publishing workflows, measure accepted changes delivered live per elapsed hour, not active slots or provider completions. Separate queue waits, useful execution, retry cooldowns, release waits and duplicated work; overlapping worker durations are not additive wall time. Missing timing or cost remains unavailable.

6. Verify the relevant unhappy paths

When changing a multi-item runner, test dependency readiness: hold one item deliberately and verify that an independent ready item advances through its eligible downstream stages before the held item settles. Also verify that genuinely dependent work and publication remain blocked until their prerequisites pass. Successful concurrent calls alone do not prove effective overlap.

Exercise the failures that the chosen provider and artifact contract make material: rate limits, timeouts, malformed/length-limited output, duplicate delivery, cancellation, restart, and cache reuse. Hand off the result coverage, notable fallbacks, cost/latency, and output/telemetry locations for significant runs.

References

© swyxio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in ai-engineering of swyxio/skills.

  • SKILL.md
  • agents/openai.yaml
  • references/failure-matrix.md
  • references/image-video-fal.md
  • references/provider-operation-notes.md

Open the folder on GitHubat commit 038ef34

Compare with similar skills

AI Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AI Engineering compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AI Engineering this skillswyxio/skills175—~2.1kAutomated safety check: PassMIT
Ag2 Middlewareag2ai/build-with-ag2252—~1.9kAutomated safety check: PassApache-2.0
Venice API Overviewveniceai/skills143—~3.5kAutomated safety check: PassMIT
Fastllm Principalsazrtydxb/Fastllm-proxy108—~865Automated safety check: PassApache-2.0
LLM Gatewaysickn33/agentic-awesome-skills47k1 repos~2.1kAutomated safety check: PassMIT
Langfuse Rate Limitsjeremylongshore/tons-of-skills-marketplace2.8k1 repos~1.8kAutomated safety check: PassMIT

Similar skills

  • Ag2 Middleware

    ag2ai/build-with-ag2

    Intercept the AG2 beta agent loop with BaseMiddleware — wrap full turns (onturn), each LLM call (onllmcall), each tool execution (ontoolexecution), or each human-input request (onhumaninput).

    252 GitHub stars~1.9k tokensUpdated 1 mo ago
    Backend & APIsAuto-check passed
  • Venice API Overview

    veniceai/skills

    High-level map of the Venice.ai API: base URL, auth modes per endpoint, endpoint categories, response headers, pricing model, error shape and versioning.

    143 GitHub stars~3.5k tokensUpdated 2 days ago
    Backend & APIsAuto-check passed
  • Fastllm Principals

    azrtydxb/Fastllm-proxy

    Manage FastLLM principals (API clients and users) and their API keys — create or delete principals, issue and revoke keys, attach roles, and set per-principal budgets and rate limits.

    108 GitHub stars~865 tokensUpdated 3 days ago
    Backend & APIsAuto-check passed
  • LLM Gateway

    sickn33/agentic-awesome-skills

    Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.

    47k GitHub starsUsed in 1 repo~2.1k tokens
    Backend & APIsAuto-check passed
  • Langfuse Rate Limits

    jeremylongshore/tons-of-skills-marketplace

    Implement Langfuse rate limiting, batching, and backoff patterns.

    2.8k GitHub starsUsed in 1 repo~1.8k tokens
    Backend & APIsAuto-check passed
  • Clade Common Errors

    jeremylongshore/tons-of-skills-marketplace

    Diagnose and fix Anthropic API errors — authentication, rate limits, Use when working with common-errors patterns.

    2.8k GitHub stars~1.6k tokensUpdated today
    Backend & APIsAuto-check passed

More from swyxio/skills

All 89 skills in this repo
  • Programmatic Agents

    swyxio/skills

    Run a selected coding-agent CLI programmatically, with latency, error, usage, cost, and trace logging.

    175 GitHub stars~2.2k tokensUpdated 3 days ago
    Auto-check passed
  • Design, implement, audit, or refresh protected username and handle namespaces for public products.

    175 GitHub stars~1.1k tokensUpdated 3 days ago
    Auto-check passed
  • New Mac Setup

    swyxio/skills

    Fully automated new Mac setup for fullstack web developers and AI engineers.

    175 GitHub stars~4.3k tokensUpdated 3 days ago
    Auto-check passed
  • Youtube API

    swyxio/skills

    Manage YouTube videos programmatically via the YouTube Data API v3 — upload video files, upload custom thumbnails, update video metadata (titles, descriptions, tags), and query video/channel info…

    175 GitHub stars~2.2k tokensUpdated 3 days ago
    Auto-check passed
  • Batch YouTube Studio upload workflow for videos sourced from Airtable, Google Drive, Loom, YouTube, or local files.

    175 GitHub stars~1.5k tokensUpdated 3 days ago
    Auto-check: warnings
  • Reconstruct and visually analyze paired agent, game, or policy trajectories to determine whether changed actions produced their intended effects.

    175 GitHub stars~1.8k tokensUpdated 3 days ago
    Auto-check passed

Questions about AI Engineering

What does AI Engineering do?

Diagnose or improve reliability of a structured, multi-request, rate-limited, or cost-sensitive AI workflow. AI Engineering is an agent skill from swyxio/skills. Diagnose or improve reliability of a structured, multi-request, rate-limited, or cost-sensitive AI workflow.

When should I use AI Engineering?

AI Engineering fits situations like: truncated outputs; retry/rate-limit failures; unreliable fan-out; cache/resume bugs.

How do I install AI Engineering in Claude Code?

Run `npx skills add swyxio/skills --skill ai-engineering -a claude-code`. Or copy the skill folder (ai-engineering in swyxio/skills) into .claude/skills/ai-engineering in your project. Claude Code loads it when a task matches its description.

How do I install AI Engineering in Codex?

Run `npx skills add swyxio/skills --skill ai-engineering -a codex`. Or copy the skill folder (ai-engineering in swyxio/skills) into .agents/skills/ai-engineering in your project. Codex loads it when a task matches its description.

Can I use AI Engineering in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add swyxio/skills --skill ai-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-engineering, .gemini/skills/ai-engineering, .github/skills/ai-engineering and .opencode/skills/ai-engineering in your project.

What does AI Engineering need to run?

SKILL.md names no scripts, command-line tools or credentials: AI Engineering is instructions for the agent only.

Does AI Engineering access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is AI Engineering safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does AI Engineering use?

AI Engineering is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AI Engineering use?

About 2.1k tokens (SKILL.md is roughly 8.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.4k tokens, read only when the agent opens those files.

What are the alternatives to AI Engineering?

Skills that share tags, products or a category with AI Engineering: Ag2 Middleware (ag2ai/build-with-ag2, 252 stars), Venice API Overview (veniceai/skills, 143 stars), Fastllm Principals (azrtydxb/Fastllm-proxy, 108 stars) and LLM Gateway (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AI Engineering?

swyxio (a GitHub user) maintains it in swyxio/skills, which has 175 GitHub stars. The repository holds 89 skills in this directory. The repository was last updated on October 5, 2026.

Source: swyxio/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.