Official agent skill

Arize Prompt Optimization

by github in github/awesome-copilot

Optimizes, improves, and debugs LLM prompts using production trace data, evaluations, and annotations.

OfficialMITAuto-check: notesAI & LLM Engineering

Install Arize Prompt Optimization

skills CLI
$ npx skills add github/awesome-copilot --skill arize-prompt-optimization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install github/awesome-copilot arize-prompt-optimization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/arize-prompt-optimization .claude/skills/arize-prompt-optimization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
arize-prompt-optimization
GitHub stars
40k
Used in
1 other repo
Token cost
~4.8k tokens
SKILL.md length
1,370 words
Files
3 (incl. references)
Skills in repo
417
Repo updated
First seen
Licence
MIT

At a glance

Optimizes, improves, and debugs LLM prompts using production trace data, evaluations, and annotations.

  • Works in 4 steps: Extract the Current Prompt → Gather Performance Data → Optimize the Prompt → …
  • The user mentions optimize prompt
  • SKILL.md covers Concepts, Prerequisites, Phase 1: Extract the Current… and Phase 2: Gather Performance Data, plus 4 more sections
  • Calls jq; needs OPENAI_API_KEY and ANTHROPIC_API_KEY

What it does

Arize Prompt Optimization is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Optimizes, improves, and debugs LLM prompts using production trace data, evaluations, and annotations. Extracts prompts from spans, gathers performance signal, and runs a data-driven optimization loop using the ax CLI. Use when the user mentions optimize prompt, improve prompt, make AI respond better, improve output quality, prompt engineering, prompt tuning, or system prompt improvement.

Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/ax-profiles.md` and `references/ax-setup.md`). Compatibility notes: Requires the ax CLI and a configured Arize profile.

It sits in AI & LLM Engineering, covering Prompt engineering. The repository describes itself as: Community-contributed instructions, agents, skills, and configurations to help you make the most of GitHub Copilot. The licence is MIT.

When your agent uses it

  • The user mentions optimize prompt
  • Make AI respond better
  • Improve output quality
  • Prompt engineering

Example prompts

  • “/arize-prompt-optimization”

Requirements

  • A credential in OPENAI_API_KEY
  • A credential in ANTHROPIC_API_KEY
  • Compatibility (from SKILL.md): Requires the ax CLI and a configured Arize profile.

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Extract the Current Prompt
  2. Gather Performance Data
  3. Optimize the Prompt
  4. Iterate

What it can do on your machine

Read from SKILL.md and the folder at commit 727ff2e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY
    • ANTHROPIC_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires the ax CLI and a configured Arize profile.

    From compatibility in the SKILL.md frontmatter.

Context cost

Arize Prompt Optimization loads about 4.8k tokens when it runs, and up to ~6.3k if it reads all its reference files. Until then it costs about 104 tokens; SKILL.md has 1,370 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~104
When it runs · the whole SKILL.md, loaded when a task matches
~4.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:63
    - **Security:** Never read `.env` files or search the filesystem for credentials. Use `ax profiles` for Arize credential

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from github/awesome-copilot at commit 727ff2e, republished under its MIT licence (© github). 1,370 words, ~4,799 tokens.

Download SKILL.mdSave it as .claude/skills/arize-prompt-optimization/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
arize-prompt-optimization
description
Optimizes, improves, and debugs LLM prompts using production trace data, evaluations, and annotations. Extracts prompts from spans, gathers performance signal, and runs a data-driven optimization loop using the ax CLI. Use when the user mentions optimize prompt, improve prompt, make AI respond better, improve output quality, prompt engineering, prompt tuning, or system prompt improvement.
compatibility
Requires the ax CLI and a configured Arize profile.
metadata.author
arize
metadata.version
1.0

Arize Prompt Optimization Skill

SPACE — All --space flags and the ARIZE_SPACE env var accept a space name (e.g., my-workspace) or a base64 space ID (e.g., U3BhY2U6...). Find yours with ax spaces list.

Concepts

Where Prompts Live in Trace Data

LLM applications emit spans following OpenInference semantic conventions. Prompts are stored in different span attributes depending on the span kind and instrumentation:

ColumnWhat it containsWhen to use
attributes.llm.input_messagesStructured chat messages (system, user, assistant, tool) in role-based formatPrimary source for chat-based LLM prompts
attributes.llm.input_messages.rolesArray of roles: system, user, assistant, toolExtract individual message roles
attributes.llm.input_messages.contentsArray of message content stringsExtract message text
attributes.input.valueSerialized prompt or user question (generic, all span kinds)Fallback when structured messages are not available
attributes.llm.prompt_template.templateTemplate with {variable} placeholders (e.g., "Answer {question} using {context}")When the app uses prompt templates
attributes.llm.prompt_template.variablesTemplate variable values (JSON object)See what values were substituted into the template
attributes.output.valueModel response textSee what the LLM produced
attributes.llm.output_messagesStructured model output (including tool calls)Inspect tool-calling responses
Finding Prompts by Span Kind
  • LLM span (attributes.openinference.span.kind = 'LLM'): Check attributes.llm.input_messages for structured chat messages, OR attributes.input.value for a serialized prompt. Check attributes.llm.prompt_template.template for the template.
  • Chain/Agent span: attributes.input.value contains the user's question. The actual LLM prompt lives on child LLM spans -- navigate down the trace tree.
  • Tool span: attributes.input.value has tool input, attributes.output.value has tool result. Not typically where prompts live.
Performance Signal Columns

These columns carry the feedback data used for optimization:

Column patternSourceWhat it tells you
annotation.<name>.labelHuman reviewersCategorical grade (e.g., correct, incorrect, partial)
annotation.<name>.scoreHuman reviewersNumeric quality score (e.g., 0.0 - 1.0)
annotation.<name>.textHuman reviewersFreeform explanation of the grade
eval.<name>.labelLLM-as-judge evalsAutomated categorical assessment
eval.<name>.scoreLLM-as-judge evalsAutomated numeric score
eval.<name>.explanationLLM-as-judge evalsWhy the eval gave that score -- most valuable for optimization
attributes.input.valueTrace dataWhat went into the LLM
attributes.output.valueTrace dataWhat the LLM produced
{experiment_name}.outputExperiment runsOutput from a specific experiment

Prerequisites

Proceed directly with the task — run the ax command you need. Do NOT check versions, env vars, or profiles upfront.

If an ax command fails, troubleshoot based on the error:

  • command not found or version error → see references/ax-setup.md
  • 401 Unauthorized / missing API key → run ax profiles show to inspect the current profile. If the profile is missing or the API key is wrong, follow references/ax-profiles.md to create/update it. If the user doesn't have their key, direct them to https://app.arize.com/admin > API Keys
  • Space unknown → run ax spaces list to pick by name, or ask the user
  • Project unclear → ask the user, or run ax projects list -o json --limit 100 and present as selectable options
  • LLM provider call fails (missing OPENAI_API_KEY / ANTHROPIC_API_KEY) → run ax ai-integrations list --space SPACE to check for platform-managed credentials. If none exist, ask the user to provide the key or create an integration via the arize-ai-provider-integration skill
  • Security: Never read .env files or search the filesystem for credentials. Use ax profiles for Arize credentials and ax ai-integrations for LLM provider keys. If credentials are not available through these channels, ask the user.

Phase 1: Extract the Current Prompt

Find LLM spans containing prompts
bash
# Sample LLM spans (where prompts live)
ax spans export PROJECT --filter "attributes.openinference.span.kind = 'LLM'" -l 10 --stdout

# Filter by model
ax spans export PROJECT --filter "attributes.llm.model_name = 'gpt-4o'" -l 10 --stdout

# Filter by span name (e.g., a specific LLM call)
ax spans export PROJECT --filter "name = 'ChatCompletion'" -l 10 --stdout
Export a trace to inspect prompt structure
bash
# Export all spans in a trace
ax spans export PROJECT --trace-id TRACE_ID

# Export a single span
ax spans export PROJECT --span-id SPAN_ID
Extract prompts from exported JSON
bash
# Extract structured chat messages (system + user + assistant)
jq '.[0] | {
  messages: .attributes.llm.input_messages,
  model: .attributes.llm.model_name
}' trace_*/spans.json

# Extract the system prompt specifically
jq '[.[] | select(.attributes.llm.input_messages.roles[]? == "system")] | .[0].attributes.llm.input_messages' trace_*/spans.json

# Extract prompt template and variables
jq '.[0].attributes.llm.prompt_template' trace_*/spans.json

# Extract from input.value (fallback for non-structured prompts)
jq '.[0].attributes.input.value' trace_*/spans.json
Reconstruct the prompt as messages

Once you have the span data, reconstruct the prompt as a messages array:

json
[
  {"role": "system", "content": "You are a helpful assistant that..."},
  {"role": "user", "content": "Given {input}, answer the question: {question}"}
]

If the span has attributes.llm.prompt_template.template, the prompt uses variables. Preserve these placeholders ({variable} or {{variable}}) -- they are substituted at runtime.

Phase 2: Gather Performance Data

From traces (production feedback)
bash
# Find error spans -- these indicate prompt failures
ax spans export PROJECT \
  --filter "status_code = 'ERROR' AND attributes.openinference.span.kind = 'LLM'" \
  -l 20 --stdout

# Find spans with low eval scores
ax spans export PROJECT \
  --filter "annotation.correctness.label = 'incorrect'" \
  -l 20 --stdout

# Find spans with high latency (may indicate overly complex prompts)
ax spans export PROJECT \
  --filter "attributes.openinference.span.kind = 'LLM' AND latency_ms > 10000" \
  -l 20 --stdout

# Export error traces for detailed inspection
ax spans export PROJECT --trace-id TRACE_ID
From datasets and experiments
bash
# Export a dataset (ground truth examples)
ax datasets export DATASET_NAME --space SPACE
# -> dataset_*/examples.json

# Export experiment results (what the LLM produced)
ax experiments export EXPERIMENT_NAME --dataset DATASET_NAME --space SPACE
# -> experiment_*/runs.json
Merge dataset + experiment for analysis

Join the two files by example_id to see inputs alongside outputs and evaluations:

bash
# Count examples and runs
jq 'length' dataset_*/examples.json
jq 'length' experiment_*/runs.json

# View a single joined record
jq -s '
  .[0] as $dataset |
  .[1][0] as $run |
  ($dataset[] | select(.id == $run.example_id)) as $example |
  {
    input: $example,
    output: $run.output,
    evaluations: $run.evaluations
  }
' dataset_*/examples.json experiment_*/runs.json

# Find failed examples (where eval score < threshold)
jq '[.[] | select(.evaluations.correctness.score < 0.5)]' experiment_*/runs.json
Identify what to optimize

Look for patterns across failures:

  1. Compare outputs to ground truth: Where does the LLM output differ from expected?
  2. Read eval explanations: eval.*.explanation tells you WHY something failed
  3. Check annotation text: Human feedback describes specific issues
  4. Look for verbosity mismatches: If outputs are too long/short vs ground truth
  5. Check format compliance: Are outputs in the expected format?

Phase 3: Optimize the Prompt

The Optimization Meta-Prompt

Use this template to generate an improved version of the prompt. Fill in the three placeholders and send it to your LLM (GPT-4o, Claude, etc.):

You are an expert in prompt optimization. Given the original baseline prompt
and the associated performance data (inputs, outputs, evaluation labels, and
explanations), generate a revised version that improves results.

ORIGINAL BASELINE PROMPT
========================

{PASTE_ORIGINAL_PROMPT_HERE}

========================

PERFORMANCE DATA
================

The following records show how the current prompt performed. Each record
includes the input, the LLM output, and evaluation feedback:

{PASTE_RECORDS_HERE}

================

HOW TO USE THIS DATA

1. Compare outputs: Look at what the LLM generated vs what was expected
2. Review eval scores: Check which examples scored poorly and why
3. Examine annotations: Human feedback shows what worked and what didn't
4. Identify patterns: Look for common issues across multiple examples
5. Focus on failures: The rows where the output DIFFERS from the expected
   value are the ones that need fixing

ALIGNMENT STRATEGY

- If outputs have extra text or reasoning not present in the ground truth,
  remove instructions that encourage explanation or verbose reasoning
- If outputs are missing information, add instructions to include it
- If outputs are in the wrong format, add explicit format instructions
- Focus on the rows where the output differs from the target -- these are
  the failures to fix

RULES

Maintain Structure:
- Use the same template variables as the current prompt ({var} or {{var}})
- Don't change sections that are already working
- Preserve the exact return format instructions from the original prompt

Avoid Overfitting:
- DO NOT copy examples verbatim into the prompt
- DO NOT quote specific test data outputs exactly
- INSTEAD: Extract the ESSENCE of what makes good vs bad outputs
- INSTEAD: Add general guidelines and principles
- INSTEAD: If adding few-shot examples, create SYNTHETIC examples that
  demonstrate the principle, not real data from above

Goal: Create a prompt that generalizes well to new inputs, not one that
memorizes the test data.

OUTPUT FORMAT

Return the revised prompt as a JSON array of messages:

[
  {"role": "system", "content": "..."},
  {"role": "user", "content": "..."}
]

Also provide a brief reasoning section (bulleted list) explaining:
- What problems you found
- How the revised prompt addresses each one
Preparing the performance data

Format the records as a JSON array before pasting into the template:

bash
# From dataset + experiment: join and select relevant columns
jq -s '
  .[0] as $ds |
  [.[1][] | . as $run |
    ($ds[] | select(.id == $run.example_id)) as $ex |
    {
      input: $ex.input,
      expected: $ex.expected_output,
      actual_output: $run.output,
      eval_score: $run.evaluations.correctness.score,
      eval_label: $run.evaluations.correctness.label,
      eval_explanation: $run.evaluations.correctness.explanation
    }
  ]
' dataset_*/examples.json experiment_*/runs.json

# From exported spans: extract input/output pairs with annotations
jq '[.[] | select(.attributes.openinference.span.kind == "LLM") | {
  input: .attributes.input.value,
  output: .attributes.output.value,
  status: .status_code,
  model: .attributes.llm.model_name
}]' trace_*/spans.json
Applying the revised prompt

After the LLM returns the revised messages array:

  1. Compare the original and revised prompts side by side
  2. Verify all template variables are preserved
  3. Check that format instructions are intact
  4. Test on a few examples before full deployment

Phase 4: Iterate

The optimization loop
1. Extract prompt    -> Phase 1 (once)
2. Run experiment    -> ax experiments create ...
3. Export results    -> ax experiments export EXPERIMENT_NAME --dataset DATASET_NAME --space SPACE
4. Analyze failures  -> jq to find low scores
5. Run meta-prompt   -> Phase 3 with new failure data
6. Apply revised prompt
7. Repeat from step 2
Measure improvement
bash
# Compare scores across experiments
# Experiment A (baseline)
jq '[.[] | .evaluations.correctness.score] | add / length' experiment_a/runs.json

# Experiment B (optimized)
jq '[.[] | .evaluations.correctness.score] | add / length' experiment_b/runs.json

# Find examples that flipped from fail to pass
jq -s '
  [.[0][] | select(.evaluations.correctness.label == "incorrect")] as $fails |
  [.[1][] | select(.evaluations.correctness.label == "correct") |
    select(.example_id as $id | $fails | any(.example_id == $id))
  ] | length
' experiment_a/runs.json experiment_b/runs.json
A/B compare two prompts
  1. Create two experiments against the same dataset, each using a different prompt version
  2. Export both: ax experiments export EXP_A and ax experiments export EXP_B
  3. Compare average scores, failure rates, and specific example flips
  4. Check for regressions -- examples that passed with prompt A but fail with prompt B
Show full SKILL.md (643 more words)Show less

Prompt Engineering Best Practices

Apply these when writing or revising prompts:

TechniqueWhen to applyExample
Clear, detailed instructionsOutput is vague or off-topic"Classify the sentiment as exactly one of: positive, negative, neutral"
Instructions at the beginningModel ignores later instructionsPut the task description before examples
Step-by-step breakdownsComplex multi-step processes"First extract entities, then classify each, then summarize"
Specific personasNeed consistent style/tone"You are a senior financial analyst writing for institutional investors"
Delimiter tokensSections blend togetherUse ---, ###, or XML tags to separate input from instructions
Few-shot examplesOutput format needs clarificationShow 2-3 synthetic input/output pairs
Output length specificationsResponses are too long or short"Respond in exactly 2-3 sentences"
Reasoning instructionsAccuracy is critical"Think step by step before answering"
"I don't know" guidelinesHallucination is a risk"If the answer is not in the provided context, say 'I don't have enough information'"
Variable preservation

When optimizing prompts that use template variables:

  • Single braces ({variable}): Python f-string / Jinja style. Most common in Arize.
  • Double braces ({{variable}}): Mustache style. Used when the framework requires it.
  • Never add or remove variable placeholders during optimization
  • Never rename variables -- the runtime substitution depends on exact names
  • If adding few-shot examples, use literal values, not variable placeholders

Workflows

Optimize a prompt from a failing trace
  1. Find failing traces:
    bash
    ax traces list PROJECT --filter "status_code = 'ERROR'" --limit 5
  2. Export the trace:
    bash
    ax spans export PROJECT --trace-id TRACE_ID
  3. Extract the prompt from the LLM span:
    bash
    jq '[.[] | select(.attributes.openinference.span.kind == "LLM")][0] | {
      messages: .attributes.llm.input_messages,
      template: .attributes.llm.prompt_template,
      output: .attributes.output.value,
      error: .attributes.exception.message
    }' trace_*/spans.json
  4. Identify what failed from the error message or output
  5. Fill in the optimization meta-prompt (Phase 3) with the prompt and error context
  6. Apply the revised prompt
Optimize using a dataset and experiment
  1. Find the dataset and experiment:
    bash
    ax datasets list --space SPACE
    ax experiments list --dataset DATASET_NAME --space SPACE
  2. Export both:
    bash
    ax datasets export DATASET_NAME --space SPACE
    ax experiments export EXPERIMENT_NAME --dataset DATASET_NAME --space SPACE
  3. Prepare the joined data for the meta-prompt
  4. Run the optimization meta-prompt
  5. Create a new experiment with the revised prompt to measure improvement
Debug a prompt that produces wrong format
  1. Export spans where the output format is wrong:
    bash
    ax spans export PROJECT \
      --filter "attributes.openinference.span.kind = 'LLM' AND annotation.format.label = 'incorrect'" \
      -l 10 --stdout > bad_format.json
  2. Look at what the LLM is producing vs what was expected
  3. Add explicit format instructions to the prompt (JSON schema, examples, delimiters)
  4. Common fix: add a few-shot example showing the exact desired output format
Reduce hallucination in a RAG prompt
  1. Find traces where the model hallucinated:
    bash
    ax spans export PROJECT \
      --filter "annotation.faithfulness.label = 'unfaithful'" \
      -l 20 --stdout
  2. Export and inspect the retriever + LLM spans together:
    bash
    ax spans export PROJECT --trace-id TRACE_ID
    jq '[.[] | {kind: .attributes.openinference.span.kind, name, input: .attributes.input.value, output: .attributes.output.value}]' trace_*/spans.json
  3. Check if the retrieved context actually contained the answer
  4. Add grounding instructions to the system prompt: "Only use information from the provided context. If the answer is not in the context, say so."

Troubleshooting

ProblemSolution
ax: command not foundSee references/ax-setup.md
No profile foundNo profile is configured. See references/ax-profiles.md to create one.
No input_messages on spanCheck span kind -- Chain/Agent spans store prompts on child LLM spans, not on themselves
Prompt template is nullNot all instrumentations emit prompt_template. Use input_messages or input.value instead
Variables lost after optimizationVerify the revised prompt preserves all {var} placeholders from the original
Optimization makes things worseCheck for overfitting -- the meta-prompt may have memorized test data. Ensure few-shot examples are synthetic
No eval/annotation columnsRun evaluations first (via Arize UI or SDK), then re-export
Experiment output column not foundThe column name is {experiment_name}.output -- check exact experiment name via ax experiments get
jq errors on span JSONEnsure you're targeting the correct file path (e.g., trace_*/spans.json)

© github, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/arize-prompt-optimization of github/awesome-copilot.

  • SKILL.md
  • references/ax-profiles.md
  • references/ax-setup.md

Open the folder on GitHubat commit 727ff2e

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in github/awesome-copilot, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Arize Prompt Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Arize Prompt Optimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Arize Prompt Optimization this skillgithub/awesome-copilot40k1 repos~4.8kAutomated safety check: NotesMIT
Prompt Improverseverity1/claude-code-prompt-improver1.9k2 repos~1.7kAutomated safety check: PassMIT
Prompt Engineering Patternsynulihao/AgentSkillOS61714 repos~1.7kAutomated safety check: PassNone
Patch CreationPiebald-AI/tweakcc2.5k—~1.6kAutomated safety check: PassMIT
LLM Application DevMoizIbnYousaf/ai-agent-skills1.1k2 repos~1.3kAutomated safety check: PassMIT
Senior Prompt Engineermaslennikov-ig/claude-code-orchestrator-kit2594 repos~1.4kAutomated safety check: PassCustom licence

Similar skills

  • Prompt Improver

    severity1/claude-code-prompt-improver

    This skill enriches vague prompts with targeted research and clarification before execution.

    1.9k GitHub starsUsed in 2 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Prompt Engineering Patterns

    ynulihao/AgentSkillOS

    Master advanced prompt engineering techniques to maximize LLM performance, reliability, and controllability in production.

    617 GitHub starsUsed in 14 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Patch Creation

    Piebald-AI/tweakcc

    Create and register new patches for tweakcc. An agent skill from Piebald-AI/tweakcc.

    2.5k GitHub stars~1.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • LLM Application Dev

    MoizIbnYousaf/ai-agent-skills

    Building applications with Large Language Models - prompt engineering, RAG patterns, and LLM integration.

    1.1k GitHub starsUsed in 2 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Senior Prompt Engineer

    maslennikov-ig/claude-code-orchestrator-kit

    Provides reference guides and Python scripts for prompt optimization, RAG evaluation, and agent orchestration when building or tuning LLM systems.

    259 GitHub starsUsed in 4 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Codex Fable5

    baskduf/FableCodex

    Apply a Claude Fable 5 inspired operating style inside Codex.

    437 GitHub stars~1.6k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed

More from github/awesome-copilot

All 417 skills in this repo
  • Acquire Codebase Knowledge

    github/awesome-copilot

    Official

    Maps an unfamiliar codebase into seven evidence-backed documents in docs/codebase/, using a scan script and templates, for onboarding or architecture write-ups.

    40k GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Azure Architecture Autopilot

    github/awesome-copilot

    Official

    Designs Azure infrastructure from a natural-language description, or diagrams an existing resource group, then refines the design through conversation and deploys it with Bicep.

    40k GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Draw.io Diagram Generator

    github/awesome-copilot

    Official

    Generates, edits and validates draw.io files with correct mxGraph XML, covering flowcharts, architecture, sequence, ER and UML class diagrams.

    40k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Credit Risk Data Cleaning

    github/awesome-copilot

    Official

    Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.

    40k GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Daily Focus Board

    github/awesome-copilot

    Official

    Builds a warm, browser-based daily focus board the user updates by talking to their agent, with Eisenhower priorities, a brain-dump box and kind not-today carryover.

    40k GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Python Pypi Package Builder

    github/awesome-copilot

    Official

    End-to-end skill for building, testing, linting, versioning, and publishing a production-grade Python library to PyPI.

    40k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Questions about Arize Prompt Optimization

What does Arize Prompt Optimization do?

Optimizes, improves, and debugs LLM prompts using production trace data, evaluations, and annotations. Arize Prompt Optimization is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Optimizes, improves, and debugs LLM prompts using production trace data, evaluations, and annotations.

When should I use Arize Prompt Optimization?

Arize Prompt Optimization fits situations like: the user mentions optimize prompt; make AI respond better; improve output quality; prompt engineering.

How do I install Arize Prompt Optimization in Claude Code?

Run `npx skills add github/awesome-copilot --skill arize-prompt-optimization -a claude-code`. Or copy the skill folder (skills/arize-prompt-optimization in github/awesome-copilot) into .claude/skills/arize-prompt-optimization in your project. Claude Code loads it when a task matches its description.

How do I install Arize Prompt Optimization in Codex?

Run `npx skills add github/awesome-copilot --skill arize-prompt-optimization -a codex`. Or copy the skill folder (skills/arize-prompt-optimization in github/awesome-copilot) into .agents/skills/arize-prompt-optimization in your project. Codex loads it when a task matches its description.

Can I use Arize Prompt Optimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add github/awesome-copilot --skill arize-prompt-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/arize-prompt-optimization, .gemini/skills/arize-prompt-optimization, .github/skills/arize-prompt-optimization and .opencode/skills/arize-prompt-optimization in your project.

What does Arize Prompt Optimization need to run?

Going by SKILL.md and its folder, Arize Prompt Optimization needs the command-line tools its instructions call (jq) and credentials named OPENAI_API_KEY and ANTHROPIC_API_KEY. Our summary lists: A credential in OPENAI_API_KEY; A credential in ANTHROPIC_API_KEY. Compatibility (from SKILL.md): Requires the ax CLI and a configured Arize profile..

Does Arize Prompt Optimization access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Arize Prompt Optimization safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Arize Prompt Optimization use?

Arize Prompt Optimization is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Arize Prompt Optimization use?

About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.5k tokens, read only when the agent opens those files.

What are the alternatives to Arize Prompt Optimization?

Skills that share tags, products or a category with Arize Prompt Optimization: Prompt Improver (severity1/claude-code-prompt-improver, 1.9k stars), Prompt Engineering Patterns (ynulihao/AgentSkillOS, 617 stars), Patch Creation (Piebald-AI/tweakcc, 2.5k stars) and LLM Application Dev (MoizIbnYousaf/ai-agent-skills, 1.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Arize Prompt Optimization?

github (a GitHub organization, an official publisher) maintains it in github/awesome-copilot, which has 39,748 GitHub stars. The repository holds 417 skills in this directory. The repository was last updated on October 7, 2026.

Source: github/awesome-copilot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.