Official agent skill

Runtime Behavior Probe

by openai in openai/openai-agents-python

Plan controlled runtime probes when explicitly invoked; execute only after the required probe approval.

OfficialMITAuto-check passedDevelopment

Install Runtime Behavior Probe

skills CLI
$ npx skills add openai/openai-agents-python --skill runtime-behavior-probe -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install openai/openai-agents-python runtime-behavior-probe --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/openai/openai-agents-python.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/runtime-behavior-probe .claude/skills/runtime-behavior-probe && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
runtime-behavior-probe
GitHub stars
30k
Token cost
~3.8k tokens
SKILL.md length
2,186 words
Files
7 (incl. references)
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

Plan controlled runtime probes when explicitly invoked; execute only after the required probe approval.

  • Works in 12 steps: Restate the investigation target in… → Do a short preflight. Check the relevant… → Define the decision signal before… → …
  • Development work in your project
  • SKILL.md covers Overview, Core Rules, Workflow and Validation Matrix, plus 3 more sections
  • Runs Python scripts from its folder; calls uv; needs OPENAI_API_KEY

What it does

Runtime Behavior Probe is an agent skill from openai/openai-agents-python, published by the product's own GitHub organization. Plan controlled runtime probes when explicitly invoked; execute only after the required probe approval.

Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including reference files (for example `agents/openai.yaml`, `references/error-cases.md` and `references/openai-runtime-patterns.md`).

It sits in Development. It works with OpenAI. The repository describes itself as: A lightweight, powerful framework for multi-agent workflows. The licence is MIT.

When your agent uses it

  • Development work in your project

Example prompts

  • “/runtime-behavior-probe”

Requirements

  • Python 3
  • A credential in OPENAI_API_KEY

Workflow steps

12 steps, taken from the first numbered list in SKILL.md.

  1. Restate the investigation target in operational terms. Name the runtime surface, the key uncertainty, and the highest-risk behaviors to…
  2. Do a short preflight. Check the relevant code or docs first, decide whether the question needs local or live validation, and note any…
  3. Define the decision signal before building the matrix. For a suspected defect, name the exact user-visible symptom, the command or probe…
  4. Create a validation matrix before executing probes. Cover both baseline behavior and the most relevant failure or drift cases. The matrix…
  5. For each case, choose an execution mode up front
  6. When the question is benchmark-like or comparative, run in phases. Start with a high-signal pilot matrix against a control, then expand…
  7. If the question is about a suspected regression or behavior change, add at least one known-good control case such as origin/main, the…
  8. For comparative probes, define parity before execution. Record prompt or input shape, tool-choice setup, model-settings parity, state…
  9. If the question asks whether one option has the same intelligence or quality as another, decide whether the matrix supports only…
  10. Plan state controls before execution when hidden state could affect the result. Record whether each case uses fresh or reused state, how…
  11. If any live case will read environment variables, list the exact variable names and purpose for each case, then ask the user for approval…
  12. Build task-specific probe scripts in a temporary location. Keep the script small, observable, and easy to discard.

What it can do on your machine

Read from SKILL.md and the folder at commit 26345c1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Runtime Behavior Probe loads about 3.8k tokens when it runs, and up to ~9.9k if it reads all its reference files. Until then it costs about 32 tokens; SKILL.md has 2,186 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~32
When it runs · the whole SKILL.md, loaded when a task matches
~3.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from openai/openai-agents-python at commit 26345c1, republished under its MIT licence (© openai). 2,186 words, ~3,839 tokens.

Download SKILL.mdSave it as .claude/skills/runtime-behavior-probe/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
runtime-behavior-probe
description
Plan controlled runtime probes when explicitly invoked; execute only after the required probe approval.

Runtime Behavior Probe

Overview

Use this skill to investigate real runtime behavior, not to restate code or documentation. Start by planning the investigation, then execute a case matrix, record observed behavior, and report both the findings and the method used to obtain them.

Core Rules

  • Treat this skill as manual-only. Do not rely on implicit invocation.
  • Invoking this skill authorizes planning only. Every runtime probe requires explicit user approval after the exact probe has been proposed. Do not infer execution approval from the skill invocation or a general request to investigate runtime behavior.
  • Before requesting approval, disclose the source identity, exact command, transitively executed material, known filesystem, environment, network, and host-service capabilities, expected side effects, and control for the proposed probe. Mark unknown capabilities as unknown rather than assuming that they are unavailable.
  • Wait for an affirmative response before executing the probe. Approval is bound to the disclosed source, command, executed material, and capability scope. Obtain new approval before changing any of those fields, adding another probe, or expanding the approved matrix.
  • A baseline success or smoke case is often the right entry point, but do not stop there when the real question involves edge cases, drift, or failure behavior.
  • Plan before running anything. Write the case matrix first, then fill it in with observed results. The matrix can live in a scratch note, a temporary file, or the probe script header.
  • Default to proposing local or read-only probes. Consider a live service only when it is clearly relevant, then apply the lightweight gates below before requesting approval.
  • Size the probe to the decision. Start with the smallest matrix that can disqualify or validate the current hypothesis, then expand only when uncertainty remains.
  • Before a live probe, apply three lightweight gates:
    • Destination gate. Use only a live destination that is clearly allowed for the task.
    • Intent gate. Run the live probe only when the user explicitly wants runtime verification on that integration, or explicitly approves it after you propose the probe.
    • Data gate. If the probe will read environment variables, mutate remote state, incur material cost, or exercise non-public or user data, name the exact variable names or data class and get explicit approval first.
  • Classify each case as read-only, mutating, or costly before execution. For mutating or costly cases, or for any live case that will read environment variables, define cleanup or rollback before running the probe.
  • Use temporary files or a temporary directory for one-off probe scripts.
  • Keep temporary artifacts until the final response is drafted. Then delete them by default unless the user asked to keep them or they are needed for follow-up. Even when artifacts are deleted, keep a short run summary of the command shape, runtime context, and artifact status in the report.
  • Before executing a live probe that will read environment variables, tell the user the exact variable names you plan to use and why, then wait for explicit approval. Examples include OPENAI_API_KEY and other expected default names for the system under test.
  • When the environment-variable approval gate is required and the request_user_input tool is available, use that tool instead of a plain-text approval question. Ask one concise question with mutually exclusive choices such as Allow once (Recommended) and Do not allow, omit autoResolutionMs, and make the approval single-probe and limited to the exact named variables and destination. If the tool is unavailable, fall back to a concise plain-text approval question and do not proceed until the user explicitly approves.
  • Never print secrets, even when they come from standard environment variables that this skill may use.
  • For OpenAI API or OpenAI platform probes in this repository, use $openai-knowledge early to confirm contract-sensitive details such as supported parameters, field names, and limits. Use runtime probing to validate or challenge the documented behavior, not to skip the documentation pass entirely. If the docs MCP is unavailable, fall back to the official OpenAI docs and say that you used the fallback in the report.
  • For benchmark or comparison probes, make parity explicit before execution. Record what is held constant, what variable is under test, which response-shape constraints keep the comparison fair, and any usage or token counters that matter for interpreting latency or cost.
  • For OpenAI hosted tool probes, remove setup ambiguity before attributing a negative result to runtime behavior:
    • Force the tool path with the matching tool_choice when the question depends on tool invocation.
    • Treat container_auto and container_reference as separate cases, not interchangeable setup details.
    • Clear unsupported model or tool options first so they do not invalidate the probe.

Workflow

  1. Restate the investigation target in operational terms. Name the runtime surface, the key uncertainty, and the highest-risk behaviors to test.
  2. Do a short preflight. Check the relevant code or docs first, decide whether the question needs local or live validation, and note any repo, baseline, or release boundary that matters.
  3. Define the decision signal before building the matrix. For a suspected defect, name the exact user-visible symptom, the command or probe that can distinguish it from correct behavior, the expected failing observation, and a known-good control. Confirm that the signal exercises the real producer and caller path rather than only an adjacent helper. Prefer a fast, deterministic local loop when one can answer the question. If no credible signal can be built, state the missing access or artifact and do not substitute a nearby behavior as proof.
  4. Create a validation matrix before executing probes. Cover both baseline behavior and the most relevant failure or drift cases. The matrix can live in a scratch note, a temporary file, or a structured header inside the probe script.
  5. For each case, choose an execution mode up front:
    • single-shot for deterministic one-run checks.
    • repeat-N for cache, retry, streaming, interruption, rate-limit, concurrency, or other run-to-run-sensitive behavior.
    • warm-up + repeat-N when first-run cold-start effects could distort the result. Use these defaults unless the task clearly needs something else:
    • Quick screen of a repeat-sensitive question: repeat-3.
    • Decision-grade latency or release recommendation: warm-up + repeat-10.
    • Costly live cases: start at repeat-3, then expand only if the answer remains unclear. If it is genuinely unclear whether extra runs are worth the time or cost, ask the user before expanding the probe.
  6. When the question is benchmark-like or comparative, run in phases. Start with a high-signal pilot matrix against a control, then expand only the surviving candidates or unresolved cases.
  7. If the question is about a suspected regression or behavior change, add at least one known-good control case such as origin/main, the latest release, or the same request without the suspected option.
  8. For comparative probes, define parity before execution. Record prompt or input shape, tool-choice setup, model-settings parity, state reuse rules, and any response-shape constraint that keeps the comparison fair. If materially different output length could bias the result, record usage or token notes too.
  9. If the question asks whether one option has the same intelligence or quality as another, decide whether the matrix supports only example-pattern parity or a broader quality claim. For broader claims, add at least one harder or more open-ended case. Otherwise say explicitly that the result is limited to the covered patterns.
  10. Plan state controls before execution when hidden state could affect the result. Record whether each case uses fresh or reused state, how cache reuse or cache busting is handled, what unique IDs isolate repeated runs, and how cleanup is verified.
  11. If any live case will read environment variables, list the exact variable names and purpose for each case, then ask the user for approval before execution. Prefer request_user_input for this gate when it is available, with no auto-resolution and choices that grant or deny only this specific probe. Keep the approval ask short and include destination, read-only versus mutating or costly risk, exact variable names, and cleanup or rollback if relevant.
  12. Build task-specific probe scripts in a temporary location. Keep the script small, observable, and easy to discard.
  13. In openai-agents-python, make the runtime context explicit:
  • Run Python probes from the repository root with uv run python when practical.
  • Record the current commit, working directory, Python executable, and Python version.
  • Avoid accidental imports from a different checkout or site-packages location. If you must deviate from uv run python, say exactly why and what interpreter or environment was used instead.
  1. Present the complete probe proposal with the disclosures required above, including the exact command for each case or approved matrix, then ask the user for explicit approval and wait.
  2. Execute only the approved matrix and capture evidence. Record request shape, setup, observation summary, unexpected or negative result, error details, timing, runtime context, approved environment-variable names, repeat counts, warm-up handling, variance when relevant, cleanup behavior, and for comparisons note what was held constant plus any response-shape or usage notes that affect interpretation.
  3. Update the matrix with actual outcomes, not guesses.
  4. Keep temporary artifacts until the final response is drafted. Then delete them unless the user asked to keep them or they are needed for follow-up. Benchmark and repeat-heavy probes often need follow-up, so keeping artifacts is normal when the result may be revisited. If deleted, retain and report a short run summary.
  5. Report findings first, with unexpected or negative findings first. Then summarize how the validation was performed and which cases were covered.
  6. If the probe isolates one clear defect, you may include a short implementation hypothesis or minimal repro direction. Do not expand into a larger next-step plan unless the user asked for it.
Show full SKILL.md (606 more words)Show less

Validation Matrix

Use a matrix that makes the news easy to scan. Start from the runtime question and the observation summary, not just from expected and pass or fail.

Use a matrix with at least these columns:

  • case_id
  • scenario
  • mode
  • question
  • setup
  • observation_summary
  • result_flag
  • evidence

Add these columns when they materially improve the investigation:

  • comparison_basis
  • variable_under_test
  • held_constant
  • output_constraint
  • status
  • confidence
  • state_setup
  • repeats
  • warm_up
  • variance
  • usage_note
  • risk_profile
  • env_vars
  • approval
  • control

Treat result_flag as a fast scan field such as unexpected, negative, expected, or blocked. Use status only when there is a credible comparison basis, baseline, or documented contract to compare against.

Always consider whether the matrix should include these categories:

  • Baseline success.
  • Control or baseline comparison when a regression is suspected.
  • Boundary input or parameter variation.
  • Invalid or unsupported input.
  • Missing or incorrect configuration.
  • Transient external failure such as timeout, network interruption, or rate limiting.
  • Retry, idempotence, or cleanup behavior.
  • Concurrency or overlapping operations when shared state or ordering may matter.
  • Open-ended quality or intelligence samples when the question is broader than pattern parity.

Open validation-matrix.md when you need a stronger prioritization model or a reusable case template.

Temporary Probe Scripts

Write one-off scripts in a temporary file or temporary directory such as one created by mktemp -d or Python tempfile. Keep the script outside the repository by default, even when it imports code from the repository.

If the probe needs repository code:

  • Run it with the repository as the working directory, or
  • Set PYTHONPATH or the equivalent import path explicitly.
  • In openai-agents-python, prefer uv run python /tmp/probe.py from the repository root.

Design the probe to maximize observability:

  • Print or log the exact scenario being exercised.
  • Capture runtime context such as git SHA, working directory, Python executable and version, relevant package versions, model or deployment name, endpoint or base URL alias, and any retry or tool options that materially affect behavior.
  • For live probes, record only the names of environment variables that were approved for use. Never print their values.
  • Capture structured outputs when possible.
  • Preserve raw error type, message, and status code.
  • For repeat-sensitive cases, capture the attempt index, warm-up status, and any stable identifiers that help compare runs.
  • For repeated or benchmark-style probes, write both raw results and a compact summary artifact when practical.
  • Keep branching minimal so each script answers a narrow question.

Before deleting the temporary script or directory, keep a short run summary of the script path, command used, runtime context, and whether the evidence was kept or deleted.

Open python_probe.py when you want a lightweight disposable Python probe scaffold.

Reporting

Report in this order:

  1. Findings. Put unexpected or negative findings first. If there was no real news, say that explicitly.
  2. Validation approach. Summarize the code used, the runtime surface exercised, the execution modes, and the case matrix coverage.
  3. Case results. Include the matrix or a condensed version of it when the case count is large.
  4. Artifact status and brief run summary. State whether temporary artifacts were deleted or kept, and provide kept paths or the retained summary.
  5. Optional implementation note. Include this only when one clear defect was isolated and a short implementation direction would help.

For comparative probes, the report should also say what was held constant, what variable was under test, and whether the result supports only pattern parity or a broader quality claim.

Open reporting-format.md for the recommended response template.

Resources

© openai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (references) in .agents/skills/runtime-behavior-probe of openai/openai-agents-python.

  • SKILL.md
  • agents/openai.yaml
  • references/error-cases.md
  • references/openai-runtime-patterns.md
  • references/reporting-format.md
  • references/validation-matrix.md
  • templates/python_probe.py

Open the folder on GitHubat commit 26345c1

Compare with similar skills

Runtime Behavior Probe next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Runtime Behavior Probe compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Runtime Behavior Probe this skillopenai/openai-agents-python30k—~3.8kAutomated safety check: PassMIT
PR Design DocOpenHands/OpenHands90k—~2.4kAutomated safety check: PassMIT
Get API Docs with chubandrewyng/context-hub14k2 repos~775Automated safety check: PassMIT
Open Code Review CLIalibaba/open-code-review44k—~3.1kAutomated safety check: PassApache-2.0
Codexskills-directory/skill-codex1.5k3 repos~1.8kAutomated safety check: PassMIT
Changeset Validationopenai/openai-agents-js3.9k—~607Automated safety check: PassMIT

Similar skills

  • PR Design Doc

    OpenHands/OpenHands

    For a non-trivial pull request, write a self-contained HTML design doc under the temporary .pr/ directory and link a visibility-appropriate preview in the PR description, so maintainers grasp the…

    90k GitHub stars~2.4k tokensUpdated today
    DevelopmentAuto-check passed
  • Get API Docs with chub

    andrewyng/context-hub

    Fetches current documentation for third-party APIs and SDKs with the chub CLI before the agent writes code against them, instead of relying on remembered API shapes.

    14k GitHub starsUsed in 2 repos~775 tokens
    DevelopmentAuto-check passed
  • Open Code Review CLI

    alibaba/open-code-review

    Runs the ocr command-line tool to review Git changes, a commit or a branch comparison with an AI model, returning line-level comments and optionally applying fixes.

    44k GitHub stars~3.1k tokensUpdated 3 days ago
    DevelopmentAuto-check passed
  • Codex

    skills-directory/skill-codex

    A skill your agent uses when the user asks to run Codex CLI (codex exec, codex resume) or references OpenAI Codex for code analysis, refactoring, or automated editing

    1.5k GitHub starsUsed in 3 repos~1.8k tokens
    DevelopmentAuto-check passed
  • Changeset Validation

    openai/openai-agents-js

    Official

    Validate changesets in openai-agents-js using LLM judgment against git diffs (including uncommitted local changes).

    3.9k GitHub stars~607 tokensUpdated today
    DevelopmentAuto-check passed
  • DevSpace Manual QA Setup

    Waishnav/devspace

    Prepares the current DevSpace checkout or worktree for isolated local manual QA, covering QA state seeding, UI asset builds and snapshot resets.

    5.2k GitHub stars~440 tokensUpdated yesterday
    DevelopmentAuto-check passed

More from openai/openai-agents-python

All 14 skills in this repo
  • Implementation Final Review

    openai/openai-agents-python

    Official

    Review completed implementation changes before final verification.

    30k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Sensitive Logging Audit

    openai/openai-agents-python

    Official

    Audit or fix sensitive-data exposure in Python SDK diagnostics, exceptions, logging, and telemetry.

    30k GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Final Release Review

    openai/openai-agents-python

    Official

    Assess a Python SDK release candidate or release plan against the previous release and recommend ship or block.

    30k GitHub stars~5.4k tokensUpdated today
    Auto-check passed
  • Release Candidate Prep

    openai/openai-agents-python

    Official

    Prepare a local Python SDK release candidate in a dedicated worktree.

    30k GitHub stars~4.9k tokensUpdated today
    Auto-check passed
  • Code Change Verification

    openai/openai-agents-python

    Official

    Run the required final formatting, lint, type, and test checks after eligible SDK changes pass review.

    30k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Examples Run Analysis

    openai/openai-agents-python

    Official

    Analyze logs and source from a completed manual examples run.

    30k GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Runtime Behavior Probe

What does Runtime Behavior Probe do?

Plan controlled runtime probes when explicitly invoked; execute only after the required probe approval. Runtime Behavior Probe is an agent skill from openai/openai-agents-python, published by the product's own GitHub organization. Plan controlled runtime probes when explicitly invoked; execute only after the required probe approval.

When should I use Runtime Behavior Probe?

Runtime Behavior Probe fits situations like: development work in your project.

How do I install Runtime Behavior Probe in Claude Code?

Run `npx skills add openai/openai-agents-python --skill runtime-behavior-probe -a claude-code`. Or copy the skill folder (.agents/skills/runtime-behavior-probe in openai/openai-agents-python) into .claude/skills/runtime-behavior-probe in your project. Claude Code loads it when a task matches its description.

How do I install Runtime Behavior Probe in Codex?

Run `npx skills add openai/openai-agents-python --skill runtime-behavior-probe -a codex`. Or copy the skill folder (.agents/skills/runtime-behavior-probe in openai/openai-agents-python) into .agents/skills/runtime-behavior-probe in your project. Codex loads it when a task matches its description.

Can I use Runtime Behavior Probe in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add openai/openai-agents-python --skill runtime-behavior-probe -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/runtime-behavior-probe, .gemini/skills/runtime-behavior-probe, .github/skills/runtime-behavior-probe and .opencode/skills/runtime-behavior-probe in your project.

What does Runtime Behavior Probe need to run?

Going by SKILL.md and its folder, Runtime Behavior Probe needs Python for the scripts in its folder, the command-line tools its instructions call (uv) and credentials named OPENAI_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY.

Does Runtime Behavior Probe access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Runtime Behavior Probe safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Runtime Behavior Probe use?

Runtime Behavior Probe is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Runtime Behavior Probe use?

About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6k tokens, read only when the agent opens those files.

What are the alternatives to Runtime Behavior Probe?

Skills that share tags, products or a category with Runtime Behavior Probe: PR Design Doc (OpenHands/OpenHands, 90k stars), Get API Docs with chub (andrewyng/context-hub, 14k stars), Open Code Review CLI (alibaba/open-code-review, 44k stars) and Codex (skills-directory/skill-codex, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Runtime Behavior Probe?

openai (a GitHub organization, an official publisher) maintains it in openai/openai-agents-python, which has 29,896 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 8, 2026.

Source: openai/openai-agents-python on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.