Agent skill

Verifying External Behavior

by kajisho5 in kajisho5/ffmpeg-skill

Confirms what a third-party library, remote API, build backend, or scraped document actually does before writing code that depends on it — throwaway probes that run in seconds, permissive clients…

MITAuto-check passed

Install Verifying External Behavior

skills CLI
$ npx skills add kajisho5/ffmpeg-skill --skill verifying-external-behavior -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install kajisho5/ffmpeg-skill verifying-external-behavior --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/kajisho5/ffmpeg-skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/verifying-external-behavior .claude/skills/verifying-external-behavior && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
verifying-external-behavior
GitHub stars
1.9k
Token cost
~2.7k tokens
SKILL.md length
1,139 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

Confirms what a third-party library, remote API, build backend, or scraped document actually does before writing code that depends on it — throwaway probes that run in seconds, permissive clients…

  • Integrating a new dependency
  • SKILL.md covers Probe the exact call you are…, Permissive clients don't…, Per-endpoint docs do not… and The response shape is part of…, plus 5 more sections
  • Calls uv, curl and docker
  • Writing a tolerated-status

What it does

Verifying External Behavior is an agent skill from kajisho5/ffmpeg-skill. Confirms what a third-party library, remote API, build backend, or scraped document actually does before writing code that depends on it — throwaway probes that run in seconds, permissive clients that forward wrong arguments instead of rejecting them, per-endpoint docs that don't generalize, response shapes that make a "fast path" always-false, fakes that encode your assumption rather than the service's behavior, and dry-runs that skip the step that fails. Use when integrating a new dependency or endpoint…

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The licence is MIT.

When your agent uses it

  • Integrating a new dependency
  • Writing a tolerated-status
  • Choosing a client argument name
  • Testing against a fake

Example prompts

  • “t generalize, response shapes that make a”
  • “Use the verifying-external-behavior skill to confirm what a third-party library, remote API, build backend, or scraped document actually does before…”
  • “/verifying-external-behavior”

Requirements

  • Python 3
  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit 008333a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • curl
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Verifying External Behavior loads about 2.7k tokens when it runs. Until then it costs about 173 tokens; SKILL.md has 1,139 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~173
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from kajisho5/ffmpeg-skill at commit 008333a, republished under its MIT licence (© kajisho5). 1,139 words, ~2,718 tokens.

Download SKILL.mdSave it as .claude/skills/verifying-external-behavior/SKILL.md (or your agent's skills folder).
name
verifying-external-behavior
description
Confirms what a third-party library, remote API, build backend, or scraped document actually does before writing code that depends on it — throwaway probes that run in seconds, permissive clients that forward wrong arguments instead of rejecting them, per-endpoint docs that don't generalize, response shapes that make a "fast path" always-false, fakes that encode your assumption rather than the service's behavior, and dry-runs that skip the step that fails. Use when integrating a new dependency or endpoint, writing a tolerated-status or error branch, choosing a client argument name, testing against a fake, or reviewing code that asserts an upstream contract.

Verifying External Behavior

Most integration bugs are not coding errors. They are a belief about someone else's system — a status code, a parameter name, a response shape, a build backend's scoping rule — that was never checked and turned out to be wrong. The code is written correctly against a contract that does not exist.

The fix is not more care. It is a probe: a throwaway command that asks the real system the exact question, before the code is written. Probes cost seconds. The bugs they prevent are silent, ship green, and are found months later.

Probe the exact call you are about to write

Not a similar call, not the documented example — the same endpoint form, the same library version, the same argument spelling.

bash
# A library's real defaults / return shapes, with no venv to create or clean up
uv run --no-project --with somelib python -c "import somelib; print(somelib.Thing())"

# The exact URL, including the collection-vs-item distinction
curl -s -o /dev/null -w '%{http_code}\n' https://api.example.com/v1/things
curl -s -o /dev/null -w '%{http_code}\n' https://api.example.com/v1/things/1

# What a build backend actually put in the artifact
uv build && unzip -l dist/*.whl

# A real service instead of a fake, for one minute
docker run --rm -p 6390:6379 redis:7-alpine

Two rules make probes worth the minute they cost:

  • Probe before you design around the answer. A probe that confirms a parameter tweak "should" fix a discrepancy is worth more than the tweak.
  • Paste the probe command and its output into the PR, comment, or test. A finding with no reproduction decays into folklore, and the next person re-derives it — or, worse, trusts it after it has gone stale.

Permissive clients don't reject wrong arguments — they ignore them

This is the highest-severity class, because the failure returns plausible data. Many HTTP client wrappers forward keyword arguments verbatim into the query string without validating them against their own documented parameter list. The remote service then drops the unknown parameter and applies its default — often "the authenticated user". A misspelled selector (owner= for owner_screen_name=, id_= for id=) does not raise; it silently returns someone else's records.

Never infer "the library would have rejected that" from the library's declared parameter list. Read the bytes that leave the process:

python
# Probe: intercept the transport and print the outbound request, then stop.
import requests

sent = {}
def spy(self, method, url, **kw):
    sent["url"], sent["params"] = url, kw.get("params")
    raise RuntimeError("probe: request intercepted")

requests.Session.request = spy
try:
    client.favorites(id_=12345)          # the call you were about to ship
except RuntimeError:
    pass
print(sent)      # {'url': '.../favorites/list.json', 'params': {'id_': 12345}}

If the parameter you passed is not in params under the name the service documents, the call is wrong no matter how healthy the response looks.

Two corollaries:

  • Pass every selector by keyword. Clients that bind positional arguments in a per-endpoint order will happily accept get_thing(owner, slug) and send slug as owner_id, while a neighbouring method with the same-looking signature is correct by luck.
  • Check that the parameter exists at all. Some endpoints have no equivalent of the selector you want. "Forward it under the right name" is not a fix when the right name does not exist — the feature has to be built differently.

Per-endpoint docs do not generalize across sibling endpoints

A status code documented for the single-item form frequently does not apply to the collection form of the same resource. Verified live against a large public REST API: GET /repos/{repo}/issues on a repository with issues disabled returns 200 with an empty array, while GET /repos/{repo}/issues/1 returns 410 Gone. Code written to "tolerate 410 when issues are disabled" therefore has a branch that never fires, and the real path — an empty 200 — falls through to whatever the generic handler does.

Probe the exact URL before writing a tolerated-status branch. If you keep a defensive branch for a status you could not reproduce, label it as defensive and name the path you did observe, so the next reader does not mistake it for verified behaviour.

python
# Observed: issues-disabled repos return 200 with []. The 410 branch is
# defensive — the item endpoint documents it, the list endpoint never sent it.
if resp.status_code == 410:
    return []

The response shape is part of the contract

List endpoints commonly embed a summary object — a handful of identity fields — rather than the full entity. A "we already have the data" fast path written against the full entity is then always false:

python
# This check is intended to skip a per-item fetch. Against summary objects that
# carry only {login, id, avatar_url}, it is False for every item — so the
# "fallback" enrichment fetch is the common path, and the loop is N+1.
if all(k in user for k in ("name", "company", "location", "followers")):
    return user
return fetch_user(user["login"])       # runs every time

Before optimizing around a response, print one real element and compare its keys to what your code reads. The same probe settles range and boundary assumptions that otherwise get "handled" defensively forever: if a per-year query provably returns exactly that year's days, the dedup pass guarding against adjacent-year leakage is dead code, and saying so in the PR is more valuable than the code.

Show full SKILL.md (507 more words)Show less

A fake proves your code calls the fake

Fakes are written by the same person as the code, from the same beliefs, so they agree with each other by construction. Stand the real thing up once and re-run the same assertions against it — a container is a minute, and it is the only thing that validates the semantics the fake asserts: TTL and expiry sentinels, whether a client factory is awaitable, key eviction, ordering, which exception type a failure raises.

Simulate the outage deterministically instead of mocking the error, so the code takes the same path production would:

python
# A dead port is a real, instant, deterministic connection failure.
client = redis.asyncio.from_url("redis://localhost:1")

The same reasoning applies to documents you parse. A hand-written fixture encodes your reading of the markup; the live page may render the same tokens across indented lines, so a regex requiring single spaces matches every fixture and never matches production. Capture one real sample, commit it, and point the parser's test at it. Prefer a machine-readable attribute (data-date="2024-01-01") over a human-readable string when the source offers both — it survives markup and locale changes that a prose regex does not.

Dry-runs skip the step that fails

Resolve-only and plan-only modes are not verification of anything the real run does after resolution. A dependency resolver's --dry-run reports success for a requirement whose presence makes the actual build hard-fail, because the dry run never builds. A build script's --check may never invoke the platform-specific tool that breaks.

Run the real operation once, on the platform that matters:

bash
uv pip install --dry-run .    # resolves; does NOT build → misses build-time errors
uv build                      # actually builds → catches them

More generally: if a mode exists specifically to be cheap, ask which step it bought that discount by skipping, and whether your bug lives there.

Some verified behaviour is not yours to fix

A probe sometimes proves the upstream system is simply wrong, or surprising, for your use case. Resist reaching for a configuration knob to make the number look right — if the underlying model or endpoint produces that output across parameter settings, tuning a parameter buries the finding instead of recording it. Write down what was observed, at which version, with the command, and choose a different approach.

Record findings with three things or they will not survive: the command, the observed output, and the date. External behaviour changes; an undated claim cannot be re-checked.

Checklist

Before shipping code that depends on an external system:
- [ ] The exact call/endpoint/version was probed, not a similar one
- [ ] Outbound request parameters inspected — names match what the service documents
- [ ] Selectors passed by keyword, never positionally
- [ ] Collection and item forms probed separately for status-code branches
- [ ] One real response element printed and compared against the fields the code reads
- [ ] Fakes validated against the real service at least once
- [ ] Failure paths exercised against a real failure (dead port, revoked token), not a mock
- [ ] Verified with the real build/install, not a dry-run
- [ ] Every tolerated-status or defensive branch is either reproduced or labelled defensive
- [ ] Findings recorded with command, output, and date

Note for this repository (ffmpeg-skill)

ffmpeg/ffprobe are exactly the "third-party system" this skill is about — their real behavior across versions and flags is the whole reason doctor exists, and this session repeatedly needed to probe ffmpeg directly rather than reason about it from documentation. The clearest example: cut.py's copy-mode -ss (before -i) seeks to the nearest preceding keyframe, and it was tempting to assume the output duration would still equal the requested -t regardless of where the seek landed. Running the actual command against a real fixture (--start 1.13 --end 5.71 --tolerance 2.0) showed a real 1.24s divergence between requested_duration and output_duration — the kind of finding this skill says to record with the command, the output, and the date, which is exactly how test_cut_copy_keyframe_snap_reports_a_real_nonzero_delta in tests/test_all.py was derived, rather than calculated from first principles.

Source: wdm0006/python-skills (MIT).

© kajisho5, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/verifying-external-behavior of kajisho5/ffmpeg-skill.

Open the folder on GitHubat commit 008333a

Compare with similar skills

Verifying External Behavior next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Verifying External Behavior compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Verifying External Behavior this skillkajisho5/ffmpeg-skill1.9k—~2.7kAutomated safety check: PassMIT
Third Party Cookiesthedaviddias/Front-End-Checklist74k—~580Automated safety check: PassMIT
Third Party Scriptsthedaviddias/Front-End-Checklist74k—~417Automated safety check: PassMIT
Audit Third Party Contractsben-manes/caffeine18k—~943Automated safety check: PassApache-2.0
Managing Third Party Vendor Riskmukul975/Anthropic-Cybersecurity-Skills34k—~2.2kAutomated safety check: PassApache-2.0
Third Party Codeopen-edge-platform/anomalib6.2k—~599Automated safety check: PassApache-2.0

Similar skills

  • Third Party Cookies

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing a website for privacy compliance, third-party resource loading, or cookie consent implementation.

    74k GitHub stars~580 tokensUpdated 2 days ago
    Legal & ComplianceAuto-check passed
  • Third Party Scripts

    thedaviddias/Front-End-Checklist

    A skill your agent uses when auditing slow page loads, heavy assets, or rendering delays related to Optimize third-party script loading.

    74k GitHub stars~417 tokensUpdated 2 days ago
    Frontend & DesignAuto-check passed
  • Audit Third Party Contracts

    ben-manes/caffeine

    Verify every third-party and sharp-edged JDK API usage against the contract the upstream documentation actually states

    18k GitHub stars~943 tokensUpdated 2 days ago
    Auto-check passed
  • Managing Third Party Vendor Risk

    mukul975/Anthropic-Cybersecurity-Skills

    Build and run a third-party/vendor risk management (TPRM) program aligned to NIST SP 800-161 C-SCRM: inventory and tier vendors, issue SIG/CAIQ questionnaires, review SOC 2/ISO 27001 evidence, set…

    34k GitHub stars~2.2k tokensUpdated 1 mo ago
    Legal & ComplianceAuto-check passed
  • Third Party Code

    open-edge-platform/anomalib

    Review/generate third-party code attribution, licensing, and notice requirements

    6.2k GitHub stars~599 tokensUpdated yesterday
    Auto-check passed
  • Verify

    asgeirtj/system_prompts_leaks

    Verify that a code change actually does what it's supposed to by exercising it end-to-end and observing behavior — drive the affected flow, not just tests or typecheck.

    69k GitHub stars~3k tokensUpdated today
    Testing & QAAuto-check passed

More from kajisho5/ffmpeg-skill

All 14 skills in this repo
  • Ffmpeg Skill

    kajisho5/ffmpeg-skill

    Edit video and audio with local FFmpeg from natural-language requests: cut, trim, join, resize/reframe (9:16, 1:1), speed change, captions and subtitles (SRT/ASS, animated, karaoke), logos and text…

    1.9k GitHub stars~7.4k tokensUpdated 3 days ago
    Auto-check passed
  • CI Pipeline Synthesizer

    kajisho5/ffmpeg-skill

    Generate GitHub Actions CI/CD pipeline configurations for automated building and testing of library and package projects.

    1.9k GitHub starsUsed in 1 repo~1.1k tokens
    Auto-check passed
  • Reviewing Ffmpeg Skill Changes

    kajisho5/ffmpeg-skill

    Review a change to the ffmpeg-skill repository for the failures its own contract makes possible — a claim in a result document that is true at one layer and false at the layer a caller reads, a new…

    1.9k GitHub stars~2.3k tokensUpdated 3 days ago
    Auto-check passed
  • Building Python MCP Servers

    kajisho5/ffmpeg-skill

    Builds robust Python MCP (Model Context Protocol) servers with FastMCP — tool design, error contracts, event-loop-safe blocking work, subprocess/CLI wrapping, single-file vs packaged distribution…

    1.9k GitHub stars~3.2k tokensUpdated 3 days ago
    Auto-check passed
  • Concurrent Branches

    kajisho5/ffmpeg-skill

    Resolve conflicts and merges when several branches are open against one repo at the same time — the hotspot files every change must touch (registry manifests, a single version field, shared tool…

    1.9k GitHub stars~2.8k tokensUpdated 3 days ago
    Auto-check passed
  • Guarding Destructive Operations

    kajisho5/ffmpeg-skill

    Add and review preconditions on operations that delete, overwrite, rewrite history, or resolve a caller-supplied name to a filesystem path — refusing instead of warning, placing the guard ahead of…

    1.9k GitHub stars~2.6k tokensUpdated 3 days ago
    Auto-check passed

Questions about Verifying External Behavior

What does Verifying External Behavior do?

Confirms what a third-party library, remote API, build backend, or scraped document actually does before writing code that depends on it — throwaway probes that run in seconds, permissive clients…. Verifying External Behavior is an agent skill from kajisho5/ffmpeg-skill. Confirms what a third-party library, remote API, build backend, or scraped document actually does before writing code that depends on it — throwaway probes that run in seconds, permissive clients that forward wrong arguments instead of rejecting them, per-endpoint docs that don't generalize, response shapes that make a "fast path" always-false, fakes that encode your assumption rather than the service's behavior, and dry-runs that skip the step that fails.

When should I use Verifying External Behavior?

Verifying External Behavior fits situations like: integrating a new dependency; writing a tolerated-status; choosing a client argument name; testing against a fake.

How do I install Verifying External Behavior in Claude Code?

Run `npx skills add kajisho5/ffmpeg-skill --skill verifying-external-behavior -a claude-code`. Or copy the skill folder (.claude/skills/verifying-external-behavior in kajisho5/ffmpeg-skill) into .claude/skills/verifying-external-behavior in your project. Claude Code loads it when a task matches its description.

How do I install Verifying External Behavior in Codex?

Run `npx skills add kajisho5/ffmpeg-skill --skill verifying-external-behavior -a codex`. Or copy the skill folder (.claude/skills/verifying-external-behavior in kajisho5/ffmpeg-skill) into .agents/skills/verifying-external-behavior in your project. Codex loads it when a task matches its description.

Can I use Verifying External Behavior in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add kajisho5/ffmpeg-skill --skill verifying-external-behavior -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/verifying-external-behavior, .gemini/skills/verifying-external-behavior, .github/skills/verifying-external-behavior and .opencode/skills/verifying-external-behavior in your project.

What does Verifying External Behavior need to run?

Going by SKILL.md and its folder, Verifying External Behavior needs the command-line tools its instructions call (uv, curl and docker). Our summary lists: Python 3; Docker.

Does Verifying External Behavior access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Verifying External Behavior safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Verifying External Behavior use?

Verifying External Behavior is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Verifying External Behavior use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Verifying External Behavior?

Skills that share tags, products or a category with Verifying External Behavior: Third Party Cookies (thedaviddias/Front-End-Checklist, 74k stars), Third Party Scripts (thedaviddias/Front-End-Checklist, 74k stars), Audit Third Party Contracts (ben-manes/caffeine, 18k stars) and Managing Third Party Vendor Risk (mukul975/Anthropic-Cybersecurity-Skills, 34k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Verifying External Behavior?

kajisho5 (a GitHub user) maintains it in kajisho5/ffmpeg-skill, which has 1,887 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 5, 2026.

Source: kajisho5/ffmpeg-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.