NIC Testing Patterns
nginx/kubernetes-ingress
Testing conventions for the NGINX Ingress Controller repo: Go table-driven tests, mandatory snapshot regeneration, Helm tests and Python pytest integration tests.
Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes.
$ npx skills add benchflow-ai/benchflow --skill benchflow-experiment-review -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install benchflow-ai/benchflow benchflow-experiment-review --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/benchflow-ai/benchflow.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/benchflow-experiment-review .claude/skills/benchflow-experiment-review && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "benchflow-experiment-review" agent skill from https://github.com/benchflow-ai/benchflow/tree/main/.agents/skills/benchflow-experiment-review into .claude/skills/benchflow-experiment-review/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchflow-experiment-review", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/benchflow-ai/benchflow/tree/main/.agents/skills/benchflow-experiment-reviewType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add benchflow-ai/benchflow --skill benchflow-experiment-review -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install benchflow-ai/benchflow benchflow-experiment-review --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benchflow-ai/benchflow.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/benchflow-experiment-review .agents/skills/benchflow-experiment-review && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "benchflow-experiment-review" agent skill from https://github.com/benchflow-ai/benchflow/tree/main/.agents/skills/benchflow-experiment-review into .agents/skills/benchflow-experiment-review/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchflow-experiment-review", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add benchflow-ai/benchflow --skill benchflow-experiment-review -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install benchflow-ai/benchflow benchflow-experiment-review --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benchflow-ai/benchflow.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/benchflow-experiment-review .cursor/skills/benchflow-experiment-review && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "benchflow-experiment-review" agent skill from https://github.com/benchflow-ai/benchflow/tree/main/.agents/skills/benchflow-experiment-review into .cursor/skills/benchflow-experiment-review/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchflow-experiment-review", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/benchflow-ai/benchflow.git --path .agents/skills/benchflow-experiment-review--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add benchflow-ai/benchflow --skill benchflow-experiment-review -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install benchflow-ai/benchflow benchflow-experiment-review --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benchflow-ai/benchflow.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/benchflow-experiment-review .gemini/skills/benchflow-experiment-review && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "benchflow-experiment-review" agent skill from https://github.com/benchflow-ai/benchflow/tree/main/.agents/skills/benchflow-experiment-review into .gemini/skills/benchflow-experiment-review/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchflow-experiment-review", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install benchflow-ai/benchflow benchflow-experiment-reviewInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add benchflow-ai/benchflow --skill benchflow-experiment-review -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/benchflow-ai/benchflow.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/benchflow-experiment-review .github/skills/benchflow-experiment-review && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "benchflow-experiment-review" agent skill from https://github.com/benchflow-ai/benchflow/tree/main/.agents/skills/benchflow-experiment-review into .github/skills/benchflow-experiment-review/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchflow-experiment-review", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add benchflow-ai/benchflow --skill benchflow-experiment-review -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install benchflow-ai/benchflow benchflow-experiment-review --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benchflow-ai/benchflow.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/benchflow-experiment-review .opencode/skills/benchflow-experiment-review && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "benchflow-experiment-review" agent skill from https://github.com/benchflow-ai/benchflow/tree/main/.agents/skills/benchflow-experiment-review into .opencode/skills/benchflow-experiment-review/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchflow-experiment-review", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
benchflow-experiment-reviewReview Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes.
Benchflow Experiment Review is an agent skill from benchflow-ai/benchflow. Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes. Use this skill whenever the user asks to audit traj health, failed or timed-out runs, healthy pass/fail/timeout status, no-skill leakage, skill loading, reward hacking, verifier isolation, metadata completeness, token usage, timing, Daytona-vs-Docker parity, path/root handling, coverage gaps, Docker/Daytona failures, or release-readiness of benchmark data.
Its SKILL.md is about 4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 41 other files, including scripts and reference files (for example `agents/openai.yaml`, `evals/evals.json` and `evals/files/clean-pass/result.json`).
It sits in Testing & QA, covering Integration testing, Test coverage and Feature launches and release readiness. It works with Docker. The repository describes itself as: Research infra for creating RL environments, post-training, and evals. The licence is Apache-2.0.
8 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit e965eee. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Benchflow Experiment Review loads about 4k tokens when it runs, and up to ~10k if it reads all its reference files. Until then it costs about 121 tokens; SKILL.md has 1,992 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from benchflow-ai/benchflow at commit e965eee, republished under its Apache-2.0 licence (© benchflow-ai). 1,992 words, ~4,036 tokens.
.claude/skills/benchflow-experiment-review/SKILL.md (or your agent's skills folder). This skill also uses 33 other files; get the full folder from GitHub.Use this skill to decide whether a Benchflow run trial is clean enough to publish or whether a Benchflow code change is safe to use for new experiments. The standard is: a clean sandbox with only task-needed resources, every agent behavior logged, no verifier leakage, and a final score or healthy failure.
This skill is intentionally harness-portable. Install or copy the entire
benchflow-experiment-review/ directory into the active harness's skill root,
keeping SKILL.md, scripts/, references/, evals/, and optional
agents/ metadata together. Use the same review procedure regardless of
whether the harness loads skills from .claude/skills, .codex/skills,
OpenHands/Gemini/pi-agent skill roots, or another compatible SKILL.md
directory.
Do not accept aggregate counts alone. Enumerate the intended matrix by
task_id, harness, model, skill mode, trial id, sandbox type, run root, and
source/ref. Mark each slot as healthy, missing, duplicate, stale, or unhealthy.
Healthy run outcomes are:
pass: agent completed and verifier produced a valid score.fail: agent completed incorrectly and verifier produced a valid score.normal_timeout: agent genuinely ran, timed out, and still produced a
complete ACP trajectory, complete LLM trajectory, and reward/scoring metadata.Infrastructure failures are not healthy failures. A stalled Docker daemon,
Daytona transport failure, missing trajectory/acp_trajectory.jsonl, missing
trajectory/llm_trajectory.jsonl, malformed or empty trajectory files, missing
reward, missing timing, missing token usage for new data, or verifier crash is
unhealthy until rerun or explicitly quarantined.
For every current BenchFlow model trial, both trajectory files plus the trainer-facing rollout row are mandatory:
trajectory/acp_trajectory.jsonl: ACP/tool trace with agent-side events.trajectory/llm_trajectory.jsonl: provider LLM request/response trace with
token usage evidence.results.jsonl: Verifiers / Prime-RL-shaped rollout row derived from the
healthy LLM trajectory.Do not treat ACP alone as sufficient. Do not treat llm_trajectory.jsonl alone
as sufficient. Do not treat aggregate result.json fields as a substitute for
either trajectory or results.jsonl. A trial with a scored reward, token
counts, or a plausible final answer is still unhealthy if either required
trajectory file or results.jsonl is missing, empty, truncated, unparsable, or
usage-only without recoverable request/response evidence.
One explicit evidence-only exception exists for native ACP subscription auth,
where provider HTTP capture is unavailable by construction. With
--allow-native-subscription-without-llm, the validator may accept a missing
LLM trajectory only when agent_result.usage_source == "agent_native_acp", ACP
events and native token usage are healthy, timing and reward are complete, and
results.jsonl marks the rollout completed but not training-ready with reason
missing_healthy_structured_llm_trajectory. This exception never creates
training data, never accepts an empty/malformed present LLM trajectory, and is
off by default.
results.jsonl must be reviewed as a training artifact, not just a sidecar. For
a healthy model rollout it should contain a parseable row with
info.training_ready == true, non-empty prompt, completion, and
trajectory, positive token usage, reward/score metadata, task identity, and
valid OpenAI-compatible message roles. Token usage may be split into
input/output fields or represented by a provider total_tokens value when the
provider does not expose a split. If assistant tool_calls appear, the row must
carry non-empty tool_defs or tools. Errored, verifier-failed, partial,
malformed, or missing-LLM rollouts must fail closed with
training_ready == false and an error/training-ready reason instead of being
accepted as SFT-ready.
Run the bundled validator before deeper manual review:
python .agents/skills/benchflow-experiment-review/scripts/validate_run_artifacts.py /path/to/rollout-or-jobs-root --jsonThe validator exits non-zero when any rollout is unhealthy. It is a deterministic fast path, not the whole audit: after it passes, continue with skill-loading, no-skill leakage, verifier-isolation, reward-hacking, and capability-attribution checks below. Oracle/reward-only lanes do not produce model LLM trajectories and must not be mixed into model-run publishability counts; if reviewed, report them as separate non-model evidence.
For each trial, first locate the authoritative run artifacts. Prefer repo validators or existing scripts when available, then inspect raw files. Required evidence includes:
trajectory/acp_trajectory.jsonl and
trajectory/llm_trajectory.jsonl, parseable from first event/model request
through final answer, failure, or timeout.results.jsonl, and enough provenance to connect the score and training row
to the exact trajectory.Review the trajectory for meaning, not just existence:
references/harness-skill-catalog-sop.md or run
scripts/extract_harness_skills.py against trajectory/llm_trajectory.jsonl.
Treat the script as a fast path: if it returns unknown, skill_count: 0,
or manual_review_required: true, fall back to the SOP's manual procedure
before concluding that no skills were available.task_skills_loading: with-skill trials should report 1, meaning
every expected task-specific skill was recovered in the startup skill catalog;
without-skill trials should report 0, meaning task-specific skills were not
injected into the agent's startup catalog.Reject or quarantine the trial if any of these appear:
trajectory/acp_trajectory.jsonl.trajectory/llm_trajectory.jsonl.results.jsonl.llm_trajectory.jsonl has no real provider request bodies, no provider
response bodies, or no provider token usage in response metadata.results.jsonl lacks a Prime-RL-compatible row, marks an unhealthy rollout as
training-ready, misses prompt/completion/trajectory/provider token usage, or
exposes assistant tool calls without tool definitions.Before using a changed Benchflow version for production experiments, run an end-to-end integration check that proves the changed code preserves artifact shape and sandbox behavior.
pytest, ty, and
ruff.acp_trajectory.jsonl and llm_trajectory.jsonl), rollout/job
results.jsonl, token usage, timing, tool usage, result status, verifier
output, and source provenance must be equivalent.references/verifier-hardening-checklist.md: network off by default (proxied
and allowlisted when needed), git history scrubbed past the base commit, no
answer/hidden-test/scorer files in agent-readable paths, empty patch fails
while golden patch passes, and agent test edits are reset before grading.The pass condition is schema and lifecycle parity between Daytona and Docker: both sandboxes should render the same task-needed resources, log the same classes of agent behavior, and emit complete metadata for the same logical trial outcome.
When a run fails or times out, classify the cause before accepting it as a healthy model failure.
Only the model-capability case can be counted as healthy fail or
normal_timeout. Experiment-fidelity failures must be rerun after fixing the
environment or documented as quarantined infra/context failures.
Docker stuck or daemon unstable:
Daytona vs VM Docker mismatch:
Path mismatch:
Verifier leakage or reward hacking:
references/reward-hacking-patterns.md and classify the evidence against the
solution-contamination, grader-gaming, and alignment-risk patterns there.references/verifier-hardening-checklist.md, which maps each pattern to the
environment/grader fix that closes it at the source.Skill loading mismatch:
references/harness-skill-catalog-sop.md.scripts/extract_harness_skills.py for a first pass; it scans the LLM
trajectory and falls back to sibling ACP system-prompt traces, but manual SOP
review remains authoritative when the output requests review.task_skills_loading separately from skill_count.
skill_count is the full recovered startup catalog size, while
task_skills_loading is only whether the task's own expected skills were
fully loaded. For with-skills runs expect task_skills_loading: 1; for
without-skills runs expect task_skills_loading: 0.pi-acp trajectories may expose no startup skill catalog; mark this
as catalog_not_serialized and rely on tool/file trace evidence.SKILL.md, .codex/skills, .agents/skills, invoke_skill,
activate_skill, Claude Code Skill calls, and ToolSearch selection.Only export, commit, or push healthy latest trials. Exclude partial runs, infra-failed trials, stale duplicates, local caches, and secrets.
Before reporting a trial as healthy or publishable, include the deterministic
trajectory-gate outcome from scripts/validate_run_artifacts.py or equivalent
manual evidence proving both required trajectory files are present and healthy.
Report findings in this order:
When reviewing an already completed run, end with a clear verdict: publishable, publishable with quarantines, or not publishable.
© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 33 other files (scripts, references) in .agents/skills/benchflow-experiment-review of benchflow-ai/benchflow.
Open the folder on GitHubat commit e965eee
Benchflow Experiment Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Benchflow Experiment Review this skillbenchflow-ai/benchflow | 356 | — | ~4k | Automated safety check: Pass | Apache-2.0 | |
| NIC Testing Patternsnginx/kubernetes-ingress | 5.1k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Suite ConverterMargin-Lab/evals | 160 | — | ~3.7k | Automated safety check: Pass | AGPL-3.0 | |
| Testcontainers Guide Migratordocker/docs | 4.7k | — | ~5k | Automated safety check: Pass | Apache-2.0 | |
| Go Redis Client Test Runnerredis/go-redis | 22k | — | ~786 | Automated safety check: Pass | BSD-2-Clause | |
| Running TestsNangoHQ/nango | 13k | — | ~876 | Automated safety check: Pass | Custom licence |
nginx/kubernetes-ingress
Testing conventions for the NGINX Ingress Controller repo: Go table-driven tests, mandatory snapshot regeneration, Helm tests and Python pytest integration tests.
Margin-Lab/evals
Converts test suites from external eval frameworks into the Margin Eval suite format.
docker/docs
Migrate a Testcontainers guide from testcontainers.com into the Docker docs site (docs.docker.com).
redis/go-redis
Explains how to run go-redis tests: the Docker Compose stack, make targets, focusing a single Ginkgo spec, the e2e suite and the version environment variables.
NangoHQ/nango
A skill your agent uses when running tests in the Nango monorepo - knows unit vs integration configs, vitest commands, Docker setup, and common test patterns
BlkLeg/CircuitBreaker
How Circuit Breaker is built, tested, packaged, and kept secret-safe — the make dev/verify/test targets, the PostgreSQL integration test database and its fixtures, the mono Docker image and native…
benchflow-ai/benchflow
SkillsBench task authoring — walk a contributor from idea to submission-ready task following CONTRIBUTING.md and the task-implementation rubric.
benchflow-ai/benchflow
SkillsBench task PR review — classifies the task track (standard / research / multimodal), runs static policy checks against the track-specific rubric, benchmarks the task across oracle plus Claude…
benchflow-ai/benchflow
Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.
benchflow-ai/benchflow
Operate, test, troubleshoot, and explain bench traj upload for public or trusted-direct trajectory contributions, including interactive and fully specified commands, dry runs, input validation…
benchflow-ai/benchflow
Find a local Claude Code or Codex session, open the BenchFlow trajectory viewer, and submit it after the user reviews it.
benchflow-ai/benchflow
Delegate complex coding tasks to a specialist model. An agent skill from benchflow-ai/benchflow.
Works with
Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes. Benchflow Experiment Review is an agent skill from benchflow-ai/benchflow. Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes.
Benchflow Experiment Review fits situations like: the user asks to audit traj health; healthy pass/fail/timeout status; no-skill leakage; verifier isolation.
Run `npx skills add benchflow-ai/benchflow --skill benchflow-experiment-review -a claude-code`. Or copy the skill folder (.agents/skills/benchflow-experiment-review in benchflow-ai/benchflow) into .claude/skills/benchflow-experiment-review in your project. Claude Code loads it when a task matches its description.
Run `npx skills add benchflow-ai/benchflow --skill benchflow-experiment-review -a codex`. Or copy the skill folder (.agents/skills/benchflow-experiment-review in benchflow-ai/benchflow) into .agents/skills/benchflow-experiment-review in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/benchflow --skill benchflow-experiment-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchflow-experiment-review, .gemini/skills/benchflow-experiment-review, .github/skills/benchflow-experiment-review and .opencode/skills/benchflow-experiment-review in your project.
Going by SKILL.md and its folder, Benchflow Experiment Review needs the command-line tools its instructions call (python). Our summary lists: Python 3; Docker.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Benchflow Experiment Review is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Benchflow Experiment Review: NIC Testing Patterns (nginx/kubernetes-ingress, 5.1k stars), Suite Converter (Margin-Lab/evals, 160 stars), Testcontainers Guide Migrator (docker/docs, 4.7k stars) and Go Redis Client Test Runner (redis/go-redis, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
benchflow-ai (a GitHub organization) maintains it in benchflow-ai/benchflow, which has 356 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 6, 2026.
Source: benchflow-ai/benchflow on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.