Agent skill

Benchflow Traj Upload Ops

by benchflow-ai in benchflow-ai/benchflow

Operate, test, troubleshoot, and explain bench traj upload for public or trusted-direct trajectory contributions, including interactive and fully specified commands, dry runs, input validation…

Apache-2.0Auto-check passedDevelopment

Install Benchflow Traj Upload Ops

skills CLI
$ npx skills add benchflow-ai/benchflow --skill benchflow-traj-upload-ops -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install benchflow-ai/benchflow benchflow-traj-upload-ops --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/benchflow-ai/benchflow.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/benchflow-traj-upload-ops .claude/skills/benchflow-traj-upload-ops && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchflow-traj-upload-ops
GitHub stars
355
Token cost
~1.1k tokens
SKILL.md length
471 words
Files
3 (incl. references)
Skills in repo
8
Repo updated
First seen
Licence
Apache-2.0

At a glance

Operate, test, troubleshoot, and explain bench traj upload for public or trusted-direct trajectory contributions, including interactive and fully specified commands, dry runs, input validation…

  • Operator wants to test
  • SKILL.md covers Choose the matching CLI build, Run the requested flow, Inspect before uploading and Verify the result at the right…
  • Calls uv
  • Debug a trajectory upload

What it does

Benchflow Traj Upload Ops is an agent skill from benchflow-ai/benchflow. Operate, test, troubleshoot, and explain bench traj upload for public or trusted-direct trajectory contributions, including interactive and fully specified commands, dry runs, input validation, local secret masking, trajectory reports and previews, manifest metadata, upload progress, idempotency, and production promotion checks. Use this skill when a maintainer or operator wants to test, inspect, or debug a trajectory upload; validate a trajectory, report, or manifest; or verify the public upload path end to end…

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `evals/evals.json` and `references/cli-contract.md`).

It sits in Development. It works with Git. The repository describes itself as: Research infra for creating RL environments, post-training, and evals. The licence is Apache-2.0.

When your agent uses it

  • Operator wants to test
  • Debug a trajectory upload
  • Validate a trajectory
  • Verify the public upload path end to end

Example prompts

  • “/benchflow-traj-upload-ops”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit e965eee. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchflow Traj Upload Ops loads about 1.1k tokens when it runs, and up to ~5.9k if it reads all its reference files. Until then it costs about 159 tokens; SKILL.md has 471 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~159
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from benchflow-ai/benchflow at commit e965eee, republished under its Apache-2.0 licence (© benchflow-ai). 471 words, ~1,082 tokens.

Download SKILL.mdSave it as .claude/skills/benchflow-traj-upload-ops/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
benchflow-traj-upload-ops
description
Operate, test, troubleshoot, and explain `bench traj upload` for public or trusted-direct trajectory contributions, including interactive and fully specified commands, dry runs, input validation, local secret masking, trajectory reports and previews, manifest metadata, upload progress, idempotency, and production promotion checks. Use this skill when a maintainer or operator wants to test, inspect, or debug a trajectory upload; validate a trajectory, report, or manifest; or verify the public upload path end to end. For helping a contributor submit their own session, use `benchflow-traj-upload` instead.

Operate BenchFlow trajectory uploads

Use the public contribution route for ordinary users. Treat --direct as a trusted-operator escape hatch, not as an equivalent public-path test.

Read references/cli-contract.md before explaining non-default modes, troubleshooting a failure, reviewing a generated manifest, or claiming an end-to-end production result. The reference is the complete behavior contract for the current CLI.

Choose the matching CLI build

Use the released CLI when validating the published user experience:

bash
uv tool install --python 3.12 --upgrade benchflow
bench traj upload

Use the repository checkout when validating unreleased PR behavior:

bash
uv sync --extra dev --locked
uv run bench traj upload

Do not substitute the installed bench binary for uv run bench while testing unreleased code. Record bench --version or the exact Git SHA so the tested artifact is unambiguous.

Run the requested flow

Prefer the guided flow when a person wants to inspect and confirm the capture:

bash
bench traj upload

The CLI asks for the path, renders a report from a locally redacted staging copy, asks for missing GitHub and email metadata (after trying gh / git inference), and defaults the upload confirmation to No.

Use the fully specified form for scripts or an intentional no-prompt upload:

bash
bench traj upload <PATH> --github-id <GITHUB_ID> --email <EMAIL>

Providing all three required inputs makes the command non-interactive: it still renders the report, then uploads without a confirmation prompt. If any one is missing, the session is interactive and asks only for missing values before a final confirmation. Identity that resolves through gh / git inference also skips the confirmation; without a TTY, unresolved identity fails with the one-line --github-id / --email fallback instead of hanging on a prompt.

Use a dry run before a real upload when testing new files or CLI changes:

bash
bench traj upload <PATH> \
  --github-id <GITHUB_ID> \
  --email <EMAIL> \
  --dry-run

A dry run validates, redacts, reports, and creates the temporary manifest, but never prompts for confirmation and never makes a network request.

Show full SKILL.md (194 more words)Show less

Inspect before uploading

Verify these invariants in the rendered report:

  • Total steps = Thinking steps + Tool-call steps + Human steps.
  • Human steps are real user messages. Tool results, status or metadata records, empty records, and invented placeholders such as Assistant response are not trajectory steps.
  • Each preview row shows up to the first 100 words of a meaningful, already redacted step. --preview-steps accepts 0 through 20 and defaults to 5.
  • Every detected secret value is replaced locally with <XXX-benchflow-key-values-XXX>. Valid JSONL containing secrets is accepted after masking; the original source files remain unchanged.
  • File count, byte size, creation time, primary file, format, step counts, masked-value count, and preview are plausible for the selected capture.

Verify the result at the right boundary

For a dry run, report only local validation. For a real public upload, verify that the trusted validator promoted the digest to sources/community/<digest>/, with manifest.json present last and bound to the uploaded artifacts. A client success message or quarantine write alone is not production end-to-end proof.

Report whether the capture was uploaded, cancelled, already present, rejected, or only dry-run validated. Never expose contributor email, detected secret values, signed upload URLs, credentials, or internal service endpoints.

© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in .agents/skills/benchflow-traj-upload-ops of benchflow-ai/benchflow.

  • SKILL.md
  • evals/evals.json
  • references/cli-contract.md

Open the folder on GitHubat commit e965eee

Compare with similar skills

Benchflow Traj Upload Ops next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchflow Traj Upload Ops compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchflow Traj Upload Ops this skillbenchflow-ai/benchflow355—~1.1kAutomated safety check: PassApache-2.0
Finishing a Development Branchobra/superpowers297k5 repos~1.9kAutomated safety check: PassMIT
Code Review ChecklistshareAI-lab/learn-claude-code78k5 repos~1.1kAutomated safety check: PassMIT
Code Design Rationale Investigatorcursor/plugins10k9 repos~2.6kAutomated safety check: PassNone
Contributor-First PR MergeHKUDS/OpenHarness16k1 repos~847Automated safety check: PassMIT
Finishing A Development Branchfarm-fe/farm5.6k34 repos~1.8kAutomated safety check: PassMIT

Similar skills

  • Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.

    297k GitHub starsUsed in 5 repos~1.9k tokens
    DevelopmentAuto-check passed
  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 5 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Official

    Digs into why code is shaped the way it is by checking git history, pull requests and connected tools in parallel, then reporting a cited read on the tradeoffs.

    10k GitHub starsUsed in 9 repos~2.6k tokens
    DevelopmentAuto-check passed
  • Merges external GitHub pull requests while keeping the original author credited, and fixes conflicts after the merge instead of rewriting the contribution.

    16k GitHub starsUsed in 1 repo~847 tokens
    DevelopmentAuto-check passed
  • A skill your agent uses when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for…

    5.6k GitHub starsUsed in 34 repos~1.8k tokens
    DevelopmentAuto-check passed
  • Moves a package from another TryGhost repository into Ghost as an internal workspace package while keeping its Git history, with checkpoints for the steps that need an administrator.

    56k GitHub stars~3.8k tokensUpdated today
    DevelopmentAuto-check passed

More from benchflow-ai/benchflow

All 8 skills in this repo
  • Task Creator

    benchflow-ai/benchflow

    SkillsBench task authoring — walk a contributor from idea to submission-ready task following CONTRIBUTING.md and the task-implementation rubric.

    355 GitHub stars~4.5k tokensUpdated 3 days ago
    Auto-check passed
  • Task Review

    benchflow-ai/benchflow

    SkillsBench task PR review — classifies the task track (standard / research / multimodal), runs static policy checks against the track-specific rubric, benchmarks the task across oracle plus Claude…

    355 GitHub stars~4.5k tokensUpdated 3 days ago
    Auto-check: notes
  • Benchflow Experiment Review

    benchflow-ai/benchflow

    Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes.

    355 GitHub stars~4k tokensUpdated 3 days ago
    Auto-check passed
  • Benchflow

    benchflow-ai/benchflow

    Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.

    355 GitHub stars~1.9k tokensUpdated 3 days ago
    Auto-check: notes
  • Benchflow Traj Upload

    benchflow-ai/benchflow

    Find a local Claude Code or Codex session, open the BenchFlow trajectory viewer, and submit it after the user reviews it.

    355 GitHub stars~1.9k tokensUpdated 3 days ago
    Auto-check: notes
  • Code Specialist

    benchflow-ai/benchflow

    Delegate complex coding tasks to a specialist model. An agent skill from benchflow-ai/benchflow.

    355 GitHub stars~225 tokensUpdated 3 days ago
    Auto-check passed

Works with

Categories

Questions about Benchflow Traj Upload Ops

What does Benchflow Traj Upload Ops do?

Operate, test, troubleshoot, and explain bench traj upload for public or trusted-direct trajectory contributions, including interactive and fully specified commands, dry runs, input validation…. Benchflow Traj Upload Ops is an agent skill from benchflow-ai/benchflow. Operate, test, troubleshoot, and explain bench traj upload for public or trusted-direct trajectory contributions, including interactive and fully specified commands, dry runs, input validation, local secret masking, trajectory reports and previews, manifest metadata, upload progress, idempotency, and production promotion checks.

When should I use Benchflow Traj Upload Ops?

Benchflow Traj Upload Ops fits situations like: operator wants to test; debug a trajectory upload; validate a trajectory; verify the public upload path end to end.

How do I install Benchflow Traj Upload Ops in Claude Code?

Run `npx skills add benchflow-ai/benchflow --skill benchflow-traj-upload-ops -a claude-code`. Or copy the skill folder (.agents/skills/benchflow-traj-upload-ops in benchflow-ai/benchflow) into .claude/skills/benchflow-traj-upload-ops in your project. Claude Code loads it when a task matches its description.

How do I install Benchflow Traj Upload Ops in Codex?

Run `npx skills add benchflow-ai/benchflow --skill benchflow-traj-upload-ops -a codex`. Or copy the skill folder (.agents/skills/benchflow-traj-upload-ops in benchflow-ai/benchflow) into .agents/skills/benchflow-traj-upload-ops in your project. Codex loads it when a task matches its description.

Can I use Benchflow Traj Upload Ops in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/benchflow --skill benchflow-traj-upload-ops -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchflow-traj-upload-ops, .gemini/skills/benchflow-traj-upload-ops, .github/skills/benchflow-traj-upload-ops and .opencode/skills/benchflow-traj-upload-ops in your project.

What does Benchflow Traj Upload Ops need to run?

Going by SKILL.md and its folder, Benchflow Traj Upload Ops needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Benchflow Traj Upload Ops access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Benchflow Traj Upload Ops safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Benchflow Traj Upload Ops use?

Benchflow Traj Upload Ops is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchflow Traj Upload Ops use?

About 1.1k tokens (SKILL.md is roughly 4.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.8k tokens, read only when the agent opens those files.

What are the alternatives to Benchflow Traj Upload Ops?

Skills that share tags, products or a category with Benchflow Traj Upload Ops: Finishing a Development Branch (obra/superpowers, 297k stars), Code Review Checklist (shareAI-lab/learn-claude-code, 78k stars), Code Design Rationale Investigator (cursor/plugins, 10k stars) and Contributor-First PR Merge (HKUDS/OpenHarness, 16k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchflow Traj Upload Ops?

benchflow-ai (a GitHub organization) maintains it in benchflow-ai/benchflow, which has 355 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 6, 2026.

Source: benchflow-ai/benchflow on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.