Agent skill

Benchflow Traj Upload

by benchflow-ai in benchflow-ai/benchflow

Find a local Claude Code or Codex session, open the BenchFlow trajectory viewer, and submit it after the user reviews it.

Apache-2.0Auto-check: notes

Install Benchflow Traj Upload

skills CLI
$ npx skills add benchflow-ai/benchflow --skill benchflow-traj-upload -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install benchflow-ai/benchflow benchflow-traj-upload --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/benchflow-ai/benchflow.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/benchflow-traj-upload .claude/skills/benchflow-traj-upload && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchflow-traj-upload
GitHub stars
353
Token cost
~1.9k tokens
SKILL.md length
965 words
Files
2
Skills in repo
8
Repo updated
First seen
Licence
Apache-2.0

At a glance

Find a local Claude Code or Codex session, open the BenchFlow trajectory viewer, and submit it after the user reviews it.

  • Works in 7 steps: Setup → Discover → Pick → …
  • Someone pastes a BenchFlow eval prize line
  • SKILL.md covers Workflow, Step 1 — Setup, Step 2 — Discover and Step 3 — Pick, plus 4 more sections
  • Calls uv and opencode

What it does

Benchflow Traj Upload is an agent skill from benchflow-ai/benchflow. Find a local Claude Code or Codex session, open the BenchFlow trajectory viewer, and submit it after the user reviews it. Use this skill whenever someone pastes a BenchFlow eval prize line, wants to submit / share / contribute / upload a trajectory, set up traj upload, view a session, or pick a session to send. Also use it when they mention the eval prize, benchflow-traj-upload, or "copy this to your agent".

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `evals/evals.json`).

The repository describes itself as: Research infra for creating RL environments, post-training, and evals. The licence is Apache-2.0.

When your agent uses it

  • Someone pastes a BenchFlow eval prize line
  • Wants to submit / share / contribute / upload a trajectory
  • Set up traj upload
  • Pick a session to send

Example prompts

  • “copy this to your agent”
  • “/benchflow-traj-upload”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Bash

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Setup
  2. Discover
  3. Pick
  4. View
  5. Confirm
  6. Submit
  7. Persist the recipe

What it can do on your machine

Read from SKILL.md and the folder at commit e965eee. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • opencode

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchflow Traj Upload loads about 1.9k tokens when it runs. Until then it costs about 108 tokens; SKILL.md has 965 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~108
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from benchflow-ai/benchflow at commit e965eee, republished under its Apache-2.0 licence (© benchflow-ai). 965 words, ~1,883 tokens.

Download SKILL.mdSave it as .claude/skills/benchflow-traj-upload/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
benchflow-traj-upload
description
Find a local Claude Code or Codex session, open the BenchFlow trajectory viewer, and submit it after the user reviews it. Use this skill whenever someone pastes a BenchFlow eval prize line, wants to submit / share / contribute / upload a trajectory, set up traj upload, view a session, or pick a session to send. Also use it when they mention the eval prize, benchflow-traj-upload, or "copy this to your agent".
allowed-tools
Read, Bash
user-invocable
true

Submit a trajectory

The human copied a line into this chat so you would do the work. They should not run BenchFlow commands. You find a local session, open the viewer, wait until they like it, then you upload.

Do not print broker URLs, Azure blob URLs, or a "run this yourself" command. Those leak private inbox paths and turn a paste-to-agent flow back into a CLI.

For operating, testing, or debugging the upload pipeline itself (dry runs, manifests, promotion checks), use benchflow-traj-upload-ops instead.

Workflow

1. setup     → ensure the latest benchflow is installed
2. discover  → list recent local Claude / Codex / trial sessions
3. pick      → user chooses one (or confirms your recommendation)
4. view      → open the viewer with --confirm and give them the localhost URL
5. confirm   → wait for the Approve button (or their chat reply)
6. submit    → you upload; report Submitted / Already submitted + digest
7. persist   → if this repo has no local copy of this skill, write one

Step 1 — Setup

Always make sure the latest BenchFlow is installed before anything else — bench traj setup and the session-JSONL viewer only exist in 0.7.1+:

bash
uv tool install --python 3.12 --upgrade benchflow

If uv reports Executables already exist, rerun with --force. Verify that bench --version reports at least 0.7.1. If the installed CLI lacks bench traj setup or cannot open a session JSONL in the viewer, upgrade first rather than working around it.

Step 2 — Discover

Prefer the listing the CLI already knows how to make:

bash
bench traj setup --list

If that command is missing, search these locations and skip nested subagents/ files unless the user asks:

  • Claude Code: ~/.claude/projects/**/*.jsonl
  • Codex: ~/.codex/sessions/**/*.jsonl and ~/.codex/archived_sessions/*.jsonl
  • Cursor: agent transcripts at ~/.cursor/projects/*/agent-transcripts/**/*.jsonl
  • OpenCode: newer versions keep sessions in a SQLite database at ~/.local/share/opencode/opencode.db (run opencode db path to confirm); older versions used JSON files under ~/.local/share/opencode/storage/session/
  • BenchFlow trials: jobs/**/trajectory/ or a directory with turn*.txt

If the user described a time window or topic (for example, sessions from the last 72 hours on a specific project), prefer sessions matching that description.

Show the 8 most recent with mtime, path, and the first user-prompt snippet. Skip sessions that clearly contain private or proprietary work unless the user names them. If the user already named a file or folder, skip discovery.

Step 3 — Pick

Recommend one. Ask which to open if more than one is plausible. Do not upload yet — the viewer is how they decide the session is the one they meant.

Step 4 — View

First stage a dry run so you can show the user what upload-time redaction would mask for them (nothing is uploaded):

bash
bench traj upload /path/to/session.jsonl --dry-run

Its output ends with a plain Masked for you: ... line (for example Masked for you: 2 API keys, 1 bearer token, or Masked for you: nothing — no secrets detected). Extract the text after Masked for you: — call it the masking summary.

Then open the viewer with the in-page confirm bar and tell the user the URL, passing the masking summary so it renders next to the Approve button:

bash
bench eval view /path/to/session.jsonl --confirm --port 8889 \
  --redaction-summary "2 API keys, 1 bearer token"

The viewer shows the ORIGINAL session (it does not redact); the --redaction-summary note tells the reviewer what the upload step will mask. If the installed CLI rejects --redaction-summary (older than 0.7.2), drop the flag and state the masking summary in chat instead.

That path may also be a trial directory. If the port is taken, pick another.

With --confirm the page shows an Approve & submit / Not this one bar. When the user clicks, the server prints one line to stdout — DECISION: approved or DECISION: rejected — and exits (exit code 0 on approve, 3 on reject). Run the command so you can wait on that output: either start it in the background and poll its output for the DECISION: line, or run it blocking with a generous timeout.

If the installed CLI predates --confirm (bench --version below 0.7.2), run the plain bench eval view /path/to/session.jsonl in the background instead and rely on the chat confirmation in Step 5.

Show full SKILL.md (395 more words)Show less

Step 5 — Confirm

Ask them to review the page and click a button in the viewer:

  • DECISION: approved (exit 0) → they approved; proceed to Step 6.
  • DECISION: rejected (exit 3) → they want a different session; go back to pick.
  • If they answer in chat before clicking anything, kill the viewer process and honor the chat answer.
  • Without --confirm (older CLI), wait until they say in chat that it looks good.

Do not upload until one of those signals says the session is right.

The upload is tagged with the repository the session was about: the CLI reads the session's recorded working directory (only the session's own cwd — never the directory you run the upload from), resolves its git origin remote, and stores repo/<owner>/<name> as the source id (it prints Repo: owner/name (from session cwd /path; use --no-repo to omit) — the local path is terminal output only; run --dry-run to see the tag beforehand). Sessions recorded outside a git repo upload untagged. Mention the detected repo when you ask for confirmation — "This session will be tagged repo/owner/name; say the word if you want it omitted" — so they can opt out for private repos.

Before upload, remind them not to submit secrets, and repeat the masking summary from the Step 4 dry run when asking for approval — "Before upload, BenchFlow masks: 2 API keys, 1 bearer token; originals never leave this machine" — so they know exactly what redaction handles for them. The CLI masks detected secret values locally before anything leaves the machine (the server independently rescans and rejects any survivor), but redaction is a safety net, not a license to upload credentials.

Step 6 — Submit

You run the upload. The CLI infers GitHub username and email from gh / git; when identity resolves, the upload runs without prompts. If inference fails, ask for both in chat and pass --github-id / --email. Do not tell the user to re-run a command.

bash
bench traj upload /path/to/session.jsonl

If the user declined the repo tag in the confirm step, add --no-repo.

If the first request times out, run the same upload again. Retries are safe because the digest is content-addressed. Report Submitted or Already submitted, plus the digest.

Step 7 — Persist the recipe

If this project does not already have .agents/skills/benchflow-traj-upload/SKILL.md (or a .claude/skills / .codex/skills copy), write this skill there so the next agent does not rediscover the loop. Do not overwrite a newer local copy.

© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .agents/skills/benchflow-traj-upload of benchflow-ai/benchflow.

  • SKILL.md
  • evals/evals.json

Open the folder on GitHubat commit e965eee

Compare with similar skills

Benchflow Traj Upload next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchflow Traj Upload compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchflow Traj Upload this skillbenchflow-ai/benchflow353—~1.9kAutomated safety check: NotesApache-2.0
File Uploadsdavila7/claude-code-templates32k4 repos~238Automated safety check: PassMIT
File Upload Accessibilitythedaviddias/Front-End-Checklist74k—~418Automated safety check: PassMIT
Upload R2remotion-dev/remotion62k—~535Automated safety check: NotesCustom licence
Hunt File Uploadsickn33/agentic-awesome-skills47k1 repos~2.4kAutomated safety check: PassMIT
Upload Element Previewsremotion-dev/remotion62k—~265Automated safety check: PassCustom licence

Similar skills

  • File Uploads

    davila7/claude-code-templates

    Expert at handling file uploads and cloud storage. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 4 repos~238 tokens
    Backend & APIsAuto-check passed
  • File Upload Accessibility

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing templates, rendered HTML, or shared components related to Make file uploads accessible.

    74k GitHub stars~418 tokensUpdated 2 days ago
    Backend & APIsAuto-check passed
  • Upload R2

    remotion-dev/remotion

    Official

    Upload large Remotion repository assets to the Cloudflare R2 bucket behind remotion.media and replace local public/ assets with hosted URLs.

    62k GitHub stars~535 tokensUpdated today
    Media & CreativeAuto-check: notes
  • Hunt File Upload

    sickn33/agentic-awesome-skills

    Hunt file upload bugs

    47k GitHub starsUsed in 1 repo~2.4k tokens
    Backend & APIsAuto-check passed
  • Upload Element Previews

    remotion-dev/remotion

    Official

    Upload a Remotion Element's preview assets to remotion.media and replace local preview URLs.

    62k GitHub stars~265 tokensUpdated today
    Media & CreativeAuto-check passed
  • R2 Upload

    zebbern/claude-code-guide

    Upload files to Cloudflare R2, AWS S3, or any S3-compatible storage (like MinIO) and generate secure, time-limited presigned download links with configurable expiration, typically set to 5 minutes.

    4.7k GitHub starsUsed in 1 repo~886 tokens
    Backend & APIsAuto-check passed

More from benchflow-ai/benchflow

All 8 skills in this repo
  • Task Creator

    benchflow-ai/benchflow

    SkillsBench task authoring — walk a contributor from idea to submission-ready task following CONTRIBUTING.md and the task-implementation rubric.

    353 GitHub stars~4.5k tokensUpdated 2 days ago
    Auto-check passed
  • Task Review

    benchflow-ai/benchflow

    SkillsBench task PR review — classifies the task track (standard / research / multimodal), runs static policy checks against the track-specific rubric, benchmarks the task across oracle plus Claude…

    353 GitHub stars~4.5k tokensUpdated 2 days ago
    Auto-check: notes
  • Benchflow Experiment Review

    benchflow-ai/benchflow

    Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes.

    353 GitHub stars~4k tokensUpdated 2 days ago
    Auto-check passed
  • Benchflow

    benchflow-ai/benchflow

    Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.

    353 GitHub stars~1.9k tokensUpdated 2 days ago
    Auto-check: notes
  • Benchflow Traj Upload Ops

    benchflow-ai/benchflow

    Operate, test, troubleshoot, and explain bench traj upload for public or trusted-direct trajectory contributions, including interactive and fully specified commands, dry runs, input validation…

    353 GitHub stars~1.1k tokensUpdated 2 days ago
    Auto-check passed
  • Code Specialist

    benchflow-ai/benchflow

    Delegate complex coding tasks to a specialist model. An agent skill from benchflow-ai/benchflow.

    353 GitHub stars~225 tokensUpdated 2 days ago
    Auto-check passed

Questions about Benchflow Traj Upload

What does Benchflow Traj Upload do?

Find a local Claude Code or Codex session, open the BenchFlow trajectory viewer, and submit it after the user reviews it. Benchflow Traj Upload is an agent skill from benchflow-ai/benchflow. Find a local Claude Code or Codex session, open the BenchFlow trajectory viewer, and submit it after the user reviews it.

When should I use Benchflow Traj Upload?

Benchflow Traj Upload fits situations like: someone pastes a BenchFlow eval prize line; wants to submit / share / contribute / upload a trajectory; set up traj upload; pick a session to send.

How do I install Benchflow Traj Upload in Claude Code?

Run `npx skills add benchflow-ai/benchflow --skill benchflow-traj-upload -a claude-code`. Or copy the skill folder (.agents/skills/benchflow-traj-upload in benchflow-ai/benchflow) into .claude/skills/benchflow-traj-upload in your project. Claude Code loads it when a task matches its description.

How do I install Benchflow Traj Upload in Codex?

Run `npx skills add benchflow-ai/benchflow --skill benchflow-traj-upload -a codex`. Or copy the skill folder (.agents/skills/benchflow-traj-upload in benchflow-ai/benchflow) into .agents/skills/benchflow-traj-upload in your project. Codex loads it when a task matches its description.

Can I use Benchflow Traj Upload in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/benchflow --skill benchflow-traj-upload -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchflow-traj-upload, .gemini/skills/benchflow-traj-upload, .github/skills/benchflow-traj-upload and .opencode/skills/benchflow-traj-upload in your project.

What does Benchflow Traj Upload need to run?

Going by SKILL.md and its folder, Benchflow Traj Upload needs the command-line tools its instructions call (uv and opencode). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Bash.

Does Benchflow Traj Upload access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Benchflow Traj Upload safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Benchflow Traj Upload use?

Benchflow Traj Upload is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchflow Traj Upload use?

About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchflow Traj Upload?

Skills that share tags, products or a category with Benchflow Traj Upload: File Uploads (davila7/claude-code-templates, 32k stars), File Upload Accessibility (thedaviddias/Front-End-Checklist, 74k stars), Upload R2 (remotion-dev/remotion, 62k stars) and Hunt File Upload (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchflow Traj Upload?

benchflow-ai (a GitHub organization) maintains it in benchflow-ai/benchflow, which has 353 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 6, 2026.

Source: benchflow-ai/benchflow on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.