A skill your agent uses when designing or auditing the experimental program of a CVPR paper, covering benchmark and baseline selection under matched-compute fairness, the ablation study reviewers…

MITAuto-check passed

Install Cvpr Experiments

skills CLI
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill cvpr-experiments -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brycewang-stanford/Awesome-Journal-Skills cvpr-experiments --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/CVPR-Skills/skills/cvpr-experiments .claude/skills/cvpr-experiments && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cvpr-experiments
GitHub stars
1.2k
Token cost
~1.6k tokens
SKILL.md length
732 words
Files
1
Skills in repo
2,387
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when designing or auditing the experimental program of a CVPR paper, covering benchmark and baseline selection under matched-compute fairness, the ablation study reviewers…

  • Works in 5 steps: List the tables and figures the argument… → Price each in GPU-hours using a pilot… → Spend on order-of-information value: the… → …
  • Auditing the experimental program of a CVPR paper
  • SKILL.md covers Does it work: main comparisons, Why it works: the ablation is…, When it fails: qualitative… and What it costs: efficiency as a…, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Cvpr Experiments is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when designing or auditing the experimental program of a CVPR paper, covering benchmark and baseline selection under matched-compute fairness, the ablation study reviewers treat as mandatory, qualitative and failure-case evidence, efficiency metrics tied to the Compute Reporting Form, and generalization tests beyond a single dataset.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Journal-specific Claude Code/Codex skill packs covering mainstream journals — AER, QJE, Nature, Cell, 管理世界, 经济研究 & 200+ more — your fast track to getting published. | 覆盖主流期刊的… The licence is MIT.

When your agent uses it

  • Auditing the experimental program of a CVPR paper
  • Covering benchmark and baseline selection under matched-compute fairness
  • The ablation study reviewers treat as mandatory
  • Qualitative and failure-case evidence

Example prompts

  • “/cvpr-experiments”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. List the tables and figures the argument needs (main comparison, decisive ablation,
  2. Price each in GPU-hours using a pilot run; the totals routinely show the "nice to
  3. Spend on order-of-information value: the ablation that could falsify the core
  4. Reserve 15–20% of budget for rebuttal week; the most common January request is a
  5. Log everything into the recipe ledger (cvpr-reproducibility) as it runs; the CRF

What it can do on your machine

Read from SKILL.md and the folder at commit 932eb23. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cvpr Experiments loads about 1.6k tokens when it runs. Until then it costs about 89 tokens; SKILL.md has 732 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brycewang-stanford/Awesome-Journal-Skills at commit 932eb23, republished under its MIT licence (© brycewang-stanford). 732 words, ~1,630 tokens.

Download SKILL.mdSave it as .claude/skills/cvpr-experiments/SKILL.md (or your agent's skills folder).
name
cvpr-experiments
description
Use when designing or auditing the experimental program of a CVPR paper, covering benchmark and baseline selection under matched-compute fairness, the ablation study reviewers treat as mandatory, qualitative and failure-case evidence, efficiency metrics tied to the Compute Reporting Form, and generalization tests beyond a single dataset.

CVPR Experiments

CVPR runs on benchmark evidence: reviewers at the 2026 edition sorted 16,092 submissions largely by asking "do the tables prove the sentence?" This skill designs an experimental program that answers the four questions every vision review implicitly asks — does it work, why does it work, when does it fail, and what does it cost.

Does it work: main comparisons

  • Benchmarks: use the datasets your subfield's last two cycles used, current versions, standard splits. A new task may justify a new benchmark, but then the benchmark itself becomes a contribution needing validation (and, if claimed as one, public release by camera-ready — verified 2026 policy).
  • Baselines: the leaderboard's current top methods plus the strongest simple baseline. The comparison that kills papers in review is the one you omitted because it was too strong.
  • Fairness: match backbones, pretraining corpora, input resolution, and training schedule wherever possible — or tabulate the mismatch explicitly. Beating a ResNet-era method with a ViT-L and calling it method innovation is the single most common CVPR review objection.
  • Provenance: mark which numbers are re-runs under your protocol vs. quoted from papers; asterisk and footnote the re-runs.

Why it works: the ablation is not optional

Vision reviewers treat ablations as the paper's proof of understanding. Structure the grid so each row removes or replaces exactly one design decision:

text
# ablation-matrix.txt — one experiment per line, one variable per experiment
A0  full method                                  (reference row)
A1  - temporal attention  → per-frame baseline    tests the core claim
A2  - our loss   → standard L1                    is the loss or the architecture doing it?
A3  - pretrain   → from scratch                   how much rides on initialization?
A4  swap: our module in baseline X                does the gain transfer?
A5  sensitivity: key hyperparameter sweep         is A0 a lucky point?

Rows A4 (transplant) and A5 (sensitivity) separate memorable ablation sections from perfunctory ones. Every ablation row cited in prose belongs in the 8-page body; the long grid goes to the supplement.

When it fails: qualitative evidence with integrity

Cherry-picked grids convince nobody at a venue that invented the genre. The credible pattern:

Qualitative elementPurpose
Random (or id-listed) sample gridShows typical, not best-case, behavior
Side-by-side vs. two strongest baselinesSame inputs, aligned crops, labeled columns
Failure cases with a taxonomy"Fails under occlusion and low light" beats silence
Video for anything temporalIn the supplement — external links are banned

State the selection rule in the caption ("first 8 validation images", "random seed 0"). A stated rule converts pretty pictures into evidence.

What it costs: efficiency as a first-class axis

The 2026 cycle made compute visible venue-wide via the mandatory Compute Reporting Form (hardware + verification sections required; deeper compute sections optional but tied to recognition badges). Align the paper with the form: report params, FLOPs, latency (named hardware, batch size, resolution), and training GPU-hours for your method and re-run baselines where you can. "Real-time" with no hardware named contradicts your own CRF and reviewers can now check.

Show full SKILL.md (316 more words)Show less

Generalization: the second-dataset rule

A method shown on one dataset is a result about that dataset. Cheap robustness evidence reviewers reward: evaluate the trained model on a second domain without retuning; report cross-dataset transfer; if your field has corruption/shift suites, run them. One honest sentence about where transfer degrades is worth more than a defensive omission — it becomes your limitations paragraph (see cvpr-writing-style).

Budgeting the program before running it

At CVPR-scale training costs, experiment selection is a resource-allocation problem. Plan the program backward from the paper's skeleton:

  1. List the tables and figures the argument needs (main comparison, decisive ablation, efficiency table, qualitative grid, one transfer result).
  2. Price each in GPU-hours using a pilot run; the totals routinely show the "nice to have" grid costing more than the flagship.
  3. Spend on order-of-information value: the ablation that could falsify the core claim runs first, not last — discovering in October that A1 ≈ A0 should redirect the project, not decorate it.
  4. Reserve 15–20% of budget for rebuttal week; the most common January request is a small matched-setting re-run, and having compute standing by converts a weakness into a mini-table.
  5. Log everything into the recipe ledger (cvpr-reproducibility) as it runs; the CRF optional sections then fill themselves.

Statistical care, vision edition

Multi-seed everything is unaffordable at modern budgets; the defensible pattern is multi-seeded cheap experiments with mean ± std, a single flagged flagship run, and no headline claims resting on differences smaller than observed seed noise (full protocol in cvpr-reproducibility). For generative work, repeat evaluation sampling; for detection, fix and disclose the exact mAP implementation.

Reverify each cycle

  • Standard splits and evaluation-server rules for your benchmarks (some test sets are withheld; see cvpr-reproducibility for submission-server discipline).
  • CRF structure and any new compute/efficiency reporting requirements.
  • Benchmark leaderboard state the month before the deadline.
  • Any dataset-ethics or human-data review requirements added to the form (待核实 each cycle).

Output format

text
[Evidence audit] works / why / fails / costs — covered?
[Comparison risks] unmatched: <backbone/pretrain/resolution/schedule>
[Ablation] rows isolating single factors: <n>; transplant + sensitivity present?
[Qualitative] selection rule stated? failures taxonomized?
[Efficiency] params/FLOPs/latency/GPU-hours vs CRF: consistent?
[Priority additions] <ordered by review impact>

© brycewang-stanford, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in CVPR-Skills/skills/cvpr-experiments of brycewang-stanford/Awesome-Journal-Skills.

Open the folder on GitHubat commit 932eb23

Compare with similar skills

Cvpr Experiments next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cvpr Experiments compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cvpr Experiments this skillbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.6kAutomated safety check: PassMIT
Design Audit Against Rams' Principlesthedotmack/claude-mem98k—~4.6kAutomated safety check: PassApache-2.0
Experiment Designeralirezarezvani/claude-skills28k1 repos~783Automated safety check: PassMIT
Experimental Designaiming-lab/AutoResearchClaw15k—~286Automated safety check: PassMIT
Experiment Auditwanshuiyin/Auto-claude-code-research-in-sleep17k1 repos~2.7kAutomated safety check: NotesMIT
Experiment Auditwanshuiyin/Auto-claude-code-research-in-sleep17k—~3.2kAutomated safety check: NotesMIT

Similar skills

  • Audits a design against Dieter Rams' ten principles of good design, scores each with evidence, and hands off a make-plan prompt for a new, refined or redesigned outcome.

    98k GitHub stars~4.6k tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • Experiment Designer

    alirezarezvani/claude-skills

    A skill your agent uses when planning product experiments, writing testable hypotheses, estimating sample size, prioritizing tests, or interpreting A/B outcomes with practical statistical rigor.

    28k GitHub starsUsed in 1 repo~783 tokens
    Research & ScienceAuto-check passed
  • Experimental Design

    aiming-lab/AutoResearchClaw

    Best practices for designing reproducible ML experiments. An agent skill from aiming-lab/AutoResearchClaw.

    15k GitHub stars~286 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Experiment Audit

    wanshuiyin/Auto-claude-code-research-in-sleep

    Audit experiment integrity before claiming results. An agent skill from wanshuiyin/Auto-claude-code-research-in-sleep.

    17k GitHub starsUsed in 1 repo~2.7k tokens
    DatabasesAuto-check: notes
  • Experiment Audit

    wanshuiyin/Auto-claude-code-research-in-sleep

    Audit experiment integrity before claiming results. An agent skill from wanshuiyin/Auto-claude-code-research-in-sleep.

    17k GitHub stars~3.2k tokensUpdated yesterday
    DatabasesAuto-check: notes
  • OpenClaw Design Audit

    openclaw/clawhub

    Audits OpenClaw frontend code and rendered pages for token misuse, reimplemented primitives, accessibility and responsive defects and off-brand copy, with an evidence-based report.

    9.5k GitHub stars~498 tokensUpdated yesterday
    Frontend & DesignAuto-check passed

More from brycewang-stanford/Awesome-Journal-Skills

All 2,387 skills in this repo
  • Aaag Data Analysis

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running and reporting the analysis for an Annals of the American Association of Geographers manuscript — spatial statistics and modeling, remote-sensing accuracy, or…

    1.2k GitHub stars~1.3k tokensUpdated 11 days ago
    Auto-check passed
  • Aaag Literature Positioning

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when positioning an Annals of the American Association of Geographers manuscript in the literature — engaging geographic scholarship across the relevant area and the…

    1.2k GitHub stars~1.3k tokensUpdated 11 days ago
    Auto-check passed
  • Aaag Rebuttal

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when responding to an Annals of the American Association of Geographers decision letter (major/minor revision) — building a point-by-point response to the subject editor and…

    1.2k GitHub stars~1.4k tokensUpdated 11 days ago
    Auto-check passed
  • Aaag Research Design

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when defending the research design of an Annals of the American Association of Geographers manuscript — spatial/quantitative analysis and GIScience, remote-sensing and…

    1.2k GitHub stars~1.4k tokensUpdated 11 days ago
    Auto-check passed
  • Aaag Review Process

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when you need to understand how the Annals of the American Association of Geographers evaluates a manuscript — double-anonymous review routed through a subject editor by…

    1.2k GitHub stars~1.3k tokensUpdated 11 days ago
    Auto-check passed
  • Aaag Submission

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running the final pre-submission preflight for the Annals of the American Association of Geographers via ScholarOne Manuscripts — area/article-type selection…

    1.2k GitHub stars~1.6k tokensUpdated 11 days ago
    Auto-check passed

Questions about Cvpr Experiments

What does Cvpr Experiments do?

A skill your agent uses when designing or auditing the experimental program of a CVPR paper, covering benchmark and baseline selection under matched-compute fairness, the ablation study reviewers…. Cvpr Experiments is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when designing or auditing the experimental program of a CVPR paper, covering benchmark and baseline selection under matched-compute fairness, the ablation study reviewers treat as mandatory, qualitative and failure-case evidence, efficiency metrics tied to the Compute Reporting Form, and generalization tests beyond a single dataset.

When should I use Cvpr Experiments?

Cvpr Experiments fits situations like: auditing the experimental program of a CVPR paper; covering benchmark and baseline selection under matched-compute fairness; the ablation study reviewers treat as mandatory; qualitative and failure-case evidence.

How do I install Cvpr Experiments in Claude Code?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill cvpr-experiments -a claude-code`. Or copy the skill folder (CVPR-Skills/skills/cvpr-experiments in brycewang-stanford/Awesome-Journal-Skills) into .claude/skills/cvpr-experiments in your project. Claude Code loads it when a task matches its description.

How do I install Cvpr Experiments in Codex?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill cvpr-experiments -a codex`. Or copy the skill folder (CVPR-Skills/skills/cvpr-experiments in brycewang-stanford/Awesome-Journal-Skills) into .agents/skills/cvpr-experiments in your project. Codex loads it when a task matches its description.

Can I use Cvpr Experiments in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill cvpr-experiments -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cvpr-experiments, .gemini/skills/cvpr-experiments, .github/skills/cvpr-experiments and .opencode/skills/cvpr-experiments in your project.

What does Cvpr Experiments need to run?

SKILL.md names no scripts, command-line tools or credentials: Cvpr Experiments is instructions for the agent only.

Does Cvpr Experiments access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cvpr Experiments safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cvpr Experiments use?

Cvpr Experiments is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cvpr Experiments use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cvpr Experiments?

Skills that share tags, products or a category with Cvpr Experiments: Design Audit Against Rams' Principles (thedotmack/claude-mem, 98k stars), Experiment Designer (alirezarezvani/claude-skills, 28k stars), Experimental Design (aiming-lab/AutoResearchClaw, 15k stars) and Experiment Audit (wanshuiyin/Auto-claude-code-research-in-sleep, 17k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cvpr Experiments?

brycewang-stanford (a GitHub user) maintains it in brycewang-stanford/Awesome-Journal-Skills, which has 1,219 GitHub stars. The repository holds 2,387 skills in this directory. The repository was last updated on September 27, 2026.

Source: brycewang-stanford/Awesome-Journal-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.