A skill your agent uses when designing or auditing the evaluation of a UIST paper — choosing among technical benchmarks, controlled comparisons, usability walkthroughs, expert sessions, and…

MITAuto-check passedFrontend & Design

Install Uist Experiments

skills CLI
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill uist-experiments -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brycewang-stanford/Awesome-Journal-Skills uist-experiments --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/UIST-Skills/skills/uist-experiments .claude/skills/uist-experiments && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
uist-experiments
GitHub stars
1.2k
Token cost
~1.7k tokens
SKILL.md length
736 words
Files
1
Skills in repo
2,387
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when designing or auditing the evaluation of a UIST paper — choosing among technical benchmarks, controlled comparisons, usability walkthroughs, expert sessions, and…

  • Auditing the evaluation of a UIST paper — choosing among technical benchmarks
  • SKILL.md covers Match evaluation to claim type, Technical evaluation discipline, Baselines that respect the… and Study mechanics for the…, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Controlled comparisons

What it does

Uist Experiments is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when designing or auditing the evaluation of a UIST paper — choosing among technical benchmarks, controlled comparisons, usability walkthroughs, expert sessions, and demonstration applications, matching evaluation shape to the systems claim, and avoiding the ritual study that proves nothing the paper asserts.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Frontend & Design, covering UX design. The repository describes itself as: Journal-specific Claude Code/Codex skill packs covering mainstream journals — AER, QJE, Nature, Cell, 管理世界, 经济研究 & 200+ more — your fast track to getting published. | 覆盖主流期刊的… The licence is MIT.

When your agent uses it

  • Auditing the evaluation of a UIST paper — choosing among technical benchmarks
  • Controlled comparisons
  • Usability walkthroughs
  • Expert sessions

Example prompts

  • “/uist-experiments”

What it can do on your machine

Read from SKILL.md and the folder at commit 932eb23. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Uist Experiments loads about 1.7k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 736 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brycewang-stanford/Awesome-Journal-Skills at commit 932eb23, republished under its MIT licence (© brycewang-stanford). 736 words, ~1,665 tokens.

Download SKILL.mdSave it as .claude/skills/uist-experiments/SKILL.md (or your agent's skills folder).
name
uist-experiments
description
Use when designing or auditing the evaluation of a UIST paper — choosing among technical benchmarks, controlled comparisons, usability walkthroughs, expert sessions, and demonstration applications, matching evaluation shape to the systems claim, and avoiding the ritual study that proves nothing the paper asserts.

UIST Experiments

UIST does not have a single evaluation orthodoxy; it has a matching rule. The evaluation must measure the claim the artifact makes, and different artifact types make categorically different claims. The classic failure is the ritual study: a 12-participant SUS-and-task-time exercise bolted onto a toolkit paper whose actual claim — "this abstraction lets developers build a class of things" — no usability score can support.

Match evaluation to claim type

Artifact claimPrimary evidenceSupporting evidence
"This technique outperforms the status quo" (input, pointing, text entry)Controlled within-subjects comparison against the real incumbentLearning-curve data over sessions
"This hardware enables new sensing/actuation"Technical characterization: accuracy, range, SNR, latency across conditionsSmall usage session showing humans can operate it
"This toolkit lowers the floor / raises the ceiling"Demonstration portfolio + developer study or expert walkthroughCode-size/API comparisons, workshop deployments
"This pipeline makes X automatic"Quantitative accuracy on held-out real inputs + failure taxonomyEnd-to-end examples spanning input diversity
"This system supports task Y better"Task-based study instrumented around YLog analysis, think-aloud excerpts

Two smaller lanes are legitimate at UIST and often misused: demonstration-only evidence carries a paper when the enabling novelty is extreme (the demo is the result), and expert sessions (3-6 domain professionals in deep sessions) beat 20 novices when the system targets professionals.

Technical evaluation discipline

The technical evaluation is the half UIST reviewers read most skeptically, because they build things too:

  • Characterize across the operating envelope, not the sweet spot: distances, angles, lighting, surface materials, user anthropometrics — whatever the physics of your system cares about.
  • Report distributions, not means: median and 95th percentile latency; per-user accuracy spread. Interactive systems fail in the tails.
  • Measure with the interaction loop closed where possible; camera-to-photon latency and recognition-in-context differ from component benchmarks.
  • Include a failure section: conditions where recognition drops, actuation stalls, or tracking drifts. A measured failure boundary reads as competence.

Baselines that respect the reader

Compare against what a skilled practitioner would actually use today — the shipping technique, the standard library, the commercial device — not a strawman reimplementation. When no incumbent exists (genuinely new capability), say so and substitute ablations: which component of the pipeline buys which capability.

Study mechanics for the systems context

  • Counterbalance technique order; interaction studies carry strong carryover effects.
  • Power the comparison for the effect you claim; 12 participants detect large effects only, so do not claim small ones from them.
  • Report effect sizes with confidence intervals alongside any test statistic.
  • Instrument the system itself — logs are free evidence and reviewers trust event-level data over recollection.
  • IRB/ethics approval, compensation, and demographic reporting are expected whenever humans touch the system; plan approval months before the March deadline (see uist-workflow).
Show full SKILL.md (294 more words)Show less

Ablations for systems: which component buys which capability

Interactive systems are pipelines, and reviewers want the credit assignment:

  • Disable or downgrade each novel component and re-measure the end-to-end behavior; report the deltas in one table.
  • Substitute the naive version of each component (the off-the-shelf recognizer, the fixed calibration) to show the custom one earns its complexity.
  • Where components interact, test the two-way combinations that your claims depend on, not the full factorial — and say that is what you did.
  • Ablations also protect the rebuttal: "does the gain come from X or from Y?" is among the most common reviewer questions at this venue, and the answer must already exist by then (see uist-author-response).

The objections this venue actually raises

Recurring objectionPre-emption
"The comparison baseline is a strawman"Use the shipping incumbent; state its version and settings
"Latency/accuracy measured under ideal conditions"Report the envelope and the tails, not the demo-day sweet spot
"The study tasks were designed around the system's strengths"Include at least one task chosen from prior work's protocol
"n = 12 cannot support this claim"Match claim size to power; concede small effects
"The applications are all variations of one scenario"Diversify the portfolio across input contexts and user types
"No failure analysis"A measured failure boundary section, always

Pre-submission evidence audit

text
For each contribution claim C1..Cn in the introduction:
  [ ] name the evaluation section that tests it
  [ ] name the measure and the condition range
  [ ] name the comparison point (incumbent / ablation / none-with-justification)
  [ ] name where dispersion is reported (SD, CI, percentiles)
Orphan claims -> weaken the claim or add evidence
Orphan studies -> cut, or promote to a claim the intro actually makes

Run this audit against the video figure too: any capability the video shows that no section measures is a question a reviewer will ask in exactly those words (rebuttal implications in uist-author-response).

Reporting

Put the strongest number in the abstract with its condition attached. In the body, one summary table of the technical characterization beats four bar charts; per-cell detail goes to the appendix, raw harness and logs to the supplement (packaging rules in uist-artifact-evaluation and uist-reproducibility).

Output format

text
[Claim type] technique / hardware / toolkit / pipeline / task-support
[Evidence match] matched / partial mismatch / ritual-study warning
[Envelope coverage] <conditions characterized vs conditions claimed>
[Baseline honesty] real incumbent / ablation / missing
[Weakest link] <the claim most exposed in review + cheapest fix>

© brycewang-stanford, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in UIST-Skills/skills/uist-experiments of brycewang-stanford/Awesome-Journal-Skills.

Open the folder on GitHubat commit 932eb23

Compare with similar skills

Uist Experiments next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Uist Experiments compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Uist Experiments this skillbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.7kAutomated safety check: PassMIT
Impeccablebestofjs/bestofjs3.1k27 repos~2.6kAutomated safety check: PassMIT
Interface Design for Dashboards and Appsholaboss-ai/holaOS11k3 repos~6kAutomated safety check: PassMIT
Animategrowupanand/ConvoForm1026 repos~1.9kAutomated safety check: PassApache-2.0
Migrate Content Iadocker/docs4.7k—~5.1kAutomated safety check: PassApache-2.0
UX WalkthroughXiaoMi/hiui878—~1.3kAutomated safety check: PassMIT

Similar skills

  • Impeccable

    bestofjs/bestofjs

    A skill your agent uses when the user wants to design, redesign, shape, critique, audit, polish, clarify, distill, harden, optimize, adapt, animate, colorize, extract, or otherwise improve a…

    3.1k GitHub starsUsed in 27 repos~2.6k tokens
    Frontend & DesignAuto-check passed
  • Pushes an agent past generic defaults when designing dashboards, admin panels, SaaS apps and tools, with attention to structure, type, navigation and how data is shown.

    11k GitHub starsUsed in 3 repos~6k tokens
    Frontend & DesignAuto-check passed
  • Animate

    growupanand/ConvoForm

    Review a feature and enhance it with purposeful animations, micro-interactions, and motion effects that improve usability and delight.

    102 GitHub starsUsed in 6 repos~1.9k tokens
    Frontend & DesignAuto-check passed
  • Official

    Handle Hugo docs information-architecture moves: discover old vs new URLs, add front matter aliases (Phase 1), update in-repo links (Phase 2), interactive List 2 resolution and fragment validation…

    4.7k GitHub stars~5.1k tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • UX Walkthrough

    XiaoMi/hiui

    体验走查 skill。适用于代码库、URL、截图三种输入,输出结构化体验问题报告,并同步生成本地 docx 报告。触发词:体验走查、UX review、交互走查、界面审查、体验问题。

    878 GitHub stars~1.3k tokensUpdated 2 mo ago
    Frontend & DesignAuto-check passed
  • Color Audit

    rome-os/rome

    Audit a design system's color palette against measurable color-science disciplines — WCAG/APCA contrast of declared token pairs, perceptual (OKLCH) ramp uniformity, color-blindness safety of…

    737 GitHub stars~2.7k tokensUpdated today
    Frontend & DesignAuto-check passed

More from brycewang-stanford/Awesome-Journal-Skills

All 2,387 skills in this repo
  • Aaag Data Analysis

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running and reporting the analysis for an Annals of the American Association of Geographers manuscript — spatial statistics and modeling, remote-sensing accuracy, or…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Literature Positioning

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when positioning an Annals of the American Association of Geographers manuscript in the literature — engaging geographic scholarship across the relevant area and the…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Rebuttal

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when responding to an Annals of the American Association of Geographers decision letter (major/minor revision) — building a point-by-point response to the subject editor and…

    1.2k GitHub stars~1.4k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Research Design

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when defending the research design of an Annals of the American Association of Geographers manuscript — spatial/quantitative analysis and GIScience, remote-sensing and…

    1.2k GitHub stars~1.4k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Review Process

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when you need to understand how the Annals of the American Association of Geographers evaluates a manuscript — double-anonymous review routed through a subject editor by…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Submission

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running the final pre-submission preflight for the Annals of the American Association of Geographers via ScholarOne Manuscripts — area/article-type selection…

    1.2k GitHub stars~1.6k tokensUpdated 12 days ago
    Auto-check passed

Questions about Uist Experiments

What does Uist Experiments do?

A skill your agent uses when designing or auditing the evaluation of a UIST paper — choosing among technical benchmarks, controlled comparisons, usability walkthroughs, expert sessions, and…. Uist Experiments is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when designing or auditing the evaluation of a UIST paper — choosing among technical benchmarks, controlled comparisons, usability walkthroughs, expert sessions, and demonstration applications, matching evaluation shape to the systems claim, and avoiding the ritual study that proves nothing the paper asserts.

When should I use Uist Experiments?

Uist Experiments fits situations like: auditing the evaluation of a UIST paper — choosing among technical benchmarks; controlled comparisons; usability walkthroughs; expert sessions.

How do I install Uist Experiments in Claude Code?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill uist-experiments -a claude-code`. Or copy the skill folder (UIST-Skills/skills/uist-experiments in brycewang-stanford/Awesome-Journal-Skills) into .claude/skills/uist-experiments in your project. Claude Code loads it when a task matches its description.

How do I install Uist Experiments in Codex?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill uist-experiments -a codex`. Or copy the skill folder (UIST-Skills/skills/uist-experiments in brycewang-stanford/Awesome-Journal-Skills) into .agents/skills/uist-experiments in your project. Codex loads it when a task matches its description.

Can I use Uist Experiments in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill uist-experiments -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/uist-experiments, .gemini/skills/uist-experiments, .github/skills/uist-experiments and .opencode/skills/uist-experiments in your project.

What does Uist Experiments need to run?

SKILL.md names no scripts, command-line tools or credentials: Uist Experiments is instructions for the agent only.

Does Uist Experiments access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Uist Experiments safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Uist Experiments use?

Uist Experiments is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Uist Experiments use?

About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Uist Experiments?

Skills that share tags, products or a category with Uist Experiments: Impeccable (bestofjs/bestofjs, 3.1k stars), Interface Design for Dashboards and Apps (holaboss-ai/holaOS, 11k stars), Animate (growupanand/ConvoForm, 102 stars) and Migrate Content Ia (docker/docs, 4.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Uist Experiments?

brycewang-stanford (a GitHub user) maintains it in brycewang-stanford/Awesome-Journal-Skills, which has 1,228 GitHub stars. The repository holds 2,387 skills in this directory. The repository was last updated on September 27, 2026.

Source: brycewang-stanford/Awesome-Journal-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.