A skill your agent uses when designing or auditing IEEE PerCom empirical evaluations, covering real human subjects, leave-one-subject-out / cross-subject evaluation, F1 and event-level metrics on…

MITAuto-check passedDevOps & Cloud

Install Percom Experiments

skills CLI
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill percom-experiments -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brycewang-stanford/Awesome-Journal-Skills percom-experiments --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/PerCom-Skills/skills/percom-experiments .claude/skills/percom-experiments && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
percom-experiments
GitHub stars
1.2k
Token cost
~1.5k tokens
SKILL.md length
570 words
Files
1
Skills in repo
2,387
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when designing or auditing IEEE PerCom empirical evaluations, covering real human subjects, leave-one-subject-out / cross-subject evaluation, F1 and event-level metrics on…

  • Auditing IEEE PerCom empirical evaluations
  • SKILL.md covers Evaluation audit, Claim-to-evidence design table, Contamination- and… and Human-subjects provenance floor, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Covering real human subjects

What it does

Percom Experiments is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when designing or auditing IEEE PerCom empirical evaluations, covering real human subjects, leave-one-subject-out / cross-subject evaluation, F1 and event-level metrics on imbalanced activity classes, deployment realism (free-living vs. lab), fair baselines, contamination-aware model ablations, and matching evidence to the shape of each pervasive-computing claim.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud. The repository describes itself as: Journal-specific Claude Code/Codex skill packs covering mainstream journals — AER, QJE, Nature, Cell, 管理世界, 经济研究 & 200+ more — your fast track to getting published. | 覆盖主流期刊的… The licence is MIT.

When your agent uses it

  • Auditing IEEE PerCom empirical evaluations
  • Covering real human subjects
  • Leave-one-subject-out / cross-subject evaluation
  • F1 and event-level metrics on imbalanced activity classes

Example prompts

  • “/percom-experiments”

What it can do on your machine

Read from SKILL.md and the folder at commit 932eb23. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Percom Experiments loads about 1.5k tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 570 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brycewang-stanford/Awesome-Journal-Skills at commit 932eb23, republished under its MIT licence (© brycewang-stanford). 570 words, ~1,471 tokens.

Download SKILL.mdSave it as .claude/skills/percom-experiments/SKILL.md (or your agent's skills folder).
name
percom-experiments
description
Use when designing or auditing IEEE PerCom empirical evaluations, covering real human subjects, leave-one-subject-out / cross-subject evaluation, F1 and event-level metrics on imbalanced activity classes, deployment realism (free-living vs. lab), fair baselines, contamination-aware model ablations, and matching evidence to the shape of each pervasive-computing claim.

PerCom Experiments

Use this before submission when the evaluation is not yet locked. PerCom reviewers are ubicomp empiricists; the evaluation is where a sensing idea is won or lost, and — because the review is a single round with a bounded rebuttal — the evaluation must be complete at submission (you cannot add experiments in the rebuttal). The organizing principle is evidence proportional to the claim, tested on people and conditions a skeptic would accept.

Evaluation audit

  • Evaluate cross-subject by default. A recognition claim about users needs leave-one-subject-out (or leave-one-session-out) results, not a pooled split that lets the same person appear in train and test. Within-subject numbers are a supporting detail, never the headline.
  • Use real human subjects, described by count and relevant characteristics, with the collection protocol stated. Report how many, doing what, wearing/placed where.
  • Report the right metric. Human activity is imbalanced (most of a day is "null"), so F1 (macro and per-class), precision/recall, or event-level metrics tell the truth where raw accuracy flatters. Say whether metrics are frame-level or event-level.
  • Test deployment realism. Distinguish lab vs. free-living and scripted vs. spontaneous behavior; a result only shown on scripted in-lab data invites the "does this survive daily life?" objection.
  • Choose fair baselines, including the strongest prior method and a simple-but-reasonable alternative, tuned with a documented, equal budget. An untuned baseline is a scored weakness.
  • Report variance: confidence intervals across subjects/folds, number of runs, and the source of stochasticity. Per-subject distributions matter more than a single mean here.
  • Design limitations in, not on: know before you collect which generalization and construct limits the study will have (subject diversity, ground-truth quality), and instrument to bound them.

Claim-to-evidence design table

Ubicomp claimMatching evidenceReject pattern avoided
"Recognizes activity for new users"Leave-one-subject-out F1 with per-subject spread"Within-subject / pooled split inflates the number"
"Works in daily life"Free-living data, event-level metrics"Only scripted in-lab sessions tested"
"Beats the prior recognizer"Same data + tuned baseline, equal budget"Baseline untuned or on a different split"
"Handles class imbalance"Macro-F1 + per-class recall, stated balance"Raw accuracy hides the rare-class collapse"
"The model adds the value"Ablation vs. classical features/heuristics"Model's marginal contribution never isolated"
"Generalizes across contexts"Diverse subjects/environments + explicit limits"One population, claimed universal"
Show full SKILL.md (201 more words)Show less

Contamination- and leakage-aware evaluation

Sensing pipelines leak in subtle ways; the reviewer's first questions are about splits and leakage:

text
[Subject leakage]  never let one participant appear in both train and test -- LOSO prevents it
[Session/time leak] windows from one recording session can leak across a naive random split
[Normalization leak] fit scalers/PCA on train only; a global normalization leaks test statistics
[Pretraining]      if a foundation model is used, report whether test subjects/data could be in its
                   training set; prefer held-out or post-cutoff data
[Ablation]         isolate the model's marginal value against a classical-feature baseline

Human-subjects provenance floor

  • State subject count and relevant demographics, the collection protocol, and IRB/consent status.
  • Pin device models, firmware, sampling rates, and sensor placement; archive the extracted, de-identified dataset, not just a description.
  • Report the labeling protocol, who labeled, and inter-annotator agreement — silent label noise skews every downstream number.

Vignette: evaluating a HAR recognizer

Suppose the paper claims a wearable recognizer beats a prior model on daily activities. The matching plan: collect from a diverse participant set over multiple days of free-living; evaluate leave-one-subject-out; report macro-F1 and per-class recall with confidence intervals across subjects; run both models on the same folds with an equal, documented tuning budget; add an ablation against classical features; and state external validity (population, device) as a bounded limitation — every number traceable to a logged run in the artifact, because the rebuttal cannot add a run.

Reporting floor

  • Cross-subject metric (F1/event-level) with confidence intervals and per-subject spread for the headline comparison; say what the intervals represent.
  • Number of runs and the source of variance for any stochastic component.
  • The compute actually consumed, not vague feasibility language.

Output format

text
[Evaluation readiness] strong / adequate / weak (remember: no new experiments in the rebuttal)
[Claim -> evidence map] <claim: subjects / split (LOSO?) / metric (F1?) / setting (free-living?)>
[Baseline fairness] <baseline -> tuned? equal budget? same split? documented?>
[Leakage check] <subject / session / normalization / pretraining leakage handled? yes/no>
[Limitations-by-design] <generalization/construct limit -> instrumentation to bound it>
[Decision-critical run to finish before submission] <one experiment>

© brycewang-stanford, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in PerCom-Skills/skills/percom-experiments of brycewang-stanford/Awesome-Journal-Skills.

Open the folder on GitHubat commit 932eb23

Compare with similar skills

Percom Experiments next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Percom Experiments compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Percom Experiments this skillbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.5kAutomated safety check: PassMIT
Monitor CInrwl/nx29k5 repos~4.7kAutomated safety check: PassMIT
Terraform and OpenTofu Guideagentscope-ai/QwenPaw35k6 repos~4.2kAutomated safety check: PassApache-2.0
Vercel Optimize Auditvercel-labs/agent-skills32k8 repos~4.3kAutomated safety check: PassNone
Analyze GitHub Action Logswithastro/astro63k1 repos~1.3kAutomated safety check: PassCustom licence
Docs Learn PR Previewnetdata/netdata81k—~2kAutomated safety check: PassGPL-3.0

Similar skills

  • Monitor CI

    nrwl/nx

    Monitor Nx Cloud CI pipeline and handle self-healing fixes. An agent skill from nrwl/nx.

    29k GitHub starsUsed in 5 repos~4.7k tokens
    DevOps & CloudAuto-check passed
  • Terraform and OpenTofu Guide

    agentscope-ai/QwenPaw

    Guidance for writing and testing Terraform and OpenTofu code: module structure, naming, test approaches, CI/CD workflows, state handling and security scanning.

    35k GitHub starsUsed in 6 repos~4.2k tokens
    DevOps & CloudAuto-check passed
  • Vercel Optimize Audit

    vercel-labs/agent-skills

    Official

    Runs a metrics-first audit of a deployed Vercel project, gating investigations on real signals to produce ranked, citation-backed cost and performance recommendations.

    32k GitHub starsUsed in 8 repos~4.3k tokens
    DevOps & CloudAuto-check passed
  • Official

    Analyze recent GitHub Actions workflow runs to identify patterns, mistakes, and improvements.

    63k GitHub starsUsed in 1 repo~1.3k tokens
    DevOps & CloudAuto-check passed
  • Docs Learn PR Preview

    netdata/netdata

    Use only when the user explicitly asks to build, run, preview, inspect, or validate learn.netdata.cloud locally using the contents of a PR or documentation branch before merge.

    81k GitHub stars~2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Repo Mirror Sources

    netdata/netdata

    Inspect Netdata-org source checkouts under NETDATAREPOSDIR, or set up and synchronize that mirror when requested.

    81k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check: notes

More from brycewang-stanford/Awesome-Journal-Skills

All 2,387 skills in this repo
  • Aaag Data Analysis

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running and reporting the analysis for an Annals of the American Association of Geographers manuscript — spatial statistics and modeling, remote-sensing accuracy, or…

    1.2k GitHub stars~1.3k tokensUpdated 10 days ago
    Auto-check passed
  • Aaag Literature Positioning

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when positioning an Annals of the American Association of Geographers manuscript in the literature — engaging geographic scholarship across the relevant area and the…

    1.2k GitHub stars~1.3k tokensUpdated 10 days ago
    Auto-check passed
  • Aaag Rebuttal

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when responding to an Annals of the American Association of Geographers decision letter (major/minor revision) — building a point-by-point response to the subject editor and…

    1.2k GitHub stars~1.4k tokensUpdated 10 days ago
    Auto-check passed
  • Aaag Research Design

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when defending the research design of an Annals of the American Association of Geographers manuscript — spatial/quantitative analysis and GIScience, remote-sensing and…

    1.2k GitHub stars~1.4k tokensUpdated 10 days ago
    Auto-check passed
  • Aaag Review Process

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when you need to understand how the Annals of the American Association of Geographers evaluates a manuscript — double-anonymous review routed through a subject editor by…

    1.2k GitHub stars~1.3k tokensUpdated 10 days ago
    Auto-check passed
  • Aaag Submission

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running the final pre-submission preflight for the Annals of the American Association of Geographers via ScholarOne Manuscripts — area/article-type selection…

    1.2k GitHub stars~1.6k tokensUpdated 10 days ago
    Auto-check passed

Categories

Questions about Percom Experiments

What does Percom Experiments do?

A skill your agent uses when designing or auditing IEEE PerCom empirical evaluations, covering real human subjects, leave-one-subject-out / cross-subject evaluation, F1 and event-level metrics on…. Percom Experiments is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when designing or auditing IEEE PerCom empirical evaluations, covering real human subjects, leave-one-subject-out / cross-subject evaluation, F1 and event-level metrics on imbalanced activity classes, deployment realism (free-living vs.

When should I use Percom Experiments?

Percom Experiments fits situations like: auditing IEEE PerCom empirical evaluations; covering real human subjects; leave-one-subject-out / cross-subject evaluation; F1 and event-level metrics on imbalanced activity classes.

How do I install Percom Experiments in Claude Code?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill percom-experiments -a claude-code`. Or copy the skill folder (PerCom-Skills/skills/percom-experiments in brycewang-stanford/Awesome-Journal-Skills) into .claude/skills/percom-experiments in your project. Claude Code loads it when a task matches its description.

How do I install Percom Experiments in Codex?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill percom-experiments -a codex`. Or copy the skill folder (PerCom-Skills/skills/percom-experiments in brycewang-stanford/Awesome-Journal-Skills) into .agents/skills/percom-experiments in your project. Codex loads it when a task matches its description.

Can I use Percom Experiments in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill percom-experiments -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/percom-experiments, .gemini/skills/percom-experiments, .github/skills/percom-experiments and .opencode/skills/percom-experiments in your project.

What does Percom Experiments need to run?

SKILL.md names no scripts, command-line tools or credentials: Percom Experiments is instructions for the agent only.

Does Percom Experiments access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Percom Experiments safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Percom Experiments use?

Percom Experiments is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Percom Experiments use?

About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Percom Experiments?

Skills that share tags, products or a category with Percom Experiments: Monitor CI (nrwl/nx, 29k stars), Terraform and OpenTofu Guide (agentscope-ai/QwenPaw, 35k stars), Vercel Optimize Audit (vercel-labs/agent-skills, 32k stars) and Analyze GitHub Action Logs (withastro/astro, 63k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Percom Experiments?

brycewang-stanford (a GitHub user) maintains it in brycewang-stanford/Awesome-Journal-Skills, which has 1,216 GitHub stars. The repository holds 2,387 skills in this directory. The repository was last updated on September 27, 2026.

Source: brycewang-stanford/Awesome-Journal-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.