Ab Testing
coreyhaines31/marketingskills
When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.
Run an explicitly requested Zephyr control/treatment benchmark on the same pre-normalized sample and compare Finelog stage metrics.
$ npx skills add marin-community/marin --skill ab-test-zephyr -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install marin-community/marin ab-test-zephyr --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/marin-community/marin.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/ab-test-zephyr .claude/skills/ab-test-zephyr && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ab-test-zephyr" agent skill from https://github.com/marin-community/marin/tree/main/.agents/skills/ab-test-zephyr into .claude/skills/ab-test-zephyr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-test-zephyr", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/marin-community/marin/tree/main/.agents/skills/ab-test-zephyrType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add marin-community/marin --skill ab-test-zephyr -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install marin-community/marin ab-test-zephyr --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/marin-community/marin.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/ab-test-zephyr .agents/skills/ab-test-zephyr && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ab-test-zephyr" agent skill from https://github.com/marin-community/marin/tree/main/.agents/skills/ab-test-zephyr into .agents/skills/ab-test-zephyr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-test-zephyr", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add marin-community/marin --skill ab-test-zephyr -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install marin-community/marin ab-test-zephyr --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/marin-community/marin.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/ab-test-zephyr .cursor/skills/ab-test-zephyr && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ab-test-zephyr" agent skill from https://github.com/marin-community/marin/tree/main/.agents/skills/ab-test-zephyr into .cursor/skills/ab-test-zephyr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-test-zephyr", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/marin-community/marin.git --path .agents/skills/ab-test-zephyr--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add marin-community/marin --skill ab-test-zephyr -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install marin-community/marin ab-test-zephyr --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/marin-community/marin.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/ab-test-zephyr .gemini/skills/ab-test-zephyr && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ab-test-zephyr" agent skill from https://github.com/marin-community/marin/tree/main/.agents/skills/ab-test-zephyr into .gemini/skills/ab-test-zephyr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-test-zephyr", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install marin-community/marin ab-test-zephyrInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add marin-community/marin --skill ab-test-zephyr -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/marin-community/marin.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/ab-test-zephyr .github/skills/ab-test-zephyr && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ab-test-zephyr" agent skill from https://github.com/marin-community/marin/tree/main/.agents/skills/ab-test-zephyr into .github/skills/ab-test-zephyr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-test-zephyr", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add marin-community/marin --skill ab-test-zephyr -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install marin-community/marin ab-test-zephyr --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/marin-community/marin.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/ab-test-zephyr .opencode/skills/ab-test-zephyr && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ab-test-zephyr" agent skill from https://github.com/marin-community/marin/tree/main/.agents/skills/ab-test-zephyr into .opencode/skills/ab-test-zephyr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-test-zephyr", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ab-test-zephyrRun an explicitly requested Zephyr control/treatment benchmark on the same pre-normalized sample and compare Finelog stage metrics.
Ab Test Zephyr is an agent skill from marin-community/marin. Run an explicitly requested Zephyr control/treatment benchmark on the same pre-normalized sample and compare Finelog stage metrics.
Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Marketing & SEO, covering A/B testing. The repository describes itself as: Open-source framework for the research and development of foundation models. The licence is Apache-2.0.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 2ea1c1d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gituvrgFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git and uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ab Test Zephyr loads about 3.3k tokens when it runs. Until then it costs about 37 tokens; SKILL.md has 1,182 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from marin-community/marin at commit 2ea1c1d, republished under its Apache-2.0 licence (© marin-community). 1,182 words, ~3,269 tokens.
.claude/skills/ab-test-zephyr/SKILL.md (or your agent's skills folder).The coordinator writes one zephyr.stage row per completed stage and
execution_id. Use these fields:
| Field | Aggregation across executions | Interpretation |
|---|---|---|
cpu_time_total | sum | Primary efficiency and compute-cost signal |
elapsed | sum, labeled as summed stage elapsed | Secondary latency signal; sensitive to scheduling and stragglers |
items | sum | Workload-equivalence check |
bytes_processed | sum | Workload-equivalence check |
mem_peak_bytes_max | max | Worst observed shard RSS and OOM guardrail |
mem_bytes_avg | weighted interpretation only | Typical shard RSS context |
cpu_pct_avg | weighted interpretation only | CPU saturation context |
item_rate, byte_rate | do not aggregate | Derived from noisy elapsed time |
cpu_time_total sums process user and system CPU-seconds across completed
shards, so worker count and queue delay do not directly change it. Use it as the
primary efficiency signal and normalize per item or byte when accepted workload
sizes differ. elapsed measures the stage barrier and includes startup, I/O,
concurrency, queueing, and stragglers; repeat an elapsed-only result under
comparable scheduling conditions.
Keep CPU and elapsed time as separate outcomes:
Calibrate thresholds from same-code repeats for the selected sample and pool shape. A new OOM, application failure, or memory peak above the worker limit is a regression regardless of CPU.
Start at Collect execution IDs when control and treatment jobs already exist. Confirm that the control and every treatment used the same immutable sample, stage range, sources, worker resources, concurrency, parallelism, cluster, region, and priority.
A scheduled baseline is usable only when its report contains the same workload fingerprint and its Finelog execution IDs remain queryable. Otherwise, launch a matching control. Do not compare a standalone benchmark treatment with a differently shaped ferry baseline.
Default to one control at the branch/PR merge base and one treatment at the branch/PR head. Add treatments only when the requester explicitly names each additional commit or configuration. Record a stable name plus the exact SHA and configuration difference for every extra arm; do not infer or invent arms.
Run on GCP in europe-west4 with
gs://marin-eu-west4/datakit/sample_100b_8ae7a94f unless the requester
selects another sample or backend. The us-central1 GCS sample is available for
us-central1 runs. CoreWeave remains available for S3-local runs; select it
explicitly with the matching S3 sample and target cluster.
For a PR, read the diff and select the smallest stage range that exercises the changed behavior:
| Change | Minimum coverage |
|---|---|
| Stage-local map, serialization, or tokenization path | The affected stage on enough shards to amortize startup |
| Shuffle, partitioning, spill, merge, or buffer behavior | Exact or MinHash through fuzzy dedup on skewed or production-shaped data |
| Shared-pool lifecycle, scheduling, or pipeline concurrency | All affected stages with representative concurrent sources |
| Documentation, tests, types, or log text only | Skip the remote benchmark with reviewer agreement |
Confirm the sample size, pool shape, stage range, cluster, and expected cost before launching an expensive or production-scale comparison. Run local Zephyr and Datakit tests before paying for remote workers.
For a PR, use the merge base as the control and the PR head as the first treatment:
git fetch origin main
BASELINE_SHA=$(git merge-base origin/main HEAD)
TREATMENT_SHA=$(git rev-parse HEAD)
WORKTREE_ROOT=$(mktemp -d /tmp/zephyr-ab.XXXXXX)
git worktree add --detach "$WORKTREE_ROOT/control" "$BASELINE_SHA"
git worktree add --detach "$WORKTREE_ROOT/treatment" "$TREATMENT_SHA"Record both SHAs. Preserve configuration-only arms in separate worktrees or commits. Add arms only when explicitly requested.
experiments.datakit.zephyr_benchmark accepts an existing normalized sample
and routes outputs to a seven-day temporary prefix. Use an immutable,
region-local sample. Its default input is the GCS 100B sample in europe-west4.
All arguments except --run-tag must match across the control and treatments.
Set exactly one data-locality argument before launching:
# Default: GCS input and GCP compute in europe-west4.
SAMPLE_PREFIX=gs://marin-eu-west4/datakit/sample_100b_8ae7a94f
DATA_LOCALITY_ARGS=(--region europe-west4)
# GCP opt-in: use the existing us-central1 sample with us-central1 compute.
# SAMPLE_PREFIX=gs://marin-us-central1/datakit/sample_100b_8ae7a94f
# DATA_LOCALITY_ARGS=(--region us-central1)
# CoreWeave opt-in: S3 input and CoreWeave compute in cw-us-east-02a.
# SAMPLE_PREFIX=s3://marin-us-east-02a/marin/datakit/sample_100b_8ae7a94f
# DATA_LOCALITY_ARGS=(--target-cluster cw-us-east-02a)Set the cluster or region from the actual sample prefix. If the mapping is
unknown, stop before launching. The benchmark passes source_prefix to
marin_temp_bucket, which keeps temporary outputs with the sample. Do not
override the output location or launch compute in a different region.
Launch each arm from its worktree:
cd <CONTROL_OR_TREATMENT_WORKTREE>
uv run iris --config=lib/iris/config/marin.yaml job run --no-wait \
--job-name zephyr-ab-<RUN_TAG>-<ARM> \
"${DATA_LOCALITY_ARGS[@]}" --memory=2G --disk=5G --cpu=1 --extra=cpu \
--priority batch \
-- python -m experiments.datakit.zephyr_benchmark \
--sample-prefix "$SAMPLE_PREFIX" \
--sources <COMMA_SEPARATED_SOURCES_OR_ALL> \
--run-tag <FRESH_RUN_TAG>-<ARM> \
--pool-workers <WORKERS> \
--pool-cpu <CPU_PER_WORKER> \
--pool-ram <RAM_PER_WORKER> \
--pool-disk <DISK_PER_WORKER> \
--first-stage <exact|tokenize|minhash|fuzzy> \
--last-stage <exact|tokenize|minhash|fuzzy> \
--max-concurrent <PIPELINES> \
--dedup-max-parallelism <SHARDS>Record this workload fingerprint for every arm:
Use fresh run tags so no arm cache-hits. One matching control can be reused for explicitly requested treatments launched in the same scheduling window. If the decision depends on elapsed time, interleave additional control trials among the treatments to measure scheduling noise.
If the request includes continuous monitoring, use babysit-zephyr. A failed
or preempted arm measures infrastructure reliability and carries no performance
result. Use debug only for a stated repeated fault.
A benchmark job can run many Zephyr pipelines on one shared pool. Collect every
YYYYMMDD-HHMMSS-<hex> execution ID from the control and each treatment's root
job and descendant logs:
uv run iris --cluster marin job logs <IRIS_JOB_ID> \
--max-lines 200000 --no-tail --level info | \
rg -o '[0-9]{8}-[0-9]{6}-[0-9a-f]{8}' | sort -uSee lib/zephyr/OPS.md for child-job naming when a missing execution needs a
specific coordinator log. Preserve the control and treatment ID lists with the
workload fingerprint.
Authenticate with uv run iris --cluster marin login when needed. Query the
zephyr.stage namespace through the cluster's Finelog deployment:
uv run finelog query marin --format table '
SELECT execution_id, stage_name, status, cpu_time_total, elapsed,
items, bytes_processed, mem_peak_bytes_max, mem_bytes_avg, cpu_pct_avg
FROM "zephyr.stage"
WHERE execution_id IN (<CONTROL_AND_TREATMENT_IDS>)
ORDER BY execution_id, stage_name'Every expected row must have status = 'END'. A FAILED row invalidates that
arm. Missing rows usually mean an execution ID was omitted or Finelog emission
failed; resolve the gap before reporting a pass.
Aggregate all executions in the control and one treatment, then compare by
stage_name:
WITH tagged AS (
SELECT CASE
WHEN execution_id IN (<CONTROL_IDS>) THEN 'control'
WHEN execution_id IN (<TREATMENT_IDS>) THEN 'treatment'
END AS arm,
stage_name, cpu_time_total, elapsed, items, bytes_processed,
mem_peak_bytes_max
FROM "zephyr.stage"
WHERE status = 'END'
AND execution_id IN (<CONTROL_AND_TREATMENT_IDS>)
), aggregated AS (
SELECT arm, stage_name,
SUM(cpu_time_total) AS cpu_time_total,
SUM(elapsed) AS elapsed,
SUM(items) AS items,
SUM(bytes_processed) AS bytes_processed,
MAX(mem_peak_bytes_max) AS mem_peak_bytes_max
FROM tagged
GROUP BY arm, stage_name
)
SELECT b.stage_name,
b.cpu_time_total AS control_cpu,
t.cpu_time_total AS treatment_cpu,
(t.cpu_time_total - b.cpu_time_total) / NULLIF(b.cpu_time_total, 0) AS cpu_delta,
b.elapsed AS control_elapsed,
t.elapsed AS treatment_elapsed,
(t.elapsed - b.elapsed) / NULLIF(b.elapsed, 0) AS elapsed_delta,
t.items - b.items AS items_delta,
t.bytes_processed - b.bytes_processed AS bytes_delta,
b.mem_peak_bytes_max AS control_mem_peak,
t.mem_peak_bytes_max AS treatment_mem_peak
FROM aggregated b
JOIN aggregated t USING (stage_name)
WHERE b.arm = 'control' AND t.arm = 'treatment'
ORDER BY cpu_delta DESC;Use this SQL once per treatment, reusing the same control IDs. Replace each ID placeholder with comma-separated, single-quoted execution IDs. Keep each raw query output with the report. Keep repeated trials separate; do not merge different variants or unequal trial counts into one ID set.
Before interpreting deltas:
items and bytes_processed, within a
fraction of a percent. Explain and normalize any accepted mismatch.iris job describe <IRIS_JOB_ID> for OOMs and peak task memory.Different work, a failed stage, or material infrastructure churn makes the comparison inconclusive. Re-run before assigning a performance verdict.
For a PR, update one sentinel-marked comment so reruns do not accumulate stale verdicts:
<!-- zephyr-ab-test -->
🤖 ## Zephyr A/B test
Verdict: pass | regression | tradeoff | inconclusive
Workload: <sample, stage range, sources, pool shape, concurrency, cluster>
Control: <sha>, <job>, <execution count>
Treatments: <name, sha, job, and execution count for each>
| Treatment | Stage | CPU control | CPU treatment | CPU change | Elapsed control | Elapsed treatment | Elapsed change | Peak memory change |
|---|---|---:|---:|---:|---:|---:|---:|---:|
| ... | ... | ... | ... | ... | ... | ... | ... | ... |
Data check: <items and bytes comparison>
Infrastructure: <preemptions, retries, failures, stragglers, or none>
Interpretation: <efficiency result, latency result, and any tradeoff>Lead with CPU change, then elapsed time and memory. State whether elapsed came from one comparison or repeated interleaved trials and label summed stage elapsed. Launcher duration and task wall time do not replace stage metrics.
Remove temporary worktrees after preserving the SHAs, job IDs, execution IDs, workload fingerprints, and Finelog output. Repeat the treatment command for each additional worktree:
git worktree remove "$WORKTREE_ROOT/control"
git worktree remove "$WORKTREE_ROOT/treatment"Benchmark outputs expire under their seven-day temporary prefix.
babysit-zephyr monitors every control and treatment job through terminal
state.debug investigates repeated failures or unexplained infrastructure churn.lib/zephyr/OPS.md documents coordinator queries and straggler diagnosis.lib/iris/OPS.md documents job summaries, task attempts, and Finelog access.© marin-community, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/ab-test-zephyr of marin-community/marin.
Open the folder on GitHubat commit 2ea1c1d
Ab Test Zephyr next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ab Test Zephyr this skillmarin-community/marin | 3.9k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | |
| Ab Testingcoreyhaines31/marketingskills | 54k | 3 repos | ~3.1k | Automated safety check: Pass | MIT | |
| AnalyticsNexus-JPF/note-companion | 870 | 7 repos | ~2.2k | Automated safety check: Pass | MIT | |
| Ad Test Designeraaron-he-zhu/aaron-marketing-skills | 2.9k | 2 repos | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Ab Test Analyzeririnabuht12-oss/marketing-skills | 4.1k | — | ~1.4k | Automated safety check: Pass | None | |
| Ab Test Store Listingappeeky/aso-skills | 2.2k | — | ~1.8k | Automated safety check: Pass | MIT |
coreyhaines31/marketingskills
When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.
Nexus-JPF/note-companion
When the user wants to set up, improve, or audit analytics tracking and measurement.
aaron-he-zhu/aaron-marketing-skills
A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…
irinabuht12-oss/marketing-skills
Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.
appeeky/aso-skills
When the user wants to A/B test App Store product page elements to improve conversion rate.
freekmurze/dotfiles
When the user wants to plan, design, or implement an A/B test or experiment.
marin-community/marin
Deslop, simplify, or review low-value tests and prose only when explicitly requested for a branch or diff.
marin-community/marin
Use Iris to submit, inspect, debug, monitor, or recover jobs and tasks; diagnose scheduling and federation; deploy controllers; or reserve dev GPUs and TPUs.
marin-community/marin
Define, validate, submit, or restart a Marin SkyRL experiment through its artifact main.
marin-community/marin
Build, validate, publish, update, inspect, query, roll back, or archive a dynamic Marina applet.
marin-community/marin
Query Finelog logs and telemetry for Iris tasks, workers, profiles, training, vLLM, and cross-cluster forwarding.
marin-community/marin
Run a read-only preview for a specified Marin infra/pulumi stack and trace each pending resource change to merged pull requests since its latest successful update when that update records a clean…
Categories
Run an explicitly requested Zephyr control/treatment benchmark on the same pre-normalized sample and compare Finelog stage metrics. Ab Test Zephyr is an agent skill from marin-community/marin. Run an explicitly requested Zephyr control/treatment benchmark on the same pre-normalized sample and compare Finelog stage metrics.
Ab Test Zephyr fits situations like: tasks that involve A/B testing.
Run `npx skills add marin-community/marin --skill ab-test-zephyr -a claude-code`. Or copy the skill folder (.agents/skills/ab-test-zephyr in marin-community/marin) into .claude/skills/ab-test-zephyr in your project. Claude Code loads it when a task matches its description.
Run `npx skills add marin-community/marin --skill ab-test-zephyr -a codex`. Or copy the skill folder (.agents/skills/ab-test-zephyr in marin-community/marin) into .agents/skills/ab-test-zephyr in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add marin-community/marin --skill ab-test-zephyr -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ab-test-zephyr, .gemini/skills/ab-test-zephyr, .github/skills/ab-test-zephyr and .opencode/skills/ab-test-zephyr in your project.
Going by SKILL.md and its folder, Ab Test Zephyr needs the command-line tools its instructions call (git, uv and rg). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use git and uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Ab Test Zephyr is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Ab Test Zephyr: Ab Testing (coreyhaines31/marketingskills, 54k stars), Analytics (Nexus-JPF/note-companion, 870 stars), Ad Test Designer (aaron-he-zhu/aaron-marketing-skills, 2.9k stars) and Ab Test Analyzer (irinabuht12-oss/marketing-skills, 4.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
marin-community (a GitHub organization) maintains it in marin-community/marin, which has 3,925 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 11, 2026.
Source: marin-community/marin on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.