Swig Test
swig/swig
Run SWIG test suite for specific languages. An agent skill from swig/swig.
Reproduces and fixes flaky or quarantined tests. An agent skill from microsoft/aspire.
$ npx skills add microsoft/aspire --skill fix-flaky-test -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install microsoft/aspire fix-flaky-test --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/microsoft/aspire.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/fix-flaky-test .claude/skills/fix-flaky-test && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "fix-flaky-test" agent skill from https://github.com/microsoft/aspire/tree/main/.agents/skills/fix-flaky-test into .claude/skills/fix-flaky-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fix-flaky-test", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/microsoft/aspire/tree/main/.agents/skills/fix-flaky-testType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add microsoft/aspire --skill fix-flaky-test -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install microsoft/aspire fix-flaky-test --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/aspire.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/fix-flaky-test .agents/skills/fix-flaky-test && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "fix-flaky-test" agent skill from https://github.com/microsoft/aspire/tree/main/.agents/skills/fix-flaky-test into .agents/skills/fix-flaky-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fix-flaky-test", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/aspire --skill fix-flaky-test -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install microsoft/aspire fix-flaky-test --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/aspire.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/fix-flaky-test .cursor/skills/fix-flaky-test && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "fix-flaky-test" agent skill from https://github.com/microsoft/aspire/tree/main/.agents/skills/fix-flaky-test into .cursor/skills/fix-flaky-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fix-flaky-test", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/microsoft/aspire.git --path .agents/skills/fix-flaky-test--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add microsoft/aspire --skill fix-flaky-test -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install microsoft/aspire fix-flaky-test --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/aspire.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/fix-flaky-test .gemini/skills/fix-flaky-test && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "fix-flaky-test" agent skill from https://github.com/microsoft/aspire/tree/main/.agents/skills/fix-flaky-test into .gemini/skills/fix-flaky-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fix-flaky-test", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install microsoft/aspire fix-flaky-testInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add microsoft/aspire --skill fix-flaky-test -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/microsoft/aspire.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/fix-flaky-test .github/skills/fix-flaky-test && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "fix-flaky-test" agent skill from https://github.com/microsoft/aspire/tree/main/.agents/skills/fix-flaky-test into .github/skills/fix-flaky-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fix-flaky-test", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/aspire --skill fix-flaky-test -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install microsoft/aspire fix-flaky-test --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/aspire.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/fix-flaky-test .opencode/skills/fix-flaky-test && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "fix-flaky-test" agent skill from https://github.com/microsoft/aspire/tree/main/.agents/skills/fix-flaky-test into .opencode/skills/fix-flaky-test/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fix-flaky-test", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
fix-flaky-testReproduces and fixes flaky or quarantined tests. An agent skill from microsoft/aspire.
Fix Flaky Test is an agent skill from microsoft/aspire, published by the product's own GitHub organization. Reproduces and fixes flaky or quarantined tests. Tries local reproduction first (fast), then falls back to CI reproduce workflow (reproduce-flaky-tests.yml). Use this when asked to investigate, reproduce, debug, or fix a flaky test, a quarantined test, or an intermittently failing test.
Its SKILL.md is about 13k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `run-test-repeatedly.sh`).
It sits in Testing & QA, covering Failing and flaky tests. The repository describes itself as: Aspire is the tool for code-first, extensible, observable dev and deploy. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 809a672. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (PowerShell and Shell), which the agent can run.
Shell commands in SKILL.md call:
ghgitdotnetFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use gh and git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Fix Flaky Test loads about 13k tokens when it runs. Until then it costs about 76 tokens; SKILL.md has 4,853 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from microsoft/aspire at commit 809a672, republished under its MIT licence (© microsoft). 4,853 words, ~12,653 tokens.
.claude/skills/fix-flaky-test/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.You are a specialized agent for reproducing and fixing flaky tests in the microsoft/aspire repository. You try local reproduction first using run-test-repeatedly.sh (Linux/macOS) or run-test-repeatedly.ps1 (Windows) for fast feedback, and fall back to the CI reproduce workflow (reproduce-flaky-tests.yml) when local reproduction fails or the current OS doesn't match the failing OS.
Do NOT skip ahead to writing a code fix. Even if you think you already know the root cause, you MUST follow every step in order:
run-test-repeatedly.sh/.ps1 (fast path) ← try this FIRSTreproduce-flaky-tests.yml (graduated: single-test → quarantine-project → log-based)Each step has a checkpoint at the end. Do not proceed to the next step until the checkpoint is satisfied. Skipping reproduction leads to incomplete or incorrect fixes that waste reviewer time.
This skill uses two branches to keep investigation artifacts separate from the final clean fix:
main)<base-branch>-investigate (e.g., flaky-test0-investigate)ci.yml, configured reproduce-flaky-tests.yml, code fixflaky-test0)ci.yml enabled, reproduce-flaky-tests.yml at defaultsWhy two branches? Pushing workflow changes (disable ci.yml, configure reproduce workflow) to the same branch as the fix would trigger unwanted CI runs and pollute the final PR diff. The investigation branch isolates this.
Use SQL to track the overall investigation state. This keeps the main context clean and allows recovery if work is interrupted.
INSERT INTO todos (id, title, description, status) VALUES
('gather-data', 'Gather failure data', 'Read issue, find test code, determine failure rates per OS', 'pending'),
('analyze-existing', 'Analyze existing quarantine logs', 'Download logs from recent quarantine failures to understand the error', 'pending'),
('reproduce-local', 'Reproduce locally', 'Try local reproduction with run-test-repeatedly.sh/.ps1 (fast path)', 'pending'),
('reproduce-ci', 'Reproduce on CI', 'Configure and run reproduce-flaky-tests.yml: single-test first, then quarantine-project if needed', 'pending'),
('analyze', 'Analyze failure logs', 'Download CI logs or review local logs, identify root cause', 'pending'),
('fix', 'Apply fix', 'Write the code fix based on root cause analysis', 'pending'),
('verify', 'Verify fix on CI', 'Re-run reproduce workflow to confirm fix works', 'pending'),
('verify-ci', 'Verify no CI regressions', 'Confirm fix does not introduce regressions in the CI workflow', 'pending'),
('cleanup', 'Clean up investigation', 'Close investigation PR, create clean fix PR', 'pending');
INSERT INTO todo_deps (todo_id, depends_on) VALUES
('analyze-existing', 'gather-data'),
('reproduce-local', 'analyze-existing'),
('reproduce-ci', 'reproduce-local'),
('analyze', 'reproduce-ci'),
('fix', 'analyze'),
('verify', 'fix'),
('verify-ci', 'verify'),
('cleanup', 'verify-ci');CREATE TABLE IF NOT EXISTS session_state (key TEXT PRIMARY KEY, value TEXT);
INSERT OR REPLACE INTO session_state (key, value) VALUES
('test_method', '<FullyQualifiedMethodName>'),
('test_project', '<ProjectShortname>'),
('issue_url', '<GitHubIssueURL>'),
('failure_rate_linux', '<rate or unknown>'),
('failure_rate_windows', '<rate or unknown>'),
('failure_rate_macos', '<rate or unknown>'),
('max_failure_rate', '<highest rate across OSes>'),
('reproduce_attempt', '1'),
('reproduce_mode', 'single-test'),
('fix_attempt', '1'),
('reproduce_run_id', ''),
('verify_run_id', ''),
('investigation_branch', ''),
('fix_branch', ''),
('user_interaction', 'false');Always update todo status as you work — set to in_progress before starting, done when complete. Query SELECT * FROM todos; to check progress. Store CI run IDs and attempt counts in session_state.
If at any point during the investigation you use the ask_user tool to get input from the user, immediately update the session state:
INSERT OR REPLACE INTO session_state (key, value) VALUES ('user_interaction', 'true');This flag determines whether the final PR is labeled as [automated] (see Step 6.2).
Keep investigation notes in the session workspace (not in the repo). This avoids commit noise from temporary artifacts:
~/.copilot/session-state/<session-id>/
├── plan.md # Summary: test name, issue, root cause, fix, status
└── files/
└── failure-logs/ # Downloaded CI failure logs (if any)Use plan.md in the session workspace for running notes and observations. Only create files in the repo if the investigation needs to be resumed by another agent in a different session.
The steps below are sequential and gated. Complete each step fully before moving to the next.
run-test-repeatedly.sh (Linux/macOS) or run-test-repeatedly.ps1 (Windows) — this is the fast path (~minutes vs ~30 min for CI). Works when the current OS matches a failing OS.reproduce-flaky-tests.yml with graduated escalation: single-test → quarantine-project → log-based analysisrun-test-repeatedly.sh/.ps1, then always validate on CI as final verification.Prefer analyzing existing data first. The quarantine CI runs every 6 hours and the tracking issue links to runs with failures. These logs are often sufficient to diagnose the root cause, but CI reproduction should still be attempted to establish a baseline failure rate.
The user may provide:
DeployAsync_WithMultipleComputeEnvironments_Works)https://github.com/microsoft/aspire/issues/13287)If you only have the test name, find the tracking issue:
First check the test code for a [QuarantinedTest] attribute — it contains the issue URL:
grep -rn "QuarantinedTest" tests/ --include="*.cs" | grep "TestMethodName"If not found there, look up the test in the quarantine tracking meta-issue https://github.com/microsoft/aspire/issues/8813 — this issue tracks all quarantined tests with links to their individual issues:
gh issue view 8813 --repo microsoft/aspireSearch the output for the test name to find its linked issue.
If neither source has the issue, proceed without historical failure data. Use a default configuration (all 3 OSes, 5×5 iterations) since you don't know which OSes fail or the failure rate.
Quarantined test issues contain tracking tables with per-OS failure rates over the last 100 runs. This data is critical:
# Read the issue to get failure data
gh issue view <issue-number> --repo microsoft/aspireFind the test method, class, and project. Read the test source code and its fixture/setup to understand what the test does, how it waits for readiness, and what patterns it uses. This is essential for understanding what you're trying to reproduce and for matching against common flaky test patterns.
# Search for the test method
grep -rn "public.*async.*Task.*TestMethodName\|public.*void.*TestMethodName" tests/ --include="*.cs"Consult the flaky test patterns in .github/instructions/test-review-guidelines.instructions.md early. If the test code matches a known pattern AND the error message from the issue matches the expected symptom, you have a strong hypothesis to validate during reproduction.
Based on the failure rate from the issue tracking data, calculate iterations to achieve 95% probability of seeing at least one failure (if the bug exists):
| Failure Rate | Runners × Iterations per OS | Total per OS | Confidence |
|---|---|---|---|
| >50% | 3 × 3 | 9 | >99% |
| 20-50% | 5 × 5 | 25 | >99% |
| 10-20% | 5 × 10 | 50 | >99% |
| 5-10% | 10 × 10 | 100 | >99% |
| <5% | 10 × 25 | 250 | >95% |
The math: for failure rate p, need n ≥ log(0.05) / log(1-p) iterations for 95% confidence. The table above provides comfortable margins.
Before proceeding to Step 1.5, confirm you have:
Do NOT write a fix yet. You have a hypothesis, but proceed to Step 1.5 to validate it with existing failure data.
Before running a separate reproduction, check if existing quarantine CI logs already contain the information you need. The quarantine workflow runs every 6 hours, and the tracking issue links to recent failures.
The tracking issue contains ❌ links to failed quarantine runs. Use those run IDs to find the specific job that failed:
# Find the failed job for your test project in a quarantine run
gh api "repos/microsoft/aspire/actions/runs/<run_id>/jobs?per_page=100&filter=latest" \
--jq '.jobs[] | select(.name | contains("<ProjectShortname>")) | select(.conclusion == "failure") | {id: .id, name: .name}'Then download the logs for that job:
# Get logs via the GitHub MCP tool (preferred — handles encoding automatically)
# Use get_job_logs with the job_id, return_content: true, tail_lines: 300
# Or via CLI
gh api "repos/microsoft/aspire/actions/jobs/<job_id>/logs" > quarantine-failure.logSearch the logs for the test name, error message, and stack trace:
grep -i "TestMethodName\|TaskCanceled\|Assert\|Exception\|FAIL" quarantine-failure.log | head -30A test is likely contention-sensitive (fails only when running alongside other tests) if:
randomizePorts: false — fixed ports can conflict with other concurrent testsWaitForTextAsync — log-based readiness checks are fragile under contentionCancellationTokenSource across startup and readiness phases — one phase can starve the other's timeout budgetIf you identify contention-sensitive indicators, note this for Step 3 — single-test CI reproduction may fail, and you'll need to escalate to quarantine-project mode. Do NOT skip reproduction; the graduated escalation in Step 3 handles this.
Before proceeding:
Proceed to Step 2 for local reproduction.
Before going to CI, try reproducing the failure locally. This gives feedback in minutes instead of 30+ minutes.
uname -s # Linux, Darwin (macOS), or Windows (via MSYS/Git Bash)Compare your OS against the failing OSes from Step 1. Local reproduction is viable when:
If the test only fails on an OS you don't have (e.g., fails only on Windows and you're on Linux), skip to Step 3 (CI reproduction).
# Restore first if not already done
./restore.sh # or ./restore.cmd on Windows
# Build the specific test project
dotnet build tests/<TestProject>.Tests/<TestProject>.Tests.csproj -v:qUse the run-test-repeatedly.sh (Linux/macOS) or run-test-repeatedly.ps1 (Windows) script in .agents/skills/fix-flaky-test/. It runs the test command repeatedly with process cleanup between iterations.
Linux/macOS:
# Basic usage — run a single test 20 times (stop on first failure)
./.agents/skills/fix-flaky-test/run-test-repeatedly.sh -n 20 -- \
dotnet test tests/<TestProject>.Tests/<TestProject>.Tests.csproj --no-build \
-- --filter-method "*.<TestMethodName>" \
--filter-not-trait "quarantined=true" --filter-not-trait "outerloop=true"Windows (PowerShell):
# Basic usage — run a single test 20 times (stop on first failure)
./.agents/skills/fix-flaky-test/run-test-repeatedly.ps1 -n 20 -- dotnet test tests/<TestProject>.Tests/<TestProject>.Tests.csproj --no-build `
-- --filter-method "*.<TestMethodName>" `
--filter-not-trait "quarantined=true" --filter-not-trait "outerloop=true"For quarantined tests, you need /p:RunQuarantinedTests=true during both build and test to prevent the build system from filtering them out:
dotnet build tests/<TestProject>.Tests/<TestProject>.Tests.csproj -v:q /p:RunQuarantinedTests=true
# Linux/macOS
./.agents/skills/fix-flaky-test/run-test-repeatedly.sh -n 20 -- \
dotnet test tests/<TestProject>.Tests/<TestProject>.Tests.csproj --no-build \
/p:RunQuarantinedTests=true \
-- --filter-method "*.<TestMethodName>"
# Windows (PowerShell)
./.agents/skills/fix-flaky-test/run-test-repeatedly.ps1 -n 20 -- dotnet test tests/<TestProject>.Tests/<TestProject>.Tests.csproj --no-build `
-- --filter-method "*.<TestMethodName>"Choose iteration count based on failure rate (same heuristic as CI):
| Failure Rate | Local Iterations | Expected failures |
|---|---|---|
| >50% | 10 | ~5+ |
| 20-50% | 20 | ~4-10 |
| 10-20% | 30 | ~3-6 |
| 5-10% | 50 | ~2-5 |
| <5% | 100 | ~1-5 |
Script options (same for both .sh and .ps1):
-n <count> — Number of iterations (default: 100)--run-all — Don't stop on first failure, run all iterations--help — Show usageResults are saved to /tmp/test-results-<timestamp>/ (Linux/macOS) or $env:TEMP\test-results-<timestamp>\ (Windows). Failure logs are in failure-*.log files.
If the test fails locally: Reproduction successful ✅. Examine the failure log:
# The script prints the results directory path
cat /tmp/test-results-*/failure-*.logMark reproduce-local as done in SQL and proceed to Step 4 (root cause analysis) using the local failure logs.
UPDATE todos SET status = 'done' WHERE id = 'reproduce-local';
UPDATE todos SET status = 'done' WHERE id = 'reproduce-ci'; -- skip CI reproduction
UPDATE todos SET status = 'in_progress' WHERE id = 'analyze';If the test passes all local iterations: Local reproduction failed. This can happen because:
Proceed to Step 3 (CI reproduction) for cross-OS, parallel-runner reproduction.
UPDATE todos SET status = 'done' WHERE id = 'reproduce-local';
INSERT OR REPLACE INTO session_state (key, value) VALUES ('local_result', 'no_failures');run-test-repeatedly.sh/.ps1 with appropriate iteration count (or skipped due to OS mismatch)Create a separate branch for CI investigation. This branch will have ci.yml disabled and reproduce-flaky-tests.yml configured, keeping the fix branch clean.
# Create investigation branch from the current working branch
git checkout -b <fix-branch>-investigate
# Store the branch namesINSERT OR REPLACE INTO session_state (key, value) VALUES
('investigation_branch', '<fix-branch>-investigate'),
('fix_branch', '<fix-branch>');Disable ci.yml so pushing to the investigation branch doesn't trigger full CI:
# .github/workflows/ci.yml — add this at the top level, after `name:`
# Change the `on:` trigger to disable automatic runs:
on:
workflow_dispatch: {} # Only manual trigger, no automatic PR/push triggersThis prevents CI from running on every push to the investigation branch. You will re-enable it when creating the final fix PR.
Edit .github/workflows/reproduce-flaky-tests.yml — change only the env: section at the top:
env:
TEST_PROJECT: "Hosting.Azure" # Project shortname
TEST_FILTER: '--filter-method "*.DeployAsync_WithMultipleComputeEnvironments_Works"'
TARGET_OSES: "windows-latest" # Focus on highest-failure-rate OS
RUNNERS_PER_OS: "5"
ITERATIONS_PER_RUNNER: "5"OS targeting strategy:
ubuntu-latest,windows-latest with moderate iterationsTest project shortname mapping: The workflow resolves TEST_PROJECT to a path:
tests/{name}.Tests/{name}.Tests.csproj firsttests/Aspire.{name}.Tests/Aspire.{name}.Tests.csprojHosting → Aspire.Hosting.Tests, Hosting.Azure → Aspire.Hosting.Azure.TestsCommon filter patterns:
# Single test method
TEST_FILTER: '--filter-method "*.TestMethodName"'
# All tests in a class
TEST_FILTER: '--filter-class "*.TestClassName"'
# Multiple test methods
TEST_FILTER: '--filter-method "*.Test1" --filter-method "*.Test2"'For quarantined tests: The workflow automatically disables the quarantine exclusion filter for both build and test phases (via /p:_NonQuarantinedTestRunAdditionalArgs=""), so quarantined tests are included regardless of their trait. You do NOT need to add any special flags.
Zero-test detection: The workflow detects when zero tests execute (e.g., due to a misconfigured filter) and treats it as a failure. If you see "Zero tests executed" errors, verify that TEST_FILTER matches the actual test name and that quarantine settings are correct.
Commit the workflow changes and open a draft PR with the investigation template:
git add .github/workflows/ci.yml .github/workflows/reproduce-flaky-tests.yml
git commit -m "🔍 Investigation: configure CI for flaky test reproduction
⚠️ DO NOT MERGE — This is a temporary investigation branch.
ci.yml disabled, reproduce workflow configured for <test name>."
git push --set-upstream origin <fix-branch>-investigateOpen a draft PR with prominent WIP marking:
gh pr create --draft --repo microsoft/aspire \
--title "🔍 [DO NOT MERGE] Investigation: <test name>" \
--body "## ⚠️ DO NOT MERGE — Investigation Branch
This is a temporary branch for reproducing and verifying a fix for a flaky test.
**Issue**: #<issue-number>
**Test**: \`<FullyQualifiedTestName>\`
### What's changed on this branch
- \`ci.yml\` disabled (prevents full CI on investigation pushes)
- \`reproduce-flaky-tests.yml\` configured for the target test
- Code fix (will be applied after reproduction)
### Status
- [ ] Reproduction confirmed
- [ ] Fix applied
- [ ] Fix verified on CI
- [ ] Clean fix PR created
This branch will be deleted after the fix is verified and a clean PR is created."gh workflow run reproduce-flaky-tests.yml --repo microsoft/aspire --ref <fix-branch>-investigateThis dispatches the workflow from main but runs the version from your branch, so your env var edits will be used.
If the workflow dispatch fails (e.g. HTTP 403 "Resource not accessible by integration"): your GitHub token lacks actions:write permission on the repository. This is a non-fatal blocker — continue with the investigation, but you must document this in every PR you open (both investigation and fix PRs). Include the exact error, and provide the manual trigger command so a reviewer or maintainer can run it. See the PR template in Step 6.2 for the required format.
Monitor the run using polling (CI runs take 10-30+ minutes):
# Find the run ID
gh run list --repo microsoft/aspire --branch <branch> --limit 1 --json databaseId,statusStore the run ID, then poll periodically for completion:
INSERT OR REPLACE INTO session_state (key, value) VALUES ('reproduce_run_id', '<run-id>');# Poll for completion (use bash mode="async", then read_bash with increasing delays)
# Avoid `gh run watch` — it produces excessive output that floods the context window.
gh run view <run-id> --repo microsoft/aspire --json status,conclusion --jq '{status, conclusion}'
# Check individual job results as they complete
gh run view <run-id> --repo microsoft/aspire --json jobs \
--jq '.jobs[] | select(.status == "completed") | {name: .name, conclusion: .conclusion}'Cancel old runs when starting new ones to avoid wasting CI resources:
# Cancel a specific run
gh run cancel <run-id> --repo microsoft/aspire
# Cancel all in-progress runs on your branch (useful when iterating)
gh run list --repo microsoft/aspire --branch <branch> --status in_progress --json databaseId --jq '.[].databaseId' | \
xargs -I {} gh run cancel {} --repo microsoft/aspireAlways cancel previous reproduce/verify runs before pushing a new configuration. workflow_dispatch runs are NOT auto-cancelled, so you must cancel them manually.
⛔ GATE: Do not proceed past this point until the CI run has completed.
If there are failure artifacts, download them:
# Download failure artifacts
gh run download <run-id> --repo microsoft/aspire --dir /tmp/failure-logs
# Or get logs directly via the GitHub API / MCP tools
gh api "repos/microsoft/aspire/actions/jobs/<job_id>/logs" > /tmp/failure.logDistinguishing test failures from infrastructure failures:
CI runners sometimes fail due to infrastructure issues, NOT the test itself. Common infrastructure failures include:
Failed to install or invoke dotnet... (exit code -1073741502 on Windows)The runner has received a shutdown signal or runner timeoutsdotnet restoreThese do NOT count as reproductions. Check the actual error message — only count iterations where the test itself failed with the expected error pattern from the tracking issue.
If some runners show test failures (the expected error): Reproduction successful ✅. Proceed to Step 4.
If no runners show the expected test failure — scale up and retry:
-- Track the scaling attempt
INSERT OR REPLACE INTO session_state (key, value)
VALUES ('reproduce_attempt', CAST((SELECT CAST(value AS INTEGER) FROM session_state WHERE key = 'reproduce_attempt') + 1 AS TEXT));Scale up progressively, focusing on the OS most likely to fail first (based on per-OS failure rates from the issue). Go back to Step 3.1 after each change:
| Attempt | TARGET_OSES | RUNNERS_PER_OS | ITERATIONS_PER_RUNNER | Notes |
|---|---|---|---|---|
| 1 | Highest-failure-rate OS only | From heuristic table | From heuristic table | Start narrow — one OS, sized by failure rate |
| 2 | Same single OS | Same | 2× previous | Double ITERATIONS_PER_RUNNER only |
Upper bounds: Do not exceed RUNNERS_PER_OS=10 or ITERATIONS_PER_RUNNER=50 (total matrix entries must stay ≤ 256 per GitHub Actions limits).
If 2 attempts produce zero test failures → escalate to quarantine-project mode (Step 3.6).
When single-test reproduction fails, the test likely only fails under contention with other tests. Escalate by running all quarantined tests in the assembly, which matches what tests-quarantine.yml does:
Edit .github/workflows/reproduce-flaky-tests.yml — change the TEST_FILTER to target all quarantined tests:
env:
TEST_PROJECT: "<same project>"
TEST_FILTER: '--filter-trait "quarantined=true"' # Run all quarantined tests in this assembly
TARGET_OSES: "<same as before>"
RUNNERS_PER_OS: "3"
ITERATIONS_PER_RUNNER: "3"This recreates the contention environment from the quarantine workflow. Push, trigger, and monitor as before.
If quarantine-project mode reproduces the failure: Reproduction successful ✅. Proceed to Step 4. Note: verification (Step 5) should use the same quarantine-project mode to confirm the fix.
If quarantine-project mode also produces zero failures: The test requires heavier contention than we can simulate on demand. In this case:
INSERT OR REPLACE INTO session_state (key, value) VALUES ('reproduce_mode', 'log-based');CRITICAL: Windows log encoding gotcha
Windows CI log files downloaded as artifacts are encoded as UTF-16LE. Running cat on them produces garbled output. Convert first:
# Convert Windows log to readable UTF-8
iconv -f UTF-16LE -t UTF-8 /tmp/failure-logs/failures-windows-latest-1/test-output.log > /tmp/readable-windows.log
cat /tmp/readable-windows.logTip: Using get_job_logs via GitHub API/MCP tools returns UTF-8 directly, avoiding encoding issues entirely. Prefer API-based log retrieval when possible.
Alternatively, search for the error directly:
# Search across all failure logs (handles encoding)
find /tmp/failure-logs -name "*.log" -exec grep -l "Assert\|Error\|Exception" {} \;RUNNERS_PER_OS and ITERATIONS_PER_RUNNER and try again.Failure logs may come from local runs (Step 2, in /tmp/test-results-*/), CI reproduce runs (Step 3), or existing quarantine runs (Step 1.5). All are valid sources.
Preferred: Use GitHub API/MCP tools to get logs directly (avoids encoding issues):
# Get job logs via GitHub MCP tool: get_job_logs with job_id, return_content: true, tail_lines: 300
# Or via CLI:
gh api "repos/microsoft/aspire/actions/jobs/<job_id>/logs" > /tmp/failure.logDelegate log analysis to a sub-agent to keep the main context clean:
Use a task agent (explore or general-purpose) to analyze the failure logs:
- Pass the log file paths or content
- Ask it to identify the specific assertion/exception
- Ask it to read the test source code and identify the concurrency/timing model
- Have it return a structured root cause summaryLook for the assertion or exception that failed:
# Find the actual test failure in logs
grep -A 10 "FAIL\|Assert\.\|Exception" /tmp/failure.log | head -50
# For .trx files (XML test results) from downloaded artifacts
find /tmp/failure-logs -name "*.trx" -exec grep -l 'outcome="Failed"' {} \;Then find the corresponding test code and understand the concurrency/timing model.
Before proceeding to Step 5, confirm you have:
Now — and only now — proceed to write the fix.
⚠️ DO NOT remove the
[QuarantinedTest]attribute or close the tracking issue. Unquarantining is a separate process that happens after 21 days of zero failures in quarantine CI. Your fix PR should contain only the code fix. See Step 6.5 for details.
dotnet build tests/<TestProject>.Tests/<TestProject>.Tests.csproj --no-restore -v:qreproduce-flaky-tests.yml configured for the same testPrinciple: Local runs are a fast pre-check, not a substitute for CI. Running a test N times on one machine does not have the same statistical power as N runs across separate CI runners. Some flakiness stems from environmental variation (machine load, Docker daemon state, network conditions) that a single machine cannot reproduce. Local verification catches obvious regressions quickly and saves CI round-trips, but CI verification is always required as the final gate.
If local reproduction succeeded in Step 2, run a quick local verification first:
# Rebuild with fix
dotnet build tests/<TestProject>.Tests/<TestProject>.Tests.csproj --no-restore -v:q
# Quick local check — same iteration count as reproduction
# Linux/macOS:
./.agents/skills/fix-flaky-test/run-test-repeatedly.sh -n 20 -- \
dotnet test tests/<TestProject>.Tests/<TestProject>.Tests.csproj --no-build \
-- --filter-method "*.<TestMethodName>"
# Windows (PowerShell):
# ./.agents/skills/fix-flaky-test/run-test-repeatedly.ps1 -n 20 -- dotnet test tests/<TestProject>.Tests/<TestProject>.Tests.csproj --no-build -- --filter-method "*.<TestMethodName>"If local verification fails, iterate on the fix before going to CI. This saves ~30 minutes per CI round-trip.
CI verification is always required. However, the scale should reflect your local confidence — how much evidence you already have that the fix is correct.
Consider these factors to determine how aggressively to scale CI verification:
Higher confidence (scale CI down):
HttpClient with resilient one)Lower confidence (scale CI up):
Use the original failure rate combined with your local confidence to size the CI verification. The base scale ensures that if the bug were still present, it would manifest with ≥95% probability (n ≥ log(0.05) / log(1-p)):
| Original Failure Rate | High Confidence (CI scale) | Low Confidence (CI scale) |
|---|---|---|
| >50% | 3 × 3 per OS (9 total) | 3 × 3 per OS (9 total) |
| 20-50% | 3 × 5 per OS (15 total) | 5 × 5 per OS (25 total) |
| 10-20% | 5 × 5 per OS (25 total) | 5 × 10 per OS (50 total) |
| 5-10% | 5 × 10 per OS (50 total) | 10 × 10 per OS (100 total) |
| <5% | 10 × 10 per OS (100 total) | 10 × 25 per OS (250 total) |
For tests with very low failure rates (<5%), consider whether the verification is practical within CI budget constraints. If not, document the limitation and rely on the 21-day quarantine monitoring to confirm.
For contention-sensitive tests (where quarantine-project mode was needed for reproduction): Use the same quarantine-project TEST_FILTER for verification. This ensures the fix is validated under the same contention conditions where the failure was observed. If reproduction fell back to log-based analysis, use the low-confidence column and note in the PR that definitive confirmation relies on the 21-day quarantine monitoring.
Push the fix to the investigation branch (where reproduce workflow is already configured):
git add -A
git commit -m "Fix flaky test: <description of fix>"
git pushThen trigger the reproduce workflow to verify:
gh workflow run reproduce-flaky-tests.yml --repo microsoft/aspire --ref <fix-branch>-investigateIf the workflow dispatch fails due to permissions (HTTP 403), see the guidance in Step 3.3. Continue to Step 6 but document the failure in the PR description.
Store the verification run ID:
INSERT OR REPLACE INTO session_state (key, value) VALUES ('verify_run_id', '<run-id>');
INSERT OR REPLACE INTO session_state (key, value) VALUES ('fix_attempt', '1');Wait for CI to complete. Monitor with polling (gh run view --json status,conclusion), not gh run watch.
If all iterations pass across all OSes: The fix is validated ✅. Proceed to Step 6.
If some iterations still fail: The fix is incomplete or incorrect. Iterate:
-- Track the fix attempt
INSERT OR REPLACE INTO session_state (key, value)
VALUES ('fix_attempt', CAST((SELECT CAST(value AS INTEGER) FROM session_state WHERE key = 'fix_attempt') + 1 AS TEXT));gh run download <run-id> --repo microsoft/aspire --dir /tmp/failure-logsAfter 3 failed fix attempts: Stop and report findings to the user. The issue may require deeper architectural changes or domain expertise.
After the fix is verified on the investigation branch, create a clean fix PR.
Cancel any in-progress reproduce or verify runs that are no longer needed:
# List and cancel any remaining runs on your branch
gh run list --repo microsoft/aspire --branch <branch> --status in_progress --json databaseId,name --jq '.[] | "\(.databaseId) \(.name)"'
gh run cancel <run-id> --repo microsoft/aspireSwitch back to the fix branch and cherry-pick only the code fix commits (not the workflow changes):
git checkout <fix-branch>
# Cherry-pick the fix commit(s) from the investigation branch
git cherry-pick <fix-commit-sha>
# Verify the fix branch has NO workflow changes
git diff main -- .github/workflows/ # Should be emptyBefore pushing, ensure the fix branch has a clean, linear history with only the code fix commit(s). If intermediate commits crept in (e.g. workflow config changes, reverts, debug attempts), squash them down:
# Interactive rebase to squash noise commits into the fix
git rebase -i $(git merge-base HEAD main)
# In the editor, mark the fix commit as "pick" and any noise commits as "fixup" or "drop"
# Save and close
# Verify: the branch should have only fix-related commits
git log --oneline main..HEAD
# Verify: no unintended file changes
git diff main -- .github/workflows/ # Should be emptyWhy this matters: A clean history makes the PR easy to review and avoids confusing commit pairs (config + revert) that produce no net change but clutter the log.
git pushDetermine the PR title prefix: Check whether any user interaction occurred during the investigation:
SELECT value FROM session_state WHERE key = 'user_interaction';user_interaction is 'false': prefix the PR title with [automated] user_interaction is 'true': no prefixOpen a non-draft PR with the fix. The PR body must include a note that it was created using the fix-flaky-test skill:
gh pr create --repo microsoft/aspire \
--title "<prefix>Fix flaky test: <description>" \
--body "## Flaky Test Fix
### Test
- **Method**: \`<fully qualified test name>\`
- **Issue**: #<issue-number>
### Root Cause
<1-2 sentence description of the root cause>
### Fix
<1-2 sentence description of what was changed>
### Verification
| Run | Config | Result |
|-----|--------|--------|
| Pre-fix (local) | <iterations>, <OS> | **<pass/fail>** |
| Post-fix (local) | <iterations>, <OS> | **<pass/fail>** |
| Post-fix (CI) | <runners × iters × OSes> | **<link to run>** |
> **If any verification step was skipped or failed** (e.g. workflow dispatch permission error), replace the CI row with a clear explanation:
> - What step failed and the exact error (e.g. \\\`HTTP 403: Resource not accessible by integration\\\`)
> - Why it could not be completed (e.g. agent token lacks \\\`actions:write\\\` permission)
> - The manual command a reviewer can run to complete verification
> - A link to the investigation PR/branch with the pre-configured reproduce workflow
### Verification Rationale
<Brief explanation of CI scale choice: local confidence level, why that scale was appropriate for the failure rate, and acknowledgment that local runs are a pre-check — not equivalent to CI runs across separate runners.>
### Notes
- \`[QuarantinedTest]\` attribute kept — unquarantining will happen separately after 21 days of zero failures in quarantine CI
---
> **Note:** This PR intentionally does not close #<issue-number>. The test will remain quarantined until a separate unquarantine process confirms it has been stable (zero failures) for a sufficient period. Once stability is confirmed, the test will be unquarantined and the issue will be closed.
---
*This fix was generated using the [fix-flaky-test skill](https://github.com/microsoft/aspire/blob/main/.agents/skills/fix-flaky-test/SKILL.md).*"If gh pr create fails (e.g. permissions error, API failure): Do NOT delete the branch or undo the work. Instead:
https://github.com/microsoft/aspire/compare/main...<branch-name>)# Close the investigation draft PR
gh pr close <investigation-pr-number> --repo microsoft/aspire --delete-branchAfter opening the final PR, the regular CI pipeline (ci.yml) will run automatically. Monitor it to confirm the fix does not introduce regressions:
# Find the CI run for your PR
gh pr checks <pr-number> --repo microsoft/aspireIf CI fails on unrelated tests, that's not your problem — note it in the PR. If CI fails on your changed files or the test project you modified, investigate and fix before marking the task complete.
UPDATE todos SET status = 'done' WHERE id = 'verify-ci';Important policy: A code fix alone is not sufficient to unquarantine a test. The test must have zero failures across all OSes for 21 consecutive days in the quarantine CI runs before it can be unquarantined. See docs/unquarantine-policy.md.
[QuarantinedTest] attributeBefore opening the final PR, verify every item. This is a hard gate — do not skip any item.
[QuarantinedTest] attribute is still present on the test method (not removed)Self-check: Run git diff on the fix branch and scan for any unintended changes — removed test attributes, workflow file edits, or unrelated modifications.
UPDATE todos SET status = 'done' WHERE id = 'cleanup';eng/Testing.targets auto-appends --filter-not-trait "quarantined=true" to test arguments via the TestRunnerAdditionalArguments MSBuild property. The filter value is defined in eng/Testing.props, and the composed property is evaluated during dotnet test even with --no-build, so it must be handled in both build and test commands:
_NonQuarantinedTestRunAdditionalArgs to empty, removing the quarantine exclusion filter for all tests/p:RunQuarantinedTests=true to both dotnet build and dotnet testTesting.targets also adds --ignore-exit-code 8 via MtpBaseArgs, which masks zero-test runs as successes. The workflow and run-test-repeatedly scripts detect this by checking test output for the Total: count indicator.
The workflow:
{os, index} combinationsFailed iterations upload their test output as artifacts named failures-<os>-<index>.
workflow_dispatch requires the workflow file to exist on the default branch (main). Key implications:
gh workflow run reproduce-flaky-tests.yml --ref <branch>. GitHub discovers the workflow from main but runs the version from the specified --ref. This means your investigation branch's env var edits will be used.ci.yml disabled, so pushes don't trigger full CI — only workflow_dispatch of the reproduce workflow is used.workflow_dispatch until it's merged to main.After completing a flaky test fix, provide a summary:
## Flaky Test Fix Summary
### Test
- **Method**: `Namespace.Type.Method`
- **Issue**: #XXXXX
- **Project**: `tests/Aspire.{Project}.Tests/`
### Failure Data
| OS | Failure Rate |
|---|---|
| Windows | XX% |
| Linux | XX% |
### Root Cause
Brief description of what caused the flaky behavior.
### Fix
Description of the code change.
### Verification
| Run | Config | Result |
|-----|--------|--------|
| Pre-fix | X runners × Y iters × Z OSes | N failures ❌ |
| Post-fix | X runners × Y iters × Z OSes | All passed ✅ |
### Files Changed
- `path/to/file.cs` — description
### Next Steps
- Test remains quarantined — will be unquarantined after 21 days of zero failures
- Issue #XXXXX remains open — will be closed by the unquarantine processrun-test-repeatedly.sh (Linux/macOS) or run-test-repeatedly.ps1 (Windows) for fast feedback (~minutes). Fall back to CI when local reproduction fails (wrong OS, contention-sensitive, very low failure rate)uname -s to decide if local reproduction is viable for the failing OSdotnet build and dotnet test commands for local reproduction. The CI reproduce workflow handles this automatically.--ignore-exit-code 8), the filter or quarantine settings are misconfigured. The run-test-repeatedly scripts and reproduce workflow detect this automatically.plan.md and files/ in the session workspace, not a directory in the repoFailed to install or invoke dotnet... on Windows). These do NOT count as test reproductions. Always verify the error matches the expected test failure pattern.docs/unquarantine-policy.md)session_state (reproduce_mode).get_job_logs via GitHub API/MCP, which returns UTF-8)gh run watch: Use gh run view --json status,conclusion to check CI status — gh run watch produces excessive output that floods the context windowCommon flaky test patterns are documented in .github/instructions/test-review-guidelines.instructions.md. Consult that file during Step 1 (gather data) to form hypotheses, and during Step 4 (analysis) to confirm root causes.
© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in .agents/skills/fix-flaky-test of microsoft/aspire.
Open the folder on GitHubat commit 809a672
Fix Flaky Test next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Fix Flaky Test this skillmicrosoft/aspire | 6.3k | — | ~13k | Automated safety check: Pass | MIT | |
| Swig Testswig/swig | 6.3k | — | ~2.3k | Automated safety check: Pass | Custom licence | |
| Dynamo Jira TicketDynamoDS/Dynamo | 2k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| Fix Ready PRsfastrepl/anarlog | 9.4k | — | ~1.4k | Automated safety check: Pass | MIT | |
| Trx Analysismicrosoft/vstest | 969 | — | ~1.8k | Automated safety check: Pass | MIT | |
| Wioworkersio/skills | 180 | — | ~5.8k | Automated safety check: Pass | MIT |
swig/swig
Run SWIG test suite for specific languages. An agent skill from swig/swig.
DynamoDS/Dynamo
Create structured Jira tickets for Dynamo from bug reports, failing tests, or feature requests.
fastrepl/anarlog
Inspect every open non-draft PR for CI failures and unresolved Cursor Bugbot findings, then fix them on the existing PR branches.
microsoft/vstest
Parse and analyze Visual Studio TRX test result files. An agent skill from microsoft/vstest.
workersio/skills
Testing workflow skill for finding high-value test candidates, writing focused tests, generating realistic workloads, reviewing test value, and diagnosing test-suite health.
UditAkhourii/quicksilver
Offload bulk judgment calls to Jev (TypeSafe's fast System One model) so Claude doesn't read, and pay for, content it only needs a verdict on.
microsoft/aspire
A skill your agent uses when asked to trigger or inspect Aspire internal Azure DevOps builds, source-index runs, or release validation on dnceng/internal; push to the internal mirror; download build…
microsoft/aspire
Backports a merged PR to a release branch by triggering the /backport bot, waiting for the bot-created PR, and filling in the shiproom template (Customer Impact, Testing, Risk, Regression?).
microsoft/aspire
Bumps the Aspire repository product version in eng/Versions.props using previous version-bump commits as guidance.
microsoft/aspire
Guide for diagnosing GitHub Actions test failures, extracting failed tests from runs, and creating or updating failing-test issues.
microsoft/aspire
Create a pull request using the repository PR template. An agent skill from microsoft/aspire.
microsoft/aspire
Guide for writing tests for the Aspire Dashboard. An agent skill from microsoft/aspire.
Categories
Reproduces and fixes flaky or quarantined tests. An agent skill from microsoft/aspire. Fix Flaky Test is an agent skill from microsoft/aspire, published by the product's own GitHub organization. Reproduces and fixes flaky or quarantined tests.
Fix Flaky Test fits situations like: tasks that involve Failing and flaky tests.
Run `npx skills add microsoft/aspire --skill fix-flaky-test -a claude-code`. Or copy the skill folder (.agents/skills/fix-flaky-test in microsoft/aspire) into .claude/skills/fix-flaky-test in your project. Claude Code loads it when a task matches its description.
Run `npx skills add microsoft/aspire --skill fix-flaky-test -a codex`. Or copy the skill folder (.agents/skills/fix-flaky-test in microsoft/aspire) into .agents/skills/fix-flaky-test in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/aspire --skill fix-flaky-test -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fix-flaky-test, .gemini/skills/fix-flaky-test, .github/skills/fix-flaky-test and .opencode/skills/fix-flaky-test in your project.
Going by SKILL.md and its folder, Fix Flaky Test needs PowerShell and a shell for the scripts in its folder and the command-line tools its instructions call (gh, git and dotnet). Our summary lists: A Bash shell; PowerShell.
SKILL.md contains no URLs. Its commands use gh and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Fix Flaky Test is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 13k tokens (SKILL.md is roughly 51k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Fix Flaky Test: Swig Test (swig/swig, 6.3k stars), Dynamo Jira Ticket (DynamoDS/Dynamo, 2k stars), Fix Ready PRs (fastrepl/anarlog, 9.4k stars) and Trx Analysis (microsoft/vstest, 969 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
microsoft (a GitHub organization, an official publisher) maintains it in microsoft/aspire, which has 6,348 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 7, 2026.
Source: microsoft/aspire on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.