Agent skill

Test Fixer

by mitchdenny in mitchdenny/hex1b

Agent for diagnosing and fixing flaky terminal UI tests in the Hex1b test suite.

MITAuto-check passedTesting & QA

Install Test Fixer

skills CLI
$ npx skills add mitchdenny/hex1b --skill test-fixer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mitchdenny/hex1b test-fixer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mitchdenny/hex1b.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/test-fixer .claude/skills/test-fixer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-fixer
GitHub stars
178
Token cost
~6.5k tokens
SKILL.md length
1,351 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Agent for diagnosing and fixing flaky terminal UI tests in the Hex1b test suite.

  • Works in 4 steps: Identify the Failure Pattern → Check Local vs CI Behavior → Examine the Test Pattern → …
  • Tests pass locally but fail in CI
  • SKILL.md covers Quick Reference: Common Flaky…, Pattern 1: Snapshot Captured…, Pattern 2: Missing WaitUntil… and Pattern 3: Race Condition with…, plus 9 more sections
  • Calls dotnet and gh

What it does

Test Fixer is an agent skill from mitchdenny/hex1b. Agent for diagnosing and fixing flaky terminal UI tests in the Hex1b test suite. Use when tests pass locally but fail in CI, or when tests exhibit timing-sensitive behavior.

Its SKILL.md is about 6.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Test generation. It works with .NET and Linux. The repository describes itself as: The .NET Terminal Application Stack. The licence is MIT.

When your agent uses it

  • Tests pass locally but fail in CI
  • Tests exhibit timing-sensitive behavior

Example prompts

  • “/test-fixer”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Identify the Failure Pattern
  2. Check Local vs CI Behavior
  3. Examine the Test Pattern
  4. Apply the Fix

What it can do on your machine

Read from SKILL.md and the folder at commit 98d8766. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • dotnet
    • gh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Fixer loads about 6.5k tokens when it runs. Until then it costs about 46 tokens; SKILL.md has 1,351 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~6.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mitchdenny/hex1b at commit 98d8766, republished under its MIT licence (© mitchdenny). 1,351 words, ~6,503 tokens.

Download SKILL.mdSave it as .claude/skills/test-fixer/SKILL.md (or your agent's skills folder).
name
test-fixer
description
Agent for diagnosing and fixing flaky terminal UI tests in the Hex1b test suite. Use when tests pass locally but fail in CI, or when tests exhibit timing-sensitive behavior.

Test Fixer Skill

This skill provides guidelines for AI agents to diagnose and fix flaky tests in the Hex1b TUI library test suite. These tests use MSTest 4 (MSTest.Sdk/4.2.3, OutputType=Exe) with Microsoft.Testing.Platform and Hex1bTerminalInputSequenceBuilder to simulate user interactions with terminal applications. global.json configures dotnet test to use MTP, and test projects can also be run as executables with dotnet run --project tests/SomeProject/.

Quick Reference: Common Flaky Test Patterns

PatternSymptomFix
Snapshot After ExitTest passes locally, fails on Linux CIMove WaitUntil before Capture, ensure Capture is before exit
Missing WaitUntilIntermittent assertion failuresAdd WaitUntil for expected state before Capture
Race with Ctrl+CSnapshot missing expected contentAdd WaitUntil between last action and Capture
Task.WhenAny RaceTest sometimes times outReplace with proper WaitUntil or increase timeout
Test InterferencePass isolated, fail in suiteCheck for shared state, file locks, or parallel execution issues
Platform-SpecificFails consistently on Windows/LinuxAdd TestCategory/Ignore or fix platform-specific code
Task.Delay for Async EventsFlaky on slower CI runnersReplace Task.Delay with TaskCompletionSource signal
Helper Partial WaitTests using multi-line helpers fail intermittentlyWait for all/last content, not just first line

Pattern 1: Snapshot Captured After App Exit

⚠️ This is the most common flaky test issue
Symptoms
  • Test passes consistently on Windows
  • Test fails on Linux CI (GitHub Actions ubuntu-latest)
  • Assertion fails with content that "should" be there
  • Error messages like Assert.IsTrue failed or An item should be selected with indicator
Root Cause

The ApplyWithCaptureAsync method returns terminal.CreateSnapshot() after all steps complete, not at the point where .Capture() is called. When the sequence includes Ctrl+C to exit the app, the terminal buffer may be cleared before the final snapshot is taken.

Platform difference: Windows terminal buffers persist longer after app exit; Linux clears them more aggressively.

Example: Broken Test
csharp
// ❌ BROKEN: Snapshot is taken AFTER Ctrl+C exits the app
var snapshot = await new Hex1bTerminalInputSequenceBuilder()
    .WaitUntil(s => s.ContainsText("Counter: 3"), TimeSpan.FromSeconds(2))
    .Capture("final")           // This saves SVG but doesn't capture for return!
    .Ctrl().Key(Hex1bKey.C)     // App exits, terminal buffer may be cleared
    .Build()
    .ApplyWithCaptureAsync(terminal, TestContext.Current.CancellationToken);
// snapshot is from AFTER Ctrl+C, not from Capture step!
Assert.IsTrue(snapshot.ContainsText("Counter: 3")); // May fail on Linux!
How It Works
  1. Hex1bTerminalInputSequenceBuilder builds a sequence of TestStep objects
  2. ApplyWithCaptureAsync executes all steps, then calls terminal.CreateSnapshot() at the end
  3. CaptureStep only saves SVG/HTML files—it does NOT store the snapshot for return
  4. When Ctrl+C triggers app exit, the terminal may clear its buffer before the final snapshot
Fix Strategy

Option A: Ensure WaitUntil is immediately before Capture (Recommended)

csharp
// ✅ FIXED: WaitUntil confirms state, then Capture, then exit
var snapshot = await new Hex1bTerminalInputSequenceBuilder()
    .Key(Hex1bKey.A)
    .Key(Hex1bKey.B)
    .WaitUntil(s => s.ContainsText("Counter: 3"), TimeSpan.FromSeconds(2))
    .Capture("final")
    .Ctrl().Key(Hex1bKey.C)
    .Build()
    .ApplyWithCaptureAsync(terminal, TestContext.Current.CancellationToken);
await runTask;

// If the WaitUntil passed, we know the content was there at Capture time
// The assertion is now checking what WaitUntil already verified
Assert.IsTrue(snapshot.ContainsText("Counter: 3"));

Option B: Don't use snapshot for content assertions

csharp
// ✅ ALTERNATIVE: Use render counters or other non-snapshot assertions
var runTask = app.RunAsync(TestContext.Current.CancellationToken);
await new Hex1bTerminalInputSequenceBuilder()
    .Key(Hex1bKey.A)
    .Key(Hex1bKey.B)
    .Capture("final")
    .Ctrl().Key(Hex1bKey.C)
    .Build()
    .ApplyWithCaptureAsync(terminal, TestContext.Current.CancellationToken);
await runTask;

// Assert on counters/state captured during execution, not snapshot content
Assert.AreEqual(1, staticRenderCount);

Option C: Use WaitUntil as the assertion itself

csharp
// ✅ ALTERNATIVE: WaitUntil IS the assertion - if it passes, test passes
await new Hex1bTerminalInputSequenceBuilder()
    .WaitUntil(s => s.ContainsText("> Item 15"), TimeSpan.FromSeconds(2))
    .Ctrl().Key(Hex1bKey.C)
    .Build()
    .ApplyAsync(terminal, TestContext.Current.CancellationToken);
await runTask;
// No additional assertions needed - WaitUntil already verified the content

Pattern 2: Missing WaitUntil Before Assertion

Symptoms
  • Test sometimes fails, sometimes passes (even on same platform)
  • Assertion checks for content that "should" appear after an action
  • Content appears when running manually but not in fast automated tests
Root Cause

User input (key press, mouse click) triggers async rendering. Without WaitUntil, the snapshot may be captured before the render completes.

Example: Broken Test
csharp
// ❌ BROKEN: No wait after Down() for selection to update
var snapshot = await new Hex1bTerminalInputSequenceBuilder()
    .WaitUntil(s => s.ContainsText("First"), TimeSpan.FromSeconds(2))
    .Down()  // Triggers async render
    .Capture("final")  // May capture before render completes!
    .Ctrl().Key(Hex1bKey.C)
    .Build()
    .ApplyWithCaptureAsync(terminal, TestContext.Current.CancellationToken);
Assert.IsTrue(snapshot.ContainsText("> Second")); // Flaky!
Fix
csharp
// ✅ FIXED: Wait for expected state after action
var snapshot = await new Hex1bTerminalInputSequenceBuilder()
    .WaitUntil(s => s.ContainsText("First"), TimeSpan.FromSeconds(2))
    .Down()
    .WaitUntil(s => s.ContainsText("> Second"), TimeSpan.FromSeconds(2))  // Wait for render!
    .Capture("final")
    .Ctrl().Key(Hex1bKey.C)
    .Build()
    .ApplyWithCaptureAsync(terminal, TestContext.Current.CancellationToken);
Assert.IsTrue(snapshot.ContainsText("> Second")); // Reliable

Pattern 3: Race Condition with Ctrl+C

Symptoms
  • Test passes most of the time
  • Occasionally fails with assertion that content is missing
  • More common on slower CI runners
Root Cause

Ctrl+C can be processed before the previous action's render completes, especially if the action triggers complex layout recalculation.

Example: Broken Test
csharp
// ❌ BROKEN: Scroll may not complete before Ctrl+C
var snapshot = await new Hex1bTerminalInputSequenceBuilder()
    .WaitUntil(s => s.ContainsText("Item 01"), TimeSpan.FromSeconds(2))
    .ScrollDown(5)
    .Capture("after_scroll")
    .Ctrl().Key(Hex1bKey.C)  // May interrupt scroll processing!
    .Build()
    .ApplyWithCaptureAsync(terminal, TestContext.Current.CancellationToken);
Assert.IsTrue(snapshot.ContainsText("> Item 06")); // Flaky!
Fix
csharp
// ✅ FIXED: Verify scroll completed before exiting
var snapshot = await new Hex1bTerminalInputSequenceBuilder()
    .WaitUntil(s => s.ContainsText("Item 01"), TimeSpan.FromSeconds(2))
    .ScrollDown(5)
    .WaitUntil(s => s.ContainsText("> Item 06"), TimeSpan.FromSeconds(2))  // Verify scroll!
    .Capture("after_scroll")
    .Ctrl().Key(Hex1bKey.C)
    .Build()
    .ApplyWithCaptureAsync(terminal, TestContext.Current.CancellationToken);
Assert.IsTrue(snapshot.ContainsText("> Item 06")); // Reliable

Test Infrastructure Overview

Key Components
ComponentLocationPurpose
Hex1bTerminalInputSequenceBuildersrc/Hex1b/Automation/Fluent builder for test sequences
Hex1bTerminalInputSequencesrc/Hex1b/Automation/Executes steps and returns final snapshot
CaptureSteptests/Hex1b.Tests/CaptureStep.csSaves SVG/HTML at capture point
WaitUntilStepsrc/Hex1b/Automation/WaitUntilStep.csPolls until condition met or timeout
TestSequenceExtensionstests/Hex1b.Tests/TestSequenceExtensions.csAdds .Capture() extension method
Understanding ApplyWithCaptureAsync
csharp
// From Hex1bTerminalInputSequence.cs
public async Task<Hex1bTerminalSnapshot> ApplyAsync(Hex1bTerminal terminal, CancellationToken ct = default)
{
    foreach (var step in _steps)
    {
        ct.ThrowIfCancellationRequested();
        await step.ExecuteAsync(terminal, _options, ct);
    }
    return terminal.CreateSnapshot();  // ⚠️ Snapshot is ALWAYS at the end!
}

The CaptureStep only saves files—it does NOT affect what ApplyAsync returns:

csharp
// From CaptureStep.cs
internal override Task ExecuteAsync(...)
{
    TestCaptureHelper.Capture(terminal, Name);  // Saves SVG/HTML only
    return Task.CompletedTask;
}

Diagnostic Workflow

Step 1: Identify the Failure Pattern
bash
# Get CI logs
gh run view <run-id> --log-failed

# Look for test names and error messages
# Common patterns:
# - Assert.IsTrue failed
# - Expected:<True> or equivalent / Actual:<False>
# - "An item should be selected with indicator"
Step 2: Check Local vs CI Behavior
powershell
# Run specific test locally
dotnet test tests/Hex1b.Tests --filter "FullyQualifiedName~<TestName>"

# Run multiple times to check for flakiness
for ($i = 1; $i -le 10; $i++) { 
    dotnet test tests/Hex1b.Tests --filter "FullyQualifiedName~<TestName>" --no-build 2>&1 | 
    Select-String -Pattern "(Passed|Failed)" 
}
Step 3: Examine the Test Pattern

Look for these anti-patterns:

csharp
// Anti-pattern 1: Snapshot assertion after Ctrl+C
.Capture("final")
.Ctrl().Key(Hex1bKey.C)
// ... later ...
Assert.IsTrue(snapshot.ContainsText("expected"));  // ❌

// Anti-pattern 2: Action without WaitUntil before Capture
.Down()
.Capture("final")  // ❌ No WaitUntil after Down()

// Anti-pattern 3: Complex action right before Ctrl+C
.ScrollDown(10)
.Capture("final")
.Ctrl().Key(Hex1bKey.C)  // ❌ No WaitUntil to verify scroll completed
Step 4: Apply the Fix
  1. Add WaitUntil after any action that changes state
  2. Ensure WaitUntil checks for the exact condition being asserted
  3. Place WaitUntil immediately before Capture
  4. Consider if the assertion is even needed (WaitUntil already verified it)

Safe Test Patterns

Pattern A: Full Integration Test with Exit
csharp
[TestMethod]
public async Task Navigation_DownArrow_SelectsNextItem()
{
    using var workload = new Hex1bAppWorkloadAdapter();
    using var terminal = Hex1bTerminal.CreateBuilder()
        .WithWorkload(workload)
        .WithHeadless()
        .WithDimensions(40, 10)
        .Build();
    
    using var app = new Hex1bApp(
        ctx => Task.FromResult<Hex1bWidget>(ctx.List(["First", "Second", "Third"])),
        new Hex1bAppOptions { WorkloadAdapter = workload }
    );

    var runTask = app.RunAsync(TestContext.Current.CancellationToken);
    
    var snapshot = await new Hex1bTerminalInputSequenceBuilder()
        .WaitUntil(s => s.ContainsText("> First"), TimeSpan.FromSeconds(2))  // Initial state
        .Down()
        .WaitUntil(s => s.ContainsText("> Second"), TimeSpan.FromSeconds(2)) // After action
        .Capture("after_down")
        .Ctrl().Key(Hex1bKey.C)
        .Build()
        .ApplyWithCaptureAsync(terminal, TestContext.Current.CancellationToken);
    
    await runTask;
    
    // This assertion is technically redundant since WaitUntil verified it
    Assert.IsTrue(snapshot.ContainsText("> Second"));
}
Pattern B: Render Count Test (No Snapshot Assertions)
csharp
[TestMethod]
public async Task StaticWidget_OnlyRendersOnce()
{
    using var workload = new Hex1bAppWorkloadAdapter();
    using var terminal = Hex1bTerminal.CreateBuilder()
        .WithWorkload(workload)
        .WithHeadless()
        .Build();
    
    var renderCount = 0;
    var staticWidget = new TestWidget().OnRender(_ => renderCount++);

    using var app = new Hex1bApp(
        ctx => Task.FromResult<Hex1bWidget>(staticWidget),
        new Hex1bAppOptions { WorkloadAdapter = workload }
    );

    var runTask = app.RunAsync(TestContext.Current.CancellationToken);
    
    await new Hex1bTerminalInputSequenceBuilder()
        .Key(Hex1bKey.A)
        .Key(Hex1bKey.B)
        .Capture("final")
        .Ctrl().Key(Hex1bKey.C)
        .Build()
        .ApplyWithCaptureAsync(terminal, TestContext.Current.CancellationToken);
    
    await runTask;
    
    // Assert on counter, not snapshot content
    Assert.AreEqual(1, renderCount);
}
Pattern C: Unit Test Without App Exit
csharp
[TestMethod]
public async Task Render_AllItems_AreVisible()
{
    using var workload = new Hex1bAppWorkloadAdapter();
    using var terminal = Hex1bTerminal.CreateBuilder()
        .WithWorkload(workload)
        .WithHeadless()
        .Build();
    
    var context = CreateContext(workload);
    var node = new ListNode { Items = ["Item 1", "Item 2", "Item 3"] };
    node.Arrange(new Rect(0, 0, 20, 5));
    node.Render(context);
    
    // No app, no Ctrl+C needed
    var snapshot = await new Hex1bTerminalInputSequenceBuilder()
        .WaitUntil(s => 
            s.ContainsText("Item 1") && 
            s.ContainsText("Item 2") && 
            s.ContainsText("Item 3"), 
            TimeSpan.FromSeconds(2))
        .Capture("final")
        .Build()
        .ApplyWithCaptureAsync(terminal, TestContext.Current.CancellationToken);
    
    Assert.IsTrue(snapshot.ContainsText("Item 1"));
    Assert.IsTrue(snapshot.ContainsText("Item 2"));
    Assert.IsTrue(snapshot.ContainsText("Item 3"));
}

Pattern 4: Task.WhenAny Race Condition

Symptoms
  • Test uses Task.WhenAny to check if app exited
  • Test sometimes times out even though behavior is correct
  • Inconsistent pass/fail across runs
Root Cause

Using Task.WhenAny(runTask, Task.Delay(...)) creates a race condition. The delay may win even when the app is about to exit, or the exit may happen but not be detected in time.

Example: Broken Test
csharp
// ❌ BROKEN: Race condition with Task.WhenAny
var runTask = app.RunAsync(TestContext.Current.CancellationToken);

await renderTest.Task.WaitAsync(TimeSpan.FromSeconds(1), TestContext.Current.CancellationToken);

await new Hex1bTerminalInputSequenceBuilder()
    .Key(Hex1bKey.C, Hex1bModifiers.Control)
    .Build()
    .ApplyAsync(terminal);

var completed = await Task.WhenAny(runTask, Task.Delay(2000));
Assert.IsTrue(completed == runTask, "Expected CTRL-C to exit");  // Flaky!
Fix
csharp
// ✅ FIXED: Use WaitAsync with timeout instead of Task.WhenAny
var runTask = app.RunAsync(TestContext.Current.CancellationToken);

await new Hex1bTerminalInputSequenceBuilder()
    .WaitUntil(s => s.ContainsText("expected content"), TimeSpan.FromSeconds(2))
    .Ctrl().Key(Hex1bKey.C)
    .Build()
    .ApplyAsync(terminal, TestContext.Current.CancellationToken);

// Use WaitAsync - throws TimeoutException if it takes too long
await runTask.WaitAsync(TimeSpan.FromSeconds(5), TestContext.Current.CancellationToken);

Pattern 5: Test Interference

Symptoms
  • Test passes when run in isolation: dotnet test --filter "FullyQualifiedName~TestName"
  • Test fails when run with full suite
  • Failures appear random or depend on test execution order
Root Cause

Tests may interfere with each other through:

  1. Shared static state - Singletons, static fields
  2. File system conflicts - Temp files with same names
  3. Port conflicts - Multiple tests binding to same port
  4. Resource exhaustion - Too many terminals/processes open
Diagnosis
powershell
# Run test in isolation
dotnet test --filter "FullyQualifiedName~TestName" --no-build

# Run test with suspected interfering tests
dotnet test --filter "FullyQualifiedName~TestName|FullyQualifiedName~OtherTest" --no-build
Fix Strategies
  1. Use unique temp paths - Include test name or GUID in temp file paths
  2. Properly dispose resources - Ensure using statements on terminals/workloads
  3. Avoid static state - Use instance fields or dependency injection
  4. Add [DoNotParallelize] on the test class - tests parallelize by default, and this forces sequential execution for conflicting tests
csharp
// Use unique temp file names
var tempPath = Path.Combine(Path.GetTempPath(), $"hex1b_test_{Guid.NewGuid()}.cast");

// Or include test name
var tempPath = Path.Combine(Path.GetTempPath(), $"hex1b_{nameof(MyTestMethod)}.cast");

Show full SKILL.md (524 more words)Show less

Pattern 6: Platform-Specific Failures

Symptoms
  • Test always fails on Windows but passes on Linux (or vice versa)
  • Error mentions platform-specific features (PTY, console mode, etc.)
  • Test timeout with empty terminal state on one platform
Common Platform Differences
FeatureWindowsLinux
PTY supportLimited (ConPTY)Native
Terminal buffer cleanupDelayedImmediate
File lockingStrictFlexible
Path separators\/
Fix: Add Platform Skip Category or Ignore
csharp
// Skip on Windows - PTY tests require Unix
[TestMethod]
[TestCategory("LinuxOnly")]
public async Task WithPtyProcess_ExecutesProcess()
{
    if (RuntimeInformation.IsOSPlatform(OSPlatform.Windows))
    {
        // Skip on Windows - PTY not fully supported
        return;
    }
    // ... test code
}

// Or use Ignore attribute
[TestMethod]
[Ignore("Requires Unix PTY support")]
public async Task PtyTest() { }

// Or conditional skip with MSTest
public sealed class LinuxOnlyTestMethodAttribute : TestMethodAttribute
{
    public override async Task<TestResult[]> ExecuteAsync(ITestMethod testMethod)
    {
        if (RuntimeInformation.IsOSPlatform(OSPlatform.Windows))
        {
            return [new TestResult
            {
                Outcome = UnitTestOutcome.Inconclusive,
                TestFailureException = new AssertInconclusiveException("This test requires Linux")
            }];
        }

        return await base.ExecuteAsync(testMethod);
    }
}
Tests Known to Require Linux
  • Hex1bTerminalBuilderTests.WithPtyProcess_* - Require PTY
  • NanoExploratoryTests.* - Require nano and PTY
  • Tests using WithPtyProcess() builder method

Pattern 7: Task.Delay for Async Events

Symptoms
  • Test passes locally most of the time
  • Test fails intermittently on CI, especially on slower runners
  • Test waits a fixed time (e.g., Task.Delay(100)) for an async event like input binding firing
  • Error message suggests the action never occurred (e.g., "Expected binding to fire")
Root Cause

Using Task.Delay to wait for async events like input bindings firing is unreliable. The fixed delay may not be long enough on slower CI runners, or may be unnecessarily long on fast machines.

Example: Input binding tests that send a key and wait for the binding callback to fire.

Example: Broken Test
csharp
// ❌ BROKEN: Fixed delay may not be long enough on slow CI runners
var bindingFired = false;

using var app = new Hex1bApp(
    ctx =>
    {
        var vstack = new VStackWidget([ctx.Test()])
            .InputBindings(bindings =>
            {
                bindings.Shift().Key(key).Action(_ =>
                {
                    bindingFired = true;  // Sets flag when binding fires
                    return Task.CompletedTask;
                }, $"Test Shift+{key}");
            });

        return Task.FromResult<Hex1bWidget>(vstack);
    },
    new Hex1bAppOptions { WorkloadAdapter = workload }
);

// ... wait for render ...

await new Hex1bTerminalInputSequenceBuilder()
    .Shift().Key(key)
    .Build()
    .ApplyAsync(terminal, TestContext.Current.CancellationToken);

await Task.Delay(100);  // ❌ Fixed delay - may not be long enough!

Assert.IsTrue(bindingFired);  // Flaky!
Fix

Replace the boolean flag and Task.Delay with a TaskCompletionSource that signals when the event occurs:

csharp
// ✅ FIXED: Use TaskCompletionSource to wait for the async event
var bindingFired = new TaskCompletionSource(TaskCreationOptions.RunContinuationsAsynchronously);

using var app = new Hex1bApp(
    ctx =>
    {
        var vstack = new VStackWidget([ctx.Test()])
            .InputBindings(bindings =>
            {
                bindings.Shift().Key(key).Action(_ =>
                {
                    bindingFired.TrySetResult();  // Signal completion
                    return Task.CompletedTask;
                }, $"Test Shift+{key}");
            });

        return Task.FromResult<Hex1bWidget>(vstack);
    },
    new Hex1bAppOptions { WorkloadAdapter = workload }
);

// ... wait for render ...

await new Hex1bTerminalInputSequenceBuilder()
    .Shift().Key(key)
    .Build()
    .ApplyAsync(terminal, TestContext.Current.CancellationToken);

// Wait for the event with a timeout - will fail fast if binding doesn't fire
await bindingFired.Task.WaitAsync(TimeSpan.FromSeconds(2), TestContext.Current.CancellationToken);

// If we got here, the binding fired (the wait would have timed out otherwise)
Assert.IsTrue(bindingFired.Task.IsCompleted);  // Reliable
Key Points
  1. Use TaskCompletionSource instead of a boolean flag
  2. Call TrySetResult() in the event handler to signal completion
  3. Use WaitAsync with timeout instead of Task.Delay
  4. Use TaskCreationOptions.RunContinuationsAsynchronously to avoid potential deadlocks
  5. The timeout should be generous (2+ seconds) to account for slow CI runners

Pattern 8: Test Helper Partial Wait

Symptoms
  • Tests using shared helper methods that write multiple lines fail intermittently
  • Assertions fail for content on lines other than the first
  • Helper waits for initial content but snapshot misses later content
  • Test name often includes directional movement (Up, Down) or multi-line content
Root Cause

Test helper methods that write multiple lines of content may only wait for the first line to appear before taking a snapshot. On faster CI systems, the snapshot may be captured before all lines are processed by the output pump.

Example: A helper writes lines ["A", "B"] but only waits for "A" to appear. Tests that expect "B" on a separate line may fail because the snapshot is taken before "B" is processed.

Example: Broken Helper
csharp
// ❌ BROKEN: Only waits for first line
private static async Task<Hex1bTerminalSnapshot> CreateSnapshotAsync(string[] lines)
{
    using var workload = new Hex1bAppWorkloadAdapter();
    using var terminal = Hex1bTerminal.CreateBuilder().WithWorkload(workload).Build();

    foreach (var line in lines)
    {
        workload.Write(line + "\r\n");
    }

    // BUG: Only waits for first line!
    var firstLine = lines.Length > 0 ? lines[0] : "";
    await new Hex1bTerminalInputSequenceBuilder()
        .WaitUntil(s => s.ContainsText(firstLine), TimeSpan.FromSeconds(1))
        .Build()
        .ApplyAsync(terminal);

    return terminal.CreateSnapshot();  // May miss content after first line!
}
Fix
csharp
// ✅ FIXED: Wait for last line to ensure all content is processed
private static async Task<Hex1bTerminalSnapshot> CreateSnapshotAsync(string[] lines)
{
    using var workload = new Hex1bAppWorkloadAdapter();
    using var terminal = Hex1bTerminal.CreateBuilder().WithWorkload(workload).Build();

    foreach (var line in lines)
    {
        workload.Write(line + "\r\n");
    }

    // Wait for the LAST line to ensure all lines are written
    var lastLine = lines.Length > 0 ? lines[^1] : "";
    await new Hex1bTerminalInputSequenceBuilder()
        .WaitUntil(s => string.IsNullOrEmpty(lastLine) || s.ContainsText(lastLine), 
                   TimeSpan.FromSeconds(1), "last line content")
        .Build()
        .ApplyAsync(terminal);

    return terminal.CreateSnapshot();  // Now contains all lines
}
How to Identify
  1. Look for test helpers that write multiple pieces of content (arrays, lists)
  2. Check if the WaitUntil condition only checks for initial/first content
  3. Tests affected often have names suggesting multi-line or positional patterns

Checklist for Test Review

Before committing test changes, verify:

  • Every action that changes UI state is followed by WaitUntil
  • WaitUntil condition matches what will be asserted
  • Capture is placed after WaitUntil, before Ctrl+C
  • Snapshot assertions don't rely on post-exit terminal state
  • Consider if assertion is redundant (WaitUntil already verified it)
  • Test runs reliably 10+ times locally
  • Test doesn't have timing dependencies (uses WaitUntil, not Wait)
  • No Task.WhenAny race conditions - use WaitAsync instead
  • No Task.Delay for async events - use TaskCompletionSource instead
  • Test passes both in isolation and in full suite
  • Platform-specific tests have appropriate TestCategory, Ignore, or conditional TestMethodAttribute handling
  • Test helpers writing multiple lines wait for all/last content

© mitchdenny, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .github/skills/test-fixer of mitchdenny/hex1b.

Open the folder on GitHubat commit 98d8766

Compare with similar skills

Test Fixer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Fixer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Fixer this skillmitchdenny/hex1b178—~6.5kAutomated safety check: PassMIT
Find Untested Sourcesdotnet/skills5.6k1 repos~3.3kAutomated safety check: PassMIT
ScottPlot Test RunnerScottPlot/ScottPlot6.8k—~308Automated safety check: PassMIT
Raven Test Triagemarinasundstrom/raven108—~1.4kAutomated safety check: PassMIT
Run Testsrunceel/ReactiveProperty944—~3.6kAutomated safety check: PassMIT
MAUI UI Test Writerdotnet/maui23k—~3kAutomated safety check: PassMIT

Similar skills

  • Official

    Statically pairs source files with test files to list code that no test references, using Roslyn for C# or tree-sitter for many languages, with no build.

    5.6k GitHub starsUsed in 1 repo~3.3k tokens
    Testing & QAAuto-check passed
  • ScottPlot Test Runner

    ScottPlot/ScottPlot

    Run or add ScottPlot 5 tests. Use for unit-test and cookbook-test work; unless explicitly asked otherwise, restrict manual test execution to the Unit Tests…

    6.8k GitHub stars~308 tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Raven Test Triage

    marinasundstrom/raven

    Testing and stabilization workflow for the Raven compiler test suite.

    108 GitHub stars~1.4k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Run Tests

    runceel/ReactiveProperty

    Runs .NET tests with dotnet test. An agent skill from runceel/ReactiveProperty.

    944 GitHub stars~3.6k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Official

    Writes UI tests that reproduce a GitHub issue in .NET MAUI and keeps iterating until the tests actually fail, proving they catch the bug.

    23k GitHub stars~3k tokensUpdated today
    Testing & QAAuto-check passed
  • Official

    Writes a paired .xaml and .xaml.cs unit test for a .NET MAUI issue that is about XAML behavior itself, such as parsing, XamlC output or generated code.

    23k GitHub stars~742 tokensUpdated today
    Testing & QAAuto-check passed

More from mitchdenny/hex1b

  • API Reviewer

    mitchdenny/hex1b

    Guidelines for reviewing API design in the Hex1b codebase. An agent skill from mitchdenny/hex1b.

    178 GitHub stars~4k tokensUpdated 6 days ago
    Auto-check passed
  • Surface Benchmarker

    mitchdenny/hex1b

    Guidelines for running and interpreting Surface API performance benchmarks.

    178 GitHub stars~3.1k tokensUpdated 6 days ago
    Auto-check passed
  • Doc Tester

    mitchdenny/hex1b

    Agent for validating Hex1b documentation against actual library behavior.

    178 GitHub stars~9.5k tokensUpdated 6 days ago
    Auto-check passed
  • Doc Writer

    mitchdenny/hex1b

    Guidelines for producing accurate and maintainable documentation for the Hex1b TUI library.

    178 GitHub stars~8.7k tokensUpdated 6 days ago
    Auto-check passed
  • Widget Creator

    mitchdenny/hex1b

    Step-by-step guide for creating new widgets in the Hex1b TUI library.

    178 GitHub stars~7.6k tokensUpdated 6 days ago
    Auto-check passed
  • Writing Unit Tests

    mitchdenny/hex1b

    Guidelines for writing unit tests in the Hex1b TUI library. An agent skill from mitchdenny/hex1b.

    178 GitHub stars~6.6k tokensUpdated 6 days ago
    Auto-check passed

Works with

Categories

Questions about Test Fixer

What does Test Fixer do?

Agent for diagnosing and fixing flaky terminal UI tests in the Hex1b test suite. Test Fixer is an agent skill from mitchdenny/hex1b. Agent for diagnosing and fixing flaky terminal UI tests in the Hex1b test suite.

When should I use Test Fixer?

Test Fixer fits situations like: tests pass locally but fail in CI; tests exhibit timing-sensitive behavior.

How do I install Test Fixer in Claude Code?

Run `npx skills add mitchdenny/hex1b --skill test-fixer -a claude-code`. Or copy the skill folder (.github/skills/test-fixer in mitchdenny/hex1b) into .claude/skills/test-fixer in your project. Claude Code loads it when a task matches its description.

How do I install Test Fixer in Codex?

Run `npx skills add mitchdenny/hex1b --skill test-fixer -a codex`. Or copy the skill folder (.github/skills/test-fixer in mitchdenny/hex1b) into .agents/skills/test-fixer in your project. Codex loads it when a task matches its description.

Can I use Test Fixer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mitchdenny/hex1b --skill test-fixer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-fixer, .gemini/skills/test-fixer, .github/skills/test-fixer and .opencode/skills/test-fixer in your project.

What does Test Fixer need to run?

Going by SKILL.md and its folder, Test Fixer needs the command-line tools its instructions call (dotnet and gh).

Does Test Fixer access the network?

SKILL.md contains no URLs. Its commands use gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Test Fixer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Fixer use?

Test Fixer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Fixer use?

About 6.5k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test Fixer?

Skills that share tags, products or a category with Test Fixer: Find Untested Sources (dotnet/skills, 5.6k stars), ScottPlot Test Runner (ScottPlot/ScottPlot, 6.8k stars), Raven Test Triage (marinasundstrom/raven, 108 stars) and Run Tests (runceel/ReactiveProperty, 944 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Fixer?

mitchdenny (a GitHub user) maintains it in mitchdenny/hex1b, which has 178 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 2, 2026.

Source: mitchdenny/hex1b on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.