Agent skill

Surface Benchmarker

by mitchdenny in mitchdenny/hex1b

Guidelines for running and interpreting Surface API performance benchmarks.

MITAuto-check passed

Install Surface Benchmarker

skills CLI
$ npx skills add mitchdenny/hex1b --skill surface-benchmarker -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mitchdenny/hex1b surface-benchmarker --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mitchdenny/hex1b.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/surface-benchmarker .claude/skills/surface-benchmarker && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
surface-benchmarker
GitHub stars
178
Token cost
~3.1k tokens
SKILL.md length
796 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Guidelines for running and interpreting Surface API performance benchmarks.

  • Works in 4 steps: Warms up the JIT by running the… → Runs multiple iterations to get… → Reports mean, median, standard deviation… → …
  • Modifying code in src/Hex1b/Surfaces/ to ensure performance is not regressed
  • SKILL.md covers Quick Reference, When to Run Benchmarks, Benchmark Harness Architecture and Benchmark Categories, plus 5 more sections
  • Calls dotnet

What it does

Surface Benchmarker is an agent skill from mitchdenny/hex1b. Guidelines for running and interpreting Surface API performance benchmarks. Use when modifying code in src/Hex1b/Surfaces/ to ensure performance is not regressed.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It works with .NET. The repository describes itself as: The .NET Terminal Application Stack. The licence is MIT.

When your agent uses it

  • Modifying code in src/Hex1b/Surfaces/ to ensure performance is not regressed

Example prompts

  • “Use the surface-benchmarker skill to guideline for running and interpreting Surface API performance benchmarks”
  • “/surface-benchmarker”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Warms up the JIT by running the benchmark multiple times before measuring
  2. Runs multiple iterations to get statistically significant results
  3. Reports mean, median, standard deviation for each benchmark
  4. Tracks memory allocations (when [MemoryDiagnoser] is enabled)

What it can do on your machine

Read from SKILL.md and the folder at commit 98d8766. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • dotnet

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Surface Benchmarker loads about 3.1k tokens when it runs. Until then it costs about 46 tokens; SKILL.md has 796 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mitchdenny/hex1b at commit 98d8766, republished under its MIT licence (© mitchdenny). 796 words, ~3,056 tokens.

Download SKILL.mdSave it as .claude/skills/surface-benchmarker/SKILL.md (or your agent's skills folder).
name
surface-benchmarker
description
Guidelines for running and interpreting Surface API performance benchmarks. Use when modifying code in src/Hex1b/Surfaces/ to ensure performance is not regressed.

Surface Benchmarker Skill

This skill provides guidelines for AI agents to run and interpret performance benchmarks for the Hex1b Surface API. Run benchmarks whenever you modify code in src/Hex1b/Surfaces/ to ensure performance is not regressed.

Quick Reference

ActionCommand
Run all Surface benchmarksdotnet run -c Release --project benchmarks/Hex1b.Benchmarks -- --filter "Surface*"
Run specific benchmarkdotnet run -c Release --project benchmarks/Hex1b.Benchmarks -- --filter "*WriteText*"
Quick dry-run (sanity check)dotnet run -c Release --project benchmarks/Hex1b.Benchmarks -- --filter "Surface*" --job dry

When to Run Benchmarks

ALWAYS run benchmarks after modifying:

File/AreaCritical Benchmarks
SurfaceCell.csAll (cells are used everywhere)
Surface.csCreateSurface_*, WriteText_*, Fill_*, Clone_*
CompositeSurface.csCompositeSurface_*
SurfaceComparer.csCompare_*, ToTokens_*, ToAnsiString_*
SurfaceDiff.csCompare_*
ComputeContext.csCompositeSurface_* (computed cells use this)

Benchmark Harness Architecture

Project Structure
benchmarks/Hex1b.Benchmarks/
├── Hex1b.Benchmarks.csproj    # Console app with BenchmarkDotNet
├── Program.cs                  # Entry point
└── SurfaceBenchmarks.cs        # Surface API benchmarks
How BenchmarkDotNet Works

BenchmarkDotNet is a .NET performance benchmarking library that:

  1. Warms up the JIT by running the benchmark multiple times before measuring
  2. Runs multiple iterations to get statistically significant results
  3. Reports mean, median, standard deviation for each benchmark
  4. Tracks memory allocations (when [MemoryDiagnoser] is enabled)
Benchmark Lifecycle
┌─────────────────────────────────────────────────────────────────┐
│ 1. GlobalSetup                                                  │
│    - Runs once before all benchmarks                            │
│    - Creates pre-allocated surfaces, diffs, etc.                │
│    - Sets up shared test data                                   │
└─────────────────────────────────────────────────────────────────┘
                              ▼
┌─────────────────────────────────────────────────────────────────┐
│ 2. For each [Benchmark] method:                                 │
│    a. Warmup phase (JIT compilation, cache warming)             │
│    b. Pilot phase (determine optimal iteration count)           │
│    c. Actual measurements (multiple iterations)                 │
│    d. Results aggregation                                       │
└─────────────────────────────────────────────────────────────────┘
                              ▼
┌─────────────────────────────────────────────────────────────────┐
│ 3. Report Generation                                            │
│    - Console output with statistics                             │
│    - Optional: HTML, CSV, JSON exports                          │
└─────────────────────────────────────────────────────────────────┘
Key Benchmark Attributes
AttributePurpose
[MemoryDiagnoser]Reports memory allocations (Gen0/1/2 collections, allocated bytes)
[Benchmark]Marks a method as a benchmark
[GlobalSetup]Runs once before all benchmarks to set up shared state
[Params(80, 160, 320)]Parameterizes benchmarks with multiple values

Benchmark Categories

Surface Creation (CreateSurface_*)

Measures allocation and initialization cost for surfaces of various sizes.

csharp
[Benchmark]
public Surface CreateSurface_Small() => new Surface(80, 24);    // Standard terminal
public Surface CreateSurface_Medium() => new Surface(160, 48);  // Large terminal
public Surface CreateSurface_Large() => new Surface(320, 96);   // Very large
public Surface CreateSurface_4K() => new Surface(480, 135);     // 4K equivalent

What to watch for:

  • Memory allocation should scale linearly with cell count
  • No unexpected allocations from internal data structures
Text Writing (WriteText_*)

Measures grapheme parsing and wide character handling.

csharp
[Benchmark]
public void WriteText_Short();      // "Hello"
public void WriteText_Medium();     // Standard sentence
public void WriteText_Long();       // Long paragraph
public void WriteText_WideChars();  // Chinese characters (2 cells each)
public void WriteText_FillScreen(); // 24 lines of text

What to watch for:

  • Wide characters should not be drastically slower than ASCII
  • FillScreen should scale linearly
Fill Operations (Fill_*)

Measures bulk cell assignment.

csharp
[Benchmark]
public void Fill_SmallRect();   // 20x10 region
public void Fill_FullScreen();  // 80x24
public void Fill_LargeScreen(); // 320x96

What to watch for:

  • Should approach memory bandwidth limits for large fills
  • Minimal overhead per cell
Compositing (Composite_*, CompositeSurface_*)

Measures layer merging and transparency resolution.

csharp
[Benchmark]
public void Composite_SmallOntoSmall();         // 20x10 onto 80x24
public void Composite_MediumOntoLarge();        // 80x24 onto 320x96
public Surface CompositeSurface_Flatten_3Layers();  // Resolve 3 layers
public SurfaceCell CompositeSurface_GetCell_Resolved();  // Single cell lookup
public void CompositeSurface_GetAllCells();     // 80x24 = 1920 lookups

What to watch for:

  • Flatten should be O(layers × cells)
  • Single cell lookup should be fast for interactive rendering
Diff Comparison (Compare_*)

Measures change detection.

csharp
[Benchmark]
public SurfaceDiff Compare_FullDiff();    // 100% of cells changed
public SurfaceDiff Compare_SparseDiff();  // ~10% of cells changed
public SurfaceDiff Compare_NoDiff();      // 0% changed
public SurfaceDiff CompareToEmpty();      // Initial render

What to watch for:

  • NoDiff should be very fast (early exit when possible)
  • Memory allocation should be proportional to changed cell count
Token Generation (ToTokens_*, ToAnsiString_*)

Measures ANSI escape sequence generation.

csharp
[Benchmark]
public IReadOnlyList<AnsiToken> ToTokens_FullDiff();  // Generate tokens
public string ToAnsiString_FullDiff();                // Tokens → string
public string ToAnsiString_SparseDiff();              // Fewer changes

What to watch for:

  • String building should be efficient
  • SGR optimization should reduce output size

Running Benchmarks

Full Benchmark Suite
bash
cd /home/midenn/Code/hex1b
dotnet run -c Release --project benchmarks/Hex1b.Benchmarks -- --filter "Surface*"

Expected runtime: 5-15 minutes depending on hardware.

Quick Validation (Dry Run)
bash
dotnet run -c Release --project benchmarks/Hex1b.Benchmarks -- --filter "Surface*" --job dry

This runs fewer iterations to quickly verify benchmarks work without waiting for full statistical accuracy.

Specific Benchmark Groups
bash
# Only WriteText benchmarks
dotnet run -c Release --project benchmarks/Hex1b.Benchmarks -- --filter "*WriteText*"

# Only Diff/Token generation
dotnet run -c Release --project benchmarks/Hex1b.Benchmarks -- --filter "*Compare*|*Token*|*Ansi*"

# Single benchmark
dotnet run -c Release --project benchmarks/Hex1b.Benchmarks -- --filter "*Fill_FullScreen*"

Interpreting Results

Reading the Output
| Method                | Mean       | Error    | StdDev   | Gen0   | Allocated |
|-----------------------|------------|----------|----------|--------|-----------|
| CreateSurface_Small   | 12.34 μs   | 0.15 μs  | 0.14 μs  | 1.2345 | 15.36 KB  |
| CreateSurface_Large   | 98.76 μs   | 1.23 μs  | 1.15 μs  | 9.8765 | 122.88 KB |
ColumnMeaning
MeanAverage execution time
ErrorHalf of 99.9% confidence interval
StdDevStandard deviation (lower = more consistent)
Gen0Gen0 garbage collections per 1000 operations
AllocatedBytes allocated per operation
Performance Expectations
BenchmarkExpected RangeAlert If
CreateSurface_Small (80×24)5-20 μs> 50 μs
CreateSurface_4K (480×135)50-200 μs> 500 μs
WriteText_Short0.5-2 μs> 5 μs
WriteText_WideChars1-5 μs> 10 μs
Fill_FullScreen2-10 μs> 25 μs
Compare_NoDiff5-20 μs> 50 μs
Compare_SparseDiff10-50 μs> 100 μs
ToAnsiString_SparseDiff20-100 μs> 250 μs

Note: These are rough guidelines. Actual values depend on hardware.

Show full SKILL.md (296 more words)Show less
Comparing Before/After

When making changes:

  1. Before changes: Run benchmarks and save output
  2. Make changes
  3. After changes: Run benchmarks again
  4. Compare: Look for significant regressions (>20% slower)
bash
# Save before results
dotnet run -c Release --project benchmarks/Hex1b.Benchmarks -- \
    --filter "Surface*" --exporters json > before.json

# Make changes...

# Compare (manual inspection)
dotnet run -c Release --project benchmarks/Hex1b.Benchmarks -- \
    --filter "Surface*" --exporters json > after.json

Common Performance Issues

Issue 1: Excessive Allocations

Symptom: High Allocated column, many Gen0 collections.

Common causes:

  • Creating new arrays/lists inside hot loops
  • Boxing value types (e.g., storing SurfaceCell in object)
  • String concatenation without StringBuilder

Fix strategies:

  • Use ArrayPool<T>.Shared for temporary buffers
  • Use Span<T> and stack allocation for small buffers
  • Pre-allocate and reuse collections
Issue 2: Poor Cache Locality

Symptom: Large surfaces are disproportionately slower than small ones.

Common causes:

  • Column-major access pattern on row-major data
  • Jumping around in memory instead of sequential access

Fix strategies:

  • Access cells in row-major order (for y, then for x)
  • Process contiguous memory regions when possible
Issue 3: Redundant Work

Symptom: Operations are slower than expected for the amount of data.

Common causes:

  • Recalculating values that could be cached
  • Not short-circuiting when result is known
  • Unnecessary defensive copies

Fix strategies:

  • Cache computed values (e.g., display width for text)
  • Early exit when no changes detected
  • Use readonly and in parameters to avoid copies

Adding New Benchmarks

When adding new functionality to the Surface API, add corresponding benchmarks:

csharp
[Benchmark]
public void NewFeature_TypicalCase()
{
    // Benchmark the common case
    _surface.NewMethod(typicalInput);
}

[Benchmark]
public void NewFeature_WorstCase()
{
    // Benchmark the worst case for regression detection
    _surface.NewMethod(worstCaseInput);
}
Guidelines for New Benchmarks
  1. Use pre-allocated data in [GlobalSetup] - don't measure setup time
  2. Return or use the result - prevent dead code elimination
  3. Benchmark realistic scenarios - not just micro-benchmarks
  4. Include edge cases - empty inputs, large inputs, worst-case patterns

CI Integration (Future)

Benchmarks can be integrated into CI to catch regressions:

yaml
# .github/workflows/benchmarks.yml (example)
- name: Run benchmarks
  run: |
    dotnet run -c Release --project benchmarks/Hex1b.Benchmarks -- \
      --filter "Surface*" --exporters json
    
- name: Compare with baseline
  run: |
    # Compare against stored baseline, fail if regression > 20%

Checklist for Surface API Changes

Before merging changes to src/Hex1b/Surfaces/:

  • Run relevant benchmarks before and after changes
  • No benchmark regressed by more than 20%
  • Memory allocations did not significantly increase
  • Add benchmarks for any new public methods
  • Document any intentional performance tradeoffs

© mitchdenny, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .github/skills/surface-benchmarker of mitchdenny/hex1b.

Open the folder on GitHubat commit 98d8766

Compare with similar skills

Surface Benchmarker next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Surface Benchmarker compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Surface Benchmarker this skillmitchdenny/hex1b178—~3.1kAutomated safety check: PassMIT
Minimax DOCXpoco-ai/poco-claw1.4k7 repos~3.9kAutomated safety check: PassMIT
Microsoft Skill CreatorMicrosoftDocs/mcp1.9k3 repos~2.1kAutomated safety check: PassCC-BY-4.0
Speckit ConstitutionWeihanLi/WeihanLi.Common24211 repos~2.1kAutomated safety check: PassApache-2.0
Copilot Session Failure Analysisdotnet/maui23k—~3.4kAutomated safety check: PassMIT
Microsoft Code ReferenceMicrosoftDocs/mcp1.9k4 repos~1.1kAutomated safety check: PassCC-BY-4.0

Similar skills

  • Minimax DOCX

    poco-ai/poco-claw

    Professional DOCX document creation, editing, and formatting using OpenXML SDK (.NET).

    1.4k GitHub starsUsed in 7 repos~3.9k tokens
    Documents & OfficeAuto-check passed
  • Microsoft Skill Creator

    MicrosoftDocs/mcp

    Official

    Create agent skills for Microsoft technologies using official documentation.

    1.9k GitHub starsUsed in 3 repos~2.1k tokens
    Agent WorkflowsAuto-check passed
  • Speckit Constitution

    WeihanLi/WeihanLi.Common

    Create or update the project constitution from interactive or provided principle inputs, ensuring all dependent templates stay in sync.

    242 GitHub starsUsed in 11 repos~2.1k tokens
    DevelopmentAuto-check passed
  • Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.

    23k GitHub stars~3.4k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Microsoft Code Reference

    MicrosoftDocs/mcp

    Official

    Find working code samples, verify API signatures, and fix Microsoft SDK errors using official docs.

    1.9k GitHub starsUsed in 4 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Official

    Audits and updates os-packages.json files listing the Linux packages each .NET release needs per distro, then regenerates the Markdown from the JSON.

    22k GitHub stars~2.3k tokensUpdated today
    DevelopmentAuto-check passed

More from mitchdenny/hex1b

  • API Reviewer

    mitchdenny/hex1b

    Guidelines for reviewing API design in the Hex1b codebase. An agent skill from mitchdenny/hex1b.

    178 GitHub stars~4k tokensUpdated 6 days ago
    Auto-check passed
  • Doc Tester

    mitchdenny/hex1b

    Agent for validating Hex1b documentation against actual library behavior.

    178 GitHub stars~9.5k tokensUpdated 6 days ago
    Auto-check passed
  • Doc Writer

    mitchdenny/hex1b

    Guidelines for producing accurate and maintainable documentation for the Hex1b TUI library.

    178 GitHub stars~8.7k tokensUpdated 6 days ago
    Auto-check passed
  • Test Fixer

    mitchdenny/hex1b

    Agent for diagnosing and fixing flaky terminal UI tests in the Hex1b test suite.

    178 GitHub stars~6.5k tokensUpdated 6 days ago
    Auto-check passed
  • Widget Creator

    mitchdenny/hex1b

    Step-by-step guide for creating new widgets in the Hex1b TUI library.

    178 GitHub stars~7.6k tokensUpdated 6 days ago
    Auto-check passed
  • Writing Unit Tests

    mitchdenny/hex1b

    Guidelines for writing unit tests in the Hex1b TUI library. An agent skill from mitchdenny/hex1b.

    178 GitHub stars~6.6k tokensUpdated 6 days ago
    Auto-check passed

Works with

Questions about Surface Benchmarker

What does Surface Benchmarker do?

Guidelines for running and interpreting Surface API performance benchmarks. Surface Benchmarker is an agent skill from mitchdenny/hex1b. Guidelines for running and interpreting Surface API performance benchmarks.

When should I use Surface Benchmarker?

Surface Benchmarker fits situations like: modifying code in src/Hex1b/Surfaces/ to ensure performance is not regressed.

How do I install Surface Benchmarker in Claude Code?

Run `npx skills add mitchdenny/hex1b --skill surface-benchmarker -a claude-code`. Or copy the skill folder (.github/skills/surface-benchmarker in mitchdenny/hex1b) into .claude/skills/surface-benchmarker in your project. Claude Code loads it when a task matches its description.

How do I install Surface Benchmarker in Codex?

Run `npx skills add mitchdenny/hex1b --skill surface-benchmarker -a codex`. Or copy the skill folder (.github/skills/surface-benchmarker in mitchdenny/hex1b) into .agents/skills/surface-benchmarker in your project. Codex loads it when a task matches its description.

Can I use Surface Benchmarker in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mitchdenny/hex1b --skill surface-benchmarker -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/surface-benchmarker, .gemini/skills/surface-benchmarker, .github/skills/surface-benchmarker and .opencode/skills/surface-benchmarker in your project.

What does Surface Benchmarker need to run?

Going by SKILL.md and its folder, Surface Benchmarker needs the command-line tools its instructions call (dotnet).

Does Surface Benchmarker access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Surface Benchmarker safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Surface Benchmarker use?

Surface Benchmarker is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Surface Benchmarker use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Surface Benchmarker?

Skills that share tags, products or a category with Surface Benchmarker: Minimax DOCX (poco-ai/poco-claw, 1.4k stars), Microsoft Skill Creator (MicrosoftDocs/mcp, 1.9k stars), Speckit Constitution (WeihanLi/WeihanLi.Common, 242 stars) and Copilot Session Failure Analysis (dotnet/maui, 23k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Surface Benchmarker?

mitchdenny (a GitHub user) maintains it in mitchdenny/hex1b, which has 178 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 2, 2026.

Source: mitchdenny/hex1b on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.