Official agent skill

Test Guidelines

by getsentry in getsentry/sentry-dart

Enforce Sentry Dart/Flutter SDK test conventions for naming, structure, and fixtures.

OfficialMITAuto-check passedTesting & QA

Install Test Guidelines

skills CLI
$ npx skills add getsentry/sentry-dart --skill test-guidelines -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install getsentry/sentry-dart test-guidelines --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/getsentry/sentry-dart.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/test-guidelines .claude/skills/test-guidelines && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-guidelines
GitHub stars
873
Token cost
~3.1k tokens
SKILL.md length
1,047 words
Files
4 (incl. references)
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Enforce Sentry Dart/Flutter SDK test conventions for naming, structure, and fixtures.

  • Modifying tests
  • SKILL.md covers Test-First Loop, File Structure, Test Naming and Depth Rules, plus 8 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Reviewing test code

What it does

Test Guidelines is an agent skill from getsentry/sentry-dart, published by the product's own GitHub organization. Enforce Sentry Dart/Flutter SDK test conventions for naming, structure, and fixtures. Use when writing tests, adding tests, modifying tests, reviewing test code, fixing failing tests, adding test coverage, TDD, test-first / red-green, reproducing bugs with tests, regression tests, or test refactoring in any package in this Melos monorepo.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/async.md`, `references/fixtures.md` and `references/mocking.md`).

It sits in Testing & QA, covering Test-driven development, Cross-platform mobile apps and Failing and flaky tests. It works with Sentry, Flutter and Dart. The repository describes itself as: Sentry SDK for Dart and Flutter. The licence is MIT.

When your agent uses it

  • Modifying tests
  • Reviewing test code
  • Fixing failing tests
  • Adding test coverage

Example prompts

  • “/test-guidelines”

What it can do on your machine

Read from SKILL.md and the folder at commit 96d6367. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are dart).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Guidelines loads about 3.1k tokens when it runs, and up to ~4.7k if it reads all its reference files. Until then it costs about 89 tokens; SKILL.md has 1,047 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from getsentry/sentry-dart at commit 96d6367, republished under its MIT licence (© getsentry). 1,047 words, ~3,133 tokens.

Download SKILL.mdSave it as .claude/skills/test-guidelines/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
test-guidelines
description
Enforce Sentry Dart/Flutter SDK test conventions for naming, structure, and fixtures. Use when writing tests, adding tests, modifying tests, reviewing test code, fixing failing tests, adding test coverage, TDD, test-first / red-green, reproducing bugs with tests, regression tests, or test refactoring in any package in this Melos monorepo.

Apply these conventions to all new and modified tests across every package in this monorepo. Existing tests may not follow these conventions — do not refactor them unless asked.

Tests are easiest to write against code designed to accept its dependencies — when implementing the code under test, load design-first (where the seams go) and code-guidelines (the rules).

Test-First Loop

Work in vertical slices, not horizontal ones. One failing test → the minimal code that makes it pass → repeat. Each test is a tracer bullet: it proves one thin path end-to-end, and what you learn from it shapes the next.

Do not write all the tests first and then all the implementation. That horizontal slicing produces tests of imagined behavior — they assert the shape you guessed at, pass when the real behavior breaks, and commit you to a structure before you understand it. Write one test at a time, against behavior you can already reason about.

Fixing a bug? Reproduce it with a failing test first — see diagnosing-bugs for the loop.

File Structure

  • One test file per source file. Mirror the source path: lib/src/hub.dart → test/hub_test.dart.
  • Every test file has a single void main() { ... } entry point.
  • Use a single top-level group() matching the class or unit name. When the subject is a class or enum, write it with $ interpolation ('$SentryClient') so the name tracks renames; for other units (top-level functions, extensions) use a plain string — see Style rules.
dart
// GOOD: Mirrors source path, single main, single top-level group
// test/sentry_client_test.dart (for lib/src/sentry_client.dart)
void main() {
  group('$SentryClient', () {
    // all tests here
  });
}

// AVOID: Multiple top-level groups or no group wrapper
void main() {
  group('SentryClient capture', () { });
  group('SentryClient close', () { });
}

Test Naming

Nested group() + test() names MUST read as a sentence when concatenated.

Pattern: [Subject] [Context] [Variant] [Behavior]

DepthRoleStyleExample
Group 1SubjectNounClient
Group 2Contextwhen / in / duringwhen connected
Group 3Variantwith / given / usingwith valid input
TestBehaviorVerb phrasesends message
Style rules
  • Use plain verb phrases: returns scope, throws ArgumentError, sends event to transport.
  • Do not prefix test names with should — prefer direct phrasing.
  • For a class or enum subject group, use $ interpolation in the string (group('$Hub', ...)) so a rename updates the test name. Never pass a bare Type literal — group(Hub, ...) is not allowed.
  • Interpolation only works for classes and enums. Extensions and top-level functions cannot be interpolated ('$MyExtension' is a compile error; '$myFunction' yields a closure string) — use a plain descriptive string there.
dart
// GOOD: Interpolated class subject, plain verb phrase
group('$Hub', () {
  test('returns the scope', () { });
});

// AVOID: "should" prefix, bare Type literal as group name
group(Hub, () {
  test('should return the scope', () { });
});
One-line check

If it doesn't read like a sentence, rename the groups:

dart
// GOOD: "Hub returns the scope" reads as a sentence
group('$Hub', () {
  test('returns the scope', () { });
});

// GOOD: "Hub when capturing sends event to transport"
group('$Hub', () {
  group('when capturing', () {
    test('sends event to transport', () { });
  });
});

// AVOID: Doesn't read as a sentence ("Hub capture test event sent")
group('$Hub', () {
  group('capture test', () {
    test('event sent', () { });
  });
});

Depth Rules

  • Maximum depth: 3 groups. More nesting is a code smell — suggest refactoring the implementation.
  • Fold simple variants into the test name instead of adding a group. Use a group only when multiple tests share the same variant setup.
  • Drop a group whose context applies to every sibling test — it adds depth without discriminating. Move the context into the parent group name or the test names instead.
dart
// GOOD: Simple variant in test name (2 groups + test)
group('$Hub', () {
  group('when bound to client', () {
    test('with valid DSN initializes correctly', () { });
    test('with empty DSN throws ArgumentError', () { });
  });
});

// AVOID: Unnecessary nesting (3 groups + test)
group('$Hub', () {
  group('when bound to client', () {
    group('with valid DSN', () {
      test('initializes correctly', () { });
    });
  });
});
dart
// GOOD: Behavior carries the context; no redundant wrapper
group('$SentryAttribute', () {
  test('string serializes value with string type', () { });
  test('int serializes value with integer type', () { });
});

// AVOID: Wrapper context true of every test — pure noise
group('$SentryAttribute', () {
  group('when serializing to JSON', () {            // every test serializes
    test('string serializes value with string type', () { });
    test('int serializes value with integer type', () { });
  });
});

Negative Tests

Use clear verb phrases indicating absence or failure:

PatternWhenExample
does not <verb>Behavior intentionally skippeddoes not send event
throws <ExceptionType>Expecting an exceptionthrows ArgumentError
returns nullNull result expectedreturns null when missing
ignores <thing>Input deliberately ignoredignores empty breadcrumbs
dart
// GOOD: Clear verb phrases indicating absence or failure
group('$Client', () {
  group('when disabled', () {
    test('does not send events', () { });
    test('returns null for captureEvent', () { });
    test('throws StateError', () { });
  });
});

// AVOID: Vague negations or "should not" phrasing
group('$Client', () {
  group('when disabled', () {
    test('should not work', () { });
    test('fails', () { });
    test('no events', () { });
  });
});

Fixtures and Setup

Encapsulate setup in a Fixture class at the bottom of each test file, exposing a getSut() that builds the System Under Test with configurable, injectable dependencies. Initialize it in setUp() within the narrowest group that needs it. Always build options via defaultTestOptions() from test_utils.dart — never construct SentryOptions directly.

dart
class Fixture {
  final transport = MockTransport();
  final options = defaultTestOptions();

  SentryClient getSut({bool attachStacktrace = true}) {
    options.attachStacktrace = attachStacktrace;
    options.transport = transport;
    return SentryClient(options);
  }
}

Full rules — Fixture placement, setUp/tearDown scoping, setUpAll caveats, and the defaultTestOptions() rule — in references/fixtures.md.

Show full SKILL.md (487 more words)Show less

What to Test

Test the behavior owned by your change.

Prefer tests that would fail if your change's intended contract were broken: user-visible behavior, public API behavior, meaningful branching logic, data transformations, integration wiring, precedence rules, error handling, and regressions your change could realistically introduce.

Avoid tests that merely re-prove guarantees owned somewhere else, such as a shared helper, base class, framework, serializer, collection type, generated model, or value object that already has focused coverage. A caller test should not exist just to show that its dependencies still work.

Before adding a test, ask:

  • What behavior would fail if my change were wrong?
  • Is this contract owned by this code, or by something it delegates to?
  • Would this test catch a plausible regression in this change?
  • Is this asserting an outcome, or just mirroring implementation details?

Do test delegated behavior when the delegation is load-bearing for your change's own contract. For example, preserving user input, choosing precedence between sources, wiring the correct helper, enforcing a public API promise, or covering a past regression can all deserve caller-level tests even if a helper implements part of the behavior.

Good tests make the intended contract harder to break. Noisy tests make refactors harder without improving confidence.

dart
// GOOD: asserts the new behavior this code path introduces
test('adds sentry.trace_lifecycle stream attribute', () async {
  final span = fixture.createRecordingSpan();
  await fixture.pipeline.captureSpan(span, scope: fixture.scope);
  expect(span.attributes[SemanticAttributesConstants.sentryTraceLifecycle]?.value, 'stream');
});

// AVOID in this feature's tests: re-proves that SentryAttribute.string
// stores its value, which is the value object's own contract
test('SentryAttribute.string stores its value', () {
  final attribute = SentryAttribute.string('value');
  expect(attribute.value, 'value');
});

Assertions

  • Use expect() with matchers from package:test.
  • Prefer specific matchers (throwsArgumentError, isA<SentryException>()) over generic ones (throwsException, isA<Exception>()).
  • One logical assertion per test. Multiple expect() calls are fine if they verify a single behavior.
  • Avoid the tautological test. Assert the literal expected value, not the same constant the production code uses to produce it — a test whose expected value is computed the way the code computes it passes by construction and can never disagree with the code. Sharing one constant across production and test makes the assertion tautological: it still passes if the constant holds the wrong value. Using the constant as the lookup key is fine; pin the expected value as a literal.
dart
// GOOD: Specific matchers, single logical assertion
test('captures exception with stacktrace', () {
  expect(event.exceptions, hasLength(1));
  expect(event.exceptions!.first, isA<SentryException>());
  expect(event.exceptions!.first.stackTrace, isNotNull);
});

// AVOID: Generic matchers, testing unrelated behaviors
test('captures exception', () {
  expect(event.exceptions, isNotNull); // too vague
  expect(event.exceptions!.first, isA<Object>()); // too generic
  expect(event.breadcrumbs, isEmpty); // unrelated assertion
});
dart
// GOOD: pin the expected value as a literal
expect(span.data[SentryDatabase.dbSystemKey], 'sqlite');

// AVOID (tautological): asserting against the same constant the production code uses to set it
expect(span.data[SentryDatabase.dbSystemKey], SentryDatabase.dbSystem);

Mocking

Prefer fakes over mocks — hand-written implementations that capture state, resilient to refactoring and readable as documentation. Reach for a mock only when faking a large third-party interface isn't worth it. A test that's hard to fake usually signals the code under test should accept its dependencies rather than construct them (a design-first concern).

Full guidance, including designing for mockability (dependency injection, SDK-style interfaces, mocking only at real boundaries), in references/mocking.md.

Async

Never fire-and-forget: return the Future or mark the callback async. Use expectLater with stream matchers for streams, and fakeAsync for timer/microtask-dependent code. Examples in references/async.md.

General

  • Keep tests deterministic. No reliance on real clocks, network, or filesystem unless writing an integration test.
  • Do not duplicate test utilities. If you need test utilities across different packages then add them to packages/_sentry_testing
dart
// GOOD: Deterministic clock
final clock = DateTime.utc(2024, 1, 15, 12, 0, 0);
options.clock = () => clock;

// AVOID: Real clock — flaky on slow CI
final now = DateTime.now();
expect(event.timestamp!.difference(now).inSeconds, lessThan(1));

Integration / E2E Tests

  • Integration tests live in packages/flutter/example/integration_test.
  • JNI and FFI bindings cannot be mocked or faked — integration tests are required when working with native interop.
  • Prefer integration tests for any behavior that depends on native platform APIs.

© getsentry, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in .agents/skills/test-guidelines of getsentry/sentry-dart.

  • SKILL.md
  • references/async.md
  • references/fixtures.md
  • references/mocking.md

Open the folder on GitHubat commit 96d6367

Compare with similar skills

Test Guidelines next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Guidelines compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Guidelines this skillgetsentry/sentry-dart873—~3.1kAutomated safety check: PassMIT
Test Guidelinesgetsentry/sentry-react-native1.8k—~1.3kAutomated safety check: PassMIT
Green GateVeryGoodOpenSource/vgv-ai-flutter-plugin169—~4.7kAutomated safety check: NotesMIT
Test Driven Developmentaiskillstore/marketplace430—~1.9kAutomated safety check: PassNone
Flutter TestingMADTeacher/mad-agents-skills110—~1.5kAutomated safety check: PassMIT
Testingevanca/flutter-ai-rules647—~1.1kAutomated safety check: PassMIT

Similar skills

  • Test Guidelines

    getsentry/sentry-react-native

    Official

    Enforce Sentry React Native SDK test conventions for naming, structure, mocking, and fixtures with Jest.

    1.8k GitHub stars~1.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Green Gate

    VeryGoodOpenSource/vgv-ai-flutter-plugin

    Drives a Dart or Flutter package fully green through a verify-fix-rerun loop across four gates: analyze, format, test, coverage.

    169 GitHub stars~4.7k tokensUpdated yesterday
    MobileAuto-check: notes
  • Test Driven Development

    aiskillstore/marketplace

    Red-green-refactor development methodology requiring verified test coverage.

    430 GitHub stars~1.9k tokensUpdated today
    Testing & QAAuto-check passed
  • Flutter Testing

    MADTeacher/mad-agents-skills

    Write, fix, review, debug, and validate Flutter tests for apps, packages, and plugins.

    110 GitHub stars~1.5k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Testing

    evanca/flutter-ai-rules

    A skill your agent uses when writing or reviewing Flutter/Dart tests (unit, widget, golden), fixing flaky tests, adding coverage, or choosing between unit and widget tests.

    647 GitHub stars~1.1k tokensUpdated 23 days ago
    Testing & QAAuto-check passed
  • Code Guidelines

    getsentry/sentry-react-native

    Official

    Enforce Sentry React Native SDK code guidelines for implementation, refactoring, and review.

    1.8k GitHub stars~3.2k tokensUpdated today
    DevelopmentAuto-check passed

More from getsentry/sentry-dart

  • Code Guidelines

    getsentry/sentry-dart

    Official

    Enforce Sentry Dart/Flutter SDK code guidelines for implementation, refactoring, and review.

    873 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Design First

    getsentry/sentry-dart

    Official

    Shape non-trivial work before writing it — decide the modules, the seams, and the public API surface up front.

    873 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Review

    getsentry/sentry-dart

    Official

    Three-axis review of the branch diff — Standards (this repo's documented standards + public API surface), Spec (the originating Linear issue / PR), and Correctness (runtime bugs + the SDK threat…

    873 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Diagnosing Bugs

    getsentry/sentry-dart

    Official

    A discipline for hard bugs, flaky tests, CI hangs, and performance regressions in this SDK.

    873 GitHub stars~1.7k tokensUpdated today
    Auto-check passed

Questions about Test Guidelines

What does Test Guidelines do?

Enforce Sentry Dart/Flutter SDK test conventions for naming, structure, and fixtures. Test Guidelines is an agent skill from getsentry/sentry-dart, published by the product's own GitHub organization. Enforce Sentry Dart/Flutter SDK test conventions for naming, structure, and fixtures.

When should I use Test Guidelines?

Test Guidelines fits situations like: modifying tests; reviewing test code; fixing failing tests; adding test coverage.

How do I install Test Guidelines in Claude Code?

Run `npx skills add getsentry/sentry-dart --skill test-guidelines -a claude-code`. Or copy the skill folder (.agents/skills/test-guidelines in getsentry/sentry-dart) into .claude/skills/test-guidelines in your project. Claude Code loads it when a task matches its description.

How do I install Test Guidelines in Codex?

Run `npx skills add getsentry/sentry-dart --skill test-guidelines -a codex`. Or copy the skill folder (.agents/skills/test-guidelines in getsentry/sentry-dart) into .agents/skills/test-guidelines in your project. Codex loads it when a task matches its description.

Can I use Test Guidelines in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add getsentry/sentry-dart --skill test-guidelines -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-guidelines, .gemini/skills/test-guidelines, .github/skills/test-guidelines and .opencode/skills/test-guidelines in your project.

What does Test Guidelines need to run?

SKILL.md names no scripts, command-line tools or credentials: Test Guidelines is instructions for the agent only.

Does Test Guidelines access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Test Guidelines safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Guidelines use?

Test Guidelines is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Guidelines use?

About 3.1k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to Test Guidelines?

Skills that share tags, products or a category with Test Guidelines: Test Guidelines (getsentry/sentry-react-native, 1.8k stars), Green Gate (VeryGoodOpenSource/vgv-ai-flutter-plugin, 169 stars), Test Driven Development (aiskillstore/marketplace, 430 stars) and Flutter Testing (MADTeacher/mad-agents-skills, 110 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Guidelines?

getsentry (a GitHub organization, an official publisher) maintains it in getsentry/sentry-dart, which has 873 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 7, 2026.

Source: getsentry/sentry-dart on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.