Official agent skill

Write Vibe Tests

by mistralai in mistralai/mistral-vibe

Guides writing or refactoring tests for the Mistral Vibe CLI agent so they check behavior through stable boundaries, like tool invocation or saved session shape, instead of internal calls.

OfficialApache-2.0Auto-check passedTesting & QA

Install Write Vibe Tests

skills CLI
$ npx skills add mistralai/mistral-vibe --skill write-vibe-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mistralai/mistral-vibe write-vibe-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mistralai/mistral-vibe.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.vibe/skills/write-vibe-tests .claude/skills/write-vibe-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
write-vibe-tests
GitHub stars
5.1k
Token cost
~1.2k tokens
SKILL.md length
535 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
Apache-2.0

At a glance

Guides writing or refactoring tests for the Mistral Vibe CLI agent so they check behavior through stable boundaries, like tool invocation or saved session shape, instead of internal calls.

  • Works in 5 steps: Find the smallest seam: function… → Add characterization tests through the… → Capture current observable behavior,… → …
  • Adding test coverage for a new feature in the Vibe codebase
  • SKILL.md covers Core Principle, Vibe Test Boundaries, Prefer Fakes Over Mocks and Legacy Or Refactor Workflow, plus 3 more sections
  • Calls uv

What it does

The core principle is that a refactor which changes no observable behavior should not break many tests; if it does, the tests were coupled to implementation details. A table sets the preferred test boundary per area: core services through their public API and persisted state, tools through BaseTool.invoke or run with typed results and permission behavior, LLM orchestration through AgentLoop events and fake backends, config through Pydantic validation and layer merge, sessions through saved JSONL shape and resume behavior, CLI widgets through Textual snapshots, and ACP through its session updates rather than core internals.

Fakes are preferred over mocks: reusable in-memory test doubles live in tests/stubs named Fake*, kept small and behavior-oriented, with mocks reserved for hard process, network, time or third-party boundaries, and assertions like "method X was called with Y" used only when the call itself is the observable contract.

For legacy or hard-to-test code, the workflow finds the smallest seam, such as a constructor dependency or module boundary, adds characterization tests through the nearest public entry point to capture current behavior even if awkward, refactors behind that seam in small steps, then replaces the characterization checks with clearer behavior tests. The stack is pytest with pytest-asyncio, pytest-textual-snapshot and respx, using descriptive test names without docstrings.

When your agent uses it

  • Adding test coverage for a new feature in the Vibe codebase
  • Replacing brittle mocks with fakes for a port or adapter
  • Adding characterization tests before refactoring hard-to-test code
  • Deciding where a test boundary should sit for tools, sessions or CLI widgets

Example prompts

  • “Write tests for the new session resume logic using a fake backend instead of mocks.”
  • “Add characterization tests around this legacy config loader before we refactor it.”
  • “Replace the mock-heavy tests in tests/tools/test_search.py with a Fake adapter.”

Requirements

  • pytest with pytest-asyncio, pytest-textual-snapshot and respx

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Find the smallest seam: function boundary, constructor dependency, protocol/port, wrapper, composition root, feature flag, or module…
  2. Add characterization tests through the nearest public entry point.
  3. Capture current observable behavior, even if awkward.
  4. Refactor behind the seam in small steps.
  5. Replace broad characterization checks with clearer behavior/spec tests as the design improves.

What it can do on your machine

Read from SKILL.md and the folder at commit 7cb9189. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Write Vibe Tests loads about 1.2k tokens when it runs. Until then it costs about 87 tokens; SKILL.md has 535 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~87
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mistralai/mistral-vibe at commit 7cb9189, republished under its Apache-2.0 licence (© mistralai). 535 words, ~1,195 tokens.

Download SKILL.mdSave it as .claude/skills/write-vibe-tests/SKILL.md (or your agent's skills folder).
name
write-vibe-tests
description
Write or refactor Mistral Vibe tests with proper decoupling. Use when adding behavior coverage, testing ports/adapters, replacing brittle mocks, creating fakes, adding characterization tests before refactors, or changing tests under tests/ for vibe/core, vibe/cli, vibe/acp, tools, config, sessions, skills, hooks, MCP, or setup.
metadata.display-name
Write Vibe Tests
metadata.short-description
Write decoupled behavior tests for Vibe
metadata.default-prompt
Use $write-vibe-tests to add or refactor Mistral Vibe tests around behavior and stable boundaries.

Write Vibe Tests

Use this skill to make Vibe architecture testable, not just well-shaped. Tests should protect observable behavior while allowing internals to move.

Core Principle

Test behavior through stable boundaries. Do not couple tests to private methods, internal call choreography, or temporary structure.

If a refactor changes no observable behavior but breaks many tests, the tests are probably coupled to implementation details.

Vibe Test Boundaries

Code areaPreferred test boundary
Core use cases and servicesPublic function/class API, typed events, model outputs, persisted state
ToolsBaseTool.invoke/run, typed args/results, permission behavior, user-facing errors
LLM/backend orchestrationAgentLoop events and fake backend outputs
ConfigPydantic model validation, layer merge outputs, migration results
SessionsSaved JSONL/metadata shape, resume/loader behavior, migration behavior
CLI widgetsTextual snapshots, posted messages, rendered user-visible state
ACPACP session updates and protocol-facing content, not core internals

Prefer Fakes Over Mocks

Prefer in-memory implementations and fake adapters that implement real contracts.

  • Put reusable test doubles in tests/stubs/ and name them Fake*.
  • Make fakes small and behavior-oriented.
  • Mock only at hard process, network, time, or third-party boundaries when a fake would be more complex than the behavior under test.
  • Avoid assertions like "method X was called with Y" unless the call itself is the observable contract.

Legacy Or Refactor Workflow

When code is hard to test:

  1. Find the smallest seam: function boundary, constructor dependency, protocol/port, wrapper, composition root, feature flag, or module boundary.
  2. Add characterization tests through the nearest public entry point.
  3. Capture current observable behavior, even if awkward.
  4. Refactor behind the seam in small steps.
  5. Replace broad characterization checks with clearer behavior/spec tests as the design improves.

Prefer "make it testable" refactors first: isolate I/O, extract pure functions, introduce ports/adapters where useful, or move construction out of business logic.

Show full SKILL.md (241 more words)Show less

Test Shape

  • Stack: pytest + pytest-asyncio + pytest-textual-snapshot + respx.
  • Use descriptive test names; do not add test docstrings. Pytest displays docstrings instead of node IDs when present, which hurts.
  • Arrange, act, and assert clearly, but optimize for readability over ceremony.
  • Keep tests deterministic, fast, and explicit about failure.
  • Use autouse fixtures from tests/conftest.py (config_dir, tmp_working_directory) for config/home/working-directory isolation.
  • Mark async tests with @pytest.mark.asyncio.
  • Mock outbound HTTP with respx.
  • Use the narrowest relevant test first, then broaden when shared contracts are touched.
  • Tests are exempt from the ANN and PLR ruff rules (see per-file-ignores).
  • Changes to managed shell tools, managed shell backends, shell prompts, or shell dependencies require focused POSIX/common validation and passing Windows shell CI. If Windows shell CI is unavailable, the PR notes must include manual native Windows commands and results.
  • After app-server session behavior, backend adapters, or shared contract-fixture changes, run both backend contract commands. A failing Unified frontier is expected before parity; include its complete pytest summary in parity-related PRs.

Avoid

  • Testing private methods as the primary coverage for behavior.
  • Asserting intermediate internal state when user-visible output, emitted events, files, or return values can be asserted.
  • Building mocks that mirror the implementation.
  • Adding abstractions only to satisfy a test.
  • Writing tests that require a specific internal file split or call order when the domain behavior is unchanged.

Verification

After Python test/code changes, run:

bash
uv run ruff format .
uv run ruff check --fix .

Then run the targeted tests:

bash
uv run pytest <test-path-or-node-id>

Run uv run pyright when signatures, models, protocols, or shared contracts changed.

© mistralai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .vibe/skills/write-vibe-tests of mistralai/mistral-vibe.

Open the folder on GitHubat commit 7cb9189

Compare with similar skills

Write Vibe Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Write Vibe Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Write Vibe Tests this skillmistralai/mistral-vibe5.1k—~1.2kAutomated safety check: PassApache-2.0
Designing TestsCloudAI-X/claude-workflow-v21.4k1 repos~1.5kAutomated safety check: PassMIT
Robotics Testingarpitg1304/robotics-agent-skills368—~4.7kAutomated safety check: PassApache-2.0
Python Testingaffaan-m/ECC275k6 repos~4.7kAutomated safety check: PassMIT
Python Testing Patternsjh941213/my-cc-harness12617 repos~5.4kAutomated safety check: PassNone
Python Testingaffaan-m/ECC275k—~2.7kAutomated safety check: PassMIT

Similar skills

  • Designing Tests

    CloudAI-X/claude-workflow-v2

    Designs and implements testing strategies for any codebase. An agent skill from CloudAI-X/claude-workflow-v2.

    1.4k GitHub starsUsed in 1 repo~1.5k tokens
    Testing & QAAuto-check passed
  • Robotics Testing

    arpitg1304/robotics-agent-skills

    Testing strategies, patterns, and tools for robotics software.

    368 GitHub stars~4.7k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Python Testing

    affaan-m/ECC

    Python testing strategies using pytest, TDD methodology, fixtures, mocking, parametrization, and coverage requirements.

    275k GitHub starsUsed in 6 repos~4.7k tokens
    Testing & QAAuto-check passed
  • Python Testing Patterns

    jh941213/my-cc-harness

    Implement comprehensive testing strategies with pytest, fixtures, mocking, and test-driven development.

    126 GitHub starsUsed in 17 repos~5.4k tokens
    Testing & QAAuto-check passed
  • Python Testing

    affaan-m/ECC

    Python testing best practices using pytest including fixtures, parametrization, mocking, coverage analysis, async testing, and test organization.

    275k GitHub stars~2.7k tokensUpdated 3 days ago
    Testing & QAAuto-check passed
  • Python Testing Strategies

    c0x12c/ai-toolkit

    Testing patterns for FastAPI with pytest-asyncio, httpx AsyncClient, fixtures, and test data factories.

    106 GitHub stars~747 tokensUpdated 3 mo ago
    Testing & QAAuto-check passed

More from mistralai/mistral-vibe

All 14 skills in this repo
  • Mistral Vibe Plugin Creator

    mistralai/mistral-vibe

    Official

    Shows how to build a Vibe plugin package in the Agent Plugins 1.0 format, with a plugin.json manifest and optional skills, MCP servers, hooks and other components.

    5.1k GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Create Vibe Feature

    mistralai/mistral-vibe

    Official

    Guides feature work in the Mistral Vibe Python CLI so each change lands in the right module and matches the project's architecture decision records.

    5.1k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Official

    Plans which analytics events and properties a new feature needs, checks them against the existing event registry, and verifies them per environment.

    5.1k GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Vibe Worktree Manager

    mistralai/mistral-vibe

    Official

    Creates, reuses and cleans up git worktrees under a shared vibe home directory, with per-repo buckets, claim records and dirty-state checks before removal.

    5.1k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Write Vibe ADR

    mistralai/mistral-vibe

    Official

    Creates or updates concise Architecture Decision Records for the Mistral Vibe CLI and registers each one in the AGENTS.md decisions table.

    5.1k GitHub stars~942 tokensUpdated yesterday
    Auto-check passed
  • Mistral Vibe CLI Reference

    mistralai/mistral-vibe

    Official

    Reference for Mistral Vibe, the CLI agent it runs inside: config files, env vars, agents, skills, tools, hooks and MCP servers, so the agent can explain and troubleshoot its own setup.

    5.1k GitHub stars~14k tokensUpdated yesterday
    Auto-check: notes

Questions about Write Vibe Tests

What does Write Vibe Tests do?

Guides writing or refactoring tests for the Mistral Vibe CLI agent so they check behavior through stable boundaries, like tool invocation or saved session shape, instead of internal calls. The core principle is that a refactor which changes no observable behavior should not break many tests; if it does, the tests were coupled to implementation details.invoke or run with typed results and permission behavior, LLM orchestration through AgentLoop events and fake backends, config through Pydantic validation and layer merge, sessions through saved JSONL shape and resume behavior, CLI widgets through Textual snapshots, and ACP through its session updates rather than core internals.

When should I use Write Vibe Tests?

Write Vibe Tests fits situations like: adding test coverage for a new feature in the Vibe codebase; replacing brittle mocks with fakes for a port or adapter; adding characterization tests before refactoring hard-to-test code; deciding where a test boundary should sit for tools, sessions or CLI widgets.

How do I install Write Vibe Tests in Claude Code?

Run `npx skills add mistralai/mistral-vibe --skill write-vibe-tests -a claude-code`. Or copy the skill folder (.vibe/skills/write-vibe-tests in mistralai/mistral-vibe) into .claude/skills/write-vibe-tests in your project. Claude Code loads it when a task matches its description.

How do I install Write Vibe Tests in Codex?

Run `npx skills add mistralai/mistral-vibe --skill write-vibe-tests -a codex`. Or copy the skill folder (.vibe/skills/write-vibe-tests in mistralai/mistral-vibe) into .agents/skills/write-vibe-tests in your project. Codex loads it when a task matches its description.

Can I use Write Vibe Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mistralai/mistral-vibe --skill write-vibe-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/write-vibe-tests, .gemini/skills/write-vibe-tests, .github/skills/write-vibe-tests and .opencode/skills/write-vibe-tests in your project.

What does Write Vibe Tests need to run?

Going by SKILL.md and its folder, Write Vibe Tests needs the command-line tools its instructions call (uv). Our summary lists: pytest with pytest-asyncio, pytest-textual-snapshot and respx.

Does Write Vibe Tests access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Write Vibe Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Write Vibe Tests use?

Write Vibe Tests is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Write Vibe Tests use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Write Vibe Tests?

Skills that share tags, products or a category with Write Vibe Tests: Designing Tests (CloudAI-X/claude-workflow-v2, 1.4k stars), Robotics Testing (arpitg1304/robotics-agent-skills, 368 stars), Python Testing (affaan-m/ECC, 275k stars) and Python Testing Patterns (jh941213/my-cc-harness, 126 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Write Vibe Tests?

mistralai (a GitHub organization, an official publisher) maintains it in mistralai/mistral-vibe, which has 5,073 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 7, 2026.

Source: mistralai/mistral-vibe on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.