Agent skill

Error Handling And E2E

by home-operations in home-operations/kopiur

How Kopiur does strongly-typed, actionable error handling and end-to-end testing.

AGPL-3.0Auto-check passedTesting & QA

Install Error Handling And E2E

skills CLI
$ npx skills add home-operations/kopiur --skill error-handling-and-e2e -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install home-operations/kopiur error-handling-and-e2e --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/home-operations/kopiur.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/error-handling-and-e2e .claude/skills/error-handling-and-e2e && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
error-handling-and-e2e
GitHub stars
113
Token cost
~2.7k tokens
SKILL.md length
1,305 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
AGPL-3.0

At a glance

How Kopiur does strongly-typed, actionable error handling and end-to-end testing.

  • Works in 4 steps: One thiserror enum per crate/domain,… → Exhaustive classification, not… → pub type Result alias per crate;… → …
  • Changing any error type
  • SKILL.md covers Part 1 — Strongly-typed,…, Part 2 — e2e testing… and Checklist before claiming an…
  • Calls kubectl, mise and helm

What it does

Error Handling And E2E is an agent skill from home-operations/kopiur. How Kopiur does strongly-typed, actionable error handling and end-to-end testing. Use when adding or changing any error type, error path, or fallible operation (kopia calls, kube IO, OTLP/telemetry init, validators, the mover), or when writing/extending tests — especially e2e scenarios in crates/e2e. Encodes the thiserror exhaustive-enum + classification pattern, the what/why/fix message rule, degrade-not-crash for non-critical subsystems, and the crates/e2e harness conventions (feature = "e2e", mise…

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering End-to-end testing. It works with OpenTelemetry and Kubernetes. The repository describes itself as: A Kopia-native Kubernetes backup operator written in Rust. The licence is AGPL-3.0.

When your agent uses it

  • Changing any error type
  • Fallible operation (kopia calls
  • OTLP/telemetry init
  • Writing/extending tests — especially e2e scenarios in crates/e2e

Example prompts

  • “/error-handling-and-e2e”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. One thiserror enum per crate/domain, #[derive(Debug, thiserror::Error)].
  2. Exhaustive classification, not catch-alls. If the error drives behavior
  3. pub type Result alias per crate; reconcile/fallible fns
  4. Test the classification AND the message. Unit-test that each variant maps

What it can do on your machine

Read from SKILL.md and the folder at commit 379efb7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kubectl
    • mise
    • helm
    • cargo

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use kubectl and helm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Error Handling And E2E loads about 2.7k tokens when it runs. Until then it costs about 164 tokens; SKILL.md has 1,305 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~164
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from home-operations/kopiur at commit 379efb7, republished under its AGPL-3.0 licence (© home-operations). 1,305 words, ~2,725 tokens.

Download SKILL.mdSave it as .claude/skills/error-handling-and-e2e/SKILL.md (or your agent's skills folder).
name
error-handling-and-e2e
description
How Kopiur does strongly-typed, actionable error handling and end-to-end testing. Use when adding or changing any error type, error path, or fallible operation (kopia calls, kube IO, OTLP/telemetry init, validators, the mover), or when writing/extending tests — especially e2e scenarios in crates/e2e. Encodes the thiserror exhaustive-enum + classification pattern, the what/why/fix message rule, degrade-not-crash for non-critical subsystems, and the crates/e2e harness conventions (feature = "e2e", mise //crates/e2e:test, World::ensure(Need) provisioning, wait_phase/wait_until, assert real operator output, never a real cluster).

Errors & e2e testing in Kopiur

Two linked disciplines. Strong typed errors make invalid states unrepresentable and tell the operator exactly how to recover; e2e tests prove the whole pipeline reaches the user-visible success condition. Both exist because this is data-protection software — a silently-swallowed error or an untested path can lose backups.

Read alongside [[kopiur-design]] (the type-safety thesis) and CLAUDE.md.

Part 1 — Strongly-typed, actionable errors

The pattern (mirror crates/controller/src/error.rs)
  1. One thiserror enum per crate/domain, #[derive(Debug, thiserror::Error)]. One variant per distinct failure mode. Use #[from] to wrap upstream errors (kube::Error, KopiaError, serde_json::Error) so ? just works.
  2. Exhaustive classification, not catch-alls. If the error drives behavior (requeue timing, fatal-vs-degrade), express that as a method with an exhaustive match and no _ => arm — a new variant must fail to compile until it is classified. See Error::class() -> ErrorClass (Transient/Structural → backoff). A _ arm here is a latent data-loss bug.
  3. pub type Result<T, E = Error> alias per crate; reconcile/fallible fns return it.
  4. Test the classification AND the message. Unit-test that each variant maps to the right class/behavior (error.rs tests) and — for any message a human acts on — assert on to_string() so the actionable text can't silently rot.
The message rule: what failed, why, how to fix

Every #[error("…")] an operator might read states what failed, the likely why, and the concrete fix (expected value, env var, command). This is the type-safety thesis extended to UX: homelabbers run this; a cryptic error wastes time and can mask a data-protection gap.

rust
#[error(
    "OTEL_EXPORTER_OTLP_ENDPOINT='{value}' is not a valid URL. Use scheme+host+port, \
     e.g. http://otel-collector:4317 (OTLP/gRPC) or :4318 (OTLP/HTTP); unset it to disable OTLP"
)]
InvalidOtlpEndpoint { value: String, #[source] source: url::ParseError },

Prefer {named} fields over positional {0} when the message interpolates user input — it reads better and survives reordering. Keep the underlying cause via #[source]/#[from] so the chain is inspectable; don't bury it inside a string.

Degrade, don't crash, for non-critical subsystems

Reconciliation and data movement are critical; observability, OTLP export, and similar side-channels are not. A misconfigured non-critical subsystem must log the actionable error and continue, never abort the operator:

  • init_telemetry() returns Result; on error the binary logs it at error and falls back (fmt-only logs + the always-on Prometheus pull). An optional KOPIUR_OTEL_STRICT=true makes it fail-fast for those who want it.
  • Critical paths still fail loudly: a bad CRD spec is Structural (long requeue, surfaced on .status + an Event), a kopia/API outage is Transient (short requeue). Both go through error_policy_for, which records controller_reconcile_errors_total{kind,class} and requeues by class.
Surfacing errors to users

Reconcile errors should reach the user, not just the log: set a Condition / phase on .status and emit a kube Event via the Recorder (Context already carries one). The metric + log + status + Event together make a failure visible in kubectl get, dashboards, and alerts.

Part 2 — e2e testing (crates/e2e)

When to write an e2e test

Both bug fixes and new features ship with e2e coverage ([[kopiur-design]] — "Every change ships with tests"). Per [[regression-test-every-bugfix]], a fix ships with a test at the cheapest tier that exercises the broken path; a new feature additionally gets a full-pipeline e2e scenario whenever it has any runtime-observable behavior. Reach for e2e (not just a unit test) when the behavior only appears against a live operator: a reconcile that must reach its terminal phase, a runtime dependency, an RBAC/SA gap, an option that survives serialization but acts at runtime, a new backend/volume source the mover must actually mount, or metrics that must be emitted and scrapable.

No e2e is too large to include — this tool is data-protection software and must be thoroughly tested. If a feature needs new cluster infrastructure to exercise (an NFS server, an SFTP server, a fresh namespace, a second backend), provision it in the World rather than skipping the test:

  • Add a Need variant + an ensure_* method (idempotent Fixture apply), plus a builders::* Deployment/Service for any in-cluster server. Mirror the existing servers (ensure_sftp/ensure_webdav/ensure_nfs).
  • Pin every server image (never :latest) and verify it actually serves the mover on a real kind cluster before trusting a green run — a server the client can't talk to fails slowly and confusingly, not fast (see [[sftp-openssh10-go-ssh-hang]]). Note in the test/PR if it's pinned-but-not- yet-cluster-verified.
  • A multi-minute run, an extra built image, or a privileged helper pod are all acceptable costs. The bar is "is the feature exercised end to end," not "is the test cheap."

The only features that legitimately skip e2e are those with no runtime-observable behavior (a pure type change/refactor) — and then say so explicitly.

Show full SKILL.md (590 more words)Show less
Harness conventions (mirror crates/e2e/tests/lifecycle.rs)
  • Gate: #![cfg(all(unix, feature = "e2e"))] at the top of the test file, and #[ignore = "requires the e2e harness (mise run //crates/e2e:test)"] on each #[tokio::test]. So the suite compiles everywhere and is skipped without a cluster; it only runs under the harness.
  • Run it: mise run //crates/e2e:test. The host-level steps are mise tasks in crates/e2e/mise.toml (a monorepo subproject): build + load the images into kind, seed the node's hostPath dirs, and helm upgrade --install with deploy/e2e/values.yaml (webhook disabled — covered by unit/integration tiers).
  • Declare cluster prerequisites as data. Each scenario opens with let Some(world) = World::connect().await else { return; }; then world.ensure(&[Need::Filesystem /* | Need::Minio | Need::WorkloadNs */]).await?. World (crates/e2e/src/world.rs) provisions namespaces/Secrets/PV-PVCs/MinIO/ buckets idempotently via the type-safe Fixture apply dispatch — never via kubectl in a shell. Add a new fixture kind by extending the Need/Fixture enums (exhaustive match).
  • NEVER target a real cluster. scripts/with-kind.sh (integration tier) and the //crates/e2e:* tasks pin an isolated kubeconfig under target/e2e/ and tear the cluster down; they never touch the homelab kubecontext. Do not add a test that reads the ambient KUBECONFIG.
  • Helpers live in kopiur_e2e (crates/e2e/src/lib.rs): E2E_NAMESPACE, try_client() (returns None when no cluster → skip gracefully), wait_until(...), default_timeout(), poll_interval(). Reuse them; don't re-roll polling.
Assert real operator output, and make the test fail on the bug

Assert the user-visible success condition, not an intermediate detail:

  • A Repository reaching Ready; a Snapshot reaching Succeeded with a real kopiaSnapshotID; a Restore Completed; a schedule actually creating a Snapshot; a finalizer deleting the snapshot; a Maintenance lease claimed.
  • Use the wait_phase(&api, name, "Succeeded") pattern (poll status to a target phase with a timeout). A regression test must be written so it times out / fails on the buggy code and passes on the fix — see cluster_repository_backup_lifecycle (added for the "ClusterRepository refs ignored" bug: it would hang at wait_phase(... "Succeeded") before the fix).
Validating metrics in e2e (observability)

When the assertion is "a metric is emitted," drive the real lifecycle, then scrape /metrics and parse the exposition rather than trusting a code path:

  • curl the controller :8081/metrics (and webhook /metrics); assert the expected families/labels are present with sane values (kopiur_controller_reconciliations_total{kind=…} > 0, kopiur_resource_phase{kind=Snapshot,phase=Succeeded} == 1, backup size/files/duration > 0, error/consecutive-failure counters reflect an induced failure). kopiur_resource_phase and the per-CR/per-policy Snapshot gauges are store-backed observable gauges: only the CR's active phase is ever emitted (never a phase=…}=0 line for the inactive phases), and a series disappears from the next scrape once its CR is deleted — so a regression test for a lifecycle fix should also assert absence after deletion, not just presence while the CR is live.
  • Assert the body still parses as valid Prometheus text — a regression guard for the OTel→Prometheus name rewrite (_total suffixes, otel_scope_*, target_info).
  • Also assert /healthz + /readyz return 200 (real endpoints, not the old any-path listener).

Checklist before claiming an error/test change done

  • New error variants classified in an exhaustive match (no _ =>).
  • User-facing messages say what/why/fix; #[source] preserved.
  • Non-critical subsystem degrades-and-logs; critical path fails loud + surfaces on status/Event/metric.
  • Message text unit-tested where a human acts on it.
  • Bug fix has a test that fails without the fix (unit if possible, e2e if the bug is pipeline-only).
  • New feature has tests at every tier it touches: unit/serde for the leaf logic + controller glue, AND a full-pipeline e2e scenario (provisioning any new World infra needed). No e2e skipped for being "too large"; only skipped when the feature has no runtime-observable behavior — stated explicitly.
  • e2e gated feature = "e2e" + #[ignore], uses kopiur_e2e helpers, asserts the user-visible success condition, never a real cluster.
  • cargo test --workspace + clippy -D warnings + fmt --check green; record the bug + guard in the operator-bugs-fixed-by-e2e memory.

© home-operations, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/error-handling-and-e2e of home-operations/kopiur.

Open the folder on GitHubat commit 379efb7

Compare with similar skills

Error Handling And E2E next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Error Handling And E2E compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Error Handling And E2E this skillhome-operations/kopiur113—~2.7kAutomated safety check: PassAGPL-3.0
Run E2E Testkubernetes-sigs/cloud-provider-azure294—~3.8kAutomated safety check: PassApache-2.0
Must Gather Investigationscylladb/scylla-operator401—~2.4kAutomated safety check: PassApache-2.0
Test Pyramidkubernetes-sigs/agent-sandbox4.2k—~1.6kAutomated safety check: PassApache-2.0
E2E Testgrafana/agento11y127—~1.7kAutomated safety check: NotesApache-2.0
Kaniop Developmentpando85/kaniop130—~3.1kAutomated safety check: PassAGPL-3.0

Similar skills

  • Run E2E Test

    kubernetes-sigs/cloud-provider-azure

    Official

    Parse a Go e2e test from tests/e2e/, translate each step to kubectl and az CLI commands, and interactively replay the test against a live cluster.

    294 GitHub stars~3.8k tokensUpdated today
    Testing & QAAuto-check passed
  • Must Gather Investigation

    scylladb/scylla-operator

    Investigate failed e2e tests from Ginkgo JSON reports and must-gather artifacts, systematically analyzing logs, events, and resource states to identify root causes.

    401 GitHub stars~2.4k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Test Pyramid

    kubernetes-sigs/agent-sandbox

    Official

    Analyze the repo's unit and E2E tests and propose rebalancing toward a test pyramid — which E2E tests (or assertions inside them) can be covered by unit tests, which unit-level gaps genuinely need…

    4.2k GitHub stars~1.6k tokensUpdated today
    Testing & QAAuto-check passed
  • E2E Test

    grafana/agento11y

    Official

    Optional credential-free Hermes integration checks using an explicit loopback model provider and local telemetry receivers.

    127 GitHub stars~1.7k tokensUpdated yesterday
    Testing & QAAuto-check: notes
  • Kaniop Development

    pando85/kaniop

    Kaniop architecture, Rust controller conventions, commands, testing, and repository workflows.

    130 GitHub stars~3.1k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Debug E2E Pipeline

    kubernetes-sigs/cloud-provider-azure

    Official

    Fetch and analyze Prow e2e pipeline failures for cloud-provider-azure.

    294 GitHub stars~3.4k tokensUpdated today
    Testing & QAAuto-check passed

More from home-operations/kopiur

  • Documentation

    home-operations/kopiur

    How Kopiur writes and maintains user-facing docs — the MkDocs Material site under docs/ and the example manifests under deploy/examples/.

    113 GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Kopiur Design

    home-operations/kopiur

    Design norms and locked decisions for the Kopiur Kopia-native Kubernetes backup operator (Rust/kube-rs).

    113 GitHub stars~2.4k tokensUpdated today
    Auto-check passed

Categories

Questions about Error Handling And E2E

What does Error Handling And E2E do?

How Kopiur does strongly-typed, actionable error handling and end-to-end testing. Error Handling And E2E is an agent skill from home-operations/kopiur. How Kopiur does strongly-typed, actionable error handling and end-to-end testing.

When should I use Error Handling And E2E?

Error Handling And E2E fits situations like: changing any error type; fallible operation (kopia calls; OTLP/telemetry init; writing/extending tests — especially e2e scenarios in crates/e2e.

How do I install Error Handling And E2E in Claude Code?

Run `npx skills add home-operations/kopiur --skill error-handling-and-e2e -a claude-code`. Or copy the skill folder (.claude/skills/error-handling-and-e2e in home-operations/kopiur) into .claude/skills/error-handling-and-e2e in your project. Claude Code loads it when a task matches its description.

How do I install Error Handling And E2E in Codex?

Run `npx skills add home-operations/kopiur --skill error-handling-and-e2e -a codex`. Or copy the skill folder (.claude/skills/error-handling-and-e2e in home-operations/kopiur) into .agents/skills/error-handling-and-e2e in your project. Codex loads it when a task matches its description.

Can I use Error Handling And E2E in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add home-operations/kopiur --skill error-handling-and-e2e -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/error-handling-and-e2e, .gemini/skills/error-handling-and-e2e, .github/skills/error-handling-and-e2e and .opencode/skills/error-handling-and-e2e in your project.

What does Error Handling And E2E need to run?

Going by SKILL.md and its folder, Error Handling And E2E needs the command-line tools its instructions call (kubectl, mise, helm and cargo).

Does Error Handling And E2E access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Error Handling And E2E safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Error Handling And E2E use?

Error Handling And E2E is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Error Handling And E2E use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Error Handling And E2E?

Skills that share tags, products or a category with Error Handling And E2E: Run E2E Test (kubernetes-sigs/cloud-provider-azure, 294 stars), Must Gather Investigation (scylladb/scylla-operator, 401 stars), Test Pyramid (kubernetes-sigs/agent-sandbox, 4.2k stars) and E2E Test (grafana/agento11y, 127 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Error Handling And E2E?

home-operations (a GitHub organization) maintains it in home-operations/kopiur, which has 113 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 8, 2026.

Source: home-operations/kopiur on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.