Agent skill

Observability Driven Testing

by petrkindlmann in petrkindlmann/qa-skills

Use production telemetry as INPUT to design new tests. An agent skill from petrkindlmann/qa-skills.

MITAuto-check passedDevOps & Cloud

Install Observability Driven Testing

skills CLI
$ npx skills add petrkindlmann/qa-skills --skill observability-driven-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install petrkindlmann/qa-skills observability-driven-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/observability-driven-testing .claude/skills/observability-driven-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
observability-driven-testing
GitHub stars
163
Token cost
~5.5k tokens
SKILL.md length
2,262 words
Files
3 (incl. references)
Skills in repo
45
Repo updated
First seen
Licence
MIT

At a glance

Use production telemetry as INPUT to design new tests. An agent skill from petrkindlmann/qa-skills.

  • Works in 10 steps: Production data informs test priorities → Traces are test evidence → Observability gaps equal test gaps → …
  • : trace-based testing
  • SKILL.md covers Quick Route, Discovery Questions, Core Principles and Traces as Test Evidence, plus 7 more sections
  • Calls npx

What it does

Observability Driven Testing is an agent skill from petrkindlmann/qa-skills. Use production telemetry as INPUT to design new tests. Covers OpenTelemetry integration with tests, trace-based assertions, log-informed test creation, production-error analysis for coverage gaps, and telemetry-driven test prioritization. Use when: "trace-based testing," "design tests from logs," "OpenTelemetry assertions," "production errors point to test gaps," "telemetry-driven testing." Not for: safe rollout techniques (flags, canary) during release — use testing-in-production. Not for: scheduled post-deploy…

Its SKILL.md is about 5.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/log-and-error-pipeline.md` and `references/trace-assertions.md`).

It sits in DevOps & Cloud, covering Observability, Deployment and Issue triage. It works with OpenTelemetry and Datadog. The repository describes itself as: 50 QA and test-automation skills for Claude Code, Codex, Cursor, and any Agent Skills Standard runtime. The licence is MIT.

When your agent uses it

  • : trace-based testing
  • Design tests from logs
  • OpenTelemetry assertions
  • Production errors point to test gaps

Example prompts

  • “trace-based testing,”
  • “design tests from logs,”
  • “OpenTelemetry assertions,”
  • “/observability-driven-testing”

Requirements

  • Node.js

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. Production data informs test priorities
  2. Traces are test evidence
  3. Observability gaps equal test gaps
  4. Close the feedback loop
  5. Ignoring production signals
  6. Testing only what is easy to observe
  7. No feedback loop between production and testing
  8. Over-instrumenting tests without acting on data
  9. Using traces only for debugging, not for assertions
  10. Asserting against a probabilistically sampled trace

What it can do on your machine

Read from SKILL.md and the folder at commit b3bb61b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Observability Driven Testing loads about 5.5k tokens when it runs, and up to ~8.2k if it reads all its reference files. Until then it costs about 178 tokens; SKILL.md has 2,262 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~178
When it runs · the whole SKILL.md, loaded when a task matches
~5.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from petrkindlmann/qa-skills at commit b3bb61b, republished under its MIT licence (© petrkindlmann). 2,262 words, ~5,546 tokens.

Download SKILL.mdSave it as .claude/skills/observability-driven-testing/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
observability-driven-testing
description
Use production telemetry as INPUT to design new tests. Covers OpenTelemetry integration with tests, trace-based assertions, log-informed test creation, production-error analysis for coverage gaps, and telemetry-driven test prioritization. Use when: "trace-based testing," "design tests from logs," "OpenTelemetry assertions," "production errors point to test gaps," "telemetry-driven testing." Not for: safe rollout techniques (flags, canary) during release — use testing-in-production. Not for: scheduled post-deploy probes — use synthetic-monitoring. Not for: triaging CI failures — use ai-bug-triage. Related: testing-in-production, synthetic-monitoring, qa-metrics, ai-bug-triage.
license
MIT
metadata.author
kindlmann
metadata.version
2.0
metadata.category
production
<objective>
Production is the richest source of test design input: every error log, slow trace, and latency spike tells you where tests are missing. This skill closes the feedback loop between production observability and test creation, and makes trace structure a test assertion. A `200 OK` that silently hit the database on a path meant to be cache-only passes an HTTP assertion — a trace assertion catches it. Output: instrumented test runners, trace-based assertions, and a production-error-to-test pipeline.
</objective>

Quick Route

SituationGo to
Make test execution emit traces correlated with the appOTel test-runner setup (references/trace-assertions.md)
Assert which services were called / no error spans / latencyTraces as Test Evidence
Turn a Sentry/Datadog error into a testProduction Error to Test Pipeline
Decide which endpoints need tests nextTelemetry-Driven Test Prioritization
A trace assertion is flaky or a span never arrivesFailure Modes

Discovery Questions

Check .agents/qa-project-context.md first. If it exists, use it as context and skip questions already answered there.

Observability stack:

  • What APM/tracing tool is in place? (Datadog, New Relic, Honeycomb, Splunk Observability/SignalFx, ServiceNow Cloud Observability — formerly Lightstep, Dash0, Jaeger, Grafana Tempo, OpenTelemetry-native) — determines how you pull traces and which query syntax the diagnosis workflow uses.
  • Is OpenTelemetry instrumented in the application, and which services? — un-instrumented services are invisible and untestable via traces.
  • What logging infrastructure exists? (ELK, Loki, CloudWatch, Datadog Logs) — sets where log-by-trace-ID correlation happens.
  • Are structured logs used, or free-form text? — structured logs are parseable into test gaps; free-form needs a fingerprinting step first.

Tracing maturity:

  • Are distributed traces available across service boundaries? — without them, only single-service span assertions are possible.
  • What is the trace sampling rate? (100%, 10%, head-based, tail-based) — probabilistic sampling will randomly drop the trace a test asserts on; you must force-sample test traffic (see Failure Modes).
  • Can you search traces by error status, latency threshold, or custom attributes?
  • Are traces correlated with logs and metrics? — enables exemplars (metric → representative trace ID), which makes prioritization concrete.

Production error tracking:

  • What error tracking tool is used? (Sentry, SmartBear Insight Hub — formerly Bugsnag, Rollbar, Datadog Error Tracking, LaunchDarkly Observability — incl. session replay, formerly Highlight.io)
  • How are production errors triaged? (Automated, manual, ignored)
  • Is there a process for turning production errors into test cases?
  • What was the last production error that a test should have caught?

Test infrastructure:

  • Can tests emit telemetry? (Traces, custom metrics, structured logs)
  • Are test results correlated with application telemetry?
  • Do you have a test-to-code coverage mapping? — required to compute the error-rate-to-coverage matrix below.

Core Principles

1. Production data informs test priorities

The most valuable tests prevent real production errors — not theoretical edge cases, not contrived scenarios. Production error logs are a pre-prioritized backlog of tests you should have written, ordered by what real users actually hit.

2. Traces are test evidence

"The API returned 200" proves the endpoint responded. "The request hit the cache, skipped the database, and returned in <50ms" proves the system behaved correctly at every layer. Traces make tests deeper without making them more brittle.

3. Observability gaps equal test gaps

A code path with no traces, no logs, and no metrics is invisible — untestable in production and unverifiable during incidents. Observability coverage and test coverage are two views of the same problem.

4. Close the feedback loop

The complete cycle: error detected → analyzed → test written → deployed → recurrence prevented. If your team finds production errors but does not systematically create tests, the same class of error recurs.


Traces as Test Evidence

Pin @opentelemetry/semantic-conventions to an exact version and treat sem-conv bumps as breaking. Trace assertions reference attribute names by string; those names drift across releases and your assertions silently break. v1.41.0 (April 2026) shipped GenAI breaking changes and a process.executable entity split, and moved graphql.document from Recommended to Opt-In. Pin the literal version and bump deliberately:

json
// package.json — exact pin, no caret
"@opentelemetry/semantic-conventions": "1.41.1"
bash
npm install --save-exact @opentelemetry/semantic-conventions@1.41.1

Do not introduce new OpenTracing shims. The OTel spec deprecated OpenTracing compatibility in March 2026 (removal no earlier than March 2027); new instrumentation should target native OTel APIs and OTLP.

Three patterns, all in references/trace-assertions.md:

  • OpenTelemetry integration in test infrastructure — instrument the test runner (test-setup/tracing.ts) so test execution correlates with application traces via service.name, test.suite, and test.run_id resource attributes. Flush from the runner's global teardown with an awaited sdk.shutdown() — not process.on('beforeExit'), which drops trailing spans.
  • Trace-based assertions — assert on trace structure, span attributes, and timing (which services were called, no ERROR spans, root-span latency, DB operations) instead of only the HTTP status. For unit-level span checks, use an in-process InMemorySpanExporter + SimpleSpanProcessor and read getFinishedSpans() synchronously — no network, no waitForTrace, no timeout flake. Reserve the real collector + waitForTrace path for cross-process traces.
  • Distributed trace validation across services — an assertTraceStructure helper that verifies a request flowed through the expected services in order, with per-span attribute and maxDuration checks.

For declarative trace-based assertions (YAML/UI-driven instead of hand-rolled span queries), the OSS Tracetest project (kubeshop/tracetest) is still available, but the last public OSS release is v1.7.1 (Oct 2024) with low recent activity — evaluate maintenance before adopting. Tracetest's commercial Cloud offering was end-of-lifed October 2024; do not set up Tracetest Cloud, users will hit a dead product.


Log-Informed Test Design

Analyze production error logs for test gaps

Production errors are the highest-priority input for test creation. Each unhandled error is a missing test. See references/log-and-error-pipeline.md for the analyze-production-errors.ts script that maps each production error to test coverage, assigns a priority by frequency and recency, and suggests a test layer (unit/integration/e2e) from the error characteristics.

Categorize errors: covered vs. uncovered
1. Export production errors from error tracker (Sentry, Insight Hub, etc.)
   - Filter: last 30 days, count > 5 (ignore one-off errors)
   - Group by: error message fingerprint

2. For each error group:
   a. Does a test exist that would catch this error?
      → Yes: the test is either not running or has a gap (investigate)
      → No: this is a test gap (create a test)

   b. What layer should the test live at?
      → TypeError, null reference → unit test
      → Timeout, connection error → integration test with fault injection
      → UI rendering error → E2E test
      → Data inconsistency → contract test or database test

3. Output: prioritized list of tests to create, ordered by:
   error frequency × user impact × recency
Prioritize test creation by error frequency and impact

Prioritize using a 2×2 of frequency (high/low) vs. impact (high/low): P0 = high-frequency + high-impact (fix now), P1 = low-frequency + high-impact (next sprint), P2 = high-frequency + low-impact (this sprint), P3 = both low (backlog). Impact indicators: high = payment/auth failure, data loss, crash; low = UI glitch, slow-but-functional response.


Telemetry-Driven Test Prioritization

Score endpoints by error-weighted gap

Invest test effort proportional to real usage and real failure. Gap Score is the canonical formula used throughout this skill:

Gap Score = (error_rate × requests_per_day) / max(test_count, 1)

This is error-weighted: it ranks an endpoint by the absolute volume of failing requests it produces, divided by how much test coverage already guards it. (If you instead want a volume-weighted lens that surfaces high-traffic-but-healthy endpoints, multiply by (1 + error_rate) rather than error_rate — a different question, not the matrix's labels.)

Error rate by endpoint to test coverage mapping
Endpoint           | Requests/day | Error Rate | Test Count | Gap Score
POST /api/orders   | 50,000       | 0.3%       | 2          | 75   CRITICAL
PUT  /api/profile  | 5,000        | 1.2%       | 1          | 60   CRITICAL
DELETE /api/items  | 2,000        | 0.8%       | 0          | 16   HIGH
POST /api/auth     | 80,000       | 0.1%       | 8          | 10   OK
GET  /api/search   | 200,000      | 0.05%      | 15         | 6.7  OK

Gap Score = (error_rate × requests_per_day) / max(test_count, 1)
Labels: CRITICAL ≥ 50, HIGH 12–49, OK < 12.

Action: create tests for endpoints at HIGH or above, highest score first.

Every label above is derived from the formula and the stated thresholds — copy the formula and you reproduce the matrix exactly. Pick your own thresholds, but state them; never hand-label.

Exemplars close the metric → trace → test loop. When a high-error endpoint surfaces in this matrix, OTel exemplars let you jump straight from the error-rate metric to a representative failing trace ID, then walk that trace (below) to write the test — instead of hunting for a matching trace by hand.

Hot path analysis

Identify the most-traversed code paths in production and ensure they have proportional test coverage.

1. Extract top 20 endpoints by request volume from APM data
2. For each endpoint, trace the code path through services
3. Map each service-level span to test coverage data
4. Identify hot paths with zero or low test coverage

Output:
  /api/checkout → cart-service → pricing-service → payment-service
  Coverage: cart-service (82%) → pricing-service (45%) → payment-service (91%)
  Gap: pricing-service discount calculation has 45% coverage on a critical path
  Action: Add tests for discount edge cases in pricing-service

Pair endpoint-level traffic data with continuous profiling to find CPU and allocation hot paths inside endpoints, not just at the boundary. The OTel profiling signal entered public alpha on 2026-03-26 (OTLP path /v1development/profiles), with GA targeted for Q3 2026 — treat it as not-yet-production. Production-ready alternatives today: Pyroscope, Parca, Polar Signals, Datadog Profiling. eBPF zero-instrumentation profilers (no SDK changes): Polar Signals, Parca, Grafana Beyla.

Zero-instrumentation observability — when adding the OTel SDK isn't feasible, eBPF tools capture HTTP/gRPC traces from kernel syscalls without code changes: Beyla (Grafana), Cilium Tetragon, Pixie, Coroot. Useful for legacy or polyglot services where SDK rollout takes quarters.

OTel Weaver generates type-safe instrumentation code from semantic-convention YAML — keeping trace assertions in sync with sem-conv bumps. Worth adopting if you maintain custom conventions or hit attribute drift between versions.


Production Error to Test Pipeline

The most important workflow in this skill: turning production errors into tests that prevent recurrence.

1. ERROR DETECTED
   Source: Sentry, Datadog, CloudWatch, or any error tracker
   Capture: error message, stack trace, request context, trace ID, user impact

2. REPRODUCE
   - Pull the trace from the observability platform (exemplar → trace ID if available)
   - Identify the exact request parameters and state that triggered the error
   - Reproduce locally or in staging with equivalent input
   - If not reproducible: add targeted logging and wait for recurrence

3. WRITE TEST
   - Choose the right layer (unit for logic bugs, integration for service interactions)
   - Test must fail before the fix (red-green verification)
   - Document the originating production error in the test name or a comment

4. FIX AND DEPLOY
   - Fix the bug; verify the test passes with the fix
   - Deploy fix + test together

5. VERIFY ELIMINATION
   - Monitor the same error in production after deploy
   - Confirm error count drops to zero
   - If it recurs: the fix was incomplete, repeat from step 2

See references/log-and-error-pipeline.md for a full test built from Sentry issue PROJ-4521 (null shipping address → null reference), asserting either a 400 or 422 (whichever your contract uses) at the API layer plus the E2E checkout prompt. The test name and a comment document the originating error, frequency, and context — the convention to follow when creating tests from production signals.

Show full SKILL.md (907 more words)Show less
Establish the team feedback loop
  • Weekly error review (30 min): pull the top 10 new errors by frequency from the error tracker. For each: assign an owner, create a test, or mark as known/acceptable. An error tracker with thousands of unresolved entries that nobody reads is the anti-pattern.
  • Incident close gate: add "What test would have prevented this?" to every postmortem; the test is created (or the gap is explicitly recorded) before the incident is closed. Tie this to a checklist item so it is auditable, not aspirational.

Diagnosis Workflows

Trace a failing request end-to-end

When a test fails or a production error occurs, use the trace to understand exactly what happened.

1. Get the trace ID (from test output, error tracker, or user report)

2. Open the trace in your APM tool
   - Jaeger: /trace/{traceId}
   - Datadog: /apm/traces?traceId={traceId}
   - Honeycomb: query by trace.trace_id

3. Walk the span tree
   - Root span: what did the user request?
   - Child spans: which services were called?
   - Error spans: where did it fail? (which span FIRST shows an error)
   - Slow spans: where did latency accumulate?

4. Correlate with logs
   - Filter logs by trace ID to see every log entry for this request
   - Look for warnings or errors that precede the failure

5. Identify the root cause
   - Is the error in your code, a dependency, or infrastructure?
   - Transient failure or persistent bug?
Correlate test failures with production telemetry

When a test fails, query your observability platform: (1) search production errors for matching messages (last 7 days); (2) search traces for the same HTTP route with ERROR status. Matches exist → the bug is real and affecting users, prioritize the fix. No matches → likely a test-only issue or a new bug not yet in production. This turns "probably flaky" into "confirmed production impact" or "test-only issue."


Anti-Patterns

1. Ignoring production signals

The error tracker has 500 unresolved errors nobody looks at; the suite passes, so the team assumes quality is fine. Fix: run the weekly 30-minute error review above — top 10 new errors, each assigned an owner, a test, or a known/acceptable mark.

2. Testing only what is easy to observe

Teams assert HTTP status and response time while ignoring data consistency, background-job completion, and cache coherence. Fix: add spans to background jobs, cache ops, and async workflows, then assert on them. If it runs in production, it should produce telemetry.

3. No feedback loop between production and testing

SRE handles errors, QA writes tests, neither shares systematically, the same class of bug recurs. Fix: establish the production-error-to-test pipeline and the incident-close gate.

4. Over-instrumenting tests without acting on data

Thousands of metrics and logs emitted, nobody analyzes them — cost with no benefit. Fix: start with three specific questions you want test telemetry to answer; build those dashboards; add instrumentation only when you have a new question.

5. Using traces only for debugging, not for assertions

Traces treated as a post-break debugging tool rather than a source of assertions that prevent breaks. Fix: add trace-based assertions to integration tests — correct services called, efficient queries, expected cache hits. These catch regressions HTTP-level assertions miss.

6. Asserting against a probabilistically sampled trace

Head/probabilistic sampling randomly drops the trace the test is asserting on, producing intermittent failures. Fix: force-sample test traffic (OTEL_TRACES_SAMPLER=always_on or a per-request override) so every asserted trace is recorded.


Failure Modes

SymptomLikely causeFix or check
waitForTrace times out, span never arrivesSampling dropped it, or exporter didn't flush before assertSet OTEL_TRACES_SAMPLER=always_on for the test run; flush via awaited sdk.shutdown() in global teardown
Trailing spans from the last test missingFlushed from process.on('beforeExit') (doesn't fire on exit/signal)Move shutdown to the runner's global teardown hook; await sdk.shutdown()
App spans not part of the test's tracetraceparent header not propagated by the appConfirm the app reads/forwards W3C traceparent; check the OTel propagator is configured
Assertion on db.system/graphql.document/GenAI attrs suddenly failssem-conv version bump renamed/moved the attributePin @opentelemetry/semantic-conventions exact; diff the release notes; update assertions deliberately
No spans reach the collector in CIOTEL_EXPORTER_ENDPOINT unreachable from the CI networkPoint at the in-CI collector address; smoke-test with the Verification step below
status?.code === 'ERROR' matches nothing despite real errorsCollector serializes status as 2 / 'STATUS_CODE_ERROR', not 'ERROR'Match what your collector actually emits (see note in references/trace-assertions.md)

Verification

Prove the telemetry path works before trusting any trace assertion, smallest first:

bash
# 1. Start a local collector, point the runner at it, run one instrumented test.
OTEL_EXPORTER_ENDPOINT=http://localhost:4318/v1/traces \
OTEL_TRACES_SAMPLER=always_on \
  npx playwright test --grep @trace

# 2. Confirm a span with service.name=integration-tests arrived at the collector
#    (check the collector's debug/logging exporter output, or query your APM).

Then, in code, assert a known trace ID resolves before relying on any structural assertion: await collector.waitForTrace(traceId, { timeout: 10_000 }) must return spans — if it times out, fix sampling/flush/propagation (Failure Modes) before adding more assertions. For unit-level span checks, the InMemorySpanExporter path returns spans synchronously with no collector at all.

Done When

  • Every one of the top-20-by-traffic endpoints (from the hot-path matrix) resolves to at least one span in a sampled trace — no high-traffic endpoint is invisible.
  • Trace-based assertions exist for at least one key user journey, verifying service calls and span attributes (not just HTTP status), and pass under OTEL_TRACES_SAMPLER=always_on.
  • Log-informed test cases exist for the known failure modes surfaced by the production error analysis.
  • The error-rate/Gap-Score matrix has been computed and has produced at least one prioritized set of untested code paths.
  • @opentelemetry/semantic-conventions is pinned to an exact version in package.json (no caret).
  • A recorded post-deploy review (checklist item or postmortem entry) confirms observability signals were checked before the release was marked stable.

Reference Files (in references/)

  • trace-assertions.md — OTel test-runner setup (with correct global-teardown flush and the in-memory exporter alternative), force-sampling note, trace-based assertions, and the distributed assertTraceStructure helper.
  • log-and-error-pipeline.md — the analyze-production-errors.ts test-gap script and a worked production-error-to-test example (400-or-422 assertion).
  • testing-in-production — safe rollout techniques (flags, canary, guardrail metrics) during a release; this skill instead uses the telemetry those releases produce as input to design tests after.
  • synthetic-monitoring — scheduled probes that run after release and themselves emit telemetry; that telemetry feeds the analysis here.
  • qa-metrics — turns telemetry-derived numbers (error rates, latency, Gap Score) into quality dashboards and KPIs; this skill produces the raw signals, qa-metrics aggregates them.
  • ai-bug-triage — when the input is a pile of CI/production failures to classify and route; use it to feed the error-categorization step here, then return to write the tests.

© petrkindlmann, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/observability-driven-testing of petrkindlmann/qa-skills.

  • SKILL.md
  • references/log-and-error-pipeline.md
  • references/trace-assertions.md

Open the folder on GitHubat commit b3bb61b

Compare with similar skills

Observability Driven Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Observability Driven Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Observability Driven Testing this skillpetrkindlmann/qa-skills163—~5.5kAutomated safety check: PassMIT
Temps Best Practicesgotempsh/temps822—~2.9kAutomated safety check: PassApache-2.0
Frontmcp Observabilityagentfront/frontmcp146—~4.6kAutomated safety check: PassApache-2.0
Monitoring Observabilityahmedasmar/devops-claude-skills203—~3.9kAutomated safety check: PassNone
Observability Architecturemajiayu000/litellm-rs116—~1.3kAutomated safety check: PassMIT
Error HandlerEliasOulkadi/shokunin114—~3.6kAutomated safety check: NotesMIT

Similar skills

  • Temps Best Practices

    gotempsh/temps

    Best-practices reference for preparing and instrumenting applications on Temps.

    822 GitHub stars~2.9k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Frontmcp Observability

    agentfront/frontmcp

    A skill your agent uses when adding tracing, structured logging, metrics, or monitoring to a FrontMCP server.

    146 GitHub stars~4.6k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Monitoring Observability

    ahmedasmar/devops-claude-skills

    Monitoring and observability strategy, implementation, and troubleshooting.

    203 GitHub stars~3.9k tokensUpdated 5 mo ago
    DevOps & CloudAuto-check passed
  • Observability Architecture

    majiayu000/litellm-rs

    LiteLLM-RS Observability Architecture. An agent skill from majiayu000/litellm-rs.

    116 GitHub stars~1.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Error Handler

    EliasOulkadi/shokunin

    Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead…

    114 GitHub stars~3.6k tokensUpdated 3 days ago
    DevOps & CloudAuto-check: notes
  • Enforcing Nophi Logging

    maziyarpanahi/openmed

    Add a logging and telemetry guard that scrubs or blocks PHI from logs, traces, and error reports around an OpenMed deployment.

    5.5k GitHub stars~1.9k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed

More from petrkindlmann/qa-skills

All 45 skills in this repo
  • Accessibility Testing

    petrkindlmann/qa-skills

    Test for WCAG 2.2 AA compliance with axe-core + Playwright, keyboard navigation audits, screen reader testing, ARIA pattern validation, and legal compliance mapping (ADA, EAA, Section 508).

    163 GitHub stars~4.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Agentic Browser Testing

    petrkindlmann/qa-skills

    Goal-driven E2E testing where a browser agent (Playwright MCP / computer-use) reads a natural-language goal and explores the app via the accessibility tree to assert outcomes — no pre-written script.

    163 GitHub stars~4.5k tokensUpdated 3 mo ago
    Auto-check passed
  • AI Test Generation

    petrkindlmann/qa-skills

    Use AI to write NEW test code from specs, PRDs, user stories, code diffs, bug reports, or OpenAPI specs.

    163 GitHub stars~4.8k tokensUpdated 3 mo ago
    Auto-check passed
  • API Testing

    petrkindlmann/qa-skills

    Test REST and GraphQL APIs with Playwright APIRequestContext, Supertest, or standalone HTTP clients.

    163 GitHub stars~2.7k tokensUpdated 3 mo ago
    Auto-check passed
  • CI CD Integration

    petrkindlmann/qa-skills

    Design CI/CD pipelines that run test suites. An agent skill from petrkindlmann/qa-skills.

    163 GitHub stars~4.8k tokensUpdated 3 mo ago
    Auto-check passed
  • Compliance Testing

    petrkindlmann/qa-skills

    Test for regulatory compliance: GDPR/CMP consent verification, Google Consent Mode v2, Global Privacy Control (GPC), CCPA/US state opt-out, EU AI Act Article 50 transparency, Better Ads Standards…

    163 GitHub stars~4.6k tokensUpdated 3 mo ago
    Auto-check passed

Categories

Questions about Observability Driven Testing

What does Observability Driven Testing do?

Use production telemetry as INPUT to design new tests. An agent skill from petrkindlmann/qa-skills. Observability Driven Testing is an agent skill from petrkindlmann/qa-skills. Use production telemetry as INPUT to design new tests.

When should I use Observability Driven Testing?

Observability Driven Testing fits situations like: : trace-based testing; design tests from logs; openTelemetry assertions; production errors point to test gaps.

How do I install Observability Driven Testing in Claude Code?

Run `npx skills add petrkindlmann/qa-skills --skill observability-driven-testing -a claude-code`. Or copy the skill folder (skills/observability-driven-testing in petrkindlmann/qa-skills) into .claude/skills/observability-driven-testing in your project. Claude Code loads it when a task matches its description.

How do I install Observability Driven Testing in Codex?

Run `npx skills add petrkindlmann/qa-skills --skill observability-driven-testing -a codex`. Or copy the skill folder (skills/observability-driven-testing in petrkindlmann/qa-skills) into .agents/skills/observability-driven-testing in your project. Codex loads it when a task matches its description.

Can I use Observability Driven Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add petrkindlmann/qa-skills --skill observability-driven-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/observability-driven-testing, .gemini/skills/observability-driven-testing, .github/skills/observability-driven-testing and .opencode/skills/observability-driven-testing in your project.

What does Observability Driven Testing need to run?

Going by SKILL.md and its folder, Observability Driven Testing needs the command-line tools its instructions call (npx). Our summary lists: Node.js.

Does Observability Driven Testing access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Observability Driven Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Observability Driven Testing use?

Observability Driven Testing is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Observability Driven Testing use?

About 5.5k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.6k tokens, read only when the agent opens those files.

What are the alternatives to Observability Driven Testing?

Skills that share tags, products or a category with Observability Driven Testing: Temps Best Practices (gotempsh/temps, 822 stars), Frontmcp Observability (agentfront/frontmcp, 146 stars), Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars) and Observability Architecture (majiayu000/litellm-rs, 116 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Observability Driven Testing?

petrkindlmann (a GitHub user) maintains it in petrkindlmann/qa-skills, which has 163 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on June 10, 2026.

Source: petrkindlmann/qa-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.