Agent skill

Tracking Quality Trends

by jaktestowac in jaktestowac/awesome-copilot-for-testers

Turns point-in-time quality readings into a trend: archives each run, diffs against the previous one, and reports direction per metric - practices newly present or regressed, coverage movement…

MITAuto-check passed

Install Tracking Quality Trends

skills CLI
$ npx skills add jaktestowac/awesome-copilot-for-testers --skill tracking-quality-trends -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaktestowac/awesome-copilot-for-testers tracking-quality-trends --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaktestowac/awesome-copilot-for-testers.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tracking-quality-trends .claude/skills/tracking-quality-trends && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tracking-quality-trends
GitHub stars
116
Token cost
~2.4k tokens
SKILL.md length
1,287 words
Files
3
Skills in repo
13
Repo updated
First seen
Licence
MIT

At a glance

Turns point-in-time quality readings into a trend: archives each run, diffs against the previous one, and reports direction per metric - practices newly present or regressed, coverage movement…

  • Works in 5 steps: Fix the metric set → Archive each run → Diff against the previous run → …
  • Quality reporting is a series of disconnected snapshots
  • SKILL.md covers When to Use, Operating Principles, Workflow and Cadence, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Tracking Quality Trends is an agent skill from jaktestowac/awesome-copilot-for-testers. Turns point-in-time quality readings into a trend: archives each run, diffs against the previous one, and reports direction per metric - practices newly present or regressed, coverage movement, flake rate, waivers expiring, eval scores - using limit/current/goal framing. Use when quality reporting is a series of disconnected snapshots, when a team needs to show improvement over a quarter, when a number is quoted with no baseline, or when a regression in the quality system itself should be visible.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `resources/trend-metrics.md` and `resources/trend-report-template.md`).

The repository describes itself as: 👨💻 Instructions, prompts, and chat modes to help You with test automation for GitHub Copilot 🤖. The licence is MIT.

When your agent uses it

  • Quality reporting is a series of disconnected snapshots
  • A team needs to show improvement over a quarter
  • A number is quoted with no baseline
  • A regression in the quality system itself should be visible

Example prompts

  • “Use the tracking-quality-trends skill to turn point-in-time quality readings into a trend: archives each run, diffs against the previous one, and…”
  • “/tracking-quality-trends”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Fix the metric set
  2. Archive each run
  3. Diff against the previous run
  4. Establish the noise band
  5. Report

What it can do on your machine

Read from SKILL.md and the folder at commit 8910672. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tracking Quality Trends loads about 2.4k tokens when it runs. Until then it costs about 132 tokens; SKILL.md has 1,287 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~132
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaktestowac/awesome-copilot-for-testers at commit 8910672, republished under its MIT licence (© jaktestowac). 1,287 words, ~2,400 tokens.

Download SKILL.mdSave it as .claude/skills/tracking-quality-trends/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
tracking-quality-trends
description
Turns point-in-time quality readings into a trend: archives each run, diffs against the previous one, and reports direction per metric - practices newly present or regressed, coverage movement, flake rate, waivers expiring, eval scores - using limit/current/goal framing. Use when quality reporting is a series of disconnected snapshots, when a team needs to show improvement over a quarter, when a number is quoted with no baseline, or when a regression in the quality system itself should be visible.
argument-hint
Where previous run artifacts live, which metrics are tracked, the reporting cadence, and the limit and goal for each metric
user-invocable
true

Use this skill when quality gets reported as a number, and nobody can say whether it is better or worse than last time.

A snapshot is almost useless on its own. "Coverage is 71%" prompts an argument about whether 71 is good. "Coverage on changed lines has moved 62 → 71 over three releases, limit 75, goal 85" prompts a decision. Direction is the finding; the absolute value is context.

The second thing this catches is regression in the quality system - a threshold lowered, a job made non-blocking, a waiver renewed for the fourth time. Those never show up in a snapshot, because a snapshot reports what is measured, not what stopped being measured.

When to Use

  • quality reporting is a series of unconnected numbers
  • a metric is quoted with no baseline and no target
  • a team needs to show a quarter of improvement, or explain a quarter without it
  • a contract has been re-derived and the question is what moved
  • gates are being quietly weakened and nobody has noticed
  • a release decision needs direction, not just current state

Operating Principles

  • Limit / current / goal, always three numbers. Limit is the threshold that triggers action; current is measured; goal is the target. A metric with only a current value cannot be acted on.
  • Direction and magnitude, not just the delta. "−3pp" needs "inside the noise band" or "third consecutive fall" beside it to mean anything.
  • Archive the run, do not recompute history. Store each reading with its date, commit, and how it was measured. Recomputed history changes when the method changes, and then the trend is fiction.
  • A method change breaks the series. Say so, and start a new one rather than pretending the numbers are comparable.
  • Track the quality system, not only the code. Thresholds, gate blocking-ness, waiver counts and ages, suppression counts. A repo whose coverage rose while its threshold fell has got worse.
  • Fewer metrics, honestly measured. Six metrics a team acts on beat twenty nobody reads. Every metric needs a named owner and an action if it crosses its limit.
  • Every metric carries its caveat. Coverage does not prove correctness; a rising pass rate can mean weaker tests. Report the caveat inline, not in a footnote.
  • Never trend a metric that can be gamed without saying so. Coverage, test count, and defect count all move under pressure without quality moving.

Workflow

Phase 1: Fix the metric set

Start from what the contract already implies, and cap it at six to eight. From ./resources/trend-metrics.md:

MetricLimit / goal exampleCaveat to print
Diff coverage≥ 75% / ≥ 85%proves execution, not assertion quality
Repo coverage directionmust not decreasea big denominator absorbs new gaps
Flake rate< 1% / < 0.3%only measurable if retries are recorded
Suite duration (p95)< 10 min / < 5 minshortcuts appear when this rises
Contracted practices PRESENT- / all MUSTPRESENT means configured and enforced
Blockers open0 / 0a MUST practice missing
Waivers: count, expired, oldest0 expiredcount alone hides age
Suppressions totaltrending downincludes disables, ignores, skips
Escaped defectstrending downdepends on consistent triage
Eval pass rate per capabilityno regressionsper capability, never aggregated

Each metric needs an owner and a stated action at the limit. A metric with neither is a dashboard decoration.

Phase 2: Archive each run
.qa/
  quality-contract.md          # current
  trends.md                    # the report
  history/
    2026-05-02/{contract.md,metrics.json}
    2026-06-13/{contract.md,metrics.json}
    2026-08-21/{contract.md,metrics.json}

Each metrics.json records the value, the date, the commit, and how it was measured - the tool, the version, and the scope. That last field is what lets a future reader tell a real improvement from a method change.

Phase 3: Diff against the previous run

Per metric: previous, current, delta, direction, and position against limit and goal. Then, separately, the structural diff nobody else reports:

  • practices that moved into PRESENT - real wins, name them
  • practices that regressed out of PRESENT - the most important line in the report
  • practices newly N/A or deferred - usually a contract edit, so check it was deliberate
  • thresholds that changed, in either direction
  • gates that became non-blocking
  • waivers added, expired, renewed
  • suppressions added and removed

A threshold quietly lowered from 75 to 60 will otherwise appear as a coverage improvement. Reading the config diff alongside the metric diff is what catches it, and it is the single highest-value habit in this skill.

Phase 4: Establish the noise band

Before reporting a movement as a finding, know what movement is normal. Flake rate and suite duration move on their own; coverage moves with the size of the release; eval scores move with model variance.

Rule: a movement inside the noise band is not a finding, and a movement in the same direction three periods running is a finding regardless of size. Slow drift is what a threshold-based alert never catches.

Show full SKILL.md (506 more words)Show less
Phase 5: Report

Use ./resources/trend-report-template.md. Written to .qa/trends.md, with:

  1. Direction summary - improving, flat, or degrading, with the two or three metrics driving it
  2. Metric table - limit, previous, current, goal, direction, and the caveat
  3. Structural changes - practices gained and lost, thresholds and gates changed
  4. Governance - waiver and suppression movement
  5. Findings - regressions, three-period drifts, limits crossed
  6. Series breaks - where a method changed and comparison stops being valid

Then the honest closing paragraph: what this report cannot see. Untracked metrics, unmeasured practices, and anything where the tooling changed.

Cadence

Tie it to something that already happens - a release, a sprint boundary, a monthly review. A report with no cadence gets written once and admired.

Two rules: re-derive the contract on the same cadence, and read the previous report before writing the new one. A trend report that does not reference its predecessor's findings is a snapshot with a date on it.

Common Failure Modes

  • Hand-maintained numbers in prose. "We have 803 tests" in a README, wrong three months later. If it is worth tracking, generate it.
  • Recomputing history. New tool, new method, retroactively applied - the trend now shows a change that never happened.
  • Trending only code metrics. Missing the lowered threshold and the disabled job, which are the changes that made the code metrics look better.
  • Reporting deltas without noise bands. Every run has a finding, so nobody reads the findings.
  • Twenty metrics. Nobody acts on any of them.
  • Aggregating what should be split. One eval score across capabilities; one coverage number across a monorepo.
  • Celebrating a rising pass rate. It rises when tests get weaker, too. Pair it with flake rate and assertion quality.
  • No owner per metric. Nothing happens when a limit is crossed.

Resource Map

  • ./resources/trend-metrics.md - the metric set with limits, goals, how to measure each in a JS/TS repo, caveats, and gaming risks
  • ./resources/trend-report-template.md - the report structure, the history layout, the metrics.json shape, and worked findings
  • analyzing-quality-metrics - metric definitions, anti-metrics, and how to interpret each honestly; this skill adds history and direction
  • deriving-a-quality-contract - supplies the practice list, the thresholds, and the gap matrix each run diffs
  • governing-quality-waivers - waiver count, age and expiry are tracked metrics here
  • verifying-change-coverage - the source of the diff-coverage reading
  • assessing-comprehension-debt - its band is a trended, advisory metric
  • testing-llm-features - eval pass rate per capability, trended against a baseline
  • assessing-release-readiness - consumes direction, not just current state

Definition of Done

This skill is complete when:

  • the metric set is six to eight metrics, each with a limit, a goal, an owner, and a stated action at the limit
  • each run is archived with its value, date, commit, and measurement method
  • the report shows direction per metric, not only current values
  • structural changes are reported: practices gained and lost, thresholds changed, gates made non-blocking
  • movements are judged against a noise band, and three-period drifts are findings regardless of size
  • each metric's caveat is printed with it, not footnoted
  • series breaks caused by method changes are stated, and comparison stops there
  • the report names what it cannot see

© jaktestowac, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in skills/tracking-quality-trends of jaktestowac/awesome-copilot-for-testers.

  • SKILL.md
  • resources/trend-metrics.md
  • resources/trend-report-template.md

Open the folder on GitHubat commit 8910672

Compare with similar skills

Tracking Quality Trends next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tracking Quality Trends compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tracking Quality Trends this skilljaktestowac/awesome-copilot-for-testers116—~2.4kAutomated safety check: PassMIT
Cost Trackingaffaan-m/ECC276k1 repos~1.3kAutomated safety check: PassMIT
Apify Trend Analysissickn33/agentic-awesome-skills47k2 repos~1.2kAutomated safety check: NotesMIT
TrackingBuilderIO/agent-native7.1k—~8kAutomated safety check: PassNone
Matlab Read Write Point Cloud Filematlab/matlab-agentic-toolkit1.1k—~3.5kAutomated safety check: PassCustom licence
Time Trackingsickn33/agentic-awesome-skills47k1 repos~3.8kAutomated safety check: PassMIT

Similar skills

  • Cost Tracking

    affaan-m/ECC

    Track and report Claude Code token usage, spending, and budgets from the local ECC cost-tracker metrics log.

    276k GitHub starsUsed in 1 repo~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Apify Trend Analysis

    sickn33/agentic-awesome-skills

    Discover and track emerging trends across Google Trends, Instagram, Facebook, YouTube, and TikTok to inform content strategy.

    47k GitHub starsUsed in 2 repos~1.2k tokens
    Data & AnalyticsAuto-check: notes
  • Tracking

    BuilderIO/agent-native

    Server-side analytics tracking with pluggable providers. An agent skill from BuilderIO/agent-native.

    7.1k GitHub stars~8k tokensUpdated today
    Backend & APIsAuto-check passed
  • Matlab Read Write Point Cloud File

    matlab/matlab-agentic-toolkit

    Read and write 3-D point cloud data using Lidar Toolbox file I/O.

    1.1k GitHub stars~3.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Time Tracking

    sickn33/agentic-awesome-skills

    Time entry register: employee, project, client, task, hours, billable flag, rate and amount, invoice and approver.

    47k GitHub starsUsed in 1 repo~3.8k tokens
    Productivity & AutomationAuto-check passed
  • Horizon Track

    ruvnet/ruflo

    Track long-horizon objectives across multiple sessions with milestone checkpoints, progress persistence, and drift detection

    74k GitHub stars~744 tokensUpdated today
    Product & Project ManagementAuto-check: notes

More from jaktestowac/awesome-copilot-for-testers

All 13 skills in this repo
  • API Playwright Test Developer

    jaktestowac/awesome-copilot-for-testers

    Writes and reviews API automation tests with Playwright Test, covering setup/teardown, assertions, data management, and hybrid API+UI flows.

    116 GitHub stars~2k tokensUpdated 1 mo ago
    Auto-check passed
  • Assessing Comprehension Debt

    jaktestowac/awesome-copilot-for-testers

    Measures the risk that code shipped without anyone understanding it: a teach-back attestation on high-risk changes, a risk band from changed-code complexity, diff size and whether a human…

    116 GitHub stars~2.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Creating Orchestration Packs

    jaktestowac/awesome-copilot-for-testers

    Creates agent orchestration packs: cooperating .agent.md files with an orchestrator, subagents, matched handoffs, minimal tool grants, and a shared handoff packet contract.

    116 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Creating Plugins

    jaktestowac/awesome-copilot-for-testers

    Packages repository skills as installable Copilot plugins: marketplace registration, plugin.json manifests, generated skill copies, and the sync check CI enforces.

    116 GitHub stars~3.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Governing Quality Waivers

    jaktestowac/awesome-copilot-for-testers

    Turns "we will skip this check for now" into a dated, attributed, expiring waiver with a stated reason and owner, inventories the silent skips already hiding in a repo - skipped tests, disabled lint…

    116 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Recording Change Intent

    jaktestowac/awesome-copilot-for-testers

    Requires an externalised rationale for high-risk changes - new public exports, new endpoints, auth edits, migrations, removed guards - recorded as an Intent commit trailer, an ADR reference, or a…

    116 GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Tracking Quality Trends

What does Tracking Quality Trends do?

Turns point-in-time quality readings into a trend: archives each run, diffs against the previous one, and reports direction per metric - practices newly present or regressed, coverage movement…. Tracking Quality Trends is an agent skill from jaktestowac/awesome-copilot-for-testers. Turns point-in-time quality readings into a trend: archives each run, diffs against the previous one, and reports direction per metric - practices newly present or regressed, coverage movement, flake rate, waivers expiring, eval scores - using limit/current/goal framing.

When should I use Tracking Quality Trends?

Tracking Quality Trends fits situations like: quality reporting is a series of disconnected snapshots; A team needs to show improvement over a quarter; A number is quoted with no baseline; A regression in the quality system itself should be visible.

How do I install Tracking Quality Trends in Claude Code?

Run `npx skills add jaktestowac/awesome-copilot-for-testers --skill tracking-quality-trends -a claude-code`. Or copy the skill folder (skills/tracking-quality-trends in jaktestowac/awesome-copilot-for-testers) into .claude/skills/tracking-quality-trends in your project. Claude Code loads it when a task matches its description.

How do I install Tracking Quality Trends in Codex?

Run `npx skills add jaktestowac/awesome-copilot-for-testers --skill tracking-quality-trends -a codex`. Or copy the skill folder (skills/tracking-quality-trends in jaktestowac/awesome-copilot-for-testers) into .agents/skills/tracking-quality-trends in your project. Codex loads it when a task matches its description.

Can I use Tracking Quality Trends in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaktestowac/awesome-copilot-for-testers --skill tracking-quality-trends -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tracking-quality-trends, .gemini/skills/tracking-quality-trends, .github/skills/tracking-quality-trends and .opencode/skills/tracking-quality-trends in your project.

What does Tracking Quality Trends need to run?

SKILL.md names no scripts, command-line tools or credentials: Tracking Quality Trends is instructions for the agent only.

Does Tracking Quality Trends access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Tracking Quality Trends safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Tracking Quality Trends use?

Tracking Quality Trends is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tracking Quality Trends use?

About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Tracking Quality Trends?

Skills that share tags, products or a category with Tracking Quality Trends: Cost Tracking (affaan-m/ECC, 276k stars), Apify Trend Analysis (sickn33/agentic-awesome-skills, 47k stars), Tracking (BuilderIO/agent-native, 7.1k stars) and Matlab Read Write Point Cloud File (matlab/matlab-agentic-toolkit, 1.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tracking Quality Trends?

jaktestowac (a GitHub user) maintains it in jaktestowac/awesome-copilot-for-testers, which has 116 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on August 26, 2026.

Source: jaktestowac/awesome-copilot-for-testers on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.