Agent skill

Assessing Comprehension Debt

by jaktestowac in jaktestowac/awesome-copilot-for-testers

Measures the risk that code shipped without anyone understanding it: a teach-back attestation on high-risk changes, a risk band from changed-code complexity, diff size and whether a human…

MITAuto-check passed

Install Assessing Comprehension Debt

skills CLI
$ npx skills add jaktestowac/awesome-copilot-for-testers --skill assessing-comprehension-debt -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaktestowac/awesome-copilot-for-testers assessing-comprehension-debt --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaktestowac/awesome-copilot-for-testers.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/assessing-comprehension-debt .claude/skills/assessing-comprehension-debt && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
assessing-comprehension-debt
GitHub stars
116
Token cost
~2.3k tokens
SKILL.md length
1,234 words
Files
3
Skills in repo
13
Repo updated
First seen
Licence
MIT

At a glance

Measures the risk that code shipped without anyone understanding it: a teach-back attestation on high-risk changes, a risk band from changed-code complexity, diff size and whether a human…

  • Works in 4 steps: Establish scope and whether provenance… → Compute the risk band per change or module → Check teach-back attestation on… → …
  • An AI-assisted codebase grows faster than the team reads it
  • SKILL.md covers When to Use, Operating Principles, Workflow and What This Skill Cannot Do, plus 4 more sections
  • Calls git

What it does

Assessing Comprehension Debt is an agent skill from jaktestowac/awesome-copilot-for-testers. Measures the risk that code shipped without anyone understanding it: a teach-back attestation on high-risk changes, a risk band from changed-code complexity, diff size and whether a human explanation accompanied it, and optional AI-authorship provenance. Findings stay advisory by design. Use when an AI-assisted codebase grows faster than the team reads it, when reviews are rubber-stamped, when nobody can explain a module that ships weekly, or when leadership asks how much of the code the team can actually maintain.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `resources/risk-band-rubric.md` and `resources/teach-back-protocol.md`).

The repository describes itself as: 👨💻 Instructions, prompts, and chat modes to help You with test automation for GitHub Copilot 🤖. The licence is MIT.

When your agent uses it

  • An AI-assisted codebase grows faster than the team reads it
  • Reviews are rubber-stamped
  • Nobody can explain a module that ships weekly
  • Leadership asks how much of the code the team can actually maintain

Example prompts

  • “Use the assessing-comprehension-debt skill to measure the risk that code shipped without anyone understanding it: a teach-back attestation on…”
  • “/assessing-comprehension-debt”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Establish scope and whether provenance exists
  2. Compute the risk band per change or module
  3. Check teach-back attestation on high-risk surface
  4. Report the picture

What it can do on your machine

Read from SKILL.md and the folder at commit 8910672. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Assessing Comprehension Debt loads about 2.3k tokens when it runs. Until then it costs about 137 tokens; SKILL.md has 1,234 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~137
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaktestowac/awesome-copilot-for-testers at commit 8910672, republished under its MIT licence (© jaktestowac). 1,234 words, ~2,347 tokens.

Download SKILL.mdSave it as .claude/skills/assessing-comprehension-debt/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
assessing-comprehension-debt
description
Measures the risk that code shipped without anyone understanding it: a teach-back attestation on high-risk changes, a risk band from changed-code complexity, diff size and whether a human explanation accompanied it, and optional AI-authorship provenance. Findings stay advisory by design. Use when an AI-assisted codebase grows faster than the team reads it, when reviews are rubber-stamped, when nobody can explain a module that ships weekly, or when leadership asks how much of the code the team can actually maintain.
argument-hint
Commit range or release scope, whether AI-provenance trailers are in use, and which modules matter most
user-invocable
true

Assessing Comprehension Debt

Use this skill when code is arriving faster than anyone is reading it, and every other quality signal still looks fine.

Comprehension debt is the gap between the code that exists and the code the team actually understands. Tests pass, coverage holds, lint is clean - and nobody can explain why the pricing module works. When an agent writes the implementation and the tests, both signals go green without a single human forming a mental model. That is the debt this skill makes visible.

It is advisory, permanently, on purpose. Understanding lives in people's heads and cannot be proven by any signal. What can be measured is the risk of comprehension debt and the absence of evidence of understanding. A gate that claims to measure understanding is lying, and once someone notices, every finding it ever produced loses credibility.

When to Use

  • an AI-assisted or agent-generated codebase is growing faster than the team reads it
  • reviews are fast, approvals are frequent, and nobody asks questions
  • a module ships weekly and no one volunteers to explain it
  • a bus-factor conversation needs evidence instead of anecdote
  • a quarterly quality review should cover more than test coverage
  • onboarding is slow in a codebase that looks well-tested

Operating Principles

  • Never blocks. Findings cap below the project's blocking threshold. State that ceiling in the report, every time.
  • Measure absence of evidence, not understanding. Say it in those words. The band is a proxy; treat it as a prompt for a conversation, not a score.
  • The risk band is directional, not precise. It says "this change is the shape of one nobody understands", not "nobody understands this change".
  • AI provenance is a signal, never a penalty. The useful figure is AI-authored and unattested. AI-authored and well-understood is a good outcome, and if trailers are not in use the honest answer is "unknown".
  • A human explanation lowers the band. An Intent: trailer or an ADR is evidence that somebody thought about it. That is exactly the behaviour to reward.
  • Teach-back beats approval. "I approve" is a click. "I can explain what happens when the provider retries this webhook" is comprehension.
  • Never blame individuals. A high band on a module is a system outcome: review load, delivery pressure, tooling. Naming people converts a useful signal into something nobody will run twice.
  • Trend over snapshot. One reading is noise. Direction over a quarter is the finding.

Workflow

Phase 1: Establish scope and whether provenance exists

Pick the range: a release, a quarter, a module's history. Then check what signals are available:

bash
git log <range> --format='%h %an %s'
git log <range> --format='%(trailers:key=Intent,valueonly)'
git log <range> --format='%(trailers:key=Assisted-by,valueonly)'
git log <range> --format='%(trailers:key=Comprehension-Attested-by,valueonly)'
git log <range> --numstat

If Assisted-by: trailers are not in use, provenance is unknown - report it as unknown and move on. Do not infer AI authorship from commit size, style, or timing; those inferences are wrong often enough to poison the whole report.

Phase 2: Compute the risk band per change or module

Three inputs, from ./resources/risk-band-rubric.md:

InputSignal
Complexity addednew branches, new conditionals, nesting depth, new cross-module calls
Sizechanged lines, files touched, whether it lands as one commit or a reviewable sequence
Explanation presentan Intent: trailer, an ADR link, a module register entry, or a substantive review discussion

Bands: high (large, branch-heavy, no explanation), medium (one of the three), low (small, or well explained, or both).

A high band means: if nobody understands this, we would not be able to tell. That is all it means, and saying so plainly is what keeps the metric usable.

Phase 3: Check teach-back attestation on high-risk surface

For changes on high-risk surface (recording-change-intent has the rules), look for a record that a named human can explain it:

  • a Comprehension-Attested-by: <name> trailer, author-side or reviewer-side
  • an entry in the attestation register (attesting-manual-verification)

Status is FULL (attested), NONE (high-risk surface, no record), or N/A (not high-risk surface).

Where a record is missing and it matters, run the teach-back in ./resources/teach-back-protocol.md: four questions, ten minutes, and the answers tell you more than the band ever will. The protocol's value is not the record - it is that the conversation happens.

Phase 4: Report the picture
  • Risk bands by module, with the inputs that produced each one
  • Attestation status on high-risk surface, and the specific changes with none
  • Provenance, if known: AI-authored share, and the share that is AI-authored and unattested
  • Trend against the previous assessment, per module
  • Concentration - comprehension debt clusters. A module where one person authored everything and nobody reviewed it is the finding, regardless of any band
  • The ceiling, stated explicitly: these findings are advisory and do not fail the build

Then the only recommendations worth making: which specific modules deserve a teach-back session, which deserve a walkthrough written down, and which deserve a second pair of eyes on the next change.

Show full SKILL.md (453 more words)Show less

What This Skill Cannot Do

Worth stating in the report, because the temptation to over-read the number is strong:

  • It cannot tell you whether a person understands code. Only whether evidence exists.
  • It cannot distinguish elegant-and-obvious from complex-and-opaque. A high band on genuinely well-written code is a false positive; check before acting.
  • It cannot detect understanding that lives in a conversation, a diagram, or a head. Absence of a trailer is not absence of thought.
  • It cannot be used for performance assessment, and using it that way guarantees the trailers become theatre.

Common Failure Modes

  • Letting it block. The single failure that discredits the whole practice.
  • Treating the band as a verdict. "High band" is a prompt to have a conversation, not a defect.
  • Inferring AI authorship. From commit size, from phrasing, from the hour of the day. Wrong often, and corrosive when it is.
  • Penalising Assisted-by:. Teams stop using the trailer, and the one honest signal disappears.
  • Naming individuals. Converts a system signal into a personnel matter, and the next assessment never happens.
  • Confusing approval with comprehension. A merged PR with an approving click is not evidence of understanding.
  • Reporting one number for the repo. Comprehension debt concentrates; an average hides exactly the module you needed to see.
  • Skipping the teach-back. The band is the cheap proxy; the conversation is the actual value.

Resource Map

  • ./resources/risk-band-rubric.md - the three inputs, how to compute each, band boundaries, worked examples, and known false positives
  • ./resources/teach-back-protocol.md - the four questions, how to run a ten-minute session, what a weak answer looks like, and how to record the outcome
  • recording-change-intent - an Intent: record is an input to the band, and the sibling debt this one pairs with
  • attesting-manual-verification - where a teach-back record lives and how it expires
  • governing-quality-waivers - the third governance record: why a check is off
  • tech-debt-analysis - code and architecture health, as opposed to whether anyone understands it
  • documenting-test-suites - the usual remediation when a suite is unmaintainable by anyone who did not build it
  • code-review-advanced - where a teach-back conversation naturally belongs
  • analyzing-quality-metrics - for trending the band honestly, with its caveats attached

Definition of Done

This skill is complete when:

  • the scope is stated and the available signals are established, with provenance marked unknown where trailers are not in use
  • every change or module in scope has a band with the three inputs that produced it
  • teach-back status on high-risk surface is FULL / NONE / N/A, with the specific unattested changes named
  • concentration is reported, not just averages
  • the trend against a previous assessment is included where one exists
  • the advisory ceiling is stated explicitly in the report
  • recommendations are specific sessions and walkthroughs, not "improve documentation"
  • no individual is named as the cause of a band

© jaktestowac, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in skills/assessing-comprehension-debt of jaktestowac/awesome-copilot-for-testers.

  • SKILL.md
  • resources/risk-band-rubric.md
  • resources/teach-back-protocol.md

Open the folder on GitHubat commit 8910672

Compare with similar skills

Assessing Comprehension Debt next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Assessing Comprehension Debt compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Assessing Comprehension Debt this skilljaktestowac/awesome-copilot-for-testers116—~2.3kAutomated safety check: PassMIT
Conducting Cyber Risk Assessment With Nist 800 30mukul975/Anthropic-Cybersecurity-Skills34k—~2.5kAutomated safety check: PassApache-2.0
Feature Risk Assessmentanthropics/claude-for-legal9.6k2 repos~2.2kAutomated safety check: PassApache-2.0
Risk Measurement and Stress TestingHKUDS/Vibe-Trading35k—~3.8kAutomated safety check: PassMIT
Climate Risk Assessmentmohitagw15856/pm-claude-skills1.4k—~1.6kAutomated safety check: PassMIT
Legal Risk AssessmentTHUYRan/Legal-Skills-Chinese873—~3.1kAutomated safety check: PassNone

Similar skills

  • Conducting Cyber Risk Assessment With Nist 800 30

    mukul975/Anthropic-Cybersecurity-Skills

    Conduct a defensible cybersecurity risk assessment using the NIST SP 800-30 Rev 1 methodology: prepare scope and a risk model, identify threat sources and threat events, identify vulnerabilities and…

    34k GitHub stars~2.5k tokensUpdated 1 mo ago
    Legal & ComplianceAuto-check passed
  • Feature Risk Assessment

    anthropics/claude-for-legal

    Official

    Deeper risk assessment for a single feature or product area when the launch review found something that needs more than a line item.

    9.6k GitHub starsUsed in 2 repos~2.2k tokens
    Legal & ComplianceAuto-check passed
  • Measures portfolio and backtest risk with VaR, CVaR, maximum drawdown, Monte Carlo simulation, tail modeling and stress tests, using one tested risk module.

    35k GitHub stars~3.8k tokensUpdated yesterday
    Business, Finance & HRAuto-check passed
  • Climate Risk Assessment

    mohitagw15856/pm-claude-skills

    Assess physical and transition climate risk for a site, product, or portfolio with scenario-based structure.

    1.4k GitHub stars~1.6k tokensUpdated yesterday
    Legal & ComplianceAuto-check passed
  • Legal Risk Assessment

    THUYRan/Legal-Skills-Chinese

    Assess an enterprise’s regulatory penalty risk across four dimensions: licensing/qualifications, compliance with regulatory rules, and historical penalty/credit records.

    873 GitHub stars~3.1k tokensUpdated 1 mo ago
    Legal & ComplianceAuto-check passed
  • Risk Assessment Creator

    jeremylongshore/tons-of-skills-marketplace

    Create risk assessment creator operations. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~586 tokensUpdated yesterday
    Auto-check passed

More from jaktestowac/awesome-copilot-for-testers

All 13 skills in this repo
  • API Playwright Test Developer

    jaktestowac/awesome-copilot-for-testers

    Writes and reviews API automation tests with Playwright Test, covering setup/teardown, assertions, data management, and hybrid API+UI flows.

    116 GitHub stars~2k tokensUpdated 1 mo ago
    Auto-check passed
  • Creating Orchestration Packs

    jaktestowac/awesome-copilot-for-testers

    Creates agent orchestration packs: cooperating .agent.md files with an orchestrator, subagents, matched handoffs, minimal tool grants, and a shared handoff packet contract.

    116 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Creating Plugins

    jaktestowac/awesome-copilot-for-testers

    Packages repository skills as installable Copilot plugins: marketplace registration, plugin.json manifests, generated skill copies, and the sync check CI enforces.

    116 GitHub stars~3.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Governing Quality Waivers

    jaktestowac/awesome-copilot-for-testers

    Turns "we will skip this check for now" into a dated, attributed, expiring waiver with a stated reason and owner, inventories the silent skips already hiding in a repo - skipped tests, disabled lint…

    116 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Recording Change Intent

    jaktestowac/awesome-copilot-for-testers

    Requires an externalised rationale for high-risk changes - new public exports, new endpoints, auth edits, migrations, removed guards - recorded as an Intent commit trailer, an ADR reference, or a…

    116 GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Running Visual Regression Tests

    jaktestowac/awesome-copilot-for-testers

    Sets up and maintains visual regression testing: what to snapshot, baseline strategy, masking dynamic regions, threshold tuning, containerized baselines, and the review-and-update workflow.

    116 GitHub stars~2.6k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Assessing Comprehension Debt

What does Assessing Comprehension Debt do?

Measures the risk that code shipped without anyone understanding it: a teach-back attestation on high-risk changes, a risk band from changed-code complexity, diff size and whether a human…. Assessing Comprehension Debt is an agent skill from jaktestowac/awesome-copilot-for-testers. Measures the risk that code shipped without anyone understanding it: a teach-back attestation on high-risk changes, a risk band from changed-code complexity, diff size and whether a human explanation accompanied it, and optional AI-authorship provenance.

When should I use Assessing Comprehension Debt?

Assessing Comprehension Debt fits situations like: an AI-assisted codebase grows faster than the team reads it; reviews are rubber-stamped; nobody can explain a module that ships weekly; leadership asks how much of the code the team can actually maintain.

How do I install Assessing Comprehension Debt in Claude Code?

Run `npx skills add jaktestowac/awesome-copilot-for-testers --skill assessing-comprehension-debt -a claude-code`. Or copy the skill folder (skills/assessing-comprehension-debt in jaktestowac/awesome-copilot-for-testers) into .claude/skills/assessing-comprehension-debt in your project. Claude Code loads it when a task matches its description.

How do I install Assessing Comprehension Debt in Codex?

Run `npx skills add jaktestowac/awesome-copilot-for-testers --skill assessing-comprehension-debt -a codex`. Or copy the skill folder (skills/assessing-comprehension-debt in jaktestowac/awesome-copilot-for-testers) into .agents/skills/assessing-comprehension-debt in your project. Codex loads it when a task matches its description.

Can I use Assessing Comprehension Debt in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaktestowac/awesome-copilot-for-testers --skill assessing-comprehension-debt -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/assessing-comprehension-debt, .gemini/skills/assessing-comprehension-debt, .github/skills/assessing-comprehension-debt and .opencode/skills/assessing-comprehension-debt in your project.

What does Assessing Comprehension Debt need to run?

Going by SKILL.md and its folder, Assessing Comprehension Debt needs the command-line tools its instructions call (git).

Does Assessing Comprehension Debt access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Assessing Comprehension Debt safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Assessing Comprehension Debt use?

Assessing Comprehension Debt is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Assessing Comprehension Debt use?

About 2.3k tokens (SKILL.md is roughly 9.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Assessing Comprehension Debt?

Skills that share tags, products or a category with Assessing Comprehension Debt: Conducting Cyber Risk Assessment With Nist 800 30 (mukul975/Anthropic-Cybersecurity-Skills, 34k stars), Feature Risk Assessment (anthropics/claude-for-legal, 9.6k stars), Risk Measurement and Stress Testing (HKUDS/Vibe-Trading, 35k stars) and Climate Risk Assessment (mohitagw15856/pm-claude-skills, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Assessing Comprehension Debt?

jaktestowac (a GitHub user) maintains it in jaktestowac/awesome-copilot-for-testers, which has 116 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on August 26, 2026.

Source: jaktestowac/awesome-copilot-for-testers on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.