Agent skill

PII Annotation and Agreement Scorer

by glebis in glebis/claude-skills

Runs a human-first workflow for labeling PII spans in a transcript, then scores inter-annotator agreement and drafts an adjudicated gold set.

MITAuto-check passedLegal & Compliance

Install PII Annotation and Agreement Scorer

skills CLI
$ npx skills add glebis/claude-skills --skill annotate -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install glebis/claude-skills annotate --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/confide/skills/annotate .claude/skills/annotate && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
annotate
GitHub stars
389
Token cost
~1.3k tokens
SKILL.md length
588 words
Files
6 (incl. scripts, references, assets)
Skills in repo
91
Repo updated
First seen
Licence
MIT

At a glance

Runs a human-first workflow for labeling PII spans in a transcript, then scores inter-annotator agreement and drafts an adjudicated gold set.

  • Works in 6 steps: Open the tool. Double-click… → Read the rules. Open… → Set your annotator id and load the… → …
  • Building a human-labeled PII gold set for a de-identification system
  • SKILL.md covers Privacy invariants (do not…, Bundled assets, FOR THE ANNOTATOR (no coding… and FOR THE COORDINATOR (measure +…, plus 2 more sections
  • Runs Python scripts from its folder; calls python3

What it does

A zero-install, offline browser tool lets each human annotator open a transcript, read a ten-type PII codebook covering direct and quasi-identifiers like person, location, phone, email and medication, and label the minimal span of text that identifies someone, without rewriting or redacting anything; an uncertain call gets logged as a question rather than guessed at silently. Everything an annotator labels, including the real surface text spans, stays in their own browser and exported file on their own machine until they choose to export it.

A bundled scorer then computes Cohen's or Fleiss' kappa and span-level F1 across annotators' exported label files, builds a queue of where they disagreed, and drafts an adjudicated gold set from the results; a companion script can also turn an existing gold set into a synthetic reference annotator for testing the scorer alone. The privacy rule running through the whole skill is that only synthetic or consented transcripts are ever loaded, and only the resulting agreement statistics and disagreement clusters travel between people, never the original transcript text or the PII itself.

When your agent uses it

  • Building a human-labeled PII gold set for a de-identification system
  • Measuring inter-annotator agreement on sensitive-span labeling
  • Adjudicating disagreements between two or more PII annotators

Example prompts

  • “Set up the annotator tool for our two labelers on this synthetic transcript.”
  • “Score inter-annotator agreement once both annotators finish labeling.”
  • “Draft an adjudicated gold set from these disagreeing label files.”

Requirements

  • A synthetic or consented transcript
  • A modern browser for the annotation tool
  • Python (standard library) for the scoring script

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Open the tool. Double-click assets/annotator.html (or open it in Chrome/Firefox/
  2. Read the rules. Open references/codebook.md first. It defines the 10 types
  3. Set your annotator id and load the transcript in the tool (e.g. A, B, or your name).
  4. Label every PII span. Select the minimal text that identifies a real person (the client
  5. When unsure, log it — don't guess silently. Add a note starting with QUESTION: on the
  6. Export. Click Export → you get labels...json

What it can do on your machine

Read from SKILL.md and the folder at commit 7524dff. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PII Annotation and Agreement Scorer loads about 1.3k tokens when it runs, and up to ~5.1k if it reads all its reference files. Until then it costs about 162 tokens; SKILL.md has 588 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~162
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from glebis/claude-skills at commit 7524dff, republished under its MIT licence (© glebis). 588 words, ~1,342 tokens.

Download SKILL.mdSave it as .claude/skills/annotate/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
annotate
description
Build and verify a PII gold set with HUMAN annotators (first-class). Launch the browser annotator, label spans per the codebook, export per-annotator label files, then compute inter-annotator agreement (Cohen's/Fleiss' kappa) and draft an adjudicated gold. Use when the user says "annotate PII", "label this transcript", "build a gold set", "inter-annotator agreement", "review annotations", "adjudicate labels", or wants to measure/defend a de-identification gold standard. Local-only: synthetic or consented data only; annotators' names and transcript text stay on the machine — only labels/stats are collected, nothing PII is re-shared.

confide:annotate — human PII gold set + inter-annotator agreement

Humans label PII spans in a transcript; you measure how much they agree (κ) and draft an adjudicated gold from their labels. Annotators are first-class here — most of this skill is plain instructions FOR a person doing the labelling, plus a coordinator path to score it.

Privacy invariants (do not violate)

  • Synthetic or consented data only. Never load a real client transcript the person did not consent to share. When in doubt, anonymize first (confide:anon) and annotate the GREEN copy.
  • Names stay local. The annotator's labels (which contain real surface text spans) live in their browser and the exported JSON file on their own machine. Collect label files locally.
  • Nothing PII is re-shared. Only κ / F1 / disagreement clusters travel between people if needed. The transcript text and the original PII are never re-distributed by this skill.

Bundled assets

  • assets/annotator.html — zero-install browser annotation tool (EN/RU, runs offline).
  • references/codebook.md — the labelling rulebook (10 PII types, direct/quasi, harm).
  • references/tool-guide.md — how to drive the tool + scorer step by step.
  • scripts/score_iaa.py — Cohen's/Fleiss' κ, span-F1, disagreement queue, draft gold (stdlib).
  • scripts/gold_to_labels.py — turn an existing gold into a "reference annotator" to test solo.

FOR THE ANNOTATOR (no coding needed)

  1. Open the tool. Double-click assets/annotator.html (or open it in Chrome/Firefox/ Safari). It runs entirely in your browser — nothing is uploaded; labels stay on your machine until you Export.
  2. Read the rules. Open references/codebook.md first. It defines the 10 types (PERSON, LOCATION, ORG, PHONE, EMAIL, ID, DATE, MEDICATION, AGE, PROFESSION), what counts as a span (the minimal identifying text), and direct vs. quasi-identifier.
  3. Set your annotator id and load the transcript in the tool (e.g. A, B, or your name). Use only synthetic or consented text.
  4. Label every PII span. Select the minimal text that identifies a real person (the client or third parties they mention) and assign its type. Record direct/quasi, entity id, role, and harm as the codebook describes. Do not rewrite or redact — only label.
  5. When unsure, log it — don't guess silently. Add a note starting with QUESTION: on the span (e.g. QUESTION: gym or city?). These flow straight into the adjudication queue.
  6. Export. Click Export → you get labels.<doc>.<annotator>.json (schema: {doc_id, annotator, text, spans:[{start,end,text,type,...}]}). Keep it local and hand only this file to the coordinator. Two+ people should label the same doc independently (blind) for a meaningful κ.
Show full SKILL.md (200 more words)Show less

FOR THE COORDINATOR (measure + adjudicate)

  1. Collect every labels.<doc>.<annotator>.json into one folder, e.g. labels/.
  2. Score IAA:
    bash
    python3 skills/annotate/scripts/score_iaa.py --labels-dir labels/ --out-dir results/
    It writes (per doc + overall): Cohen's κ (pairwise), Fleiss' κ (3+ annotators), span-F1, a disagreement queue (*-iaa-disagreements.json: every cluster annotators don't fully agree on, plus any QUESTION: spans), and a draft adjudicated gold (*-adjudicated-gold-draft.json: majority span per overlap-cluster, ties/questions marked needs_review:true). Character-level κ sidesteps tokenization disputes.
  3. Target κ ≥ 0.80 = a defensible gold. Lower usually means an unclear codebook rule, not a careless annotator — fix the rule and re-label, don't just discard.
  4. Adjudicate. Walk the disagreement queue with a human adjudicator; resolve each needs_review cluster. The resulting label set is the published gold; report post-adjudication κ too. Nothing is ever auto-finalised.

Test the loop solo (no second person yet)

Treat an existing gold JSONL as one "reference annotator", label the same doc yourself in annotator.html as another, then score the pair:

bash
python3 skills/annotate/scripts/gold_to_labels.py --gold GOLD.jsonl --name gold --out-dir labels/
# label the same doc yourself in annotator.html as "me" -> drop labels.<doc>.me.json into labels/
python3 skills/annotate/scripts/score_iaa.py --labels-dir labels/ --out-dir results/

(--sessions-dir DIR lets gold_to_labels.py read transcript text from disk so char offsets match the gold exactly.)

Output

IAA results (κ, F1) + a disagreement list + a draft adjudicated gold — labels/stats only. Transcript text and original PII stay local; only what's needed to adjudicate is shared.

© glebis, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references, assets) in confide/skills/annotate of glebis/claude-skills.

  • SKILL.md
  • assets/annotator.html
  • references/codebook.md
  • references/tool-guide.md
  • scripts/gold_to_labels.py
  • scripts/score_iaa.py

Open the folder on GitHubat commit 7524dff

Compare with similar skills

PII Annotation and Agreement Scorer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PII Annotation and Agreement Scorer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PII Annotation and Agreement Scorer this skillglebis/claude-skills389—~1.3kAutomated safety check: PassMIT
De-identification Leakage Auditmaziyarpanahi/openmed5.5k—~1.9kAutomated safety check: PassApache-2.0
Data Lineage Trackingmukul975/Privacy-Data-Protection-Skills295—~1.8kAutomated safety check: PassApache-2.0
Implementing Cloud Dlp For Data Protectionmukul975/Anthropic-Cybersecurity-Skills34k—~4.2kAutomated safety check: PassApache-2.0
Cursor Compliance Auditjeremylongshore/tons-of-skills-marketplace2.8k—~2.3kAutomated safety check: NotesMIT
Posthog Data Handlingjeremylongshore/tons-of-skills-marketplace2.8k—~2.7kAutomated safety check: PassMIT

Similar skills

  • De-identification Leakage Audit

    maziyarpanahi/openmed

    Scans text that has already been de-identified for leftover identifiers such as SSNs, card numbers, emails and dates, and blocks release if anything turns up.

    5.5k GitHub stars~1.9k tokensUpdated today
    Legal & ComplianceAuto-check passed
  • Data Lineage Tracking

    mukul975/Privacy-Data-Protection-Skills

    Implements data lineage tracking for privacy compliance including origin tracking, transformation logging, access auditing, deletion verification, and cross-system lineage graphs.

    295 GitHub stars~1.8k tokensUpdated 6 mo ago
    Legal & ComplianceAuto-check passed
  • Implementing Cloud Dlp For Data Protection

    mukul975/Anthropic-Cybersecurity-Skills

    Implement cloud DLP using Amazon Macie, Google Cloud DLP API, Microsoft Purview, Azure Information Protection, and Nightfall AI to discover, classify, label, de-identify, and protect sensitive data…

    34k GitHub stars~4.2k tokensUpdated 1 mo ago
    Legal & ComplianceAuto-check passed
  • Cursor Compliance Audit

    jeremylongshore/tons-of-skills-marketplace

    Compliance and security auditing for Cursor IDE usage: SOC 2, GDPR, HIPAA assessment, evidence collection, and remediation.

    2.8k GitHub stars~2.3k tokensUpdated today
    Legal & ComplianceAuto-check: notes
  • Posthog Data Handling

    jeremylongshore/tons-of-skills-marketplace

    Implement consent-aware PostHog collection, PII minimization, masking, retention, and deletion workflows.

    2.8k GitHub stars~2.7k tokensUpdated today
    Legal & ComplianceAuto-check passed
  • Cometchat Compliance

    cometchat/cometchat-skills

    Data governance & compliance for CometChat — pick the data-residency region, satisfy GDPR/CCPA (right-to-erasure and data export), plan message retention & purge, and produce audit / eDiscovery…

    130 GitHub stars~1.7k tokensUpdated 3 days ago
    Legal & ComplianceAuto-check passed

More from glebis/claude-skills

All 91 skills in this repo
  • Automates a dedicated, logged-in Chrome instance per profile without ever closing the user's own open tabs or browser windows.

    389 GitHub stars~973 tokensUpdated 12 days ago
    Auto-check passed
  • Deep Research

    glebis/claude-skills

    This skill should be used when conducting comprehensive research on any topic using the OpenAI Deep Research API.

    389 GitHub stars~2.6k tokensUpdated 12 days ago
    Auto-check: notes
  • Elimination Research

    glebis/claude-skills

    This skill should be used for elimination-style research where the user wants to choose from a shortlist of products, tools, services, vendors, or other options using explicit criteria, numeric…

    389 GitHub stars~1.6k tokensUpdated 12 days ago
    Auto-check passed
  • Narrated HTML Presentations

    glebis/claude-skills

    Generates a self-contained HTML presentation with article and slides modes, ElevenLabs voiceover narration and optional GPT Image 2 illustrations.

    389 GitHub stars~2.3k tokensUpdated 12 days ago
    Auto-check: notes
  • Writes fictional but realistic coaching or therapy session transcripts for evals, demos and few-shot examples, in several modalities and export formats.

    389 GitHub stars~2.9k tokensUpdated 12 days ago
    Auto-check passed
  • The Goal Automation Diagnostic

    glebis/claude-skills

    Walks through Goldratt's Five Focusing Steps to find the real bottleneck in your work, then recommends one automation aimed at it and a list of what not to automate.

    389 GitHub stars~2k tokensUpdated 12 days ago
    Auto-check passed

Questions about PII Annotation and Agreement Scorer

What does PII Annotation and Agreement Scorer do?

Runs a human-first workflow for labeling PII spans in a transcript, then scores inter-annotator agreement and drafts an adjudicated gold set. A zero-install, offline browser tool lets each human annotator open a transcript, read a ten-type PII codebook covering direct and quasi-identifiers like person, location, phone, email and medication, and label the minimal span of text that identifies someone, without rewriting or redacting anything; an uncertain call gets logged as a question rather than guessed at silently. Everything an annotator labels, including the real surface text spans, stays in their own browser and exported file on their own machine until they choose to export it.

When should I use PII Annotation and Agreement Scorer?

PII Annotation and Agreement Scorer fits situations like: building a human-labeled PII gold set for a de-identification system; measuring inter-annotator agreement on sensitive-span labeling; adjudicating disagreements between two or more PII annotators.

How do I install PII Annotation and Agreement Scorer in Claude Code?

Run `npx skills add glebis/claude-skills --skill annotate -a claude-code`. Or copy the skill folder (confide/skills/annotate in glebis/claude-skills) into .claude/skills/annotate in your project. Claude Code loads it when a task matches its description.

How do I install PII Annotation and Agreement Scorer in Codex?

Run `npx skills add glebis/claude-skills --skill annotate -a codex`. Or copy the skill folder (confide/skills/annotate in glebis/claude-skills) into .agents/skills/annotate in your project. Codex loads it when a task matches its description.

Can I use PII Annotation and Agreement Scorer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add glebis/claude-skills --skill annotate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/annotate, .gemini/skills/annotate, .github/skills/annotate and .opencode/skills/annotate in your project.

What does PII Annotation and Agreement Scorer need to run?

Going by SKILL.md and its folder, PII Annotation and Agreement Scorer needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: A synthetic or consented transcript; A modern browser for the annotation tool; Python (standard library) for the scoring script.

Does PII Annotation and Agreement Scorer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is PII Annotation and Agreement Scorer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does PII Annotation and Agreement Scorer use?

PII Annotation and Agreement Scorer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PII Annotation and Agreement Scorer use?

About 1.3k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.7k tokens, read only when the agent opens those files.

What are the alternatives to PII Annotation and Agreement Scorer?

Skills that share tags, products or a category with PII Annotation and Agreement Scorer: De-identification Leakage Audit (maziyarpanahi/openmed, 5.5k stars), Data Lineage Tracking (mukul975/Privacy-Data-Protection-Skills, 295 stars), Implementing Cloud Dlp For Data Protection (mukul975/Anthropic-Cybersecurity-Skills, 34k stars) and Cursor Compliance Audit (jeremylongshore/tons-of-skills-marketplace, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PII Annotation and Agreement Scorer?

glebis (a GitHub user) maintains it in glebis/claude-skills, which has 389 GitHub stars. The repository holds 91 skills in this directory. The repository was last updated on September 26, 2026.

Source: glebis/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.