Agent skill

Pii Safe Documents

by danyuchn in danyuchn/pii-guard

Processes sensitive local documents through PII Guard and a local Ollama model into a reversible redacted copy, without letting the main agent read the original or restored contents.

MITAuto-check passedAI & LLM Engineering

Install Pii Safe Documents

skills CLI
$ npx skills add danyuchn/pii-guard --skill pii-safe-documents -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install danyuchn/pii-guard pii-safe-documents --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/danyuchn/pii-guard.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/pii-safe-documents .claude/skills/pii-safe-documents && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pii-safe-documents
GitHub stars
249
Token cost
~2.7k tokens
SKILL.md length
1,478 words
Files
4 (incl. scripts)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Processes sensitive local documents through PII Guard and a local Ollama model into a reversible redacted copy, without letting the main agent read the original or restored contents.

  • Works in 5 steps: Path-only preflight → Create the redacted working copy → Work only on the redacted copy → …
  • Tasks that involve LLM inference and serving
  • SKILL.md covers Non-negotiable isolation rules, Supported inputs, Workflow and Safe status language, plus 1 more section
  • Runs Python scripts from its folder; calls python3 and git

What it does

Pii Safe Documents is an agent skill from danyuchn/pii-guard. Processes sensitive local documents through PII Guard and a local Ollama model into a reversible redacted copy, without letting the main agent read the original or restored contents.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts (for example `evals/evals.json`, `scripts/pii_safe_workflow.py` and `tests/test_pii_safe_workflow.py`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with Ollama. The repository describes itself as: 繁體中文(台灣)個人資料去識別化工具,讓業務文件可以安全地送進 AI 處理 / Reversible Traditional-Chinese (Taiwan) PII de-identification for LLM workflows — fully local. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM inference and serving

Example prompts

  • “Use the pii-safe-documents skill to process sensitive local documents through PII Guard and a local Ollama model into a reversible redacted copy…”
  • “/pii-safe-documents”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Path-only preflight
  2. Create the redacted working copy
  3. Work only on the redacted copy
  4. Restore without reading the result
  5. Retain or purge the private map

What it can do on your machine

Read from SKILL.md and the folder at commit ff36613. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pii Safe Documents loads about 2.7k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 1,478 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from danyuchn/pii-guard at commit ff36613, republished under its MIT licence (© danyuchn). 1,478 words, ~2,689 tokens.

Download SKILL.mdSave it as .claude/skills/pii-safe-documents/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
pii-safe-documents
description
Processes sensitive local documents through PII Guard and a local Ollama model into a reversible redacted copy, without letting the main agent read the original or restored contents.

PII Safe Documents

This skill creates a reversible, locally redacted working copy with a strict main-agent workflow boundary:

  • The main agent is untrusted for raw data. It receives paths and safe receipts only.
  • The bundled open-source local wrapper may read the original solely to run PII Guard and local Ollama.
  • No hook is required. Normal wrapper output suppresses raw content.

This is strong protection against accidental model exposure, not an OS security boundary against a malicious process running as the same macOS user. A Skill cannot revoke its own filesystem tools. Hostile-agent isolation requires a separately permissioned local broker or OS account. Never describe this Skill alone as mathematically or technically impossible to bypass.

Non-negotiable isolation rules

When this skill is active, the main agent MUST NOT:

  1. Read, preview, search, summarize, diff, upload, attach, or otherwise inspect the original file.
  2. Use cat, sed, head, tail, grep, rg, strings, Python, a document reader, a browser, or any other tool on the original file.
  3. Read or reveal mapping.private.json, private worker files, raw logs, or the restored output.
  4. Pass the original path to a subagent, cloud model, MCP server, website, or third-party API.
  5. Run PII Guard directly. Its libraries may echo source text in warnings; only the bundled wrapper may invoke it.
  6. Debug a failure by opening the original, mapping, worker output, or restored document.
  7. Run git diff, content scans, or indexing over a directory that contains the original or restored output.
  8. Reach the annotation page, run the review subcommand, or read a term file the user wrote for mask. All three carry unredacted values. The annotation URL is deliberately withheld from you; do not reconstruct it, scan for the port, or ask the user to paste it.

These rules still apply if the user asks the main agent to “check quickly.” If raw inspection is genuinely required, stop this workflow and obtain explicit permission for a different trust model.

Supported inputs

Use this version for UTF-8 plain-text files up to 64 KiB only: .txt, .md, .csv, .tsv, .log, and .dat.

Do not rename a binary document to bypass this restriction. For .docx, .xlsx, .pdf, images, or audio, report that this version does not yet provide a verified isolation-preserving parser.

This is a PII redactor, not a general confidentiality classifier. Amounts, health details, schedules, contract terms, business strategy, and other sensitive facts may remain visible when they do not identify a person. Do not use this skill alone to claim that an entire document is safe for external disclosure.

Workflow

1. Path-only preflight

The user may provide the original path. You may check path metadata such as existence, suffix, and file size, but never its contents. Do not autocomplete or glob inside a sensitive directory.

Choose allowed terms only when the user explicitly wants them preserved, such as a company or product name. An allowed term is visible to the main agent because the user supplied it; do not discover allowed terms from the original.

1b. 第一階段快速模式(不啟動 Ollama)

需要只做決定性本機處理時,可明確使用 quick:

bash
python3 <skill-dir>/scripts/pii_safe_workflow.py quick \
  --input "/absolute/path/to/private-file.txt"

quick 只呼叫共用的 PII Guard 核心(Presidio、台灣規則與 CKIP),不啟動、連線或探測 Ollama。它和 repo 的 pii_guard quick CLI 及 localhost 網頁使用同一個 ~/.local/share/pii-safe-documents/jobs/<job_id>/ 私有工作目錄、mapping、快照與還原邏輯。回執只含 job ID、去識別化檔案路徑、數量、摘要與 roundtrip_verified,不含原文或 mapping 值。

quick 回執成功後,主 agent 只能讀 redacted_path;仍應讓使用者在本機網頁人工快審,因為決定性偵測可能漏掉或誤遮。還原前保留 job ID,完成後以 purge 手動清除,不會自動 TTL。

2. Create the redacted working copy

Run:

bash
python3 <skill-dir>/scripts/pii_safe_workflow.py redact \
  --input "/absolute/path/to/private-file.txt" \
  --allow "company name the user explicitly supplied"

Repeat --allow as needed. The wrapper runs deterministic PII Guard detection plus the same chunked, three-sample local Ollama audit used by the localhost enhanced mode, captures all raw output, creates a private job directory, and prints only a safe JSON receipt. The verified default model is ornith-1.5:9b; override it only after a representative local accuracy and speed test.

If the receipt says both redaction_checks_passed: true and agent_may_read_redacted: true, the main agent may read only redacted_path. The receipt also provides safe replacement counts, audit-pass count, and the local model name. Keep job_id for restoration. Never infer or probe the mapping path.

If the command fails, report its safe error code and stop. In particular, NO_PII_CONFIDENCE means the detector found no reversible replacements and therefore withheld the copy instead of calling an unchanged file safe. ADVERSARIAL_INPUT_REVIEW_REQUIRED means instruction-like document text could interfere with the local model, so the wrapper refused automated release. Do not inspect hidden files or rerun lower-level commands.

2b. Offer the user a manual pass over the redacted copy

The detector and the audit both miss things, and both over-redact. The user is the backstop, and this step is where they act on what they see. Offer it whenever the redacted copy will be used for anything that matters; do not skip it silently.

Run:

bash
python3 <skill-dir>/scripts/pii_safe_workflow.py annotate --job-id "<job_id>"

This opens a page in the user's browser and blocks until they close it out. Tell them it has opened and what to do there; then wait.

On that page the user can:

  • Select any still-visible text and mask it. Every occurrence is masked, not just the selected one.
  • Click a marker to see the value behind it and put it back. This is for text that should never have been redacted — typically a court, hospital, or company name whose removal makes the document unusable.

Neither action requires comparing against the original document.

The page is not addressable by you. The URL carries a single-use token minted inside the private worker and passed only to the browser it opens; it is never printed, and the receipt you get back contains counts, not a URL. Do not attempt to discover the port, reconstruct the URL, or fetch the page. Do not ask the user to paste the URL, the page, or any value from it — ask only for what they want done, or let them do it themselves on the page.

When the user finishes, the command returns a receipt with terms_masked and markers_restored. Every edit is persisted and re-verified as it happens, so closing the browser early loses only unmade edits, never made ones.

Re-read redacted_path afterwards; its contents and redacted_sha256 have changed.

For a headless machine with no browser, the same two operations exist as mask --terms <file> and unmask --marker TYPE-N, with review to list markers and values. review prints unredacted values, refuses when its output is not a terminal, and must be run by the user, never by you.

Show full SKILL.md (446 more words)Show less
3. Work only on the redacted copy

Read and edit only the redacted working copy. Preserve placeholders exactly, including brackets, capitalization, and job namespace. Never normalize, translate, renumber, or combine them.

Save the edited redacted document as another UTF-8 text file. Prefer the same private job directory or a user-approved destination. Before restoration, verify mechanically that every placeholder from the redacted working copy is still present; do not open the mapping to do this.

In Markdown or Obsidian files, a placeholder inserted into a person-bearing link slug can temporarily make that link nonfunctional. Preserve the placeholder and surrounding link syntax exactly; restoration recreates the original link.

4. Restore without reading the result

Run:

bash
python3 <skill-dir>/scripts/pii_safe_workflow.py restore \
  --job-id "<job_id from receipt>" \
  --input "/absolute/path/to/edited-redacted-file.txt" \
  --output "/absolute/path/chosen/by/user/restored-file.txt"

The wrapper prints a safe receipt. After success, tell the user the output path, but do not read, preview, diff, hash through a content-printing tool, or summarize the restored file. A digest and roundtrip_equal boolean shown by the wrapper are safe to relay. roundtrip_equal: true is expected only when the redacted working copy was not intentionally edited.

5. Retain or purge the private map

The mapping is required for later restoration and is stored with restrictive permissions. Keep it by default. Purging is destructive, so do it only after the user explicitly confirms that no further restoration is needed:

bash
python3 <skill-dir>/scripts/pii_safe_workflow.py purge --job-id "<job_id>"

Safe status language

You may report:

  • job ID;
  • readable redacted path;
  • restored output path;
  • whether the local audit passed;
  • the redacted-file digest and permission checks emitted by the wrapper.

Never report original values, mapping entries, raw model output, raw warning text, or excerpts from the original/restored document.

Security notes

  • The wrapper refuses network Ollama endpoints; only loopback addresses are accepted.
  • Job directories use mode 0700; sensitive files use mode 0600.
  • Existing placeholder-like text is protected before redaction to prevent restoration collisions.
  • Allowed terms are protected before detection rather than restored afterward.
  • The local audit uses bounded overlapping chunks, a system/user role boundary, schema-constrained JSON, and repeated residual passes to catch aliases and contextual identifiers missed by rule-based detection.
  • Local-model guesses are replaceable only when they match exactly or normalize to one unique source span in both the original and current redacted document; ambiguous or hallucinated values fail closed and never enter the restoration map.
  • Inputs are copied through a single-open, non-symlink private snapshot before a worker reads them, preventing path swaps during processing.
  • The local model connection bypasses system proxies and verifies that port 11434 belongs to this user's Ollama process.
  • No automated detector is perfect. agent_may_read_redacted: true means the configured local checks passed, not that zero privacy risk is mathematically guaranteed.
  • Document text can try to mislead an LLM. The local audit is a supplemental detector, not a proof against adversarial prompt injection.

© danyuchn, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in .agents/skills/pii-safe-documents of danyuchn/pii-guard.

  • SKILL.md
  • evals/evals.json
  • scripts/pii_safe_workflow.py
  • tests/test_pii_safe_workflow.py

Open the folder on GitHubat commit ff36613

Compare with similar skills

Pii Safe Documents next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pii Safe Documents compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pii Safe Documents this skilldanyuchn/pii-guard249—~2.7kAutomated safety check: PassMIT
Subwave LLM Benchperminder-klair/subwave1.4k—~2.4kAutomated safety check: NotesMIT
Visiongridaco/grida2.7k—~1.5kAutomated safety check: PassApache-2.0
Domodomo Local AI Maintenancedarknecrocities/DomoDomo---All-in-one-Tool239—~17kAutomated safety check: PassNone
Cc Ollamamathruffian-dot/claude-code-lazy-packs254—~118Automated safety check: PassMIT
Bdistill Knowledge Extractionsickn33/agentic-awesome-skills47k2 repos~926Automated safety check: PassMIT

Similar skills

  • Subwave LLM Bench

    perminder-klair/subwave

    Benchmark and compare LLM models for SUB/WAVE's on-air calls — track picks, segments, listener requests, DJ scripts, banter, and programme beats — in both candidate-pool and agent modes, using…

    1.4k GitHub stars~2.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Vision

    gridaco/grida

    Query images with a local Ollama vision model without loading the image into the main agent context.

    2.7k GitHub stars~1.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Domodomo Local AI Maintenance

    darknecrocities/DomoDomo---All-in-one-Tool

    Maintain DomoDomo private local AI features, Ollama connections, browser inference, streaming UX, embeddings, RAG, memory, and agent interfaces.

    239 GitHub stars~17k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Cc Ollama

    mathruffian-dot/claude-code-lazy-packs

    Claude Code 安裝本地 AI Ollama。說「安裝 Ollama」「本地 AI」時載入. An agent skill from mathruffian-dot/claude-code-lazy-packs.

    254 GitHub stars~118 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Bdistill Knowledge Extraction

    sickn33/agentic-awesome-skills

    Extract structured domain knowledge from AI models in-session or from local open-source models via Ollama.

    47k GitHub starsUsed in 2 repos~926 tokens
    AI & LLM EngineeringAuto-check passed
  • Pudu Task Telemetry

    davila7/claude-code-templates

    Measure local AI task latency, token usage, errors and verified outcomes using Pudu AI hardware evidence and installed Ollama models.

    32k GitHub stars~1.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

Works with

Questions about Pii Safe Documents

What does Pii Safe Documents do?

Processes sensitive local documents through PII Guard and a local Ollama model into a reversible redacted copy, without letting the main agent read the original or restored contents. Pii Safe Documents is an agent skill from danyuchn/pii-guard. Processes sensitive local documents through PII Guard and a local Ollama model into a reversible redacted copy, without letting the main agent read the original or restored contents.

When should I use Pii Safe Documents?

Pii Safe Documents fits situations like: tasks that involve LLM inference and serving.

How do I install Pii Safe Documents in Claude Code?

Run `npx skills add danyuchn/pii-guard --skill pii-safe-documents -a claude-code`. Or copy the skill folder (.agents/skills/pii-safe-documents in danyuchn/pii-guard) into .claude/skills/pii-safe-documents in your project. Claude Code loads it when a task matches its description.

How do I install Pii Safe Documents in Codex?

Run `npx skills add danyuchn/pii-guard --skill pii-safe-documents -a codex`. Or copy the skill folder (.agents/skills/pii-safe-documents in danyuchn/pii-guard) into .agents/skills/pii-safe-documents in your project. Codex loads it when a task matches its description.

Can I use Pii Safe Documents in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add danyuchn/pii-guard --skill pii-safe-documents -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pii-safe-documents, .gemini/skills/pii-safe-documents, .github/skills/pii-safe-documents and .opencode/skills/pii-safe-documents in your project.

What does Pii Safe Documents need to run?

Going by SKILL.md and its folder, Pii Safe Documents needs Python for the scripts in its folder and the command-line tools its instructions call (python3 and git). Our summary lists: Python 3.

Does Pii Safe Documents access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Pii Safe Documents safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Pii Safe Documents use?

Pii Safe Documents is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pii Safe Documents use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Pii Safe Documents?

Skills that share tags, products or a category with Pii Safe Documents: Subwave LLM Bench (perminder-klair/subwave, 1.4k stars), Vision (gridaco/grida, 2.7k stars), Domodomo Local AI Maintenance (darknecrocities/DomoDomo---All-in-one-Tool, 239 stars) and Cc Ollama (mathruffian-dot/claude-code-lazy-packs, 254 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pii Safe Documents?

danyuchn (a GitHub user) maintains it in danyuchn/pii-guard, which has 249 GitHub stars. The repository was last updated on October 2, 2026.

Source: danyuchn/pii-guard on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.