Agent skill

Pre Commit Audit

by flonat in flonat/flonat-research

Deliver a fast pre-commit safety scan: file size, anonymity (author / affiliation strings in tex/bib), hardcoded secrets, and invisible-Unicode carriers.

MITAuto-check: notesDocuments & Office

Install Pre Commit Audit

skills CLI
$ npx skills add flonat/flonat-research --skill pre-commit-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install flonat/flonat-research pre-commit-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/flonat/flonat-research.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/pre-commit-audit .claude/skills/pre-commit-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pre-commit-audit
GitHub stars
145
Token cost
~2.8k tokens
SKILL.md length
1,098 words
Files
2 (incl. scripts)
Skills in repo
83
Repo updated
First seen
Licence
MIT

At a glance

Deliver a fast pre-commit safety scan: file size, anonymity (author / affiliation strings in tex/bib), hardcoded secrets, and invisible-Unicode carriers.

  • Works in 5 steps: Size → Anonymity → Secrets → …
  • The user requests a fast pre-commit safety scan: file size
  • SKILL.md covers Hard Rules, When to Use, When NOT to Use and Modes, plus 8 more sections
  • Runs Python scripts from its folder; calls git and uv

What it does

Pre Commit Audit is an agent skill from flonat/flonat-research. Deliver a fast pre-commit safety scan: file size, anonymity (author / affiliation strings in tex/bib), hardcoded secrets, and invisible-Unicode carriers. Use when the user requests a fast pre-commit safety scan: file size, anonymity (author / affiliation strings in tex/bib), hardcoded secrets, and invisible-Unicode carriers. Triggers: 'audit before commit', 'check before push', 'pre-commit scan', 'safety check'.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/check_invisible_unicode.py`).

It sits in Documents & Office, covering LaTeX and Secrets management. The repository describes itself as: Shareable Claude Code + Codex infrastructure for PhD researchers — skills, agents, hooks, and rules for academic workflows. The licence is MIT.

When your agent uses it

  • The user requests a fast pre-commit safety scan: file size
  • Anonymity (author / affiliation strings in tex/bib)
  • Hardcoded secrets
  • Invisible-Unicode carriers

Example prompts

  • “audit before commit”
  • “check before push”
  • “pre-commit scan”
  • “/pre-commit-audit”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Bash(git*), Bash(grep*), Bash(find*), Bash(stat*), Bash(numfmt*), Bash(wc*), Bash(head*), Bash(uv*), AskUserQuestion

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Size
  2. Anonymity
  3. Secrets
  4. Invisible Unicode
  5. Consolidate + ask

What it can do on your machine

Read from SKILL.md and the folder at commit da27600. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash(git*)
    • Bash(grep*)
    • Bash(find*)
    • Bash(stat*)
    • Bash(numfmt*)
    • Bash(wc*)
    • Bash(head*)
    • Bash(uv*)
    • AskUserQuestion

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • git
    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pre Commit Audit loads about 2.8k tokens when it runs. Until then it costs about 108 tokens; SKILL.md has 1,098 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~108
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:109
    | `.env`-style assignments | `^[A-Z_]+=[a-zA-Z0-9+/=_-]{20,}` outside `.env.example`/`.env.template` |

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from flonat/flonat-research at commit da27600, republished under its MIT licence (© flonat). 1,098 words, ~2,768 tokens.

Download SKILL.mdSave it as .claude/skills/pre-commit-audit/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
pre-commit-audit
description
Deliver a fast pre-commit safety scan: file size, anonymity (author / affiliation strings in tex/bib), hardcoded secrets, and invisible-Unicode carriers. Use when the user requests a fast pre-commit safety scan: file size, anonymity (author / affiliation strings in tex/bib), hardcoded secrets, and invisible-Unicode carriers. Triggers: 'audit before commit', 'check before push', 'pre-commit scan', 'safety check'.
allowed-tools
Read, Bash(git*), Bash(grep*), Bash(find*), Bash(stat*), Bash(numfmt*), Bash(wc*), Bash(head*), Bash(uv*), AskUserQuestion
argument-hint
[--staged | --unstaged | --all]

Pre-Commit Audit — File Size, Anonymity, Secrets

Three-pass safety scan before commit. Blocks pain-points: oversized commits (1.3GB parquet incident), CCS-style anonymity breaches, leaked credentials. Each pass is a hard gate; user OKs proceed.

Hard Rules

Existential — block proceed
  1. Size pass blocks anything >10MB unless gitignored OR user explicitly approves. Threshold matches block-large-files.sh hook for consistency.
  2. Anonymity pass runs only on paper-relevant paths: paper-*/, *.tex, *.bib, *.md inside paper directories. Scope is data-driven — we don't false-flag README authorship.
  3. Secrets pass uses an entropy heuristic + known-prefix list. No regex-only matching that misses high-entropy keys. Block on confirmed secret; warn on suspicious-but-uncertain.
  4. All four passes run on the same file set (default: staged for commit). Don't mix staged-only size check with all-files anonymity check.
  5. The Unicode pass is the one pass that is always in scope, including infra-only commits. It is identity-blind — it reads codepoints, never names — so it cannot false-flag authorship the way the anonymity pass can.
Format — catch in review
  1. Output a single consolidated table (file × pass × verdict) — not four separate reports.
  2. Severity tiers: BLOCK (must fix), WARN (proceed with confirm), OK.
  3. Always show the file path and line number when flagging — make it copy-paste-fixable.

When to Use

  • Before any git commit on a paper / project repo
  • Pre-submission checklist (research projects)
  • Before pushing to anonymous-artifact repo (the highest-risk anonymity surface)
  • After a sub-agent has written files (verify before committing their work)
  • Wired as a session-close Phase 4 sub-step (alternate path to running the hook)

When NOT to Use

  • Before commits on infrastructure-only changes (skills, hooks, rules) — anonymity scan would false-flag README author lines. Exception: Phase 4 (Unicode) stays valid there and can be run alone: pre-commit-audit --unicode-only
  • Auto-generated commits (CI, dependabot) — they don't need the scan
  • When the change is purely metadata (.gitignore, LICENSE, etc.)

Modes

InvocationScope
pre-commit-auditStaged files (default)
pre-commit-audit --stagedSame as default
pre-commit-audit --unstagedModified-but-not-staged files
pre-commit-audit --allWorking tree + staged
pre-commit-audit --unicode-onlyPhase 4 only — safe on infra commits

Architecture

Phase 1 (size)        → find files >10MB; check .gitignore status
Phase 2 (anonymity)   → grep paper paths for author/affiliation/email
Phase 3 (secrets)     → entropy + prefix heuristics on all staged files
Phase 4 (unicode)     → invisible-carrier scan via check_invisible_unicode.py
Phase 5 (consolidate) → single table; user OK to proceed

Phase 1: Size

bash
THRESHOLD_BYTES=$((10 * 1024 * 1024))

git diff --cached --name-only --diff-filter=ACM | while read f; do
    if [ -f "$f" ]; then
        size=$(stat -f%z "$f" 2>/dev/null || stat -c%s "$f" 2>/dev/null || echo 0)
        if [ "$size" -gt "$THRESHOLD_BYTES" ]; then
            ignored=$(git check-ignore "$f" 2>/dev/null && echo "ignored" || echo "tracked")
            echo "$f|$size|$ignored"
        fi
    fi
done

Verdict: BLOCK if any tracked file >10MB. WARN if gitignored but staged anyway.

Phase 2: Anonymity

Single source of truth: skills/_shared/double-blind-anonymity-checklist.md holds the authoritative pattern list (paper-side P1-P8, artifact-side A1-A6) plus the CCS 2026 #1328 incident report that motivates the checks. Both pre-commit-audit (this skill) and anonymous-artifact reference it; do not duplicate the patterns inline.

Scope for pre-commit-audit: only files staged for commit at paths matching paper-*/, **/*.tex, **/*.bib, **/*.cls, **/*.sty. The full checklist also covers structured-metadata files (pyproject.toml, package.json, Cargo.toml, CITATION.cff, etc.) — those are checked by anonymous-artifact since they only matter when pushing the artifact, not on every commit.

Apply checks P1-P8 from the shared checklist:

  • P1 — title page authors blinded
  • P2 — no \thanks{} / \acknowledgements / \funded / \grant revealing identity
  • P3 — first-person self-reference uses third person (no "our prior work")
  • P4 — self-citation bib entries blinded if author overlap
  • P5 — body text doesn't name authors of self-cited works
  • P6 — no identifying URLs (personal sites, lab pages, repos with handles)
  • P7 — PDF metadata sanitized (pdfinfo paper.pdf)
  • P8 — figures contain no identifying watermarks/captions

For each match: print path:line and the matched text (truncated). Verdict: BLOCK on any match in a path containing paper-*/. WARN on matches outside paper directories.

Exception: a frontmatter-style Made by the user in friends-repo / public-repo READMEs is allowed (intentional attribution).

Phase 3: Secrets

Patterns by category:

CategoryDetection
API keys with prefixsk-, sk-proj-, sk-ant-, hf_, gh[a-zA-Z]_, xoxb-, AIza
AWSAKIA[0-9A-Z]{16}, aws_secret_access_key
Generic high-entropy stringslength ≥32, base64-ish character set, no whitespace
.env-style assignments^[A-Z_]+=[a-zA-Z0-9+/=_-]{20,} outside .env.example/.env.template

Skip .env.template, .env.example, *.example, comment lines starting with #.

Verdict: BLOCK on any match in a non-template file.

Show full SKILL.md (498 more words)Show less

Phase 4: Invisible Unicode

Catches characters that are invisible on screen but change what a machine reads — zero-width carriers, bidi overrides, tag characters, and space homoglyphs. These arrive by paste (Word, PDF, web page, LLM output) and fail silently: a U+200B inside a citekey makes BibTeX drop the entry with no error, a U+00A0 makes an exact-match string edit fail on a line that looks identical.

Deterministic — delegate to the script rather than grepping by hand:

bash
git diff --cached --name-only --diff-filter=ACM \
  | while read -r f; do [ -f "$f" ] && printf '%s\n' "$f"; done \
  | xargs -r uv run --no-sync python "<skill-dir>/scripts/check_invisible_unicode.py" --format json

Map the script's severities onto this skill's tiers:

Script severityMeaningVerdict
criticalZero-width / bidi / tag char, or any invisible inside a citekey, crossref, path, url, bib_entry_key, bib_doiBLOCK
warningSpace homoglyph or soft hyphen in prose — renders fine, breaks exact-match editing and grepWARN
rescuedLegitimate emoji ZWJ / VS16 sequencenot reported

Report path:line:col, the U+XXXX label, and the containing construct when the JSON carries a context. The context is most of the signal: the same U+00A0 is a WARN in prose and a BLOCK inside \ref{}.

Fixes are mechanical, so offer them — but never apply silently (Anti-Patterns below):

bash
uv run --no-sync python "<skill-dir>/scripts/check_invisible_unicode.py" <path> --diff    # preview
uv run --no-sync python "<skill-dir>/scripts/check_invisible_unicode.py" <path> --apply   # write

--apply refuses Overleaf-canonical paths unless given --yes-live-surface, per reconcile-before-rewriting.md; if a flagged file is a live Overleaf surface, surface the finding and let the user fix it in the editor rather than passing that flag on their behalf.

Phase 5: Consolidate + ask

Single table:

File                                   Size   Anon   Secrets  Uni    Verdict
────────────────────────────────────────────────────────────────────────────
paper-ccs/paper/main.tex               5KB    OK     OK       OK     OK
paper-ccs/paper/references.bib         12KB   FAIL   OK       FAIL   BLOCK   (line 23: the user; line 8: U+200B in bib_entry_key)
data/raw/results.parquet               1.3GB  N/A    N/A      N/A    BLOCK   (size; not gitignored)
scripts/sync-creds.sh                  2KB    OK     FAIL     OK     BLOCK   (line 15: hf_<...>)
notes/from-word.md                     4KB    OK     OK       WARN   WARN    (4x U+00A0)

If any BLOCK rows: ask the available structured-question mechanism:

  • Fix the issues and rerun — recommended
  • Override and proceed — risky, requires explicit confirmation
  • Cancel commit — abort

If only WARN rows: ask whether to proceed. If all OK: print "Pre-commit clean ✓" and exit silently.

Cross-References

Skill / Hook / RuleRelationship
block-large-files.sh hookRuntime enforcement of Phase 1 (this skill is the manual / verbose counterpart)
scripts/check_invisible_unicode.pyBundled deterministic engine behind Phase 4. Codepoint tables adapted from guillaumemeyer/watermarks-remover (MIT); severity tiers, LaTeX-context detection, and emoji rescue are local
reconcile-before-rewriting.md ruleGoverns Phase 4's --apply on live Overleaf surfaces
anonymous-artifactThe anonymity-on-publish workflow; this skill is the pre-commit equivalent
session-close Phase 4Can invoke this skill as part of pre-commit verification
mark-unverified.md ruleVerbal counterpart for content claims; this skill handles file content

Anti-Patterns

  • Don't scan binary files for anonymity strings — false positives on PDFs, images.
  • Don't rely solely on regex for secret detection — entropy heuristics catch keys with novel prefixes.
  • Don't auto-fix issues — surface them for the user to handle. Auto-replacement of names risks breaking acknowledgements.
  • Don't skip the .gitignore check — large gitignored files staged anyway are a signal of git add -A mistakes.
  • Don't require this skill on infra-only commits. Add a path filter or let the user opt out for those — but --unicode-only remains safe and useful there.
  • Don't run Phase 4's --apply as part of the audit. The audit reports; the user decides. Auto-rewriting bytes in a file the user is mid-edit on is the exact failure reconcile-before-rewriting.md exists to prevent.
  • Don't flag visible non-ASCII. Accented names (López-Ibáñez), em dashes, and → are legitimate and the script deliberately ignores them; a pass that flags them will be muted within a week.

© flonat, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in skills/pre-commit-audit of flonat/flonat-research.

  • SKILL.md
  • scripts/check_invisible_unicode.py

Open the folder on GitHubat commit da27600

Compare with similar skills

Pre Commit Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pre Commit Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pre Commit Audit this skillflonat/flonat-research145—~2.8kAutomated safety check: NotesMIT
Evidence Ledgerwanshuiyin/Anti-Autoresearch160—~11kAutomated safety check: NotesMIT
Research Writingalfonso0512/research-writing-skill4861 repos~818Automated safety check: PassMIT
Paper WritingMLNLP-World/Paper-Writing-Tips4.7k—~630Automated safety check: PassNone
Evomath TaoEvoScientist/EvoSkills4742 repos~3.8kAutomated safety check: PassApache-2.0
Math Modeling to EI Conference Paperjihe520/MathModelAgent6.2k—~688Automated safety check: PassNone

Similar skills

  • Evidence Ledger

    wanshuiyin/Anti-Autoresearch

    Build the deterministic evidence ledger (artifactmanifest.json + claims.json) that every other Anti-Autoresearch auditor reads.

    160 GitHub stars~11k tokensUpdated yesterday
    Documents & OfficeAuto-check: notes
  • Research Writing

    alfonso0512/research-writing-skill

    科研论文写作助手,提供 30 个 Prompt 模板覆盖论文写作全流程. An agent skill from alfonso0512/research-writing-skill.

    486 GitHub starsUsed in 1 repo~818 tokens
    Documents & OfficeAuto-check passed
  • Paper Writing

    MLNLP-World/Paper-Writing-Tips

    学术论文写作检查与优化助手。基于 MLNLP-World 社区整理的论文写作技巧,帮助检查和优化学术论文。Use when: (1) 检查论文 LaTeX 格式和排版, (2) 优化公式符号使用, (3) 改进图表设计, (4) 润色英文学术表达, (5) 检查参考文献格式, (6) 投稿前终稿检查, (7) 用户询问论文写作技巧或规范。

    4.7k GitHub stars~630 tokensUpdated 11 days ago
    Documents & OfficeAuto-check passed
  • Evomath Tao

    EvoScientist/EvoSkills

    A skill your agent uses whenever the user submits a non-trivial mathematical claim that needs a rigorous proof or audit.

    474 GitHub starsUsed in 2 repos~3.8k tokens
    Documents & OfficeAuto-check passed
  • Rewrites a math modeling competition paper into an English EI conference submission, using a bundled IEEE LaTeX template, with evidence tracked back to the original paper.

    6.2k GitHub stars~688 tokensUpdated 4 days ago
    Documents & OfficeAuto-check passed
  • Paperjury

    Spark-To-Paper-Skills/paperjury

    Three modes for CS-conference papers (CVPR/ICCV/ECCV vision, ACL/EMNLP/NAACL NLP, ICLR/NeurIPS/ICML/AAAI ML).

    1.2k GitHub stars~5.3k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed

More from flonat/flonat-research

All 83 skills in this repo
  • Latex Posters

    flonat/flonat-research

    Create a large-format academic poster in LaTeX using beamerposter, tikzposter, or baposter.

    145 GitHub stars~1.5k tokensUpdated 8 days ago
    Auto-check: notes
  • Skill Creator

    flonat/flonat-research

    Create, revise, and evaluate reusable AI workflow skills, including trigger-quality tests.

    145 GitHub stars~4.4k tokensUpdated 8 days ago
    Auto-check passed
  • DOCX

    flonat/flonat-research

    Create, read, edit, or convert Microsoft Word documents while preserving professional document structure.

    145 GitHub stars~1.2k tokensUpdated 8 days ago
    Auto-check passed
  • PDF

    flonat/flonat-research

    Read, create, combine, split, rotate, OCR, watermark, secure, or extract content from PDF files.

    145 GitHub stars~488 tokensUpdated 8 days ago
    Auto-check passed
  • Init Project Orchestration

    flonat/flonat-research

    Create or migrate project-level agents, repeatable project workflows, and planning state from one client-neutral contract, then render repository-scoped adapters for both Claude Code and Codex.

    145 GitHub stars~1.6k tokensUpdated 8 days ago
    Auto-check passed
  • Skill Extract

    flonat/flonat-research

    Extract reusable knowledge from the current session into a persistent skill.

    145 GitHub stars~2.4k tokensUpdated 8 days ago
    Auto-check passed

Questions about Pre Commit Audit

What does Pre Commit Audit do?

Deliver a fast pre-commit safety scan: file size, anonymity (author / affiliation strings in tex/bib), hardcoded secrets, and invisible-Unicode carriers. Pre Commit Audit is an agent skill from flonat/flonat-research. Deliver a fast pre-commit safety scan: file size, anonymity (author / affiliation strings in tex/bib), hardcoded secrets, and invisible-Unicode carriers.

When should I use Pre Commit Audit?

Pre Commit Audit fits situations like: the user requests a fast pre-commit safety scan: file size; anonymity (author / affiliation strings in tex/bib); hardcoded secrets; invisible-Unicode carriers.

How do I install Pre Commit Audit in Claude Code?

Run `npx skills add flonat/flonat-research --skill pre-commit-audit -a claude-code`. Or copy the skill folder (skills/pre-commit-audit in flonat/flonat-research) into .claude/skills/pre-commit-audit in your project. Claude Code loads it when a task matches its description.

How do I install Pre Commit Audit in Codex?

Run `npx skills add flonat/flonat-research --skill pre-commit-audit -a codex`. Or copy the skill folder (skills/pre-commit-audit in flonat/flonat-research) into .agents/skills/pre-commit-audit in your project. Codex loads it when a task matches its description.

Can I use Pre Commit Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add flonat/flonat-research --skill pre-commit-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pre-commit-audit, .gemini/skills/pre-commit-audit, .github/skills/pre-commit-audit and .opencode/skills/pre-commit-audit in your project.

What does Pre Commit Audit need to run?

Going by SKILL.md and its folder, Pre Commit Audit needs Python for the scripts in its folder and the command-line tools its instructions call (git and uv). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Bash(git*), Bash(grep*), Bash(find*), Bash(stat*), Bash(numfmt*), Bash(wc*), Bash(head*), Bash(uv*), AskUserQuestion.

Does Pre Commit Audit access the network?

SKILL.md contains no URLs. Its commands use git and uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Pre Commit Audit safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Pre Commit Audit use?

Pre Commit Audit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pre Commit Audit use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Pre Commit Audit?

Skills that share tags, products or a category with Pre Commit Audit: Evidence Ledger (wanshuiyin/Anti-Autoresearch, 160 stars), Research Writing (alfonso0512/research-writing-skill, 486 stars), Paper Writing (MLNLP-World/Paper-Writing-Tips, 4.7k stars) and Evomath Tao (EvoScientist/EvoSkills, 474 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pre Commit Audit?

flonat (a GitHub user) maintains it in flonat/flonat-research, which has 145 GitHub stars. The repository holds 83 skills in this directory. The repository was last updated on September 29, 2026.

Source: flonat/flonat-research on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.