Agent skill

Anonymizer

by NVIDIA-NeMo in NVIDIA-NeMo/Anonymizer

A skill your agent uses when the user wants to anonymize a text dataset, redact PII, de-identify free-text data, or rewrite text to remove sensitive or inferable identifying information.

Apache-2.0Auto-check passed

Install Anonymizer

skills CLI
$ npx skills add NVIDIA-NeMo/Anonymizer --skill anonymizer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA-NeMo/Anonymizer anonymizer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA-NeMo/Anonymizer.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/anonymizer .claude/skills/anonymizer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
anonymizer
GitHub stars
123
Token cost
~5.1k tokens
SKILL.md length
1,432 words
Files
6 (incl. references)
Skills in repo
1
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when the user wants to anonymize a text dataset, redact PII, de-identify free-text data, or rewrite text to remove sensitive or inferable identifying information.

  • The user wants to anonymize a text dataset
  • Calls pip and python; reaches nvidia-nemo.github.io; needs OPENROUTER_API_KEY
  • De-identify free-text data
  • Rewrite text to remove sensitive

What it does

Anonymizer is an agent skill from NVIDIA-NeMo/Anonymizer. Use when the user wants to anonymize a text dataset, redact PII, de-identify free-text data, or rewrite text to remove sensitive or inferable identifying information. Produces a runnable Python script that calls the NeMo Anonymizer pipeline (detection → replace or rewrite).

Its SKILL.md is about 5.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `BENCHMARK.md`, `evals/evals.json` and `references/interactive.md`).

It works with Python and NVIDIA AI Platform. The repository describes itself as: 🕵️ NeMo Anonymizer: Detect and protect PII through context-aware replacement and rewriting. The licence is Apache-2.0.

When your agent uses it

  • The user wants to anonymize a text dataset
  • De-identify free-text data
  • Rewrite text to remove sensitive
  • Inferable identifying information

Example prompts

  • “/anonymizer”

Requirements

  • Python 3
  • A credential in OPENROUTER_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 630002d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • nvidia-nemo.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENROUTER_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Anonymizer loads about 5.1k tokens when it runs, and up to ~6.6k if it reads all its reference files. Until then it costs about 71 tokens; SKILL.md has 1,432 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~5.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA-NeMo/Anonymizer at commit 630002d, republished under its Apache-2.0 licence (© NVIDIA-NeMo). 1,432 words, ~5,101 tokens.

Download SKILL.mdSave it as .claude/skills/anonymizer/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
anonymizer
description
Use when the user wants to anonymize a text dataset, redact PII, de-identify free-text data, or rewrite text to remove sensitive or inferable identifying information. Produces a runnable Python script that calls the NeMo Anonymizer pipeline (detection → replace or rewrite).
license
Apache-2.0
metadata.author
Aaron Gonzales <aagonzales@nvidia.com>

Before You Start

Do not explore the workspace first. The workflow's data-inspection step shows you what you need.

Goal

Anonymize a text dataset using NeMo Anonymizer in the way the user describes:

$ARGUMENTS

The output is a single runnable Python script that builds an AnonymizerConfig, previews results on a few rows, inspects failures and quality metrics, optionally scores output with LLM-as-judge evaluation (Replace and Rewrite modes), and (on user approval) runs the full pipeline. The script is the durable artifact — the user keeps it for re-runs, version control, and production.

Workflow

Read references/interactive.md and follow it. Anonymization is high-stakes, so there is no autopilot mode. Even when the user says "you decide" or "be opinionated", ask the minimum questions needed to choose risk_tolerance and phrase privacy_goal. The user must make those choices based on their regulatory and business context.

Rules

  • Always preview before running the full pipeline. Preview is cheap; a full run can be expensive and slow.
  • If result.failed_records is non-empty after preview, fix that before tweaking strategy. Dropped rows are a model/provider/infra problem (rate limits, auth, etc.), not a config problem. Strategy knobs won't help. See docs/troubleshooting.md "Did the run actually complete cleanly?" or the published troubleshooting guide.
  • Ask the user which mode. Briefly describe both: Replace detects entities and replaces each in place (faster, cheaper, keeps shape); Rewrite transforms the full text to also remove inferable identifiers (more expensive, may restructure). Use the data shape as a hint — free-text with implicit identifiers (clinical notes, biographies, depositions) leans Rewrite; structured records / log lines lean Replace — but the user picks.
  • For cross-record consistency (same value → same replacement everywhere), use Hash, not Substitute. Substitute is consistent within a row only.
  • In Replace mode, default to Substitute if the user hasn't specified a strategy. It's the most general-purpose choice and matches the bulk of production usage.
  • Annotate is for inspection, not production. Its output keeps the original entity text and is not privacy-safe. Use it during iteration to confirm detection is working, then switch.
  • Evaluation is opt-in and runs as a separate step (Replace and Rewrite modes). After preview() / run(), call anonymizer.evaluate(result) to score the output with LLM-as-judge. Entity coverage always runs in both modes — it reports detection recall over the judge's unique candidate values (entity_coverage + missed_entities). On top of that: Replace Substitute adds three quality judges (type fidelity, relational consistency, attribute fidelity); Rewrite adds the holistic privacy/quality/style judge. Detection validity is opt-in via EvaluateConfig(compute_detection_validity=True) (off by default). Evaluation is diagnostic — it scores quality, it does not change the anonymized output.
  • Always set AnonymizerInput.data_summary, even briefly. It is the single cheapest quality lever and it improves both detection and rewrite.
  • Never claim privacy guarantees. Anonymizer is best-effort. Outputs may need human review depending on risk_tolerance. Tell the user this when you finalize.

Usage Tips and Common Pitfalls

  • Detect.entity_labels=None (the default) is permissive — the augmenter LLM may invent labels not in DEFAULT_ENTITY_LABELS. Setting an explicit list switches to strict mode where only the listed labels are detected.
  • Label terminology: A default label is present in DEFAULT_ENTITY_LABELS, while any label not present in this list is called a non-default label. An explicit label set is supplied through entity_labels and may contain either kind.
  • Detect.entity_label_examples provides configured positive examples for detection. For a default label, configured examples are appended to its built-in examples. Every non-default label referenced by entity_label_examples must also appear in the explicit entity_labels set. To keep every default while adding one, use entity_labels=[*DEFAULT_ENTITY_LABELS, "clinical_facility"]. Configured examples are soft detection guidance: they do not limit detection to the listed value formats, and they are not passed to substitution or evaluation. Keep configured examples concise and lists short to limit prompt growth, token cost, latency, and context-window pressure. Use synthetic values because configured examples are included in prompts, exported builders, provider requests, and explicitly enabled raw message traces.
  • Detect.excluded_entity_labels removes specific label types from the active detection scope and final results. Use it when a label type is systematically noisy for your data or should never be anonymized (e.g. Detect(excluded_entity_labels=["occupation", "gender"])). Exclusions take precedence over labels and examples; configured examples for excluded labels are ignored with a warning. If exclusions empty the default or explicit label set, Detect raises a ValueError.
  • GLiNER is zero-shot — entity labels are natural-language concept names (e.g. "clinical_facility", "internal_project_codename"), not codes or enum values. Any concept you can name in English is a label GLiNER can detect.
  • Rewrite.instructions is a dead field today — it exists on the model but the rewrite engine never reads it. Do not use it. Put rewriter guidance in privacy_goal.protect / privacy_goal.preserve instead.
  • risk_tolerance only applies to Rewrite mode, not Replace.
  • PrivacyGoal.protect and .preserve must each be 10–1000 chars and at least 3 words. Be specific (categories, named identifiers, structural facets); avoid generic phrasing like "preserve meaning".
  • Validator pool is the only model role with built-in load-spreading. Set entity_validator: [a, b, c] in models.yaml if rate limits drop rows. Other roles (rewriter, evaluator, etc.) are single-alias.
  • Self-hosted GLiNER2: When detection must stay local (PHI on-prem, air-gapped, latency), run the reference server from a source checkout with python tools/serve_gliner.py. The server is not installed by pip install nemo-anonymizer. Add a provider with endpoint: http://localhost:8001/v1, then route entity_detector through a gliner-pii-detector alias whose model is fastino/gliner2-privacy-filter-PII-multi, whose provider points to that endpoint (the bundled name is local-gliner2), and whose skip_health_check is true. Match any custom --port or --host in the provider endpoint. model_configs is a complete model pool, not an overlay. Copy src/anonymizer/config/default_model_configs/models.yaml and change only the provider name or endpoint as needed, keeping the LLM entries. See docs/concepts/self-hosting-gliner.md or the published self-hosting guide.
  • The evaluation judges use their own model roles (entity_coverage_judge, detection_validity_judge, replace_type_fidelity_judge, replace_relational_consistency_judge, replace_attribute_fidelity_judge, rewrite_judge), configured in the evaluate section of models.yaml. They are not consumed by preview() / run(), so a config that anonymizes fine can still fail validation at evaluate() if those roles are unset. Defaults ship in src/anonymizer/config/default_model_configs/evaluate.yaml (entity_coverage_judge defaults to nemotron-super).
  • Verdict columns are null when the judge was unavailable — None means "unscored", never a pass. entity_coverage is a 0–1 float (1.0 = no missed candidate values or no PII found) or None; missed_entities lists unique candidate values the anonymizer failed to detect. Replace verdict columns (type_fidelity_valid, etc.) are True / False / None. Rewrite detection_valid is a 0–1 float fraction (or None if unscored). Inspect verdicts per record with evaluated.display_record(i).
  • EvaluateConfig has one knob today: compute_detection_validity (default False). Plain anonymizer.evaluate(result) runs entity coverage + the mode's quality judges; pass EvaluateConfig(compute_detection_validity=True) only to additionally score detection validity (an internal-facing tag-precision metric).
Show full SKILL.md (357 more words)Show less

Reference Docs

The agent should consult these as it goes — do not try to enumerate field reference inline:

Troubleshooting

This section covers environment-level issues. For quality and pipeline issues, read docs/troubleshooting.md or the published troubleshooting guide.

  • anonymizer not installed: Tell the user nemo-anonymizer is not in this Python environment (requires Python ≥ 3.11). Ask if they want you to install it (pip install nemo-anonymizer) or do it themselves. Do not install without permission.
  • Model/provider setup: Plain Anonymizer() ships with bundled models.yaml and providers.yaml (see src/anonymizer/config/default_model_configs/). For the default path, confirm OPENROUTER_API_KEY is set. Pass custom model_configs or model_providers only for non-default endpoints or model pools. See docs/concepts/models.md or the published models guide.
  • LLM calls failing at preview: Check for a missing or invalid API key, a network problem, or a wrong endpoint URL. See docs/troubleshooting.md "Validation passed but preview errors at LLM call" or the published troubleshooting guide.
  • Local / on-prem GLiNER2: Clone or download tools/serve_gliner.py from the Anonymizer repo, start the server, add a provider with endpoint: http://localhost:8001/v1, and point the fastino/gliner2-privacy-filter-PII-multi detector config at that provider (local-gliner2 in the bundled defaults) with skip_health_check: true. Preflight errors about missing aliases usually mean model_configs lists only the detector. Include the full default pool. A wrong endpoint or stopped server raises an actionable configuration error before detection starts. See docs/concepts/self-hosting-gliner.md or the published self-hosting guide.

Output Template

Write a Python script to the current directory. Name it after the dataset (for example, anonymize_clinical_notes.py or anonymize_support_logs.py). Fill in the TODO markers in this template and remove unused sections.

python
"""Anonymize <dataset> using NeMo Anonymizer.

Generated by the anonymizer agent skill.

Usage:
    python <this_script>.py                 # preview on 5 rows (fast, cheap)
    python <this_script>.py --full          # run on the full dataset
    python <this_script>.py --evaluate      # preview 5 rows, then LLM-judge-score those rows
    python <this_script>.py --full --evaluate  # run full dataset, then score the full output
"""

from __future__ import annotations

import argparse
import sys

from anonymizer import (
    Anonymizer,
    AnonymizerConfig,
    AnonymizerInput,
    Detect,
    # Pick what you need:
    # Replace mode:
    Substitute, Redact, Annotate, Hash,
    # Rewrite mode:
    Rewrite, PrivacyGoal,
)


def build_config() -> tuple[AnonymizerInput, AnonymizerConfig]:
    """Single source of truth for what we anonymize and how."""
    data = AnonymizerInput(
        source="TODO: path to .csv / .parquet / .jsonl",
        text_column="TODO: name of the text column",
        data_summary="TODO: one-line description of the data (domain, genre, anything non-obvious)",
    )

    detect = Detect(
        # Every non-default label referenced by entity_label_examples must also be in this explicit label set.
        # entity_labels=["clinical_facility", "diagnosis_code"],
        # entity_label_examples={
        #     "clinical_facility": ["North Valley Oncology Center"],
        #     "diagnosis_code": ["C50.919"],
        # },
        gliner_threshold=0.3,  # default; lower (0.2) for recall, raise (0.5) for cost savings
    )

    # ---- Pick ONE of the two strategies below ----

    # Replace mode (Substitute | Redact | Annotate | Hash):
    # config = AnonymizerConfig(detect=detect, replace=Substitute(
    #     instructions="TODO: short hint about the domain (e.g. names should remain plausible "
    #                  "for the original cultural context)",
    # ))

    # Rewrite mode (free-text de-identification with inferable-identifier suppression):
    config = AnonymizerConfig(
        detect=detect,
        rewrite=Rewrite(
            privacy_goal=PrivacyGoal(
                protect="TODO: what must not appear in the output, even by inference",
                preserve="TODO: what must be kept so the rewritten text is still useful",
            ),
            risk_tolerance="low",          # minimal | low | moderate | high
            strict_entity_protection=False, # True = force every detected entity into a protective disposition
            max_repair_iterations=3,
        ),
    )
    return data, config


def main() -> None:
    parser = argparse.ArgumentParser(description=__doc__)
    parser.add_argument("--full", action="store_true", help="Run on full dataset (default: preview 5 rows)")
    parser.add_argument("--num-records", type=int, default=5, help="Rows to preview (ignored with --full)")
    parser.add_argument(
        "--evaluate",
        action="store_true",
        help="LLM-judge-score the output produced this run (preview rows, or full output with --full)",
    )
    args = parser.parse_args()

    anonymizer = Anonymizer()
    data, config = build_config()

    if args.full:
        result = anonymizer.run(config=config, data=data)
        out_path = "output.parquet"  # TODO: change path/format (.csv, .jsonl) as needed
        result.dataframe.to_parquet(out_path)
        print(f"Wrote {len(result.dataframe)} rows to {out_path}")
    else:
        result = anonymizer.preview(config=config, data=data, num_records=args.num_records)
        print(f"Previewed {len(result.dataframe)} rows.")

        # Save preview output so you can investigate without re-running.
        # trace_dataframe is a superset of dataframe — it has the user-facing
        # columns plus internal columns (validation decisions, sensitivity
        # dispositions, etc.) that explain why entities were kept, dropped,
        # or rewritten.
        result.trace_dataframe.to_parquet("preview.parquet")
        print("Saved: preview.parquet (load with pd.read_parquet)")

    # Failure-first protocol: dropped rows are infra issues, not strategy issues.
    if result.failed_records:
        print(f"\n⚠️  {len(result.failed_records)} record(s) failed:")
        for fr in result.failed_records[:3]:
            print(f"   - record_id={fr.record_id} step={fr.step} reason={fr.reason}")
        print("\nFix dropped rows before tweaking strategy. See docs/troubleshooting.md or https://nvidia-nemo.github.io/Anonymizer/dev/troubleshooting/.")
        sys.exit(1)

    # Optional LLM-as-judge evaluation (Replace and Rewrite modes). Opt-in, separate
    # step — scores quality without changing the anonymized output.
    # Both modes: entity_coverage (judge-anchored recall) always runs.
    # Replace: Substitute adds type fidelity, relational consistency, attribute fidelity.
    # Rewrite: adds the holistic privacy/quality/style judge.
    # Detection validity is opt-in (EvaluateConfig(compute_detection_validity=True)).
    # Needs the `evaluate` model roles in models.yaml
    # (see src/anonymizer/config/default_model_configs/evaluate.yaml).
    if args.evaluate:
        result = anonymizer.evaluate(result)
        df = result.dataframe
        # entity_coverage is a per-record 0–1 float (1.0 = no missed candidate values or no PII found by judge); aggregate mean shown below.
        if "entity_coverage" in df.columns:
            scored = int(df["entity_coverage"].notna().sum())
            mean_cov = df["entity_coverage"].mean()
            print(f"entity_coverage: mean={mean_cov:.2f}  scored={scored}/{len(df)}")
        if config.replace is not None:
            for col in (
                "type_fidelity_valid",
                "relational_consistency_valid",
                "attribute_fidelity_valid",
                "detection_valid",  # present only with compute_detection_validity=True
            ):
                if col in df.columns:
                    passed = int(df[col].eq(True).sum())  # None = unscored, never a pass
                    scored = int(df[col].notna().sum())
                    print(f"{col}: {passed}/{scored} passed ({len(df) - scored} unscored)")
        else:
            # Rewrite: detection_valid is a 0–1 fraction (present only when opted in).
            if "detection_valid" in df.columns:
                scored = int(df["detection_valid"].notna().sum())
                mean_val = df["detection_valid"].mean()
                print(f"detection_valid: mean={mean_val:.2f}  scored={scored}/{len(df)}")
            if "judge_evaluation" in df.columns:
                scored = int(df["judge_evaluation"].notna().sum())
                print(f"judge_evaluation: {scored}/{len(df)} scored")
        # In a notebook, inspect per-record verdicts visually:
        #   result.display_record(0)

    # Rewrite-mode quality summary (skip for Replace mode).
    if config.rewrite is not None:
        df = result.dataframe
        print(f"\nleakage_mass:   mean={df['leakage_mass'].mean():.3f}  max={df['leakage_mass'].max():.3f}")
        print(f"utility_score:  mean={df['utility_score'].mean():.3f}  min={df['utility_score'].min():.3f}")
        print(f"flagged for review: {int(df['needs_human_review'].sum())} / {len(df)}")


if __name__ == "__main__":
    main()

© NVIDIA-NeMo, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/anonymizer of NVIDIA-NeMo/Anonymizer.

  • SKILL.md
  • BENCHMARK.md
  • evals/evals.json
  • references/interactive.md
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 630002d

Compare with similar skills

Anonymizer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Anonymizer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Anonymizer this skillNVIDIA-NeMo/Anonymizer123—~5.1kAutomated safety check: PassApache-2.0
Refactor OpCVCUDA/CV-CUDA2.7k—~1.5kAutomated safety check: PassCustom licence
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Gds DiagNVIDIA/MagnumIO125—~1.6kAutomated safety check: PassApache-2.0
Nsight Graphics AnalyzerLuna5ama/Alpha-Piscium156—~4.7kAutomated safety check: PassGPL-3.0
Optimize OpCVCUDA/CV-CUDA2.7k—~834Automated safety check: PassCustom licence

Similar skills

  • Refactor Op

    CVCUDA/CV-CUDA

    Find and safely apply per-operator refactoring / redundancy-reduction opportunities in a CV-CUDA operator (near-duplicate Tensor/VarShape kernels, reinvented shared utilities, dead code).

    2.7k GitHub stars~1.5k tokensUpdated 22 days ago
    DevelopmentAuto-check passed
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Gds Diag

    NVIDIA/MagnumIO

    Official

    A skill your agent uses when diagnosing NVIDIA GPUDirect Storage with this repository: choose and run the right gds-diag.py subcommand, interpret its output, and explain operator next steps without…

    125 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Nsight Graphics Analyzer

    Luna5ama/Alpha-Piscium

    Drive NVIDIA Nsight Graphics 2026.1+ from the command line for GPU performance analysis, frame capture, frame trace inspection, draw-call inspection, NVTX/D3DPERF stage timing, replay metadata…

    156 GitHub stars~4.7k tokensUpdated 3 days ago
    Documents & OfficeAuto-check passed
  • Optimize Op

    CVCUDA/CV-CUDA

    Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.

    2.7k GitHub stars~834 tokensUpdated 22 days ago
    AI & LLM EngineeringAuto-check passed
  • Deep Researcher Research

    NVIDIA-AI-Blueprints/deep-researcher-agent

    A skill your agent uses when asked to run deep research or Deep Researcher Agent research through a reachable NVIDIA Deep Researcher Agent Blueprint backend.

    883 GitHub stars~4.4k tokensUpdated 2 days ago
    Research & ScienceAuto-check: notes

Questions about Anonymizer

What does Anonymizer do?

A skill your agent uses when the user wants to anonymize a text dataset, redact PII, de-identify free-text data, or rewrite text to remove sensitive or inferable identifying information. Anonymizer is an agent skill from NVIDIA-NeMo/Anonymizer. Use when the user wants to anonymize a text dataset, redact PII, de-identify free-text data, or rewrite text to remove sensitive or inferable identifying information.

When should I use Anonymizer?

Anonymizer fits situations like: the user wants to anonymize a text dataset; de-identify free-text data; rewrite text to remove sensitive; inferable identifying information.

How do I install Anonymizer in Claude Code?

Run `npx skills add NVIDIA-NeMo/Anonymizer --skill anonymizer -a claude-code`. Or copy the skill folder (skills/anonymizer in NVIDIA-NeMo/Anonymizer) into .claude/skills/anonymizer in your project. Claude Code loads it when a task matches its description.

How do I install Anonymizer in Codex?

Run `npx skills add NVIDIA-NeMo/Anonymizer --skill anonymizer -a codex`. Or copy the skill folder (skills/anonymizer in NVIDIA-NeMo/Anonymizer) into .agents/skills/anonymizer in your project. Codex loads it when a task matches its description.

Can I use Anonymizer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA-NeMo/Anonymizer --skill anonymizer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/anonymizer, .gemini/skills/anonymizer, .github/skills/anonymizer and .opencode/skills/anonymizer in your project.

What does Anonymizer need to run?

Going by SKILL.md and its folder, Anonymizer needs the command-line tools its instructions call (pip and python) and credentials named OPENROUTER_API_KEY. Our summary lists: Python 3; A credential in OPENROUTER_API_KEY.

Does Anonymizer access the network?

SKILL.md names 1 domain. In commands or code: nvidia-nemo.github.io; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Anonymizer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Anonymizer use?

Anonymizer is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Anonymizer use?

About 5.1k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.5k tokens, read only when the agent opens those files.

What are the alternatives to Anonymizer?

Skills that share tags, products or a category with Anonymizer: Refactor Op (CVCUDA/CV-CUDA, 2.7k stars), Dstack Prototyping (dstackai/dstack, 2.3k stars), Gds Diag (NVIDIA/MagnumIO, 125 stars) and Nsight Graphics Analyzer (Luna5ama/Alpha-Piscium, 156 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Anonymizer?

NVIDIA-NeMo (a GitHub organization) maintains it in NVIDIA-NeMo/Anonymizer, which has 123 GitHub stars. The repository was last updated on October 8, 2026.

Source: NVIDIA-NeMo/Anonymizer on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.