PII filtering: strip emails, phone numbers, SSNs from document body before indexing.

MITAuto-check passedLegal & Compliance

Install Glean Data Handling

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill glean-data-handling -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace glean-data-handling --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/glean-data-handling .claude/skills/glean-data-handling && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
glean-data-handling
GitHub stars
2.8k
Token cost
~1.8k tokens
SKILL.md length
534 words
Files
1
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

PII filtering: strip emails, phone numbers, SSNs from document body before indexing.

  • Works in 5 steps: Inventory fields and source ACLs before… → Run redaction and schema validation on a… → Compare source and proposed index ACLs… → …
  • Legal & Compliance work in your project
  • SKILL.md covers Overview, Data Classification, Data Import and Data Export, plus 9 more sections
  • Reaches customer-be.glean.com; needs GLEAN_INDEXING_TOKEN and GLEAN_API_TOKEN

What it does

Glean Data Handling is an agent skill from jeremylongshore/tons-of-skills-marketplace. PII filtering: strip emails, phone numbers, SSNs from document body before indexing. Trigger: "glean data handling", "data-handling".

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Designed for Claude Code

It sits in Legal & Compliance. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Legal & Compliance work in your project

Example prompts

  • “glean data handling”
  • “data-handling”
  • “/glean-data-handling”

Requirements

  • A credential in GLEAN_INDEXING_TOKEN
  • A credential in GLEAN_API_TOKEN
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash(npm:*), Bash(curl:*), Grep

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Inventory fields and source ACLs before indexing; default any unknown permission or classification to deny and quarantine the record.
  2. Run redaction and schema validation on a bounded staging batch, logging only record counts, rule versions, and opaque document identifiers.
  3. Compare source and proposed index ACLs for both an allowed and denied test principal before approving a production batch.
  4. Send idempotent, bounded batches to the approved datasource and retain a reversible manifest of submitted opaque IDs and timestamps.
  5. On a failed validation, stop the batch, delete only the staged artifacts, and escalate to the data owner rather than weakening redaction…

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash(npm:*)
    • Bash(curl:*)
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are typescript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • customer-be.glean.com

    Also links to:

    • developers.glean.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GLEAN_INDEXING_TOKEN
    • GLEAN_API_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Glean Data Handling loads about 1.8k tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 534 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~38
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 534 words, ~1,775 tokens.

Download SKILL.mdSave it as .claude/skills/glean-data-handling/SKILL.md (or your agent's skills folder).
name
glean-data-handling
description
PII filtering: strip emails, phone numbers, SSNs from document body before indexing. Trigger: "glean data handling", "data-handling".
allowed-tools
Read, Write, Edit, Bash(npm:*), Bash(curl:*), Grep
compatibility
Designed for Claude Code
version
1.8.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, enterprise-search, glean

Glean Data Handling

Overview

Glean enterprise search ingests documents from dozens of connectors (Google Drive, Confluence, Slack, Jira, Salesforce, etc.) and builds a unified search index with permission-aware access control. Data types include indexed document content, connector metadata, user permission maps, query logs, and search analytics. All document content must be PII-filtered before indexing, permission boundaries must be preserved to prevent data leakage across teams, and retention policies must be enforced to comply with corporate governance and GDPR/CCPA obligations.

Data Classification

Data TypeSensitivityRetentionEncryption
Indexed document contentHigh (may contain PII)Per source retention policyAES-256 at rest
User permission mapsHigh (access control)Sync lifecycleTLS + at rest
Connector metadataMediumUntil connector removedAES-256 at rest
Search query logsMedium (reveals intent)90 days defaultAES-256 at rest
Search analytics/aggregatesLow1 yearTLS in transit

Data Import

typescript
interface GleanDocument {
  id: string; datasource: string; title: string;
  body: string; permissions: { allowedUsers?: string[]; allowAnonymousAccess?: boolean };
  updatedAt: string; url: string;
}

async function indexDocuments(docs: GleanDocument[], datasource: string) {
  // PII strip before indexing
  const sanitized = docs.map(doc => ({
    ...doc,
    body: stripPII(doc.body),
  }));
  // Batch upload with pagination (max 100 per request)
  for (let i = 0; i < sanitized.length; i += 100) {
    const batch = sanitized.slice(i, i + 100);
    await fetch(`https://customer-be.glean.com/api/index/v1/bulkindexdocuments`, {
      method: 'POST',
      headers: { Authorization: `Bearer ${process.env.GLEAN_INDEXING_TOKEN}`, 'Content-Type': 'application/json' },
      body: JSON.stringify({ datasource, documents: batch }),
    });
  }
}

function stripPII(text: string): string {
  return text
    .replace(/\b[\w.+-]+@[\w-]+\.[\w.]+\b/g, '[EMAIL_REDACTED]')
    .replace(/\b\d{3}[-.]?\d{3}[-.]?\d{4}\b/g, '[PHONE_REDACTED]')
    .replace(/\b\d{3}-\d{2}-\d{4}\b/g, '[SSN_REDACTED]');
}

Data Export

typescript
async function exportSearchAnalytics(startDate: string, endDate: string) {
  const res = await fetch(`https://customer-be.glean.com/api/v1/analytics`, {
    method: 'POST',
    headers: { Authorization: `Bearer ${process.env.GLEAN_API_TOKEN}`, 'Content-Type': 'application/json' },
    body: JSON.stringify({ startDate, endDate, metrics: ['query_count', 'click_through', 'zero_results'] }),
  });
  const data = await res.json();
  // Redact user identifiers from analytics export
  return data.results.map((r: any) => ({ ...r, userId: undefined, query: r.query?.length > 3 ? r.query : '[SHORT_QUERY_REDACTED]' }));
}

Data Validation

typescript
function validateDocument(doc: GleanDocument): string[] {
  const errors: string[] = [];
  if (!doc.id || doc.id.length > 512) errors.push('Invalid document ID');
  if (!doc.datasource) errors.push('Missing datasource identifier');
  if (!doc.title || doc.title.length > 1000) errors.push('Title missing or exceeds 1000 chars');
  if (!doc.body || doc.body.length === 0) errors.push('Empty document body');
  if (!doc.permissions) errors.push('Missing permissions — defaults to deny-all');
  if (doc.updatedAt && isNaN(Date.parse(doc.updatedAt))) errors.push('Invalid updatedAt timestamp');
  return errors;
}

Compliance

  • PII stripped from document body before indexing (emails, phones, SSNs)
  • Permission boundaries enforced: allowedUsers scope matches source system ACLs
  • Connector credentials stored in secret manager, rotated quarterly
  • Search query logs retained max 90 days, purged via automated job
  • GDPR right-to-erasure: delete all indexed content referencing a specific user on request
  • CCPA: honor do-not-sell signals for search analytics data
  • SOC 2 Type II audit trail for all indexing and deletion operations

Error Handling

IssueCauseFix
403 on bulk indexExpired or insufficient indexing tokenRotate token, verify datasource permissions
Permission mismatch in searchStale ACL sync from connectorForce re-sync connector permissions via admin API
PII detected in indexed contentNew PII pattern not in strip regexAdd pattern to stripPII, re-index affected datasource
Zero-result queries spikeConnector sync failure, stale indexCheck connector health dashboard, trigger manual re-crawl
Rate limit 429 on indexingBatch size too large or too frequentReduce batch to 50 docs, add 500ms delay between batches
Show full SKILL.md (226 more words)Show less

Prerequisites

  • A documented data owner and approved classification/retention policy for the source being indexed.
  • A non-production sample containing fictitious identities, plus an allowlisted destination and credentials scoped only to that datasource.
  • A reviewed redaction and deletion plan; pattern matching is a safeguard, not proof that sensitive data is absent.

Instructions

  1. Inventory fields and source ACLs before indexing; default any unknown permission or classification to deny and quarantine the record.
  2. Run redaction and schema validation on a bounded staging batch, logging only record counts, rule versions, and opaque document identifiers.
  3. Compare source and proposed index ACLs for both an allowed and denied test principal before approving a production batch.
  4. Send idempotent, bounded batches to the approved datasource and retain a reversible manifest of submitted opaque IDs and timestamps.
  5. On a failed validation, stop the batch, delete only the staged artifacts, and escalate to the data owner rather than weakening redaction or ACL rules.

Output

Return a handling receipt with source classification, approved destination, policy and redaction-rule versions, input/accepted/quarantined counts, ACL comparison result, retention date, and deletion or rollback reference. Do not emit bodies, queries, emails, phone numbers, or access tokens.

Examples

Example receipt: source=staging-wiki; classification=confidential; accepted=96; quarantined=4; acl_probe=allow+deny pass; retention=2026-10-01; redaction_rules=v3. The sample is fictional and identifies no employee or document.

Resources

Next Steps

See glean-security-basics.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/.curated/glean-data-handling of jeremylongshore/tons-of-skills-marketplace.

Open the folder on GitHubat commit cfae287

Compare with similar skills

Glean Data Handling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Glean Data Handling compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Glean Data Handling this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.8kAutomated safety check: PassMIT
Paper to Chinese Patent DrafterYuan1z0825/nature-skills47k1 repos~1.1kAutomated safety check: PassApache-2.0
C15tc15t/c15t1.9k1 repos~1.6kAutomated safety check: PassApache-2.0
Contract Reviewevolsb/claude-legal-skill4641 repos~3.6kAutomated safety check: PassMIT
Legal Clinic Client Intakeanthropics/claude-for-legal9.6k3 repos~3.2kAutomated safety check: PassApache-2.0
Paper To Cn Patentsnipp-zha/Paper-to-patent-Skill1071 repos~959Automated safety check: PassNone

Similar skills

  • Paper to Chinese Patent Drafter

    Yuan1z0825/nature-skills

    Drafts Chinese invention patent applications and technical disclosures from research papers or inventor materials, tying each claim feature to source evidence.

    47k GitHub starsUsed in 1 repo~1.1k tokens
    Legal & ComplianceAuto-check passed
  • C15t

    c15t/c15t

    Work with c15t consent management docs, APIs, and integrations for Next.js, React, and JavaScript.

    1.9k GitHub starsUsed in 1 repo~1.6k tokens
    Legal & ComplianceAuto-check passed
  • Contract Review

    evolsb/claude-legal-skill

    Review legal contracts, NDAs, employment agreements, SaaS terms, and M&A documents.

    464 GitHub starsUsed in 1 repo~3.6k tokens
    Legal & ComplianceAuto-check passed
  • Legal Clinic Client Intake

    anthropics/claude-for-legal

    Official

    Structures a legal clinic client intake interview and produces a case summary with cross-area issue spotting, conflict flags and triage classification.

    9.6k GitHub starsUsed in 3 repos~3.2k tokens
    Legal & ComplianceAuto-check passed
  • Paper To Cn Patent

    snipp-zha/Paper-to-patent-Skill

    Convert scientific papers, theses, technical reports, source code, figures, or research manuscripts into evidence-grounded Chinese invention patent drafts.

    107 GitHub starsUsed in 1 repo~959 tokens
    Legal & ComplianceAuto-check passed
  • Employment Contract Templates

    ynulihao/AgentSkillOS

    Create employment contracts, offer letters, and HR policy documents following legal best practices.

    618 GitHub starsUsed in 12 repos~4.1k tokens
    Legal & ComplianceAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Questions about Glean Data Handling

What does Glean Data Handling do?

PII filtering: strip emails, phone numbers, SSNs from document body before indexing. Glean Data Handling is an agent skill from jeremylongshore/tons-of-skills-marketplace. PII filtering: strip emails, phone numbers, SSNs from document body before indexing.

When should I use Glean Data Handling?

Glean Data Handling fits situations like: legal & Compliance work in your project.

How do I install Glean Data Handling in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill glean-data-handling -a claude-code`. Or copy the skill folder (skills/.curated/glean-data-handling in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/glean-data-handling in your project. Claude Code loads it when a task matches its description.

How do I install Glean Data Handling in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill glean-data-handling -a codex`. Or copy the skill folder (skills/.curated/glean-data-handling in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/glean-data-handling in your project. Codex loads it when a task matches its description.

Can I use Glean Data Handling in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill glean-data-handling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/glean-data-handling, .gemini/skills/glean-data-handling, .github/skills/glean-data-handling and .opencode/skills/glean-data-handling in your project.

What does Glean Data Handling need to run?

Going by SKILL.md and its folder, Glean Data Handling needs credentials named GLEAN_INDEXING_TOKEN and GLEAN_API_TOKEN. Our summary lists: A credential in GLEAN_INDEXING_TOKEN; A credential in GLEAN_API_TOKEN. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(npm:*), Bash(curl:*), Grep. Compatibility (from SKILL.md): Designed for Claude Code.

Does Glean Data Handling access the network?

SKILL.md names 2 domains. In commands or code: customer-be.glean.com; the agent is likely to contact it when it follows the instructions. As links in the text: developers.glean.com. This is read from the text; nothing was executed.

Is Glean Data Handling safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Glean Data Handling use?

Glean Data Handling is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Glean Data Handling use?

About 1.8k tokens (SKILL.md is roughly 7.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Glean Data Handling?

Skills that share tags, products or a category with Glean Data Handling: Paper to Chinese Patent Drafter (Yuan1z0825/nature-skills, 47k stars), C15t (c15t/c15t, 1.9k stars), Contract Review (evolsb/claude-legal-skill, 464 stars) and Legal Clinic Client Intake (anthropics/claude-for-legal, 9.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Glean Data Handling?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.