Agent skill

Glean Migration Deep Dive

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Migrate from Elasticsearch/Algolia: 1) Export all documents from source, 2) Transform to Glean document schema (id, title, url, body, permissions), 3) Create datasource with adddatasource, 4) Bulk…

MITAuto-check passedBackend & APIs

Install Glean Migration Deep Dive

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill glean-migration-deep-dive -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace glean-migration-deep-dive --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/glean-migration-deep-dive .claude/skills/glean-migration-deep-dive && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
glean-migration-deep-dive
GitHub stars
2.8k
Token cost
~2k tokens
SKILL.md length
476 words
Files
1
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Migrate from Elasticsearch/Algolia: 1) Export all documents from source, 2) Transform to Glean document schema (id, title, url, body, permissions), 3) Create datasource with adddatasource, 4) Bulk…

  • Works in 3 steps: Prepare → Migrate → Validate
  • Tasks that involve Search implementation
  • SKILL.md covers Overview, Migration Assessment, Step-by-Step Migration and Rollback Plan, plus 8 more sections
  • Calls curl; needs GLEAN_API_TOKEN

What it does

Glean Migration Deep Dive is an agent skill from jeremylongshore/tons-of-skills-marketplace. Migrate from Elasticsearch/Algolia: 1) Export all documents from source, 2) Transform to Glean document schema (id, title, url, body, permissions), 3) Create datasource with adddatasource, 4) Bulk index with bulkindexdocuments, 5) Validate search quality with test queries, 6) Switch search UI to use Glean Client API. Trigger: "glean migration deep dive", "migration-deep-dive".

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Designed for Claude Code

It sits in Backend & APIs, covering Search implementation. It works with Elasticsearch and Algolia. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Tasks that involve Search implementation

Example prompts

  • “glean migration deep dive”
  • “migration-deep-dive”
  • “/glean-migration-deep-dive”

Requirements

  • A credential in GLEAN_API_TOKEN
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash(npm:*), Bash(curl:*), Grep

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Prepare
  2. Migrate
  3. Validate

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash(npm:*)
    • Bash(curl:*)
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • developers.glean.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GLEAN_API_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Glean Migration Deep Dive loads about 2k tokens when it runs. Until then it costs about 101 tokens; SKILL.md has 476 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 476 words, ~2,000 tokens.

Download SKILL.mdSave it as .claude/skills/glean-migration-deep-dive/SKILL.md (or your agent's skills folder).
name
glean-migration-deep-dive
description
Migrate from Elasticsearch/Algolia: 1) Export all documents from source, 2) Transform to Glean document schema (id, title, url, body, permissions), 3) Create datasource with adddatasource, 4) Bulk index with bulkindexdocuments, 5) Validate search quality with test queries, 6) Switch search UI to use Glean Client API. Trigger: "glean migration deep dive", "migration-deep-dive".
allowed-tools
Read, Write, Edit, Bash(npm:*), Bash(curl:*), Grep
compatibility
Designed for Claude Code
version
1.8.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, enterprise-search, glean

Glean Migration Deep Dive

Overview

Comprehensive guide for migrating enterprise search from Elasticsearch or Algolia to Glean. Covers connector migration (replacing custom crawlers with Glean's push indexing API), permission model changes (mapping ACLs to Glean's datasource-level permissions), and full index rebuilds using bulkindexdocuments. Typical timeline is 2-4 weeks for a mid-size deployment with 100K-1M documents across multiple datasources.

Migration Assessment

typescript
// Scan current integration for deprecated patterns and index health
const assessment = {
  source: process.env.SEARCH_PROVIDER ?? 'elasticsearch',
  indices: await sourceClient.cat.indices({ format: 'json' }),
  totalDocs: 0, connectors: [] as string[], permissionModel: '',
};
for (const idx of assessment.indices) {
  assessment.totalDocs += parseInt(idx['docs.count'] ?? '0', 10);
  assessment.connectors.push(idx.index);
}
assessment.permissionModel = assessment.source === 'elasticsearch' ? 'index-level' : 'api-key';
console.log(`Source: ${assessment.source}`);
console.log(`Indices: ${assessment.connectors.length} | Total docs: ${assessment.totalDocs}`);
console.log(`Permission model: ${assessment.permissionModel} → Glean datasource ACLs`);

Step-by-Step Migration

Phase 1: Prepare

Export all documents from the current search provider and map them to Glean's document schema. Each document needs id, title, url, body, and permissions.

typescript
interface GleanDocument {
  id: string;
  datasource: string;
  title: string;
  url: string;
  body: { mimeType: 'text/plain' | 'text/html'; content: string };
  permissions: { allowedUsers?: string[]; allowedGroups?: string[] };
  updatedAt: string;
}

async function exportAndTransform(sourceIndex: string): Promise<GleanDocument[]> {
  const docs: GleanDocument[] = [];
  let scrollId: string | undefined;
  do {
    const batch = scrollId
      ? await sourceClient.scroll({ scroll: '2m', scroll_id: scrollId })
      : await sourceClient.search({ index: sourceIndex, scroll: '2m', size: 500 });
    scrollId = batch._scroll_id;
    for (const hit of batch.hits.hits) {
      docs.push({
        id: hit._id, datasource: 'custom_' + sourceIndex,
        title: hit._source.title, url: hit._source.url,
        body: { mimeType: 'text/plain', content: hit._source.body },
        permissions: { allowedGroups: hit._source.acl_groups ?? ['everyone'] },
        updatedAt: hit._source.updated_at ?? new Date().toISOString(),
      });
    }
  } while (scrollId && docs.length < 1_000_000);
  return docs;
}
Phase 2: Migrate

Create the Glean datasource and bulk-index all transformed documents.

typescript
// Create datasource
await gleanClient.post('/api/index/v1/adddatasource', {
  name: 'custom_knowledge_base',
  displayName: 'Knowledge Base',
  homeUrl: 'https://kb.example.com',
  objectDefinitions: [{ name: 'Article', propertyDefinitions: [] }],
});

// Bulk index in batches of 200
const BATCH_SIZE = 200;
for (let i = 0; i < docs.length; i += BATCH_SIZE) {
  const batch = docs.slice(i, i + BATCH_SIZE);
  await gleanClient.post('/api/index/v1/bulkindexdocuments', {
    datasource: 'custom_knowledge_base',
    documents: batch.map(d => ({
      id: d.id, title: d.title, url: d.url, body: d.body,
      permissions: d.permissions, updatedAt: d.updatedAt,
    })),
  });
  console.log(`Indexed ${Math.min(i + BATCH_SIZE, docs.length)}/${docs.length}`);
}
Phase 3: Validate

Run test queries against both old and new systems and compare result quality.

typescript
const testQueries = ['onboarding process', 'security policy', 'API authentication'];
for (const query of testQueries) {
  const result = await gleanClient.post('/api/client/v1/search', {
    query, pageSize: 5, requestOptions: { datasourceFilter: 'custom_knowledge_base' },
  });
  console.log(`"${query}" → ${result.results.length} results, top: ${result.results[0]?.title}`);
}

Rollback Plan

bash
# Delete the Glean datasource (removes all indexed documents)
curl -X DELETE "https://your-instance.glean.com/api/index/v1/deletedatasource" \
  -H "Authorization: Bearer $GLEAN_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"name": "custom_knowledge_base"}'

# Re-enable old search provider
export SEARCH_PROVIDER=elasticsearch
echo "Rollback complete — old search provider re-enabled"

Migration Checklist

  • Inventory all source indices and document counts
  • Map permission model (index-level ACLs to Glean datasource permissions)
  • Export and transform documents to Glean schema
  • Create datasource via adddatasource API
  • Bulk index all documents in batches of 200
  • Run validation queries and compare relevance
  • Update search UI to use Glean Client API
  • Decommission old search crawlers and indices
  • Monitor search quality metrics for 7 days post-migration
  • Remove old provider SDK dependencies from codebase

Error Handling

IssueCauseFix
401 UnauthorizedInvalid Glean API tokenRegenerate token in admin console
409 Datasource existsName collisionDelete or rename existing datasource
413 Payload Too LargeBatch exceeds 10 MBReduce batch size from 200
422 Invalid documentMissing required fieldValidate id, title, url before indexing
Stale search resultsIndex not refreshedWait 5 min or trigger manual refresh

Prerequisites

  • A signed migration plan with data owners, source ACL authority, retention requirements, and a cutover/rollback decision owner.
  • An inventory of opaque source IDs, counts, and checksums rather than exported document bodies or user lists.
  • A sandbox rehearsal and a proven method to freeze writes, disable the new connector, and restore the prior search path.
Show full SKILL.md (165 more words)Show less

Instructions

  1. Baseline source counts, ACL mappings, freshness, and synthetic allow/deny probes before any migration.
  2. Rehearse a bounded, idempotent migration in sandbox; quarantine schema or permission mismatches rather than coercing them.
  3. Migrate one source cohort at a time, compare aggregate counts/checksums and access outcomes, and log opaque correlation IDs.
  4. Cut over only after owner approval and a monitoring window; keep the prior path available until recovery criteria are met.
  5. Roll back on an ACL, integrity, or freshness failure, then document the exact boundary that failed before retrying.

Output

Create a migration receipt with source cohort, baseline/target counts, ACL and probe results, checkpoint, owner approval, cutover status, retention action, and rollback reference. Never attach source content or credentials.

Examples

cohort=sandbox-hr-01; baseline=420; indexed=420; checksum=match; allow=pass; deny=pass; cutover=held; rollback=old-search-r8 documents a safe rehearsal.

Resources

Next Steps

After migration, configure Glean's built-in connectors for Google Drive, Confluence, and Slack to replace custom crawlers. See glean-core-workflow-a for search integration.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/.curated/glean-migration-deep-dive of jeremylongshore/tons-of-skills-marketplace.

Open the folder on GitHubat commit cfae287

Compare with similar skills

Glean Migration Deep Dive next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Glean Migration Deep Dive compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Glean Migration Deep Dive this skilljeremylongshore/tons-of-skills-marketplace2.8k—~2kAutomated safety check: PassMIT
Product Full-Text Searchlobehub/lobehub83k—~4.1kAutomated safety check: PassCustom licence
Hackernewssigcli/sigcli293—~2kAutomated safety check: PassMIT
Foundatio Repositoriesexceptionless/Exceptionless2.5k—~1.9kAutomated safety check: PassApache-2.0
Elasticsearch Authnaspectrr/deer405—~1.2kAutomated safety check: NotesMIT
Elasticsearch Authzaspectrr/deer405—~1.8kAutomated safety check: PassMIT

Similar skills

  • Guides work on LobeHub's own product search: the shared search repository, provider choice, Elasticsearch mappings, change syncing and reindexing.

    83k GitHub stars~4.1k tokensUpdated today
    Backend & APIsAuto-check passed
  • Hackernews

    sigcli/sigcli

    Interact with Hacker News (news.ycombinator.com) — browse top, new, and best stories, read item details and comment threads, look up user profiles, and search posts via Algolia.

    293 GitHub stars~2k tokensUpdated 13 days ago
    Backend & APIsAuto-check passed
  • Foundatio Repositories

    exceptionless/Exceptionless

    Query, aggregate, patch, or paginate Exceptionless data through its Elasticsearch repository abstractions.

    2.5k GitHub stars~1.9k tokensUpdated 3 days ago
    Backend & APIsAuto-check passed
  • Elasticsearch Authn

    aspectrr/deer

    Authenticate to Elasticsearch using native, file-based, LDAP/AD, SAML, OIDC, Kerberos, JWT, or certificate realms.

    405 GitHub stars~1.2k tokensUpdated 5 mo ago
    Backend & APIsAuto-check: notes
  • Elasticsearch Authz

    aspectrr/deer

    Manage Elasticsearch RBAC: native users, roles, role mappings, document- and field-level security.

    405 GitHub stars~1.8k tokensUpdated 5 mo ago
    Backend & APIsAuto-check passed
  • Ingest and transform data files (CSV/JSON/Parquet/Arrow IPC) into Elasticsearch with stream processing and custom transforms.

    405 GitHub stars~684 tokensUpdated 5 mo ago
    Backend & APIsAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Categories

Questions about Glean Migration Deep Dive

What does Glean Migration Deep Dive do?

Migrate from Elasticsearch/Algolia: 1) Export all documents from source, 2) Transform to Glean document schema (id, title, url, body, permissions), 3) Create datasource with adddatasource, 4) Bulk…. Glean Migration Deep Dive is an agent skill from jeremylongshore/tons-of-skills-marketplace. Migrate from Elasticsearch/Algolia: 1) Export all documents from source, 2) Transform to Glean document schema (id, title, url, body, permissions), 3) Create datasource with adddatasource, 4) Bulk index with bulkindexdocuments, 5) Validate search quality with test queries, 6) Switch search UI to use Glean Client API.

When should I use Glean Migration Deep Dive?

Glean Migration Deep Dive fits situations like: tasks that involve Search implementation.

How do I install Glean Migration Deep Dive in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill glean-migration-deep-dive -a claude-code`. Or copy the skill folder (skills/.curated/glean-migration-deep-dive in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/glean-migration-deep-dive in your project. Claude Code loads it when a task matches its description.

How do I install Glean Migration Deep Dive in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill glean-migration-deep-dive -a codex`. Or copy the skill folder (skills/.curated/glean-migration-deep-dive in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/glean-migration-deep-dive in your project. Codex loads it when a task matches its description.

Can I use Glean Migration Deep Dive in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill glean-migration-deep-dive -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/glean-migration-deep-dive, .gemini/skills/glean-migration-deep-dive, .github/skills/glean-migration-deep-dive and .opencode/skills/glean-migration-deep-dive in your project.

What does Glean Migration Deep Dive need to run?

Going by SKILL.md and its folder, Glean Migration Deep Dive needs the command-line tools its instructions call (curl) and credentials named GLEAN_API_TOKEN. Our summary lists: A credential in GLEAN_API_TOKEN. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(npm:*), Bash(curl:*), Grep. Compatibility (from SKILL.md): Designed for Claude Code.

Does Glean Migration Deep Dive access the network?

SKILL.md names 1 domain. As links in the text: developers.glean.com. This is read from the text; nothing was executed.

Is Glean Migration Deep Dive safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Glean Migration Deep Dive use?

Glean Migration Deep Dive is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Glean Migration Deep Dive use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Glean Migration Deep Dive?

Skills that share tags, products or a category with Glean Migration Deep Dive: Product Full-Text Search (lobehub/lobehub, 83k stars), Hackernews (sigcli/sigcli, 293 stars), Foundatio Repositories (exceptionless/Exceptionless, 2.5k stars) and Elasticsearch Authn (aspectrr/deer, 405 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Glean Migration Deep Dive?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.