Agent skill

Knowledge Bootstrap

by ai-analyst-lab in ai-analyst-lab/ai-analyst

Initialize session context, resolve the active dataset and context source, load resident instructions, and inventory the context available for question-specific selection.

MITAuto-check passed

Install Knowledge Bootstrap

skills CLI
$ npx skills add ai-analyst-lab/ai-analyst --skill knowledge-bootstrap -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai-analyst-lab/ai-analyst knowledge-bootstrap --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai-analyst-lab/ai-analyst.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/knowledge-bootstrap .claude/skills/knowledge-bootstrap && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
knowledge-bootstrap
GitHub stars
304
Token cost
~3.4k tokens
SKILL.md length
1,428 words
Files
1
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

Initialize session context, resolve the active dataset and context source, load resident instructions, and inventory the context available for question-specific selection.

  • Works in 11 steps: Setup State → Active Dataset → User Profile → …
  • SKILL.md covers Purpose, When to Use, Instructions and User Profile Template, plus 2 more sections
  • Calls python

What it does

Knowledge Bootstrap is an agent skill from ai-analyst-lab/ai-analyst. Initialize session context, resolve the active dataset and context source, load resident instructions, and inventory the context available for question-specific selection. Run at the start of every session and again after /connect-data or /switch-dataset. Handles missing files gracefully, so running it when unsure is harmless.

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: AI Product Analyst — Claude Code-powered data analysis toolkit. The licence is MIT.

Example prompts

  • “/knowledge-bootstrap”

Requirements

  • Python 3

Workflow steps

11 steps, taken from the step headings in SKILL.md.

  1. Setup State
  2. Active Dataset
  3. User Profile
  4. User Integrations
  5. Organization Context
  6. Corrections
  7. Learnings
  8. Query Archaeology
  9. Analysis Archive
  10. Mark Bootstrap Complete
  11. Report Readiness

What it can do on your machine

Read from SKILL.md and the folder at commit 52c0744. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Knowledge Bootstrap loads about 3.4k tokens when it runs. Until then it costs about 87 tokens; SKILL.md has 1,428 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~87
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai-analyst-lab/ai-analyst at commit 52c0744, republished under its MIT licence (© ai-analyst-lab). 1,428 words, ~3,416 tokens.

Download SKILL.mdSave it as .claude/skills/knowledge-bootstrap/SKILL.md (or your agent's skills folder).
name
knowledge-bootstrap
description
Initialize session context, resolve the active dataset and context source, load resident instructions, and inventory the context available for question-specific selection. Run at the start of every session and again after /connect-data or /switch-dataset. Handles missing files gracefully, so running it when unsure is harmless.

Skill: Knowledge Bootstrap

Purpose

Initialize the knowledge subsystems for a new session. Resolve the active context source, load the small resident layer, and inventory the selected and compiled context that can be supplied after the user asks a question.

When to Use

  • At the start of any session
  • After /connect-data or /switch-dataset
  • When the system detects missing or stale knowledge files

Instructions

Load each subsystem in order. An absent optional file can be reported as "not yet populated." A configured context source that is missing, invalid or inaccessible is different: report the problem and repair the connection before relying on its business definitions. Do not silently substitute local context.

Step 1: Setup State

Read .knowledge/setup-state.yaml.

  • Parse setup_complete and count phases with status: "complete".
  • If setup_complete: false, note incomplete phases to offer /setup.
  • If missing: Note "Setup: not initialized -- offer /setup".
Step 2: Active Dataset

Read .knowledge/active.yaml.

  • If active_dataset is null or missing, note "No active dataset" and continue.

  • Resolve the context source first. Call resolve_context_dir(active, project_root) from helpers/knowledge/context_sync.py -> (ctx_dir, source). source: path reads the visible external store directly, without a cache. source: git uses the legacy Git cache; source: local reads .knowledge/datasets/{active}/. Load dataset knowledge (semantic/, metrics/, schema.md, quirks.md) from ctx_dir either way - the same loader, the source just differs. Report the source ("context: local" or "context: team repo @ {ref}") in the readiness summary.

  • Inventory from ctx_dir. Load context-policy.yaml and custom_instructions.md as the resident layer. Do not load every metric, relationship, query, and correction into the prompt by default.

  • Confirm these components are available:

FileRequiredIf Missing
manifest.yamlYesNote "manifest missing -- not usable"
schema.mdYesGenerate via schema_to_markdown() or profiling
quirks.mdNoCreate empty template
metrics/index.yamlNoCount as 0
custom_instructions.md (root) else semantic/custom_instructions.mdNoSkip
verified_queries.yaml (root) else semantic/verified_queries.yamlNoSkip
corrections.md (root)NoSkip
semantic/entities.yamlNoNote "no semantic layer"
semantic/relationships.yamlNoSkip
semantic/dimensions.yamlNoSkip
semantic/measures.yamlNoSkip
semantic/filters.yamlNoSkip

Store layout — root or semantic/ (backward-compatible). Three of these files can live at EITHER the dataset root ({ctx_dir}/) OR under semantic/ ({ctx_dir}/semantic/), depending on the store's layout: custom_instructions.md, verified_queries.yaml, and corrections.md. Reconciled stores keep them at the dataset root; older stores keep the first two under semantic/. For each of the three, check the dataset root first; if present, load it from there, ELSE fall back to semantic/. Do not require one layout over the other, and do not skip the file just because it is absent from semantic/ — it may be at the root, and vice versa. The five pure-semantic YAMLs (entities, relationships, dimensions, measures, filters) always live under semantic/ and are not root-or-semantic.

corrections.md is the store-level communal corrections home — a human-curated, cross-session list of standing corrections that ship WITH the dataset context (root-or-nothing; there is no semantic/ fallback for it). It is DISTINCT from the per-session correction log at .knowledge/corrections/index.yaml loaded in Step 6 — that one is the local session log, this one is the communal store file. Load both; they are different subsystems.

Question-specific context before SQL. Once the exact analytical question is known, run /context-trace, or call helpers.knowledge.context_manifest directly. Load the selected items from that manifest, not the whole context store. Stop on a blocking conflict. Name stale or missing review evidence before relying on it. Resolve the question's metric, authoritative entities and relationships, real filter values, relevant verified queries, and applicable corrections from the selected bundle. A manifest proves what was supplied. It does not prove the worker used it. Reconcile cited items and SQL-use evidence after the analysis.

The context store separates three delivery modes:

  • resident context is small and broadly applicable;
  • selected context is chosen for the question and worker;
  • compiled context is executable, deterministic context such as a metric compile block.

Workspace guidance and task guides. Call python -m helpers.connected_context --dataset DATASET catalog for version-2 guides, query entries and semantic resources. Follow docs/CONNECTED-CONTEXT.md for typed links, loading and execution. Do not reinterpret a draft or failed dependency as missing optional context. Legacy guide discovery remains available below.

For legacy guides, call helpers.knowledge.context_guides.guide_catalog(project_root, dataset=active). Apply the small workspace_guidance to this session. Inspect guide descriptions and scope once the question is known; do not preload every guide body. For a relevant guide call load_guide with its ID, catalog hash, question, selection reason and the current analysis ID. Read the returned content before querying. The helper logs that delivered content in working/context_loads_<analysis_id>.jsonl. If two sources conflict or scope is unclear, ask rather than silently choosing a definition. Draft, expired and other-dataset guides are listed as excluded. This is a simple agent-selected catalog, not vector search or Hex's proprietary retrieval algorithm. A loading record proves delivery, not correct application.

Schema generation if schema.md is missing (REQUIRED):

The schema is critical for SQL queries and analysis — never proceed without it. Follow this sequence:

  1. Check data/schemas/{active}.yaml — if found, import schema_to_markdown() from helpers/data/schema_profiler.py and generate schema.md
  2. If no YAML schema file exists, use get_connection_for_profiling() to query the live database and generate schema.md from introspection
  3. For CSV datasets, read the first 1000 rows of each file with pandas, infer dtypes, and write schema.md with table/column/type info
  4. Staleness check: if last_profile.md exists and is newer than schema.md, regenerate

After generation, write schema.md to .knowledge/datasets/{active}/schema.md so future sessions can load it directly.

System variables from manifest:

Extract these variables for use in SQL queries and agent prompts:

  • {{SCHEMA}} — Schema prefix for external warehouses (e.g., "analytics", "prod")
  • {{DISPLAY_NAME}} — User-friendly dataset name for status messages
  • {{DATE_RANGE}} — Available date range (e.g., "2024-01-01 to 2026-03-31")
  • {{DATABASE}} — Database name or connection string

For Snowflake use manager.table_reference(table) to construct DATABASE.SCHEMA.TABLE; schema alone is insufficient. Other data warehouses have their own naming rules. For local DuckDB/CSV a schema prefix is typically absent.

Show full SKILL.md (484 more words)Show less
Step 3: User Profile

Read .knowledge/user/profile.md.

  • If exists: Apply Detail level, Chart preference, Narrative style.
  • If missing: Create from template (see below), note "Profile: new".

On explicit user corrections during session, update the profile: append YYYY-MM-DD | Assumed [X] | User prefers [Y] to the Corrections Log section. Never infer from silence.

Step 4: User Integrations

Read .knowledge/user/integrations.yaml.

  • Extract preferred_export_format, communication.detail_level.
  • Count configured channels (configured: true).
  • If missing: Note "Integrations: not configured -- defaults apply".
Step 5: Organization Context

Check for org ID in setup-state.yaml (phases.phase_3_business.data.organization_id) or in the active dataset manifest's organization field.

If an org ID exists and is not _example:

  • Resolve helpers.knowledge.context_snapshot.knowledge_root(project_root) first.
  • Read {resolved_root}/organizations/{org_id}/manifest.yaml for name, industry.
  • Read {resolved_root}/organizations/{org_id}/business/index.yaml for section counts (glossary terms, products, metrics, objectives, teams).
  • If org dir missing: Note "Org: linked but not found".

If no org linked: Note "Org: not configured".

Step 6: Corrections

Read .knowledge/corrections/index.yaml.

  • Extract total_corrections and by_severity counts.
  • If total_corrections > 0, highlight critical/high counts so agents check the full log before writing SQL.
  • If missing: Note "Corrections: not yet populated".
Step 7: Learnings

Read .knowledge/learnings/index.md.

  • Scan for category headings (### N. Category Name).
  • Note which categories have content entries vs are empty.
  • Do NOT load full content -- just category presence.
  • If missing: Note "Learnings: not yet populated".
Step 8: Query Archaeology

Read .knowledge/query-archaeology/curated/index.yaml.

  • Extract cookbook_entries, table_cheatsheets, join_patterns counts.
  • If missing: Note "Archaeology: not yet populated".
Step 9: Analysis Archive

Read .knowledge/analyses/index.yaml:

  • Extract total_analyses and last 5 entries (title, date, findings count, level).
  • If most recent analysis was <24h ago: Add to user-facing status as "Recent work: [title] from [date]" and suggest "Want to build on your recent analysis?" This helps users pick up where they left off.

Read .knowledge/analyses/_patterns.yaml:

  • Count patterns[] entries and note pattern names if any.
  • If missing: Note "Patterns: not yet populated".
Step 10: Mark Bootstrap Complete

Write a completion signal so agents can check if bootstrap already ran this session:

python
import yaml
from datetime import datetime

timestamp = datetime.now().isoformat()
with open('.knowledge/.bootstrap_timestamp', 'w') as f:
    yaml.dump({'last_bootstrap': timestamp, 'status': 'complete'}, f)

This prevents redundant re-runs mid-session. To check if bootstrap is needed, read this file and compare timestamps — if <5 minutes old, skip re-running.

Step 11: Report Readiness

Compile an internal context summary (held in working memory, not shown raw):

Setup: {complete (N/M phases) | incomplete (list missing) | not initialized}
Dataset: {display_name} ({source_type}, {N} tables, ~{rows} rows, {date_range}) | not configured
Profile: {role}, {detail_level} | new
Integrations: {preferred_format}, {N} channels | not configured
Org: {company} ({industry}), {N} glossary, {N} products, {N} metrics | not configured
Corrections: {N} logged ({N} critical, {N} high) | none
Learnings: {N}/{6} categories populated | not yet populated
Archaeology: {N} cookbook, {N} cheatsheets, {N} join patterns | not yet populated
Archive: {N} analyses, {N} recurring patterns | none

Then output the user-facing status:

Dataset: {display_name} ({source_type})
Tables: {N} tables, ~{row_count} rows
Date range: {date_range}
Metrics: {M} defined
Profile: {loaded | new}
Status: Ready for analysis

If a critical subsystem is missing (no dataset, no manifest), adjust the status and suggest /connect-data or /setup.


User Profile Template

markdown
# User Profile

Auto-created by knowledge bootstrap. Updated as the system learns preferences.

## Role & Expertise
- **Role:** _[auto-detected or user-specified]_
- **Technical level:** _[beginner | intermediate | advanced]_
- **SQL comfort:** _[none | basic | intermediate | advanced]_
- **Statistics comfort:** _[none | basic | intermediate | advanced]_
- **Domain:** _[e-commerce | fintech | saas | marketplace | other]_

## Communication Preferences
- **Detail level:** _[executive-summary | standard | deep-dive]_
- **Chart preference:** _[minimal | standard | chart-heavy]_
- **Narrative style:** _[bullet-points | prose | mixed]_

## Corrections Log
_Records of times the user corrected the system's assumptions._
<!-- Format: YYYY-MM-DD | What was wrong | What was right -->

Edge Cases

  • No .knowledge/ dir: Create full tree and prompt /connect-data.
  • Empty schema.md: Regenerate via profiling.
  • No data files: Suggest checking connection or falling back to CSV.
  • Multiple datasets: Report active, remind about /switch-dataset.
  • Setup incomplete: Note phases, do not block. Suggest /setup.

Anti-Patterns

  1. Never skip bootstrap. Always read manifest -- details may have changed.
  2. Never hardcode dataset names. Resolve from active.yaml.
  3. Never modify manifest during bootstrap. Bootstrap is read-only.
  4. Never dump raw YAML to the user. Show the brief status, not the load.
  5. Distinguish absent optional context from a broken configured source. Never hide a failed connection by substituting a different definition.

© ai-analyst-lab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/knowledge-bootstrap of ai-analyst-lab/ai-analyst.

Open the folder on GitHubat commit 52c0744

Compare with similar skills

Knowledge Bootstrap next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Knowledge Bootstrap compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Knowledge Bootstrap this skillai-analyst-lab/ai-analyst304—~3.4kAutomated safety check: PassMIT
DatasetsArize-ai/phoenix12k—~1.6kAutomated safety check: PassCustom licence
Bootstrap Soulbytedance/deer-flow83k2 repos~1.2kAutomated safety check: PassMIT
Modeling Activation MetricsPostHog/posthog40k—~1.4kAutomated safety check: PassCustom licence
Initiate SetupNousResearch/hermes-agent252k—~2.3kAutomated safety check: PassMIT
Dataset Curationwshobson/agents40k—~2kAutomated safety check: PassMIT

Similar skills

  • Datasets

    Arize-ai/phoenix

    Understand what a Phoenix dataset is and reason well about its examples, outputs, splits, and how it feeds evaluators and experiments.

    12k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Bootstrap Soul

    bytedance/deer-flow

    Runs a short adaptive conversation with you and writes a personalized SOUL.md that defines your AI partner's name, personality, communication style and boundaries.

    83k GitHub starsUsed in 2 repos~1.2k tokens
    Agent WorkflowsAuto-check passed
  • Official

    Build reusable activation models — an activation-rate metric and a per-user/per-account activated flag — on either PostHog data-warehouse views (HogQL) or an external dbt project.

    40k GitHub stars~1.4k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Initiate Setup

    NousResearch/hermes-agent

    Run the first-run setup chat in the Hermes desktop app. An agent skill from NousResearch/hermes-agent.

    252k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Dataset Curation

    wshobson/agents

    Prepare, format, and validate datasets for supervised fine-tuning and preference training.

    40k GitHub stars~2k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Arize Dataset

    github/awesome-copilot

    Official

    Creates, manages, and queries Arize datasets and examples. An agent skill from github/awesome-copilot.

    40k GitHub starsUsed in 1 repo~3.9k tokens
    Testing & QAAuto-check: notes

More from ai-analyst-lab/ai-analyst

All 43 skills in this repo
  • Always Compare

    ai-analyst-lab/ai-analyst

    Never present a metric or number in isolation; anchor every number to a comparison (prior period, benchmark, or another segment) or state that none is available.

    304 GitHub stars~1.4k tokensUpdated 7 days ago
    Auto-check passed
  • Archaeology

    ai-analyst-lab/ai-analyst

    Retrieve proven SQL patterns, table cheatsheets, and join patterns from .knowledge/query-archaeology/ so past work gets reused.

    304 GitHub stars~1.3k tokensUpdated 7 days ago
    Auto-check passed
  • Archive Analysis

    ai-analyst-lab/ai-analyst

    Save completed analyses to the knowledge system's analysis archive for future reference.

    304 GitHub stars~2.7k tokensUpdated 7 days ago
    Auto-check passed
  • Auth Preflight

    ai-analyst-lab/ai-analyst

    Verify Google Workspace MCP authentication at the start of any session that needs Google APIs (Docs, Slides, Drive).

    304 GitHub stars~3.1k tokensUpdated 7 days ago
    Auto-check passed
  • Causal

    ai-analyst-lab/ai-analyst

    Causal inference toolkit for when experiments are not possible: estimate treatment effects from observational data with assumption checks and mandatory caveats.

    304 GitHub stars~1.8k tokensUpdated 7 days ago
    Auto-check passed
  • Chart To Drive

    ai-analyst-lab/ai-analyst

    Standardized workflow for uploading local chart PNGs to Google Drive and making them available for insertion into Google Docs and Slides.

    304 GitHub stars~1.4k tokensUpdated 7 days ago
    Auto-check passed

Questions about Knowledge Bootstrap

What does Knowledge Bootstrap do?

Initialize session context, resolve the active dataset and context source, load resident instructions, and inventory the context available for question-specific selection. Knowledge Bootstrap is an agent skill from ai-analyst-lab/ai-analyst. Initialize session context, resolve the active dataset and context source, load resident instructions, and inventory the context available for question-specific selection.

How do I install Knowledge Bootstrap in Claude Code?

Run `npx skills add ai-analyst-lab/ai-analyst --skill knowledge-bootstrap -a claude-code`. Or copy the skill folder (.claude/skills/knowledge-bootstrap in ai-analyst-lab/ai-analyst) into .claude/skills/knowledge-bootstrap in your project. Claude Code loads it when a task matches its description.

How do I install Knowledge Bootstrap in Codex?

Run `npx skills add ai-analyst-lab/ai-analyst --skill knowledge-bootstrap -a codex`. Or copy the skill folder (.claude/skills/knowledge-bootstrap in ai-analyst-lab/ai-analyst) into .agents/skills/knowledge-bootstrap in your project. Codex loads it when a task matches its description.

Can I use Knowledge Bootstrap in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-analyst-lab/ai-analyst --skill knowledge-bootstrap -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/knowledge-bootstrap, .gemini/skills/knowledge-bootstrap, .github/skills/knowledge-bootstrap and .opencode/skills/knowledge-bootstrap in your project.

What does Knowledge Bootstrap need to run?

Going by SKILL.md and its folder, Knowledge Bootstrap needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Knowledge Bootstrap access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Knowledge Bootstrap safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Knowledge Bootstrap use?

Knowledge Bootstrap is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Knowledge Bootstrap use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Knowledge Bootstrap?

Skills that share tags, products or a category with Knowledge Bootstrap: Datasets (Arize-ai/phoenix, 12k stars), Bootstrap Soul (bytedance/deer-flow, 83k stars), Modeling Activation Metrics (PostHog/posthog, 40k stars) and Initiate Setup (NousResearch/hermes-agent, 252k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Knowledge Bootstrap?

ai-analyst-lab (a GitHub organization) maintains it in ai-analyst-lab/ai-analyst, which has 304 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on September 30, 2026.

Source: ai-analyst-lab/ai-analyst on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.