Agent skill

Bi Clustering

by agentscope-ai in agentscope-ai/QwenPaw-Data

对用户、产品等业务对象做分群:用波士顿矩阵法做象限分群,或用分层聚类、K-means、DBSCAN 等聚类技术分群。当需要做客群/产品分群、象限策略、画像或密度型子结构发现时调用。触发条件:当对话中出现“分群”、“分类”、“聚类”、“不同类型”、“不同场景”等体现分群分析词语时触发。

Apache-2.0Auto-check passed

Install Bi Clustering

skills CLI
$ npx skills add agentscope-ai/QwenPaw-Data --skill bi-clustering -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentscope-ai/QwenPaw-Data bi-clustering --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentscope-ai/QwenPaw-Data.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/qwenpaw-data-skills/skills/atomic/bi-clustering .claude/skills/bi-clustering && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bi-clustering
GitHub stars
127
Token cost
~1.6k tokens
SKILL.md length
356 words
Files
4 (incl. scripts)
Skills in repo
29
Repo updated
First seen
Licence
Apache-2.0

At a glance

对用户、产品等业务对象做分群:用波士顿矩阵法做象限分群,或用分层聚类、K-means、DBSCAN 等聚类技术分群。当需要做客群/产品分群、象限策略、画像或密度型子结构发现时调用。触发条件:当对话中出现“分群”、“分类”、“聚类”、“不同类型”、“不同场景”等体现分群分析词语时触发。

  • Works in 5 steps: 分群方法的选择 → 特征维度的确定 → 分群数据准备 → …
  • SKILL.md covers 执行流程, 输出结果 and 注意事项
  • Runs Python scripts from its folder; calls python

What it does

Bi Clustering is an agent skill from agentscope-ai/QwenPaw-Data. 对用户、产品等业务对象做分群:用波士顿矩阵法做象限分群,或用分层聚类、K-means、DBSCAN 等聚类技术分群。当需要做客群/产品分群、象限策略、画像或密度型子结构发现时调用。触发条件:当对话中出现“分群”、“分类”、“聚类”、“不同类型”、“不同场景”等体现分群分析词语时触发。

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `scripts/boston_quadrant.py`, `scripts/clustering.py` and `scripts/json_groups.py`).

The repository describes itself as: Agentic enterprise data analytics: governed facts (DataBridge), reusable methodology (Skill-Hub), and controllable execution (Host). The licence is Apache-2.0.

Example prompts

  • “/bi-clustering”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. 分群方法的选择
  2. 特征维度的确定
  3. 分群数据准备
  4. 数据分群
  5. 报道结果(两类路径共用)

What it can do on your machine

Read from SKILL.md and the folder at commit e0bae36. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bi Clustering loads about 1.6k tokens when it runs. Until then it costs about 39 tokens; SKILL.md has 356 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~39
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from agentscope-ai/QwenPaw-Data at commit e0bae36, republished under its Apache-2.0 licence (© agentscope-ai). 356 words, ~1,568 tokens.

Download SKILL.mdSave it as .claude/skills/bi-clustering/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bi-clustering
description
对用户、产品等业务对象做分群:用波士顿矩阵法做象限分群,或用分层聚类、K-means、DBSCAN 等聚类技术分群。当需要做客群/产品分群、象限策略、画像或密度型子结构发现时调用。触发条件:当对话中出现“分群”、“分类”、“聚类”、“不同类型”、“不同场景”等体现分群分析词语时触发。

bi-clustering

将特征相似或策略上需区分的业务对象(用户、产品、门店等)划为若干群组。根据分群分析的目的。

执行流程

1. 分群方法的选择

分析数据分析的目的和两种方法的适用场景,明确是使用波士顿矩阵法对分析对象进行分群还是应该使用聚类技术对分析对象完成分群。

路径更合适_when说明
波士顿矩阵法策略上需要 2×2 象限(如「增长×份额」「客流×客单」)、维度含义清晰、便于与经典业务框架对齐每条轴上把对象划为「高/低」两档,得到四象限;解释成本低,适合汇报与策略分派
聚类需要 多特征综合、簇数/形状事先不明确、或簇非球形/需标出噪声点用距离与密度在特征空间划分;可选 K-means、分层聚类、DBSCAN(参见下文)

若仅能在「简单四象限」与「多簇细划」之间二选一:优先波士顿当业务叙事依赖两个主轴;优先聚类当维度多、需数据驱动定簇结构。

2. 特征维度的确定

确定分群所依据的维度:若采用波士顿矩阵法,则确定两个象限用以分群;若采用聚类方法,则按需选择合适的分群特征。

实务上,波士顿路径需选定横轴、纵轴各一个指标且与目标一致;聚类路径需注意缺失、异常、量纲与类别编码(高基数类别慎用无约束 one-hot)。

3. 分群数据准备

查看数据,确认数据中包含在步骤 2 中确定的特征维度;然后,整理数据为 CSV 格式,数据包含一个分析对象列,和分群所需的特征维度,每个特征应该对应一个数据列。

补充:对象为行粒度,按需完成对象级汇总(如事件级先聚合到用户/产品)、缺失与异常处理;聚类路径下数值特征由脚本的 --scale 处理缩放。

4. 数据分群
a) 波士顿矩阵法

若象限对应数据是离散的,根据离散值将其分成两个区间;若象限对应数据是连续的,在连续轴上选定分界统计量(平均数或中位数),按该值将数据分成「≤ 分界 / > 分界」两个区间。完成象限划分后,将分析对象划分至各个象限。

说明:平均数对极端值敏感、中位数更稳健;横轴与纵轴可分别指定(脚本参数 --x-continuous-split / --y-continuous-split,值为 mean 或 median,默认均为 mean)。离散轴仍按有序取值前半/后半分为两档;auto 模式下由列类型与去重个数判定连续/离散(见脚本 --discrete-max-uniques)。结果中 stderr 的 axes: 一行在连续轴上会附带 :mean 或 :median。可向业务侧标注象限名称(如 Q1–Q4)并统计规模与指标概要。

分群结果保存为 json 文件。

使用 <skill-dir>/scripts/boston_quadrant.py 脚本完成数据的划分,如

bash
python scripts/boston_quadrant.py \
  --input-file data.csv \
  --id-col user_id \
  --x-col 市场份额 \
  --y-col 增长率 \
  --x-continuous-split mean \
  --y-continuous-split median \
  --output-json result.json

参数说明:

参数说明默认值
--input-file输入 CSV 路径(必填)
--id-col分析对象唯一标识列(必填)
--x-col / --y-col横轴、纵轴特征列各一列(必填)
--x-mode / --y-mode该轴划分方式:auto | continuous | discreteauto
--discrete-max-uniquesauto 时:数值列去重个数 ≤ 此阈值则按离散轴处理12
--x-continuous-split横轴为连续时:用 mean(均值)或 median(中位数)作分界,低为 ≤、高为 >mean
--y-continuous-split纵轴为连续时:同上mean
--output-json象限分群结果 JSON(格式见「输出结果」)(必填)
b) 聚类分析

聚类方法选择:

方法适用场景适合的数据类型 / 形态
K-meansK 可预估或可试算;簇大致球形、规模相近;样本量大、需快速迭代主要为连续数值(脚本内会按 --scale 处理);对离群点敏感
分层聚类需要树状结构或 K 不固定、多层解读;样本量中等数值矩阵 + 选定距离/连接法(脚本中欧氏 + linkage)
DBSCAN簇数未知;形状任意、密度不均;需显式噪声/未分类调节 eps、min_samples;高维时距离区分度可能下降

使用聚类脚本:

使用 <skill-dir>/scripts/clustering.py执行聚类,分群结果保存为 json 文件。脚本调用示例如下,

bash
python scripts/clustering.py \
  --input-file data.csv \
  --id-col user_id \
  --feature-cols 年龄 消费金额 访问次数 \
  --method kmeans \
  --n-clusters 5 \
  --output-json result.json

调优示例(K-means 按轮廓系数选 K;--output-json 仍必填,写出最终簇划分):

bash
python scripts/clustering.py \
  --input-file data.csv \
  --id-col user_id \
  --feature-cols 年龄 消费金额 访问次数 \
  --method kmeans \
  --tune \
  --k-min 2 \
  --k-max 10 \
  --output-json result.json \
  --tuning-report-file ./out/tune.json

DBSCAN 调优示例(--tune 时必须同时提供 --tune-eps 与 --tune-min-samples):

bash
python scripts/clustering.py \
  --input-file data.csv \
  --id-col user_id \
  --feature-cols f1 f2 \
  --method dbscan \
  --tune \
  --tune-eps 0.3,0.5,0.8,1.2 \
  --tune-min-samples 3,5,10 \
  --max-noise-ratio 0.35 \
  --output-json result.json \
  --tuning-report-file ./out/tune.json

聚类脚本参数说明

参数含义默认
--input-file输入 CSV 路径(必填)
--id-col分析对象唯一标识列名(必填)
--feature-cols参与聚类的数值特征列名,多个列名以空格分隔(脚本内 pd.to_numeric)(必填)
--methodkmeans | hierarchical | dbscan(必填)
--output-json聚类分群结果 JSON 路径(必填)
--scale聚类前特征缩放:standard | minmax | nonestandard
--random-stateK-means 随机种子42
--n-initK-means n_init10
--n-clustersK-means / 分层聚类的簇数 K;未使用 --tune 时与上述方法搭配为必填无
--linkage分层聚类连接法:ward、complete、average、singleward
--epsDBSCAN 邻域半径;未使用 --tune 时与 --min-samples 同时必填无
--min-samplesDBSCAN min_samples;未使用 --tune 时与 --eps 同时必填无
--skip-drop-na含 NaN 的特征行不丢弃(需 --fill-mean 或数据已无 NaN)默认会丢弃含 NaN 行
--fill-mean用列均值填补特征中的 NaN关闭
--centroids-file若指定且为 K-means:另写出质心 CSV(列为 --feature-cols,与聚类一致的缩放后空间;并带 cluster_label)无
--tune开启超参搜索:K-means/分层按轮廓系数选 K(分层可同时试多种 --tune-linkages);DBSCAN 在非噪声点上算轮廓且受 --max-noise-ratio 约束关闭
--k-min / --k-maxK-means / 分层在 --tune 时的 K 搜索范围2 / 10
--tune-linkages分层 --tune 时要尝试的连接法列表(每项为 ward 等)仅用当前 --linkage
--tune-epsDBSCAN --tune 必填:逗号分隔的 eps 候选,如 0.3,0.5,0.8无
--tune-min-samplesDBSCAN --tune 必填:逗号分隔的 min_samples 候选,如 3,5,10无
--max-noise-ratioDBSCAN --tune 时:噪声占比超过该值的参数组合被淘汰0.35
--tuning-report-file可选:写出调优过程与最终选中超参数的 JSON(不是分群 ID 列表文件)无
Show full SKILL.md (73 more words)Show less

拟合后结合写出的 JSON 汇报各簇(及噪声)规模、方法与全部关键超参数,并做业务解读。

5. 报道结果(两类路径共用)

必备内容:

  • 输出物:主结果 --output-json 文件路径(及调优 JSON 若使用);可得自脚本的 stderr 相对路径与 stdout 中的 JSON 内容
  • 数据与特征:对象粒度、样本量、特征列表、预处理(缺失、标准化、编码)
  • 方法与参数:波士顿两轴定义与分割规则(连续轴的 mean/median 与离散分档);或聚类方法理由与全部关键超参数
  • 群组规模:各象限或各簇的样本量、占比;DBSCAN 单独汇报噪声数量与占比
  • 群组画像:相对总体或跨组的业务指标/特征对比;避免仅有算法指标而无业务解释
  • 稳定性与局限(视情况):参数微扰或随机种子对标签的影响;高维、小样本对结论的影响

质量检查

  • 行粒度与「一个待分群对象」一致
  • 波士顿:两维分割规则在文档中一致且可复现
  • 聚类:连续特征缩放策略明确;方法与数据形态匹配
  • 各象限/簇规模可解释,结论能指向可执行动作

输出结果

分群/聚类结果以 JSON 格式呈现,包含以下字段和对应的聚类结果:

  • 聚类(clustering.py):键为 "cluster 1"、"cluster 2"、…(对应 sklearn 簇标签 0、1、…);DBSCAN 中标签 -1 的样本归入 "noise"(仅当存在噪声点时才有此键)。仅当簇内至少有一个对象时才出现对应键,不会出现空数组的簇键。
  • 波士顿象限(boston_quadrant.py):恒包含 "cluster 1"~"cluster 4" 四键(象限含义见该脚本文件头注释),某象限无对象时值为 []。连续轴分界可为 均值(mean) 或 中位数(median);运行结束后 stderr 中 axes: 一行对连续轴会附带 :mean 或 :median,便于核对所用规则。

聚类另可选用 --tuning-report-file 写出调优网格与选中超参数的 JSON,与上述「ID 分组」主结果文件相互独立。

注意事项

  1. 可解释性:无监督分群的组是否业务上有意义,需结合画像与验证;不宜单凭轮廓系数定成败。
  2. 高维:维度过高时距离趋于均匀;可考虑特征选择或领域驱动子集。
  3. 公平与合规:分群若用于差异化策略,需遵守隐私与反歧视相关规范。

© agentscope-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in packages/qwenpaw-data-skills/skills/atomic/bi-clustering of agentscope-ai/QwenPaw-Data.

  • SKILL.md
  • scripts/boston_quadrant.py
  • scripts/clustering.py
  • scripts/json_groups.py

Open the folder on GitHubat commit e0bae36

Compare with similar skills

Bi Clustering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bi Clustering compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bi Clustering this skillagentscope-ai/QwenPaw-Data127—~1.6kAutomated safety check: PassApache-2.0
Vector Clusterruvnet/ruflo74k—~540Automated safety check: NotesMIT
Exploring LLM ClustersPostHog/posthog40k—~3.1kAutomated safety check: PassCustom licence
SEO Keyword ClusteringAgriciDaniel/claude-seo19k2 repos~3.3kAutomated safety check: PassMIT
Exploring MCP Intent ClustersPostHog/posthog40k—~1.9kAutomated safety check: PassCustom licence
Running Clustering Algorithmsjeremylongshore/tons-of-skills-marketplace2.8k—~1kAutomated safety check: PassMIT

Similar skills

  • Vector Cluster

    ruvnet/ruflo

    Cluster code by graph community detection via npx ruvector@0.2.25 hooks graph-cluster (spectral / Louvain)

    74k GitHub stars~540 tokensUpdated today
    Auto-check: notes
  • Exploring LLM Clusters

    PostHog/posthog

    Official

    Investigate AI observability clusters — understand usage patterns in AI/LLM traffic, compare cluster behavior, compute cost/latency metrics, and drill into individual traces within clusters.

    40k GitHub stars~3.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • SEO Keyword Clustering

    AgriciDaniel/claude-seo

    Clusters keywords by how much their search results overlap and designs a hub-and-spoke content plan with an internal link matrix and an interactive cluster map.

    19k GitHub starsUsed in 2 repos~3.3k tokens
    Marketing & SEOAuto-check passed
  • Official

    Explore PostHog MCP intent clusters — agent goals grouped by semantic similarity, with each cluster's tool distribution and error rates, plus the tool-centric pivot (capture rate per intent…

    40k GitHub stars~1.9k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Running Clustering Algorithms

    jeremylongshore/tons-of-skills-marketplace

    Analyze datasets by running clustering algorithms (K-means, DBSCAN, hierarchical) to identify data groups.

    2.8k GitHub stars~1k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Official

    Trigger on mention of GKE cluster autoscaler, node autoscaling, node pool auto-creation / node auto-provisioning.

    21k GitHub stars~3k tokensUpdated today
    DevOps & CloudAuto-check passed

More from agentscope-ai/QwenPaw-Data

All 29 skills in this repo
  • Bi Report Generation

    agentscope-ai/QwenPaw-Data

    将 BI 数据分析结果组织成可视化 HTML 报告。当分析完成、需要生成报告时调用. An agent skill from agentscope-ai/QwenPaw-Data.

    127 GitHub stars~1.2k tokensUpdated 5 days ago
    Auto-check passed
  • Fetch Data

    agentscope-ai/QwenPaw-Data

    取数 / 查数据 / 拉数据 / 跑 SQL。把自然语言取数需求转为 SQL,经数据湖仓执行后返回查询结果供下游分析。任何需要业务数据的任务在工作区缺少对应文件时都必须先调用此技能——覆盖 BI 业务分析、留存 / 转化 / 同期群分析、数据探索 EDA、统计建模、定量计算、元数据查询、数据查询。命中任一即触发:(1) 直接索要指标或记录,如「DAU 多少」「上月销售额」「3…

    127 GitHub stars~2.5k tokensUpdated 5 days ago
    Auto-check passed
  • Bi Adaptive Threshold

    agentscope-ai/QwenPaw-Data

    通过量化历史数据的自然波动幅度,自适应计算判定阈值。当需要从数据本身确定阈值(如波动阈值、影响度阈值等)、而非使用固定值时调用。仅适用于日/周粒度阈值确定。

    127 GitHub stars~726 tokensUpdated 5 days ago
    Auto-check passed
  • Bi Anomaly Detection

    agentscope-ai/QwenPaw-Data

    基于阈值检测时间序列中的显著异常波动点。当需要找出指标异常波动日期、识别数据异动时调用. An agent skill from agentscope-ai/QwenPaw-Data.

    127 GitHub stars~637 tokensUpdated 5 days ago
    Auto-check passed
  • Bi Attribution Analysis

    agentscope-ai/QwenPaw-Data

    计算各维度(组)值对指标变动的贡献度,支持可加型量值指标和加权平均型/率值指标。当需要计算贡献度、解释指标"为什么涨/跌"时调用。

    127 GitHub stars~1.2k tokensUpdated 5 days ago
    Auto-check passed
  • Bi Causal Attribution

    agentscope-ai/QwenPaw-Data

    从运营周报、活动文档、对话输入或文档工具 API 中提取业务事件,与指标异常时间窗口对齐,生成有证据支撑的因果归因假设并排序。当已知指标存在异常波动、需要从外部文档证据中解释"为什么"时调用。

    127 GitHub stars~1.3k tokensUpdated 5 days ago
    Auto-check passed

Questions about Bi Clustering

What does Bi Clustering do?

对用户、产品等业务对象做分群:用波士顿矩阵法做象限分群,或用分层聚类、K-means、DBSCAN 等聚类技术分群。当需要做客群/产品分群、象限策略、画像或密度型子结构发现时调用。触发条件:当对话中出现“分群”、“分类”、“聚类”、“不同类型”、“不同场景”等体现分群分析词语时触发。. Bi Clustering is an agent skill from agentscope-ai/QwenPaw-Data.

How do I install Bi Clustering in Claude Code?

Run `npx skills add agentscope-ai/QwenPaw-Data --skill bi-clustering -a claude-code`. Or copy the skill folder (packages/qwenpaw-data-skills/skills/atomic/bi-clustering in agentscope-ai/QwenPaw-Data) into .claude/skills/bi-clustering in your project. Claude Code loads it when a task matches its description.

How do I install Bi Clustering in Codex?

Run `npx skills add agentscope-ai/QwenPaw-Data --skill bi-clustering -a codex`. Or copy the skill folder (packages/qwenpaw-data-skills/skills/atomic/bi-clustering in agentscope-ai/QwenPaw-Data) into .agents/skills/bi-clustering in your project. Codex loads it when a task matches its description.

Can I use Bi Clustering in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentscope-ai/QwenPaw-Data --skill bi-clustering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bi-clustering, .gemini/skills/bi-clustering, .github/skills/bi-clustering and .opencode/skills/bi-clustering in your project.

What does Bi Clustering need to run?

Going by SKILL.md and its folder, Bi Clustering needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Bi Clustering access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bi Clustering safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Bi Clustering use?

Bi Clustering is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bi Clustering use?

About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bi Clustering?

Skills that share tags, products or a category with Bi Clustering: Vector Cluster (ruvnet/ruflo, 74k stars), Exploring LLM Clusters (PostHog/posthog, 40k stars), SEO Keyword Clustering (AgriciDaniel/claude-seo, 19k stars) and Exploring MCP Intent Clusters (PostHog/posthog, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bi Clustering?

agentscope-ai (a GitHub organization) maintains it in agentscope-ai/QwenPaw-Data, which has 127 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 5, 2026.

Source: agentscope-ai/QwenPaw-Data on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.