Agent skill

Taobao MCP Benchmark

by LeoYeAI in LeoYeAI/openclaw-master-skills

淘宝桌面版MCP工具评测框架。用于系统化测试MCP工具的各项功能,生成专业的技术评测报告。Use when 需要对淘宝MCP工具进行评测、测试、验收、迭代验证。

MITAuto-check passedAgent Workflows

Install Taobao MCP Benchmark

skills CLI
$ npx skills add LeoYeAI/openclaw-master-skills --skill taobao-mcp-benchmark -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install LeoYeAI/openclaw-master-skills taobao-mcp-benchmark --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/taobao-mcp-benchmark .claude/skills/taobao-mcp-benchmark && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
taobao-mcp-benchmark
GitHub stars
2.2k
Token cost
~3k tokens
SKILL.md length
566 words
Files
9 (incl. scripts)
Skills in repo
1,235
Repo updated
First seen
Licence
MIT

At a glance

淘宝桌面版MCP工具评测框架。用于系统化测试MCP工具的各项功能,生成专业的技术评测报告。Use when 需要对淘宝MCP工具进行评测、测试、验收、迭代验证。

  • Works in 4 steps: 初始化评测任务 → 执行评测任务 → 生成评测报告 → …
  • 需要对淘宝MCP工具进行评测、测试、验收、迭代验证
  • SKILL.md covers 概述, ⚠️ 执行原则(必须遵守), 适用场景 and 评测任务清单, plus 4 more sections
  • Runs Shell and JavaScript scripts from its folder

What it does

Taobao MCP Benchmark is an agent skill from LeoYeAI/openclaw-master-skills. 淘宝桌面版MCP工具评测框架。用于系统化测试MCP工具的各项功能,生成专业的技术评测报告。Use when 需要对淘宝MCP工具进行评测、测试、验收、迭代验证。

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including scripts (for example `_meta.json`, `history/benchmark_history.md` and `scripts/generate_report.js`).

It sits in Agent Workflows, covering MCP servers. It works with Model Context Protocol. The repository describes itself as: 🧠 Curated collection of 1209+ best OpenClaw skills — weekly updated by MyClaw.ai. The licence is MIT.

When your agent uses it

  • 需要对淘宝MCP工具进行评测、测试、验收、迭代验证
  • Tasks that involve MCP servers

Example prompts

  • “/taobao-mcp-benchmark”

Requirements

  • Node.js
  • A Bash shell

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. 初始化评测任务
  2. 执行评测任务
  3. 生成评测报告
  4. 更新评测记录

What it can do on your machine

Read from SKILL.md and the folder at commit e5199b5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Shell and JavaScript), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Taobao MCP Benchmark loads about 3k tokens when it runs. Until then it costs about 25 tokens; SKILL.md has 566 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~25
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from LeoYeAI/openclaw-master-skills at commit e5199b5, republished under its MIT licence (© LeoYeAI). 566 words, ~3,018 tokens.

Download SKILL.mdSave it as .claude/skills/taobao-mcp-benchmark/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
taobao-mcp-benchmark
description
淘宝桌面版MCP工具评测框架。用于系统化测试MCP工具的各项功能,生成专业的技术评测报告。Use when 需要对淘宝MCP工具进行评测、测试、验收、迭代验证。
version
1.4.1

淘宝桌面版MCP评测框架

概述

本skill提供一套系统化的评测框架,用于测试淘宝桌面版MCP工具的各项功能,并生成专业的技术评测报告。

⚠️ 执行原则(必须遵守)

原子性执行

评测任务一旦开始,必须完整执行完成,不可中断。

原则说明
不可中断开始评测后,必须完成所有5个任务 + 生成报告
完整流程初始化 → 任务1-5 → 截图收集 → 报告生成 → 清理
状态跟踪每个任务完成后记录 checkpoint,便于恢复
用户提醒如用户试图中断,提醒"评测任务未完成,是否继续?"
任务状态管理

评测开始时创建状态文件 ~/.copaw/tasks/benchmark_YYYYMMDD_HHMMSS/status.json:

json
{
  "benchmark_id": "20260317_145034",
  "version": "1.2.0",
  "start_time": "2026-03-17 14:50:00",
  "status": "running",
  "current_task": 1,
  "tasks": [
    {"id": 1, "name": "淘金币签到", "status": "pending", "score": null},
    {"id": 2, "name": "商品搜索+对比+加购", "status": "pending", "score": null},
    {"id": 3, "name": "订单管理", "status": "pending", "score": null},
    {"id": 4, "name": "获取购物车以及降价信息", "status": "pending", "score": null},
    {"id": 5, "name": "客服咨询对话", "status": "pending", "score": null}
  ],
  "screenshots": [],
  "report_generated": false
}

每个任务完成后立即更新状态:

bash
# 任务完成后更新
echo '{"id": 1, "status": "completed", "score": 9, "end_time": "..."}' >> status.json
中断恢复机制

如果会话中断,下次用户询问评测时:

  1. 检查 status.json 是否存在
  2. 如果存在未完成任务:
    • 提示用户:"发现未完成的评测任务(任务X/Y),是否继续?"
    • 用户确认后,从 current_task 继续执行
  3. 如果已完成但未生成报告:
    • 直接生成报告
执行流程图
开始评测
    │
    ▼
创建任务目录 + status.json
    │
    ▼
┌─────────────────────────────┐
│  任务1:淘金币签到           │◄─── 记录截图、耗时、结果
│  任务2:商品搜索+对比+加购   │◄─── 记录截图、耗时、结果
│  任务3:订单管理            │◄─── 记录截图、耗时、结果
│  任务4:获取购物车以及降价信息 │◄─── 记录截图、耗时、结果
│  任务5:客服咨询对话        │◄─── 记录截图、耗时、结果
└─────────────────────────────┘
    │
    ▼
收集所有截图
    │
    ▼
生成 Word 报告(含截图)
    │
    ▼
更新 status.json → completed
    │
    ▼
输出评测结果摘要
禁止操作
禁止行为原因
❌ 任务中途停止导致评测数据不完整
❌ 跳过任务影响总分计算
❌ 跳过截图报告缺失关键证据
❌ 不生成报告用户无法查看结果
用户中断处理

如果用户在评测过程中说"停"、"不做了"等:

AI:⚠️ 评测任务尚未完成(已完成 X/5 个任务)。
    中断将导致评测数据不完整,无法生成完整报告。
    是否继续完成评测?(建议选择"继续")
    
    - 继续:继续执行剩余任务
    - 中断:停止评测,生成不完整报告(不推荐)

适用场景

  • MCP工具版本更新后的回归测试
  • 新功能发布前的验收测试
  • 定期质量检查和稳定性监控
  • 问题复现和性能基准测试

评测任务清单

任务1:淘金币签到(权重 25%)

测试目标:验证导航、元素识别、点击操作的稳定性

测试步骤:

  1. navigate → 首页
  2. scan_page_elements → 识别淘金币入口
  3. click_element → 进入淘金币页面
  4. read_page_content → 读取金币数量
  5. 完成签到任务(逛商品等)
  6. 验证金币增加

评分标准:

指标分值
导航成功2分
元素识别准确2分
点击操作成功2分
金币增加验证2分
流程顺畅度2分

任务2:商品搜索+对比+加购(权重 30%)

测试目标:验证搜索、详情查看、SKU选择、加购流程

测试步骤:

  1. search_products → 搜索关键词(如"保温杯")
  2. read_page_content → 读取搜索结果
  3. 筛选前3个商品进行对比
  4. click_element → 进入商品详情页
  5. read_page_content → 读取商品信息
  6. add_to_cart → 加入购物车(带SKU参数)

评分标准:

指标分值
搜索返回结果2分
商品详情页导航2分
信息提取完整2分
SKU选择准确2分
加购成功2分

任务3:订单管理(权重 20%)

测试目标:验证订单页面导航、状态筛选功能

测试步骤:

  1. navigate → 订单页面
  2. scan_page_elements → 识别筛选标签
  3. 依次测试:待付款、待发货、待收货、待评价
  4. read_page_content → 读取订单列表
  5. 验证筛选功能正常

评分标准:

指标分值
订单页面导航2分
筛选标签识别2分
筛选功能正常2分
订单信息读取2分
页面切换流畅2分

任务4:获取购物车以及降价信息(权重 20%)

测试目标:验证购物车导航、商品列表读取、降价信息提取

测试步骤:

  1. navigate → 购物车页面
  2. read_page_content → 读取商品列表
  3. 统计购物车商品总数
  4. 点击"降价"标签筛选降价商品
  5. read_page_content → 读取降价商品详情
  6. 记录降价商品数量和降价金额

评分标准:

指标分值
购物车导航成功2分
商品列表读取完整2分
降价标签点击成功2分
降价信息提取准确2分
数据记录完整2分

输出数据:

  • 购物车商品总数
  • 降价商品数量
  • 每个降价商品的:商品名、原价、券后价、降价金额

任务5:客服咨询对话(权重 15%)

测试目标:验证搜索商品、发起客服咨询、多轮对话功能

测试步骤:

  1. 随机选择一个商品主题(如:鼠标、键盘、台灯等)
  2. search_products → 搜索商品
  3. open_chat_from_search → 进入商家客服对话
  4. 发起第一轮咨询:"你好,请问这个商品今天下单,3天后能到杭州吗?"
  5. 等待客服回复(最多60秒)
  6. send_chat_message → 发起第二轮追问:"好的,那发什么快递呢?可以发顺丰吗?"
  7. 等待客服回复(最多60秒)
  8. 记录两轮对话内容

评分标准:

指标分值
商品搜索成功1分
进入客服对话1分
第一轮对话发送成功1.5分
客服第一次回复接收1.5分
第二轮追问发送成功2分
客服第二次回复接收2分
对话记录完整1分

工具调用:

bash
# 搜索商品
search_products keyword="鼠标"

# 通过搜索进入客服对话
open_chat_from_search query="鼠标" message="你好,请问这个商品今天下单,3天后能到杭州吗?"

# 发送第二轮追问(等待客服回复后)
send_chat_message message="好的,那发什么快递呢?可以发顺丰吗?"

注意事项:

  • 优先选择官方旗舰店或高销量店铺
  • 如果客服回复较慢,等待时间不超过60秒
  • 必须完成两轮对话才算任务完成
  • 记录两轮客服回复内容用于验证
  • 如果客服长时间未回复,可主动发送追问(不算失败)

评测流程

1. 初始化评测任务
bash
# 创建评测任务目录
mkdir -p ~/.copaw/tasks/benchmark_$(date +%Y%m%d_%H%M%S)/screenshots

# 记录评测开始时间
echo "评测开始时间: $(date '+%Y-%m-%d %H:%M:%S')" > ~/.copaw/tasks/benchmark_*/timing.log
2. 执行评测任务

必须严格遵守以下规范:

截图规范(每个任务必须)
截图时机文件命名说明
任务开始XX_task_start.png任务开始时的页面状态
关键操作前XX_step_N_操作名_before.png操作前的页面状态
关键操作后XX_step_N_操作名_after.png操作后的页面状态
任务完成XX_task_end.png任务完成时的页面状态
异常/问题XX_issue_N.png发现问题时的截图

截图命令:

bash
screencapture -x ~/.copaw/tasks/benchmark_*/screenshots/01_task_start.png
耗时统计(每个操作必须)
bash
# 操作开始
START_TIME=$(date +%s)

# 执行操作(如 navigate、click 等)

# 操作结束,计算耗时
END_TIME=$(date +%s)
echo "navigate_home: $((END_TIME - START_TIME))秒" >> timing.log
工具调用记录

每次工具调用必须记录:

  • 工具名称
  • 调用参数
  • 返回结果摘要
  • 是否成功
  • 耗时
bash
echo "$(date '+%H:%M:%S') | navigate | page=home | success | 2.3s" >> calls.log
3. 生成评测报告

报告命名规范(必须遵守):

项目格式示例
报告标题淘宝桌面版MCP评测报告 {YYYY-MM-DD}淘宝桌面版MCP评测报告 2026-03-17
Word文件名淘宝桌面版MCP评测报告 {YYYY-MM-DD}.docx淘宝桌面版MCP评测报告 2026-03-17.docx
Markdown文件名report_{YYYY-MM-DD}.mdreport_2026-03-17.md

Word 报告必须包含以下内容:

第一部分:整体小结(必须)
  • 评测概览表格(版本、时间、环境)
  • 总体评分和等级
  • 任务完成度统计表
  • 工具调用总览表
  • 耗时分布图/表
  • 发现问题汇总表
  • 关键结论
第二部分:分任务详情(每个任务必须包含)

每个任务需包含:

  1. 任务概要

    • 任务名称和目标
    • 开始/结束时间
    • 耗时统计
    • 评分和完成状态
  2. 执行流程表

    • 步骤编号
    • 操作描述
    • 工具名称
    • 调用参数
    • 返回结果
    • 是否成功
    • 耗时
  3. 过程截图

    • 每个关键步骤的截图(嵌入文档)
    • 截图说明文字
  4. 数据结果

    • 具体的数据(如金币数、商品数等)
    • 对比表格
  5. 问题分析

    • 发现的问题列表
    • 问题截图和标注
    • 影响评估
    • 建议解决方案
  6. 评价与建议

    • 优点总结
    • 可优化点
Show full SKILL.md (227 more words)Show less
第三部分:技术分析
  • 工具调用统计表(工具名、调用次数、成功率、平均耗时)
  • 性能指标表
  • 问题清单(编号、描述、影响范围、优先级、状态)
第四部分:附录
  • 完整截图清单
  • 工具调用日志
  • 相关文件路径
4. 更新评测记录

将评测结果追加到 benchmark_history.md


工具调用规范

导航操作
bash
# 优先使用专用导航
mcporter call taobao-native.navigate --args '{"target":"home"}' --output json
mcporter call taobao-native.navigate --args '{"target":"cart"}' --output json
mcporter call taobao-native.navigate --args '{"target":"order"}' --output json
元素扫描
bash
# 使用filter参数缩小范围
mcporter call taobao-native.scan_page_elements --args '{"filter":"淘金币"}' --output json
mcporter call taobao-native.scan_page_elements --args '{"filter":"保温杯"}' --output json
内容读取
bash
# 使用scope参数限定范围
mcporter call taobao-native.read_page_content --args '{"maxLength":3000}' --output json
截图保存
bash
# 使用screencapture命令
screencapture -x ~/.copaw/tasks/benchmark_*/screenshots/01_step_name.png

评分计算

总分 = 任务1得分 × 0.20 + 任务2得分 × 0.30 + 任务3得分 × 0.15 + 任务4得分 × 0.20 + 任务5得分 × 0.15

任务权重:

任务权重
1. 淘金币签到20%
2. 商品搜索+对比+加购30%
3. 订单管理15%
4. 获取购物车以及降价信息20%
5. 客服咨询对话15%

评分等级:

  • 9-10分:优秀 ⭐⭐⭐⭐⭐
  • 7-8分:良好 ⭐⭐⭐⭐
  • 5-6分:及格 ⭐⭐⭐
  • 3-4分:需改进 ⭐⭐
  • 0-2分:不合格 ⭐

常见问题与解决方案

问题1:搜索结果页停留在首页

现象:search_products 返回结果,但页面仍在首页

解决方案:

  1. 检查当前页面URL
  2. 使用 scan_page_elements 确认搜索结果
  3. 必要时重新导航
问题2:元素点击失败

现象:click_element 返回失败

解决方案:

  1. 检查元素是否可见
  2. 尝试滚动页面后再点击
  3. 使用text参数模糊匹配
问题3:SKU选择失败

现象:add_to_cart 提示SKU参数错误

解决方案:

  1. 先进入商品详情页
  2. 使用 scan_page_elements 获取可用SKU选项
  3. 按文本匹配选择

评测报告结构

Word 报告采用总分结构,面向技术团队,聚焦评测过程和问题分析。

报告大纲
淘宝桌面版MCP评测报告 {YYYY-MM-DD}
│
├── 一、整体小结 ⭐ 必须首先呈现
│   ├── 1.1 评测概览
│   │   └── 表格:评测日期、版本、环境、总耗时
│   ├── 1.2 总体评分
│   │   └── 大字号评分 + 等级 + 雷达图(可选)
│   ├── 1.3 任务完成度
│   │   └── 表格:任务名、权重、评分、状态、完成率
│   ├── 1.4 工具调用总览
│   │   └── 表格:工具名、调用次数、成功率、平均耗时
│   ├── 1.5 耗时分布
│   │   └── 表格:任务名、耗时、占比
│   ├── 1.6 问题汇总
│   │   └── 表格:问题编号、描述、影响范围、优先级
│   └── 1.7 关键结论
│       └── 3-5条核心结论
│
├── 二、分任务详情
│   ├── 2.1 任务一:淘金币签到
│   │   ├── 2.1.1 任务概要
│   │   │   └── 表格:目标、时间、耗时、评分
│   │   ├── 2.1.2 执行流程
│   │   │   └── 详细表格:每步操作、工具、参数、结果、耗时
│   │   ├── 2.1.3 过程截图 ⭐ 必须嵌入
│   │   │   ├── 图1:首页淘金币入口
│   │   │   ├── 图2:淘金币页面
│   │   │   └── ... 每个关键步骤
│   │   ├── 2.1.4 数据结果
│   │   │   └── 金币数、签到天数等具体数据
│   │   ├── 2.1.5 问题分析
│   │   │   ├── 问题描述 + 截图标注
│   │   │   └── 影响评估 + 建议方案
│   │   └── 2.1.6 评价与建议
│   │
│   ├── 2.2 任务二:商品搜索+对比+加购
│   │   ├── 2.2.1 任务概要
│   │   ├── 2.2.2 执行流程
│   │   ├── 2.2.3 过程截图 ⭐
│   │   │   ├── 搜索结果页
│   │   │   ├── 商品详情页
│   │   │   ├── SKU选择
│   │   │   └── 加购成功
│   │   ├── 2.2.4 数据结果
│   │   ├── 2.2.5 问题分析
│   │   └── 2.2.6 评价与建议
│   │
│   ├── 2.3 任务三:订单管理
│   │   └── (同上结构)
│   │
│   ├── 2.4 任务四:获取购物车以及降价信息
│   │   └── (同上结构)
│   │
│   └── 2.5 任务五:客服咨询对话
│       └── (同上结构)
│
├── 三、技术分析
│   ├── 3.1 工具调用统计
│   │   └── 详细表格:工具、调用次数、成功、失败、成功率、总耗时、平均耗时
│   ├── 3.2 性能指标
│   │   └── 表格:总任务数、成功率、总耗时、平均耗时、截图数、调用总数
│   ├── 3.3 问题清单
│   │   └── 表格:编号、问题描述、复现步骤、影响范围、优先级、建议方案
│   └── 3.4 改进建议
│       ├── 短期(1周内)
│       ├── 中期(1个月内)
│       └── 长期(3个月内)
│
└── 四、附录
    ├── 4.1 完整截图清单
    │   └── 表格:序号、文件名、说明、对应任务
    ├── 4.2 工具调用日志
    │   └── 完整的调用记录
    └── 4.3 相关文件
        └── Markdown报告、Word报告、截图目录路径
报告要点
要点要求说明
总分结构必须先整体小结,再分任务详情
截图嵌入必须每个关键步骤必须有截图,嵌入Word文档
耗时统计必须每个操作、每个任务、总体都要有耗时
问题标注必须发现问题必须在截图上标注,并说明影响
工具调用日志必须完整记录每次工具调用的参数和结果
数据具体化必须用具体数字代替模糊描述(如"返回48个商品"而非"返回多个商品")
面向技术团队必须使用专业术语,聚焦技术细节和问题分析

迭代记录

版本日期变更内容
v1.4.12026-03-17报告标题和文件名增加日期,便于识别
v1.4.02026-03-17任务4改名"获取购物车以及降价信息",任务5要求至少两轮对话
v1.3.02026-03-17新增原子性执行原则:任务不可中断、状态管理、中断恢复机制
v1.2.02026-03-17优化报告结构:总分结构、详细截图规范、耗时统计、问题标注
v1.1.02026-03-17新增任务5:客服咨询对话,调整任务权重
v1.0.02026-03-17初始版本,完成首次评测(4个任务)
v1.4.1 更新内容

报告命名优化:

  • 报告标题格式:淘宝桌面版MCP评测报告 {YYYY-MM-DD}
  • Word文件名格式:淘宝桌面版MCP评测报告 {YYYY-MM-DD}.docx
  • Markdown文件名格式:report_{YYYY-MM-DD}.md
  • 目的:便于识别和管理多次评测记录
v1.4.0 更新内容

任务4调整:

  • 原名称:购物车比价
  • 新名称:获取购物车以及降价信息
  • 优化评分标准:聚焦购物车商品统计和降价信息提取

任务5调整:

  • 要求:必须完成至少两轮对话
  • 第一轮:发起咨询(如发货时间)
  • 第二轮:追问(如快递方式)
  • 评分标准更新:两轮对话各占2分,回复接收各占2分
v1.3.0 更新内容

原子性执行原则:

  • 评测任务一旦开始,必须完整执行完成,不可中断
  • 完整流程:初始化 → 任务1-5 → 截图收集 → 报告生成 → 清理

状态管理机制:

  • 创建 status.json 跟踪任务进度
  • 每个任务完成后立即更新状态
  • 支持中断恢复:下次询问时检测未完成任务

用户中断处理:

  • 用户尝试中断时提醒"评测任务未完成"
  • 提供"继续"或"中断"选项
  • 中断后生成不完整报告(不推荐)

禁止操作清单:

  • ❌ 任务中途停止
  • ❌ 跳过任务
  • ❌ 跳过截图
  • ❌ 不生成报告
v1.2.0 更新内容

报告结构优化:

  • 采用总分结构:先整体小结,再分任务详情
  • 面向技术团队,聚焦评测过程和问题分析

新增规范:

  • 截图规范:每个关键步骤必须截图并嵌入文档
  • 耗时统计:每个操作、每个任务、总体都要有耗时记录
  • 问题标注:发现问题必须在截图上标注
  • 工具调用日志:完整记录每次调用的参数和结果
  • 数据具体化:用具体数字代替模糊描述

报告内容强化:

  • 整体小结新增:任务完成度表、工具调用总览表、耗时分布表、问题汇总表
  • 分任务详情新增:执行流程详细表、过程截图嵌入、问题分析章节
  • 技术分析强化:工具调用统计表增加成功/失败/平均耗时列
v1.1.0 更新内容

新增任务:客服咨询对话(权重15%)

  • 随机选择商品主题进行搜索
  • 通过搜索进入商家客服对话
  • 发起至少两轮客服咨询
  • 记录客服回复内容

权重调整:

任务v1.0.0v1.1.0v1.4.0
1. 淘金币签到25%20%20%
2. 商品搜索+对比+加购30%30%30%
3. 订单管理20%15%15%
4. 获取购物车以及降价信息25%20%20%
5. 客服咨询对话-15%15%(新增)

文件结构

~/.copaw/active_skills/taobao-mcp-benchmark/
├── SKILL.md                    # 本文档
├── templates/
│   ├── task_template.json      # 任务配置模板
│   └── report_template.md      # 报告模板
├── scripts/
│   └── generate_report.js      # Word报告生成脚本
└── history/
    └── benchmark_history.md    # 评测历史记录

快速开始

用户:帮我评测一下淘宝MCP工具
AI:好的,开始执行淘宝桌面版MCP评测...
    [执行4个评测任务]
    [生成评测报告]
    评测完成!总分:8.3/10

最后更新:2026-03-17 v1.4.1

© LeoYeAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts) in skills/taobao-mcp-benchmark of LeoYeAI/openclaw-master-skills.

  • SKILL.md
  • _meta.json
  • history/benchmark_history.md
  • scripts/generate_report.js
  • scripts/generate_report.sh
  • scripts/run_benchmark.sh
  • templates/report_template.md
  • templates/status_template.json
  • templates/task_template.json

Open the folder on GitHubat commit e5199b5

Compare with similar skills

Taobao MCP Benchmark next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Taobao MCP Benchmark compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Taobao MCP Benchmark this skillLeoYeAI/openclaw-master-skills2.2k—~3kAutomated safety check: PassMIT
MCP Server Builderanthropics/skills180k63 repos~2.3kAutomated safety check: PassApache-2.0
MCP Server BuildershareAI-lab/learn-claude-code78k4 repos~1.2kAutomated safety check: PassMIT
MCP Integration for Pluginsanthropics/claude-plugins-official38k11 repos~3.1kAutomated safety check: PassApache-2.0
Crush Configurationcharmbracelet/crush29k—~3.7kAutomated safety check: PassCustom licence
Context Mode Output Sandboxmksglu/context-mode26k—~4.1kAutomated safety check: PassCustom licence

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 63 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • MCP Server Builder

    shareAI-lab/learn-claude-code

    Walks through building MCP servers in Python or TypeScript that expose tools, resources and prompts to Claude, with templates, registration and testing.

    78k GitHub starsUsed in 4 repos~1.2k tokens
    Agent WorkflowsAuto-check passed
  • MCP Integration for Plugins

    anthropics/claude-plugins-official

    Official

    Explains how to bundle Model Context Protocol servers in a Claude Code plugin, covering config files, stdio, SSE, HTTP and WebSocket server types, and authentication.

    38k GitHub starsUsed in 11 repos~3.1k tokens
    Agent WorkflowsAuto-check passed
  • Crush Configuration

    charmbracelet/crush

    Explains how to configure the Crush coding agent with crushrc or crush.json, covering providers, models, LSPs, MCP servers, hooks, permissions and config precedence.

    29k GitHub stars~3.7k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Context Mode Output Sandbox

    mksglu/context-mode

    Routes large command, file, API and browser output through context-mode tools so only the needed result enters the agent's context, instead of dumping it via Bash.

    26k GitHub stars~4.1k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Migrates the compatible subset of settings and global file-based MCP servers from the Warp desktop app into Warp Agent CLI without exposing credentials or state.

    65k GitHub starsUsed in 1 repo~2.1k tokens
    Agent WorkflowsAuto-check passed

More from LeoYeAI/openclaw-master-skills

All 1,200 skills in this repo
  • DevOps Pipeline Management

    LeoYeAI/openclaw-master-skills

    Manages pipelines on a DevOps quality and efficiency platform through its OpenAPI: list workspaces and templates, create, update, run and cancel pipelines, and read run records.

    2.2k GitHub stars~4.2k tokensUpdated 2 mo ago
    Auto-check: notes
  • Feishu Document Collaboration

    LeoYeAI/openclaw-master-skills

    Patches OpenClaw's Feishu extension so an edited document triggers an isolated agent session that reads the doc and replies inline, turning it into a live chat space.

    2.2k GitHub stars~2k tokensUpdated 2 mo ago
    Auto-check passed
  • Files Memory System

    LeoYeAI/openclaw-master-skills

    Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.

    2.2k GitHub stars~3.8k tokensUpdated 2 mo ago
    Auto-check passed
  • GEO-Claw AI Visibility Agent

    LeoYeAI/openclaw-master-skills

    Runs a brand's AI-search visibility work end to end: diagnosing how AI platforms represent it, repositioning it, producing AI-optimized content and monitoring ongoing mentions.

    2.2k GitHub stars~4.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Google Workspace CLI

    LeoYeAI/openclaw-master-skills

    Installs and authenticates the gws CLI, then automates Gmail, Drive, Sheets, Calendar, Docs, Chat and Tasks with ready-made recipes, persona bundles and security audits.

    2.2k GitHub stars~2.6k tokensUpdated 2 mo ago
    Auto-check: notes
  • HealthFit Health Advisors

    LeoYeAI/openclaw-master-skills

    Runs four advisor roles, a fitness coach, nutritionist, data analyst and TCM practitioner, to build a health profile and track workouts, diet and wellness over time.

    2.2k GitHub stars~4.4k tokensUpdated 2 mo ago
    Auto-check passed

Categories

Questions about Taobao MCP Benchmark

What does Taobao MCP Benchmark do?

淘宝桌面版MCP工具评测框架。用于系统化测试MCP工具的各项功能,生成专业的技术评测报告。Use when 需要对淘宝MCP工具进行评测、测试、验收、迭代验证。. Taobao MCP Benchmark is an agent skill from LeoYeAI/openclaw-master-skills.

When should I use Taobao MCP Benchmark?

Taobao MCP Benchmark fits situations like: 需要对淘宝MCP工具进行评测、测试、验收、迭代验证; tasks that involve MCP servers.

How do I install Taobao MCP Benchmark in Claude Code?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill taobao-mcp-benchmark -a claude-code`. Or copy the skill folder (skills/taobao-mcp-benchmark in LeoYeAI/openclaw-master-skills) into .claude/skills/taobao-mcp-benchmark in your project. Claude Code loads it when a task matches its description.

How do I install Taobao MCP Benchmark in Codex?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill taobao-mcp-benchmark -a codex`. Or copy the skill folder (skills/taobao-mcp-benchmark in LeoYeAI/openclaw-master-skills) into .agents/skills/taobao-mcp-benchmark in your project. Codex loads it when a task matches its description.

Can I use Taobao MCP Benchmark in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LeoYeAI/openclaw-master-skills --skill taobao-mcp-benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/taobao-mcp-benchmark, .gemini/skills/taobao-mcp-benchmark, .github/skills/taobao-mcp-benchmark and .opencode/skills/taobao-mcp-benchmark in your project.

What does Taobao MCP Benchmark need to run?

Going by SKILL.md and its folder, Taobao MCP Benchmark needs a shell and JavaScript for the scripts in its folder. Our summary lists: Node.js; A Bash shell.

Does Taobao MCP Benchmark access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Taobao MCP Benchmark safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Taobao MCP Benchmark use?

Taobao MCP Benchmark is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Taobao MCP Benchmark use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Taobao MCP Benchmark?

Skills that share tags, products or a category with Taobao MCP Benchmark: MCP Server Builder (anthropics/skills, 180k stars), MCP Server Builder (shareAI-lab/learn-claude-code, 78k stars), MCP Integration for Plugins (anthropics/claude-plugins-official, 38k stars) and Crush Configuration (charmbracelet/crush, 29k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Taobao MCP Benchmark?

LeoYeAI (a GitHub user) maintains it in LeoYeAI/openclaw-master-skills, which has 2,161 GitHub stars. The repository holds 1,235 skills in this directory. The repository was last updated on July 20, 2026.

Source: LeoYeAI/openclaw-master-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.