Agent skill

Regex Vs LLM Structured Text

by xu-xiang in xu-xiang/everything-claude-code-zh

在解析结构化文本时,用于在正则表达式(Regex)和大型语言模型(LLM)之间进行选择的决策框架——优先使用正则表达式,仅针对低置信度的边界情况引入 LLM。

MITAuto-check passed

Install Regex Vs LLM Structured Text

skills CLI
$ npx skills add xu-xiang/everything-claude-code-zh --skill regex-vs-llm-structured-text -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install xu-xiang/everything-claude-code-zh regex-vs-llm-structured-text --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/xu-xiang/everything-claude-code-zh.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/regex-vs-llm-structured-text .claude/skills/regex-vs-llm-structured-text && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
regex-vs-llm-structured-text
GitHub stars
2k
Token cost
~1.3k tokens
SKILL.md length
107 words
Files
1
Skills in repo
78
Repo updated
First seen
Licence
MIT

At a glance

在解析结构化文本时,用于在正则表达式(Regex)和大型语言模型(LLM)之间进行选择的决策框架——优先使用正则表达式,仅针对低置信度的边界情况引入 LLM。

  • Works in 4 steps: 正则解析器 (处理绝大多数情况) → 置信度评分 (Confidence Scoring) → LLM 验证器 (仅处理边界情况) → …
  • SKILL.md covers 何时激活(When to Activate), 决策框架(Decision Framework), 架构模式(Architecture Pattern) and 实现参考(Implementation), plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Regex Vs LLM Structured Text is an agent skill from xu-xiang/everything-claude-code-zh. 在解析结构化文本时,用于在正则表达式(Regex)和大型语言模型(LLM)之间进行选择的决策框架——优先使用正则表达式,仅针对低置信度的边界情况引入 LLM。

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: everything-claude-code 中文翻译项目:完整的 Claude Code 配置集合(agents, skills, hooks, commands, rules, MCPs)。源自 Anthropic 黑客松获胜者的实战配置,助力中文工程师高效理解与使用 Claude Code。 The licence is MIT.

Example prompts

  • “/regex-vs-llm-structured-text”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. 正则解析器 (处理绝大多数情况)
  2. 置信度评分 (Confidence Scoring)
  3. LLM 验证器 (仅处理边界情况)
  4. 混合流水线 (Hybrid Pipeline)

What it can do on your machine

Read from SKILL.md and the folder at commit dfbf946. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Regex Vs LLM Structured Text loads about 1.3k tokens when it runs. Until then it costs about 27 tokens; SKILL.md has 107 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~27
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from xu-xiang/everything-claude-code-zh at commit dfbf946, republished under its MIT licence (© xu-xiang). 107 words, ~1,276 tokens.

Download SKILL.mdSave it as .claude/skills/regex-vs-llm-structured-text/SKILL.md (or your agent's skills folder).
name
regex-vs-llm-structured-text
description
在解析结构化文本时,用于在正则表达式(Regex)和大型语言模型(LLM)之间进行选择的决策框架——优先使用正则表达式,仅针对低置信度的边界情况引入 LLM。
origin
ECC

结构化文本解析:正则表达式(Regex)与大语言模型(LLM)的抉择

这是一个用于解析结构化文本(测验、表单、发票、文档)的实用决策框架。核心见解是:正则表达式(Regex)能以极低的成本确定性地处理 95-98% 的情况。应将昂贵的 LLM 调用保留给剩余的边界情况(Edge Cases)。

何时激活(When to Activate)

  • 解析具有重复模式的结构化文本(问题、表单、表格等)
  • 在文本提取任务中权衡使用正则表达式还是 LLM
  • 构建结合这两种方法的混合流水线(Hybrid Pipelines)
  • 在文本处理中优化成本与准确性的平衡

决策框架(Decision Framework)

文本格式是否一致且重复?
├── 是 (>90% 遵循某种模式) → 从正则表达式(Regex)开始
│   ├── 正则表达式处理了 95%+ → 完成,无需 LLM
│   └── 正则表达式处理率 <95% → 仅针对边界情况添加 LLM
└── 否 (非格式化,高度多变) → 直接使用 LLM

架构模式(Architecture Pattern)

源文本(Source Text)
    │
    ▼
[正则解析器 (Regex Parser)] ─── 提取结构 (95-98% 准确率)
    │
    ▼
[文本清洗器 (Text Cleaner)] ─── 去除噪声 (标记、页码、人工痕迹)
    │
    ▼
[置信度评分器 (Confidence Scorer)] ─── 标记低置信度提取结果
    │
    ├── 高置信度 (≥0.95) → 直接输出
    │
    └── 低置信度 (<0.95) → [LLM 验证器 (LLM Validator)] → 输出

实现参考(Implementation)

1. 正则解析器 (处理绝大多数情况)
python
import re
from dataclasses import dataclass

@dataclass(frozen=True)
class ParsedItem:
    id: str
    text: str
    choices: tuple[str, ...]
    answer: str
    confidence: float = 1.0

def parse_structured_text(content: str) -> list[ParsedItem]:
    """使用正则表达式模式解析结构化文本。"""
    pattern = re.compile(
        r"(?P<id>\d+)\.\s*(?P<text>.+?)\n"
        r"(?P<choices>(?:[A-D]\..+?\n)+)"
        r"Answer:\s*(?P<answer>[A-D])",
        re.MULTILINE | re.DOTALL,
    )
    items = []
    for match in pattern.finditer(content):
        choices = tuple(
            c.strip() for c in re.findall(r"[A-D]\.\s*(.+)", match.group("choices"))
        )
        items.append(ParsedItem(
            id=match.group("id"),
            text=match.group("text").strip(),
            choices=choices,
            answer=match.group("answer"),
        ))
    return items
2. 置信度评分 (Confidence Scoring)

标记可能需要 LLM 审查的项目:

python
@dataclass(frozen=True)
class ConfidenceFlag:
    item_id: str
    score: float
    reasons: tuple[str, ...]

def score_confidence(item: ParsedItem) -> ConfidenceFlag:
    """对提取置信度进行评分并标记问题。"""
    reasons = []
    score = 1.0

    if len(item.choices) < 3:
        reasons.append("few_choices") # 选项过少
        score -= 0.3

    if not item.answer:
        reasons.append("missing_answer") # 缺失答案
        score -= 0.5

    if len(item.text) < 10:
        reasons.append("short_text") # 文本过短
        score -= 0.2

    return ConfidenceFlag(
        item_id=item.id,
        score=max(0.0, score),
        reasons=tuple(reasons),
    )

def identify_low_confidence(
    items: list[ParsedItem],
    threshold: float = 0.95,
) -> list[ConfidenceFlag]:
    """返回低于置信度阈值的项目。"""
    flags = [score_confidence(item) for item in items]
    return [f for f in flags if f.score < threshold]
3. LLM 验证器 (仅处理边界情况)
python
def validate_with_llm(
    item: ParsedItem,
    original_text: str,
    client,
) -> ParsedItem:
    """使用 LLM 修复低置信度的提取结果。"""
    response = client.messages.create(
        model="claude-haiku-4-5-20251001",  # 使用最便宜的模型进行验证
        max_tokens=500,
        messages=[{
            "role": "user",
            "content": (
                f"Extract the question, choices, and answer from this text.\n\n"
                f"Text: {original_text}\n\n"
                f"Current extraction: {item}\n\n"
                f"Return corrected JSON if needed, or 'CORRECT' if accurate."
            ),
        }],
    )
    # 解析 LLM 响应并返回修正后的项目...
    return corrected_item
4. 混合流水线 (Hybrid Pipeline)
python
def process_document(
    content: str,
    *,
    llm_client=None,
    confidence_threshold: float = 0.95,
) -> list[ParsedItem]:
    """完整流水线:正则提取 -> 置信度检查 -> 针对边界情况调用 LLM。"""
    # 步骤 1: 正则提取 (处理 95-98% 的情况)
    items = parse_structured_text(content)

    # 步骤 2: 置信度评分
    low_confidence = identify_low_confidence(items, confidence_threshold)

    if not low_confidence or llm_client is None:
        return items

    # 步骤 3: LLM 验证 (仅针对被标记的项目)
    low_conf_ids = {f.item_id for f in low_confidence}
    result = []
    for item in items:
        if item.id in low_conf_ids:
            result.append(validate_with_llm(item, content, llm_client))
        else:
            result.append(item)

    return result

真实世界指标 (Real-World Metrics)

来自一个生产环境的测验解析流水线(410 个项目):

指标 (Metric)数值 (Value)
正则表达式成功率98.0%
低置信度项目数8 (2.0%)
需要的 LLM 调用次数~5
相比全 LLM 方案节省的成本~95%
测试覆盖率93%

最佳实践 (Best Practices)

  • 从正则表达式开始 —— 即使是不完美的正则也能为你提供改进的基准。
  • 使用置信度评分 —— 以编程方式识别哪些内容需要 LLM 协助。
  • 使用最便宜的 LLM 进行验证 —— Haiku 级别的模型通常已经足够。
  • 切勿修改(Mutate)原始解析项 —— 在清洗/验证步骤中应返回新的实例。
  • 测试驱动开发(TDD)非常适用 —— 先为已知模式编写测试,然后再处理边界情况。
  • 记录指标 —— 记录正则表达式成功率、LLM 调用次数等,以跟踪流水线的健康状况。

应避免的反模式 (Anti-Patterns to Avoid)

  • 当正则表达式能处理 95% 以上情况时,仍将所有文本发送给 LLM(既昂贵又缓慢)。
  • 对非格式化、高度多变的文本使用正则表达式(在这种情况下 LLM 表现更好)。
  • 跳过置信度评分并寄希望于正则表达式“刚好能用”。
  • 在清洗/验证步骤中直接修改已解析的对象。
  • 不测试边界情况(格式错误的输入、缺失字段、编码问题等)。

适用场景 (When to Use)

  • 测验/考试题目解析
  • 表单数据提取
  • 发票/收据处理
  • 文档结构解析(标题、章节、表格)
  • 任何具有重复模式且对成本敏感的结构化文本处理

© xu-xiang, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/regex-vs-llm-structured-text of xu-xiang/everything-claude-code-zh.

Open the folder on GitHubat commit dfbf946

Compare with similar skills

Regex Vs LLM Structured Text next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Regex Vs LLM Structured Text compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Regex Vs LLM Structured Text this skillxu-xiang/everything-claude-code-zh2k—~1.3kAutomated safety check: PassMIT
Regex Vs LLM Structured Textaffaan-m/ECC277k5 repos~1.7kAutomated safety check: PassMIT
Regex Vs LLM Structured Textaffaan-m/ECC277k3 repos~1.2kAutomated safety check: PassMIT
Regex Vs LLM Structured Textaffaan-m/ECC276k—~1.3kAutomated safety check: PassMIT
Regex ExpertRightNow-AI/openfang18k—~792Automated safety check: PassApache-2.0
Regex Buildermergisi/awesome-openclaw-agents4k—~256Automated safety check: PassMIT

Similar skills

  • Decision framework for parsing structured text (quizzes, forms, invoices, receipts, tables) with a hybrid regex-first pipeline — regex extraction handles 95%+ cheaply, a confidence scorer flags…

    277k GitHub starsUsed in 5 repos~1.7k tokens
    EducationAuto-check passed
  • 选择在解析结构化文本时使用正则表达式还是大型语言模型的决策框架——从正则表达式开始,仅在低置信度的边缘情况下添加大型语言模型。

    277k GitHub starsUsed in 3 repos~1.2k tokens
    Auto-check passed
  • 構造化テキストの解析に正規表現と大規模言語モデルのどちらを使うかを選択するための意思決定フレームワーク——まず正規表達式から始め、信頼度の低いエッジケースにのみ大規模言語モデルを追加する。

    276k GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Regex Expert

    RightNow-AI/openfang

    Regular expression expert for crafting, debugging, and explaining patterns

    18k GitHub stars~792 tokensUpdated 3 mo ago
    DevelopmentAuto-check passed
  • Regex Builder

    mergisi/awesome-openclaw-agents

    Describe a text pattern in plain English and get a working regular expression with an explanation.

    4k GitHub stars~256 tokensUpdated 14 days ago
    Writing & ContentAuto-check passed
  • Phy Regex Audit

    LeoYeAI/openclaw-master-skills

    Static ReDoS (Regular Expression Denial of Service) vulnerability scanner and regex quality auditor for codebases.

    2.2k GitHub stars~5.1k tokensUpdated 2 mo ago
    SecurityAuto-check passed

More from xu-xiang/everything-claude-code-zh

All 78 skills in this repo
  • Configure Ecc

    xu-xiang/everything-claude-code-zh

    Everything Claude Code 的交互式安装程序 — 引导用户选择并安装技能和规则到用户级或项目级目录,验证路径,并可选择优化已安装文件。

    2k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Continuous Learning V2

    xu-xiang/everything-claude-code-zh

    基于本能(Instinct)的学习系统,通过钩子(hooks)观察会话,创建带有置信度评分的原子本能,并将其演化为技能(Skills)、命令(Commands)或智能体(Agents)。v2.1 版本增加了项目作用域(project-scoped)的本能,以防止跨项目污染。

    2k GitHub stars~2.1k tokensUpdated 7 mo ago
    Auto-check passed
  • API Design

    xu-xiang/everything-claude-code-zh

    生产级 API 的 REST API 设计模式,包括资源命名、状态码、分页、过滤、错误响应、版本控制和速率限制. An agent skill from xu-xiang/everything-claude-code-zh.

    2k GitHub stars~2.7k tokensUpdated 7 mo ago
    Auto-check passed
  • Backend Patterns

    xu-xiang/everything-claude-code-zh

    后端架构模式、API 设计、数据库优化以及适用于 Node.js、Express 和 Next.js API 路由的服务端最佳实践。

    2k GitHub stars~3.2k tokensUpdated 7 mo ago
    Auto-check passed
  • Backend Patterns

    xu-xiang/everything-claude-code-zh

    后端架构模式、API 设计、数据库优化以及 Node.js、Express 和 Next.js API 路由的服务端最佳实践。

    2k GitHub stars~3.1k tokensUpdated 7 mo ago
    Auto-check passed
  • Backend Patterns

    xu-xiang/everything-claude-code-zh

    后端架构模式、API 设计、数据库优化以及针对 Node.js、Express 和 Next.js API 路由的服务端最佳实践。

    2k GitHub stars~3.2k tokensUpdated 7 mo ago
    Auto-check passed

Questions about Regex Vs LLM Structured Text

What does Regex Vs LLM Structured Text do?

在解析结构化文本时,用于在正则表达式(Regex)和大型语言模型(LLM)之间进行选择的决策框架——优先使用正则表达式,仅针对低置信度的边界情况引入 LLM。. Regex Vs LLM Structured Text is an agent skill from xu-xiang/everything-claude-code-zh.

How do I install Regex Vs LLM Structured Text in Claude Code?

Run `npx skills add xu-xiang/everything-claude-code-zh --skill regex-vs-llm-structured-text -a claude-code`. Or copy the skill folder (skills/regex-vs-llm-structured-text in xu-xiang/everything-claude-code-zh) into .claude/skills/regex-vs-llm-structured-text in your project. Claude Code loads it when a task matches its description.

How do I install Regex Vs LLM Structured Text in Codex?

Run `npx skills add xu-xiang/everything-claude-code-zh --skill regex-vs-llm-structured-text -a codex`. Or copy the skill folder (skills/regex-vs-llm-structured-text in xu-xiang/everything-claude-code-zh) into .agents/skills/regex-vs-llm-structured-text in your project. Codex loads it when a task matches its description.

Can I use Regex Vs LLM Structured Text in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add xu-xiang/everything-claude-code-zh --skill regex-vs-llm-structured-text -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/regex-vs-llm-structured-text, .gemini/skills/regex-vs-llm-structured-text, .github/skills/regex-vs-llm-structured-text and .opencode/skills/regex-vs-llm-structured-text in your project.

What does Regex Vs LLM Structured Text need to run?

SKILL.md names no scripts, command-line tools or credentials: Regex Vs LLM Structured Text is instructions for the agent only. Our summary lists: Python 3.

Does Regex Vs LLM Structured Text access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Regex Vs LLM Structured Text safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Regex Vs LLM Structured Text use?

Regex Vs LLM Structured Text is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Regex Vs LLM Structured Text use?

About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Regex Vs LLM Structured Text?

Skills that share tags, products or a category with Regex Vs LLM Structured Text: Regex Vs LLM Structured Text (affaan-m/ECC, 277k stars), Regex Vs LLM Structured Text (affaan-m/ECC, 277k stars), Regex Vs LLM Structured Text (affaan-m/ECC, 276k stars) and Regex Expert (RightNow-AI/openfang, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Regex Vs LLM Structured Text?

xu-xiang (a GitHub user) maintains it in xu-xiang/everything-claude-code-zh, which has 1,978 GitHub stars. The repository holds 78 skills in this directory. The repository was last updated on March 5, 2026.

Source: xu-xiang/everything-claude-code-zh on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.