Agent skill

Regex Vs LLM Structured Text

by affaan-m in affaan-m/ECC

选择在解析结构化文本时使用正则表达式还是大型语言模型的决策框架——从正则表达式开始,仅在低置信度的边缘情况下添加大型语言模型。

MITAuto-check passed

Install Regex Vs LLM Structured Text

skills CLI
$ npx skills add affaan-m/ECC --skill regex-vs-llm-structured-text -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install affaan-m/ECC regex-vs-llm-structured-text --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/docs/zh-CN/skills/regex-vs-llm-structured-text .claude/skills/regex-vs-llm-structured-text && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
regex-vs-llm-structured-text
GitHub stars
277k
Used in
3 other repos
Token cost
~1.2k tokens
SKILL.md length
86 words
Files
1
Skills in repo
683
Repo updated
First seen
Licence
MIT

At a glance

选择在解析结构化文本时使用正则表达式还是大型语言模型的决策框架——从正则表达式开始,仅在低置信度的边缘情况下添加大型语言模型。

  • Works in 4 steps: 正则表达式解析器(处理大多数情况) → 置信度评分 → LLM 验证器(仅用于边缘情况) → …
  • SKILL.md covers 何时使用, 决策框架, 架构模式 and 实现, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Regex Vs LLM Structured Text is an agent skill from affaan-m/ECC. 选择在解析结构化文本时使用正则表达式还是大型语言模型的决策框架——从正则表达式开始,仅在低置信度的边缘情况下添加大型语言模型。

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. The licence is MIT.

Example prompts

  • “/regex-vs-llm-structured-text”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. 正则表达式解析器(处理大多数情况)
  2. 置信度评分
  3. LLM 验证器(仅用于边缘情况)
  4. 混合管道

What it can do on your machine

Read from SKILL.md and the folder at commit 2d515e4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Regex Vs LLM Structured Text loads about 1.2k tokens when it runs. Until then it costs about 23 tokens; SKILL.md has 86 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~23
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from affaan-m/ECC at commit 2d515e4, republished under its MIT licence (© affaan-m). 86 words, ~1,212 tokens.

Download SKILL.mdSave it as .claude/skills/regex-vs-llm-structured-text/SKILL.md (or your agent's skills folder).
name
regex-vs-llm-structured-text
description
选择在解析结构化文本时使用正则表达式还是大型语言模型的决策框架——从正则表达式开始,仅在低置信度的边缘情况下添加大型语言模型。
origin
ECC

正则表达式 vs LLM 用于结构化文本解析

一个用于解析结构化文本(测验、表单、发票、文档)的实用决策框架。核心见解是:正则表达式能以低成本、确定性的方式处理 95-98% 的情况。将昂贵的 LLM 调用留给剩余的边缘情况。

何时使用

  • 解析具有重复模式的结构化文本(问题、表单、表格)
  • 决定在文本提取时使用正则表达式还是 LLM
  • 构建结合两种方法的混合管道
  • 在文本处理中优化成本/准确性权衡

决策框架

文本格式是否一致且重复?
├── 是 (>90% 遵循某种模式) → 从正则表达式开始
│   ├── 正则表达式处理 95%+ → 完成,无需 LLM
│   └── 正则表达式处理 <95% → 仅为边缘情况添加 LLM
└── 否 (自由格式,高度可变) → 直接使用 LLM

架构模式

[正则表达式解析器] ─── 提取结构(95-98% 准确率)
    │
    ▼
[文本清理器] ─── 去除噪声(标记、页码、伪影)
    │
    ▼
[置信度评分器] ─── 标记低置信度提取项
    │
    ├── 高置信度(≥0.95)→ 直接输出
    │
    └── 低置信度(<0.95)→ [LLM 验证器] → 输出

实现

1. 正则表达式解析器(处理大多数情况)
python
import re
from dataclasses import dataclass

@dataclass(frozen=True)
class ParsedItem:
    id: str
    text: str
    choices: tuple[str, ...]
    answer: str
    confidence: float = 1.0

def parse_structured_text(content: str) -> list[ParsedItem]:
    """Parse structured text using regex patterns."""
    pattern = re.compile(
        r"(?P<id>\d+)\.\s*(?P<text>.+?)\n"
        r"(?P<choices>(?:[A-D]\..+?\n)+)"
        r"Answer:\s*(?P<answer>[A-D])",
        re.MULTILINE | re.DOTALL,
    )
    items = []
    for match in pattern.finditer(content):
        choices = tuple(
            c.strip() for c in re.findall(r"[A-D]\.\s*(.+)", match.group("choices"))
        )
        items.append(ParsedItem(
            id=match.group("id"),
            text=match.group("text").strip(),
            choices=choices,
            answer=match.group("answer"),
        ))
    return items
2. 置信度评分

标记可能需要 LLM 审核的项:

python
@dataclass(frozen=True)
class ConfidenceFlag:
    item_id: str
    score: float
    reasons: tuple[str, ...]

def score_confidence(item: ParsedItem) -> ConfidenceFlag:
    """Score extraction confidence and flag issues."""
    reasons = []
    score = 1.0

    if len(item.choices) < 3:
        reasons.append("few_choices")
        score -= 0.3

    if not item.answer:
        reasons.append("missing_answer")
        score -= 0.5

    if len(item.text) < 10:
        reasons.append("short_text")
        score -= 0.2

    return ConfidenceFlag(
        item_id=item.id,
        score=max(0.0, score),
        reasons=tuple(reasons),
    )

def identify_low_confidence(
    items: list[ParsedItem],
    threshold: float = 0.95,
) -> list[ConfidenceFlag]:
    """Return items below confidence threshold."""
    flags = [score_confidence(item) for item in items]
    return [f for f in flags if f.score < threshold]
3. LLM 验证器(仅用于边缘情况)
python
def validate_with_llm(
    item: ParsedItem,
    original_text: str,
    client,
) -> ParsedItem:
    """Use LLM to fix low-confidence extractions."""
    response = client.messages.create(
        model="claude-haiku-4-5-20251001",  # Cheapest model for validation
        max_tokens=500,
        messages=[{
            "role": "user",
            "content": (
                f"Extract the question, choices, and answer from this text.\n\n"
                f"Text: {original_text}\n\n"
                f"Current extraction: {item}\n\n"
                f"Return corrected JSON if needed, or 'CORRECT' if accurate."
            ),
        }],
    )
    # Parse LLM response and return corrected item...
    return corrected_item
4. 混合管道
python
def process_document(
    content: str,
    *,
    llm_client=None,
    confidence_threshold: float = 0.95,
) -> list[ParsedItem]:
    """Full pipeline: regex -> confidence check -> LLM for edge cases."""
    # Step 1: Regex extraction (handles 95-98%)
    items = parse_structured_text(content)

    # Step 2: Confidence scoring
    low_confidence = identify_low_confidence(items, confidence_threshold)

    if not low_confidence or llm_client is None:
        return items

    # Step 3: LLM validation (only for flagged items)
    low_conf_ids = {f.item_id for f in low_confidence}
    result = []
    for item in items:
        if item.id in low_conf_ids:
            result.append(validate_with_llm(item, content, llm_client))
        else:
            result.append(item)

    return result

实际指标

来自一个生产中的测验解析管道(410 个项目):

指标值
正则表达式成功率98.0%
低置信度项目8 (2.0%)
所需 LLM 调用次数~5
相比全 LLM 的成本节省~95%
测试覆盖率93%

最佳实践

  • 从正则表达式开始 — 即使不完美的正则表达式也能提供一个改进的基线
  • 使用置信度评分 来以编程方式识别需要 LLM 帮助的内容
  • 使用最便宜的 LLM 进行验证(Haiku 类模型已足够)
  • 切勿修改 已解析的项 — 从清理/验证步骤返回新实例
  • TDD 效果很好 用于解析器 — 首先为已知模式编写测试,然后是边缘情况
  • 记录指标(正则表达式成功率、LLM 调用次数)以跟踪管道健康状况

应避免的反模式

  • 当正则表达式能处理 95% 以上的情况时,将所有文本发送给 LLM(昂贵且缓慢)
  • 对自由格式、高度可变的文本使用正则表达式(LLM 在此处更合适)
  • 跳过置信度评分,希望正则表达式“能正常工作”
  • 在清理/验证步骤中修改已解析的对象
  • 不测试边缘情况(格式错误的输入、缺失字段、编码问题)

适用场景

  • 测验/考试题目解析
  • 表单数据提取
  • 发票/收据处理
  • 文档结构解析(标题、章节、表格)
  • 任何具有重复模式且成本重要的结构化文本

© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in docs/zh-CN/skills/regex-vs-llm-structured-text of affaan-m/ECC.

Open the folder on GitHubat commit 2d515e4

Used in 3 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in affaan-m/ECC, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Regex Vs LLM Structured Text next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Regex Vs LLM Structured Text compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Regex Vs LLM Structured Text this skillaffaan-m/ECC277k3 repos~1.2kAutomated safety check: PassMIT
Regex ExpertRightNow-AI/openfang18k—~792Automated safety check: PassApache-2.0
Regex Buildermergisi/awesome-openclaw-agents4k—~256Automated safety check: PassMIT
Phy Regex AuditLeoYeAI/openclaw-master-skills2.2k—~5.1kAutomated safety check: PassApache-2.0
Regex Buildermohitagw15856/pm-claude-skills1.4k—~704Automated safety check: PassMIT
Regex DebuggerOneWave-AI/claude-skills336—~1.3kAutomated safety check: PassMIT

Similar skills

  • Regex Expert

    RightNow-AI/openfang

    Regular expression expert for crafting, debugging, and explaining patterns

    18k GitHub stars~792 tokensUpdated 3 mo ago
    DevelopmentAuto-check passed
  • Regex Builder

    mergisi/awesome-openclaw-agents

    Describe a text pattern in plain English and get a working regular expression with an explanation.

    4k GitHub stars~256 tokensUpdated 14 days ago
    Writing & ContentAuto-check passed
  • Phy Regex Audit

    LeoYeAI/openclaw-master-skills

    Static ReDoS (Regular Expression Denial of Service) vulnerability scanner and regex quality auditor for codebases.

    2.2k GitHub stars~5.1k tokensUpdated 2 mo ago
    SecurityAuto-check passed
  • Regex Builder

    mohitagw15856/pm-claude-skills

    Build a regular expression from a plain-English description, or explain an existing one.

    1.4k GitHub stars~704 tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Regex Debugger

    OneWave-AI/claude-skills

    Debug regex patterns with visual breakdowns, plain English explanations, test case generation, and flavor conversion.

    336 GitHub stars~1.3k tokensUpdated 9 days ago
    Testing & QAAuto-check passed
  • Regex Builder

    Mathews-Tom/armory

    DEPRECATED: The base model generates, explains, and tests regex patterns natively with high accuracy.

    329 GitHub stars~2k tokensUpdated 5 days ago
    Auto-check passed

More from affaan-m/ECC

All 682 skills in this repo
  • Skill Stocktake

    affaan-m/ECC

    Audits your installed Claude skills and commands for quality, with a quick mode for recently changed skills and a full mode that evaluates all of them through subagents.

    277k GitHub starsUsed in 5 repos~3.1k tokens
    Auto-check passed
  • Ingests, indexes, searches, edits and monitors video, audio and live streams through the VideoDB Python SDK, returning stream links, clips and timestamps.

    277k GitHub starsUsed in 3 repos~3.5k tokens
    Auto-check: notes
  • Docs Governance

    affaan-m/ECC

    Route broad documentation-governance requests to existing ECC skills and run an opt-in, read-only audit of mapped documentation roles, links, ADR indexes, and evidence references.

    277k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Rules Distillation

    affaan-m/ECC

    Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.

    277k GitHub starsUsed in 2 repos~2.3k tokens
    Auto-check passed
  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    277k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Set an ECC-specific frontend design direction for production UI work.

    277k GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed

Questions about Regex Vs LLM Structured Text

What does Regex Vs LLM Structured Text do?

选择在解析结构化文本时使用正则表达式还是大型语言模型的决策框架——从正则表达式开始,仅在低置信度的边缘情况下添加大型语言模型。. Regex Vs LLM Structured Text is an agent skill from affaan-m/ECC.

How do I install Regex Vs LLM Structured Text in Claude Code?

Run `npx skills add affaan-m/ECC --skill regex-vs-llm-structured-text -a claude-code`. Or copy the skill folder (docs/zh-CN/skills/regex-vs-llm-structured-text in affaan-m/ECC) into .claude/skills/regex-vs-llm-structured-text in your project. Claude Code loads it when a task matches its description.

How do I install Regex Vs LLM Structured Text in Codex?

Run `npx skills add affaan-m/ECC --skill regex-vs-llm-structured-text -a codex`. Or copy the skill folder (docs/zh-CN/skills/regex-vs-llm-structured-text in affaan-m/ECC) into .agents/skills/regex-vs-llm-structured-text in your project. Codex loads it when a task matches its description.

Can I use Regex Vs LLM Structured Text in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill regex-vs-llm-structured-text -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/regex-vs-llm-structured-text, .gemini/skills/regex-vs-llm-structured-text, .github/skills/regex-vs-llm-structured-text and .opencode/skills/regex-vs-llm-structured-text in your project.

What does Regex Vs LLM Structured Text need to run?

SKILL.md names no scripts, command-line tools or credentials: Regex Vs LLM Structured Text is instructions for the agent only. Our summary lists: Python 3.

Does Regex Vs LLM Structured Text access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Regex Vs LLM Structured Text safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Regex Vs LLM Structured Text use?

Regex Vs LLM Structured Text is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Regex Vs LLM Structured Text use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Regex Vs LLM Structured Text?

Skills that share tags, products or a category with Regex Vs LLM Structured Text: Regex Expert (RightNow-AI/openfang, 18k stars), Regex Builder (mergisi/awesome-openclaw-agents, 4k stars), Phy Regex Audit (LeoYeAI/openclaw-master-skills, 2.2k stars) and Regex Builder (mohitagw15856/pm-claude-skills, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Regex Vs LLM Structured Text?

affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 276,673 GitHub stars. The repository holds 683 skills in this directory. The repository was last updated on October 11, 2026.

Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.