Agent skill

Regex Vs LLM Structured Text

by affaan-m in affaan-m/ECC

構造化テキストの解析に正規表現と大規模言語モデルのどちらを使うかを選択するための意思決定フレームワーク——まず正規表達式から始め、信頼度の低いエッジケースにのみ大規模言語モデルを追加する。

MITAuto-check passed

Install Regex Vs LLM Structured Text

skills CLI
$ npx skills add affaan-m/ECC --skill regex-vs-llm-structured-text -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install affaan-m/ECC regex-vs-llm-structured-text --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/docs/ja-JP/skills/regex-vs-llm-structured-text .claude/skills/regex-vs-llm-structured-text && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
regex-vs-llm-structured-text
GitHub stars
276k
Token cost
~1.3k tokens
SKILL.md length
59 words
Files
1
Skills in repo
683
Repo updated
First seen
Licence
MIT

At a glance

構造化テキストの解析に正規表現と大規模言語モデルのどちらを使うかを選択するための意思決定フレームワーク——まず正規表達式から始め、信頼度の低いエッジケースにのみ大規模言語モデルを追加する。

  • Works in 4 steps: 正規表現パーサー(大半のケースを処理) → 信頼度スコアリング → LLM バリデーター(エッジケースのみ) → …
  • SKILL.md covers 使用場面, 意思決定フレームワーク, アーキテクチャパターン and 実装, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Regex Vs LLM Structured Text is an agent skill from affaan-m/ECC. 構造化テキストの解析に正規表現と大規模言語モデルのどちらを使うかを選択するための意思決定フレームワーク——まず正規表達式から始め、信頼度の低いエッジケースにのみ大規模言語モデルを追加する。

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. The licence is MIT.

Example prompts

  • “/regex-vs-llm-structured-text”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. 正規表現パーサー(大半のケースを処理)
  2. 信頼度スコアリング
  3. LLM バリデーター(エッジケースのみ)
  4. ハイブリッドパイプライン

What it can do on your machine

Read from SKILL.md and the folder at commit 4eb71d9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Regex Vs LLM Structured Text loads about 1.3k tokens when it runs. Until then it costs about 31 tokens; SKILL.md has 59 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~31
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from affaan-m/ECC at commit 4eb71d9, republished under its MIT licence (© affaan-m). 59 words, ~1,301 tokens.

Download SKILL.mdSave it as .claude/skills/regex-vs-llm-structured-text/SKILL.md (or your agent's skills folder).
name
regex-vs-llm-structured-text
description
構造化テキストの解析に正規表現と大規模言語モデルのどちらを使うかを選択するための意思決定フレームワーク——まず正規表達式から始め、信頼度の低いエッジケースにのみ大規模言語モデルを追加する。
origin
ECC

構造化テキスト解析における正規表現 vs LLM

構造化テキスト(クイズ、フォーム、請求書、ドキュメント)を解析するための実用的な意思決定フレームワーク。核心的な洞察:正規表現は低コストかつ決定論的に95〜98%のケースを処理できる。コストのかかるLLM呼び出しは残りのエッジケースに留める。

使用場面

  • 繰り返しパターンを持つ構造化テキスト(設問、フォーム、表)の解析
  • テキスト抽出に正規表現とLLMのどちらを使うかの判断
  • 両方のアプローチを組み合わせたハイブリッドパイプラインの構築
  • テキスト処理におけるコスト/精度のトレードオフの最適化

意思決定フレームワーク

テキスト形式は一貫していて繰り返しがあるか?
├── はい (>90% が何らかのパターンに従う) → 正規表現から始める
│   ├── 正規表現が 95%+ を処理 → 完了、LLM は不要
│   └── 正規表現が <95% を処理 → エッジケースのみ LLM を追加
└── いいえ (自由形式、高度に可変) → LLM を直接使用

アーキテクチャパターン

[正規表現パーサー] ─── 構造を抽出(95〜98% の精度)
    │
    ▼
[テキストクリーナー] ─── ノイズを除去(マーカー、ページ番号、アーティファクト)
    │
    ▼
[信頼度スコアラー] ─── 信頼度の低い抽出結果にフラグを立てる
    │
    ├── 高信頼度(≥0.95)→ 直接出力
    │
    └── 低信頼度(<0.95)→ [LLM バリデーター] → 出力

実装

1. 正規表現パーサー(大半のケースを処理)
python
import re
from dataclasses import dataclass

@dataclass(frozen=True)
class ParsedItem:
    id: str
    text: str
    choices: tuple[str, ...]
    answer: str
    confidence: float = 1.0

def parse_structured_text(content: str) -> list[ParsedItem]:
    """Parse structured text using regex patterns."""
    pattern = re.compile(
        r"(?P<id>\d+)\.\s*(?P<text>.+?)\n"
        r"(?P<choices>(?:[A-D]\..+?\n)+)"
        r"Answer:\s*(?P<answer>[A-D])",
        re.MULTILINE | re.DOTALL,
    )
    items = []
    for match in pattern.finditer(content):
        choices = tuple(
            c.strip() for c in re.findall(r"[A-D]\.\s*(.+)", match.group("choices"))
        )
        items.append(ParsedItem(
            id=match.group("id"),
            text=match.group("text").strip(),
            choices=choices,
            answer=match.group("answer"),
        ))
    return items
2. 信頼度スコアリング

LLMによるレビューが必要かもしれない項目にフラグを立てる:

python
@dataclass(frozen=True)
class ConfidenceFlag:
    item_id: str
    score: float
    reasons: tuple[str, ...]

def score_confidence(item: ParsedItem) -> ConfidenceFlag:
    """Score extraction confidence and flag issues."""
    reasons = []
    score = 1.0

    if len(item.choices) < 3:
        reasons.append("few_choices")
        score -= 0.3

    if not item.answer:
        reasons.append("missing_answer")
        score -= 0.5

    if len(item.text) < 10:
        reasons.append("short_text")
        score -= 0.2

    return ConfidenceFlag(
        item_id=item.id,
        score=max(0.0, score),
        reasons=tuple(reasons),
    )

def identify_low_confidence(
    items: list[ParsedItem],
    threshold: float = 0.95,
) -> list[ConfidenceFlag]:
    """Return items below confidence threshold."""
    flags = [score_confidence(item) for item in items]
    return [f for f in flags if f.score < threshold]
3. LLM バリデーター(エッジケースのみ)
python
def validate_with_llm(
    item: ParsedItem,
    original_text: str,
    client,
) -> ParsedItem:
    """Use LLM to fix low-confidence extractions."""
    response = client.messages.create(
        model="claude-haiku-4-5-20251001",  # Cheapest model for validation
        max_tokens=500,
        messages=[{
            "role": "user",
            "content": (
                f"Extract the question, choices, and answer from this text.\n\n"
                f"Text: {original_text}\n\n"
                f"Current extraction: {item}\n\n"
                f"Return corrected JSON if needed, or 'CORRECT' if accurate."
            ),
        }],
    )
    # Parse LLM response and return corrected item...
    return corrected_item
4. ハイブリッドパイプライン
python
def process_document(
    content: str,
    *,
    llm_client=None,
    confidence_threshold: float = 0.95,
) -> list[ParsedItem]:
    """Full pipeline: regex -> confidence check -> LLM for edge cases."""
    # Step 1: Regex extraction (handles 95-98%)
    items = parse_structured_text(content)

    # Step 2: Confidence scoring
    low_confidence = identify_low_confidence(items, confidence_threshold)

    if not low_confidence or llm_client is None:
        return items

    # Step 3: LLM validation (only for flagged items)
    low_conf_ids = {f.item_id for f in low_confidence}
    result = []
    for item in items:
        if item.id in low_conf_ids:
            result.append(validate_with_llm(item, content, llm_client))
        else:
            result.append(item)

    return result

実際のメトリクス

本番のクイズ解析パイプライン(410項目)より:

メトリクス値
正規表現の成功率98.0%
低信頼度項目8 (2.0%)
必要なLLM呼び出し回数~5
全件LLM比のコスト節約~95%
テストカバレッジ93%

ベストプラクティス

  • 正規表現から始める — 不完全な正規表現でも改善のベースラインになる
  • 信頼度スコアリングを使用して、LLMの助けが必要なものをプログラムで特定する
  • 最も安価なLLMを使用して検証する(Haikuクラスのモデルで十分)
  • 解析済み項目を変更しない — クリーニング/検証ステップから新しいインスタンスを返す
  • TDDは解析器に効果的 — まず既知のパターンのテストを書き、次にエッジケースを書く
  • メトリクスを記録(正規表現の成功率、LLM呼び出し回数)してパイプラインの健全性を追跡する

避けるべきアンチパターン

  • 正規表現が95%以上を処理できる場合に全テキストをLLMに送る(コスト高・低速)
  • 自由形式で高度に可変なテキストに正規表現を使用する(LLMの方が適切)
  • 信頼度スコアリングをスキップして正規表現が「うまくいく」ことを期待する
  • クリーニング/検証ステップで解析済みオブジェクトを変更する
  • エッジケースをテストしない(不正な入力、欠損フィールド、エンコーディング問題)

適用場面

  • クイズ/試験問題の解析
  • フォームデータの抽出
  • 請求書/レシートの処理
  • ドキュメント構造の解析(見出し、セクション、表)
  • 繰り返しパターンがあり、コストが重要なあらゆる構造化テキスト

© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in docs/ja-JP/skills/regex-vs-llm-structured-text of affaan-m/ECC.

Open the folder on GitHubat commit 4eb71d9

Compare with similar skills

Regex Vs LLM Structured Text next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Regex Vs LLM Structured Text compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Regex Vs LLM Structured Text this skillaffaan-m/ECC276k—~1.3kAutomated safety check: PassMIT
Regex ExpertRightNow-AI/openfang18k—~792Automated safety check: PassApache-2.0
Regex Buildermergisi/awesome-openclaw-agents4k—~256Automated safety check: PassMIT
Phy Regex AuditLeoYeAI/openclaw-master-skills2.2k—~5.1kAutomated safety check: PassApache-2.0
Regex Buildermohitagw15856/pm-claude-skills1.4k—~704Automated safety check: PassMIT
Regex DebuggerOneWave-AI/claude-skills336—~1.3kAutomated safety check: PassMIT

Similar skills

  • Regex Expert

    RightNow-AI/openfang

    Regular expression expert for crafting, debugging, and explaining patterns

    18k GitHub stars~792 tokensUpdated 3 mo ago
    DevelopmentAuto-check passed
  • Regex Builder

    mergisi/awesome-openclaw-agents

    Describe a text pattern in plain English and get a working regular expression with an explanation.

    4k GitHub stars~256 tokensUpdated 14 days ago
    Writing & ContentAuto-check passed
  • Phy Regex Audit

    LeoYeAI/openclaw-master-skills

    Static ReDoS (Regular Expression Denial of Service) vulnerability scanner and regex quality auditor for codebases.

    2.2k GitHub stars~5.1k tokensUpdated 2 mo ago
    SecurityAuto-check passed
  • Regex Builder

    mohitagw15856/pm-claude-skills

    Build a regular expression from a plain-English description, or explain an existing one.

    1.4k GitHub stars~704 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Regex Debugger

    OneWave-AI/claude-skills

    Debug regex patterns with visual breakdowns, plain English explanations, test case generation, and flavor conversion.

    336 GitHub stars~1.3k tokensUpdated 9 days ago
    Testing & QAAuto-check passed
  • Regex Builder

    Mathews-Tom/armory

    DEPRECATED: The base model generates, explains, and tests regex patterns natively with high accuracy.

    329 GitHub stars~2k tokensUpdated 5 days ago
    Auto-check passed

More from affaan-m/ECC

All 682 skills in this repo
  • Skill Stocktake

    affaan-m/ECC

    Audits your installed Claude skills and commands for quality, with a quick mode for recently changed skills and a full mode that evaluates all of them through subagents.

    277k GitHub starsUsed in 5 repos~3.1k tokens
    Auto-check passed
  • Ingests, indexes, searches, edits and monitors video, audio and live streams through the VideoDB Python SDK, returning stream links, clips and timestamps.

    277k GitHub starsUsed in 3 repos~3.5k tokens
    Auto-check: notes
  • Docs Governance

    affaan-m/ECC

    Route broad documentation-governance requests to existing ECC skills and run an opt-in, read-only audit of mapped documentation roles, links, ADR indexes, and evidence references.

    277k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Rules Distillation

    affaan-m/ECC

    Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.

    277k GitHub starsUsed in 2 repos~2.3k tokens
    Auto-check passed
  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    277k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Set an ECC-specific frontend design direction for production UI work.

    277k GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed

Questions about Regex Vs LLM Structured Text

What does Regex Vs LLM Structured Text do?

構造化テキストの解析に正規表現と大規模言語モデルのどちらを使うかを選択するための意思決定フレームワーク——まず正規表達式から始め、信頼度の低いエッジケースにのみ大規模言語モデルを追加する。. Regex Vs LLM Structured Text is an agent skill from affaan-m/ECC.

How do I install Regex Vs LLM Structured Text in Claude Code?

Run `npx skills add affaan-m/ECC --skill regex-vs-llm-structured-text -a claude-code`. Or copy the skill folder (docs/ja-JP/skills/regex-vs-llm-structured-text in affaan-m/ECC) into .claude/skills/regex-vs-llm-structured-text in your project. Claude Code loads it when a task matches its description.

How do I install Regex Vs LLM Structured Text in Codex?

Run `npx skills add affaan-m/ECC --skill regex-vs-llm-structured-text -a codex`. Or copy the skill folder (docs/ja-JP/skills/regex-vs-llm-structured-text in affaan-m/ECC) into .agents/skills/regex-vs-llm-structured-text in your project. Codex loads it when a task matches its description.

Can I use Regex Vs LLM Structured Text in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill regex-vs-llm-structured-text -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/regex-vs-llm-structured-text, .gemini/skills/regex-vs-llm-structured-text, .github/skills/regex-vs-llm-structured-text and .opencode/skills/regex-vs-llm-structured-text in your project.

What does Regex Vs LLM Structured Text need to run?

SKILL.md names no scripts, command-line tools or credentials: Regex Vs LLM Structured Text is instructions for the agent only. Our summary lists: Python 3.

Does Regex Vs LLM Structured Text access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Regex Vs LLM Structured Text safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Regex Vs LLM Structured Text use?

Regex Vs LLM Structured Text is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Regex Vs LLM Structured Text use?

About 1.3k tokens (SKILL.md is roughly 5.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Regex Vs LLM Structured Text?

Skills that share tags, products or a category with Regex Vs LLM Structured Text: Regex Expert (RightNow-AI/openfang, 18k stars), Regex Builder (mergisi/awesome-openclaw-agents, 4k stars), Phy Regex Audit (LeoYeAI/openclaw-master-skills, 2.2k stars) and Regex Builder (mohitagw15856/pm-claude-skills, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Regex Vs LLM Structured Text?

affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 276,111 GitHub stars. The repository holds 683 skills in this directory. The repository was last updated on October 10, 2026.

Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.