Agent skill

Regex Vs LLM Structured Text

by affaan-m in affaan-m/ECC

Decision framework for parsing structured text (quizzes, forms, invoices, receipts, tables) with a hybrid regex-first pipeline — regex extraction handles 95%+ cheaply, a confidence scorer flags…

MITAuto-check passedEducation

Install Regex Vs LLM Structured Text

skills CLI
$ npx skills add affaan-m/ECC --skill regex-vs-llm-structured-text -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install affaan-m/ECC regex-vs-llm-structured-text --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/regex-vs-llm-structured-text .claude/skills/regex-vs-llm-structured-text && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
regex-vs-llm-structured-text
GitHub stars
277k
Used in
5 other repos
Token cost
~1.7k tokens
SKILL.md length
284 words
Files
1
Skills in repo
683
Repo updated
First seen
Licence
MIT

At a glance

Decision framework for parsing structured text (quizzes, forms, invoices, receipts, tables) with a hybrid regex-first pipeline — regex extraction handles 95%+ cheaply, a confidence scorer flags…

  • Works in 4 steps: Regex Parser (Handles the Majority) → Confidence Scoring → LLM Validator (Edge Cases Only) → …
  • Choosing between regex and LLM for text extraction
  • SKILL.md covers When to Activate, Decision Framework, Architecture Pattern and Implementation, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Regex Vs LLM Structured Text is an agent skill from affaan-m/ECC. Decision framework for parsing structured text (quizzes, forms, invoices, receipts, tables) with a hybrid regex-first pipeline — regex extraction handles 95%+ cheaply, a confidence scorer flags low-confidence items, and an LLM validator fixes only the edge cases. Use when choosing between regex and LLM for text extraction, building a cheap document parser, or optimizing extraction cost and accuracy.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Education, covering Quizzes and assessments and Forms and invoices. The repository describes itself as: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. The licence is MIT.

When your agent uses it

  • Choosing between regex and LLM for text extraction
  • Building a cheap document parser
  • Optimizing extraction cost and accuracy

Example prompts

  • “/regex-vs-llm-structured-text”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Regex Parser (Handles the Majority)
  2. Confidence Scoring
  3. LLM Validator (Edge Cases Only)
  4. Hybrid Pipeline

What it can do on your machine

Read from SKILL.md and the folder at commit 2d515e4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Regex Vs LLM Structured Text loads about 1.7k tokens when it runs. Until then it costs about 108 tokens; SKILL.md has 284 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~108
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from affaan-m/ECC at commit 2d515e4, republished under its MIT licence (© affaan-m). 284 words, ~1,663 tokens.

Download SKILL.mdSave it as .claude/skills/regex-vs-llm-structured-text/SKILL.md (or your agent's skills folder).
name
regex-vs-llm-structured-text
description
Decision framework for parsing structured text (quizzes, forms, invoices, receipts, tables) with a hybrid regex-first pipeline — regex extraction handles 95%+ cheaply, a confidence scorer flags low-confidence items, and an LLM validator fixes only the edge cases. Use when choosing between regex and LLM for text extraction, building a cheap document parser, or optimizing extraction cost and accuracy.
metadata.origin
ECC

Regex vs LLM for Structured Text Parsing

A practical decision framework for parsing structured text (quizzes, forms, invoices, documents). The key insight: regex handles 95-98% of cases cheaply and deterministically. Reserve expensive LLM calls for the remaining edge cases.

When to Activate

  • Parsing structured text with repeating patterns (questions, forms, tables)
  • Deciding between regex and LLM for text extraction
  • Building hybrid pipelines that combine both approaches
  • Optimizing cost/accuracy tradeoffs in text processing

Decision Framework

Is the text format consistent and repeating?
├── Yes (>90% follows a pattern) → Start with Regex
│   ├── Regex handles 95%+ → Done, no LLM needed
│   └── Regex handles <95% → Add LLM for edge cases only
└── No (free-form, highly variable) → Use LLM directly

Architecture Pattern

Source Text
    │
    ▼
[Regex Parser] ─── Extracts structure (95-98% accuracy)
    │
    ▼
[Text Cleaner] ─── Removes noise (markers, page numbers, artifacts)
    │
    ▼
[Confidence Scorer] ─── Flags low-confidence extractions
    │
    ├── High confidence (≥0.95) → Direct output
    │
    └── Low confidence (<0.95) → [LLM Validator] → Output

Implementation

1. Regex Parser (Handles the Majority)
python
import re
from dataclasses import dataclass

@dataclass(frozen=True)
class ParsedItem:
    id: str
    text: str
    choices: tuple[str, ...]
    answer: str
    confidence: float = 1.0

def parse_structured_text(content: str) -> list[ParsedItem]:
    """Parse structured text using regex patterns."""
    pattern = re.compile(
        r"(?P<id>\d+)\.\s*(?P<text>.+?)\n"
        r"(?P<choices>(?:[A-D]\..+?\n)+)"
        r"Answer:\s*(?P<answer>[A-D])",
        re.MULTILINE | re.DOTALL,
    )
    items = []
    for match in pattern.finditer(content):
        choices = tuple(
            c.strip() for c in re.findall(r"[A-D]\.\s*(.+)", match.group("choices"))
        )
        items.append(ParsedItem(
            id=match.group("id"),
            text=match.group("text").strip(),
            choices=choices,
            answer=match.group("answer"),
        ))
    return items
2. Confidence Scoring

Flag items that may need LLM review:

python
@dataclass(frozen=True)
class ConfidenceFlag:
    item_id: str
    score: float
    reasons: tuple[str, ...]

def score_confidence(item: ParsedItem) -> ConfidenceFlag:
    """Score extraction confidence and flag issues."""
    reasons = []
    score = 1.0

    if len(item.choices) < 3:
        reasons.append("few_choices")
        score -= 0.3

    if not item.answer:
        reasons.append("missing_answer")
        score -= 0.5

    if len(item.text) < 10:
        reasons.append("short_text")
        score -= 0.2

    return ConfidenceFlag(
        item_id=item.id,
        score=max(0.0, score),
        reasons=tuple(reasons),
    )

def identify_low_confidence(
    items: list[ParsedItem],
    threshold: float = 0.95,
) -> list[ConfidenceFlag]:
    """Return items below confidence threshold."""
    flags = [score_confidence(item) for item in items]
    return [f for f in flags if f.score < threshold]
3. LLM Validator (Edge Cases Only)
python
def validate_with_llm(
    item: ParsedItem,
    original_text: str,
    client,
) -> ParsedItem:
    """Use LLM to fix low-confidence extractions."""
    response = client.messages.create(
        model="claude-haiku-4-5-20251001",  # Cheapest model for validation
        max_tokens=500,
        messages=[{
            "role": "user",
            "content": (
                f"Extract the question, choices, and answer from this text.\n\n"
                f"Text: {original_text}\n\n"
                f"Current extraction: {item}\n\n"
                f"Return corrected JSON if needed, or 'CORRECT' if accurate."
            ),
        }],
    )
    # Parse LLM response and return corrected item...
    return corrected_item
4. Hybrid Pipeline
python
def process_document(
    content: str,
    *,
    llm_client=None,
    confidence_threshold: float = 0.95,
) -> list[ParsedItem]:
    """Full pipeline: regex -> confidence check -> LLM for edge cases."""
    # Step 1: Regex extraction (handles 95-98%)
    items = parse_structured_text(content)

    # Step 2: Confidence scoring
    low_confidence = identify_low_confidence(items, confidence_threshold)

    if not low_confidence or llm_client is None:
        return items

    # Step 3: LLM validation (only for flagged items)
    low_conf_ids = {f.item_id for f in low_confidence}
    result = []
    for item in items:
        if item.id in low_conf_ids:
            result.append(validate_with_llm(item, content, llm_client))
        else:
            result.append(item)

    return result

Real-World Metrics

From a production quiz parsing pipeline (410 items):

MetricValue
Regex success rate98.0%
Low confidence items8 (2.0%)
LLM calls needed~5
Cost savings vs all-LLM~95%
Test coverage93%

Best Practices

  • Start with regex — even imperfect regex gives you a baseline to improve
  • Use confidence scoring to programmatically identify what needs LLM help
  • Use the cheapest LLM for validation (Haiku-class models are sufficient)
  • Never mutate parsed items — return new instances from cleaning/validation steps
  • TDD works well for parsers — write tests for known patterns first, then edge cases
  • Log metrics (regex success rate, LLM call count) to track pipeline health

Anti-Patterns to Avoid

  • Sending all text to an LLM when regex handles 95%+ of cases (expensive and slow)
  • Using regex for free-form, highly variable text (LLM is better here)
  • Skipping confidence scoring and hoping regex "just works"
  • Mutating parsed objects during cleaning/validation steps
  • Not testing edge cases (malformed input, missing fields, encoding issues)

When to Use

  • Quiz/exam question parsing
  • Form data extraction
  • Invoice/receipt processing
  • Document structure parsing (headers, sections, tables)
  • Any structured text with repeating patterns where cost matters

© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/regex-vs-llm-structured-text of affaan-m/ECC.

Open the folder on GitHubat commit 2d515e4

Used in 5 other repositories

We found 11 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 5 other GitHub owners. This page covers the copy in affaan-m/ECC, which our catalogue first saw on October 9, 2026.

Compare with similar skills

Regex Vs LLM Structured Text next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Regex Vs LLM Structured Text compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Regex Vs LLM Structured Text this skillaffaan-m/ECC277k5 repos~1.7kAutomated safety check: PassMIT
Task Creatorbenchflow-ai/benchflow356—~4.5kAutomated safety check: PassApache-2.0
Weak Agent Testkklimuk/docx-cli216—~6.7kAutomated safety check: NotesMIT
Ccar F Examprep Coachsarveshtalele/claude-architect-exam-guide175—~5.8kAutomated safety check: PassNone
Value Mining LengthybooksLeoYeAI/openclaw-master-skills2.2k—~5.8kAutomated safety check: PassMIT
Canvas Reading AnnotationX-isdoingreat/canvas-pilot125—~8.4kAutomated safety check: NotesAGPL-3.0

Similar skills

  • Task Creator

    benchflow-ai/benchflow

    SkillsBench task authoring — walk a contributor from idea to submission-ready task following CONTRIBUTING.md and the task-implementation rubric.

    356 GitHub stars~4.5k tokensUpdated 5 days ago
    EducationAuto-check passed
  • Weak Agent Test

    kklimuk/docx-cli

    Run the weak-agent adversarial test harness against docx-cli.

    216 GitHub stars~6.7k tokensUpdated today
    Documents & OfficeAuto-check: notes
  • Ccar F Examprep Coach

    sarveshtalele/claude-architect-exam-guide

    A personalized study-coach skill for the Claude Certified Architect – Foundations (CCAR-F) certification.

    175 GitHub stars~5.8k tokensUpdated 28 days ago
    EducationAuto-check passed
  • Value Mining Lengthybooks

    LeoYeAI/openclaw-master-skills

    Extract actionable insights from books using Four-Layer Methodology: (1) Skeleton - conceptual frameworks and mental models, (2) Flesh - 2-3 detailed case studies including original examples…

    2.2k GitHub stars~5.8k tokensUpdated 2 mo ago
    EducationAuto-check passed
  • Canvas Reading Annotation

    X-isdoingreat/canvas-pilot

    Generic reading-annotation handler for academic-writing courses — annotates reading PDFs with color-coded highlights + margin notes + filled answer blanks per the instructor's rubric.

    125 GitHub stars~8.4k tokensUpdated 2 mo ago
    EducationAuto-check: notes
  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.8k tokensUpdated 3 days ago
    EducationAuto-check passed

More from affaan-m/ECC

All 682 skills in this repo
  • Skill Stocktake

    affaan-m/ECC

    Audits your installed Claude skills and commands for quality, with a quick mode for recently changed skills and a full mode that evaluates all of them through subagents.

    277k GitHub starsUsed in 5 repos~3.1k tokens
    Auto-check passed
  • Ingests, indexes, searches, edits and monitors video, audio and live streams through the VideoDB Python SDK, returning stream links, clips and timestamps.

    277k GitHub starsUsed in 3 repos~3.5k tokens
    Auto-check: notes
  • Docs Governance

    affaan-m/ECC

    Route broad documentation-governance requests to existing ECC skills and run an opt-in, read-only audit of mapped documentation roles, links, ADR indexes, and evidence references.

    277k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Rules Distillation

    affaan-m/ECC

    Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.

    277k GitHub starsUsed in 2 repos~2.3k tokens
    Auto-check passed
  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    277k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Set an ECC-specific frontend design direction for production UI work.

    277k GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed

Questions about Regex Vs LLM Structured Text

What does Regex Vs LLM Structured Text do?

Decision framework for parsing structured text (quizzes, forms, invoices, receipts, tables) with a hybrid regex-first pipeline — regex extraction handles 95%+ cheaply, a confidence scorer flags…. Regex Vs LLM Structured Text is an agent skill from affaan-m/ECC. Decision framework for parsing structured text (quizzes, forms, invoices, receipts, tables) with a hybrid regex-first pipeline — regex extraction handles 95%+ cheaply, a confidence scorer flags low-confidence items, and an LLM validator fixes only the edge cases.

When should I use Regex Vs LLM Structured Text?

Regex Vs LLM Structured Text fits situations like: choosing between regex and LLM for text extraction; building a cheap document parser; optimizing extraction cost and accuracy.

How do I install Regex Vs LLM Structured Text in Claude Code?

Run `npx skills add affaan-m/ECC --skill regex-vs-llm-structured-text -a claude-code`. Or copy the skill folder (skills/regex-vs-llm-structured-text in affaan-m/ECC) into .claude/skills/regex-vs-llm-structured-text in your project. Claude Code loads it when a task matches its description.

How do I install Regex Vs LLM Structured Text in Codex?

Run `npx skills add affaan-m/ECC --skill regex-vs-llm-structured-text -a codex`. Or copy the skill folder (skills/regex-vs-llm-structured-text in affaan-m/ECC) into .agents/skills/regex-vs-llm-structured-text in your project. Codex loads it when a task matches its description.

Can I use Regex Vs LLM Structured Text in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill regex-vs-llm-structured-text -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/regex-vs-llm-structured-text, .gemini/skills/regex-vs-llm-structured-text, .github/skills/regex-vs-llm-structured-text and .opencode/skills/regex-vs-llm-structured-text in your project.

What does Regex Vs LLM Structured Text need to run?

SKILL.md names no scripts, command-line tools or credentials: Regex Vs LLM Structured Text is instructions for the agent only. Our summary lists: Python 3.

Does Regex Vs LLM Structured Text access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Regex Vs LLM Structured Text safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Regex Vs LLM Structured Text use?

Regex Vs LLM Structured Text is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Regex Vs LLM Structured Text use?

About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Regex Vs LLM Structured Text?

Skills that share tags, products or a category with Regex Vs LLM Structured Text: Task Creator (benchflow-ai/benchflow, 356 stars), Weak Agent Test (kklimuk/docx-cli, 216 stars), Ccar F Examprep Coach (sarveshtalele/claude-architect-exam-guide, 175 stars) and Value Mining Lengthybooks (LeoYeAI/openclaw-master-skills, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Regex Vs LLM Structured Text?

affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 276,673 GitHub stars. The repository holds 683 skills in this directory. The repository was last updated on October 11, 2026.

Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.