Regex Vs Llm Structured Text

affaan-m/ECC/skills/regex-vs-llm-structured-text

作者 affaan-mef648e01899ba3e8dc6371642deaaf64b4477775無授權條款275K 個星標收錄於 2026年10月9日更新於 2026年10月9日儲存庫4 天前更新

Decision framework for parsing structured text (quizzes, forms, invoices, receipts, tables) with a hybrid regex-first pipeline — regex extraction handles 95%+ cheaply, a confidence scorer flags low-confidence items, and an LLM validator fixes only the edge cases. Use when choosing between regex and LLM for text extraction, building a cheap document parser, or optimizing extraction cost and accuracy.

AI 產生的概覽

針對結構化文字的混合解析決策框架與程式碼模式:以正規表達式優先,僅對低信賴度項目呼叫 LLM 驗證。

功能
說明在從測驗題、表單、發票、收據等重複性文字中擷取結構時,何時使用正規表達式、何時使用 LLM。它提供決策樹、架構圖,以及正規表達式解析器、信賴度評分器、LLM 驗證器與組合流程的 Python 程式碼。另列出正式環境指標、最佳實務與反模式,用於權衡成本與準確率。
適用情境
適用於在結構化文字擷取中選擇正規表達式或 LLM、建置低成本文件解析器,或在維持準確率的同時減少 LLM 呼叫。適合測驗題、表單欄位、發票、收據、表格等重複格式。不適用於自由格式、變化很大的文字,框架建議這類情況直接使用 LLM。
執行需求
此技能不附帶指令碼,僅為說明性內容。所述流程使用 Python 的 re 與 dataclasses 模組;可選的 LLM 驗證步驟需要 Anthropic 風格的用戶端以及存取 Haiku 等級模型的 API 權限。

Regex vs LLM for Structured Text Parsing

A practical decision framework for parsing structured text (quizzes, forms, invoices, documents). The key insight: regex handles 95-98% of cases cheaply and deterministically. Reserve expensive LLM calls for the remaining edge cases.

When to Activate

  • Parsing structured text with repeating patterns (questions, forms, tables)
  • Deciding between regex and LLM for text extraction
  • Building hybrid pipelines that combine both approaches
  • Optimizing cost/accuracy tradeoffs in text processing

Decision Framework

Is the text format consistent and repeating?├── Yes (>90% follows a pattern) → Start with Regex│   ├── Regex handles 95%+ → Done, no LLM needed│   └── Regex handles <95% → Add LLM for edge cases only└── No (free-form, highly variable) → Use LLM directly

Architecture Pattern

Source Text    │    ▼[Regex Parser] ─── Extracts structure (95-98% accuracy)    │    ▼[Text Cleaner] ─── Removes noise (markers, page numbers, artifacts)    │    ▼[Confidence Scorer] ─── Flags low-confidence extractions    │    ├── High confidence (≥0.95) → Direct output    │    └── Low confidence (<0.95) → [LLM Validator] → Output

Implementation

1. Regex Parser (Handles the Majority)

python
import refrom dataclasses import dataclass
@dataclass(frozen=True)class ParsedItem:    id: str    text: str    choices: tuple[str, ...]    answer: str    confidence: float = 1.0
def parse_structured_text(content: str) -> list[ParsedItem]:    """Parse structured text using regex patterns."""    pattern = re.compile(        r"(?P<id>\d+)\.\s*(?P<text>.+?)\n"        r"(?P<choices>(?:[A-D]\..+?\n)+)"        r"Answer:\s*(?P<answer>[A-D])",        re.MULTILINE | re.DOTALL,    )    items = []    for match in pattern.finditer(content):        choices = tuple(            c.strip() for c in re.findall(r"[A-D]\.\s*(.+)", match.group("choices"))        )        items.append(ParsedItem(            id=match.group("id"),            text=match.group("text").strip(),            choices=choices,            answer=match.group("answer"),        ))    return items

2. Confidence Scoring

Flag items that may need LLM review:

python
@dataclass(frozen=True)class ConfidenceFlag:    item_id: str    score: float    reasons: tuple[str, ...]
def score_confidence(item: ParsedItem) -> ConfidenceFlag:    """Score extraction confidence and flag issues."""    reasons = []    score = 1.0
    if len(item.choices) < 3:        reasons.append("few_choices")        score -= 0.3
    if not item.answer:        reasons.append("missing_answer")        score -= 0.5
    if len(item.text) < 10:        reasons.append("short_text")        score -= 0.2
    return ConfidenceFlag(        item_id=item.id,        score=max(0.0, score),        reasons=tuple(reasons),    )
def identify_low_confidence(    items: list[ParsedItem],    threshold: float = 0.95,) -> list[ConfidenceFlag]:    """Return items below confidence threshold."""    flags = [score_confidence(item) for item in items]    return [f for f in flags if f.score < threshold]

3. LLM Validator (Edge Cases Only)

python
def validate_with_llm(    item: ParsedItem,    original_text: str,    client,) -> ParsedItem:    """Use LLM to fix low-confidence extractions."""    response = client.messages.create(        model="claude-haiku-4-5-20251001",  # Cheapest model for validation        max_tokens=500,        messages=[{            "role": "user",            "content": (                f"Extract the question, choices, and answer from this text.\n\n"                f"Text: {original_text}\n\n"                f"Current extraction: {item}\n\n"                f"Return corrected JSON if needed, or 'CORRECT' if accurate."            ),        }],    )    # Parse LLM response and return corrected item...    return corrected_item

4. Hybrid Pipeline

python
def process_document(    content: str,    *,    llm_client=None,    confidence_threshold: float = 0.95,) -> list[ParsedItem]:    """Full pipeline: regex -> confidence check -> LLM for edge cases."""    # Step 1: Regex extraction (handles 95-98%)    items = parse_structured_text(content)
    # Step 2: Confidence scoring    low_confidence = identify_low_confidence(items, confidence_threshold)
    if not low_confidence or llm_client is None:        return items
    # Step 3: LLM validation (only for flagged items)    low_conf_ids = {f.item_id for f in low_confidence}    result = []    for item in items:        if item.id in low_conf_ids:            result.append(validate_with_llm(item, content, llm_client))        else:            result.append(item)
    return result

Real-World Metrics

From a production quiz parsing pipeline (410 items):

MetricValue
Regex success rate98.0%
Low confidence items8 (2.0%)
LLM calls needed~5
Cost savings vs all-LLM~95%
Test coverage93%

Best Practices

  • Start with regex — even imperfect regex gives you a baseline to improve
  • Use confidence scoring to programmatically identify what needs LLM help
  • Use the cheapest LLM for validation (Haiku-class models are sufficient)
  • Never mutate parsed items — return new instances from cleaning/validation steps
  • TDD works well for parsers — write tests for known patterns first, then edge cases
  • Log metrics (regex success rate, LLM call count) to track pipeline health

Anti-Patterns to Avoid

  • Sending all text to an LLM when regex handles 95%+ of cases (expensive and slow)
  • Using regex for free-form, highly variable text (LLM is better here)
  • Skipping confidence scoring and hoping regex "just works"
  • Mutating parsed objects during cleaning/validation steps
  • Not testing edge cases (malformed input, missing fields, encoding issues)

When to Use

  • Quiz/exam question parsing
  • Form data extraction
  • Invoice/receipt processing
  • Document structure parsing (headers, sections, tables)
  • Any structured text with repeating patterns where cost matters

來源與署名

來源:affaan-m/ECC位於skills/regex-vs-llm-structured-text提交ef648e0

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架