You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于正则表达式提取OCR损坏文件中的姓名与地址

Extracting LASTNAME, FIRSTNAME & Addresses from Garbled OCR Data

Got it, let's break this down—dealing with messed-up OCR output is a pain, but regex can cut through most of the noise if you target the right patterns. Here's how I'd approach it:

Step 1: Target the LASTNAME, FIRSTNAME Pattern (Even with OCR Garbage)

OCR often munges separators, so we need a regex that handles messy variations of the last-name-comma-first-name format. Here's a robust starting point:

\b([A-Za-z'-]+)\s*[,;:\s]\s*([A-Za-z'-]+)\b

What each part does:

  • \b: Word boundary to avoid partial matches (like picking up parts of addresses)
  • ([A-Za-z'-]+): Captures last name—handles hyphenated names (like Smith-Jones) and apostrophes (like O'Connor)
  • \s*[,;:\s]\s*: Matches any messy separator OCR might have inserted (comma, semicolon, colon, extra spaces, or even a mix)
  • ([A-Za-z'-]+): Captures first name, same logic as last name

If your OCR has more extreme character mix-ups (like 1 instead of l, 0 instead of O), tweak it to include those:

\b([A-Za-z0-9'-]+)\s*[,;:\s]\s*([A-Za-z0-9'-]+)\b

Once you've got the names, you need to grab the address that follows each name until the next name (or end of file). Here's a regex that pairs names with their address blocks:

(\b([A-Za-z'-]+)\s*[,;:\s]\s*([A-Za-z'-]+)\b)\s*(.*?)(?=\b[A-Za-z'-]+\s*[,;:\s]\s*[A-Za-z'-]+\b|$)

Breakdown:

  • The first group is our name pattern from Step 1
  • \s*: Skips any whitespace between name and address
  • (.*?): Non-greedy capture of the address (stops at the next valid name or end of text)
  • (?=\b[A-Za-z'-]+\s*[,;:\s]\s*[A-Za-z'-]+\b|$): Positive lookahead to detect the next name or end of file—this prevents capturing multiple addresses into one block

Step 3: Clean Up OCR Artifacts

After extraction, you'll probably need to post-process to fix common OCR issues:

  • Remove random separators (like |, _) from addresses
  • Fix character substitutions (replace 1 with l, 0 with O where it makes sense)
  • Trim extra whitespace from name and address fields

Example Workflow (Python)

If you're processing a large file, use Python's re module to automate extraction:

import re

with open("garbled_ocr.txt", "r") as f:
    text = f.read()

# Regex to capture name + address pairs
pattern = r'(\b([A-Za-z'-]+)\s*[,;:\s]\s*([A-Za-z'-]+)\b)\s*(.*?)(?=\b[A-Za-z'-]+\s*[,;:\s]\s*[A-Za-z'-]+\b|$)'
matches = re.finditer(pattern, text, re.DOTALL)

for match in matches:
    full_name = match.group(1).strip()
    last_name = match.group(2).strip()
    first_name = match.group(3).strip()
    address = re.sub(r'[|_]', '', match.group(4).strip())  # Clean OCR junk
    print(f"Name: {full_name}")
    print(f"Address: {address}\n")

A quick tip: Test your regex on a sample of your messy data first—tweak the character sets or separators to match the specific garbage your OCR produced. Small adjustments go a long way!

内容的提问来源于stack exchange,提问作者hapax

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 11:12:18