You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python Regex或NLTK提取街道名、人名及门牌号的结构化列表?

Solution for Parsing Unstructured Address Entries into Structured Lists

Approach Overview

Your problem is super common when dealing with messy OCR-extracted address data—especially when multiple entries are grouped under a single street name. A regex-based strategy paired with light preprocessing is going to be your best bet here, since the text follows a loose but predictable pattern:

  • Street names are typically followed by an orientation tag (N/S/E/W/SE/etc.), often separated by an em dash or comma.
  • Multiple entries for the same street are listed with dashes, each linking a person’s name to a house number.

Step 1: Preprocess the Text First

First, clean up the input to standardize separators and reduce noise from OCR errors:

  • Replace em dashes (—) with regular dashes (-) to simplify splitting.
  • Strip extra spaces and ignore obvious garbage text (like "Widdicombe- co ee" which is clearly an OCR mistake).
  • Normalize punctuation to make regex matching more reliable.

Step 2: Regex Patterns to Target Key Entities

We’ll use two core regex patterns to break down the data:

  1. Street Block Pattern: Identifies the street name and orientation, then captures all associated entries for that street:

    ([A-Za-z\s]+),?\s*([NSEW][EW]?)[.,-]\s*(.*?)(?=\d+\.|[A-Z][a-z]+[\s,]|$)
    

    This captures:

    • Group 1: Full street name (e.g., "Cannon Street Road")
    • Group 2: Orientation tag (e.g., "E")
    • Group 3: All entries linked to this street
  2. Entry Pattern: Extracts the person’s name and house number from each individual entry:

    ([A-Z][\w\s.]+?)\s*(\d+)\.?
    

    Captures:

    • Group 1: Person’s name (including titles like M.D.)
    • Group 2: House number

Step 3: Full Python Implementation

Here’s a ready-to-use script that combines these patterns with post-processing to generate your desired structured list:

import re

def parse_addresses(raw_text):
    # Preprocess to clean up OCR artifacts and standardize separators
    cleaned = raw_text.replace("—", "-").replace("  ", " ").strip()
    
    # Regex to match entire street blocks
    street_block_re = re.compile(
        r'([A-Za-z\s]+),?\s*([NSEW][EW]?)[.,-]\s*(.*?)(?=\d+\.|[A-Z][a-z]+[\s,]|$)',
        re.DOTALL
    )
    # Regex to extract name and number from each entry
    entry_re = re.compile(r'([A-Z][\w\s.]+?)\s*(\d+)\.?')
    
    structured_data = []
    
    # Iterate over each detected street block
    for street_match in street_block_re.finditer(cleaned):
        street_name = street_match.group(1).strip()
        entries_raw = street_match.group(3).strip()
        
        # Split entries by dashes, skipping empty strings from extra spaces
        entries = [e.strip() for e in entries_raw.split("-") if e.strip()]
        
        # Parse each entry for name and number
        for entry in entries:
            entry_match = entry_re.search(entry)
            if entry_match:
                person_name = entry_match.group(1).strip()
                house_number = entry_match.group(2).strip()
                structured_data.append([street_name, person_name, house_number])
    
    return structured_data

# Example input from your dataset
input_text = """Camden Row,Camberwell, S.E—A. Massey, M.D.4. Campden Hill, Kensington. (Hornton House). Campden Hill Road, Kensington. James, M.D. 6. Canning Town. E—R. J. Carey (Widdicombe- co ee Cannon Street. E.C.—R. Cresswell, 151. Cannon Street Road. E.—R. W. Lammiman, 106. —J. R. Morrison, 57.—B. R. Rygate, 126.— J. J. Rygate, M.B. 126. Canonbury N. (see foot note)—J. Cheetham, M.D. (Springjield House), Canonbury Lane, N.—H. Bateman, Roberts, 10.—J. Rose, 3."""

# Run the parser and print results
parsed = parse_addresses(input_text)
for item in parsed:
    print(item)

Step 4: Handling Edge Cases

  • OCR Errors: Garbage text like "Widdicombe- co ee" is automatically ignored since it doesn’t match the name-number regex pattern.
  • Parenthetical Notes: Phrases like "(Hornton House)" or "(see foot note)" are skipped because they don’t contain a valid name-number pair.
  • Multiple Entries per Street: The code splits entries by dashes, so every entry under the same street gets its own structured list item.

Sample Output

For your example input, the script will produce exactly the entries you want, like:

['Cannon Street Road', 'R. W. Lammiman', '106']
['Cannon Street Road', 'J. R. Morrison', '57']
['Cannon Street Road', 'B. R. Rygate', '126']
['Cannon Street Road', 'J. J. Rygate, M.B', '126']

Optional: Enhance with NLTK

If you want to boost accuracy, you can pair this regex approach with NLTK’s Named Entity Recognition (NER) to validate that extracted names are actually people. For example, after extracting a name, use nltk.ne_chunk() to confirm it’s a person entity—this helps filter out false positives from OCR mistakes.

内容的提问来源于stack exchange,提问作者jamespaulphelan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:29:10