You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python 3.x识别OCR残缺的1910-20年代美国专利地址信息

Alright, let's break down how to solve this problem—extracting city and state details for inventors and assignees from 1 million+ early 20th century US patents with messy OCR data is a tough lift, but Python gives us the tools to make it work reliably. I’ve worked with historical OCR datasets before, so I know the biggest pain points are broken headers, random character corruption, and redundant punctuation—we can’t rely on rigid regex alone; we need a hybrid approach of rule-based parsing, fuzzy matching, and entity recognition.

Problem Overview

Let’s recap the core requirements to stay aligned:

  • Target: 1910s–1920s US patents (Google-provided OCR text)
  • Extract: City + US State for:
    • Inventors (usually in headers, sometimes first paragraph)
    • Assignees (in headers; if no assignee exists, the inventor is a business entity)
  • Edge Cases: Multiple inventors/assignees, broken OCR (missing headers, split text, redundant periods, character corruption)
  • Scale: ~1M documents
Core Strategy

Our approach needs to be flexible enough to handle OCR chaos while efficient enough for large-scale processing:

  1. Preprocess to clean up noisy text
  2. Locate inventor/assignee sections (headers first, fall back to first paragraph if needed)
  3. Extract city/state entities using fuzzy matching (to fix typos) and regex (to capture patterns)
  4. Handle edge cases like business inventors and multiple entities
Step-by-Step Implementation

1. Preprocess Noisy OCR Text

First, we’ll strip out the worst of the OCR noise to make parsing easier:

import re
from rapidfuzz import fuzz  # Faster alternative to fuzzywuzzy for large datasets

def preprocess_ocr(text):
    # Normalize messy whitespace (replace tabs/newlines/multiple spaces with single spaces)
    text = re.sub(r'\s+', ' ', text).strip()
    # Remove redundant periods that don't end words/abbreviations (e.g., ". New York" → "New York")
    text = re.sub(r'(?<!\w)\.(?!\w)', '', text)
    # Fix common OCR typos for state abbreviations (e.g., "N Y" → "NY")
    state_fixes = {
        r'\bN\s*Y\b': 'NY', r'\bC\s*A\b': 'CA', r'\bI\s*L\b': 'IL',
        r'\bO\s*H\b': 'OH', r'\bM\s*I\b': 'MI', r'\bP\s*A\b': 'PA'
    }
    for pattern, replacement in state_fixes.items():
        text = re.sub(pattern, replacement, text, flags=re.IGNORECASE)
    return text

2. Locate Inventor/Assignee Sections

Inventors and assignees are almost always in the top of the document—we’ll target the first 10 lines as header candidates, and fall back to the first paragraph if headers are missing:

def extract_sections(text):
    lines = re.split(r'\n|\. ', text)  # Split text into logical lines (handle OCR line breaks)
    header_text = ' '.join(lines[:10])  # Grab top 10 lines as header candidates

    # Extract assignee section using keyword matches
    assignee_match = re.search(
        r'(Assignee|Assigned to|By):\s*(.*?)(?=(Inventor|Inventors|Application|$))',
        header_text, re.IGNORECASE | re.DOTALL
    )
    assignee_text = assignee_match.group(2).strip() if assignee_match else None

    # Extract inventor section: check header first, then fall back to first paragraph
    inventor_match = re.search(
        r'(Inventor|Inventors|Application of):\s*(.*?)(?=(Assignee|Assigned to|By|$))',
        header_text, re.IGNORECASE | re.DOTALL
    )
    if not inventor_match:
        # Check first 5 lines if header has no inventor info
        first_paragraph = ' '.join(lines[:5]) if len(lines) >=5 else ' '.join(lines)
        inventor_match = re.search(
            r'(Inventor|Inventors):\s*(.*?)(?=\.|$)',
            first_paragraph, re.IGNORECASE | re.DOTALL
        )
    inventor_text = inventor_match.group(2).strip() if inventor_match else None

    return {'inventors': inventor_text, 'assignees': assignee_text}

3. Extract City + State Entities

We’ll use a reference list of US states and early 20th century major cities, plus fuzzy matching to handle OCR typos (like "Detroi" → "Detroit"):

# Reference data: expand this with 1910s-1920s specific cities (e.g., Akron, OH; Buffalo, NY)
US_STATES = {
    'AL': 'Alabama', 'AK': 'Alaska', 'AZ': 'Arizona', 'AR': 'Arkansas', 'CA': 'California',
    'CO': 'Colorado', 'CT': 'Connecticut', 'DE': 'Delaware', 'FL': 'Florida', 'GA': 'Georgia',
    'IL': 'Illinois', 'IN': 'Indiana', 'IA': 'Iowa', 'KS': 'Kansas', 'KY': 'Kentucky',
    'LA': 'Louisiana', 'ME': 'Maine', 'MD': 'Maryland', 'MA': 'Massachusetts', 'MI': 'Michigan',
    'MN': 'Minnesota', 'MS': 'Mississippi', 'MO': 'Missouri', 'MT': 'Montana', 'NE': 'Nebraska',
    'NV': 'Nevada', 'NH': 'New Hampshire', 'NJ': 'New Jersey', 'NM': 'New Mexico', 'NY': 'New York',
    'NC': 'North Carolina', 'ND': 'North Dakota', 'OH': 'Ohio', 'OK': 'Oklahoma', 'OR': 'Oregon',
    'PA': 'Pennsylvania', 'RI': 'Rhode Island', 'SC': 'South Carolina', 'SD': 'South Dakota',
    'TN': 'Tennessee', 'TX': 'Texas', 'UT': 'Utah', 'VT': 'Vermont', 'VA': 'Virginia',
    'WA': 'Washington', 'WV': 'West Virginia', 'WI': 'Wisconsin', 'WY': 'Wyoming'
}
STATE_FULL_TO_ABBR = {v: k for k, v in US_STATES.items()}
MAJOR_CITIES_1920s = [
    'New York', 'Chicago', 'Philadelphia', 'Detroit', 'Boston', 'St. Louis',
    'Cleveland', 'Pittsburgh', 'Baltimore', 'San Francisco', 'Cincinnati', 'Milwaukee'
]

def extract_city_state(entity_text):
    if not entity_text:
        return []
    
    results = []
    # Split multiple entities (e.g., "John Doe, Chicago, IL; Jane Smith, Detroit, MI")
    entities = re.split(r';|, and | and ', entity_text)
    
    for entity in entities:
        entity = entity.strip()
        if not entity:
            continue
        
        # Find matching state (handle abbreviations and full names with fuzzy matching)
        state_abbr = None
        for abbr, full_name in US_STATES.items():
            # Exact match for abbreviations, fuzzy match for full names (85% threshold)
            if re.search(r'\b' + abbr + r'\b', entity, re.IGNORECASE):
                state_abbr = abbr
                break
            if fuzz.partial_ratio(entity.lower(), full_name.lower()) >= 85:
                state_abbr = abbr
                break
        
        if not state_abbr:
            continue
        
        # Extract city: remove inventor names and isolate text before the state
        state_pattern = re.compile(
            r'\b' + US_STATES[state_abbr] + r'\b|\b' + state_abbr + r'\b',
            re.IGNORECASE
        )
        city_part = state_pattern.split(entity)[0].strip()
        # Strip out names (usually come before city in patent text)
        city_part = re.sub(r'^[A-Za-z\s]+,?\s*', '', city_part).strip()
        
        # Fuzzy match city against 1920s major cities
        city = None
        for candidate_city in MAJOR_CITIES_1920s:
            if fuzz.partial_ratio(city_part.lower(), candidate_city.lower()) >= 80:
                city = candidate_city
                break
        # If no fuzzy match, clean up remaining text and use as city
        if not city:
            city = re.sub(r'[^\w\s]', '', city_part).strip()
        
        if city and state_abbr:
            results.append({'city': city, 'state': state_abbr})
    
    return results

4. Handle "Inventor is Business" Edge Case

If there’s no assignee, check if the inventor text contains business keywords (e.g., "Co.", "Corporation"):

def is_business(entity_text):
    if not entity_text:
        return False
    business_keywords = ['Co.', 'Company', 'Corp.', 'Corporation', 'Inc.', 'Incorporated']
    for keyword in business_keywords:
        if re.search(r'\b' + keyword + r'\b', entity_text, re.IGNORECASE):
            return True
    return False

5. End-to-End Pipeline

Put all the pieces together into a single function for processing patents:

def process_patent(text):
    cleaned_text = preprocess_ocr(text)
    sections = extract_sections(cleaned_text)
    
    inventor_locations = extract_city_state(sections['inventors'])
    assignee_locations = extract_city_state(sections['assignees'])
    
    # Mark inventor as business if no assignee exists and text looks like a business
    inventor_is_business = is_business(sections['inventors']) if not sections['assignees'] else False
    
    return {
        'inventor_locations': inventor_locations,
        'assignee_locations': assignee_locations,
        'inventor_is_business': inventor_is_business
    }
Optimization Tips for 1M Documents
  • Multiprocessing: Use Python’s multiprocessing to parallelize processing across CPU cores:
    import multiprocessing
    
    def batch_process(patent_texts):
        with multiprocessing.Pool(processes=multiprocessing.cpu_count()) as pool:
            results = pool.map(process_patent, patent_texts)
        return results
    
  • Expand Reference Data: Add more 1910s–1920s specific cities and towns to improve fuzzy matching accuracy.
  • Adjust Fuzzy Thresholds: Tweak the fuzz.partial_ratio values based on your OCR quality—lower thresholds for extremely corrupted text, higher for cleaner OCR.
  • Precompile Regex: Compile regex patterns once outside functions to save processing time.

内容的提问来源于stack exchange,提问作者MERose

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:27:02