使用Python处理Excel地址字段:修正序数词与方位词大小写格式
The issue you're facing is that standard title-case conversion (like str.title()) incorrectly capitalizes ordinal suffixes (turning 1st into 1St) and lowercases direction abbreviations (turning SW into Sw). Here's a robust Python solution that handles these exceptions properly:
Custom Proper Case Function
This function processes each word in the address, applying title case by default but adjusting ordinals and direction abbreviations to their correct formats:
import pandas as pd def custom_propercase(address): if pd.isna(address): return address # Handle missing values gracefully words = address.split() fixed_words = [] # Define ordinal suffixes to check for ordinal_suffixes = {'st', 'nd', 'rd', 'th'} # Define direction abbreviations (lowercase for case-insensitive matching) directions = {'n', 's', 'e', 'w', 'ne', 'nw', 'se', 'sw'} for word in words: word_lower = word.lower() # Check if the word is an ordinal (ends with a known suffix after numbers) if any(word_lower.endswith(suffix) for suffix in ordinal_suffixes): # Split into numeric part and suffix num_part = '' suffix_part = '' for idx, char in enumerate(word): if char.isdigit(): num_part += char else: suffix_part = word[idx:] break # Keep number as-is, make suffix lowercase fixed_word = f"{num_part}{suffix_part.lower()}" fixed_words.append(fixed_word) # Check if the word is a direction abbreviation elif word_lower in directions: fixed_words.append(word.upper()) # For all other words, apply standard title case else: fixed_words.append(word.title()) return ' '.join(fixed_words) # Apply to your DataFrame addr_df['ADDRESS1'] = addr_df['ADDRESS1'].apply(custom_propercase)
How It Works:
- Ordinals: Detects words ending with
st,nd,rd, orthafter a number, then converts the suffix to lowercase (e.g.,1St→1st,2Nd→2nd). - Directions: Converts common direction abbreviations (N/S/E/W, NE/NW/SE/SW) to uppercase (e.g.,
Sw→SW,Nw→NW). - Other Words: Applies standard title case for street names, cities, and other address components.
- Missing Values: Handles
NaNentries to avoid runtime errors.
Alternative: Post-Processing with Regex
If you want to keep your original proper case function and just fix the errors afterward, use these regex substitutions:
import pandas as pd import re def fix_address_case(address): if pd.isna(address): return address # Fix ordinals: 1St → 1st, 2Nd → 2nd address = re.sub(r'(\d+)(St|Nd|Rd|Th)', lambda m: f"{m.group(1)}{m.group(2).lower()}", address) # Fix directions: Sw → SW, Nw → NW address = re.sub(r'\b(Sw|Nw|Ne|Se|N|S|E|W)\b', lambda m: m.group(1).upper(), address) return address # Apply after your initial proper case conversion addr_df['ADDRESS1'] = addr_df['ADDRESS1'].apply(fix_address_case)
Both solutions will correctly format your addresses while preserving the intended capitalization for ordinals and directions.
内容的提问来源于stack exchange,提问作者spys

