Pandas技术问题:如何用动态正则匹配更新匹配列的值?
Got it, let's break down why your code isn't working and fix it. The core issue is that you can't pass an entire Series directly into a vectorized str.replace regex pattern—rf'{df.city}$' doesn't reference each row's city value, it converts the whole city column into a messy single string (like "0 London\n1 Detroit\n...") that never matches anything.
The Fix: Row-Wise Replacement with apply
We'll use apply to handle each row individually, which lets us use the exact city value from the same row as our regex pattern. We'll also add re.escape to safely handle any special characters in city names (like periods or parentheses) that might break regex matching.
First, let's set up your sample data to test:
import pandas as pd import re import numpy as np # Sample DataFrame df = pd.DataFrame({ 'name': ['Danny London', 'Tim Detroit', 'Keith New Orleans', 'Mary Jane'], 'city': ['London', 'Detroit', 'The city of New Orleans', 'London'] }) df['name_cleaned'] = df['name'].copy()
Now apply the row-wise replacement logic:
df['name_cleaned'] = df.apply( lambda row: row['name_cleaned'].replace( rf'{re.escape(row["city"])}$', # Use current row's city as escaped regex '', regex=True ) if row['name_cleaned'].endswith(row['city']) else row['name_cleaned'], axis=1 )
What This Does:
- Row-wise check: For each row, we first verify if
name_cleanedends with the correspondingcityvalue. - Safe regex replacement: If the check passes, we strip the trailing city name (escaped to avoid regex conflicts) from
name_cleaned. - Fallback: If no match is found, we keep the original
name_cleanedvalue.
Result:
After running the code, your DataFrame will match your expected output exactly:
name city name_cleaned 0 Danny London London Danny 1 Tim Detroit Detroit Tim 2 Keith New Orleans The city of New Orleans Keith New Orleans 3 Mary Jane London Mary Jane
Alternative: Use np.where for Structured Logic
If you prefer sticking with a pattern similar to your original np.select approach, you can wrap the row-wise checks and replacements in apply calls with np.where:
# Define the match condition condition = df.apply(lambda x: x['name_cleaned'].endswith(x['city']), axis=1) # Define the replacement logic replacement = df.apply( lambda x: x['name_cleaned'].replace(rf'{re.escape(x["city"])}$', '', regex=True), axis=1 ) # Update the column df['name_cleaned'] = np.where(condition, replacement, df['name_cleaned'])
This achieves the same result, just with more explicit separation of condition and replacement logic.
内容的提问来源于stack exchange,提问作者venv

