You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas技术问题:如何用动态正则匹配更新匹配列的值?

Solution: Fix Dynamic Regex Replacement for Pandas Columns

Got it, let's break down why your code isn't working and fix it. The core issue is that you can't pass an entire Series directly into a vectorized str.replace regex pattern—rf'{df.city}$' doesn't reference each row's city value, it converts the whole city column into a messy single string (like "0 London\n1 Detroit\n...") that never matches anything.

The Fix: Row-Wise Replacement with apply

We'll use apply to handle each row individually, which lets us use the exact city value from the same row as our regex pattern. We'll also add re.escape to safely handle any special characters in city names (like periods or parentheses) that might break regex matching.

First, let's set up your sample data to test:

import pandas as pd
import re
import numpy as np

# Sample DataFrame
df = pd.DataFrame({
    'name': ['Danny London', 'Tim Detroit', 'Keith New Orleans', 'Mary Jane'],
    'city': ['London', 'Detroit', 'The city of New Orleans', 'London']
})
df['name_cleaned'] = df['name'].copy()

Now apply the row-wise replacement logic:

df['name_cleaned'] = df.apply(
    lambda row: row['name_cleaned'].replace(
        rf'{re.escape(row["city"])}$',  # Use current row's city as escaped regex
        '',
        regex=True
    ) if row['name_cleaned'].endswith(row['city'])
    else row['name_cleaned'],
    axis=1
)

What This Does:

  1. Row-wise check: For each row, we first verify if name_cleaned ends with the corresponding city value.
  2. Safe regex replacement: If the check passes, we strip the trailing city name (escaped to avoid regex conflicts) from name_cleaned.
  3. Fallback: If no match is found, we keep the original name_cleaned value.

Result:

After running the code, your DataFrame will match your expected output exactly:

name                      city      name_cleaned
0    Danny London                    London             Danny
1     Tim Detroit                   Detroit              Tim
2  Keith New Orleans  The city of New Orleans  Keith New Orleans
3         Mary Jane                    London         Mary Jane

Alternative: Use np.where for Structured Logic

If you prefer sticking with a pattern similar to your original np.select approach, you can wrap the row-wise checks and replacements in apply calls with np.where:

# Define the match condition
condition = df.apply(lambda x: x['name_cleaned'].endswith(x['city']), axis=1)

# Define the replacement logic
replacement = df.apply(
    lambda x: x['name_cleaned'].replace(rf'{re.escape(x["city"])}$', '', regex=True),
    axis=1
)

# Update the column
df['name_cleaned'] = np.where(condition, replacement, df['name_cleaned'])

This achieves the same result, just with more explicit separation of condition and replacement logic.

内容的提问来源于stack exchange,提问作者venv

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 09:27:29