Python:如何用正则匹配替换DataFrame字符串及删除指定行
Hey there! Let's fix up your DataFrame cleaning task step by step. You've got two key goals: removing rows with specific values, and stripping those unwanted suffixes using regex. Let's break this down clearly.
第一步:修正删除指定行的代码
Your original approach has a couple of issues (like missing df['value'] reference and inefficient looping). Instead, use Pandas' built-in isin() method for a cleaner, faster solution:
# 定义要删除的内容列表 exclude_values = ['moreinfo', 'GoCats'] # 筛选出不在排除列表中的行(~表示取反) df = df[~df['value'].isin(exclude_values)]
This avoids looping and directly filters your DataFrame in one go.
第二步:用正则表达式去除后缀
Looking at your examples, the unwanted suffixes all start with a comma and continue to the end of the string. We can target this pattern with a simple regex:
# 替换掉从第一个逗号到行尾的所有内容 df['value'] = df['value'].str.replace(r',.*$', '', regex=True)
Let's break down the regex:
,: Matches the starting comma of the suffix.*: Matches any number of any characters (except newlines)$: Anchors the match to the end of the string
This ensures we strip everything after the first comma, which exactly matches your desired outcome.
完整代码示例
Putting it all together (including resetting the index to fix gaps from deleted rows):
import pandas as pd # 原始DataFrame数据 data = {'value': [ 'A067-M4FL-CAA-020', 'MRF2-050A-TFC,60 ,R-12,HT', 'moreinfo', 'MZF8-050Z-AAB', 'GoCats', 'MZA2-0580-TFD,60 ,R-669,LT' ]} df = pd.DataFrame(data) # 1. 删除指定行 exclude = ['moreinfo', 'GoCats'] df = df[~df['value'].isin(exclude)] # 2. 去除后缀 df['value'] = df['value'].str.replace(r',.*$', '', regex=True) # 重置索引(可选但推荐) df = df.reset_index(drop=True) print(df)
Running this will give you exactly the expected result:
value 0 A067-M4FL-CAA-020 1 MRF2-050A-TFC 2 MZF8-050Z-AAB 3 MZA2-0580-TFD
内容的提问来源于stack exchange,提问作者BenT

